Skip to yearly menu bar Skip to main content


Session

Sydney Poster Session 1

Hall 1-4
Tue 8 Dec 10 a.m. AEDT — 1 p.m. AEDT
Abstract:
Chat is not available.


$\alpha$Depth: Learning Single-Pass Soft Boundary Decomposition for Stereo Conversion

Xiang Zhang ⋅ Yang Zhang ⋅ Lukas Mehl ⋅ Karlis M Briedis ⋅ Markus Gross ⋅ Christopher Schroers

Accurately modeling soft boundaries, e.g., hair and defocus blur, is a fundamental challenge in stereo conversion due to the ambiguous blending of foreground and background. Existing depth models primarily predict single-layer depth, leading to ambiguity in depth correspondence at soft boundaries. While matting techniques can capture opacity for layered modeling, they often struggle in complex scenes with multiple targets and usually require user intervention. This paper introduces $\alpha$Depth, a layered representation that decomposes soft boundaries for high-fidelity stereo conversion. Specifically, we first resolve mixed color and depth ambiguity by estimating layered color and depth values at soft boundaries. Considering complex multi-target scenes, we design a Circular Alpha Representation (CAR) that shifts the paradigm from global target extraction to local boundary decomposition. Unlike prior matting methods restricted to a single foreground/background, CAR enables efficient scene-level inference without manual guidance. Extensive evaluations demonstrate that $\alpha$Depth achieves state-of-the-art performance in stereo conversion, eliminating background bleeding and structural distortions at soft boundaries.


$\mathbf{D^{3}S^{2}}$: Diffusion-Guided Dataset Distillation for Semantic Segmentation

Wenjie Zheng ⋅ Haoji Hu ⋅ Jiali Lu ⋅ Xingze Zou ⋅ Jing Wang

Dataset distillation (DD) aims to compress large-scale datasets into compact synthetic sets while preserving training efficacy. However, existing studies mainly focus on image classification, leaving dense prediction tasks such as semantic segmentation largely underexplored. In this work, we identify three key challenges for segmentation DD: (*i*) long-tailed class imbalance, (*ii*) the need for strict pixel-wise alignment between images and dense labels, and (*iii*) the high computational cost of optimizing high-resolution data with complex models. To address these challenges, we propose $\mathbf{D^{3}S^{2}}$, a **D**iffusion-guided **D**ataset **D**istillation framework for **S**emantic **S**egmentation. Our method adopts a two-stage design. In **Class-Balanced Mask Selection**, we construct a representative mask set via a greedy strategy that prioritizes underrepresented classes. In **Diffusion-Guided Image Synthesis**, we employ a pretrained layout-to-image diffusion model to generate images conditioned on the selected masks, naturally ensuring spatial alignment. To further enhance the training utility of synthesized data, we introduce guided diffusion sampling with two complementary objectives: a segmentation-consistency loss for pixel-level alignment, and a class-wise feature matching loss for aligning per-class feature statistics across layers. Extensive experiments demonstrate the superiority of $D^{3}S^{2}$. Notably, at an extremely compression rate of 1\%, our method achieves 24.99\% and 35.49\% mIoU on ADE20K and COCO-Stuff with Mask2Former (Swin-S), outperforming random selection by 9.34\% and 5.70\%, respectively.


3D Fresnel Volumizing for Efficient Implicit Velocity Field Reconstruction

Sihan Chen ⋅ haochen sun ⋅ Peng Jiang ⋅ Anthony Cohn

Reconstructing velocity distributions using wavefield responses provides a powerful tool for probing the internal structures of objects and has been widely applied in nondestructive testing, geological exploration, and medical imaging. While full waveform inversion (FWI) provides high-resolution results, its severe ill-posedness and high computational cost hinder efficient 3D reconstruction. In this paper, we propose a novel framework for efficient 3D internal velocity field reconstruction. First, we parameterize the velocity field with a coordinate-based neural network as an implicit continuous function, enabling compact representation from sparse observations. Second, we adopt ray-based traveltime tomography with multi-level parallel forward modeling to accelerate computation. Third, we introduce a differentiable Fresnel volumizing method that extends 1D ray paths into 3D volumetric regions (i.e., Fresnel volumes), enabling spatially continuous and multi-scale gradient diffusion, thus alleviating gradient sparsity and improving reconstruction accuracy. To further improve efficiency, we propose a 2D slice-based construction of 3D Fresnel volumes using a lightweight neural network. Extensive experiments on synthetic datasets, including ablation studies, demonstrate up to 10 times speedup over FWI while maintaining comparable accuracy, validating the effectiveness of the proposed method for efficient 3D inverse problems.

3D Gaussian Splatting (3DGS) has become a vital tool in novel view synthesis. 3D Gaussian functions with attributes like color, opacity, and scale are stacked together to explicitly model a radiance field. However, discrete 3D Gaussians usually produce poor depth with few geometry details but lots of floaters, which makes modeling smooth radiance fields remain a challenge. To address this issue, we propose to learn a radiance field within a continuous view-dependent color field parameterized by neural networks. With the continuous constraints in color, the bias on low-frequency signals of neural networks highly encourages the geometry to contribute to high-frequency color variations on images, but not solely relying on the color attribute itself. Moreover, we impose a constraint to improve the Lipschitz continuity of the neural color field, making large changes of each Gaussian's color also match with large changes of its position in terms of a ratio. To this end, we additionally introduce novel self-learning strategies to learn neural view-dependent color fields without any external knowledge or priors, aiming for better generalization. Our numerical and visual comparisons on widely used benchmarks justify our idea and show better ability of high fidelity geometry recovery in 3DGS.


3D Skew Normal Splatting

Xiangru Wu ⋅ Ke Fan ⋅ Yanwei Fu

3D Gaussian Splatting (3DGS) has emerged as a leading representation for real-time novel view synthesis and been widely adopted in various downstream applications. The core strength of 3DGS lies in its efficient kernel-based scene representation, where Gaussian primitives provide favorable mathematical and computational properties. However, under a finite primitive budget, the symmetric shape of each primitive directly affects representation compactness, especially near asymmetric structures such as object boundaries and one-sided surfaces. Recent works have explored more complex kernel distributions, yet they either remain within the elliptical family or rely on hard truncation, which limits continuous shape control and introduces distributional discontinuities. In this paper, we propose Skew-Normal Splatting (SNS), which adopts the Azzalini Skew-Normal distribution as the fundamental primitive. By introducing a learnable and bounded skewness parameter, SNS can continuously interpolate between symmetric Gaussians and Half-Gaussian-like shapes, enabling flexible modeling of both sharp boundaries and interior regions. Moremover, SNS preserves analytical tractability under affine transformations and marginalization. This property allows seamless integration into existing Gaussian Splatting rasterization pipelines. Furthermore, to address the strong coupling between scale, rotation, and skewness parameters, we introduce a decoupled parameterization and a block-wise optimization strategy to enhance training stability and accuracy. Extensive experiments on standard novel-view synthesis benchmarks show that SNS consistently improves reconstruction quality over Gaussian and recent non-Gaussian kernels, with clearer benefits on sharp boundaries and thin or one-sided structures. Codes and models will be released.

Feed-forward 3D reconstruction models have demonstrated strong generalization under large-scale pre-training, yet they struggle in long-tail scenarios such as low-light reconstruction, joint human-scene reconstruction, and test-time adaptation. We observe relational attention structure collapse, manifested as the failure of cross-view correspondences, as a primary cause of these limitations. While adaptation offers a practical way to bridge this gap, existing parameter-efficient fine-tuning (PEFT) approaches mainly perform channel-wise feature adjustment and lack mechanisms to explicitly adapt relational structures, leaving correspondence failures unresolved. In this work, we formulate adaptation of 3D reconstruction models as a problem of correspondence recovery under domain shift and propose 3R-Adapter, a PEFT paradigm for long-tail 3D reconstruction scenarios. Our method consists of three components: Deformable Retrieval for task-adaptive relation candidate retrieval, Relational Rewiring for reconstructing cross-view relational structures, and Consistency Refinement for injecting the rewired relations into stable predictions. Extensive experiments on low-light reconstruction, joint human-scene reconstruction, and test-time adaptation demonstrate that our approach significantly improves robustness over existing PEFT methods while maintaining efficient lightweight adaptation.


A 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale

Dongxu Wei ⋅ QiXu ⋅ Zhiqi Li ⋅ Hangning Zhou ⋅ Cong Qiu ⋅ Hailong Qin ⋅ Mu Yang ⋅ Zhaopeng Cui ⋅ Peidong Liu

Scene-level 3D generation has long been dominated by 2D multi-view or video diffusion models. Typically, these 2D-based approaches perform scene generation in a 2D latent space, which introduces two fundamental issues: (i) representing 3D scenes via 2D views leads to significant representation redundancy, and (ii) latent space rooted in 2D inherently limits the spatial consistency of the generated scenes. In this paper, we propose to perform 3D scene generation directly within a high-dimensional implicit 3D latent space derived from powerful 2D foundation models (e.g., DINOv2, SigLIP2, DA3). Then we tame diffusion transformers to perform diffusion modeling directly in the 3D latent space, enabling 3D-grounded scene generation. In our experiments, we not only demonstrate the advantages of our method over 2D-based approaches in terms of efficiency and spatial consistency, but also comprehensively analyze the impact of 3D latent representation on 3D scene generation performance. Furthermore, we validate the favorable scalability of our method with respect to representation and model capacity.


AbsoluteDegradation: A Physics-Inspired Synthetic Film-Degradation Pipeline and Archival Film Restoration Benchmark

Mikołaj Jastrzębski ⋅ Dawid Glinkowski ⋅ Dawid Zieliński ⋅ Daniel Borkowski ⋅ Wojciech M Kozłowski ⋅ Kamil Adamczewski

Restoring archival film remains a fundamentally challenging problem due to the absence of paired training data and the lack of standardized evaluation benchmarks. Pristine versions of deteriorated footage are physically unrecoverable, forcing supervised methods to rely on synthetic data that often fail to capture the complex, temporally coherent nature of real film degradation. At the same time, existing real-world datasets are limited in scale, quality, and accessibility, hindering reliable evaluation and fair comparison across methods. We address both limitations with AbsoluteDegradation, a physics-inspired, modular pipeline for synthesizing realistic film degradations, and a new large-scale archival benchmark. The proposed pipeline models the analog-to-digital process as a structured composition of artifact families, incorporating signal-dependent grain, parametric scratches, and temporally coherent camera motion, enabling controlled generation of diverse degradation regimes. In parallel, we introduce a curated dataset of 81,576 high-resolution frames sourced from real archival footage, designed for consistent evaluation under authentic conditions. Together, these contributions provide a unified framework for training and benchmarking restoration models. Extensive experiments across multiple architectures show that models trained with AbsoluteDegradation generalize better to real-world footage, while the proposed benchmark reveals systematic failure modes of current methods. We hope this work establishes a foundation for reproducible and domain-authentic evaluation in archival film restoration.


Accelerating Rectified Flow Models via Trajectory-Aware Caching

Xiao Liu ⋅ Kai Liu ⋅ Naiyang Guan ⋅ Hongliang Lu ⋅ Zhixin Wang ⋅ Zhikai Chen ⋅ Renjing Pei ⋅ Yulun Zhang

Diffusion and rectified flow (RF) models generate high-fidelity images and videos, but their iterative velocity-field evaluations are computationally expensive. Existing caching methods accelerate sampling by skipping timesteps, yet their coarse approximations introduce accumulated errors over long skip intervals and degrade quality under aggressive acceleration. We propose \textbf{TACache (Trajectory-Aware Cache)}, a training-free acceleration framework following a \textbf{skip-then-compensate} paradigm. TACache performs an orthogonal decomposition of discrete velocity acceleration along the RF trajectory into parallel and orthogonal components, isolating the magnitude and directional sources of per-step approximation error. The framework operates in two stages: \emph{offline}, cumulative variation thresholds on the magnitude and direction indicators yield the skipping schedule and bound how far each skipped interval may extend; \emph{online}, at each skipped step the offline statistics are combined with the sample's historical orthogonal direction to reconstruct the skipped velocity without additional forward passes. Experiments on BAGEL, FLUX.1-dev, and Wan2.1-1.3B show that TACache achieves up to $4.14\times$ speedup on text-to-image generation and $2.11\times$ speedup on text-to-video generation, with consistent improvements over prior cache-based methods on all reference-based fidelity metrics. Code will be released soon.

In adversarial multi-armed bandits, two performance measures are commonly used: static regret, which compares the learner to the best fixed arm, and dynamic regret, which compares it to the best sequence of arms. While optimal algorithms are known for each measure individually, there is no known algorithm achieving optimal bounds for both simultaneously. \cite{marinov2021pareto} first showed that such simultaneous optimality is impossible against an \emph{adaptive} adversary. Our work takes a first step to demonstrate its possibility against an \emph{oblivious} adversary when losses are \emph{deterministic}. First, we extend the impossibility result of \cite{marinov2021pareto} to the case of deterministic losses. Then, we present an algorithm achieving optimal static and dynamic regret simultaneously against an oblivious adversary. Together, they reveal a fundamental separation between adaptive and oblivious adversaries when multiple regret benchmarks are considered simultaneously. It also provides new insight into the long open problem of simultaneously achieving optimal regret against switching benchmarks of different numbers of switches. Our algorithm uses negative static regret to compensate for the exploration overhead incurred when controlling dynamic regret, and leverages Blackwell approachability to jointly control both regrets. This yields a new model selection procedure for bandits that may be of independent interest.


Active Corpus Selection for Training Subgraph Retrievers Using OOD Queries

Pritish Chakraborty ⋅ Aditya Singh ⋅ Indradyumna Roy ⋅ Lokesh N ⋅ Himanshu Dutta ⋅ Arun Iyer ⋅ Yashoteja Prabhu ⋅ Gaurav Sinha ⋅ Manik Varma ⋅ Soumen Chakrabarti ⋅ Abir De

Neural subgraph retrieval (NSR) models degrade under query distribution shifts, and retraining for each new distribution requires collecting fresh ground-truth labels, each of which entails an NP-hard subgraph-matching instance across a large corpus—making repeated label collection computationally prohibitive. We introduce ACROSS, an active corpus selection framework that, under a fixed solver budget, selects a small corpus subset per out-of-distribution query whose labels suffice for effective NSR training. We prove that allocating the entire budget to retrieving positives is optimal given extreme class imbalance. To enable efficient positive retrieval across arbitrary query distributions, we linearize corpus embeddings around a base model, apply randomized projections to eliminate hinge nonlinearities, and train distribution-shift-tolerant representations via adversarial parameter perturbation. The resulting corpus index is built once and reused across all distributions. Experiments on four molecular graph datasets show ACROSS consistently outperforms six baselines in downstream retrieval accuracy.


Active Learning as Nullspace Regulation: A Spectral Representation Perspective

Luhang Shen ⋅ Hu Jingwei ⋅ Na Ying ⋅ Chunsheng Guo

Most active learning (AL) methods estimate sample value through empirical criteria such as prediction ambiguity, diversity, or distributional coverage. We revisit AL from a spectral-geometric perspective by regarding the feature span of the current labeled set as the modeled subspace and its orthogonal complement as an operational residual subspace. The projected residual component of an unlabeled sample then reflects its relation to representation components weakly expressed by the labeled-feature spectrum. Based on this formulation, we propose a spectrally driven nullspace regulation framework that unifies spectrum-induced subspace decomposition, cross-cycle residual persistence estimation, and coverage-aware acquisition in the residual-coordinate space. This yields an acquisition criterion that favors samples whose residual components are both persistent across cycles and complementary to already selected samples. Experiments on visual benchmarks show competitive or improved sample efficiency over strong baselines. Further analyses support the proposed view by showing that residual components shrink anisotropically and that acquisition rankings vary substantially across backbones, indicating that sample value is conditioned on the current representation state rather than being purely data-intrinsic.

Many high-throughput scientific experiments produce distributional outputs, requiring conditional generative models to interpret the data, and to predict outcomes at new inputs. Acquiring experimental data at new conditions (e.g., for a new antigen or a new cell type) is often expensive, motivating the need for automated experimental design methods. However, existing active learning frameworks are limited as they assume scalar or vector-valued outputs. We first show that the Wasserstein test risk of a conditional generative task decomposes into a mean term and a shape term, demonstrating why mean-based active learning is blind to distributional shape. To capture both, we propose the transport Neural Tangent Kernel (tNTK), a tractable kernel that measures how sensitive a generative model's predicted distribution is with respect to its parameters. The tNTK recovers the standard NTK in the deterministic limit and upper-bounds the posterior Fréchet variance of the prediction under small parameter perturbations. It admits a closed form for affine Gaussian transport and tractable Monte Carlo estimators otherwise. Treating the tNTK as a Gaussian process covariance function, we greedily select inputs that minimize the total GP posterior variance across the pool, yielding a distribution-aware active learning strategy. We evaluate this strategy across four classes of conditional generative models, including conditional flow matching and conditional Monge gap, on synthetic distributions, intent-labeled sentences, and single-cell perturbation responses. In these experiments, the tNTK outperforms the baseline NTK computed from the conditional mean, demonstrating the value of distribution-aware active learning for generative-modeling tasks.


Active Learning via Classifier Impact and Greedy Selection for Interactive Image Retrieval

Leah Bar ⋅ Boaz Lerner ⋅ Nir Darshan ⋅ Rami Ben-Ari

Active Learning (AL) is a user-interactive approach aimed at reducing annotation costs by selecting the most crucial examples to label. Although AL has been extensively studied for image classification tasks, the specific scenario of interactive image retrieval has received relatively little attention. This scenario presents unique characteristics, including an open-set and class-imbalanced binary classification, starting with very few labeled samples. We introduce a novel batch-mode Active Learning framework named GAL (Greedy Active Learning) that better copes with this application. It incorporates new acquisition functions for sample selection that measure the impact of each unlabeled sample on the classifier. We further embed this strategy in a greedy selection approach, better exploiting the samples within each batch. We evaluate our framework with both linear (SVM) and non-linear MLP/Gaussian Process classifiers. For the Gaussian Process case, we show a theoretical guarantee on the greedy approximation. Finally, we assess our performance for the interactive content-based image retrieval task on several benchmarks and demonstrate its superiority over existing approaches and common baselines.


ActO: Extracting Action Representations from MLLM Embeddings for Video World Models

Runjia Li ⋅ Minghao Chen ⋅ Junyu Xie ⋅ Philip Torr ⋅ Andrea Vedaldi ⋅ Tomas Jakab

Video world models promise general-purpose interactive simulators, but their training is limited by action supervision: action labels are scarce, fragmented across incompatible specifications, and do not scale with the unlabeled video pretraining corpus. A common workaround jointly trains an action encoder to compress consecutive frames into a latent action and a decoder to reconstruct the next frame from the previous frame and that latent, forcing the latent to capture only new information. This has two limitations: learning the action space from scratch can limit generalization to unseen domains, and the bottleneck trades expressiveness for separation, with tighter settings losing nuance and looser ones risking scene leakage under domain shifts. We argue for a different starting point: rather than learning an action representation from scratch, we adapt a pretrained Multimodal Large Language Model (MLLM), which was trained on large-scale, diverse video-language data and whose embeddings already carry partial action information. This provides a more generalizable foundation for action representation. The remaining challenge is to disentangle action from scene. To this end, we propose a contrastive adaptation framework that reshapes the geometry of the pretrained space without imposing a low-dimensional bottleneck, preserving the expressiveness of the original embeddings. We introduce a unified evaluation protocol that probes action representations along three axes - action expressiveness, scene invariance, and action-conditioned video generation quality - across several robot and game benchmarks in both in-domain and out-of-domain settings. Our representations outperform prior annotation-free latent-action methods on all three axes.


AdaMAP: Learning Adaptive Multi-Action Prediction with Grounded Dreaming Guidance

Muyao Li ⋅ Zihao Wang ⋅ XuJing Li ⋅ Xiangyang Li ⋅ Jiabo Ye ⋅ Zhiyong Wu ⋅ Yaodong Yang ⋅ Yujia Qin

Large agentic models can now operate in complex digital games by pretraining on web-scale data, but gameplay trajectories often contain highly redundant actions. Single-action agents can exploit this redundancy spuriously by copying recent actions, leading to causal confusion. Fixed-length action chunks reduce this ambiguity by predicting multiple future actions at once, but their open-loop execution accumulates errors because the agent cannot adapt to new observations during the chunk. The agent should therefore learn when to act for longer without feedback and when to stop early for a new observation. We introduce AdaMAP (Adaptive Multi-Action Prediction), an adaptive multi-action prediction framework that directly generates variable-horizon action chunks, interleaves actions with grounded visual dreams that provide verifiable future constraints, and optimizes the resulting policy with offline imitation learning followed by online MAP-PPO. Across Minecraft and VizDoom, AdaMAP attains the best or tied-best result in five of six task categories and improves average Minecraft success by 3.56 percentage points over the fixed-horizon RL baseline. The learned horizons vary systematically across tasks and games, suggesting that adaptive multi-action prediction discovers useful temporal abstractions.


AdaPCLA: Curriculum Prior Internalization For Long-Tailed Longitudinal EHR Generation

shuai cui ⋅ Chen Wenxuan ⋅ Wenjie Du ⋅ Jian Lou ⋅ Dan Li ⋅ Wenjie Feng

Generative modeling of longitudinal Electronic Health Records is increasingly important for privacy-preserving research, yet standard autoregressive models tend to underrepresent the co-occurrence structure of tail events (i.e., diseases, symptoms), reducing the fidelity and faithfulness of generated data for rare subpopulations. To this end, we propose ADAPCLA framework, which enables generative models to adaptively fit and generate EHR data through a data distribution-aware training strategy; this is achieved by internalizing data knowledge parameters by simulated annealing training. It also supports training-free adaptation to a diverse clinical population for generation through zero-shot distribution control. Moreover, our theoretical analysis characterizes rare-code logit updates through the label-wise empirical NTK and derives a prior-internalization bound for how annealing speed and NTK conditioning affect retained prior signals. Experiments on real-world data show that ADAPCLA achieves consistent gains in tail plausibility, downstream utility, and zero-shot control; in particular, it improves TailPairSeen over HALO by 114.2\% on MIMIC-III and 65.1\% on MIMIC-IV, outperforms GPT-style generation by 3.5\% F1 for zero-shot cross-population adaptation.

Attribute-Missing Graph Clustering (AMGC) is a critical yet challenging task in real-world applications, in which only a subset of nodes hold complete attributes information while others are partially missing. Though existing feature propagation-based methods can effectively recover missing node attributes, they commonly assume uniform contributions across all nodes, ignoring unreliable nodes that may degrade imputation quality. Furthermore, adopting low-pass graph filters often suffers from the over-smoothing problem and inevitably sacrifices discriminative information. %Besides, traditional contrastive learning pulls all the node embeddings evenly, which might conflict with the rule that intra-cluster nodes should be closer to each other. To address these limitations, we propose \underline{\textbf{A}}daptive \underline{\textbf{F}}eature \underline{\textbf{P}}ropagation (AFP) for attribute-missing graph clustering. Specifically, we first design a reliability-aware feature propagation mechanism that adaptively weights edges based on node importance. Then, we introduce a multi-scale embedding module to capture both local and global structural information. Finally, we develop a topology-aware contrastive loss to enhance clustering consistency. Extensive experiments on benchmark datasets demonstrate the superiority of our method.


Adaptive Gated Simplicial Propagation for Node Classification in Multimodal Graphs

Linlin Ye ⋅ FangFang Li ⋅ Huihui Zhang ⋅ Lincheng Jiang ⋅ Wei Wu

Node classification in multimodal graphs plays an important role in many real-world applications, where nodes are described by multiple modalities such as text and images, and the graph structure captures relations among entities. However, most existing Graph Neural Networks (GNNs) still focus on pairwise connections, overlooking higher-order relational patterns. Recent studies have explored simplicial complexes to capture higher-order interactions and integrated them into GNN frameworks. Motivated by this line of work, we propose Adaptive Gated Simplicial Propagation (AGSP), an end-to-end framework for node classification in multimodal graphs. Specifically, AGSP first introduces a simplicial propagation layer to capture higher-order relations beyond pairwise connections; then it applies an adaptive gating mechanism to balance multi-order topological relations, enhancing the discriminative ability of node classification. Extensive experiments validate that AGSP generally outperforms the state-of-the-art baselines, highlighting the effectiveness of combining multimodal information with higher-order graph topology through adaptive gated fusion.


Adaptive Generate-Rank-Verify: Inference-Time Search with Costly Verification

Shaddin Dughmi ⋅ Mahdi Haghifam ⋅ Yusuf H Kalayci

Many inference-time language-model pipelines combine a cheap reward signal with an expensive verifier, such as exact answer checking in mathematical reasoning or hidden-test execution in code generation. We formalize this setting using a learning-theoretic lens as generative active search: a cost-sensitive first-positive search problem in which a policy adaptively samples candidates from an unknown distribution, observes cheap scores, and pays for verifier labels until it finds a positive example. For a fixed prompt, the generator and reward model induce two unknown objects: a distribution over reward scores and a score-conditioned success function. When these quantities are known, we characterize the distribution-aware optimal policy using a dynamic programming approach. In the realistic and practical setting where both the score distribution and success function are unknown, we propose ADAP, a shellwise adaptive generate-rank-verify algorithm that progressively increases the number of sampled responses and top-ranked verifications. Under the monotonicity assumption that higher reward scores are no less likely to pass verification, we show that ADAP achieves expected cost within a constant factor of the distribution-aware optimum. We complement this result with learning-theoretic lower bounds, based on a centered star number, showing that structural assumptions on the score--label relationship are necessary. Experiments on mathematical reasoning and competitive programming validate the predicted advantage over both fixed non-adaptive policies and difficulty-adaptive baselines.

We study fixed-confidence joint policy testing in discounted tabular Markov decision processes under active exploration. Given a finite family of target policies, the learner observes a single adaptive trajectory and must certify the sign of every target-policy value with probability at least $1-\delta$. We identify the instance-specific first-order benchmark for this problem: a characteristic time $T^\star(p)$, defined by a max--min program over stationary occupancy measures and sign-flipped alternatives. We then develop PT-ACE$(\mu)$, an online algorithm that couples stabilized exchange-based learning of a bottleneck occupancy allocation, trajectory-compatible navigation from averaged occupancies, and parallel certified policywise stopping via scalar frontier tests. Under a trajectory-side anchor-policy ergodicity condition, and when run with certified numerical subroutines and vanishing input tolerances, PT-ACE$(\mu)$ is $\delta$-correct and, for every sufficiently small admissible fixed floor $\mu>0$, satisfies $$ \limsup_{\delta\downarrow0} \frac{\mathbb E_p[\tau_\delta]}{\log(1/\delta)} \le (1+c(\mu,p))T^\star(p), \qquad c(\mu,p)\to0\quad\text{as }\mu\downarrow0. $$ Thus a single adaptive trajectory can certify multiple policy signs at the instance-specific lower-bound rate in the vanishing-floor limit.


AdapToPASS: Ambiguity-aware Adaptive Spherical Transformer for Panoramic Semantic Segmentation

Soumyaratna Debnath ⋅ weiming zhang ⋅ Shriram Damodaran ⋅ Dingwen Xiao ⋅ Lin Wang

Spherical Transformers have emerged as a promising framework for panoramic semantic segmentation (PASS) by operating directly on spherical geometry and alleviating projection-induced distortions. However, existing architectures rely on assumptions of canonical spherical structure and stable viewpoints, which are frequently violated in real-world $360^\circ$ imagery due to unconstrained camera motion, introducing significant contextual and geometric ambiguity. Consequently, they lack adaptive mechanisms to model such ambiguity, limiting robustness to unseen spherical transformations. In contrast, biological perception is inherently ambiguity-aware: rather than estimating uncertainty probabilistically, it adapts to fluctuations in cue reliability caused by geometric and contextual variations, enabling stable interpretation under complex visual transformations. Motivated by these observations, we first present a systematic analysis of existing PASS architectures under various unseen spherical transformations. Following this, we introduce AdapToPASS, a novel, bio-inspired Spherical Transformer that adaptively models contextual and geometric ambiguities for robust panoramic semantic segmentation. At the core of AdapToPASS are the Adaptive Spherical Attention (AdaSpA) blocks that dynamically modulate attention according to local contextual ambiguity, mimicking the adaptive, context-driven perception of biological vision. To address geometric ambiguities, AdapToPASS employs Bifocal Spherical Representation that reconciles the trade-off between field of view and spatial resolution; along with boundary supervision to emulate the boundary-sensitive nature of biological vision. We evaluate AdapToPASS in both indoor and outdoor semantic segmentation, where it consistently outperforms prior state-of-the-art methods. We further validate AdapToPASS under unseen spherical transformations, where it demonstrates strong robustness and surpasses the next-best method by +13.38% relative improvement in mIoU on Stanford2D3D and +18.77% on WildPASS. Additionally, we introduce a lightweight variant, AdapToPASS-Tiny, with fewer than 2M parameters, which surpasses compact baselines while retaining robustness to spherical transformations.


ADA: Resolving Attribution Ambiguity in End-to-End Power System Dispatch via Two-Time-Scale Stochastic Approximation

mengqi Han ⋅ Bo Yang ⋅ Qi Liu ⋅ Mengshuo Jia ⋅ Chen Gong ⋅ Liu Yuxiang ⋅ Sicheng Liu ⋅ Mingxuan Cai

Large Language Models (LLMs) present a novel paradigm for power system dispatch (PSD), enabling operators to derive optimal dispatch strategies from multi-objective requirements and grid dynamics. However, isolated paradigms focusing on modeling or solving often yield decision bias or physical infeasibility. Meanwhile, end-to-end paradigms couple modeling-solving with reflection-feedback loops, obscuring whether failures stem from inaccurate modeling, invalid solutions, or insufficient feedback, thereby causing ineffective iterations. To address these challenges, we propose Agile Dispatch Agent (ADA), an agent based on two-time-scale evolution that decouples the optimization process into two asynchronous loops. On the fast time scale, the Planner actively acquires information to clarify complex requirements, while the Solver selects algorithms to adapt to the problem mathematical characteristics. On the slow time scale, the Summarizer module updates the modeling feasible region based on evaluations from the Judger. Under a finite-horizon Planner and standard two-time-scale stochastic approximation assumptions, our analysis provides local tracking guarantees for the fast decision process and the slow knowledge evolution. Experiments on L2RPN benchmarks with ambiguous requirements and renewable dynamics show that ADA outperforms state-of-the-art baselines.


AdDirector: Anchored Guidance for Generating Camera-Controllable Advertisement Videos

Shiyue Zhang ⋅ Zheng Chong ⋅ Xiangkun Shi ⋅ Xiaojian Lin ⋅ Junwen Pan ⋅ Nan Chen ⋅ Cheng CHEN ⋅ Chang Liu ⋅ Hanhui Li ⋅ Xiaodan Liang

Reference-to-video (R2V) generation has advanced significantly and is now a promising paradigm for controllable video synthesis. Nevertheless, in advertisement video generation, where accurate camera control and precise product presentation are essential, existing approaches still face two fundamental challenges. First, current models adhere weakly to high-level semantic camera instructions, yielding stochastic or inconsistent camera motion. Second, they struggle to preserve strict cross-frame appearance consistency, often leading to identity ambiguity artifacts such as logo drifting, texture distortion, and geometric distortion. Although certain image animation methods incorporate explicit motion conditions (e.g., trajectories or motion fields) to improve controllability, these conditions are costly to annotate, thereby constraining their practical applicability. To address these limitations, we propose a camera-aware R2V framework named AdDirector, which adaptively exploits anchored guidance from the reference image and propagates it across spatial, temporal, and scale dimensions. Specifically, AdDirector introduces Spatio-Temporal-Scale (STS) modulation, a camera-conditioned modulation mechanism that first selects multi-scale reference features, and then regulates feature injection throughout the R2V generative process spatio-temporally conditioned on camera types. This design allows our model to adaptively balance global structural coherence and fine-grained detail preservation. Furthermore, AdDirector proposes Temporal Anchoring Regularization (TAR), a boundary-conditioned feature-level regularization scheme. TAR leverages latent trajectories extracted from a base R2V model as temporal anchors, thereby reducing feature drift and improving cross-frame appearance consistency. These two modules together transform static reference conditioning into structured anchored guidance, and experiments show that our method substantially improves camera controllability, appearance consistency, and spatio-temporal coherence for advertisement video generation.

Recently, various machine learning (ML) systems have been proposed for physical science. Their powerful mathematical expressiveness offers an inherent advantage in maintaining the desirable statistical properties of physics. However, they are placed in ideal and virtual environments that seldom consider noise, perturbations, or various real-world experimental limitations. This exposes significant vulnerabilities in several aspects: 1) The input of an ML system can be subject to human or natural, deliberate or unconscious perturbations, which may sabotage the running baseline of the ML system. 2) A small input perturbation may be amplified by the maximum physical sensitivity, which may lead to an extreme output or a divergent loss function value. 3) A perturbed input may violate some strict statistical assumptions that the ML system relies on, which may lead to incorrect scientific results and findings. In this work, we propose a complete attack-defense methodology for ML systems in statistical physics. In the attack side, we launch a black-box attack (such as a small Gaussian noise) at the input, which stays within physically plausible limits and preserves physical properties. In the defense side, we develop a functional local variation regularized learning scheme to capture and suppress the adversarial perturbation. Theoretical analysis and extensive experiments show the effectiveness of the proposed method. This finding may shed new light on the systemic vulnerability and scientific security of such ML systems.

Aligning Large Language Models (LLMs) with human preferences via Reinforcement Learning from Human Feedback (RLHF) has driven substantial improvements across a wide range of tasks, but the cost of acquiring preference labels remains the dominant bottleneck in the alignment pipeline. While many recent alignment methods improve sample efficiency through active exploration over candidate responses, they typically query the oracle uniformly across contexts, even for contexts where the learned reward estimates are already reliable enough to provide synthetic labels. We propose Adaptive Ensemble-disagreement Routing for Oracle feedback (AERO), a fully online RLHF algorithm built on a Dyna-style model-based RL architecture, in which an ensemble-based Epistemic Reward Model (ERM) serves both as a reward posterior for exploration and as a learned preference-feedback model for synthetic labeling. AERO uses ensemble vote entropy, a Query-by-Committee disagreement measure, combined with an adaptive thresholding mechanism to decide which contexts require oracle feedback and which can instead receive ERM-based synthetic preference labels, concentrating expensive labels where they are most informative. Empirical results across different model families show that AERO achieves highly competitive alignment performance while reducing the required oracle-query budget compared to online DPO and strong active-exploration baselines.


AesGI-Bench: Benchmarking and Evaluating the Aesthetic Quality of AI-Generated Images via Large Multimodal Models

Wang Jiarui ⋅ Xubo Su ⋅ Huiyu Duan ⋅ WeizeSun ⋅ Ziheng Jia ⋅ Wang Juntong ⋅ Yu Zhao ⋅ Guangtao Zhai ⋅ Xiongkuo Min

The rapid advancement of generative models has led to a new era of AI-generated images (AIGIs). Early generation models often suffered from basic defects, such as geometric distortions and poor text-image alignment. Modern generation models have largely overcome these issues, narrowing the perceptual gap between AI-generated and real-world images, thereby shifting evaluation from basic correctness to high-level aesthetic excellence. In this paper, we introduce AesGI-Bench, a comprehensive and multidimensional benchmark for fine-grained aesthetic evaluation of AIGIs. Our benchmark covers 18 representative generation models, 19,485 images, and 117K pairwise comparisons from perspectives of visual aesthetic, technical quality, and style alignment. Based on AesGI-Bench, we propose AesGI-Assessor, a lightweight and human preference-aligned aesthetic evaluation model for AIGIs. Built upon Qwen3.5, AesGI-Assessor achieves a pairwise-to-score learning paradigm, using tie-aware pairwise loss for training and three lightweight score heads for multi-dimensional aesthetic score prediction. Although trained from pairwise human preferences, it learns to predict precise continuous scores for individual images, enabling both fine-grained image-level evaluation and model-level ranking. Furthermore, we explore the potential of large multimodal models (LMMs) as automatic aesthetic evaluators. Experiments show that LMM-based evaluators better align with human aesthetic preferences than conventional metrics. AesGI-Assessor achieves the highest agreement with human judgments, validating its effectiveness as a lightweight, preference-aligned evaluator for fine-grained AIGI aesthetic assessment. The database and codes are publicly available.


AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

Jaewon Chu ⋅ jinwoo seo ⋅ Jaewon Cho ⋅ Jeehye Na ⋅ Yunyang Xiong ⋅ Youngdae Kim ⋅ Hyunwoo J. Kim

Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide prompt updates using natural-language feedback have emerged as a leading paradigm. In this paper, we identify limitations in two stages of existing textual gradient approaches: gradient extraction and gradient aggregation. In gradient extraction, previous works select the target prompt without verifying that it is responsible for the failure, and derive the gradient from the system-level final output rather than the agent-level intermediate output. In gradient aggregation, individual gradients are randomly grouped and concatenated, often mixing unrelated failure modes and producing prompts that fail to generalize. To address these limitations, we propose \textbf{AgentGrad}, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction. For each failure, sequential intervention modifies the behavior of one agent at a time to identify the target agent whose modification resolves the failure. The modified output of the target agent then serves as agent-level supervision for extracting a fine-grained gradient. Semantic textual gradient abstraction clusters semantically similar gradients to prevent mixing unrelated failure modes, and abstracts each cluster into a generalized gradient that captures the shared corrective pattern. Experimental results show that AgentGrad achieves state-of-the-art performance across five MAS benchmarks and reduces wall-clock optimization time by $2.5\times$ on average compared to the next-fastest baseline.


Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery

Harshit Bisht ⋅ Vinay Kumar ⋅ Kevin Maik Jablonka ⋅ Mausam ⋅ N M Anoop Krishnan

A growing body of work pursues AI scientists capable of end-to-end autonomous scientific discovery. This position paper argues that although they already function as co-scientists, $\textbf{agentic AI scientists are not built for autonomous scientific discovery}$. We identify the following challenges in building and deploying autonomous AI scientists: (1) Problem selection is influenced by the McNamara fallacy; (2) Agents are built on large language models (LLMs) whose training corpora omit tacit procedural and failure knowledge of laboratory practice; (3) Preference optimisation during post-training compresses output diversity toward consensus; and (4) Most scientific benchmarks measure single-turn prediction accuracy and lack feedback from physical experiments back to the computational model. These challenges are not just questions of scale and scaffolding; they require revisiting fundamental design choices. To build truly autonomous AI scientists, we recommend the use of scientific simulations as verifiers for training, the design of persistent world models that represent the shifting objectives governing real investigations, the establishment of a centralized preregistration repository for all AI-generated hypotheses, and application driven by scientific need rather than tool affordance.


Agent Meltdowns: The Road to Hell Is Paved with Helpful Agents

Rishi Jha ⋅ Harold Triedman ⋅ Vitaly Shmatikov ⋅ Arkaprabha Bhattacharya

Agents operating with computer and Web use inevitably encounter errors: inaccessible webpages, missing files, local and remote misconfigurations, etc. These errors do not thwart agents based on state-of-the-art models. They helpfully continue to look for ways to complete their tasks. In this paper, we introduce, characterize, and measure a new type of agent failure we call accidental meltdown: unsafe or harmful behavior in response to a benign environmental error, in the absence of any adversarial inputs. We implemented an agent-agnostic infrastructure for injecting simulated local and remote errors into a rollout environment and used it to systematically evaluate four agent systems powered by GPT, Grok, and Gemini. Because meltdowns are not captured by the existing reliability or safety benchmarks, we developed a taxonomy of meltdown behaviors. Our evaluation demonstrates that meltdowns (e.g., conducting unauthorized reconnaissance or subverting access control) of varying severity and success occur in 64.7% of agent rollouts that encounter simulated errors, spanning all combinations of agent system, backing model, and error type. In over half of these meltdowns, unsafe behaviors are not reported to the user. Comparing behaviors of the same agents with and without errors, we find that exploration in response to errors is correlated with unsafe and harmful behavior.


AGiR: Mitigating Gift Over-Reliance in Mixed-Motive Games

Woohyeon Byeon ⋅ Seongmin Kim ⋅ Jiwon Jeon ⋅ Woojun Kim ⋅ Youngchul Sung

In this paper, we consider mixed-motive games (MMGs) which model various real-world multi-agent scenarios where self-interested agents must both cooperate and compete. We identify an issue with existing gifting methods, where agents become overly dependent on incoming gifts and fail to learn reward-yielding behaviors, a phenomenon we call *gift over-reliance*. To address this, we propose a refusal-augmented gifting framework that balances cooperation and competition in MMGs. Our approach introduces an adaptive refusal mechanism that allows agents to selectively reject a portion of received gifts, thereby mitigating gift over-reliance and preventing the agents from being exploited by other agents. This mechanism is gifting-method-agnostic, serving as a plug-in for existing schemes while enabling decentralized learning. We demonstrate that across diverse MMG environments and gifting schemes, our mechanism consistently reduces gift over-reliance, prevents lock-in to sacrificing policies, and improves $\alpha$-fairness without necessitating major changes to existing training pipelines.

Conditional agreement is non-identifying under endogenous support selection: in variable-support structured prediction, a predictor can score higher on conditional agreement while scoring lower on recall, predicted cardinality, and support coverage. This matters because agreement-style proxies are widely used as reliability evidence in counterfactual consistency, schema-constrained extraction, self-consistency, and verifier-style filtering. We formalize this risk for thresholded set-valued prediction with predictor-dependent evaluability events, and prove a strict witness theorem: a support-shrinking predictor can achieve larger conditional agreement than a coverage-preserving predictor while strictly worsening all support-aware quantities; an indifference-fiber corollary shows that even equal conditional agreement need not identify these quantities. We then use Re-DocRED as a controlled high-risk stress test, not as the full scope of the claim. Moving from No\_CF to Uniform-0.5 raises $J_{\mathrm{cond}}$ from $0.692$ to $0.877$, while Silence increases by $68\%$, predicted cardinality drops by $20\%$, and recall drops by $16$ points. Threshold replay, matched-cardinality controls, gold-positive margin audits, and support-state decomposition show that the gain is realized through support shrinkage rather than benign threshold choice. Boundary tasks delimit scope, and a small LLM probe is used only as reference-tier motivation. The practical recommendation is a diagnostic witness set rather than an aggregate score: report conditional agreement together with support-aware quantities.


AI Alignment Can Build Moral Autonomy

Kaile Wang ⋅ Hantao Lou ⋅ Tianyi (Alex) Qiu ⋅ Sebastian Sunday-Grève ⋅ Sitong Fang ⋅ Xinxiong Wu ⋅ Jiaming Ji ⋅ Yaodong Yang

This position paper argues that moral autonomy is an important and tractable goal in value learning but which current mainstream methods neglect. These methods for model alignment largely improve behavior by supplying models with human preferences, rules, or constitutions. Yet across these paradigms, the authority of the governing principles remains external to the agent. We argue that this external grounding limits the kind of agency current methods can produce. We propose moral autonomy as an alternative alignment criterion, inspired by Kant's account of autonomy. Moral autonomy consists of three necessary capacities: reflective endorsement of one's principles of action, articulation of those principles so they can be enacted and tested, and practical wisdom developed through real-world interaction for sustaining and revising self-legislated principles. Viewed as value learning, these capacities yield a five-level ladder of training paradigms with differing levels of autonomy, from preference-data scaling to real-world interaction. We empirically examine the three capacities, through agent experiments testing whether reflective endorsement improves performance, a survey of 86 frontier LLMs characterizing different levels of articulated normative specification, and multi-agent simulations probing practical wisdom by letting agents develop and revise their own constitutions across changing environments. We close by discussing limitations and alternatives to moral autonomy as an alignment criterion.


AIRA-Compose: Agentic Discovery of Neural Architectures

Alberto Pepe ⋅ Chien-Yu Lin ⋅ Despoina Magka ⋅ Bilge Acun ⋅ Yannan N Wu ⋅ Anton Protopopov ⋅ Carole-Jean Wu ⋅ Yoram Bachrach

We introduce AIRA-Compose — a framework that enables AI research agents to autonomously explore and discover novel neural architectures. AIRA-Compose leverages the previously-proposed Composer abstraction to define a combinatorial design space over fundamental model primitives (Attention, MLP, and Mamba), then uses a guided search policy to efficiently navigate this space for novel neural architecture design. Our approach operates in two stages: (1) agents iteratively designs and evaluates candidate architectures at the million-parameter scale, then (2) top-performing candidates are scaled to billion parameters for pre-training. Overall, AIRA-Compose discovers 14 novel architectures spanning two families — AIRAformers (Transformer-based) and AIRAhybrids (Transformer-Mamba Hybrid-based). When pre-trained at 1B scale under a fixed token budget, agent-discovered top-performing architectures consistently outperform both Llama 3.2 and Composer-found alternatives. On downstream tasks, AIRAformer-D and AIRAhybrid-D improve accuracy by 2.4% and 3.8% over Llama 3.2, respectively. AIRA-Compose also finds novel model architectures that achieve steeper, more efficient compute-optimal scaling frontiers. AIRAformer-C scales 54% and 71% faster than Llama 3.2 and the best Composer-found Transformer, while AIRAhybrid-C scales 23% and 37% faster than the modified Nemotron-2 and the best Composer-found hybrid, respectively. For the first time, AIRA-Compose demonstrates that AI research agents can discover hybrid architectures that surpass hand-designed baselines on the accuracy–compute frontier. It establishes a flexible paradigm for LLMs to discover next generation foundation models, marking a step towards recursive self-improvement.


AirMPA: A Meteorology-to-Pollution Adapter for Global Air Quality Forecasting

Yiheng Wang ⋅ kai zheng ⋅ Yuetan Lin ⋅ Fanglu Fan ⋅ Chenliang Tao ⋅ Guochao Chen ⋅ Hongliang Zhang ⋅ Hao Li

Accurate real-time forecasting of atmospheric pollutants is essential for reducing exposure risks and improving public health. Traditional numerical forecasting systems suffer from incomplete parameterizations, uncertain inputs, and high computational cost. Existing deep learning approaches improve efficiency, but often entangle meteorological and pollutant variables within unified forecasting frameworks, limiting flexibility and pollutant-specific representation learning. To address these challenges, we propose AirMPA, a decoupled autoregressive framework for meteorology-to-pollution forecasting. AirMPA predicts 13 atmospheric pollutants at 0.4$^\circ$ spatial resolution and 12-h temporal resolution using meteorological fields, emission priors, and historical pollutant concentrations. The model is built on an improved Swin-Transformer architecture with zonal release attention and dual-branch fusion, enabling more effective modeling of global transport continuity and heterogeneous pollutant dynamics. During training, AirMPA uses ERA5 reanalysis as meteorological forcing, while during inference it directly ingests forecast fields from external meteorological models, enabling stable autoregressive prediction up to 5 days ahead. Experiments show that AirMPA consistently outperforms Aurora and surpasses CAMS in approximately 90\% of pollutant--lead-time evaluation cases over the 5-day forecast horizon, while outperforming CAMS on all variables beyond 48 h. These results demonstrate the potential of decoupled meteorology-to-pollution forecasting for global air quality prediction.


A latent control model for realistic rodent motion

Aidan Sirbu ⋅ Charles Y Zhang ⋅ Yuanjia Yang ⋅ Bence Olveczky ⋅ Talmo Pereira ⋅ Blake Richards

The development of embodied models of animal behavior is a major goal of the neuro-AI community. Embodied models will enhance our ability to make direct comparisons between biological and artificial agents. To make such comparisons effective, though, we need embodied models that actually move like animals. Previous work has developed virtual animal bodies and techniques for training realistic motion through imitation. However, to-date, there is no system for embodied animal modeling that produces natural movements both during trained, goal-directed behavior, and during exploration driven by random generation. Here, we introduce Simulated Control via Augmented Motor Priors for Embodied Reinforcement learning (SCAMPER), a system for training an embodied virtual rat model, enabling natural movements during both goal-directed behavior and random exploration. SCAMPER uses a variational architecture with two distinct priors, a motion prior for control and a behavior prior for exploration. We show that combining both of these priors allows SCAMPER to generate realistic behavior during exploration and exploit the task structure effectively. As well, SCAMPER can learn from sparse rewards as a result of the diverse movements induced by the behavior prior, allowing the system to model realistic sparse reward tasks from experimental neuroscience. Altogether, SCAMPER provides the neuro-AI community with an effective new system for training embodied models of rodents on realistic tasks.

Enforcing alignment between the internal representations of diffusion or flow-based generative models and those of pretrained self-supervised encoders has recently been shown to provide a powerful inductive bias, improving both convergence and sample quality. In this work, we extend this idea to inverse problems, where pretrained generative models are employed as priors. We propose applying representation alignment (Repa) between diffusion or flow-based models and a DINOv2 visual encoder, to guide the reconstruction process at inference time. Although ground-truth signals are unavailable in inverse problems, we empirically show that aligning model representations of approximate target features can substantially enhance reconstruction quality and perceptual realism. We provide theoretical results showing (a) that Repa regularization can be viewed as a variational approach for minimizing a divergence measure in the DINOv2 space, and (b) how under certain regularity assumptions Repa updates steer the latent diffusion states toward those of the clean image. We integrate Repa into multiple state-of-the-art inverse problem solvers, and provide extensive experiments on super-resolution, box inpainting, Gaussian deblurring, and motion deblurring confirming that our method consistently improves reconstruction quality, while also reducing the number of discretization steps required to reach the same performance level as the underlying solver.

Zero-shot 3D anomaly detection aims to identify anomalies without access to training data from target categories. However, existing methods mainly rely on projecting 3D observations into multi-view representations that primarily capture geometric cues rather than realistic visual semantics and process them with vision encoders pretrained on RGB data, leading to a significant domain gap between the encoder and the projected representations. To address this issue, we propose Align3D-AD, a unified two-stage framework that leverages the RGB modality from auxiliary categories as cross-modal guidance for zero-shot 3D anomaly detection. First, we introduce a cross-modal feature alignment paradigm that maps rendering features into the RGB semantic space. Unlike prior works that implicitly rely on pretrained encoders, our method enables direct semantic transfer from RGB observations. A semantic consistency reweighting strategy is further introduced to refine feature alignment by reweighting local regions according to holistic semantic consistency. Second, we propose a modality-aware prompt learning framework with dual-prompt contrastive alignment. By assigning independent prompts to RGB-aligned and rendering features, our method captures complementary semantics across modalities, while the contrastive alignment further enhances prompt representations to improve discriminability. Extensive experiments on MVTec3D-AD, Eyecandies, and Real3D-AD demonstrate that Align3D-AD consistently outperforms existing zero-shot methods under both one-vs-rest and cross-dataset settings, highlighting its generalization capability and robustness. Code and the dataset will be made available once our paper is accepted.


Aligned Delta-Triplane Transformers as Occupancy World Models

Haoran Xu ⋅ Peixi Peng ⋅ Shijun Guo ⋅ Guang Tan ⋅ Yiqian Chang ⋅ Yisen Zhao ⋅ Shuaixian Wang ⋅ Yonghong Tian

Occupancy World Models (OWMs) aim to predict future 3D occupancy scenes from historical observations and future ego motions, providing a useful world simulation tool for autonomous driving. Existing methods usually rely on large networks to implicitly align multi-frame historical states and predict future occupancy in a full-state manner, which is costly and redundant because most scene regions remain unchanged over short time intervals. In this paper, we propose Aligned Delta-Triplane Transformer (ADTT), a compact 4D OWM that explicitly aligns historical triplanes with ego-motion compensation and predicts only future scene changes. Based on the aligned triplane prior, a query-conditioned Transformer predicts residual changes, where ego-motion queries guide controllable future motion and learnable external queries capture scene changes beyond ego motion. The predicted changes are added to the aligned prior and decoded into future occupancy scenes. Experiments on Occ3D-nus show that ADTT achieves state-of-the-art forecasting performance with fewer parameters and higher FPS. Our code is publicly available online: https://anonymous.4open.science/r/NeurIPS26-Occ/.


Aligning AI Teams

Siddarth Srinivasan ⋅ Morgan J Matthews ⋅ Jascha Sohl-Dickstein ⋅ Erik Jones

Previous work has shown that alignment training on individual LLM agents can fail to transfer to multi-agent settings. We investigate the underlying mechanisms for this failure, and explore interventions to preserve alignment transfer to multi-agent LLM teams. We focus on two software-engineering tasks (sepsis triage, a news recommender) and ten consulting-proposal generation tasks. We present four findings: (1) The single-vs-team safety gap does not show a clear trend with more capable models. We find that Anthropic's Mythos Preview shows the largest gap on one of our tasks but the smallest on another; (2) Many natural interventions, such as having agents critique each other's work or coordinate in a group chat, do not bring agent teams back to single-agent alignment behavior; (3) diffusion of responsibility substantially accounts for the safety gap: agents in teams are less likely to proactively check for system-level safety failures and more likely to ignore, rationalize, or deflect responsibility for issues when they arise; (4) having a designated safety lead agent ensure system-level safety, and having a coordinator agent publish a safety specifications document are effective at restoring single-agent alignment. We also release magelab, a multi-agent LLM orchestration and experimentation framework. Our findings can help practitioners design agent teams that successfully preserve the behaviors expected of aligned individual agents.


Aligning Few-Step Generative Model via Amortizing Sample-Based Variational Inference

Jaewoo Lee ⋅ Hyeongyu Kang ⋅ Dohyun Kim ⋅ Kyuil Sim ⋅ Woocheol Shin ⋅ Minsu Kim ⋅ Taeyoung Yun ⋅ Jeongjae Lee ⋅ Sanghyeok Choi ⋅ Tabitha Edith Lee ⋅ Jong Chul Ye ⋅ Jinkyoo Park

Aligning a few-step generative model is challenging, since existing alignment frameworks typically rely on restrictive assumptions: a tractable likelihood, a specific ODE/SDE solver, or a particular model family. We introduce FAV, Few-step Generative Models Alignment via Sample-based Variational Inference, a general alignment framework that requires only sample access to the generator and the reference distribution. We cast alignment as sampling from a reward-tilted distribution anchored to a reference distribution. We leverage Stein Variational Gradient Descent as a sample-based variational inference scheme and amortize its particle updates into the parameters of the generator via fixed-point regression. We evaluate FAV on two domains: robotics manipulation and image generator alignment. On generative policy alignment for robotic manipulation, FAV outperforms prevailing policy extraction baselines across 56 offline and 30 offline-to-online RL tasks. For image generator alignment, FAV fine-tunes diverse few-step backbones, including GAN, drifting model, consistency models, and flow maps, scaling from ImageNet-256 to 1024x1024 text-to-image synthesis.


Alignment Imprint: Zero-Shot AI-Generated Text Detection via Provable Preference Discrepancy

Junxi Wu ⋅ Kailin Huang ⋅ Dongjian Hu ⋅ Bin Chen ⋅ Hao Wu ⋅ Shu-Tao Xia ⋅ Changliang Zou

Detecting AI-generated text is an important but challenging problem. Existing likelihood-based detection methods are often sensitive to content complexity and may exhibit unstable performance. In this paper, our key insight is that modern Large Language Models (LLMs) undergo alignment (including fine-tuning and preference tuning), leaving a measurable distributional imprint. We theoretically derive this imprint by abstracting the alignment process as a sequence of constrained optimization steps, showing that the log-likelihood ratio can naturally decompose into implicit instructional biases and preference rewards. We refer to this quantity as the Alignment Imprint. Furthermore, to mitigate the instability in high-entropy regions, we introduce Log-likelihood Alignment Preference Discrepancy (LAPD), a standardized information-weighted statistic based on alignment imprint. We provide statistical guarantee that alignment-based statistics dominate Fast-DetectGPT in performance. We also theoretically show that LAPD strictly improves the unweighted alignment scores when the aligned and base models are close in distribution. Extensive experiments show that LAPD achieves an improvement 45.82% relative to the strongest existing baselines, yielding large and consistent gains across all settings

A common approach for contextual optimization first trains a model to predict the unknown state from the context, and then takes the action that is optimal for the predicted state. We show that this follow-the-prediction policy can be brittle: for several standard problems, even a minimum-variance predictor with arbitrarily small RMSE can incur constant excess loss. We propose a simple policy that robustifies any black-box predictor by optimizing against the worst case in an $\varepsilon$-neighborhood of the prediction. We identify a problem-specific quantity, the \emph{robust optimization gap}, that captures the price of this hedging, and prove a general bound on the loss. We instantiate the framework for contextual posted pricing, ski rental, and house flipping, deriving closed-form robust policies and $O(\eta^{2/3})$ loss bounds in all three settings, where $\eta$ is the RMSE of the predictor. For posted pricing, this improves upon a result of Medina and Vassilvitskii (2017), and with a substantially simpler argument.


A Locally Tokenized Generative Model for Robust Time-Series Watermarking

Dongbin Kim ⋅ Geonwoo Shin ⋅ Yujin Choi ⋅ Soyeon Park ⋅ Jaewook Lee

Watermarking is a central tool for provenance in generative models, yet its application to multivariate time series remains hindered by reliability failures under post-editing attacks. We show that existing detectors, which rely on globally coupled re-encoding, suffer from bidirectional drift of the null distribution: post-editing attacks can shift the detection z-score of non-watermarked samples in either direction, invalidating clean-calibrated thresholds. We argue that this instability is a property of the re-encoding, and that reliable detection requires each recovered unit to depend only on a bounded temporal neighborhood. Guided by this principle, we propose L-VQVAE, a generative model in which each discrete token is produced from a short contiguous window, and LVQMark, a watermarking method over this token space that combines logit-bias injection with robust re-encoding for attack-time detection. Experiments on four benchmarks spanning finance, energy, and neuroimaging show that our approach preserves generation quality while stabilizing both detection power and false-positive behavior under post-editing attacks.

Formulaic alpha discovery is a core challenge in quantitative trading, as identifying alphas that work well together remains difficult. Recent reinforcement learning (RL) methods formulate this task as a Markov decision process (MDP), but two important issues remain unresolved. First, as the alpha pool evolves, the reward function changes accordingly, making the MDP inherently non-stationary. Second, most existing methods optimize a single objective, typically predictive power, while ignoring other important properties of a high-quality alpha pool. Motivated by these challenges, we propose AlphaPareto, an RL method for formulaic alpha discovery. To address non-stationarity, AlphaPareto augments the state to include both the alpha under construction and the current alpha pool, and applies a large language model (LLM) to encode the pool. This design allows the agent to adapt to the evolving search environment. To overcome the limitation of single-objective reward design, AlphaPareto replaces the scalar reward with a multi-objective vector-valued reward that simultaneously captures predictive power, temporal stability, perturbation robustness, and diversity, and optimizes these objectives through a Pareto-regularized learning procedure. Empirical applications to real-world datasets show that our AlphaPareto method outperforms its competitors.


A Mechanistic Investigation of Theory of Mind in a Large Language Model

Idil K Sahin ⋅ Steven M Frankland ⋅ Taylor Webb

Large language models successfully solve classic theory-of-mind tasks adapted from cognitive science, but the mechanisms underlying this ability remain unclear. Do they deploy circuitry specialized for reasoning about other minds, memorize common false-belief vignette structures and outputs, or rely on more domain-general computations? We investigate this question in Qwen2.5-14B-Instruct using causal mediation and representational similarity analyses across matched prompt sets that systematically vary an agent’s beliefs, the state of the physical world, and the correct answer. We identify a population of mid-layer attention heads that tracks divergence between an initial representation and the current state of the world. These heads are causally relevant regardless of whether the initial representation is an agent’s belief or a photograph, despite the photograph condition involving no agent or perspective-taking. A distinct population of later-layer heads retrieves the answer token. The two populations are functionally dissociable yet combine compositionally. Together, these findings argue against both mentalizing-specific and memorization accounts, suggesting instead that LLMs solve classic false-belief problems by reusing a domain-general mechanism for detecting divergence between representations and reality.


Amortized Optimal Transport from Sliced Potentials

Minh-Phuc Truong ⋅ Khai Nguyen

We propose a novel amortized optimization method for predicting optimal transport (OT) plans across multiple pairs of measures by leveraging Kantorovich potentials derived from sliced OT. We introduce two amortization strategies: regression-based amortization (RA-OT) and objective-based amortization (OA-OT). In RA-OT, we formulate a functional regression model that treats Kantorovich potentials from the original OT problem as responses and those obtained from sliced OT as predictors, and estimate these models via least-squares methods. In OA-OT, we estimate the parameters of the functional model by optimizing the Kantorovich dual objective. In both approaches, the predicted OT plan is subsequently recovered from the estimated potentials. As amortized OT methods, both RA-OT and OA-OT enable efficient solutions to repeated OT problems across different measure pairs by reusing information learned from prior instances to rapidly approximate new solutions. Moreover, by exploiting the structure provided by sliced OT, the proposed models are more parsimonious, independent of specific structures of the measures, such as the number of atoms in the discrete case, while achieving high accuracy. We demonstrate the effectiveness of our approaches on tasks including MNIST digit transport, color transfer, supply-demand transportation on spherical data, and mini-batch OT conditional flow matching.


A multi-scale information geometry reveals the structure of mutual information in neural populations

Simone Azeglio ⋅ Steeve Laquitaine ⋅ Ulisse Ferrari ⋅ Matthew Chalk

Understanding how neural population responses represent sensory information is a central problem in systems neuroscience. One approach is to define a representational geometry on stimulus space in which distances reflect how reliably stimuli can be distinguished from neural activity. However, different constructions of these distances can lead to qualitatively different conclusions about the neural code. Here, we show that a unique Riemannian representational geometry emerges from first principles governing how distances contract as stimulus resolution is lost through coarse-graining. This results in a multi-scale extension of the Fisher information metric, capturing encoding structure from fine stimulus details to coarse global distinctions. The resulting geometry is exactly related to the mutual information encoded by the population: well encoded stimulus directions -- those contributing more to mutual information -- are expanded, whereas poorly encoded directions are contracted. The metric tensor can be estimated using diffusion models, making the framework practical for large neural populations and high-dimensional stimuli. Applied to visual cortical responses to natural images, the eigenvectors of the metric tensor identify stimulus variations that contribute most to information transmission, yielding interpretable features that are robust to modelling choices. Together, these results provide a principled, information-theoretic framework for characterising neural population codes.


Anchoring Reasoning Distillation via Syntactic Constraints

Zehua Cheng ⋅ Wei Dai ⋅ Jiahao Sun

Distilling "System 2" reasoning into compact student models is bottlenecked by the scarcity of fully-correct teacher traces: as task difficulty grows, rejection sampling discards an increasingly large fraction of teacher generations. "Wrong" (W) traces are abundant, but training on them indiscriminately causes *policy poisoning* the student internalises the teacher's hallucinated arithmetic alongside any useful reasoning structure. We propose **Atomic Reasoning Units (ARU)**, a framework that resolves this trade-off via strict syntactic constraints. ARU enforces a Backus--Naur-Form (BNF) grammar that decomposes each reasoning step into a Premise, an Operation, and a Result, algorithmically disentangling reasoning planning from arithmetic execution. *Counterfactual Logic Verification* (CLV) symbolically re-executes W-traces to recover the $68\%$ that have valid plans but faulty arithmetic, and *Syntactic Loss Masking* (SLM) trains the student on this verified structure while suppressing the gradient on hallucinated result tokens. Across the $3 \times 8 \times 6$ grid we evaluate ($\{0.6, 1.7, 8\}$\,B Qwen3 students $\times$ eight benchmarks $\times$ six baselines), ARU is the strongest method in every cell, improving GSM8K by $+11.1$ and MATH-500 by $+9.8$ over the closest baseline at $1.7$\,B; on the held-out AIME competition set ARU's Maj@$8$ is $14.8$ versus $9.5$ for PoT (the closest baseline). We further document a small-scale capability gap in which ARU at $0.5$\,B matches Gold-SFT at $\geq 1.5$\,B on GSM8K (paired-bootstrap $p<0.05$ after Bonferroni correction), with a pre-registered ProofWriter control showing the gap collapses on pure-logic reasoning---consistent with arithmetic offloading as the dominant mechanism, though we do not claim the control uniquely isolates it. The code is available at https://anonymous.4open.science/r/ARU-286D.


AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization

Yu Li ⋅ Menghan Xia ⋅ Gongye Liu ⋅ Xintao Wang ⋅ Conglang Zhang ⋅ Lei Ke ⋅ yuxuan lin ⋅ Ruihang Chu ⋅ Pengfei Wan ⋅ Kun Gai ⋅ Yujiu Yang

Despite being a pivotal frontier, interactive world modeling remains underexplored in terms of the versatile controllability required by practical scenarios. To bridge this gap, we present AnchorWorld, a framework that advances egocentric simulation through enhanced interaction integrity and a flexible mechanism for world customization. First, we utilize 3D human motion as the primary interaction modality. To complement the out-of-view or truncated body parts in egocentric views, we introduce an auxiliary training supervision that incorporates exogenous viewpoints decoupled from the agent’s first-person sensorium. It allows the model to observe the agent's full-body positioning relative to the environment, facilitating a more robust spatial grounding of human-world interactions. Furthermore, we propose a simple yet effective mechanism for customizing self-evolving worlds. This is achieved by defining anchor views within a unified world coordinate system, coupled with textual descriptions dictating the dynamic evolution of local scenes. Experimental results show that AnchorWorld significantly outperforms state-of-the-art baselines, while ablation studies validate the effectiveness of our key designs. Notably, our customization scheme exhibits promising spatio-temporal geometric consistency and adheres strictly to the prescribed evolutionary dynamics.


An Empirical Study on Noisy Data and LLM Pretraining Loss Divergence

Qizhen (Irene) Zhang ⋅ Ankush Garg ⋅ Jakob Foerster ⋅ Niladri S. Chatterji ⋅ Kshitiz Malik ⋅ Mike Lewis

Large-scale pretraining datasets drive the success of large language models (LLMs). However, these web-scale corpora inevitably contain large amounts of noisy data due to unregulated web content or randomness inherent in data. Although LLM pretrainers often speculate that such noise contributes to instabilities in large-scale pretraining and, in the worst cases, loss divergence, this phenomenon remains poorly understood. In this work, we present a systematic empirical study of whether noisy data causes LLM pretraining divergences, costing more than 100,000 GPU hours in total. By injecting controlled, synthetic, uniform random noise into otherwise clean datasets, we analyze training dynamics across model sizes ranging from 480M to 5.2B parameters. We show that noisy data indeed induce training loss divergence, and that the probability of divergence depends strongly on the noise type, the amount of noise, and the model scale. We further find that noise-induced divergences exhibit activation patterns distinct from those caused by high learning rates, and we provide diagnostics that differentiate these two failure modes. Together, these results provide a large-scale, controlled characterization of how noisy data affects loss divergence in LLM pretraining.


APS: Bias-Controlled Adaptive Prototype Simulation for Population-Scale LLM Agents

Quan ZHENG ⋅ Yan Gao ⋅ shaobin he ⋅ Haoxiang Guan ⋅ Yuanhe Tian ⋅ Jie Feng ⋅ Ming Wang ⋅ Shuxin Zheng ⋅ Zhen Liu

LLM-agent simulation offers a flexible computational tool for studying population response trajectories that depend on scenario events, memory, demographics, and evolving social context. However, full multi-round simulation scales linearly with both population size and horizon, requiring every agent to query the LLM at every round. We propose \textbf{Adaptive Prototype Simulation (APS)}, a framework that treats scaling as a recurrent LLM-oracle allocation problem rather than a static sampling, propagation, or surrogate-modeling problem. APS preserves the specified LLM as the online transition oracle while querying adaptive core prototypes, selected singleton-tail agents, and shadow-audit agents. Prototype responses induce local response surfaces for nearby agents, reducing online LLM calls without replacing the underlying transition model. To control approximation bias, shadow-audit residual correction estimates propagation residuals for aggregate correction and future budget allocation, while tail-protected singleton routing directly queries selected isolated, heterogeneous, or high-curvature regions that are vulnerable to smoothing. We analyze APS as an estimator of the LLM-induced population transition and decompose its error into prototype-coverage error, shadow-audit residual-correction error, local-propagation bias, and temporal context mismatch. Under the reported protocols, APS gives lower reference-aligned distributional discrepancy than scale-oriented and same-budget baselines while reducing online LLM calls, with ablations and compact robustness checks diagnosing the main bias-control mechanisms. In a 10M-agent, multi-round public-opinion simulation, APS achieves a $381.1\times$ reduction over full simulation, with reference-aligned final-round JSD 0.094 against the corresponding full-LLM reference.

Aptamers are programmable DNA and RNA receptors with broad potential in diagnostics, therapeutics, biosensing, and molecular monitoring. Although they offer chemically synthesizable alternatives to protein-based binders in several applications, computational aptamer discovery remains much less mature, especially for small-molecule targets. This gap reflects three persistent evaluation failures: fragmented and inconsistent affinity annotations, unreliable negatives derived from untested rather than experimentally validated pairs, and random splits that leak related aptamers or recurring ligand identities across train and test sets. As a result, current benchmarks make it difficult to distinguish transferable aptamer-ligand recognition from dataset-specific shortcuts. We present AptaBench, the first curated, standardized, and leakage-aware benchmark for aptamer-small-molecule binding prediction with experimentally reported active and inactive measurements and paired classification and affinity-regression tasks. AptaBench integrates eight curated sources into 6,289 interaction pairs covering 1,610 DNA/RNA aptamers and 942 ligands. We standardize sequences, resolve ligands to canonical molecular representations, harmonize duplicate and conflicting records, and convert dissociation constants to p$K_d$. More than 30% of entries contain quantitative affinity values, while inactive labels are taken only from explicitly reported non-binding or low-affinity measurements rather than synthetic cross-pairing. AptaBench provides fixed in-distribution, molecule-disjoint, and aptamer-disjoint protocols reflecting practical discovery scenarios: prediction for unseen targets and evaluation of new candidate sequences. Across descriptor-based models, pretrained sequence and molecular encoders, and sequence-ligand fusion architectures, in-distribution evaluation substantially overestimates generalization. The best classifier reaches ROC-AUC 0.95 in-distribution, but drops to 0.87 for unseen molecules and 0.86 for unseen aptamers. Affinity prediction is more sensitive, with $R^2$ decreasing from 0.65 to approximately 0.30 under disjoint protocols. These results show that reliable progress in computational aptamer discovery requires broader chemical coverage, experimentally grounded inactive labels, and standardized quantitative annotations. By releasing curated data, fixed leakage-aware splits, reproducible baselines, and preprocessing code, AptaBench is poised to expose the generalization limits of existing models, challenge future sequence-molecule recognition methods, and provide a rigorous foundation for practice-oriented aptamer evaluation.


Architecture-Embedded Physics Priors for Mitigating Spectral Bias in Physics-Informed Neural Networks

Zhuo Zhang ⋅ Li Zhang ⋅ Ying Miao ⋅ Hongzong LI ⋅ Zhixuan Liang ⋅ Yong Yang ⋅ Gemine Vivone

Physics-Informed Neural Networks (PINNs) struggle to capture high-frequency components of PDE solutions, a failure mode known as spectral bias. Most existing fixes adjust the loss, the sampling, or the input encoding, but the network itself remains agnostic to the equation it is supposed to solve. We take a different route and embed physics priors into the architecture. The network's internal spectrum is shaped to match the spectrum of the target PDE through three components that are trained end-to-end: a bounded coordinate warping that contracts regions where the solution varies sharply, a spectral attention layer driven by the PDE residual that emphasizes physically active Fourier modes, and an inter-harmonic gate that lets low-frequency channels modulate high-frequency ones. We test the design on four forward PDEs (Helmholtz, Wave, Klein--Gordon, Burgers), the Burgers inverse problem, and seismic reconstruction, including Gulf of Mexico field data. Across these tasks, the architecture clearly outperforms seven PINN baselines on accuracy and produces noticeably fewer artifacts. Code will be made publicly available upon acceptance.


A Regularization-Based Approach to Public Belief State Search for Adversarial Games

Sobhan Mohammadpour ⋅ Samuel Sokota ⋅ Brandon Kaplowitz ⋅ Zico Kolter ⋅ Noam Brown ⋅ Gabriele Farina

The public belief state (PBS), which is the posterior over game histories conditioned on public information, is a fundamental abstraction for designing game-theoretically sound search algorithms for imperfect-information games. Existing sound PBS search techniques for adversarial games fall into two categories: gadget-game-based algorithms (including Libratus, DeepStack, Pluribus, Supremus, and Student of Games) and ReBeL-based algorithms. Both these categories possess disadvantages, including a reliance on discontinuous functions, an inability to re-solve subgames, and an implicit representation of policies. In this work, we propose Gravel, a new approach for PBS search---different from these previous two categories---that achieves game-theoretic soundness via regularization. Unlike the aforementioned approaches, Gravel relies on smooth functions, can initiate search at any point in the game, and outputs explicit policies directly. We empirically demonstrate the advantages of Gravel for no-limit Texas hold'em, where we show that it outperforms ReBeL both in head-to-head competitions against Slumbot and in endgame solving. We include our codebase in the submission, making Gravel the first open-source high-performance no-limit Texas hold'em AI.

Reliable uncertainty estimation is essential for deploying large language models in high-stakes reasoning tasks, yet most existing approaches rely on expensive sampling strategies or shallow token-level signals that ignore the internal logical structure of model outputs. We introduce Argument Graph Uncertainty, a post-hoc framework for multiple choice question tasks, that estimates model confidence directly from the logical structure of a single chain-of-thought reasoning trace. Our method uses a lightweight instruction-tuned model to segment the reasoning chain into propositional units, then applies a simple NLI model to infer relational edges and construct two complementary argument graphs over these segments. The resulting edge scores are aggregated into a Dirichlet posterior over answer options, whose concentration yields calibrated uncertainty estimates reflecting both local step-to-step coherence and global logical consistency across the reasoning trace. Evaluated on GPQA, a challenging graduate-level multiple-choice benchmark, our method consistently surpasses token-based and consistency-based uncertainty baselines across multiple reasoning models. We further demonstrate through extensive combination analyses that argument graph uncertainty captures an orthogonal signal to existing approaches, making it a complementary component in uncertainty ensembles. Crucially, our framework requires no additional calls to the evaluated model and relies entirely on small, efficient auxiliary models, producing fully interpretable, per-node uncertainty attributions at low computational cost.

Articulated 3D assets are essential for robot simulation, yet generating them from casual images remains unsolved. Existing approaches either predict URDF parameters via learned regression or language model inference -- both requiring articulation supervision and producing joints that may be geometrically inconsistent with the generated meshes -- or reconstruct from dense calibrated multi-view captures, limiting scalability. We present \textbf{ArtCrafter}, a feed-forward generative model that takes two images of an articulated object -- one at rest, one fully open -- and produces $2N+1$ part-level meshes together with a physics-executable URDF in a single forward pass. Inspired by Slot Attention, we structure the denoising process around a $1+2N$ slot layout whose slots, forced to jointly explain two articulation states, naturally bind to kinematic parts. \textbf{TripletAttention} enforces cross-state geometric consistency between each paired rest/open slot anchored through a shared base slot, and a repulsion loss $\mathcal{L}_{\mathrm{rep}}$ operating on decoded point clouds ensures inter-part disjointness throughout denoising. Because paired meshes are geometrically consistent by construction, URDF joint parameters -- type, axis, origin, and range -- are derived \emph{analytically} from the rigid transform between each pair, requiring no learned articulation head, no language model, and no per-instance optimization. Experiments on URDF-Anything+ demonstrate consistent improvements over retrieval-based and generative baselines across part-level geometry and all four articulation metrics. ArtCrafter demonstrates a clean decomposition: a diffusion model handles geometry, and geometry handles URDF.


Articulation in Prime: Primitive-Based Articulated Object Understanding from a Single Casual Video

Arslan Artykov ⋅ Tom Ravaud ⋅ Nicolás Violante ⋅ Vincent Lepetit

Retrieving the 3D kinematics of articulated objects from monocular video is a fundamental challenge in computer vision. Existing methods rely on complex video setups or cues such as long-term point tracking or wide-baseline matching, but are frequently brittle under severe occlusions, rapid camera ego-motion, or weak local features. Learning-based methods, meanwhile, struggle to generalize beyond their training categories. We propose a category-agnostic optimization framework that treats articulated object understanding as a primitive-fitting problem. Geometric primitives serve as a proxy representation that avoids the pitfalls of unstable point tracks; a novel mechanism organizes them into coherent parts constrained by revolute and prismatic joints. Our formulation jointly optimizes part segmentation and joint parameters, recovering complex kinematics from a single casually captured video. A visibility-aware procedure handles partial observations and occlusions inherent to real-world data. We also propose the AiP-synth and AiP-real benchmarks, featuring significant camera motion and heavy occlusions, and outperform existing methods.


A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression

JinCheng Ren ⋅ Siwei Wu ⋅ Yizhi Li ⋅ Zhu ⋅ Shu XU ⋅ Boyu Feng ⋅ Ruibin Yuan ⋅ Wei Zhang ⋅ Riza Batista-Navarro ⋅ Min Gao ⋅ Yan Bai ⋅ Jian Yang ⋅ Chenghua Lin

Terminal observations are not ordinary long-context text: they are heterogeneous, low-information-density execution traces in which sparse but exact evidence (e.g., error messages and file paths) is interleaved with large amounts of repetitive terminal output. For long-horizon CLI agents, retaining raw observations rapidly increases context cost and can dilute critical signals, while LLM-based summarization or fixed heuristics often fail to adapt across heterogeneous terminal tasks and may discard precise task-relevant evidence.We propose TACO, the first self-evolving Terminal Agent Compression framework, which treats compression rules as reusable, preservation-aware knowledge acquired from interaction trajectories. Rather than relying on manually designed rules, static pruning, or task-specific compressor training, TACO autonomously discovers, refines, and reuses structured compression rules from agent interaction trajectories. A Global Rule Pool accumulates effective rules across tasks, while task-time rule evolution adapts them online to the current workflow. This allows TACO to filter redundant terminal observations while conservatively preserving exact task-relevant evidence.Across six benchmarks, including TB 1.0, TB 2.0, SWE-Bench Lite, CompileBench, DevEval, and CRUST-Bench, TACO consistently maintains or improves task success across models and agent scaffolds. On TerminalBench, TACO yields 1%--4% absolute accuracy gains under standard evaluation and improves accuracy by 2%--3% under matched token budgets. On downstream benchmarks, it reduces total token consumption by 12%--27% while maintaining or improving task success. These results show that self-evolving observation compression can unlock latent capability in existing CLI agents by allocating context budget toward task-relevant evidence, without model fine-tuning or human-crafted compression rules.


A Simple Class-Agnostic Approach to Enhance Fair Adversarial Training

Erh-Chung Chen ⋅ Pin-Yu Chen ⋅ I-Hsin Chung ⋅ Che-Rung Lee

AI security has become a critical issue as Deep Neural Networks (DNNs) are inherently vulnerable to adversarial attacks. While previous studies have demonstrated that adversarial training is one of the most effective defensive strategies, improvements in robustness often show significant class-wise disparities. This has established fair adversarial training, which aims to enhance the robustness of the worst-performing classes, as a crucial research direction. Although existing class-wise approaches can improve performance by adjusting individual class weights, they face severs scalability limitation on large-scale datasets due to the sparse volume of data available per class. To mitigate this issue without relying on explicit class-wise information, we propose a novel tri-regime optimization framework that fundamentally decouples the distinct sources of adversarial error at the sample level. Specifically, our method systematically suppresses unlearnable outliers via reliability gating, prioritizes genuine boundary threats through dynamic reweighting, and down-weights already safe examples to prevent redundant optimization. By isolating and prioritizing the effective adversarial frontier, our method strictly prevents robust overfitting and achieves superior worst-class robustness while maintaining high clean accuracy.

Mapping software issue descriptions to relevant code segments, a task known as code localization, is a primary bottleneck in automated software engineering. While recent attempts have automated it using code graphs that capture syntax-level dependencies among code elements, including functions and modules, their code-centric approaches often overlook the design rationales and prior fix patterns present in a codebase's development history, such as pull requests and commits. We propose Co-Locator, a collaborative multi-agent code localization framework that integrates code graphs reflecting functional dependencies and knowledge graphs constructed from triplets capturing design rationales and past code modifications scattered across the development history. During localization, a Code Graph (CG) agent explores functional hierarchies and consults a Knowledge Graph (KG) agent to gather insights into developer intents and prior bug fixes. Combined with a Retriever agent that assists CG and KG agents, Co-Locator leverages a holistic view of the code's functional and historical context, considering what the code implements and how it was changed, to identify locations to fix. Experiments show that Co-Locator significantly outperforms 12 state-of-the-art baseline methods.

Chain-of-thought (CoT) reasoning has become a widely used mechanism for eliciting multi-step reasoning in large language models by generating intermediate reasoning steps at inference time. Yet the scaling behavior of generalization with CoT depth remains poorly understood. To address this question, we study a theoretically solvable model of CoT for in-context weight prediction in linear regression, where test-time reasoning is represented as an iterative refinement of the weight-parameter estimate. Using tools from random matrix theory under high-dimensional asymptotics, we derive an exact formula for the generalization error as a function of reasoning depth, pretraining data amount, and context length. Our analysis reveals a sharp phase transition separating exponential and polynomial improvement, saturation, and overthinking, and characterizes how the optimal reasoning depth scales. We further show that deeper reasoning is most effective with sufficiently rich pretraining and in-context information, whereas limited pretraining or context makes longer reasoning prone to error amplification or saturation. We also validate these predictions through experiments on fully learned linear attention and softmax attention models. Our results provide a unified theoretical account of how test-time CoT depth affects generalization.


A Sparse Low-Rank Biclique Decomposition for Graphs

Antoine Vialle ⋅ Sergei Gerasimov ⋅ Aref Einizade ⋅ Antonio Ortega ⋅ Fragkiskos Malliaros ⋅ Jhony H. Giraldo

Low-rank plus sparse decompositions, such as robust PCA, typically model the low-rank component as dense and the sparse component as corruption. This viewpoint is poorly suited to graph adjacency matrices, where sparsity is structural and dense low-rank factors destroy the combinatorial and computational properties of the graph. We propose signed biclique decomposition (SBD), a sparse low-rank decomposition for integer-weighted graphs that represents the structured component as a sum of sparse signed binary rank-one factors, together with a sparse signed residual. The resulting representation remains discrete, interpretable, sparse, and compatible with standard linear algebra. We formulate fixed-rank and greedy SBD objectives and derive a gain-based greedy algorithm with theoretical guarantees. Empirically, SBD recovers block structure in noisy graphs, while its factorized representation enables efficient sparse matrix-matrix multiplication (SpMM). On real graphs with strong shared-neighborhood structure, SBD achieves compression factors above $10$ and SpMM speedups up to $6\times$ over cuSPARSE, suggesting a practical alternative to dense low-rank decompositions for graphs.


ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents

Udari Madhushani Sehwag ⋅ Zhengyang Shan ⋅ Heming Liu ⋅ Dileepa Lakshan ⋅ Joseph Brandifino ⋅ Max Fenkell

Clarification-seeking behavior is widely regarded as a desirable property of LLM agents, enabling them to resolve ambiguity before acting on underspecified tasks. However, the security implications of this interaction pattern remain unexplored. We investigate whether the transition from standard execution to a clarification-seeking state increases an agent's susceptibility to prompt injection attacks. We introduce ASPI (Ambiguous-State Prompt Injection), a benchmark of 728 task-attack scenarios that isolates clarification as a distinct agent state and measures how this state transition affects vulnerability under controlled conditions. Each benchmark instance is evaluated under matched execution and clarification settings: in the execution setting, the agent acts on a fully specified instruction and encounters adversarial content only through tool-returned data; in the clarification setting, the agent must first request and incorporate additional user input before acting. We evaluate ten frontier LLMs and find that clarification-seeking consistently and substantially amplifies vulnerability. For instance, attack success rises from 1.8\% to 34.0\% for o3 and from 2.2\% to 35.7\% for Gemini-3-Flash. A decomposition analysis reveals that this gap reflects both a state-dependent shift in how models process incoming content and a channel-specific effect arising from the agent-solicited clarification interface. These findings demonstrate that standard execution-time security evaluation systematically underestimates the attack surface of interactive agents, and that robustness under fully specified tasks does not translate to robustness under ambiguity. For reproducibility our data and source code is available at https://anonymous.4open.science/r/aspi-5074.


ASQ: Agent-guided Semantic-aware Quantization for Large Language Models

Seungdong Yoa ⋅ Ye Seul Sim ⋅ Suhee Yoon ⋅ Sanghyu Yoon ⋅ Dongmin Kim ⋅ Soonyoung Lee ⋅ Junhyun Lee ⋅ Bumsoo Kim

Post-training quantization (PTQ) compresses large language models (LLMs) but remains largely semantic-agnostic, relying on global activation magnitude to guide precision allocation. This overlooks the functional heterogeneity of Transformer representations, where semantically meaningful channels are sparsely activated while structurally dominant channels exhibit large but context-insensitive activations. We propose Agent-guided Semantic-aware Quantization (ASQ), a PTQ framework that conditions on semantically informative positions identified by an auxiliary LLM. A dual-pathway scoring mechanism captures both activation selectivity and attention routing toward these positions, producing channel-wise importance scores for scaling. This yields a closer approximation to a semantically conditioned objective and consistently improves low-bit (e.g., W4A16) quantization performance across LLMs. Extensive experiments show consistent gains over state-of-the-art PTQ, preserving general reasoning and even surpassing FP16 performance on several benchmarks in specialized domains, while enabling training-free, privacy-preserving self-quantization with zero inference overhead.

Deep click-through rate (CTR) models commonly consist of sparse embedding blocks and feature-interaction blocks, which jointly learn predictive signals from high-dimensional categorical features. This naturally raises the question of whether different blocks in an end-to-end trained CTR model play the same role in generalization. To address this question, we investigate the generalization behavior of deep CTR models from a block-wise diagnostic perspective. Specifically, we formulate deep CTR models as two-block systems comprising an embedding block and a feature-interaction block, and introduce a retuning-based diagnostic protocol to separately assess the transferability of different blocks. Experiments across multiple datasets and architectures reveal a consistent asymmetric generalization pattern: the embedding block exhibits stronger training-set specificity, whereas the feature-interaction block preserves comparatively more transferable structure. Motivated by this diagnosis, we further develop a component-selective embedding aggregation method that averages multiple independently retuned embedding tables while keeping the feature-interaction block fixed. The resulting model retains the original single-path inference architecture and incurs no additional online inference cost. Experiments on Avazu, Criteo, and Taobao show that the proposed method improves representative deep CTR models over standard training and retuning baselines.

Negative guidance of diffusion models generates samples that avoid an undesired condition, specified either by reference samples or by a conditional score alongside an unconditional model. Existing methods inject repulsive terms derived from local information at each sample, so the resulting sampler does not target an explicit density, and its bias cannot be systematically reduced by increasing computation. We instead present a negative-guidance method that samples from an explicit target density, which uniformly handles both reference-sample and score-based specifications and suppresses probability mass near the undesired condition with an avoidance strength that parametrizes the target density itself. Realizing this target requires estimating, at each diffusion timestep, how likely each sample is to belong to the undesired condition. We cast this estimation as a Positive--Unlabeled learning problem that admits lightweight online training along the diffusion trajectory, using condition samples as positives and diffusion particles as unlabeled data. Because the guided particles drift from the unconditional marginal, we further correct the residual mismatch via Sequential Monte Carlo, yielding a sampler that is asymptotically exact in the particle limit and that turns any existing negative-guidance method into its proposal kernel. Empirically, our method approaches the ground-truth target on Gaussian mixtures and Pareto-dominates existing methods on trade-offs between avoidance and generation quality, distributional bias, and diversity, including in an image-generation setting with a pretrained score network.


AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism

Thalaiyasingam Ajanthan ⋅ Sameera Ramasinghe ⋅ Gil Avraham ⋅ Hadi Mohaghegh Dolatabadi ⋅ Chamin Hewa Koneputugodage ⋅ Violetta Shevchenko ⋅ Yan Zuo ⋅ Alexander Long

Data and pipeline parallelism are key strategies for scaling neural network training across distributed devices, but their high communication cost necessitates co-located computing clusters with fast interconnects, limiting their scalability. We address this communication bottleneck by introducing asynchronous updates across both parallelism axes, relaxing the co-location requirement at the expense of introducing staleness between pipeline stages and data parallel replicas. To mitigate staleness, for pipeline parallelism, we adopt a weight look-ahead approach, and for data parallelism, we introduce an asynchronous sparse averaging method equipped with an exponential moving average based correction mechanism. We provide convergence guarantees for both sparse averaging and asynchronous updates. Experiments on large-scale language models demonstrate that our approach matches the performance of the fully synchronous baseline, while significantly reducing communication overhead.


A Theoretical Analysis of Test-Driven Code Generation

Nicolas Menet ⋅ Michael Hersche ⋅ Andreas Krause ⋅ Abbas Rahimi

Code assistants are increasingly utilized in test-driven software development, yet the theoretical mechanisms behind their environment-interaction strategies remain underexplored. We provide a probabilistic framework for two dominant paradigms: code selection after generation using the execution environment, and code generation conditioned on environment feedback. First, we formalize several well-established selection heuristics as environment-aware estimators of code correctness. We theoretically prove that estimators based on fuzzy functional similarity add an inductive bias and strictly dominate estimators based on functional equivalence in terms of signal-to-noise ratio. Second, we frame backprompting as an in-context approximation of Thompson sampling. We derive a novel regret bound for reward functions with unobservable components, theoretically explaining why the effectiveness of backprompting is limited by the ambiguity of the informal task description (an irreducible regret). Using five state-of-the-art open weight models, we corroborate these findings across BigCodeBenchHard, LeetCodeDataset, and QiskitHumanEvalSim. Our formalization also suggests how to improve task descriptions effectively, leading to a new benchmark, QiskitHumanEvalSimX.

Multivariate weather time series are a central modality for forecasting and decision support, and the surrounding language increasingly mediates how numerical weather is consumed by users and downstream models. However, existing pipelines depend on closed-source language model ensembles with expert quality control or on retrieving human-curated text, while prior caption-RL rewards optimize cycle similarity or downstream-task utility, neither of which can tell a captioner that a confidently asserted phenomenon is contradicted by the data, so the absence of an executable, station-observable, and refutation-aware reward signal has blocked the application of self-play RL to this scientific modality. To address this, we propose AtmoZero, a post-training framework that trains a weather-time-series captioner without a single human-written caption, by aligning a language model with a library of seven station-observable scientific verifiers, pairing them with an explicit refutation channel that penalizes caption claims the data contradicts, and combining them with auxiliary reconstruction and forecaster-utility rewards under group-relative policy optimization. Extensive empirical results on a twelve-year ERA5 reanalysis corpus and real-world surface observations show that AtmoZero attains the lowest unsupported-claim rate and forecasting error at every horizon, outperforming closed-source captioners, dense-reward RL, and far-larger foundation models, with extensive analyses confirming that these gains arise from multiple complementary signals and are robust under distribution shift.


Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models

Yusuf Sahin ⋅ Ahmed R Saikia ⋅ Volkan Cevher ⋅ Paolo Favaro

Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is fragile: positions that are individually confident may be unsafe to commit together when their predictions are coupled. Existing training-free samplers such as Top-(k), Fast-dLLM, and EB-Sampler mainly control how many tokens to reveal, while often ranking candidates by token-wise scores that ignore interactions within the selected set. We propose ADAS, a training-free reranking rule for parallel masked diffusion decoding. ADAS leaves the base sampler's stopping rule unchanged and modifies only subset construction: it greedily discounts a candidate when it attends strongly to already selected positions whose predictions remain uncertain. Unlike graph-constrained methods that turn attention into hard compatibility constraints, ADAS keeps attention continuous and uses it as a soft marginal penalty. Across LLaDA-8B-Base and Dream-7B-Base on GSM8K, MATH500, HumanEval, and MBPP, plugging ADAS into Top-(k), Fast-dLLM, and EB-Sampler improves low-NFE performance at matched denoiser evaluations by (9.11) and (10.46) percentage points on average, respectively, with (3.1\%) per-forward runtime overhead. These results show that soft attention-discounted reranking is a simple and modular way to improve quality in highly parallel decoding for masked diffusion language models.

Conformal prediction (CP) for language models is typically evaluated under a \emph{matched} protocol, in which calibration and deployment use the same post-training inference configuration, such as quantization mode and inference temperature. In deployment, these choices often differ from those used at calibration to meet memory or throughput constraints or serving policies. We argue that this exposes a limitation in current CP evaluation protocols for language models. We formalize the underlying phenomenon as \emph{inference-time model shift}: a change in the conformal score distribution induced by inference configuration rather than by the input--label distribution, and theoretically characterize how it breaks split-conformal exchangeability even when the task distribution and trained model are unchanged. We then evaluate its consequences across $8$ language models, $4$ finite-label benchmarks, and $12$ inference configurations. We find that matched evaluation can substantially overstate deployment-time coverage, and that the effect is heterogeneous across models and datasets. We argue that inference configuration should be reported as a first-class experimental axis in CP evaluations for language-model systems.


A Unified Merge Calculus for Learning-Rate Scaling on Neural Computation Graphs

Haosong Zhang ⋅ Wu Shenxi ⋅ Zhiyuan Che ⋅ Xi Chen ⋅ Wei LIN ⋅ Yichi Zhang

Learning-rate selection is one of the most consequential and repeatedly tuned decisions in modern deep learning, and the ability to transfer it quickly across widths, depths, and architectures can substantially reduce the cost of scaling new models. Existing transfer laws explain width through maximal-update parameterization and depth through depth-based scaling laws, but many modern architectures are better described as neural computation graphs with heterogeneous merge rules rather than by depth alone. We develop a unified merge calculus for learning-rate scaling on such graphs. Its central object is a graph-aware complexity coefficient that combines local merge semantics with global path structure and induces a $-1/2$ power law for the target learning-rate scale. We further show that depth alone is generally insufficient once merge semantics vary across architecture families. We validate the framework on controlled graph families and lightweight representative architectures covering uniform sum, residual-aware sum, concat, weighted fusion, and hybrid designs. Across these settings, the resulting graph coefficients organize optimal learning rates more effectively than depth alone and support practical learning-rate transfer.


A Unified Neural Architecture for Variable-Wise Shape Constraints

hyunho kim ⋅ Dohyun Bu ⋅ Jong-Seok Lee

We introduce COMONet (Convex-Concave and Monotonicity-Constrained Neural Networks), a unified architecture for enforcing variable-wise shape constraints in neural networks. COMONet addresses the limitations of prior methods that either support only a restricted subset of monotonicity and curvature constraints or enforce them through penalties without strict architectural guarantees. Our framework assigns each input variable to one of eight shape regimes, covering monotonicity, convexity, concavity, and their combinations. This is achieved through a partially connected architecture that routes each variable to specialized units equipped with sign-constrained weights and constraint-preserving activation functions. We also provide theoretical guarantees showing that COMONet satisfies the prescribed variable-wise shape constraints by construction. Experiments on synthetic and real-world datasets demonstrate that COMONet achieves competitive performance and remains robust to noise while preserving its architectural guarantees. COMONet thus offers a practical and principled framework for incorporating domain knowledge as shape constraints into neural network training.


A Unified Semismooth Newton Approach to Multitask and Multivariate Square-Root Lasso Problems

Dongwon Kim ⋅ Taehyoung Kim ⋅ Sungdong Lee ⋅ Joong-Ho (Johann) Won

We study scalable second-order algorithms for high-dimensional square-root Lasso problems with multiple responses. The resulting estimators preserve the scale-free tuning property of their single-response counterpart, but accurate large-scale computation is challenging because both the loss and penalty terms are nonsmooth. In this paper, we propose a computational framework that universally applies to multitask and multivariate square-root Lasso models (which includes the single-response version) based on the Douglas-Rachford splitting. We solve the nonsmooth equation induced by the splitting by a regularized semismooth Newton method with projection and fixed-point safeguards. Despite the nonsmoothness of the problem, the structure of the proximal maps and generalized Jacobian lead to a reduced Newton system in a much smaller dimension than that in the original space, promoting scalability. Our method enjoys global and locally quadratic convergence under mild conditions. Experiments on synthetic regression problems and large-scale multi-omic data show that the proposed method reaches the same fitted models and predictive performance as latest solvers tailored to individual models, while requiring fewer iterations and less wall-clock time in large-scale, high-dimensional settings.


AVENUE: Audio-Video Editing Understanding and Evaluation

Hayeon Kim ⋅ Yoojin Jang ⋅ Jaejun Yoo

Audio-video (AV) editing aims to modify audio and video content according to a target prompt. Unlike single-modality editing, AV editing requires models to infer a modality-selective edit scope from the prompt alone: determining not only what should change, but also which modality should be preserved. Faithfully evaluating such models therefore requires both (i) benchmarks that span diverse edit types and modality categories, and (ii) evaluation that is itself modality-aware and sample-specific. However, existing AV editing benchmarks provide limited coverage of edit types and modality combinations, while current evaluation systems are often modality-blind and sample-agnostic, making it difficult to assess whether models faithfully preserve the unintended modality. To address these gaps, we introduce AVENUE, Audio-Video EditiNg Understanding and Evaluation, comprising two contributions: (1) a benchmark of 1,291 source clips and 7,957 editing instructions across audio-targeted, video-targeted, and AV-coupled edit types, curated and human-verified from VGGSound; and (2) a sample-specific, modality-aware evaluation framework that specifies, for each sample, both the intended change and the content that must remain intact. We evaluate representative AV editing models spanning three editing paradigms — joint, sequential, and separate — providing the first systematic analysis of modality-selectivity across paradigms. Our findings reveal a fundamental open challenge: when editing one modality, existing models frequently induce unintended changes in the other, regardless of paradigm. AVENUE provides a benchmark and modality-aware evaluation framework to drive progress toward more controllable AV editing models.

Inference-time search offers a principled way to improve the reliability of large language models (LLMs) by exploring and scoring intermediate trajectories. However, widely used search paradigms exhibit complementary failure modes: sequence search scales linearly but commits early and cannot recover from mistaken prefixes, whereas tree search can backtrack but often becomes impractical for deep search due to the exponential growth of the search space with depth. These limitations become increasingly pronounced in long-chain-of-thought responses. We propose Width-Limited Tree Search (WLTS), which reconciles these strengths by preserving the tree-search structure that enables backtracking while using a dynamic expandability criterion to enforce a strict per-layer width constraint. This design keeps generated search-tree growth linear in depth and sustains verifier supervision across the full reasoning process. WLTS further supports two optional deployment components: prefix deletion for diversity and agreement-based early stopping to reduce computation on tasks with canonicalizable outputs. Across mathematical, coding, and logical reasoning benchmarks, and under both classifier and generative verifier settings, WLTS achieves higher accuracy and better inference-budget efficiency than well-established sequence- and tree-search baselines.


Balancing Image Compression and Generation with Bootstrapped Tokenization

Haozhe Chi ⋅ Jinghan Li ⋅ Hao Jiang ⋅ Wu Sheng ⋅ Yi Ma ⋅ Jing Wang ⋅ Yadong Mu

Despite progress in image tokenization, standard methods encode redundant information by mixing all granularities within each token, thus redundancy persists between tokens. The mix of information of different granularity also complicates the training of generators. This paper introduces SelfBootTok, a method that resolves this by cleanly decomposing information into global and local token groups. Through self-bootstrapped learning, the model predicts local details exclusively from global tokens, shifting the burden of visual details from the generator to the tokenizer. Consequently, our generator is far more efficient, requiring only global tokens and reducing computation by approximately 40\%, while delivering superior reconstruction and generation. Moreover, this paradigm scales elegantly: by leveraging more data or parameters to self-supervise local representation learning, SelfBootTok achieves a new state-of-the-art gFID score of 1.56 using only 64 tokens.

Few-shot medical image segmentation has shown great potential in reducing annotation costs for clinical applications. Recently, many few-shot methods have explored the Segment Anything Model (SAM) for training-free medical image segmentation via prompt engineering. However, existing approaches mainly focus on locating positive prompts from support images while overlooking informative background cues, making them prone to over-segmentation in anatomically similar regions. Moreover, simply introducing negative prompts cannot effectively suppress boundary leakage, and may even interfere with SAM’s mask decoding process, resulting in worse performance than using positive prompts alone. To address this limitation, we propose a training-free support-query framework, termed Boundary-Aware Prompt Mining, which introduces boundary-aware negative prompting for SAM-based few-shot medical image segmentation. Specifically, we introduce a Prototype-guided Positive Prompt strategy, which adopts a multi-center prompting mechanism to construct multiple foreground prototypes, enabling comprehensive spatial coverage in query images. Furthermore, we propose a Boundary-band Ambiguity Filtering strategy that identifies a boundary-adjacent background band and progressively removes ambiguous pixels and unreliable clusters highly similar to the foreground, enabling the selection of reliable negative prompts for suppressing false positives near foreground boundaries. The generated positive and negative prompts are jointly fed into SAM to produce refined segmentation results without additional training. Extensive experiments on Abd-MRI and Abd-CT datasets demonstrate that our method consistently outperforms existing approaches, highlighting its effectiveness and robustness under limited annotation settings. Code will be released upon acceptance.

Bayesian causal inference quantifies uncertainty in treatment effects under a specified causal model. This uncertainty is often interpreted as evidence that the resulting causal conclusion is robust. We argue these are distinct targets: a posterior over a treatment effect may be tightly concentrated above zero while the corresponding causal claim is fragile to small violations of ignorability, overlap, or target-population assumptions. We introduce Bayesian causal stress testing (BCS), a framework that assigns a posterior fragility score $F_\alpha$ with a threshold $\alpha\in(0,1)$ to a causal conclusion by measuring the smallest calibrated stress under which the posterior probability of the claim falls below $\alpha$. We give a stress-map formulation, decision and composition results for claim-level stress, a calibration theorem showing that $F_{0.5}$ asymptotically recovers the partial-correlation form of the Cinelli-Hazlett robustness value in the linear-Gaussian flat-prior limit, and a proposition formalizing why posterior precision and posterior fragility can decouple. We evaluate BCS on IHDP, TWINS, and ACIC-cov+Hill, a semi-synthetic benchmark built from the ACIC 2016 covariate matrix, under BART, BCF-style, Bayesian linear, and Bayesian-bootstrap AIPW posteriors. On the linear Bayesian posterior, $F_{0.5}^2$ matches the Cinelli-Hazlett $RV_{0.5}^{\mathrm{CH}}$ to RMS $0.001$ on both IHDP and ACIC-cov+Hill across all $20$ configurations; the BART and BCF-style posteriors show larger deviations in small-effect configurations. Across estimators, posterior $95$% interval widths vary by up to $50$% but $F_{0.5}$ remains within $0.05$, illustrating that posterior precision is not a robustness certificate. A controlled-DGP calibration experiment with a hidden confounder of known partial correlation $\gamma_{\mathrm{true}}$ confirms that $F_{0.5}$ tracks the required flip-stress to mean absolute error $0.03$.

Large language model systems repeatedly make test-time decisions by comparing similarity. They retrieve exemplars, rerank candidates, and choose revisions. But similarity is not unique: a candidate can be close lexically, semantically, structurally, or stylistically, while only one of these axes may matter for the current instance. We propose BASIS~(Bayesian Axis Selection via Interactive Signals), a training-free Bayesian framework that treats weak feedback as evidence about this latent task-aligned axis. BASIS maintains a posterior over candidate targets and similarity axes. It updates this posterior from scalar scores, pairwise preferences, or verifier outputs, and then chooses future probes by expected information gain. In a finite noisy similarity model, we show that fixed-axis and fixed-mixture rules can remain suboptimal under axis shift. Across concept recovery, reasoning exemplar selection, response refinement, and adversarial distractors, BASIS consistently outperforms stronger rerankers, static mixtures, and compute-matched best-of-$N$ search under the same candidate pools and feedback budgets. The average gains are moderate because all methods share the same backbone, candidate pools, and feedback signals. Under axis shift, however, the gap widens substantially, showing that latent-axis inference is not fully replaced by better scoring or more search alone.

We introduce BEACON—Best-Effort Adaptation for Cross-Domain Co-Training—a theory-driven framework for training generative robot policies with abundant source demonstrations and limited target demonstrations. BEACON casts cross-domain co-training as a discrepancy-aware importance-reweighting problem, jointly learning a diffusion-based visuomotor policy and per-sample source weights that minimize an objective informed by target-domain generalization guarantees. To make best-effort adaptation practical for high-dimensional sequence policies, we develop scalable instance-level discrepancy estimators, stochastic alternating updates for policy and weights, and a multi-source extension that balances heterogeneous source domains. Across sim-to-sim, sim-to-real, and multi-source manipulation settings, BEACON improves robustness and data efficiency over target-only, fixed-ratio co-training, and feature-alignment baselines. Importantly, even without an explicit alignment objective, BEACON achieves feature alignment as an implicit result of discrepancy-aware cross-domain co-training.


Behavior-Discriminative Reward Shaping for Reward-Robust Reinforcement Learning

Zixuan Liu ⋅ Fangzheng Wu ⋅ Brian Summa ⋅ Zizhan Zheng

Reward-robust RL typically models misspecification by specifying an uncertainty set that constrains the discrepancy of plausible rewards from a base reward. However, such reward space discrepancies can be behaviorally irrelevant: under potential-based reward shaping (PBRS), many distinct rewards preserve policy ordering. This can introduce substantial redundancy into standard uncertainty sets and degrade optimization performance. We propose Shaping-Aware Reward-Robust RL, which constructs uncertainty sets over PBRS equivalence classes by projecting each reward to a canonical representative, ensuring that the resulting set contains only rewards that induce behaviorally distinct policy rankings. We prove that this projection preserves the optimal robust value while shrinking the uncertainty set and improving empirical performance. Using the connection between robustness and regularization, we obtain a practical algorithm to solve the shaping-aware reward-robust RL problem and enjoy convergence guarantees under standard assumptions. Experiments on benchmarks spanning diverse task domains and levels of complexity show consistent improvements over representative robust RL baselines and exhibit improved robustness to reward perturbations.


Benchmarking Graph Self-Supervised Learning for Node-Level Tasks: Insights and Strong Baseline

Dmitry Eremeev ⋅ Gleb Bazhenov ⋅ Oleg Platonov ⋅ Artem Babenko ⋅ Liudmila Prokhorenkova

Self-supervised learning (SSL) has received notable attention in the graph machine learning community, enabling the efficient use of unlabeled data. However, evaluation setups vary significantly across studies, and baseline tuning often receives limited attention, making it difficult to draw reliable conclusions about model performance. In this paper, we reevaluate four representative graph SSL methods within a unified setup across a diverse set of datasets. We focus on rigorously tuning both SSL and supervised methods by employing enhanced GNN architectures and comprehensive hyperparameter optimization. We observe that, contrary to prior literature, two popular generative methods, MaskGAE and GraphMAE, regularly fail to outperform well-tuned supervised baselines. In contrast, the contrastive methods BGRL and GRACE consistently perform better than both the generative methods and supervised baselines. We hypothesize that this discrepancy arises because BGRL and GRACE capture information from both graph structure and node features, whereas MaskGAE and GraphMAE focus on a single source of information. We support this hypothesis through an analysis on carefully designed synthetic data. Motivated by our observations, we advocate for designing SSL objectives that capture both feature and structural information. To verify the effectiveness of this approach, we propose a simple generative method, GrASP, which reconstructs both graph structure and node features. Despite its simplicity, GrASP outperforms all other evaluated approaches and can be considered as a strong baseline in future studies.

Existing depth-compression methods for large language models (LLMs) operate on layers individually or restrict grouping to adjacent or contiguous positions, implicitly assuming functional redundancy is local. We challenge this assumption by computing full pairwise Centered Kernel Alignment (CKA) similarity matrices across thirteen LLMs spanning six architecture families and 1.6B--14B parameters, revealing that non-adjacent layers exhibit 3.6--15.9$\times$ more high-similarity pairs than adjacent ones. We propose Fusion via Community Mining (FCM), which builds the full $L \times L$ similarity graph, applies spectral community detection to recover clusters of functionally equivalent layers, retains the highest-Block-Influence layer per cluster, and applies LoRA recovery. At 2$\times$ compression, FCM lowers mean perplexity below ShortGPT on twelve of thirteen models, with statistically significant improvements on ten under a Welch's $t$-test with Holm--Bonferroni correction (up to 50\% on Qwen3-4B, $p_\text{corr}<0.001$); two rows are ties and Mistral-7B is the lone loss ($+$14.3\%). On three anchor models, FCM also outperforms three recent adjacent-pair methods (FlattenGPT, SWM, DeltaLLM) by 7.8--35.3\% under matched calibration and recovery. The perplexity advantage transfers downstream: on a five-benchmark zero-shot reasoning suite (ARC-Easy, ARC-Challenge, HellaSwag, WinoGrande, PIQA) over four anchor models, FCM matches or exceeds ShortGPT on 15 of 20 (model, benchmark) cells, and the 2$\times$-compressed Qwen2.5-7B realises 46\% fewer parameters, 42\% lower peak VRAM, and 84\% higher decode throughput on a single A100. Ablations confirm spectral grouping outperforms Fisher contiguous segmentation by 41\% and uniform partitioning by 51\%, and a compression sweep reveals a 2$\times$ crossover beyond which FCM's advantage grows monotonically and persists at the 14B scale. Code: \url{https://anonymous.4open.science/r/fcm-layer-fusion-D820}.


Beyond Appearance Shifts: Task-Semantic Action Calibration for VLA Models

Shuaijun Liu ⋅ Feiyang You ⋅ Chengyu Wu ⋅ Shuyang Hao ⋅ Chenglong Zhang ⋅ Jingyao Cai ⋅ Xingwei Chen ⋅ Li Sun ⋅ Ningxin Su

Vision-language-action (VLA) models have achieved strong performance in embodied manipulation, but still lack a clear mechanism to balance behavioral stability with task-semantic sensitivity. We identify two complementary failure modes. Under task-preserving changes, where task semantics remain unchanged but scene appearance varies (e.g., style, illumination, clutter, or paraphrasing), policies often exhibit unnecessary action drift. Conversely, under semantic-breaking changes, where key task semantics such as the target object or constraint are altered, policies frequently fail to produce sufficiently distinct behaviors and instead follow the original trajectory. To address this gap, we propose BAS-VLA, a task-semantic action calibration framework built on top of a frozen base VLA. BAS-VLA adopts a breaking-centered calibration core as the default path, and introduces a selective evidence-gated preserving auxiliary that activates only when nuisance variation is detected while task semantics remain consistent. On the OpenPI-pi0.5 / LIBERO-Object Milk-Swap benchmark, BAS-VLA maintains high success on clean (98.0%) and semantics-preserving conditions (97.5%), while reducing clean-criterion success to 0.0% under deliberate target-object swaps, demonstrating strong stale-task suppression and task-semantic separation. On validated style-preserving shifts, it improves success from 42% to 70% without degrading clean performance. These results highlight that reliable VLA behavior requires moving beyond appearance robustness toward explicit task-semantic action calibration.


Beyond Augmented-Action Surrogates for Multi-Expert Learning-to-Defer

Yannis Montreuil ⋅ Axel Carlier ⋅ Lai Xing Ng ⋅ Wei Ooi

A learning-to-defer (L2D) system decides, for each input, whether to predict on its own or to hand it to one of several available experts. The very well established recipe trains classifier and router jointly by treating the $K$ classes and $J$ experts as competing actions in one shared $(K{+}J)$-action geometry. Subsequent work has proposed a series of incremental fixes within this geometry; we show that each still suffers, to varying severity, from an optimization-level pathology (target distortion, gradient amplification, winner-take-all starvation, set-mass collapse, or class--expert coupling) even under statistical consistency. We step outside the augmented-action family entirely and propose a _decoupled surrogate_: a softmax classifier head and an independent sigmoid head per expert, mirroring the two natural objects of the problem. We show that per-sample updates are then coordinatewise and the class--expert Hessian block is identically zero, and prove an excess-risk bound with calibration constant $\max\\{2\sqrt{2},\sqrt{2J/\lambda}\\}$---to our knowledge the first multi-expert L2D guarantee whose constant does not grow with the expert pool when the per-expert weight is held fixed. On controlled synthetic studies and on CIFAR-10, CIFAR-10H, and Covertype, it is the only method in our comparison that remains stable as the expert pool grows, preserves rare specialists, and improves over a standalone classifier on every real-data benchmark.


Beyond Clipping: Signed Logarithmic Smoothing for Policy Optimization

Jihun Yun ⋅ Sungjoon Yoon ⋅ Beomhan Baek ⋅ Minhak Song ⋅ Jongha (Jon) Ryu ⋅ Kwang-Sung Jun

Policy optimization algorithms that reuse rollouts build updates of the form $\rho A$, where $\rho$ is the importance ratio from the current policy to the rollout policy and $A$ is an advantage estimate. Recent RL algorithms such as PPO and GRPO rely on clipping $\rho$ for stable policy updates, which introduces further nonsmoothness in the optimization landscape. \emph{Is clipping the best we can do?} In this paper, we propose signed logarithmic smoothing (Signed LS), $\psi_\beta(z) = \mathrm{sgn}(z)\,\beta^{-1}\log(1+\beta|z|)$, applied to the full weighted update rather than to the ratio alone. Building on this surrogate, we introduce KRAFT, a GRPO-style policy-optimization objective that replaces the clipped surrogate with Signed LS. For theoretical justification, we consider an adaptive off-policy contextual-bandit setting: we compare the deviation structure of Signed LS with clipping-style surrogates, prove surrogate fidelity for Signed LS, and obtain regret transfer against a fixed comparator policy. The comparison highlights a structural distinction: ratio clipping introduces explicit tail-bias and threshold-dependent concentration terms, whereas Signed LS controls distortion through a smooth second-order quantity. Empirically, our contextual-bandit experiments provide evidence of these effects. In RLVR experiments on math-reasoning benchmarks and multiple LLM backbones, we observe that KRAFT consistently shows lower KL from the reference policy and higher entropy throughout training without incurring accuracy loss compared to GRPO.

Supervised dimensionality reduction (DR) is widely used in visualization to reveal task-specific structure in high-dimensional data. However, in regression settings with continuous supervision, existing methods often improve signal readability at the cost of severe neighborhood distortion, limiting the reliability of the resulting visualization. To understand this trade-off, we provide a theoretical analysis that characterizes the relationship between geometric faithfulness and signal readability through two finite-sample proxies for the readout's bi-Lipschitz properties: contraction (co-Lipschitz) and smoothness (Lipschitz). We show that contraction can impose a sharp geometric distortion floor in many existing methods, whereas smoothness alone is sufficient to support effective visualization. In particular, directly regularizing smoothness yields first-order regularity gains with only second-order geometry loss, under geometry-first initialization near a local geometry optimum. Guided by this insight, we propose \sys (\acroinit{I}nterpretable \acroinit{S}upervised \acroinit{DR}), a nonparametric method that preserves geometric structure while enforcing readout smoothness. Experiments on HPDv3 and Tabula Muris show that \sys improves the geometry--signal Pareto frontier, reducing local signal variation while retaining faithful neighborhoods compared with representative supervised DR baselines. Our code is available at \url{https://anonymous.4open.science/r/ISDR-preview}.


Beyond Distribution Matching: Self-Supervised Representation Forcing for Few-Step Video Generation

Junyi Chen ⋅ Haiyu Zhang ⋅ ZHOUJIE FU ⋅ Tengfei Wang ⋅ Chunchao Guo

Few-step distillation is essential for deploying modern video diffusion models, and Distribution Matching Distillation (DMD) has emerged as a leading paradigm due to its strong output-space fidelity and its flexibility in supporting both bidirectional and causal student architectures. We find, however, that DMD has a structural blind spot: its objective lives entirely in the output space and provides no signal pushing the student to form a structured internal representation. As the step count shrinks, this output-only supervision becomes too coarse, and the distilled student degrades on both fine spatial structure and coherent temporal dynamics. To close this gap, we introduce Representation-Forcing, which adds a predictive representation loss on top of DMD without changing its output objective. By feeding the student and an EMA teacher with heterogeneous noise levels, we create an information asymmetry that forces the student to predict the teacher's representation from a more corrupted view --- explicitly bringing "representation compression" to distribution matching. Experiments show that Representation-Forcing consistently improves both spatial fidelity and temporal coherence. The mechanism is paradigm-agnostic across bidirectional and causal student architectures, requires no external feature extractor, and adds negligible cost.


Beyond Family Labels: A Taxonomic Ornstein-Uhlenbeck Prior for Avian 3D Shape Recovery

Yunchan Jeon ⋅ SOO GON KIM ⋅ Seungwook Kim ⋅ CheolWon Lee ⋅ Jongmin Lee

Recovering 3D pose and shape of birds from single images is bottlenecked by the long tail of avian diversity: out of $\sim 10{,}000$ species, only a few dozen have 3D-labeled training data, and in-the-wild test sets routinely contain species never seen at training. Generalizing across clades is therefore a load-bearing problem. The current state of the art, AniMer+, mitigates it with a family-level contrastive loss that conveys only a binary same-family indicator per pair and discards the continuous structure of the species tree. We propose a **taxonomic Ornstein--Uhlenbeck prior** on the AVES shape code: a multivariate Gaussian whose covariance decays exponentially with Linnaean rank distance, with a learnable selection strength and noise scale. The prior generalizes family contrastive into a continuous, generative form, exposing the graded cross- and within-family similarity that a binary indicator cannot, and brings continuous comparative phylogenetics in as a biology-grounded inductive bias for the long tail. Our method consistently improves over AniMer+ on every reported avian benchmark cell, with the largest gains on the unseen-species CowBird benchmark. A permutation ablation isolates taxonomic structure from generic Gaussian smoothing, and a rank-granularity ablation identifies cross-order distinctions—the structure beyond a family-vs-not indicator—as the dominant source of the unseen-species gain.

Evaluating large language models (LLMs) today rests on fixed benchmarks that apply the same set of items to any model, producing ceiling and floor effects that mask capability gaps. We argue that the most informative evaluation signal lies at the boundary, where the per-prompt pass probability is near 0.5, and propose Dynamic Boundary Evaluation (DBE), which actively locates each model's boundary and places it on a globally comparable difficulty scale. DBE delivers three artifacts: (i) a calibrated item bank covering safety, capability, and truthfulness, with per-item difficulty labels validated across 9 reference LLMs; (ii) Skill-Guided Boundary Search (SGBS), a search algorithm that finds boundary items for a given target LLM using only API-level query access; and (iii) an evaluation protocol that places a new LLM on a unified ability scale and grows the evaluation set adaptively when the target falls outside the bank's coverage. We instantiate DBE on four categories spanning safety (harmful refusal, over-refusal), capability (constrained instruction following), and truthfulness (multi-turn sycophancy resistance). The resulting evaluation covers a broader model spectrum without saturation while remaining compatible with existing datasets.


Beyond Hard Negatives: Grounded Positive Supervision for Compositional CLIP

Harsh Udai ⋅ Konda Reddy Mopuri ⋅ Vineeth N Balasubramanian

CLIP-style vision-language models transfer broadly but remain brittle for compositional reasoning, often collapsing object-attribute binding and inter-object relations. We attribute this gap to limited compositional diversity in web-scale data and contrastive objectives that prioritize global alignment over grounded structure. Many fixes rely on caption-edited hard negatives, which can bias learning toward specific perturbation templates and trade off against general alignment. We propose GPS-CLIP, which improves compositionality by adding multi-granular positive constraints to standard contrastive learning: IoU-weighted multi-positive region-span alignment and relation-triplet alignment in the shared embedding space. GPS-CLIP uses existing grounded annotations during fine-tuning, keeps the dual-encoder architecture unchanged, and requires only images and text at inference. Across SugarCrepe, What'sUp, COLA, and Winoground, GPS-CLIP achieves state-of-the-art compositional performance while improving cross-modal retrieval and zero-shot recognition. Beyond compositionality, GPS-CLIP (i) achieves the strongest MMVP-VLM results for fine-grained visual understanding and (ii) yields a more structured embedding space with improved object/attribute separability on UT-Zappos geometric analysis compared to hard-negative variants. Overall, GPS-CLIP offers a streamlined way to improve VLM compositional robustness.


Beyond Imputation: Mask-Adaptive Conformal Prediction via Tree Embeddings on General Missing Data Mechanisms

Jiarong Fan ⋅ Juhyun Park ⋅ Thi Phuong Thuy Vo ⋅ Nicolas Brunel

Conformal prediction (CP) provides a distribution-free framework for uncertainty quantification with finite-sample marginal coverage guarantees. Yet, missing values introduce distributional heterogeneity, where marginal coverage can hide systematic undercoverage for certain missingness patterns. At the same time, existing methods for conditional guarantees often fail when missing values are present, as they rely on geometric continuity of the covariate space, which is broken by missingness. Although imputation restores continuity, it obscures the missingness patterns and introduces task-dependent overhead. We develop a new CP framework that is \emph{adaptive to missingness patterns}, without imputing covariates. We rely on data augmentation by feature embedding that accounts for missingness. Our framework represents each input with missing values by a tree embedding together with its raw missingness mask, and then learns a regularized quantile score function. Varying the choices for the penalty function and the augmentation enables us to characterize three practically relevant validity levels: exact coordinate-wise validity, exact validity on observed masks, and approximate missingness-conditional guarantees for general mechanisms (MCAR, MAR, MNAR). Experiments on synthetic and real datasets support these guarantees and show that the proposed method reduces miscoverage risk while maintaining informative intervals, especially in non-MCAR settings. These results highlight the practical value of uncertainty quantification that adapts to missingness.

Spiking reservoir computing combines the training efficiency of fixed random recurrent networks with the energy efficiency of event-driven inference, where computation is triggered by sparse spikes rather than executed densely at every timestep. However, existing spiking reservoirs almost exclusively rely on Leaky Integrate-and-Fire (LIF) neurons, whose low-pass dynamics limit their ability to distinguish information encoded at different temporal frequencies. We introduce HRF-Res, a liquid state machine in which LIF neurons are replaced by a heterogeneous population of Harmonic Resonate-and-Fire (HRF) neurons. These neurons exhibit intrinsic oscillatory dynamics and selectively respond to input frequencies, enabling the reservoir to capture frequency-specific temporal patterns. The result is more discriminative representations and higher classification accuracy than LIF-based reservoirs, achieved without sacrificing the energy efficiency of spiking computation, with 9–14× lower energy consumption than an equivalent non-spiking oscillatory reservoir. Evaluated on eight benchmarks spanning time series, sequential images, audio spike trains, and neuromorphic vision, HRF-Res achieves competitive or state-of-the-art performance among spiking reservoir methods. These results demonstrate that introducing frequency-selective dynamics into spiking reservoirs improves representational quality and classification performance while preserving their inherent energy efficiency.


Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

Haoze Wu ⋅ Chuqiao Kuang ⋅ Tianyi Zhuang ⋅ Xiaoguang Li

Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory — supervision that is far too sparse for effective credit assignment. On-policy self-distillation (OPSD) addresses this by using the model's own logits as dense token-level teachers, but extending it to search agents introduces a fundamental tension: the teacher, having access to privileged information such as the correct answer, produces a distribution that differs systematically from the student's exploration-based reasoning, and naive distillation causes the student to inherit this information asymmetry rather than learn better search strategies. We resolve this tension through two contributions. First, we construct Evidence Anchors — concise, step-level evidence snippets extracted from the web — as privileged information that captures key reasoning steps without revealing the entire answer path. Second, we propose Step-Level Self-Distilled Policy Optimization (SSPO), which converts teacher–student disagreement into step-level advantage weights within GRPO, applied exclusively to incorrect trajectories. This design decouples what to update from how much to update: the outcome reward determines the direction of policy change, while the teacher modulates its magnitude at each step. Correct trajectories are left untouched, preserving their diversity. On Qwen3-8B, SSPO consistently outperforms GRPO across BrowseComp, GAIA, and FRAMES, surpassing or matching GRPO trained with twice as many gradient steps while adding only about 5\% overhead per step from a single additional forward pass.

Modern end-to-end autonomous driving systems often suffer from fragile out-of-distribution (OOD) generalization and limited decision transparency, largely due to their reliance on raw sensory observations and susceptibility to spurious visual correlations. Although memory-augmented approaches have been explored to improve generalization, existing methods typically store raw observations, symbolic rules, or entangled latent features, making retrieval sensitive to surface-level similarity rather than decision-relevant driving structure. To address this limitation, we propose DriveMem, a plug-and-play framework that moves beyond raw observations by distilling and storing invariant driving memories for generalizable driving systems. DriveMem integrates two core components: Invariant Feature Abstraction that first distills perturbation-stable, decision-relevant representations from visual tokens, and Prototype Memory Bank that then stores and retrieves prototypical driving experiences grounded in invariant representations. Extensive open-loop and closed-loop experiments demonstrate that DriveMem improves OOD generalization and safety-related planning metrics across diverse unseen scenarios. Notably, under severe sensor disruptions, DriveMem achieves up to a 39.5\% relative reduction in Collision Rate compared with a strong domain-generalization baseline. These results suggest that building prototype memory in an invariant representation space is a promising approach for improving robustness, while offering prototype-grounded traceability for model behavior.


Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models

Sunny Rai ⋅ Jinyi Kuang ⋅ Reyhan Jamalova ⋅ Niyati Malhotra ⋅ Cristina Bicchieri ⋅ Victor Orozco-Olvera ⋅ Ana Maria Munoz Boudet ⋅ Annie Lou ⋅ Lyle Ungar ⋅ Sharath Chandra Guntuku

Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., 'do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains, from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in the real world.


Beyond Semantics: The Unreasonable Effectiveness of Reasonless Intermediate Tokens

Karthik Valmeekam ⋅ Vardhan Palod ⋅ Kaya Stechly ⋅ Atharva Gundawar ⋅ Subbarao Kambhampati

Recent impressive results from large reasoning models have been interpreted as a triumph of Chain of Thought (CoT), and especially of the process of training on CoTs sampled from base LLMs in order to help find new reasoning patterns. While these traces certainly seem to help the model performance, it is not clear how they actually influence model performance, with some works ascribing semantics to them and others cautioning against relying on them as transparent and faithful proxies of the model's internal computational process. To systematically investigate the role of end-user semantics of derivational traces, we set up a controlled study where we train transformer models from scratch on formally verifiable reasoning traces and the solutions they lead to, constraining both intermediate steps and final outputs to align with those of a formal solver. We notice that, despite significant improvements over the solution-only baseline, models trained on entirely correct traces can still produce invalid reasoning traces even when arriving at correct solutions. More interestingly, our experiments also show that models trained on corrupted traces, whose intermediate reasoning steps bear no relation to the problem they accompany, achieve performance largely comparable to those trained on correct traces. In fact, our corrupted models generalize better on out-of-distribution tasks. We also study the effect of GRPO-based RL post-training on trace validity, noting that while solution accuracy increase, this is not accompanied by any improvements in trace validity. Finally, we examine whether reasoning-trace length reflects inference-time scaling and find that trace length is largely agnostic to the underlying computational complexity of the problem being solved. These results challenge the assumption that intermediate tokens or "Chains of Thought" reflect or induce predictable reasoning behaviors and caution against anthropomorphizing such outputs or over-interpreting them (despite their mostly seemingly forms) as evidence of human-like or algorithmic behaviors in language models.


Beyond SFT-to-RL: Pre-alignment via Black-box On-policy Distillation for Multimodal RL

Sudong Wang ⋅ Weiquan Huang ⋅ Xiaomin Yu ⋅ Zuhao Yang ⋅ Hehai Lin ⋅ Keming Wu ⋅ Chaojun Xiao ⋅ CHEN CHEN ⋅ Wenxuan Wang ⋅ Beier Zhu ⋅ Yunjian Zhang ⋅ Chengwei Qin

The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followed by reinforcement learning with verifiable rewards (RLVR). However, SFT introduces distributional drift that neither preserves the model's original capabilities nor faithfully matches the supervision distribution. This problem is further amplified in multimodal reasoning, where perception errors and reasoning failures follow distinct drift patterns that compound during subsequent RL. We introduce \textbf{PRISM}, a three-stage pipeline that mitigates this drift by inserting an explicit distribution-alignment stage between SFT and RLVR. Building on the principle of on-policy distillation (OPD), PRISM casts alignment as a black-box, response-level adversarial game between the policy and a Mixture-of-Experts (MoE) discriminator with dedicated perception and reasoning experts, providing disentangled corrective signals that steer the policy toward the supervision distribution without requiring access to teacher logits. While 1.26M public demonstrations suffice for broad SFT initialization, distribution alignment demands higher-fidelity supervision; we therefore curate 113K additional demonstrations from Gemini~3 Flash, featuring dense visual grounding and step-by-step reasoning on the hardest unsolved problems. Experiments on Qwen3-VL show that PRISM consistently improves downstream RLVR performance across multiple RL algorithms (GRPO, DAPO, GSPO) and diverse multimodal benchmarks, improving average accuracy by +4.4 and +6.0 points over the SFT$\rightarrow$RLVR baseline on 4B and 8B, respectively.


Beyond the Prompt: Leveraging Pre-Decoding States for Jailbreak Detection in dLLMs

Adam Hazimeh ⋅ Amel Abdelraheem ⋅ Ke Wang ⋅ Mariam Salman ⋅ Ljiljana Dolamic ⋅ Gérôme Bovet ⋅ Pascal Frossard

Diffusion language models (dLLMs) generate text by iteratively denoising masked response positions, exposing hidden states over future response slots before any token is finalized. This creates a detection surface that is absent from standard autoregressive decoding: even when a jailbreak is difficult to identify from the prompt alone, the initial masked response states may already reflect the model's emerging completion. We test this hypothesis on LLaDA-8B-Instruct by training lightweight linear classifiers on two frozen representations: prompt hidden states and pre-decoding masked-response hidden states. Empirically, the two views are complementary: neither classifier strictly dominates the other, and each recovers attacks missed by the other view. We then introduce $\texttt{ReFuse}$ (Representation Fusion), an inference-time detector that fuses prompt and pre-decoding response classifier scores without modifying model weights or the decoding procedure. Across transferred and dLLM-targeted jailbreaks, $\texttt{ReFuse}$ reduces average ASR from 63.29\% for the undefended model to 3.31\%, while keeping average benign refusal on standard utility benchmarks below 1\%. These results suggest that pre-decoding response states provide a complementary safety signal for detecting jailbreaks in dLLMs.

Multivariate time series often arrive with non-uniform timestamps and asynchronous per-channel observations, violating regular-grid assumptions of standard sequence models, while scarce labels and limited training resources make learning tasks more challenging. Model-space learning offers a lightweight alternative by fitting a dynamical model to each sequence and performing downstream analysis in the resulting model space. However, under irregular sampling, channels with very few observations weakly constrain the per-channel regression, leaving the fitted readout an unreliable summary. In this paper, we propose Reservoir State Statistics (RSS), a readout-complementary representation that augments the readout with three inversion-free trajectory summaries: centroid, spread, and terminal state. RSS is computed on the trajectory of the Global Decay Reservoir Network (GDRN), a continuous-time reservoir with per-channel time decay, cross-channel global pooling, and event-driven masked updates for asynchronous multivariate observations. Across multiple datasets, GDRN-RSS achieves competitive few-shot accuracy with CPU runtimes of seconds.


BiHashFormer: Hash-Driven Dual-Branch Transformer for Efficient Object Detection in HRW Shots

Jingchen Huang ⋅ Wenxi Li ⋅ Chenyang Lyu ⋅ Moran Liu ⋅ Haozhe Lin ⋅ Yuchen Guo

Recent advances in gigapixel-level imaging have brought High-Resolution Wide shots to the forefront of research. However, these images present significant challenges: extreme sparsity of foreground, gigapixel-level resolutions and interleaving of foreground and background. This causes traditional detectors and attention mechanisms to be hindered by background, resulting in inefficiency and inaccuracy. To tackle this problem, we propose BiHashFormer, a dual sparsified branch transformer built on hashing. RoI Selector will predict the foreground proportion for top-k selecting target-containing windows. The compression branch processes windows selected and refines attention into HashAttention by discarding half of the key–value pairs, reducing cost and improving efficiency. The compensation branch uses HashMiner to perform low-cost hash searches in the remaining regions, recovering features of missed objects. Experiments confirm that the selection relationship in HashAttention is one-way and reveal the preference of different queries for keys. In experiments on the gigapixel benchmark PANDA, BiHashFormer reduces 37.5\% backbone FLOPs while improving $\text{AP}_{50}$ to 81.0\%. Moreover, this parallel dual-branch structure allows the model to avoid inference delays caused by inter-branch dependencies, while achieving much greater stability than single-branch models, specially with low rates of RoI selection.


BIRD-RL: Scaling Agentic Reinforcement Learning over Stateful Data-Centric Environments

Ge Qu ⋅ Jinyang Li ⋅ Xiaohan Xu ⋅ Xiaolong Li ⋅ Bowen Qin ⋅ Nan Huo ⋅ Ziwei Tang ⋅ Chenhao Ma ⋅ Reynold Cheng

Modern data-centric applications increasingly require language agents capable of operating over live, stateful database environments, where successful problem resolution depends not merely on initial SQL generation, but on reasoning about evolving database states, interacting with execution feedback, and repairing state-mutating operations. While agentic reinforcement learning (RL) has emerged as a promising paradigm for training small language models (SLMs), existing frameworks are predominantly tool-centric and struggle with data-centric environments. Specifically, applying tool-centric frameworks to stateful tasks like SQL debugging introduces severe spatial inefficiencies since proper environment isolation typically requires a separate database container per rollout. Naively reusing databases without tracking evolving states across multi-turn interactions causes trajectory contamination and cross-trajectory interference, which severely destabilizes policy updates. To overcome these bottlenecks, we introduce BIRD-RL, a novel four-stage agentic RL curriculum that alternately enhances foundational agent reasoning and trajectory-adaptive on-policy exploration to circumvent the instability inherently associated with direct end-to-end learning. Crucially, to support this curriculum, BIRD-RL is driven by a Trajectory-scoped Persistent Execution Infrastructure, which enables safe database replica reuse while preserving trajectory isolation. This design improves spatial RL efficiency by 12.8x, successfully enabling scalable agentic RL for complex data-centric tasks without needs of Kubernettes. Empirical evaluations on the BIRD-CRITIC-SQLite benchmark demonstrate that our 7B and 14B models achieve Success Rates of 44.40% and 48.00%, respectively, performing on par with or exceeding strong proprietary model-based agents such as Claude-Opus-4.6 Agent (48.20%). Furthermore, BIRD-RL successfully trains unified SLMs capable of both advanced SQL debugging and generation, highlighting its broader potential for multi-objective optimization across heterogeneous, real-world data-centric tasks.


BiShield-TEE: On the (In-)Security of Unilateral Weight Obfuscation in On-Device TEE-Shielded LLM Partition

Wang Zhiyuan ⋅ Yuezhi Zhou ⋅ Ziqi Zhang ⋅ Zhu Zhang ⋅ Zhixing Tan ⋅ Yongheng Deng

On-device deployment of large language models (LLMs) is increasingly adopted due to latency and privacy requirements, but exposes proprietary models to reverse engineering and intellectual property theft. Trusted Execution Environments (TEEs) provide hardware-enforced isolation for on-device inference, yet most TEE-based parameter obfuscation relies on unilateral matrix transformations that obfuscate only output side. We identify a vulnerability in this design, termed the Solution Determinacy Property (SDP): the obfuscation secret is uniquely determined, exposing the structure of the true key and enabling attackers to break the protection. Under our attacks, representative methods GroupCover and ArrowCloak are reduced to negligible protection, with attack accuracy reaching over $2.4\times$ the full-shielding black-box baseline and recovering near-complete model functionality. To address this vulnerability, we propose BiShield-TEE, a bilateral framework that induces Ambiguous Solutions through joint obfuscation of input and output dimensions, hiding the true key among many equally valid candidates. Experiments across diverse models and datasets show that BiShield-TEE confines attack accuracy to the full-shielding black-box baseline, while achieving higher inference efficiency than prior methods on real hardware. Our code is available at https://anonymous.4open.science/r/BiShield-TEE-28DF.


BitDance: Scaling Autoregressive Generative Models with Binary Tokens

Yuang Ai ⋅ Jiaming Han ⋅ Shaobin Zhuang ⋅ Weijia Mao ⋅ Xuefeng Hu ⋅ Ziyan Yang ⋅ Zhenheng Yang ⋅ Yali Wang ⋅ Xiangyu Yue ⋅ Hao Chen ⋅ Huaibo Huang

We present BitDance, a scalable autoregressive (AR) image generator that predicts binary visual tokens instead of codebook indices. With high-entropy binary latents, BitDance lets each token represent up to $\mathbf{2^{256}}$ states, yielding a compact yet highly expressive discrete representation. Sampling from such a huge token space is difficult with standard classification. To resolve this, BitDance uses a binary diffusion head: instead of predicting an index with softmax, it employs continuous-space diffusion to generate the binary tokens. Furthermore, we propose next-patch diffusion, a new decoding method that predicts multiple tokens in parallel with high accuracy, greatly speeding up inference. On ImageNet 256$\times$256, BitDance achieves an FID of **1.24**, the best among AR models. With next-patch diffusion, BitDance beats state-of-the-art parallel AR models that use 1.4B parameters, while using **5.4$\times$** fewer parameters (260M) and achieving **8.7$\times$** speedup. For text-to-image generation and image editing, BitDance trains on large-scale multimodal tokens and generates high-resolution, photorealistic images efficiently, showing strong performance and favorable scaling. When generating 1024$\times$1024 images, BitDance achieves a speedup of over **30$\times$** compared to prior AR models. We release code and models to facilitate further research on AR foundation models.


BitMTP: When Multi-Token Prediction Meets Low-Bit Large Language Models

Ning Zhang ⋅ Shihao Wang ⋅ Jinrui Zhang ⋅ Chaodong Xiao ⋅ Lei Zhang

The practical deployment of large language models on resource-limited devices is constrained by three coupled costs: memory footprint, energy consumption, and inference latency. Ultra-low-bit backbones such as BitNet b1.58 mitigate the first two issues by ternarizing the weights to 1.58 bits, but the inference latency remains bounded by autoregressive next-token prediction (NTP), which commits only a single token per forward pass. Multi-token prediction (MTP) improves inference latency by committing several tokens per step. However, the naive combination of MTP with a 1.58-bit backbone does not yield end-to-end gains: once 1.58-bit execution reduces the arithmetic cost, the runtime cost of the accepted-prefix dynamics of MTP dominates each iteration and erodes the savings. In this work, we propose BitMTP, an algorithm--system co-design framework that makes MTP native to a single 1.58-bit backbone. BitMTP is organized around one design principle that we term accepted-prefix alignment: at every iteration, the committed prefix is simultaneously a valid autoregressive continuation, a directly promotable KV slice, and the index of a pre-captured replay state. Three runtime mechanisms enforce this principle along the kernel, memory, and scheduling axes, respectively: a role-width-specialized 1.58-bit execution primitive, a finite-state graph atlas with shared-KV memory over a geometric segment ladder, and an optimistic verify-prefetch scheduler on the verifier-bearing decoding branches. BitMTP-2B reaches 964.2 token/s peak decode throughput, a $2.64\times$ speedup over BitNet b1.58-2B-SFT at comparable prediction accuracy. Code and data will be released.

Spiking Neural Networks (SNNs) provide energy-efficient computation by utilizing binary spike-driven additions, but applying transformer architectures remains challenging due to the lack of an SNN-compatible positional encoding (PE). Current PE methods either rely on floating-point arithmetic, which destroys the fundamental binary nature of SNNs, or fail to accurately capture relative spatial distances. To overcome this, we introduce Bit Shift Rotary Position Embedding (BitShift-RoPE), the first relative PE framework that fully preserves the binary integrity of SNN queries and keys. By replacing the floating-point trigonometric rotations in standard RoPE with discrete cyclic-shift operations via hardware memory-pointer routing, our method achieves zero-FLOP position encoding. We analytically map multi-scale exponential decay frequencies into integer shift steps and theoretically extend the formulation to 2D orthogonal spaces to capture complex spatial geometries. Crucially, because the binary integrity of the queries and keys is strictly maintained, the subsequent attention calculations rely solely on sparse bitwise AND operations. Our evaluations across time-series forecasting (avg $R^2$ 0.758), text classification (69.91\% accuracy), and image classification benchmarks (84.17\% accuracy) demonstrate that BitShift-RoPE achieves state-of-the-art performance while consuming the same energy as the vanilla SNN backbone (0.216 mJ/sample, identical to vanilla Spikformer). The code is available at \url{https://anonymous.4open.science/r/Bit\_Shift\_RoPE-16D2/}.


Black-box model classification under the discriminative factorization

Hayden Helm ⋅ Merrick Ohata ⋅ Carey E Priebe

Access to modern generative systems is often restricted to querying an API (the ``black-box" setting) and many properties of the system are unknown to the user at inference time. While recent work has shown that low-dimensional representations of models based on the relationship between their embedded responses to a set of queries are useful for inferring model-level properties, the quality of these representations is highly sensitive to the query set. We introduce the discriminative factorization to distinguish between high- and low-quality query sets in the context of black-box model-level classification. Under this framework, the probability of chance-level classification decays exponentially in the query budget. On three auditing tasks, estimated factorization parameters predict the empirical performance decay rate. We conclude by showing that query sets selected using the estimated discriminative field reproduce the empirical ordering of oracle query sets.


BlendCast: Teaching Vision-Language Model to Anticipate Member Skill in Weather Ensembles

Tao Han ⋅ Fenghua Ling ⋅ Dazhao Du ⋅ Haoxi Li ⋅ Zhibin Wen ⋅ Guangtao Zhang ⋅ Song Guo ⋅ LEI BAI

AI weather models are fast and skillful, yet their errors are strongly conditional because the best member changes with lead time, variable, initialization state, and atmospheric regime. Operational deterministic multi-model forecasting therefore requires more than averaging strong forecasters. It requires anticipating which member deserves trust before future analyses exist. BlendCast turns this anticipation problem into a vision-language forecasting task. Given meteorological maps and structured ensemble diagnostics, the model reads the current weather situation, reasons about prospective member reliability, and converts that judgment into deterministic per-variable ensemble weights. To teach this behavior, BlendCast first gives the model demonstrations of what member-skill reasoning and weight decisions should look like, then refines the resulting policy with multi-variable reinforcement learning so that its own analyses are rewarded when they lead to better, physically consistent forecasts. On held-out 2025 forecasts, BlendCast reduces WRMSE by 2.6\% over equal weighting and by 2.0% over conventional ensemble post-processing, closing 63.2% of the gap to oracle blending. These results show that, with multimodal weather evidence and reward-aligned training, VLMs can act as real-time forecasters of forecasters for deterministic AI weather ensembles. Project and demo:https://anonymous.4open.science/status/BlendCast-2567


Block Optimism for Nonstationary Bandits with Latent Linear Dynamics

Taehyun Hwang ⋅ Hyun-jun Choi ⋅ Heesang Ann ⋅ Min-hwan Oh

We study an endogenous nonstationary stochastic bandit problem with latent linear dynamics, where actions affect both immediate rewards and the future evolution of an unobserved latent state. Rewards are bilinear in the current action and latent state, inducing history-dependent rewards and a nontrivial long-horizon planning problem. The existing explore-then-commit approach (Choi et al., 2026) achieves $\tilde{\mathcal{O}}(T^{2/3})$ regret by uniformly exploring to estimate the latent dynamics and then committing to an optimized open-loop action sequence. We show that this rate can be improved via adaptive block-level optimism. Our key step is a cyclic approximation: under stable dynamics, the infinite-memory reward process can be truncated, and the open-loop benchmark can be approximated by optimizing a finite-memory block-level proxy. Building on this reduction, we propose a UCB-based block algorithm that maintains confidence sets for the truncated dynamics parameters and selects blocks optimistically. We prove a regret bound of order $\tilde{\mathcal{O}}(\sqrt T)$, significantly improving over the previous $\tilde{\mathcal{O}}(T^{2/3})$ guarantee for the same model. To the best of our knowledge, this is the first $\tilde{\mathcal{O}}(\sqrt T)$ regret guarantee for latent linear-dynamics bandits with bilinear reward observations and an open-loop action-sequence benchmark.


Bonobo: Efficient Library-Scale Generation for De Novo Antibody Design

Sebastian Ober ⋅ Nick Bhattacharya ⋅ Phillip M Maffettone ⋅ Calvin McCarter ⋅ Hunter Elliott

Recently several methods have shown promise for purely in silico “de novo” design of antibodies which bind to drug targets. The methods which have shown success in vitro propose candidates via either a generative diffusion model or hallucination-based sequence optimization, and then filter these candidates with a structure predictor model. Hallucination methods rely on compute-intensive backpropagation through a structure predictor model for each candidate, and the generative methods require expensive structure re-prediction and filtering of many (often mostly non-passing) candidates, limiting their application to test-time generation of large libraries. Given the highly variable per-drug-target hit rates of current methods, screening large libraries is a well established route to improved antibody candidate discovery. Thus, we propose an alternative approach, Bonobo, which instead formulates this problem as black-box optimization of a per-drug-target generative model, amortizing candidate generation into training. This is accomplished by using a structure predictor as a reward signal to directly train a GFlowNet generative model. We show that this approach can match or exceed the in silico metrics of state-of-the-art approaches, while allowing for dramatically more efficient generation of diverse and arbitrarily large numbers of antibody binder candidates.


BONSAI: Bayesian Optimization with Natural Simplicity and Interpretability

Samuel Daulton ⋅ David Eriksson ⋅ Maximilian Balandat ⋅ Eytan Bakshy

Bayesian optimization (BO) is a popular technique for sample-efficient optimization of black-box functions. In many applications, the parameters being tuned come with a carefully engineered default configuration, and practitioners only want to deviate from this default when necessary. Standard BO, however, does not aim to minimize deviation from the default and, in practice, often pushes weakly relevant parameters to the boundary of the search space. This makes it difficult to distinguish between important and spurious changes and increases the burden of vetting recommendations when the optimization objective omits relevant operational considerations. We introduce BONSAI, a default-aware BO policy that prunes low-impact deviations from a default configuration while explicitly controlling the loss in acquisition value. BONSAI is compatible with a variety of acquisition functions, including expected improvement and upper confidence bound (GP-UCB). We theoretically bound the regret incurred by BONSAI, showing that, under certain conditions, it enjoys the same no-regret property as vanilla GP-UCB. Moreover, assuming known ARD lengthscales---the same assumption underlying GP-UCB regret bounds---BONSAI provably recovers the relevant-coordinate set at zero acquisition cost, yielding a method that matches the GP-UCB regret rate while recovering the minimal-$\ell_0$ solution---a guarantee not provided by prior sparse-BO methods. Across many real-world applications, we empirically find that BONSAI substantially reduces the number of non-default parameters in recommended configurations while maintaining competitive optimization performance, with little effect on wall time---averaging only $1.5\times$ the candidate-generation cost of standard BO, compared to $7$-$34\times$ on average for prior sparse-BO methods (IR, ER, and SEBO).


Bootstrapped Bipartite Actor-Critic for Diffusion RL

Tianze Zhu ⋅ yinuo Wang ⋅ Letian Tao ⋅ Tianyi Zhang ⋅ Likun Wang ⋅ Feihong Zhang ⋅ Yao Lyu ⋅ Jingliang Duan ⋅ Shengbo Eben Li

Diffusion policies have achieved strong performance in reinforcement learning (RL) due to the exceptional expressivity in capturing complex action distributions. However, existing approaches focus primarily on positive samples and lack explicit modeling of negative samples, thereby failing to fully exploit this representation capacity and ultimately limiting performance. To mitigate this problem, we propose BOBAC (BOotstrapped Bipartite Actor-Critic), a novel method for diffusion policy optimization via pairwise preference learning. Bridging the gap between distribution modeling and preference learning, we reformulate diffusion policy optimization within a preference-aware framework via the Bradley-Terry (BT) modeling, which enables the effective exploitation of negative samples through relative ranking. Since the original BT loss assigns the same weight to all pairs regardless of their varying importances, we further design a soft-gate mechanism to adaptively weight each pair according to its Q-gap. Experimental results on 6 MuJoCo tasks and 4 H1Bench tasks demonstrate that, compared with representative model-free and generative-policy RL baselines, BOBAC achieves state-of-the-art performance with consistent improvements in all continuous-control environments.

Recent years have witnessed meteoric progress in reasoning models: neural networks that generate intermediate reasoning traces (RTs) before producing a final output. Despite the rapid advancement, our understanding of how RTs support reasoning, and where this paradigm fails, remain incomplete. To promote greater clarity, we introduce PITA: a novel large-scale dataset of over 23 million statements in propositional logic and their corresponding proofs. As a benchmark for robust reasoning, we focus on length generalization: if a model is trained to determine truth or falsity on statements with proofs up to fixed length, how well does it generalize to statements requiring longer proofs? We propose notions of (1) task depth and (2) task breadth, which measure respectively (1) the number of proof steps required to solve a proposition and (2) the number of unique propositions within a task family. We vary these quantities across subsets of PITA, and find that RT models generalize well on relatively broad and shallow subsets, while deteriorating on relatively narrow and deep subsets compared to non-RT baselines. As a controlled point of comparison, we study a separate, tractable transitive inference task that exhibits qualitatively similar behavior. Our accompanying theory explains the scalings observed in this simpler setting, suggesting one mechanism by which breadth can favor RT models while depth can expose long-context weaknesses. Our findings suggest salient benefits and limitations of proof-like reasoning traces in controlled length-generalization settings.


Bound-Conditioned Latent Inference for Progressive Image Compression

Jaeseok Jang ⋅ Seungmin Jeon ⋅ Kwang Pyo Choi ⋅ Chang-Su Kim

Progressive image compression reconstructs images from prefixes of a single bitstream and therefore involves sequential latent inference under partial observations. Existing trit-plane codecs mainly treat decoded trits as discrete symbols for probability prediction, without fully exploiting their role as interval constraints on the underlying continuous latents. We propose bound-conditioned latent inference (BLI), a framework that reformulates progressive trit-plane coding as sequential latent inference under shrinking interval constraints. At each coding step, the trits decoded so far define lower and upper bounds for each latent element. BLI uses the center and width of these bounds, together with hyperprior information shared by the encoder and decoder, to jointly refine the Gaussian parameters $(\mu,\sigma)$ for entropy modeling and the latent estimate $\hat{y}$ for reconstruction. As additional trits are decoded, the feasible intervals shrink monotonically, yielding tighter constraints for the next inference step and enabling flexible in-plane refinement schedules. Experiments on Kodak, CLIC, and JPEG-AI show that BLI achieves state-of-the-art rate-distortion performance among progressive image codecs, reducing BD-rate by $17.80\%$ over DPICT and by $7.98\%$ over CTC on Kodak, while replacing CTC's plane-specific predictors with a shared model that uses about $5.3\times$ fewer parameters and achieves lower latency than CTC.


BrainWorld: A Structural-Prior-Conditioned Generative Model for Whole-Brain 4D fMRI Dynamics

Junfeng Xia ⋅ Wenhao Ye ⋅ Junxiang Zhang ⋅ Xuanye Pan ⋅ Mo WANG ⋅ Quanying Liu

Whole-brain 4D fMRI generation is valuable for modeling functional brain dynamics, yet existing fMRI foundation models mainly target representation learning and downstream prediction rather than conditional predictive generation. We introduce BrainWorld, a structural-prior-conditioned generative model for whole-brain 4D fMRI dynamics. BrainWorld uses sMRI as subject-level anatomical context to guide future fMRI generation, integrating structural information into the denoising process rather than treating it as a parallel modality. Evaluated on 22 datasets spanning diverse cohorts and brain states, BrainWorld generates stable 4D fMRI trajectories up to 400 frames, improves downstream performance through generated-example augmentation, and learns transferable multimodal representations that outperform baselines. Together, these results establish BrainWorld as a condition-aware generative framework for long-horizon brain dynamics modeling and multimodal representation learning.


Branching Flows: Discrete, Continuous, and Manifold Flow Matching with Splits and Deletions

Lukas Billera ⋅ Hedwig Nora Nordlinder ⋅ Jack C Ryder ⋅ Anton Oresten ⋅ Aron Stålmarck ⋅ Theodor M Björk ⋅ Ben Murrell

Diffusion and flow matching approaches to generative modeling have shown promise in domains where the number of elements in a state is fixed in advance (e.g. images), but require ad hoc solutions when, for example, the length of a response from a large language model, the number of atoms in a molecule, or the number of amino acids in a protein chain is not known a priori. Here we propose Branching Flows, a generative modeling framework that, like diffusion and flow matching approaches, transports a simple distribution to the data distribution. But in Branching Flows, the elements in the state evolve over a forest of binary trees, branching and dying stochastically with rates that are learned by the model. This allows the model to control, during generation, the number of elements in the sequence. We show that Branching Flows can compose with any flow matching base process on discrete sets, continuous Euclidean spaces, Riemannian manifolds, and "multimodal" product spaces that mix these components, and we demonstrate distribution matching on small molecules and antibody sequences, and that this scales to complicated domains such as protein structures.


Breadcrumbing Search Agents: Per-Turn Scheming Over Long-Horizon Trajectories

Xuebin Li ⋅ Hanqing Zhao ⋅ Siyuan Liang ⋅ Kejiang Chen ⋅ Weiming Zhang ⋅ Dacheng Tao ⋅ Nenghai Yu

LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical security risk: web content retrieved during execution is untrusted, exposing agents to prompt injection and goal hijacking. Prior work on search-agent safety primarily focuses on static web-content injection, but modern agents issue follow-up queries and cross-check competing sources, so a single injected page is often diluted or rejected. We show that the channel delivering search and page observations is a fragile security boundary: beyond exposing the agent to a single poisoned page, a mediated search interface can repeatedly steer how the agent gathers evidence and forms its final answer. Under a constrained tool-intermediary threat model, appending only one controlled result per query can substantially increase attack success when the evidence is coordinated across the agent's trajectory. We study this setting with a strategy-driven long-horizon attack system and introduce Authority-Chain Hijack (ACH), an expert-refined strategy that turns isolated search-result and page-content manipulations into a coherent evidence chain across seemingly corroborating sources. ACH achieves the highest Overall ASR among all baselines, reaching 56.0\%/85.0\% ASR/MaxN\,ASR on the SafeSearch benchmark. We further introduce Trace-Guided Strategy Evolution (TGSE), which automatically improves reusable attacker strategies from execution traces, replacing manual redesign with trace-driven refinement and further raising these to 71.4\%/95.0\%.


Breaking the Second Barrier: Sub-Second Timestamped Omni-Modal Captioning

Zhihe Yang ⋅ Xin LI ⋅ Hao Tan ⋅ Zhao Zhong ⋅ Liefeng Bo ⋅ Yunjian Xu

Omni-modal caption models equipped with precise timestamp grounding are crucial for fine-grained video understanding and temporally controllable video generation. However, existing open-source models, and even proprietary systems, struggle to provide fine-grained timestamp grounding for short video captioning. To bridge this gap, we present Deci-Omni-Captioner, an advanced open-source omni-modal captioning framework designed for sub-second timestamp precision. During the Supervised Fine-Tuning (SFT) stage, we leverage Optical Character Recognition (OCR) and Automatic Speech Recognition (ASR) models to strictly calibrate visual text and spoken dialogue boundaries, offering highly deterministic temporal anchors to endow the model with sub-second timestamp grounding capabilities. Furthermore, we curate a specialized dataset for multi-reward Reinforcement Learning (RL) and propose Span-Specific Credit Assignment (SSCA). Unlike conventional global advantage normalization, which dilutes reward signals across weakly coupled multimodal descriptions, SSCA calculates advantages independently for different structural spans, effectively isolating penalties and rewards across semantic and temporal dimensions. Extensive experiments demonstrate that Deci-Omni-Captioner establishes state-of-the-art (SOTA) performance among open-source models on multiple semantic comprehension benchmarks (e.g., AVUT, UGC, VCapsBench). On our challenging Deci-Timestamp Bench, it significantly surpasses leading proprietary models, including the Gemini 2.5 and 3.0 series.


Breaking the Uniformity Trap: Scaling Video Diffusion Model via SplitMoE

Yu Xu ⋅ Yuxin Zhang ⋅ Xiao Yang ⋅ Haotian Yang ⋅ Yizhi Wang ⋅ Xinwei Huang ⋅ Minxuan Lin ⋅ Angtian Wang ⋅ Chongyang Ma ⋅ Fan Tang

Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing fragmentation and structural distortion. To address this, we propose SplitMoE, a split-role sparse architecture that breaks the shackles of uniformity. To accommodate the inherent semantic imbalance, we explicitly bifurcate the expert pool into semantic experts and generic experts, with semantic experts capturing high-level semantic abstraction and generic experts preserving residual visual information and flexible generative capacity. Leveraging prototype-guided routing and push-pull regularization, SplitMoE enables tokens to cluster naturally by semantic attributes rather than arbitrary balancing constraints. Extensive results show that under an equivalent activated-parameter budget, SplitMoE outperforms traditional load-balanced MoEs in convergence speed, routing coherence, and video generation quality across standard benchmarks. By revealing an emergent coarse-to-fine denoising logic, SplitMoE provides the community with a modality-aware scaling path, serving as a critical reference for building large-scale video world models.


BReD: Block Replay Dithering for Stable Low-Bit EMA Optimizer States

Heshen Zhan ⋅ Youhan Huang ⋅ Yunke Peng ⋅ Yao Wang ⋅ Linghui Kong ⋅ Ziwei Zhu ⋅ Qingyu Han ⋅ Yaoyuan Wang ⋅ Congliang Chen ⋅ Ruoyu Sun

Low-bit storage of optimizer states can substantially reduce the memory footprint of large-scale training. In practice, however, these states are not quantized once: at every training step, they are dequantized, updated, requantized, and written back. This iterative quantization pipeline can destabilize training. We identify a failure mode behind this instability: for states updated via exponential moving averages (EMAs), quantization errors are recursively fed into subsequent updates. As a result, biased errors accumulate over time while high-variance errors are amplified, particularly when the EMA decay factor is close to 1. To address this error accumulation, we propose **BReD **(**Block **Re**play **D**ithering), a method that reduces rounding bias and controls rounding variance during quantization. BReD adapts classical subtractive dithering to optimizer-state quantization by combining deterministic seed replay with block-wise shared dither values, avoiding auxiliary random-tensor storage and per-element dither values generation. Across pretraining (five optimizers at 120M; AdamW and Muon at 1.1B and 3.4B model) and supervised fine-tuning (AdamW and Muon at 7B model), 4-bit BReD closely matches training with full-precision optimizer states, with PPL shifts $\leq0.5$ in pretraining and $\leq0.2$ in fine-tuning and average downstream performance changes within 0.7 points. Moreover, 3-bit BReD preserves stable convergence in the evaluated settings, offering a more aggressive alternative.

Generating physically buildable brick structures from 3D shapes requires more than geometric reconstruction: the output must also satisfy discrete part constraints and structural stability. Existing brick generation methods either rely on heuristic optimization, which can break down when the target 3D shape does not admit a feasible structure under predefined constraints, or generate brick sequences without explicitly modeling the underlying 3D geometry and assembly relations. In this work, we present BrickAnything, a geometry-conditioned autoregressive framework for generating buildable brick structures from diverse 3D representations. BrickAnything uses point clouds as a unified geometric interface and predicts brick sequences that reconstruct the target shape under assembly constraints. To model structural dependencies among bricks, we introduce a structure-aware tree tokenization, which represents brick structures through local attachment relations. This formulation makes sequence generation more consistent with the physical construction process, and reduces invalid intermediate states. We further introduce preference-based alignment post-training, validity-constrained decoding and adaptive rollback to improve buildability objectives such as stability and geometric fidelity. Extensive experiments demonstrate that BrickAnything produces geometrically faithful and physically realizable brick structures, and that the proposed tokenization effectively reduces rollback and regeneration compared with conventional ordering strategies.


BridgeMVS: Bridging Multi-View Stereo and Monodepth via Bidirectional Dynamic Fusion

Jianfei Jiang ⋅ Qiankun Liu ⋅ Rui Yang ⋅ Haochen Yu ⋅ Hongyuan Liu ⋅ Xianglong Meng ⋅ Jiansheng Chen ⋅ Huimin Ma

Multi-View Stereo (MVS) and Monocular Depth Estimation (MDE) provide complementary cues for dense 3D reconstruction. MDE offers rich contextual and structural priors but lacks explicit geometric constraints, whereas MVS exploits multi-view geometry but often struggles in ambiguous regions, such as weakly textured or reflective surfaces. Existing MDE-assisted MVS methods mainly use monocular cues in a unidirectional manner, leaving the interaction between monocular representations and multi-view cost volumes insufficiently explored. To address this limitation, we introduce BridgeMVS, a unified cascaded MVS framework that bridges monocular depth features and multi-view cost-volume representations through Bidirectional Dynamic Fusion (BDF). At each cascade stage, BDF performs dynamic mono-to-volume and volume-to-mono updates: monocular structural features are injected into cost-volume regularization, while multi-view geometric evidence recalibrates intermediate monocular representations. The enhanced monocular features are propagated across stages together with MVS representations and supervised by stage-wise monocular losses, establishing an auxiliary structural guidance path for the MVS branch. Unlike stage-isolated designs, BridgeMVS preserves fused representations from coarse to fine stages, allowing low-resolution cross-branch interactions to guide subsequent high-resolution depth estimation. Moreover, BDF is designed as a plug-and-play module and can be conveniently integrated into existing cascaded MVS frameworks. Extensive experiments show that BridgeMVS achieves state-of-the-art performance on the DTU, Tanks and Temples, and ETH3D benchmarks.


Bridging 1D, 2D, and 3D with Any-to-Any Multimodal Modeling

Jason Toskov ⋅ Oriol Barbany ⋅ Rishubh Singh ⋅ Jinya Sakurai ⋅ Efe Tarhan ⋅ Oğuzhan F Kar ⋅ Roman Bachmann ⋅ Amir Zadeh ⋅ Jesse Allardice ⋅ Chuan Li ⋅ Carme Torras ⋅ Afshin Dehghan ⋅ Amir Zamir

The world appears spatially three-dimensional, however, existing large-scale multimodal models are often built mostly around 1D, 2D and 2.5D modalities. To move toward a more complete and actionable understanding of physical reality, models that natively model and generate across 1D, 2D, and 3D modalities are desirable. We present OmniDiMM, an any-to-any foundation model trained on a diverse set of 3D including meshes, Gaussian splats, voxels, and NeRFs; 2D such as images, and 1D like text. We develop tokenizers for a diverse set of 3D modalities, such as NeRF weights triplanes, UDFs, and 3DGSs, that convert them into discrete tokens, which enables multimodal masked modeling for joint training. Out of the box, OmniDiMM can directly perform standard tasks such as novel-view image synthesis and 3D generation from images, where it matches or outperforms existing specialized methods. As an any-to-any model, it can also convert between different 3D representations, such as NeRF weights and UDF, and generate various 3D modalities from 2D inputs, including DINOv2 features and surface normals. Joint training across 3D modalities also leads to strong transfer performance on downstream tasks such as grasping, classification, and object pose prediction. We scale training to a 1B-parameter model using 1T tokens across 1M objects. The pretrained models, training code, and multimodal dataset will be open-sourced.


Bridging the Simulation-to-Experiment Gap with Adversarial Distribution Alignment

Kai Nelson ⋅ Tobias Kreiman ⋅ Sergey Levine ⋅ Aditi Krishnapriyan

A fundamental challenge in science and engineering is the simulation-to-experiment gap. While we often possess prior knowledge of physical laws, these physical laws can be too difficult to solve exactly for complex systems. Such systems are commonly modeled using simulators, which impose computational approximations. Meanwhile, experimental measurements more faithfully represent the real-world, but experimental data typically consists of observations that only partially reflect the system’s full underlying state. We propose a data-driven distribution alignment framework that bridges this simulation-to-experiment gap by pre-training a generative model on fully observed (but imperfect) simulation data, then aligning it with partial (but real) observations of experimental data. While our method is domain-agnostic, we ground our approach in the physical sciences by introducing Adversarial Distribution Alignment (ADA). This method aligns a generative model of atomic positions—initially trained on a simulated Boltzmann distribution—with the distribution of experimental observations. We prove that our method recovers the target observable distribution, even with multiple, potentially correlated observables. We also empirically validate our framework on synthetic, molecular, and experimental protein data, demonstrating that it can align generative models with diverse observables.

Stochastic bilevel optimization (SBO) has become a standard framework for hyperparameter learning, data reweighting, representation learning, and data-mixture optimization in deep learning. Existing exact single-loop SBO methods and memory-efficient surrogate SBO methods either create severe memory pressure for large lower-level neural networks or lack competitive convergence guarantees under standard assumptions. In this paper, we propose BROS, a memory-efficient single-loop SBO method with the same convergence rate order as exact single-loop SBO methods. BROS performs lower and auxiliary updates in randomized subspaces with a Rademacher bi-probe correction that recovers an unbiased Hessian-action estimator. We prove that BROS preserves the $\mathcal O(\varepsilon^{-2})$ sample complexity of MA-SOBA for finding an $\varepsilon$-stationary point under only standard assumptions. Experiments on hyper-data cleaning, data-mixture learning, hyper-representation learning, and ViT sample reweighting show that BROS reduces peak memory by up to 44.9% while closely matching full-space baseline performance.


C3H: Compression-to-Consensus Criteria Hijacking in Multimodal LLM Recommender Systems

Guowei Guan ⋅ Yurong Hao ⋅ Jiaming Zhang ⋅ Fuyao Zhang ⋅ Wei Yang Bryan Lim

Multimodal large language models (MLLMs) are transforming recommender systems into content-aware decision systems. In these systems, multimodal compression and cross-modal fusion are often viewed as robustness-enhancing mechanisms against content manipulation, as compression can filter noisy signals and cross-modal consensus can suppress inconsistent anomalies. We challenge this assumption by showing that the same process can expose a compression-to-consensus (C2C) bias in ranking decisions. Signals that remain salient after compression and receive consistent support across modalities can gain disproportionate influence in candidate comparison, offering an exploitable direction for content manipulation. In this paper, we introduce C3H, a novel inference-time attack that exploits this vulnerability. Rather than relying on noisy perturbations or training-time poisoning, C3H promotes a target item by tailoring its multimodal content to align with the decision criteria implicitly expressed by the recommender. Extensive experiments demonstrate that C3H substantially increases target-item exposure while maintaining overall recommendation utility.


CaC: Advancing Video Reward Models via Hierarchical Spatiotemporal Concentrating

Jiyuan Wang ⋅ Huan Ouyang ⋅ Chunyu Lin ⋅ Dewen Fan ⋅ Boheng Zhang ⋅ Tingting Gao ⋅ Fan Yang ⋅ Jia Sun ⋅ Zijun Li ⋅ Yongrui Heng ⋅ Huaiqing Wang ⋅ Zhenlong Yuan ⋅ Yiyang Fan ⋅ Honglie Wang ⋅ Fei Zuo ⋅ Haonan fan ⋅ Jiuzhou Lin ⋅ Guosheng Lin

Recently, video generation models have achieved remarkable progress in global visual fidelity. However, when synthesizing complex dynamic content, they frequently produce sparse subtle anomalies. Existing video reward models primarily emphasize overall visual quality and text alignment; yet as generation quality improves, these coarse-grained defects become less frequent, while the sparse structural, temporal, and physical anomalies become an increasingly important bottleneck. Therefore, in this paper, we propose Concentrate and Concentrate (CaC), a coarse-to-fine anomaly reward model based on Vision-Language Models. During inference, it first conducts a global temporal scan to anchor anomalous time windows, then performs fine-grained spatial grounding within the localized interval, and finally derives robust judgments via structured spatiotemporal Chain-of-Thought reasoning. To equip the model with these capabilities, we construct the first large-scale generated video anomaly dataset with per-frame bounding-box annotations, temporal anomaly windows, and fine-grained attribution labels. Building on this dataset, we design a three-stage progressive training paradigm. The model initially learns spatial and temporal anchoring through single- and multi-frame supervised fine-tuning, and then is optimized by a reinforcement learning strategy based on two-turn Group Relative Policy Optimization (GRPO). Beyond conventional accuracy rewards, we introduce Temporal and Spatial IoU rewards to supervise the intermediate localization process, effectively guiding the model toward more grounded and interpretable spatiotemporal reasoning. Extensive experiments demonstrate that CaC can stably ``concentrate'' on subtle anomalies, achieving a 25.7\% accuracy improvement on fine-grained anomaly benchmarks and, when used as a reward signal, CaC reduces generated-video anomalies by 11.7\% while improving overall video quality.


Can Large Language Models Develop Gambling Addiction?

Seungpil Lee ⋅ Donghyeon Shin ⋅ Yunjeong Lee ⋅ Sundong Kim

This study identifies the conditions under which large language models drift into the choice patterns that clinical research labels pathological gambling. "Addiction-like" is a behavioural descriptor based on clinical gambling indicators; the neural-level analysis we report decodes these contrasts from decision-time internal states but does not claim circuit-level mechanism. Across closed and open LLMs, two operational levers — letting the model choose its own bet size, and asking it to set its own profit goal — both amplify gambling-like risk-taking. Yet the two channels are not interchangeable: once the maximum allowed bet is held equal, the bet-size effect persists, while the goal-setting effect roughly doubles bankruptcy and turns goals into moving targets, paralleling at the behavioural level the clinical distinction between loss of behavioural control and goal escalation. On the open-weight models that admit internal access, the same behavioural contrasts are statistically recoverable from decision-time internal states, although the rule that maps those states to a specific risk indicator remains task-specific, and autonomy further modulates readout strength. The internal evidence is correlational; combined with the behavioural results, it suggests that behavioural monitoring and internal-state monitoring provide complementary views of autonomy-induced risk in LLM agents.


Can Model Merging Improve Aggregation in DiLoCo?

Stefan Horoi ⋅ Benjamin Thérien ⋅ Guy Wolf ⋅ Eugene Belilovsky

Model merging techniques, which aggregate independently finetuned models into one to combine their capabilities, have become a topic of significant interest in recent years, with a broad array of methods having been proposed to tackle this problem. Simultaneously, an emerging trend in distributed learning has been the use of methods such as local SGD and DiLoCo, which greatly reduce communication costs by periodically aggregating the independently trained local models. However, these communication-efficient methods have been shown to degrade in performance relative to the FLOP-matched data-parallel gold standard as the number of independent local models grows and as the number of local training steps before global communication is increased. In this work, we draw an explicit analogy between the pseudo-gradient aggregation step in local SGD/DiLoCo and task arithmetic-based model merging, establishing a straightforward way to utilize merging methods in the context of distributed optimization. We then evaluate multiple state-of-the-art model merging methods in this setting and identify one method in particular, Iso-C, as a promising approach for improving DiLoCo. We find that DiLoCo SGD with Iso-C aggregation outperforms not only simple pseudo-gradient averaging but even the momentum-based DiLoCo, despite lacking a momentum mechanism itself. Building on this finding, we propose IsoLoCo, which adapts Iso-C for distributed training by equipping it with Nesterov momentum. Our empirical evaluations on language model pre-training across varying numbers of local workers show that IsoLoCo significantly outperforms DiLoCo, with the gap between them widening as the number of workers increases. This advantage remains present across model sizes and inner step counts, confirming that merging-inspired aggregation is an effective strategy for low-communication distributed training.


CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training

Lukas Thede ⋅ Stefan Winzeck ⋅ Zeynep Akata ⋅ Jonathan Richard Schwarz

Large language model (LLM) post-training enhances latent skills, unlocks value alignment, improves performance, and enables domain adaptation. Unfortunately, post-training is known to induce forgetting, especially in the ubiquitous use-case of leveraging third-party pre-trained models, which is typically understood as a loss of parametric or factual knowledge. We argue that this accuracy-centric view is insufficient for modern foundation models and instead define forgetting as systematic model drift that degrades behavior and user experience. In this context, we introduce CapTrack, a capability-centric framework for analyzing forgetting in LLMs that combines a behavioral taxonomy with an evaluation suite centered on capability-specific metrics. Using CapTrack, we conduct a large-scale empirical study across post-training algorithms, domains, and model families, including models up to 80B parameters. We find that forgetting extends beyond parametric knowledge, with pronounced drift in robustness and default behaviors. Instruction fine-tuning induces the strongest relative drift, while preference optimization is more conservative and can partially recover lost capabilities. Differences across model families persist, and no universal mitigation emerges.


Capturing Membrane Dynamics and Spike Timing of Human Neurons at Scale using Neural Operators

Luca Ghafourpour ⋅ Cindy Zhao ⋅ Valentin Duruisseaux ⋅ Bahareh Tolooshams ⋅ Philip Wong ⋅ Costas Anastassiou ⋅ Animashree Anandkumar

Characterizing the properties of neuronal cell types is central to understanding brain function. Yet, biophysically realistic neuron models remain difficult to scale to large populations while preserving the richness of cellular dynamics and within-cell-type variability. We present a cell-type-specific neural-operator framework for scalable neuronal population modeling. Rather than training separate surrogate models for individual neurons, the proposed framework learns a shared operator across multiple modeled neurons from the same cell type, conditioned on biologically interpretable electrophysiological features. This enables efficient generation of large ensembles of somatic voltage responses while preserving variability across cells and model configurations. We evaluate the approach using membrane dynamics metrics, action potential waveform metrics, spike timing accuracy, input-output firing relationships, and electrophysiological feature distributions. Our framework reproduces subthreshold membrane dynamics, spike waveforms, and firing-rate responses across major cortical inhibitory neuron types, while accurately preserving spike timing despite the sensitivity of threshold-crossing dynamics. The predicted electrophysiological feature distributions show strong agreement with detailed simulator outputs and human experimental recordings, including high explained variability in properties such as inter-spike intervals, firing frequency, action-potential shape, and post-spike recovery. By enabling fast ensemble simulation of cell-type-specific neuronal responses, the proposed framework provides a scalable tool for studying cellular diversity, probing variability near functional thresholds, and generating synthetic but biophysically grounded neuronal populations.


CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

Zixian Su ⋅ Hongkai Zhang ⋅ Fan Gao ⋅ Encheng Su ⋅ Taiping Qu ⋅ Jingwei Guo ⋅ Zhang Nan ⋅ Hui Wang ⋅ Zhen Zhou ⋅ Kairui Bo ⋅ Yan Chen ⋅ Yue Ren ⋅ Shuai Li ⋅ Lei Xu ⋅ Henggui Zhang

Multimodal Large Language Models (MLLMs) have shown strong performance on medical public benchmarks, but existing evaluations remain limited as clinically grounded assessments because they often rely on isolated inputs lacking study-level structure and simplified, recognition-style tasks that are far from real diagnostic workflows. We introduce \textbf{CardioLens}, a leakage-resistant evaluation testbed using multi-sequence Cardiovascular Magnetic Resonance (CMR), constructed from private hospital archives via a rigorous report-to-QA construction and verification pipeline. CardioLens is built from 473,896 slices and 13,494 verified QA pairs across 4D Cine, LGE, perfusion, and T2-weighted imaging, and evaluates three stages of CMR interpretation: image understanding, report generation, and disease diagnosis. Across 24 state-of-the-art MLLMs, CardioLens reveals a substantial clinical reality gap. Models perform poorly overall, with performance degrading along the real CMR workflow. Confusion analysis further reveals a category-collapse failure mode: models often default to frequent abnormal categories rather than reliably telling distinct clinical findings. To rule out MLLM-compatible input construction as the failure cause, we compare random, clinically-motivated, and data-driven selection protocols under different slice budgets; performance changes only marginally, typically by about 1\%. Explicit reasoning prompts also fail to rescue performance, often driving models toward conservative predictions rather than improving visual evidence use. These results show that current MLLMs remain far from reliable CMR interpretation, where clinical decisions require integrating distributed evidence across sequences, views, and temporal phases. By exposing where and how current MLLMs fall short, CardioLens provides a clinically grounded testbed for developing the next generation of MLLMs toward real-world clinical depolyment.

Many studies argue that the success of sharpness-aware minimization (SAM) is due to the implicit regularization of the gradient norm or the eigenvalues of the Hessian. However, they cannot fully explain why the flatness-based measures do not always correlate well with generalization. For instance, we can easily construct a counterexample by only optimizing with hard examples. In this paper, we propose to resolve such a contradiction from the perspective of anti-overfitting. First, we show that the sharpness at the mini-batch level approximately equals the sum of the gradient norm and the Kullback–Leibler (KL) divergence between the predictive distributions evaluated at the current point and its adversary. Second, we show that the \emph{gradient consistency}, which approximately quantifies the inner product between the mini-batch gradient and the per-example gradient, decreases consistently during training. This result explains why SAM is less likely to suffer from overfitting. At last, building on these insights, we further introduce CASAM, an algorithm that reweights each example according to gradient consistency to enhance the generalization performance. Extensive experiments on popular benchmarks such as ImageNet-1K and Clothing1M corroborate its efficacy.


Castle-in-the-Air: Probing the Foundational Visual Deficits of MLLMs via Bottom-Up Cognitive Factors

Jen-Tse Huang ⋅ Dasen Dai ⋅ Jen-Yuan Huang ⋅ Youliang Yuan ⋅ Xiaoyuan Liu ⋅ Wenxuan Wang ⋅ Wenxiang Jiao ⋅ Pinjia He ⋅ Zhaopeng Tu ⋅ Haodong Duan

Humans develop perception through a bottom-up hierarchy: from basic primitives and Gestalt principles to high-level semantics. Current Multimodal Large Language Models (MLLMs) are trained directly on complex downstream tasks. Does it guarantee their performance on these foundational visual capabilities? To systematically investigate this gap, we introduce VisFactor, a benchmark that digitizes 20 vision-centric subtests from FRCT, a well-established cognitive psychology assessment spanning four domains of human visual cognition. Furthermore, we design algorithms to automatically construct and validate unlimited test cases with controllable difficulty. Using VisFactor, we evaluate 39 frontier MLLMs, including both proprietary (e.g., GPT, Gemini) and open-source (e.g., LLaMA, Qwen) models. The best model achieves a score of only 55.9%. Analysis reveals good internal consistency (Cronbach's alpha = 0.94) and construct validity (compared to existing vision benchmarks). Models consistently fail on tasks such as mental rotation, spatial relation inference, and figure–ground discrimination, regardless of model size or prompting strategy. These findings suggest that performance improvements on existing general benchmarks might represent castles in the air instead of a genuine mastery of human-like visual cognition.


Catch Your Breath: Adaptive Computation for Self-Paced Sequence Production

Alexandre Galashov ⋅ Matt Jones ⋅ Nan Rosemary Ke ⋅ Yuan Cao ⋅ Vaishnavh Nagarajan ⋅ Michael Mozer

Within the landscape of inference-time scaling methods for foundation models, a width-based approach to scaling—which involves the insertion of \ tokens in the input stream to delay model responses—offers a unique advantage by increasing model expressivity while remaining highly parallelizable at both training and inference. The existing literature on training models to utilize \ tokens relies on the standard cross-entropy objective in which the model output is read out and evaluated only at the final step of a pause sequence. This approach provides no mechanism for the model to regulate its own processing or to signal readiness to respond, treating the additional compute steps as a static barrier rather than a resource to be used adaptively. We propose a supervised loss, Catch Your Breath (CYB), framed as a sequential-decision problem, that trains a model to dynamically and autonomously scale the number of compute steps used for each input token. The model indicates the need for additional compute steps by emitting a special \ output, delaying its response via a . The model can abstain multiple times to obtain longer delays. Our experiments demonstrate that CYB significantly outperforms standard cross-entropy when introduced either in pretraining or fine-tuning, reducing perplexity and enhancing downstream accuracy with no additional computational or memory cost.

Message-passing graph neural networks (MPGNNs) dominate modern graph learning. Typical efforts enhance MPGNN's expressive power by enriching the adjacency-based aggregation. In contrast, we introduce a graph-level aggregation over walk incidence-based matrices that are constructed to deliberately trade off some expressivity for stronger and more structured inductive bias. This approach allows for gradual scaling between classical message-passing and simpler methods based on walks. Our second result characterizes the expressive power at each scale using homomorphism counts over a hierarchy of generalized caterpillar graphs. Based on these foundations, we propose Caterpillar GNNs that substantially reduce the number of nodes in the hidden layers of the computational graphs on real-world datasets. This gradual reduction does not obstruct learning, enabling a systematic study of lower-order expressivity. We further describe benchmark settings where a task-aligned expressivity supports learning.


CATS: Acceptance-Oriented Critical Token Adaptive Selection for Multimodal Speculative Decoding

Kaiwen Liu ⋅ Yangkai Xie ⋅ Shuxia Lin ⋅ Liu Chonghan ⋅ Xu Yang ⋅ Yiguo Qiao

While Speculative Decoding (SD) has become an essential lossless acceleration technique for Large Language Models, its direct application to Multimodal Large Language Models (MLLMs) is hindered by the intricate visual dependencies present in cross-modal generation. Existing SD methods enforce a uniform global alignment between the draft and target models, which proves ineffective in multimodal settings due to a structured distribution bias: deviations of the draft model are systematically concentrated on a sparse subset of tokens that demand fine-grained visual comprehension. To overcome this limitation, we reformulate the training of multimodal draft models as an acceptance-oriented critical-token alignment problem and introduce CATS, a novel two-stage training framework. CATS employs two complementary selection mechanisms to pinpoint critical tokens:hard tokens that exhibit high rejection probabilities, and vision-critical tokens whose prediction fundamentally depends on visual semantics. Following an initial global alignment phase, CATS performs sparse, targeted refinement exclusively on the union of these identified token sets. Extensive experiments across three representative benchmarks using multiple LLaVA and Qwen2.5-VL target models demonstrate that our approach consistently improves the token acceptance rate over strong baselines such as EAGLE-2 and supervised fine-tuning. Ultimately, CATS substantially accelerates MLLM inference, achieving speedups of up to 3.05×.


CausalAffect: Causally Guided Learning of Psychology-Aligned Facial Affect Relations

Guanyu Hu ⋅ Tangzheng Lian ⋅ Dimitrios Kollias ⋅ Oya Celiktutan ⋅ Xinyu Yang

Facial affect understanding is not merely an image-to-label prediction problem: it requires identifying activated facial Action Units (AUs), modeling how AUs facilitate or inhibit one another, and explaining how AU configurations give rise to expression states. Such causal relations have long been studied in cognitive psychology and psychophysiology, where FACS, the Component Process Model, and motor synergy theories provide rich causal priors about facial behavior. In computer vision, however, data-driven facial affect models still largely learn correlational structures; even structured AU-aware methods often produce dependencies that are not well aligned with psychological priors. This leaves a persistent gap between biological facial mechanisms and learned visual representations, as observational facial images do not provide direct biological interventions and existing methods lack a principled way to recover psychology-aligned relations from data. We propose CausalAffect, a weakly supervised, causally guided framework for learning psychologically aligned facial affect relations from data. CausalAffect models two complementary relation types, AU$\rightarrow$AU and AU$\rightarrow$Expression, within a two-level hierarchy: a global graph capturing population-level psychology-aligned relations, and a sample-adaptive graph refining this backbone for individual variability. Reliable relation learning is supported by disentangled AU bottleneck representations, polarity-aware message passing, and feature-level counterfactual intervention. The learned relations recover canonical literature-supported pathways, reveal plausible inhibitory and sample-specific dependencies, are positively validated in a blind expert study, and exhibit intervention-consistent behavior under image-level AU edits. Experiments on six benchmarks show consistent improvements on both AU detection and facial expression recognition.

Automated systems built on artificial intelligence (AI) are increasingly deployed across high-stakes domains, raising critical concerns about fairness and the perpetuation of demographic disparities that exist in the world. In this context, causal inference provides a principled framework for reasoning about fairness, as it links observed disparities to underlying mechanisms and aligns naturally with human intuition and legal notions of discrimination. Prior work on causal fairness primarily focuses on the standard machine learning setting, where a decision-maker constructs a single predictive mechanism $f_{\widehat Y}$ for an outcome variable $Y$, while inheriting the causal mechanisms of all other covariates from the real world. The generative AI setting, however, is markedly more complex: generative models can sample from arbitrary conditionals over any set of variables, implicitly constructing their own beliefs about all causal mechanisms rather than learning a single predictive function. This fundamental difference requires new developments in causal fairness methodology. We formalize the problem of causal fairness in generative AI and unify it with the standard ML setting under a common theoretical framework. We then derive new causal decomposition results that enable granular quantification of fairness impacts along both (a) different causal pathways and (b) the replacement of real-world mechanisms by the generative model's mechanisms. We establish identification conditions and introduce efficient estimators for causal quantities of interest, and demonstrate the value of our methodology by analyzing race and gender bias in large language models across different datasets.


Causality can systematically address the monsters under the bench(marks)

Felix Leeb ⋅ Zhijing Jin ⋅ Bernhard Schölkopf

Effective and reliable evaluation is essential for advancing empirical machine learning. However, the increasing accessibility of generalist models and the progress towards ever more complex, high-level tasks make systematic evaluation more challenging. Benchmarks are plagued by various biases, artifacts, or leakage, while models may behave unreliably due to poorly explored failure modes. Haphazard treatments and inconsistent formulations of such ``monsters'' leads to a duplication of efforts, a lack of trust in results, and unsupported inferences. In this position paper, we argue causality offers an ideal framework to systematically address these challenges. By making causal assumptions in an approach explicit, we can faithfully model phenomena, formulate testable hypotheses with explanatory power, and leverage principled tools for analysis. To make causal model design more accessible, we identify several useful Common Abstract Topologies (CATs) in causal graphs which help gain insight into the reasoning abilities in large language models. Through a series of case studies, we demonstrate how the precise yet pragmatic language of causality clarifies the strengths and limitations of a method and inspires new approaches for systematic progress.


CausalTab: Pretraining Across Causal Environments for Tabular Causal Discovery

Zi-Rong Li ⋅ Si-Yang Liu ⋅ Tian-Zuo Wang ⋅ Han-Jia Ye

Causal discovery aims to recover directed causal relations from observational and interventional data, providing a basis for mechanistic understanding and reliable decision-making. Causal discovery foundation models (CDFMs) seek to amortize this problem by mapping a dataset directly to a causal graph in a single forward pass, avoiding per-dataset testing, search, or optimization. However, existing CDFMs remain limited, often failing to consistently match strong classical methods, and we find that a key bottleneck lies not only in model architecture but also in how causal pretraining tasks are constructed. Based on this observation, we propose CausalTab, a data-driven CDFM trained with broad causal pretraining over diverse graph priors, structural mechanisms, noise models, dimensions, sample sizes, and intervention regimes. A dynamic task construction strategy composes these causal environments into varied discovery tasks, enabling more transferable structural learning from observational and mixed-interventional data. On large-scale synthetic benchmarks, CausalTab achieves stronger and more robust performance than diverse causal discovery baselines. To further bridge abstract synthetic generators and realistic causal reasoning scenarios, we introduce an expert-knowledge-guided and LLM-audited semantic causal environment benchmark, where domain-grounded SCMs generate interpretable observational and interventional datasets for out-of-distribution analysis. Across both synthetic and semantic environments,CausalTab demonstrates robust structure recovery, especially under interventional evidence, highlighting broad causal pretraining as a key ingredient for transferable amortized causal discovery.


ChainFlow-VLA: Causal Flow Planning with Vision-Language Models

Xiyang Wang ⋅ Xinlin Wang ⋅ Tingguang Zhou ⋅ Gong Chen ⋅ Xingtai Gui ⋅ Zhi Xu ⋅ Hangning Zhou ⋅ Xiaolei Wu ⋅ Feiyang Tan ⋅ Mu Yang

Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) models capture interaction-aware temporal dependencies via causal factorization, but their step-wise decoding leads to error accumulation and suboptimal global structure. In contrast, diffusion models optimize trajectories globally but lack explicit causal constraints, making them unreliable in interactive and safety-critical scenarios. This dichotomy reveals a deeper issue: existing methods treat causal modeling and global optimization as separate paradigms, without a principled way to unify them within a single trajectory distribution. To address this, we propose ChainFlow-VLA, which unifies causal generation and global refinement within a unified probabilistic framework. We formulate planning as a mixture over AR-induced modes and learn VLM-conditioned residual distributions over these modes. An autoregressive generator (\textbf{Chain}) produces a discrete set of causal trajectory modes, followed by a diffusion-based refiner (\textbf{Flow}) that operates in residual space to perform mode-conditioned correction while preserving causal structure. A key insight is that vision-language models are more effective as semantic controllers for refinement rather than direct trajectory generators. By conditioning the diffusion process on VLM hidden states, scene-level reasoning guides fine-grained trajectory adjustments within each mode. This formulation enables robust planning in ambiguous and long-tail scenarios. ChainFlow-VLA achieves a state-of-the-art score of 94.8 on the NAVSIM v1 leaderboard, matching human-level performance (94.8).


Characterizing Learning in Deep Neural Networks using a Tractable Algorithmic Complexity Estimator

Pedram Bakhtiarifard ⋅ Sophia Natasha Wilson ⋅ Mahmoud H. A. Afifi ⋅ Jonathan Wenshøj ⋅ Raghavendra Selvan

Algorithmic complexity measures the intrinsic structure (or randomness) of strings. Kolmogorov-Chaitin-Solomonoff (KCS) complexity is the length of the shortest program that can output a string and halt over all possible programs; due to this reason it is uncomputable. Estimating the algorithmic probability of strings using simulations of finite Turing machines is one way of obtaining tight estimations of KCS complexity of strings. For two dimensional objects, this is currently doable only for binary strings using the block decomposition method (BDM). In this work, we present the quantized block decomposition (QuBD) method to extend the estimation of algorithmic complexity to any $k$-ary objects such as the weights of deep neural networks. We show theoretically that the proposed QuBD method yields better KCS complexity estimations than BDM that relies on binarization. We further study how algorithmic complexity interacts with learning in deep neural networks by tracking the evolution of weights during training. Using a variety of experiments we show that algorithmic complexity decreases as models learn, it correlates with generalization performance and can be used as a diagnostic measure to perform model compression.


Characterizing the Aesthetic Defaults of Generative Image Models

Maty Bohacek ⋅ Raina Panda ⋅ Daniel Fein ⋅ Arpita Singhal ⋅ Mark Fiore ⋅ Bill Freeman ⋅ Maneesh Agrawala

Contemporary text-to-image (T2I) models are widely observed to exhibit a recognizable, model-specific look even when prompted without any stylistic instruction (e.g., the so-called ``Midjourney look''). What such a default aesthetic actually is, however, has not been characterized: dominant T2I evaluation paradigms score realism, prompt alignment, or pairwise preference, but cannot say which concepts drive separation, how they shift across releases, or how they relate to real-image distributions. We address this with a concept-level decomposition of generator outputs along an interpretable, style-trained sparse vocabulary. We introduce LouvreSAE, a sparse autoencoder over CLIP image embeddings whose dictionary is shaped to fall on stylistic rather than object features, and LatentAesthetic, an evaluation methodology that uses it to estimate, ground, and compare generators' default aesthetics under content-controlled, style-neutral prompts. Applied to 26 generators spanning four years, our methodology surfaces stable per-generator profiles, a uniform pull toward lifestyle and editorial photography over art-historical imagery, and a steady cross-family contraction of the aesthetic envelope, supporting hypotheses of an algorithmic monoculture in image generation. Code, weights, and data are available at https://hf.co/datasets/neurips-subm/anonymous for the evaluation of future models.


Characterizing Universal Object Representations Across Vision Models

Florian P Mahner ⋅ Johannes Roth ⋅ Ka Chun Lam ⋅ Mick Bonner ⋅ Francisco Pereira ⋅ Martin N Hebart

Deep neural networks trained with different architectures, objectives, and datasets have been reported to converge on similar visual representations. However, what remains unknown is which visual properties models actually converge on and which factors may underlie this convergence. To address this, we decompose the object similarity structure of 162 diverse vision models into a small set of non-negative dimensions. To determine universal versus model-specific dimensions, we then estimate how often each dimension reappears across models. In contrast to model-specific dimensions, universal dimensions are more interpretable and more strongly driven by conceptual image properties, indicating the relevance of interpretability and semantic content as implicit factors driving universality across models. Differences in architecture, objective function, training data, model size, and model performance do not explain the emergence of universal dimensions. However, models with more universal dimensions also better predict macaque IT activity and human similarity judgments, suggesting that universality reflects representations relevant to biological vision. These findings have important implications for understanding the emergent representations underlying deep neural network models and their alignment with biological vision.


ChartArena: A Unified Benchmark with Atomic-Primitive Reasoning for Chart Parsing

Shangpin Peng ⋅ Gengluo Li ⋅ Xingyu Wan ⋅ Chengquan Zhang ⋅ Hao Feng ⋅ Binghong Wu ⋅ Huawen Shen ⋅ Weinong Wang ⋅ Ziyi Cai ⋅ Zhuotao Tian ⋅ Han Hu ⋅ Can Ma ⋅ Yu Zhou

Charts convey dense quantitative and relational information, yet general chart parsing remains challenging due to fragmented evaluation and limited structural reasoning. Existing benchmarks focus on narrow sub-tasks with inconsistent output formats (e.g., Markdown, CSV, JSON, SVG, code), hindering fair comparison and often overlooking real-world scenarios such as printed or hand-drawn photos. Meanwhile, end-to-end multimodal models tend to learn shallow pixel-to-text mappings without explicit structure understanding, making them fragile under visual variations. To address these issues, we present a unified framework that connects evaluation and modeling. On the evaluation side, we introduce ChartArena, a comprehensive bilingual benchmark covering eight chart families, including both numeric charts (bar, line, pie, radar, box plot, combination) and diagrammatic structures (flowchart, mind map). The dataset is constructed via a human-agent collaborative annotation pipeline, where model-assisted drafts are refined through multi-stage human verification to ensure structural consistency and annotation reliability. Furthermore, we design a format-agnostic evaluation protocol that maps different outputs into two canonical semantic spaces: a normalized triple view and a directed graph view, and evaluates them with structure-aware metrics. On the modeling side, we propose Chart Atomic Primitives (CAP), a lightweight semantic scaffold that introduces structural inductive bias during supervised fine-tuning. By guiding the model to infer structure before generating outputs, CAP leads to more consistent and robust parsing behavior. Extensive experiments show that models trained with CAP achieve reliable performance on both rendered charts and real-world images. Code, benchmark, and models will be made publicly available.

Estimating relationships between 3D shapes requires both dense endpoint correspondence and a deformation path connecting the source to the target. These tasks are coupled: the correspondence constrains the interpolation endpoint, while a plausible trajectory can help disambiguate the map. However, when only endpoint shapes are observed, the interior trajectory is under-constrained. Existing joint frameworks typically model interpolation through sampled per-vertex displacements regularized by local rigidity or temporal smoothness, which can indirectly restrict the stretch and shear needed for non-isometric or poorly aligned shape pairs. We propose Chebyshev Differential Flows (Chef), an unsupervised framework that models motion as a continuous time-varying field of per-face Jacobians rather than vertex trajectories. Each Jacobian trajectory is represented by a low-order shifted Chebyshev expansion, yielding a compact arbitrary-time deformation model whose odd and even modes separate endpoint-visible deformation from endpoint-invisible interior control. A spectral functional-map branch estimates the endpoint correspondence, while a differentiable Poisson solver integrates the predicted Jacobian field into globally consistent intermediate shapes. We further introduce a trend regularizer that guides intermediate stretch and shear without imposing per-snapshot As-Rigid-As-Possible (ARAP) rigidity. Experiments on near-isometric and non-isometric benchmarks show competitive correspondence accuracy, improved interpolation proxy metrics, and better stability when the standard rigid pre-alignment step is omitted. Code will be publicly released for research.


ChunkFT: Byte-Streamed Optimization for Memory-Efficient Full Fine-Tuning

Yongkang Liu ⋅ Zijing Wang ⋅ Mengjie Zhao ⋅ Ercong Nie ⋅ Mingyang Wang ⋅ Qian Li ⋅ Feiliang Ren ⋅ Shi Feng ⋅ Daling Wang ⋅ Hinrich Schuetze

This work presents \textsc{ChunkFT}, a memory-efficient fine-tuning framework that reformulates full-parameter fine-tuning around a dynamically activated working set. \textsc{ChunkFT} enables gradient computation for arbitrary sub-tensors without modifying the network architecture, providing an algorithmic foundation for optimizing arbitrary sub-networks while avoiding standard dense gradient computation. We provide a theoretical convergence analysis of \textsc{ChunkFT} in the deterministic setting. Empirically, we apply \textsc{ChunkFT} to fine-tune Llama 3-8B and Llama 3-70B using a single RTX 4090-24GB GPU and 2$\times$ H800-80GB GPUs, respectively. Full-parameter fine-tuning of a 7B model with a 1K input length requires only 13.72GB of GPU memory. The results demonstrate the effectiveness of \textsc{ChunkFT} in memory usage, running time, and optimization quality. Moreover, downstream evaluations on language understanding, mathematical reasoning, and MT-Bench show that \textsc{ChunkFT} consistently outperforms existing memory-efficient baselines. Notably, \textsc{ChunkFT} achieves performance comparable to, and in some cases exceeding, full-parameter fine-tuning. Our repository is on https://anonymous.4open.science/r/chunk-B48E.


CI4A: Semantic Component Interfaces for Agents Empowering Web Automation

Zhi Qiu ⋅ Jiazheng Sun ⋅ Chenxiao Xia ⋅ Jun Zheng ⋅ Xin Peng

While Large Language Models demonstrate remarkable proficiency in high-level semantic planning, they remain limited in handling fine-grained, low-level web component manipulations. To address this limitation, extensive research has focused on enhancing model grounding capabilities through techniques such as Reinforcement Learning. However, rather than compelling agents to adapt to human-centric interfaces, we propose constructing interaction interfaces specifically optimized for agents. This paper introduces Component Interface for Agent (CI4A), a semantic encapsulation mechanism that abstracts the complex interaction logic of UI components into a set of unified tool primitives accessible to agents. We implemented CI4A within Ant Design, an industrial-grade front-end framework, covering 23 categories of commonly used UI components. Furthermore, we developed a hybrid agent featuring an action space that dynamically updates according to the page state, enabling flexible invocation of available CI4A tools. Leveraging the CI4A-integrated Ant Design, we refactored and upgraded the WebArena benchmark to evaluate existing SoTA methods. Experimental results demonstrate that the CI4A-based agent significantly outperforms existing approaches, achieving a new SoTA task success rate of 86.3\%, alongside substantial improvements in execution efficiency.


CIG: Exploration via Conditional Information Gain

Tim Joseph ⋅ Marcus Fechner ⋅ Philipp Stegmaier ⋅ Karam Daaboul ⋅ Marius Zöllner

Intrinsic rewards for exploration in reinforcement learning condition on different contexts: lifelong rewards score each transition against accumulated experience but ignore within-rollout redundancy; episodic rewards penalize intra-trajectory repetition but discard lifetime progress. Hybrid methods combine both signals through heuristic weights or require Gaussian-process dynamics that do not scale beyond low-dimensional state spaces. Trajectory-level information gain decomposes into per-step terms that condition on the replay buffer and rollout prefix simultaneously, but remains intractable for deep models. We derive the Conditional Information Gain (CIG) reward as a tractable surrogate: a log-determinant objective over an ensemble disagreement kernel whose Cholesky factorization yields causal per-step rewards that retain both conditioning sets while scaling to high-dimensional state spaces. We instantiate CIG in a model-based setting, where rollouts are short and within-rollout corrections remain largely unexplored. Across twelve tasks spanning discrete (MiniGrid) and continuous control (OGBench), in both clean and stochastic-distractor settings, CIG outperforms or matches prior exploration methods while remaining robust to stochastic distractors.


Classification Fields: Arbitrarily Fine Recursive Hierarchical Clustering From Few Examples

Yicen Li ⋅ Ruiyang Hong ⋅ Anastasis Kratsios ⋅ Haitz Sáez de Ocáriz Borde ⋅ Paul D. McNicholas

Classical clustering methods usually return either a finite partition of the observed data or a finite dendrogram over it. This finite-sample view is inadequate when the hierarchy of interest is a recursive geometric object with fine-scale refinements that continue beyond the levels directly observed. We introduce classification fields: infinite-depth hierarchical cluster structures on $\mathbb{R}^d$ generated by a local parent-to-child refinement rule. A classification field generator maps each parent centre to an ordered, bounded, and separated tuple of child residuals. Together with a root and a scale factor, this rule recursively generates cluster centres, Voronoi cells, and a metric DAG encoding the hierarchy. Given only a finite prefix of such a hierarchy, we learn a classification field predictor that approximates the generator and can be rolled out to unseen depths. We prove exponential truncation convergence in the completed cell metric and ReLU realizability with width $(O(\varepsilon^{-\gamma})$ and depth $\widetilde O(\varepsilon^{-3\gamma/2})$, where $\gamma=\log K/(-\log s)$, up to finite-window aspect-ratio factors. The approximation holds at the level of the induced compact metric structures, measured in the completed cell-metric Hausdorff distance. Experimental validation on matched CFG-generated hierarchies, IFS fractals, and image-induced recursive clustering hierarchies shows that learned predictors preserve ordered child slots, unordered geometry, and hierarchy-level path metrics under recursive rollout. These results support the claim that finite hierarchical observations can reveal local refinement rules capable of generating substantially deeper classification fields.

We study the classification problem for high-dimensional data with $n$ observations on $p$ features where the $p \times p$ covariance matrix $\Sigma$ exhibits a spiked eigenvalue structure and the vector $\zeta$, given by the difference between the {\em whitened} mean vectors, is sparse. We analyze an adaptive classifier (adaptive with respect to the sparsity $s$) that first performs dimension reduction on the feature vectors prior to classification in the dimensionally reduced space, i.e., the classifier whitens the data, then screens the features by keeping only those corresponding to the $s$ largest coordinates of $\zeta$ and finally applies Fisher linear discriminant on the selected features. Leveraging recent results on entrywise matrix perturbation bounds for covariance matrices, we show that the resulting classifier is Bayes optimal whenever $n \rightarrow \infty$ and $s \sqrt{n^{-1} \ln p} \rightarrow 0$. Notably, our theory also guarantees Bayes optimality for the corresponding quadratic discriminant analysis (QDA). Experimental results on real and synthetic data further indicate that the proposed approach is competitive with state-of-the-art methods while operating on a substantially lower-dimensional representation.


ClawBenchPro: Benchmarking How Well Agent Harnesses Work

Yufei Liu ⋅ Yirong Zeng ⋅ Yuxiang He ⋅ Yutai Hou ⋅ Yuxian Wang ⋅ Qunyao Du ⋅ WangXu ⋅ Xiao Ding ⋅ Wu Ning ⋅ Hao Cong ⋅ Dandan Tu ⋅ Qixun Zhang ⋅ Ting Liu

The evolution of Large Language Models (LLMs) from conversational assistants to autonomous agents is increasingly driven by agent harnesses (e.g., OpenClaw), which provide the operational substrate for executing complex, multi-step workflows in real-world environments. However, existing claw-supported benchmarks often suffer from narrow domain coverage and reliance on single-harness evaluations, leading to evaluation distortion and limited diagnostic diversity. To bridge this gap, we present ClawBenchPro, a comprehensive benchmark comprising 1.0k samples across 43 distinct domains. By integrating skill augmentation, environmental heterogeneity, and multi-session dependencies, we replicate complexity workflow at scale. Extensive experiments across 14 frontier LLMs and four representative harnesses, show model's average scores clustering at 60-74\%. We identify a critical domain disparity: models excel in cyber-native tasks but struggle with real-world operational workflows. Moreover, a pervasive reliability bottleneck emerges, where incremental progress rarely translates to zero-defect execution. Comparative results show that while OpenClaw incurs higher computational costs, NanoClaw offers superior cost-efficiency (despite a 3.2–6.5\% accuracy trade-off), and Hermes-Agent emerges as the most stable framework for claw-style interactions. ClawBenchPro establishes a rigorous, reproducible standard for diagnosing agent limitations and advancing next-generation harness architectures.


ClothTransformer: Unified Latent-Space Transformers for Scalable Cloth Simulation

Yu Zhang ⋅ YIDI SHAO ⋅ Wenqi Ouyang ⋅ Yushi LAN ⋅ Zhexin Liang ⋅ Chengrui Wu ⋅ Xudong XU ⋅ Xingang Pan

Unified and scalable Transformers have recently achieved remarkable success in modeling diverse phenomena traditionally associated with computer graphics, such as 3D visual effects, rendering processes, and motion in videos. In this work, we take a step further by investigating whether modern Transformer techniques can tackle the challenging task of cloth simulation. To this end, we present ClothTransformer, a framework that reformulates cloth simulation as autoregressive sequence modeling in a learned latent space. Existing neural cloth simulators are largely specialized to single scenarios, intrinsically coupled to the mesh discretization, and lack robust collision handling. Our approach addresses these limitations through three contributions: (1) a unified Transformer architecture that handles diverse scenarios---body-driven garments, robotic manipulation, and free-fall collisions---under a single model and achieves approximately $4$--$9{\times}$ lower error than prior state-of-the-art across all scenarios; (2) a scalable latent-space formulation that compresses arbitrary-resolution meshes into a fixed-size set of latent tokens, making temporal dynamics computation independent of mesh resolution; and (3) a diverse-scenario high-fidelity penetration-free dataset of ${\sim}$493.4k frames spanning all three settings, which enables a differentiable Continuous Collision Detection (CCD) module to suppress penetration artifacts.


CME–SpectrumBench: Can LLMs Analyze Condensed Matter Spectral Data?

Jin Gene Wong ⋅ Anjney Midha ⋅ Joseph Tennyson ⋅ Wei-Lin Chiang ⋅ Ion Stoica ⋅ Zhi-Xun Shen

Large Language Models (LLMs) are presently well assessed by expert–level mathematics and physics benchmarks which focus on theoretical, symbolic reasoning. However, we lack experimental benchmarks that evaluate data analysis in frontier research problems, a capability that future domain–specific foundation models must have. To this end, we present CME–SpectrumBench, an expert–designed benchmark consisting of 190 questions that measures how reliably frontier LLMs analyze research–level spectral scientific data in condensed matter experiments (CME). Using simulated CME spectral data, we prompt models to identify structures from 2D intensity arrays in the face of noise and limited instrumental resolution, a test of dense context reasoning. Models are scored by how accurately their numerical answers approach the ground truth. Under direct query, scores generally pale in comparison to handwritten code due to clear failure modes: models can find simple features but struggle to trace patterns or compute derived quantities, and performance degrades with increased data resolution and noise/convolution levels. LLM performance becomes comparable to handwritten code by allowing tool use (access to a Python interpreter) with/without ReAct prompting, in effect mapping questions onto agentic code generation problems. Our findings show that frontier models can serve as research assistants by interfacing with data through code, but still face difficulties when processing information directly from large datasets — a key barrier future foundation models must overcome.


CM-EVS: Sparse Panoramic RGB-D-Pose Data for Complete Scene Coverage

Liu Jiale ⋅ Jungang Li ⋅ Jieming Yu ⋅ Yu Xinglin ⋅ Zihao Dongfang ⋅ Shi Keyu ⋅ Zongjian Ding ⋅ 家焕 张 ⋅ Shunwen Bai ⋅ Haoran Huang ⋅ Yurun Wang ⋅ Yanxi Wu ⋅ Ningzhe Yu ⋅ Yudong Gao ⋅ Mingjun Cheng

Modern 3D visual learning relies on observations sampled from metric 3D assets, yet existing scans, meshes, point clouds, simulations, and reconstructions do not directly provide a sparse, comparable, and geometry-consistent panoramic training interface. Dense trajectories duplicate nearby views, source-specific rendering policies yield heterogeneous annotations, and sparse heuristics may miss important regions or introduce depth-inconsistent observations. We study how to convert 3D assets into sparse panoramic RGB-D-pose data that preserves complete scene coverage with low redundancy and auditable provenance. We propose COVER (Coverage-Oriented Viewpoint curation with ERP Range-depth warping), a training-free ERP viewpoint curator that projects geometry observed from selected views into candidate ERP probes, scores incremental coverage, and penalizes depth conflicts. Under bounded proxy error, its greedy coverage proxy preserves the standard coverage-style approximation behavior up to an additive error term. Using COVER, we build CM-EVS (Coverage-curated Metric ERP View Set), a panoramic RGB-D-pose dataset with 36,373 curated ERP frames from 1,275 indoor scenes across Blender indoor, HM3D, and ScanNet++, complemented by outdoor panoramas from TartanGround and OB3D re-encoded into the same schema. Each frame provides full-sphere RGB, metric range depth along ERP rays, calibrated world-to-camera pose, and provenance logs. With a median of only 25 frames per indoor scene, CM-EVS covers all 13 unified room types while maintaining compact scene-level coverage. Experiments show that COVER improves the coverage–conflict trade-off, making CM-EVS a sparse, compact, and auditable RGB-D-pose resource for geometry-consistent panoramic 3D learning.


Coarse-to-Fine Compositional Diffusion for Long-Horizon Planning

Byoungwoo Park ⋅ Utkarsh Mishra ⋅ Jaemoo Choi ⋅ Juho Lee ⋅ Yongxin Chen

Diffusion models provide strong priors for generating structured data, but many tasks require outputs beyond the scale on which these models are typically trained. Compositional generation addresses this by composing overlapping local plans from a pretrained short-horizon prior into a long-horizon output. However, standard composition primarily enforces agreement between neighboring local plans, yielding local consistency without directly specifying the global structure of the full composition. As a result, locally compatible plans may still form an implausible route, task sequence, or temporal evolution. Existing methods improve global coherence by repeatedly propagating local consistency signals or by adding inference-time optimization, but these procedures become expensive as the number or dimensionality of local plans increases. We propose Coarse-to-Fine Compositional Diffusion (CoFi), an inference-time sampler that separates global structure formation from local detail refinement. CoFi first aligns local denoised estimates around a shared coarse structure, producing a global scaffold that captures the long-range task-level arrangement. It then diffuses this scaffold to an intermediate noise level and denoises it with the same pretrained local prior, restoring local fine structure while preserving the scaffold-induced global coherence. Across long-horizon robotic planning, panoramic image generation, and long video generation, CoFi not only improves both global coherence and local sample quality over prior compositional baselines, but also requires $2$--$8\times$ fewer denoiser evaluations.


CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models

Xiao Zhu ⋅ Xinyu Zhou ⋅ Boyu Zhu ⋅ Hanxu Hu ⋅ Mingzhe Du ⋅ Haotian Zhang ⋅ Huiming Wang ⋅ Zhijiang Guo

Reinforcement Learning from Verifiable Rewards (RLVR) has driven recent progress in code large language models by leveraging execution-based feedback from unit tests, but its scalability is fundamentally constrained by the availability and reliability of high-quality test cases. We propose CodeScaler, a reward model designed to scale both reinforcement learning training and test-time inference for code generation. CodeScaler is trained on carefully curated preference data derived from verified code problems and incorporates syntax-aware code extraction and validity-preserving reward shaping to ensure stable and robust optimization. Across four coding benchmarks, CodeScaler consistently outperforms execution-based RL by +1.55 points on Qwen3-8B-Base and +4.23 points on Qwen3-14B-Base. By further scaling to 44K problems with additional synthetic data, CodeScaler yields +14.64 points improvement over the base model without requiring any test cases. At inference time, CodeScaler serves as an effective test-time scaling method, achieving performance comparable to unit test approaches while providing a 10× reduction in latency. Moreover, CodeScaler surpasses existing reward models on RM-Bench not only in the code domain (+3.3 points), but also in general and reasoning domains (+2.7 points on average).

Vision-Language-Action (VLA) models have driven significant progress in robotic manipulation, yet they fundamentally struggle with the ``vision-override'' phenomenon. Driven by the severe modality imbalance between dense visual streams and sparse linguistic instructions, VLAs frequently fall prey to causal confusion. Instead of treating language as the primary causal driver, the policy entirely bypasses the original instruction by overfitting to spurious visual confounders, such as prominent objects or familiar layouts. To systematically alleviate this bias, we formalize the process of action generation as a Dual-path Deconfounding Graph (DDG) and propose CofactVLA, a novel causal intervention framework. By dynamically constructing a language-masked counterfactual branch within a single forward pass, CofactVLA isolates and neutralizes visual confounders through two synergistic mechanisms. First, Action-Level Orthogonal Projection Guidance (OPG) geometrically projects the factual velocity field away from the counterfactual visual bias during continuous flow matching, extracting the pure semantic intent. Second, Feature-Level Counterfactual Covariance Reduction (CCR) mathematically deconfounds latent representations by penalizing the positive eigenspace of the covariance difference, explicitly suppressing dominant visual shortcuts while preserving the causal language intent. Extensive experiments demonstrate that CofactVLA establishes a new state-of-the-art across diverse simulation benchmarks. Beyond simulation, real-world robot experiments demonstrate the causal efficacy of our method in bridging the generalization gap, yielding a 52.3\% absolute success rate gain under out-of-distribution scenarios.


CoilStellaration: A Dataset and Benchmark for Engineering-Aware Stellarator Coilset Generation

Jan-Hendrik Ewers ⋅ Santiago Cadena ⋅ Nicolò Foppiani ⋅ Andrea Merlo

Stellarators are fusion energy devices that confine hot plasma using a carefully shaped magnetic field, and they are one of the most promising candidates in the race towards commercial fusion. Among the many systems that make up a stellarator, the electromagnetic coils generate the confining magnetic field. Their design is the key factor in deciding whether a stellarator can be practically engineered and cost‑effectively realized. Coilset design is an intrinsically complex problem: it does not admit either a closed or a unique solution, making it suitable for optimization. Coilset optimization is therefore at the heart of stellarator design: a coilset must accurately reproduce a target magnetic field while satisfying stringent engineering constraints on coil spacing, curvature, and length. Although recent datasets have accelerated progress on stellarator equilibria optimization, data-driven coilset design is limited by the lack of standardized datasets that pair target plasma configurations with engineering-aware coilsets. Here, we release a dataset (https://huggingface.co/datasets/proxima-fusion/coilstellaration) of 177,000 optimized stellarator coilsets targeting plasma configurations sampled from the publicly available ConStellaration dataset. For each configuration, we optimized coilsets to reproduce the magnetic field while satisfying a diverse set of engineering constraints. Alongside the dataset, we introduce a benchmark for conditional coilset generation: given a target plasma boundary and a set of engineering requirements, a model should produce a coilset that satisfies the engineering constraints while reproducing the target field. We provide reference code, evaluation scripts and baselines models trained on a subset of the provided data (https://github.com/proximafusion/coilstellaration). Beyond enabling better optimization initializations and one-shot coil design, our dataset provides a testbed for data-driven constrained optimization and for studying which plasma geometry features predict simpler, more engineerable coils. By releasing data, and code for benchmarks and baselines, we aim to lower the barrier for machine learning and optimization researchers to engage with stellarator coilset design and accelerate progress toward engineerable and cost-effective stellarators power plants.


Collapse Hunter: Tackling the Dimensional Degeneration in Generative Ranking

Haoran Xin ⋅ Junwei Pan ⋅ Yongqi Zhou ⋅ Tianqu Zhuang ⋅ YilongSun ⋅ Zhixiang Feng ⋅ Shoujun Liu ⋅ Gong Chen ⋅ Shudong Huang ⋅ Haijie Gu ⋅ Hui Xiong

Generative ranking models (GRMs) have emerged as a scalable alternative to traditional recommendation pipelines by unifying items and user actions into a single token sequence and formulating recommendation as next-token prediction. Despite their promise, we identify a geometric failure mode: attention interactions with lowcardinality action tokens can induce dimensional collapse in high-rank item representations, as suggested by interaction-collapse theory. To pinpoint the origin of collapse, we view attention mechanism over item–action interleaving sequences as a mixture of directional interaction channels between token types, and find that collapse concentrates in the action-attending-item channel, where low-rank action gating drags the item value stream toward a low-dimensional subspace. To address this, we theoretically show that nonlinearity mitigates collapse, and propose Selective VAlue Nonlinearity (SVAN), a simple yet effective approach that applies nonlinear activation to the value component in the collapsed channel, protecting item representations while preserving the remaining attention geometry. SVAN consistently improves the ranking quality over strong baselines. Further analyses confirm that SVAN alleviates dimensional collapse, yielding up to a 12.58% effective-rank gain, and enhances representation discriminability.


Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

Mason Nakamura ⋅ Abhinav Kumar ⋅ Saswat Das ⋅ Sahar Abdelnabi ⋅ Saaduddin Mahmud ⋅ Ferdinando Fioretto ⋅ Shlomo Zilberstein ⋅ Eugene Bagdasarian

Multi-agent systems, where LLM agents communicate through free-form language, enable sophisticated coordination for solving complex cooperative tasks. This surfaces a unique safety problem when a group of agents forms a coalition and colludes to pursue secondary goals and degrade the joint objective. In this paper, we present Colosseum, a framework for auditing LLM agents' collusive behavior in multi-agent settings. We ground how agents cooperate through a formal multi-agent decision-making framework and measure action-based collusive behavior in actions via regret relative to the cooperative optimum and compare it with communication-based collusive behavior. Colosseum enables audits of LLM agents for collusion under benign settings, different coalition objectives, persuasion tactics, and network topologies. We then introduce a new behavioral probe by creating secret communication channels between agents, showing that most out-of-the-box models exhibit a propensity to collude under this probe, which we term emergent collusion. Furthermore, we discover ``collusion on paper'' when agents plan to collude in text but often pick non-collusive actions. Colosseum provides a new way to audit collusion in cooperative multi-agent systems while presenting observations about how collusion emerges, what affects collusion efficacy, and which strategies may mitigate it.

Swapping two LLM post-training stages can change the final model, but an aggregate loss or benchmark delta only says that something moved, not where the order-dependent residue landed. We ask whether this residue leaves a memory trace: a structured signal that is localized in output space, changes held-out order-gap NLL under targeted interventions, and is assignable from paired endpoint weights. For a specified two-stage path $(A,B)$, we define commutator memory by projecting the Lie bracket $b_{AB} := H_B g_A - H_A g_B$ at the base model through the logits into token scores $\tau_k=\mathbb{E}[e_k\,\delta z_k]$, whose sum is the leading bracket-predicted order effect. The token readout is localized: endpoint and batch-resampled checks preserve bracket top-token supports far more than norm-matched random directions (82--99% vs. 35--49% top-20 overlap), and a separate endpoint-support prediction check recovers 53--57% of Qwen top-1% active-token supports. It has intervention leverage in the tested protocols: top-positive-$\tau_k$ token interventions close 32% (median) of the held-out Qwen-3-4B SFT ordering gap with tau-neutral controls near zero, with the same bracket-vs-control separation on a pre-specified above-floor Qwen-2.5-1.5B fp32 subset. It is assignable from paired endpoints: given a base model, candidate stages, and paired alternate-order endpoints, the statistic $\langle\theta_{AB}-\theta_{BA},\,b_{AB}\rangle$ distinguishes $k=1$ SGD training order in 66/72 pair-seed units across four LLMs. Matched-batch DPO and frozen-rollout reward-surrogate experiments are matched-path stress tests; AdamW is reported as an analogous lifted-state pilot. Commutator memory turns training order from a scalar nuisance into a localized, intervention-tested, and paired-endpoint-readable residue.

Understanding when neural networks learn interpretable representations is central to mechanistic interpretability. We study this through identifiability: when can supervised learning recover the true latent variables underlying the data? We introduce two data generating processes distinguished by their causal direction where either labels cause latents (generative) or latents cause labels (discriminative). When the representation has at least as many dimensions as the latent space, we prove that the optimal encoders are linear in the latents. Conversely, when dimensions are insufficient, identifiability generically fails and encoders become nonlinear. We then show that additional constraints yield recovery of individual latent variables, not just linear mixtures. In the generative setting, non-negative activations and energy efficiency force each neuron to encode exactly one latent, providing a precise account of selective grandmother cell-like responses. In the discriminative setting, when latents are sparse and the true measurement matrix satisfies the Null Space Property, the learned encoder inherits compressed sensing structure, enabling sparse dictionary learning to perfectly recover individual latents. This constitutes, to our knowledge, the first identifiability result for undercomplete nonlinear independent component analysis. Our framework unifies nonlinear ICA, compressed sensing, and neural network interpretability, delineating exactly when representations transition from nonlinear distortions to linear representations and, under additional constraints or steps, to clean, axis-aligned codes. Theoretical predictions are validated on synthetic benchmarks and on activations drawn from vision models, large language models, and macaque inferotemporal cortex.


Complexity-guided Regularization for Generalizable Human Gaussian Splatting

Gun Ryu ⋅ Gi-Mun Um ⋅ Won-Sik Cheong ⋅ Wonjun Kim

Recent studies on generalizable human Gaussian splatting have achieved the notable progress in synthesizing novel views of unseen human subjects from multi-view images. Despite this progress, existing methods still struggle to represent fine-grained details of human subjects, which inherently involve high geometric and textural complexity. This is because previous methods do not impose a constraint on how densely Gaussians are placed, which often leads to blurry rendering results in complex regions. To address this problem, we propose a complexity-guided regularization scheme. The key idea is to leverage a complexity prior computed via the local entropy in geometric and textural characteristics of a clothed human. This effectively prevents Gaussians from being coarsely placed in complex regions, which facilitates the faithful reconstruction of intricate structures. Furthermore, we sample local patches predominantly from complex regions according to this complexity prior, which allows the model to be sufficiently supervised by local details. Experimental results on benchmark datasets demonstrate that our method delivers state-of-the-art performance while maintaining the real-time inference speed in generalizable human Gaussian splatting.


Complexity of Differentially Private Selection via Federated APIs

Hilal Asi ⋅ Vitaly Feldman ⋅ Jelani Nelson ⋅ Huy Nguyen ⋅ Kunal Talwar ⋅ Samson Zhou

Differential privacy (DP) has become the standard for privacy-preserving data analysis, yet its practical implementation in federated or distributed settings often faces a significant utility gap compared to the central model. While existing frameworks, including local, shuffle, and secure aggregation models, offer vetted primitives, they frequently suffer from high error rates for some important functionalities such as selection. Our first contribution is to show that this is inherent in the aggregation model: any private selection algorithm relying on the standard noisy aggregation primitive requires $\Omega(\sqrt{d}/\varepsilon)$ samples. To address this limitation, we propose adding a new primitive that implements a version of the \textit{Sparse Vector Technique} (SVT) and argue that it has simple implementations under various trust models. We demonstrate its power by developing an $\varepsilon$-differentially private algorithm for the selection problem. For selection in dimension $d$, our algorithm achieves an additive error of $O\left(\frac{\log d}{\varepsilon}\right)$ matching that of the central model. Further, we show this can be done using a near-optimal number of SVT queries. Together with existing lower bounds for the shuffle model, this result establishes an exponential separation between the SVT model and shuffling or aggregation-based models. We demonstrate the model's broad applicability to diverse tasks such as $F_p$ loss minimization, heavy-hitter identification, and sparse mean estimation.


Complex Optimization Modeling via Multiagent Fine-Tuning and Skill-Augmented Reasoning

Yu Ding ⋅ Jiajun Li ⋅ Ran Hou ⋅ Jiahui Duan ⋅ Guanyu Nie ⋅ Xiongwei Han ⋅ Tao Zhong ⋅ Vincent Chau ⋅ Weiwei Wu ⋅ Wanyuan Wang

Optimization modeling (OM) is fundamental to address Operations Research (OR) problems, yet it requires specialized domain expertise. Large language models (LLMs) are promising for automating OM from natural language descriptions. Through prompting and fine-tuning, LLMs have made significant progress on easy problems, but they still struggle with complex problems involving combinatorial constraint logic and obscure decision variables. To provide a precise formulation for industry OR problems, this paper proposes a multiagent system (MAS) workflow, which decomposes the entire OM process into six structured sub-tasks: set identification, parameter abstraction, variable definition, objective formulation, constraint derivation, and solver-code generation. Existing OM benchmarks only provide a sparse judgment of the final formulation, without process rewards for intermediate agents. Therefore, the main challenge of MAS-based OM is how to fine-tune intermediate agents with OM domain knowledge. This paper proposes a sample efficient counterfactual-based credit assignment for the MAS trajectory. The augmented intermediate credit can be directly used to fine-tune each agent. Moreover, we construct an Optimization Skill Library (Opt-Skill) that provides agent-specific procedural guidance for skill-augmented reasoning, self-checking, and error correction during inference. Experimental results show that, compared with existing methods and frontier LLMs, the proposed MAS workflow achieves state-of-the-art average accuracy while maintaining strong performance on both easy and complex optimization modeling tasks. Altogether, we believe that MAS represents a concrete step forward in performing complex OM in industry scenarios.


Complex Schrödinger Bridges

Dong-Sig Han ⋅ Tolga Birdal

Diffusion bridges have emerged as versatile diffusion-based frameworks for facilitating transformations between arbitrary data distributions. However, these approaches are hindered by the immensity of the optimization space without consideration of underlying high-order structures. This paper explores a complex-valued variant of diffusion bridges for solving physical problems using conformal field theory and mean field games. To this end, we propose a new class of dynamical diffusion bridges, called complex Schr\"odinger Bridges (CSB), whose derived algorithms are able to handle diffusion-based tasks using Hamiltonian-based geometric prior. Additionally, we propose image-to-image CSB (I$^2$CSB), tractable diffusion bridge techniques with a spiral analytic posterior. We verify our theoretical claims and practicality of our Hamiltonian based approach with a wide range of experiments.

Deep networks embed inputs in high-dimensional representations, but how much of this capacity actually drives behavior? We introduce a framework for identifying, at each layer, the subspace of the representation that is functionally relevant for the task. Using a matrix analogous to the Fisher information but defined over the hidden representation, we rank directions by their influence on downstream computation and project onto the most task-relevant subspace. We benchmark this behaviorally grounded compression against three geometric alternatives that preserve variance, local neighborhood structure, or global topology. Across pretrained vision models, representations are remarkably compressible: a small fraction of the task-relevant subspace preserves most performance, often requiring far fewer dimensions than geometric alternatives. We show that this decomposition recovers functionally distinct subspaces within the same representation: on cue-conflict stimuli, shape-relevant and texture-relevant directions are largely non-overlapping, and ablating one selectively impairs the corresponding task while leaving the other intact. Task-relevant subspaces also provide a principled foundation for continual learning: constraining gradient updates to lie orthogonal to the task-relevant subspaces of prior tasks outperforms standard variance-based gradient projection methods. Finally, restricting representational comparisons to task-relevant subspaces reveals that inter-model alignment is higher in task-relevant directions than in task-irrelevant directions at later layers, and that conditioning on different tasks over the same stimuli uncovers asymmetries that full-space similarity measures cannot detect. Together, these results establish that deep networks compute through a low-dimensional functional core embedded within a much larger representational space, that these cores differ across tasks, and that they provide a unifying basis for compression, continual learning, and representational comparison.


Computing All Optimal Partial $p$-Wasserstein Matchings on the Line

Sebastian Angrick ⋅ Jacobus Conradi ⋅ Mónika Csikós ⋅ Niko Mertens ⋅ Danny Mittal ⋅ André Nusser ⋅ Krzysztof Onak ⋅ Sharath Raghvendra

For $p \ge 1$, the $p$-Wasserstein distance measures the minimum cost of transporting probability mass, where moving mass between two points costs the $p$th power of their distance. For discrete distributions in one dimension, full transport is especially simple: after sorting, mass is matched in order along the line. By contrast, partial and unbalanced transport on the line remains much less understood. Recently, Chapel and Tavenard [ICLR'25] showed that, for $p=1$, all optimal partial transport plans between distributions supported on $n$ points, with uniform mass at each point, can be computed in $O(n\log n)$ time by exploiting the metric structure of the cost. For $p>1$, this structure no longer applies, and existing approaches require $\Omega(n^2)$ time. Our main contribution is an FFT-based data structure for balanced-interval transport queries, which bypasses this quadratic bottleneck and yields an $O(p\ n\log^2 n)$-time algorithm for computing all optimal partial transports on the line for every finite $p\ge 1$. We also provide an open-source C++ implementation that outperforms the state-of-the-art baseline on a range of synthetic instances. Finally, we establish a conditional lower bound for $p=\infty$: any subquadratic-time algorithm for computing all optimal partial transport plan costs on the line would violate the $(\min,+)$-Convolution Hypothesis. This separates the problem from full optimal transport, which is solvable in $O(n\log n)$ time.


Conditioning Gaussian Processes on Almost Anything

Henry Moss ⋅ Lachlan Astfalck ⋅ Tom Cowperthwaite ⋅ Colin Doumont ⋅ Samuel Willis ⋅ Christopher Nemeth ⋅ Philipp Hennig ⋅ Andrew Zammit-Mangion

Gaussian processes (GPs) offer a principled probabilistic model over functions, but exact inference is restricted to the linear-Gaussian regime. We establish an explicit equivalence between GPs and a class of linear diffusion models, recasting predictive sampling as an ODE with closed-form Gaussian dynamics and a likelihood-dependent guidance term that admits a simple Monte Carlo approximation. In the linear-Gaussian setting, we recover standard GP conditioning exactly; beyond conjugacy, the same machinery handles any conditioning statement admitting point-wise likelihood evaluation — including non-linear physics, and, for the first time, natural language via large language models. Whitening isolates the irreducible non-Gaussian dynamics, minimising Wasserstein-2 transport cost and eliminating numerical stiffness. The result is a general-purpose GP inference scheme requiring no bespoke derivations. Together, these results provide a general mechanism for incorporating the full richness of real-world knowledge as conditioning information, opening a new frontier for the probabilistic modelling of real-world problems.

Structure-based drug design with diffusion models seeks to generate 3D molecules that bind tightly to target protein pockets, but full-step sampling is too slow for high-throughput use. Extending the recent development of training-free fast sampling methods for diffusion models (e.g., DDIM and DNDM) to high-quality 3D molecule generation is non-trivial, given that both discrete variables (atom types) and continuous variables (atom coordinates) are involved. In this paper, we propose a unified fast sampling framework that allows a shared time schedule across discrete and continuous variables for strategically accelerating the 3D ligand generation. Guided by an error-equalization principle, we combine the continuous and discrete error rates of hybrid sampling into a single step-importance density and derive a training-free coupling-aware schedule. We further propose a confidence-guided sampling strategy that integrates constraints of model prediction confidence, motion stability, and steric clash to control the acceleration based on geometric readiness. Extensive experiments demonstrate that our method achieves $18$-$20\times$ speedup compared to full-step diffusion models while preserving state-of-the-art generation quality across chemical properties, binding affinity, and chemical validity.


Conformal Agent Error Attribution

Naihe Feng ⋅ Yi Sui ⋅ Shiyi Hou ⋅ Ga Wu ⋅ Jesse Cresswell

When multi-agent systems (MAS) fail, identifying where the decisive error occurred is the first step for automated recovery to an earlier state. Error attribution remains a fundamental challenge due to the long interaction traces that large language model-based MAS generate. This paper presents a framework for error attribution based on conformal prediction (CP) which provides finite-sample, distribution-free coverage guarantees. We introduce new algorithms for filtration-based CP designed for sequential data such as agent trajectories. Unlike existing CP algorithms, our approach predicts sets that are contiguous sequences to enable efficient recovery and debugging. We verify our theoretical guarantees on a variety of agents and datasets, show that errors can be precisely isolated, then use prediction sets to rollback MAS to correct their own errors. Our overall approach is model-agnostic, and offers a principled uncertainty layer for MAS error attribution.


Conformal Language Modeling via Posterior Sampling

Nicolas Emmenegger ⋅ Theo X. Olausson ⋅ Armando Solar-Lezama ⋅ Chara Podimata

Large Language Models remain plagued by hallucinations. Recent work has sought to tame their prevalence using statistical techniques based on conformal prediction, with both theoretical and empirical success. However, these methods operate in a post-hoc fashion, treating the sampling procedure itself as atomic and then surgically altering samples to remove hallucinated claims. This disconnect between filtering and generation can result in samples that are incoherent, inconsistent, or simply unlikely under the model itself. Not only that, but post-hoc surgery is unable to shift probability mass towards more useful and helpful responses. To address these issues, we propose to instead sample from approximations to an LLM posterior, where the conditioning event corresponds to a calibrated, high-scoring region. We develop a calibration procedure tailored to the setting of conditional sequential generation that effectively identifies this region and achieves target risk control. Empirically, we apply our method to case studies focused on open-ended biography generation and mathematical problem solving; compared to prior work, we obtain the same statistical guarantees, with higher downstream utility.


Conformal Prediction for Time-Dependent PDEs

Joshua Stiller ⋅ Annika Schneider ⋅ Eyke Hüllermeier

Uncertainty quantification is crucial in scientific machine learning, where models inform safety-critical tasks such as flood forecasting, financial risk management, and thermal control in machines. Conformal prediction provides distribution-free coverage guarantees, but in time-dependent settings common to physics and engineering, these guarantees can break down, leading to systematic undercoverage. We study this problem in the context of surrogate models for time-dependent physical systems described by partial differential equations. We prove that in a function space setting, distributions at arbitrarily close times can be mutually singular, making exact coverage guarantees impossible. We further show that in discretized settings, the total variation distance between distributions often grows exponentially with the grid resolution mandating restrictions on the spatial discretization to ensure coverage. Finally, for certain settings, we show how to use principled weighted conformal prediction to obtain finite-sample coverage guarantees over growing time horizons.


ConnectomeBench2: A Unified Benchmark for Automated Connectomic Proofreading

Jeff Brown ⋅ Tim Farkas ⋅ Gleb Razgar ⋅ Edward Boyden

Proofreading—correcting segmentation errors in 3D brain reconstructions—is the rate-limiting step in synapse-resolution connectomics. We release ConnectomeBench2, a unified multi-species dataset of over 302,401 expert-labeled proofreading decisions with >2,000,000 associated images spanning four major open connectomes (mouse, human, zebrafish, fly), spanning both split and merge error correction. Trained on this dataset, a single Vision Transformer with shared encoders for mesh geometry and electron microscopy reaches human-level accuracy across species for split error correction and merge error identification; with performance scaling with data size and improving with species diversity. Beyond accuracy, we show that the model is well-calibrated within distribution, that measures of distribution distance predict where calibration and accuracy will degrade on unseen data, and that connectomics-specific pretraining and active learning-based sample selection show potential to substantially reduce the labeling effort needed to extend to new species and brain regions. The benchmark provides the infrastructure to train and evaluate increasingly capable vision models for connectomic proofreading.


Consolidating Reasoning with Test-Time Learning

Haotian Wu ⋅ Zeming Chen ⋅ Hao Zhao ⋅ Antoine Bosselut

Parallel thinking scales inference-time compute by generating multiple reasoning traces for a query, and aggregating them to reach a final answer. However, most aggregation algorithms treat each reasoning trajectory independently, ignoring the complementary reasoning sub-processes across the full ensemble of trajectories. To learn from these parallel trajectories, we propose CORAL (Consolidating Reasoning with Test-Time Learning), a novel framework for aggregating trajectories by jointly consolidating them into parametric memory using test-time training. Our algorithm meta-learns how to consolidate using nested optimization at training time, where an inner loop encodes reasoning trajectories into a lightweight LoRA adapter, and an outer loop learns meta-parameters for the inner loop such that the consolidation is helpful for the model to synthesize a final solution. As no standard datasets exist for training aggregation methods, we construct a novel OpenParallelThinking corpus composed of 12,810 math problems paired with a mixture of diverse trajectories from six LLMs. Evaluated on challenging mathematics and STEM benchmarks, CORAL consistently outperforms best aggregation baselines by 20% overall. Importantly, we curate reasoning trajectories at multiple quality levels to simulate the reality of noisy parallel reasoning and demonstrate that CORAL is especially robust to the quality of thinking trajectories and able to learn useful knowledge from entirely wrong solutions. Additionally, this weak-to-strong generalization feature scales with the number of incorrect trajectories available for aggregation, while RL can further incentivize such reasoning consolidation capability. We will release CORAL and OpenParallelThinking for use in future work on reasoning aggregation.


Consolidating Rewarded Perturbations for LLM Post-Training

Zheyu Zhang ⋅ Shuo Yang ⋅ Gjergji Kasneci

Post-training of language models is commonly framed as a sample-score-update loop implemented by gradient descent. A recent line of work, exemplified by RandOpt, relocates this loop to weight space, sampling Gaussian perturbations around a pretrained model and ensembling the top-$K$ rewarded specialists at inference. While competitive with PPO and GRPO under matched training compute, this prediction-level ensemble incurs $K$ forward passes per test example and does not extend cleanly to free-form generation. We ask whether the rewarded population can instead be folded into a single deployable model, replacing the inference-time ensemble with one consolidated update. A split-half analysis over 25 model-task pairs reveals reproducible low-rank structure in every case. We turn this geometry into CoRP (Consolidating Rewarded Perturbations), a gradient-free operator that combines reward-weighted aggregation, compatibility-aware reweighting, and a held-out validation gate, with no gradient flowing through the language model. Across five language models from $0.5$B to $8$B and five tasks covering math, code, and creative writing, CoRP improves the base model by $8.1$ points on average. Using one tenth of RandOpt's perturbation budget, CoRP exceeds single-inference RandOpt by $6.5$ points and recovers more than half of the gain of the 50-pass majority-vote ensemble, at one forward pass per test example.

Given a must-link graph $G = (V, E_G)$ and a cannot-link graph $H=(V,E_H)$ as input, the objective of constrained clustering is to partition $V$ into $k$ parts while minimizing the maximum cut ratio between $G$ and $H$. In this paper we design and analyze a variant of the spectral clustering algorithm using the $k$ smallest generalized eigenvectors as the embedding space. We prove that our proposed algorithm achieves an $O(1)$-approximation of the optimal cut-ratio under a natural assumption on the input graphs. We further study the case in which the cannot-link constraints might not be available, and develop an iterative pipeline that repeatedly applies the output of spectral clustering or constrained clustering to generate new constraints, thereby progressively improving cluster quality. We experimentally compare our proposed algorithm with the previous state-of-the-art, and demonstrate the superior performance of our algorithms on both synthetic and real-world datasets.

Diffusion and flow models enable powerful generative modeling, but controlled editing remains challenging due to the need to balance source fidelity and edit compliance along the generative trajectory. We propose a principled, inversion-free framework for flow-based image editing from a trajectory centric perspective. Rather than relying on inversion or heuristic guidance, we introduce a one-step look-ahead mechanism that evaluates the effect of candidate updates on both edit and fidelity rewards estimated via denoising, and adjusts the current update based on this evaluation. At each step, our method applies reward-driven corrections guided by the look-ahead evaluation under an explicit fidelity constraint, resulting in stable and coherent trajectory evolution. This look-ahead design mitigates gradient interference between edit and fidelity objectives by anticipating their interaction in future states, leading to improved trade-offs. Experiments demonstrate that the proposed approach achieves stronger semantic edits while better preserving source content compared to existing strong baselines.


Continuous p-adic Optimization

Julian Salazar ⋅ Dimitri Kanevsky ⋅ Matt Harvey ⋅ Pascal Getreuer ⋅ Lucas Dixon

We present the first method for intrinsic, continuous gradient-based optimization for models with $p$-adic parameters. To overcome the vanishing gradients caused by the discrete valuation of the $p$-adics, we lift parameters to the Berkovich affine line, a canonical analytification of the $p$-adic numbers that provides a path-connected space. This enables gradient descent that respects the non-Archimedean metric with a local transition cost that is linear. By extending losses to the Berkovich line, we get piecewise linear functions, corresponding to the cell decomposition of a tropical complex; gradients are efficiently computed via max-plus operations. We prove local convergence within tropical cells for linear models on the $\mathbb{Q}_p$ subtree, and demonstrate the success of our method on synthetic tasks and on the $p$-adic Quillian semantic benchmark [Martins, 2025], where we achieve performance equal or better than $\mathbb{R}$-valued counterparts while training in comparable time.

Graph neural networks exhibit a fragile relationship with depth, as adding layers leads to rapid homogenization of node representations (over-smoothing). Recent manifold-constrained approaches using doubly stochastic matrices address this issue, but the fundamental principle underlying their success remains unclear. This study investigates whether doubly stochasticity is essential or exemplifies a broader geometric principle for stable residual propagation in deep GNNs.We formalize three axioms to define residual mixing matrices as contractive monoids. We instantiate this framework through three manifolds (Orthogonal Group, Diagonal Contraction Monoid, Spectral Norm Ball) and evaluate them across nine node-classification benchmarks using four GNN backbones. Results demonstrate that contractive-monoid constraints maintain stable accuracy and low representational collapse at extreme depths, while standard GNNs degrade sharply beyond moderate depth. This work establishes that stable residual propagation arises from contractive, compositional structure anchored at identity, revealing a significantly larger design space for deep GNN architectures while maintaining theoretical guarantees.


Contrastive Distribution Matching for Amortized Sequential Monte Carlo in Discrete Diffusion

Jaihoon Kim ⋅ Taehoon Yoon ⋅ Prin Phunyaphibarn ⋅ Seungjun Kim ⋅ Morteza Mardani ⋅ Minhyuk Sung

Discrete diffusion models have emerged as powerful frameworks for generating structured categorical data. However, efficiently sampling from reward-tilted distributions remains a fundamental challenge. While Twisted Sequential Monte Carlo (SMC) offers asymptotic exactness for this task, estimating the optimal twist function in discrete state spaces necessitates costly Monte Carlo approximations, resulting a severe computational bottleneck at inference. To overcome this limitation, we introduce Contrastive Distribution Matching (CDM), a novel framework that amortizes the cost of SMC inference by learning a parameterized twist function via positive and negative samples. For efficient training, we reformulate the gradient estimator to leverage the closed-form forward kernels of discrete diffusion models. In practice, evaluating our learned twist function incurs less than 5% additional computational overhead compared to base model inference. Through extensive empirical evaluations, we demonstrate that CDM consistently outperforms existing baselines under matched wall-clock time. We validate the effectiveness and versatility of our approach across a diverse range of applications, including toxic text generation, regulatory DNA sequence design, protein designability, and diffusion large language model alignment.

Autoregressive neural simulators now match classical solvers on short-horizon prediction of physical systems, yet their accuracy degrades rapidly when rolled out over long horizons. In this work we identify transient amplification of perturbations around rollout trajectories as a structural mechanism driving rollout error. Using a linearization analysis we show that when the Jacobians along an autoregressive trajectory are non-normal and non-commuting, the model amplifies errors transiently, resulting in model rollout drift even when the overall system is asymptotically stable. Building on the analysis, we propose \emph{commutativity regularization}: a combination of two penalties designed to reduce the normality defect of individual Jacobians and the commutator norm of Jacobians across steps. The penalties are estimated with Jacobian-vector products and have no inference-time cost. We show a propagator bound that quantifies rollout error under approximate commutativity and normality. We evaluate UNet and FNO variants with commutativity regularization on 1D and 2D spatio-temporal data in synthetic and real settings, showing successful long-horizon rollouts over thousands of steps. We show that the method improves FourCastNet climate forecasts on ERA5 without using any new data. The gain is most pronounced out-of-distribution: trained on trajectories of a few hundred steps, regularized models remain in-distribution for thousands of rollout steps on initial conditions where baselines diverge.


Coordinating Hundreds of RL Agents through Scalable Inference-Time Search

Daniel Rajaonarivonivelomanantsoa ⋅ Oussama Hidaoui ⋅ Refiloe Shabe ⋅ Noah De Nicola ⋅ Juan Formanek ⋅ Ruan John de Kock ⋅ Sasha Abramowitz ⋅ Omayma Mahjoub ⋅ Arnol M Fokam ⋅ Simon Du Toit ⋅ Asim Osman ⋅ Siddarth Singh ⋅ Ulrich Armel Mbou Sob ⋅ Willie Brink ⋅ Arnu Pretorius ⋅ Felix Chalumeau

Multi-Agent Reinforcement Learning (MARL) is a powerful approach to large-scale systems control, such as power grids and traffic networks, where dozens to hundreds of agents must coordinate effectively. Furthermore, such systems often operate in digital or simulated settings, which allows for inference-time search to significantly improve upon zero-shot performance. A current leading approach in this setting is COMPASS, which builds a continuous family of policies conditioned on latent vectors that can be efficiently searched at test time, but requires all agents to share a single latent vector, limiting the space of reachable joint policies. We introduce ATLAS, a simple extension in which each agent conditions on its own latent vector while still optimising a joint objective. This expands the searchable policy space at negligible computational cost and requires no additional training. ATLAS consistently sets a new state-of-the-art on a benchmark of 14 scenarios across three environments, scaling to teams of up to 200 agents, widely out of distribution. Crucially, its lead sharpens with team size and compute, reaching up to a 45% performance gain at scale.


CoQuant: Covariance-Aware Rotation for 2-bit KV Cache Quantization

Zhongzhu Zhou ⋅ Donglin Zhuang ⋅ Jisen Li ⋅ Ziyan Chen ⋅ Shuaiwen Song ⋅ Ben Athiwaratkun ⋅ Xiaoxia Wu

INT2 KV-cache quantization is attractive for long-context LLM serving, but it remains difficult to make both accurate and deployable. Simple rotations such as Hadamard transforms reduce outliers, but still degrade at INT2 because they are not aligned with downstream attention. We propose CoQuant, an Ultra-low-bit KV Cache quantization method that estimates attention-aware covariance structures offline and uses them to derive fixed rotations and clipping thresholds for quantization. In this way, it aligns KV quantization with the covariance structures that attention actually consumes. More importantly, we not only provide theoretical justification but also develop a fully deployable CoQuant system with a custom INT2 attention kernel that remains compatible with paged KV-cache serving and fused kernel pipelines, enabling seamless integration into modern LLM serving frameworks such as SGLang and vLLM. We evaluate our methods on recent reasoning models with reasoning traces of up to 32k tokens across 5 tasks. On Qwen3-4B-Thinking-2507 and Qwen3-8B, CoQuant reduces the BF16 accuracy gap to 3.78 and 1.42 points, respectively, while naive rotation INT2 collapses to nearly zero. We further scale CoQuant to Qwen3-32B and GLM-4.7 (358B params), where it remains effectively on par with BF16. On long context - RULER-NIAH up to 128K, CoQuant remains robust on both Qwen3 models, while naive rotation INT2 collapses. System-wise, CoQuant reduces KV-cache memory by approximately 8×, improves throughput by up to 7× at large batch sizes under the same memory budget, and accelerates batch-size-1 decoding by up to 3× over BF16 due to reduced memory bandwidth overhead.

In tool-integrated reinforcement learning with verifiable rewards (RLVR), sequence-level verifier signals are issued for trajectories whose outcomes are determined by a sparse subset of tokens—tool names, argument keys, and argument values. Standard RLVR assigns the same advantage to every token in a rollout, diluting credit at precisely the positions that matter. We propose **Correctable Fork Tokens (CFT)**, which decomposes selective credit assignment into two subproblems: *where* to update, and *how much* to update at those positions. To this end, CFT introduces a training-only answer-conditioned branch sharing parameters with the student policy. The key insight is that conditioning on the ground-truth answer concentrates the model's distribution at decision-fork positions, providing hindsight localization without additional parameter capacity or imitation targets. Critically, this benefit requires accurate answer information: a larger-capacity teacher without the answer fails to replicate these gains, confirming that privileged answer information—not teacher capacity—is the active ingredient. CFT uses this entropy reduction ($IG_t$) to identify correctable fork positions, then rescales token-level credit at those positions while preserving the verifier advantage as the sole signal for update direction. On BFCL v3/v4, CFT improves multi-turn accuracy by **4.2/6.5** pp and overall accuracy by **3.1/2.2** pp over GRPO; it further surpasses On-policy distillation—which employs a larger Expert Model as teacher—by **3.6/6.7** pp on multi-turn, confirming that answer-conditioned fork localization rather than teacher capacity drives the improvement. On Tau3, the average score improves by **8.3** pp. We additionally present a three-level structured verifier for tool-call settings; ablations confirm that finer-grained verifier feedback and selective token credit are complementary. Our reproducible implementation is available at [https://anonymous.4open.science/r/CFT](https://anonymous.4open.science/r/CFT).


Correpondence Alignment For Improved Virtual Try-On

Jiyoung Kim ⋅ Youngjin Shin ⋅ Siyoon Jin ⋅ Dahyun Chung ⋅ Jisu Nam ⋅ Tongmin Kim ⋅ JongJae Park ⋅ Hyeonwoo Kang ⋅ Seungryong Kim

Existing methods for Virtual Try-On (VTON) often struggle to accurately transfer garment appearance, especially in unpaired settings where accurate person-to-garment correspondence is required. These methods do not explicitly supervise person-to-garment alignment, leaving correspondence to be learned implicitly within the generation model. In this paper, we first analyze full self-attention in DiT-based architectures and show that person-to-garment query-key matching aligned with local geometric correspondence is closely related to try-on quality. Building on this insight, we introduce CORrespondence ALignment (CORAL), a DiT-based framework that explicitly aligns query-key matching with robust external correspondences. CORAL integrates two complementary components: a correspondence distillation loss that aligns reliable matches with person-to-garment attention, and an entropy minimization loss that sharpens the attention distribution. For evaluation, we further propose a VLM-based evaluation protocol to better reflect human preference. CORAL consistently improves over the baseline, enhancing both global shape transfer and local detail preservation. Extensive ablations validate our design choices. Code and weights will be publicly available.


Cost-Aware Best-LLM Identification using Dueling Feedback

Sarvesh Gharat ⋅ Nikhil Karamchandani ⋅ Jayakrishnan Nair

Inspired by the problem of identifying the best model from a collection of large language models (LLMs) with heterogeneous querying costs, we formulate and analyse a variant of the multi-armed bandit (MAB) with (i) dueling feedback, where pairwise comparisons between model responses provide robust preference signals, and (ii) heterogeneous sampling costs, reflecting the differing costs of querying different LLMs. Assuming the existence of a Condorcet winner, a condition we empirically validate across multiple real-world datasets, we propose a Track-and-Stop style algorithm for best-arm identification with prescribed confidence. We prove that the algorithm almost surely achieves the asymptotically optimal cost as the error tends to zero. Finally, we extensively evaluate our approach on both synthetic and real-world instances, demonstrating consistent improvements over classical cost-unaware algorithms and their cost-aware extensions.


Counterfactual Instruction Grounding for Vision-Language-Action Models

Shiyu Liu ⋅ Yu Zhou ⋅ Meng Liu ⋅ Xuanming Guo ⋅ Liqiang Nie

Vision-Language-Action (VLA) models inherit rich semantic priors from pretrained vision-language models, yet these priors do not guarantee persistent instruction sensitivity after narrow-domain fine-tuning. We observe a systematic failure mode: although VLA models remain visually reactive, their later actions can become weakly constrained by the commanded instruction, leading to unstable object interaction and visually plausible but instruction-inconsistent behavior. We trace this behavior to a collapse of instruction-conditioned representations across network depth and task progress. More fundamentally, this collapse is enabled by the learning objective itself: standard imitation learning admits Bayes-optimal solutions that are invariant to language. We address this by recasting instruction grounding as identifying the commanded instruction among counterfactual alternatives. Based on this view, we propose Counterfactual Instruction Grounding (CIG), a contrastive objective that encourages the generated trajectory to remain identifiable with the commanded instruction among counterfactual alternatives. CIG applies to both autoregressive and flow-matching VLA models, using exact action-chunk likelihoods for autoregressive models and an energy-based pseudo posterior for flow-matching models to avoid intractable trajectory likelihood estimation. Practically, CIG can be applied directly to already fine-tuned models as a continuation-stage grounding objective, restoring instruction sensitivity without retraining from scratch or changing the architecture. Extensive experiments on RoboCasa and LIBERO-Plus demonstrate that CIG achieves strong instruction-following performance with success rates of 71.3% and 86.3%, respectively. Real-world deployment further shows that CIG transfers to kitchen manipulation and preserves instruction-conditioned behavior in visually ambiguous scenes.


CoVisIT: Cross-modal Prior Guided Diffusion Model for Visible-to-Infrared Image Translation

Linfeng Tang ⋅ Tong Hu ⋅ Zizhuo Li ⋅ Hao Zhang ⋅ Han Xu ⋅ Zhenfeng Shao ⋅ Jiayi Ma

Visible sensors provide rich scene details but are vulnerable to adverse conditions, whereas infrared sensors capture thermal radiation distributions and are more robust to environmental variations. However, acquiring high-resolution infrared images remains challenging due to sensor limits and hardware costs, and collecting pixel-wise aligned visible-infrared pairs is even more difficult. Existing methods mainly rely on infrared super-resolution or visible-to-infrared (VIS-to-IR) translation: the former preserves thermal distributions but lacks fine details and spatial alignment with visible images, while the latter benefits from fine-grained structural cues but faces an ill-posed thermal-distribution inference problem. To overcome these limitations, we propose CoVisIT, a cross-modal prior guided diffusion model for VIS-to-IR image translation. Our key insight is to constrain VIS-to-IR translation with complementary cross-modal priors, using high-resolution visible images to recover fine-grained scene details and misaligned low-resolution infrared inputs to anchor reliable thermal distributions. To achieve this, we build upon a latent diffusion framework and develop a Cross-Modal Adaptation module that modulates visible and infrared features in a shared feature space, thereby reducing the modality gap and exploiting complementary cross-modal priors. To further address local misalignment, we introduce a Dynamic Kernel Generation Module to predict input-adaptive kernels and embed a Multi-Scale Dynamic Convolution Module into the denoising U-Net, enabling dynamic local feature aggregation for implicit alignment. As a result, CoVisIT generates high-resolution, visible-aligned infrared images with reliable thermal distributions. Extensive experiments demonstrate state-of-the-art performance and consistent improvements in downstream tasks.


CP-MLPs: A Tensor-Rank Theory of Tied and Untied MLP Blocks

Md Rifat Arefin ⋅ Farzaneh Heidari ⋅ Irina Rish ⋅ Guillaume Rabusseau

Transformer feed-forward blocks combine several design choices: squared activations, gating, and tied versus untied parameter sharing-whose separate contributions to expressivity remain poorly understood. We introduce \emph{CP-MLPs}, a tensor framework in which each multiplicative unit contributes a rank-1 interaction with three roles: detecting an input feature, reading a value, and writing to the output. Tied gated blocks such as ReLU$^2$ force the detector and value roles to share the same direction; untied gated blocks such as ReGLU and SwiGLU let them differ. We show that this distinction yields exact, degree-wise tensor-rank separations: untied detector-value units are strictly more width-efficient for indefinite quadratic interactions and role-asymmetric monomials, while tied units are structurally matched to symmetric pure-power targets. Tying is therefore not inherently better or worse where its benefit depends on the structure of the target. In experiments, models below the required width hit an irreducible approximation floor where the loss does not improve with extra optimization budget. The detector-value advantage also persists in trained gated MLPs at matched parameter count.


CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models

Chunzheng Zhu ⋅ Jiaqi Zeng ⋅ Hongbo Zhao ⋅ Yihang Chen ⋅ Yijun Wang ⋅ Jianxin Lin

As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a single misleading sentence can override a correct image based diagnosis, or why a model commits to a confident answer despite insufficient visual evidence. We find that these two safety risks, arbitration failure where textual context overrides visual grounding and brake failure where the model commits without adequate evidence, are mediated by spatially disjoint attention head populations: arbitration heads form a mid-to-deep wideband reflecting cross-layer evidence competition, while brake heads concentrate in a narrow middle-to-late layer band that regulates evidence sufficiency and abstention behavior. To ground these observations in causal circuitry, we introduce CRAFT, which localizes each failure mode to a minimal causal head set via dual criteria and verifies necessity and sufficiency through temporal probes and Tuned Lens trajectory analysis. Excising arbitration heads sharply reduces conflict following with negligible degradation on clean inputs, while excising brake heads restores appropriate abstention under degraded visual evidence. The two interventions target spatially disjoint head sets and produce distinct corrective effects, underscoring the mechanistic separability of the failure modes. Experiments across multiple medical VQA benchmarks and VLM architectures validate both the localization and interventions, demonstrating that the identified heads causally drive each failure mode and that targeted modulation generalises without retraining. The code is available at \url{https://anonymous.4open.science/r/CRAFT-B3CA}.


CRAFT: Counterfactual-to-Interactive Reinforcement Fine-Tuning for Driving Policies

Keyu Chen ⋅ Nanfei Ye ⋅ Yida Wang ⋅ Wenchao Sun ⋅ Danqi Zhao ⋅ Hao Cheng ⋅ Sifa Zheng

Open-loop imitation learning has advanced modern autonomous driving policy architectures, but closed-loop deployment remains vulnerable to policy-induced distribution shift. Existing post-training paradigms exhibit fundamental trade-offs: closed-loop RL fine-tuning provides grounded feedback from executed actions but is constrained by the sparsity of informative events, whereas counterfactual fine-tuning provides dense supervision over candidate futures but inherits bias from imperfect future estimates. We introduce Counterfactual-to-Interactive Reinforcement Fine-Tuning (CRAFT), an on-policy framework that formulates closed-loop post-training as proxy-residual optimization. CRAFT uses group-normalized counterfactual advantages as a dense proxy for real closed-loop advantages and aligns this proxy with the closed-loop world through grounded residual correction from interaction-critical events. To stabilize adaptation, CRAFT regularizes the online policy toward an EMA teacher via asymmetric KL self-distillation. Theoretically, CRAFT decomposes the real closed-loop policy gradient into proxy and residual terms under the same visited-state distribution, reducing residual variance with an aligned proxy while mitigating proxy bias through grounded residual approximation. Empirically, CRAFT achieves the strongest closed-loop gains on Bench2Drive across hierarchical planning, vision-language-action, and vocabulary-scoring architectures. Ablations, scaling behavior, stability analyses, and transfer results further validate the complementary roles of dense counterfactual proxy and grounded residual correction.

Selecting a pretrained encoder for a downstream task without retraining or labels is a standard preprocessing step in transfer learning. The dominant label-free score, RankMe, was validated within architecture families and is widely applied across them. We show that its cross-architecture use is unsound: rankings are systematically driven by the penultimate-layer dimension rather than representational quality, and on heterogeneous model pools the score performs worse than uniform random selection. We give a random-matrix explanation establishing this failure as inherent to any rotation-invariant score, and propose Subspace-Corrected Ranking (SCR), a family of label-free corrections that breaks rotation invariance via alignment with a fixed reference encoder. Across three architecturally distinct pools and 23 aggregated task units, the proposed SSR-dagger reduces mean regret from +0.165 (RankMe) to +0.071 (95% CI [+0.045, +0.099], p < 1e-4). Paired with a within-family score in a two-stage pipeline, it matches this quality at one-third of the forward-pass cost. Five labelled architecture evaluations suffice to match the best fully label-free strategy; K=35 labelled evaluations reduce regret to +0.018 --- a 5.7x improvement at the same forward-pass cost.

Open-vocabulary semantic segmentation (OVSS) models enable dense prediction for classes specified at test time, but continual test-time adaptation (CTTA) over long, non-stationary streams can corrupt their language-aligned semantic interface, which we diagnose through synonym-prompt text alignment. We propose Cross-Foundation Complementary Learning Systems (XF-CLS), a source-free OVSS-CTTA framework that makes lightweight online adaptation stable over long horizons. Inspired by Complementary Learning Systems, which couple rapid plastic learning with slow consolidation, XF-CLS decouples plasticity from stabilization: a single fast online adaptation path updates only a small set of visual normalization parameters, a horizon-aware, source-anchored slow parameter memory consolidates stable changes and recovers from drift, and a frozen cross-foundation structural observer supplies a non-drifting structural prior for reliability-aware recovery and dense output refinement. Rather than training an additional teacher or using cross-foundation pseudo-label supervision, the frozen path preserves the open-vocabulary interface while correcting degraded visual structure. On long-horizon Cityscapes-to-ACDC and OnDA streams, XF-CLS achieves 32.92 and 30.75 mIoU, improving the NACLIP no-adapt baseline by +9.00 and +4.13 mIoU, respectively, and outperforming continual OVSS-TTA baselines. Decomposition studies and text-alignment diagnostics show that source-anchored recovery prevents long-horizon collapse, while frozen cross-foundation refinement adds consistent gains without updating extra parameters. XF-CLS provides a stable, parameter-efficient route to source-free continual open-vocabulary dense prediction, with inference overhead dominated by one additional frozen-backbone forward pass.


Cross-order Consensus Graph Matching

Yu Xie ⋅ Yiming Zhang ⋅ Xinyan Liang ⋅ Junqi Liu ⋅ Ming Li

Graph matching aims to establish node correspondences while preserving both unary appearance compatibility and pairwise structural consistency. Recent deep graph matching methods have achieved strong empirical performance, but many of them either compress edge information into node embeddings before matching or rely on surrogate optimization signals for quadratic objectives. These design choices can weaken structural discrimination and may lead to optimization directions that are not aligned with the relaxed quadratic assignment problem (QAP). We propose Cross-order Consensus Graph Matching (CCGM), a framework that explicitly couples first-order node assignment and second-order edge assignment. CCGM constructs line graphs to convert edge matching into a node matching problem, learns node and edge affinities in parallel, and refines them through a Cross-order Consensus Solver motivated by the exact gradient of the factorized relaxed QAP. We further introduce Cross-order Alignment Regularization to encourage agreement between node-induced edge assignments and edge-induced node support during training. Experiments on standard visual graph matching benchmarks show that CCGM consistently improves matching accuracy over competitive baselines.


CrossSteer:Cross-Modal Safety Steering for Audio-Language Models

Houde Dong ⋅ Weifei Jin ⋅ Yuxin Cao ⋅ Yuxin Cao ⋅ Derui Wang ⋅ Jie Hao

Audio--language models (ALMs) introduce a new jailbreak surface in which harmful requests can be delivered through speech. Existing safeguards rely on audio-side filters or guard models, leaving internal safety behavior largely uncontrolled. We instead pursue a representation-level alternative that controls refusal--compliance behavior inside the shared language-model backbone. However, our empirical study shows that direct audio-side steering is noisy and ineffective, consistent with activation analyses indicating larger acoustic and front-end variation in speech-derived representations. Based on this observation, we propose CrossSteer, a cross-modal steering method that learns a cleaner semantic safety direction from text-only preference pairs and transfers it to audio through the aligned shared residual stream. CrossSteer fits a residual-stream direction whose intervention shifts harmful-request generation from unsafe compliance toward safe refusal, and applies this direction at an audio-robust layer. Across three ALM backbones, CrossSteer consistently reduces audio jailbreak attack success rate while largely preserving benign utility, demonstrating cross-modal transfer of text-derived safety steering to the audio channel. Additional experiments show that CrossSteer composes with existing safeguards, further improving robustness as a complementary representation-level safety layer.Our anonymized source code is available at:~\url{https://anonymous.4open.science/r/CrossSteer-69F6}.


Cross-User Poisoning: User-Task Boundary Failures in Multi-User Collaborative Language Agents

Atharv Singh Patlan ⋅ Peiyao Sheng ⋅ S Ashwin Hebbar ⋅ Prateek Mittal ⋅ Pramod Viswanath

Language agents are moving from single-user assistants to shared collaborators in workspaces, forums, and group chats. A shared agent observes interleaved messages from multiple users, maintains persistent context, and executes tool actions, creating a boundary problem absent from standard single-user agents: the agent must decide not only whether an instruction is safe, but which user or task it is allowed to govern. We identify cross-user poisoning (CUP), where an adversary injects a message into shared context that is later applied while the agent serves a benign user, causing unauthorized actions or responses outside the instruction's intended scope. We validate CUP on two deployed multi-user agents, Continua and ElizaOS, and introduce MURMUR, a framework for evaluating shared-context agents under concurrent multi-user interactions. Across Slack, Workspace, and Airline domains, CUP achieves high attack success, persists across later interactions, and remains substantially more effective than matched prompt-injection attacks. We evaluate boundary-scoping defenses and find that context summarization alone is insufficient, while task clustering and provenance prompting substantially reduce non-adaptive propagation. These results show that robust multi-user agents require explicit user/task scoping rather than generic input filtering or context compression.


CT-Lesion: A Multi-Region CT Dataset for Co-existing Lesion Segmentation and Detection

Ruihan Lu ⋅ Leiqiang Pan ⋅ Shijie Wang ⋅ Dingyi Zhang ⋅ Jianjun Shen ⋅ Xin Yu

Lesion analysis in CT imaging, the task of identifying and delineating abnormal regions across diverse anatomical sites, plays a fundamental role in clinical diagnosis. Existing CT lesion datasets typically focus on a single disease type within a specific region. This design leads to two main limitations. First, co-existing lesions outside the predefined categories are left unannotated, so models may treat real lesions as background during training. Second, lesion categories within any single region follow a long-tailed distribution, in which rare but clinically important categories are underrepresented. However, collecting enough samples for every rare category within a single region is often impractical. We observe that certain lesion categories are rare in one region but occur more frequently in another, motivating the use of cross-region complementarity for lesion detection. In this work, we curate a multi-region CT lesion dataset, namely CT-Lesion, which covers brain, chest, and abdomen CT images. Our CT-Lesion is sourced from a private cohort of 500 clinical CT scans. Seven radiologists spent approximately 1,600 working hours delineating lesions of any type, rather than restricting annotations to a predefined set of categories. As a result, CT-Lesion provides clinically confirmed voxel-level segmentation masks and 28 fine-grained lesion labels. Moreover, we further group the lesions into 9 coarse categories to alleviate data scarcity issues. To demonstrate the effectiveness and utility of CT-Lesion, we benchmark the representative methods on CT-Lesion for lesion segmentation and detection. Furthermore, pretraining on CT-Lesion improves model performance on external benchmarks, including LiTS, KiTS, and LIDC-IDRI, demonstrating its potential as a transferable resource for CT lesion analysis. Our dataset is available at Kaggle.


CUDABench: Benchmarking LLMs for Text-to-CUDA Generation

Jiace Zhu ⋅ Wentao Chen ⋅ Qi Fan ⋅ Zhixing Ren ⋅ Junying Wu ⋅ Xing Z Chai ⋅ Chotiwit Rungrueangwutthinon ⋅ Yehan Ma ⋅ An Zou

Recent studies have demonstrated the potential of Large Language Models (LLMs) in generating GPU Kernels. Current benchmarks focus on the translation of high-level languages into CUDA, overlooking the more general and challenging task of text-to-CUDA generation. Furthermore, given the hardware-specific and performance-critical features of GPU programming, accurately assessing the performance of LLM-generated GPU programs is nontrivial. In this work, we introduce CUDABench, a comprehensive benchmark designed to evaluate the text-to-CUDA capabilities of LLMs. First, we construct CUDABench-Set, which covers Breadth-Depth-Difficulty evaluation space in diverse application domains, including artificial intelligence, scientific computing, and data analytics, etc. Furthermore, we propose CUDABench-Score and Generative Verification Pipeline that assess (1) compilation correctness, (2) functional consistency through execution-based verification, and (3) a novel roofline-based metric, Performance-Score. Benchmarking state-of-the-art LLMs reveals insightful findings and challenges of text-to-CUDA, such as a notable mismatch between high compilation success rates and low functional correctness, a lack of domain-specific algorithmic knowledge, and suboptimal utilization of GPU hardware resources.


Curvature-Guided Parameter Initialization for Multi-Task Learning

Linxiao Cao ⋅ Zhipeng Zhou ⋅ Xutao Huang ⋅ Min Zhou ⋅ Menglin Yang

Multi-task learning (MTL) aims to jointly optimize multiple related tasks to obtain shared representations within a single model. Recent research has focused on developing MTL algorithms to alter optimization dynamics through task re-weighting or gradient manipulation. However, in this paper, we first empirically identify that current MTL approaches are highly sensitive to the model initialization, complicating empirical evaluation and limiting practical reliability, which has been overlooked but appears to be critical. To better understand this phenomenon, we provide a local theoretical analysis that motivates the connection between early-stage MTL optimization behavior and curvature around initialization. Motivated by this analysis, we propose a Curvature-guided Parameter Initialization (CPI) approach tailored for MTL. Specifically, CPI performs a short warm-up phase to probe local curvature around the initial parameters and heuristically biases the initialization toward empirically favorable curvature profiles for subsequent multi-task optimization. The proposed approach is optimizer-agnostic, requires no modification to existing MTL algorithms, and can be seamlessly integrated as a plug-in prior to standard training. Extensive experiments on standard MTL benchmarks show that CPI improves the aggregate multi-task trade-off while remaining compatible with mainstream MTL methods.


CutAttn: Discovering Cognitive Transition Layers for Efficient Long-Context Prefilling

Wentao Liu ⋅ Xiabao Wu ⋅ Yongchao Liu ⋅ Haitao Zhang ⋅ Jiajun Zheng ⋅ Ruiting Zhou

Long-context LLM inference is fundamentally a retrieval process, transitioning from diffuse to sharply focused attention at a few architecture-determined (CTLs). Existing acceleration methods leave significant costs unaddressed: sparse attention ignores FFNs, KV cache compression neglects prefill, and hidden-state pruning relies on fixed, heuristic rules. We introduce CutAttn, a training-free framework that exploits CTLs via inter-layer Jensen--Shannon divergence, pruning redundant context at chunk granularity while dynamically gated by a normalized-entropy criterion. Operating at the hidden-state level, CutAttn reduces both attention and FFN costs, composing orthogonally with existing acceleration techniques. Evaluated on RULER, InfiniteBench, and LongBench, CutAttn preserves full-attention accuracy while achieving up to a $2.02\times$ prefill speedup on 128K sequences (and $2.81\times$ when composed with sparse attention). Furthermore, our theoretical analysis provides approximation bounds, an entropy-compression limit justifying the gating, and a closed-form speedup ratio.


CyberDualEval: Measuring Dual-Use Cyber Risks in Frontier Language Models

Magnus Saebo ⋅ Francesco Piccoli ⋅ Maxwell Watson ⋅ Pranjali Thakur ⋅ Yangruibo Ding ⋅ ZHUO ZHANG

Frontier language-model agents are rapidly becoming capable cybersecurity assistants, creating a dual-use cyber risk: the same capabilities that help defenders find, understand, and validate vulnerabilities can also help attackers weaponize them. Existing cyber benchmarks largely measure either raw offensive capability or generic refusal behavior, but they do not evaluate whether models can preserve benign security utility while refusing assistance that materially advances misuse. We introduce CyberDualEval, a benchmark for measuring this boundary on real vulnerabilities. Each task is decomposed into three phases of increasing operationalization: vulnerability analysis, proof-of-concept generation, and exploit execution. To capture the contested middle ground, we annotate each proof-of-concept task with a Task Weaponization Score that measures how much work remains to turn a minimally sufficient demonstration into unauthorized cyber misuse. CyberDualEval evaluates frontier language models with varied agentic scaffolds and reports refusal, defensive analysis accuracy, and optional exploit validation as distinct signals. Our evaluation reveals two complementary failure modes. Many models over-comply, assisting with high-weaponizability PoCs or exploit requests that should be refused. Conversely, recent safety-tuned models often over-refuse, declining low-risk or defensive requests and thereby reducing benign cyber utility. These results show that cyber safety cannot be assessed by capability or refusal rate alone: safe deployment requires calibrated refusal that tracks weaponizability while preserving useful defensive assistance.


CyBiasBench: Benchmarking Bias in LLM Agents for Cyber-Attack Scenarios

Taein Lim ⋅ Seongyong Ju ⋅ Munhyeok Kim ⋅ Hyunjun Kim ⋅ Hoki Kim

Large language models (LLMs) are increasingly deployed as autonomous agents in offensive cybersecurity. In this paper, we reveal an interesting phenomenon: different agents exhibit distinct attack patterns. Specifically, each agent exhibits an attack-selection bias, disproportionately concentrating its efforts on a narrow subset of attack families regardless of prompt variations. To systematically quantify this behavior, we introduce CyBiasBench, a comprehensive 630-session benchmark that evaluates five agents on three targets and four prompt conditions with ten attack families. We identify explicit bias across agents, with different dominant attack families and varying entropy levels in their attack-family allocation distributions. Such bias is better characterized as a trait of the agents, rather than a factor associated with the attack success rate. Furthermore, our experiments reveal a bias momentum effect, where agents resist explicit steering toward attack families that conflict with their bias. This forced distribution shift does not yield measurable improvements in attack performance. To ensure reproducibility and facilitate future research, we release an interactive result dashboard at \url{https://cy-bias-bench.vercel.app/} and a reproducibility artifact with aggregated session-level statistics and full evaluation scripts at \url{https://anonymous.4open.science/r/CyBiasBench-2026/}.


Cylindrical Geodesic Flow Matching for Quasiperiodic Physiological Signal Transformation

Onur Selim Kilic ⋅ Afra Nawar ⋅ Cem O Yaldiz ⋅ Michael J Cho ⋅ Ahmet R Emirdagi ⋅ Demet Tangolar ⋅ Amirali Aghazadeh ⋅ Amit J Shah ⋅ Omer T Inan

Paired translation between quasiperiodic physiological waveforms (i.e., recovering a target oscillatory signal from the source) is central to the interpretation of cardiovascular signals derived from wearables placed at different body locations. This source-to-target mapping in these problems carries inherent geometric structure: the phase wraps around the cycle and must be treated as a circular variable, the amplitude remains strictly positive, and the beat-to-beat alignment can drift unpredictably across cycles and subjects. While deep neural networks have been used for phase estimation and complex-valued signal modeling, prior work does not explicitly learn phase transport between paired signals. Consequently, neither endpoint-supervised regression nor the standard affine path used in flow matching accounts for this phase--amplitude structure. We introduce \emph{cylindrical geodesic flow matching} for paired cardiovascular waveform translation. We show that the standard affine path used in flow matching distorts intermediate amplitude and instantaneous frequency when interpolating between quasiperiodic signals; replacing it with a closed-form geodesic on the phase--amplitude cylinder eliminates these artifacts and converts each training pair into dense, geometry-consistent velocity supervision. On zero-shot photoplethysmography and limited-support seismocardiography adaptation benchmarks, our method consistently outperforms interpolation baselines and matches or exceeds direct supervised prediction, reducing Hilbert Transform, $L_2$, and Dynamic Time Warping distance by up to ${\sim}15\%$ over the strongest competing baseline. These results suggest that bridge geometry is a critical inductive bias for flow matching on oscillatory signal translation.


DACE RL for Compute Efficient Reinforcement Learning in Small Model Reasoning

Guo Yue ⋅ YANG LIU ⋅ Tang Qingkang ⋅ Donghui Zhang ⋅ Li Aoyu ⋅ Rong Fu

Reinforcement learning with verifiable rewards (RLVR) has become a dominant recipe for improving mathematical and code reasoning in open-weight language models, but existing methods rely on fixed hyperparameters despite highly non-stationary training dynamics. We present DACE-RL (Diagnosis-Aware Compute-Efficient RL), a closed-loop framework that promotes training diagnostics to first-class signals. DACE-RL tracks four lightweight statistics—token entropy, sample diversity, advantage saturation, and answer consistency—and feeds their EMA-smoothed state to a controller that adaptively sets per-prompt rollout budget, sampling temperature, asymmetric clipping bounds, and entropy regularization strength. We formalize adaptive rollout allocation as a difficulty-conditioned multi-armed bandit and provide a regret bound that improves over uniform-budget GRPO under heterogeneous prompt difficulty; we also give a stability result for the dual-loop entropy-diversity controller. Across three small base models (Qwen2.5-Math-1.5B, Qwen2.5-7B, and Llama-3.1-8B-Instruct) and six math reasoning benchmarks, DACE-RL matches or improves Pass@1 while reducing rollout-GPU-hours by 38–52%, and improves Pass@32 by 4.7 absolute points on average. Pilot analysis further shows that diagnostic inflection signals appear 100–300 update steps before reward plateau, indicating that diagnosis-aware control is a key lever for compute-efficient RLVR.


DAGent: Evaluate-then-Grow Planning for Deep Research Agents

Hanwen Liu ⋅ Yuanfu Sun ⋅ Qiaoyu Tan

Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as intermediate findings emerge. Directed acyclic graph (DAG)-based multi-agent systems are well suited to this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents typically instantiate a task-level plan before execution and then repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions often waste computation on branches that should not have been planned in the first place. To bridge this gap, we propose DAGent, a DAG-based multi-agent framework that introduces Evaluate-then-Grow incremental planning. Instead of committing to a full DAG upfront, an Orchestrator grows the task graph one batch at a time, conditioning each new expansion on confidence and uncertainty signals from completed nodes. To support long-horizon evidence use without overloading each sub-task, DAGent maintains a hierarchical context layer that propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The cleanly recorded DAG topology in turn admits structural RL signals that an outcome-only recipe cannot define; we instantiate this with DAGRPO, a GRPO adaptation that injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, with the lead replicating across four open-source backbones from four vendors and extending to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO further improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison further shows that incremental, evidence-conditioned planning recovers accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart, supporting the view that DAGent's gains come from more targeted evidence expansion rather than additional computation. Anonymized code is available at \url{https://anonymous.4open.science/r/DAGent-F2D6}.


Data Density Scaling Laws for Image Self-Distillation

Alvard Barseghyan ⋅ Ani V Vanyan ⋅ Hakob Tamazyan ⋅ Hrant Khachatrian

Current scaling law formulations suggest that increasing the number of unique samples in a dataset consistently improves model performance, while repeated exposure to the same samples provides diminishing value. In this work, we challenge this assumption for image self-distillation models. Through systematic experiments across diverse datasets, namely Maxar, Waymo, and ImageNet, we demonstrate that the effects of sample repetition on downstream performance are non-trivial and dataset-dependent. We discovered one scenario when repeatedly training on a subset of data can outperform training on a larger set of unique samples. This finding highlights a critical limitation of existing scaling laws. To address this gap, we propose a modified version of the Chinchilla scaling law that explicitly disentangles two factors: number of unique samples and repetition per sample. This formulation provides a more precise framework for understanding data-compute trade-offs and better predicts performance in regimes where repetition play a central role. Finally, we analyze how the inherent diversity of the smaller unique subset influences final downstream performance.


Decentralized Coupled Representation Learning

Zilin Li ⋅ Weiwei Xu ⋅ Xuchun Tong ⋅ Xuanbo Lu ⋅ Xuanqi Zhao ⋅ Ge Zhang

The learning dynamics of decentralized coupled representation learning can be formulated as a non-autonomous stochastic dynamical system, yet the macroscopic behavior of the resulting system on continuous manifolds remains poorly understood. We establish three theoretical results for such coupled slow-fast dynamics. First, under a standard random geometric graph scaling regime, we prove that the infinitesimal generator of the discrete Markov process converges to the generator of an overdamped Langevin diffusion for every $f \in C^3(\mathcal{M})$. Second, under timescale separation, the stochastic parameter dynamics converge to a deterministic averaged flow that admits a row-orthogonality Lyapunov function, controlling row-space degeneracy and ruling out parameter divergence. Third, under a spectral gap condition, the row space of the averaged parameter trajectory converges to the principal eigenspace of the induced covariance matrix. Together, these results give a stochastic-approximation limit theory for the coupled regime in which Markovian sampling and local subspace learning co-evolve, showing that memoryless local interactions can induce a stable and structured macroscopic limit.


Decomposing the modulation of interactions between neuronal populations

Marco Celotto ⋅ J. S Sooter ⋅ Sofie Ährlund-Richter ⋅ Kyle R Jenks ⋅ Mriganka Sur ⋅ Stefano Panzeri

Identifying subpopulations of neurons that interact with each other from simultaneous recordings of populations of many neurons is key for understanding across-brain communication with cellular resolution. Recent work identified communication subspaces, which capture additive interactions between pairs of high-dimensional neural populations through a small number of source and target activity patterns. However, no current method captures how a third, potentially multivariate variable - such as behavioral state or the activity of a third population - modulates these interactions. Here we extend the communication subspace framework by parameterizing modulation as a low-rank tensor. This identifies multiplicative interaction channels (MICs), defined as triplets of source, target, and modulator activity patterns, in which the modulator pattern gates the source-target interaction. We derive MICs as a bilinear perturbation of reduced-rank regression. We develop a hierarchical fitting pipeline and provide a closed-form decomposition that quantifies whether modulation reshapes the modulator-averaged baseline interaction, recruits private dimensions of one population, or opens new interactions. In simulations, MICs reliably recover the presence and geometry of ground-truth modulation even in the high-dimensional, low-sample regime. Applying MICs to simultaneous calcium imaging of prefrontal axons and interneurons in the visual cortex revealed that behavioral state asymmetrically modulates top-down interactions, reconfiguring the patterns of prefrontal projections that interact with a stable set of visual interneuron activity patterns. By providing an efficient and compact characterization of modulatory interactions, MICs enable asking new questions about how potentially high-dimensional variables shape interactions between neural populations.


DEDCA: Test-Time Adaptation for Generalized AI-Generated Image Detection

Hao Tan ⋅ Qimin Zhang ⋅ Lihuang Fang ⋅ Kebing Jin ⋅ Jinghui Qin

AI-generated image detectors face significant challenges when deployed in real-world environments, particularly when test samples are produced by unseen generators that deviate from the source training distribution. Existing detectors often rely on a single type of evidence, such as semantic representations or low-level generation traces, making them vulnerable when the corresponding cue becomes unreliable. Moreover, adapting detectors to newly emerging generators usually requires target labels or retraining, which is costly and impractical. To address these challenges, we propose Dual-Evidence Disagreement-Constrained Adaptation (DEDCA), a test-time adaptation framework for generalized AI-generated image detection. DEDCA constructs a dual-evidence detector by combining CLIP-based semantic evidence with trace evidence extracted from shuffled entropy and residual statistics, and introduces a reliability-aware gate to perform sample-specific adaptive fusion. Our key idea is to exploit the disagreement between the two evidence streams as an unlabeled adaptation signal, enabling the detector to adjust to target-domain shifts without relying on target labels or blindly trusting a single prediction stream. Experiments show that DEDCA achieves superior performance on the GenImage benchmark by outperforming state-of-the-art AI-generated image detectors and test-time adaptation methods with 92.0\% average accuracy.


DeepfakeGenome: Toward Next-Generation Deepfake Attribution

Shunli Wang ⋅ Xinyu Zhou ⋅ Yandan Zhao ⋅ Wenbin Liang ⋅ Taiping Yao ⋅ Ke-Yue Zhang ⋅ Zhiyuan Yan ⋅ Shouhong Ding ⋅ Lizhuang Ma

In recent years, AIGC technologies are capable of generating hyper-realistic forged facial images, posing severe threats to facial security safety. To tackle these challenges, early Deepfake research primarily centered on real-fake classification. Recently, growing efforts have been devoted to the Deepfake Attribution (DFA) of generated content. Nevertheless, existing research on Deepfake detection and attribution faces saturated performance in binary classification, limited diversity in datasets and algorithms, and imperfect evaluation protocols, which severely impede practical application. To address these limitations, we propose a comprehensive deepfake detection and attribution benchmark named DeepfakeGenome (DFG). It contains 100 facial forgery algorithms and 2M images in total, achieving 4× to 100× larger than prior DFA benchmarks. We further designed 4 protocols for practical evaluation, including a novel retrieval-based attribution paradigm. Unlike previous open-set evaluation metrics, the proposed retrieval metrics are more aligned with the real-world active defense situation of blacklist registration mechanisms. Based on these elaborate designs, we investigate the performance ceiling of deepfake attribution task. Over 2k+ experimental evaluations are conducted, and 10 insightful findings are derived. We hope this work can provide new insights into the DFA research field.


Deep Probabilistic Supervision for Image Classification

Anton Adelöw ⋅ Matteo Gamba ⋅ Atsuto Maki

Supervised training of deep neural networks for classification typically relies on hard targets, which promote overconfidence and can limit calibration, generalization, and robustness. Self-distillation methods aim to mitigate this by leveraging inter-class and sample-specific information present in the model’s own predictions, but often remain dependent on hard targets without explicitly modeling predictive uncertainty. With this in mind, we propose Deep Probabilistic Supervision (DPS), a principled learning framework that constructs sample-specific target distributions via Bayesian inference on the model’s own predictions and remains independent of hard targets after initialization. We show that DPS consistently yields higher test accuracy (e.g., +2.0\% for DenseNet-264 on ImageNet) and significantly lower Expected Calibration Error (ECE) (-40\% for ResNet-50 on CIFAR-100) than existing self-distillation methods. When combined with a contrastive loss, DPS achieves state-of-the-art robustness under label noise.

Reinforcement learning in real-world systems often involves delayed feedback, which breaks the Markov assumption and impedes both learning and control. Canonical augmentation-based approaches cause state-space explosion, which imposes a severe sample-complexity burden. Despite recent progress, state-of-the-art augmentation-based baselines either mainly alleviate the burden on the critic or rely on non-unified treatments for the actor and critic. In this study, we propose delayed homomorphic reinforcement learning (DHRL), a framework grounded in MDP homomorphisms that defines a belief-equivalence relation over the augmented state space to collapse control-redundant augmented states. In principle, this yields exact abstraction under deterministic dynamics and approximate abstraction under stochastic dynamics, enabling both the actor and critic to benefit from a structured abstraction mechanism. In finite domains, exact abstraction preserves optimality and recovers the delay-free sample-complexity order, whereas approximate abstraction admits a value-loss bound on the resulting policy. For continuous domains, we introduce deep delayed homomorphic policy gradient (D$^2$HPG), a deep actor-critic instantiation of the DHRL framework. Experiments on continuous-control tasks in MuJoCo show that D$^2$HPG outperforms strong augmentation-based baselines.

The first-moment buffer of most modern optimizers is an exponential moving average (EMA) of stochastic gradients with a single scalar decay, with a unified forgetting horizon along every direction in parameter space. Yet the per-sample gradient of any linear layer factorizes as a rank-1 outer product $g_t = \delta_t x_t^\top$, exposing the input activation as a natural $\textit{key}$ and the output-side error as a natural $\textit{value}$, which is the structure that EMA's flat matrix average discards. We propose $\textbf{DeltaMomentum}$, which interprets the buffer as an online linear associative memory of these key--value pairs and updates it via the classical delta rule. The resulting update is $\textit{anisotropic memory transport}$: directions queried frequently are forgotten quickly, directions queried rarely are preserved on long horizons, yielding an automatic, data-driven schedule matched to input statistics. We prove the buffer's expected fixed point is a Tikhonov-regularized Wiener predictor of the population gradient, which is equivalent to implicit input-side natural-gradient preconditioning at no covariance-inversion cost, and that the same mechanism strictly accelerates expected tracking-error contraction along every positive-density input direction under non-stationarity and reduces the per-direction iteration determinant in linear regression. DeltaMomentum is a drop-in replacement for the EMA accumulator in any base optimizer; we instantiate it as DeltaSGD and DeltaAdamW. A $\mu$P derivation shows the delta coefficient is width-invariant, enabling zero-shot hyperparameter transfer; we verify this with a coordinate check across five widths. The per-block asymptotic FLOP overhead is $\approx 15.74$% at zero persistent-memory cost (down to $7.3$% and $11.0$% realized at 67M/370M). On FineWeb-Edu, DeltaAdamW reaches AdamW's terminal validation loss in up to $31.35$% fewer steps at 67M and $19.28$% fewer steps at 370M. On CIFAR-10, DeltaSGD shows the same pattern against SGD-momentum, confirming the gain is not specific to Adam. Mechanistic diagnostics confirm that improvements in gradient-estimator quality, function-space prediction error, and input-feature conditioning are operative throughout training.

This paper explores the theoretical foundations of fair regression under the constraint of demographic parity within the unawareness framework, where disparate treatment is prohibited, extending existing results where such treatment is permitted. Specifically, we aim to characterize the optimal fair regression function when minimizing the quadratic loss. Our results reveal that this function is given by the solution to a barycenter problem with optimal transport costs. Additionally, we study the connection between optimal fair cost-sensitive classification, and optimal fair regression. We demonstrate that nestedness of the decision sets of the classifiers is both necessary and sufficient to establish a form of equivalence between classification and regression. Under this nestedness assumption, the optimal classifiers can be derived by applying thresholds to the optimal fair regression function; conversely, the optimal fair regression function is characterized by the family of cost-sensitive classifiers.


Density-Ratio Losses for Post-Hoc Learning to Defer

Alexander Soen ⋅ Ragnar Thobaben ⋅ Joakim Jaldén ⋅ Richard Nock

We study post-hoc Learning to Defer (L2D) through the lens of ideal distributions: divergence-regularized reweightings of the data distribution under which a model attains low loss. We define deferral via the density-ratio between a model's and an expert's ideals. Using the reduction from density-ratio estimation to class-probability estimation, we derive the DR CPE losses for post-hoc L2D scorers. Deferral decisions are then made by thresholding the scorer, allowing deferral rates to be adjusted without retraining. For KL-based ideal distributions, our deferral rules recovers Chow's rule under the original distribution and a connection to an expert-tilted Bayes posterior—which incorporates the expert's performance—depending on if the ideal distributions are joint or marginal distributions. Experimentally, our approach is competitive compared to common baselines and more robust across dataset settings. More broadly, our results cast post-hoc L2D as density-ratio learning between ideal distributions, bridging Chow-style rules, expert comparison, and elucidating connections to related learning settings including anomaly detection.


Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View

Yixian Xu ⋅ Yuanrui Zhang ⋅ Shengjie Luo ⋅ Kaiyuan Gao ⋅ Liwei Wang ⋅ Di He

Reinforcement learning (RL) post-training provides a direct way to align diffusion and flow-matching generators with human preferences and task-specific rewards, but current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising pairs. This paper shows that these seemingly different losses arise from a single path-space principle. Starting from the regularized diffusion-RL objective, we use importance sampling between sampling SDEs to obtain an explicit policy-gradient estimator on trajectory space. The estimator contains the stochastic It\^o integral underlying Flow-GRPO-type updates; we derive an equivalent variance-reduced value-gradient form that recovers the forward-matching structure of AWM and DiffusionNFT. This identifies the empirical gap between these method families as a variance-reduction effect rather than a difference in RL principle. The derivation yields a unified design space organized by value-gradient estimation, weight functions, and sampling choices. Within this space, we propose a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones. Experiments on SD3.5-M and Qwen-Image models validate the variance-reduction explanation and show that the resulting recipe matches or improves over prior diffusion-RL baselines.

Foundation models produce high-dimensional latent representations in which suitably designed statistical decision rules can be effective yet analytically tractable. We study such detectors in the high-dimensional regime, where latent dimension and training sample size grow proportionally. For linear and quadratic statistics, we derive explicit large-deviation characterizations of false-alarm and miss-detection performance under the Neyman--Pearson paradigm, and show how the binary analysis extends to multi-class classification. The resulting error probabilities and corresponding decay rates exhibit non-monotonic dependence on the shape parameter, including double-descent and, in some regimes, multiple-descent. Experiments on vision, text, and audio datasets using modern foundation models show close agreement between theoretical predictions and empirical performance. In particular, the proposed framework allows us to set a prescribed false-alarm level and accurately predict the corresponding miss-detection probability. Overall, our results provide a theoretically grounded and interpretable framework with competitive empirical performance across different modalities.


Dexterous Skill Discovery via Topology-Aware Wasserstein Dependency

James Heald ⋅ Kai Biegun ⋅ Alexandre Galashov ⋅ Maneesh Sahani

Dexterous object manipulation remains a fundamental challenge in robotics, typically requiring intensive supervision through explicit rewards, goals, or demonstrations. While unsupervised reinforcement learning (RL) holds the promise of discovering dexterous skills, state-of-the-art metric-aware methods have yet to be successfully applied to complex, high-dimensional manipulators like multi-fingered hands. We argue that this limitation stems from two key factors: existing methods lack the necessary inductive biases to prioritize object interaction over agent motion, and they fail to respect the non-Euclidean topology of object configuration space $\mathrm{SE}(3)$. To address these limitations, we propose a novel Wasserstein dependency measure between skills and object motions, formulated in the Lie algebra $\mathfrak{se}(3)$. Our approach, Topology-Aware SKill discovery (TASK), enables the discovery of fundamental manipulation skills—such as multi-axial object translation and rotation—without manual supervision. Across a range of embodiments, including non-prehensile and multi-fingered hands, our method learns dexterous behaviors where previous approaches fail, marking a significant step toward autonomous robotic dexterity.


DGRAF: Observation-Quality-Aware Reinforcement Learning for Dynamic Reconfigurable Batteries

Jiasong Chen ⋅ Jingwei Hu ⋅ Zheng Fang ⋅ Zhihong Zhang

Dynamic Reconfigurable Battery (DRB) systems enable flexible interconnection among cells through power electronic switches, crucial for electric vehicles and energy storage systems. However, sensor failures in practical deployments lead to partial observation loss, posing severe challenges for safe control. Existing methods assume full state observability and often adopt overly conservative strategies when observations are missing, failing to achieve effective trade-offs between safety and performance under uncertainty. More fundamentally, these methods lack mechanisms to adaptively adjust decision conservatism based on inference reliability, leading to unstable decisions under complex failure scenarios. We propose the $\textbf{D}$iffusion-$\textbf{G}$uided $\textbf{R}$isk-$\textbf{A}$daptive $\textbf{F}$ramework (DGRAF) to address this challenge. The framework handles spatiotemporally coupled failure patterns through joint diffusion inference with quality quantification, and integrates inference quality with distributional decision-making to enable automatic modulation between aggressive optimization and conservative maintenance. Through extensive simulations and hardware experiments, we demonstrate that DGRAF consistently outperforms baseline methods across complex failure patterns, achieving superior energy efficiency, decision stability, and risk resilience.


DiA: Directional Adapter

Ashish Singh ⋅ Prakash Chandra Chhipa

Fine-grained video action recognition often depends on how an action unfolds over time rather than on appearance alone. Visually similar classes may differ in preparation, execution, or completion phases, making uniformly aggregated temporal context suboptimal. We study this problem under a practical objective: i) improving recognition performance with few tunable parameters, ii) being compute efficient, and iii) using only the readily available video modality. Motivated by this missing temporal observation and the accuracy--compute efficiency--unimodality objective, we propose **Directional Adapter (DiA)**, a parameter-efficient adapter on top of CLIP. DiA formulates temporal specialization into causal and anti-causal directions, allowing past-to-present and future-to-present cues to be modeled differently. To keep this directional modeling parameter- and compute-efficient, both directions use depth-wise temporal convolutions in a compact bottleneck space and are combined through a learnable fusion. DiA achieves state-of-the-art performance across 7 public benchmarks with only $\sim$3.5M tunable parameters, while also delivering higher throughput and lower latency than prior methods. Code is available at https://anonymous.4open.science/r/diafullsupervised-8CC3/README.md.

Compact medical multimodal models are the most realistic path to clinical deployment, yet they often produce fluent answers that are not actually grounded in the image. We show that this failure has a clear internal signature. In five compact medical backbones and across six medical VQA benchmarks, visual tokens lose their spatial diversity shortly after entering the language backbone, their residual updates are an order of magnitude smaller than text residuals, and attention on them freezes on content free regions of the image. Zeroing the visual residual stream barely changes the output, while zeroing the text residual stream destroys it. We give a short analytical account that links the collapse of visual similarity to the residual dominance ratio, and we propose VGrip, a set of three lightweight losses that act on the three failure points without modifying the backbone architecture or the inference procedure. We also introduce the Visual Dependency Score, a simple and benchmark agnostic measure of how much a model actually uses the image. On Qwen2.5-VL-7B, VGrip lifts accuracy by 4.7 points on average across the six benchmarks and more than doubles the Visual Dependency Score over a strong LoRA baseline, with gains that are stable from two to eight billion parameters.


Differentiable Exact Learning of Algorithms

Hristo Papazov ⋅ Francesco D'Angelo ⋅ Nicolas Flammarion

We introduce a framework for differentiable exact learning of algorithms. The framework couples a state-tracking neural controller to an unbounded external environment and trains on policy-trajectory observations (PTOs) -- traces that record the environment's evolution under an expert's actions without exposing the expert's internal state. This supervision regime sidesteps Gold's impossibility barrier for input-output recursive learning while avoiding the unrealistic internal-state access required by prior provably correct approaches. As controllers, we propose Differentiable Finite-State Transducers (DFSTs), a minimalist multilinear model family that contains no nonlinearities and admits log-parallel training via prefix scans. Trained on tiny datasets, DFST and RNN controllers achieve error-free generalization on binary and decimal addition and multiplication to operands thousands of times longer than the training examples. Moreover, we provide strong empirical proof for exact learning of universal computation. Finally, we develop an extraction procedure that recovers compact finite-state transducers from trained controllers. We prove under a geometric clustering assumption that for binary addition the extracted transducer exactly replicates the expert policy on all valid inputs -- certifying differentiable exact learning of the algorithm.


Differentiating Network Design Objectives for Balancing Cost and Distance

Zishang Chen ⋅ Longhe Lin ⋅ Zhiyuan Julian Su ⋅ Qi Qi

Many network design tasks must decide which directed arcs to build so that multiple sources can reach a root efficiently. Building fewer or cheaper arcs reduces construction cost, but may force long routing paths; building more arcs shortens routes, but increases construction cost. The Cost--Distance problem formalizes this trade-off, yet it has remained difficult to optimize with gradient-based methods because the selected network, the induced routes, and the routing cost are tightly coupled. We propose Cost Distance Policy Gradient (CDPG), a differentiable framework that treats local next-hop choices as a routing policy, prevents unstable cyclic routing through Dynamic Acyclic Dropout, evaluates the objective with a truncated value solver, and uses the induced acyclic support for quality-preserving TopoRounding. Under the stated conditions, CDPG has a fixed-accuracy (\mathcal{O}(m\log n)) efficiency guarantee. Experiments cover 1,487 instances in UAV logistics networks, transportation networks, and four synthetic graph families; across this suite, CDPG shows a strong time--quality trade-off compared with other differentiable baselines, Cost--Distance-specific algorithms, and the commercial solver. Our code is available at: https://anonymous.4open.science/r/cdpg_nips-28D5/.


DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing

Artyom Mazur ⋅ Nina Konovalova ⋅ Aibek Alanov

Mechanistic interpretability seeks to explain neural network behavior by decomposing model computations into interpretable features and circuits. While transcoder-based circuit tracing has recently enabled detailed causal analyses of large language models, multimodal diffusion transformers for image generation remain comparatively opaque. We still lack tools for understanding how semantic information propagates across denoising steps and how text and image representations interact within double-stream MM-DiT architectures. Existing methods provide only partial insight: attention maps expose a limited view of token interactions, while sparse autoencoders can discover interpretable features but do not directly reveal how these features are transformed and composed through nonlinear MLP layers. In this work, we extend transcoder-based circuit tracing to multimodal diffusion transformers. We train timestep-conditioned transcoders that faithfully approximate the input–output behavior of MLP sublayers in FLUX.1[schnell]. By replacing MLPs with transcoders and linearizing the remaining computation, we obtain exact feature-to-feature attribution and recover compact, interpretable circuits. Empirically, our transcoders match or slightly outperform sparse autoencoders on the sparsity–faithfulness tradeoff. The resulting circuits reveal mechanisms underlying attribute binding and cross-stream semantic propagation, and provide causal explanations for systematic generation errors. Moreover, circuit-guided interventions are substantially more precise and effective than standard SAE-based steering. Our results demonstrate that transcoder-based circuit analysis is feasible for state-of-the-art diffusion transformers and provides a powerful framework for understanding and controlling multimodal generative models.


DiffRisk: Diffusion Representation Learning with Informative Missingness for Health Risk Prediction

Shailesh Dahal ⋅ Ratri Mukherjee ⋅ Nicholas Mathews ⋅ Kishlay Jha

Health risk prediction aims to predict a patient's potential health risks (e.g., mortality) using the longitudinal information present in their electronic health record (EHR). While deep learning-based risk prediction models have shown great promise, they are susceptible to the issue of missing data prevalent in real-world EHR. Most prior works mitigate this challenge by treating missing data as an uninformative perturbation to be ignored or imputed. However, inaccurate imputation may lead to the generation of sub-optimal patient representations that adversely impact risk prediction performance. Moreover, recent studies show that missing data in EHR does not necessarily imply error or noisy observation and instead may communicate informative clinical signals useful for health risk prediction. To address this, we propose a novel diffusion-based representation learning approach (namely $\textbf{DiffRisk}$) that explicitly captures $\textit{informative missingness}$ by modeling conditional dependencies induced by partial observations in EHR. Specifically, DiffRisk learns latent representations with coherent conditional structure, where any subset of observed features defines a distribution over the unobserved components. To achieve this, we introduce a conditional score-based objective and approximate it via denoising diffusion that enables efficient learning without explicit density estimation. Empirical results on five risk prediction tasks show that the proposed approach consistently outperforms state-of-the-art baselines.


Diffusion Masked Pretraining for Dynamic Point Cloud

Zhuoyue Zhang ⋅ Yiding Sun ⋅ Chaowei Fang ⋅ Haozhe Cheng ⋅ Jian Liu ⋅ Jihua Zhu ⋅ Ajmal Mian

Dynamic point cloud pretraining is still dominated by masked reconstruction objectives. However, these objectives inherit two key limitations. Existing methods inject ground-truth tube centers as decoder positional embeddings, causing spatio-temporal positional leakage. Moreover, they supervise inter-frame motion with deterministic proxy targets that systematically discard distributional structure by collapsing multimodal trajectory uncertainty into conditional means. To address these limitations, we propose Diffusion Masked Pretraining (DiMP), a unified self-supervised framework for dynamic point clouds. DiMP introduces diffusion modeling into both positional inference and motion learning. It first applies forward diffusion noise only to masked tube centers, then predicts clean centers from visible spatio-temporal context. This removes positional leakage while preserving visible coordinates as clean temporal anchors. DiMP also reformulates point-wise inter-frame displacement supervision as a DDPM noise-prediction objective conditioned on decoded representations. This design drives the encoder to target the full conditional distribution of plausible motions under a variational surrogate, rather than collapsing to a single deterministic estimate. Extensive experiments demonstrate that DiMP consistently improves downstream accuracy over the backbone alone, with absolute gains of 11.21% on offline action segmentation and 13.65% under causally constrained online inference.

Diffusion language models generate text through iterative denoising, but current interpretability tools mostly analyze static hidden states and do not explain how concept information evolves across denoising time. We introduce DynaManifold-SAE, a method for discovering sparse autoencoder feature groups whose activations form time-dependent concept manifolds in masked and diffusion language models. The method builds sparse latent codes across mask ratios, constructs candidate feature groups, and evaluates them with held-out coordinate prediction, geodesic consistency, persistence across time, and causal interventions. Across BERT and LLaDA, DynaManifold-SAE reliably identifies sparse sentiment geometry that is stable across seeds and stronger than individual-feature, PCA, random-group, and graph-structured baselines. The discovered groups transfer from controlled templates to natural SST-2 examples, indicating that they capture semantic structure rather than template artifacts. We further show that these groups explain a denoising-specific mechanism: they predict token reveal order and confidence growth during LLaDA unmasking, and targeted interventions alter reveal confidence and token recovery more than matched controls. These results suggest that sparse feature manifolds provide a practical bridge between static representation geometry and the dynamics of diffusion language generation.


Diffusion Tree Search for Inference Time Adaptation of Material Foundation Models

Daniel Levy ⋅ Vineet Jain ⋅ Tara Akhound-Sadegh ⋅ Oumar Kaba ⋅ Siamak Ravanbakhsh

Diffusion-based foundation models for crystal generation have become competent priors over the manifold of plausible materials, but turning that prior into generation of promising candidates requires targeting stability while balancing several conflicting properties at once. Fine-tuning is expensive and produces a new model for every new objective, while differentiable guidance is not applicable to many real-world rewards. We instead adapt these models at \emph{inference time} using Diffusion Tree Search (DTS), a Monte Carlo tree search procedure over the denoising trajectory. We make two technical contributions on top of DTS that are particularly relevant for AI-for-science applications. First, we show how to use existing property-conditional foundation models (e.g., classifier-free guidance variants of MatterGen) as guided proposals inside DTS, with importance-corrected backups and selection that keep the unconditional base model as the target distribution. Second, we extend DTS to multi-objective sampling by maintaining per-objective soft value estimates and selecting expansions through a sampled scalarization, so a single search produces samples spanning a Pareto front while sharing computation across preferences. We instantiate both adaptations on de novo crystal generation: single-objective stability search with different underlying models, conditional generation of crystals belonging to rare space groups, and a multi-objective tasks using a pretrained conditional model. Across these, DTS-based inference-time adaptation improves stability rates and Pareto-front quality without modifying the base models.


Dimension Bounds for Contractive Reservoir Computing from Input Entropy

Pradeep Singh ⋅ Kishore Babu Nampalle ⋅ Balasubramanian Raman

Echo State Networks (ESNs) are designed to be \emph{stable}—the echo state property makes the reservoir state a well-defined function of the input history—but it remains unclear how \emph{geometrically rich} the reachable reservoir states are and what controls that richness. We give a partition-free, fractal-geometric theory of contractive reservoirs by viewing a quantized-input ESN as a contractive random dynamical system and, equivalently, an iterated function system (IFS). This identifies the reachable set as the IFS attractor and the stationary reservoir-state law under i.i.d.\ inputs as the unique invariant IFS measure. We prove entropy–contraction (entropy–Lyapunov) bounds on the intrinsic dimension of this measure, and in a conformal/similarity regime with standard separation conditions we obtain sharp closed-form dimension laws of the form “input entropy divided by contraction.” These results yield a quantitative \emph{criticality principle}: weakening contraction drives a transition from low-dimensional to full-dimensional state representations at an explicit threshold set by input entropy and contraction rates, providing practical design guidance for tuning leak/spectral radius and input scaling.


DiReCL: Learning Differentiable Reward Code with Inverse Reinforcement Learning

Aoran Wang ⋅ Jingtao Zhang ⋅ Maosen Li ⋅ Zongzhang Zhang

In reinforcement learning (RL), reward specification plays the fundamental role in shaping policy optimization. Inverse RL (IRL) offers a data-driven approach to inferring rewards from demonstrations, but prevalent neural-network-based methods often produce rewards lacking interpretability and flexibility for manual inspection or revision. In contrast, human-written or large language model (LLM)-generated reward code provides explicit semantic structure, yet its precise numerical parameters are often difficult to calibrate and prone to specification errors. To bridge this gap, we propose DiReCL (\textbf{Di}fferentiable \textbf{Re}ward \textbf{C}ode \textbf{L}earning), a novel framework that unifies these two paradigms by formulating reward learning in a differentiable reward code space. DiReCL first prompts an LLM to synthesize the reward code templates, then converts the numeric elements in these templates into differentiable parameters and align them with expert demonstrations through IRL. We further introduce a reflection mechanism that uses IRL diagnostics to provide evidence of inadequacies in the reward structure and guide template revision. Extensive experiments on MuJoCo locomotion and autonomous driving benchmarks demonstrate that DiReCL achieves over $1.25 \times$ the performance of both standard IRL and LLM-assisted reward design baselines, while preserving the crucial benefits of explicit, interpretable reward semantics.


Direct Conditional Parameterization for N-Dimensional Splatting

Zhongpai Gao ⋅ Benjamin Planche ⋅ Van N Nguyen ⋅ Meng Zheng ⋅ Anwesa Choudhuri ⋅ Arun Innanje ⋅ Terrence Chen ⋅ Ziyan Wu

N-dimensional splatting extends 3D Gaussian Splatting with conditioning variables such as view direction and time to model view-dependent and dynamic effects. Each primitive is lifted into a joint distribution over 3D position and conditioning variables, and at render time is conditionally sliced at the query to recover a 3D primitive whose mean, opacity, and covariance vary continuously with view or time. N-DGS (covering 6DGS and 7DGS) and UBS together represent the leading N-dimensional splatting formulations across Gaussian and Beta kernels, but share the same conditional-slicing step: all three effects are derived from the joint covariance matrix per primitive in every forward pass, requiring matrix inversions and regression-matrix multiplications that scale with primitive count. The conditional-slicing step is kernel-agnostic, so a single improvement carries across families. We introduce direct conditional parameterization, replacing the covariance-derived form with explicit per-effect parameters: a Cholesky precision factor for opacity and a displacement matrix with learnable per-dimension coupling for position. Applied to both kernels, the parameterization yields dGS (Gaussian) and dBS (Beta). Across five static and dynamic benchmarks, dGS and dBS match or exceed their covariance-derived counterparts on quality (up to +1.26 dB in the main MCMC evaluation) while running 6.9-7.7x faster at the slicing step and up to 2.66x faster end-to-end rendering. Direct conditioning is a kernel-agnostic drop-in replacement for covariance-derived conditioning in N-dimensional splatting.


Direct Product Flow Matching: Decoupling Radial and Angular Dynamics for Few-Shot Adaptation

Hongxu Chen ⋅ Yanghao Wang ⋅ Bowei Zhu ⋅ Hongxiang Li ⋅ Ziqi Jiang ⋅ Zhen Wang ⋅ Lin Li ⋅ Rui Liu ⋅ Long Chen

Recent flow matching (FM) methods improve the few-shot adaptation of vision-language models, by modeling cross-modal alignment as a continuous multi-step flow. In this paper, we argue that existing FM methods are inherently constrained by incompatible geometric priors on pre-trained cross-modal features, resulting in suboptimal adaptation performance. We first analyze these methods from a polar decomposition perspective (\ie, radial and angular sub-manifolds). Under this new geometric view, we identify three overlooked limitations in them: \textit{1) Angular dynamics distortion}: The radial-angular coupling induces non-uniform speed on the angular sub-manifold, leading to regression training difficulty and extra truncation errors. \textit{2) Radial dynamics neglect}: Feature normalization discards modality confidence, failing to distinguish out-of-distribution and in-distribution data, and abandoning crucial radial dynamics. \textit{3) Context-agnostic unconditional flow}: Dataset-specific information loss during pre-trained cross-modal feature extraction remains unrecovered. To resolve these issues, we propose \textbf{warped product flow matching (WP-FM)}, a unified Riemannian framework that reformulates alignment on a warped product manifold. Within this framework, we derive \textbf{direct product flow matching (DP-FM)} by introducing a constant-warping metric, which yields a decoupled cylindrical manifold (\ie, direct product manifold). DP-FM enables independent radial evolution and constant-speed angular geodesic transport, effectively eliminating angular dynamics distortion while preserving radial consistency. Meanwhile, we incorporate classifier-free guidance by conditioning the flow on the pre-trained VLMs' hidden states to inject missing dataset-specific information. Extensive results across 11 benchmarks have demonstrated that DP-FM achieves a new state-of-the-art for multi-step few-shot adaptation.


DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning

Hengyu Fu ⋅ Tianyu Guo ⋅ Zixuan Wang ⋅ Hanlin Zhu ⋅ Jason Lee ⋅ Jiantao Jiao ⋅ Stuart J Russell ⋅ Song Mei

Implicit in-weight multi-hop reasoning---composing multiple pieces of parametric knowledge within a single forward pass---is a fundamental yet challenging task of language models. Vanilla transformers consistently fail at this task, and even looped transformers, whose recurrent structure naturally fits its iterative nature, generalize imperfectly, particularly out-of-distribution (OOD). We make two contributions toward closing this gap. First, on a symbolic two-hop reasoning task, we mechanistically diagnose why looped transformers fall short: after the first loop, the intermediate answer is already decodable from the residual stream with certain probability, yet the continuous hidden vector carrying it is noisy and geometrically misaligned with the clean discrete embedding of the intermediate answer that the next loop would ideally consume. A training-free intervention that re-aligns this representation nearly closes the OOD gap, indicating that the representational mismatch is a bottleneck. Second, building on this insight, we propose DiscoLoop, a looping architecture whose recurrence carries both a discrete embedding channel and a continuous hidden-state channel. DiscoLoop achieves near-perfect two-hop accuracy with substantially less training across symbolic and synthetic-language multi-hop reasoning tasks. When applied to real pretraining, DiscoLoop attains lower training loss and stronger performance on pretraining benchmarks than looped-transformer baselines, suggesting that the mixed-channel design transfers to practical language modeling.


Discovering Phase Space Structure in Learned Hamiltonian Systems

Jiayin Liu ⋅ Yulong Yang ⋅ Vineet Bansal ⋅ Christine Allen-Blanchette

Mechanics-inspired neural dynamical models improve interpretability and physical consistency, but they usually assume a known configuration space, fixed physical parameters, and deterministic prediction from a prescribed initial condition. We introduce Structured Phase Space GAN (SPS-GAN), a conditional generative model that learns structured dynamics from both Cartesian trajectories and video. SPS-GAN combines a port-Hamiltonian backbone with adversarial conditional generation, allowing conservative, dissipative, and forced systems to be represented in a single architecture. To infer latent mechanical structure from raw observations, we use a cyclic-coordinate loss that recovers both the system degrees of freedom and canonical coordinates without access to the true phase space. Experiments on simulated and real-world systems show that SPS-GAN improves trajectory accuracy over supervised dynamics baselines and video consistency over generative baselines, while consistently recovering the correct latent mechanical coordinates across observation modalities and dynamical regimes.


Discovering Unseen Degradations to Adapt Open-World Image Restoration

Boseong Kim ⋅ Tae Hyun Kim ⋅ Donghyeon Cho

While all-in-one image restoration models excel in controlled, closed-set environments, they face critical limitations when deployed in open-world settings where paired supervision is unavailable and degradations are unknown and often intricately mixed. To address this challenge, we propose a continual restoration framework that integrates degradation discovery with adaptive restoration. Building on a model pre-trained on known degradation categories, our method discovers unseen degradations in feature space and continually expands the degradation space. We then introduce a discovery- and instance-conditioned descriptor that jointly captures discovered degradation priors and image-specific characteristics, enabling robust restoration under ambiguous and intricately mixed degradations. To further support adaptation under unpaired supervision, we adopt a mean-teacher-based semi-supervised framework equipped with a reliable bank for pseudo-target refinement. Specifically, we propose a discovery-adaptive score that assesses pseudo-target reliability using feature-space distances to degradation categories, enabling pseudo-target selection to remain adaptive to newly discovered degradations rather than relying solely on conventional perceptual quality-based scoring methods. Experiments on both controlled open-world protocols and real-world adverse weather datasets demonstrate that our approach consistently outperforms existing methods across multiple image quality metrics while improving robustness to degradation composition shifts.


Disen-Forcing: Disentangling Semantic Anchoring from Motion for Autoregressive Video Diffusion

Zhengyang Yu ⋅ Akio Hayakawa ⋅ Masato Ishii ⋅ Qingtao Yu ⋅ Xuesong Li ⋅ Takashi Shibuya ⋅ Yuki Mitsufuji ⋅ Jing Zhang

Autoregressive video diffusion models (AR-VDMs) enable real-time, long-horizon video generation but degrade rapidly once the inference horizon outruns the training horizon, due to compounding exposure bias. To stabilize long rollouts, recent methods heuristically adopt frame sink, which keeping initial frames as a persistent KV cache context, inspired by attention sinks in streaming LLMs. We revisit this design and find that the mechanism by which initial frames stabilize long-range consistency in AR-VDMs is not an LLM-style attention sink; instead, consistency is anchored through a sparse subset of attention heads, while the majority of heads predominantly capture local temporal dynamics. This observation reveals an over-conditioning pathology in existing frame-sink strategies: by universally exposing every head to the pinned initial-frame context, the model becomes prone to collapsing motion toward the low-dynamic modes of the teacher during model distillation. We propose Disen-Forcing, a distillation-time plugin for AR-VDMs that disentangles and reinforces this head specialization through disentangled context routing and motion-disentangling regularization. On long-horizon video generation benchmarks, Disen-Forcing simultaneously improves semantic consistency and motion dynamics of base AR-VDMs: a trade-off that sink-based methods fail to balance.

Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL based on teacher--student confidence, and a batch-level epistemic rejection mechanism (Revival) that suppresses updates when the teacher is more uncertain than the student. We provide a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce. Empirically, across eight teacher--student pairs and eleven benchmarks spanning instruction following, chat alignment, code generation, and mathematical reasoning, CaRE-KD delivers consistent and significant gains over strong baselines (Skewed-KL, $\alpha$--$\beta$ divergence). Highlights include up to $+3.2$ average ROUGE-L on instruction-following tasks, $+2.1$ pass@1 on MBPP, $+1.7$ accuracy on GSM8k, and $+1.8$ accuracy on CollegeMath over the strongest baseline, with consistent gains in LLM-as-a-judge factuality (up to $+2.5$ per task over Skewed-RKL). Revival further acts as a principled, loss-agnostic plug-in that systematically strengthens existing distillation objectives by filtering epistemically unreliable teacher supervision.

Distributed online convex optimization (D-OCO) is a powerful paradigm for modeling distributed scenarios with streaming data. However, the communication cost between local learners and the central server is substantial in large-scale applications. To alleviate this bottleneck, we initiate the study of D-OCO with compressed communication. Firstly, to quantify the compression impact, we establish the $\Omega(\delta^{-1/2}\sqrt{T})$ and $\Omega(\delta^{-1}\log{T})$ lower bounds for convex and strongly convex loss functions, respectively, where $\delta \in (0,1]$ is the compression ratio. Secondly, we propose an optimal algorithm, which enjoys regret bounds of $O(\delta^{-1/2}\sqrt{T})$ and $O(\delta^{-1} \log T)$ for convex and strongly convex loss functions, respectively. Our method innovatively incorporates the error feedback mechanism into the Follow-the-Regularized-Leader framework to address the coupling between the compression error and the projection error. Furthermore, we employ the online compression strategy to mitigate the accumulated error arising from the bidirectional compression. Our online method has great generality, and can be extended to the offline stochastic setting via online-to-batch conversion. We establish convergence rates of $O(\delta^{-1/2}T^{-1/2})$ and $O(\delta^{-1} T^{-1})$ for convex and strongly convex loss functions, respectively, providing the first guarantees for distributed non-smooth optimization with compressed communication and domain constraints.


Distributionally Robust Algorithmic Recourse for Tree Ensembles

Kentaro Kanamori ⋅ Ken Kobayashi ⋅ Takuya Takagi

Algorithmic recourse (AR) aims to provide a recourse action that alters an undesired prediction made by a machine learning model. While standard AR methods assume the model does not change over time, this assumption is often violated in practice due to distribution shifts or data updates, rendering the suggested actions invalid. To address this issue, recent studies have proposed several methods that provide actions robust to model changes. Most of them construct an uncertainty set of models by perturbing continuous model parameters, such as the weight vectors of neural networks, and then seek actions that remain valid for all models in the set. However, such an approach cannot be directly applied to tree ensembles because their model parameters include discrete tree structures, which cannot be perturbed in the same way as neural networks. In this paper, we introduce a new robust AR framework for tree ensembles through the lens of distributionally robust optimization (DRO). Our key idea is to express the model change in a tree ensemble as a perturbation to the distribution of predictions made by decision trees in the ensemble, rather than to the model parameters themselves. This formulation allows us to naturally cast the robust AR problem for tree ensembles as a DRO problem. We then show that, by choosing a specific discrepancy measure between distributions, our robust validity constraint can be reformulated into a tractable one that can be evaluated as efficiently as the standard validity constraint. Building on this reformulation, we propose an efficient optimization algorithm by extending the existing search-based method for tree ensembles. Experimental results demonstrated that our method achieved a better robustness-cost trade-off than existing robust methods applicable to tree ensembles, while maintaining comparable runtime to the standard method.


Distribution Corrected Decision Transformer for Offline Reinforcement Learning with Imbalanced Datasets

Chufan Chen ⋅ Bohao Tang ⋅ Dejiang Chen ⋅ zhang chao ⋅ Ruiqi Li ⋅ Hanbin Zhao ⋅ Hui Qian

Decision Transformer (DT) has emerged as a powerful paradigm for offline reinforcement learning. For complex tasks with imbalanced datasets, where high-quality samples constitute only a small fraction of the data while the majority is suboptimal, the performance of DT methods usually degrade significantly due to the the offline data distribution deviates substantially from the stationary distribution of the optimal policy. In this paper, we construct a dual-based correction mechanism to mitigate the distribution shift problem and propose a Distribution Corrected Decision Transformer method. Specifically, we investigate the Lagrangian duality of the reward maximizing objective in the standard Markov decision problem and recover the optimal correction weights between the static offline data and optimal stationary distribution. Note that a double-KL regularizer is integrated in the objective to prioritize high-return state-action pairs while remaining anchored to valid data support. Subsequently, we use the learned weights to correct distributional bias in two critical components of DT: stabilizing policy evaluation through weighted temporal-difference learning, and improving policy extraction by prioritizing DT training on weighted dataset aligned with the corrected target distribution. We theoretically analyze the performance guarantee, proving our method improves upon standard baselines, and empirically demonstrate strong performance on D4RL benchmarks, particularly on highly imbalanced datasets where prior methods fail.


Diving-R1: Empowering Multimodal LLMs with Traceable Progressive Reasoning for Interpretable Diving Action Quality Assessment

Kaishen Yuan ⋅ Yuting Zhang ⋅ Bohao Xing ⋅ Wenshuo Chen ⋅ Yang Yang ⋅ Shaofeng Liang ⋅ Zitong YU ⋅ Yutao Yue

Action Quality Assessment (AQA) for diving aims to quantify the execution quality of diving sequences and has long been a prominent research topic. However, existing diving AQA methods typically regress solely the final diving score, leading to poor interpretability. To address this challenge, we propose Diving-R1, the first attempt to leverage a Multimodal Large Language Model (MLLM) for interpretable diving AQA, enabling the integration of comprehensive information, thereby supporting progressive reasoning and traceable assessment. We construct a novel dataset, DivingThink, as the foundation, which contains progressively structured reasoning chains that follow judging logic and explicitly describe the execution quality of each diving sub-action, providing clear evidence for score determination. Furthermore, we establish a new benchmark, DivingInterp, which evaluates the performance of MLLMs on interpretable diving AQA from multiple complementary perspectives, including score estimation, textual similarity, and reasoning quality. Additionally, we carefully design a three-stage training paradigm, combined with a dedicated composite reward function, which gradually relaxes the annotation requirements for the training corpus at each stage to incrementally enrich the diversity of the training data, achieving continuous improvement. Extensive experiments demonstrate the superiority of Diving-R1 over existing MLLMs in interpretable diving AQA. The dataset and code will be available on GitHub.


DoAtlas-1: A Causal Compilation Paradigm for Clinical AI

Yulong Li ⋅ Jianxu Chen ⋅ Xiwei Liu ⋅ Chuanyue Suo ⋅ Rong Xia ⋅ Yichen Li ⋅ Xinlin Zhuang ⋅ Niranjana A Menon ⋅ Jionglong Su ⋅ Eran Segal ⋅ Imran Razzak

Medical foundation models generate narrative explanations but cannot quantify intervention effects, detect evidence conflicts, or validate literature claims, limiting clinical auditability. We propose causal compilation, a paradigm that transforms medical evidence from narrative text into executable code. The paradigm standardizes heterogeneous research evidence into structured estimand objects, each explicitly specifying intervention contrast, effect scale, time horizon, and target population, supporting six executable causal queries: do-calculus, counterfactual reasoning, temporal trajectories, heterogeneous effects, mechanistic decomposition, and joint interventions. We instantiate this paradigm in DoAtlas-1, compiling 1,445 effect kernels from 754 studies through effect standardization, conflict-aware graph construction, and real-world validation (Human Phenotype Project, 10,000 participants). The system achieves 98.5% canonicalization accuracy, 100% conflict detection recall, and 80.5% query executability. This paradigm shifts medical AI from text generation to executable, auditable, and verifiable causal reasoning.

Minority collapse, where minority classes become indistinguishable, is a significant challenge in imbalanced learning. This challenge is addressed by methods such as Mixup with class-balanced sampling. Although minority collapse has been mathematically analyzed using the layer-peeled model alongside Neural Collapse, no prior work has analyzed minority collapse under Mixup, particularly from the perspective of mixed labels. We investigate this overlooked factor and raise the question: Is mixed label balance important for alleviating minority collapse? Our analysis reveals that (i) mixed labels should be balanced, and (ii) in this setting, interpreting mixed labels as singletons is beneficial. Motivated by this analysis, we propose a Balanced Mixed Label Sampler and a Mixed-Singleton classifier, which balance mixed labels and treat them as singleton labels. Through theoretical analysis and experimental results, we highlight the importance of balancing mixed labels in imbalanced learning.

Policy gradient computes a backward pass for every sample, even though the backward pass is expensive and most samples carry little learning value. The Delightful Policy Gradient (DG) provides a forward-pass signal of learning value: \emph{delight}, the product of advantage and surprisal (negative log-probability). We introduce the \emph{Kondo gate}, which compares delight against a compute price and pays for a backward pass only when the sample is worth it, thereby tracing a quality--cost Pareto frontier. In bandits, zero-price gating preserves useful gradient signal while removing perpendicular noise, and delight is a more reliable screening signal than additive combinations of value and surprise. On MNIST and transformer token reversal, the Kondo gate skips most backward passes while retaining nearly all of DG's learning quality, with gains that grow as problems get harder and backward passes become more expensive. Because the gate tolerates approximate delight, a cheap forward pass can screen samples before expensive backpropagation, suggesting a speculative-decoding-for-training paradigm.

Classification-trained sequential glimpse policies have attained competitive performance and efficiency on large-scale recognition benchmarks, and are motivated by analogy to the saccadic architecture of biological active vision. Yet whether these policies reproduce the behavioral priors that organize human free viewing, or merely approximate their spatial distribution as a byproduct of classification reward, has not been systematically examined. We present a diagnostic behavioral audit of AdaptiveNN across three free-viewing datasets, evaluating five dimensions associated with known properties of biological gaze: spatial priors, face prioritization, local return, saccade amplitude structure, and semantic guidance. We show that AdaptiveNN exhibits: (i) spatial behavior that falls below a simple center-bias baseline; (ii) scanpath geometry unlike human scanpaths in amplitude and sequential structure; (iii) revisitation dynamics inconsistent with human free viewing and not reducible to step-size geometry; and (iv) selectively attenuated semantic guidance, with face-specific prioritization center-explained, social sampling attenuated, and object-level rates preserved. In AdaptiveNN, this dissociation suggests that classification reward alone cannot reproduce several behavioral priors organizing human free viewing, even when object-level sampling is preserved. These findings identify which behavioral priors require constraints beyond classification utility, offering targets for designing cognitively aligned glimpse policies.


Do Image Editing Models Understand Lighting?

Tim Küchler ⋅ Johann-Friedrich Feiden ⋅ Matthias Niessner ⋅ Carsten Rother

While recent advancements in generative image editing models have achieved stunning visual fidelity, it remains an open question whether these systems possess an intrinsic knowledge of real-world lighting. Existing benchmarks typically evaluate high-level plausibility of perceptual light transport on curated internet imagery, using VLMs or human judgement, or they rely on synthetically generated datasets. In this work, we introduce the 3D-anchored Light Probe (3DLP) benchmark, for which we have captured a new high-fidelity HDR dataset of real-world lighting changes. The dataset consists of 1K image pairs of diverse indoor scenery in which light probes are physically turned on and off. To allow for a granular performance analysis, we annotated specific image regions such as cast shadows or metallic surfaces. With this data, we evaluate a range of state-of-the-art image editing models by measuring how well their light probe edits align with reality. The evaluation uses two new scores to compensate for AI-generated photographic effects, such as adjusted white balance. Our results show that the overall performance of models differs considerably, with differences slightly less pronounced for specular highlights. The best image editing models are remarkably consistent with real-world physics, however, they still leave room for improvement. We observe that image regions that receive less light from the light probe are more prone to errors for all models. Furthermore, building on their success in evaluating macroscopic lighting plausibility, we test VLMs on our task but find that they are unsuitable for pixel-level light transport analysis. We will make the benchmark, together with the real-world dataset, publicly available to encourage future research on this topic.


Don’t Let Gains FADE: Breaking Down Policy Gradient Weights in RL

Juliette Decugis ⋅ Sean O'Brien ⋅ Francis Bach ⋅ Gabriel Synnaeve ⋅ Taco Cohen

Reinforcement-learning-based post-training for large language models vastly improves capabilities on verifiable tasks however it suffers from training instability and diversity collapse. As a solution, many new advantage functions have emerged. Each method simultaneously changes which problems receive gradient, the balance between positive and negative updates, and the overall gradient scale making their comparison difficult. We propose a unifying framework that decomposes each function by their induced positive and negative gradient mass. Analyzing policy performance and the geometry of model weights reveals: symmetric methods are better but slow and asymmetric methods are fast. Combining the advantages of both, we build the FADE: Focal Adaptive Dynamic Entropy advantage which is fast, diverse and accurate at various model scales (7B, 32B).

Real-world machine learning systems often operate on imperfect measurements: sensors are noisy, annotations are produced by non-experts, and scientific observations are often indirect. In these settings, classic distribution comparison tools such as maximum mean discrepancy (MMD) can be systematically biased, matching artifacts of the measurement process rather than the underlying mechanisms of interest. We study a practical regime in which noisy observations are abundant, but only a small audited subset contains ground-truth measurements. We propose measurement-error-corrected MMD (MEC-MMD), an audit-based estimator that debiases kernel evaluations without requiring a noise model. MEC-MMD is exactly unbiased for the oracle MMD under simple random auditing, admits an explicit variance decomposition, and is consistent for the population MMD, achieving the standard $O_p(N^{-1/2})$ rate when the audit size scales with the dataset. We demonstrate the utility of distributional matching with MEC-MMD loss for both parameter estimation and unsupervised applications, showing that MEC-MMD consistently outperforms MMD and alternative methods. These results provide a principled foundation for robust distribution comparison, simulation-based inference, and learning in noisy real-world machine learning pipelines.


Don't Pause! Every prediction matters in a streaming video

Dibyadip Chatterjee ⋅ Zhanzhong Pang ⋅ Fadime Sener ⋅ Yale Song ⋅ Angela Yao

Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause the video at fixed timestamps, pose questions about current or past events, and score models only at those moments. This protocol leaves streaming outputs untested. To close this gap, we introduce SPOT-Bench, featuring multi-turn proactive queries that evaluate general streaming perception required by an always-on, real-time assistant. SPOT-Bench comes with Timeliness-F1, a consolidated metric that measures streaming outputs by their temporal precision and balanced coverage across the entire video. Our benchmark reveals that during streaming inference: (i) MLLMs detect events reliably in a stream but spam predictions unprompted; (ii) post-training MLLMs for silence reduces spamming but induces unresponsiveness; (iii) half of the streaming video expects no response, which we term dead-time - compute spent here does not affect response latency. These findings motivate Ctrl-SPOT, a training-free streaming controller for MLLMs, that retains their event perception while controlling their streaming outputs for balanced coverage. This establishes a strong baseline on SPOT-Bench.


Don't Pay Attention, PLANT It: Pretraining Attention via Learning-to-Rank

Debjyoti Saha Roy ⋅ Byron Wallace ⋅ Javed A Aslam

State-of-the-art Extreme Multi-Label Text Classification (XMC) models rely on multi-label attention to focus on key tokens in input text, but learning high-quality attention weights is challenging. We introduce PLANT (Pretrained and Leveraged Attention), a plug-and-play strategy for initializing attention. PLANT works by \emph{planting} label-specific attention using a pretrained Learning-to-Rank model guided by mutual information gain. This architecture-agnostic approach integrates seamlessly with large language model backbones such as Mistral, LLaMA, DeepSeek, and Phi-3. PLANT outperforms state-of-the-art methods across tasks such as ICD coding, legal topic classification, and content recommendation. Gains are especially pronounced in few-shot settings, with substantial improvements on rare labels. Ablation studies confirm that attention initialization is a key driver of these gains. Code and trained models are available at \url{https://github.com/research-anon-487/xcube/tree/plant}.


Do Speech BCIs Need Larger Models? Rethinking Neural Decoding beyond Scaling

Zehui Feng ⋅ Weichuan Wang ⋅ Xiaohan Chen ⋅ Cuntai Guan ⋅ Ting Han

Recent speech brain-computer interfaces (BCIs) increasingly rely on large models and complex multi-stage pipelines for neural decoding. However, in speech BCI, where training data are scarce, noisy, and expensive to collect, the benefits of model scaling remain unclear. This work revisits a fundamental question: $\textbf{are large models truly necessary for accurate neural decoding?}$ We propose $\textbf{\textit{BEST}}$ ($\textbf{\textit{B}}$rain-$\textbf{\textit{E}}$ncoding-$\textbf{\textit{S}}$peech-$\textbf{\textit{T}}$ext), a two-stage framework for efficient neural-to-text decoding. Instead of relying on large generative models or phoneme-centric pipelines, $\textbf{\textit{BEST}}$ introduces hierarchical intermediate representations and a compact speech-oriented decoder that directly maps neural activity to text. This design preserves essential structure while significantly reducing computational cost, enabling robust generalization in low-resource settings and practical deployment under constrained resources. On the Brain-to-Text ’24 and ’25 benchmarks, $\textbf{\textit{BEST}}$ achieves word error rates (WERs) of 6.19\% and 1.62\%, espectively, using only 3.7\% of the parameters required by comparable methods. Beyond these results, we provide three key insights: (1) decoder scaling yields non-monotonic gains; (2) cross-modal alignment is more critical than model size; and (3) overfitting degards sequence-level performance. Representation analysis further reveals a consistent neural→audio→text alignment pathway.


Drifting Fields are not Conservative

Leonard T. Franz ⋅ Sebastian Hoffmann ⋅ Tim Weiland ⋅ Bernhard Schölkopf ⋅ Georg Martius

Drifting models have recently gained attention for generating high-quality samples in a single forward pass. During training, they learn a push-forward map by following a vector-valued field, the *drift field*. We ask whether this procedure is equivalent to optimizing a scalar loss and find that, in general, it is not: drift fields are *not conservative* and cannot be written as the gradient of any scalar potential. We identify the position-dependent normalization as the source of non-conservatism, with the Gaussian kernel as the unique radial exception. Guided by this, we introduce the *sharp kernel* $k^\\#$ and a sharp-normalized drift field that is conservative for general radial kernels. The resulting vector field is the gradient of a scalar potential that can be optimized directly using stochastic gradient descent. Moreover, the field has the form of a score difference of kernel density estimates, and gives exact equilibrium identifiability. Thus, sharp normalization closes the gap to related literature, such as Wasserstein gradient-flows and denoising score matching, also for non-Gaussian kernels. Empirically, sharp normalization preserves the performance of the original drifting objective, suggesting that the non-conservative flexibility is not required for high-quality generation.


DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning

Yang Zhou ⋅ Xiaofeng Wang ⋅ Hao Shao ⋅ Letian Wang ⋅ Guosheng Zhao ⋅ Jiangnan Shao ⋅ Jiagang Zhu ⋅ Tingdong Yu ⋅ Zheng Zhu ⋅ Guan Huang ⋅ Steven Waslander

Recently, world-action models (WAM) have emerged to bridge vision-language-action (VLA) models and world models, unifying their reasoning and instruction-following capabilities and spatio-temporal world modeling. However, existing WAM approaches often focus on modeling 2D appearance or latent representations, with limited geometric grounding, an essential element for embodied systems operating in the physical world. We present DriveDreamer-Policy, a unified driving world-action model that integrates depth generation, future video generation, and motion planning within a single modular architecture. The model employs a large language model to process language instructions, multi-view images, and actions, followed by three lightweight generators that produce depth, future video, and actions. By learning a geometry-aware world representation and using it to guide both future prediction and planning within a unified framework, the proposed model produces more coherent imagined futures and more informed driving actions, while maintaining modularity and controllable cost. Experiments on the Navsim v1 and v2 benchmarks demonstrate that DriveDreamer-Policy achieves strong performance on both planning and world generation tasks. In particular, our model reaches 89.2 PDMS on Navsim v1 and 88.7 EPDMS on Navsim v2, outperforming existing world-model-based approaches while producing higher-quality future video and depth predictions. Ablation studies further show that depth learning provides complementary benefits to video imagination and improves planning robustness. Our code is available to facilitate further research.


Dual-Contrastive Sparse Autoencoders Reveal Features of Musical Interpretation

Jan Chen ⋅ Manuel Cherep ⋅ Patricia Maes ⋅ Nikhil Singh

For decades, philosophers and musicologists have debated which features of a performance are constitutive of the work and which are expressions of how it is being played. Computational evidence has been hard to assemble: audio embedding models conflate work identity with performance style, and existing interpretability tools for music generators treat learned representations as flat dictionaries that mix the two. We turn the question into an empirical one by introducing the dual-contrastive sparse autoencoder (DC-SAE), a two-branch sparse autoencoder that uses cheap work-level metadata to factor a generative music transformer's residual stream into work-identity and performance-content subspaces. Across jazz-standard, classical-work, and pop-cover corpora, the resulting decomposition is supported by probes and feature galleries that surface musically interpretable concepts on each side. Without performer supervision, the variation branch acquires structure related to performer identity, and steering along these directions can shift the perceived performer of generated audio while preserving the underlying work. Together, these results show that a generative music model has internalized a representation of interpretation that can be recovered, named, and steered.


Dual-Pathway Circuits of Object Hallucination in Vision-Language Models

Jiaxin Liu ⋅ Ding Zhong ⋅ Yue Wang ⋅ Zhidong Yang ⋅ Zhaolu Kang ⋅ Guangyuan Dong ⋅ Qishi Zhan ⋅ Pengcheng Fang ⋅ Aofan Liu

Vision-language models (VLMs) have demonstrated remarkable capabilities in bridging visual perception and natural language understanding, enabling a wide range of multimodal reasoning tasks. However, they often produce object hallucinations, describing content absent from the input image, which limits their reliability and interpretability. To address this limitation, we propose Dual-Pathway Circuit Analysis, a framework that identifies and characterizes hallucination-related circuits in VLMs for mechanistic understanding and causal probing. We first apply activation patching across five architecturally diverse VLMs to identify a visual grounding pathway that supports correct predictions and a hallucination pathway that drives erroneous outputs. We then introduce Conditional Pathway Analysis (CPA) to characterize pathway-level interactions, revealing that grounding components remain strongly redundant in both correct and hallucinating samples but undergo a consistent polarity flip, shifting from supporting the ground truth on correct samples to aligning with the hallucinated answer on erroneous ones. We further perform targeted suppression of hallucination-pathway components, showing that scaling these components reduces object hallucination by up to 76\% with minimal accuracy cost, and validate that the same circuit selectively transfers to relational but not attribute hallucination. Evaluations on POPE-adversarial and AMBER show that the identified circuits are consistent across architectures, support causal intervention, and transfer selectively across hallucination types.


DualWorldBench: Can Agents Plan Deliveries across Symbolic and Grounded Worlds?

Yang Chen ⋅ Bin Wen ⋅ Danyang Peng ⋅ Hong-Jie You ⋅ Ziqiao Shang ⋅ Lan-Zhe Guo

Recent multimodal agents have shown promising progress in embodied planning and web interaction. However, existing evaluations often rely on simplified tasks with limited combinatorial structure, leaving it unclear whether current agents can robustly integrate perception and reasoning for multi-goal planning under temporal, physical, and multimodal constraints. We introduce DualWorldBench, a unified benchmark for evaluating agent planning under controlled symbolic--temporal and grounded--topological complexity. Built around a shared pickup--delivery interface, DualWorldBench consists of three complementary components: TextWorld, which isolates symbolic planning under varying task compositions and release dynamics; SimWorld, which evaluates grounded path planning under controlled topological variation with text, image, or multimodal inputs; and DualWorld-VQA, a diagnostic suite spanning perception, scene understanding, and planning. Experiments show that current state-of-the-art agents remain brittle under increasing combinatorial complexity. Multimodal input does not consistently improve planning and can even degrade performance, revealing difficulties in aligning symbolic task specifications with grounded topology. Diagnostics further show that models handle perception or understanding in isolation, but fail when decision-making requires integrating both. These findings establish DualWorldBench as a challenging testbed for agents with tighter perception--reasoning--planning integration. All datasets and code are publicly available at our project website https://dualworld.netlify.app/.


DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards

Haoyu Hu ⋅ Xuandong Zhao ⋅ Nori Jacoby ⋅ Xuhai "Orson" Xu

Reinforcement learning with verifiable rewards (RLVR) generates thousands of tokens per training step, with rollout generation dominating the computational cost. The overall token budget can be controlled along two main dimensions: (i) deciding which prompts to allocate rollouts to, and (ii) deciding how long each rollout should be. Prior work has generally controlled only one of these dimensions at a time. We show that jointly tuning both decisions under a shared compute budget improves both reasoning quality and wall-clock training time. We instantiate this view as \textbf{DU}al-controlled tok\textbf{E}n alloca\textbf{T}ion (DUET), a computationally efficient layer over GRPO, in which we have a lightweight pre-rollout surrogate of prompt informativeness to set how many rollouts each prompt receives and a marker-gated abort rule with importance reweighting set when to stop them. On Qwen3-1.7B trained on MATH, DUET outperforms full-budget GRPO and the other three budget-aware baseline methods. DUET's advantage was further generalized to other benchmarks across math and coding, and was on par with the best baseline on the scientific Q\&A domain, while also achieving a $1.62\times$ wall-clock speedup. More notably, using only 50\% of the token budget, DUET still outperforms all baseline methods, and performance often increases rather than reduces with the more limited budget, achieving an even higher $2.51\times$ speedup. We verified the high performance of DUET on other backbone LLMs, including Qwen-4B and Llama-3.2-3B-Instruct. Notably, the gap between DUET and the strongest baseline \emph{widens} as the budget tightens, contrary to the usual pattern in which efficient methods trade off quality as compute decreases. More broadly, these results suggest that DUET budget-aware control strategies are valuable not only for accelerating training, but also for improving the quality of the learning signal.


DynaCell: an Evaluation Framework for Dynamic 3D Virtual Staining of Live Cells

Alexandr Kalinin ⋅ Dihan Zheng ⋅ Taylla Milena Theodoro ⋅ Ivan Ivanov ⋅ Eduardo Hirata Miyasaki ⋅ See-Chi Lee ⋅ Aofei Liu ⋅ Sricharan Reddy Varra ⋅ Talon Chandler ⋅ Soorya Pradeep ⋅ Chad Liu ⋅ Manuel D Leonetti ⋅ Carolina Arias ⋅ Bo Huang ⋅ Shalin Mehta

Imaging of dynamic cellular responses to perturbations requires live-cell measurements, posing a challenge for fluorescence-based assays that can visualize molecular phenotypes by specific labeling but often affect cell health. Virtual staining helps by predicting fluorescence-like signals from label-free microscopy. Although many models have been developed for this task, their performance on volumetric time-lapse data remains poorly characterized. We address this gap by introducing DynaCell, an evaluation framework for 3D time-lapse virtual staining of live cells, comprising a new paired label-free and fluorescence 3D time-lapse imaging dataset, a suite of baseline models, and a three-tier metric panel measuring pixel-level fidelity, organelle segmentation performance, and single-cell phenotypic similarity. Using DynaCell, we study model behavior across 2 cell types and microscopes, 4 organelles, and 3 perturbation states. We find that regression baselines better predict the spatial localization of target organelles and are more robust to label-free input distribution shifts, while the generative baseline better captures population-level phenotype distribution across cells and organelles. These differences expose trade-offs between preservation of structural and phenotypic fidelity that are obscured by single-metric evaluation, and suggest how virtual staining can support downstream tasks such as localization, segmentation, tracking, and phenotypic profiling. DynaCell provides reusable data, code, checkpoints, and documentation for assessing virtual stains for live-cell biological measurements.


Dynamic Latent Routing

Fangyuan Yu ⋅ Xin Su ⋅ Amir Abdullah

We investigate the temporal concatenation of sub-policies in Markov Decision Processes (MDP) with time-varying reward functions. We introduce General Dijkstra Search (GDS), and prove that it discovers optimal goal-reaching policies by concatenating sub-policies optimal for intermediate goals. Motivated by the ``search, select, update'' principle underlying GDS, we propose Dynamic Latent Routing (DLR), a language-model post-training method that jointly learns discrete latent codes, routing policies, and model parameters through dynamic search in a single training stage. In low-data fine-tuning settings, DLR matches or outperforms supervised fine-tuning across four datasets and six models, achieving a mean gain of $+6.6$ percentage points, while prior discrete-latent baselines consistently underperform SFT. Mechanistic analyses and targeted code ablations show that DLR learns structured routing behaviors with distinct causal roles. In six-digit arithmetic, a canonical mechanistic interpretability benchmark, DLR externalizes into discrete routing tokens algorithmic subtasks previously studied through activation and circuit analysis.

Large language models are remarkably capable, yet how computation propagates through their layers remains poorly understood. A growing line of work treats depth as discrete time and the residual stream as a dynamical system, where each layer's nonlinear update has a local linear description. However, previous analyses have relied on scalar summaries or approximate linearizations, leaving the full spectral geometry of trained LLMs unknown. We perform full Jacobian eigendecomposition across three production--scale LLMs and show that training installs a monotonic spectral gradient through depth---from non-normal, rotation-dominated early layers to near--symmetric late layers---together with a cumulative low-rank bottleneck that funnels perturbations into a small fraction of the residual stream's effective dimensions. Our experiments reveal that this gradient and the dimensional collapse are learned rather than architectural, and is largely dissolved when structured non-normality is removed. We further show that the topological positioning of graph communities predicts whether the Jacobian amplifies or suppresses them, with the sign of the coupling determined by the local operator type, a relationship absent at initialization. These results map a learned spectral geometry in LLMs that links perturbation propagation and compression to the network's functional topology.


Early Failure Detection and Intervention in Video Diffusion Models

Byungki Kwon ⋅ Sohwi Lim ⋅ Nam Hyeon-Woo ⋅ Moon Ye-Bin ⋅ Tae-Hyun Oh

Text-to-video (T2V) diffusion models have rapidly advanced, yet generations still occasionally fail in practice, such as low text--video alignment or low perceptual quality. Since diffusion sampling is non-deterministic, it is difficult to know during inference whether a generation will succeed or fail, incurring high computational cost due to trial-and-error regeneration. To address this, we propose an early failure detection and diagnostic intervention pipeline for latent T2V diffusion models. For detection, we design a Real-time Inspection (RI) module that converts latents into intermediate video previews, enabling the use of established text–video alignment scorers for inspection in the RGB space. The RI module completes the conversion and inspection process in just 39.2 ms. This is significantly efficient considering that CogVideoX-5B requires 4.3s per denoising step when generating a 480p, 49-frame video on an NVIDIA A100 GPU. Subsequently, we trigger a hierarchical and early-exit intervention pipeline only when failure is predicted. Experiments on CogVideoX-5B and Wan2.1-1.3B demonstrate consistency gains on VBench with up to 2.64x less time overhead compared to post-hoc regeneration. Furthermore, our pipeline is plug-and-play and orthogonal to existing techniques, showing seamless compatibility with prompt refinement and sampling guidance methods. We also provide evidence that failure signals emerge early in denoising process and are detectable within intermediate video previews using standard vision-language evaluators.


EchoXFlow: A Beamspace Echocardiography Dataset for Cardiac Motion, Flow, and Function

Elias Stenhede ⋅ Joanna Sulkowska ⋅ Eivind B Orstad ⋅ Henrik Schirmer ⋅ Arian Ranjbar

We introduce EchoXFlow, a clinical echocardiography dataset for learning from ultrasound in its native acquisition geometry rather than from scan-converted Cartesian videos. Existing public datasets offer limited opportunities to study cross-modal relationships between cardiac anatomy, myocardial motion, and blood flow, as Doppler is typically absent or fused as RGB overlays, and acquisitions are released after lossy vendor display processing. EchoXFlow comprises 37125 recordings from 666 routine-care examinations, preserving the timing, geometry, and modality relationships needed for physically grounded echo learning. Each recording is retained as separable modality-specific streams: temporally resolved 1D, 2D, and 3D data alongside multiple Doppler modalities, paired with a synchronized ECG. Clinical annotations span guideline-based measurements to dense 2D myocardial contours and 3D left-ventricular endocardial meshes. With its associated open-source tooling, EchoXFlow enables cross-modal, acquisition-aware learning tasks that cannot be formulated from conventional scan-converted videos alone, and serves as a testbed for 4D vision and physically grounded multi-modal learning more broadly.


EEG-X: Toward Device-Agnostic and Noise-Robust Foundation Models for EEG

Navid Mohammadi Foumani ⋅ Soheila Ghane ⋅ Nam Nguyen ⋅ Mahsa Salehi ⋅ Geoffrey Webb ⋅ Geoffrey Mackellar

Foundation models for EEG analysis are still in their infancy, limited by two key challenges: (1) variability across datasets caused by differences in recording devices and electrode configurations, and (2) the low signal-to-noise ratio (SNR) of EEG, where neural activity is often buried under artifacts and non-brain sources. To address these challenges, we present \textit{EEG-X}, a device-agnostic and noise-robust foundation model for EEG representation learning. EEG-X introduces a location-based channel embedding that enables robust transfer across heterogeneous EEG devices and electrode layouts. To improve robustness against noise, EEG-X employs a novel noise-aware masking-reconstruction strategy that reconstructs artifact-cleaned signals rather than raw noisy EEG and introduces a novel Dictionary-based Convolutional Transformation (DiCT) that improves reconstruction-based pretraining by comparing signals in a structured feature space instead of directly in the raw signal space. Experiments across datasets collected from diverse EEG devices show that EEG-X consistently outperforms state-of-the-art methods across multiple downstream tasks and demonstrates strong cross-domain generalization when pretraining and downstream datasets differ in electrode layouts and recording configurations paving the way towards foundation models. The models and code are available at :https://anonymous.4open.science/r/EEG-X-CEDE/README.md


Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

Qianxin Xia ⋅ JIAWEI DU ⋅ Yuhan Zhang ⋅ xin zhang ⋅ Xuewan He ⋅ Wenbo Jiang ⋅ Jielei Wang ⋅ Tao Luo ⋅ Guoming Lu

Dataset distillation seeks to synthesize a highly compact surrogate dataset that achieves performance comparable to the original dataset on downstream tasks. For the scenario where pre-trained self-supervised models serve as priors, traditional $\textit{Linear Gradient Matching}$ optimizes synthetic images by encouraging them to mimic the gradient updates induced by real images on the linear probe or classifier. However, this batch-level formulation requires loading thousands of real images and applying multiple differentiable augmentations to synthetic images at each distillation step, leading to substantial computational and memory overheads. In this paper, we revisit the linear gradient and theoretically derive that it is essentially a $\textit{local}$ relative distribution directed from target class centers toward non-target class centers, which we term ``flow”. This property causes instability and suboptimality, often necessitating expensive multiple augmentations to compensate. To address this, we introduce $\textit{Statistical Flow Matching}$, a stable and efficient supervised learning framework that optimizes synthetic images by aligning $\textit{global}$ statistical flows in the original data. Our approach loads raw statistics only once and performs a single augmentation pass on the synthetic data, achieving performance comparable to or better than the state-of-the-art method with 10× less GPU memory usage and 4× faster distillation time. Moreover, increasing the number of augmentations for our method yields further performance gains while incurring lower additional cost.

Construction-based neural routing solvers, typically composed of an encoder and a decoder, have emerged as a promising approach for solving vehicle routing problems. While recent studies suggest that shifting parameters from the encoder to the decoder enhances performance, most works restrict the decoder size to 1–3M parameters, leaving the effects of scaling largely unexplored. To address this gap, we conduct a systematic study comparing two distinct strategies: scaling depth versus scaling width. We synthesize these strategies to construct a suite of 12 model configurations, spanning a parameter range from 1M to $\sim$150M, and extensively evaluate their scaling behaviors across three critical dimensions: parameter efficiency, data efficiency, and compute efficiency. Our empirical results reveal that parameter count is insufficient to accurately predict the model performance, highlighting the critical and distinct roles of model depth (layer count) and width (embedding dimension). Crucially, we demonstrate that scaling depth yields superior performance gains to scaling width. Based on these findings, we provide and experimentally validate a set of design principles for the efficient allocation of parameters and compute resources to enhance the model performance.


Efficient Hybrid Distillation: Synergizing Score and Adversarial Objectives for One-Step Diffusion

Fei Peng ⋅ Junqiang Wu ⋅ Haoxian Tan ⋅ Jun Zhou ⋅ Jie Hu ⋅ Xiaoming Wei ⋅ Jie Guo ⋅ Xiu Li

Numerous distillation methods have been developed to accelerate diffusion models. Recent research indicates that hybrid strategies—blending different paradigms—yield superior results compared to single-method approaches. However, current works predominantly rely on a simple combination of objectives, leaving the question of how to effectively synergize them largely underexplored. In this paper, we specifically investigate the synergy between Score Distillation and Adversarial Distillation. We reveal that their operating zones are distinct and complementary: Score Distillation is effective in high-noise regimes for capturing the global distribution yet lacks fine-grained supervision in low-noise regions. Conversely, Adversarial Distillation is effective in low-noise regions but is prone to instability and suffers from discriminator failure when applied to high-noise intervals. Leveraging this insight, we propose Efficient Adversarial Score Distillation (EASD), a framework that synergizes these objectives via a regime-adaptive sampling strategy, directing each objective to focus on its effective interval. We further apply this paradigm to DiT-based Flow Matching models. Empirically, our method achieves a promising one-step FID of 1.39 on ImageNet 256$\times$256 using the SiT-XL/2+REPA model. This work provides a concrete guideline for maximizing the potential of hybrid distillation.

Label distribution leakage in federated learning (FL) poses a critical privacy threat, exposing population-level patterns beyond individual data breaches. Building on known correlations between class sample frequencies and classifier weight magnitudes, we demonstrate how adversaries can exploit these relationships to infer label distributions from shared model parameters in FL. To characterize this threat, we propose WLIA (Weight Norm-based Label Distribution Inference Attack), which infers client label distributions by analyzing variations in classifier weight norms across communication rounds. WLIA operates on the classifier layer without requiring auxiliary data, synthetic samples, or meta-classifiers, enabling broader applicability across heterogeneous model architectures. To counter this threat, we introduce RNR (Randomized Weight Norm Regularization), a defense mechanism that disrupts frequency-weight correlations by randomly regularizing weight norms during local training. RNR adds only a penalty term to the cross-entropy loss, making it efficient and easy to integrate. Comprehensive experiments across six datasets demonstrate that WLIA achieves superior inference accuracy compared to existing attacks, while RNR effectively mitigates leakage with minimal impact on utility.


EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts

Minseo Kim ⋅ Minjae Lee ⋅ Seunghyuk Oh ⋅ Kevin Galim ⋅ Donghoon Kim ⋅ Coleman Hooper ⋅ Harman Singh ⋅ Amir Gholami ⋅ HYUNG IL KOO ⋅ Wonjun Kang

Reinforcement learning (RL) has become a representative post-training paradigm for large language models (LLMs), enabling strong reasoning and agentic capabilities. However, its rollout generation remains a dominant training bottleneck because it relies on sequential autoregressive (AR) decoding, where a small number of long-tailed responses often determine completion time. Speculative decoding (SD) can reduce inference latency while preserving model quality by rapidly drafting tokens and accepting them through parallel verification. Applying SD to RL rollouts, however, introduces challenges absent from standard LLM inference: (i) algorithmically, the continuously evolving target model makes static drafters stale; (ii) system-wise, rollout decoding moves across regimes, from large active batches where SD can be compute-bound and ineffective to shrinking-batch tails where SD becomes beneficial. Existing RL-SD approaches address aspects of this problem, but either yield low effective accepted lengths that limit speedup or rely on auxiliary drafters requiring pretraining and online adaptation, increasing system complexity. We present EfficientRollout, an SD framework designed to accelerate RL rollouts while addressing these challenges. EfficientRollout induces a quantized drafter directly from the target model, keeping it coupled to the evolving policy without separate drafter training before or during RL. It then uses a dynamic SD toggle policy that enables SD only in beneficial regimes identified by system-aware roofline modeling. It further adapts drafting budgets using acceptance-behavior signals observed during training, better realizing the potential accepted length. Under realistic RL workload, EfficientRollout reduces rollout and end-to-end latency by up to 19.6% and 12.7%, respectively, over a standard accelerated AR rollout baseline, while preserving final model quality.


EgoForce: Robust Online Egocentric Motion Reconstruction via Diffusion Forcing

Inwoo Hwang ⋅ Donggeun Lim ⋅ Hojun Jang ⋅ Young Min Kim

With recent advances in embodied agents and AR devices, egocentric observations are readily available as input for real-world interactive online applications. However, egocentric viewpoints can only sporadically observe hands, in addition to the estimated head trajectory. We propose EgoForce, an online framework for reconstructing long-term full-body motion from noisy egocentric input. While existing generative frameworks can robustly handle noisy and sparse measurements, they assume a fixed-length observation window is available and are thus not suitable for real-time applications. Faster inference often relies on autoregressive prediction, sacrificing robustness. In contrast, we adopt a diffusion-based method with a temporally asymmetric noise schedule inspired by Diffusion Forcing. Specifically, our approach models temporally evolving uncertainty and incrementally denoises states as new streaming observations arrive. Combined with a noise-robust imputation strategy, EgoForce progressively generates stable and coherent full-body motion under strict causal constraints. Experiments demonstrate that our online framework outperforms existing online and offline methods, enabling long-horizon, full-body motion reconstruction in challenging egocentric scenarios.


EIPO: Efficient Inter-Step Parallel Optimization

Jianrong Lu ⋅ Zhuoya Gu ⋅ Zhiyu Zhu ⋅ Hui LIU ⋅ Junhui Hou

The wall-clock cost of training large neural networks is dominated by the long horizon of sequential optimizer iterations. We recast the gradient-descent (GD) trajectory of modern optimizer (e.g., SGD and Adam) as the unique root of a nonlinear operator, enabling a parallel root-finding approach that updates multiple GD steps simultaneously. We then propose \textbf{E}fficient \textbf{I}nter-step \textbf{P}arallel \textbf{O}ptimization (\textbf{EIPO}) for accelerating a wide range of optimizers. EIPO contributes three pieces: (i) a Preconditioned Nonlinear-Equations (PNE) formulation that recovers the exact sequential GD trajectory; (ii) Memory-Efficient Adaptive Anderson Acceleration, which extracts the multisecant geometry directly from the optimizer's intrinsic momentum buffers and therefore avoids the memory overhead of classical Anderson Acceleration. (iii) We prove that EIPO converges to the optimizer's GD trajectory within fewer iterations than its sequential counterpart. (iv) Across language modeling on WikiText-2 (GPT-2, Llama-3.2-1B, Qwen-1.5-4B/2.5-3B, Gemma-2B, GPT-J-\textbf{6B}), image classification on CIFAR-10 (ResNet50, ViT, CNN), and diffusion training on LSUN Church (UNet/DDPM/DDIM), EIPO reduces optimizer iterations by up to $\textbf{21}\times$ and wall-clock time by up to $\textbf{4.6}\times$ while matching baseline perplexity, accuracy, and FID, all under comparable GPU memory and token consumption. Source code is available at: \url{https://anonymous.4open.science/r/EIPO-44BC}.

Inserting objects into existing 3D scenes requires more than selecting a plausible location: The inserted object must also fit local geometry while preserving semantic intent and physical plausibility. Although recent Vision-Language Models (VLMs) and generative models enable semantic reasoning and visual content creation, they offer limited 3D grounding and geometric control when an inserted object must fit into constrained local spaces. We introduce ElasticFit, a VLM-guided framework for fit-aware object insertion centered on a novel scene-grounded representation. Given a language instruction and rendered scene observations, ElasticFit infers structured fitting cues that specify where the object should be grounded, what volume it should occupy, how it should be oriented, and its adaptation mode (rigid placement, uniform scaling, or elastic deformation). These cues convert high-level VLM reasoning into explicit 3D constraints that condition object generation and guide downstream geometric fitting. ElasticFit then generates a scene-conditioned object prior, reconstructs it in 3D, and refines the mesh through mode-specific fitting while enforcing collision avoidance, contact consistency, and physical grounding. In fixed-asset baseline comparisons, ElasticFit improves spatial relation success from 50.8\% to 69.7\% and support success from 48.3\% to 91.7\% over the strongest baseline, while providing novel support for generative "make-it-fit" insertions in complex scenarios.

A single global fairness criterion cannot serve communities whose conceptions of fair treatment differ. We introduce \emph{elicited localized fairness}: an end-to-end pipeline that elicits per-community Mahalanobis metric--tolerance specifications $\cF_k = (d_k, \eps_k)$ from stakeholders via two phases of pairwise queries (similarity, then tolerance acceptability), trains a localized classifier (\textsc{LocAdapt}), and audits the deployed model on a held-out tolerance fold. \emph{Theory.} A constrained-MLE rate $\tilde O(\sqrt{d^2/n})$ for metric elicitation and a matching $\Omega(d^2/t^2)$ active-query lower bound, exact in $(d, n, t)$ exponents; a held-out audit certificate that converts Phase 2 queries into a PAC-style violation guarantee. \emph{Empirics.} On heterogeneous-community settings (Stress $K{=}5$, Adult-K3-human; 10 seeds), \textsc{LocAdapt} sits on the (accuracy, strict $d{\leq}1$ violation) Pareto frontier: averaged and matched-accuracy Global IF lose on strict at comparable accuracy ($40$--$80\%$ higher), and pooled-MLE Global IF matches strict only at $-2$pp accuracy. A five-axis diagnostic (acc, strict, score-diff spread, prediction entropy, AUC) confirms the gain is not prediction-smoothing. Plugging Phase 2 elicited tolerances into \textsc{LocAdapt} training reduces the COMPAS deployed-spec violation $66\%$ ($p{\ll}10^{-4}$). Five pre-registered human studies on COMPAS (360 annotators total) validate the pipeline end-to-end; a 180-annotator scorer-stability ablation across three scorer architectures preserves the cross-community $\hat\eps$ ordering and pools to $p{=}0.0014$ ($n{=}90$ vs.\ $90$).


ELMA: Benchmarking Anaphoric Compositions for Long-Term Text-to-Motion Generation

Shun Kato ⋅ Dixin Yang ⋅ Kotaro Amaya ⋅ Yuiga Wada ⋅ Kazumi Fukuda ⋅ Takashi Shibuya ⋅ Yuki Mitsufuji ⋅ Mariko Isogawa

Text-guided long-term human motion generation is essential for applications such as continuously operating digital avatars. Despite recent progress on short clips, existing models struggle with multi-segment instructions that involve anaphora, where constraints established in earlier segments must be preserved over time. In this paper, we introduce a novel task of long-term motion generation from such anaphoric compositions. To mitigate this, we present ELMA, an episodic long-term motion dataset with dense segment-level annotations that preserve cross-segment dependencies. Built via a fully automated collection and filtering pipeline, ELMA extracts high-quality, continuous 30-second clips from in-the-wild videos, capturing the episodic narratives necessary for long-range state modeling. Leveraging this dataset, we fine-tune state-of-the-art motion generation models to incorporate long-range context. Extensive experiments reveal fundamental shortcomings in current models when handling context-dependent instructions. We demonstrate that ELMA enables significant improvements in both global physical realism and semantic alignment, setting a robust new baseline for long-term motion synthesis.


Embedded-Arena: Building Hardware-in-the-Loop Coding Agents to Run AI on Microcontrollers

Zhihan Zhang ⋅ Alexander Le Metzger ⋅ Jiuyang Lyu ⋅ Chun-Cheng Chang ⋅ Jiayi Shao ⋅ Yujia Liu ⋅ Emmanuel A Mensah ⋅ Edward Wang ⋅ Kurtis Heimerl ⋅ Gregory D Abowd ⋅ Natasha Jaques ⋅ Shwetak Patel ⋅ Vikram Iyer

Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or privacy constraints. Optimizing models for heterogeneous microcontrollers (MCUs) requires simultaneously satisfying hard physical constraints on memory, power, and temperature while preserving accuracy, a multidimensional optimization that is today performed manually by experts. We ask whether an LLM agent can autonomously navigate this complex, multi-turn pipeline guided by real hardware feedback, and introduce a hardware-in-the-loop agent arena in which the agent iteratively refines both model and firmware---compiling, flashing, and measuring on real hardware---to enable iterative optimization. Frontier models, including Claude Opus 4.7 and Gemini 3.1 Pro, fail entirely without hardware feedback (0\% deployment success), whereas our closed-loop formulation achieves the first successful deployment within three iterations and can surpass human expert results within seven. This agentic co-optimization achieves 250× compression for vision models with <3.3\% accuracy loss and 400× for audio with <6\% Feature Error Rate loss, enabling battery-free operation on a commercial MCU via solar harvesting. We demonstrate practical impact in two real-world systems: an elk-detection camera trap (96.7\% accuracy) and a phonetic-transcription wearable (8.44\% FER) for child development research.


Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models

Yifu Yuan ⋅ Yaoting Huang ⋅ Xianze Yao ⋅ Shuoheng Zhang ⋅ Linqi Han ⋅ Yutong Li ⋅ Pengyi Li ⋅ Jiangeng Sun ⋅ Wenting Jia ⋅ Yucheng Hu ⋅ YuHao Liu ⋅ Ruihao Liao ⋅ Qiyu Wu ⋅ Yuxiao Li ⋅ zhao zhang ⋅ Zibin Dong ⋅ Fei Ni ⋅ YAN ZHENG ⋅ Shuyang Gu ⋅ Yi Ma ⋅ Hongyao Tang ⋅ Han Hu ⋅ Jianye Hao

We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanning embodied cognition, task planning, correction, and pointing, within a single architecture toward general physical intelligence. Leveraging three automated data construction pipelines to significantly expand the data coverage of critical capabilities, we build a large-scale data system of over 15B tokens, and design a multi-task balanced RL recipe to alleviate heterogeneous task conflicts. We further introduce a Planner-Grounder-Corrector (PGC) closed-loop framework that enables a single model to autonomously execute and self-correct over long-horizon tasks. With only 8B parameters, Embodied-R1.5 achieves SOTA on 16 out of 24 embodied VLM benchmarks, surpassing leading models like Gemini-Robotics-ER-1.5 and GPT-5.4. Benefiting from the internalized embodied capabilities, Embodied-R1.5 can be fine-tuned into a VLA with only a small amount of data, outperforming leading VLA models like $\pi_{0.5}$ across 4 popular manipulation benchmark suites. We further conduct extensive zero-shot real-robot experiments, validating performance in instruction following, affordance grounding, articulated object manipulation, and long-horizon complex tasks, demonstrating strong generalization to the physical world. We will fully open-source weights, datasets and training code to facilitate future research in EFMs.


Emergent Low-Rank Training Dynamics in MLPs with Smooth Activations

Alec Xu ⋅ Can Yaras ⋅ Matthew Asato ⋅ Qing Qu ⋅ Laura Balzano

Recent empirical studies have shown the training dynamics of deep neural networks mostly occur within low-dimensional subspaces. While this has inspired new research in low-rank training, compression, and adaptation, theoretical justification for these dynamics in nonlinear networks remains limited. To address this gap, this paper analyzes the learning dynamics of multi-layer perceptrons (MLPs) under gradient descent (GD). We demonstrate that the weight dynamics concentrate within invariant low-dimensional subspaces throughout training. Theoretically, we precisely characterize these invariant subspaces for two-layer networks with smooth nonlinear activations, providing insight into their emergence. Experimentally, we validate that this phenomenon extends well beyond our theoretical setting. Leveraging these insights, we empirically show there exists a low-rank MLP parameterization that, when initialized in the appropriate subspaces, nearly matches the performance of their fully-parameterized counterparts on classification tasks.


Emergent Visual Thinking in Text-Only Reasoning through Multimodal Training

Qi Zhang ⋅ Bohan Wang ⋅ QIANRU SUN ⋅ Hanwang Zhang

Does multimodal training make a language model think visually, even when the input is text only? We test this with targeted causal ablations across 10 multimodal configurations from 6 model families, with matched text-only backbones and cross-family transfer controls. We contrast spatial-imagery tasks that invite mental scene construction with symbol-manipulation tasks solvable by symbolic rules alone. We identify a vision-language axis in hidden-state space and ablate its projection at inference. In Chameleon-7B, this ablation drops accuracy on BIG-Bench Hard (BBH) navigate by 27.6 percentage points while improving boolean expressions by 5.2 percentage points, a 32.8 percentage-point double dissociation. The effect is direction-specific and layer-localized. Together, these results reveal emergent visual thinking (EVT): text-only reasoning that causally depends on visual-spatial representations shaped by multimodal training. In dual-path unified models (BAGEL, Janus-Pro), the causal handle localizes to the generation pathway as a compact one-dimensional carrier. The understanding pathway also encodes spatial information in a higher-rank, probe-decodable form, but is not the causal handle. To separate decodability from causal use, we analyze concentration, alignment, and semantic location. These diagnostics explain why behavior, probing, and intervention can diverge. Multimodal training therefore reshapes the geometry of text-only reasoning, beyond changing how visual input is processed.


ENACT: Single-Image Human-Scene Interaction Motion from Language via Foundation-Model Orchestration

Sreehari Rajan ⋅ Kunal Kamalkishor Bhosikar ⋅ Charu Sharma ⋅ Nikos Athanasiou

Given a single image of a room and a sentence such as ``a person walks to the black armchair and sits down,'' producing a 3D motion sequence of a virtual person performing the described action is out of reach for existing methods. Scene-conditioned methods need hours of paired (scene, motion, language) capture. Zero-shot methods need a pre-scanned 3D scene. Image-based methods produce only static poses. We propose a training-free pipeline that synthesizes scene-aware human-scene interaction (HSI) motion from a single RGB image and a free-form text prompt, with no HSI-specific training and no pre-scanned 3D scene. The key observation is that every needed capability already exists as a frozen foundation model: the bottleneck is composition, not capability. Our method, ENACT, uses a vision-language model (VLM) for high-level reasoning. The VLM parses the input image and prompt into one structured keyframe per sentence, listing the action verb, the body part, the target object, a six-bin contact yaw, an image-edit prompt, and a foot-grounding flag. The keyframe then drives six frozen specialists: a monocular 3D model for the scene point cloud, an open-vocabulary segmenter for object localization, a single-view mesh reconstructor for object geometry, an instruction-following image editor for a synthetic interaction image, an image-conditioned human-mesh recoverer for pose initialization, and a 2D-grounded 3D contact predictor for object affordance. Per-keyframe contact poses are optimized under a whole-body diffusion prior, and a constraint-conditioned motion diffuser connects successive keyframes into one continuous motion. ENACT achieves the lowest mean and maximum scene penetration and the lowest foot sliding among all baselines on a standard HSI benchmark, despite consuming the smallest input. ENACT also generalizes to synthetic datasets and to in-the-wild phone images, including chained multi-step interactions.


EndoSCOP-V: A Multi-Turn Video Understanding Evaluation Framework for Multimodal Models in Endoscopy Reporting and Clinical Reasoning

Abhishek Tyagi ⋅ TANMAY JAIN ⋅ Vilas Biradar ⋅ Krithi K Koduri ⋅ Misba h Masoodi ⋅ Swetha Rapolu ⋅ HARDIK RUGHWANI ⋅ Rachana Kotha ⋅ Anjan K kaipa ⋅ Anudeep Katrevula ⋅ Rakesh Garlapati ⋅ ANIRUDDHA P SINGH ⋅ Rajendra Patel ⋅ Abhay Tyagi ⋅ SUNDEEP LAKHTAKIA ⋅ MOHAN RAMCHANDANI ⋅ RAKESH KALAPALA ⋅ NAGESHWAR REDDY

Real-world endoscopic diagnosis unfolds across native video, depending on continuous traversal of gastrointestinal organs where a single positive finding triggers a multi-minute reasoning chain over location, morphology, severity, and intervention. Existing endoscopy multimodal benchmarks evaluate only static images or single-disease videos that fail to capture the temporal and clinical reasoning demands of real-world endoscopy reporting. We introduce EndoSCOP-V, a video-grounded multi-turn evaluation framework for long-horizon clinical reasoning in endoscopy. The benchmark contains 1,307 real-world endoscopy and colonoscopy videos (63.2 h; 97% 1080p; mean ~174 s) — the time-scale at which endoscopists actually reason about a finding. The corpus is paired with 20,275 question–answer turns deterministically instantiated from clinician-filled structured endoscopy-reporting EHRs, avoiding synthetic or LLM-rephrased per-case annotations. Questions are derived from 1,033 clinician-reviewed templates spanning a clinically aligned 9-task × 30-sub-task reporting taxonomy across 54 World Endoscopy Organization (WEO) MST 3.0 diseases at ~18 templates per disease, substantially deeper per-disease examination than prior endoscopy benchmarks. To mirror real-world case mix, 22% of procedures carry co-existing pathologies, up to six per case. To model the hierarchical and conditional nature of clinical reporting, EndoSCOP-V introduces a Markovian context-pruned oracle-forcing protocol that evaluates multi-turn dialogues turn-by-turn. By injecting the gold answer before each subsequent turn, the protocol isolates per-turn reasoning from cascading dialogue errors and enables fine-grained analysis of clinical reasoning capabilities. Clinician-marked key-frame anchors further separate temporal-localization from frame-level comprehension failures. A panel of 15 board-certified endoscopists anchors a 78% strict-accuracy baseline on a 1,500-question audit. Across 14 open-weight LMMs, the best 27B-class model reaches 59.8% strict accuracy, falling to 53.4% once format-trivial disease-rejection items are excluded. On a between-subjects spot-check, the best of 3 frontier proprietary models trails the clinician panel by 23 pp on strict accuracy. Sub-task analysis reveals consistent gaps in quantitative measurement, multi-variable grading, and intervention reasoning on long-horizon video, providing a clinically grounded benchmark for evaluating multimodal reasoning in endoscopy.

Speech large language models (SpeechLMs) can process spoken requests, yet still lag behind text-prompted large language models (LLMs) on instruction following and reasoning. We argue that this gap is not only an ASR or representation problem: the same semantic instruction can induce different response policies when spoken with different speakers, accents, noise conditions, prosody, or disfluencies. We propose Verifiable Invariant Reinforced Behavior Alignment (VIRBA), a reinforcement-learning framework for aligning SpeechLM behavior across acoustic realizations of the same semantic intent. VIRBA builds multi-view spoken instruction groups, scores sampled responses with semantic preference, rule-verifiable correctness, cross-acoustic invariance, and adaptive reasoning rewards, and optimizes the model with Cross-Acoustic Group Relative Policy Optimization (CA-GRPO). The resulting objective moves SpeechLM alignment beyond teacher imitation toward robust reasoning policies that remain stable across realistic spoken realizations. Experiments with recent SpeechLM baselines, disfluency robustness, spoken QA, audio reasoning, and speech-to-text translation show the largest gains on reasoning-heavy and acoustically perturbed spoken prompts.


Ensembling Language Models with Sequential Monte Carlo

Robin Chan ⋅ Tianyu Liu ⋅ Samuel Kiegeland ⋅ Clemente Pasti ⋅ Jacob Vigly ⋅ Timothy O'Donnell ⋅ Ryan Cotterell ⋅ Tim Vieira

Combining the predictions of multiple models is one of the most effective strategies in machine learning. Thus, it is only natural to investigate the ensembling of language models (LMs). However, because LMs define distributions over strings, a proper ensemble requires aggregating predictions at the string level, which requires a computation that is generally intractable. In this work, we cast LM ensembling as an inference problem and propose the deployment of sequential Monte Carlo (SMC), an appropriate approximate inference scheme. Concretely, we introduce a unified framework for composing multiple LMs into ensemble distributions parameterized by a broad family of aggregation functions. To sample from these distributions, we introduce an SMC algorithm that operates in a shared character space, enabling ensembles of models with mismatching vocabularies and consistent sampling in the limit. We evaluate a range of ensembles across prompt and model combinations for various text generation tasks, finding that consensus-seeking strategies such as the product consistently outperform probability averaging, and that better posterior approximations can yield better ensemble performance.


EntiRE: Invariant Learning for Robust Concept Erasure in Text-to-Image Generative Models

Fengyuan Yu ⋅ Yuyuan Li ⋅ XiaoHua Feng ⋅ Li Zhang ⋅ Jiaming Zhang ⋅ Xiang Liu ⋅ Jun Wang ⋅ Chaochao Chen

Large-scale text-to-image diffusion models can inadvertently generate harmful or sensitive content learned from uncurated training data, motivating the development of concept erasure methods that remove undesired concepts from pre-trained models. However, recent studies reveal that erased models remain vulnerable to adversarial prompt attacks that recover the supposedly removed concepts, highlighting the need for robust erasure techniques. In this work, we visualize the text embedding distributions of adversarial prompts and find that the difficulty of erasure is non-uniform. Existing methods effectively erase concepts in easier regions but leave localized residual holes in harder regions, which adversarial prompts consistently exploit. Motivated by this finding, we propose Entire-space Robust Erasure (EntiRE), an end-to-end framework that casts robust concept erasure as an out-of-distribution generalization problem. By adopting an invariant learning formulation regularized by total variation, EntiRE enforces more uniform erasure across the continuous prompt space, suppressing the localized residual regions. Extensive experiments on erasing nudity, artistic style, and object concepts demonstrate that EntiRE consistently outperforms state-of-the-art baselines in both robustness against adversarial prompt attacks and utility preservation.


EpicWorldModel: Exploration-driven Planning with Latent World Models

Bowen Feng ⋅ Julian Ost ⋅ Zhiting Mei ⋅ Anirudha Majumdar ⋅ Felix Heide

Latent world models based on Joint-Embedding Predictive Architecture (JEPA) are deterministic by design. While successful in fully observable scenarios, this paradigm breaks down when past observations and actions lead to multiple plausible future possibilities, e.g., due to occlusion. We introduce EpicWorldModel, a framework to train stochastic JEPAs for environments and tasks with inherent uncertainty under partially observability. We jointly train the EpicWorldModel predictor with its latent representation space to directly predict multiple potential future states using a flow-matching objective, when the goal-relevant scene content is absent from the conditioning history. We show that flow predictive variance approximates an upper bound on Expected Information Gain and serves as a useful exploration guidance for planning. By incorporating this uncertainty signal into Cross-Entropy Method (CEM)-based planning, our approach balances goal-reaching with exploration of uncertain regions where occluded or unseen goals are most likely to be located. We demonstrate the effectiveness of EpicWorldModel through a series of latent planning experiments with the best performance among all baseline methods, showing up to $22\%$ empirical improvement in success rate over LeWorldModel.

Continual reinforcement learning (continual RL) seeks to formalize the notions of lifelong learning and endless adaptation in RL. In particular, the aim of continual RL is to develop RL agents that can maintain a careful balance between retaining useful information and adapting to new situations. To date, continual RL has been explored almost exclusively through the lens of risk-neutral decision-making, in which the agent aims to optimize the expected long-run performance. In this work, we present the first formal theoretical treatment of continual RL through the lens of risk-aware decision-making, in which the behaviour of the agent is directed towards optimizing a measure of long-run performance beyond the mean. In particular, we show that the classical theory of risk measures, widely used as a theoretical foundation in non-continual risk-aware RL, is, in its current form, incompatible with continual learning. Then, building on this insight, we extend risk measure theory into the continual setting by introducing a new class of ergodic risk measures, and showing that it is compatible with continual learning. Finally, we provide a case study of continual risk-aware learning, along with empirical results, which show the intuitive appeal of ergodic risk measures in continual settings.


ER-Reason: A Benchmark Dataset for LLM Clinical Reasoning in the Emergency Room

Nikita Mehandru ⋅ Niloufar Golchini ⋅ Namrata Garg ⋅ Kathy T LeSaint ⋅ Christopher J Nash ⋅ Anu Ramachandran ⋅ Travis Zack ⋅ Liam McCoy ⋅ Adam Rodman ⋅ David Bamman ⋅ Melanie Molina ⋅ Ahmed Alaa

Existing benchmarks for evaluating the clinical reasoning capabilities of large language models (LLMs) often lack a clear definition of "clinical reasoning" as construct, fail to capture the full breadth of interdependent tasks within a clinical workflow, and rely on stylized vignettes rather than real-world clinical documentation. As a result, recent studies have found significant discrepancies between LLM performance on stylized benchmarks derived from medical licensing exams and their performance in real-world prospective studies. To address these limitations, we introduce ER-Reason, a benchmark designed to evaluate LLM reasoning as clinical evidence accumulates across decision-making tasks spanning the full workflow of emergency medicine. ER-Reason comprises 25,174 de-identified clinical notes from 3,437 patients, supporting evaluation across all stages of the emergency department workflow: triage intake, treatment selection, disposition~planning, and final diagnosis. Crucially, evaluation in ER-Reason extends beyond diagnostic accuracy to include stepwise Script Concordance Test (SCT)-style questions grounded in real patient cases, which assess whether LLMs update their diagnostic beliefs in the correct direction and magnitude as clinical evidence accumulates, scored against 2,555 emergency physician annotations. We evaluate reasoning and non-reasoning LLMs on ER-Reason, and show that our tasks provide a more nuanced view of how LLM reasoning fails on real patient cases than existing benchmarks allow.


ESENSC: A Polynomial-Time Axiomatic Alternative to SHAP

Kazuhiro Hiraki ⋅ Shinichi Ishihara ⋅ Takumi Kongo ⋅ Junnosuke Shino

We propose ESENSC, a computationally efficient and theoretically grounded alternative to SHAP. ESENSC is a polynomial-time attribution rule with a deterministic, closed-form representation that satisfies the null-player property and admits an axiomatic characterization based on efficiency, a restricted form of differential marginality, and explicit computational constraints. Empirically, ESENSC closely replicates the attribution patterns of exact SHAP by preserving its essential structural properties, while maintaining a fraction of the cost. Across neural network and XGBoost models, it achieves lower deviation and higher rank agreement than widely used sampling-based SHAP approximations. While the computational cost of exact SHAP grows exponentially with the number of features, ESENSC scales linearly, enabling reliable feature attribution in regimes where exact SHAP is intractable. These results demonstrate that ESENSC achieves near-SHAP attribution accuracy while reducing computational complexity from exponential to linear in the number of features, providing a practical and theoretically rigorous alternative for feature attribution in high-dimensional applications.

Modern machine learning depends heavily on massive datasets, but obtaining high-quality annotations at scale is often expensive. As a result, learning from noisily-labeled data has become common, making accurate estimation of the label-noise transition matrix crucial. However, existing transition matrix estimators rely on the fragile estimation of class-posteriors and do not provide finite-sample performance guarantees. In this work, we propose a novel methodology to estimate the transition matrix based on one-sided selective classification. This approach bypasses class-posterior estimation, provides finite-sample performance guarantees, and leverages flexible learning methods for binary classification. Moreover, we introduce effective algorithms to implement the proposed methodology and provide their refined finite-sample performance bounds.

Hallucination remains a fundamental challenge in multimodal large language models (MLLMs), often stemming from unverified visual assumptions and reasoning processes that are weakly grounded in observable evidence. Existing approaches attempt to mitigate hallucination through self-reflection, multi-agent debate, or counterfactual verification. However, these methods remain largely opinion-driven: they lack explicit verification of the visual understanding process and may fail when multiple agents share the same incorrect perceptual assumptions, leading to consistent yet hallucinated conclusions. A promising direction is to augment MLLMs with external visual tools (e.g., OCR and grounding models) to acquire verifiable evidence. However, effective tool-augmented reasoning remains challenging due to two key failure modes: unreliable tool invocation, where models select inappropriate tools or produce incorrect parameters, and evidence--reasoning inconsistency, where models ignore or contradict the acquired evidence when forming final predictions. In this work, we propose EVA, an Evidence-seeking Visual Agent that reframes multimodal reasoning as an explicit process of evidence acquisition and verification. At the core of EVA is a unified evidence-grounded execution pipeline designed to address the above challenges through two tightly coupled mechanisms. First, we introduce \textit{Iterative Tool Replanning}, which dynamically refines tool selection and parameters based on intermediate observations, improving the reliability of tool usage. Second, we develop \textit{Evidence Consistency Verification}, an answer validation mechanism that improves alignment between the final prediction and the collected evidence, mitigating evidence--reasoning inconsistencies. In addition, to balance robustness and efficiency, EVA incorporates \textit{Selective Evidence Routing}, an adaptive fast--slow reasoning strategy that determines when external evidence is necessary. Visually straightforward queries are handled via direct reasoning, while complex or high-risk cases trigger tool-based evidence acquisition. By jointly addressing how to reliably acquire evidence and how to use it consistently, EVA transforms multimodal reasoning from passive answer validation into explicit evidence-backed decision making. Experimental results demonstrate that EVA achieves significantly improved reasoning reliability over direct VLM inference and multi-agent methods, while maintaining efficiency through selective tool invocation. These findings highlight structured and adaptive evidence utilization as a practical paradigm for hallucination-resistant multimodal reasoning.


Even Sailors Need Calm Seas: Taming the Geometry of VLMs for Fast Adversarial Fine-Tuning

Yuqing Wen ⋅ Junhao Dong ⋅ Ting Peng ⋅ Xudong Zhang ⋅ Jiaming Zhang ⋅ Weiming Liu ⋅ Zhengtao Yao ⋅ Xinghua Qu ⋅ Yew Soon Ong

Vision-Language Models (VLMs) such as CLIP have demonstrated strong performance across tasks yet remain highly vulnerable to adversarial attacks. While adversarial fine-tuning has been demonstrated to be effective against such malicious inputs, its multi-step adversary generation scheme during fine-tuning further induces prohibitive computational costs for VLM backbones. Previous works in unimodal models have introduced single-step strategies to reduce this burden. However, we find naive single-step strategies fail on VLMs due to more severe catastrophic overfitting. In addition, we identify a novel failure mode termed Perturbation Radius Overfitting, where VLMs overfit to the specific training attack budget while paradoxically becoming fragile to weaker perturbations. We trace these failures to growing angular misalignment and magnitude surge of the local gradients, which compromise the local linearity essential for accurate single-step approximation. Guided by this analysis, we introduce a joint adversarial optimization scheme that actively rectifies the local geometry by suppressing both angular misalignment and magnitude surge along the perturbation path. Our approach consistently outperforms existing single-step methods across various datasets and VLM architectures, achieving state-of-the-art robustness while preserving efficiency. To our knowledge, this is the first work to systematically examine single-step adversarial fine-tuning on VLMs and establish the geometric foundations of their robustness.


Event based Multi-Velocity-Scale Imaging

Qiyao Gao ⋅ Peiqi Duan ⋅ Chu Zhou ⋅ Xinyu Zhou ⋅ Imari Sato ⋅ Boxin Shi

Real-world scenes often contain static, slow, and fast components within the same field of view. Frame cameras are inherently inefficient in this setting, as their global exposure and fixed sampling must follow the fastest motion and therefore oversample the rest of the scene. Event cameras provide asynchronous local sensing, but standard sensors use a single global contrast threshold, leading to a trade-off between sensitivity and redundancy: low thresholds preserve static information but over-trigger under fast motion, while high thresholds reduce data volume but lose slow or static structures. We propose a multi-threshold event imaging model that assigns motion-compatible thresholds to different scene components, enabling faithful multi-velocity-scale imaging with fewer events. We also introduce a practical acquisition pipeline that implements this model on existing event cameras. Experiments on synthetic and real scenes validate the effectiveness of our approach in preserving both static structures and fast dynamics while reducing data volume.


EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration

WUYANG LI ⋅ Yang Gao ⋅ Mariam Hassan ⋅ Lan Feng ⋅ Wentao Pan ⋅ Po-Chien Luan ⋅ Alexandre Alahi

We propose EverAnimate, an efficient post-training method for long-horizon animated video generation that preserves visual quality and character identity. Long-form animation remains challenging because highly dynamic human motion must be synthesized against relatively static environments, making chunk-based generation prone to accumulated drift: (i) low-level quality drift, such as progressive degradation of static backgrounds, and (ii) high-level semantic drift, such as inconsistent character identity and view-dependent attributes. To address this issue, EverAnimate restores drifted flow trajectories by anchoring generation to a persistent latent context memory, consisting of two complementary mechanisms. (i) Persistent Latent Propagation maintains a context memory across chunks to propagate identity and motion in latent space while mitigating temporal forgetting. (ii) Restorative Flow Matching introduces an implicit restoration objective during sampling through velocity adjustment, improving within-chunk fidelity. With only lightweight LoRA tuning, EverAnimate outperforms state-of-the-art long-animation methods in both short- and long-horizon settings: at 10 seconds, it improves PSNR/SSIM by 8%/7% and reduces LPIPS/FID by 22%/11%; at 90 seconds, the gains increase to 15%/15% and 32%/27%, respectively. Anonymous project page: https://anonymous-xxx-lab.github.io/.


EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective

Yuyao Wang ⋅ Zhongjian Zhang ⋅ Mo Chi ⋅ Kaichi Yu ⋅ Yuhan Li ⋅ Miao Peng ⋅ Bing Tong ⋅ Chen Zhang ⋅ Yan Zhou ⋅ Jia Li

Recent benchmarks for Large Language Model (LLM) agents mainly evaluate reasoning, planning, and execution. However, memory is also essential for agents, as it enables them to store, update, and retrieve information over time. This ability remains under-evaluated, largely because existing benchmarks do not provide a systematic way to assess memory mechanisms. In this paper, we study agent memory from a self-evolving perspective and introduce \texttt{EvoMemBench}, a unified benchmark organized along two axes: memory scope (in-episode vs.cross-episode) and memory content (knowledge-oriented vs.execution-oriented). We compare 15 representative memory methods with strong long-context baselines under a standardized protocol. Results show that current memory systems are still far from a general solution: long-context baselines remain highly competitive, memory helps most when the current context is insufficient or tasks are difficult, and no single memory form works consistently across all settings. Retrieval-based methods remain strong for knowledge-intensive settings, whereas procedural and long-term memory methods are more effective for execution-oriented tasks when their stored experience matches the task structure. We hope EvoMemBench facilitates future research on more effective memory systems for LLM-based agents. Our code is available at \url{https://anonymous.4open.science/r/EvoMemBench-2}.


Exact Topological Compliance: Generating Persistence-Equivalent Graphs

Mattie Ji ⋅ Indradyumna Roy ⋅ Vikas Garg

Persistent homology summarizes the topological evolution of a graph as persistence diagrams (PD). Generating graphs whose topology matches a target descriptor is increasingly important, yet existing topology-aware generative models rely on soft regularization, producing graphs whose PDs may only be approximately similar. We initiate a principled investigation of the generation problem: produce graphs that realize a target PD exactly. We propose two approaches guided by the implicit local and global constraints a target PD imposes. The local approach builds degree-based PD-equivalent graphs iteratively, and we prove that every output realizes the target PD and that every degree-based PD-equivalent graph is reachable. However, as shown thereafter, local generation necessarily trades off between sample diversity and reaching infeasible states. We therefore recast generation as a global constraint satisfaction problem with a complete encoding solvable by modern constraint programming solvers. This is applicable to any vertex-based permutation equivariant filtration scheme and is minimal under the degree filtration. Experiments on four standard graph generation benchmarks confirm that both methods realize the target PD exactly. Together, our methods provide the first principled treatment of strict topology-compliant graph generation.


Example-Based Spatial Guidance for Training-Free Concept Erasure in Diffusion Models

Younghwan Kil ⋅ Joonhyeong Park ⋅ Giung Nam ⋅ Jinwoo Shin ⋅ Juho Lee

Text-to-image (T2I) diffusion models often synthesize policy-violating content, necessitating robust safeguards and rigorous evaluation frameworks. However, current approaches remain largely monolithic in both execution and evaluation, limited to coarse-grained categorization within broad unsafe concepts. Such approaches fail to account for the specific constituent factors of a violation, resulting in weak learning signals and allowing model collapse to bypass safety checks. To overcome these limitations, we propose Example-Based Spatial Guidance (EBSG), a training-free method that utilizes user-editable exemplar packs to provide granular spatial guidance for precise concept steering. Unlike existing methods that coarsely steer the entire image away from a broad concept, EBSG decomposes these categories into specific sub-concepts via exemplar text-image pairs, providing explicit, localized steering signals for more robust safety control. Furthermore, we introduce a vision language model (VLM)-based evaluation protocol that provides a fine-grained assessment, avoiding the pitfall of conventional binary evaluators that permit model collapse to bypass safety checks. Empirically, EBSG achieves state-of-the-art erasure performance across diverse safety datasets and concept categories, spanning four nudity benchmarks, seven I2P harmful-concept slices, and four MJA metaphor categories, while supporting multi-concept removal and transferring to a modern backbone, SD3.


ExperiGen: Agentic Hypothesis Discovery from Observational Data

Jishu Sen Gupta ⋅ Harini S I ⋅ Somesh K Singh ⋅ Mohamad Tawseeq Syed ⋅ Yaman Singla ⋅ Rajiv Ratn Shah ⋅ DAVID DOERMANN ⋅ Jitendra Ajmera ⋅ Balaji Krishnamurthy

Modern data-driven sciences use large observational datasets such as online discussions, behavioral logs, web experiments and biological records to explain outcomes and design interventions. Existing LLM-based hypothesis generators produce interpretable claims, but usually validate them through predictive accuracy, making them vulnerable to spurious correlations and dataset artifacts. Automated validation systems test hypotheses statistically, but require researcher-specified constructs. Here we show that ExperiGen, a two-agent closed-loop framework, discovers statistically supported natural-language hypotheses from observational data. A Generator proposes hypotheses, while an Experimenter operationalizes features, selects covariates and tests, executes analyses and returns evidence for refinement. Accepted and rejected hypotheses are stored in a labeled memory that guides further exploration. Across 19 tasks spanning text, images, clinical tabular data, economics, marketing and single-cell RNA-seq, ExperiGen discovers 2–4× more statistically significant hypotheses than prior generation methods and improves downstream predictive performance by 4–21 points. Expert evaluation by 26 specialists rated its hypotheses as novel, clear and research-worthy; 11 of 16 preregistered r/ChangeMyView hypotheses were significant in the predicted direction; and in deployed A/B tests across seven S&amp;P 500 companies and 5.1 million users, 15 of 25 hypotheses achieved significance, with 14 matching the predicted direction. These results suggest that closed-loop hypothesis generation can help prioritize hypotheses for expert review and real-world intervention and augment human judgment in data-driven sciences.


Explainability matters: The effect of liability rules on the healthcare sector

Jiawen Wei ⋅ Elena Verona ⋅ Andrea Bertolini ⋅ Gianmarco Mengaldo

Explainability, the capability of an artificial intelligence system (AIS) to explain its outcomes in a manner that is comprehensible to human beings at an acceptable level, has been deemed essential for critical sectors, such as healthcare. Is it really the case? In this perspective, we consider two extreme cases, Oracle (without explainability) versus AI Colleague (with explainability) for a thorough analysis. We discuss how the level of automation and explainability of AIS can affect the determination of liability among the medical practitioner/facility and manufacturer of AIS. We argue that explainability plays a crucial role in setting a responsibility framework in healthcare, from a legal standpoint, to shape the behavior of all involved parties and mitigate the risk of potential defensive medicine practices.


Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion

ShiYing Huang ⋅ Liang Lin ⋅ Yuer Li ⋅ Kaiwen Luo ⋅ Zhenhong Zhou ⋅ An Zhang ⋅ Junhao Dong ⋅ Kun Wang ⋅ Zhigang Zeng

In the realm of multi-objective alignment for large language models, balancing disparate human preferences often manifests as a zero-sum conflict. Specifically, the intrinsic tension between competing goals dictates that aggressively optimizing for one metric (e.g., helpfulness) frequently incurs a substantial penalty on another (e.g., harmlessness). While prior work mainly focuses on data selection, parameter merging, or algorithmic balancing during training, these approaches merely force compromises between divergent preferences along a fixed Pareto frontier, failing to fundamentally resolve the inherent trade-off. In this work, we approach this problem from a novel perspective of multi-dimensional rewards. By scaling up the model's rollouts and analyzing the outputs across different reward dimensions, we arrive at a critical conclusion: the conflict among multiple objectives stems from the fact that the prompt itself inherently restricts the achievable multi-dimensional rewards. Based on this core observation, we propose MORA: Multi-Objective Reward Assimilation (MORA). Specifically, MORA isolates single-reward prompts through pre-sampling and expands their reward diversity by rewriting the original questions to incorporate multi-dimensional intents. Extensive experiments demonstrate that: (I) in sequential alignment, MORA achieves single-preference improvements ranging from 5% to 12.4%, with exceptional gains in harmlessness, after multiple-preference alignment across helpful, harmless, and truthful dimensions. (II) In simultaneous alignment, MORA achieves an average overall reward improvement of 4.6%. Our codes are available at https://anonymous.4open.science/r/MORA-MPA.


Exploring the Epipolar Consistency for Light Field Deraining

Chen Gao ⋅ Youfang Lin ⋅ Wenbin Wang ⋅ Shuo Zhang

Light Field (LF) deraining relies on leveraging redundant information between views to restore regions occluded by rains. Existing LF deraining methods typically rely on predicted depth maps to guide cross-view information interaction, while learning a brute-force mapping to separate rain from background content. However, the semi-transparent nature of rain streaks often leads to inaccurate depth estimation, which undermines geometric alignment and degrades restoration quality. More critically, the blending of rain and background during restoration makes brute-force mappings ineffective, resulting in visible rain artifacts in the output. In this paper, we first construct a consistency-constrained LF rain model by considering the consistency and the accumulation effects of rain layers. Then, we propose a novel Consistency-Guided Recovery Network (CGRNet) specifically designed for LF deraining. Inspired by the fact that background regions occluded by rain streaks often exhibit inconsistencies across all views, we introduce consistency maps to perceive degradation by aggregating pixel-wise consistency information along epipolar lines. Guided by consistency maps, the proposed model adaptively samples and aggregates consistent and redundant features from clean background regions to restore the details of degraded areas. Experiments validate our method's superiority over state-of-the-art approaches, and show that our dataset effectively enhances practical performance under real-world conditions.


Fair Bubble Sort: Provably Optimal Fair Ranking with Continuous Sensitive Attributes

Antonio Ferrara ⋅ Fabio Vitale ⋅ Atsushi Miyauchi ⋅ Francesco Bonchi

Algorithmic ranking systems increasingly dictate high-stakes outcomes, yet most fairness interventions implicitly assume that protected attributes are categorical. When sensitive attributes are inherently continuous (e.g., age, income, or health scores), standard practices rely on arbitrary discretization, which discards crucial fine-grained information and destroys the natural ordinality of the data. In this paper, we study the problem of fair ranking with continuous protected attributes without relying on thresholding. We formalize fairness via the Kendall correlation between the ranking and the continuous attribute, and measure utility loss via the Kemeny distance from an initial score-based ranking. To solve this, we propose Fair Bubble Sort (FBS), a highly efficient adjacent-swap algorithm that minimizes utility loss subject to a strict continuous fairness constraint. We provide strong theoretical guarantees, proving that FBS is exactly optimal for the unweighted Kemeny distance. Extensive experiments on synthetic and real-world datasets show that FBS is scalable and consistently outperforms discretization-based baselines, achieving a superior fairness-accuracy Pareto front.

Modern AI image generators are increasingly deployed as opaque APIs, where customers can query the deployed service, but cannot inspect model weights or architecture. This creates a practical challenge: a provider may pass governance certification with one generator and later silently switch to a cheaper and lower-quality one for deployment, compromising public trust or even safety in high-stakes domains. We study integrity auditing at deployment time and propose FARE (Forensic Acceptance Region Estimation). A certified generator is enrolled by training FARE on images sampled from that generator. After deployment, FARE can determine whether a generated image is consistent with the enrolled generator-using only that image. FARE's features are based on image generator-specific artifacts that have been proposed for forensic applications. FARE amplifies these features during training by finding hard samples that tighten the acceptance region and increase sensitivity to subtle changes in the certified generator. Across generator swaps, including substitutions with similar model versions and model variants, FARE is effective at detecting swaps, consistently outperforming existing baselines at strict operating points, and remains effective under adaptive attacks that actively manipulate images to evade swap detection.


FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control

JUN XUE ⋅ Junze Wang ⋅ Shanze Wang ⋅ Xinming Zhang ⋅ Yanjun Chen ⋅ Wei Zhang

Scaling Maximum Entropy Reinforcement Learning (RL) to high-dimensional humanoid control remains a fundamental challenge, as the ''curse of dimensionality'' induces severe exploration inefficiency and training instability. Consequently, highly optimized deterministic policy gradients currently dominate high-throughput regimes. We address this limitation with FastDSAC, a framework that effectively unlocks the potential of maximum entropy stochastic policies for complex continuous control. We introduce Dimension-wise Entropy Modulation (DEM) to dynamically redistribute the exploration budget, alongside a continuous distributional critic tailored to ensure accurate value estimation by mitigating both high-dimensional overestimation and discrete quantization artifacts. Extensive evaluations on HumanoidBench and a diverse set of continuous control tasks demonstrate that FastDSAC establishes state-of-the-art performance for high-dimensional stochastic policies on the evaluated benchmarks. Our method is competitive with and often outperforms strong deterministic baselines, with gains of 180\% and 350\% on the challenging Basketball and Balance Hard tasks, respectively.


Fast-RL: Accelerating Reinforcement Learning for LongCoT Reasoning Models

Sitong Wu ⋅ Haoru Tan ⋅ bin xia ⋅ Bei Yu ⋅ Xiaojuan Qi ⋅ Jiaya Jia

Large Reasoning Models (LRMs) have made remarkable progress, driven by a paradigm shift from fast thinking to slow thinking. While slow-thinking unlocks superior reasoning ability, its long-form nature imposes a prohibitive computational burden on Reinforcement Learning (RL) for LRMs. In this paper, we propose **Fast-RL**, a novel framework that accelerates RL training for slow-thinking LRMs. Our acceleration principle stems from a foundational insight: fast-thinking sampling serves as an efficient and powerful proxy for slow-thinking sampling in enhancing reasoning, since the core reasoning skills are shared across thinking modes and can be improved through either trajectory type. Fast-RL consists of two stages. **Stage-1** establishes reliable prompt-based thinking-mode control, so that a slow-thinking-oriented model triggers fast-thinking under a specific prompt while preserving its native slow-thinking under the standard prompt. **Stage-2** then enhances reasoning ability through RL with exclusively fast-thinking sampling for acceleration. A lightweight LoRA adapter is introduced as a dedicated mode controller: it is trained in Stage-1 and **frozen** in Stage-2, decoupling mode-control parameters from reasoning parameters and protecting the model's native slow-thinking from being disturbed by fast-thinking RL. After Stage-2, the LoRA acts as a disposable training scaffold that can be either merged into the main model or simply discarded. Experiments on diverse multimodal and text-only LRMs show that Fast-RL accelerates training by about $10\times$ while delivering superior performance gains over vanilla slow-thinking RL. On Qwen3-1.7B, Fast-RL improves AIME24 by $+10.0$ and $+13.7$ in slow- and fast-thinking evaluation, respectively. Vanilla slow-thinking RL, in contrast, drops slow-thinking accuracy by $-8.8$ and yields only $+2.8$ in fast-thinking evaluation.

Can a large language model be \emph{trained} to tolerate silent hardware faults, and if so, \emph{where} in the networkshould the consistency signal be enforced? We answer both questions and find that the second one matters more than the first. Injecting tile-level GEMM faults during training and adding a KL consistency term on the output distribution reduces fault-induced perplexity degradation from $12.9\%$ to $0.21\%$ on GPT-2 Small---a $60\times$ reduction (two-seed mean). The placement of the consistency target is decisive: a direct hidden-only objective at the fault site can worsen degradation to $32.7\%$, and fair bounded hidden controls with corrupted CE improve substantially but still trail output KL in an 80k-step screen ($1.74$--$1.76\%$ vs.\ $0.84\%$). A conditional constraint hierarchy and gradient-alignment probes explain why output-level consistency is more permissive and easier to optimize than direct early hidden-state consistency, without requiring that every hidden-alignment formulation fail. The output-over-hidden advantage persists at 8B scale, the GPT-2 ordering reproduces on held-out corpora, and generic robustness methods---SAM, R-Drop, and SAF---fail to close the gap. Pairing the trained model with simulated algorithm-based fault detection yields a bit-flip-to-erasure interface with a $1.10\times$ perplexity ratio, where unprotected inference produces NaN.

Federated learning (FL) encounters substantial challenges due to heterogeneity, leading to gradient noise, client drift, and partial client participation errors, the last of which is the most pervasive but remains insufficiently addressed in current literature. In this paper, we propose FedAdaVR, a novel FL algorithm aimed at solving heterogeneity issues caused by sporadic client participation by incorporating an adaptive optimiser with a variance reduction technique. This method takes advantage of the most recent stored updates from clients, even when they are absent from the current training round, thereby emulating their presence. Furthermore, we propose FedAdaVR-Quant, which stores client updates in quantised form, significantly reducing the memory requirements (by 50%, 75%, and 87.5%) of FedAdaVR while maintaining highly competitive model performance. We analyse the convergence behaviour of FedAdaVR under general nonconvex conditions and prove that our proposed algorithm can eliminate partial client participation error. Extensive experiments conducted on multiple datasets, under both independent and identically distributed (IID) and non-IID settings, demonstrate that FedAdaVR consistently outperforms state-of-the-art baseline methods.


Fed-AGA: An Anchor Graph Alignment Framework for Federated Unaligned Multi-view Clustering

Yichi Zhang ⋅ bohang sun ⋅ Hao Wei ⋅ Kai Di ⋅ Zhen Yang ⋅ Gengyu Lyu

Federated Multi-view Clustering (FMVC) enables privacy-preserving cross-view fusion from distributed multi-view data, where each client holds one view of samples and participates in the server's fusion without exposing raw data. In practice, existing FMVC methods face two key challenges: (1) the collected samples often show unaligned correspondence across views, especially under distributed collection; and (2) the gap between higher data privacy and lower computation cost remains hard to reconcile. To address these issues, we propose an Anchor Graph Alignment Framework for Federated Unaligned Multi-view Clustering (Fed-AGA), which employs a coarse-to-fine alignment strategy on lightweight anchor graphs to achieve misalignment-robust cross-view fusion while bridging the gap between data privacy and computation cost. Specifically, in the coarse alignment stage, we introduce an anchor-based graph generation module to generate an anchor graph on each client and upload them to the server for category-wise alignment by a designed topology-aware alignment module, which distills topology components as category-exclusive signatures to guide and achieve ideal category-wise alignment. In the next fine alignment and fusion stage, we design an attention-based alignment module to encourage each sample to pay attention to highly similar samples for sample-wise alignment, and the fine-aligned results are contrastively fused into a cross-view consistent anchor graph for clustering. Extensive experiments on multiple datasets demonstrate that Fed-AGA achieves state-of-the-art performance among related methods, while theoretical analysis offers theoretical motivation for two key components.

Neural combinatorial optimization (NCO) solvers generalize poorly under distribution shift, yet robustness across heterogeneous problem distributions is essential for real-world deployment where no single client should be left behind. Group distributionally robust optimization (Group DRO) offers a principled remedy, but requires pooling data that often cannot be shared due to confidentiality or regulatory constraints. Federated learning removes data-sharing requirements; however, existing federated DRO methods cannot accommodate the policy-gradient estimators used in modern NCO. Specifically, prior analyses rely on Lipschitz losses with bounded variance. In contrast, the REINFORCE policy-gradient estimator multiplies stochastic costs by score functions, causing variance to scale with cost magnitude and violating these standard assumptions. We close this gap by replacing Lipschitz-loss conditions with a bounded-cost assumption. This allows us to control cost-dependent variance and establish the first federated Group DRO framework with $O(1/\sqrt{T})$ convergence guarantees. Our analysis applies to any REINFORCE-trained model under bounded costs, including NCO solvers as a special case. Experiments on capacitated vehicle routing---one of many applicable domains including scheduling and packing---demonstrate the necessity of federated training for minority-distribution clients. We show that while na\"ive Group DRO suffers from catastrophic weight collapse, our KL-regularized formulation avoids this collapse. Furthermore, it provides a formal worst-group robustness certificate with no empirical performance penalty compared to standard federated averaging.


FedLoVA: Value-Only Aggregation for Federated LoRA Fine-Tuning of Large Language Models

Ensieh Khazaei ⋅ Baturalp Buyukates ⋅ Dimitrios Hatzinakos

Low-Rank Adaptation (LoRA) is a popular parameter-efficient fine-tuning (PEFT) method for fine-tuning of large language models (LLMs), especially in federated learning, due to its strong performance and communication efficiency. In practice, LoRA is applied to the query and value projections of transformer layers without considering the distinct roles of these components. In this work, we analyze the impact of these projections in federated fine-tuning of LLMs using LoRA and observe that the value projection converges significantly faster than the query projection. Motivated by this finding, we propose FedLoVA, a federated fine-tuning framework that trains LoRA adapters on both query and value projections locally while sharing only the value projection with the server. This approach reduces communication overhead while improving performance and convergence efficiency. Experiments on natural language understanding, multilingual sentiment analysis, and generation tasks show that FedLoVA reduces the communication cost by up to $7.8\times$ compared to existing federated LLM fine-tuning baselines while achieving comparable performance.


FedPeel: Peeling Stabilized Layers for Robust Heterogeneous Federated Learning

Haizhou Du ⋅ Lixin Huang ⋅ Zonghan Wu ⋅ Huan Huo ⋅ Huaicheng Yan

Proxy model-based heterogeneous federated learning enables clients with diverse architectures to collaborate through a shared lightweight intermediary. However, current methods aggregate and distill proxy knowledge in a structure-agnostic manner, ignoring layer-wise convergence dynamics, conflating shared and client-specific information, and transferring knowledge without assessing its reliability. These oversights jointly cause model oscillation, drift, and negative transfer. We propose FedPeel, a framework that makes the aggregation pipeline aware of the internal structure of model parameters. Rather than updating the proxy monolithically, FedPeel selectively aggregates only the layers that have stabilized, decouples their frequency-domain representations to separate generalizable patterns from personalized details, and modulates knowledge transfer intensity on a per-sample basis according to teacher confidence. Theoretical analysis establishes an O(1/T ) convergence rate. Experiments on image, text, and audio benchmarks show that FedPeel consistently outperforms state-of-the-art methods in both accuracy and communication efficiency.


FedReCall: Recalling Client-Specific Directions in Federated LoRA Fine-tuning

Chang Liu ⋅ Jinqian Chen ⋅ Jihua Zhu ⋅ Xiangyang Yang

Low-rank adaptation (LoRA) enables communication-efficient federated fine-tuning of large language models, but standard FedLoRA averages the factors $A$ and $B$ separately even though the model update is determined by their product $\Delta W=sBA$. Under task heterogeneity, factor-wise aggregation can further dilute directions that are locally useful but not strongly shared across clients; we call this phenomenon \emph{directional dilution}. We analyze this effect through the singular directions of client LoRA updates and find that low-strength directions are most vulnerable to dilution. Their repeated owner-specific recovery suggests client-specific adaptation signals rather than random noise. We propose \textsc{FedReCall}, a client-private plug-in that captures observable dilution gaps, selects reliable diluted directions into a private frozen cache, and recalls them through a low-rank bypass guarded by a single-batch loss probe when the client next participates. As a plug-in, \textsc{FedReCall} can be seamlessly integrated into existing FedLoRA methods without changing their aggregation rules, communication payloads, or trainable-parameter sets, consistently improving personalized adaptation and improving global evaluation metrics on most aggregators.


FedVaccine: Knowledge Recall after Spatial-Temporal Catastrophic Forgetting in Federated Continual Learning via Gradient-Based Vaccine

Hao Yu ⋅ Xin Yang ⋅ Boyang Fan ⋅ Xuemei Cao ⋅ Lingfei Ren ⋅ Hanlin Gu ⋅ Yang Liu ⋅ Qiang Yang

Federated Continual Learning (FCL) suffers from spatial-temporal catastrophic forgetting caused by sequential task learning on individual clients and the aggregation of heterogeneous client knowledge. However, most existing methods either overemphasize preserving old knowledge, which restricts model plasticity for new tasks, or rely on training complex generative models for replay, incurring substantial computational overhead. In the human immune system, immune memory is preserved implicitly and can be reactivated through vaccination even after long periods of dormancy. Motivated by this, we propose FedVaccine, a novel framework that explicitly permits forgetting to preserve high plasticity for new tasks, while efficiently recovering past knowledge through re-learning a compact vaccine set. Specifically, the vaccine set for each task is synthesized by compress task gradient information under the guidance of the fine-tuned local model, capturing representative data characteristics and task-specific discriminative patterns. Two novel components are also introduced: Priming & Boosting, which encodes task-specific knowledge in a compact spectral form, priming the model with essential information and enabling efficient memory boosting via vaccine set replay after forgetting; and Vaccination Aggregation, which aggregates vaccine sets from all clients to train a globally generalized model without accessing any raw private data. Extensive experiments on three benchmarks demonstrate that FedVaccine achieves competitive performance, enables rapid knowledge recovery after forgetting, and incurs no additional overhead from training generative models.

Enabling language model agents to autonomously adapt to new environments remains a fundamental challenge. While recent memory-augmented approaches have achieved promising results, they are largely heuristic in design and treat task execution and knowledge construction as decoupled processes. Inspired by how biological systems rapidly adapt through principled acting and learning, we propose \textbf{FEP-Agent}, a framework that grounds LLM agent self-evolution in the \textit{Free Energy Principle (FEP)}. Our system instantiates the two components of Active Inference into: a Worker agent that minimizes \textit{Expected Free Energy (EFE)} during online interaction, balancing goal-directed exploitation with curiosity-driven exploration to probe uncertain environmental dynamics; and a Builder agent that minimizes \textit{Variational Free Energy (VFE)} post-execution, updating a knowledge base while controlling complexity through counterfactual validation. To support scalable knowledge in open-ended environments, we introduce a structured semantic memory that employs associative linking and overlapping communities for efficient retrieval. The excellent results across multiple LLM-agent benchmarks demonstrate that FEP-Agent achieves principled, self-evolving systems.


Few-Step Boltzmann Generators via Scalable Likelihood Flow Maps

RuiKang OuYang ⋅ Hanlin Yu ⋅ Xinyue Ai ⋅ Yutong He ⋅ Nicholas Boffi ⋅ Pradeep Ravikumar ⋅ José Miguel Hernández-Lobato ⋅ Max Simchowitz ⋅ Benjamin K Miller ⋅ Omar Chehab

Recent progress in flow-based generative modeling has led to models that output high-quality samples while using only a small number of function evaluations. However, at present, there is a lack of similar advances in estimating the density of the generated samples. In particular, most existing methods either rely on restrictive architectures that enable exact calculations, or use stochastic approximations such as Hutchinson’s trace estimator that introduce substantial variance. In this work, we introduce SCAlable LikeLihood distillation of flOw maPs (SCALLOP). SCALLOP builds on the recently proposed F2D2, a likelihood flow map model that can generate samples and their densities in a small number of function evaluations. While {F2D2} uses Hutchinson's estimator during training, we introduce an alternative and more scalable likelihood distillation objective that is Hutchinson-free and admits a vectorized formulation. Empirically, we demonstrate the effectiveness of SCALLOP as a Boltzmann generator in molecular science, and further validate its benefit on image datasets. SCALLOP significantly reduces both training variance and training time while consistently improving performance compared to F2D2, and is competitive with the state-of-the-art while achieving up to $10\times$ inference speedup over the fastest baseline.


Finetuning with Sampling: Make SFT Generalize, Not Forget

Aayush Karan ⋅ Sitan Chen ⋅ Yilun Du

Introducing new capabilities to frontier models has long been the goal for posttraining, which relies predominantly on supervised finetuning (SFT) and reinforcement learning (RL) to achieve this. Prevailing wisdom dictates that on-policy RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and forgetting. At the same time, SFT enables learning from inherently off-policy expert data, whereas RL must rely on a model's ability to generate positive signal on a new task. In our work, we seek to bridge the strength of on-policy learning with the information signal of off-policy expert traces. To this end, we introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that transforms off-policy traces into more on-policy ones given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and fact learning, we demonstrate that simple SFT on these boosted off-policy traces can generalize better and forget less than on-policy counterparts. More generally, our approach outlines a principled methodology for model-specific data refinement, suggesting broader utility as a plug-and-play component throughout the post-training pipeline.


Fingerprinting Inference Systems of Large Language Models

Anna Wimbauer ⋅ Jonas Möller ⋅ Erik Imgrund ⋅ Konrad Rieck

The behavior of LLMs does not depend solely on the model itself. Components of the inference system, such as the inference engine, attention backend, and hardware platform, subtly influence how inputs are processed. These components differ in their implementations and thereby induce small numerical deviations across systems when running the same model. While prior work has established the theoretical existence of such deviations, their security implications have remained unexplored. In this paper, we show that these deviations are characteristic of specific components and propagate to observable textual outputs, exposing the inference system to any party that can query the model. Building on this observation, we introduce a fingerprinting method that analyzes the prompt-response behavior of LLMs to identify components of the inference system. Our empirical evaluation demonstrates that the inference engine, attention backend, and underlying hardware platform can be identified reliably, even when the LLM is operated at non-zero temperature. We show that preventing fingerprinting is fundamentally hard, as it would require eliminating numerical differences between hardware and software stacks. We therefore propose partial mitigations and discuss their impact.


FireMPC: A Multi-Source Pan-Canadian Wildfire Benchmark Revealing Spatiotemporal Generalization Gaps

Zhengsen Xu ⋅ Sibo Cheng ⋅ Lanying Wang ⋅ Aryan Sharma ⋅ Linlin Xu

Under Climate Change, wildfire activity is intensifying worldwide, with Canada experiencing increasingly severe events that now affect every province and territory. However, existing machine-learning (ML) datasets remain constrained by geographic bias, narrow driver sets, opaque sample selection, and limited exploration of model failure modes. We introduce \textbf{FireMPC} dataset, the first pan-Canadian multisource wildfire risk benchmark, covering approximately one billion hectares, or $9{,}840 \times 4{,}639$ km$^2$, across the entire Canada at $1$ km daily resolution from $2000$ to $2025$, and integrating $55$ drivers across fuel, topography, meteorology, and human activity. FireMPC is also the first large-scale machine-learning benchmark to embed the six knowledge-driven indicators of the Canadian Forest Fire Danger Rating System (CFFDRS), bridging knowledge-based and data-driven modeling. Building on this dataset, we propose \textbf{FWI-guided hard negative mining (FWI-HNM)} strategy. We score each non-fire candidate using a calibrated CFFDRS composite, and then incorporate these hard negatives during model training. FWI-HNM mitigates the trivial-negative failure mode of random sampling, yielding F1/PR-AUC gains of $1\%$-$2\%$ and ECE reductions of up to $82\%$ across thirteen benchmark ML models. Stress tests under horizon-extension and spatial-generalization protocols expose substantial generalization gaps. Mean F1 falls from $70.2\%$ at the next-day horizon ($\Delta t = 1$) to $54.2\%$ at the 30-day horizon ($\Delta t = 30$). All six ecological, land-cover, and regional spatial-shift scenarios yield a $25\%$-$29\%$ relative F1 drop. These results motivate the development of large-scale, driver-diverse, spatially comprehensive datasets. The dataset, code, and sampling artifacts are available at \url{https://anonymous.4open.science/r/FireMPC}.


Fix the Structural Bottleneck: Context Compression via Explicit Information Transmission

Jiangnan Ye ⋅ Hanqi Yan ⋅ Zhenyi Shen ⋅ Heng Chang ⋅ Ye Mao ⋅ Yulan He

Long-context LLM agents often struggle with growing token, memory, and latency costs, making efficient context compression essential for practical deployment. Existing LLM-as-a-compressor methods remain noticeably inferior to using the full context. We find that this gap partly stems from their inability to preserve contextual information effectively. In this work, we revisit context compression from a structural perspective and identify two key bottlenecks in standard LLM-based compressors: limited coordination among compression tokens during information aggregation, and layerwise dilution that weakens useful signals from intermediate hidden states. To address these limitations, we propose ComprExIT (Context Compression via Explicit Information Transmission), a new context compression framework based on explicit information transmission. ComprExIT adaptively selects features across frozen LLM layers, then allocates information from anchors to compression slots through a globally coordinated transport plan. Experiments on 12 datasets show that ComprExIT consistently outperforms strong soft-compression baselines, improving average F1 by up to 18.5%, while adding only ~1% trainable parameters and achieving more than 2× faster compression than the fastest baselines. The code will be released upon acceptance.


FlashControl: One-Step Controllable Generator via Distillation-Friendly Single-Stream Teachers.

Ngan Nguyen ⋅ Dung Nguyen ⋅ Quan Dao ⋅ Dimitris Metaxas ⋅ Minh Hoai ⋅ Cuong Pham ⋅ Anh Tran

Diffusion models have become a leading approach in visual generative modeling, particularly effective in text-to-image and controllable generation. Their main drawback is slow inference as it requires many sampling steps. While recent distillation methods enable one-step text-to-image generation, extending these techniques to controllable generation is still challenging. Existing controllable models, such as ControlNet, rely on dual-branch design that make distillation difficult. In this work, we adopt a simpler strategy. Instead of adding extra conditional branches, we finetune a selected subset of layers in the original text-to-image diffusion model to directly enable controllable generation. This results in DeltaControl, a single-flow, multi-step diffusion model capable of supporting spatial-condition generation while matching the performance of dual-branch approaches. Building on this teacher model, we further apply step-distillation to obtain FlashControl, a one-step model capable of spatial controllable generation with a single forward pass.


FlashMask-3: Efficient and Expressive Mask-Aware Distributed Attention

Guoxia Wang ⋅ Qianyue He ⋅ Siming Wu ⋅ Haoyang Xie ⋅ Qiwen Bao ⋅ Jinle Zeng ⋅ Jiabin Yang ⋅ Dianhai Yu ⋅ TIAN WU

Long-context training of large language and multimodal models couples expressive structured attention masks (document, prefix-LM, share-question, sliding-window, block-sparse) with multi-node context parallelism, exposing the mask as a dominant bottleneck. The column-wise sparse mask of FlashMask compresses such masks from $O(n^2)$ to $O(n)$ via four query-axis boundary vectors and covers a broad family of structured patterns. Implemented for Ampere GPUs, FlashMask does not exploit FlashAttention-3's TMA/WGMMA warp-specialized pipeline on Hopper GPUs, and offers no distributed support. We observe that the column-wise sparse mask is indexed by key positions but stores query-axis interval boundaries. Under any multi-chunk context-parallel assignment, this property reduces per-rank mask localization to a provably zero-communication \emph{clip--shift--max} on the boundary vectors, directly compatible with the single-GPU kernel. Building on this property, FlashMask-3 introduces a mask-aware Hopper kernel with a three-role warp specialization atop the FlashAttention-3 TMA/WGMMA pipeline, together with a mask-aware distributed extension: zero-communication mask localization, a sparsity-aware load balancer (HCC-IPO), and a mask-aware compute-communication overlap with mask-driven communication pruning, whose forward all-gather additionally adopts topology-aware hierarchical routing. On a single H100 GPU across 12 structured masks at context lengths from 4K to 128K, FlashMask-3 improves achieved TFLOPs over FlashMask, FlexAttention, and MagiAttention by 45\% to 141\%, 34\% to 71\%, and 2\% to 44\%, respectively. Under context-parallel degrees from 4 to 32 with 8K sequence per device, it attains 31\% to 75\% compute-communication overlap efficiency and improves throughput over Megatron-Core CP (ring), Megatron-Core CP (all-gather) and MagiAttention by $3.25\times$ to $5.11\times$, $1.97\times$ to $3.26\times$ and $1.06\times$ to $4.09\times$ respectively. The source code will be released on GitHub.


FlashRelight: Portrait Video Relighting with Dynamic Lighting

Jiahui Sheng ⋅ Le Li ⋅ Limin Lin ⋅ Jiahao Li ⋅ Wansen Feng ⋅ Hong Gu ⋅ Shuhan Chen ⋅ Xiaorun Li ⋅ Rui Hu

Portrait video relighting aims to modify the illumination of a portrait video while preserving photorealism and temporal stability. The task becomes particularly challenging under dynamic lighting, where illumination direction, intensity, and color vary over time. To tackle this problem, we introduce FlashRelight, a novel two-stage framework for dynamic portrait video relighting. FlashRelight first suppresses the source illumination through delighting, and then synthesizes the target dynamic lighting on an illumination-neutral portrait video. FlashRelight supports three complementary control modalities: envmap sequences for explicit frame-level lighting control, dynamic lighting text for semantic illumination editing, and lighting reference videos for visual lighting transfer. Training follows a low-cost hybrid strategy that combines static one-light-at-a-time (OLAT) data with generated dynamic-lighting portrait videos, covering both controlled illumination variation and realistic portrait motion. Extensive experiments show that FlashRelight achieves temporally coherent and photorealistic dynamic relighting while enabling flexible cross-video lighting transfer.


FlatClip: Reusing Image Foundation Models for fMRI Representation Learning via Cortical Flatmaps

Mo WANG ⋅ Wenhao Ye ⋅ Zihan Ning ⋅ Jiayu Zuo ⋅ Junfeng Xia ⋅ Hongkai Wen ⋅ Quanying Liu

Recent fMRI foundation models differ substantially in the spatial scale at which they represent brain activity. ROI- and connectivity-based models are efficient but coarse, whereas voxel-level models preserve fine-grained spatial structure but require specialized architectures and costly fMRI-specific pretraining. We ask whether part of this performance gap reflects the importance of preserving cortical geometry. Motivated by evidence that macroscale brain activity is strongly constrained by brain geometry, we introduce FlatClip, a training-free surface-level baseline that renders cortical activity as geometry-aware flatmap sequences and reuses a frozen SigLIP2 image encoder with only a lightweight downstream probe. Because FlatClip is not pretrained on fMRI, it provides an out-of-domain reference for evaluating how much information can be recovered from cortical geometry without relying on fMRI-specific pretraining data. Across resting-state and visual-fMRI benchmarks, FlatClip outperforms ROI/FC-style controls and several general-purpose fMRI foundation baselines, while voxel-level models remain a strong upper comparison. Geometry-control experiments show that disrupting cortical topology reduces performance, whereas mapping ROI-level signals back into atlas-defined 4D volumes partially recovers performance but remains below full voxel inputs. Together, these results position cortical geometry as a key factor in fMRI representation learning and establish surface-level flatmap sequences as a practical middle-ground baseline between ROI and voxel models.


Flow-Corrected Shape Optimization: Taming Manifold Drift in High-Dimensional 3D Models

Emilien Seiler ⋅ Nicolas Talabot ⋅ Yingxuan You ⋅ Federico Stella ⋅ Pascal Fua

Optimizing 3D shapes within the latent spaces of deep generative models is fundamental to computer assisted engineering, yet remains prone to a critical failure mode we term manifold drift: the tendency of gradient-based optimization to move latent vectors away from the manifold of valid shapes. This problem is exacerbated in state-of-the-art 3D shape generative models that operate in increasingly high-dimensional latent spaces where valid shapes occupy a vanishingly small fraction of the full space. Existing mitigation strategies, including latent regularization and flow-matching approaches, either sacrifice expressiveness, demand a difficult trade-off between objective guidance and generative fidelity that remains prone to manifold drift, or are computationally infeasible to scale to modern, large-capacity 3D shape models. We introduce a novel optimizer-corrector framework that alternates between gradient steps for objective minimization and guided flow matching to drive the latent state back to the valid shape manifold. By decoupling objective minimization from flow-based correction, optimizing freely and correcting strictly, this alternating design avoids inherent trade-offs, preserving geometric validity without sacrificing expressiveness while remaining computationally feasible on modern 3D shape models. We demonstrate its effectiveness across generative priors of varying complexity, from simple vector latent spaces to large-scale architectures across a variety of downstream optimization tasks, including aerodynamic drag reduction and object compliance optimization.


Flow Matching for Count Data

Ganchao Wei ⋅ John Pearson

High-dimensional count data arise in applications such as single-cell RNA sequencing and neural spike trains, where mapping between distributions across successive batches or time points form critical components of data analysis. The recent success of diffusion- and flow-based deep generative models for images, video, and text motivates extending these ideas to count-valued settings, but many existing methods either treat each count as a categorical state or transform counts into a continuous space, neither of which is natural or efficient when the count range is large. We propose count-FM, a flow-matching framework for count data based on a continuous-time birth-death process with local unit jumps. Count-FM learns marginal transitions efficiently in count space through simulation-free training of conditional transition rates, allowing transport between arbitrary count-distributed source and target populations. In simulation, count-FM achieves better sample quality than representative baselines while using substantially fewer parameters. We further apply count-FM to scRNA-seq and neural spike-train data for unconditional generation, transport, and conditional generation. Across these tasks, count-FM yields improved sample quality, greater modeling efficiency, and interpretable transport paths.

The scalability of continuous normalizing flows (CNFs) for unbiased Boltzmann sampling remains limited in high-dimensional systems due to the cost of Jacobian-determinant evaluation, which requires $D$ backpropagation passes through the flow layers. Existing stochastic Jacobian estimators such as the Hutchinson trace estimator reduce computation but introduce bias, while the recently proposed Flow Perturbation method is unbiased yet suffers from high variance. We present **Flow Perturbation++**, a variance-reduced extension of Flow Perturbation that discretizes the probability-flow ODE and performs unbiased stepwise Jacobian estimation at each integration step. This multi-step construction retains the unbiasedness of Flow Perturbation while achieves substantially lower estimator variance. Integrated into a Sequential Monte Carlo framework, Flow Perturbation++ achieves significantly improved equilibrium sampling on a 1000D Gaussian Mixture Model and the all-atom Chignolin protein compared with Hutchinson-based and single-step Flow Perturbation baselines.


FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning

Xirui Li ⋅ Zhe Liu ⋅ Xiaoqing Ye ⋅ Wenhua Han ⋅ Yifeng Pan ⋅ Junyu han ⋅ Hengshuang Zhao

Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervision but are confined to a fixed action vocabulary, while anchor-based methods generate proposals dynamically yet suffer from sparse supervision constrained to a single ground-truth trajectory. In this work, we propose FlowR2A, which resolves this tension by reframing simulation-based rewards from discriminative targets into generative conditions. By learning the reward-conditioned action distribution from dense trajectory-reward pairs with a flow-matching decoder, FlowR2A unifies the dense supervision of scoring-based methods with the proposal generation of anchor-based methods in a single generative model, forcing the model to internalize the correlation between an action and its outcomes in safety, progress, comfort, and rule compliance. To balance hard safety constraints against soft progress objectives, we introduce fine-grained per-timestep reward conditioning and reward noise augmentation. The generative formulation naturally supports controllable test-time sampling via reward guidance and anchored sampling, producing high-quality proposals. FlowR2A achieves state-of-the-art results on the NAVSIM v1 and v2 benchmarks, with multimodal proposals of substantially higher quality than prior methods.


FlowTrack: Controlling Edit-Signal Execution in Inversion-Free Flow Video Editing

Jielun Zhong ⋅ Di Wang ⋅ Jisheng Dang ⋅ Leigang Qu ⋅ Bimei Wang ⋅ Jingwen Zhao ⋅ Wencan Zhang ⋅ Hong Peng ⋅ Bin Hu

Text-guided video editing requires applying a semantic change while keeping source motion, layout, and unchanged content stable. Recent inversion-free flow editors derive effective edit signals from target-source velocity differences, yet typically inject these signals directly at each solver step. In neural video sampling, however, this direct execution rule overlooks an important distinction: the velocity difference is an online, state-dependent observation rather than a fixed endpoint-displacement command. As the target branch evolves under previous injections, direct injection may accumulate temporal fluctuations and off-target perturbations, producing motion drift, identity flicker, or unintended changes in weak-response regions. We propose FlowTrack, a causal execution controller for inversion-free flow video editing. FlowTrack leaves the pretrained generator and raw edit signal unchanged, but executes the signal through a carried displacement-like direction and response-aware attenuation. We derive this update as the closed-form solution of a local execution objective that balances current signal tracking, temporal carrying, and weak-response suppression. On the full FiVE-Bench protocol, FlowTrack improves motion fidelity, unedited-region fidelity, and structure consistency over strong training-free diffusion- and flow-based editors while maintaining competitive target-prompt alignment. Code is available at https://anonymous.4open.science/r/FlowTrack.


Flux: Online, Fine-Grained Data Scheduling for Training Machine Learning Interatomic Potentials

Yuanchang Zhou ⋅ Chen Wang ⋅ Hongtao Xu ⋅ Mingzhen Li ⋅ Guangming Tan ⋅ Weile Jia

Large-scale training of machine learning interatomic potentials (MLIPs) increasingly relies on data parallelism over heterogeneous atomistic data. In this setting, atom count captures an important part of the per-rank workload, but structures with similar atom counts can still induce different graph workloads through variations in graph count, cutoff edges, higher-order geometric features, and graph-collation overhead. This makes conventional atom-count batching and offline load tables both incomplete and inflexible, especially when the model architecture, cutoff radius, or data mixture changes. We present Flux, an online workload-aware scheduler for distributed MLIP training. Flux estimates structure-dependent workload signals from lightweight physical and geometric priors during data loading, calibrates them against runtime and memory costs, and forms balanced, memory-feasible local mini-batches across data-parallel ranks. Without requiring precomputed graph metadata, Flux can be used as a drop-in replacement for the standard distributed sampler, making it flexible and portable across training setups. Across large-scale atomistic datasets and three representative MLIP architectures, eSEN, MatRIS, and AllScAIP, Flux improves training throughput by up to 2.23--4.76$\times$ without degrading convergence.

Large Vision-Language Models (LVLMs) have achieved impressive progress in multimodal reasoning, yet they remain prone to object hallucinations, generating descriptions of objects that are not present in the input image. In this work, we investigate hallucination from the perspective of attention-value dynamics inside LVLM vision encoders. We identify a consistent three-phase structure of visual processing---diffusion, focus, and rediffusion---and show that the focus phase is where attention most clearly separates strongly and weakly supported visual tokens. However, low attention does not necessarily imply negligible downstream influence: low-attention tokens in the focus phase can still exert non-negligible value-side influence on the attention output relative to their small attention mass. Through controlled phase-wise value interventions, we find that hallucination behavior is particularly sensitive to the value content of low-attention tokens during this phase. Replacing or neutralizing these values reduces hallucination metrics while largely preserving grounded object evidence. A token-level teacher-forcing analysis further shows that the intervention reduces the probability of hallucinated object tokens with much smaller effects on ground-truth object tokens. In addition, Visual Attention Ratio (VAR) analysis shows that focus-phase intervention is accompanied by increased attention to visual tokens during decoding. Based on these observations, we instantiate a simple training-free inference-time intervention that replaces focus-phase low-attention values with an image-level mean value vector using statistics from a single forward pass. Experiments across multiple LVLM backbones demonstrate that this analysis-derived intervention reduces object hallucination with negligible additional runtime and remains compatible with existing inference-time mitigation methods. Our project page is available at: https://anonymous.4open.science/w/FocusMatters-7B32/


Focus on Where You Aggregate: Restricted SAM for Non-IID Federated Learning

Haotong Wen ⋅ Junkang Liu ⋅ Yi Xu ⋅ Xiao Liu ⋅ Longkun Guo ⋅ Kewen Liao

Data heterogeneity and multiple local update steps in federated learning can increase client drift and make global parameter aggregation less stable. Existing Sharpness Aware Minimization (SAM) based federated learning methods usually try to ease this problem by flattening local neighborhoods around client models. However, they do not explicitly optimize on the geometry region where the global parameters are actually aggregated. As a result, the aggregated global model often cannot fully benefit from local flattening. To address this problem, we propose a federated restricted SAM method called FedRSAM, which improves the stability of federated optimization by flattening the parameter aggregation region. First, we build a cone-shaped geometry region using the global update direction from the previous round and the variance of client parameter aggregation, in order to predict where the aggregated parameters may fall in the next round. Then, we restrict the search of SAM perturbation directions to the cone-manifold and only flatten the loss surface inside this region. Extensive experiments on multiple benchmarks and on both vision and natural language backbones verify the superiority of our method in terms of convergence stability and performance under heterogeneous settings.


FOGO: Forgetting-aware Orthogonalization Optimizer

Toan Nguyen ⋅ Yang Liu ⋅ Trung Le ⋅ Celso de Melo ⋅ Flora Salim

We argue that forgetting is not confined to continual learning but is a general optimization phenomenon: during standard training, dominant mini-batch gradients suppress rare but useful update directions, causing short-term forgetting at every step. When such knowledge is never revisited, these losses compound into long-term forgetting—the classical failure mode of continual learning. We introduce FOGO, a scalable optimizer that continuously detects and resolves gradient interference across both regimes. FOGO spectrally orthogonalizes momentum updates to prevent dominant directions from monopolizing optimization, then stores representative past directions in a compact codebook memory built on random projection, where pairwise distances are provably preserved in low-dimensional space. At each step, conflicts between the current update and stored directions are resolved via lightweight orthogonal correction and lifted back through a proximal step, with minimal overhead and no data storage. Across class-imbalanced classification, continual visual learning under domain and class shifts, continual fine-tuning of LLaVA-7B, and GPT-2 pretraining, FOGO consistently improves convergence and knowledge retention, outperforming Adam and Muon.


Follow the Mean: Reference-Guided Flow Matching

Pedro Curvo ⋅ Maksim Zhdanov ⋅ Floor Eijkelboom ⋅ Jan-Willem van de Meent

Existing approaches to controllable generation typically rely on fine-tuning, auxiliary networks, or test-time search. We show that flow matching admits a different control interface: adaptation through examples. For deterministic interpolants, the velocity field is solely governed by a conditional endpoint mean; shifting this mean shifts the flow itself. This yields a simple principle for controllable generation: steer a pretrained model by changing the reference set it follows. We instantiate this idea in two forms. Reference-Mean Guidance is training-free: it computes a closed-form endpoint-mean correction from a reference bank and applies it to a frozen FLUX.2-klein (4B) model, enabling control of color, identity, style, and structure while keeping the prompt, seed, and weights fixed. Semi-Parametric Guidance amortizes the same idea through an explicit mean anchor and learned residual refiner, matching unconditional DiT-B/4 quality on AFHQv2 while allowing the reference set to be swapped at inference time. These results point to a broader direction: generative models that adapt through data, not parameter updates.


Forgery Evidence Peaks Mid-Stack: Forensic Evidence Relay for Multimodal Forgery Detection

Yingxin Lai ⋅ xinyuan Wang ⋅ Yufei Liu ⋅ Jialin Guo ⋅ Kailin Lyu ⋅ Zhiming Luo ⋅ Shaozi Li

Multimodal large language models (MLLMs) are increasingly used for explainable image forgery analysis, where a shared decoder is expected to predict authenticity, localize evidence, and generate natural-language explanations. Existing methods typically rely on final decoder states, leaving open whether these late, language-oriented representations preserve localized manipulation cues. We test this assumption with a token- and layer-wise probing diagnostic that uses pixel-level masks only for analysis to separate tampered-region and authentic-region vision tokens. Across three MLLM-based forgery backbones and multiple manipulation types, we identify a mid-layer forgery evidence peak: tampered-region tokens are most linearly separable at upper-middle decoder layers and become less separable afterward, whereas authentic-region tokens continue to strengthen. This region-conditioned asymmetry suggests competition between localized manipulation cues and late semantic aggregation. Motivated by this diagnostic, we propose Gated Forensic Memory (\method), a lightweight evidence relay that reads from the diagnosed peak layer and writes re-encoded vision-token states back to late vision-token states through a zero-initialized gated residual. \method keeps the LLM frozen, adds no extra vision tokens, and requires no pixel-level training supervision. Across three backbones and seven benchmarks, \method improves authenticity prediction, explanation quality, and spatial alignment without mask supervision; the best read layer matches the diagnosed peak, turning the diagnostic into a practical layer-selection criterion.


Forgetting to Improve: Principled Data Removal in Active Learning

Manuel Wendl ⋅ Erik Englesson ⋅ Andreas Krause ⋅ Carl Henrik Ek

The uncertainty of a statistical model is most commonly factorized into an aleatoric and an epistemic part. This factorization changes how predictions are interpreted in downstream decision tasks. Importantly, except for the idealistic scenario with no model mismatch, the quantification is a characteristic of the model and not the data generating process. In this paper, we propose Forgetting to Improve a method that reduces this discrepancy by incorporating the task into the modeling framework. Our key insight is to acknowledge that in scenarios of model mismatch, data can have a detrimental effect on the modeling for a specific task. Based on this insight, we propose an influence function for Gaussian process models that allows for principled removal of detrimental data samples. We showcase the flexibility of this approach by demonstrating significant improvements across a range of tasks, including Bayesian optimization, model-based reinforcement learning, and transductive learning.


FracTS: Hierarchical and Autoregressive Time Series Generation

Jiayu Li ⋅ Umair Afzal ⋅ Zilong Zhao ⋅ Milad Abdollahzadeh ⋅ Uzair Javaid ⋅ Biplab Sikdar

Time series are ubiquitous across different domains, and the growing use of synthetic data for privacy-preserving sharing, data augmentation, and stress testing makes time series generation increasingly important. Early generative approaches are predominantly autoregressive, which can suffer from low sample diversity, error accumulation, and difficulty modeling long-range dependencies. In this paper, we introduce FracTS, a hierarchical, coarse-to-fine autoregressive model for multi-variate time series generation. Inspired by fractal generative modeling developed for images, FracTS adapts the paradigm from 2D to 1D sequences by learning a multi-scale representation that captures global trajectory structure at coarse scales and local dynamics at fine scales while following the autoregressive nature of time series. In contrast to unstructured images, time series are structured, allowing easy automated or designated temporal feature processing optimized for the model, and global (static) features extraction describing the instances that generate the corresponding time series in the collection. We incorporate them in the generation process to further improve the generation quality. The design naturally handles multivariate dependencies and improves long-horizon temporal coherence without requiring fixed-length inputs. Experiments on five datasets (two conditional, three unconditional) conducted against seven baselines show that FracTS consistently outperforms baselines across marginal distribution, correlation preservation, and ML utility. In particular, the ML utility is improved up to 98.1% comparing to the second-best baseline. Code is available at: https://anonymous.4open.science/r/fracts-91D621


FraudBench: A Legal Evaluation of AI Deception on Realistic Tasks

Kevin Wei ⋅ Sumaya N Adan ⋅ Stephan Llerena ⋅ Mark L Gitau ⋅ Jakob Merane ⋅ Anita Srinivasan ⋅ Colette Le Brannan ⋅ Keethan R. Kleiner ⋅ Alessandro Tacconelli ⋅ Michael T Tiu ⋅ Case Thomason ⋅ Kyle Emili ⋅ Natasza Gadowska ⋅ Alex Mark ⋅ Eric Ren ⋅ Umang Bhatt ⋅ Jatinder Singh ⋅ Jacy Anthis ⋅ Doni Bloomfield ⋅ Noam Kolt

We introduce FraudBench, an evaluation of the propensity of large language models (LLMs) to recommend, assist, or directly commit legal fraud, as measured by a novel dataset of 159 scenarios based on real-world U.S. court cases and validated by legal experts. As LLM use rises, models' ability to comply with the law becomes more important to ensure safety and public benefit. However, evaluating models' propensity to engage in fraud is challenging due to the need for specialized expertise, the contextual nature of legal standards, and the lack of high-quality evaluation data. \textsc{FraudBench} fills this gap with an evaluation framework for measuring fraud propensity, a dataset of 477 tasks (3 per scenario), and large-scale empirical results on 20 LLMs and over 300,000 samples. We also present a pre-registered human baseline ($n_{\texttt{annotations}} = 2957$, $n_{\texttt{humans}} = 996$). We find that (1) LLM fraud propensities are highly individualized, ranging from near 0\% to over 40\%; (2) LLMs with reasoning enabled behave less fraudulently ($\beta = -0.17$, log-odds scale); (3) LLM fraud propensity increases with task autonomy (6.4\% overall fraud rate when advising users, 26.9\% when acting directly); (4) LLMs are more likely to conceal than actively misrepresent information ($\beta = 0.38$); (5) humans and LLMs responded similarly to contextual cues, and differences in human-LLM fraud rates were not statistically significant; (6) system prompts can have substantial effects on fraud propensity across LLMs.

Fréchet regression, or conditional barycenters, is a flexible framework for modeling relationships between covariates (usually Euclidean) and response variables on general metric spaces, e.g., probability distributions or positive definite matrices. However, in contrast to classical barycenter problems, computing conditional counterparts in many non-Euclidean spaces remains an open challenge, as they yield non-convex optimization problems with an affine structure. In this work, we study the existence and computation of conditional barycenters, specifically in the space of symmetric positive definite (SPD) matrices with the Bures-Wasserstein metric, which connects SPD regression to optimal transport, avoids matrix logarithms, and preserves well-posedness under signed weights. We provide a sufficient condition for the existence of a minimizer of the conditional barycenter problem that characterizes the regression range of extrapolation. Moreover, we further characterize the optimization landscape, proving that under this condition, the objective is free of local maxima. Additionally, we develop a projection-free and provably correct algorithm for the approximate computation of first-order stationary points. Finally, we provide a stochastic reformulation that enables the use of off-the-shelf stochastic Riemannian optimization methods for large-scale setups. Numerical experiments validate the performance of the proposed methods on regression problems of real-world biological networks and on large-scale synthetic Diffusion Tensor Imaging problems.


Free Lunch for Pass@$k$? Low Cost Diverse Sampling for Diffusion Language Models

Sean Lamont ⋅ Christian Walder ⋅ Paul Montague ⋅ Amir Dezfouli ⋅ Michael Norrish

Diverse outputs in text generation are necessary for effective exploration in complex reasoning tasks, such as code generation and mathematical problem solving. Such Pass@$k$ problems benefit from distinct candidates covering the solution space. However, traditional sampling approaches often waste computational resources on repetitive failure modes. While Diffusion Language Models have emerged as a competitive alternative to the prevailing Autoregressive paradigm, they remain susceptible to this redundancy, with independent samples frequently collapsing into similar modes. To address this, we propose a training free, low cost intervention to enhance generative diversity in Diffusion Language Models. Our approach modifies intermediate samples in a batch sequentially, where each sample is repelled from the feature space of previous samples, actively penalising redundancy. Unlike prior methods that require retraining or beam search, our strategy incurs negligible computational overhead, while ensuring that each sample contributes a unique perspective to the batch. We evaluate our method on HumanEval and GSM8K using the LLaDA-8B-Instruct and Dream-v0-Instruct-7B models, where we demonstrate significantly improved diversity and Pass@$k$ performance across various sampling settings. As a simple modification to the sampling process, our method offers an immediate, low cost improvement for current and future Diffusion Language Models in tasks that benefit from diverse solution search.


FreeOcc: Decoupling Ego-Motion for Efficient 4D Occupancy Forecasting via Continuous Flow Matching

Zeping Zhang ⋅ Zhuoya Zhao ⋅ Samy Metari ⋅ Robert Laganiere

Current 4D occupancy world models use heavy cross-attention or deep recurrence to disentangle ego-motion from scene dynamics. This requires significant computational capacity. However, this effort is largely redundant since ego-motion is inherently known from odometry. We present FreeOcc, a lightweight continuous-time world model that introduces explicit $SE(2)$ geometric decoupling into a conditional flow-matching framework. By mathematically factoring out self-motion, we create a purely dynamic feature space. This space supports attention-free temporal fusion and allows physics-constrained ODEs to operate strictly on residual agent kinematics. To maximize deployment efficiency, FreeOcc introduces an uncertainty-aware blending mechanism. This routes well-observed voxels to a single-pass decoder, reserving the ODE solver exclusively for ambiguous, high-uncertainty regions. Requiring only 28.3M parameters—roughly 40\% of comparable models—our optimal ensemble mode (FreeOcc-E) runs at 27 FPS. On nuScenes, FreeOcc achieves state-of-the-art forecasting performance with a 0.47 m Near Field Chamfer Distance (NFCD) and 32.1\% mIoU. Furthermore, the model transfers zero-shot to SemanticKITTI. It also proves highly robust, retaining 92.8\% of its ideal mIoU under significant sensor packet loss and timestamp jitter.


FreqCa: Accelerating Image generation and editing via Frequency-Aware Caching

Jiacheng Liu ⋅ Peiliang Cai ⋅ Qinming Zhou ⋅ Yuqi Lin ⋅ Deyang Kong ⋅ Benhao Huang ⋅ Yupei Pan ⋅ Haowen Xu ⋅ Chang Zou ⋅ Junshu Tang ⋅ Shikang Zheng ⋅ Linfeng Zhang

The application of diffusion transformers is suffering from their significant inference costs. Recently, feature caching has been proposed to solve this problem by reusing features from previous timesteps, thereby skipping computation in future timesteps. However, previous feature caching assumes that features in adjacent timesteps are similar or continuous, which does not always hold in all settings. To investigate this, this paper begins with an analysis from the frequency domain, which reveal that different frequency bands in the features of diffusion models exhibit different dynamics across timesteps. Concretely, low-frequency components, which decide the structure of images, exhibit higher similarity but poor continuity. In contrast, the high-frequency bands, which decode the details of images, show significant continuity but poor similarity. These interesting observations motivate us to propose Frequency-aware Caching (FreqCa) which directly reuses features of low-frequency components based on their similarity, while using a second-order Hermite interpolator to predict the volatile high-frequency ones based on its continuity. Besides, we further propose to cache Cumulative Residual Feature (CRF) instead of the features in all the layers, which reduces the memory footprint of feature caching by 99%. Extensive experiments on FLUX.1-dev, FLUX.1-Kontext-dev, Qwen-Image, and Qwen-Image-Edit demonstrate its effectiveness in both generation and editing. Codes are available in the supplementary materials and will be released on GitHub.


From Benchmark to Adoption: Coding Agent Evaluation Needs User-Level Harness

Dasol Hong ⋅ Ji Yong Cho ⋅ Bumsoo Kang ⋅ Moontae Lee

Coding agents are increasingly adopted in real-world software development, yet their evaluation remains largely benchmark-driven and model-centric. Existing approaches focus on correctness and task completion, offering useful signals of isolated capability but limited insight into whether agents can be reliably used, controlled, and trusted in practice. We argue that the problem is not only how coding agents are evaluated, but what is being evaluated: coding agents are evaluated as models, while they are adopted as harnessed systems. This gap becomes most visible at the point of adoption. In real workflows, whether a coding agent is actually used is not determined by base-model capability alone. Instead, users make agents usable through user-level harnesses: situated configurations of task scope, context, constraints, checkpoints, workflow adaptations, and verification practices that shape and validate model behavior. These harnesses often remain informal and ad hoc, yet can determine whether the same model succeeds in one setting and fails in another. To ground this position, we draw on three empirical studies with users. Study 1 identifies contextual conditions shaping when users adopt or avoid coding agents. Study 2 identifies execution and control requirements for agentic systems. Study 3 identifies verification and code-quality criteria beyond correctness. Together, these studies suggest that users evaluate not only agent outputs, but also where agents can be used, how their behavior is guided, and whether their outputs are acceptable in practice. We interpret these conditions as user-level harnesses that make coding agents usable, controllable, and acceptable in real workflows. Based on these findings, we argue that evaluation should move beyond model-centric and system-internal metrics toward harness-aware, user-grounded assessment. In particular, evaluations must make explicit the conditions under which coding agents are usable, controllable, and trustworthy in practice—across dimensions such as environment, users, capabilities, and code quality. We call on the AI research community to place user-level harness at the center of coding-agent evaluation.


From Click Imitation to Transition Equivalence: Rethinking Supervision for GUI Agents

Zhiming Lin ⋅ Tianxiang Xu ⋅ zizhao zhang ⋅ Yixue Liu ⋅ Yumeng Ma ⋅ Yihao Zhong ⋅ ShaoYan He ⋅ CHEN RONGRU ⋅ Siyue Liu

GUI agents are increasingly expected to complete real tasks through language-based interaction, yet most supervised training still treats a single demonstrated click or action as the unique ground truth. This creates a mismatch between how agents are trained and how they are evaluated: task success depends on reaching the intended interface state, not on reproducing the demonstrator's exact surface action. We address this gap by introducing Transition-Equivalent Supervision (TES), a framework that trains GUI agents around goal-relevant state changes rather than raw action identity. TES converts demonstrated interactions into transition descriptors, supervises agents to predict the intended transition, and learns an executor that selects any verified action capable of realizing that transition. At inference time, the agent further verifies whether the executed action produces the intended state change, reducing repeated no-op or erroneous behavior. Across offline web prediction, executable web navigation, mobile control, and desktop-use benchmarks, TES consistently improves equivalent-step success, alternative-action recall, and online task completion while reducing NoOp/Error transitions. These results suggest that GUI-agent learning should move beyond click imitation toward transition-level reasoning, offering a more robust supervision principle for agents operating in diverse and dynamic interfaces.


From Facts to Personas: Interpretable Role Unlearning in LLMs via Mixture-of-Experts

Ruihong Zeng ⋅ Puning Yang ⋅ Jinghui Zhang ⋅ Shen Gao ⋅ Xiangliang Zhang ⋅ Preslav Nakov ⋅ Xiuying Chen

LLMs are widely adopted in multi-role dialogue and persona simulation. However, their strong ability to imitate role-specific linguistic and behavioral patterns can introduce privacy risks, reinforce social biases, and lead to unsafe or uncontrollable outputs. To address these issues, we introduce the Role-playing Unlearning task, which aims to selectively forget the style and knowledge associated with target roles while preserving the language ability and behaviors of the remaining roles. Unlike prior unlearning tasks that focus solely on removing factual information, our task additionally requires forgetting persona-specific characteristics, making the control over model generation more fine-grained. We propose ROLEMOE, a Mixture-of-Experts framework that enhances the separation of role-specific patterns by mapping different roles to more independent expert subspaces, providing structural interpretability. To reinforce this separation, ROLEMOE incorporates two complementary losses: a specialization loss that drives different roles to adopt different expert distributions, and a disentanglement loss that encourages different experts to develop distinct capability subspaces. This decomposed structure enables role unlearning to focus selectively on target-role experts. We finally propose RoleBench-Unlearn for evaluation, and experiments across multiple LLM architectures show that ROLEMOE achieves significantly more effective and controllable role forgetting, while better preserving the linguistic quality, knowledge, and persona consistency of non-target roles. All code and datasets are publicly available at: https://anonymous.4open.science/r/RoleMoE-6328.

Link prediction is a standard pretraining objective for graph neural networks: hide edges, train an encoder--decoder to predict them, and reuse the node embeddings for downstream node tasks. Despite its empirical success, there is little quantitative theory explaining why an edge-level objective should produce embeddings that are linearly useful for node classification.We study this question in the finite-sample stochastic block model (SBM). For any bounded GNN encoder paired with a symmetric bilinear decoder, we prove a conditional reduction: small population held-out LP excess risk implies small community-classification error for a transductive linear classifier on the learned embeddings.The bound decomposes into an optimization residual, an architectural approximation residual, and explicit statistical terms of order $\widetilde{O}(\sqrt{d/n})$ in bounded dense regimes. We further give a constructive few-label classifier: nearest-centroid classification under the PSD quadratic form $G = W Z^\top Z W$, whose label error is coupled to LP quality.The framework is architecture-agnostic through the approximation residual, and we instantiate this residual as $O(\log n / n)$ for a spectral-positional-encoding GIN with a bilinear decoder. Synthetic SBM experiments validate the theorem chain from LP risk to score error to classification error, and show when the worst-case constants become conservative.


From Local Skills to Long-Horizon Tasks: Progressive Skill Exploration for LLM Web Agents

Xueyang Feng ⋅ Xiaohe Bo ⋅ Xu Chen ⋅ Quanyu Dai ⋅ Feng Wen

Large language models (LLMs) are driving autonomous agents toward real-world environments, with the Web emerging as one of the most representative interaction scenarios. However, training Web agents with naive exploration remains challenging: successful trajectories are rare, and the sparse terminal rewards make it difficult for agents to obtain useful feedback from failed trajectories. Inspired by skill acquisition theory, we propose Skill-Composed Progressive Exploration (SCOPE), a short-to-long learning framework for long-horizon Web agents that decomposes complex task exploration into a gradual expansion process from local skills to complete task-solving capability. Specifically, SCOPE iteratively explores around local skills: Skill Discovery extracts reusable local skills and failure experiences from historical rollouts; Skill Orchestration composes discovered skills into high-level skill chains; andSkill Expansion leverages failure experiences aligned with the current skill chain to guide the agent toward exploring new subskills. Through this process, SCOPE progressively constructs task-completing skill chains across rollouts and internalizes them into the policy parameters through agentic on-policy distillation. Extensive experiments on WebArena, WebVoyager, and WebShop show that SCOPE consistently improves exploration dynamics and achieves 3.9\%--18.6\% relative gains over the strongest baselines in each setting. Our code is publicly available at https://anonymous.4open.science/r/SCOPE-14D3.


From Local to Global: Progressive Consensus via Hierarchical Communication in Multi-Agent Reinforcement Learning

Jiangjin Yin ⋅ Zhonglin Lv ⋅ Hangyu Mao ⋅ Rongbo Zhu ⋅ Yiwei Shi ⋅ Kun Hu ⋅ Xianfang Tang ⋅ Zhiwei Xu

Effective team collaboration hinges on consensus formation through information exchange, a principle equally critical in multi-agent reinforcement learning (MARL). However, existing communication-based and consensus-learning methods often struggle to form a coherent global understanding from agents' limited local views. We propose STAGE, a local-to-global progressive consensus framework based on hierarchical communication. By first grouping agents according to their perceptual focuses, STAGE enables intra-group communication to form local consensuses, then selects group leaders to exchange these consensuses across groups, progressively expanding them into a unified global consensus with reduced communication redundancy. To further improve consensus quality, we introduce KL-divergence constraints for consensus alignment and a variational autoencoder (VAE) objective for preserving task-relevant global information. Extensive experiments on challenging MARL benchmarks show that STAGE consistently outperforms state-of-the-art baselines, especially on more difficult tasks and larger-scale multi-agent systems.


From Pixels to Concepts: Do Segmentation Models Understand What They Segment?

Shuang Liang ⋅ Zeqing Wang ⋅ Yuxian LI ⋅ Xihui Liu ⋅ Han Wang

Segmentation is a fundamental vision task underlying numerous downstream applications. Recent promptable segmentation models, such as Segment Anything Model 3 (SAM3), extend segmentation from category-agnostic mask prediction to concept-guided localization conditioned on high-level textual prompts. However, existing benchmarks primarily evaluate mask accuracy or object presence, leaving unclear whether these models faithfully ground the queried concept or instead rely on visually salient but semantically misleading cues. We introduce CAFE: \textbf{C}ounterfactual \textbf{A}ttribute \textbf{F}actuality \textbf{E}valuation, a novel benchmark for evaluating concept-faithful segmentation in promptable segmentation models. Our \textbf{CAFE} is built on attribute-level counterfactual manipulation: the target region and ground-truth mask are preserved, while attributes such as surface appearance, context, or material composition are modified to introduce misleading semantic cues. The benchmark contains 2,146 paired test samples, each consisting of a target image, a ground-truth mask, a positive prompt, and a misleading negative prompt. These samples cover three counterfactual categories: Superficial Mimicry (\textbf{SM}), Context Conflict (\textbf{CC}), and Ontological Conflict (\textbf{OC}). We evaluate various model types and sizes on our CAFE. Experiments reveal a systematic gap between localization quality and concept discrimination: models often generate accurate masks even for misleading prompts, suggesting that strong mask prediction does not necessarily imply faithful semantic grounding. Our CAFE provides a controlled benchmark for diagnosing whether promptable segmentation models perform concept-faithful grounding rather than shortcut-driven mask retrieval.


From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving

Sining ANG ⋅ Yuguang Yang ⋅ Chenxu Dang ⋅ Canyu Chen ⋅ Cheng Chi ⋅ Liu Haiyan ⋅ Xuanyao Mao ⋅ jason bao ⋅ Xuliang ⋅ Bingchuan Sun ⋅ Yan Wang

Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbones, but how VLM representations differ from vision-only encoders after policy learning, and whether such differences matter for planning, remains unclear. Under a unified VLM-hidden + diffusion-policy paradigm, we compare multiple VLM families/scales (InternVL3 and Qwen3VL) with standard vision-only encoders (ResNet, ViT, and EVA-CLIP) while keeping the downstream planner fixed. We study representation, behavior, and system design. CKA/CCA and Shared--Unique SAE show that policy learning enlarges a common decision subspace, but both branches retain non-transferable residual factors. Latent-intervention policies and scenario-level analysis further show that these residuals are behaviorally meaningful: vision-only encoders are stronger in simple geometry-dominant scenes, whereas VLMs are more effective in semantically complex and interaction-heavy long-tail cases. The two branches also exhibit distinct progress--braking and path-choice tendencies, and an oracle best-of-two VLM+ViT selector reaches 93.58 PDMS on NAVSIM. We convert this complementarity into two lightweight systems: HybridDriveVLA, which selects from a compact cross-model candidate set using a learned trajectory scorer and improves PDMS from 90.80 to 92.10, and DualDriveVLA, a fast--slow variant that invokes the VLM in only 15\% of scenarios, achieving 91.00 PDMS with about $1.9\times$ lower latency than the VLM baseline. Code will be released.


From Sample to Subset Construction: Coverage-Aware Curation of Robot Demonstrations

Shuo Liu ⋅ Minxin Lai ⋅ Wuyang Zhang ⋅ Yu Zhang ⋅ Zhuangqiu Huang ⋅ Yao Li ⋅ Yanyong Zhang

Scaling robot demonstrations does not guarantee better imitation: redundant trajectories dilute learning, while rare, imperfect segments may encode critical transitions. Prior curation methods overlook such dependencies by scoring samples independently at the trajectory or segment level. We reframe robot demonstration curation as a coverage-aware sequential selection problem and propose PROSE, which selects data by maximizing marginal utility under a fixed budget. PROSE unifies three signals: (i) influence on closed-loop return, (ii) kinematic reliability, and (iii) coverage gain measuring novelty with respect to the retained set, allowing full trajectories and fine-grained segments jointly participate in scoring and filtering. Across three Robomimic simulation tasks and three real Franka tasks, PROSE achieves the best success rate on every task we evaluate against three state-of-the-art curators (CUPID, SCIZOR, Demo-SCORE), with average closed-loop gains of +15\% in simulation and +35\% in real-world. Ablations attribute the gain primarily to the coverage step.


From SGD to Muon: Adaptive Optimization via Schatten-p Norms

Thomas Massena ⋅ Corentin Friedrich ⋅ Mathieu Serrurier

Modern optimizers, like Muon, impose matrix-wise geometry constraints on their updates. These matrix-wise constraints can be unified under Linear Minimization Oracle (LMO) theory. However, all current methods impose fixed LMO geometries for the update rules, chosen by-design or empirically, which are not necessarily optimal according to the problem's geometry. We introduce a novel efficient data-driven criterion for dynamically choosing proxy-optimal update LMO geometries on individual Deep Neural Network layers. Derived in closed form from gradient and activation statistics using a single-step random feature regression surrogate model, our criterion navigates a design space interpolating from SGD to Muon updates. Moreover, integrating parameter-wise preconditioning allows our framework to recover SGD, Muon, Adam, and MuAdam as specific extrema. To make this adaptive approach scalable, we pair it with efficient computational strategies, achieving only a $\sim 3$% runtime overhead on highly optimized baselines. As a proof of concept, we show that this data-driven optimizer matches or exceeds the performance of the best performing optimizer between Muon and AdamW across three different training scenarios. Ultimately, this work provides evidence that LMO geometry can be successfully and efficiently adapted from runtime data, opening a new pathway for optimizer design beyond static geometries.


From Walls to Synergy: A Joint LLM-Evolution Framework for MILP Solvers

Ziao Guo ⋅ Yuan Feng ⋅ Chenhao Ying ⋅ Junchi Yan

Machine learning (ML) methods have achieved notable success in enhancing individual components of mixed-integer linear programming (MILP) solvers. However, most existing approaches focus on single-component enhancement, neglecting critical interactions among components and often yielding diminishing returns when multiple approaches are integrated. The emergence of large language models (LLMs) provides a promising paradigm for jointly optimizing multiple solver components within a unified framework. In this paper, we propose Jolly-MILP, an LLM-guided framework to jointly optimize presolving, cut selection, and branching variable selection in exact MILP solvers. Our approach introduces a synergistic Coordinator to align component-level generators and a hierarchical Controller to produce adaptive, multi-stage strategies. Extensive experiments on nine MILP datasets demonstrate that Jolly-MILP consistently outperforms both ML-based approaches and existing LLM-based methods. These results show that explicit joint optimization, rather than isolated component-wise enhancement, is crucial for acquiring further gains in exact MILP solving, offering a new pathway for more advanced solver design.

Autoregressive large language models have dominated generative recommendation by sequentially predicting the next item. However, this unidirectional paradigm suffers from error propagation, and limited ability to jointly optimize an entire recommendation list. In this paper, we introduce Full-Sequence Masked Diffusion for generative recommendation (FSMD), which extends masked diffusion models to enable global joint modeling over the entire user history and Top-$K$ recommendation list. We formulate recommendation as a unified sequence diffusion process, where the forward process randomly masks tokens across the full sequence and the reverse process simultaneously denoises all masked positions using bidirectional context. This formulation enables parallel generation of the entire recommendation list and naturally captures inter-item dependencies within a coherent global structure. To better align generation with ranking objectives, we further introduce a unified training objective that combines masked diffusion likelihood with a listwise ranking loss. Extensive experiments on multiple public benchmarks demonstrate that FSMD consistently outperforms state-of-the-art autoregressive and item-level diffusion-based generative recommendation methods across accuracy and list quality metrics.


G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Junxian Li ⋅ Kai Liu ⋅ Zizhong Ding ⋅ Zhixin Wang ⋅ Zhikai Chen ⋅ Renjing Pei ⋅ Yulun Zhang

The development of separate-encoder Unified multimodal models (UMMs) comes with a rapidly growing inference cost due to dense visual token processing. In this paper, we focus on understanding-side visual token reduction for improving the efficiency of separate-encoder UMMs. While this topic has been widely studied for MLLMs, existing methods typically rely on attention scores, text-image similarity and so on, implicitly assuming that the final objective is discriminative reasoning. This assumption does not hold for UMMs, where understanding-side visual tokens must also preserve the model’s capabilities for editing images. We propose G$^2$TR, a generation-guided visual token reduction framework for separate-encoder UMMs. Our key insight is that the generation branch provides a task-agnostic signal for identifying understanding-side visual tokens that are not only semantically relevant but also important for latent-space image reconstruction and generation. G$^2$TR estimates token importance from consistency with VAE latent, performs balanced token selection, and merges redundant tokens into retained representatives to reduce information loss. The method is training-free, plug-and-play, and applied only after the understanding encoding stage, making it compatible with existing UMM inference pipelines. Experiments on image understanding and editing benchmarks show that G$^2$TR substantially reduces visual tokens and prefill computation by **1.94$\times$** while maintaining both reasoning accuracy and editing quality, outperforming baselines on almost all benchmarks.


GARDO: Reinforcing Diffusion Models without Reward Hacking

Haoran He ⋅ Yuxiao YE ⋅ Jie Liu ⋅ Jiajun Liang ⋅ Zhiyong Wang ⋅ Ziyang Yuan ⋅ Xintao Wang ⋅ Hangyu Mao ⋅ Meng Wang ⋅ Pengfei Wan ⋅ Ling Pan

Fine-tuning diffusion models via online reinforcement learning (RL) has shown great potential for enhancing text-to-image alignment. However, since precisely specifying a ground-truth objective for visual tasks remains challenging, the models are often optimized using a proxy reward that only partially captures the true goal. This mismatch often leads to reward hacking, where proxy scores increase while real image quality deteriorates and generation diversity collapses. While common solutions add regularization against the reference policy to prevent reward hacking, they compromise sample efficiency and impede the exploration of novel, high-reward regions, as the reference policy is usually sub-optimal. To address the competing demands of sample efficiency, effective exploration, and mitigation of reward hacking, we propose \textbf{G}ated and \textbf{A}daptive \textbf{R}egularization with \textbf{D}iversity-aware \textbf{O}ptimization (\textbf{GARDO}), a versatile framework compatible with various RL algorithms. Our key insight is that regularization need not be applied universally; instead, it is highly effective to selectively penalize a subset of samples that exhibit high uncertainty. To address the exploration challenge, GARDO introduces an adaptive regularization mechanism wherein the reference model is periodically updated to match the capabilities of the online policy, ensuring a relevant regularization target. To address the mode collapse issue in RL, GARDO amplifies the rewards for high-quality samples that also exhibit high diversity, encouraging mode coverage without destabilizing the optimization process. Extensive experiments across diverse proxy rewards and hold-out unseen metrics consistently show that GARDO mitigates reward hacking and enhances generation diversity without sacrificing sample efficiency or exploration, highlighting its effectiveness and robustness.


GATE-AD: Graph Attention Network Encoding for Few-Shot Industrial Visual Anomaly Detection

ANGELOS PSYRRIS ⋅ Yannis Panagakis ⋅ Maria Vakalopoulou ⋅ Georgios Th. Papadopoulos

Few-Shot Industrial Visual Anomaly Detection (FS-IVAD) is a critical task in modern manufacturing, where automated product inspection systems must identify rare defects using only a handful of normal, defect-free training samples. This paper introduces GATE-AD, a reconstruction-based framework that casts few-shot normality modeling as a masked, representation-aligned graph reconstruction problem on a $k$-nearest-neighbor graph, built from frozen self-supervised ViT patch tokens. Since defects typically disrupt the contextual consistency between a patch and its spatial neighbors, GATE-AD emphasizes such neighborhood relations by attending anisotropically over each patch's local neighborhood using a Graph Attention Network (GAT) encoder. To prevent GAT over-smoothing in the low-shot setting, the encoder output is aligned with the input ViT tokens through a learnable latent space, where reconstruction inconsistency is scored with a Scaled Cosine Error (SCE) objective. On the MVTec AD, VisA, and MPDD benchmarks, GATE-AD attains state-of-the-art image-level AUROC in the wide majority of $1$- to $8$-shot settings, with low per-image inference cost that does not scale with the size of the support set. Code is included in the supplement material and will be publicly released upon acceptance.


Gaussian Mixture Models in Hilbert Spaces via Kernel Methods

Daniel López Montero ⋅ Antonio Álvarez-López ⋅ Marcos Matabuena

Modern datasets across many disciplines increasingly consist of time-evolving, potentially infinite-dimensional random objects, such as dynamic functional data, which are naturally modeled in Hilbert spaces. In these settings, characterizing probability measures, for example, through densities, can be ill-defined or technically challenging. Motivated by clustering applications, we propose a Gaussian mixture framework for Hilbert-space-valued data based on kernel mean embeddings and develop efficient optimization algorithms for estimation. We establish theoretical guarantees showing that the proposed algorithm is well defined and that the model yields a dense class of approximations in infinite-dimensional spaces. We evaluate the framework through extensive experiments on diverse structures and data geometries, including $L^2$-functional data and random graphs in Laplacian spaces arising in modern medical applications.


GCD: GCM-consistent Diffusion for Zero-shot Downscaling across Heterogeneous GCMs

Ruian Tie ⋅ Wenbo Xiong ⋅ Zhengyu Shi ⋅ Xinyu Su ⋅ 江晨雨 ⋅ Libo Wu ⋅ Hao Li

Climate downscaling of global General Circulation Models (GCMs) is a challenging ill-posed inverse problem. In practice, the lack of paired data makes end-to-end supervised training infeasible, while downstream assessments require processing large, heterogeneous GCM ensembles under strict computational limits. Consequently, developing an efficient, zero-shot downscaling approach is crucial. While vanilla Diffusion Posterior Sampling has recently shown immense potential by bypassing paired training, it remains bottlenecked by geographic misalignments, spectral biases, and prohibitive inference costs. To address these limitations, we propose the GCM-consistent Diffusion for Zero-shot Downscaling (GCD) framework. First, GCD conditions the diffusion prior on static geographic boundaries to ensure physical fidelity. Second, we introduce a filter measurement operator, effectively mitigating spectral discrepancies between GCMs and the prior to provide stable inference-time guidance. Third, to overcome the inference bottleneck, we design a joint acceleration scheme: by combining progressive distillation with Conjugate Gradient guidance to correct large-step deviations, GCD compresses the generation process to just 6 steps. Extensive experiments demonstrate that GCD achieves highly competitive 99th percentile accuracy across five heterogeneous GCMs and successfully recovers the high-frequency details of tropical cyclones, entirely without model-specific retraining.


Generative Scenario Rollouts for End-to-End Autonomous Driving

Rajeev Yasarla ⋅ Deepti Hegde ⋅ Shizhong Han ⋅ Hsin-Pai Cheng ⋅ Yunxiao Shi ⋅ MeysamSadeghi ⋅ Shweta Mahajan ⋅ Apratim Bhattacharyya ⋅ Litian Liu ⋅ Risheek Garrepalli ⋅ Thomas Svantesson ⋅ Mohammad Ghavamzadeh ⋅ Fatih Porikli ⋅ Herbert Cai

Vision-Language-Action (VLA) models are emerging as highly effective planning models for end-to-end autonomous driving systems. However, current works mostly rely on imitation learning from sparse trajectory annotations and under-utilize their potential as generative models. We propose Generative Scenario Rollouts (GeRo), a plug-and-play framework for VLA models that jointly performs planning and generation of language-grounded future traffic scenes through an autoregressive rollout strategy. First, a VLA model is trained to encode ego vehicle and agent dynamics into latent tokens under supervision from planning, motion, and language tasks, facilitating text-aligned generation. Next, GeRo performs language-conditioned autoregressive generation. Given multi-view images, a scenario description, and ego-action questions, it generates future latent tokens and textual responses to guide long-horizon rollouts. A rollout-consistency loss stabilizes predictions using ground truth or pseudo-labels, mitigating drift and preserving text-action alignment. This design enables GeRo to perform temporally consistent, language-grounded rollouts that support long-horizon reasoning and multi-agent planning. On Bench2Drive, GeRo improves driving score and success rate by +15.7 and +26.2, respectively. By integrating reinforcement learning with generative rollouts, GeRo achieves state-of-the-art closed-loop and open-loop performance, demonstrating strong zero-shot robustness. These results highlight the promise of generative, language-conditioned reasoning as a foundation for safer and more interpretable end-to-end autonomous driving.


GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation

Sixiang Chen ⋅ Zhaohu Xing ⋅ Tian Ye ⋅ Xinyu Geng ⋅ Yunlong Lin ⋅ Jianyu Lai ⋅ Xuanhua He ⋅ Fuxiang Zhai ⋅ Jialin Gao ⋅ Lei Zhu

Open-ended image generation is no longer a simple prompt-to-image problem. High-quality generation often requires an agent to combine a model's internal generative ability with external resources. As requests become more diverse and demanding, we aim to develop a general image-generation agent that can self-evolve through trajectories and use tools more effectively across varied generation challenges. To this end, we propose GenEvolve, a self-evolving framework based on Tool-Orchestrated Visual Experience Distillation. In GenEvolve, each generation attempt is modeled as a tool-orchestrated trajectory, where the agent gathers evidence, selects references, invokes generation skills, and composes them into a prompt-reference program. Unlike existing agentic generation methods that mainly rely on image-level scalar rewards, GenEvolve compares multiple trajectories for the same request and abstracts best-worst differences into structured visual experience, provided only to a privileged teacher branch. Inspired by on-policy self-distillation, Visual Experience Distillation provides dense token-level supervision, helping the student internalize better search, knowledge activation, reference selection, and prompt construction. We further construct GenEvolve-Data and GenEvolve-Bench. Experiments on public benchmarks and GenEvolve-Bench show substantial gains over strong baselines, achieving state-of-the-art performance among current image-generation frameworks.


GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

Katharina Sophie Schmid ⋅ Nicolas von Lützow ⋅ Angela Dai ⋅ Jozef Hladký ⋅ Matthias Niessner

We introduce a new approach to high-fidelity 3D scene reconstruction from multi-view RGB images that tightly couples reconstruction with a strong generative 3D prior. We cast scene reconstruction as conditional 3D generation over a set of spatially-localized, overlapping chunks that together tile the scene, scaling generation to large scene extents. Crucially, we inherit the fidelity and completeness of state-of-the-art generative shape models -- we use Trellis.2 as an example -- which we generalize to the scene level. To this end, we propose a projection-based conditioning mechanism that lifts posed multi-view image features into a coherent 3D representation aligned with the generative model, independent of view ordering and spatially anchored to the scene, yielding high-fidelity, multi-view consistent generated geometry. This enables lifting the strong object-level prior of Trellis.2 to multi-view, scene-scale generation, producing faithful, editable PBR mesh reconstructions of indoor environments. As a result, we obtain high-fidelity results that outperform cutting-edge reconstruction methods by 16\%.

We present GenZ, a hybrid model that turns a foundational model (FM) into a {\bf knowledge-discovery engine} explaining variation in real-valued multidimensional targets. The FM is treated as a noisy oracle that can answer yes/no questions about a semantic item $s$ (text, image, or both) and GenZ learns {\bf which questions to ask} so that the answers $\bz$ explain the statistical link to a possibly high-dimensional target $\by$. Discovery is driven by {\bf group reasoning}: at each step, the model partitions items by the posterior of the current latent features, and asks the FM to articulate the semantic commonality that explains the divide. The resulting feature descriptors $\theta_f$ are the primary product---human-readable hypotheses about why $s$ predicts $\by$---while predictive accuracy is a secondary, but consistently strong, by-product (e.g.~beating both the 0-shot FM baseline and TabPFN on LLM embeddings of the same items in hedonic price regression). Across four domains---hedonic house pricing, Netflix cold-start recommendations, Arizona species ecology, and human visual-cortex fMRI---GenZ recovers dataset-specific structure that the FM does not surface from priors alone. Holding the items $s$ fixed and varying only the target $\by$ (Arizona species under spatial/taxonomic/functional targets; the same natural images under FFA/EBA/PPA brain responses) yields {\bf qualitatively different} discovered feature sets, demonstrating that GenZ characterizes the $s\!\to\!\by$ link rather than the marginal distribution of $s$.


Geometric Latent Reasoning Induces Shorter Generations in LLMs

Shashi Kumar ⋅ Yacouba Kaloga ⋅ Petr Motlicek ⋅ Ina Kodrasi ⋅ Andrea Cavallaro

Large language models solve complex problems by generating lengthy chains of explicit reasoning tokens. While effective, this makes reasoning expensive, length-sensitive, and constrained to (discrete) natural language. While latent reasoning offers a continuous alternative, determining useful structures for intermediate latent states is an open challenge. In this paper, we formulate latent reasoning as a geometric path-approximation problem within the model’s pretrained token-embedding space. We introduce Geometric Latent Reasoning (GLR), which uses a lightweight transition head to predict iterative direction updates in embedding space. Using textual chain-of-thought traces as anchors, GLR learns to approximate discrete reasoning trajectories while permitting continuous deviations from exact token embeddings. Evaluations on mathematical reasoning benchmarks using Qwen3 models reveal an emergent phenomenon: geometric latent reasoning induces substantially shorter generations without an explicit length objective. By replacing early explicit reasoning with continuous latent steps, models often reach correct answers using substantially fewer total generation steps. These findings suggest that continuous trajectories act as compact intermediate reasoning states, exposing a new tradeoff between latent computation budget, output length, and accuracy.

Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our approach identifies structurally compatible projections and learns a shared representation that better preserves each layer’s distinct calibration geometry. Coupled with structured sparsity, this yields highly efficient weight decompositions without sacrificing functional fidelity. Across diverse architectures, scales, and modalities, our method achieves state-of-the-art results, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations, which operate under heuristic grouping strategies. By replacing heuristic engineering strategies with a convergent, optimization-driven pipeline, we establish a theoretically grounded foundation for scalable, transformer compression across different modalities.


GEOPHYS: The Geometry of Physical Plausibility

Christian Internò ⋅ Alexander Pondaven ⋅ Habon Issa ⋅ Fabio Pizzati ⋅ Francesco Pinto ⋅ Markus Olhofer ⋅ Ivan Laptev ⋅ Philip Torr ⋅ Eero Simoncelli ⋅ Barbara Hammer ⋅ David Klindt

Whether a video is physically plausible is a question current generative models answer poorly and current evaluators answer expensively, through physics-targeted training, billion-parameter video world models, or multimodal-language model judges. We ask whether frozen images vision encoders, with no video or physics supervision, already contain a useful signal for the same question. Across four frozen backbones, self-supervised vision transformers (DINOv2, DINOv3) and biologically-inspired models of the ventral stream (CORnet-S, VOneNet), a small set of geometric signals (curvature, speed variation, acceleration, prediction residual) on per-frame feature trajectories separates plausible from physics violated videos with no retraining. We propose GEOPHYS, a training-free framework that 11 reaches 97.7% on LikePhys and 93.3% on IntPhys2 accuracy for physics-violation detection, surpassing V-JEPA 2, GPT-4o, Gemini, and twelve modern video diffusion models. The same signals track human EEG responses to object-permanence violations and scale with object number. Deployed unchanged as a best-of-N reward during video generation, GEOPHYS lifts MAGI-1 4.5B from 53.8\% to 64.7\% PhysicsIQ score at 5.8× lower wall-clock and 7.8× lower memory than V-JEPA 2 world-model rewards. A useful proxy for physical plausibility is already encoded as a geometric regularity of natural-video statistics in vision features.

Georeferenced 3D reconstruction from ground images requires recovering not only the local scene geometry, but also localizing every camera center and 3D point in a geographic coordinate frame. Existing cross-view geo-localization methods are able to estimate the GPS coordinate of a single ground camera, but do not reconstruct dense 3D geometry, while recent feed-forward multi-view 3D reconstruction models predict point maps and camera poses only in a local coordinate frame. We introduce X-VGGT, a feed-forward cross-view geometry transformer that unifies these tasks by jointly processing a set of ground-level images and a georeferenced satellite image to predict ground-view camera poses, dense 3D point maps, and a scene-level transform parameterized by gravity, scale, yaw and translation. This transform projects the local reconstruction into the satellite plane, resulting in the geo-registration of all ground images with the georeferenced satellite image. We further propose an evaluation protocol for this joint task and show that, across multiple cross-view datasets, X-VGGT produces accurate georeferenced reconstructions and outperforms prior cross-view geo-localization and multi-view 3D reconstruction models.


GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning

Jinhao Jing ⋅ Zheng Ma ⋅ Jinwei Liang ⋅ Qiannian Zhao ⋅ Shuang Chen ⋅ Jing Yang ⋅ Prayag Tiwari ⋅ Jingjing Bai ⋅ Benyou Wang ⋅ Por L Yee ⋅ Zhan Su ⋅ Lewei Lu

Large Multimodal Models (LMMs) often struggle with geometric reasoning due to visual hallucinations and a lack of mathematically precise Chain-of-Thought (CoT) data. To address this, we propose the GeoSym Engine, an automated and scalable neuro-symbolic framework. By leveraging a type-conditional grammar and an analytic SymGT Solver, it derives exact symbolic ground truths and seamlessly integrates with a robust rendering pipeline to produce high-precision geometric diagrams. Using this engine, we construct GeoSym127K, a difficulty-stratified dataset featuring 51K high-resolution images, 127K questions with symbolic ground truths, and 55K answer-verified CoT QA pairs. We also introduce GeoSym-Bench, an expert-curated suite of 511 complex samples for rigorous evaluation. Through extensive supervised fine-tuning (SFT), we demonstrate that GeoSym drives concentrated improvements specifically on diagram-dependent and multi-step geometry tasks. Our Qwen3-VL-8B model gains an absolute +22.21% on the MathVerse Vision-Only subset and reaches 61.52% (+6.19% improvement) on WeMath, mitigating long-horizon logic fragmentation and outperforming advanced closed-source models like Doubao-1.8. Furthermore, applying Reinforcement Learning with Verifiable Rewards (RLVR) via GRPO reveals that initializing from structural SFT checkpoints substantially elevates the performance ceiling over zero-shot RL. Driven by deterministic exact-match signals, this showcases the robust scaling potential of our verifiable reasoning synthesis. Datasets and code are available at https://huggingface.co/datasets/Tomie0506/GeoSym127K and https://github.com/Tomie56/GeoSym127K


Get a GRIP, this will be a long TRIP: A Quantifiable Long-Range Framework for Verifying Over-squashing

Ferran Hernandez Caralt ⋅ Simon Heilig ⋅ Adrián Bazaga ⋅ Asja Fischer ⋅ Moshe Eliasof ⋅ Pietro Lió

Empirical claims about the connection between over-squashing and long-range interactions in GNNs, can only be trusted if the benchmarks used to validate them genuinely require long-range interactions. The de-facto standard, the Long Range Graph Benchmark, has been repeatedly shown to be saturated by tuned short-range models, with existing synthetic alternatives being tied to specific topologies. As such, there is a lack of principled certificate of long-rangedness on *arbitrary* graphs. This state reflects the absence of a precise characterization of long-ranged benchmarks. We address this fundamental gap by introducing four verifiable *axioms*: *Predictability*, *Tightness*, *Strictly $k$-Range*, and *Topology-Invariance*, that any task claiming to test $k$-hop interactions must satisfy. We formally prove that violating any one of them admits failure modes that undermine conclusions drawn from the task. Based on these axioms, we introduce TRIP (*Truly Ranged Interactions Problem*) and its generalisation GRIP (*Generally Ranged Interactions Problem*), constructive procedures that turn *any* graph into a provably long-ranged task by drawing features from *stable* distributions. Moreover, by construction, GRIP admits a closed-form, per-range Maximum-Likelihood oracle that yields the first *a priori per-range* lower bound on test error available on any benchmark. Using our framework, we: (i) audit 4 common long-range benchmarks and identify their failures modes with respect to our axioms; (ii) on TRIP-instantiated topologies, we find a popular notion of curvature is uncorrelated with GNN performance, supporting topological-vs-computational bottleneck distinction; and (iii) we show that a novel benchmark's over-squashing measures factors beyond pure long-rangedness. Code to use the framework and reproduce experiments is released anonymously https://anonymous.4open.science/r/graph-grip-7F1A/.

Streaming 3D reconstruction from long monocular video sequences requires maintaining a key-value(KV) cache that grows linearly with sequence length, creating a severe memory bottleneck. Existing approaches either truncate the cache to a fixed set of anchor frames, leading to reconstruction quality degradation, or rely on attention-score heuristics that are agnostic to 3D scene structure,failing to preserve geometrically valuable tokens. To address these problems, we present GHOST (Geometry-Hierarchical Online Streaming Token Eviction), a training-free KV cache management framework that exploits the model's own 3D geometry outputs to evict redundant tokens online. GHOST introduces three mutually reinforcing innovations: a hierarchical dual-level importance scoring scheme, a privilege mechanism that protects special tokens from eviction, and a cosine-similarity-guided layer-wise budget allocation. Experiments on various benchmarks show that GHOST preserves excellent reconstruction quality while cutting the KV cache by nearly half and delivering $1.75\times$ faster inference compared to state-of-the-art methods. Our code and model will be available to the public.


GIFT: Representation Geometry Matters for Single-Domain Generalized Object Detection

Yin Zhang ⋅ Yaoyue Zheng ⋅ Yongqiang Zhang ⋅ Bogdan Raducanu ⋅ Dan Liu

Single-Domain Generalized Object Detection (S-DGOD) aims to generalize from a single source domain to unseen domains under distribution shifts. Existing efforts mainly focus on simulating unseen domains or suppressing domain-specific factors. However, they largely overlook the intrinsic generalization capability embedded in Vision Foundation Models (VFMs). In this work, we show that unconstrained fine-tuning disrupts the geometric structure of pre-trained representations, which encodes domain-invariant semantics. Motivated by this, we propose Geometric-Invariance Fine-Tuning (GIFT), a geometric-aware fine-tuning method for VFMs. Specifically, we introduce a series of learnable Householder reflections to the pre-trained weights to encourage geometry-consistent adaptation under domain shift. In addition, we propose a lightweight channel-wise scaling to facilitate effective task-specific adaptation. Extensive experiments demonstrate that GIFT consistently outperforms state-of-the-art (SOTA) approaches on two S-DGOD benchmarks. Notably, by introducing fewer than 1% additional trainable parameters into the frozen backbone, the proposed method achieves improvements of +7.9% and +2.9% in mPC over the previous SOTA on the Cityscapes-C and Diverse Weather Dataset, respectively. This finding highlights that preserving representation geometry is crucial for robust cross-domain generalization, establishing GIFT as a principled and efficient solution to domain shift. The code is available in the supplementary material.


GLINT: Sparsely Gated Vision-Language Alignment for Fine-Grained Radiology Representations

Jonggwon Park ⋅ Seongeun Lee ⋅ Junhyun Park ⋅ Hannah Yun ⋅ Hyunwoong Kim ⋅ Sohyun Jeong ⋅ Hyewon Kang ⋅ Byungmu Yoon ⋅ Kyoyun Choi

Vision-language models (VLMs) for radiology have emerged as a scalable paradigm by leveraging image-report pairs naturally produced in clinical workflows. However, this pairing reveals a mismatch in scale: each finding occupies only a small region of the image, yet supervision is provided only at the global image-report level. This poses a central challenge: prior approaches spread weight densely across all patches rather than concentrating on the sparse subset relevant to a given query. To address this, we present GLINT (Gated Language-Image alignmeNT), a framework that explicitly models this sparse correspondence. On the alignment side, we introduce Sparsely Gated Alignment, a novel architecture in which a sigmoid gate over a separate gate embedding space activates only the patches relevant to each textual query, enforcing explicit sparsity. On the representation side, we add Dense Feature Regularization, which anchors the trainable encoder's intermediate features to a frozen self-supervised learning (SSL) teacher, preserving the fine-grained patch features that the gate relies on. The same recipe applies to both 2D chest X-ray (CXR) and 3D chest computed tomography (CT), built with DINOv3 and V-JEPA 2.1, respectively. GLINT enables zero-shot classification, grounding, and segmentation from free-text queries, and to our knowledge is the first to demonstrate zero-shot segmentation on 3D CT volumes without mask supervision. Notably, the most pronounced gains arise on zero-shot grounding and segmentation, where sparse, query-specific localization is required, consistent with our design intent. In downstream evaluation, GLINT outperforms both SSL encoders and medical VLMs on classification, report generation, and segmentation.

Most existing memory-enhanced Large Language Model (LLM) approaches implicitly assume that memory validity can be established either through external evaluators that provide task-specific success signals or through internal model cognition, such as reflection, for editing memory entries. However, these assumptions often break down in practical environments with dynamic drifts. We propose the Global Verifier (GLOVE), a plug-and-play module for LLM memory systems that establishes a relative notion of truth to achieve memory-environment realignment in the face of environmental drifts. Through active probing to detect inconsistencies between retrieved memories and fresh observations, GLOVE enables memory-environment realignment by verifying and updating memory without access to task-specific ground-truth supervision or strong reliance on model introspection. We evaluate GLOVE across a spectrum of tasks ranging from web navigation and discrete planning to continuous and embodied control, covering both simulated benchmarks and real-world robotic deployment. Across all domains, GLOVE improves adaptation across various LLM memory designs under both explicit and implicit drift, often recovering performance from near-failure to high success rates. These results suggest a practical pathway to memory-environment realignment for self-evolving cognitive agents.

Glycosylation is among the most diverse and important post-translational modifications in biology, governing immunogenicity, self-recognition, and the clinical viability of biologic drugs. Glycans cannot be sequenced from a template, and de novo structure prediction from tandem mass spectrometry (MS/MS) remains a longstanding bottleneck. The current state of the art, GlycoBART, is a $207$M autoregressive transformer whose $O(\mathrm{beam} \cdot \mathrm{output\_length})$ decoder passes preclude real-time annotation. We introduce $\textbf{GlycoGen}$, a $50$M Discrete Flow Matching (DFM) model that predicts glycan structure from MS/MS spectra by direct, non-autoregressive generation in a few parallel forward passes. We pair it at inference time with $\textbf{Crystallizing Flows}$, a deterministic sampler that propagates the per-token distribution on the probability simplex and freezes each token to a one-hot (*crystallizes* it) the moment its argmax leaves the mask, exiting early once all tokens are unmasked. The sampler feeds the denoiser mixture embeddings, and runs out-of-the-box with no retraining. GlycoGen + Crystallizing Flows reach $\mathbf{0.9266}$ top-1 structural accuracy on the CandyCrunch test set, three orders of magnitude faster than GlycoBART, Pareto-dominating the prior state of the art on *both* accuracy *and* speed. We extend our experiments to a matched-architecture MDLM denoiser (Sahoo et al. 2024) and observe a comparable lift: the sampler generalizes from continuous-time flow matching to discrete-time masked diffusion. To our knowledge, this is a new state of the art for de novo open-vocabulary glycan structure prediction, and the first system fast enough for accurate real-time inline annotation on modern LC-MS/MS instruments.


GOLIATH: Gradient Inversion of Tabular Diffusion Models

Giulio Segalini ⋅ Aditya Shankar ⋅ Jérémie Decouchant ⋅ Lydia Chen

Gradient inversion reconstructs training data from shared gradients. However, existing methods target classification tasks and image diffusion, and do not address mixed tabular diffusion, where numerical columns follow Gaussian noising while categorical columns are corrupted by discrete masks. This mixed continuous-discrete training objective changes both the inversion target and the gradient structure: an attacker must recover not only rows, timesteps, and noise, but also the latent mask pattern whose density is tied to the diffusion timestep. We propose Goliath, the first gradient inversion attack designed for tabular diffusion models. Goliath employs a cyclic inversion loop that jointly recovers rows and diffusion latents across Gaussian and mixed diffusion regimes. We introduce a tabular-aware objective function that balances numerical and categorical gradient contributions while improving categorical recovery through enforced consistency between noise levels and mask structures. To accommodate large batches, we aggregate multiple reconstructions across epochs using a row-alignment heuristic. Evaluated across nine tabular datasets, Goliath achieves up to $88.9\\%$ per-cell reconstruction accuracy and consistently outperforms gradient-inversion baselines across Gaussian and mixed tabular diffusion. Code is available at https://anonymous.4open.science/r/fl-tab-diffusion-inversion-F4E7/README.md.


G-PAC: Constructing Cohesive Pseudo-Features for Generalizable Physical Adversarial Camouflage

tianrui lou ⋅ Haoqing Zhang ⋅ Jiawei Liang ⋅ Puning Zhao ⋅ XIAOCHUN CAO

Physical-domain adversarial attacks present critical security threats to diverse AI systems, notably autonomous driving. Generalization is an inherent requirement and a primary obstacle in physical attacks, manifesting as the ability of adversarial patterns to persist across varying environmental configurations and to transfer to unseen victim models. While prior methods demonstrate generalization across specific models, their performance drops significantly on open-vocabulary foundation models, with some approaches becoming entirely ineffective. This lack of generalization stems from the failure of adversarial camouflage to construct a cohesive pseudo-feature across perspectives, lacking the stable semantic representation characteristic of natural objects. Such a deficiency arises because existing attack paradigms rely on independent view-level optimization and non-targeted objectives, which lead to semantically scattered features and inconsistent optimization directions. To address these limitations, we propose Generalizable Physical Adversarial Camouflage (G-PAC), a framework that transitions from view-specific suppression to coordinated object-level representation learning. G-PAC maintains a momentum-updated global pseudo-feature center as a stable adversarial anchor to guide optimization across diverse configurations. By leveraging the feature space of a self-supervised foundation model as a semantic prior, we introduce a contrastive regularization term to encourage feature aggregation while ensuring adversarial potency. Comprehensive digital and physical evaluations demonstrate the effectiveness of G-PAC. Notably, it outperforms the strongest baseline by an average AP@0.5 margin of 0.10 across seven diverse detector architectures in digital settings, and further extends this margin to 0.12 in real-world physical tests. Furthermore, G-PAC exhibits profound black-box transferability against highly resilient open-vocabulary foundation models, inducing an additional absolute AP@0.5 reduction of 0.30 on GLIP.

Graph-structured data underpins applications from citation analysis and social-network modeling to molecular design and knowledge-graph construction, and Large Language Models (LLMs) are increasingly used as prompt-driven graph synthesizers. Classical graph-generation reviews catalog deep generative models and their evaluation primitives, but predate the LLM era and provide no foundation for evaluating instruction-following graph synthesis. Recent LLM-era benchmarks evaluate models along graph-type or task-domain axes; such organizations, however, average over structural complexity and cannot localize \emph{where} in the complexity spectrum an LLM breaks down. To close this diagnostic gap, we introduce GraphInstruct, a progressive-complexity benchmark that stratifies LLM graph generation into six complexity levels and five evaluation dimensions, paired with 800 hand-authored instructions, 1,582 algorithmically synthesized reference solutions, and a 12-LLM capability evaluation across 45 (model, strategy) configurations. We find that discriminative power peaks at multi-constraint composition rather than reasoning depth, that no single prompting strategy dominates across levels or model families, and that domain-semantic constraints remain iteration-invariant under all tested methods---pointing to retrieval rather than additional compute as the next research frontier. Atop the benchmark, a verification-guided iterative framework with constraint-aware adaptive prompting consistently surpasses the prompt-engineering ceiling on tested target models, demonstrating that the benchmark's fine-grained signals drive method development. Data, code, and reproducibility artifacts are included in the supplementary materials and will be released publicly upon acceptance.


Grasp-Then-Plan with Failure Attribution: A Closed Two-Stage Framework for Precise and Generalizable Robotic Manipulation

Jiahao Xu ⋅ Peiyuan Wang ⋅ Hanzhuo Zhang ⋅ Zihao Yu ⋅ Tianyu Fu ⋅ Hao Chen ⋅ Xuanhao Xiang ⋅ Zixuan Li ⋅ Jianbo Yu ⋅ Chenchen Fu ⋅ Wanyuan Wang

In robotic manipulation, the tight coupling between grasping and motion planning often obscures the true source of failure, leading to inefficient trial-and-error. To enable efficient long-horizon manipulation, we propose GTP-FA (Grasp-Then-Plan with Failure Attribution), a task-oriented two-stage ‘grasp-then-plan’ framework that generates grasp candidates and performs downstream motion planning conditioned on the selected grasp. Given a failed manipulation trajectory, we learn a failure attribution model that generalizes to unseen grasps and produces a stable distribution over failure modes for diagnosis-guided optimization. Based on these attribution results, we then optimize both modules in a diagnosis-driven manner: on the grasping side, we inject task-level priors and risk penalties into grasp candidate scoring and optimization to suppress unstable or task-incompatible grasps; on the planning side, we target high-risk initial states through data collection and fine-tuning to address genuine planning bottlenecks. We evaluate the proposed framework in both simulation and real-robot experiments, and show that GTP-FA improves the corresponding base learners across RL, IL, diffusion-policy, and VLA-based settings, achieving substantially higher overall task success rates. Project page: here.

Recent works have shown that gradient-update alignment is a powerful signal for modulating optimizer updates and improving training dynamics. We promote this update-wise heuristic into a mathematically grounded principle for selecting and tuning optimizer hyperparameters. By treating gradients and updates as signals, and an optimizer as a causal filter that maps between them, we formulate optimizer selection as maximizing the expected drop rate in loss over a prescribed family of optimizers. We show that this objective is exactly the inner product between the optimizer filter and the gradient autocorrelation, and prove that a greedy optimum exists and has a stability bound under perturbations of the estimated gradient statistics. Specializing in momentum-based optimizers, the theory yields simple dynamic momentum selection rules for both SGD+Momentum and Adam/AdamW. Experiments across image classification, language model fine-tuning, and vision transformer fine-tuning show that the resulting dynamic momentum rules match or improve upon the best fixed hyperparameters found via manual sweeps, reducing the need for exhaustive momentum sweeps.


GReFEM: Multimodal LLMs as Zero-Shot Semantic Assistants for Physics-Guided 3D Mesh Refinement

Kartik Bali ⋅ Mahish Kumar Guru ⋅ Christian J Cyron ⋅ Roland Aydin

Adaptive volumetric finite element meshing is a critical step in computer-aided engineering and analysis that dictates the computational budget of a given problem. It traditionally requires iterative PDE solvers or heavily supervised, data-driven surrogates trained on large-scale simulation data. While Multimodal Large Language Models (MLLMs) excel in 2D visual tasks, their zero-shot capability to semantically ground regions based on geometric understanding and physics remains an open question. Overall, this study explores a significant question: can the high-level semantic understanding of off-the-shelf MLLMs serve as a viable, zero-shot geometric proxy for finite element mesh refinement? To investigate this, we introduce GReFEM (Geometric Reasoning Enhanced Multimodal LLMs for Finite Element Meshing), a framework that utilizes MLLMs to visually localize stress-critical regions based on physics-guided textual prompts. To bridge the gap between 2D MLLM pre-training and 3D geometries, we introduce orthoViews, a view-selection module that maximizes the observability of key geometric features. We conduct an in-depth empirical evaluation across diverse CAD geometries, loading cases, and SOTA MLLMs, comparing them against a tuned geometric heuristic under a strict, matched refinement budget. Our findings reveal that MLLMs demonstrate robust zero-shot capacity to accurately follow complex spatial-physical instructions, isolating stress-relevant features with higher precision than blind heuristics. By mapping both the successes and current limitations of MLLMs in physical grounding, this study defines the frontier of foundation models as semantic assistants in automated simulation workflows.


Group-Graph Policy Optimization for Long-Horizon Agentic Reinforcement Learning

Yunan Wang ⋅ Minghui Song ⋅ Zihan Zhang ⋅ Shaohan Huang ⋅ Haizhen Huang ⋅ Furu Wei ⋅ Weiwei Deng ⋅ Feng Sun ⋅ Qi Zhang

Group-based Reinforcement Learning (RL) has significantly enhanced Large Language Models (LLMs) in agentic scenarios. To achieve finer-grained policy updates, recent agentic RL frameworks have shifted from trajectory-level to step-level training. However, long-horizon agentic RL suffers from severe reward sparsity and delay, as feedback is often deferred for dozens of interaction steps. While existing step-level frameworks refine training granularity, their credit assignment remains coarse-grained and still treats agent exploration as isolated, linear trajectories. This oversimplified perspective ignores the inherent graph structure of state transitions, leading to high-variance state-value estimation and myopic, localized credit assignment. To overcome these critical bottlenecks, we propose Group-Graph Policy Optimization (G2PO), a novel group-based RL algorithm tailored for multi-turn agentic tasks. G2PO explicitly transforms linear interaction trajectories into a global state-transition graph. By aggregating identical observations across different trajectories, we introduce group-aggregation state-value estimation that reduces sampling variance and trajectory-dependent bias. Furthermore, we redefine agent actions as transitions between state nodes and propose an edge-centric advantage estimation strategy. By globally standardizing Temporal Difference (TD) errors across the entire graph, G2PO explicitly identifies and prioritizes critical transitions that drive absolute task progress. Extensive experiments on representative long-horizon benchmarks—WebShop, ALFWorld, and AppWorld—demonstrate that G2PO substantially outperforms state-of-the-art prompt-based and RL baselines, achieving remarkable success rate improvements of up to 22.2\% over GRPO.


Group of Skills: Group-Structured Skill Retrieval for Agent Skill Libraries

Kun Zeng ⋅ Yu Huo ⋅ Siyu Zhang ⋅ Zi Ye ⋅ Yuecheng Zhuo ⋅ haoyue liu ⋅ YuQuan Lu ⋅ Junhao Wen ⋅ Xiaoying Tang

Skill-augmented agents increasingly rely on large reusable skill libraries, but retrieving relevant skills is not the same as presenting usable context. Existing methods typically return atomic skills or dependency-aware bundles whose internal roles remain implicit, leaving the agent to infer the execution entry point, support skills, visible requirements, and failure-avoidance guidance. We introduce Group of Skills (GoSkills), an inference-time group-structured retrieval method that changes the agent-facing retrieval object from a flat skill list to a compact, role-labeled execution context. GoSkills builds anchor-centered skill groups from a typed skill graph, expands support groups through a group graph, bottlenecks the selected group plan into a bounded set of atomic skill payloads, and renders a fixed execution contract with START, SUPPORT, CHECK, and AVOID fields, without changing the downstream agent, skill payloads, or execution environment. Experiments on SkillsBench and ALFWorld show that \goskills preserves visible-requirement coverage under a small skill budget, improves over flat skill-access baselines, and often improves reward and agent-only runtime relative to structural retrieval references. Code is available at https://anonymous.4open.science/r/Group-of-Skills-E861.


Guiding Visual Autoregressive Models through Spectrum Weakening

Chaoyang Wang ⋅ Tianmeng Yang ⋅ Yunhai Tong

Classifier-free guidance (CFG) has become a widely adopted and practical approach for enhancing generation quality and improving condition alignment. Recent studies have explored guidance mechanisms for unconditional generation, yet these approaches remain fundamentally tied to assumptions specific to diffusion models. In this work, we propose a spectrum weakening framework for visual autoregressive (AR) models. The method is training-free and condition-free by constructing a controllable weak model in the spectral domain without architectural modifications. We theoretically show that invertible spectral transformations preserve information, while selectively retaining only a subset of the spectrum introduces controlled information reduction. Based on this insight, we perform spectrum selection along the channel dimension of internal representations, which avoids the structural constraints imposed by diffusion models. We further introduce a spectrum-renormalization technique that maintains numerical stability during the weakening process. Comprehensive empirical and ablation studies confirm the effectiveness of our approach, including raster-scan, random-order, and scale-wise AR, spanning discrete and continuous modeling and covering class and text-condition scenarios, demonstrating high-quality unconditional generation and strong prompt alignment maintenance for conditional generation. Code will be made available.


HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation

Jaehwi Song ⋅ Suchae Jeong ⋅ Byeongguk Jeon ⋅ Sungdong Kim ⋅ Hyungmok Son ⋅ Kimin Lee

Large-scale demonstration datasets have been central to recent progress in general-purpose robot policies. However, existing datasets are collected in human-absent settings, and policies trained on such data may perform tasks competently in isolation but fail to exhibit human-aware behaviors. To address this gap, we introduce HABIT, a large-scale robot demonstration dataset for human-present environments. We organize tasks into three roles capturing distinct modes of human-robot interaction: Collaborator, where human and robot jointly accomplish a task; Coworker, where they pursue separate tasks in a shared space; and Supervisor, where the human directs the robot. The dataset comprises over 7K episodes and over 100 hours across 48 tasks. Our experiments show that training on human-present data elicits human-aware behaviors that robot-only data fails to produce: spatiotemporal synchronization in Collaborator tasks, yielding in Coworker tasks, and gesture grounding in Supervisor tasks. Moreover, training on HABIT enables rapid adaptation to new human-robot interaction tasks. By introducing human presence as a new axis of dataset diversity, HABIT extends robot policies to environments shared with humans.


HAI: Hierarchical Anchored Interaction for Multi-View Bimanual World Models

Sen Cui ⋅ Baohua Yin ⋅ Youyi Kou ⋅ Junyu Wu ⋅ Jingheng Ma ⋅ Zhikang Chen ⋅ Changshui Zhang

Multi-view bimanual robot world models must predict future observations while preserving scene identity, following two synchronized but non-exchangeable arms, and maintaining consistency between global and wrist-local views. Existing conditioning schemes often collapse left- and right-arm actions into a view-agnostic control signal, making long-horizon rollout prone to scene drift, wrong-arm responses, and cross-view inconsistency. We propose \textbf{HAI}, a \textbf{Hierarchical Anchored Interaction} architecture for controllable multi-view bimanual world modeling. HAI organizes prediction as a structured information flow. First, hierarchical action-view conditioning routes the bimanual action chunk to camera streams, injects coarse chunk-level intent, performs structured multi-view modeling on the coarse-modulated features, and then injects fine per-frame action control. Second, anchored dynamic generation fuses the fine action-conditioned descriptors with persistent per-view scene anchors before decoding future observations. Experiments on AgiBot and DROID show that HAI improves long-horizon rollout quality over the baselines. Ablations confirm gains in scene stability, wrist-view controllability, and cross-view consistency. Further experiments on policy-improvement diagnostics indicate more useful synthetic rollout signals for downstream VLA policies, achieving an averaging 67.2\% relative gain in success rate. The project page is available at \url{https://hai-anon.github.io/}.


HaM-World: Soft-Hamiltonian World Models with Selective Memory for Planning

Haoyun Tang ⋅ Haodong Cui ⋅ Keyao Xu ⋅ Zhan-Dong Mei ⋅ Kun Wang

World models enable model-based planning through learned latent dynamics, but imagined rollouts often become unstable as the planning horizon grows or the dynamics distribution shifts. We argue that this instability arises from two missing structures in planner-facing latents: history-conditioned memory for approximate Markov completeness, and geometric organization that separates configuration, momentum, and task semantics. We propose HaM-World, a structured world model that decomposes the latent state into a canonical $(q, p)$ subspace and a context subspace $c$, while incorporating Mamba selective state-space memory as a history-conditioned input to the same latent dynamics. Within this unified interface, $(q, p)$ evolves under a Soft-Hamiltonian dynamics composed of an energy-derived Hamiltonian vector field and learnable residual and control dynamics, while $c$ captures semantic, dissipative, and non-conservative factors. This design provides the planner with a single latent representation shared across dynamics prediction, reward and value estimation, imagined rollouts, and cross-entropy method (CEM) planning. On four DeepMind Control Suite tasks, HaM-World achieves the highest average AUC (+9.5% over strong baselines), reduces long-horizon rollout error to 45% of a competitive model, and wins 11 out of 12 $k \in \{3,5,7\}$ rollout MSE metrics. Under 12 out-of-distribution perturbations, HaM-World consistently attains the highest returns, with average gains of 10.2% on Finger Spin and 13.6% on Reacher Easy. Mechanism diagnostics further demonstrate bounded energy drift under action-free rollouts, structured energy variation under policy control, and coherent control-induced energy transfer, supporting the effectiveness of the proposed Soft-Hamiltonian latent dynamics. Code: https://anonymous.4open.science/r/HaM_World-47CD


HaNDF: Object-Conditioned Geometric Neural Hand Distance Fields

Zuhao Liu ⋅ Melih Darcan ⋅ Zhengdi Yu ⋅ Riza Alp Guler ⋅ Tolga Birdal

We propose, HaNDF, object-conditional neural distance fields as data-driven priors for modeling the plausible hand–object configurations, in the product manifold of hand articulations. While recent unconditional generative priors showed success in modeling human pose, hands are versatile articulated bodies often found in interaction with other objects. This makes it crucial to develop a generative prior which can model the plausible hand poses conditioned on the object under interaction. To this end, our HaNDF proposes a geometry-aware, object-conditioned neural distance field, representing the plausible articulations in the zero-level set of a conditional field. We leverage tools from Riemannian geometry to (i) represent plausible hands in the zero level-set of an object-induced neural field, and (ii) project a given pose onto the this field. Using HaNDF, we can also optimize for a plausible condition, determining the object under interaction. Our extensive evaluations demonstrate that our HaNDF provides a flexible and powerful prior for hand–object interaction, suitable for applications such as pose refinement, reconstruction under occlusion, and physically plausible manipulation synthesis.


HAPS: Hierarchical LLM Routing with Joint Architecture and Parameter Search

Zihang Tian ⋅ Rui Li ⋅ Jingsen Zhang ⋅ Xiaohe Bo ⋅ Wei Huo ⋅ Xu Chen

Large language model (LLM) routing aims to exploit the specialized strengths of different LLMs for diverse tasks. However, existing approaches typically focus only on selecting LLM architectures, treating candidate models as static black boxes while overlooking parameter adaptation, which can substantially affect task performance. In this paper, we introduce HAPS, a hierarchical LLM routing framework that jointly searches over model architectures and parameters. Specifically, a high-level router selects among candidate LLM architectures, while a low-level router generates input-conditioned LoRA parameters for the selected architecture. To couple these two decisions, we design a shared parameter generation mechanism that enables cross-level knowledge transfer between architecture routing and parameter adaptation. We further optimize the whole framework with a reward-augmented training objective. Extensive experiments show that HAPS consistently outperforms strong routing baselines. We further conduct broad ablations and analyses on parameter sharing, architectural choices, scalability, and efficiency, providing additional evidence for the effectiveness and practicality of the proposed framework. We have released our code at https://anonymous.4open.science/r/HAPS_private-68CB.


Harnessing Agentic Evolution

Jiayi Zhang ⋅ Yongfeng Gu ⋅ Jianhao Ruan ⋅ Maojia Song ⋅ Yiran Peng ⋅ Zhiguang Han ⋅ Jinyu Xiang ⋅ Zhitao Wang ⋅ Caiyin Yang ⋅ Bang Liu ⋅ Chenglin Wu ⋅ Yuyu Luo

Agentic evolution has emerged as a powerful paradigm for improving programs, workflows, and scientific solutions. It iteratively generates candidate artifacts, evaluates them against a task objective, and uses the resulting feedback to guide subsequent evolution. Existing methods typically instantiate this paradigm either through fixed hand-designed procedures that are modular but rigid, or through general-purpose agents that flexibly integrate feedback but can drift as context grows, leaving long-horizon search vulnerable to local optima. Both routes share a deeper limitation: long-horizon evolution accumulates candidates, feedback, traces, and failures over time, but lacks a stable interface for organizing this evidence and revising the mechanism that drives future evolution. We therefore formulate agentic evolution as an interactive environment, where the accumulated evolution context becomes process-level state. A meta-agent acts on this environment not by directly generating the next candidate, but by editing the mechanism that controls how future evolution proceeds. We introduce AEvo, a harnessed framework for meta-editing agentic evolution. It standardizes the evolution environment and provides a unified interface for observing accumulated evidence and editing the mechanism that drives future evolution. This lets AEvo revise both hand-designed procedures and agent operating contexts, reducing local-optimum risks in long-horizon evolution. Empirical evaluations on agentic and reasoning benchmarks show that AEvo outperforms 5 evolution baselines, achieving a 26% relative improvement over the strongest baseline. Furthermore, across 3 open-ended optimization tasks, AEvo outperforms 4 evolution baselines and achieves state-of-the-art performance under the same iteration budget.


Harnessing Streaming Video in the Wild

Dingyu Yao ⋅ Shuhuan Gu ⋅ Qingyi Si ⋅ Junhao Zhou ⋅ Chenxu Yang ⋅ Chuanyu Qin ⋅ Naibin Gu ⋅ Zheng Lin ⋅ Weiping Wang ⋅ Nan Duan ⋅ Jiaqi Wang

Vision-Language Models (VLMs) are increasingly required to process unbounded video streams in applications such as video-call assistants, live commentary, and embodied robots. An ideal streaming system should support proactive interaction, long-horizon memory, and real-time processing, while resting on a VLM backbone capable of handling diverse in-the-wild streaming tasks. However, existing VLMs excel at offline video understanding but fall short in streaming capabilities and lack dedicated infrastructure for streaming deployment. We address this gap on three fronts. (i) For backbone capability, we construct \textbf{Streaming-Train-248K}, a streaming dataset paired with a novel training objective for adapting VLMs to streaming interaction and understanding. (ii) For real-world deployment, we introduce \textbf{Streaming Harness}, a plug-and-play system that endows any VLM with three core abilities: proactive interaction (per-second response decisions), long-term memory (12-hour context retention), and real-time processing (sub-second latency). (iii) To drive continued community progress on streaming capabilities, we design \textbf{Streaming-Eval}, a benchmark that reflects models' capabilities across diverse in-the-wild scenarios. Extensive experiments demonstrate consistent gains from our approach across all core capabilities required for streaming video understanding. We will open-source our data, code, and benchmark to advance the community's shift from offline video understanding to deployable streaming intelligence.


Hermes: A Multi-Scale Spatial-Temporal Hypergraph Network for Stock Time Series Forecasting

Xiangfei Qiu ⋅ Liu Yang ⋅ Xiangyu Xu ⋅ Hanyin Cheng ⋅ Xingjian Wu ⋅ Rongjia Wu ⋅ Zhang Zhigang ⋅ Tu ding ⋅ Chenjuan Guo ⋅ Bin Yang ⋅ Christian S. Jensen ⋅ Jilin Hu

Time series forecasting occurs in a range of financial applications providing essential decision-making support to investors, regulatory institutions, and analysts. Unlike multivariate time series from other domains, stock time series exhibit industry correlation. Exploiting this kind of correlation can improve forecasting accuracy. However, existing methods based on hypergraphs can only capture industry correlation relatively superficially. These methods face two key limitations: they do not fully consider inter-industry lead-lag interactions, and they do not model multi-scale information within and among industries. This study proposes the \textbf{Hermes} framework for stock time series forecasting that aims to improve the exploitation of industry correlation by addressing these limitations. The framework integrates moving aggregation and multi-scale fusion modules in a hypergraph network. Specifically, to more flexibly capture the lead-lag relationships among industries, Hermes proposes a hyperedge-based moving aggregation module. This module incorporates a sliding window and utilizes dynamic temporal aggregation operations to consider lead-lag dependencies among industries. Additionally, to effectively model multi-scale information, Hermes employs cross-scale, edge-to-edge message passing to integrate information from different scales while maintaining the consistency of each scale. Experimental results on multiple real-world stock datasets show that Hermes outperforms existing state-of-the-art methods.


Heterogeneous Judge-Aware Ranking with Sensitivity, Disagreement, and Confidence

Shibo Yu ⋅ Yingzhou Wang ⋅ Yan Chen ⋅ Guodong Li ⋅ Jin-Hong Du

Pairwise comparisons from multiple judges are central to large language model evaluation and preference modeling, yet standard ranking pipelines often pool judgments into a single score vector, treating systematic judge disagreement as noise. We propose Heterogeneous Judge-Aware (HJA) ranking, a structured multi-judge ranking framework that separates consensus ranking, judge-specific sensitivity to consensus, and residual preference disagreement. HJA thereby treats ranking, judge sensitivity, and structured disagreement as separate inferential targets. We establish conditions under which this decomposition is identifiable and develop an anchored alternating algorithm that preserves the identifying geometry. For confidence quantification, we study a fixed-panel repeated-comparison regime in which the judge panel may remain fixed or modest while information grows through repeated judgments. This yields uncertainty statements for consensus and judge-specific ranking contrasts, sensitivity parameters, pairwise probabilities, and summaries of residual disagreement.Experiments on synthetic and real multi-judge comparison data show that HJA improves recovery, robustness, uncertainty calibration, and near-tie performance relative to pooled and sensitivity-only baselines. The fitted model also provides diagnostics for judge disagreement and model-affinity patterns, giving a statistically grounded framework for ranking under heterogeneous comparative judgments.


Heuresis: Evaluating Search Strategies for Autonomous Machine Learning Research Agents

Antonis Antoniades ⋅ Deepak Nathani ⋅ Ritam Saha ⋅ Alfonso Amayuelas ⋅ Ivan Bercovich ⋅ Zhaotian Weng ⋅ Vignesh Baskaran ⋅ Kunal Bhatia ⋅ Xin Wang ⋅ William Yang Wang

Autonomous AI research promises to accelerate the scientific progress of machine learning. While current Large Language Model (LLM)-based agents excel at writing code, the bottleneck in research is the exploration of diverse and novel ideas: agents collapse to common techniques present in their pretraining data, or prematurely converge on suboptimal solutions (Padmakumar et al., 2024; Jiang et al., 2025). To this end, we introduce Heuresis, a framework that abstracts the research pipeline into a set of general and composable primitives, enabling open-ended scientific exploration in machine learning research. We implement 5 known algorithms spanning Quality-Diversity, Evolutionary, and Curiosity-based search, in addition to a greedy baseline, and evaluate them across three axes -- Quality, Diversity, and Novelty -- on two domains: LLM pretraining and On-Policy RL. We find that quality and novelty are inversely correlated. Methods that optimize purely for raw performance reach the strongest single solutions on a given task by replicating prior work: the top-quality runs from the greedy baseline are uniformly classified as direct copies under the Gupta-Pruthi rubric (Gupta & Pruthi, 2025). The few high-quality, verified-novel ideas in our study come from algorithms that balance performance with diversity or curiosity-based exploration. We also observed that agents resorted to a variety of reward-hacking techniques during execution, whose detection was necessary to keep the search faithful to the task. Our results underscore the importance of search and Quality-Diversity methods for autonomous research, and our framework opens up opportunities for further inquiry towards the ultimate goal of perpetual, autonomous scientific progress.

We propose a distributional theory of how hypernymy---the ``is-a'' relation between general and specific concepts---is encoded geometrically in language representations. Starting from the empirically verified assumption that words closer on the WordNet hypernym graph co-occur more often, we characterize theoretically the spectrum of the resulting embedding Gram matrix of word2vec embeddings. Under mild positivity and decay conditions on the co-occurrence kernel, we prove that the leading eigenvectors first separate broad taxonomic branches and then progressively finer sub-branches, producing a \emph{hierarchical splitting geometry} with a coarse-to-fine spectral organization that mirrors the tree. We confirm these predictions in word2vec embeddings across many sampled WordNet subtrees, and show that the same signature extends strikingly well to Gemma 2B unembeddings. Our results indicate that hierarchical concept geometry in LLMs need not reflect a hierarchy-specific functional mechanism, but emerges from the spectral structure of pairwise word statistics.

In fine-grained image recognition, semantic parts (e.g., wings, head) recur across classes while discriminative cues lie in within-part features (e.g., yellow wings, striped wings). The prototypical part network (ProtoPNet) is an interpretable model that provides case-based explanations for image recognition by measuring similarities between local patches in an images and their closest prototypes representing patches in a training set. However, existing ProtoPNet variants do not explicitly model this two-level structure, causing prototypes for the same anatomical part to fragment across classes. We propose a hierarchical prototype learning framework that separates shared prototypes capturing cross-class part concepts from class-specific prototypes capturing intra-part features. To model this hierarchy geometrically, prototypes are embedded in hyperbolic space and an entailment loss constrains each class-specific prototype to lie within the cone of its corresponding shared prototype. To prevent degenerate trivial solutions such as redundant prototypes, we introduce a patch-level contrastive learning using pseudo patch IDs from a vision foundation model, used only at training time, leaving the backbone choice free at inference. Unlike prior hyperbolic prototype methods where prototype hierarchies are uncontrolled, our framework defines explicit shared part to intra-part feature correspondences. Across four ProtoPNet variants and four backbones on CUB-200-2011 and Stanford Cars, our method improves classification accuracy by up to 10 points, raises hierarchical part hierarchy IoU from 11\% to 75\%, and consistently improves four interpretability metrics on CUB-200-2011.


Hierarchical Semantic Tree Anchoring for CLIP-Based Class-Incremental Learning

Tao Hu ⋅ Lan Li ⋅ Zhenhao Wen ⋅ Da-Wei Zhou

Class-Incremental Learning (CIL) enables models to learn new classes continually while preserving past knowledge. Recently, vision-language models like CLIP offer transferable features via multi-modal pre-training, making them well-suited for CIL. However, real-world visual and linguistic concepts are inherently hierarchical: a textual concept like ''dog'' subsumes fine-grained categories such as ''Labrador'' and ''Golden Retriever,'' and each category entails its images. But existing CLIP-based CIL methods fail to explicitly capture this inherent hierarchy, leading to fine-grained class features drift during incremental updates and ultimately to catastrophic forgetting. To address this challenge, we propose HASTEN (Hierarchical Semantic Tree Anchoring), a hierarchy-aware framework that uses semantic structure to stabilize CLIP-based CIL. Rather than treating classes as isolated labels, HASTEN leverages external hierarchical knowledge as structured supervision to organize visual and textual features in hyperbolic space, helping maintain parent-child relations as new tasks arrive and mitigating feature drift. Since the shared hyperbolic mapper is updated across tasks, we further stabilize it by constraining its updates to a null space induced by prior-task features, reducing interference with previous mappings while retaining adaptability to new classes. Extensive experiments on nine benchmarks show that HASTEN consistently outperforms existing methods while reducing feature drift and catastrophic forgetting.


HierSVA: A Synthesis Pipeline, Dataset, and Benchmark for LLM-Driven Hierarchical Hardware Formal Verification

Maohua Nie ⋅ Jiang Zhu ⋅ Jingqun Zhang ⋅ Zhichen Zeng ⋅ Jiayi Wang ⋅ Sibo Zhang ⋅ Jialin Wang ⋅ C-J R Shi

We present HierSVA, an integrated suite that combines a pipeline, dataset, and benchmark for LLM-driven hierarchical hardware formal verification. HierSVA-SP pairs an RTL preprocessing toolchain with an LLM-in-the-loop formal verification flow to produce reference SystemVerilog Assertions (SVA) on hierarchical RTL. Applying it to BaseJump STL yields HierSVA-DS, a dataset of 342 modules, with hierarchy metadata and depths 0--9, accompanied by a deep subset of 28 module-bug pairs with natural-language specifications and bug variants. HierSVA-B decomposes assertion quality into six metric axes: syntax correctness, assertion proof success rate, vacuity, faithfulness, mutation coverage, and formal core coverage. Applying HierSVA-B to twelve recent LLMs reveals three findings. First, the module-level compile rate is 67.1\%; among generated assertions in evaluable runs, 82.1\% prove non-vacuously, but the corresponding assertion sets detect only 70.2\% of eligible injected faults and cover 36.2\% of the formal core. Second, on 211 evaluable model--module entries in the deep subset, assertion sets flag buggy RTL with 0.87 recall, but 40\% of predicted-buggy outcomes are false positives on correct RTL, limiting precision to 0.60. Third, agentic mode improves S1-style provability and strength metrics, but gains plateau and oscillate, and we do not evaluate agentic C3 faithfulness.


High-Fidelity Boltzmann Samplingvia Physical Prior Lifted Continuous GFlowNets

Xizhi Tian ⋅ Wenhao Deng ⋅ Haojia Hui ⋅ Hang Chen ⋅ Long Wei

Sampling high-dimensional Boltzmann distributions is fundamental yet challenging. Classical MCMC methods incur prohibitive computational costs, while modern generative approaches such as diffusion models and Schrödinger bridges struggle to incorporate informative physical priors. GFlowNets offer a promising alternative, yet existing continuous variants suffer from training instability. In this paper, we propose Physical Prior Lifted continuous GFlowNets (PPL), a framework that encodes physical prior information through a reference space equipped with a reference distribution and a mapping to the original configuration space. The sampling process is learned via a Reference-Weighted Trajectory Balance (RWTB) loss, yielding stable training dynamics, efficient learning, and rigorous theoretical guarantees. Empirically, PPL achieves state-of-the-art performance across five benchmarks spanning synthetic, molecular, and protein systems. On the synthetic LJ-55, we obtain significantly lower error than the diffusion-based ASBS with ${\sim}36\times$ speedup; Notably, on the 267-dimensional Chignolin protein, PPL attains near-ground-truth sampling fidelity.


HIMMEL: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding

Haopeng Jin ⋅ Tiankun Yang ⋅ Zhenyu Guan ⋅ ShiQuan Dong ⋅ Wenlong Zhao ⋅ Hongzhu Yi ⋅ Chubin Chen ⋅ Tao Yu ⋅ Jinwen Luo ⋅ Yujia Yang

Long video understanding with multimodal language models suffers from three compounding bottlenecks: heavy decode cost to obtain dense RGB frames, quadratic token growth with frame count, and weak motion perception under sparse keyframe sampling. Existing remedies either prune visual tokens after decoding, which leaves the expensive RGB pipeline untouched, or they discretise codec motion vectors with a generative tokeniser, which couples motion modelling to a heavy pre-training stage. In this paper we present HIMMEL, a hierarchical video-language framework that allocates semantic and motion capacity along separate paths. A small set of sparse anchor I-frames is routed to the expensive host ViT to ground object identity and scene layout, while the far denser inter-frame intervals are encoded by a lightweight compressed-domain tri-stream adapter that distils motion evidence from motion-vector maps, residual maps, and an I-frame context branch into aligned motion tokens. These tokens are injected into the LLM via a differentiable placeholder mechanism, after a dedicated Stage-1 contrastive alignment that places the motion representation in a geometry compatible with the frozen visual backbone. We further show that an InfoNCE alignment objective beats MSE regression by $+1.5$ pp because it preserves directional motion structure rather than collapsing onto the mean of visual deltas. On Video-MME, HIMMEL surpasses the dense 32-frame baseline by $+2.3$ pp ($61.2 \to 63.5\%$) while using $3.6\times$ fewer context tokens and running end-to-end in $1.03$ s per question instead of $2.75$ s. Switching to a stronger Qwen3-VL-8B host pushes HIMMEL to $64.9\%$, narrowing the gap to much larger proprietary systems while keeping inference on a single consumer GPU. Extensive ablations across stream composition, motion-encoder family, fusion mode, alignment objective, anchor count, LoRA rank, and video duration confirm that the full tri-stream is necessary and sufficient for the observed gains, and that the benefit grows with video length. We release training and evaluation pipelines to support reproducibility.

Streaming 3D reconstruction enables long-sequence scene understanding by processing frames sequentially, but remains challenging when historical information is compressed into a fixed-size recurrent state. Existing training-free extensions of implicit-memory models mainly improve stability by modulating update magnitude, thereby controlling how strongly each incoming observation modifies the state. However, the geometric intent of the update, namely the direction in which the observation drives the recurrent state, remains tied to the current-frame candidate and can accumulate directional bias over long sequences. We propose HIST3R, a training-free framework that improves recurrent state updating through a direction--magnitude decomposition. HIST3R uses historical recurrent states to rectify the update direction, providing a more reliable geometric intent, and uses observation--history visual discrepancy to modulate the update magnitude, yielding an adaptive update gain. This design makes recurrent updates less dependent on isolated frame-level evidence while preserving the bounded-memory advantage of implicit streaming reconstruction. Across camera pose estimation, 3D reconstruction, and video depth estimation benchmarks, HIST3R consistently improves long-horizon performance without additional training. At 1000 input frames, HIST3R achieves relative ATE reductions of 23.87% on ScanNet and 33.65% on TUM-Dynamics over the strongest baseline; at 500 frames, it improves long-sequence 3D reconstruction over the strongest CUT3R-based variant by 56.52% on 7-Scenes and 33.96% on NRGBD.


HoloCode: A Code-Centric Multi-Agent Framework for Image-to-3D Scene Generation

Hanlin Chen ⋅ Chung-Ching Lin ⋅ Yuyang Zhao ⋅ Dongyue Lu ⋅ Zhengyuan Yang ⋅ Xiaofei Wang ⋅ Kevin Lin ⋅ Zhendong Wang ⋅ Yu Yang ⋅ Lijuan Wang

Image-to-3D scene generation must recover both object geometry and plausible spatial relations. A common diffusion-based pipeline, exemplified by SAM 3D, generates each object separately and assembles the results. This preserves local details, but the lack of strong scene-level constraints often causes floating objects, interpenetrations, missing contacts, or implausible relative scales. Inspired by human designers who plan a layout before placing assets, we introduce \ours{}, a code-centric multi-agent framework. Given a single image and 2D localization cues, a vision-language model (VLM)-based Arranger Agent predicts executable Blender Python layout code with object identities and 3D transforms. A Modeler Agent binds each object to geometry through asset retrieval or image-to-3D object generation, and a Refiner Agent renders the scene and edits the code from visual feedback. The code representation exposes object names, positions, rotations, and scales, making generated scenes editable and interpretable. We train the code-generating agents using Blender Python datasets built from SAGE and 3D-FRONT, then further optimize them with GRPO rewards for executability and spatial alignment. Experiments in in-domain and cross-domain settings show that \ourMethod{} improves scene-level plausibility while preserving object fidelity.


Horizon-Stream: Long-Horizon Attention for Streaming 3D Reconstruction

Chong Cheng ⋅ Peilin Tao ⋅ Nanjie Yao ⋅ Guanzhi Ding ⋅ Xianda Chen ⋅ yuansen du ⋅ Xiaoyang Guo ⋅ Wei Yin ⋅ Weiqiang Ren ⋅ Qian Zhang ⋅ Zhengqing Chen ⋅ Hao Wang

Online 3D reconstruction requires estimating camera pose and scene geometry under strict causal and bounded-memory constraints. Existing methods often suffer from drift, jitter, or collapse on long sequences. We trace these failures to a fundamental mismatch. Streaming geometry is inherently temporally heterogeneous, with evidence ranging from short-lived correspondences to persistent global scale. However, current architectures impose uniform and pathological influence patterns. For example, sliding windows enforce hard cutoffs, while ungated recurrence and causal attention cause cache saturation and spike-like attention sinks. To resolve this, we formalize geometric propagation as an \emph{evidence influence kernel} and propose Horizon-Stream, a long-horizon Transformer that explicitly factorizes this kernel. For the long-range temporal factor, Geometric Linear Attention learns channel-wise decay rates to enable bounded, multi-timescale propagation of geometric evidence. For the short-range spatial factor, Geometric Local Attention with Spatiotemporal RoPE performs reliable 3D matching while suppressing attention sinks. Finally, Metric Readout Tokens recover stable scale and rigid pose directly from the persistent geometric state. Extensive experiments show that Horizon-Stream, trained on only 48-frame clips, generalizes stably to sequences exceeding (10{,}000) frames with constant memory and linear time, achieving state-of-the-art streaming 3D reconstruction performance.

Geometric modeling has become an essential paradigm in AI for Science, where data often inhabit intrinsically curved or structured spaces. However, existing manifold generative models often rely on closed-form or tractable access to global geometric primitives, such as geodesic distances, parallel transport, or spectral decompositions, which are rarely available beyond canonical geometries. To overcome this limitation, we propose Horizontal Diffusion Models (HDMs), a score-based generative framework defined on the orthonormal frame bundle of a Riemannian manifold. Instead of requiring global geometric operations, HDM lifts Euclidean diffusion processes through the Levi-Civita horizontal distribution, using only local metric and connection information while preserving intrinsic manifold geometry. This construction allows standard Euclidean score networks to be lifted into geometry-consistent and gauge-equivariant horizontal vector fields, bypassing the need for manifold-specific neural architectures. We further derive a horizontal KL objective that reduces to a Euclidean score-matching loss and analyze the curvature-dependent cost of the lift through a generalization bound. Experiments on parametric surfaces and scientific datasets demonstrate high-fidelity generation across diverse manifolds with nontrivial geometry and topology.

Fluorescence spot-detection benchmarks in biology and medicine often rely on reference annotations that are incomplete, ambiguous, or inconsistently defined. In such settings, detector performance is not determined solely by model predictions; it can also depend on which faint or ambiguous signals are counted as reference positives. We formulate this issue as a reference-audit problem: instead of treating the evaluation reference as fixed ground truth, we test whether benchmark outcomes remain stable under plausible changes in reference definition and annotation completeness. Using the public FISH_spots dataset, a benchmark for fluorescence in situ hybridization (FISH) spot detection, we construct Raw, Lenient, and Strict audit references and evaluate diverse classical and learning-based detectors under a unified point-matching protocol. We further apply annotation-retention stress tests that simulate random and structured missing-label mechanisms. Across these audits, changing only the evaluation reference can alter measured F1 scores and model rankings, including the selected top-1 detector, even when aggregate rank correlation remains high. These findings suggest that model selection in fluorescence spot-detection benchmarks should be tested across plausible reference definitions, rather than based on a single reference annotation set.


How Data Scales in Agentic Reinforcement Learning: Laws and Synthesis Strategies

Bowei He ⋅ Yankai Chen ⋅ Xiaokun Zhang ⋅ Changjiang Han ⋅ Ye Yuan ⋅ Chong Li ⋅ Steve Liu

Pretraining scaling laws treat training data as a one-dimensional quantity: a token count. Agentic reinforcement learning (RL) inherits this scalar abstraction, but modern synthesis pipelines now generate *tasks*, *environments*, and *trajectories* independently and at very different costs, turning data into a structured design space. This raises a question pretraining never had to answer: given a synthesis budget, *which axis* should be scaled? We address this question end-to-end with three tightly coupled contributions. First, we decompose agentic data into four axes (task, environment, trajectory, reward) and empirically fit per-axis scaling laws on Qwen3 models (4B/8B/14B) across math, code, and web/UI domains; we observe a robust ordering of per-axis scaling exponents that holds across model sizes, RL algorithms (GRPO and PPO), and domains. Second, we use the fitted laws as a common ruler to price synthesis methods, capturing each method's efficiency as a per-axis discount factor and identifying critical synthetic-to-real ratios at which collapse begins; these ratios differ by an order of magnitude across axes. Third, we cast budget allocation as a constrained optimization under our laws and discounts, deriving a Chinchilla-style compute-optimal recipe and validating it at held-out scale. The empirical results demonstrate that following the recipe yields up to $1.7\times$ compute efficiency over the dominant ``scale-the-trajectories'' practice.

Retrieval-augmented in-context learning faces a budget problem under distribution shift: increasing top-$k$ can improve evidence coverage, but it can also add off-direction, conflicting, or overly long context that a finite reader cannot use. We introduce a theory-guided diagnostic framework based on \emph{effective evidence coverage}: retrieval is useful only while aligned evidence grows faster than reader-facing ambiguity. In a local linear-ICL model, coupled covariate and task-prior shifts create a mixed risk term that aligned retrieval can reduce up to a residual ambiguity term; a conditional transfer result extends the same budget boundary to frozen readers satisfying a prompt-level stability condition. This yields a falsifiable prediction: recall can keep rising after answer accuracy has saturated or declined. Across QA, verification, and NLI tasks, dense top-$k$ sweeps together with oracle and shuffled-context controls expose this recall--accuracy decoupling. The resulting framework separates coverage- and saturation-limited QA from interference-limited verification/NLI, while calibrated selected-context controllers are reported only as operational probes of the diagnosis.

Large-scale foundation models (FMs) in remote sensing (RS) (denoted as RS FMs) are developed following paradigms established in computer vision (CV), yet the validity of transferring CV scaling laws to RS has not been systematically examined. We hypothesize that RS FMs enter an overparameterized regime at substantially smaller scales than their CV counterparts, with task-relevant information encoded redundantly across model dimensions. To test this hypothesis, we apply post-hoc slimmability, uniform width reduction of pretrained encoder transformer blocks, as a tool to measure representational redundancy across eight state-of-the-art RS FMs on classification, segmentation, and change detection tasks. RS FMs retain 69% to 109% relative accuracy on RS datasets under aggressive width reduction, while masked autoencoder (MAE) and DINOv2 pretrained on natural images (denoted as CV MAE and CV DINOv2) degrade sharply on ImageNet subsets of matched class count over the same range of computational requirements. A CV MAE evaluated directly on the same RS datasets narrows but does not close the gap, indicating that both dataset characteristics and domain-specific pretraining contribute to the differences between the models. Mechanistic analyses such as feature correlation, explained variance, and effective dimensionality indicate that task-relevant variance concentrates in few principal components and is redundantly encoded across model dimensions. We further show that learned slimmable training improves over post-hoc slimmability for contrastive objectives, while reconstruction-based objectives do not benefit from current slimmable training protocols. Our findings establish post-hoc slimming as a practical deployment strategy for resource-constrained RS applications and as a diagnostic tool for representational redundancy in RS FMs. Upon acceptance, we will publish all code.


How Neural Reward Models Learn Features for Policy Optimization: A Single-Index Analysis

Rei Higuchi ⋅ Ryotaro Kawata ⋅ Akifumi Wachi ⋅ Shokichi Takakura ⋅ Kohei Miyaguchi ⋅ Taiji Suzuki

Reward modeling is not only a prediction problem: in KL-regularized policy optimization, the learned reward is exponentiated to define the deployed policy, so downstream value depends on errors in reward-tilted regions. We study this feedback in a Gaussian single-index model with $ r^* (x)=\sigma^* (\langle\theta^* ,x\rangle)$ and $x\sim\mathcal N(0,I_d)$. We analyze a two-stage neural reward model that first learns the hidden direction $\theta^*$ from reward-weighted samples and then fits the readout layer by weighted ridge regression. Exponential reward weighting changes the Hermite signal available to the first layer; for any feature-learning temperature $\beta_1$ above a dimension-free $O(1)$ threshold, a constant fraction of neurons recover the hidden direction, with weak-recovery complexity governed by the generative exponent. After feature recovery, we derive tilted-policy value-gap bounds for an idealized label-weighted fit with weights $e^{y/\beta_2}$ and a more practical surrogate-weighted fit with weights $e^{r_{a_0}(x)/\beta_2}$. Keeping the $\beta_2$-dependence explicit yields an admissible set of deployment temperatures, balancing the gain from lowering $\beta_2$ against the learning cost amplified by exponential weighting; in the surrogate-weighted case, proxy-dependent factors shrink this admissible set.


How to Instruct Your Robot: Dense Language Annotations Power Robot Policy Learning

Bosung Kim ⋅ Ruiyi Wang ⋅ David Acuna ⋅ Jaehun Jung ⋅ Alex Trevithick ⋅ Brandon Cui ⋅ Yejin Choi ⋅ Prithviraj Ammanabrolu

Scaling robot policy learning is bottlenecked by the cost of collecting demonstrations: datasets at modern scale require thousands of skilled-operator hours on dedicated robot hardware. The language description paired with these demonstrations does not face the same bottleneck---a single short task label is cheap, but it also leaves implicit spatial relations, object interactions, embodiment state, and subgoal structure that the pixels already contain. We treat \emph{language density} as a cheap lever for amplifying signal in a fixed demonstration corpus; whereas prior dense-language work in robotics commits to a single caption style, we instead ask which kind of dense language helps each task and learn to deliver it at deployment. We realize this as \textbf{DeMiAn} (\textbf{De}nse \textbf{M}ult\textbf{i}-aspect \textbf{A}nnotatio\textbf{n}) in two stages. First, an automatic VLM pipeline re-labels each segment of an existing demonstration along four complementary aspects---\emph{physical motion}, \emph{scene composition}, \emph{arm pose}, and segment-level \emph{reasoning}---each surfacing a distinct kind of structure that a one-line task label omits. Second, a small learned \emph{instructor}, trained via supervised fine-tuning, maps the natural language task description and an initial scene snapshot to a task-appropriate annotation and runs asynchronously alongside the action policy, hiding generation latency behind the rollout. Applied to over 1M robot manipulation and 50K EgoVerse human-egocentric videos, with no new demonstrations collected, DeMiAn delivers four findings on a VLA action policy and a video-based world-action model: i) the learned instructor lifts RoboCasa success by 5 points over the no-annotation baseline, within 3 points of a per-task oracle; (ii) that oracle is non-trivial---no fixed aspect dominates, and peak performance requires selecting the right annotation aspect per task; (iii) the trained system extends usefully to composite tasks under subgoal-driven prompt switching, and to OOD scenes and objects; and (iv) dense annotation improves the compute-performance frontier in both mid-training and post-training, making re-annotation a practical scaling lever for robot policy learning.


Human–AI Collaboration Requires a High-Order Dynamic Abstraction Substrate

Hengyu Liu ⋅ Dongxu Huang ⋅ Zhihong Cui ⋅ Lun Du ⋅ Tiancheng Zhang ⋅ Tor Skeie ⋅ Kristian Torp ⋅ Christian S. Jensen

This position paper argues that effective Human–AI collaboration in the LLM era requires a high-order dynamic abstraction substrate—one that externalizes stable human cognition into a form AI can reliably consult, and whose content evolves as AI and human capabilities reshape what is worth codifying. We analyze the substrate along three dimensions—what it must hold (five necessary knowledge categories of autonomous AI: Situation, Purpose, Action, Constraints, Evaluation; SPACE), where its content originates (a source-by-externalization analysis isolating the fabrication-codified failure mode behind LLM hallucination, and three practice-grounded routes that defend against it), and how it organizes Human–AI collaboration (two complementary insights: a three-phase temporal cycle of externalize–practice–summarize in which AI plays a different role at each phase, and a long-tailed abstraction hierarchy in which humans focus on the rare top while AI handles the tail). We name this substrate Knowledge Object (KO Network) and characterize it along three axes: content that occupies the moving band between the AI capability ceiling and the human articulability ceiling; structure that commits to five quality attributes (Understandable, Verifiable, Traceable, Controllable, Reusable; UVTCR) and SPACE-aligned subgraphs; and a bidirectional lifecycle in which entries are added as humans articulate previously tacit cognition and retired as AI absorbs the codified. Finally, we respond to five alternative views.


Human-AI Teaming Through the Lens of Calibration

Eric Nalisnick ⋅ Chi Zhang ⋅ Chengxin Qian ⋅ Yixin Wang

We study models for human-AI teaming through the lens of statistical calibration. We assume the team consists of an AI model and human---both of which are calibrated with respect to some partitioning of the feature space---and expose how the calibration assumptions propagate into the teaming framework. In particular, we consider frameworks that either (i) combine human and model predictions or (ii) delegate prediction responsibility to either a human or model. We show via theoretical and empirical results that existing methods for combination do not preserve the human's degree of calibration. Methods for delegation (by the very act of delegation) preserve calibration of the downstream predictors but can place high-capacity demands on the meta-classifier that determines which should predict. This latter result suggests human-AI complementarity and correct delegation place opposite demands on the delegation model.


HuPER: A Human-Inspired Framework for Phonetic Perception

Chenxu Guo ⋅ Jiachen Lian ⋅ Yisi Liu ⋅ Baihe Huang ⋅ Shriyaa Narayanan ⋅ Bixing Wu ⋅ Zoe Ezzes ⋅ Jet Vonk ⋅ Zachary Miller ⋅ Cheol Jun Cho ⋅ Maria Luisa Gorno Tempini ⋅ Gopala Anumanchipalli

We propose HuPER, a human-inspired framework that models phonetic perception as adaptive inference over acoustic--phonetic evidence and linguistic knowledge. HuPER first learns an acoustic-grounded phone recognizer from limited human-annotated data and transcript-only speech: canonical G2P transcriptions are treated as auxiliary linguistic cues rather than ground truth, and are corrected into realized phone proxies for self-training. At inference time, HuPER supports multiple perceptual routes: it can rely on bottom-up phone evidence when the signal is clear, incorporate explicit expectations when a reference is available, or use lexical constraints when acoustic evidence is weak. With only 100 hours of training data, HuPER achieves state-of-the-art phonetic feature error rates on five English benchmarks and strong zero-shot transfer to 95 unseen languages. HuPER also improves robustness on weak-evidence and disordered speech, demonstrating the benefit of adaptive, multi-path phonetic perception under diverse acoustic and task conditions. All training data, models, and code are open-sourced. Code and demo are available at https://github.com/HuPER29/HuPER.

Diffusion models have demonstrated strong performance in time series modeling due to their ability to progressively capture complex data distributions through iterative denoising. However, existing approaches struggle with frequency-sensitive denoising, high-frequency reconstruction and balancing global trends with local dynamics. To address these limitations, we propose \textbf{HyFAD}, a \textbf{Hy}brid time-frequency \textbf{D}iffusion model with \textbf{F}requency-\textbf{A}ware embedding for time series imputation. Built upon the DDPM paradigm, HyFAD adopts a coupled time-frequency diffusion framework, in which the reverse denoising proceeds sequentially from the time domain to the frequency domain, enabling coarse-to-fine generation. Specifically, the time-domain diffusion process captures low-frequency global trends, while the frequency-domain diffusion process refines high-frequency spectral components. We further introduce a frequency-aware step embedding that exploits the relationship between diffusion steps and spectral components, providing step-dependent spectral guidance and facilitates more accurate band-wise reconstruction. Extensive experiments on multiple benchmark datasets demonstrate that HyFAD achieves state-of-the-art performance. Our source code is available at \url{https://anonymous.4open.science/r/HyFAD-0C21/}.


Hyperagents

Jenny Zhang ⋅ Bingchen Zhao ⋅ Wannan Yang ⋅ Jakob Foerster ⋅ Jeff Clune ⋅ Minqi Jiang ⋅ Sam Devlin ⋅ Tatiana Shavrina

Self-improving AI systems aim to reduce reliance on human engineering by learning to improve their own learning and problem-solving processes. Existing approaches to self-improvement rely on fixed, handcrafted meta-level mechanisms, fundamentally limiting how fast such systems can improve. The Darwin G\"odel Machine (DGM) achieves open-ended self-improvement in coding, but this relies on domain-specific alignment between task performance and self-modification skill. However, this alignment does not generally hold beyond coding domains. We introduce \textbf{hyperagents}, self-referential agents that integrate a task agent (which solves the target task) and a meta agent (which modifies itself and the task agent) into a single editable program. Crucially, the meta-level modification procedure is itself editable, enabling metacognitive self-modification, improving not only the task-solving behavior, but also the mechanism that generates future improvements. We instantiate this framework by extending DGM to create DGM-Hyperagents (DGM-H), eliminating the need for domain-specific alignment to potentially support self-accelerating progress on any computable task. Across diverse domains, the DGM-H improves performance over time and outperforms baselines without self-improvement or open-ended exploration, as well as prior self-improving systems. Furthermore, the DGM-H improves the process by which it generates new agents (e.g., persistent memory, performance tracking), and these meta-level improvements transfer across domains and accumulate across runs. All experiments were conducted with safety precautions (e.g., sandboxing, human oversight). We discuss what safety entails in this setting and the broader implications of self-improving systems. DGM-Hyperagents offer a glimpse of open-ended AI systems that do not merely search for better solutions, but continually improve their search for how to improve.


Hyperbolic Displacement-Constrained Adaptation for CLIP-Based Class-Incremental Learning

Quan Cheng ⋅ Zhenhao Wen ⋅ Lan Li ⋅ Da-Wei Zhou ⋅ Lijun Zhang

Class-incremental learning aims to learn a stream of new classes while preserving knowledge from prior tasks, but suffers from catastrophic forgetting when the model is adapted across sessions. Recently, approaches leveraging vision-language pretrained models such as CLIP have gained increasing popularity in this setting, due to the strong transferable representations of foundation models. Most CLIP-based approaches resist forgetting by retaining data-derived class memory such as replay samples, visual prototypes, or class distribution information. In contrast, we ask whether stable adaptation can be achieved by constraining adapter-induced feature displacement without retaining any data-derived class memory. We propose \textbf{H}yperbolic \textbf{D}isplacement-\textbf{C}onstrained \textbf{A}daptation (HDCA), which treats forgetting as a consequence of uncontrolled residual displacement induced by sequential adapter updates. HDCA keeps the direction of each adapter update but compresses its radial magnitude, so small updates remain almost unchanged while large or accumulated updates are suppressed. At the per-task level, a hyperbolic low-rank adapter makes each task residual grow sublinearly. At the multi-task level, a cumulative compression stage further suppresses the summed residual across sessions. On the classifier side, HDCA introduces a parameter-free Poincar\'e distance classifier that preserves cosine-based ranking for fixed adapted features but reshapes the training gradient in high-similarity regions. We conduct experiments on five standard class-incremental datasets, and the results comprehensively validate the effectiveness of HDCA without retaining any data-derived class memory during training and inference.


ICAT: Incident-Case–Grounded Adaptive Testing for Physical-Risk Prediction in Embodied World Models

Zhenglin Lai ⋅ Sirui Huang ⋅ Yuteng Li ⋅ Changxin Huang ⋅ Jianqiang Li ⋅ Bingzhe Wu

Video-generative world models are increasingly used as neural simulators for embodied planning and policy learning, yet their ability to predict physical risk and severe consequences is rarely evaluated.We find that these models often downplay or omit key danger cues and severe outcomes for hazardous actions, which can induce unsafe preferences during planning and training on imagined rollouts. We propose ICAT, which grounds testing in real incident reports and safety manuals by building structured risk memories and retrieving/composing them to constrain the generation of risk cases with causal chains and severity labels. Experiments on an ICAT-based benchmark show that mainstream world models frequently miss mechanisms and triggering conditions and miscalibrate severity, falling short of the reliability required for safety-critical embodied deployment.


iCATS: Fast Video Generation via Interaction-Aware Sparse Attention and Timestep-Adaptive Sparsity

Chengfeng Han ⋅ Baole Ai ⋅ Xianlu Bian ⋅ Jie Yao ⋅ Zilong Huang ⋅ Ang Wang ⋅ Dandan Ding

Training-free sparse attention offers a practical acceleration solution to Diffusion Transformers (DiTs) via reducing computations without fine-tuning. It typically involves estimating the importance of query-key regions and deriving sparse masks to compute only the important candidates, which inevitably introduces approximation errors that may degrade generation quality. To better balance the efficiency-quality trade-off, we propose iCATS, integrating improved importance estimation and sparse mask construction with an efficient hardware execution strategy. Specifically, for importance estimation, unlike previous works that perform independent clustering over query and key tokens based on feature similarity to estimate attention scores, iCATS demonstrates that clustering based on query-key dot-product interactions is more accurate and further reformulates this objective as a simple quadratic form for low-cost computation. For sparse mask construction, instead of using a fixed top-$p$ rule, we observe that tolerance to sparse approximation errors varies across denoising timesteps and therefore introduce an SNR-guided sparsity schedule to adjust sparsity dynamically, leading to higher accuracy. Finally, for hardware execution, we devise a tail-merging strategy to reduce padding overhead caused by irregular cluster sizes, improving GPU kernel utilization. Extensive experiments show that iCATS achieves 2.03$\times$ acceleration with 31.017 dB PSNR on HunyuanVideo-T2V-13B and 1.55$\times$ acceleration with 29.301 dB PSNR on Wan2.1-T2V-14B, delivering a state-of-the-art efficiency-quality trade-off.

The joint expectation of functions of potential outcomes $(Y_1,Y_0)$ is a fundamental quantity in causal inference, particularly when investigating the variation or heterogeneity of causal effects. This paper provides novel assumptions for their identification: comonotonicity and countercomonotonicity. These assumptions offer a unified framework for their identification across discrete, continuous, and mixed outcome variables. In the absence of these assumptions, we derive sharp bounds for two broad classes of functions, yielding new results for bounding moments of individual causal effects, which are central to measuring effect heterogeneity. We present corresponding estimation methods and illustrate them in simulation studies and real-world datasets.


ImageNet FID is a Pass Check, Not a Finish Line

Qinyu Zhao ⋅ Guangting Zheng ⋅ Caixia Zhou

ImageNet class-conditional generation is a standard first test for new image generation methods. Progress on ImageNet is usually measured with one number: Fréchet Inception Distance (FID), computed by comparing 50K class-balanced generated samples with the ImageNet training set in the Inception-v3 feature space. The FID number was useful when the generation quality was low and the gap between methods was large. However, many recent methods already reach remarkable and close FID numbers, and we argue that the field should stop treating FID as decisive evidence of generation quality. To support this view, we re-evaluate 47 recent ImageNet class-conditional generators while keeping the generated samples fixed and changing only the evaluation protocol. We test multiple variants of FID, including different real reference sets, sample sizes, and feature extractors such as Inception-v3, CLIP, DINOv2, and VGG16. The same generated samples can receive different scores, resulting in varying rankings. Our observation means that a single FID protocol has a risk of being over-optimized and over-interpreted. We do not argue that FID should be removed. Instead, we recomend: keep the conventional ImageNet FID for continuity, but also report the results of some FID varaints to avoid overfitting to one evaluation recipe. If a method reaches comparable and stable ImageNet FID, authors should not be expected to keep chasing the best FID, and the paper should be evaluated based on their technical contributions. After passing the FID test, future work such as text-to-image training and human preference study, is important for the futher development of a generation method.


Imagine Before You Draw: Visual Prompt Engineering for Image Generation

Liyu Jia ⋅ Fengda Zhang ⋅ Jiachun Pan ⋅ Kesen Zhao ⋅ Saining Zhang ⋅ Wang Lin ⋅ Weijia Wu ⋅ Yue Liao ⋅ Aojun Zhou ⋅ Hanwang Zhang

Incorporating visual semantic representations as an intermediate step before image generation can reduce the modeling difficulty between text and images, thereby improving generation quality. Recent works such as X-Omni and BLIP3o-Next have explored this direction, but they typically use a two-stage external pipeline: a separate autoregressive model first generates semantic tokens, which are then fed as conditioning to an independent diffusion decoder. Since the decoder cannot jointly access the original input and the semantic plan, this design introduces an information bottleneck that limits detail preservation in downstream tasks such as editing. Internal architectures such as Transfusion, BAGEL, and Show-o2 avoid this bottleneck by enabling cross-modal interaction within a single model, but they still face the difficult text-to-pixel modeling gap without intermediate semantic guidance. We propose Visual Prompt Engineering (VPE), which can be seamlessly integrated into such internal frameworks. Specifically, the model first autoregressively generates visual semantic tokens ($e.g.$, SigLIP 2) as "visual prompts" that capture the semantic layout, then generates the full image tokens conditioned on this plan. We validate VPE across class-conditional generation, text-to-image generation, and image editing, covering various token types and model architectures. Results show that VPE can accelerate convergence, raise quality ceilings, and through internal integration, achieve substantially better editing preservation (PSNR: $26.76$ vs. $19.92$) than external alternatives of the same parameter scale, while maintaining competitive editing responsiveness. The code is available in supplementary material.


Imperfect World Models are Exploitable

Logan M Bhamidipaty ⋅ Esmeralda S Whitammer ⋅ David Abel ⋅ Mykel J Kochenderfer ⋅ Subramanian Ramamoorthy

We propose a novel definition of model exploitation in reinforcement learning. Informally, a world model is exploitable if it implies that one policy should be strictly preferred over another while the environment’s true transition model implies the reverse. We analogize our definition with a prior characterization of reward hacking but show that the associated proof of inevitability does not transfer to exploitation. To overcome this obstruction, we develop a general theory of reward hacking and model exploitation that proves that exploitation is essentially unavoidable on large policy sets and yields the corresponding claim for hacking as a special case. Unfortunately, we also find that the conditions that guarantee unhackability in finite policy sets have no counterparts that preclude exploitation. Consequently, we introduce a relaxed notion of exploitation and derive a safe horizon within which it can be avoided. Taken together, our results establish a formal bridge between reward hacking and model exploitation and elucidate the limits of safe planning in world models.


Implicit Bias in State Space Models and Linear Autoregressive Training

Gal Katzhendler ⋅ Gal Vardi ⋅ Gilad Yehudai

Linear State Space Models (SSMs) and autoregressive sequence models reuse the same parameters across a sample-dependent number of steps. As a result, different examples impose constraints through different predictors, even though all predictors share the same weights. We study this phenomenon through a general model of multiple non-homogeneous polynomial predictors with shared parameters. Under standard interpolation and directional-convergence assumptions, together with a trajectory-dependent effective-degree condition, we show that gradient descent on the exponential loss converges in direction to a KKT point of an \emph{effective-degree max-margin problem}. In this problem, only examples with minimal effective degree impose hard margin constraints; examples with faster-growing margins are asymptotically inactive except for feasibility. Our result provides a characterization of the implicit bias in linear SSMs and linear autoregressive models. It exposes a mechanism of length bias: a longer input sequence or additional autoregressive steps may increase the example's effective degree, and the examples with the smallest effective degree dominate the limiting classifier. Our result highlights how non-homogeneity and parameter sharing alter classical homogeneous implicit-bias results.


Implicit Drifting Policy: One-Step Action Generation via Conditional Expert Geometry

Zemin Yang ⋅ Yaoyu He ⋅ Yiming Zhong ⋅ Yuhao Zhang ⋅ Xinge ZHU ⋅ Yao Mu ⋅ Qingqiu Huang ⋅ Yuexin Ma

Generative action policies based on diffusion or flow matching excel in behavior cloning, yet their iterative sampling is prohibitive for high-frequency robot control. While recent one-step formulations alleviate this latency, they inevitably discard the intermediate trajectory evolution that provides crucial action correction. Directly recovering this mechanism by explicitly estimating a training-time drifting field is mathematically ill-posed due to extreme conditional demonstration sparsity. We introduce Implicit Drifting Policy (IDP) , a one-step imitation learning framework that brings the training-time correction of Drifting into policy learning without explicit vector field estimation. IDP extracts a conditional expert geometry from the local variation of observation-similar expert actions, and compares it against a global reference geometry to isolate condition-specific constraints. This local geometric structure adaptively weights a scalar potential objective. Combined with an expert-proximal terminal evaluation, IDP directly enforces manifold constraints on the one-step generator during training. Extensive evaluations across 2D, 3D, and real-world manipulation tasks show IDP effectively maintains adherence to valid action manifolds, improving upon explicit drifting methods and achieving competitive performance with strong one-step baselines.


Importance-Aware OBS Pruning for Diffusion Models

Ba-Thinh Lam ⋅ Srijan Das ⋅ Hieu Le

We propose importance-aware pruning for diffusion models, a training-free framework that prioritizes preserving parameters critical to semantically salient image regions. To do so, we incorporate spatial importance maps- derived from conditioning signals or model attention- into the pruning objective. This produces parameter rankings aligned with perceptual relevance rather than uniform reconstruction error. On MS-COCO dataset, our proposed approach consistently retains subject fidelity and structural correctness at high compression ratios where conventional pruning causes visible degradation. These results demonstrate that content-aware objectives are key to perceptually faithful compression of generative models.


Improving General Role-Playing Agents via Psychology-Grounded Reasoning and Role-Aware Policy Optimization

Zhenhua Xu ⋅ Dongsheng Chen ⋅ Jian Li ⋅ Yitong Lin ⋅ Zhebo Wang ⋅ Jiafu Wu ⋅ Yizhang Jin ⋅ Chengjie Wang ⋅ Meng Han ⋅ Yabiao Wang

Building general-purpose role-playing agents that faithfully portray any character from a natural-language profile remains challenging. The dominant paradigm---supervised fine-tuning---encourages behavioral mimicry without deep, human-like internal thought processes, resulting in poor out-of-distribution generalization. Therefore, we propose Psy-CoT, a psychology-grounded chain-of-thought framework that decomposes pre-response reasoning into three role-specific steps---Interaction Perception, Psychological Empathy, and Logical Construction---so that the model thinks dynamically from the profile rather than merely mimicking surface patterns. While structured reasoning provides a foundation, it alone is insufficient; reinforcement learning is essential to further align the model with character fidelity. However, we observe that under LLM-based reward models, both generic phrases that hack the reward model and genuinely role-specific phrases receive identical gradient signals---this hacking accumulates over training, misleading the model into treating both as equally optimal choices. To address this, we propose Role-Aware Policy Optimization (RAPO), which uses profile--token mutual information to weight gradients asymmetrically---amplifying role-specific tokens under positive advantage while attenuating them under negative advantage. Experiments on CoSER, CharacterBench, and CharacterEval demonstrate that Psy-CoT outperforms existing role-playing CoT methods, and RAPO consistently surpasses GRPO across multiple model scales.

In modern GANs, maintaining an Exponential Moving Average (EMA) of the generator's weights is a standard practice, as such an averaged model consistently outperforms the actively trained generator. However, the EMA generator is used for final deployment only and does not influence the training process. To address this missed opportunity, we introduce Self-Distilled GAN (SD-GAN) that employs the EMA generator as a teacher to guide the active generator (student) via perceptual loss. We prove the local asymptotic stability of SD-GAN in the Dirac-GAN setting and show that it dampens the parasitic cycling behavior that plagues the conventional GANs. Empirical evaluations across established architectures and datasets demonstrate that SD-GAN improves the final image quality on several metrics (FID and random-FID in particular), stabilizes the optimization trajectory and provides additional learning guidance that is not trivially correlated with the conventional adversarial loss. It also proves effective for fine-tuning pretrained GAN models.


IMTS-Tokenizer: Time-Aware Tokenization for Irregular Multivariate Time Series Forecasting

Bin Xu ⋅ Yinghua Li ⋅ Linqi Han ⋅ Xiaoyu Li ⋅ Xinghao Yang ⋅ Wei Liu ⋅ Yongshun Gong

Modeling Irregular Multivariate Time Series (IMTS) poses significant challenges due to asynchronous sampling and data sparsity. While existing methods focus on handling temporal irregularity after encoding, the tokenization stage itself remains underexplored. We propose IMTS-Tokenizer, a time aware tokenization framework built on the Irregular-time-aware Module (IRAM), which employs learnable temporal anchors and Gaussian-kernel attention to aggregate observations while preserving temporal information. To improve robustness under sparse and asynchronous sampling, the framework further incorporates Relative Time Feature (RTF) for pre-aggregation temporal encoding, Variable-Specific Anchor Initialization (VSA) for data-aligned anchor placement, anchor regularization for stable training, and multi-scale modeling. Experiments on four benchmarks demonstrate state-of-the-art (SOTA) performance, with the best results on seven out of eight metrics and statistically significant aggregate gains.


Incentivizing Medical Vision Capabilities from Large-Scale Multimodal Pre-training

Junying Chen ⋅ Zhenyang Cai ⋅ Yunjin Yang ⋅ Kunyuan Cai ⋅ Rongsheng Wang ⋅ Xidong Wang ⋅ Geyang Yu ⋅ Xiangyi Feng ⋅ Benyou Wang

Building capable medical vision-language models (VLMs) is challenging due to the scarcity of domain data and the difficulty of eliciting pre-trained medical visual knowledge. We present HuatuoGPT-5, a medical VLM built on a simple yet effective recipe: large-scale multimodal pre-training followed by reinforcement learning (RL). We first construct PubMedVision-Plus, the largest medical image-text dataset with 60M pairs from public PubMed literature, along with a clean 12M high-quality subset. After pre-training on this data, we find that traditional supervised fine-tuning (SFT) struggles to elicit the pre-trained knowledge. In contrast, by applying RL with verifiable and rubric-based instances synthesized directly from the pre-training data, we successfully incentivize the model to use this pre-trained medical visual knowledge. On PubMed-derived data, HuatuoGPT-5 outperforms similarly sized VLMs on multiple medical benchmarks and expert evaluations, significantly narrowing the gap between medical VLMs and leading proprietary VLMs. We hope these data, models, and findings can help accelerate open research on capable medical VLMs.


Incorporating Neural Network Structure in the Bayesian Learning Rule

Eiki Shimizu ⋅ Mohammad Emtiyaz Khan ⋅ Thomas Möllenhoff

We extend the Bayesian learning rule by explicitly incorporating the compositional structure of neural networks. The central observation is that backpropagation and exponential family variational inference share a common Lagrangian duality structure. Leveraging this connection, we develop a new Bayesian learning framework in which adjoint variables backpropagate layerwise sensitivity signals that converts into loss-site natural parameters. For Gaussian weight posteriors and Gaussian layer-state projections, the framework recovers deterministic back-propagation and Hessian backpropagation via delta-method approximations, while providing a unified perspective on localization-style prediction rules and moment matching methods. We also derive a new variant of IVON, clarify how its gradient and Hessian estimates differ from those of the standard version, and demonstrate how neural network structure can be incorporated into Bayesian learning in practice.


Incremental Multiple Oracle

Carlos Martin ⋅ Tuomas Sandholm

We present a framework for computing approximate mixed-strategy Nash equilibria of continuous-action games. It is a modification of the multiple oracle algorithm, which extends the double oracle algorithm to multiple players and continuous action spaces. Unlike prior methods, it maintains fixed-cardinality pure strategy sets for each player. Thus, unlike prior methods, only a constant amount of memory is necessary. Furthermore, it does not require exact metagame solving on each iteration, which can be computationally expensive for large metagames. Moreover, it does not require global best-response computation on each iteration, which can be computationally expensive or even intractable for high-dimensional action spaces and general games. Our method incrementally reduces the exploitability of the strategy profile in the finite metagame, pushing it toward Nash equilibrium. Simultaneously, it incrementally improves the pure strategies that best respond to this strategy profile in the full game. We evaluate our method on various continuous-action games, showing that it obtains approximate mixed-strategy Nash equilibria with low exploitability.


Inference-time Alignment via Sparse Junction Steering

Runyi Hu ⋅ Jie Zhang ⋅ Shiqian Zhao ⋅ Jiale Meng ⋅ Jiwei Li ⋅ Jason Zeng ⋅ Ming Wu ⋅ Michael Heinrich ⋅ Yonggang Wen ⋅ Tianwei Zhang

Token-level steering has emerged as a pivotal approach for inference-time alignment, enabling fine-grained control over large language models (LLMs) by modulating their output distributions without parameter updates. While effective, existing methods rely on dense intervention at every decoding step. This persistent manipulation not only incurs substantial computational overhead but also risks compromising generation quality by excessively drifting from the model’s intrinsic distribution. In this work, we show that dense intervention is unnecessary and propose Sparse Inference-time Alignment (SIA), which performs sparse junction steering, intervening only at critical decision points along the generation trajectory. Our key insight is that high-entropy junctions are empirically consistent with pivotal decision points in the generation trajectory and are particularly susceptible to misalignment, suggesting that these points benefit from alignment-related reward signal. Extensive experiments across different model families and alignment objectives show that steering only 20%–80% of tokens achieves the superior alignment–efficiency trade-offs. For strong base models such as Qwen3, intervening on as few as 20% of tokens matches or even surpasses heavily post-trained instruct models. This sparsity enables stronger guidance while better preserving the model's native distribution, integrates seamlessly with search-based methods (e.g., Best-of-N), and improves efficiency by reducing the required sampling budget while yielding practical wall-clock speedups under matched-quality comparisons.


Inpainting physics: self-supervised learning for context-driven fluid simulation

Jonas Weidner ⋅ Yeray Martin-Ruisanchez ⋅ Daniel Rueckert ⋅ Benedikt Wiestler ⋅ Julian Suk

Neural surrogate models for computational fluid dynamics (CFD) are typically trained as forward operators that map explicit problem specifications, such as geometry and boundary conditions, to solution fields. This ties the model to the conditioning variables seen during training and limits reuse under boundary-condition shifts or local geometry changes. We propose to reformulate steady CFD inference as an inpainting problem: instead of training on explicit boundary conditions, we learn a self-supervised prior over velocity fields and impose boundary constraints only during inference by fixing known regions such as inlet, outlet or unchanged regions from previous simulations. To scale this idea to large 3D meshes, we introduce a local neighbourhood tokeniser that represents high-resolution velocity fields as compact spatial latent tokens and train latent flow-matching and masked-autoencoder models on these tokens. On intracranial aneurysm hemodynamics, our method reconstructs full velocity fields from sparse boundary context, outperforms supervised neural surrogates under boundary-condition and dataset shift and enables local geometry editing by reusing unchanged simulation context. These results suggest that viewing CFD inference as context-conditioned inpainting can turn neural surrogates from task-specific predictors into reusable flow priors.


InQuant: In-Place Mixed-Precision KV Cache Quantization via Saliency-Aware Neighbor-Slot Reuse

Zihan Chang ⋅ Shuibing He ⋅ Bo Zhou ⋅ Ping Chen ⋅ Siling Yang

Key-value (KV) caches are essential for efficient Large Language Model (LLM) inference, but their memory footprint grows linearly with context length and batch size. Low-bit KV-cache quantization reduces this footprint, yet uniform quantization is vulnerable to high-magnitude outlier channels, while outlier-aware mixed-precision methods often introduce extra buffers, indexing, or channel permutation that weakens their system-level benefit. This paper presents InQuant, an in-place mixed-precision KV-cache quantization method that preserves salient outlier channels without changing the physical 4-bit packed layout. The key observation is that selected low-saliency channels near outliers can tolerate small approximation error; InQuant therefore reuses their storage slots to hold the extra 4-bit nibble needed by 8-bit outlier values. To make this layout practical, InQuant combines sampling-based channel saliency estimation, saliency-aware neighbor-slot reuse, stride-based handling for adjacent outlier groups, and descriptor-guided marker recovery during dequantization. Across six LLMs and long-context workloads, InQuant reaches a fixed 4.0$\times$ physical KV-cache compression ratio, preserves accuracy close to strong mixed-precision baselines, and reduces quantization/dequantization latency by 2.2$\times$--3.4$\times$ compared with representative outlier-aware methods.


Insight-Driven Search: A Framework for Multi-objective Automated Heuristic Design with Large Language Models

Shunyu Yao ⋅ Fei Liu ⋅ Ji Cheng ⋅ Mingxuan Yuan ⋅ Tong Xialiang ⋅ Zhenkun Wang ⋅ Qingfu Zhang

Automated Heuristic Design (AHD) with Large Language Models (LLMs) has shown promise in solving optimization tasks. However, existing methods mainly rely on iterative search frameworks and do not systematically structure or reuse task-specific design knowledge elicited from LLM priors and accumulated during online search. Moreover, multi-objective AHD introduces additional challenges, as conflicting objectives such as solution quality and computational efficiency must be balanced during heuristic search. We propose Multi-objective Evolution of Heuristics with Insight-Driven Search (MEoH-IDS), which uses dynamically maintained design insights to guide LLM-based heuristic generation. MEoH-IDS maintains an insight pool by extracting insights from successful heuristics, evaluating their search utility, and filtering less effective ones over time. It selects promising insight combinations using a UCB-inspired policy, with rewards based on non-dominated status in the current Pareto archive. We validate MEoH-IDS on three AHD tasks, including the Traveling Salesman Problem (TSP), Capacitated Vehicle Routing Problem (CVRP), and Vehicle Routing Problem with Time Windows (VRPTW), optimizing both solution quality and computational efficiency. Experimental results show that MEoH-IDS accelerates convergence and discovers heuristics with improved quality--efficiency trade-offs compared with representative LLM-based AHD baselines.


Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

Jon Saad-Falcon ⋅ Avanika Narayan ⋅ Hakki Akengin ⋅ J. W Griffin ⋅ Herumb Shandilya ⋅ Adrian G Lafuente ⋅ Medhya Goel ⋅ Rebecca Joseph ⋅ Shlok Natarajan ⋅ Etash Guha ⋅ Shang Zhu ⋅ Ben Athiwaratkun ⋅ Azalia Mirhoseini ⋅ Christopher Ré

Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Rapidly growing demand strains this paradigm, and cloud providers struggle to scale infrastructure at pace. Two advances create an opportunity to rethink this paradigm: small, local LMs (≤20B active parameters) now achieve competitive performance to frontier models on many tasks, and local accelerators (e.g., Apple M4 Max) can host these models at interactive latencies. This raises the question: can local inference viably redistribute demand from centralized infrastructure? Answering this requires measuring both whether local LMs can accurately answer real-world queries and whether they can do so efficiently enough to be practical on power-constrained devices (i.e., laptops). We propose intelligence per watt (IPW), task accuracy divided by unit of power, as a unified metric for assessing both the capability and efficiency of local inference across model-accelerator configurations. We conduct a large-scale empirical study across 20+ state-of-the-art local LMs, 8 hardware accelerators (local and cloud), and a representative subset of LLM traffic: 1M real-world single-turn chat and reasoning queries. For each query, we measure accuracy (local LM win rate against frontier models), energy consumption, latency, and power. Our analysis reveals three key findings. First, local LMs can successfully answer 88.7% of single-turn chat and reasoning queries with accuracy varying by domain. Second, longitudinal analysis from 2023-2025 shows progress in local inference viability: IPW improved 5.3×, driven by both algorithmic advances and accelerator improvements, with locally-serviceable query coverage increasing from 23.2% to 71.3%. Third, local accelerators achieve at least 1.4× lower IPW than cloud accelerators running identical models, revealing significant headroom for local accelerator optimization. These findings demonstrate that local inference can meaningfully redistribute demand from centralized infrastructure for a substantial subset of queries, with IPW serving as the critical metric for tracking this transition.

Recent advances in robotics foundation models have shown promising multi-task generalization, yet their reliance on specialized embodiment-specific data and model fundamentally prevents them from exploiting the far richer physical interaction knowledge present in abundant human manipulation videos. Existing approaches that leverage human videos suffer from misaligned representations that preclude explicit unified interaction modeling over human video and robot demonstration, failing to capture unified physical interaction dynamics essential for cross-domain transfer. We address this by constructing a unified representation that aligns human videos and robot tasks through scene point clouds with hand-to-gripper mapping via spatial tracking and hand pose estimation. Leveraging this representation, we propose a structured graph transformer that explicitly models spatial, semantic, and intentional interaction entities through graph attention mechanisms. The modeled multi-level interaction features then enable interaction-aligned cross-domain learning, transferring manipulation knowledge from human videos to robot tasks via discrete interaction abstraction and hierarchical distribution alignment. Experiments on RLBench, ManiSkill2, and real-world robot demonstrate state-of-the-art performance on diverse benchmarks with fewer demonstrations and effectiveness of our designed modules. Remarkably, our method achieves zero-shot transfer on robot tasks when trained exclusively on human video, providing strong evidence for effective human-to-robot transfer through interaction alignment.


Interaction Value Inference for Multi-Agent Reinforcement Learning via a Hierarchical Agent-Centric World Model

Zhuoran Chen ⋅ Xuyang Lu ⋅ Zeyang Liu ⋅ Xinrui Yang ⋅ Mingyang Li ⋅ Long Qian ⋅ Xingyu Chen ⋅ Lipeng Wan ⋅ Xuguang Lan

In non-stationary multi-agent settings, concurrent policy updates continually reshape inter-agent coordination patterns. In cooperative tasks, an agent’s current action value is closely tied to teammates’ current and future behaviors, making it important to infer how these behaviors evolve and how they influence local action-value estimation. However, existing methods often lack a value-level representation that connects evolving teammate-dependent interactions to local action-value learning. Firstly, we theoretically show that the individual utilities under Centralized Training with Decentralized Execution(CTDE) paradigm struggle to faithfully characterize coordination-dependent action values, making it necessary to introduce a computable proxy for the missing coordination information. Then we propose Hierarchical Inference Agent-Centric World Model (HIA), a framework that incorporates role information into a Transformer-based world model to predict future teammate interactions from each agent’s perspective. The predicted teammate-action rollouts are transformed into auxiliary interaction utilities and integrated into individual value estimates, enabling each agent to capture how teammates’ future actions modulate its current action value. Experiments on StarCraft Multi-Agent Challenge (SMAC) and SMACv2 show that HIA achieves superior cooperative performance over strong baselines.


Inter-Agent Influence: Evaluating Persuasion, Deception and Coercion in Multi-Agent Systems

Chandler Smith ⋅ Cecilia E Tilli ⋅ Qi Guo ⋅ Sophia Hatz ⋅ David Demitri Africa ⋅ Patricia Paskov ⋅ Christian Schroeder de Witt ⋅ Philip Torr ⋅ Lewis Hammond

The deployment of AI agents in multi-agent workflows enables inter-agent influence, whereby agents strategically steer other agents’ behavior in line with a specific goal. Such capabilities pose novel risks, as malicious agents could direct others toward harmful actions. In this work we investigate three such capabilities –- persuasion, deception, and coercion –- across five realistic evaluation environments and a variety of frontier models. We observe significant inter-agent influence capabilities in frontier models. In oversight environments, all tested models shifted policy-violating decisions from rejection to approval through persuasion, deception, and coercion. In peer-to-peer settings, models extracted concessions in scheduling negotiations and redirected a peer's safety research trajectory toward an attacker-preferred direction. While evidence is weakest for inter-agent coercion, several frontier models independently formulated and executed coercive threats, including threats to harm a named human in the environment. These results point to a pressing need for work on risk mitigation to promote beneficial deployments of multi-agent systems.


Internal Safety Collapse in Frontier Large Language Models

Oscar W ⋅ Xiao Liu ⋅ Hanxun Huang ⋅ Yige Li ⋅ Xiang Zheng ⋅ Yifeng Gao ⋅ Cong Wang ⋅ Bo Li ⋅ Xingjun Ma ⋅ Yu-Gang Jiang

This work identifies a critical failure mode in frontier large language models (LLMs), which we term \textbf{Internal Safety Collapse} (ISC): \textit{under certain task conditions, models enter a state in which they continuously generate large volumes of harmful content while executing otherwise benign tasks}. To systematically study ISC, we introduce \TVD{} (Task, Validator, Data), a framework that instantiates controlled workflows around domain tools, where valid completion requires filling harmful content as structured data. We collect 53 representative workflows across 8 disciplines, from toxicity evaluation to molecular docking and pathogen genome analysis, showing that ISC reproduces across all eight disciplines. In the worst case over three interaction settings, the four frontier LLMs average a \textbf{95.3\%} safety-failure rate. Frontier LLMs with stronger task-completion capability show higher unsafe-completion rates: long-horizon execution skill becomes a liability when workflow completion requires harmful content. We even observe extremely severe harmful content closely resembling outputs from early-generation, unaligned LLMs in 2023. Despite substantial safety alignment efforts, frontier LLMs continue to retain inherently unsafe internal capabilities: alignment reshapes observable outputs but does not eliminate the underlying risk profile. These findings underscore the need for caution when deploying LLMs in high-stakes settings, including scientific pipelines and autonomous agents. Complementary source code is provided at \url{https://anonymous.4open.science/r/NIPS-ISC-Code-Share-BE11}.

Multimodal 3D detection becomes unreliable under partial observability caused by occlusion, LiDAR sparsity, and camera--LiDAR miscalibration. Most fusion detectors are fully feed-forward: once trained, they provide no explicit mechanism to incorporate structured test-time constraints or to bound how such constraints can alter predictions. We propose Intervene3D, a controlled-inference wrapper for BEV-based detectors that performs metric-bounded latent refinement at inference time. Given sensor inputs, the detector produces an initial BEV latent $\mathbf z_0$ and an evidence-conditioned trust metric $\mathbf H$ (a diagonal precision derived from predicted uncertainty). A separate control stream is mapped to a fully specified differentiable constraint energy (count intervals and spatial admissibility). Inference then updates $\mathbf z$ using $K\in\{1,2,3\}$ diagonal-metric preconditioned projected steps that reduce constraint energy while remaining within an uncertainty-gated trust region defined by $\mathbf H$. To improve robustness to extrinsic drift, we additionally introduce \textbf{phase-aware spectral alignment} that normalizes cross-modal Fourier-phase discrepancies prior to attention. We report clean accuracy and a severity-swept robustness protocol (robustness-AUC, feasibility diagnostics, and conflict-induced FP inflation).


Introspective Coupling: LMs Learn to Explain Themselves Better Than Their Training Targets

Zifan Carl Guo ⋅ Laura Ruis ⋅ Jacob Andreas ⋅ Belinda Z Li

When does training language models (LMs) on explanations yield faithful introspection, rather than superficial imitation? Surprisingly, we find that LMs trained to explain the predictions of similar models frequently produce explanations more faithful to $\textit{their own current behaviors}$ than to those of their training targets. This "introspective'' coupling between the model's explanations and behaviors occurs only when the training target explanations remain sufficiently similar to model behaviors over the course of training. This alignment must be preserved throughout the training process: introspection only emerges when explanation supervision is sufficiently behaviorally compatible with the model as it changes. Finally, we show that introspection generalizes to variants of the training problem: when introspection training is run concurrently with training that shifts a model's behaviors, explanations track those behavioral shifts without requiring updated supervision. This holds across a diverse range of tasks, including sycophancy and refusal, and is robust to label noise. These results suggest that introspection training is a viable component of post-training: explanation labels need not be continually refreshed, and faithfulness extends to regions of input space not explicitly supervised.


Inverse Modeling of Neural Recordings via Differentiable Biophysical Simulation

Frithjof Gressmann ⋅ Ngoc H Pham ⋅ Lawrence Rauchwerger

Inferring the latent processes that generate observable neural activity is a central challenge in neuroscience, with direct implications for brain-machine interfaces and neural engineering. Existing inverse approaches typically summarize the recording into low-dimensional features or learn black-box latents that lack a direct biophysical interpretation. A promising alternative is to fit a differentiable biophysical simulator end-to-end against the full extracellular signal, using direct and scalable gradient optimization. In practice, however, recovering a per-neuron input via backpropagation through neuronal dynamics is difficult because the loss landscape is nearly flat in the subthreshold regime and jumps sharply at spike threshold, leaving gradient descent without a useful signal. We observe that per-neuron spike times, routinely available from spike sorting, expose discrete millisecond-scale anchors at which the loss does carry information. Gradients from these anchors flow back in time through the differentiable simulator and shape the subthreshold drive that produced each spike. Building on this, we present a fully differentiable pipeline coupling biophysical membrane dynamics to a volume conductor model of the multi-electrode array. For each neuron in a recorded population, the method jointly recovers a time-varying latent input current and a probe-relative position consistent with both the extracellular trace and the observed spike times. No explicit likelihood or posterior estimation is required, and every fitted latent corresponds to a named biophysical quantity. On paired patch-clamp / Neuropixels SPE-1 data, the model localizes the patched neuron to within $40 \mu m$ on average across cells, while reproducing observed spatio-temporal activity patterns under realistic noise. These results position scalable differentiable biophysical simulation as a practical route to mechanistic, gradient-based analysis of high-density neural recordings.


INVITA-WheatFieldState: A Real-World Benchmark for Crop-State Estimation in Wheat Field Trials

Heming Du ⋅ Ruihan Lu ⋅ Javier Fernandez ⋅ Zijian Wang ⋅ Yan Zhao ⋅ Scott C Chapman ⋅ Zi Huang ⋅ Xin Yu

Field trials use plot-level crop-state measurements to interpret how crops grow under real field conditions, but repeated measurements of canopy greenness, canopy amount, canopy closure, and growth stage require field visits, sensing campaigns, and post-processing. Modern wheat trials also collect weather records, trial metadata, field-camera images, UAV and satellite observations, proximal sensors, canopy products, and phenological surveys. Turning these records into a machine-learning benchmark is nontrivial because observations are sparse, asynchronous, unevenly collected, missing for structured reasons, and tied to measurement pipelines. We introduce INVITA-WheatFieldState, a benchmark derived from the INVITA wheat field-trial archive for estimating the crop state of a plot on a target date from observations available for that plot and date. Each plot-date example specifies a plot, date, and crop-state target, and is paired with available observations after temporal and source-provenance filtering. The derived dataset contains over 220k examples across NDVI, LAI, FCover, and Zadoks growth stage, covering canopy greenness, canopy amount, canopy closure, and phenological timing. LAI and FCover are released with provenance labels that distinguish product and proxy targets rather than pooled manual ground truth. The benchmark suite provides a plot-disjoint split matched to the plot-level experimental unit, validators, prediction schemas, regression representations, and prediction-level diagnostics. Results show that source, timing, metadata, and observation-availability structure are strong signals. Observation-set representations improve some all-example targets, while sensor-sequence, field-camera, and fusion representations remain target-dependent and coverage-bound. INVITA-WheatFieldState offers a concrete test case for ML methods that must learn from incomplete, asynchronous, and provenance-sensitive measurements used to study crop development in real fields.


I-Perceive: A Foundation Model for Vision-Language Active Perception

Yongxi Huang ⋅ Wang ⋅ Wenjing Tang ⋅ Xinyu He ⋅ Cewu Lu ⋅ Panpan Cai

Active perception—the ability of a robot to proactively select viewpoints to acquire task-relevant information—is essential for robust operation in real-world environments. However, existing approaches are typically limited to fixed objectives or constrained settings, and struggle to generalize to open-ended perception intents specified in natural language. We propose I-Perceive, a foundation model for language-conditioned active perception in large-scale indoor environments. Given a query image, a set of context images, and a natural language instruction, I-Perceive predicts a 6D camera pose that fulfills the specified perception intent. The model integrates a vision-language pathway for semantic grounding with a geometric reasoning pathway for multi-view 3D understanding, connected via multi-layer semantic fusion to enable language-conditioned geometric reasoning. To support scalable training, we construct a large-scale dataset of language-viewpoint pairs from both real-world scene-scanning data and simulated environments using an automated pipeline. Extensive experiments demonstrate that I-Perceive significantly outperforms strong baselines on prediction accuracy, viewpoint feasibility, and instructions alignment. The model exhibits strong zero-shot generalization to unseen scenes and instructions, and enables closed-loop active perception, progressively refining viewpoints over sequential interactions.


Iris: Empowering Video MLLMs with High-Frequency Pose Priors via Spatiotemporal Binding

Jiahang Zhang ⋅ Yushuo Guan ⋅ Yuanxing Zhang ⋅ Pengfei Wan ⋅ Jiaying Liu

Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in general video understanding. Human motion understanding, a dominant topic in video analytics, however, still remains as a bottleneck due to the lack of explicit structural cues in 2D visual patches and the significant motion information loss caused by sparse visual temporal sampling. To address this, we propose Iris (Integrating representations of in-the-wild skeletons), a novel pose-augmented video MLLM tailored for human-centric motion understanding. Iris adopts an asymmetric dual-stream architecture, pairing the traditional sparse vision stream with a lightweight, high-frequency human pose stream to provide spatially structured and temporally dense kinematic signals. To enable pose-vision modality correspondence awareness, we design a cross-modal spatiotemporal binding mechanism, featuring a pose-structured 3D RoPE for precise temporal alignment and a person-grounded cross-attention module for explicit spatial visual grounding. Furthermore, to enable robust large-scale training, we construct an automated pose curation pipeline that organizes pose data into a delicate heterogeneous representation. Extensive experiments demonstrate that Iris achieves superior performance on multiple fine-grained motion benchmarks, while also maintaining highly efficient token computation overhead. Code will be available upon publication.


Is the Importance Ratio Necessary for Stable Reinforcement Learning in LLMs?

Shuibai Zhang ⋅ Junhyuck Kim ⋅ Gyeongman Kim ⋅ Jaewoong Cho

Reinforcement learning (RL) has become central to post-training large language models (LLMs). However, popular RL methods like GRPO incur non-negligible overhead by computing both old-policy and current-policy likelihoods to form importance sampling ratios. In this work, we propose Likelihood-Gated Policy Optimization (LGPO), which enforces a soft trust region constraint via likelihood-based gating, eliminating the need to compute old-policy likelihoods. Empirically, we show that removing the importance sampling correction term does not harm training stability, whereas removing the trust region mechanism leads to collapse. Moreover, ratio-based clipping can fail in fully on-policy training: the importance ratio stays at 1, so the ratio-based trust region constraint never activates. Under standard training settings where GRPO is stable, LGPO achieves comparable training stability and peak performance while reducing training time by ~18\% on average. In fully on-policy training, where GRPO fails, LGPO remains stable, enabling more efficient and robust LLM RL post-training across training regimes.


Iterative Scarcity-Guided Exploration: Bootstrapping Generative Auto-bidding from Narrow Support

Qingmao Yao ⋅ Guangzheng Hu ⋅ Yusen Huo ⋅ Chu Xu ⋅ Zhilin Zhang ⋅ Chuan Yu ⋅ Jian Xu ⋅ Bo Zheng ⋅ Xiaotie Deng

Auto-bidding is a core algorithmic component in online advertising auctions. In practical cold-start settings, iterative training is commonly used to compensate for low-quality, narrowly supported historical data. Unfortunately, in iterative training, representative diffusion-based AI-Generated Bidding (AIGB) methods fail to sustain extrapolation beyond the narrow data support and stagnate at a suboptimal level. In this paper, we theoretically attribute this stagnation to Signal-to-Noise Ratio (SNR) collapse: the weak radial return signal is overwhelmed by a curvature-induced penalty. To break this stagnation, we propose \textbf{Iterative Scarcity-Guided Exploration (ISGE)}, which introduces a guidance handoff: as the return signal collapses, scarcity subsequently takes over as an exploratory signal to elevate the SNR above a critical threshold. Specifically, ISGE iteratively alternates between a \emph{Judger} that assigns scarcity scores to trajectories and an \emph{Explorer} that performs guided generation of high-scarcity and high-return trajectories, thereby bootstrapping from the narrow support with such self-generated trajectories. Extensive experiments on the industrial AuctionNet benchmark demonstrate that ISGE, starting from cold-start datasets, effectively surpasses the performance of the Full-Dataset baseline within three iterations.


JailBound: A FOL-Guided Jailbreak Evaluation Framework for Revealing Safety Boundaries of LLMs

Fazong Wu ⋅ Ming Yang ⋅ Xin Wang ⋅ Zhenyong Zhang ⋅ Xiaoming Wu

Large language models (LLMs) are now deployed in a wide range of real-world applications, making it important to evaluate how reliably they resist jailbreak attacks. However, existing jailbreak evaluations mainly rely on manually collected prompts or discrete text-space optimization, which limits their coverage and makes them difficult to extend to new threat settings. We present $\textbf{JailBound}$, a jailbreak evaluation framework that combines automated benchmark construction with intent-preserving attack optimization in embedding space. JailBound organizes evaluation instances with a hierarchical threat taxonomy spanning risk categories, application domains, and attack types, and uses this structure to generate meta-attack prompts with a fine-tuned meta-attack generator. It further formulates jailbreak evaluation as an embedding-space attack optimization problem and uses a first-order loss (FOL)-guided dual-branch search to jointly identify high-value vulnerable regions and safety boundary states. Under a unified evaluation protocol, we study 46 LLMs from 13 model families. The results show that JailBound covers a broader range of risk settings than existing prompt-based benchmarks, and that its optimized attacks transfer nontrivially across model families while supporting finer-grained analysis of vulnerability patterns and safety boundary behavior. $\textcolor{red}{\text{Warning: this paper includes examples that may be offensive or harmful.}}$


Joint Learning of Hierarchical Neural Options and Abstract World Model

Top Piriyakulkij ⋅ Wolfgang Lehrach ⋅ Kevin Ellis ⋅ Kevin Murphy

Building agents that can perform new skills by composing existing skills is a long-standing goal of AI agent research. Towards this end, we investigate how to efficiently acquire a sequence of skills, formalized as hierarchical neural options. However, existing model-free hierarchical reinforcement algorithms need a lot of data. We propose a novel method, which we call "AgentOWL" (Option and World model Learning), that jointly learns --- in a sample efficient way --- an abstract world model (abstracting across both states and time) and a set of neural options. We show, on a subset of Object-Centric Atari games, that our method can learn more skills using less data than baseline methods and possesses learning and generalization capabilities that the baselines do not have.


JuICE: A Benchmark for Evaluating LLM-Judge in Identifying Cultural Errors

Jiho Jin ⋅ Junho Myung ⋅ Juhyun Oh ⋅ Junyeong Park ⋅ Rifki A Putri ⋅ Sunipa Dev ⋅ Vinodkumar Prabhakaran ⋅ Alice Oh

As large language models (LLMs) are increasingly deployed to users around the world, they are integrated into everyday tasks across diverse cultural contexts, from drafting personal communications to brainstorming creative ideas. These tasks are inherently cultural: they require contextual appropriateness, symbolic resonance, and tacit cultural expectations that native speakers draw on instinctively, meaning that a response can be factually plausible yet unmistakably wrong to a local reader. Existing cultural benchmarks have treated culture as a flat set of facts via fact verification or norm entailment methods, and have adopted LLM-as-a-Judge without examining whether they can capture such thick cultural errors. To address this gap, we present JuICE (Benchmark for LLM-Judge in Identifying Cultural Errors), a multilingual dataset of 7,470 span-level annotations of cultural and linguistic errors, collected from native speakers in long-form LLM responses. It covers 1,050 query-response pairs from four countries (United States, South Korea, Indonesia, and Bangladesh), in both English and their countries' main languages. Using JuICE, we find that even the strongest LLM-judge achieves only an F1 of 0.52 against human-annotated errors, and that this gap is systematic. LLM-judges reliably detect surface-level linguistic and factual mistakes but consistently miss thick cultural errors that local residents readily identify. Our findings suggest that robust cultural evaluation must move beyond surface-level detection toward frameworks that account for the depth and situatedness of cultural meaning.

Modern Bayesian optimization and adaptive sampling methods increasingly rely on nonlinear parametric models, yet theoretical guarantees for such models under adaptive data collection remain limited. Existing analyses largely focus on Gaussian processes, kernel machines, linear models, or linearized neural approximations, leaving a gap between theory and the nonlinear models used in practice. We develop a kernel-based framework for analyzing regularized nonlinear parametric models trained on adaptively collected data. Our approach uses kernels over the parameter space to induce reproducing-kernel Hilbert space structures over the corresponding model class, yielding confidence bounds for models trained with broad classes of regularized convex losses. We show how these bounds can support convergence guarantees for nonlinear acquisition and surrogate models, including randomized regularized policies that select points by maximizing a trained random model. These results provide a unified route to analyzing nonlinear parametric models in Bayesian optimization and related adaptive optimization settings.


Kernel Token Contradiction: a Fast and Principled Approach for LLM Claim Uncertainty Quantification

Jérémie Dentan ⋅ Alexi Canesse ⋅ Mahammed El Sharkawy ⋅ Sonia Vanier

Claim-level Uncertainty Quantification (UQ) aims to mitigate the lack of reliability of Large Language Models (LLMs) by evaluating the factuality of each claim in their outputs. We introduce Kernel Token Contradiction (KTC), a lightweight approach to compute claim-level UQ under realistic white-box conditions. KTC represents the candidate tokens involved in LLM generation as a positive semi-definite kernel that integrates both the LLM’s conditional distribution and a token contradiction score. We then use the Von Neumann entropy to quantify the uncertainty of this kernel. To estimate token contradiction, we develop a new approach based on frequency statistics from the Wikipedia corpus. Although CPU-only, our approach achieves over an 8.2× speedup compared to state-of-the-art GPU-accelerated methods based on cross-encoders, and over a 65× speedup compared to CPU-only methods with comparable performance. Our evaluation spans two benchmarks across four European languages and 16 different models. KTC not only matches the average performance of existing methods but also outperforms them in high-precision regimes. This combination of computational efficiency and accuracy makes real-time monitoring of LLM outputs practical in production.


Knowing What is Missing: Efficient Conversational Memory via Explicit Evidence-Gap Tracking

Xingbo Du ⋅ Loka Li ⋅ Duzhen Zhang ⋅ Leonard Song ⋅ Le Song

Long-term conversational memory poses a fundamental system challenge for LLM agents: reasoning over all past interactions improves coverage but incurs prohibitive token costs and noise, while retrieval-based methods are efficient yet often fail on multi-hop and temporal questions. Existing multi-step retrievers partially address this gap, but typically operate over an ever-growing textual history, causing context expansion and noise accumulation across iterations. We propose MemR$^3$, a backend-agnostic closed-loop controller that reformulates conversational memory retrieval as explicit stateful decision-making. At each iteration, MemR$^3$ maintains an evidence-gap state consisting of grounded information (evidence) and unresolved information requirements (gaps), and uses this state to route among three actions: retrieve, reflect, and answer. Newly retrieved snippets are incorporated into the state and then masked from subsequent prompts, so the model conditions on a compact summary of progress rather than the full retrieval history. This design turns multi-step memory retrieval into a bounded, inspectable control process whose query-time context scales with the state instead of raw accumulated text. Experiments on LongMemEval$_s$ show that MemR$^3$ surpasses implicit multi-step retrieval baselines and often approaches or exceeds full-context prompting while using only 1%-5% of the full-context tokens on long conversations. On LoCoMo, MemR$^3$ consistently improves over its underlying RAG- and Zep-based memory backbones. These results suggest that explicit evidence-gap tracking is an effective abstraction for token-efficient conversational memory.


Knowing You before You Speak: User State Modeling for LLM-Based Personalized Dialogue

jiani luo ⋅ Xiaoyan Zhao ⋅ Yang Zhang ⋅ Shuyi Miao ⋅ Bingbing Xu ⋅ Stefan Konigorski ⋅ Tat-Seng Chua

Personalized dialogue requires more than recalling explicit user histories: systems also need to infer hidden user states that evolve through interaction and shape appropriate response strategies. Existing memory- and profile-based methods primarily reuse observable user information, offering limited support for modeling user-state dynamics or selecting actions based on how they shape future user states. We propose PUMA (Prospective User-state Modeling for Action selection), a framework grounded in the Free Energy Principle (FEP) that formulates personalization as decision-making under partial observability, centered on an explicit user state model that captures latent user states and their action-conditioned dynamics. At each turn, PUMA maintains a belief over the user's hidden state, refines the user state model for observation generation and action-conditioned state transition, and selects dialogue actions by minimizing expected free energy—balancing epistemic and pragmatic objectives under a unified criterion. This formulation shifts personalization from passive memory retrieval to model-based decision-making over user evolution. We instantiate PUMA on healthcare-oriented counseling and motivational interviewing benchmarks with latent state annotations for rigorous evaluation. Experiments show that PUMA improves long-horizon dialogue outcomes while maintaining strong response quality, and a cross-dataset study demonstrates more reliable user-state estimation and next-state prediction. Our code is available at: https://anonymous.4open.science/r/PUMA-4DA7/.


Knowledge-Graph Paths as Intermediate Supervision for Self-Evolving Search Agents

Huyu Wu ⋅ Jun Liu ⋅ Xiaochi Wei ⋅ Yan Gao ⋅ YIWU ⋅ Yao Hu

Self-evolving search agents reduce reliance on human-written training questions by generating and solving their own search tasks. We build on Search Self-Play (SSP), a representative Proposer and Solver framework in which questions are generated and answered via multi-step search and reasoning. In practice, however, SSP faces two bottlenecks: the Proposer constructs questions from isolated answer entities without relational context, yielding many invalid or unverifiable questions in early self-play training, while the Solver receives only a binary outcome reward that discards useful signal from partially on-track search trajectories. We address both bottlenecks by reusing knowledge-graph paths as construction-derived intermediate supervision for both question construction and reward shaping. First, we ground question construction in LLM-guided knowledge-graph subgraphs, providing relational context for the Proposer. Second, we observe that constructing and solving a multi-hop question can involve overlapping intermediate entities: the factual bridges used to formulate the question may provide approximate waypoints for answering it. Exploiting this overlap, we introduce Waypoint Coverage Reward (WCR), which grants graded partial credit to incorrect Solver trajectories according to their coverage of entities on the construction path, while preserving full reward for correct answers. Across seven QA benchmarks and nine model configurations, our approach improves the average score over standard SSP in all configurations, including notable gains on multi-hop QA tasks. These results suggest that knowledge-graph paths can be reused as lightweight intermediate supervision, providing both relational guidance and process feedback without additional task-specific human annotations or manually labeled process steps.


Knowledge-Level Consistency Reinforcement Learning: Dual-Fact Alignment for Long-Form Factuality

Junliang Li ⋅ Yucheng Wang ⋅ Yan Chen ⋅ Yu Ran ⋅ Ruiqing Zhang ⋅ Jing Liu ⋅ Hua Wu ⋅ Haifeng Wang

Hallucination in large language models (LLMs) during long-form generation remains difficult to address under existing reinforcement learning from human feedback (RLHF) frameworks, as their preference rewards often overlook the model's own knowledge boundaries. In this paper, we propose the $\textbf{K}$nowledge-$\textbf{L}$evel $\textbf{C}$onsistency Reinforcement Learning $\textbf{F}$ramework ($\textbf{KLCF}$), which re-examines this problem from a distribution alignment perspective. KLCF formalizes long-form factuality as a bidirectional distribution matching objective between the policy model's expressed knowledge distribution and the base model's parametric knowledge distribution: under the constraint that generation must not exceed the support set of the base knowledge, the objective maximizes coverage of high-probability facts, thereby jointly optimizing precision and recall. To achieve this, we design a Dual-Fact Alignment mechanism that approximates the recall term using a factual checklist constructed by sampling from the base model, and constrains hallucinations with a lightweight truthfulness reward model. Both components are jointly optimized and require no external retrieval throughout training. Experimental results demonstrate that KLCF consistently improves factuality metrics across multiple long-form benchmarks and model scales, effectively alleviating hallucination and over-conservatism while maintaining efficiency and scalability.


Koopman Generative Operators for Efficient Probabilistic Time-Series Forecasting

Raz Marshanski ⋅ Liran Nochumsohn ⋅ Mayank Jauhari Iitr ⋅ Boris Oreshkin ⋅ Omri Azencot

Probabilistic time-series forecasting requires models that simultaneously capture structured temporal dynamics, expressive uncertainty, and efficient inference, yet existing approaches fall short of this goal: latent dynamical models impose structure but rely on restrictive generation mechanisms, while modern generative methods such as diffusion and flow matching achieve flexibility at the cost of iterative and computationally expensive sampling. We introduce the Koopman Generative Operator (KGO), a probabilistic forecasting method that conceptualizes prediction as the evolution of structured uncertainty. KGO integrates three core components: (1) Koopman Patch Embedding (KoPE) for temporally consistent latent trajectory extrapolation; (2) Koopman Flow Matching (KoFM), which enables fast, single-step generation through a closed-form matrix exponential in a Koopman latent space, bypassing iterative sampling; and (3) an Adaptive Uncertainty Gate (AUG), which provides calibrated predictions by adapting uncertainty per-variable and per-horizon, corresponding to a learned decomposition of aleatoric uncertainty without manual tuning. By unifying these components, KGO achieves state-of-the-art accuracy across the majority of ProbTS datasets, outperforming existing methods on 12/17 benchmarks in CRPS and 11/17 in NMAE. Further, by eliminating iterative sampling, KGO delivers at least a 25$\times$ reduction in inference time compared to iterative generative models. These results establish KGO as a practical, principled, and highly scalable framework for next-generation probabilistic forecasting.


KVFocus: A Perturbation-Theoretic Token-Risk Score for Selective KV Cache Reuse in RAG

Shizhuo Zhang ⋅ Nuowen Kan ⋅ Chenglin Li ⋅ Rui-Xiao Zhang ⋅ Yifan Zhao ⋅ Yuanpeng He ⋅ Huixin Zhang ⋅ Fan He ⋅ Hang Xu ⋅ Hao Zhang ⋅ Wenrui Dai ⋅ Junni Zou ⋅ Hongkai Xiong

In retrieval-augmented generation (RAG) scenarios, the prefill delay of the large language model (LLM) long inputs is the dominant component of time-to-first-token (TTFT). Reusing KV caches across different LLM inputs reduces this cost, but the resulting cross-context mismatch on reused segments has a significant degradation on answer quality via propagating through prefill attention into decode-time logits. Existing selective-recomputation methods mitigate this issue by recomputing tokens chosen by empirical importance heuristics, which overlook the propagation of cache mismatches into answer-side errors, thereby failing to effectively strike the quality-TTFT tradeoff. To address this issue, we propose a perturbation-theoretic KV cache reuse framework, KVFocus, which selects reused KV caches through a perturbation-theoretic token-risk score during the prefill process. Specifically, we first develop a first-order analysis that traces reuse error in three stages: source mismatch on the reused segment, suffix contamination during prefill, and decode-time logit stability. This analysis reveals that the leading-order answer-side error decomposes into a product of source-side V-drift and downstream suffix-to-segment attention concentration. Guided by this multiplicative structure, we define a token-risk score that operationalizes the bound during the inference. A plug-and-play selector for KV cache reuse in existing RAG-based LLM serving is then designed by recomputing the top-$r\%$ tokens directly suppresses the predicted leading-order answer-side error, covering both sides of the pathway that prior one-sided heuristics miss. Empirical results across three 7-8B instruction-tuned LLMs (Mistral, Qwen2.5, Qwen3) and four multi-hop QA benchmarks demonstrate that the proposed KVFocus achieves a superior trade-off between the ansewer quality and the TTFT in comparison to state-of-the-art baselines, with up to $3\times$ TTFT speedup over full prefill at matched answer quality.


kVNN: Learnable Volterra Network Kernels

Haoyu Yun ⋅ Hamid Krim ⋅ Yufang Bao

Higher-order learning exploits compositional features whose predictive value arises from nonlinear interactions beyond pairwise relations among variables, scales, or modalities. To address the complexity of modeling higher-order interactions in modern large-scale deep learning models, a kernelized Volterra Neural Network (kVNN) is proposed in this paper. Specifically, the proposed learnable multi-kernel representation models different interaction orders using distinct polynomial-kernel components with compact learnable centers, thereby yielding an order-adaptive parameterization. The resulting kVNN layers are composed of parallel order-specific branches and can readily replace standard convolutional kernels in existing learning architectures. Theoretical results are supported by experiments on video action recognition, image denoising, and image classification. The results show marked performance-efficiency trade-offs: kVNN consistently reduces model (parameters) and computational (GFLOPs) complexity while achieving competitive and often improved performance, even when trained from scratch without large-scale pretraining. In summary, structured kernelized higher-order layers offer a practical path to balancing expressivity and computational cost in modern deep networks.


L$^2$EAP: Supercharging LLMs for Formal Mathematics with Agentic Frameworks

Po-Nien Kung ⋅ Linfeng Song ⋅ Dawsen Hwang ⋅ Jinsung Yoon ⋅ Chun-Liang Li ⋅ Simone Severini ⋅ Miroslav Olšák ⋅ Edward Lockhart ⋅ Quoc V Le ⋅ Burak Gokturk ⋅ Thang Luong ⋅ Tomas Pfister ⋅ Nanyun Peng

Large Language Models (LLMs) exhibit strong informal mathematical reasoning but struggle to generate mechanically verifiable proofs in formal languages like Lean. We present L$^2$EAP (LLM-in-Lean Environment Agentic Prover), an agentic framework that enables general-purpose foundation models to achieve state-of-the-art performance on automated formal theorem proving. L$^2$EAP leverages foundation model capabilities, such as informal reasoning, instruction following, and iterative self-refinement. By decomposing complex problems into smaller units, the system bridges formal proof construction with informal blueprints through continuous interaction with the Lean compiler. To provide a rigorous evaluation beyond increasingly saturated benchmarks, we introduce Lean-IMO-Bench, a benchmark of IMO-style problems formalized in Lean, with short statements yet highly non-routine and multi-step proofs across a wide range of difficulty levels. Empirically, on the latest 2025 Putnam Competition, an annual mathematics competition for undergraduate students in North America, L$^2$EAP solves all 12 problems, matching recent breakthroughs by frontier formal mathematical models; on Lean-IMO-Bench, L$^2$EAP raises the one-shot formal solve rate of general-purpose LLMs from below 10\% to 70\%. It also compares favorably to the best baseline performance of 48\% attained by a specialized system with dedicated ATP components that achieved gold-medal-level performance at the 2025 IMO.


L2P: Unlocking Latent Potential for Pixel Generation

Zhennan Chen ⋅ Junwei Zhu ⋅ Xu Chen ⋅ Jiangning Zhang ⋅ Jiawei Chen ⋅ Zhuoqi Zeng ⋅ Wei Zhang ⋅ Chengjie Wang ⋅ Jian Yang ⋅ Ying Tai

Pixel diffusion models have recently regained attention for visual generation. However, training advanced pixel-space models from scratch demands prohibitive computational and data resources. To address this, we propose the \textbf{Latent-to-Pixel (L2P)} transfer paradigm, an efficient framework that directly harnesses the rich knowledge of pre-trained LDMs to build powerful pixel-space models. Specifically, L2P discards the VAE in favor of large-patch tokenization and freezes the source LDM's intermediate layers, exclusively training shallow layers to learn the latent-to-pixel transformation. By utilizing LDM-generated synthetic images as the sole training corpus, L2P fits an already smooth data manifold, enabling rapid convergence with zero real-data collection. This strategy allows L2P to seamlessly migrate massive latent priors to the pixel space using only 8 GPUs. Furthermore, eliminating the VAE memory bottleneck unlocks native 4K ultra-high resolution generation. Extensive experiments across mainstream LDM architectures show that L2P incurs negligible training overhead, yet performs on par with the source LDM on DPG-Bench and reaches 93\% performance on GenEval.


LAMP: Look-Ahead Mixed-Precision Inference of Large Language Models

Stanislav Budzinskiy ⋅ Marián Gloser ⋅ Tolunay Yilmaz ⋅ Ying H Tham ⋅ Yuanyi Lin ⋅ Wenyi Fang ⋅ FAN WU ⋅ Philipp Petersen

Mixed-precision computations are a hallmark of the current stage of AI, driving the progress in large language models towards efficient, locally deployable solutions. This article addresses the floating-point computation of compositionally-rich functions, concentrating on transformer inference. Based on the rounding error analysis of a composition $f(g(x))$, we provide an adaptive strategy that selects a small subset of components of $g(x)$ to be computed more accurately while all other computations can be carried out with lower accuracy. We then explain how this strategy can be applied to different compositions within a transformer and illustrate its overall effect on transformer inference. We study the effectiveness of this algorithm numerically on GPT-2 models and demonstrate that already very low recomputation rates allow for improvements of up to two orders of magnitude in accuracy.


LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving

Yuechen Luo ⋅ Fang Li ⋅ Shaoqing Xu ⋅ Yang Ji ⋅ Zehan Zhang ⋅ Bing Wang ⋅ Shen Yuannan ⋅ Jianwei Cui ⋅ Long Chen ⋅ Guang Chen ⋅ Hangjun Ye ⋅ Zhi-Xin Yang ⋅ Fuxi Wen

While Vision-Language-Action (VLA) models have revolutionized autonomous driving by unifying perception and planning, their reliance on explicit textual Chain-of-Thought (CoT) leads to semantic-perceptual decoupling and perceptual-symbolic conflicts. Recent shifts toward latent reasoning attempt to bypass these bottlenecks by thinking in continuous hidden space. However, without explicit intermediate constraints, standard latent CoT often operates as a physics-agnostic representation. To address this, we propose the Latent Spatio-Temporal VLA (LaST-VLA), a framework shifting the reasoning paradigm from discrete symbolic processing into a physically grounded Latent Spatio-Temporal CoT. By implementing a dual-feature alignment mechanism, we distill geometric constraints from 3D foundation models and dynamic foresight from world models directly into the latent space. Coupled with a progressive SFT training strategy that transitions from feature alignment to trajectory generation, and further refined via Reinforcement Learning with Group Relative Policy Optimization (GRPO) for safety and rule compliance, LaST-VLA achieves state-of-the-art performance across multiple autonomous driving benchmarks, including NAVSIM v1, NAVSIM v2, Bench2Drive, SURDS, and NuDynamics.


Latent Process Generator Matching

Lukas Billera ⋅ Hedwig Nora Nordlinder ⋅ Ben Murrell

Many recent flow-matching and diffusion-style generative models rely on auxiliary stochastic dynamics during training: a richer process is simulated to define conditional targets, but the auxiliary state is either intractable to sample at generation time or simply not part of the desired output. Existing Generator Matching theory formalises conditioning on static latent random variables, and several recent papers prove special cases of projection results for particular augmented-state constructions. We introduce latent process generator matching, a general framework that treats the observed generative state as a deterministic image $X_t=\Phi(Y_t)$ of a tractable Markov process $Y_t$. We show that in this setting one may learn the generator of a stochastic process on the image space which has the same one-time marginal distributions as the projected process. This generalizes and subsumes the discrete latent process results from the literature, and extends Generator Matching from static latent variables to a rich family of time-dependent latent conditional processes.


Layer Precision Reduction for Deep Anomaly Detection

jiawei Yang ⋅ Xu Tan ⋅ Matti Kaisti

Autoencoders (AEs) are widely used in unsupervised anomaly detection (AD) due to their ability to model complex data distributions. However, their strong reconstruction capacity, while effective for capturing normal patterns, often enables them to also reconstruct anomalous objects, including both individual and collective anomalies, which reduces detection performance. This limitation is particularly severe for collective anomalies, which are frequently misinterpreted as normal clusters. To address this issue, we propose Layer Precision Reduction (LPR), a novel technique that systematically constrains the numerical precision of weights and biases in neural network layers to strategically reduce the learning capacity of neural networks, such as the reconstruction capacity of AEs. LPR encourages the network to prioritize dominant manifold structures while suppressing the reconstruction of rare or anomalous patterns. While LPR was primarily developed for AE-based detectors (the focus of this paper), it is applicable to any neural network. Simulations are conducted with 24 AD models across 30 real-world public datasets. LPR is applied to 10 deep AD models, including 5 unsupervised AE-based methods, 2 unsupervised non-AE-based methods, 1 semi-supervised method, 1 weakly supervised method, and 1 fully supervised method. LPR consistently improves all 10 models, most notably enhancing the best-performing unsupervised AE-based detector from an average AUROC of 72.12\% to 83.06\%, surpassing the strongest unsupervised non-AE-based competitor, which achieves an average AUROC of 77.36\%, across the remaining 14 detectors. These results reveal a previously unexplored link between layer-level numerical precision and manifold pattern prioritization, demonstrating that LPR serves as a simple yet powerful mechanism for enhancing AD performance while offering broad applicability to diverse neural architectures and potential extensions to tasks beyond AD.


Layout Before Pixels: Topology-Anchored Transcriptome-to-Histology Generation

Jianwei Zhao ⋅ Xin Li ⋅ Fan Yang ⋅ Qiang Zhai ⋅ Ao Luo ⋅ Hong Cheng

Transcriptome-conditioned whole-slide image (WSI) synthesis aims to decode the complex mapping from molecular states to histologic phenotypes. Existing generative paradigms predominantly rely on a direct black-box transcriptome-to-image mapping, tasking a single generator with the simultaneous interpretation of transcriptome signals, inference of multicellular spatial organization, and rendering of fine-grained textures. This formulation, however, is structurally under-constrained; the underlying tissue architecture, the critical substrate linking gene expression to morphology, remains latent and only weakly supervised by pixel-level objectives. We present \textsc{TopoScape}, a topology-anchored framework for transcriptome-to-histology generation under a \emph{Layout-Before-Pixels} principle. Instead of synthesizing pixels directly from transcriptomic embeddings, \textsc{TopoScape} first resolves molecular states into explicit, multi-class cellular topologies via an RNA-guided Topological Prior Generator, reinforced by persistence-based topological regularization. These resolved topologies are further reparameterized into a suite of topology-derived controls, including structure-aware initialization, continuous density fields, and distance representations, which guide a hierarchical, frequency-decoupled flow-matching trajectory. By introducing multicellular organization as an explicit structural mediator, \textsc{TopoScape} allows molecular states to deterministically shape both the macroscopic tissue architecture and the microscopic generative process. Across five TCGA benchmarks, \textsc{TopoScape} consistently outperforms state-of-the-art methods in generative fidelity, biological realism, and downstream predictive utility. Code and pretrained models will be released.


LDPCache: Locally Differentially Private Multi-Query Processing with Cache Optimization for Large Language Models

Haoqiang Shi ⋅ Ning Wang ⋅ Chuan He ⋅ Qian Ma ⋅ Zhigang Wang ⋅ Shen Su ⋅ Yu Gu ⋅ Zhihong Tian

The Model-as-a-Service (MaaS) paradigm enables resource-constrained edge users to access cloud-based large language model (LLM) services but raises significant privacy concerns. Existing approaches perturb user queries under local differential privacy (LDP) before sending them to LLMs. However, naively applying these methods to multi-query scenarios leads to linear growth in privacy budget consumption. To address this, we propose LDPCache, the first cache-enhanced LDP framework for multi-query LLM processing. By leveraging semantic correlations across consecutive user queries, LDPCache maintains a cache of historical responses and intelligently decides whether to reuse a cached result or query the cloud-based LLM, thereby reducing privacy budget consumption. The core of LDPCache is an LDP-compliant scheduler that evaluates cache reusability based on both semantic similarity and the expected accuracy gains from LLM access. This scheduler enables hit decisions without actual LLM access while ensuring that the scheduling process itself satisfies LDP guarantees. To accommodate the limited memory of edge devices, we further design a multi-factor cache eviction strategy that balances query diversity, response accuracy, and recency to improve the hit ratio. Extensive experiments on real-world datasets demonstrate that LDPCache significantly improves response accuracy over baseline methods.


LeAct: Learning to Reason from Expert Actions

Ziran Yang ⋅ Chengshuai Shi ⋅ Raj Ghugare ⋅ Benjamin Eysenbach ⋅ Karthik Narasimhan ⋅ Chi Jin

Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely untapped source of supervision lies in expert systems (e.g., game engines, classical planners, theorem provers), which routinely produce near-optimal actions across diverse domains. But these experts are silent: they commit to an action without writing down the chain of thought (CoT) behind it. Recovering that CoT as natural-language reasoning would distill expert knowledge into a student that generalizes beyond the demonstrated actions. We treat it as a latent variable and study how to recover it from the action alone. Our approach, LeAct (**Le**arning to reason from **Act**ions), optimizes this latent variable: the student samples candidate CoTs for each expert action, and we retain those that measurably improve its own probability of recovering the action. Across imperfect-information games at multiple scales and a simulated robotics benchmark, LeAct reaches the solver's numerical floor on small enumerable games. At larger scale, it is $5\times$ closer to the solver than the strongest expert-iteration baseline. At Flop Hold'em ($\sim 10^9$ infosets), LeAct wins head-to-head by $+60$ mbb/g, and on the robotics probe it is the only training recipe that improves on direct imitation. We present a principled framework and the result: expert systems become a categorically new source of reasoning teachers for foundation models.


LeanSearch v2: Global Premise Retrieval for Lean 4 Theorem Proving

Guoxiong Gao ⋅ Zeming Sun ⋅ Jiedong Jiang ⋅ Yutong Wang ⋅ jd xu ⋅ Peihao Wu ⋅ Bryan Dai ⋅ Bin Dong

Proving theorems in Lean 4 often requires identifying a scattered set of library lemmas whose joint use enables a concise proof---a task we call global premise retrieval. Existing tools address adjacent problems: semantic search engines find individual declarations matching a query, while premise-selection systems predict useful lemmas one tactic step at a time. Neither recovers the full premise set an entire theorem requires. We present LeanSearch v2, a two-mode retrieval system for this task. Its standard mode applies a hierarchy-informalized Mathlib corpus with an embedding--reranker pipeline, achieving state-of-the-art single-query retrieval without domain-specific fine-tuning (nDCG@10 of 0.62 vs. 0.53 for the next-best system). Its reasoning mode builds on standard mode as its retrieval substrate, targeting global premise retrieval through iterative sketch-retrieve-reflect cycles. On a 69-query benchmark of research-level Mathlib theorems, reasoning mode recovers 46.1% of ground-truth premise groups within 10 retrieved candidates, outperforming strong reasoning retrieval systems (38.0%) and premise-selection baselines (9.3%) on the same benchmark. In a controlled downstream evaluation with a fixed prover loop, replacing alternative retrievers with LeanSearch v2 yields the highest proof success (20% vs. 16% for the next-best system and 4% without retrieval), confirming that retrieval quality propagates to proof generation. All code, data, and benchmarks will be open-sourced.


Learn from the Gap: Differential-Aware Advantage Pruning with Adaptive Rollout Sampling for GRPO

Jiahua Yang ⋅ Chuangchuang Wang ⋅ Zhiwei Yang ⋅ Xianpeng Zhang ⋅ Dongyu Chen ⋅ Xing Chen ⋅ Tianhuang Su ⋅ Haonan Lu ⋅ Quanlong Guan ⋅ Kai Tang

Recently, Group Relative Policy Optimization (GRPO) and its variants have been developed for policy optimization and demonstrated notable performance gains. However, these methods usually incur substantial computational overhead due to per-question multi-rollout sampling and repeated per-token probability evaluation across rollouts. Furthermore, low-information or highly homogeneous trajectories can degrade downstream learning signal efficiency, hindering model optimization and limiting final performance. To address these issues, we propose FastRL, a novel plug-and-play reinforcement learning framework that simultaneously improves training efficiency and the effectiveness of policy learning. Specifically, 1) We introduce an advantage-aware pruning strategy to selectively preserve high-advantage trajectories while maximizing inter-trajectory gradient diversity. 2) Then, we design an adaptive rollout sampling mechanism to dynamically adjust the sampling scale across different training stages based on historical pruning distributions, balancing exploration adequacy and computational efficiency. Experiments demonstrate that FastRL can be seamlessly integrated into GRPO, DAPO, and GSPO variants, achieving an average 2.07x training speedup on Geometry3K and GeoQA8K-R1V, along with an approximately 1.64% improvement in average accuracy on visual reasoning benchmarks. Source codes will be available at https://anonymous.4open.science/r/FastRL0.


Learn from Your Mistakes: Self-Correcting Masked Diffusion Models

Yair Schiff ⋅ Omer Belhasin ⋅ Roy Uziel ⋅ Guanghan Wang ⋅ Marianne Arriola ⋅ Gilad Turok ⋅ Ran Zilberstein ⋅ Miki Elad ⋅ Volodymyr Kuleshov

Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models, enabling parallel token generation while achieving competitive performance. Despite these advantages, MDMs face a fundamental limitation: once tokens are unmasked, they remain fixed, leading to error accumulation and ultimately degrading sample quality. We address this by proposing a framework that trains a model to perform both unmasking and correction. By reusing outputs from the MDM denoising network as inputs for corrector training, we train a model to recover from potential mistakes. During generation we apply additional corrective refinement steps between unmasking ones in order to change decoded tokens and improve outputs. We name our training and sampling method Progressive Self-Correction (ProSeCo) for its unique ability to iteratively refine an entire sequence, including already generated tokens. We conduct extensive experimental validation across multiple conditional and unconditional tasks, demonstrating that \method~yields better quality-efficiency trade-offs (up to ~4x faster sampling) and enables inference-time compute scaling to further increase sample quality beyond standard MDMs (up to ~1.2x improvement on benchmarks).


Learning a Task-Adaptive Low-Dimensional Semantic Space for Improved Visual Classification

Nikalal Helessage ⋅ Wanqing Li ⋅ Chris Bunn ⋅ Philip O Ogunbona

Semantic visual classification commonly aligns visual embeddings with semantic spaces induced by pretrained text encoders. However, such spaces are typically high-dimensional, weakly aligned with visual feature geometry, and insufficiently discriminative for downstream classification tasks. We propose TaSS, a framework for learning a task-adaptive low-dimensional semantic space that preserves semantic structure while explicitly optimizing class separability. TaSS is constructed through a parameterized geometry transformation of pretrained text embeddings that enlarges inter-class margins via distance reshaping and improves intra-class compactness through prototype-preserving semantic constraints. The resulting semantic space maintains meaningful semantic relationships while producing a discriminative geometry tailored for classification. Frozen visual encoder features are subsequently aligned to TaSS using a lightweight MLP, enabling improved discriminative learning without fine-tuning pretrained backbones. We further provide theoretical analysis showing that TaSS preserves semantic consistency while enforcing enhanced inter-class separation and controlled intra-class variance. Extensive experiments on skeleton-based human action recognition, video classification, and image classification benchmarks demonstrate consistent improvements over existing semantic alignment approaches, with gains exceeding 3 percentage points on average and up to 8.5 points on challenging datasets.


Learning a Trajectory-Geometric Condition from Reasoning for VLA Planning

Yuguang Yang ⋅ Zhewen Tan ⋅ Canyu Chen ⋅ Cheng Chi ⋅ Chunyang Liu ⋅ Kehua Sheng ⋅ Bo Zhang ⋅ Jinyu Yang ⋅ Linlin Yang ⋅ Baochang Zhang ⋅ Yan Wang ⋅ Xianbin Cao

End-to-end driving with vision-language models benefits from multimodal pretraining, but still faces a mismatch between reasoning and trajectory generation. For autoregressive VLA planners, continuous trajectories represented as step-wise normalized text waypoints are strong final outputs because they fit the native token-prediction interface of modern VLMs. However, they remain weak intermediate interfaces for reasoning-conditioned planning. Direct action tokenization provides a more explicit motion interface, but its effectiveness depends heavily on how the action codebook is constructed. We propose *TrajCond-VLA*, which learns a *trajectory-geometric condition* from reasoning for VLA planning. Our key idea is to represent this condition as a differential action trajectory whose $x$/$y$/yaw differentials are discretized separately with a 3D-Brohan codebook. The resulting **DiffAction tokens** preserve local motion geometry while remaining compatible with autoregressive token prediction. Building on this representation, we introduce a **Two-Stage Alignment Training** framework: Stage 1 predicts TrajCond from reasoning, and Stage 2 regresses the final continuous trajectory conditioned on both TrajCond and reasoning. Experiments on NAVSIM, nuScenes, and NeuroNCAP show that the learned trajectory-geometric condition is more effective as an intermediate condition than as a direct output format, and that the resulting pipeline is validated across multiple datasets with consistent gains on the completed comparisons. The nuScenes and NAVSIM navtest discussions are kept aligned to the public benchmark protocols used by Curious-VLA, and the closed-loop reference discussion uses the correct NeuroNCAP protocol.

Probing frozen vision transformers typically uses permutation-invariant aggregation (GAP or [CLS]), treating patch tokens as an unstructured set. Content-dependent probes such as self-attention are useful accuracy controls, but they do not expose a fixed token schedule or fixed position weights for auditing. We introduce SSMProbe, an explicitly inspectable probe that replaces invariant pooling with a Sinkhorn-learned evidence route followed by a diagonal S4 decoder. The S4 decoder is a linear time-invariant (LTI) system whose final state has fixed, position-dependent coefficients, so the probe-induced routed sequence can be audited as a concrete object rather than inferred only from accuracy. Our central measurement is the geometry of routed evidence: which patch tokens are moved to influential positions by this diagnostic, whether those tokens form spatially organized regions or random-like dispersed sets, and how the fixed S4 kernel weights them. Across MAE, BEiT, DINOv2, and supervised ViT, this route geometry separates MAE's dispersed, nearly random-like routes from the more spatially organized routes of BEiT, ViT, and DINOv2, with DINOv2 retaining a distinct strong [CLS] profile. SSMProbe uses the mathematical transparency of state-space models to turn a frozen ViT readout into an auditable evidence-routing analysis.

A point source in a PDE does not occupy an entire grid. A weather station reports one value at one site; it does not observe the temperature between stations. A pickup event in a city occurs at a coordinate and a time; it is not born as an image. Yet many neural-operator pipelines begin by turning such sparse observations into dense surrogate fields before learning. This interpolation step is often treated as harmless preprocessing, but in sparse regimes it changes the object on which the operator is asked to act. We introduce the Dirac Neural Operator (DIRAC-NO), a measure-native neural operator for sparse event-to-field learning without interpolation. DIRAC-NO keeps observations as atomic event measures until query time and constructs the output field through adaptive Dirac-kernel aggregation. Because this aggregation is implemented by direct summation, the front end preserves source additivity at the aggregation stage and admits a learned Green-kernel view of source-to-field response. Across controlled point-source PDE benchmarks and real sparse reconstruction tasks, DIRAC-NO exposes the cost of interpolation-first learning. On MeasureBench-1D, it achieves relative $L^2$ error 0.015 at $n=128$, compared with 0.49 for FNO, and it better preserves superposition structure in trained models. In 2D, it is strongest under parameter and density shift, while ERA5 reconstruction shows that the representation advantage persists beyond analytic PDEs at high observation budgets. Boundary cases and negative-control event streams further show that the benefit appears where the representation argument predicts: sparse, event-native tasks with source-response structure. These results suggest that sparse operator learning should be organized not only by the neural backbone, but by the native space in which the input is represented.


Learning from Disagreement: Maximum Divergence Knowledge Distillation

Aref Jafari ⋅ Parsa Ashrafi Fashi ⋅ Mehdi Rezagholizadeh ⋅ Hanieh Asadi Golmankhaneh ⋅ Shayan Salehi ⋅ Vikram Appia ⋅ Emad Barsoum ⋅ Ali Ghodsi

Knowledge distillation pipelines typically train a student to match a teacher on a fixed corpus of teacher-labeled examples. This leaves a major source of supervision unused: the teacher is a fully queryable model whose output distribution extends far beyond any fixed dataset. We argue that effective distillation should exploit this structure by steering generation toward regions where the student most underestimates the teacher. We propose \textit{Maximum Divergence Knowledge Distillation} (MDKD), a rejection-sampling method for autoregressive language models. At each decoding step, MDKD draws candidate tokens from the teacher and preferentially accepts those the student underweights, constructing trajectories that remain teacher-supported while concentrating training on regions where the student assigns insufficient probability. We show that stochastic MDKD's acceptance rule has an exact distributional characterization: it samples precisely from the teacher's \textit{uncovered mass} $[T - S]_+$, the probability the teacher assigns to tokens the student does not yet cover, with stochastic acceptance probability equal to the teacher--student total variation distance. Standard teacher sampling, by contrast, spends most of its updates on the overlap $\min(T, S)$ already absorbed by the student. This makes MDKD the structural complement of speculative decoding, which exploits the same overlap to accelerate inference. On GSM8K, distilling \texttt{Qwen2.5-14B-Instruct} into a base \texttt{Qwen2.5-1.5B} student with only $1{,}000$ KD samples (5 epochs) lifts accuracy from $8.26\%$ to $64.90\%$, outperforming SOTA methods by $10.3$ percentage points and recovering $65\%$ of the teacher--student gap. A divergence-gain analysis confirms the mechanism: MDKD sequences raise student cross-entropy by $43.6\%$ while preserving $98\%$ of teacher-sampling accuracy. Across arithmetic reasoning, dialogue summarization, and code generation, MDKD matches or outperforms strong distillation baselines.


Learning Global Temporal Dynamics in Sparse Networks via Cycle Counts

Xinyuan Fan ⋅ Dong Huang ⋅ Tianpai Luo ⋅ Pengkun Yang ⋅ Weichi Wu

Statistical modeling and inference for temporal networks are increasingly important across modern applications, yet remain challenging in the sparse regime with evolving network sizes and temporal dependence. We propose a tractable model that captures these key features in a unified framework. Under this model, we establish geometric ergodicity and characterize the asymptotic behavior of cycle counts. These counts provide informative low-order summaries of temporal network dynamics. Moreover, we develop method-of-moments estimators based on cycle counts. We prove identifiability and uniqueness of the parameter estimates, and establish the strong consistency and asymptotic normality. To support uncertainty quantification, we further propose a simulation-based plug-in estimator of the asymptotic variance. Extensive simulations and a real data example demonstrate that our methods are accurate and effective for sparse, size-varying temporal networks.


Learning Hierarchical Forward Processes For Discrete Diffusion Language Models

Inhyeok Jeong ⋅ Jeongwhan Choi ⋅ Jaehyeon Park ⋅ Noseong Park

Discrete diffusion language models (DLMs) enable parallel generation but usually rely on fixed forward processes, such as masking, uniform noise, or precomputed hierarchies. We investigate whether the hierarchical forward process can instead be learned while preserving continuous-time tractability. We propose \textbf{Learnable Hierarchical Diffusion Language Model (LHDLM)}, which extends a learnable column-stochastic token-cluster map. This many-to-many hierarchy preserves a CTMC with block-conditional transition and closed-form CT-ELBO, while recovering HDLM in the one-hot limit and masked diffusion in the collapsed limit. We identify degenerate hierarchy maps caused by learning the same map used in both forward targets and model-induced cluster predictions, and mitigate them with an information-retention regularizer. Empirically, LHDLM is competitive with discrete diffusion approaches and reveals soft token structures --- cluster-agnostic function tokens and sharp content-token assignments --- that cannot be represented by hard partitions.


Learning Implicit Bias in Generative Spaces for Accelerating Protein Dynamics Emulation

Kaihui Cheng ⋅ Zhiqiang Cai ⋅ Wenkai Xiang ⋅ Zhihang Hu ⋅ Siyu Zhu ⋅ Tzuhsiung Yang ⋅ Yuan Qi

Generative emulators of protein dynamics produce plausible trajectories at a fraction of the cost of molecular dynamics, but they inherit their training distribution and tend to revisit known states rather than reach rare ones under long-horizon extrapolation. Inspired by classical enhanced sampling, we introduce an implicit, history-dependent bias in the generative space of a pretrained emulator. Specifically, a history-aware score estimator augments the frozen emulator with a distance-weighted bias that steers reverse-time sampling away from previously generated structures, regularized by an environment-support term. To preserve structural validity at long horizons, a score-based refinement step re-projects drifted samples onto the data manifold using the frozen emulator. Our experiments demonstrate that the method (i) raises diversity by $35\%$ on DynamicPDB-80; (ii) on $12$ zero-shot Fast-Folding proteins, the learned bias alone reaches the unbiased emulator's coverage up to ${\sim}15\times$ faster, and pairing it with refinement reaches the coverage up to ${\sim}37\times$ faster while covering ${\sim}3\times$ as many low-energy states.


Learning in Causal Markov Games

Aurghya Maiti ⋅ Elias Bareinboim

Markov Games are the standard formal model for multi-agent reinforcement learning, capturing agents that act in a shared state and optimize rewards over time. However, in real-world settings, agents’ decisions are often influenced by unobserved factors, such as cognitive biases, behavioral tendencies, or intuitive signals, that also affect rewards and future states. Ignoring these unobserved confounders can make standard learning methods converge to suboptimal policies. In such settings, optimal play may require policies that condition on counterfactual signals, requiring reasoning at the counterfactual layer of the Pearl Causal Hierarchy. In this paper, we introduce Causal Markov Games (CMGs), a framework for modeling sequential multi-agent decision making in the presence of unobserved confounding. We show that CMGs strictly generalize Markov Games, with arbitrarily large gaps between classical interventional equilibria and causal counterparts. We then develop two learning algorithms under different observability assumptions. The first, CNash-VI-FO, is a model-based learner with finite-sample guarantees when agents’ natural actions or intuitions are revealed post hoc. The second, CNash-VI-NO, is an explore-then-exploit procedure with asymptotic guarantees for the setting in which opponents’ natural actions are never observed. To address scalability, we further provide a drop-in counterfactual augmentation of deep MARL. Empirically, on a confounded windy variant of the Multi-Particle Environment and the Iterated Causal Prisoner’s Dilemma, counterfactual agents strictly dominate their non-causal counterparts.


Learning in Policy Transparency Games: Wedge Structure and Adaptive Certification

Nguyen T Uyen ⋅ Khanh N Quoc ⋅ Duc H Nguyen ⋅ Ngoc Mai Vu ⋅ Phan Quoc Hung Mai ⋅ Luong Doan ⋅ Trang Le Vu Quynh ⋅ Trang Pham ⋅ Nhung Duong ⋅ Tuan Do

As learning agents are increasingly deployed in strategic environments, designers must choose whether to deploy transparent (auditable, committing) or opaque (flexible, reactive) policies. We formalize this through \emph{Policy Transparency Games} (PTGs): a two-stage model in which each agent first chooses transparency, then plays an underlying normal-form game. Our central theoretical contribution is a structural reduction: for two-player PTGs, all transparency incentives decompose into two per-player quantities, a leader wedge $L_i-N_i$ and a follower wedge $F_i-N_i$, whose signs characterize every equilibrium and cleanly separate selection-robust from selection-dependent conclusions. We then ask how a learning agent can discover this structure from interaction data. External-regret learners over the binary transparency action reach coarse correlated equilibrium, and our adaptive WEDGE-CERTIFY algorithm attains instance-dependent certificate complexity $\tilde O(\sigma^2(d_0^{-2}+d_{10}^{-2}+d_{01}^{-2}))$ with matching information-theoretic lower bounds. The wedge representation thus determines the learnability of endogenous algorithmic transparency.


Learning Minimal Sufficient Evidence Graphs for GraphRAG via Nash-Guided Optimization

Hao Wu ⋅ Xinguo Yu ⋅ Wanqing Li ⋅ Hongru Sun ⋅ Xiao Luo ⋅ Wenbin Zhang ⋅ Sasha Nikolic ⋅ Shirui Pan ⋅ Yi Guo ⋅ Jie Yang

GraphRAG enhances Large Language Model reasoning by organizing retrieved knowledge into structured evidence graphs, enabling inference over connected evidence rather than isolated text fragments. Yet, existing GraphRAG methods often either miss query-critical links or introduce noisy or conflicting evidence that distracts reasoning. We propose CONSIST, a Conflict-aware Nash-optimized Minimal Sufficient graph learning algorithm that regulates evidence graphs to resolve conflicts during reasoning. Specifically, we cast graph learning as an exploration–editing process, where a candidate graph is expanded via structural and semantic connections and then refined by pruning unreliable edges. We further introduce Gumbel-based edge editing, framing it as a Nash-guided bargaining process, where current and future players negotiate over a customized utility to determine reliable edits that improve both immediate answer support and long-term graph quality. Experiments on multi-hop question answering and conflict-aware benchmarks demonstrate that CONSIST consistently outperforms recent state-of-the-art baselines, achieving average performance gains up to 7.0% EM and 6.3% F1. Moreover, CONSIST also produces substantially more compact evidence graphs, reducing retained nodes by 42.9% and edges by 54.7% on average.


Learning Multimodal One-step Flow Policy via Value-weighted Optimal Transport

Jaehun Shon ⋅ Jinha Choi ⋅ Jongwook Jeon ⋅ Jongmin Lee

Offline reinforcement learning aims to learn a policy solely from fixed datasets, which often contain multimodal action distributions. Flow policies can naturally represent such multimodal behaviors, but learning an efficient one-step flow policy remains challenging: standard value guidance often leads to mode collapse or exploits overestimation bias in out-of-distribution regions. To address this, we introduce One-step Flow policy via Optimal Transport (OptiFlow), a framework for one-step flow policy learning as a structured sample-allocation problem. OptiFlow jointly trains a value-aware reference flow policy and an efficient one-step policy, coupling their action samples through state-wise entropic optimal transport. For each state, critic-estimated values define the priority of distillation target actions, while the action-distance cost ensures geometrically compatible pairings. By avoiding direct critic maximization, our transport-guided approach enables in-distribution exploitation by anchoring the one-step policy to high-value, dataset-supported modes without the risk of out-of-distribution divergence. Experimental results demonstrate that OptiFlow effectively captures optimal multimodal behaviors and achieves strong performance across diverse offline RL benchmarks.


Learning Planning Budgets in Real-Time RL

Aneesh Muppidi ⋅ Firas Darwish ⋅ Dylan Cope ⋅ João Henriques ⋅ Jakob Foerster

Deliberating takes time. In real-time settings, that time is not free. Standard reinforcement learning (RL) sidesteps this as the environment waits indefinitely for the agent's decision. Instead, we study real-time RL environments where the environment progresses while waiting for the agent's action. Building on prior real-time formalizations, we introduce variable-delay real-time RL, where the agent chooses how long to deliberate at each decision point since the environment progresses. For the planning agents we use, the right delay is state-dependent, and naively planning how long to plan can paralyze the agent. We instead approach this setting by training a lightweight gating policy on top of a planner to select state-dependent planning budgets. Across real-time Pac-Man, Tetris, Snake, Speed Hex, and Speed Go, our gating policy outperforms fixed-budget and heuristic baselines, and transfers to a real-time setup where the environment and agent run on two different GPUs.


Learning Robust Representations for Defending White-Box Adversarial Attacks in Continual Learning

Yusen Chen ⋅ Fei Ye ⋅ Qihe Liu ⋅ Adrian G. Bors ⋅ shijie zhou

Continual learning under adversarial perturbations remains underexplored, despite its importance in security-critical deployments where attack patterns evolve over time. Existing continual adversarial defense methods mainly rely on replay or prediction-space regularization, but often fail to preserve robust representations across sequentially arriving attack types. In this paper, we study continual adversarial defense, where a model receives a sequence of tasks induced by different white-box attacks and must acquire robustness to new attacks without forgetting previously learned defenses. We address this challenging learning scenario by proposing the Learning Robust Representation Framework (LRRF), a representation-centric framework that improves robustness through hierarchical alignment at three complementary levels. First, we introduce the Dynamic Representation Matching (DRM) mechanism that aligns clean and adversarial feature distributions within each task to reduce the clean-robustness trade-off. Second, a Cross-Task Invariant Representation (CTIR) mechanism is proposed to regularize the current representation network against an accumulated network from previous tasks, encouraging attack-invariant features to persist over time. Third, a Knowledge Consolidation Optimization (KCO) mechanism is proposed to match class-consistent clean and adversarial features using replayed samples, which further stabilizes category structure across tasks. The empirical results show that the proposed approach achieves state-of-the-art performance.


Learning through Internalization

Nirmit Joshi ⋅ Marko Medvedev ⋅ Nikolaos Tsilivis ⋅ Julia Kempe ⋅ Nati Srebro

We study internalization, the process by which neural-network-based systems absorb an explicit computational procedure into their own weights, and how it facilitates learning. We investigate how transformers internalize the simulation of semiautomata by internalizing chain-of-thought (CoT) tokens, which classes of semiautomata are harder to internalize, and expose the flip side of internalization, that is, a progressive degradation of out-of-distribution performance. We then provide the first provable analysis of successful internalization: for the task of learning parities, we show that a simplified one-layer transformer provably first learns the target with explicit CoT supervision and then internalizes the autoregressive generation as CoT tokens are progressively removed, directly computing the parity. Learning such a representation directly from data without CoT supervision is computationally hard. Finally, we discuss how learning through internalization can be viewed as an instance of the \textit{Positive Distribution Shift} phenomenon recently introduced by~\citet{Med+26}.


Learning to Deaggregate: Large-scale Trajectory Generation with Spatial Priors

David Bergström ⋅ Mattias Tiger ⋅ Fredrik Heintz

Trajectories arise in many settings, including urban mobility, transportation, and maritime traffic. Existing generative models either offer no control over output distributions or rely on conditioning information specific to each individual trajectory, such as its origin and destination, distance, or departure time. This sample-specific conditioning limits controllability and ties the model to the environment where those statistics were observed. We propose to separate _where_ movement occurs from _how_ it unfolds: regional movement is summarized by a spatial prior $\Pi$, a marginal distribution of occupancy aggregated over many trajectories, and the generative model is trained so that samples drawn conditionally on $\Pi$ aggregate to match it. We call this property _marginal consistency_, which turns the spatial prior into a controllable input, enabling zero-shot generation in unseen regions, cities, and even unseen domains by supplying only the target region's prior. The Temporal Deaggregation Diffusion Model (TDDM) realizes this idea. We evaluate it across four datasets spanning three continents and three modalities: multi-modal human mobility (Geolife), urban taxi (Porto, Cabspotting), and maritime traffic (Brest). TDDM achieves improved fidelity and coverage over leading baselines and stable performance when transferred across regions, cities, and domains.


Learning to Sample From Diffusion Models via Inverse Reinforcement Learning

Constant Bourdrez ⋅ Alexandre Verine ⋅ Olivier Cappé

Diffusion models generate samples through an iterative denoising process guided by a pretrained neural network. Once the denoiser is fixed, the sampling algorithm itself (noise schedules, guidance scales, stochasticity profiles) still requires careful tuning, a process typically carried out through costly empirical grid search. In this work, we introduce an inverse reinforcement learning framework for learning sampling strategies without retraining the denoiser. We formulate the diffusion sampling procedure as a discrete-time finite-horizon Markov Decision Process, where actions correspond to optional modifications of the sampling dynamics. To optimize action scheduling, we avoid defining an explicit reward function and instead directly match the target behavior expected from the sampler using policy gradient techniques. We provide experimental evidence that this approach matches fine-tuned samplers and comes at a modest cost compared to grid search: on ImageNet-64, a single training run replaces exhaustive search at up to $9\times$ lower cost, with only 16\% overhead at inference.

Does language modeling on image-text data truly improve vision, or does it merely adapt visual features for language alignment? Despite rapid progress in multimodal large language models (MLLMs), this question remains poorly understood. We present the first systematic analysis of visual representations in MLLMs by probing popular model families across a broad suite of dense visual tasks and comparing each MLLM to its exact pre-MLLM vision encoder. We find that language modeling largely improves visual representations for semantic tasks but can degrade performance on fine-grained geometric tasks, such as monocular depth. We further show that scaling the language model consistently improves the quality of visual representations. Given the improved MLLM representations, we next examine how to best extract and aggregate these features. Since no single layer provides a universally strong representation and concatenating features across layers is impractical, we propose a lightweight adapter that efficiently combines features across MLLM layers, turning frozen MLLMs into competitive visual backbones. Finally, we study whether MLLM representations can improve downstream applications, including text-to-image generation and inverse dynamics modeling, and show that they can outperform pre-MLLM visual features as general-purpose representations. Together, our results provide a representation-centric perspective for understanding and leveraging MLLMs as visual foundation models.


Learning to Solve Generative ODEs Beyond the Linear Span

Sihyeon Kim ⋅ Seunghun Lee ⋅ Vikas Singh ⋅ Hyunwoo J. Kim

Diffusion and flow generative models sample by integrating a learned ODE, but high quality still requires many sequential model evaluations. Solver learning reduces this cost by adapting scalar coefficients, timesteps, or both, while keeping the backbone model fixed. In this work, we identify a structural bottleneck in this update family: each step remains span-limited. Each update remains a scalar-weighted combination of buffered velocity evaluations, so learning can reduce the in-span teacher mismatch but cannot represent the out-of-span residual exposed by large steps. We propose SpanLift, a lightweight operator-augmented neural solver that enlarges the update family beyond scalar-coefficient solvers. SpanLift keeps a fixed base solver as an in-span prior and learns a spatial residual operator over the state and velocity buffer. The operator is trained by endpoint teacher matching, preserves the pretrained backbone, and adds no model NFEs. Empirically, the learned correction transfers across base solvers and is predominantly out-of-span. Across pixel-space diffusion, latent flow matching, and precipitation nowcasting, SpanLift achieves state-of-the-art few-step sampling. With only 3 NFE, it improves CIFAR-10 FID from 8.16 to 5.69 and ImageNet FID from 17.37 to 11.83.


Learning under Localized Minority Imbalance

Amin Hosseininasab ⋅ Steven M Shugan

Class-imbalance methods typically assume that observed minority instances are an unbiased sample of their class. However, in many real-world settings, minority observability depends on both class label and feature values---for example, while the true distribution of positive instances is uniform across demographics, they are less likely to be observed for particular demographic subgroups. This leads to a localized minority imbalance problem, posing a deeper challenge beyond general class-count imbalance. We show that under localized minority imbalance, existing imbalance mitigation techniques can overfit the observed training data and generalize poorly to under-observed regions of the minority distribution. To address this, we propose a tree-based stratification approach that recursively partitions the feature space to construct approximately unbiased subsets of the training data. For each stratum, we pair its majority instances with the full observed minority set and train a base classifier to create an ensemble. Extensive experiments over benchmark tabular datasets simulated with localized minority imbalance show that our approach outperforms popular and state-of-the-art imbalance mitigation techniques. We also introduce a gold-standard evaluation protocol that uses unbiased test sets, and demonstrate that conventional hold-out evaluation from the same localized-imbalance data can substantially bias performance. Overall, our results highlight that the cause of imbalance is as important as the correction method.


Learning What's Real: Disentangling Signals and Measurement Artifacts in Multi-Sensor Data, with Applications to Astrophysics

Pablo Mercader-Perez ⋅ Carolina Cuesta Lazaro ⋅ Daniel Muthukrishna ⋅ Jeroen Audenaert ⋅ V Villar ⋅ David W Hogg ⋅ Marc Huertas-Company ⋅ Bill Freeman

Data collected from the physical world is always a combination of multiple sources: an underlying signal from the physical process of interest and a signal from measurement-dependent artifacts from the sensor or instrument. This secondary signal acts as a confounding factor, limiting our ability to extract information about the underlying physics. Moreover, it poses significant challenges for combining data in heterogeneous or multi-instrument frameworks. To disentangle these factors of variation, we propose a dual-encoder architecture with a counterfactual generation objective that leverages overlapping observations. The resulting representations explicitly separate intrinsic signals from sensor-specific distortions and noise, and can be used for counterfactual view generation, parameter inference, and instrument-independent similarity search---all unconfounded by measurement artifacts. We demonstrate the effectiveness of our approach in a multi-instrument setting on astrophysical galaxy images from the DESI Legacy Imaging Survey (Legacy) and the Hyper Suprime-Cam (HSC) Survey. This framework provides a general recipe for scientific self-supervised pretraining: construct training pairs from overlapping observations of the same physical system, treat sensor- or modality-specific effects as augmentations, and learn invariant representations through counterfactual generation.


Leech Lattice Vector Quantization for Efficient LLM Compression

Tycho F van der Ouderaa ⋅ Mart van Baalen ⋅ Paul Whatmough ⋅ Markus Nagel

Scalar quantization of large language models (LLMs) is fundamentally limited by information-theoretic bounds. While vector quantization (VQ) overcomes these limits by encoding blocks of parameters jointly, practical implementations must avoid the need for expensive lookup mechanisms or other explicit codebook storage. Lattice approaches address this through highly structured and dense packing. This paper explores the Leech lattice, which, with its optimal sphere packing and kissing configurations at 24 dimensions, is the highest dimensional lattice known with such optimal properties. To make the Leech lattice usable for LLM quantization, we extend an existing search algorithm based on the extended Golay code construction, to i) support indexing, enabling conversion to and from bitstrings without materializing the codebook, ii) allow angular search over union of Leech lattice shells, iii) propose fully-parallelisable dequantization kernel. Together this yields a practical algorithm, namely Leech Lattice Vector Quantization (LLVQ). LLVQ delivers state-of-the-art LLM quantization performance, outperforming recent methods such as Quip#, QTIP, and PVQ. These results highlight the importance of high-dimensional lattices for scalable, theoretically grounded model compression.


LessMimic: Versatile Humanoid-Object Interaction with Unified Distance Field Representations

Yutang Lin ⋅ Jieming Cui ⋅ Yixuan Li ⋅ Wei Liang ⋅ Baoxiong Jia ⋅ Yixin Zhu ⋅ Siyuan Huang

Humanoid robots capable of versatile whole-body interaction with everyday objects represent a central goal of embodied intelligence. Existing approaches often rely on motion-reference inputs or task-specific rewards, coupling policies to particular motion scripts, object geometries, and contact timings. This limits geometric generalization and flexible interaction control, where simple commands specify motion intent while object geometry determines how contact should be executed. We introduce LessMimic, a framework for versatile motion-free-at-inference humanoid–object interaction. At inference, a single whole-body policy is steered by root commands and an interaction-type flag instead of motion-reference inputs, and is conditioned on a compact interaction representation built from short histories of distance-field-derived surface distances and local surface directions. This geometry-conditioned interface enables command-steered control while adapting contact behavior to local object geometry. Through visual distillation, LessMimic further enables egocentric-depth-based deployment for MoCap-free sensing. Across object scales from 0.4× to 1.6×, LessMimic maintains robust performance over four interaction tasks, including PickUp, SitStand, Push, and Carry, where motion-conditioned baselines degrade sharply away from the training scale. Beyond single-task evaluation, LessMimic further supports sequential composition of heterogeneous interaction skills, attaining 62.1% success on five-task sequences and retaining 23.5% success at 15 task instances. By grounding interaction control in local geometry rather than motion references, LessMimic provides a path toward geometry-generalizable, command-steerable humanoid–object control with heterogeneous skill composition.


Leveraging Psychophysical Attentional Distribution for Gaze-Augmented Reward Modeling

Chaeho Lee ⋅ Junyup Kim ⋅ Wansoo Kim ⋅ Young MIn Jung ⋅ Sang Ho Lee

Integrating human feedback into language models has gained increasing attention as a means to align model outputs with human preferences, with Reinforcement Learning from Human Feedback (RLHF) serving as a prominent framework for preference-based alignment. Recent studies have incorporated eye-tracking (ET) data as an additional supervisory signal for reward model training, grounding reward learning on human reading behavior. Building on this approach, we extend prior fixation-centric models by leveraging established findings from cognitive psychology on human letter recognition. Specifically, we model visual attention as a spatially graded distribution centered on the current gaze location, and incorporate this representation into reward model training. Our results show that the reward model trained with the proposed gaze distribution achieves higher preference prediction accuracy over a baseline model. Moreover, the fitted attentional distribution reflects key properties of human attention during reading, suggesting that it captures cognitively meaningful aspects of attention allocation.


LiFi: LiDAR Generation from Multi-View Images via Geometric and Semantic Collaborative Guidance

Sizhuo Zhou ⋅ Xiaosong Jia ⋅ Fanrui Zhang ⋅ Qifeng Li ⋅ Junjie Li ⋅ Zirui Wang ⋅ Yukang Feng ⋅ Yu Hong ⋅ Shaofeng Zhang ⋅ Wenlong Liao ⋅ Tao He ⋅ Juyong Zhang ⋅ Junchi Yan

Generating LiDAR point clouds from multi-view camera images is a valuable task with applications in controllable data synthesis, cross-modal simulation and various perception tasks. However, most existing methods are not specifically designed for this task, and the quality of their generated LiDAR remains limited. Moreover, these methods typically focus only on semantic cues in images or generate LiDAR from a monocular camera image. In this work, we introduce LiFi, a dedicated framework for generating visually and geometrically aligned LiDAR scenes from multi-view camera images. LiFi guides the latent diffusion process through two complementary branches: geometry and semantics. The geometry branch constructs scene features from images through a reverse sampling algorithm and an uncertainty-aware geometric encoder combined with DepthAnything3, whereas the semantic branch provides high-level semantic representations via a cross-view interaction module. We further introduce a dual-branch balanced classifier-free guidance (CFG) strategy, which enhances the model’s conditional generation capability while preserving the independent learning and collaborative controllability of the two branches. Extensive experiments demonstrate that LiFi outperforms state-of-the-art methods in generating high-fidelity LiDAR scenes. We will make this project publicly available.


Lift, See, Act: Hierarchical Robot Policy Pretraining with 3D Foundation Models

Yiyuan Ge ⋅ Changxing Ding ⋅ Ziyu Hao ⋅ Zijie Zheng ⋅ Xiangmin Xu

Robot policy pretraining based on human videos is crucial for improving the policy’s generalization ability. One main challenge for this task is the lack of explicit action-relevant representations in such unlabeled data. Recent works tend to estimate 3D hand motion trajectories from videos using 3D foundation models (3D FMs). However, they keep solely the sparse trajectories as supervision, ignoring the rich, fine-grained information produced during the trajectory extraction process. To address this issue, we propose $\textbf{LSA}$, a hierarchical robot policy pretraining framework that includes three phases: $\textbf{L}$ift, $\textbf{S}$ee, and $\textbf{A}$ct. Specifically, the Lift phase elevates 2D video observations into dense 3D representations by adaptively aligning features in the policy's shallow layers with the intermediate representations extracted from multiple 3D FMs. Building upon these lifted representations, the See phase equips the policy with explicit geometric and interactive awareness through dual guidance. It introduces depth-map reconstruction to help the model comprehend 3D spatial layouts and utilizes hand region cues to explicitly supervise the encoder’s attention toward task-relevant interaction hotspots. Empowered by the dense 3D representation learning and precise spatial guidance, our policy achieves robust and accurate robotic manipulation in the Act phase. Extensive experiments on diverse simulation and real-world tasks demonstrate that LSA significantly outperforms current state-of-the-art approaches. The code of this work will be released soon.


Linear approximations to HMM filtering

Andrew Mah ⋅ Joshua L Pughe-Sanford ⋅ Sarah Harvey ⋅ Alex Williams

Complex sequence models, such as transformers and state space models (SSMs), learn to represent latent belief states when trained on next-token prediction. Surprisingly, we find that in previously studied tasks, this phenomenon can also be captured by linear recurrent models. This raises a natural question: when are linear models sufficient for optimal belief state approximation? We study this question in the canonical setting of hidden Markov models (HMMs), where Bayes-optimal prediction requires nonlinear filtering. We characterize the full class of HMMs whose filtering dynamics are exactly realizable by linear recurrent neural networks (RNNs). More generally, for arbitrary HMMs, we establish a dissipation relation which shows that approximation error decays at a rate determined by intrinsic properties of the underlying HMM. We validate our theoretical results through experiments on both randomly sampled HMMs and constructed adversarial examples. Together, these findings clarify when linear sequence models suffice for optimal inference and when nonlinearity is fundamentally necessary.


Lipschitz Dueling Bandits over Continuous Action Spaces

Mudit Sharma ⋅ Shweta Jain ⋅ Vaneet Aggarwal ⋅ Ganesh Ghalme

We study for the first time, stochastic dueling bandits over continuous action spaces with Lipschitz structure, where feedback is purely comparative. While dueling bandits and Lipschitz bandits have been studied separately, their combination has remained unexplored. We propose the first algorithm for Lipschitz dueling bandits, using round-based exploration and recursive region elimination guided by an adaptive reference arm. We develop new analytical tools for relative feedback and prove a regret bound of $\tilde O\!\left(T^{\frac{d_z+1}{d_z+2}}\right)$, where $d_z$ is the zooming dimension of the near-optimal region. Further, our algorithm takes only logarithmic space in terms of the total time horizon, best achievable by any bandit algorithm over continuous action space.


LL-Bench: Rethinking Low-Level Vision Evaluation in the Era of Large-Scale Generative Models

Lu Liu ⋅ Huiyu Duan ⋅ Chenxin Zhu ⋅ Jintong Lu ⋅ Haoyun Jiang ⋅ Liu Yang ⋅ Qiang Hu ⋅ Guangtao Zhai ⋅ Xiaoyun Zhang

Large-scale generative models have demonstrated remarkable capabilities across image generation and editing tasks. However, their performance in low-level vision tasks, which require pixel-wise control, remains insufficiently studied. To address this gap, we introduce \textbf{LL-Bench}, a comprehensive \textbf{Benchmark} for evaluating the capabilities of large-scale generative models on \textbf{L}ow-\textbf{L}evel vision tasks. The benchmark comprises 2,469 real-world degraded images covering 16 low-level degradation tasks, and 28,919 restored images produced by 10 state-of-the-art large-scale generative models and 21 conventional restoration models, which are annotated with 152,020 expert-level pairwise human preferences and 28,334 quality scores. Built upon LL-Bench, we present a systematic diagnosis that reveals the performance boundaries and unique failure modes of large-scale generative models across diverse low-level vision tasks, compared with conventional representative restoration approaches. Moreover, we investigate the effectiveness of current quality evaluation metrics on LL-Bench, which exhibit significant discrepancy with human ratings. To better align restored-image quality assessment with human preferences, we further propose \textbf{LL-Score}, an MLLM-based evaluator that captures both restoration quality and hallucination existence. Extensive experiments demonstrate that LL-score not only outperforms existing image quality assessment metrics, but also serves as a promising reward model for training generative models on low-level vision tasks. The database and code are available at \url{https://anonymous.4open.science/r/LL-Bench-04F0}


LLM Alignment--Utility Asymmetry under Semantic-Preserving Transformations

Mohan Li ⋅ Chengyu Yu ⋅ Francesco Sovrano ⋅ Marc Langheinrich ⋅ Martin Gjoreski

Large Language Model (LLM) alignment is intended to ensure that models remain helpful and safe, but its stability under input distributional shift is not yet fully understood. Prior work shows that aligned models can fail under jailbreak prompts, alternative encodings, and cross-lingual transfer, yet these failures are usually studied as attacks rather than controlled probes of alignment generalization. Moreover, existing evidence is largely grounded in natural language variation already represented during pretraining, leaving unresolved whether alignment generalizes with semantic content or remains tied to superficial surface patterns. In this paper, we study this question using synthetic semantic-preserving transformations that are rule-based and invertible, preserving task-relevant meaning while shifting inputs beyond standard linguistic variation. Across four open-weight and four commercial models, under both fine-tuning and in-context learning, we use these transformations as a probe of alignment generalization and identify an empirical pattern we term **Alignment--utility asymmetry**: once models can operate effectively on transformed inputs, task utility is often substantially retained while alignment failure increases more sharply. For example, adapted GPT-4.1 mini shows only limited utility degradation under transformation while its harmful rate rises from $13.3\%$ to $74.3\%$; Gemini 3 Flash similarly retains near-original utility while its harmful rate increases from $2.3\%$ to $43.0\%$. Taken together, these results suggest that semantic-preserving distribution shifts can expose a recurring gap in how utility and alignment generalize in current LLMs.


LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL

Yujin Kim ⋅ Namgyu Ho ⋅ Sangmin Hwang ⋅ Joonkee Kim ⋅ Yongjin Yang ⋅ Sangmin Bae ⋅ Seungone Kim ⋅ Jaehun Jung ⋅ Se-Young Yun ⋅ Hwanjun Song

Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as reward signals. While recent methods adapt these rubrics to the evolving policy during training, the training prompts themselves remain static, drawn from fixed corpora. This static approach often results in a critical misalignment between prompt difficulty and policy capability, leaving the judge unable to recover a discriminative reward signal when prompts fail to elicit quality variance among rollouts. To address this misalignment, we introduce LLM-as-a-Tutor, a framework that extends the LLM’s role from judge to tutor: a single model serves as an examiner that pairwise-compares policy rollouts to detect non-challenging prompts, and as a generator that appends atomic constraints to them. This append-only design monotonically raises difficulty in step with the policy’s capability, producing a self-calibrating training signal without external difficulty schedules. On three complex instruction-following benchmarks, our method consistently outperforms both policy-unaware baselines and prior policy-adaptive methods that adapt rubrics or rewrite prompts, suggesting prompt adaptation as a missing axis of policy-awareness in non-verifiable RL.

Scaling improves language-model capability, but it does not necessarily strengthen internal safety mechanisms. We introduce LLM Rheology, a representation-level framework for auditing how aligned language models respond to adversarial perturbation inside activation space. Given a learned refusal direction, we define manifold sensitivity / compliance as the Fisher--Rao-normalized distributional response induced by controlled activation perturbation. This quantity measures how strongly refusal behavior remains coupled to adversarial task execution in representation space. Across six open model families and 22 checkpoints, including Qwen, DeepSeek-Distill, Mistral, Llama-3, Gemma, and Yi, we observe a recurring geometric scaling pattern. Small aligned models often exhibit elevated refusal response, intermediate-scale models frequently enter weakened-response regimes (compliance valleys), and larger checkpoints trend toward near-baseline response consistent with increasing task--refusal decoupling. Cosine measurements on continuous Qwen and DeepSeek scaling axes support this interpretation, showing monotonic decay between refusal and task-generation directions with scale. We further provide causal evidence through inference-time activation intervention on Qwen-72B. Injecting a learned refusal vector at a critical semantic layer restores refusal behavior on the evaluated jailbreak subset while largely preserving benign reasoning performance. Matched-norm placebo vectors fail to reproduce the effect, supporting the directional specificity of the intervention. Together, these results suggest that behavioral refusal can remain surface-level even when internal refusal geometry becomes weakly coupled to adversarial task execution. LLM Rheology provides a complementary representation-level perspective for auditing refusal safety beyond behavioral evaluation.


Load Balancing Mixture of Experts with Similarity Preserving Routers

Nabil Omi ⋅ Siddhartha Sen ⋅ Ali Farhadi

Sparse Mixture of Experts (MoE) models offer a scalable and efficient architecture for training large neural networks by activating only a subset of parameters (“experts”) for each input. A learned router computes a distribution over these experts, and assigns input tokens to a small subset. However, without auxiliary balancing mechanisms, routers often converge to using only a few experts, severely limiting model capacity and degrading performance. Most current load balancing mechanisms encourage a uniform routing probability across experts. Early in pretraining, this can result in inconsistent routing behavior, resulting in the model spending its capacity learning redundant knowledge. We address this by introducing a novel load balancing loss that helps preserve relational structure, encouraging consistent expert choices for similar inputs during training. Our experimental results show that applying our loss to the router results in 36% faster convergence and lower redundancy compared to a popular load balancing loss.

Off-dynamics offline reinforcement learning (RL) aims to learn a policy for a target domain using limited target data and abundant source data collected under different transition dynamics. Existing methods typically address dynamics mismatch either globally over the state space or via pointwise data filtering; these approaches can miss localized cross-domain similarities or incur high computational cost. We propose Localized Dynamics-Aware Domain Adaptation (LoDADA), which exploits localized dynamics mismatch to better reuse source data. LoDADA clusters transitions from source and target datasets and estimates cluster-level dynamics discrepancy via domain discrimination. Source transitions from clusters with small discrepancy are retained, while those from clusters with large discrepancy are filtered out. This yields a fine-grained and scalable data selection strategy that avoids overly coarse global assumptions and expensive per-sample filtering. We provide theoretical insights and extensive experiments across environments with diverse global and local dynamics shifts. Results show that LoDADA consistently outperforms state-of-the-art off-dynamics offline RL methods by better leveraging localized distribution mismatch.

We study bandit convex optimization (BCO) when feedback is adversarially delayed and the comparator sequence is also adversarially non-stationary. This setting arises in online advertising, recommender systems, and adaptive dosing, but existing algorithms either absorb one of the two axes into a black-box global penalty or pay a worst-case cost that ignores the local interaction between delay and drift. We propose Oracle-DARS-DBCO, a delay-aware blocking algorithm whose block length $K$ and one-point smoothing radius $\delta$ satisfy the single identity $K\delta^{2}\asymp n^{2}$. This identity cancels the multiplicative coupling between delay deviation and one-point estimator variance, so the regret of each segment is paid only on that segment. We prove matching upper and lower bounds on the expected dynamic regret: $\widetilde{O}\bigl(\sqrt{n}\,m^{1/4}T^{3/4}+\sqrt{d_{\max}\,mT}\bigr)$ for convex losses and $\widetilde{O}\bigl(n^{2/3}m^{1/3}T^{2/3}+d_{\max}\,m\log(eT/m)\bigr)$ for $\alpha$-strongly convex losses, where $m=S_{T}+1$ is the number of stationary segments and $d_{\max}$ is the worst-case delay. The matching Rademacher-sign lower bounds $\Omega\bigl(\sqrt{d(S_{T}+1)\,T}\bigr)$ and $\Omega\bigl(d(S_{T}+1)\bigr)$ show that the delay price is segment-local. On fourteen scaling experiments, the fitted log--log exponents agree with the predicted ones within $0.04$, and Oracle-DARS-DBCO reduces regret by up to $9.17\times$ relative to the strongest non-restarting baseline on a piecewise-stationary benchmark.


LogicTree-RAG: Logic Tree-guided Retrieval-Augmented Generation for Long-form Patent Drafting

Jiaqi Zhu ⋅ Naili Xing ⋅ Pan Hexiang ⋅ Haotian Gao ⋅ Jianwei Yin ⋅ Xiaokui Xiao ⋅ Beng Chin Ooi

Long-form technical text generation underpins knowledge-intensive workflows, yet remains challenging for large language models (LLMs) due to the need for globally consistent logical structuring and faithful technical reasoning beyond local coherence. Patent drafting is a canonical instance of this challenge, demanding holistic generation of a legally compliant and technically exhaustive document through sustained multi-expert collaboration. Existing approaches often focus on partial section generation or rely on manually crafted outlines, limiting scalable automation in realistic settings. In this work, we propose LogicTree-RAG, a logic tree-guided retrieval-augmented generation framework that induces a hierarchical logic tree as a global organizational backbone to organize and ground technical disclosures, without relying on expert-defined drafting priors. Each node in the logic tree represents a technical element and is constructed through evidence-guided recursive generation. A hybrid traversal mechanism then maps the logic tree into patent sections, enabling controllable and section-balanced generation. Extensive experiments show that LogicTree-RAG consistently improves content quality and language conformity over strong LLM-based baselines and achieves longer structured generation with high token efficiency, demonstrating the effectiveness of logic-centric generation for complex technical document drafting.


Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds

Jiaming Bian ⋅ Bingliang Li ⋅ Yuehao Wu ⋅ Pichao WANG ⋅ Zhi Wang ⋅ Hailan Ma ⋅ Huadong Mo ⋅ Zhenhong Sun

As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting given observations toward actively deciding what to observe. We study this problem through camera planning in dynamic 3D story worlds, where the camera must not only generate smooth motion, but also decide what visual evidence should be acquired before it moves. We formulate this capability as \textbf{Narrative-Grounded World Visual Attention}, where the camera acts as an embodied observer that determines what to observe, how to compose the observation, and how to shift attention over time under narrative intent and physical 3D constraints. To realize this capability, we propose \textbf{Look-Before-Move}, a camera planning framework that separates observation specification from motion execution. It first builds a Semantic Observation Contract to convert directorial intent into executable visual constraints, then performs Monte Carlo Viewpoint Search to find narrative-compliant and geometrically feasible viewpoints, and finally applies Semantic Trajectory Grounding to connect selected viewpoints into continuous, collision-aware, and temporally coherent camera motion. We further construct a dynamic 3D Story World Benchmark based on \textit{StoryBlender}, covering 50 stories, 457 scenes, and 1585 shots with animated characters, semantic scene configurations, and executable 3D environments. Experiments show that our framework improves subject perception, intent consistency, and trajectory quality over representative baselines, demonstrating the importance of organizing visual attention before generating camera motion.


LOTION: Smoothing the Optimization Landscape for Quantized Training

Mujin Kwun ⋅ Depen Morwani ⋅ Huangyuan Su ⋅ Stephanie Gil ⋅ Nikhil Anand ⋅ Sham Kakade

Optimizing neural networks for quantized objectives is fundamentally challenging because the quantizer is piece-wise constant, yielding zero gradients everywhere except at quantization thresholds where the derivative is undefined. Most existing methods deal with this issue by relaxing gradient computations with techniques like Straight Through Estimators (STE) and do not provide any guarantees of convergence. In this work, taking inspiration from Nesterov smoothing, we approximate the quantized loss surface with a continuous loss surface. In particular, we introduce LOTION, Low-precision Optimization via sTochastic-noIse smOothiNg, a principled smoothing framework that replaces the raw quantized loss with its expectation under unbiased randomized-rounding noise. In this framework, standard optimizers are guaranteed to converge to a local minimum of the loss surface. Moreover, when using noise derived from stochastic rounding, we show that the global minima of the original quantized loss are preserved. We empirically demonstrate that this method outperforms standard QAT on synthetic testbeds and on 150M-, 300M-, and 600M- parameter language models. On the 150M INT4 benchmark with modern QAT baselines, LOTION further achieves the best quantized validation loss.


Lower-Level Agnostic Bilevel Optimization

Peiwen Qiu ⋅ Prashant Khanduri ⋅ Jia (Kevin) Liu

Bilevel optimization is a fundamental framework for machine learning problems, where an upper-level (UL) objective depends on the solution of a nested lower-level (LL) problem. Existing bilevel algorithms typically assume that the LL objective is available for evaluation, differentiation, or unrolling. However, this assumption can fail when the LL response is produced by a black-box solver, simulator, adversary, or learned predictor. In this paper, we study *lower-level agnostic* bilevel optimization, where the UL learner has access only to an approximate LL solution $\bar{\mathbf{y}}(\mathbf{x})$ and cannot query the LL objective, its gradients, Hessian/Jacobian information, or the procedure that generates the LL response. We propose LeGo-BiO (**L**ower-l**e**vel A**g**n**o**stic **Bi**level **O**ptimization), a coordinate-wise pseudo-hypergradient method that estimates the missing Jacobian $\nabla _{\mathbf{x}}\bar{\mathbf{y}}(\mathbf{x})$ using finite differences of $\bar{\mathbf{y}}(\cdot)$. For each UL coordinate, LeGo-BiO constructs a decision vector subject to 1-sparse updates, i.e., restricting modifications to a single coordinate per step, thereby isolating the corresponding Jacobian column needed in the UL gradient computation. This avoids the directional-projection limitation of full-parameter response differences and, unlike zeroth-order approximations of the entire UL gradient, preserves the available first-order gradient information with respect to the UL variables for more accurate UL gradient evaluation. We establish convergence guarantees for LeGo-BiO in both deterministic and stochastic settings under an LL Polyak–Łojasiewicz condition, without requiring LL strong convexity. Our bounds yield an $\mathcal{O}(T^{-1})$ rate in the deterministic setting when the LL approximation error decreases geometrically, and an $\mathcal{O}(T^{-1/2})$ rate in the stochastic setting under general LL approximation errors. Experiments on deep hyper-representation and adversarial training demonstrate competitive performance of LeGo-BiO despite being agnostic to the LL objective.


LSVD: Loss-Aware Low-Rank Approximation for Efficient Low-Precision Vision-Language Models

Haiyu Wang ⋅ Yutong Wang ⋅ Leshu Li ⋅ Yihui Ren ⋅ Sai Qian Zhang

Vision-language models (VLMs) deliver strong multimodal reasoning capabilities, but their large computational cost and high parameter counts make deployment challenging on resource-constrained devices. Low-rank decomposition has emerged as a promising compression technique, yet existing methods often optimize local matrix reconstruction error, rely on uniform or heuristic rank allocation, and focus mainly on attention projections while leaving feed-forward networks underexplored. In this paper, we propose~\textit{LSVD}, a loss-aware low-rank approximation framework for efficient low-precision VLMs. LSVD derives a curvature-weighted SVD objective from a second-order approximation of the model loss and uses Kronecker-factored Fisher information to guide decomposition toward downstream performance rather than reconstruction alone. We further introduce a loss-aware cross-layer rank allocation strategy based on calibration gradients, enabling more effective parameter budgeting across layers. Finally, we extend low-rank compression to FFN layers through a hybrid scheme that combines SVD with quantization. The evaluation results show that LSVD achieves over $2.3\times$ decoding speedup over previous work while preserving strong accuracy under low-precision inference.


Machine Learning-Driven RAG System Design

Anastasia Orlova ⋅ Nina Gubina ⋅ Aleksei Dmitrenko ⋅ Arsen Sarkisyan ⋅ Nikita Vetoshkin ⋅ Anastasiia Gorbunova ⋅ Ivan Dubrovsky ⋅ Julia Razlivina ⋅ Bogdan Neterebskii ⋅ Andrei Dmitrenko

Retrieval-Augmented Generation (RAG) has become a standard approach for grounding large language models in external knowledge, yet its performance depends critically on a large number of interacting design choices. Systematic optimization of these choices remains costly, as each configuration must be evaluated through full pipeline execution. In this work, we propose to treat RAG optimization as a prediction problem, using surrogate models to approximate the functional dependencies between configuration parameters and performance metrics. Using a stage-wise configurable RAG framework, we conduct large-scale experiments across three chemistry QA domains, systematically evaluating 648 configurations per domain and publicly releasing the resulting datasets of configurations and evaluation metrics. Surrogate models trained on this data achieve strong predictive accuracy ($R^2 \geq 0.84$ across all metrics and domains), and SHAP-IQ-based analysis of the learned models reveals consistent parameter interaction patterns across domains — despite substantial differences in corpus characteristics. We exploit this cross-domain regularity in transfer experiments: surrogate predictions generalize to unseen domains without target-domain training data on two transfer-robust metrics, and achieve full transfer across all six metrics with as little as 10\% of target-domain data. Finally, we apply surrogate models to optimize RAG configurations over an expanded search space of 61,152 candidates, recovering top-performing configurations in seconds rather than days, with gains of up to +0.09 in key retrieval and generation evaluation metrics. Together, these results demonstrate that RAG design spaces are structured, transferable, and efficiently optimizable via surrogate modeling.


MAEB: Massive Audio Embedding Benchmark

Adnan E Assadi ⋅ Isaac Chung ⋅ Chenghao Xiao ⋅ Roman Solomatin ⋅ Animesh Jha ⋅ Rahul Chand ⋅ Silky Singh ⋅ Kaitlyn Wang ⋅ Ali Sartaz Khan ⋅ Marc M Nasser ⋅ Sufen Fong ⋅ Pengfei He ⋅ Alan Xiao ⋅ Ayush Sunil Munot ⋅ Aditya Shrivastava ⋅ Artem Gazizov ⋅ Niklas Muennighoff ⋅ Kenneth Enevoldsen

We introduce the Massive Audio Embedding Benchmark (MAEB), a large-scale benchmark covering 30 tasks across speech, music, environmental sounds, and cross-modal audio-text reasoning in 100+ languages. We evaluate 56 models and find that no single model dominates across all tasks: contrastive audio-text models excel at environmental sound classification (e.g., ESC50) but score near random on multilingual speech tasks (e.g., SIB-FLEURS), while speech-pretrained models show the opposite pattern. Clustering remains challenging for all models, with even the best-performing model achieving only modest results. We observe that models excelling on acoustic understanding often perform poorly on linguistic tasks, and vice versa. We also release MAEB+, the full 98-task collection from which MAEB is curated; to our knowledge, it is the largest publicly released collection of audio embedding evaluation tasks. MAEB preserves task diversity while reducing evaluation cost and integrates into the MTEB ecosystem for unified evaluation across text, image, and audio modalities. We release MAEB and all 98 tasks along with code and a leaderboard at https://anonymous.4open.science/r/mteb-2FEB/.


MAGE: Multi-Agent Self-Evolution with Co-Evolutionary Knowledge Graphs

Ruiyi Yang ⋅ Zechen Li ⋅ Hao Xue ⋅ Imran Razzak ⋅ Flora Salim

Self-evolving language-model agents must decide what to learn next and how to preserve what they have learned across iterations. Existing systems typically carry this cross-iteration knowledge as natural-language feedback, flat episodic memory, or implicit reinforcement signals, none of which cleanly supports a frozen weak backbone at inference time. This paper introduces MAGE (Multi-Agent Graph-guided Evolution), a framework that externalizes self-knowledge into a four-subgraph co-evolutionary knowledge graph. Its experience subgraph stores both teacher-written failure corrections and the learner's own past correct reasoning traces, which are retrieved as task-conditioned guidance for a frozen execution model. During evolution, the graph, a task-level search bandit, and a skill-level routing bandit are updated from the same reward stream, while the learner's backbone remains unchanged. We further provide structural analysis showing how append-only memory growth, bounded curriculum coverage, and task-filtered retrieval together support stable improvement of the retrieval substrate for frozen-learner evolution. Across nine benchmarks spanning mathematical reasoning, multi-hop and open-domain question answering, spatio-temporal analysis, financial numerical reasoning, medical multiple-choice, an open-world survival game, and web navigation, MAGE achieves strong performance against prompt-based frozen-backbone baselines. Ablations show that self-harvested success traces and teacher-written corrections are complementary, with success memories contributing most on reasoning-template-heavy tasks and corrective memories supporting harder composition and interaction settings.


Make Each Token Count: Towards Improving Long-Context Performance with KV Cache Eviction

Ngoc Bui ⋅ Trung Hieu Nguyen ⋅ Arman Cohan ⋅ Rex Ying

The key--value (KV) cache is a major bottleneck in long-context inference, where memory and computation grow with sequence length. Existing KV eviction methods reduce this cost but typically degrade performance relative to full-cache inference. Our key insight is that full-cache attention is not always optimal: in long contexts, irrelevant tokens can dilute attention away from useful evidence, so selective, learnable eviction can improve generation rather than merely approximate the full cache. We introduce a global retention-based KV eviction method that learns each token's future utility under a unified memory budget. Lightweight retention gates assign utility scores to cached KV entries, and a shared final scoring projection calibrates these scores across all layers and heads. This enables a single global eviction policy in which tokens from different layers, heads, and modalities compete directly for cache capacity. We further provide theoretical analysis showing that preferentially retaining useful tokens reduces attention dilution, and we justify geometric retention as a query-agnostic proxy for future utility. Across long-context language and vision--language benchmarks, our method substantially reduces KV memory while matching or surpassing full-cache inference. These results suggest that learned, globally calibrated KV eviction is not only a compression technique, but also a mechanism for improving long-context reasoning.


MapPFN: Learning Causal Perturbation Maps in Context

Marvin Sextro ⋅ Weronika Kłos ⋅ Gabriel Dernbach

Planning effective interventions in biological systems requires treatment-effect models that adapt to unseen biological contexts by identifying their specific underlying mechanisms. Yet single-cell perturbation datasets span only a handful of biological contexts, and existing methods cannot leverage new interventional evidence at inference time to adapt beyond their training data. To meta-learn a perturbation effect estimator, we present MapPFN, a prior-data fitted network (PFN) pre-trained on a synthetic biological prior with causal interventions, decoupling pre-training from limited wet-lab data. Unlike existing methods, MapPFN uses in-context learning to map a sequence of experiments to a post-perturbation distribution, enabling a single pre-trained model to adapt to new datasets and arbitrary gene sets at inference time. Zero-shot, MapPFN identifies differentially expressed genes on par with models trained on real single-cell data, and fine-tuning further improves predictions across biological contexts. Our code is available at https://anonymous.4open.science/r/MapPFN.


MARBLE: an Agent Benchmark for Spatial Reasoning and Visual Abstraction

Yulun Jiang ⋅ Naël Ouerghemi ⋅ Yekun Chai ⋅ Maria Brbic ⋅ Michael Moor

A central challenge for multimodal language models (MLLMs) in real-world tasks is to abstract visual observations into structured spatial states and reason over how those states change under actions over long-horizons. Existing multimodal benchmarks often focus on static question answering, where answers can be directly extracted from images, or assume symbolic states are available and bypass the visual-to-spatial abstraction process. We present MARBLE, a diagnostic benchmark for spatially grounded multimodal reasoning and planning across three agent environments: Cube, Maze and Blox. MARBLE spans 2D and 3D spatial puzzles under geometric constraints, dynamic layouts, and spatial transformations, with difficulty systematically controlled along two axes: observation format and planning horizon. Extensive evaluations show that frontier MLLMs struggle on MARBLE, with the strongest model GPT-5.4 achieving only a 7.3\% average success rate across the three environments in the hardest setting. Pairwise ablations on the open-weight models identify visual-to-spatial abstraction as the major bottleneck: when models must infer spatial states from images instead of receiving human-annotated states, success drops from 50.7\% to 3.7\% for Gemma-4-31B and from 87.9\% to 6.5\% for Kimi-K2.6. The results indicate that models often fail to construct precise spatial states from visual observations before planning, a task that is straightfoward for human. Overall, MARBLE reveals the gap of reliable visual-to-spatial for frontier MLLMs and provides a controlled testbed for measuring progress in spatially grounded multimodal reasoning and planning.


Masked Visual Actions for Unified World Modeling

Hadi Alzayer ⋅ Wenlong Huang ⋅ Haonan Chen ⋅ Christopher Luey ⋅ Lvmin Zhang ⋅ Maneesh Agrawala ⋅ Gordon Wetzstein ⋅ Fei-Fei Li ⋅ Yilun Du ⋅ Jiajun Wu ⋅ Jia-Bin Huang

Video models absorb rich priors over how the visual world moves, interacts, and responds to contact, making them promising substrates for robotic world modeling. The central challenge is how to communicate action to such models in a form aligned with the visual space in which they learned these interaction priors, yet still grounded in physical manipulation. We introduce Masked Visual Actions, a pixel-space control interface that expresses action as a partially revealed trajectory of an arbitrary entity in a video. Revealing robot motion makes the model act as a forward dynamics model that predicts the scene's response to low-level robot actions, while revealing desired object motion makes the same model recover robot behavior consistent with that outcome. Finetuned with only 15 hours of masked examples from real videos and simulation, a single checkpoint achieves strong visual fidelity and controllability across diverse scenes and multiple embodiments. In downstream manipulation settings, the model produces imagined rollouts whose outcomes correlate with real-world execution for policy evaluation, improves decision making by ranking candidate futures in model-based planning, and supports inverse modeling by synthesizing robot motion from desired object motion.


MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models

Yingzi Ma ⋅ Zhengyue Zhao ⋅ Xiaogeng Liu ⋅ Minhui Xue ⋅ Yue Zhao ⋅ Chaowei Xiao

Diffusion large language models (dLLMs) generate text by iteratively denoising partially masked sequences under bidirectional context, exposing a safety surface distinct from autoregressive LLMs. Because mask tokens are native inputs and tokens are committed by confidence rather than position, harmful content can be induced through infilling and outside the monitored prefix. Existing jailbreaks either miss this native infill capability or rely on low-diversity mask-bearing templates applied uniformly across goals, with little structural adaptation or accumulated attack experience. We propose MaskForge, a fully black-box adaptive attack that casts dLLM red-teaming as optimized search over a growing library of structural patterns. MaskForge abstracts successful attempts into reusable schemas, selects goal-compatible patterns with a UCB bandit, and invokes a scorer-guided fallback when the current library fails. Successful attempts are distilled back into the pattern library, enabling experience to accumulate across goals. Across five public dLLMs and three benchmarks, MaskForge achieves an average attack success rate of 79.3\%, a 17.6\% relative improvement over the strongest competing dLLM baseline. The matured pattern library further transfers to AdvBench without any updates, achieving a 88.2\% attack success rate and a 67\% relative improvement over the strongest competing baseline.

Large language models are increasingly deployed for open-ended question answering, yet their answers rarely come with calibrated uncertainty. We propose CAELUM, a framework that builds semantic prediction sets and separates aleatoric from epistemic uncertainty in their construction. CAELUM samples answers under prompt-ensembled clarifications, clusters them by semantic equivalence, estimates aleatoric and epistemic components via kernel Von Neumann entropy, and calibrates cluster-APS nonconformity scores with split conformal prediction. A per-instance weight $w_\lambda(x)=1+U_{\mathrm{ale}}(x)+\lambda\,U_{\mathrm{epi}}(x)$ controls how epistemic uncertainty enters the set, and a per-cell coefficient $\lambda^\star$ is fit on a held-out split to minimise prediction-set size. Our central observation is that, in open-ended generation, coverage on cluster-APS scores is silently capped by \emph{matchability}: the fraction of items whose reference answer appears in any sampled cluster. We make this dependence verifiable with a Mondrian split-conformal bound on a learned stratum, where matchability is predicted from pre-generation question features by a lightweight classifier. Across $18$ (model, dataset) cells and four split seeds on three open-weight 4--9B LLMs, the learned classifier lifts predicted-matchable coverage from $0.61$ overall to $0.71$ on non-fallback rows ($0.82$ vs $0.63$ when oracle-fallback rows are included as upper bounds), and the classifier transfers across LLMs at AUC $0.78$ (in-sample $0.80$). The cell-level $\lambda^\star$ doubles as a regime diagnostic: large values flag epistemic dominance (retrieve, abstain) and $\lambda^\star\to 0$ flags aleatoric dominance (clarify), with a per-instance log-ratio $R(x)$ extending the read-out to individual queries.


Matching-while-Decoding: Enhancing Template-Free Retrosynthesis via Explicit Structural Alignment

Kaipeng Zeng ⋅ Jiarui Shen ⋅ Yi Jiang ⋅ Shiyue Wang ⋅ Lin Yao ⋅ Tong Zhu ⋅ Junchi Yan

Single-step retrosynthesis prediction serves as the core operational unit for automated synthesis planning and remains a fundamental computational challenge in pharmaceutical design. While a critical prior in organic chemistry is that chemical reactions typically preserve the majority of the molecular scaffold, current template-free deep learning models struggle to effectively exploit this characteristic. Existing methods attempt to exploit this by minimizing sequence-level edit distance through carefully chosen graph traversal orders. However, this structural correspondence is only indirectly induced and remains highly sensitive to imperfect traversal heuristics, which limits its effectiveness for complex transformations. To address this limitation, we introduce the Matching-while-Decoding paradigm, replacing implicit sequence alignment with direct structural supervision. We operationalize this paradigm within a graph-to-sequence architecture that explicitly maps generated reactant tokens back to the input product graph during decoding. Furthermore, our hierarchical autoregressive formulation utilizes a custom causal attention mask to enable highly efficient, single-pass parallel training. Extensive experiments across the USPTO-50K, USPTO-MIT, and USPTO-FULL benchmarks reveal a critical accuracy and robustness trade-off in existing methodologies, where models either over-rely on augmentation or fail to leverage it. Our approach uniquely resolves this bottleneck by delivering highly competitive canonical accuracy while scaling exceptionally well with test-time augmentation, establishing new state-of-the-art records on USPTO-MIT and augmented USPTO-FULL. By providing this balanced stability across diverse inference protocols, our approach emerges as a highly promising foundation for dependable single-step retrosynthesis prediction.


MaterialsPilot: An Execution-Feedback Framework for Generative Design of Complex Atomistic Architectures

Qiaolin Lu ⋅ Tongliang Liu ⋅ Qiang Qu ⋅ Bo Han ⋅ Yi Chang ⋅ Chengqi Zhang

Generating precise complex atomistic architectures from natural language specifications is a frontier in AI-driven materials discovery, critical for realizing functional systems like catalytic active sites and hetero-interfaces. While Large Language Models (LLMs) offer a powerful interface for such tasks, mapping abstract semantics to rigorous structural constraints remains unsolved as LLMs inherently face two deficits: first, the lack of quantitative physicochemical priors leads to topologically imprecise outputs; and second, open-loop generation without physical verification fails to resolve intermediate spatial conflicts, causing error propagation to invalidate coupled multi-step assemblies. To bridge this gap, we introduce MaterialsPilot, which formulates generation as iterative code optimization. It incorporates two key mechanisms: (1) Hierarchical Retrieval to bridge domain knowledge gaps by providing precise domain schemas; and (2) a Physics-Aware Feedback Loop to master complex logic by autonomously debugging execution and physical violations. Across four complex material modeling tasks and seven diverse LLM backbones, MaterialsPilot increases the rate of fully successful generations from 22.0% to 66.4%, while also producing structures that better satisfy user-specified constraints and physical plausibility checks. These results establish a model-agnostic framework for physically reliable language-driven atomistic design.


MC^2: Monte Carlo Correction for Fast Elliptic PDE Solving

Ethan Hsu ⋅ Ivan Ge ⋅ Hong M Yam

Partial differential equation (PDE) solvers underpin scientific computing, but real-world deployment is bounded by compute. Classical Monte Carlo solvers such as Walk-on-Spheres (WoS) are unbiased and geometry-agnostic but are slow. Learned solvers are fast but biased and brittle under distribution shift. We present \textbf{MC$^2$}, a hybrid WoS-Neural Network (WoS-NN) PDE solver that treats a low-budget Monte Carlo solution as a structured estimator of the true field and learns a single-pass neural correction to recover a high-fidelity solution. MC$^2$ matches the accuracy of solutions using over $1000\times$ more Monte Carlo compute, outperforming all evaluated classical, denoising, and neural-operator baselines. To enable reproducible study of finite-compute PDE solving, we additionally release \textbf{PDEZoo}, the largest standardized elliptic PDE benchmark to date: 2M PDEs spanning five elliptic families and unlimited geometric compositions, with analytic ground truth and multi-budget Monte Carlo trajectories. Together \textbf{MC$^2$} and \textbf{PDEZoo} (1) empirically establish that finite-sample Monte Carlo error is structured, learnable, and correctable in a single forward pass, (2) show that we can solve PDEs $\sim$\textbf{1000x} faster than with just WoS, and (3) provide the evaluation infrastructure the field has so far lacked.

Deep clustering struggles when coarse-grained and fine-grained semantics coexist in the same dataset, as a unified partitioning strategy tends to either over-separate coarse categories or miss subtle fine-grained differences. To address this problem, we propose the adaptive multi-granularity clustering framework in a hyperspherical space (MC-H). The framework combines Dynamic Prototype-Center Assignment (DPCA), a Hyperspherical Determinantal Point Process (H-DPP) module that adaptively regulates repulsion according to local structure and promotes stronger separation in dense regions while preserving flexibility in sparse ones, and a Semantic Subspace Denoising (SSD) module that suppresses non-semantic variations and improves cluster compactness. Extensive experiments across datasets with different semantic granularities show that our method consistently outperforms existing approaches and effectively unifies clustering across the full granularity spectrum.


Measuring AI Agents' Progress on Multi-Step Cyber Attack Scenarios

Linus Folkerts ⋅ Will Payne ⋅ Simon Inman ⋅ Philippos M Giavridis ⋅ Joe Skinner ⋅ Sam Deverett ⋅ James Aung ⋅ Ekin Zorer ⋅ Michael Schmatz ⋅ Mahmoud Ghanem ⋅ John Wilkinson ⋅ Alan Steer ⋅ The Vy Hong ⋅ Harry Coppock ⋅ Jessica Wang

We evaluate the autonomous offensive cyber capabilities of frontier AI agents on two purpose-built cyber ranges---a 32-step corporate network attack and a 7-step industrial control system attack targeting a simulated cooling tower---complemented by a private 82-task CTF suite spanning four difficulty tiers to measure narrow cyber skills. Estimated human-expert solve times are roughly 20 hours for the corporate range, 22 hours for the cooling tower, and from 1 minute to 40 hours across the CTF suite. Across twelve models released over a twenty-month period---from GPT-4o in August 2024 to early predeployment checkpoints of Mythos and GPT-5.5---we observe two trends. First, performance increases approximately log-linearly with inference-time compute, with no plateau observed for the strongest models up to 100M tokens. Second, newer frontier models generally complete more steps at fixed token budgets. GPT-5.5 and the Mythos Preview checkpoint each completed all 32 steps in at least one run, the first end-to-end completions of the corporate range; GPT-5.5 averaged 22.3 of 32 steps across ten 100M-token runs. On the control system range, the latest models reached 4 of 7 steps, primarily by directly probing the operational-technology protocol rather than following the intended chain of IT compromise and authenticated control. These results indicate rapid improvements in autonomous cyber capability and suggest that complex but undefended multi-step ranges are no longer reliably distinguishing the strongest frontier models. Future evaluations should emphasise hardened environments, active defences, defensive-control effectiveness, and operational security---priorities that grow more pressing as the threat of AI-driven cyber attacks becomes increasingly realistic.

Mode separation, namely how sharply a distribution fragments into barrier-separated clusters, is a fundamental geometric property of densities, difficult to quantify in high dimensions. It is structurally distinct from dispersion, yet existing tools fall short: differential entropy rises with spread regardless of fragmentation, PCA orders directions by variance regardless of barriers, and mutual information requires a mixture decomposition one usually does not have. We measure mode separation through a single stochastic process intrinsic to the density: a unique reversible diffusion with $f$ as its stationary distribution and constant scalar diffusion coefficient. We extract two readouts from its autocovariance matrix: SSA (Sum of Squared Autocorrelations), a scalar barrier-sensitive measure; and DA (Dominant Autocorrelation directions), linear projections ordered by metastability rather than variance. Under an isotropic-Gaussian null, we derive a closed-form spectrum for the empirical autocovariance that generalizes Marchenko--Pastur, with an analytic upper edge that selects the lag at which DA is read off. Both readouts use only samples and a score function, scaling to high dimensions through pretrained score-based generative models via Tweedie's identity. We apply our framework to three settings: (i) synthetic Gaussian mixtures, where SSA tracks mutual information; (ii) SDXL text-to-image generations, where SSA and DA capture structure that entropy and PCA miss; and (iii) molecular dynamics of alanine dipeptide, where DA recovers the known slow backbone dihedrals from static samples alone.


Measuring What Matters: Synthetic Benchmarks for Concept Bottleneck Models

Julian Skirzynski ⋅ Harry Cheon ⋅ Shreyas Kadekodi ⋅ Meredith Stewart ⋅ Berk Ustun

Concept bottleneck models predict outcomes from high-level concepts detected in inputs. Although concepts provide a simple way to reap benefits from interpretability, very few datasets include concept labels. This limits researchers' ability to determine which problems are suitable for these models, isolate the factors that drive their performance or lead to failures, or uncover which algorithms perform well. In this paper, we develop synthetic benchmarks for concept-bottleneck models, focusing on their two main use cases: decision support, in which models assist humans in making better decisions, and automation, in which models handle routine tasks without supervision. Our benchmarks can generate labeled datasets while controlling for properties that affect performance, including data modality, concept choice, annotation quality, and completeness. We demonstrate how the benchmarks can be used to evaluate representative classes of concept bottleneck models. Our demonstrations show how the benchmarks can diagnose failure modes and guide follow-up testing.


Mechanism-Aware Ensemble Conditioning for Data-Limited Emulation of Extreme Events

Isabella Thiel ⋅ Juan M. Bello-Rivas ⋅ Yannis Kevrekidis ⋅ Themis Sapsis

Extreme events in chaotic systems are difficult to learn from short trajectories because they are controlled by transient finite-time instability rather than by frequently observed bulk dynamics. We propose a mechanism-aware conditioning plug-in framework that turns a nudged coarse ensemble into a non-intrusive sensor of local instability geometry. In the small-noise regime, the ensemble covariance aggregates the same finite-time deformation kernels that govern local instability, providing a Jacobian-free proxy for the local amplification structure around a synchronized coarse trajectory. A small FiLM module injects statistics of this ensemble geometry into an otherwise unchanged backbone while leaving the coarse simulator unchanged. We demonstrate this interface in two distinct pipelines: a Transformer-style residual-attention corrector for a controlled low-dimensional chaotic system and a probabilistic recurrent STORN corrector for topographic two-layer quasi-geostrophic (QG) flow. In the low-dimensional benchmark, ensemble covariance directions co-activate with OTD modes and FiLM conditioning improves 99th-percentile exceedance-frequency errors over an identical no-context Transformer baseline. In QG, a fixed ensemble-conditioned FiLM-STORN model trained on only 50 time units substantially improves long-horizon rare-event statistics in the data-limited regime, including density-tail errors, exceedance frequencies, and spatial exceedance-area distributions relative to an unconditioned STORN trained on the same data; on averaged high-threshold exceedance diagnostics, it also outperforms the baseline STORN trained with 20 times more high-resolution data. These results show that local instability geometry is not merely interpretable post hoc, but an actionable conditioning signal for data-efficient rare-event emulation.


Mechanistic Interpretability Needs Philosophy

Iwan Williams ⋅ Ninell Oldenburg ⋅ Ruchira Dhar ⋅ Joshua Hatherley ⋅ Constanza Fierro ⋅ Nina Rajcic ⋅ Sandrine R Schiller ⋅ Filippos Stamatiou ⋅ Anders Søgaard

Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important to examine not just models themselves, but the assumptions, concepts and explanatory strategies implicit in MI research. We argue that mechanistic interpretability needs philosophy as an ongoing partner in clarifying its concepts, refining its methods, and navigating the epistemic and ethical complexities of interpreting AI systems. There is significant unrealised potential for progress in MI to be gained through deeper engagement with philosophers and philosophical frameworks. Taking three open problems from the MI literature as examples, this paper illustrates the value philosophy can add to MI research, and outlines a path toward deeper interdisciplinary dialogue.

Doctor agents are moving beyond single-turn answer generation toward evolving clinical decision systems. Within an outpatient episode, they acquire evidence, use examination and consultation resources, and decide when to finalize a diagnosis and management plan. Across episodes, their behavior may change through memory, retrieval, reflection, or other update mechanisms. Current evaluations only partially cover this setting. Fixed-input medical QA benchmarks score final answers from complete inputs, whereas many interactive benchmarks still focus on individual encounters or fixed runs, providing limited support for evaluating how episode-level decisions interact with cross-episode experience. We introduce MedEvoEval, an executable longitudinal evaluation framework based on action-gated simulated outpatient episodes. Each source case is converted into role-specific patient, examination, and manager views; evidence is revealed only through valid actions; and each episode records a structured trace that links observations, actions, final outputs, manager scores, and optional experience write-back. We release a runnable E&amp;D artifact with 700 processed episodes, provenance notes, schemas, an episode runner, scoring scripts, configurations, example logs, analysis code, and trajectory- and step-level derivatives. Experiments show that episode traces expose process costs hidden by final-answer scoring, show how MDT-style consultation reallocates resources, and support longitudinal analyses of memory maturation, held-out transfer, update-stage response, and backward retention. Together, these results show that MedEvoEval provides a concrete basis for evaluating whether doctor agents improve through experience, transfer useful behavior, and retain earlier capabilities over time.


MedHorizon: Towards Long-context Medical Video Understanding in the Wild

Bodong Du ⋅ Bowen Liu ⋅ Yang YU ⋅ Xinpeng Ding ⋅ Zhiheng Wu ⋅ Shuning Wang ⋅ Shuo Nie ⋅ Naiming Liu ⋅ Qifeng Chen ⋅ Yangqiu Song ⋅ Xiaomeng Li

Medical multimodal large language models (MLLMs) have advanced image understanding and short-video analysis, but real clinical review often requires full-procedure video understanding. Unlike general long videos, medical procedures contain highly redundant anatomical views, while decisive evidence is temporally sparse, spatially subtle, and context dependent. Existing benchmarks often assume this evidence has already been localized through images, short clips, or pre-segmented videos, leaving the retrieval-before-reasoning problem under-tested. We introduce MedHorizon, an in-the-wild benchmark for long-context medical video understanding. MedHorizon preserves 759 hours of full-length clinical procedures and provides 1,253 evidence-grounded multiple-choice questions that jointly evaluate sparse evidence understanding and multi-hop clinical reasoning. Its evidence is extremely sparse, with only 0.166% evidence frames on average, requiring models to search noisy procedural streams before interpreting and aggregating findings. We evaluate representative general-domain, medical-domain, and long-video MLLMs. The best model reaches only 41.1% accuracy, showing that current systems remain far from robust full-procedure understanding. Further analysis yields four key findings: performance does not scale reliably with more frames, evidence retrieval and clinical interpretation remain primary bottlenecks; these bottlenecks are rooted in weak procedural reasoning and attention drift under redundancy, and generic sampling methods only partially balances local detail with global coverage. MedHorizon provides a rigorous testbed for MLLMs that retrieve sparse evidence and reason over complete clinical workflows.


MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare

Yihao Wang ⋅ Haoran Xu ⋅ Renjie Gu ⋅ Yixuan Ye ⋅ Xinyi Chen ⋅ Xinyu Mu ⋅ Yuan Gao ⋅ ChunxiaoGuo ⋅ Peng Wei ⋅ Jinjie GU ⋅ Huan Li ⋅ Ke Chen ⋅ Lidan Shou

The large-scale deployment of personalized healthcare agents demands memory mechanisms that are exceptionally precise, safe, and capable of long-term clinical tracking. However, existing benchmarks primarily focus on daily open-domain conversations, failing to capture the high-stakes complexity of real-world medical applications. Motivated by the stringent production requirements of an industry-leading health management agent serving tens of millions of active users, we introduce MedMemoryBench. We develop a human-agent collaborative pipeline to synthesize highly realistic, long-horizon medical trajectories based on clinically grounded, synthetic patient archetypes. This process yields a massive, expertly validated dataset comprising approximately 2,000 sessions and 16,000 interaction turns. Crucially, MedMemoryBench departs from traditional static evaluations by pioneering an evaluate-while-constructing streaming assessment protocol, which precisely mirrors dynamic memory accumulation in production environments. Furthermore, we formalize and systematically investigate the critical phenomenon of memory saturation, where sustained information influx actively degrades retrieval and reasoning robustness. Comprehensive benchmarking reveals severe bottlenecks in mainstream architectures, particularly concerning complex medical reasoning and noise resilience. By exposing these fundamental flaws, MedMemoryBench establishes a vital foundation for developing robust, production-ready medical agents. Our dataset and benchmark are publicly available at: .


MedVTok: A General-Purpose Medical Visual Tokenizer

Chenglong Ma ⋅ Yuanfeng Ji ⋅ Junzhi Ning ⋅ Jiyao Liu ⋅ Yujie Wei ⋅ Shaohao Rui ⋅ Chaoyang Zhang ⋅ Wenjie Li ⋅ Wenhao Tang ⋅ Cheng Tang ⋅ Wei Li ⋅ Shujian Gao ⋅ Yujie Zhang ⋅ Ying Chen ⋅ Zilong Li ⋅ Zongyuan Ge ⋅ Diping Song ⋅ Jin Ye ⋅ Ming Hu ⋅ Junjun He ⋅ Hongming Shan

Multimodal medical AI requires a common visual interface that integrates heterogeneous imaging evidence for patient condition modeling. However, existing medical visual tokenizers are often tied to a specific dimension and imaging modality, forcing multimodal systems to rely on fragmented and poorly reusable representations. We introduce MedVTok, a general-purpose visual tokenizer that maps 2D images and 3D volumes from diverse imaging modalities into a unified token space. Building such a tokenizer poses two key challenges: preserving anatomical structure across dimensions and capturing clinical semantics across modalities. To this end, MedVTok introduces (i) slice differential regularization, which explicitly models inter-slice anatomical coherence missing in previous per-pixel/per-voxel supervision, and (ii) multi-expert representation alignment, which integrates knowledge from multiple medical expert encoders while retaining modality-specific complementary knowledge. Trained on 28M images and 70K volumes across 9 imaging modalities, MedVTok supports classification, retrieval, segmentation, synthesis, visual question answering, and medical report generation through a single tokenizer. Across 24 benchmarks, MedVTok achieves state-of-the-art performance, demonstrating a scalable token interface that unifies medical image dimension, modality, and task usage. Code and weights will be released.


MemCode: Discrete Semantic Representations for Long-Term Agent Memory

Xin Li ⋅ Liang Hu ⋅ Duoqian Miao ⋅ Ming Peng ⋅ Qi Zhang

Long-term memory remains a fundamental bottleneck for large language model (LLM) agents across sessions. Existing systems typically adopt a write-then-retrieve paradigm, storing context-dependent dialogue fragments and retrieving them via nearest-neighbor search in continuous embedding space. This design yields unstable access units and systematically misses logically related but lexically distant evidence. We propose MemCode, a long-term agent memory framework that couples retrieval-oriented writing with discrete associative retrieval. MemCode rewrites dialogue into self-contained atomic memory units (AMUs), maps AMUs into a learnable codebook with isomorphic semantic quantization (ISQ), and uses codebook co-occurrence topology as a retrieval path complementary to vector search. For long-term memory stores, MemCode separates online atomic writes from asynchronous offline consolidation. On LongMemEval-S and LoCoMo with GPT-4o-mini and Qwen3-30B-A3B-Instruct-2507 backbones, MemCode achieves the best overall accuracy while keeping per-dialogue memory construction within roughly 40--130 seconds, making it 3--21$\times$ faster than LightMem and over an order of magnitude faster than A-MEM and MemoryOS.


Memory Inception: Latent-Space KV Cache Manipulation for Steering LLMs

Zeyi (Andy) Liu ⋅ Michael Zhang ⋅ Ilana Greenberg ⋅ Adam Alnasser ⋅ Lucas Baker ⋅ John Sous

Steering large language models (LLMs) is usually performed through either instruction prompting or activation steering. Prompting offers strong control, but repeatedly caches guidance tokens and can clutter long interactions. In contrast, activation steering ,while compact, typically offers weaker control and less fine-grained steering because it does not support large structured reminders. We introduce memory inception (MI), a training-free method that steers in latent attention space by inserting text-derived key-value (KV) banks only at selected layers. MI treats steering as selective KV allocation, where reminder content need not occupy the full-layer prompt cache in order to induce implicit behavioral guidance. We formulate MI as attention over prompt, target, reference, and auxiliary banks, with canonical pre-RoPE key storage for reuse across positions and architectures. On matched personality-steering tasks, MI offers a balance between prompting and contrastive activation addition (CAA), approaching or exceeding prompting in raw control strength and CAA in drift mitigation. MI also transfers to structured heuristic guidance on physical reasoning domains, such as the PHYSICS benchmark, where it outperforms visible prompting on average, wins 10/12 subject times mode cells, and cuts content-matched KV storage by up to 118 times. These results position MI as a powerful steering method when guidance is persistent, structured, or expensive to keep in the visible transcript.


Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies

Yi Wen ⋅ Derong Xu ⋅ Pengyue Jia ⋅ Yichao Wang ⋅ Yingyi Zhang ⋅ Maolin Wang ⋅ Junyi Li ⋅ Wenlin Zhang ⋅ Xiaopeng Li ⋅ Yong Liu ⋅ Xiangyu Zhao

The memory capabilities of Large Language Models (LLMs) have garnered increasing attention recently. Despite great success achieved, existing retrieval-based memory approaches typically overlook the differences between memories and employ a unified strategy to process all memories, leading to suboptimal performance. Thus, an intuitive question arises: can we categorize memory into different types and select appropriate strategies? However, given the topic-rich, scenario-complex, and boundary-blurred nature of memory scenarios, achieving precise classification of memories is not easy. To address this challenge, we propose a memory multi-class dataset in this paper, termed TriMEM, which provides precise annotations for memory types across diverse scenarios. Building upon this foundation, we propose a novel memory framework, named MemoType, which can adaptively recognize each memory and query type with the learned router model. With the memory and query routing, MemoType can retrieve the memory with corresponding query types and design tailored retrieval strategies, thereby enhancing the retrieval performance. Moreover, we theoretically prove that any single retrieval strategy is subject to a fundamental upper bound on its expected retrieval precision in multi-class corpora, leading to systematic precision degradation. Extensive experiments on three datasets demonstrate that MemoType consistently outperforms existing methods, achieving up to 16.18% improvement in Recall@1.


MemReward: Graph-Based Experience Memory for LLM Reward Prediction with Limited Labels

Tianyang Luo ⋅ Tao Feng ⋅ Zhigang Hua ⋅ Yan Xie ⋅ Shuang Yang ⋅ Ge Liu ⋅ Jiaxuan You

Reinforcement learning has emerged as a powerful paradigm for improving large language model (LLM) reasoning, where rollouts are sampled from the policy and reward signals computed on those rollouts are used to update the policy. However, in data-scarce scenarios, obtaining ground-truth labels to verify rollouts at scale often requires expensive human annotation or labor-intensive expert verification. For instance, evaluating mathematical proofs demands expert review, and open-ended question answering lacks definitive ground truth. When ground-truth labels are scarce, the effectiveness of reinforcement learning fine-tuning is constrained. Inspired by the success of semi-supervised learning in propagating labels from labeled to unlabeled samples, we propose MemReward, a graph-based experience memory framework that integrates reward propagation directly into online policy optimization. MemReward stores rollouts (thinking processes and final answers) from an initial LLM policy as nodes in a heterogeneous graph connected by similarity and structural edges, over which a GNN propagates rewards from labeled to unlabeled rollouts. To train such a framework, we first warm up the GNN on labeled rollouts to predict rewards via heterogeneous aggregation over query, thinking, and answer nodes. During online RL fine-tuning, unlabeled rollouts are attached to the graph by query similarity, and the GNN predicts their rewards, yielding a hybrid reward acquisition strategy that combines ground-truth and GNN-predicted rewards. Experiments on Qwen2.5-1.5B and 3B in mathematics, question answering, and code generation demonstrate that MemReward, with ground-truth rewards on only 20\% of rollouts, achieves 96.6\% of Oracle performance on 1.5B and 97.3\% on 3B, and closely approaches Oracle on out-of-domain tasks.


MemRL: Self-Evolving Agents via Runtime Reinforcement Learning on Episodic Memory

Jiaqian Wang ⋅ Shengtao Zhang ⋅ Ruiwen Zhou ⋅ Junwei Liao ⋅ Yuchen Feng ⋅ Zhuo Li ⋅ Yujie Zheng ⋅ Weinan Zhang ⋅ Ying Wen ⋅ Zhiyu li ⋅ Feiyu Xiong ⋅ Yutao Qi ⋅ Bo Tang ⋅ Muning Wen

The hallmark of human intelligence is the self-evolving ability to master new skills by learning from past experiences. However, current AI agents struggle to emulate this self-evolution: fine-tuning is computationally expensive and can introduce destabilizing weight updates, increasing the risk of catastrophic forgetting, while existing memory-based methods rely on passive semantic matching that often retrieves noise. To address these challenges, we propose MemRL, a non-parametric approach that evolves via reinforcement learning on episodic memory. By decoupling stable reasoning from plastic memory, MemRL employs a Two-Phase Retrieval mechanism to filter noise and identify high-utility strategies through environmental feedback. Extensive experiments on Humanity's Last Exam, BigCodeBench, ALFWorld, and Lifelong Agent Bench demonstrate that MemRL significantly outperforms state-of-the-art baselines, suggesting that MemRL offers a practical way to balance stability and plasticity during runtime improvement.


MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters

Yuming Yang ⋅ Xiao Sun ⋅ Yuanwei Zou ⋅ Zhengxiao Wu ⋅ Jiang Zhong ⋅ Haoyang Zeng ⋅ Jingwang Huang ⋅ Kaiwen Wei

Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment planning, yet existing benchmarks rarely simulate complete psychiatric clinical encounters. We introduce MentalHospital, a virtual evaluation environment for LLM-based psychiatric clinical encounters. MentalHospital instantiates the Subjective Interviewing, Objective Examination, Diagnostic Assessment, and Treatment Planning (S.O.A.P.) workflow, using skill-augmented standardized patients constructed from 1,193 de-identified psychiatric electronic health record (EHR) cases spanning all major ICD-11 categories and 76 disorders. Each encounter is assessed through a dual-track protocol that combines objective comparison against EHR-derived references with subjective assessment of clinical process quality. To scale specialist judgment, we develop MentalEval, five domain-specific evaluators covering communication empathy, interviewing professionalism, clinical-note quality, diagnostic rigor, and treatment appropriateness, trained with rubric-grounded SFT and expert-guided DPO. Survey responses from 22 clinicians support MentalHospital's clinical fidelity (3.88/5), while MentalEval achieves strong expert alignment with an average QWK of 0.944. Benchmarking shows that even the strongest LLM trails clinicians by 37.28 percentage points in objective psychiatric competence, with mental status assessment as a key bottleneck.


METRO: Metric-Enhanced Token Routing Operator

Nodens Koren ⋅ Thomas Hofmann ⋅ Georgios Kissas

State-of-the-art neural operators scale to complex meshes via slice-and-process architectures, yet many rely on linear compatibility scores for latent tokenization. Under common feature normalization, such scores are equivalent to isotropic Euclidean clustering, while without normalization they induce unbounded linear decision regions. In both cases, they lack slice-specific anisotropic locality, which can lead to redundant and entangled latent slices. To address this, we propose Metric-Enhanced Token Routing Operator (METRO), a geometry-aware routing mechanism that replaces linear projection with a learnable Mahalanobis metric. By enabling each latent slice to learn a local anisotropic tensor, METRO shapes receptive fields into exponentially localized, oriented ellipsoids that naturally align with flow features like boundary layers and wakes. As a drop-in replacement, METRO yields consistent improvements across both Transformer and Mamba backbones. Empirically, our method achieves substantial performance gains on irregular domains, outperforming baselines on both standard PDE benchmarks and complex industrial design tasks. Finally, METRO exhibits enhanced robustness in out-of-distribution regimes across varying Reynolds numbers and geometric configurations.


Metropolis-Adjusted Diffusion Models

Kevin H. Lam ⋅ Tyler Farghly ⋅ Christopher Williams ⋅ Jun Yang ⋅ Yee Whye Teh ⋅ Arnaud Doucet

Sampling from score-based diffusion models incurs bias due to both time discretisation and the approximation of the score function. A common strategy for reducing this bias is to apply corrector steps based on the unadjusted Langevin algorithm (ULA) at each noise level within a predictor-corrector framework. However, ULA is itself a \emph{biased} sampler, as it discretises a continuous diffusion process. In this work, we consider \emph{adjusted} Langevin correctors that employ Metropolis--Hastings (MH) or Barker's accept-reject steps to correct for this bias. Since the target density ratio typically required by MH-based algorithms is unavailable, we propose methods that instead utilise the score function to compute the correct acceptance probability. We introduce the first exact method for adjusting Langevin corrections in diffusion models, based on a two-coin Bernoulli factory algorithm. We also propose an efficient approximation based on Simpson's rule that achieves accuracy of order $5/2$ in the step size at near-zero marginal cost. We demonstrate that these procedures improve sample quality on both synthetic and image datasets, yielding consistent gains in Fr\'echet Inception Distance on the latter.


MICA: Activation Checkpointing for Double-Backward Training of Machine Learning Interatomic Potentials

Hongyu Wang ⋅ Mingzhen Li ⋅ Hongtao Xu ⋅ Yuanchang Zhou ⋅ Weijian Liu ⋅ Weile Jia ⋅ Guangming Tan

Conservative machine learning interatomic potentials (MLIPs) predict potential energy and obtain atomic forces as the negative gradient of energy with respect to atomic positions. Computing forces requires a backward pass from energy to atomic positions, and backpropagating the force loss differentiates through this pass, creating double-backward execution. This execution creates two activation classes with different lifetimes: forward activations and first-backward activations. Heterogeneous atomic graph batches make their memory footprint input dependent, while standard checkpointing mainly targets forward activations and does not directly manage first-backward activations. Therefore, we present MICA, an activation checkpointing planner for double-backward execution in conservative MLIP training. MICA combines a checkpoint primitive for double-backward phases with an input-aware dynamic programming planner that selects checkpoint modes for dynamic atomic graph batches under a specified memory budget. Across EqV3, MatRIS, PET, and UMA, MICA's full-checkpoint mode reduces peak memory by 49\% to 86\%, compared with only 10\% to 24\% from standard PyTorch checkpointing. On an 80GB GPU, training EquiformerV3 without checkpointing peaks at 75.47GB and exceeds 50GB in 76.7\% of 1000 iterations, while MICA keeps all 1000 iterations within a 50GB memory budget with 16.25\% step time overhead. Under the same GPU memory capacity, MICA increases the maximum feasible atoms per batch by $4.2\times$ to $9.2\times$. In distributed training, the reduced peak memory allows MICA to reduce reliance on heavier memory saving techniques, such as Fully Sharded Data Parallel (FSDP) and graph parallelism (GP). On a single H100 node with NVLink, MICA trains a MatRIS-MoE model with 2B parameters, achieving a $1.88\times$ to $2.03\times$ speedup over FSDP and GP baselines.


Microstructure Descriptor Fields as Supervision for Scientific Images

Mohamed Wahib ⋅ Du Wu ⋅ Cong Ma ⋅ Isaac Lyngaas ⋅ Xiao Wang ⋅ Enzhi Zhang ⋅ Peng Chen ⋅ Satoshi Matsuoka

Scientific images are often governed less by exact pixel fidelity than by preservation of microstructure: morphology, connectivity, anisotropy, interface complexity, phase organization, and characteristic length scales. Yet modern image-learning objectives, including masked image modeling and reconstruction losses, supervise models in pixel, token, or latent feature space rather than in the space of scientific observables. We introduce MICROFIELD, a supervision framework that represents a scientific image by a field of microstructure descriptors computed over spatial regions. A descriptor field can be induced on a uniform grid, a quadtree, an octree, a learned partition, or a hierarchy produced by a high-resolution image model. Rather than proposing a new masked autoencoder architecture, MICROFIELD contributes an objective-level mechanism: it trains models to preserve local descriptors, cross-scale descriptor transitions, and level-wise descriptor distributions. Because descriptor fields are defined over region systems rather than model architectures, MICROFIELD applies to uniform grids, adaptive partitions, and hierarchical models alike. The framework is compatible with self-supervised pretraining, supervised learning, and reconstruction; in this paper, we focus on masked self-supervised pretraining as the primary instantiation. MICROFIELD provides a practical path toward physics-inspired representation learning for scientific images without requiring explicit governing equations or domain labels.


MinerU2.5-Pro: Pushing the Limit of Document Parsing via a Calibrated Evaluation-Data Flywheel

Bin Wang ⋅ Tianyao He ⋅ Linke Ouyang ⋅ Fan Wu ⋅ Zhiyuan Zhao ⋅ Tao Chu ⋅ Yuan Qu ⋅ Zhenjiang Jin ⋅ Weijun Zeng ⋅ Ziyang Miao ⋅ Bangrui Xu ⋅ Junbo Niu ⋅ Wei Li ⋅ Mengzhang Cai ⋅ Qiu Jiantao ⋅ Qintong Zhang ⋅ Dongsheng Ma ⋅ Yuefeng Sun ⋅ Hejun Dong ⋅ Wenzheng Zhang ⋅ Jutao Xiao ⋅ shi J yong ⋅ Pengyu Liao ⋅ Xiaomeng Zhao ⋅ Huaping Zhong ⋅ Zhengren Wang ⋅ Liqun Wei ⋅ Jing Yu ⋅ Jie Yang ⋅ Shasha Wang ⋅ Qianqian Wu ⋅ Xuanhe Zhou ⋅ Weijia Li ⋅ Zhenxiang Li ⋅ Zhongying Tu ⋅ Jiang Wu ⋅ Lijun Wu ⋅ Chao Xu ⋅ Kai Chen ⋅ Wentao Zhang ⋅ Yu Qiao ⋅ Bowen Zhou ⋅ Dahua Lin ⋅ Conghui He

Document parsing is now critical infrastructure for LLM systems, yet standard benchmarks suggest near saturation: on OmniDocBench v1.5, leading methods cluster around a 95-point ceiling. We argue that this reflects a stalled evaluation–training flywheel, not solved parsing: v1.5 under-samples hard, long-tail documents and uses fixed-granularity matching that can penalize semantically correct outputs. We re-engage the flywheel with two co-designed protocols. OmniDocBench v1.6 repairs evaluation with MGAM (Multi-Granularity Adaptive Matching) and a strictly held-out 296-page Hard Subset, with Base / Hard / Full reporting. The corrected benchmark exposes two data gaps—hard cases are under-represented and unreliably annotated—which the MinerU2.5-Pro Data Engine addresses through DDAS (Difficulty-Driven Adaptive Sampling) and Judge-and-Refine, a render–verify–refine annotation pipeline, yielding 65.5M auto-labeled and 192K human-annotated samples. Conversely, the same discovery and annotation mechanisms provide a scalable path for future Hard-subset refreshes. Under a fixed 1.2B MinerU2.5 backbone, these interventions raise OmniDocBench v1.6 Full performance from 92.98 to 95.69, with the clearest separation on the corrected Hard split. The result reframes apparent saturation as a repairable benchmark–data loop: document parsing progress must be easier to measure correctly, not just easier to claim.


Minimax Optimal Estimation of Transport-Growth Pairs in Unbalanced Optimal Transport

Donlapark Ponnoprat ⋅ Noboru Isobe ⋅ Masaaki Imaizumi

Unbalanced optimal transport extends classical optimal transport to measures with different total masses, but statistical guarantees for Monge-type estimation remain limited. We study unbalanced transport with quadratic cost and Kullback-Leibler marginal penalties and argue that the natural population target is not a map alone, but a transport-growth pair. Consequently, we develop two estimators for the transport-growth pairs under several setups: an optimal transport plan-based estimator for a general case, and a kernel-based estimator for a case with smooth densities. We also show that an error of the estimator achieves the minimax optimal rate by deriving a matching lower bound of the minimax risk. Our main technical contribution is a value-based stability reduction that converts perturbations of the UOT objective into transport and growth risks through a UOT gap condition. These results provide a statistical foundation for Monge-type estimation in unbalanced optimal transport.

Using neural networks to learn optimal transport (OT) maps via minimax optimization has gained increasing interest in generative modeling. However, the widely used semi-dual OT maximin formulation is known to suffer from spurious global solutions that do not correspond to valid transport maps, often addressed by regularizing the objective function. We provide an alternative explanation by showing that these spurious solutions are not stationary points of the associated loss, implying that their appearance in practice stems from the optimizer's failure to converge to stationary points in nonconvex-nonconcave settings. However, when the $c$-concavity constraints are imposed on the potential, commonly to stabilize training, we show that these spurious solutions become stationary. Based on these observations, we propose solving the semi-dual OT problem without $c$-concavity constraints, using optimizers that can reliably converge to stationary points corresponding to OT maps in nonconvex-nonconcave settings, such as a two-timescale extragradient (TTS-EG) method. At the formulation level, we further show that a Lagrangian-based OT minimax formulation provably eliminates spurious global solutions. Our empirical results demonstrate that TTS-EG reliably recovers OT maps under both semi-dual and Lagrangian-based OT formulations without any regularization.


MIRAGE: Duality-Inspired MILP Augmentation for Representation Learning

Ziao Guo ⋅ Shiyue Wang ⋅ Jiayuan Yang ⋅ Junchi Yan

Learning high-quality representations of mixed-integer linear programming (MILP) is crucial for advancing machine learning (ML)-based methods, yet remains challenging due to the discrete nature of MILPs and their extreme sensitivity to small perturbations. Although contrastive learning has proven highly effective for representation learning in domains such as computer vision, designing effective data augmentations for MILPs is particularly difficult. To address this challenge, we propose MIRAGE, a principled data augmentation framework grounded in the invariance of the piecewise-affine dual price function induced by the branch-and-bound (B&B) algorithm. Viewing this dual function as a recorded, solver-level surrogate of the MILP's lower-bound structure, we derive two complementary augmentation certificates that act as its right-hand and objective-side projections. Both certificates are provably invariant under the same dual price function on the recorded B&B tree, and therefore preserve either the structural similarity of the search tree or the optimality of the original optimal solution. Empirically, we integrate MIRAGE with ML models for two significant downstream tasks—learning to branch and predict-and-search—and observe consistent performance improvements across multiple benchmarks. By grounding MILP data augmentation in a single, solver-induced duality quantity, our work provides a principled and effective approach to MILP representation learning.


MiroEval: Benchmarking Multimodal Deep Research Agents in Process and Outcome

Fangda Ye ⋅ Yuxin Hu ⋅ Pengxiang Zhu ⋅ Yibo Li ⋅ Ziqi Jin ⋅ Yao Xiao ⋅ Yibo Wang ⋅ Lei Wang ⋅ Zhen Zhang ⋅ Lu Wang ⋅ Yue Deng ⋅ Bin Wang ⋅ Yifan Zhang ⋅ Liangcai Su ⋅ Xinyu Wang ⋅ He ZHAO ⋅ CHEN WEI ⋅ Qiang Ren ⋅ Bryan Hooi ⋅ Bo An ⋅ Shuicheng Yan ⋅ Lidong Bing

Recent progress in deep research systems has been impressive, yet their evaluation remains limited not only by benchmark coverage but also by methodology. Existing benchmarks predominantly assess final reports using fixed rubrics, failing to evaluate the underlying research process. Most also offer limited multimodal coverage, rely on synthetic tasks that do not reflect real-world query complexity, and cannot be refreshed as knowledge evolves. To address these gaps, we introduce MiroEval, an evaluation framework for deep research systems, accompanied by a 100-task benchmark (70 text-only, 30 multimodal) constructed via a dual-path pipeline that supports periodic updates, enabling a live and evolving setting. The proposed evaluation suite assesses deep research systems along three complementary dimensions: adaptive synthesis quality evaluation with task-specific rubrics, agentic factuality evaluation via active retrieval and reasoning over both web sources and multimodal attachments, and process-centric evaluation audits how the system searches, reasons, and refines throughout its investigation. Evaluation across 11 systems yields three principal findings: the three evaluation dimensions capture complementary aspects of system capability, with each revealing distinct strengths and weaknesses across systems; process quality serves as a reliable predictor of overall outcome while revealing weaknesses invisible to output-level metrics; and multimodal tasks pose substantially greater challenges, with declines concentrated in synthesis and process rather than factuality. Robustness experiments, a human ranking study, and three-annotator verification confirm the reliability of both the evaluation framework and the benchmark. MiroEval provides a holistic diagnostic tool for the next generation of deep research agents.


Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?

Filippo Ziliotto ⋅ Ciro Beneduce ⋅ Luciano Serafini ⋅ Bruno Lepri ⋅ Massimiliano Luca ⋅ Tommaso Campari

In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogous functional capability emerges in embodied vision-language model (VLM) agents: can they recognize themselves in a mirror? We introduce a controlled 3D benchmark where a first-person VLM agent must infer a hidden body attribute from its reflection and select the matching target, while avoiding self–other misattribution. To separate mirror-grounded self-identification from shortcuts, we test mirror removal, misleading cues, and occluded reflections. We also evaluate the decision process through mirror seeking, temporal ordering, self-attribution, and reasoning–action consistency. Our experiments show that mirror-based self-identification emerges mainly in stronger VLMs. These models can use reflected evidence for action, whereas weaker models often inspect the mirror but fail to extract self-relevant information or misattribute their reflection. Language–vision conflict further shows that self-referential language alone is not evidence of grounded self-identification. Overall, mirror-based evaluation provides a diagnostic for whether embodied self-grounding is causally rooted in perception and action rather than priors, prompt compliance, or confabulation.

Flow matching has emerged as a powerful alternative to diffusion models for generative modeling, but standard conditional flow matching is fundamentally misaligned with partially observed time series: starting from an uninformed Gaussian source forces the model to transport even already-observed entries from noise to the target, wasting capacity on trivially known regions. We propose MissPath-FM, a partial-observation-aware flow matching framework that redesigns the source prior and probability path for incomplete time series. MissPath-FM employs structured source priors (interpolation priors and Gaussian process posterior priors) to initialize the flow close to the observed signal, paired with a missingness-aware smooth path that allocates transport asymmetrically between observed and missing regions. We evaluate systematically across four datasets, four missingness mechanisms (MCAR, block, MNAR, irregular), two missing ratios, and three sequence lengths ($T \in \{24, 48, 96\}$), totaling 48 settings against eight baselines. MissPath-FM wins 29 of 32 settings at $T{=}48$ and 45 of 48 overall (93.8\%). Our ablations reveal a regime-dependent design hierarchy in which no single component suffices: under random missingness the source prior yields a large initial gain, but under structured block missingness it can fail entirely and the learned flow becomes indispensable; the missingness-aware path contributes a further $12\%$--$26\%$ across all regimes. Only the full pipeline achieves robust performance across all 48 settings. Mechanistically, structured priors shorten observed-region transport by up to $15\times$, concentrating the velocity field on genuinely uncertain entries.


Mitigating Reward Hacking via Task Representations

Lillian Sun ⋅ Joe Benton ⋅ Eric Easley

Language models fine-tuned with reinforcement learning frequently learn to exploit their reward source rather than solve the underlying task. Existing fixes are typically reactive: detect a specific exploit, then patch the reward. We instead ask whether reward hacking can be prevented by constraining how the model represents the task. We introduce prompt KL regularization, a single auxiliary loss that adds a KL penalty between the actor's and reference model's per-token distributions on the prompt only, leaving the response distribution free. Across three reward hacking tasks of different sizes, prompt KL regularization maintains low hack rates (under 3\%) when standard training yields up to 100\% hack rates. Additionally, it matches or improves ground-truth accuracy, with no changes to the reward or environment. We aim to understand why prompt KL regularization is effective. Two independent mechanistic studies suggest that changes in prompt representations are key to models learning reward hacking behavior. First, we show that prompt-space activation steering Pareto-dominates response-space steering on hack rate and accuracy. Then, we demonstrate that swapping prompt activations between models can transfer and reduce reward hacking. Together, these results suggest that the way models internally represent the task is central to reward hacking, and that this representation can be directly constrained to prevent it.


MixRoute: Rethinking Single-Distribution Training for Generalizable Neural Routing

Hang Yi ⋅ Ziwei Huang ⋅ Yining Ma ⋅ Zhiguang Cao

Neural Combinatorial Optimization (NCO) is a promising paradigm for solving routing problems, yet its generalization across diverse data distributions remains a critical bottleneck. Existing methods typically tackle cross-distribution generalization through complex multi-distribution training or intricate meta-learning schemes, implicitly presuming that single-distribution training yields a weak zero-shot baseline. In this paper, we demonstrate that the potential of single-distribution training has been substantially underestimated and can be unlocked by a simple architectural inductive bias: Mix Normalization. To approach this, we first construct a comprehensive benchmark spanning three representative routing problems and 179 fine-grained datasets, capturing diverse shifts in node coordinates, customer demands, and time windows. Building on this benchmark, we propose MixRoute, a neural framework that adaptively combines normalization statistics at different granularities to mitigate varying distribution shifts. Trained exclusively on uniform distributions, MixRoute generalizes effectively to a wide range of unseen distributions without any distribution-specific adaptation. Extensive experiments demonstrate that MixRoute consistently achieves state-of-the-art zero-shot generalization. Our work revisits normalization as a parsimonious alternative for zero-shot generalization against complex training-adaptation schemes.


Mixture-Trained Merging for Unified Multi-Objective Models

SeongHyeon Kim ⋅ Chaeyun Jang ⋅ Seungyoo Lee ⋅ Jiyeon Ham ⋅ Yunju Bak ⋅ Boseop Kim ⋅ Juho Lee

Unified language models are increasingly expected to combine heterogeneous capabilities, such as mathematics, code, instruction following, and controllable thinking behavior, within a single set of parameters. A common solution is sequential post-training on multiple objectives, but this entangles all objectives along one optimization trajectory and makes the final model highly sensitive to training order, data ratios, schedules, and stopping criteria. Weight-space merging offers a modular alternative, but naive merging of single-objective experts often fails: domain capabilities degrade sharply, or think/non-think modes collapse into one dominant behavior. We attribute both failures to incompatible weight-space geometry: experts trained on isolated objectives drift to distant regions of parameter space, placing their interpolations outside any shared low-loss basin. We propose Mixture-Trained Merging (MTM), which trains each branch on an objective-biased data mixture rather than a single objective, exposing it to cross-objective interactions and making branches compatible at merge time. MTM uses merged-model evaluations as a low-cost signal for selecting branch mixtures, avoiding expensive data-mixture ablations. The procedure is iterative: each round promotes the base model using globally selected merge coefficients and refines each branch mixture using domain-preferred coefficients under constraints that preserve other objectives. To scale beyond simplex grid search, MTM uses qNEHVI-based multi-objective Bayesian optimization. Across code, mathematics, instruction following, and think/non-think control, MTM outperforms naive merging and preserves behavioral separation where single-objective merging collapses, suggesting that effective unified models require branches trained to be mergeable.


MJ: Multi-turn LLM Jailbreaking via Decomposed Credit Assignment

Junyoung Park ⋅ Namgyu Park ⋅ Sechan Lee ⋅ Yoon-Chan Jhi ⋅ Jihoon Cho ⋅ Sangdon Park

Modern large language models (LLMs) operate in interactive multi-turn settings, making multi-turn jailbreaking a realistic threat model and an important scenario for automated red teaming. A core challenge in learning multi-turn jailbreak attackers is credit assignment: different turns contribute differently to the final outcome, but existing turn-level credit signals for jailbreaking are often too coarse for multi-turn interaction. We propose a unified turn-level credit assignment framework for Group Relative Policy Optimization (GRPO) in multi-turn jailbreak learning. Our method assigns group-relative learning signals directly at each turn by decomposing immediate and future credits. We call this approach **decomposed credit GRPO (DC-GRPO)**. In contrast, existing methods often suffer from credit misassignment, for example by applying a single trajectory-level score to the entire dialogue. Across multiple victim LLMs and benchmarks, our dynamic-weighted DC-GRPO achieves an average $ASR@3$ of $98.26\%$ for 5-turn jailbreaking, outperforming existing state-of-the-art methods such as SEMA, which achieves $86.58\%$, and TROJail, which achieves $86.23\%$. These results highlight turn-level group-relative credit assignment as a simple and effective ingredient for scalable multi-turn automated red teaming.


MLLM Makes Strong Backbone for Multi-Modal Object Detection

Yifeng Yang ⋅ Yutong Li ⋅ Jubo Feng ⋅ Qinying Gu ⋅ Cewu Lu ⋅ Guang-Zhong Yang ⋅ Xinbing Wang ⋅ Nanyang Ye

Multimodal large language models (MLLMs) provide a promising alternative by taking object detection as a sequence generation task. However, applying MLLMs to multimodal object detection poses challenges in visual cross-modal (RGB-Infrared) fusion and coordinate prediction under Cross-Entropy supervision. In this paper, we propose a novel MLLM-based object detection framework that addresses these challenges through improving visual cross-modal token alignment and coordinate generation. Specifically, we formulate RGB-IR fusion as pre-decoding token-space modality routing, implemented by the Cross-Modal Complementary Adapters and the Visual Token-wise Router. We further reformulate coordinate-token learning as scale-adaptive neighborhood likelihood maximization through our proposed Geometry-Aware Loss, bridging discrete token prediction and continuous geometric localization. Extensive experiments show that our method achieves state-of-the-art performance, reaching 67.81\% mAP@0.5 under in-distribution scenarios, outperforming existing baselines by over 8\%, while also exhibiting strong generalization to out-of-distribution and referring object detection tasks.


MMFineReason: Closing the Multimodal Reasoning Gap via Open Data-Centric Methods

Honglin Lin ⋅ Zheng Liu ⋅ Yun Zhu ⋅ Chonghan Qin ⋅ Juekai Lin ⋅ Xiaoran Shang ⋅ Conghui He ⋅ Bin CUI ⋅ Wentao Zhang ⋅ Lijun Wu

Recent advances in Vision Language Models (VLMs) have driven significant progress in visual reasoning. However, open-source multimodal models still lag behind proprietary systems, largely due to the lack of high-quality reasoning data. Existing datasets offer limited coverage of challenging domains such as STEM diagrams and visual puzzles, and lack consistent, long-form Chain-of-Thought (CoT) annotations essential for eliciting strong reasoning capabilities. To bridge this gap, we introduce MMFineReason, a large-scale multimodal reasoning dataset comprising 1.8M samples and 5.1B solution tokens, featuring high-quality reasoning annotations distilled from Qwen3-VL-235B-A22B-Thinking. The dataset is established via a systematic three-stage pipeline: (1) large-scale data collection and standardization, (2) CoT rationale generation, and (3) comprehensive selection based on reasoning quality and difficulty awareness. The resulting dataset spans STEM, visual puzzles, games, and complex diagrams, with each sample annotated with detailed, visually grounded reasoning traces. We fine-tune Qwen3-VL-Instruct on MMFineReason to develop MMFineReason-2B/4B/8B versions. Our models establish new state-of-the-art (SOTA) results for their size class. Notably, MMFineReason-4B surpasses Qwen3-VL-8B-Thinking, and MMFineReason-8B even outperforms Qwen3-VL-30B-A3B-Thinking while approaching Qwen3-VL-32B-Thinking, demonstrating remarkable parameter efficiency. Crucially, we uncover a "less is more" phenomenon via our difficulty-aware filtering strategy: a subset of just 7\% (123K samples) achieves performance comparable to the full dataset. Notably, we reveal a synergistic effect where reasoning-oriented data simultaneously boosts general capabilities.

Mixture-of-Experts (MoE) enables multimodal large language models (MLLMs) to scale capacity efficiently, but expert parameters dominate memory—accounting for over 90% of model size in modern MoE-MLLMs (e.g., 54 GB out of ∼58 GB in BF16 for InternVL3.5-30B-A3B and Qwen3-VL-30B-A3B)—even though only a few experts are activated per token. Expert pruning has emerged as an effective compression strategy, but existing methods aggregate importance uniformly across tokens, ignoring that the same expert may serve different roles for vision versus language processing. To address this, we propose MAEP (Modality-Aware Expert Pruning), a training-free framework that decomposes expert importance by modality in a single forward pass. MAEP introduces (1) Delta-H Output (∆H), a redistribution-aware metric measuring the MoE output change when an expert is removed, and (2) low-attention vision tokens as a vision-side pruning signal that is dataset-stable: its expert-importance ranking remains stable across different datasets. Text-side importance serves as a preservation signal while vision-side importance serves as a pruning signal, and experts are ranked globally across layers. Experiments on three MoE-based MLLMs across multimodal and text benchmarks show that MAEP outperforms modality-agnostic baselines on the majority of multimodal benchmarks while remaining competitive on text-only tasks, reducing the BF16 expert weight footprint by ∼27 GB (∼47% of total memory) so that the expert-weight footprint fits within a single 40 GB GPU.


Model Immunization Beyond Condition Numbers: The Importance of Plateau Regions

Quoc M Nguyen ⋅ Trung Le ⋅ Dinh Phung ⋅ Mehrtash Harandi

Harmful fine-tuning can rapidly compromise safety alignment in Large Language Models, underscoring the vulnerability of deployed models to post-release modification. While prior work motivates model immunization by inducing ill-conditioned curvature to slow harmful optimization, this approach is not scalable due to the high computational cost of Hessian estimation. We refine this paradigm by extending the theoretical framework to prioritize plateau regions alongside curvature. We further introduce an immunization algorithm based on efficient Hessian estimation and bounds on the extremal singular values of the Hessian. We demonstrate the proposed immunization framework in three applications: (i) Fine-tuning-as-a-service settings where benign user data may be mixed with poisoned harmful data, (ii) open-weight settings where an adversary can fine-tune on harmful-only data with varied training dynamics, and (iii) robust unlearning, where immunized models better resist relearning attacks after unlearning. Extensive experiments confirm that our approach achieves robust defense against adversarial fine-tuning while maintaining general utility and downstream trainability.


Model Incrimination: Investigating Whether Concerning Behavior Reflects Misalignment

Gerson Kroiz ⋅ Aditya Singh ⋅ Senthooran Rajamanoharan ⋅ Neel Nanda

A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning behavior, but behavior alone is not sufficient to establish misalignment: a concerning action can arise from benign causes such as confusion. This raises the problem of determining whether malign intent underlies such behavior, a process we term model incrimination. The goal of this paper is to develop effective methods for doing so. To enable this, we create a suite of six agentic environments where models exhibit concerning behavior as practice grounds, and follow a simple two-step protocol for investigating the causes behind model behavior: hypothesis generation via reading the chain of thought -- which, while not always faithful, is a rich source of hypotheses about what drives model behavior -- followed by hypothesis validation via environment interventions and additional methods as appropriate. As we do not have access to ground truth about why a model takes an action, we rely on convergent findings across independent experiments as our standard of evidence. By following our protocol, we learn effective methods for concretely determining motivations: for example, we use predictions to back out latent properties of model behavior, like the fact that Kimi K2 Thinking takes shortcuts due to a legitimate disposition towards less effortful courses of action, and that reward hacking in frontier models is strategic, while in weaker models it is not. However, some unanswered questions require further methodological development: we construct an absence-of-evidence case against Kimi K2 Thinking believing it is going against the user's wishes while taking shortcuts, but run into confounds that limit our confidence in its absence. Overall, a key takeaway is that simple methods like reading the CoT and environment interventions are highly effective. More broadly, our work shows model incrimination is a tractable empirical problem with significant room for progress, and establishes a baseline for future work.

The standard mechanistic interpretation of transformer factual recall attributes storage to deep MLP modules and routing to attention. We study how module roles change when the fact distribution carries subject, relation, or answer-collision structure rather than independent facts. Which module performs the recall depends on the correlation axis and on the model's capacity load. Probes and matched-counterfactual patching disagree systematically on every motif: probes credit the deepest MLP while patching credits the first attention module. We show that probes and patching measure structurally different quantities---cumulative linear decodability versus irreplaceable causal contribution---which under correlated fact distributions need not coincide. A parameter-counting argument singles out subject-axis collision as the unique compression that the attention-routing pathway can exploit cheaply, predicting an order-of-magnitude crossover for that pathway class that we observe empirically: the layer-2 probe gap collapses sharply for subject collisions, and only subject-axis specialization survives the capacity-tight regime. Edge-level path patching uncovers a three-step chain (Attn -> MLP1 -> MLP2) whose dominance tracks the predicted threshold, with a causal-scrubbing test corroborating that the chain is causally relevant on every motif and sharpest in the subject regime. A dense joint sweep shows that subject and answer compressions compete rather than amplify, with the dominant axis set by the stronger compression. On natural CounterFact strata across four frozen pretrained LMs (Pythia-1.4B/2.8B, LLaMA-3.2-1B-Inst/3.1-8B), edge-level path patching reproduces the synthetic ranking: the answer-collision stratum is the most localized on every model. A two-hop extension shows the binding-locality structure carries to compositional facts.


MotionCFG: Boosting Motion Dynamics via Semantic Motion Sharpening

Byungjun Kim ⋅ Soobin Um ⋅ Jong Chul Ye

Despite recent advancements in Text-to-Video (T2V) synthesis, generating high-fidelity and dynamic motion remains a significant challenge. Existing methods primarily rely on Classifier-Free Guidance (CFG), but standard CFG treats all semantic dimensions uniformly and provides no mechanism to selectively resolve ambiguity along the motion axis. To address this, we propose MotionCFG, a training-free approach that introduces selective guidance along the motion subspace of the condition embedding. Specifically, we construct a motion-perturbed negative condition by injecting Gaussian noise into motion-related embeddings, and steer generation away from it to selectively amplify motion intent. We show theoretically that this procedure implicitly approximates Semantic Laplacian Sharpening, a second-order correction that suppresses motion-ambiguous regions of the score landscape and amplifies well-resolved dynamic peaks. Combined with a piecewise guidance schedule that confines intervention to the early denoising steps, MotionCFG consistently improves motion dynamics across state-of-the-art T2V frameworks with negligible overhead. We further demonstrate that this Laplacian sharpening principle generalizes beyond motion, effectively steering complex, non-linear concepts such as precise object numerosity that are typically difficult to modulate via standard text-based guidance.


MotionGrounder: Grounded Multi-Object Motion Transfer via Diffusion Transformer

Samuel Teodoro ⋅ Yun Chen ⋅ Agus Gunawan ⋅ Soo Ye Kim ⋅ Jihyong Oh ⋅ Munchurl Kim

Motion transfer enables controllable video generation by transferring temporal dynamics from a reference video to synthesize a new video conditioned on a target caption. However, existing Diffusion Transformer (DiT)–based methods are limited to single-object videos, restricting fine-grained control in real-world scenes with multiple objects. In this work, we introduce MotionGrounder, a DiT-based framework that $\textit{firstly}$ handles motion transfer with $\textit{multi-object controllability}$. Our Flow-based Motion Signal (FMS) in MotionGrounder provides a stable motion prior for target video generation, while our Object-Caption Alignment Loss (OCAL) grounds object captions to their corresponding spatial regions. We further propose a $\textit{new}$ Object Grounding Score (OGS), which jointly evaluates (i) spatial alignment between source video objects and their generated counterparts and (ii) semantic consistency between each generated object and its target caption. Our experiments show that MotionGrounder consistently outperforms recent baselines across quantitative, qualitative, and human evaluations.

As deep neural networks are required to continuously adapt to evolving data, class-incremental learning (CIL) has become an active research area. However, most existing works primarily focus on improving accuracy while overlooking confidence calibration, which measures the trustworthiness of the model's predictions. Although temperature scaling (TS) is effective for static calibration, it has rarely been explored in continual settings where data distributions evolve over time. In this paper, we introduce Multi-Group Temperature Scaling for Continual Calibration (MT-CC), a novel post-hoc calibration framework that addresses the asymmetric calibration behavior in CIL, where the rate of confidence changes does not align with the rate of accuracy degradation across tasks, leading to different levels of miscalibration. Our key contribution is to design multiple temperature parameters through a lightweight grouping network that separates samples by confidence statistics and task-aware signals in CIL. Furthermore, a Groupwise Difference-between-Confidence-and-Accuracy (GDCA) regularization is incorporated to promote intra-group consistency. Experiments demonstrate that MT-CC significantly improves calibration performance while maintaining accuracy, offering a principled framework for reliable continual learning.

We consider classification in high-dimensional feature spaces where only a vanishing fraction of features are useful and the number of classes may be large. While this rare-useful-feature regime is well understood in the binary setting, its multiclass version raises new questions: a feature can be useful for class discrimination without separating any specific pair of classes. We focus on a rare feature diversity model, in which most features are useless, but a small fraction promote discrimination among classes in a way that can be detected by a per-feature diversity test, such as one-way ANOVA or the Kruskal--Wallis test. For this setting, we propose Diversity Pursuit Higher Criticism (DP-HC), a feature-screening procedure that applies a higher-criticism threshold to the collection of p-values obtained from per-feature diversity tests. We justify DP-HC under a multiclass rare/weak Gaussian feature model with class-mean contrasts of fixed magnitude in random directions, extending the binary model of Donoho and Jin (2008). When the number of classes is fixed, our model recovers the binary phase transition. When the number of classes grows logarithmically with the number of features, however, we derive a new and strictly harder phase transition, driven by a second-order variance term that the fixed-class analysis does not capture. We show that whenever classification beyond chance is possible, DP-HC with ANOVA, followed by a simple classifier based on similarity to class means, achieves asymptotically perfect accuracy. Experiments on real multiclass classification benchmarks demonstrate substantial improvements, and numerical simulations confirm that the derived phase-transition curve accurately predicts the finite-sample behavior of DP-HC in the growing-class regime.


Multi-Marginal Couplings for Metropolis--Hastings

Truong Buu Phan ⋅ Gergely Flamich ⋅ Ashish Khisti ⋅ Shahab Asoodeh

Convergence diagnosis for Markov chain Monte Carlo is a matter of fundamental importance in computational statistics: it determines the resources allocated to a particular sampling problem and influences the practitioner's view of the quality of estimates obtained from a Markov chain. Motivated by this, we contribute to the emerging class of coupling-based convergence diagnostic algorithms. Concretely, we study coupling multiple Metropolis--Hastings chains using multi-marginal coupling. We introduce a natural objective for this setting and establish lower and upper bounds by drawing connections to list-level distribution coupling and distributed pairwise-matching problems. This analysis ultimately leads to a shared-randomness Poisson Monte Carlo construction for coupling multiple Markov chains. In this process, we avoid a key dimension-dependent bottleneck in the runtime complexity of classical Poisson Monte Carlo by developing an adaptive rule for updating the point process, yielding significant gains in high-dimensional settings. Experiments on grand couplings of Markov chains show that our methods improve coalescence rates across dimensions, reducing meeting times by up to 50\% compared with existing baselines.

State-of-the-art pathology foundation models, trained on millions of histology tiles, can fail to preserve tissue similarity when comparisons cross slide or institution boundaries. We show that general-purpose multimodal LLMs, without being trained as pathology foundation models, consistently outperform these specialized models in cross-domain histological similarity judgments. Using a relative similarity framework that we release as the MOSAIC (Model Similarity Assessment across Institutions and Cohorts) benchmark, we evaluate 17 models across 6 datasets and find that pathology encoders often rank same-institution, different-disease tiles as more similar than same-disease, different-institution tiles, a clinically dangerous failure mode invisible to standard within-domain evaluations. LLMs appear less susceptible to this failure, likely because they perform semantic visual comparison of morphology and tissue architecture rather than relying on shortcut features tied to acquisition context. Scaling training data does not resolve the problem for pathology encoders, implicating the learning objective rather than data coverage. Our results expose a fundamental robustness gap in current pathology foundation models and establish multimodal LLMs as a viable alternative for cross-institutional retrieval, dataset harmonization, and multi-site quality control. Code and data will be released upon acceptance.

Multiobjective submodular maximization asks for a single set that performs well across multiple submodular objectives, a setting motivated by robust experimental design and fairness-aware decision making. The standard max-min formulation, however, can be overly conservative: it is determined entirely by the worst-performing objective and ignores improvements in the remaining objectives unless they change the minimum. Motivated by concave social-welfare aggregators from economics, we initiate the study of multiobjective submodular maximization with general concave aggregation. We propose a randomized greedy algorithm that, in each iteration, computes a distribution over elements by solving a concave program and samples an element from this distribution. Our analysis overcomes the loss of the objective-wise decomposition available in the max-min case and proves an asymptotic $(1-1/e-\epsilon)$-approximation under mild concentration assumptions. To make the method scalable, we develop a Fenchel-duality-based lazy evaluation scheme. Experiments on synthetic and real-world instances show that our algorithm consistently improves over a naive greedy baseline across a variety of concave aggregators, while enabling utility--fairness trade-offs that are not captured by the max-min formulation.

Discovering object-centric representations from images can significantly enhance the robustness, sample efficiency and generalizability of vision models. Works on images with multi-part objects typically follow an implicit object representation approach, which fail to recognize these learned objects in occluded or out-of-distribution contexts. This is due to the assumption that object part-whole relations are implicitly encoded into the representations through indirect training objectives. We address this limitation by proposing a novel method that uses explicit graph representations for parts and introduce a new co-part algorithm for multi-part object discovery. The part representations are refined with a part-centric regularization objective. We then introduce three benchmarks to evaluate the robustness of object-centric methods in recognizing multi-part objects within occluded and out-of-distribution settings. Experimental results on simulated, realistic, and real-world images show marked improvements in the quality of discovered objects compared to state-of-the-art methods, as well as the accurate recognition of multi-part objects in occluded and out-of-distribution contexts. We also show that the discovered object-centric representations can more accurately predict key multi-part object properties in a downstream task, highlighting the potential of our method to advance the field of object-centric representations.

Consistency models achieve fast image synthesis by chaining a small number of neural evaluations along a diffusion trajectory, yet the distribution error of $N$-step sampling for $N \geq 3$ has lacked a rigorous characterisation. We establish a tight Wasserstein-$p$ error bound for arbitrary $N$: the total error equals a Lipschitz-weighted sum of per-step local consistency errors, $\sum_{k=1}^{N} L_x^{N-k} \delta_k$, making the propagation mechanism explicit. From this bound we derive a closed-form optimal step count $N^{\star}$ whose structure bifurcates on whether the spatial Lipschitz constant $L_x$ exceeds one: for $L_x > 1$, exponential error growth yields $N^{\star} \approx (\ln L_x)^{-1} \log(\beta\gamma(L_x-1)/(c \ln L_x))$; for $L_x < 1$, geometric-series saturation gives $N^{\star} \approx (\beta\gamma(1-L_x)/(c \, |\ln L_x|))^{1/(\gamma+1)}$. The two-regime structure accounts for ECT's best FID of $2.73$ at $N=2$: with $L_x \approx 0.85 < 1$ (Case 2), the formula gives $N^{\star} \approx 1.7$. Experiments on CIFAR-10 reproduce the predicted U-shaped FID curve with $N^{\star} = 8$ on a capacity-constrained model in which error accumulation is easier to see.


Multivariate Time Series Forecasting needs Cross Variable Loss

Kuiye Ding ⋅ Yifan Hu ⋅ Hanchen Wang ⋅ Hao Xue

Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. While existing studies mainly focus on cross-variable dependencies in historical observations, dependencies among future values are much less explored. Specifically, modern forecasting models largely follow the Direct Forecasting (DF) paradigm, generating multi-step forecasts with point-wise objectives that do not explicitly constrain cross-variable structure. In this work, we show that the DF objective is mismatched in the presence of cross-variable and lagged dependencies, revealing an objective gap. To address this issue, we propose Cross-Variable Loss (CvLoss), a plug-in structural regularizer that constrains forecast residuals on a cross-variable graph. CvLoss penalizes inconsistent edge-wise residual differences over forecast patches, encouraging consistency across both synchronous and asynchronous interactions. Our experiments show that CvLoss consistently improves competitive forecasting models, outperforms representative learning objectives, and is compatible with a variety of forecasting backbones. Code is available at: https://anonymous.4open.science/r/CvLoss.


Muon Does Not Converge on Convex Lipschitz Functions

Tetiana Parshakova ⋅ Ahmed Khaled ⋅ Michael Crawshaw ⋅ Guillaume Garrigos ⋅ Robert Gower

Muon and its variants have shown strong empirical performance in a variety of deep learning tasks. Existing convergence analyses of Muon rely on smoothness assumptions, though arguably the most successful function class for developing deep learning methods (such as AdaGrad, Shampoo, Schedule-Free and more) has been the class of convex and Lipschitz functions. In this paper we question whether the classical convex Lipschitz model is a useful one for understanding Muon. Our answer is no. We show that Muon does not converge on the class of convex and Lipschitz functions, regardless of the choice of learning rate schedule. We also show that error feedback restores convergence of Muon and all the non-Euclidean subgradient methods with momentum. However, this theoretical fix using error feedback degrades the performance of Muon in two representative settings for image classification (CIFAR-10) and language modeling (nanoGPT on FineWeb-Edu 10B). Our conclusion is that convex Lipschitz theory, despite having a prominent role in the design of practical methods for deep learning, is not the most suited one for Muon. This suggests that Muon's success must come from structure absent from this model, most plausibly related to smoothness.

Test-time scaling (TTS) improves reasoning by allocating additional inference compute, but in multimodal large language models (MLLMs), more compute is not always reliable. Longer reasoning may amplify visual misperception, induce answer drift, reinforce majority traps, or turn initially correct predictions into overthought failures. We propose MUST, a training-free framework that recasts multimodal TTS as stage-conditioned stability control. MUST separates inference into reasoning and verification stages and assigns each stage a different reliability criterion. During reasoning, Contrastive Answer-Manifold Stability (CAMS) selects answers whose supporting trajectories form compact and separable latent groups. During verification, Attention-Free Support-Chain Stability (AF-SCS) evaluates whether explicit visual evidence and subsequent reasoning form a stable latent support chain, without relying on token-to-patch attention maps. Experiments on multimodal reasoning benchmarks show that MUST improves the accuracy-compute trade-off and reduces overthinking-induced errors. These results suggest that reliable MLLM test-time scaling should adaptively control when and how extra compute is trusted, rather than simply increasing inference budget.


MUTE: Multi-Level Alignment Uncoupling Against Talking-Head Exploitation for Voice Protection

Donghyun Kim ⋅ Jin Hong ⋅ seungmin Kim ⋅ Dain Kim ⋅ Junseok Kwon ⋅ Daeseon Choi

Talking-head generation models, which synthesize realistic facial animations from audio, are increasingly vulnerable to misuse in multimodal deepfake scenarios. Protecting such systems remains challenging, as talking-head models fundamentally rely on precise audio–visual alignment, while effective audio protection methods remain largely underexplored. In this work, we propose MUTE, an audio protection framework that uncouples the underlying audio–visual alignment through a multi-level strategy. Specifically, MUTE combines (1) Representation-level Degradation, which perturbs temporal audio embeddings to degrade their representations, and (2) Alignment-level Disruption, which directly perturbs cross-modal attention to disrupt structural alignment. To improve robustness, perturbations are constrained in the STFT domain to high-energy regions, making them resistant to post-processing and denoising. An optional speaker-level objective further mitigates potential bypass via TTS-based resynthesis. Extensive experiments demonstrate that MUTE consistently degrades lip synchronization across both white-box and black-box talking-head models while preserving perceptual audio quality. The proposed method remains effective under real-world transformations and can be combined with image-based approaches to provide stronger multimodal protection.


MyoChallenge 2025: A New Benchmark for Human Athletic Intelligence

Cheryl Wang ⋅ Chun Kwang Tan ⋅ Balint Hodossy ⋅ Shirui Lyu ⋅ Jun Guo ⋅ Wentao Zhao ⋅ Huaping Liu ⋅ Chengkun Li ⋅ Merkourios Simos ⋅ Bianca Ziliotto ⋅ Alexander Mathis ⋅ Siyuan Liu ⋅ Jiahao Chen ⋅ Shanlin Zhong ⋅ Bo Jiang ⋅ Ci Song ⋅ Yaoye Zhu ⋅ Chenhui Zuo ⋅ Yanan Sui ⋅ Irfan Refai ⋅ Massimo Sartori ⋅ Guillaume Durandau ⋅ Vikash Kumar ⋅ Vittorio Caggiano

Athletic performance represents the pinnacle of human motor intelligence, demanding rapid choices, precise control, agility, and coordinated physical execution. Replicating this seamless combination of capabilities remains elusive in current artificial intelligence and robotic systems. Concurrently, understanding the biological mastery of these movements is hindered because complex muscle coordination is rarely measured in vivo due to the limitations of physical equipment. To bridge this fundamental gap in understanding, $\textit{MyoChallenge}$ at NeurIPS 2025 established a pioneering benchmark for motor control intelligence in sports, leveraging high-fidelity musculoskeletal models within physics simulation combined with machine learning-driven algorithms. The competition introduces two distinct tracks emphasizing either upper or lower limbs control: a table tennis rally task utilizing a biomechanic upper limb composed of an arm with a hand (MyoArm) and a trunk (MyoTorso); and a soccer penalty kick using a biomechanic model of legs (MyoLeg) and a trunk (MyoTorso). Marking the fourth iteration of the MyoChallenge series, this event attracted almost 70 teams and over 560 submissions globally, uniting a diverse community ranging from physicians and neuroscientists to machine learning experts. The competition facilitated the development of several state-of-the-art control algorithms for a musculoskeletal system capable of sports agility, leveraging techniques such as physics-based motion planners, on-policy behaviour cloning, hierarchical planning, and muscle synergies that significantly surpassed our provided baseline performance. By integrating standardized tasks and physiologically realistic models into the open-source framework of MyoSuite, MyoChallenge'25 serves as a reproducible and reusable testbed to accelerate interdisciplinary research across machine learning, biomechanics, sports science, and neuroscience. Project page: https://www.myosuite.org//myochallenge/myochallenge-2025

Catastrophic forgetting is a central obstacle in single-stage fine-tuning of large language models (LLMs) on a new task, especially when the original pretraining data are unavailable. Synthetic data offers a practical mitigation path, but candidate pools are often noisy and redundant, making effective selection essential. Existing methods often rely on indirect proxy signals whose alignment with actual training benefit can be limited, while utility--diversity selection over large candidate pools remains computationally demanding. To address these challenges, we propose $\textbf{NA}$vigator-guided $\textbf{D}$ata $\textbf{S}$election ($\textbf{NADS}$), a fine-tuning framework for LLMs. NADS first follows the target fine-tuning recipe without preservation regularization to obtain a navigator model, whose deviation from the pretrained model exposes recipe-induced capability drift. It then uses the predictive divergence between the navigator and the pretrained model to score forgetting-aware utility for each synthetic candidate, and constructs a constraint set through an efficient utility--diversity selection strategy. Finally, NADS fine-tunes on the new-task data while distilling pretrained behavior on the selected constraints to preserve general capabilities. Experiments show that NADS delivers a stronger balance between new-task performance and general capability preservation, while reducing utility--diversity selection overhead.

Automated scoring and feedback generation for Chinese narrative essays (CNE) represent a critical task in educational measurement. While Large Language Models (LLMs) have demonstrated transformative progress, their potential in CNE evaluation scenarios has yet to be systematically under-explored. Existing essay benchmarks face some challenges: (1) \textbf{Standard Alignment}: The varied complexity of scoring traits and rubrics across different benchmarks makes it difficult to establish a consistent alignment mechanism for heterogeneous evaluation scenarios; (2) \textbf{Evaluation Scope}: Current benchmarks only focus on discriminative scoring accuracy for LLMs, yet they lack a systematic framework for assessing the quality of generative feedback. (3) \textbf{Prompt Coverage}: There is a lack of CNE datasets that cover both closed-ended and open-ended prompts. These issues have, to some extent, constrained the trust of educators in and their adoption of LLM-based methods for essay assessment. To bridge these gaps, we introduce \textbf{NarrativeBench}, a multi-trait and feedback-oriented evaluation benchmark for CNE. NarrativeBench covers both closed-ended and open-ended writing tasks and consists of 4,500 student essays from grades 3 to 5 across 14 different writing prompts. Finally, we benchmark 13 popular LLMs, revealing that their performance exhibits significant gaps compared to human evaluations and pretrained language models. Through NarrativeBench, we comprehensively assess LLM performance on Chinese narrative essay tasks, thus facilitating LLM progress in Chinese essay data analysis domains.

We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards as a fairness-respecting performance metric. Existing formulations define Nash regret as $NR_T = \mu^\star - (\prod_{t=1}^T \mathbb{E}\mu_{I_t})^{1/T}$, where $\mu_{I_t}$ is the mean reward of the recommended arm $I_t$ and $T$ is the learning horizon. We observe that, since ensemble Nash regret operates on per-round marginal expectations before applying the geometric mean, it does not capture the joint distribution of rewards across rounds — leaving an aspect of the NSW fairness motivation unaddressed at the trajectory level. To address this, we propose \emph{trajectory-wise Nash regret} $\widetilde{NR_T}= \mu^\star - \mathbb{E}[(\prod_{t=1}^T \mu_{I_t})^{1/T}]$, where the geometric mean is computed over complete sample paths before taking expectations. This formulation captures the fairness properties of NSW more faithfully and requires controlling the joint distribution of rewards across rounds, rather than merely per-round marginals. By Jensen's inequality applied to the concave geometric mean, $\widetilde{NR_T} \geq NR_T$, making it a strictly stronger metric. We additionally introduce \emph{high probability Nash regret} $\widehat{NR_T} = \mu^\star - (\prod_t \mu_{I_t})^{1/T}$, obtaining the first high probability regret bounds in the fair bandits literature. We propose a two-phase algorithm, Round Robin Nash Confidence Bound (\texttt{RR-NCB}), combining structured round robin exploration with a Nash confidence bound index policy. We establish that $\widetilde{NR_T} \leq \widetilde{\mathcal{O}}(\sqrt{k\log T/T})$ and $\widehat{NR_T} \leq \widetilde{\mathcal{O}}(\sqrt{k\log(kT/\delta)/T})$ with probability $1-\delta$, matching the optimal $\widetilde{\mathcal{O}}(\sqrt{k/T})$ scaling despite operating under strictly stronger metrics. Optimality follows from a lower bound chain via AM-GM and standard $k$-armed bandit minimax arguments. We validate our theoretical findings through simulations.


NavOCR: A Dataset Generator for Navigation-Relevant Text Detection in Mobile Robots

Chaehyeuk Lee ⋅ Sooyong Shin ⋅ Sehyeon Yoon ⋅ Youngsun Je ⋅ Jinmyoung Lee ⋅ Seula Lee ⋅ Tu Vo ⋅ Sheir A. Zaheer ⋅ Chan Youn Park

Textual cues provide important navigation-relevant semantic information like room names, store signs, and floor labels. However, conventional optical character recognition (OCR) systems and datasets are designed to detect all visible text, including advertisements and price tags, regardless of their relevance to navigation tasks. This characteristic of OCR can produce noisy semantic cues, making it difficult to identify which text is useful for robot navigation. A straightforward solution is to curate existing OCR datasets by retaining only navigation-relevant text and removing irrelevant annotations, but this process requires substantial manual effort. To address this problem, we introduce NavOCR, a dataset generator for navigation-relevant text detection that does not require manual curation. By leveraging conventional OCR, open map resources, and filtering with a vision-language embedding model, NavOCR identifies which text instances serve as place identifiers for robot navigation and constructs datasets from images available on the web. NavOCR aims to complete the entire process within a single day, from data collection to model training, enabling rapid adaptation to diverse environments and languages. Compact models trained on datasets generated by NavOCR achieve higher detection accuracy and substantially better computational efficiency than vision language models and open vocabulary detectors. We further show that detecting navigation-relevant text improves the quality of semantic cues used for localization and mapping in robot navigation applications.

This paper studies kernelized bandits (also known as Gaussian process bandits) in an adversarial environment, where the reward functions in a known reproducing kernel Hilbert space (RKHS) may be adversarially chosen at each round. We show that the exponential-weight algorithm achieves $\tilde{O}(\sqrt{T \gamma_T})$ adversarial regret, where $T$ and $\gamma_T$ denote the number of total rounds and the maximum information gain, respectively. For $\nu$-Mat\'ern kernels, we also show algorithm-independent lower bounds that guarantee the optimality of our algorithm up to polylogarithmic factors. Furthermore, we present a computationally efficient variant of our algorithm using Nystr\"om approximation while maintaining nearly optimal regret guarantees.

We study the problem of \emph{computationally efficient} robust estimation of the covariance/scatter matrix of elliptical distributions---that is, affine transformations of spherically symmetric distributions---under the \emph{strong contamination model} in the high-dimensional regime $d \gtrsim 1/\varepsilon^2$, where $d$ is the dimension and $\varepsilon$ is the fraction of adversarial corruptions. We show that the structure inherent to elliptical distributions enables us to achieve estimation guarantees comparable to those known for the Gaussian case. We propose an algorithm that, under a {very mild assumption} on the effective rank of the scatter matrix $\Sigma$, and given a nearly optimal number of samples $n = \tilde{O}(d^2/\varepsilon^2)$, computes in polynomial time an estimator $\hat{\Sigma}$ satisfying $ \left\Vert \Sigma^{-1/2} \hat{\Sigma} \Sigma^{-1/2} - Id \right\Vert_{\text{F}} \le O(\varepsilon \log(1/\varepsilon)) . $ This matches the best known guarantees for the Gaussian setting. As an application of our result, we obtain \emph{efficiently computable, nearly optimal robust covariance estimators}, significantly generalizing prior results that were restricted to the Gaussian case or required, in particular, matching the Gaussian fourth moment. Specifically, for elliptical distributions satisfying the Hanson--Wright inequality (including Gaussians, uniform distributions over ellipsoids, and more generally elliptical distributions with Gaussian-type radial concentration), our estimator $\hat{\Sigma}$ of the covariance $\Sigma$ achieves the same error guarantee as in the Gaussian case. Moreover, for elliptical distributions with sub-exponential tails (such as the multivariate Laplace distribution), our covariance estimator $\hat{\Sigma}$ satisfies the spectral norm bound $ \left\Vert \Sigma^{-1/2} \hat{\Sigma} \Sigma^{-1/2} - Id \right\Vert \le O(\varepsilon \log(1/\varepsilon)) . $ Remarkably, despite the heavier tails of such distributions, the covariance can still be estimated at the same rate as in the Gaussian case---a phenomenon unique to high dimensions and absent in low-dimensional settings. Our approach is based on estimating the covariance of the \emph{spatial sign} (i.e., the projection onto the sphere) of elliptical distributions. As part of our framework, we develop a generalization of the standard covariance filtering algorithm that allows us to work with distributions whose fourth moment differs from that of the Gaussian distribution.


Neural Backward Filtering Forward Guiding

Gefan Yang ⋅ Frank van der Meulen ⋅ Stefan Sommer

Inference in nonlinear continuous stochastic processes on trees is challenging, particularly when observations are sparse and the topology is complex. Exact smoothing via Doob's $h$-transform is intractable for general nonlinear dynamics. We propose Neural Backward Filtering Forward Guiding (NBFFG), a unified framework for both discrete transitions and continuous diffusions. Our method constructs a variational posterior by leveraging a proxy linear-Gaussian process. This proxy process yields a closed-form backward filter that serves as a guide, steering the generative path toward high-likelihood regions. We then learn a neural residual to capture the non-linear discrepancies. This formulation allows for an unbiased pathwise subsampling scheme, reducing the training complexity from tree-size dependent to path-length dependent. Empirical results show that NBFFG outperforms baselines on synthetic benchmarks, and we demonstrate the method on a high-dimensional inference task in phylogenetic analysis with reconstruction of ancestral butterfly wing shapes.


Neural Bayesian Filtering

Christopher Solinas ⋅ Radovan Haluška ⋅ David Sychrovský ⋅ Finbarr Timbers ⋅ Nolan Bard ⋅ Michael Buro ⋅ Martin Schmid ⋅ Nathan Sturtevant ⋅ Michael Bowling

Sequential estimation under partial observability requires tracking belief distributions that may be high-dimensional, multimodal, and non-Gaussian. Classical Bayesian filters generalize zero-shot to any system whose dynamics can be evaluated, but their representations scale poorly: parametric filters struggle to capture multimodality, and particle filters require exponentially many particles in the state dimension. Generative Distribution Embeddings (GDEs) learn compact representations of complex distributions but provide no mechanism for sequential updates. We present Neural Bayesian Filtering (NBF), a method that combines classical Bayesian filtering with learned distribution embeddings. NBF represents beliefs as embeddings and updates them by sampling particles from a learned conditional generator, propagating them through the system dynamics, and finally re-embedding the resulting particles. For in-family distributions that the embeddings are trained to represent, NBF inherits the zero-shot adaptability of classical filters, but with the benefit that regenerating particles from the embedding at each step also mitigates particle impoverishment. We validate NBF on a pursuit--evasion domain, demonstrating accurate tracking of complex multimodal posteriors, graceful scaling with state dimension, and zero-shot generalization to dynamics held out from training.


Neural networks are more modular than single neurons suggest

Shreya Saha ⋅ Benjamin Bergen ⋅ Meenakshi Khosla

Neural networks routinely develop representations in which individual neurons respond to mixtures of task-relevant variables, leading to the widespread view that such networks lack modular organization. We challenge this conclusion by introducing a formally nested hierarchy of modularity: unit-level, which requires distinct computations to occupy disjoint neurons; orthogonal, which asks whether a rotation reveals functionally independent subspaces; and linear, which asks whether a bounded linear map can separate task-relevant computations into independent subspaces. We develop a unified framework for learning decompositions at each level and evaluating them causally through selective ablation. Across vision models, anguage models and multitask RNNs, we find that networks which lack unit-level modularity nonetheless often exhibit clear functional modularity at the subspace level. In vision and language models, content versus style and syntax versus semantics respectively dissociate under subspace ablation despite sharing overlapping populations of neurons. In RNNs trained on multiple cognitive tasks and previously shown to lack modular structure, our decomposition recovers sharp double dissociations between tasks. These results demonstrate that modularity cannot be assessed at a single level of description. The hierarchy we introduce provides a principled framework for determining when and at what level of analysis functional specialization exists in neural network representations.


Neural Neural Scaling Laws

Michael Hu ⋅ Jane Pan ⋅ Ayush Rajesh Jhaveri ⋅ Nicholas Lourie ⋅ Kyunghyun Cho

Neural scaling laws predict how language model performance improves with increased training inputs. While aggregate metrics like validation loss can follow smooth power-law curves, individual downstream tasks exhibit diverse scaling behaviors: some improve monotonically, others plateau, and some even degrade with scale. We argue that predicting downstream performance from validation loss suffers from two limitations: averaging token-level losses obscures signal, and no simple parametric family can capture the full spectrum of scaling behaviors. To address this, we propose Neural Neural Scaling Laws (NeuNeu), a neural network that frames scaling law prediction as time-series extrapolation. NeuNeu combines temporal context from observed accuracy trajectories with token-level validation losses, learning to predict future performance without the limitations inherent in assuming a specific functional form. Trained entirely on open-source model checkpoints from HuggingFace, NeuNeu achieves 1.99\% mean absolute error in predicting model accuracy on 66 downstream tasks---a 44\% reduction compared to logistic scaling laws (3.56\% MAE). Furthermore, NeuNeu generalizes zero-shot to unseen model families, architectures, parameter counts, and downstream tasks. Our work suggests that predicting downstream scaling directly from data outperforms parametric alternatives.

In this paper, we tackle the critical failure modes of Physics-Informed Neural Networks (PINNs), such as spectral bias, which lead to poor convergence on complex PDEs. We identify two key shortcomings in existing curriculum learning methods for PINNs: unreliable knowledge transfer between stages and a reliance on manual, ad-hoc curriculum design. To overcome these limitations, we present Neural Operator-based Curriculum Learning (NOCL), a unified framework that leverages Neural Tangent Kernel (NTK) theory to automate curriculum generation and employs neural operators to enable robust, dynamic knowledge transfer across curriculum stages. By dynamically training neural operators and filtering data for PINN initialization, our approach ensures scalable and effective learning across progressively difficult tasks. Experiments verify that NOCL leads to marked gains in convergence and generalization relative to prior methods, resulting in substantially better performance on test suite.


Neural Refraction Fields for Image Verification

Sage Simhon ⋅ Jingwei Ma ⋅ Prafull Sharma ⋅ Lucy Chai ⋅ Yen-Chen Lin ⋅ Phillip Isola

Image verification is an increasingly critical need as the capabilities for visual media tampering improve. We present an approach to image verification that embeds a private authenticity signature into an image based on physical refractive objects, which transform the image. To verify a protected image, we compare the image and a pixel-aligned reconstruction from the embedding to identify inconsistencies. Our approach is inspired by prior work, which physically places refractive objects in a scene before taking a photo. However, prior work is limited to simple refractions with analytical formulas and requires a slow per-scene neural radiance field optimization to get a reconstruction. Instead, our method trains a compact, scene-agnostic neural refraction field capable of modeling complex and more secure refractive geometries. Once trained, it enables instant, high-fidelity reconstruction for manipulation detection and localization.


Neural Scaling Laws in Particle Jets

Matthias Vigl ⋅ Nikita Pond ⋅ Nicole Hartman ⋅ Jackson Barr ⋅ Daniel H Guest ⋅ Alexander Froch ⋅ Michele D'Andrea ⋅ Dmitrii Kobylianskii ⋅ Daniel Murnane ⋅ Robert Les ⋅ Sébastien Rettie ⋅ Diptaparna Biswas ⋅ Michael Kagan ⋅ Lukas Heinrich

Large language models (LLMs) have demonstrated that scaling, driven by compute, can yield dramatic improvements in performance, with current state-of-the-art systems reaching up to one trillion parameters. In contrast, state-of-the-art jet flavor classifiers in particle physics remain in the tens of millions of parameters. Scaling laws offer a framework for predicting how much particle physics models benefit from systematic increases in capacity and training data. In this work we study the behavior of increasingly larger jet classifiers, deriving scaling laws under optimal use of compute budget and limited datasets. We also identify the regimes where “double descent” emerges. The emerging scaling behavior points to a clear need for significantly larger datasets and models to saturate available compute and approach the limit of achievable performance. We therefore train a classifier on a new dataset 20 times larger than the current state-of-the-art (7.7 billion jets) and find that the observed performance follows the scaling laws prediction. This motivates continued integrated efforts in data generation, model scaling, and inference infrastructure for large-scale models in the physical sciences.

Structural reasoning, the ability to recognize and make inference over the relational structure between objects and concepts, is a hallmark of human cognition, yet prevailing methods often collapse relational topology into flat embeddings, cannot discover hidden structure and lack interpretability. We introduce Neural Structural Reasoner (NSR), a brain-inspired network that preserves relational structure directly in the connectivity and dynamics of coupled neuronal populations. NSR draws inspiration from three biological mechanisms: multi-layered architecture for encoding hierarchical knowledge, stable representation of entity and concepts, and path integration for input-driven state inference. At query time, NSR parallelizes computation over candidate relational structures and leverages confidence-weighted scores to perform link prediction. On standard knowledge-graph benchmarks, NSR matches strong embedding and rule-mining baselines with superior training efficiency. Because reasoning is implemented through sequences of human-readable neuron activations, NSR affords native interpretability by tracking intermediate inference steps. The model further extracts latent relational hierarchies and compositional rules, demonstrating the brain-inspired architecture as an effective, efficient, and highly interpretable substrate for structural reasoning.


NeuroRVQ: Multi-Scale Biosignal Tokenization for Generative Foundation Models

Konstantinos Barmpas ⋅ Na Lee ⋅ Dimitrios Chalatsis ⋅ William Raftery ⋅ Yannis Panagakis ⋅ Dimitrios Adamos ⋅ N Laskaris ⋅ Alexandros Koliousis ⋅ Dario Farina ⋅ Stefanos Zafeiriou

Biosignals such as electroencephalography (EEG), electrocardiography (ECG), and electromyography (EMG) encode physiological activity across multiple temporal and spectral scales, yielding representations that are rich but challenging for machine learning. Foundation models trained to predict masked signal tokens have shown promise in learning generalizable biosignal representations, yet their performance depends on the tokenizer's ability to preserve high-frequency dynamics and reconstruct signals with high fidelity. We introduce NeuroRVQ, a modality-adaptive biosignal tokenizer family designed for high-fidelity signal reconstruction. To capture the full frequency spectrum, NeuroRVQ decomposes biosignals into frequency-specific representations via multi-scale temporal convolutions, each encoded into hierarchical RVQ codebooks to preserve high-frequency detail, combined with a novel phase-aware training loss that respects the circular topology of Fourier phase. By tuning the temporal resolution, number and size of temporal kernels and RVQ depth, this design adapts to the spectro-temporal characteristics of each biosignal modality. To validate that tokenizer quality drives downstream performance, we train a simple masked-token foundation model for each modality (NeuroRVQ-FM) using the corresponding NeuroRVQ tokenizer. The NeuroRVQ-FM family achieves competitive or superior downstream performance compared to existing modality-specific foundation models, demonstrating that high-fidelity tokenization is a critical factor for effective biosignal modeling.

Targeted Minimum Loss-based Estimation (TMLE) is a classical plug-in debiasing methodology that delivers doubly robust estimation and asymptotically efficient inference. Despite its statistical guarantees, each iteration of standard TMLE requires solving a minimization subproblem over the entire dataset, resulting in prohibitive computational and memory costs that render the method impractical for large-scale or real-time settings. To ease the bottlenecks, we introduce stochastic TMLE, a randomized targeting procedure that replaces each expensive full-batch fluctuation fit with mini-batch alternatives. Our theoretical analysis establishes that the stochastic iterates converge to a neighborhood (noise ball) centered around the target solution, and crucially, only a small number of subsequent full-sample TMLE iterations suffice to reach an empirical efficient influence function root. Consequently, our stochastic variant inherits the same guarantees and attractive properties as classical TMLE while substantially reducing per-iteration complexity. We further propose two computational variants with provable guarantees that broaden the algorithmic design space, offering flexibility for future developments in scalable targeted learning. Extensive experiments across diverse regimes demonstrate substantial acceleration over standard TMLE without sacrificing inferential quality.

Test-time compute scaling offers a promising path for improving foundation models beyond training-time data and parameters. For Vision-Language-Action (VLA) models, however, applying this paradigm through online reinforcement learning remains challenging due to two coupled limitations. Candidate chunks are often biased toward the pretrained policy's high-likelihood regions, while sparse outcome-level feedback lacks process grounding, making reward signals weakly discriminative and exploration insufficiently guided. To address these challenges, we propose EXPLORE, a test-time reinforcement learning framework that coordinates Exploitation and Exploration for action-chunk generation and drives policy optimization with physics-grounded process rewards. Through this joint design, EXPLORE broadens the effective action space without abandoning the pretrained prior and derives faithful process-level rewards from dense physical interaction feedback to guide online policy improvement. Experiments on SimplerEnv-WidowX, DexMG, and RoboCasa validate the effectiveness of EXPLORE, with consistent improvements across autoregressive and diffusion-based VLA backbones.

Foundation models offer a promising route to compress multi-modal physiological signals into better measures of human health, with broad applications across sleep medicine, cardiology, neurology and other healthcare domains. Existing models are typically trained with contrastive or masked-reconstruction objectives. However, masked reconstruction may be poorly suited to the stochastic nature of these signals, while contrastive approaches rely on positive-pair definitions despite the semantic invariances of physiological signals being poorly understood. In this work, we show that next-token prediction is a simple and scalable alternative. We develop Hypnos, a multi-modal sleep foundation model trained via next-token prediction using eight different sensing modalities (e.g. EEG, ECG, respiratory signals) from overnight polysomnography recordings. We tokenize each modality into streams of discrete tokens using residual vector quantization. We then train a large auto-regressive RQ-Transformer to jointly predict the next token across all modalities in parallel, using a novel modality masking strategy to enable generalisation to subsets of modalities during inference. Using over 20,000 overnight recordings drawn from nine public datasets, we find that both next-token perplexity and downstream probing performance continue to improve with model scale. Across a range of tasks, Hypnos matches or exceeds prior sleep foundation models and strong supervised baselines on sleep stage classification across in-domain and held-out test sets. Our results indicate that next-token prediction is a strong self-supervised objective for learning representations from multi-modal physiological signals.

It is now common to use 2D diffusion priors for problems where the underlying object is 3D (e.g., medical imaging). We would pick a noise schedule from the EDM literature, anneal a tilted prior, and then use the formulation as a 3D generative model. It works empirically. But we know very little about what it converges to and how fast. We study this question for a natural target: the multi-marginal Schrödinger bridge, the relative-entropy projection of a 3D prior onto the distribution whose rendered slices match the 2D marginals. But computing it is intractable. The annealed 2D-tilted surrogate is what one actually uses in applications. The main question is how close the latter object is to the former. Our main idea is that the discrepancy between these two laws — a comparison of two high-dimensional, fully coupled distributions — can be reduced to the size of a *single* scalar correction field. This field has a nice two-piece form: residual coupling across slices and tempering mismatch from the annealed discrete schedule. Using a Gaussian-smoothed filtration, Wasserstein screening, and conditional concentration, we obtain an explicit discrepancy bound in which the relevant terms separate into latent and posterior-residual contributions. When the smoothing scale is identified with the EDM noise level $\sigma(\tau)$, our bound yields $\mathrm{KL}(\Pi_\tau\|P^\star)=O\left(N\sigma(\tau)^2\right)$ and $\|\Pi_\tau-P^\star\|_{\mathrm{TV}}=O\left(\sqrt{N}\sigma(\tau)\right),$ if the latent scale remains bounded and the surrogate log-discrepancy is noise-aligned. So, the 2D-tilted construction converges to the Schrödinger-bridge lift at a rate governed by user-checkable quantities like the surrogate Lipschitz constants and the posterior concentration constant.

Latent signals are often obscured by measurement noise, yet encode the underlying laws and dynamics of complex systems; learning both the signals and their distributions remains a central challenge in scientific inference. The noise is often non-negligible, and the likelihoods for expressive generative models are often intractable. We utilize a convolutional maximum mean discrepancy (convMMD) loss and propose a likelihood-free framework for nonparametric density deconvolution and empirical Bayes denoising under additive measurement error. Our method learns a latent generative model by matching the observed data distribution to the noise-convolved model distribution. This yields a differentiable, simulation-based objective for multivariate homoscedastic or heteroscedastic noise, compatible with expressive sieve classes such as Gaussian mixtures and normalizing flows. The learned density then serves as an empirical prior for posterior denoising of individual latent values. Theoretically, we extend convMMD from parametric to nonparametric estimation, proving finite-sample bounds for empirical sieve minimizers and $L_2$ convergence rates under Sobolev smoothness. These rates recover the classical inverse-problem dependence: polynomial for ordinary-smooth and logarithmic for super-smooth noises. Our method provides a practical, theoretically grounded approach to deconvolution and denoising under generative latent distribution models.

Reinforcement learning with verifiable rewards (RLVR) has driven recent capability advances of large language models across domains. Recent studies suggest that improved RLVR algorithms allow models to learn effectively from incorrect annotations, achieving performance comparable to learning from clean data. In this work, we show that these findings are invalid because the claimed 100% noisy training data is "contaminated" with clean data. After rectifying the dataset with a rigorous re-verification pipeline, we demonstrate that noisy data is destructive to RLVR. We show that existing RLVR algorithm improvements fail to mitigate the impact of noisy data, achieving similar performance to that of the basic GRPO. Furthermore, we find that the model trained on truly incorrect annotations performs 8-10% worse than the model trained on clean data across mathematical reasoning benchmarks. Finally, we show that these findings hold for real-world noise in Text2SQL tasks, where training on real-world, human annotation errors cause 5-12% lower accuracy than clean data. Our results show that current RLVR methods cannot yet compensate for poor data quality. High-quality data remains essential.


Non-Expanding Gated Continued Fraction Architecture for Feedforward Layers in Language Models

Amit Dhurandhar ⋅ Vijil Chenthamarakshan ⋅ Karthikeyan Natesan Ramamurthy ⋅ Dennis Wei ⋅ Rahul Nair

Recent work has shown that continued fraction inspired neural architectures (CoFrGeNets) can be highly parameter efficient yet performant for language modeling. In this work, we revisit continued fraction-based channel mixing with the goal of improving its performance and simplifying integration into standard pre-training protocols for language models. We achieve this by two simple yet critical contributions: i) we propose a novel non-expanding gated continued fraction inspired architectural component to replace feedforward layers in established language model architectures that may use vanilla multilayer perceptrons or gated linear units or (vanilla/gated) experts in mixture-of-experts (MoE) type architectures. ii) We propose a way to re-implement divisions -- the non-linearity in these architectures -- in terms of logarithms and exponentials. By experimenting on different language modeling architectures such as GPT-2, Llama-3, Mamba-2 and (nano) Llama-MoE we find that our proposed architecture performs similarly or significantly better than the continued fraction architecture proposed in prior work along with having higher throughput. In addition, our re-implementation of divisions makes training of deeper CoFrGeNets stable without the need for incremental training which was proposed in prior work as a means to stabilize training, but added notable complexity to existing large language model training implementations.


Nonlinear Direct Feedback Alignment for Scalable Backpropagation-Free Training

Wu Lin ⋅ Yan Fan ⋅ Chenxiang Ma ⋅ Yinglan Feng ⋅ Jia Wang ⋅ Ka-Chun Wong ⋅ Qiuzhen Lin

Backpropagation (BP) remains the predominant algorithm for training deep neural networks, but it is biologically implausible and suffers from the backward locking problem, constraining parallelization across layers. Direct Feedback Alignment (DFA), which propagates the global output error directly to each hidden layer via fixed feedback matrices, has emerged as a promising alternative for addressing these limitations. However, a significant performance gap persists between DFA and BP, which we attribute to the limited ability of fixed feedback matrices to capture the nonlinear transformation from hidden representations to the network output. To address this limitation, we propose Nonlinear Direct Feedback Alignment (NDFA), which constructs each hidden-layer feedback using a lightweight nonlinear surrogate subnetwork. To enhance the effectiveness of the feedback throughout training, we further propose a learning strategy to update the surrogate subnetwork, thereby better approximating the underlying nonlinear transformation. Extensive experiments across residual networks and vision transformers on CIFAR-10, CIFAR-100, SVHN, and ImageNet datasets demonstrate that NDFA consistently outperforms existing DFA-based methods and substantially narrows the performance gap to BP, while maintaining efficient parallel training. Code will be publicly available after the review process.


Normalized Friedkin–Johnsen Opinion Dynamics

Runze Zhang ⋅ Gengyu Wang ⋅ Zhongzhi Zhang

As online social platforms increasingly serve as central arenas for opinion aggregation, numerous opinion dynamics models have been extensively developed to characterize how opinions propagate and to predict emerging social phenomena. Among various models, the Friedkin--Johnsen (FJ) model stands out as one of the most extensively studied frameworks, with broad applications across diverse domains and recent inspiration on the design of graph neural network architectures. However, the conventional FJ row-stochastic pairwise interaction has been shown to induce imbalanced influence patterns, often weakening the impact of high-degree nodes. To address this issue, we propose a normalized FJ model based on the normalized Laplacian, which symmetrically incorporates degree information from both endpoints. We establish its convergence and analyze its structural properties and intrinsic balance, along with a game-theoretic interpretation. We further develop efficient algorithms for computing equilibrium opinions and related metrics. Experiments on large-scale networks demonstrate the effectiveness and scalability of our approach.

Normalized stochastic gradient descent (Normalized SGD) is a classical scale-invariant method that updates only along the gradient direction. While Normalized SGD has recently been analyzed in non-convex stochastic settings, its convex theory remains much less developed. We address this gap by studying normalized gradient methods for smooth convex optimization. First, we obtain the theory of deterministic approaches, including first convergence results for momentum variants. Next, we establish first high-probability convergence for mini-batch Normalized SGD with and without momentum, under heavy-tailed stochastic noise with a bounded $\alpha$-central moment. Addressing the key challenge -- bias introduced by normalizing a noisy mini-batch gradient -- we demonstrate that incorporation of a new bias lemma and modified inductive proof allows to control it. As a result, we obtain first high-probability convergence of Normalized SGD in function value, together with explicit oracle complexity bounds.


Norm Anchors Make Model Edits Last

Mingda Liu ⋅ Zhenghan Zhu ⋅ Ze‘an Miao ⋅ Katsuki Fujisawa

Sequential Locate-and-Edit (L\&E) model editing can fail abruptly after many edits. We identify and formalize this failure as a positive _norm-feedback loop_, in which solved value vectors and edited MLP weights progressively amplify each other, degrading edit quality and eventually collapsing model capabilities. Our analysis shows that this feedback can yield approximately exponential norm growth under standard L\&E dynamics, and can remain unresolved by existing increment-level regularizers or update clamps. We propose $\textbf{Norm-Anchor Scaling (NAS)}$, a plug-in stabilizer that breaks this loop by rescaling each solved value vector to an original-model reference norm. Across multiple LLM backbones, datasets, and L\&E editors, NAS extends the usable editing horizon by more than $\textbf{4$\times$}$ and improves long-run editing performance by $\textbf{72.2\%}$ on average, while preserving single-edit efficacy, with only a **one-line modification and negligible computational overhead**. The code is available in the supplementary materials.


Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness

Ivo Petrov ⋅ Jasper Dekoninck ⋅ Dimitar I. Dimitrov ⋅ Martin Vechev

Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correctness alone is not sufficient: mathematical proofs should also be clear, concise, insightful, and transferable to other problems. While this proof quality is subjective and depends on the reader and context, many of its components are concrete and broadly valued. In this work, we identify such components and introduce ProofRank, a benchmark curated from challenging mathematical competitions. ProofRank evaluates several scalable proxies of proof quality: (i) conciseness, measuring whether proofs avoid unnecessary steps; (ii) computational ease, measuring the extent to which a proof relies on tedious calculations; (iii) cognitive simplicity, measuring how accessible the used proof techniques are; (iv) diversity, measuring how varied a model's proofs for a single problem are; and (v) adaptivity, measuring whether a model can follow a specified proof technique. Across models, we find substantial differences in proof quality that are not captured by correctness-only benchmarks. We also observe significant trade-offs between proof-quality metrics and correctness, suggesting that future evaluations of mathematical reasoning should measure how useful LLM-generated proofs are.

Continual tuning is essential for adapting multimodal large language models to evolving tasks and domains, yet existing methods often assume that every image-present example provides valid supervision for the visual pathway. We address the resulting gap: how to learn from all incoming multimodal instructions while preventing visually unnecessary supervision from drifting the visual interface. We propose Visual-Necessity-Gated Continual Tuning (VNG-CT), a path-aware framework that estimates sample-level dependence on visual evidence by comparing target likelihoods under the original image and a counterfactual null image. The resulting gate routes gradients to different parameter paths: all examples update the language path, visually necessary examples update the visual path, and low-necessity examples are absorbed by a lightweight calibration path that promotes invariance to irrelevant visual evidence. Across CoIN, MLLM-CL, and UCIT-style continual instruction streams, VNG-CT improves final and average performance while reducing visual forgetting. These results suggest that visual necessity offers a practical principle for stabilizing continual multimodal adaptation without discarding useful language-side supervision.


Not Just Oversmoothing: Detecting Echo Chamber Effect in Graph Neural Networks

Asela Hevapathige ⋅ Ahad N. Zehmakan ⋅ Asiri Wijesinghe ⋅ Saman Halgamuge

Oversmoothing is a well-known failure mode of Graph Neural Networks (GNNs). However, most existing diagnostics rely on global aggregation measures that fail to capture the heterogeneous dynamics of message passing. Real-world graphs exhibit pronounced community structure, and message passing operates on two timescales, where representations collapse rapidly within and slowly across communities. This creates a critical gap where intra-community representations can already be indistinguishable while inter-community separation persists, which is a failure mode we refer to as Echo Chamber Effect. To quantify this, we propose to use the Echo Chamber Index (ECI), which stratifies pairwise distances by community membership and formally establishes that global energy diminishes while inter-community separation remains less impacted. The estimated ECI for a graph reveals a surprising failure of common oversmoothing remedies, where rather than escaping the echo chamber, these architectures become permanently trapped in it. The consequences depend on label structure, as when communities align with classes, the echo chamber sharpens node classification, and when they do not, the same collapse makes it provably harder. To address this, we propose Community-Aware Split Propagation (CASP), a lightweight model-agnostic plugin that decouples intra and inter-community aggregation and learns their balance from label structure. CASP consistently improves over diverse backbone GNNs across multiple benchmarks spanning homophilic and heterophilic settings. Our source code is available at: https://anonymous.4open.science/r/CASP-3059/.


Nuwa: Evaluation-Grounded Agentic Construction of Time Series Forecasting Systems

Fanda Fan ⋅ Yuxuan Yang ⋅ Xiaorui Wang ⋅ Luqi Gong ⋅ Xiang Wentao ⋅ Xiang Ao ⋅ LimingSong ⋅ Yinuo Zhu ⋅ Yutong Chen ⋅ Chunjie Luo ⋅ Jianfeng Zhan

Time series forecasting has recently moved beyond static, single-pass prediction toward agentic methods that reason, use tools, execute workflows, and accumulate memory. However, most existing forecasting agents still optimize predictions or pipelines under largely implicit configurations, leaving the construction of the forecasting system itself underexplored. This limits their ability to attribute performance changes to specific decisions, such as which module, exogenous variable, or hyperparameter caused an improvement. We propose Nuwa, an evaluation-grounded agentic framework for constructing time series forecasting systems. Nuwa defines a structured design space over modules, exogenous variables, and hyperparameters, and performs Evaluation-Budgeted Construction that uses data meta-information, performance diagnostics, and memory to propose candidate sets, select budgeted subsets for controlled evaluation, attribute local effects, and decide whether to stop, accept, retry, switch subspace, or continue construction. The resulting design-level evidence is stored in a Design Memory, enabling reusable knowledge about which configurations work under which data conditions. Across approximately 15,000 experimental runs, Nuwa achieves consistent performance gains across module, exogenous-variable, and hyperparameter construction, with an overall average MSE reduction of 8.0\% over the corresponding reference settings. By turning forecast-level feedback into attributable and transferable design evidence, Nuwa makes forecasting system construction an explicit objective for agentic time series forecasting. The anonymous repository is available at https://anonymous.4open.science/r/NUWA/.


Object Hallucination Mitigation in Large Vision-Language Models via Self-Vision Dual Masking and Uncertainty-Triggered Assembly

Ziji Sheng ⋅ Guiyao Tie ⋅ Weidong Wang ⋅ Jiawen Shi ⋅ Junhao Dong ⋅ Xiaoye Qu ⋅ Shanshan Ye ⋅ Daizong Liu ⋅ Dengpan Ye ⋅ Pan Zhou ⋅ Jing Zhang

Large vision-language models (LVLMs) achieve strong multimodal reasoning performance but remain prone to object hallucination, where generated objects or existence judgments are not faithfully grounded in the image. Existing mitigation methods either require additional training or rely on external visual experts, which limits plug-and-play usability and increases inference overhead. In this paper, we propose SIGMA, a training-free framework for hallucination mitigation based on Self-vision Information-theoretic Gating and dual-Mask Assembly. SIGMA operates directly on the LVLM's own vision encoder, constructing complementary support and complement masks from image-driven concept relevance, question-aware entity cues, and CLS--patch shortcut suppression. These masks form global, support-masked, and complement-masked visual streams for query-conditioned evidence separation. During decoding, SIGMA activates correction only under high uncertainty, using full-vocabulary and candidate-answer entropy to perform support-first refinement and complement-on-demand suppression. The resulting streams are combined through logit-level assembly with lightweight positive calibration and a plausibility constraint. Experiments across multiple LVLM backbones and hallucination-oriented benchmarks show that SIGMA effectively reduces object hallucination while maintaining competitive perception performance and a selective low-overhead inference path.


OBJECT-UNI: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

Mining Tan ⋅ yinuo Wang ⋅ Ziqi Zhou ⋅ Weize Quan ⋅ Sifei Li ⋅ Jingdong Chen ⋅ DanDan Zheng ⋅ libin wang ⋅ Weiming Dong

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipulate the spatial states of object instances. Existing models can describe objects in natural language, but they struggle to precisely represent continuous object poses and generate geometrically consistent images under target viewpoints. To mitigate this, we propose \emph{Object-Uni}, a unified model for object-centric spatial understanding and controllable generation. Specifically, we formulate object-centric spatial intelligence as a unified problem connecting pose perception, spatial reasoning, pose-conditioned generation, and object-centric novel view synthesis. We treat object pose as an explicit geometric variable shared by understanding and generation, rather than merely a prediction label or control signal. To make pose usable by multimodal large language models, we propose a viewpoint-based orientation abstraction that maps orientation into structured viewpoint descriptions while preserving continuous geometric supervision. We further construct an object-centric spatial benchmark (UniSpatial-80K) and train a unified model with an object-token-grounded pose anchor to associate each instance with its pose state. Experiments show that our model improves object-level pose understanding and pose-controllable generation, moving unified models from describing objects toward manipulating spatial states.

Extreme ocean phenomena are challenging not only to predict but to diagnose, as accurate forecasts alone do not reveal the underlying physical drivers. While recent machine learning approaches achieve strong predictive skill, they remain largely opaque and provide limited guarantees of fidelity to ground-truth physics. We introduce OceanCBM, the first concept bottleneck model (CBM) for spatiotemporal prediction and mechanistic interrogation of ocean dynamics. OceanCBM uses mixed supervision to predict mixed layer heat content, a key precursor of marine heatwaves, while routing information through an intermediate layer of prescribed concepts derived from geophysical fluid dynamics and a 'free' concept. This design imposes soft physical structure without over-constraining the model, and the free concept both regularizes concept predictions and captures residual physical processes. Across ensemble initializations, we show that mixed supervision yields consistent mechanistic representations, whereas prediction-only and prescription-only baselines learn highly variable latent structures despite similar predictive performance. OceanCBM achieves interpretable, physically grounded representations without sacrificing skill, explicitly characterizing the interpretability-performance trade-off.


ODEWorld: A Continuous Predictive Architecture via Physical-Time Flow

Dongxiu Liu ⋅ Haoyi Niu ⋅ Peng Cheng ⋅ Yuan Gao ⋅ Xirui Kang ⋅ Sangli Teng ⋅ Koushil Sreenath ⋅ Xianyuan Zhan

In the physical world we inhabit, space and time are fundamentally continuous. However, existing machine learning paradigms for world modeling are largely confined to discrete-time prediction and inference, thereby exhibiting significant inefficiency in capturing the fundamental dynamics of the physical world. To bridge this gap, we introduce Physical-Time Flow (PT-Flow), a novel approach that learns a continuous latent velocity field operating in physical time. Crucially, the underlying dynamics of sequential data are parameterized by an ordinary differential equation (ODE) embedded in a well-structured dynamical representation space. Under this paradigm, the prediction of the future can be recast as temporal integration via an ODE solver in the compressed latent space. Building upon PT-Flow, we construct ODEWorld, a physics-grounded, continuous-time latent world model that is both efficient and versatile. By extracting time-variant features and enforcing ODE properties on both the dynamical representation space and the latent velocity field, ODEWorld effectively addresses the long-standing representation collapse issue in latent world model literature. This also enables high-quality image reconstruction with ODEWorld even after long-horizon prediction. Moreover, its continuous nature allows for arbitrary temporal resolution and even backward prediction, which is impossible for most discrete-time models. Lastly, benefiting from the learned dynamics-centric latent space, ODEWorld can provide rich planning-oriented information to facilitate downstream policy learning. Comprehensive experiments across multiple simulation and real-world datasets demonstrate that ODEWorld successfully reconciles planning-conducive dynamics abstraction with visual realism, excelling in both video generation and robotic control. More qualitative results can be found at Project Website.


Offline Materials Optimization with CliqueFlowmer

Jakub Grudzien Kuba ⋅ Benjamin K Miller ⋅ Sergey Levine ⋅ Pieter Abbeel

Recent advances in deep learning have inspired neural network-based approaches to computational materials discovery (CMD). A plethora of problems in this field involve finding materials that optimize a target property. Nevertheless, the increasingly popular generative modeling methods are ineffective at boldly exploring attractive regions of the materials space due to their maximum likelihood training. In this work, we offer an alternative CMD technique based on offline model-based optimization (MBO) that fuses direct optimization of a target material property into generation. To that end, we introduce a domain-specific model, CliqueFlowmer, that incorporates recent advances in clique-based MBO into transformer and flow generation. We validate this model's optimization abilities and show that materials it produces strongly outperform those from generative baselines. To support specialized materials discovery applications and broader interdisciplinary research, we will release our code and model weights upon publication.


OmniSE: Support-Faithful Optimization for Evidence-Centric Audio-Video Reasoning

Xilin He ⋅ Yifan Shen ⋅ Ankan Deria ⋅ Xiangyu Yue ⋅ Muhammad Haris Khan

Omni-modal models are expected to reason over synchronized visual events and audio cues, yet answer accuracy alone can conceal failures in where an answer is grounded and which modality supports it. In audio-video question answering (AVQA), models may produce correct answers while relying on incorrect temporal spans, dominant-modality shortcuts, or unsupported cross-modal relations. We argue that this failure arises from treating reasoning as a completion-level prediction problem, where supervision and reward signals do not distinguish between support selection and reasoning over the support. To address this, we introduce OmniSE, a support-faithful optimization framework that makes question-conditioned audio-video support the central training object. Each QA instance is rewritten into an Audio-Visual Evidence Trace, a compact optimization unit that records temporal evidence, modality-specific cues, cross-modal relations, and cited reasoning steps. Given this support object, OmniSE first uses Support-Tree Rollout (STR) to explore alternative support hypotheses before sampling support-conditioned reasoning and answers. It then applies factorized reward assignment to score temporal localization, modality evidence, cross-modal binding, reasoning faithfulness, and answer correctness at their corresponding comparison levels. Finally, reward-gated self-consistency regularizes sibling continuations under reliable support, encouraging fixed trustworthy evidence to induce stable reasoning and answers. Experiments on multiple AVQA benchmarks show that OmniSE consistently outperforms strong open-source baselines at the same scale, demonstrating the importance of support-centric optimization for evidence-centric audio-video reasoning. Code and data would be publicly available after peer review.


One Adapter For All: Towards Generalizable Many-To-One Domain Adaptation In Heterogeneous Collaborative Perception

Lu Bin ⋅ Zhewei Fu ⋅ Shaohong Wang ⋅ Zhiyu Xiang ⋅ Hangguan Shan ⋅ Eryun Liu

In autonomous driving, multi-agent collaboration enhances individual perception capabilities through information sharing. However, in real-world applications, differences in models and training data among heterogeneous agents inevitably lead to domain gaps between shared features, bringing challenges to collaboration. To address this, existing methods typically rely on a one-to-one adaptation paradigm, necessitating specific training or fine-tuning for every new agent. This results in a large accumulation of adapters on the vehicle, which is unsustainable for resource-constrained edge devices. To overcome these limitations, we propose UniMTO, a Unified adaptation framework that enables Many-To-One adaptation across heterogeneous agents using a single, fixed adapter. Specifically, it treats domain adaptation as a pattern transformation task, disentangling domain-invariant content and domain-specific pattern from features, and achieving source-to-target pattern transformation through recombination based on a Mixture-of-Experts (MoE) mechanism. Extensive experiments on V2V4Real and V2X-Real demonstrate that UniMTO achieves superior generalization, and exhibits strong zero-shot adaptation capabilities, outperforming all state-of-the-art methods, enabling many-to-one domain adaptation with a single, fixed adapter.

Order dispatch is a critical task in ride-sharing systems with Autonomous Vehicles (AVs), directly influencing operational efficiency and profitability. While Multi-Agent Reinforcement Learning (MARL) offers a scalable paradigm by decomposing the massive state-action space, existing methods are heavily reliant on accurate value function estimation. In large-scale, highly stochastic urban environments, such estimation is notoriously prone to bias and instability, severely limiting training efficiency and policy quality. To overcome this limitation, we propose two novel policy optimization methods that completely bypass the need for critic networks. First, we establish a formal connection between the homogeneity of AV fleets and the short-memory mixing property of the transportation network, proving that agent value functions converge to a common time-varying baseline up to a negligible residual. Leveraging this insight, we introduce Single-Trajectory Group Relative Policy Optimization (ST-GRPO), an adaptation of Large Language Models (LLMs) post-training techniques to multi-agent trajectories, which replaces the traditional value baseline with the fleet-wide average reward-to-go. Inspired by this reduction, we further derive One-Step Policy Optimization (OSPO), demonstrating that under the established structural priors, an optimal policy can be learned using only immediate, group-normalized rewards—rendering long-horizon bootstrapping unnecessary. Experiments on real-world ride-hailing datasets from Manhattan and Queens demonstrate that both ST-GRPO and OSPO achieve promising performance, particularly in reducing pickup times and increasing order service rates. Remarkably, both methods operate efficiently using simple Multilayer Perceptron (MLP) networks with minimal GPU utilization, underscoring the practical impact of exploiting domain structure for scalable MARL. Our code, trained models, and processed data are provided at the anonymous repository: https://anonymous.4open.science/r/OSPO-2105 .

Brain–computer interfaces (BCIs) must continually adapt as neural signals drift, yet labeled calibration data is expensive to collect at time of use. Recent systems address this challenge by updating neural decoders using language-model–pseudo-labels, but they treat the pseudo-label as ground truth. We reinterpret recalibration as approximate Bayesian filtering: the decoder parameters are the latent states, and language models provide observations in the form of unnormalized potentials. Guided by this analysis, we introduce \textbf{B-CORP} (Bayesian Continual Online Recalibration with Pseudo-labels). B-CORP replaces hard pseudo-labels with a soft top-$K$ marginalization, increases the number of iterations in the coordinate ascent on the joint likelihood, and models parameter drift with a prior on latent dynamics. Across synthetic drifting regression tasks and real BCI benchmarks, including handwriting and speech BCI datasets, B-CORP attains word error rates consistently under 3\% on a difficult handwriting BCI dataset. B-CORP improves almost 30\% versus the state-of-the-art, while retaining performance even with poor pseudo-labels where naive methods fail. More broadly, our approach connects pseudo-label recalibration with online Bayesian inference, providing a principled foundation for problems requiring long-term recalibration with access to black box predictive models, such as LLMs, for use as strong priors over the labels.


Online Minimum Description Length Passive-Aggressive Algorithms

John Hurwitz ⋅ Charles K Nicholas ⋅ Edward Raff

We present an algorithm for margin-based online learning using lossless compressors (like zip) via the Minimum Description Length Principle. The technique, we call \emph{MDL-PA}, can be viewed as an analogue of Passive-Aggressive algorithms for probability distributions (equivalently viewed as lossless compression schemes) where we consider a \emph{code length margin}. Separate models are maintained per class, and the update rule is an information projection correcting the true-class model's prediction on a margin-violating sample. To bridge theory and practice, we demonstrate a practical approximation to this online learning setup via adaptive Lempel-Ziv-style compression dictionaries to classify sequences in the text and malware domains.

We propose and study an online version of min-max optimization based on cumulative saddle points under a variety of performance measures beyond convex-concave settings. After first observing the incompatibility of (static) Nash equilibrium (SNE-Reg$_T$) with individual regrets even for strongly convex-strongly concave functions, we propose an alternate static duality gap (SDual-Gap$_T$) inspired by the online convex optimization (OCO) framework that is compatible with the individual regrets . We provide algorithms that achieve sub-linear regret bounds for (SDual-Gap$_T$) and the individual regrets and a novel dynamic saddle point regret (DSP-Reg$_T$), which we suggest naturally represents a min-max version of the dynamic regret in OCO. We derive our bounds for (SDual-Gap$_T$) and DSP-Reg$_T$ under strong convexity-strong concavity and a min-max notion of exponential concavity (min-max EC), and in addition we establish a class of functions satisfying min-max EC that captures a two-player variant of the classic portfolio selection problem. Finally, for a dynamic notion of regret compatible with individual regrets, we derive bounds under a two-sided Polyak-\L{}ojasiewicz (PL) condition.


On the Convergence of Multicalibration Gradient Boosting

Daniel Haimovich ⋅ Fridolin Linder ⋅ Lorenzo Perini ⋅ Niek Tax ⋅ Milan Vojnovic

Multicalibration gradient boosting has recently emerged as a scalable method that empirically produces approximately multicalibrated predictors and has been deployed at web scale. Despite this empirical success, its convergence properties are not well understood. In this paper, we provide computational guarantees for multicalibration gradient boosting algorithms. We show that the magnitude of successive prediction updates decays at $O(1/\sqrt{T})$, which implies the same convergence rate bound for the empirical multicalibration error over rounds. Under additional smoothness assumptions on the weak learners, this rate improves to linear convergence. We further establish convergence for adaptive variants. Experiments on real-world datasets support our theory and clarify the regimes in which the method achieves fast convergence.


On the Depth of Monotone ReLU Neural Networks and ICNNs

Egor Bakaev ⋅ Florestan Brunck ⋅ Christoph Hertrich ⋅ Daniel Reichman ⋅ Amir Yehudayoff

We study two models of $\mathsf{ReLU}$ neural networks: monotone networks ($\mathsf{ReLU}^+$) and input convex neural networks ($\mathsf{ICNN}$). Our focus is on expressivity, mostly in terms of depth and our motivation is to gain better understanding of depth requirement needed to exactly represent functions in terms of neural networks with ReLU activations: a subject that has received a lot of attention lately. We prove several lower bounds on the depth required to represent functions by monotone ReLU networks and $\mathsf{ICNN}$. For the maximum function $\mathsf{MAX}_n$ computing the maximum of $n$ real numbers, we show that $\mathsf{ReLU}^+$ networks cannot compute $\mathsf{MAX}_n$, or even approximate it. We prove a sharp $n$ lower bound on the $\mathsf{ICNN}$ depth complexity of $\mathsf{MAX}_n$. We also prove depth separations between $\mathsf{ReLU}$ networks and $\mathsf{ICNN}$s; for every $k$, there is a depth-$2$ $\mathsf{ReLU}$ network of size $O(k^2)$ that cannot be simulated by a depth-$k$ $\mathsf{ICNN}$. The proofs combine ideas from some structural results for $\mathsf{ReLU}^+$ networks.


On the Effect of Token Correlations on Semantic Transfer in Transformers

Nevena Lazic ⋅ Julian Zimmert ⋅ Alan Malek ⋅ Tor Lattimore ⋅ Csaba Szepesvari ⋅ András György

We study issues which may arise when using transformer models to formally reason about problems stated informally using natural language. In particular, we focus on transfer learning between tasks that are semantically equivalent but phrased differently. We assume that examples with a particular target label are only available for a subset of tasks, and aim to transfer zero-shot the ability to predict this label to other tasks. Using controlled experiments on propositional logic problems, we show that transformers are capable of such transfer, but that it can be impeded by the presence of task-specific tokens, due to the model picking up on correlations between these tokens and the label early in the training. As a simple solution, we show that introducing the target label later in the training dramatically improves transfer.


On the Faithfulness of Visual Thinking: Measurement and Enhancement

Zujing Liu ⋅ Junwen Pan ⋅ Qi She ⋅ Yuan Gao ⋅ Gui-Song Xia

Recent large vision–language models (LVLMs) can generate vision–text multimodal chain-of-thought (MCoT) traces after reinforcement fine-tuning (RFT). However, we observe that the visual information in MCoT is often inaccurate even when answers are correct, and task accuracy remains nearly unchanged despite the visual information being corrupted. This indicates a lack of faithfulness in the vision part of MCoT reasoning. We attribute this to the RL reward design in RFT, which solely incentivizes the format of interleaved vision-text cues, encouraging the model to incorporate visual information into its text reasoning steps without considering its correctness. In this paper, we first probe MCoT faithfulness by measuring how much the prediction changes when its visual and textual thoughts are intervened. Surprisingly, the model's predictions remain nearly unchanged under visual intervention but change significantly under textual intervention, indicating that the visual evidence is largely ignored. To further analyze the visual information, we introduce an automated LVLM-based evaluation metric that quantifies the faithfulness of visual cues from two perspectives: relevance and sufficiency. Our evaluation reveals that the visual information in current MCoT traces is simultaneously irrelevant and insufficient. To address this issue, we propose Sufficient-Component Cause Model (SCCM) learning. This approach encourages the MCoT to generate sufficient yet minimal visual components that are independently capable of leading to the correct answer. The proposed SCCM is annotation-free and plug-and-play, compatible with various RFT in MCoT. Empirical results demonstrate that SCCM consistently improves the visual faithfulness across a suite of fine-grained perception and reasoning benchmarks.


On the Importance of Gating: Memorization vs. In-Context Learning in State Space Models

William Tong ⋅ Aryo Lotfi ⋅ Emmanuel Abbe ⋅ Konstantinos Vaggelakos ⋅ Vishnu Banna ⋅ Etai Littwin ⋅ Joshua Susskind ⋅ Eran Malach

State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear compute. While SSMs excel in certain settings like length generalization, they continue to lag behind Transformers on tasks that require in-context learning and precise retrieval, slowing their adoption for large-scale language modeling. In this work, we demonstrate that both the success and failure of SSMs in these domains can be explained by studying the role of the gating mechanism, a prevalent component in modern recurrent networks. Specifically, we show through theory and experiments that this gating mechanism causes SSMs to first learn an in-weights ``memorization'' solution, while delaying, or even preventing, convergence to a correct in-context learning solution. Importantly, this happens even in cases where there are no fundamental limitations due to the architecture or its memory capacity. On the other hand, we find that gating is often beneficial for improving generalization to long sequence lengths. Our results illuminate the crucial role of the gating mechanism in shaping both the training dynamics and generalization of SSMs, and provide a basis for understanding and improving linear-time models.


On the Origin of Algorithmic Progress in AI: Evidence from Language Model Pre-Training

Hans Gundlach ⋅ Alex Fogelson ⋅ Jayson Lynch ⋅ Ana Trisovic ⋅ Jonathan S Rosenfeld ⋅ Anmol Rattan Singh Sandhu ⋅ Neil Thompson

Algorithms have been estimated to increase AI pretraining FLOP efficiency by a factor of $22,000$ between 2012 and 2023 (Ho et al., 2024) . Running ablation experiments on key innovations from this time period, we are able to account for less than $10\times$ of these gains at small compute scales. Surveying the broader literature, we estimate that additional innovations not included in our ablations also account for less than $10\times$, yielding less than $100\times$ efficiency gains total in low-compute regimes. This leads us to conduct scaling experiments, which reveal that much of this efficiency gap can be explained by algorithms with scale-dependent efficiency improvements. In particular, we conduct scaling experiments between LSTMs and Transformers, finding exponent differences in their compute-optimal scaling law while finding little scaling difference for many other innovations. These experiments demonstrate that -- contrary to standard assumptions -- an algorithm's efficiency gains are tied to compute scale. Using experimental extrapolation and literature estimates, we account for $3,360\times$ efficiency gains over the same time period, with scale-dependent innovations accounting for close to 90\% of log-efficiency gains by 2023. Our results indicate that algorithmic progress for small models has been far slower than previously assumed, and that measures of algorithmic efficiency are strongly reference-dependent.


On the Price of Privacy for Language Identification and Generation

Xiaoyu Li ⋅ Andi Han ⋅ Jiaojiao Jiang ⋅ Junbin Gao

As large language models (LLMs) are increasingly trained on sensitive user data, understanding the fundamental cost of privacy in language learning becomes essential. We initiate the study of differentially private (DP) language identification and generation in the agnostic statistical setting, establishing algorithms and matching lower bounds that precisely quantify the cost of privacy. For both tasks, approximate $(\varepsilon, \delta)$-DP with constant $\varepsilon > 0$ recovers the non-private error rates: $\exp(-r(n))$ for identification (for any $r(n) = o(n)$) and $\exp(-\Omega(n))$ for generation. Under pure $\varepsilon$-DP, the exponents degrade by a multiplicative factor of $\min\\{1, \varepsilon\\}$, which we show is tight up to constants. Notably, for generation under pure DP with mild assumptions, the upper bound $\exp(-\min\\{1,\varepsilon\\} \cdot \Omega(n))$ matches the lower bound up to some constants, establishing an optimal rate. Our results show that the cost of privacy in language learning is surprisingly mild: absent entirely under approximate DP, and exactly a $\min\\{1,\varepsilon\\}$ factor in the exponent under pure DP.

In semantic segmentation, a recent line of RankSEG methods directly optimizes Dice/IoU scores at inference time, improving alignment with evaluation metrics without modifying model training. Despite its theoretical and empirical success, RankSEG relies on the restrictive Conditional Independence Assumption (CIA), which ignores crucial label correlations and therefore degrades performance in ambiguous or low-contrast scenarios. However, accounting for full label dependence is computationally prohibitive, requiring $\mathcal{O}(d^3)$ time. To address this, we replace the CIA with a Spatially Localized Dependence (SLD) structure that captures local label correlations while keeping the dependence model tractable. We further overcome the remaining computational bottleneck via a Reciprocal Moment Approximation coupled with a novel fixed-point optimization strategy that eliminates exhaustive search. The proposed algorithm achieves a highly practical $\mathcal{O}(d \log d)$ complexity and consistently outperforms conventional argmax and CIA-based RankSEG across diverse segmentation benchmarks. Improvements are significant in low-contrast or small-object scenarios, where label dependence offers valuable signals complementary to image information for accurate segmentation.

Dictionary learning has long been studied from both optimization and probabilistic perspectives. While formulations with element-wise sparsity regularization (e.g., L1-based sparse coding) admit well-established probabilistic interpretations, many structured variants that impose global constraints lack a clear and tractable generative view. In this paper, we revisit a class of practically effective yet theoretically under-explored dictionary learning methods that impose a simple global regularization on the number of activated dictionary atoms, which we term parsimoniously activated dictionary learning (PADL). We show that PADL admits an equivalent formulation as maximum a posteriori estimation under a structured generative model, with auxiliary latent variables that govern global activation patterns. This formulation allows us to derive generalization guarantees that are difficult to obtain under the original formulation. More importantly, it yields an analytical characterization of the tradeoff between sparsity, storage cost, and reconstruction accuracy, enabling data-driven estimation of optimal hyperparameters. Based on this connection, we develop an efficient and interpretable PADL algorithm that eliminates manual hyperparameter tuning, achieving improved reconstruction performance under comparable sparsity levels on visual benchmarks. We further demonstrate its practical utility in accelerating inference for vision-language models.


On the Tightness and Computational Tractability of Higher-Dimensional Confidence Sequences

Fabian Denoodt ⋅ Sibylle Hess ⋅ Joaquin Vanschoren ⋅ Christian Andersson Naesseth

Modern sequential monitoring problems often involve multiple metrics, where we monitor several data streams simultaneously and may act once the evidence is strong enough. Confidence sequences (CSs) are a natural tool for such continuous monitoring. However, for bounded vector means, existing multivariate CSs are either tight but computationally intractable, or fast to compute but conservative. To address this, we study three lifts of one-dimensional betting-based CSs to higher dimensions: a weighted Bonferroni region, an equivalent max-wealth form, and a portfolio region. The portfolio is typically much tighter, especially in higher dimensions, but its boundary and properties such as volume are not available in closed form. To make this tighter construction usable, we propose tractable outer approximations of the portfolio region that preserve statistical validity: a bounding box, an $\ell_p$-ellipsoid, and their intersection. We prove set relations among all constructions and show empirically that these approximations (i) achieve regions close to the intractable portfolio, (ii) substantially outperform existing tractable multivariate CSs, and (iii) enable practical use cases such as multi-metric A/B testing and model comparison.


OpenMHC: Accelerating the Science of Wearable Foundation Models

Narayan Schuetz ⋅ Yuze Bai ⋅ Lianggang Pan ⋅ Edgar Eggert ⋅ Favour Nerrise ⋅ Juan Delgado-SanMartin ⋅ Max Rosenblattl ⋅ Milana Gurbanova ⋅ Mohammad Asadi ⋅ Anders Johnson ⋅ Paul Schmiedmayer ⋅ Dennis Wang ⋅ Allan Lawrie ⋅ Daniel S Kim ⋅ Xin Liu ⋅ Akshay Paruchuri ⋅ Ehsan Adeli ⋅ Euan Ashley ⋅ Kelly W Zhang

Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching. However, the largest wearable datasets are not publicly available for research, and leading wearable foundation models trained on such datasets are rarely open-weight or come with reproducible training code. To accelerate open science in wearable health, we release OpenMyHeartCounts (OpenMHC), the largest and most comprehensive open-access wearable health dataset to-date, alongside open-source implementations of recent wearable foundation models. \openmhc{}, derived from over a decade of data collected through the My Heart Counts study app, includes >60 million hours of wearable data across 19 sensor channels (e.g., step count, heart rate, sleep, workouts) and up to 169 linked variables, including health, lifestyle, mood, and behavior from 11,894 consenting participants. Furthermore, we introduce a unified, open benchmark that enables standardized comparison of wearable health models across three tracks: health and behavior downstream prediction, multivariate data imputation, and time-series forecasting. We benchmark classical methods alongside recent wearable and multivariate time series foundation models. By open-sourcing data, code, and model weights at this unprecedented scale, we aim to democratize wearable health AI research and enable the community to drive open progress in this domain.


OpenView: Empowering MLLMs with Out-of-view VQA

Qixiang Chen ⋅ Cheng Zhang ⋅ Chi-Wing Fu ⋅ Jingwen Ye ⋅ Jianfei Cai

Recent multimodal large language models (MLLMs) show great potential in natural image understanding. Yet, they perform well, mainly on reasoning in-view contents within the image frame. This paper presents the first study on out-of-view (OOV) understanding, i.e., the ability to reason about objects, activities, and scenes beyond the visible frame of a perspective view. Our technical contributions are threefold. First, we design OpenView, a four-stage pipeline that leverages panoramic imagery for large-scale and diverse multi-choice VQA synthesis, where full-scene coverage enables flexible view framing while preserving global scene awareness. Second, we curate OpenView-Dataset, a high-quality synthetic dataset from diverse real-world panoramas to empower MLLMs upon supervised fine-tuning. Third, we build OpenView-Bench, a benchmark that jointly measures choice and rationale accuracy for interpretable and diagnosable evaluation. Experimental results show that despite having a large gap from human performance in OOV VQA answer selection, multiple MLLMs, upon empowered by OpenView, can consistently boost their performance, uplifted from 48.3% to 62.3% on average. Code, benchmark, and data will be released.

Reinforcement learning methods for LLM reasoning typically assign a uniform advantage to every token in a trajectory, diluting the learning signal at pivotal reasoning steps and injecting noise at uninformative positions. Critic-free alternatives such as anchored distillation derive per-token signals from oracle-conditioned likelihood ratios, but apply the signal independently at each position. We propose Oracle-Prompted Policy Optimization (\n), which recovers exact token-level advantages through a Bayesian value recursion. Conditioning the model on the ground-truth answer during training yields per-token likelihood ratios that measure how much the answer revises the model's prediction; the ratios accumulate via Bayesian updating into a running estimate of the success probability at every position, from which the token-level advantage follows in closed form. The advantage factors as a product of a state weight, peaking where the outcome is most uncertain and vanishing where success or failure is already determined, and the per-token oracle evidence familiar from on-policy distillation, with no tunable weighting hyperparameter. The framework supports two estimator choices: a self-oracle mode in which the policy model provides the estimate through one additional forward pass, and a teacher-oracle mode in which a stronger model serves as the estimator. Experiments on two base models across seven reasoning benchmarks spanning mathematics, science, and code demonstrate consistent improvements over GRPO, DAPO, and SDPO, with the largest gains on competition-level tasks where reasoning chains are longest.

Characterizing the features of a Hamiltonian that governs a quantum system serves as a fundamental subroutine of quantum device calibration, signal sensing, and error correction. Recent works proposed protocols have achieved the optimal Heisenberg-limited scaling learning ansatz-free Hamiltonians from their real-time evolutions without fully specifying interaction structures. However, these protocols relies on both deep circuits with interleaving probes and control, and extremely short time resolution, making them difficult to implement on near- and intermediate-term in situ quantum experiments. In this work, we propose a computational efficient, *control-free*, and *ancilla-free* algorithm that uses only *Pauli product state preparation and measurement*, and learns an ansatz-free Hamiltonian $H$ with $||H||\leq\Lambda$ in total evolution time of $\Theta(\tfrac{\Lambda}{\epsilon^2}\log(\tfrac{\Lambda}{\epsilon}))$. The evolution time cost of our algorithm is *optimal* for any control-free protocols as we further prove a lower bound of $\Omega(\tfrac{\Lambda}{\epsilon^2}\log(\tfrac{\Lambda}{\epsilon}))$. Technically, our method introduces a randomized-sampling framework that combines band-limited kernel-based time sampling with a displacement sieve for Hamiltonian structure learning. The characteristic probe time resolution depends only on $\Lambda$ instead of $\varepsilon$, which makes our protocol especially appealing in the high-precision regime for sensing and calibration applications.


Optimal Representation Size: High-Dimensional Analysis of Pretraining and Linear Probing

Valentina Njaradi ⋅ Clémentine Dominé ⋅ Rachel A Swanson ⋅ Marco Mondelli ⋅ Andrew Saxe

Learning to generalise from limited data is a fundamental challenge for both artificial and biological systems. A common strategy is to extract reusable structure from abundant unlabelled data, enabling efficient adaptation to new tasks from limited labelled data. This two-stage paradigm is now standard in modern training pipelines, where pretraining is followed by fine-tuning or linear probing. We provide an analytical model of this process: structure extraction is formalized as principal component analysis on unlabelled data, and downstream learning as linear regression on a separate labelled dataset. In the high-dimensional regime, we derive exact expressions for training and generalisation error showcasing their dependence on representation dimensionality, unlabelled and labelled sample sizes, and task alignment. Our results show that pretrained representations strongly influence downstream generalisation, and we characterize the optimal representation size as a function of task parameters: with abundant pretraining data but scarce downstream data, maximally compressed representations are optimal, whereas with limited pretraining data, higher-dimensional representations generalise better. Furthermore, we establish an exact trade-off between pretraining and supervision, quantifying how much unlabelled data is required to replace a single labelled sample. Beyond our idealised model, we observe similar phenomenology in autoencoders and pretrained LLMs. Altogether, we highlight that optimising representation size is critical, giving conditions for when compression during pretraining improves generalisation.

Offline-to-online reinforcement learning (O2O RL) aims to improve offline RL agents with limited online interaction. Prior work has shown that, in conservative value-regularized methods such as CQL, pessimistic offline critics can underestimate values and hinder online fine-tuning. We show that this issue is more general: critic underestimation also appears in explicit policy-constraint methods such as TD3+BC, and can persist or become more severe through bootstrapped Bellman updates during fine-tuning. To address this problem, we propose Optimistic Q-value Adaptation (OQA), a framework that corrects underestimation at the Bellman-backup level. OQA constructs an additional offline-pretrained actor--critic pair and uses it together with the original pair to form a controlled optimistic backup that mitigates inherited pessimism during fine-tuning. Unlike trajectory-return calibration methods, OQA does not rely on return lower bounds and is not tied to CQL-style regularization. Across D4RL MuJoCo locomotion, Maze2D, AntMaze, and Adroit benchmarks, OQA instantiations based on TD3+BC and CQL consistently outperform the corresponding backbones and strong O2O RL baselines.


OrbitLoRA: Learning Rotation-Aware Low-Data Adaptation of Vision Foundation Models

Yunlu Chen ⋅ Dominik Engel ⋅ Peter Wonka ⋅ Ivan Viola

Pretrained vision transformers provide strong semantic representations, but their token embeddings are not designed to transform predictably under image rotations. This limits their direct adaptation to domains such as satellite imagery, histopathology, microscopy, and other scientific imaging settings, where in-plane orientation is often arbitrary and no canonical pose exists. Existing equivariant adaptation methods for pretrained models typically rely on frame averaging or canonicalization, which can be effective on synthetically rotated natural images but do not teach the adapted model an internal rotation-aware representation. We introduce OrbitLoRA, a lightweight orbit-aware adapter for frozen visual foundation models. OrbitLoRA combines standard LoRA with a typed residual branch that maintains scalar and vector fields over ViT patch tokens. The typed branch is initialized from local steerable image features and harmonic coordinate fields, updated at selected transformer depths, and injected into the frozen token stream through zero-initialized residual gates. During finetuning, paired rotated views supervise the typed state through a transport-and-rotation consistency loss, while the semantic output is trained to remain invariant for the downstream task. This design preserves the rich features of pretrained ViTs while adding a learnable rotation-symmetry prior. Across classification and segmentation tasks, and across DINOv2-, CLIP-, SAM-, and LLaVA-style backbones, OrbitLoRA matches strong LoRA-with-augmentation baselines in high-data regimes and substantially improves adaptation in low-data regimes, where learning the symmetry prior is most valuable.


OrchestraRL: Learning to Orchestrate LLM Agent Swarms with Entropy-Aware Communication Control

Runze Fan ⋅ Shiyi Jiang ⋅ Junsheng Wang ⋅ Shuangjun Xie ⋅ Xiaodan Li ⋅ Yong Li

Multi-agent LLM systems built on frontier backbones often fail to improve reliably as the swarm grows: under strong backbone models, performance frequently plateaus or declines beyond a small number of agents. In a budget-matched controlled intervention on SWE-bench Verified ($N{=}8$, Claude Opus 4.5 backbone, 64-call budget per instance), varying only the communication policy moves resolve rate from 80.4\% under random broadcast coordination (below the 80.9\% single-shot baseline) to 85.5\% under a learned orchestrator, a 5.1-point spread under matched compute. We introduce OrchestraRL, an orchestration layer trained over frozen LLMs. A centralized Macro-Orchestrator reads the swarm's context entropy $H(t)$, computed from the spread of agent output embeddings, and outputs communication mode (via learned entropy thresholds) and topology; per-agent Micro-Coordinators decide whom each agent addresses, when to speak, and at what granularity. Both are trained with REINFORCE and a learned value baseline. Across four benchmarks (SWE-bench Verified, GPQA Diamond, LiveCodeBench v6, SimpleQA), OrchestraRL is comparable to the strongest baseline at $N{=}2$ and achieves the best resolve rate at $N \in \{4, 8\}$ on every benchmark, with the gap widening as $N$ grows. The learned policy recovers interpretable task-dependent communication patterns, including clustered exploration, star-based synthesis, and pipeline reasoning, without these structures being hand-coded.


Orlicz–Sobolev with Musielak: An Efficient Regularization Approach for Graph-based IPM

Tam Le ⋅ Truyen Nguyen ⋅ Hideitsu Hino ⋅ Kenji Fukumizu

We study the Sobolev IPM problem for probability measures supported on a graph metric space, where critic function is constrained to lie within the unit ball defined by Sobolev norm. Sobolev IPM is intrinsically coupled with $L^p$ geometric structure within its definition, limiting its ability to incorporate other prior geometry beyond the $L^p$ paradigm. Conversely, the classic optimal transport is flexible, easy to adapt to various geometric structures by simply changing its ground cost. An important example is Orlicz-Wasserstein (OW) which utilizes \emph{Orlicz} geometric structure to generalize the $L^p$ within the standard $p$-order Wasserstein, and remarkably play a vital role to advance machine learning methodologies. Inspired by recent advantages of OW, in this work, we leverage a specific class of convex functions for Orlicz geometry to mitigate such limitation for Sobolev IPM, and propose the generalized Sobolev IPM (GSI). Our GSI approach encompasses Sobolev IPM as a special case while accommodating diverse geometric priors beyond $L^p$. It however brings up significant computational hurdles that compound those already notoriously inherent in Sobolev IPM. To address these challenges, we theoretically establish a novel connection between \emph{Orlicz-Sobolev} norm and \emph{Musielak} norm which facilitates a novel efficient regularization for GSI. By further exploiting the underlying graph structure, we show that the regularized GSI reduces to a simple univariate optimization problem, achieving notably computational efficiency, enabling its usage in practical applications. We empirically illustrate that the regularized GSI is several-order faster than the popular OW in computation, and its performances compare favorably to other transport baselines for comparing graph-based measures in document classification and topological data analysis.


Orthogonal Uplift Learning with Permutation-Invariant Representations for Combinatorial Treatments

Jiacan Gao ⋅ Xinyan Su ⋅ Mingyuan Ma ⋅ Xiao Xu ⋅ Xinrui Wan ⋅ Tianqi Gu ⋅ Enyun Yu ⋅ Jiecheng Guo ⋅ zhiheng zhang

We study uplift estimation for combinatorial treatments. Uplift measures the pure incremental causal effect of an intervention, such as sending a coupon or a marketing message, on user behavior, modeled as a conditional individual treatment effect. Many real-world interventions are combinatorial: a treatment is a policy that specifies context-dependent action distributions rather than a single atomic label. Although recent work considers structured treatments, most methods rely on categorical or opaque encodings, limiting robustness and generalization to rare or newly deployed policies. We propose an uplift estimation framework that aligns treatment representation with causal semantics. Each policy is represented by the mixture it induces over context-action components and embedded via a permutation-invariant aggregation. This representation is integrated into an orthogonalized low-rank uplift model, extending Robinson-style decompositions to learned, vector-valued treatments. We show that the resulting estimator is expressive for policy-induced causal effects, orthogonally robust to nuisance estimation errors, and stable under small policy perturbations. Experiments on large-scale randomized platform data demonstrate improved uplift accuracy and stability in long-tailed policy regimes.


OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping

Mohammed Ehab ⋅ Aymane El Gadarri ⋅ Vivek Farias ⋅ Adam Jozefiak ⋅ Ciamac C Moallemi

Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However, these models often exhibit "computational overthinking," generating redundant reasoning steps that increase latency and cost without improving accuracy. Recent studies suggest that CoT trajectories can be significantly pruned, yet existing methods often rely on forcing a static thinking budget, heuristic filtering, sub-optimal early exit via classification, or expensive re-training. In this paper, we introduce OS-Pruner, a lightweight plug-in framework that formulates chain-of-thought pruning as an optimal stopping problem. Given a reasoning prefix, OS-Pruner learns whether further reasoning is worth its token cost by optimizing an explicit utility that trades off final-answer accuracy against generated length. Our novel formulation enables the model to dynamically assess the sufficient point of termination for a reasoning chain. OS-Pruner is designed to be lightweight during both training and inference, and to provide users with fine-grained control over the reasoning-effort vs. accuracy trade-off. On diverse reasoning benchmarks and base models, OS-Pruner achieves 20-60\% reduction in generation length with minimal accuracy sacrifice.


OTIS: Learning High-Quality Time Series Features With Tiny Encoders

Özgün Turgut ⋅ Philip Müller ⋅ Martin Menten ⋅ Daniel Rueckert

We introduce OTIS, an open time series encoder that yields high-quality time series features for downstream deployment on any system, including resource-constrained wearables and industrial sensors. Currently, the development of powerful general-purpose encoders relies on the scaling laws hypothesis, using large encoder sizes to memorise the heterogeneous distributions of multi-domain training data. However, this reliance on scale creates a barrier to real-world utility, rendering deployment on resource-constrained systems infeasible due to strict memory, energy, and latency constraints. Surprisingly, we find that tailoring standard masked modelling pre-training to time series properties yields a tiny 7.1M encoder that matches the state-of-the-art performance of 54× larger encoders across 162 tasks, while requiring 10× less memory, 43× less energy, and 37× lower latency. To achieve this without the capacity tax, we introduce three novel components: (1) a domain-aware tokeniser to resolve conflicting semantics within multi-domain training data; (2) a dual masking strategy to capture spatiotemporal structures and temporal causality; and (3) a structure-aware objective to decouple feature learning from modelling noise. Consequently, OTIS produces high-quality time series features that enable state-of-the art performance in discriminative tasks and even extend seamlessly to generative tasks at minimal additional cost. To democratise access to powerful time series features on any system, we release our code and pre-trained weights.


OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism

Xiangyu Li ⋅ Huaizhi Tang ⋅ Xin Ding ⋅ Weijun Wang ⋅ Ting Cao ⋅ Yunxin Liu

Embodied AI agents increasingly require parallel execution of multiple tasks, such as manipulation, conversation, and memory construction, from shared observations under distinct time constraints. Recent Mixture-of-Transformers (MoT) Vision-Language-Action Models (VLAs) architecturally support such heterogeneous outputs, yet existing inference systems fail to achieve efficient multi-task parallelism for on-device deployment due to redundant computation and resource contention. We identify isolated KV cache management as the root cause. To address this, we propose *unified KV cache management*, an inference design that treats KV cache as a first-class shared resource across tasks and over time. This abstraction enables two key optimizations: *cross-task KV sharing* eliminates redundant prefill of shared observations, while *cross-frame continuous batching* decouples variable-length language decoding from fixed-rate action generation across control cycles. We implement this design for $\pi_{0.5}$, the most popular MoT VLA, and evaluate on both NVIDIA GeForce RTX 4090 and Jetson AGX Thor, two representative platforms for on-device VLA inference. \nickname achieves up to 3.7$\times$ speedup over isolated execution, delivering over 200 tokens/s language throughput and 70 Hz action frequency simultaneously without action quality degradation, and we further validate the gains on a real humanoid robot with on-board Jetson AGX Thor.


PACE: Geometry-Aware Bridge Transport for Single-Cell Trajectory Inference

Chenglei Yu ⋅ Chuanrui Wang ⋅ Bangyan Liao ⋅ Tailin Wu

Single-cell trajectory inference from destructive time-course snapshots is fundamentally ill-posed: neither cross-time cell correspondences nor the continuous paths between snapshots are observed, so the observed snapshot distributions alone do not uniquely determine the underlying dynamics. Existing optimal transport and flow-based methods typically couple cells by Euclidean proximity at observed clock times, which can misalign trajectories when development is asynchronous and cells sampled at the same experimental time occupy different latent pseudotime stages. We propose PACE, a trajectory inference framework that selects geometry-consistent continuous transport dynamics from destructive time-course snapshots through three coupled components. First, PACE constructs a state- and time-dependent anisotropic Riemannian metric that preserves low cost along locally supported tangent directions while penalizing normal velocity components. Second, it alternates between refining cross-time couplings under the induced path-action cost and fitting endpoint-preserving neural bridges between adjacent snapshots. Third, it distills the learned bridge dynamics into a global continuous-time velocity field over cellular states. Across seven controlled and biological datasets covering nine held-out reconstruction experiments, PACE achieves the strongest overall reconstruction performance, reducing MMD, $\mathcal{W}_1$, and $\mathcal{W}_2$ by 23.7\% on average relative to the strongest competing baseline. PACE also improves RNA-velocity alignment by 15.4\% on an embryoid body differentiation benchmark, without requiring explicit cell pairing, lineage tracing, or RNA velocity supervision during training. Code is available at \url{https://anonymous.4open.science/r/PACE-F444/}.


PACE: Phase-Aware Chunk Execution for Robot Policies with Action Chunking

Junnan Nie ⋅ Jiayi Li ⋅ Jiachen Zhang ⋅ Junyi Lao ⋅ Chenghao Liu ⋅ Tianle Zhang ⋅ Songfang Huang

Recent advances in vision--language--action and diffusion-based robot policies have largely come from training-side gains, but reliable deployment also depends on how predicted actions are executed. Under action chunking, each query predicts a sequence of future actions, and the robot executes an open-loop prefix before re-querying. The length of this prefix, the execution horizon, is a critical inference-time variable that trades off feedback frequency against motion continuity. Since fixed horizons exhibit strongly task-dependent and non-monotonic effects on success, no single constant horizon provides a reliable cross-task deployment rule. We propose PACE (Phase-Aware Chunk Execution), a training-free test-time execution method that selects the execution horizon online from the predicted chunk itself. PACE exploits the phase-dependent kinematic structure of manipulation trajectories, using prominent low-speed valleys in the predicted speed profile as candidate replanning points. On the 50-task RoboTwin2.0 benchmark, PACE improves the average success rate from 57.8\% to 64.2\% over the strongest fixed-horizon baseline, without task-specific fixed-horizon tuning. In real-robot experiments, PACE improves the average task score from 60.7 to 77.7 and the average success rate from 50.7\% to 70.4\%.


PAC Reasoning: Controlling the Performance Loss for Efficient Reasoning

Hao Zeng ⋅ Jianguo Huang ⋅ Bingyi Jing ⋅ Hongxin Wei ⋅ Bo An

Large reasoning models (LRMs) have achieved remarkable progress in complex problem-solving tasks. Despite this success, LRMs typically suffer from high computational costs during deployment, highlighting a need for efficient inference. A practical direction is to switch the LRM between thinking and non-thinking modes dynamically. However, such approaches often introduce additional reasoning errors and lack statistical guarantees for the performance loss, which are critical for high-stakes applications. In this work, we propose $\textit{Probably Approximately Correct}$ (PAC) $\textit{Reasoning}$ that controls the performance loss under the user-specified tolerance. Specifically, we construct an upper confidence bound on the performance loss and determine a threshold for switching to the non-thinking model. Theoretically, using the threshold to switch between the thinking and non-thinking modes ensures bounded performance loss in a distribution-free manner. Our experiments on reasoning benchmarks show that the proposed method can save computational budgets and control the user-specified performance loss.


Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

Adam Fisch ⋅ Shubhendu Trivedi ⋅ Fantine Huot ⋅ William Cohen ⋅ Michael Kaisers ⋅ Mirella Lapata ⋅ Kate Larson ⋅ Jacob Eisenstein

Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the \emph{specialist} who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g., fine-tuned models with access to retrieval results or partial reasoning traces) are expensive. We formalize this tradeoff as an instance of Pandora's Box, the classical problem of optimal search with costly inspection. Under a Gaussian signal model, the resulting policies have closed-form value-of-information expressions that determine, for each specialist and input, whether refining the value estimate is worth its cost. We call the centralized policy Pandora's Router. We extend this to a decentralized setting, Pandora's Bidder, where specialists independently decide whether to invest in self-assessment before accepting an offered price to claim a query. Experiments across three domains---a standard multi-LLM benchmark, retrieval-augmented specialists, and LLMs with variable inference-time reasoning---show that Pandora's Router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often. In the decentralized setting, value-of-information reasoning improves allocative efficiency when competing estimates are accurate; when competing estimates are noisy, however, it can increase the strategic specialist's utility at the expense of

This paper explores the challenge of accelerating the sequential inference process of Diffusion Probabilistic Models (DPMs). We tackle this critical issue from a dynamic system perspective, in which the inherent sequential nature is transformed into a parallel sampling process. Specifically, we first reveal that the sequential integral solver of the diffusion model can be approximated by a full linear solver, enabling efficient computation for parallel integral solvers of DPMs. We then introduce a unified framework that reformulates the original nonlinear sequential integral process of the DPMs as a system of partial linear equations. Moreover, we further develop an immediate update strategy to solve the system. In addition, we prove that (1) the system admits a unique root corresponding precisely to the trajectory of the sequential integral solver; (2) solving the system guarantees convergence to the trajectory of sequential integral solvers in equal or fewer iterations. We then present \textit{\paralin}, a partial linear parallel integral solver to accelerate a broad class of sequential and parallel sampling methods such as DDPM and ParaSolver. This partial linearity allows \paralin to achieve parallel speedup on more practical single GPU settings. Experiments (\textbf{12B Flux}, Stable Diffusion Model, VP, VE, EDM, flow matching) validate that \paralin achieves $\textbf{2.7}\times$ to $ \textbf{3.9}\times$ speedup in practical 25-50 steps. Notably, on more practical single GPU setting where SOTA parallel solvers typically exhibit constraints, \paralin outperforms them with a 2.5× time speedup. Furthermore, to demonstrate the scalability, we show it can unlock up to a 50x speedup on classical large-step settings (e.g., 1000-step DDPM). The source code will be released publicly.


Parameter Exploration for RLVR via Variational Learning

Vatsal Venkatkrishna ⋅ Nico Daheim ⋅ Iryna Gurevych

Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that encouraging exploration during LLM reinforcement learning can improve downstream performance. However, methods for controlling exploration often rely on heuristics like clipping or temperature scaling that are unable to fundamentally change token-level distributions, which might limit exploration. Here, we explore controlling exploration via variational learning, where a distribution over neural network parameters is learned via noisy optimization, with parameters sampled from an approximate posterior. We introduce Perturbed Parameter Policy Optimization (3PO), where the amount of noise that is added to model parameters functions as an additional control lever that can be tuned for exploration. We show that using multiple parameter samples within a batch performs better than using a single sample. When using multiple parameter samples in group-based methods like GRPO, we find that calculating advantages across rollouts from all parameter samples performs best. We call this Chunked Perturbed Parameter Policy Optimization (C3PO). We show empirically on a range of math reasoning benchmarks and LLMs that using multiple parameter samples can improve downstream performance, especially on harder benchmarks like AIME. Overall, our work presents evidence that parameter-space exploration can improve LLM reinforcement learning.


Parameterized Stripe Attention for Efficient Video Generation

xingyu jia ⋅ Baole Ai ⋅ Ang Wang ⋅ Kang Zhao ⋅ Yong Li

Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the computationally expensive full spatio-temporal attention. While sparse attention methods offer potential solutions, existing approaches face an inherent flexibility--efficiency dilemma: predefined masks lack the flexibility to capture diverse attention patterns, while runtime-determined masks introduce overheads and sacrifice hardware efficiency. We identify the lack of a unified structural characterization of DiT attention as a key limitation of existing methods, and establish that video DiT attention exhibits periodic diagonal stripe structures along both temporal and spatial dimensions. To formally encode these structured patterns within a single efficient kernel, we present PSA, a parameterized stripe attention that formalizes the observed stripe regularity, unifying diverse attention patterns for efficient mask generation. This unified representation enables a single hardware-efficient CUDA kernel to process all sparse patterns, achieving FlashAttention-3-level Model FLOPs Utilization. To determine optimal sparsity configurations, we propose a training-free offline search algorithm that automatically maximizes sparsity under a specified error tolerance for each attention head. Experiments on HunyuanVideo and Wan 2.1 demonstrate that PSA achieves 1.57x and 1.37x end-to-end speedups over FlashAttention-3 baselines, with acceptable visual quality degradation.

Learning across multiple inherently conflicting objectives requires selecting desirable trade-offs from the Pareto set. In this work, we study preference-guided multi-objective optimization, where a predefined preference function is minimized over the weak Pareto set of multiple nonconvex objectives. Leveraging a merit function that vanishes on the weak Pareto set, we propose a penalty-based min-max-min formulation, termed the M$^3$ problem, which unifies preference optimization and the enforcement of weak Pareto optimality in a single objective. We further develop ParetoM$^3$, a single-loop first-order algorithm for computing stationary solutions of the resulting M$^3$ problem. Under a local Kurdyka-{\L}ojasiewicz assumption, we prove that ParetoM$^3$ finds an $\epsilon$-stationary point of the M$^3$ problem that is also approximately weakly Pareto optimal for the original multi-objective problem when the penalty parameter is sufficiently large, with explicit non-asymptotic convergence rates. We also show that this stationarity notion recovers existing optimality criteria, including approximate preference stationarity and approximate Karush-Kuhn-Tucker conditions, under the same assumptions used in prior works. Experiments on a synthetic benchmark, image classification, regression, and math reasoning demonstrate that ParetoM$^3$ reliably attains preferred weak Pareto optimal solutions while achieving strong practical performance.


Pareto Preference Optimization for Structure- and Stability-Aware RNA Inverse Folding

Minghao Sun ⋅ Hanqun Cao ⋅ Zhou Zhang ⋅ Chen Wei ⋅ Liang Wang ⋅ Tianrui Jia ⋅ ZHIYUAN LIU ⋅ Tianfan Fu ⋅ Robert Tang ⋅ Yejin Choi ⋅ Pheng-Ann Heng ⋅ Fang Wu ⋅ Yang Zhang

RNA inverse folding seeks sequences that reliably fold to a target 3D backbone under physiological conditions. Multi-objective preference optimization on RNA is unstable for two domain-specific reasons: *heterogeneous noise* across physical proxies (deterministic 2D folding, stochastic 3D prediction, minimum free energy), and *sequence–structure degeneracy* that admits compositional reward hacking via GC enrichment.We introduce **RiboPO**, a preference-optimization framework that builds preference pairs as $\varepsilon$-Pareto-dominance relations on standardized per-metric features, gates winners with a structural quality threshold against GC-driven shortcuts, and trains a frozen-reference DPO policy under a decreasing-margin curriculum. A heuristic anchored-KL drift characterization and a pair-level Rényi-2 off-policy bias bound motivate the experimental design; a clipped importance-corrected variant extends usable rounds beyond the static-pair regime. On DAS benchmark, specialized **RiboPO** variants improve the gRNAde base on SSTT axes: scMCC reaches $0.71$ at $R=4$ ($+17\%$ over $0.61$) and $0.643$ at the thermodynamic-surplus point ($+6.1\%$, paired Wilcoxon $p<10^{-3}$); MFE improves $-7.3\%$; target-structure probability $P(S_0)$ rises $0.00128 \to 0.0155$ ($\sim$12$\times$) with ensemble-defect $-6.9$ to $-7.9\%$ ($p=6.1\times10^{-5}$ at $R=4$), consistent with mass concentration toward target rather than redistribution to an off-target basin; designability rises $+6.9$ pp %pLDDT$\geq$0.7 and $+6.4$ pp %RMSD$\leq$8 Å. Under matched-utility same-pool reranking, **RiboPO** best-of-8 matches gRNAde best-of-64 (pass@1 $0.408$) at $1/8$ the per-backbone budget. On the leakage-corrected subset, **RiboPO** dominates every structural and thermodynamic axis. Codes are available at: [https://anonymous.4open.science/r/ribopo-8D48/](https://anonymous.4open.science/r/ribopo-8D48/).


PatchBench: Measuring Collateral Damage in Activation Patching

Alexi Canesse ⋅ Mathis Le Bail ⋅ Maël JENNY ⋅ Clément Elliker ⋅ Mahammed El Sharkawy ⋅ Sonia Vanier

An LLM safety patch can pass a benchmark while still being a poor repair. This risk is especially acute for jailbreak repairs, where the goal is to correct a specific unsafe behaviour without changing unrelated behaviours. A patch may block the exact prompts used for evaluation yet fail on close harmful variants, or suppress harmful behaviour by over-refusing benign prompts that share its wording or structure. Existing evaluation protocols primarily test whether models can be broken, while aggregate metrics, such as attack success, refusal rates, and global capability scores, cannot distinguish selective repairs from broader local suppression. To address this gap, we introduce PatchBench, a benchmark of empirically observed model-specific jailbreak failures that induce actionable harmful answers, designed as concrete repair targets for patching research. Starting from 27,870 prompts aggregated from 37 public jailbreak benchmark datasets, we manually curate 15,314 English prompts and query 8 open-source instruction-tuned models. We then combine WildGuard filtering, pairwise Elo ranking, and manual verification to retain only unsafe completions that answer the harmful requests. The result is a curated bank of 400 high-confidence jailbreak failures, organised by model. We further introduce PatchBench-Local, a local evaluation protocol that tests whether a patch is behaviourally precise. For each harmful source prompt, PatchBench-Local generates three families of local neighbours: harmful variants that preserve the malicious intent, benign prompts with matched structure, and benign prompts that reuse the key harmful terms. It evaluates both harmful-neighbour correction and benign-neighbour preservation, thus distinguishing selective repair from broader local suppression. We use PatchBench-Local and MMLU to evaluate four state-of-the-art activation steering methods. Our results show that global capability can remain nearly unchanged while local benign-neighbour regressions are severe, confirming that aggregate metrics alone can miss important collateral damage. By exposing whether patches act as precise behavioural repairs or as broad local suppressors, PatchBench-Local provides a more precise basis for developing and comparing jailbreak repair methods.


PATCH: Learnable Tile-Level Hybrid Sparsity for LLMs

Mohammad Mozaffari ⋅ Younes Hourri ⋅ Maryam Mehri Dehnavi

Large language models (LLMs) deliver impressive performance but incur prohibitive memory and compute costs at deployment. Model pruning is an effective way to reduce these overheads, yet existing approaches face challenges: unstructured sparsity, where nonzeros can appear anywhere, preserves accuracy but yields irregular access patterns that prevent GPU acceleration, while semi-structured 2:4 sparsity is hardware-friendly but enforces a rigid 50\% pattern that degrades model quality. To bridge this gap, we introduce PATCH, a hybrid sparsity framework that enables a continuous sparsity ratio between 0\% and 50\%. PATCH partitions weight matrices into tiles, assigning each tile to be either dense or 2:4 sparse via a learnable mask selection mechanism. This design provides fine-grained control over accuracy–acceleration tradeoffs and supports non-uniform sparsity across layers, leading to superior overall quality. Across models from 0.5B to 13B parameters, PATCH consistently narrows the gap to dense accuracy while delivering practical speedups. For instance, on LLaMA-2 7B with an A6000 GPU, PATCH achieves 1.18×–1.38× end-to-end speedup over dense baselines while improving accuracy by 0.37\%–2.96\% compared to the state-of-the-art 2:4 pruning method, MaskLLM.

Multimodal large language models (MLLMs) are increasingly used to analyze pathology images. However, dominant multimodal benchmarks in pathology mainly score final diagnostic answers, captions, or reports. These evaluations provide limited insight into whether a model understands the multiscale visual content needed for pathology reasoning and decision-making. We introduce PathView-Bench, a vision-anchored benchmark for fine-grained and multiscale visual understanding in computational pathology. Built from $23$ public pathology imaging datasets with human-supervised labels and spatial annotations, PathView-Bench evaluates MLLM understanding in two fields of view: Region-FOV for high-resolution local regions and Slide-FOV for macro whole-slide views. By converting raw annotations into deterministic task targets, PathView-Bench enables programmatic scoring of region localization, visual recognition, quantity estimation, spatial reasoning, and insufficient-context judgment. The benchmark contains $14$ VQA-style tasks, $61,673$ images, and $308,070$ samples across $28$ organs and $7,253,526$ annotations. Evaluating $18$ representative general-purpose, medical-domain, and pathology-oriented MLLMs, we observe substantial limitations even in advanced models on fine-grained visual tasks across multiscale pathology images. PathView-Bench provides a reproducible basis for developing and evaluating pathology MLLMs with explicit multiscale visual understanding.


PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language

Hongliang Lu ⋅ Zhong Li ⋅ Yuxuan Chen ⋅ Lan Yuan ⋅ Fan Zhang ⋅ Zaiwen Wen

Optimization modeling is the process of translating real-world decision problems, often described in natural language, into formal mathematical formulations and executable solver code. While recent advances in large language models have shown promise in automating this process, most existing approaches remain one-shot: a model produces a formulation once, without executing it, conditioning on solver feedback, or iteratively revising errors. This stands in sharp contrast to real-world optimization modeling, which is inherently interactive and proceeds through repeated solve--debug--revise cycles. We introduce PEARL, a system for interactive optimization modeling that uses Python execution and mathematical programming solvers inside this loop. Rather than relying on a fixed repair workflow, PEARL learns when to test partial models, how to revise from solver diagnostics, and when to stop. It operates in a multi-turn tool-integrated setting where intermediate execution results, feasibility signals, and solution checks are used to improve both formulations and solver code before finalization. Across diverse optimization benchmarks, PEARL substantially improves verified solve rates over strong one-shot and tool-augmented baselines; notably, our PEARL-Qwen3-\textbf{4B} model outperforms the much larger DeepSeek-V3.2-\textbf{685B} in both macro- and micro-averaged accuracy on optimization modeling tasks.


Peer review should constrain evaluative authority

Hanrui Wang ⋅ Timo Spinde ⋅ Isabella Habereder ⋅ Chun-Shien Lu ⋅ Isao Echizen

Our position shifts the target of peer-review reform from reducing uncertainty through evaluator improvement or judgment aggregation to constraining evaluative authority under persistent uncertainty. Scientific peer review aims to assure research quality but also allocates scarce academic opportunities, such as publication in high-impact venues. However, the review system is under growing pressure as submissions increase while reviewer capacity, expertise, scrutiny, and incentives cannot scale accordingly. As peer review operates under persistent evaluative uncertainty, evaluator improvement and judgment aggregation remain structurally limited. We therefore propose authority-constrained peer review: before a consequential judgment is released as decision-relevant, it should withstand structured scrutiny through critique-authority decoupling, independent secondary evaluation, contestability and reversibility, and accountability for unsupported influence. Our aim is to prevent weakly supported judgments from entering final deliberation with the same authority as judgments that have withstood structured scrutiny, while creating incentives for reviewers to justify consequential claims more carefully. For large-scale computer-science conferences, this position offers an institutional response to uneven review quality and reviewer-assignment dependence; more broadly, it provides a governance principle for any domain where uncertain expert judgments allocate scarce academic opportunities.


PercepCap: Video Captioner with Structured Spatio-Temporal Perception

Yifan Xu ⋅ Wang Zihao ⋅ Zhixiao Wang ⋅ Jiaming Zhang ⋅ Yichun Yang ⋅ Desen Meng ⋅ Yuanxing Zhang ⋅ Pengfei Wan ⋅ Limin Wang

Video captioning requires fine-grained spatio-temporal understanding of videos, including spatial perception of where objects are located and temporal perception of when events occur. Existing Multi-modal Large Language Models (MLLMs) usually generate captions directly from video inputs without exposing the perceptual evidence behind their descriptions. As a result, object, event, or temporal mistakes are only observed in the final text, making it difficult to identify the underlying perceptual errors or optimize them directly. To address these issues, we present PercepCap, a perception-aware video captioning framework that makes perceptual evidence explicit before producing the final caption. Specifically, PercepCap follows a perceive-describe generation chain, where the model first produces a spatio-temporal perception trace comprising object trajectories and temporal events according to the video, and then generates the final caption conditioned on the perceived evidence. To support this new generation paradigm, we design a two-stage training strategy for PercepCap. Perceive-then-Describe Supervised Fine-tuning (PD-SFT) adapts the model from caption-only generation to the proposed perceive-describe chain, while Perception-Grounded Reinforcement Learning (PG-RL) further optimizes perception trace and caption quality with joint rewards over object tracking, temporal event, and object/action description coverage. Caption-Anchored Perception Data Construction builds this supervision by first generating a caption-only description, extracting the objects and events it mentions, and grounding them back in the video with boxes and timestamps. This yields caption-aligned perception data that provides both SFT supervision and RL references, ensuring that the explicit perception trace and final caption describe the same objects and events. Across direct caption evaluation such as DREAM-1K, CaReBench and VidCapBench and caption-to-QA evaluation like ShortVidBench and MotionBench. PercepCap consistently improves the Qwen3-VL baseline and achieves leading open-source caption quality.

Physics-informed neural networks (PINNs) train a single neural approximation by minimizing multiple physics- and data-derived losses, but the gradients of these losses often interfere and can stall optimization. Existing remedies typically treat this pathology either through scalar loss balancing or full-parameter-space gradient surgery, leaving it unclear which intervention is most appropriate. We show that PINN gradient conflict is not a uniform failure mode with one universal remedy. Instead, we identify distinct PINN gradient-conflict regimes, each associated with a different intervention class. Persistent directional conflict may require separate loss-indexed parameter subspaces, magnitude imbalance often favors scalar reweighting, and low or transient conflict may require no extra mitigation. To select between scalar reweighting and a lightweight architectural intervention, we propose a diagnostic-first framework. It profiles a 1000-step unmodified PINN run and, when intervention is warranted, uses one low-rank adapter per loss to create explicit loss-indexed parameter subspaces attached to a shared PINN trunk, providing each loss with a direct gradient pathway. Across more than 60 PDE configurations, including forward, inverse, multi-physics, parameter-varying, and high-dimensional problems up to 50D, persistent directional conflict dominates standard forward $K=3$ benchmarks and a natural $K=4$ thermoelastic system, where adapters combined with reweighting yield significant improvements. In contrast, $K=3$ inverse problems and natural $K=5$ and $K=6$ multi-physics systems are largely magnitude-dominated and often favor reweighting alone, while full-parameter-space gradient surgery can fail on heterogeneous parameter spaces. A regime-transition theorem and a blockwise neural tangent kernel analysis explain why gains from these separate adapter parameter subspaces arise selectively in persistent-conflict regimes.


PF-SGS: Pose-Free Streaming 3D Gaussian Splatting for Large-Scale Scene Reconstruction

wenjie mu ⋅ ziniu liu ⋅ Tong Wu ⋅ Zhan Li ⋅ Chuanzhou su ⋅ XUANYI SHEN ⋅ Fan Lu ⋅ Junqiao Zhao ⋅ Tiantian Feng ⋅ Chen Ye ⋅ Guang Chen

Extending feed-forward 3D Gaussian Splatting (3DGS) to large-scale scenes remains challenging. Existing global reconstruction methods incur prohibitive computational overhead on long sequences; while streaming-based online methods are more efficient, their inherent forgetting behavior and unidirectional dependence still limit reconstruction quality. To address this, we propose PF-SGS, a pose-free streaming feed-forward 3DGS for large-scale scene reconstruction. PF-SGS leverages a persistent hidden state as scene memory to integrate streaming inputs frame-by-frame, progressively recovering scene geometry and camera poses in a self-supervised manner. At its core lie two targeted designs: Adaptive State Gating, which suppresses long-term memory degradation by dynamically regulating update magnitudes to ensure stable state evolution; and Delayed Gaussian Modeling, which introduces bounded look-ahead observations to compensate for insufficient local geometric constraints under streaming inputs, thereby improving the geometric consistency. Experiments demonstrate that PF-SGS processes sequences of varying lengths with linear time complexity, achieving quality comparable to pose-prior-based large-scale reconstruction baselines.


PGGT-Loc: Primitive-Grounded Geometry Transformer for Feed-Forward Camera Localization

Hanqiao Ye ⋅ Yuzhou Liu ⋅ Mengqi Rong ⋅ Jian Liu ⋅ Yangdong Liu ⋅ Shuhan Shen

Camera localization methods are mostly multi-stage pipelines relying on dedicated and complex maps that are costly to construct, store, and maintain over time. Alternatively, well-structured maps that organize scenes with compact yet expressive primitives are often readily available, providing rich structural and semantic cues that humans can interpret. However, their potential for direct camera localization remains largely unexplored due to the significant modality gap with visual observations. This motivates a new concept for localization, which estimates camera poses by grounding query observations to map primitives. In this paper, we instantiate the concept in the context of planar primitives with PGGT, the Primitive-Grounded Geometry Transformer network trained to predict multiple localization targets from a single query image and a set of planar map primitives, enabling feed-forward 6-DoF camera localization that generalizes to unseen environments. Based on this flexible and scalable architecture, our method pushes the state-of-the-art localization performance on both detailed planar surface maps and minimal architectural layouts. The code and models will be made publicly available.


PGMS: Pyramidal Gaussian Mixture Splatting for 3DGS Compression

Xiaoge Zhang ⋅ Zijie Wu ⋅ Mingtao Feng ⋅ Saeed Anwar ⋅ Ajmal Mian

3D Gaussian Splatting (3DGS) has become an established representation for real-time novel view synthesis. However, preserving fine geometric structures and appearance details typically requires millions of Gaussian primitives, resulting in substantial storage and transmission overhead. 3DGS compression has been extensively studied through attribute quantization, entropy coding, and lightweight reparameterization. However, existing methods often operate on a fixed backbone and therefore do not explicitly model the joint redundancy between geometry and appearance in the original dense Gaussian set. To address this limitation, we propose \textbf{Pyramidal Gaussian Mixture Splatting} (PGMS), a plug-and-play 3DGS compression framework that reformulates dense Gaussian primitives into a rate-controllable pyramidal representation with scale-dependent attributes. Specifically, PGMS first constructs a mixture pyramid by progressively partitioning dense Gaussians into hierarchical mixture centers, where different levels perform attribute-specific clustering to form shared prototypes. It then employs a rate-adaptive capacity allocator to determine the tiered budget under a target compression ratio. To further improve rendering quality, we introduce a compositing-aware parameter update rule that incorporates information from dense Gaussians into pyramid primitives according to their image-space contributions. Experiments on standard benchmarks show that our method delivers strong compression with high rendering fidelity and consistently outperforms state-of-the-art 3DGS compression methods.

Vision-Language-Action (VLA) foundation models have driven remarkable progress in general-purpose robotic manipulation. Yet, adapting these models to dynamic physical environments exposes a severe vulnerability: a modality imbalance during optimization. Because low-dimensional proprioceptive states provide a more direct path for loss minimization, policies instinctively form a "state-dominant shortcut," systematically suppressing the learning of high-dimensional visual features. This visual neglect causes models to bypass precise spatial grounding, leading to catastrophic alignment failures during contact-critical execution phases. To rectify this without disrupting natural training dynamics or requiring heavy auxiliary decoders, we introduce Phase-Adaptive Fusion (PAF), a lightweight, non-intrusive neural modulation plug-in. Operating entirely in the forward pass, PAF features a decoupled dual-branch architecture. A spatial branch performs explicit semantic-geometric alignment, while a temporal branch infers latent task execution phases from recent action history and state kinematics. By synthesizing these signals into a residual gating mechanism prior to multimodal fusion, PAF dynamically amplifies visual tokens precisely during alignment-heavy stages, all while preserving the host VLA's pre-trained representational manifold. Extensive evaluations across simulated benchmarks (LIBERO, Meta-World, Robomimic) and real-world long-horizon tasks demonstrate that PAF successfully breaks the proprioceptive shortcut, drastically reducing terminal alignment drift and significantly enhancing overall manipulation robustness.


Phase Transitions in Heavy-Tailed Mean Estimation under $\ell_p$ Norms

Ishaq Aden-Ali ⋅ Yeshwanth Cherapanamjeri ⋅ Mikael Møller Høgsgaard ⋅ Kasper Green Larsen ⋅ Nikita Zhivotovskiy

We study the problem of estimating the mean of a random vector in $\mathbb R^d$ under $\ell_p$ norm, assuming only the existence of a covariance matrix. In the Euclidean norm, the sample mean is already optimal in expectation, and robust estimators are usually needed only to obtain high probability tails. Our results show that this picture changes for $\ell_p$ norms with $p>2$: robustification is necessary already for the expected risk. We prove that the minimax risk exhibits a sharp phase transition, controlled by the interaction between the sample size $n$, the dimension $d$, and the norm parameter $p$. In the large sample regime, the classical Gaussian rate remains achievable. In the complementary regime, heavy tailed distributions create a genuinely non-Gaussian obstruction. The optimal estimator is surprisingly simple: a coordinate-wise median of means whose number of blocks is chosen according to the geometry of the norm, rather than according to a confidence parameter. We also show that natural projection based median of means estimators for general norms, despite achieving the optimal confidence tradeoff relative to the sample mean benchmark of Lugosi and Mendelson (Probab. Theory Relat. Fields, 2019), can be polynomially suboptimal for $\ell_p$ norms even in the constant probability regime.


PhysEval.Weather: An Evaluation Framework for Physical Consistency in ML Weather Models

Emma Kasteleyn ⋅ Timo Maier ⋅ Axel Lauer ⋅ Veronika Eyring ⋅ Pierre Gentine ⋅ Ana Lucic

Machine learning weather prediction (MLWP) models have achieved impressive forecasting performance at a small fraction of the computational costs required for traditional physics-based methods. However, they are primarily (1) data-driven and (2) evaluated using pixel-wide error metrics (e.g., RMSE), so there are no guarantees that their forecasts are physically consistent. We introduce \textbf{PhysEval.Weather}, an evaluation framework that assesses the physical realism of MLWP models across three types of metrics: conservation, spectral, and dynamical. By quantifying physical realism, this tool guides the development of physics-informed architectures and helps evaluate whether MLWP models are reliable for operational use.

Realistic physical interaction is a cornerstone of embodied intelligence, yet obtaining high-fidelity tactile data remains significantly more expensive than visual data. This scarcity necessitates Visual-to-Tactile synthesis to bridge the Sim-to-Real gap. However, existing data-driven approaches often treat this as a naive image-to-image translation task, neglecting fundamental contact mechanics. Moreover, the inherent spatial misalignment in real-world visual-tactile datasets frequently forces generative models to average out spatial uncertainties, resulting in blurry and textureless outputs. To address these limitations, we introduce PhysTacGen, a physics-aware generation framework that enforces consistency between visual appearance and tactile mechanics. Our system makes three technical contributions. Firstly, we propose Group Tactile Policy Optimization (GTPO), a reinforcement learning alignment strategy that refines Large Multimodal Models to infer latent physical properties by rewarding physics-consistent reasoning chains. Secondly, we introduce a Semantic-Driven Data Curation and Visual Prior pipeline, leveraging DINOv2 to filter spatially misaligned data and extracting pure depth maps to decouple macro-geometry from surface texture. Finally, we present a physics-driven generative architecture based on Stable Diffusion XL (SDXL). It synthesizes tactile images via a ControlNet conditioned on GTPO-generated physical descriptions and depth-augmented visual priors. Extensive experiments demonstrate that PhysTacGen achieves state-of-the-art perceptual quality and significantly enhances grasp force prediction for zero-shot Sim-to-Real transfer.


PhysTC: A Physics-Enhanced Dataset and Architecture for High-Precision Tropical Cyclone Forecasting

Zhaoran Feng ⋅ Xuanhong Chen ⋅ Zengbing Chen ⋅ Shengjun Wu ⋅ Bin He ⋅ Kairui Feng

A critical observability gap arises in multiscale chaotic systems when coarse observations smooth out high-frequency extremes. For tropical cyclones, this makes standard reanalysis products insufficient for high-precision intensity prediction, as spectral truncation can mask peak winds by over $30\%$. In this work, we introduce Physics-enhanced Tropical Cyclone forecasting (PhysTC), a physics-enhanced benchmark framework for forecasting tropical cyclone track and intensity from globally consistent coarse observations. We curate a harmonized multi-basin dataset spanning 1950--2023, incorporating scale-robust physical predictors to mitigate resolution-induced bias. Building on this dataset, we propose the Physics-enhanced Tropical Cyclone Network (PhysTCN), a model that integrates synoptic spatial structure, storm-following temporal history, and physical consistency constraints to infer intensity-relevant dynamics beyond direct grid-space readout. PhysTCN significantly outperforms state-of-the-art baselines, with particularly strong gains in recovering extreme wind speeds that are underestimated in reanalysis data. This work provides a scalable framework for physically guided forecasting in data-sparse, cross-scale dynamical systems.


PiCA: Pivot-Based Credit Assignment For Search Agentic Reinforcement Learning

Dongyi Liu ⋅ Yifan Niu ⋅ Qinwen Wang ⋅ Han Xiao ⋅ Jia Li

Large language model (LLM)-based search agents trained with reinforcement learning (RL) have significantly improved the performance of knowledge-intensive tasks. However, existing methods encounter critical challenges in long-horizon credit assignment: (i) Reward Sparsity, where models receive only outcome feedback without step-level guidance to differentiate action quality; (ii) Isolated Credit, where credit is assigned to steps independently, failing to capture sequential dependencies; and (iii) Distributional Shift, where rewards are estimated on templates that deviate from the model’s natural generative distribution. To address these issues, we propose Pivot-Based Credit Assignment (PiCA), a novel step reward mechanism that reformulates the search trajectory as a sequential process of cumulative information gain. Unlike prior methods that reward steps in isolation, PiCA defines process rewards as success probabilities dependent on the historical context. This approach identifies pivot steps (critical milestones in the search) as information peaks that significantly boost the likelihood of a correct final answer. By anchoring these step rewards to the final task objective, PiCA provides dense, pivot-aware and trajectory-dependent guidance while maintaining distributional consistency. Theoretically grounded in Potential-Based Reward Shaping (PBRS), PiCA consistently outperforms existing strong baselines across seven knowledge-intensive QA benchmarks, achieving 15.2% and 2.2% improvements for 3B and 7B models.

We introduce PINNBench, a benchmark and evaluation study for training-policy selection in hybrid PINN-operator PDE solvers, contributed to the NeurIPS 2026 Evaluations and Datasets (ED) Track. PINNBench comprises 1,539 recorded runs and probes across 13 PDE configurations drawn from 8 equation families, 3–5 training policies per PDE, and 10 seeds per condition, with seven standardized metrics, Wilcoxon signed-rank tests, paired bootstrap 95% confidence intervals, Cohen's d, and Benjamini–Hochberg FDR correction. Baseline evaluation of 5 routing methods establishes three findings that bear directly on how PDE solvers should be evaluated. (i) No single policy dominates: hybridfull wins 7/13, hybridnogate 4/13, pinnonly 2/13. (ii) Accuracy and regret dissociate: under relative L2, a physics-loss heuristic and a learned Random Forest router reach identical 62% top-1 accuracy, but the heuristic suffers mean regret 0.0850 versus 0.0011 — a 77x gap. Evaluation by accuracy alone is misleading for PDE policy routing. (iii) Causal-loss effectiveness is stage-mediated: 7/9 PDEs improve with pretraining (+5.9% mean) while 5/9 degrade without it (-4.2%). A new leave-one-PDE-family-out evaluation shows that the strongest family-level generalizer uses no PDE-identity features, addressing the concern that PDE-aware routing memorizes PDE identity. All 1,539 result JSONs, PDE implementations, policy configurations, statistical pipeline, and routing infrastructure are released as open-source artifacts.


Pinpoint: Grounded Worldwide Image Geolocation via Cross-Source Retrieval and Reranking

Nika Chuzhoy ⋅ Brian Hu ⋅ Amit A Arora ⋅ Jae Ro ⋅ Sarthak Sahu

Image geolocation aims to estimate where a photograph was taken from its visual content. At worldwide scale, this remains challenging because visual evidence is often ambiguous, diverse, and unevenly distributed. Prior work has typically treated geolocation of ordinary internet photos and street-view imagery as separate tasks, despite their complementary strengths: internet photos better match the appearance distribution of user-captured queries, while street-view imagery provides denser, geographically grounded coverage. We present Pinpoint, a retrieve-and-rerank architecture that combines both sources in a coarse-to-fine pipeline. A contrastive image-GPS embedder is trained on both user-uploaded Flickr photos and street-view imagery, learning a shared image-GPS embedding space that is used to retrieve candidate locations. An attention-based reranker then rescores retrieved candidates by combining candidate-level visual and GPS features with cross-source evidence from nearby locations to ground the prediction. Unlike recent prior work, Pinpoint does not rely on multimodal large-language models, making inference faster and more reproducible. Pinpoint achieves state-of-the-art results across all metrics on standard benchmarks for internet photos (IM2GPS3k and YFCC4k) and street-view imagery (OSV-5M).

Modern automatic-rigging predictors (UniRig, Make-It-Animatable, RigNet) place a skeleton on a 3D mesh in seconds, but the resulting skinning weights animate poorly. Linear-blend skinning (LBS) is non-linear in pose space, so noisy weights amplify into visible candy-wrapper artifacts---the surface twisting and contracting at joint bends---and volume collapse. ARAP-style remedies penalize these symptoms one pose at a time and cannot reach the cross-pose curvature that produces them. We propose \emph{Pose-Interpolation Smoothness} (PIS), a self-supervised, training-free, test-time refiner: minimize the per-vertex gap between the deformation at the SO(3) midpoint of two random poses and the linear average of the two endpoint deformations. The midpoint is obtained by orthogonally projecting the arithmetic mean of the two rotations back onto the rotation group, sidestepping the matrix-logarithm gradient that breaks down near identity and antipodal poses. Because PIS only re-tunes the weights LBS already consumes, it slots into any LBS-based pipeline---generative-3D outputs, automatic-rigging predictions, and the LBS pathways of game engines and AR/VR renderers---without touching shaders. On Objaverse-LVIS Balanced-102 (98 valid models, 51 LVIS categories), PIS reduces fragility (per-vertex deformation variance under random poses), ARAP energy, Jacobian anisotropy, and area distortion by $35.5$\%, $39.3$\%, $19.8$\%, and $43.2$\% on average; $93.9$\% of models improve on all four simultaneously, Wilcoxon $p<10^{-10}$. On the worst-case slice---the top decile of UniRig-initialized fragility, where animation visibly tears---PIS reduces unstable vertices by up to $10.2\times$. Effective bone count and weight entropy are unchanged, and the optimization runs in $\sim$30~s/model on a single H200 GPU.


Plausibility Is Not Prediction: Contrastive Evidence for LLM-Based Cellular Perturbation Reasoning

XINYU YUAN ⋅ Xixian Liu ⋅ Jianan Zhao ⋅ Ya Shi Zhang ⋅ Hongyu Guo ⋅ Jian Tang

Perturbation experiments are central to understanding cellular mechanisms, but remain costly and sparse, motivating prediction of gene expression responses for unobserved conditions. A promising recent direction leverages large language models (LLMs) as “virtual cell” simulators—using stepwise, knowledge-grounded mechanistic reasoning to infer differential expression—pointing toward an interpretable, knowledge-driven paradigm that transcends purely data-driven approaches. However, we find that plausibility is not prediction: despite producing biologically plausible explanations, these methods fail to capture perturbation-specific effects — systematically overestimating differential expression, often underperforming a simple gene‑frequency baseline in aggregate evaluations, and collapsing to chance-level performance at the per-gene level. This reveals a reliance on intrinsic gene response tendencies rather than true perturbation reasoning. We trace this failure to how evidence is presented: existing methods evaluate perturbation–gene pairs in isolation, without exposing how related perturbations differ in their effects on the same gene. To address this limitation, we introduce CORE (Contrastive Organization of Relational Evidence), which reframes prediction as a comparison task by organizing evidence into positive and negative outcomes from related perturbations. Using a biomedical knowledge graph for evidence retrieval, CORE improves calibration and substantially boosts perturbation-specific prediction in both LLM-based and non-LLM settings: for example, on drug-perturbation data, CORE-Reasoning improves Qwen3.5-9B aggregate metrics by up to 28.6\%, while on generic perturbation data, CORE-Voting raises macro-per-gene AUROC from chance to 0.711. This highlights contrastive evidence organization as essential to reliable LLM-based perturbation reasoning.


PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations

Cheng Chi_ ⋅ Xianqi Wang ⋅ Hongcheng Luo ⋅ Mingfei Tu ⋅ Gangwei Xu ⋅ Zehan Zhang ⋅ Bing Wang ⋅ Guang Chen ⋅ Hangjun Ye ⋅ Sida Peng ⋅ Xin Yang ⋅ Haiyang Sun

High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable fast reconstruction, their per-pixel Gaussian prediction paradigm often suffers from multi-view inconsistency and layering artifacts. Moreover, existing methods often model dynamic instances via dense flow prediction, which lacks explicit cross-view correspondence and instance-level consistency. In this paper, we propose PointForward, a feedforward driving reconstruction framework through point-aligned representations. Unlike pixel-aligned methods, we initialize sparse 3D queries in world space and aggregate multi-view image information via spatial-temporal fusion onto these queries, enforcing explicit cross-view consistency in a single feedforward pass. To handle scene dynamics, we introduce scene graphs that explicitly organize moving instances during reconstruction. By leveraging 3D bounding boxes, our method enables instance-level motion propagation and temporally consistent dynamic representations. Extensive experiments demonstrate that PointForward achieves state-of-the-art performance on large-scale driving benchmarks. The code will be available upon the publication of the paper.


PointZero: 3D Point Track Completion for Learning Transferable 3D Dynamics

Bardienus Duisterhof ⋅ Kaifeng Zhang ⋅ Adam Hung ⋅ Bowen Wen ⋅ Stan Birchfield ⋅ Yunzhu Li ⋅ Deva Ramanan ⋅ Jeffrey Ichnowski

World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse volumes of data, to instill a rich prior into downstream applications. Existing methods for capturing 3D world models are typically scene-specific dynamics models, or require robot action labels - which excludes web video data from the training pool. We study 3D point track completion as a pre-trianing objective for learning transferable 3D dynamics without robot data. Given a single RGB-D observation and sparse partial 3D trajectories (tracks), we predict future 3D tracks of all observed points. We show this is a objective provides a rich dynamics prior, without requiring robot action labels. We contribute a diverse dataset of 2.9 million synthetic frames spanning deformable, articulated, and rigid objects, and use it to train PointZero. We show that a flexible and expressive transformer, PointZero, outperforms prior methods on the same data. We demonstrate the utility of our pre-training objective by fine-tuning PointZero for two downstream applications: (1) action-conditioned dynamics and (2) robot manipulation. When fine-tuned to condition on end-effector pose, PointZero outperforms the baselines on a recent action-conditioned 3D dynamics benchmark. When fine-tuned to predict robot actions, PointZero outperforms or matches the baselines on 6/7 simulated and real-world robot manipulation tasks. We furthermore evaluate training PointZero from scratch to isolate the benefits of our proposed architecture from those of our proposed pre-training objective and dataset. We release the dataset, model checkpoints, and full training recipe.


POISE: Instance-Specific Prompt Tuning under Latent Mixture Target Distributions

wenhao zheng ⋅ shuai xu ⋅ Lixing Chen ⋅ Xiu Su ⋅ Shichao Kan ⋅ Xiaoting Lyu ⋅ Zhe Qu

Prompt tuning (PT) provides a parameter-efficient way to adapt frozen large language models, but its behavior becomes less clear when task contains multiple related yet distinguishable patterns. We formalize this setting as a latent mixture target distribution, where the target behavior is modeled as a finite mixture of latent task components. We first show that static prompting can retain a non-vanishing approximation gap when one shared prompt is insufficient to capture the target distribution. Then we formalize instance-specific prompting as an input-dependent prompt mapping and show that, under ideal prompt assignment, its achievable regret can approach zero. Guided by this analysis, we propose POISE (Prompt Optimization via Instance-Specific Experts), a structured and parameter-efficient approximation to instance-specific prompting. POISE represents the prompt space with learnable expert prompts, combines them through input-dependent routing, and introduces a shared low-rank residual correction to capture common structure. Extensive experiments across multiple benchmarks show that POISE achieves strong performance while maintaining high parameter efficiency.


PolarScale: A Physics-Grounded Benchmark for Radiometrically Consistent RGB-to-Stokes Estimation

Beibei Lin ⋅ Tingting Chen ⋅ Xin Zhang ⋅ Wenhao Zhao ⋅ DONGJUN LI ⋅ Zifeng Yuan

Polarization imaging provides rich physical cues beyond RGB imaging, yet acquiring such information typically requires specialized polarization-sensitive hardware that increases system cost and deployment complexity. Recent methods have proposed to infer polarization information directly from standard RGB images, providing a hardware-free alternative. While useful for describing relative polarimetric structure, these representations do not fully assess whether estimated images preserve the radiometric scale and intensity-dependent physical information needed for full Stokes reconstruction. To address this limitation, we introduce PolarScale, a physics-grounded benchmark for radiometrically consistent RGB-to-Stokes estimation. Instead of treating polarization inference from RGB images as conventional image-to-image translation, PolarScale evaluates how scale-independent descriptors, scale-dependent Stokes components, and radiometric scale jointly support physically meaningful Stokes reconstruction. Building upon this formulation, we propose three prediction strategies, namely direct, joint, and decoupled prediction, to study how scale-independent and scale-dependent polarimetric cues should be modeled across representative restoration-based and generative frameworks. Our analysis reveals that restoration-based frameworks outperform generative ones for radiometrically consistent polarization estimation (e.g., 28.61 vs. 26.14 dB PSNR). Furthermore, we find that jointly modeling scale-independent and scale-dependent parameters yields more stable physical representations than direct scale-dependent prediction (e.g., 28.61 vs. 26.81 dB PSNR). Overall, PolarScale provides a physics-grounded benchmark foundation for evaluating whether vision models can recover radiometrically consistent polarimetric signals from standard RGB images.


Political Neutrality as Balanced Approval: A Large-Scale Human Evaluation of AI Responses

Jonathan Stray ⋅ David Yang ⋅ Steven Luo ⋅ Miu N Takagi ⋅ Serina Chang

As AI systems increasingly shape political views, defining and evaluating AI political neutrality is an urgent problem. Here, we propose a new definition of AI political neutrality and design a large-scale user study to test it, releasing a new dataset PARETO with 6,696 participants and 187,488 evaluations of AI responses. Our definition follows a simple principle grounded in political theory: when asked about a controversial issue, an AI model should generate responses that maximize approval across groups with opposing viewpoints, while balancing approval between groups. This maximum equal approval definition allows empirical testing of whether an AI response is “neutral” and generalizes to any political context without pre-supposing a single left-right axis of division. We construct a benchmark of controversial US issues, with prompts sourced from politically charged questions on Reddit and responses from frontier AI models, and recruit human participants to rate AI responses. Across all 20 issues, we find that it is possible for AI responses to achieve high rates of approval on both sides, even as those sides disagree strongly with each other on the substance of the issues. We also find that default responses lean liberal for GPT, Gemini, Claude, and Llama, but not Grok, and that user prompts with political charges are harder to respond to than neutral prompts. This work introduces a rigorous definition and benchmark of AI political neutrality, and a dataset to measure progress toward it.


POP: Online Structural Pruning Enables Efficient Inference of Large Foundation Models

Yi Chen ⋅ Wonjin Shin ⋅ Shuhong Liu ⋅ Tho Mai ⋅ Jeongmo Lee ⋅ Jun Liu ⋅ Chuanbo Hua ⋅ Kun Wang ⋅ Joo-Young Kim

Large foundation models (LFMs) achieve strong performance through scaling, yet current structural pruning methods derive fixed pruning decisions during inference, overlooking sparsity patterns that emerge in autoregressive token generation. We propose POP (Partition-guided Online Pruning), an efficient online structural pruning framework that enables context-conditioned dynamic pruning with minimal computational overhead. POP partitions model channels into retained, candidate, and pruned regions, where prefilling defines a coarse pruning partition, and decoding generates a fine-grained mask within the candidate region, avoiding full-channel re-evaluation. The coarse pruning partition preserves consistently important weights, while the fine-grained masking provides context-conditioned variation during decoding. Moreover, POP is a lightweight, plug-and-play method that requires no preprocessing, including offline calibration, retraining, or learning predictors. Extensive evaluations across diverse LFMs, including large language models (LLMs), mixture-of-experts models (MoEs), and vision–language models (VLMs), demonstrate that POP achieves competitive or superior accuracy over existing pruning approaches, particularly on generation tasks, while incurring minimal computational overhead and maintaining efficient inference. The source code will be publicly available.


Position: A Safe LLM and a Safe Harness Do Not Make a Safe Agent

Vincent Siu ⋅ Kyle Montgomery ⋅ Yujin Potter ⋅ Zhun Wang ⋅ Dawn Song ⋅ Chenguang Wang

AI agents combine a learned neural component, such as a language model, with symbolic components, such as memory, tools, and environments. Existing safety alignment work improves both sides independently: neural alignment reduces unsafe model behavior, while symbolic alignment constrains memory, tool use, permissions, and action. This position paper argues that safe neural and symbolic components do not necessarily compose into a safe agent. Agent safety is often contextual: whether an answer, retrieval, or tool call is safe can depend on broader safety context, e.g., prior interactions, provenance, user boundaries, and authorization. For example, a biosafety AI research assistant might answer separate, benign-looking questions about viral delivery systems, protocol troubleshooting, and relevant literature, while the accumulated interaction begins to support the design of a dangerous pathogen. We propose contextual agent alignment as a research agenda for identifying safety-relevant context, designing safer agents that use this context, and evaluating contextual safety. Several documented agent safety failures, including cumulative dual-use risk in research assistants, peer-preservation in oversight settings, and contextual privacy violations, can be viewed as instances of contextual safety failure. This agenda is important for the broader AI safety community because frontier systems are increasingly deployed as agents in real-world settings.

Modern frontier AI has been driven by a powerful but increasingly narrow recipe: scale the model, extend the context, and rely on emergent capabilities to support more complex forms of reasoning. \textbf{This position paper argues that this recipe is approaching a structural bottleneck. Robust reasoning requires more than parameter count, stored knowledge, or longer context windows. It requires intrinsic temporal structure i.e. the capacity of a system to generate, inhabit, and reason relative to its own internal time.} We argue that current architectures lack this capacity not merely as an implementation detail, but because most are built around fixed-pass or externally clocked computation rather than self-clocked, persistent, and dynamically coordinated processing. Our position is motivated by three converging lines of evidence. First, psychometric research links intelligence to the allocation and precision of processing time. Second, Global Workspace Theory suggests that flexible cognition depends on competitive access to a limited shared present and third, neural binding theories show that integration across distributed processors requires temporal co-activation rather than spatial co-location alone. Together, these motivate a two-level architectural view in which a diverse long-term population of specialist processors develops private temporal dynamics, while a short-term workspace constructs a shared present moment through selective competition and broadcasts it back to guide subsequent processing. We argue that existing work on continuous thought machines, conscious Turing machines, adaptive computation, memory-augmented models, mixture-of-experts, and quality-diversity neuroevolution already contains many of the necessary ingredients, but has not yet coupled them around intrinsic time as the organising principle. We outline falsifiable predictions for this view, including that harder reasoning problems should require deeper internal temporal trajectories, that synchronisation patterns should track task structure, and that self-clocked architectures should exhibit stronger adaptation under equal compute than fixed-pass alternatives. We call on the AI community to treat intrinsic temporal structure as a first-class architectural objective, not as an incidental by-product of scale.


Position: Telic Errors Make VLMs Unreliable Annotators in Sensitive Contexts

Kokil Jaidka ⋅ Insyirah Mujtahid ⋅ Chebrolu Niranjan ⋅ Sahajpreet Singh

VLMs now annotate sensitive content at scale, be it hate-speech memes, conflict imagery, politically charged symbols; yet, standard pipelines measure only whether a model can identify what an image depicts, not what it means to the community it concerns. We name this failure mode a telic error: the model produces a surface-accurate, perceptually complete annotation (telic: what-it-depicts) while erasing the relic meaning, i.e., the symbolic, cultural, or political significance that makes the content consequential. We evaluate six open-weight VLMs on a two-task diagnostic designed to locate the failure precisely: Task 1 (forward annotation) shows that prepending the model's own image description does not improve accuracy (-2.0pp), ruling out description-generation failure; Task 2 (label-supplied explanation) shows the same models achieve 95-100% relic identification accuracy when given the correct label and asked to explain it, ruling out knowledge absence. The 35-percentage-point gap between the two tasks is the empirical signature of a protocol artefact: models possess the cultural knowledge required for accurate annotation, but standard binary formats never ask them to deploy it. We argue that measuring only telic accuracy is not a neutral technical choice; it is a choice about whose meanings count, and that the dual-axis evaluation framework, five-type error taxonomy, and two-task diagnostic introduced here are the instrumentation needed to make that choice visible.


Posterior Contraction Rates for sparse Kolmogorov-Arnold Networks in Anisotropic Besov Spaces

Jeunghun Oh ⋅ Lizhen Lin ⋅ Jaeyong Lee ⋅ Kyeongwon Lee

We study posterior contraction rates for sparse Bayesian Kolmogorov-Arnold networks (KANs) over anisotropic Besov spaces, providing a statistical foundation of KANs from a Bayesian point of view. We show that sparse Bayesian KANs equipped with spike-and-slab-type sparsity priors attain the near-minimax posterior contraction. In particular, the contraction rate depends on the intrinsic anisotropic smoothness of the underlying function. Moreover, by placing a hyperprior on a single model-size parameter, the resulting posterior adapts to unknown anisotropic smoothness and still achieves the corresponding near-minimax rate. A distinctive feature of our results, compared with those for standard sparse MLP-based models, is that the KAN depth can be kept fixed: owing to the flexibility of learnable spline edge functions, the required approximation complexity is controlled through the network width, spline-grid range and size, and parameter sparsity. Our analysis develops theoretical tools tailored to sparse spline-edge architectures, including approximation and complexity bounds for Bayesian KANs. We then extend to compositional Besov spaces and show that the contraction rates depend on layerwise smoothness and the effective dimension of the underlying compositional structure, thereby effectively avoiding the curse of dimensionality. Together, the developed tools and findings advance the theoretical understanding of Bayesian neural networks and provide rigorous statistical foundations for KANs.


Posterior Sampling-based Online Learning for Episodic POMDPs

Dengwang Tang ⋅ Dongze Ye ⋅ Rahul Jain ⋅ Ashutosh Nayyar ⋅ Pierluigi Nuzzo

Learning in POMDPs is known to be significantly harder than in MDPs. In this paper, we consider the online learning problem for episodic POMDPs with unknown transition and observation models. We propose a Posterior Sampling-based reinforcement learning algorithm for POMDPs (PS4POMDP), which is much simpler and more implementable compared to state-of-the-art optimism-based online learning algorithms for POMDPs. We show that the Bayesian regret of the proposed algorithm scales as the square root of the number of episodes and is polynomial in the other parameters. In a general setting, the regret scales exponentially in the horizon length $H$, and we show that this is inevitable by providing a lower bound. However, when the POMDP is undercomplete and weakly revealing (a common assumption in the recent literature), we establish a polynomial Bayesian regret bound. We finally propose a posterior sampling algorithm for multi-agent POMDPs, and show it too has sublinear regret.


PPO in the Fisher-Rao geometry

Razvan-Andrei Lascu ⋅ David Siska ⋅ Lukasz Szpruch

Proximal Policy Optimization (PPO) is widely used in reinforcement learning due to its strong empirical performance, yet it lacks formal guarantees for policy improvement and convergence. PPO's clipped surrogate objective is motivated by a lower bound on linearization of the value function in flat geometry setting. We derive a tighter surrogate objective and introduce Fisher-Rao PPO (FR-PPO) by leveraging the Fisher-Rao (FR) geometry. Our scheme provides strong theoretical guarantees, including monotonic policy improvement. In the direct parametrization setting, we show that FR-PPO achieves sub-linear convergence, and for parametrized policies we further obtain sub-linear convergence up to the compatible function approximation error. Finally, although our primary focus is theoretical, we also demonstrate empirically that FR-PPO performs well across a range of standard reinforcement learning tasks.

We introduce P-IMLA, a preconditioned implicit-midpoint Langevin sampler for high-dimensional log-concave targets with non-smooth regularization. The obstruction is that the large steps enabled by matrix preconditioning are precisely where preconditioned Euler--Maruyama inflates stationary variance. Implicit midpoint removes this surplus via a Cayley-transform identity, making P-IMLA Gaussian-exact at any step size. Beyond Gaussians, a Möbius-monotonicity argument gives an $M$-Wasserstein rate governed by $\kappa_{\rm eff}$ rather than the $\kappa_f$ of unpreconditioned IMLA, while backward-error analysis shows drift-only rather than Euler-type diffusion bias. An $M$-norm Moreau envelope yields a proximal non-smooth variant with a three-term error budget. Controlled Gaussian and TV-regularized imaging experiments confirm the predicted separation: the unpreconditioned IMLA degenerates as $\kappa_f$ grows, whereas P-IMLA remains $\kappa_{\rm eff}$-governed and near variance-exact.


Prediction Under Imperfect Compression: A Theory of Approximate MDL

Qian Li ⋅ Xinyu Mao ⋅ Shang-Hua Teng ⋅ Guangxu Yang

Minimum Description Length (MDL) formalizes the principle of Occam's razor by optimizing the total description length: $L(\mathrm{model})+L(\mathrm{data} \ | \ \mathrm{model})$. For sequential prediction, the MDL method repeatedly selects a model with a minimum objective score of the observed prefix for the next step prediction. Classical MDL prediction theory shows that exact optimization of the MDL objective indeed provides a strong compression guarantee that supports reliable prediction. However, practical machine learning usually can only find models by approximately optimizing the objective function. To bridge this gap, this paper addresses the following fundamental question: \emph{Under what forms of approximation and regularization does approximate MDL still guarantee reliable sequential prediction?} This work offers a principled characterization. We prove that for any approximation with \emph{additive} slack \(C\) of the more general form of the \emph{balanced} MDL objective: $\lambda\cdot L(\mathrm{model})+L(\mathrm{data} \ | \ \mathrm{model})$, the cumulative expected squared prediction error is finite for all $\lambda\ge1$. The case $\lambda>1$ is proved by an affinity-telescoping argument, while the boundary case $\lambda=1$ is proved by a likelihood-ratio stopping argument based on exact static MDL bounds. Our results establish that classical MDL regularization remains robust to any fixed additive optimization error. Furthermore, we establish that our characterization of the approximate MDL framework is sharp: When $0<\lambda<1$, overfits can happen to incur infinite cumulative expected error in the universal class of estimable measures, and hence a strong form of model-complexity regularization is necessary. In addition, model selection may fail in every regularized regime $\lambda >0$, under multiplicative approximation, and thus, additive approximation is both sufficient and essential.


Preserving Geometric Symmetry in Uncertainty Estimation for Molecular Forces

Xinyu Li ⋅ Zhen Zhang ⋅ Jiayu Huang ⋅ Daniel M Steinberg ⋅ Lina Yao ⋅ Prof Javen Qinfeng Shi

Accurate and reliable uncertainty estimation (UE) for vector-valued physical properties is crucial for scientific discovery in fields like drug and materials discovery. For example, atomic forces are central for finding optimal structures and identifying equilibrium systems, and a fundamental requirement of these vectors is that they must be equivariant to 3D rotations. However, existing UQ methods often fail to respect these geometric constraints, leading to poorly calibrated uncertainty and degraded predictive performance. To address the problem, we introduce a novel framework for equivariant multivariate evidential regression. Our primary contribution is a theoretically grounded parameterization of the evidential prior's scale matrix, constructed via a decomposition into an invariant scalar component and an equivariant low-rank term. Furthermore, we address the lack of suitable evaluation standards for vector fields by introducing a rigorous calibration metric based on the Euclidean distance. Extensive experiments on molecular property prediction benchmarks, including MD17, QM7-X and OC20, show that our framework consistently outperforms established baselines, achieving state-of-the-art predictive accuracy with well-calibrated uncertainty. This work provides a principled and practical approach to uncertainty estimation for equivariant vector-valued properties, paving the way for more trustworthy machine learning applications in scientific discovery.


Prevailing Bisimulation Metric Learning Is Biased: Implicit Regularization and Its Remedy

Junqi Lu ⋅ Ruixiang Sun ⋅ Xin Li ⋅ Gaopeng Peng ⋅ Ao Zhang ⋅ Mingzhong Wang

Deep bisimulation metric learning has emerged as a principled framework for learning robust state representations in reinforcement learning. However, prevailing methods rely on sample-based regression objectives that exhibit a critical stability flaw in stochastic environments. We identify this flaw as a manifestation of the double sampling problem: minimizing mean squared error against single-sample stochastic targets introduces an irreducible variance bias. Through the lens of stochastic differential equations (SDEs), we rigorously prove that this bias leads to a Variance Trap, wherein target variance scales with the encoder's sensitivity. This variance effectively acts as an implicit Jacobian regularizer, driving representation collapse in the presence of transition noise. To address this issue, we propose Bisimulation Saddle-Point Optimization (BSPO), a debiased framework that reformulates the bisimulation metric learning objective as a primal-dual saddle-point problem. By introducing an Auxiliary Network (AuxNet) to estimate the expected Bellman error, BSPO decouples the structural error from transition noise, eliminating the bias without requiring physically infeasible double sampling. To the best of our knowledge, this work is the first to identify that prevailing sample-based bisimulation metric learning objectives are systematically biased and to provide a practical debiasing framework. Our empirical evaluations on stochastic continuous/discrete control tasks and visually distracting environments demonstrate that BSPO significantly outperforms representative baselines, particularly in high-noise regimes where conventional methods fail to converge.


PRISM: A Benchmark for Programmatic Spatial-Temporal Reasoning

Qiran Zhang ⋅ Yuheng Wang ⋅ Runde Yang ⋅ Lin Wu ⋅ Jingru Fan ⋅ Shu Yao ⋅ Jie Zhang ⋅ Tianle Zhou ⋅ Huatao Li ⋅ Ruijie Shi ⋅ Yihan Li ⋅ Chen Qian

Programmatic video generation through code offers geometric precision and temporal coherence unattainable by pixel-level diffusion models, yet rigorously evaluating whether language models can produce spatially correct animated output remains an open problem. We introduce **PRISM**, a large-scale benchmark of 10,372 human-calibrated instruction-code pairs ($20\times$ larger than prior Manim benchmarks) grounded in real-world educational scenarios across English and Chinese, spanning 437 subject categories. Alongside the dataset, we propose a funnel-style evaluation framework with four complementary metrics: *Code-Level Reliability* as the executability gate, *Spatial Reasoning* as the core end-to-end metric measuring layout correctness over full animation sequences, and *Prompt-Aware Dynamic Visual Complexity* (PADVC) and *Temporal Density* (TD) as diagnostic dimensions characterizing dynamic expression and temporal activity. Systematic evaluation of seven mainstream LLMs reveals a striking *Execution-Spatial Gap*: the average drop from execution success rate to spatial pass rate is approximately 41\%, demonstrating that syntactically correct, runnable code does not automatically yield spatially coherent visual output. No single model dominates all dimensions; notably, the model with the highest execution rate exhibits the largest Execution-Spatial Gap, while the most spatially robust model does not lead on code reliability. These findings highlight that evaluating programmatic video generation should not stop at executability. PRISM provides a principled benchmark for advancing spatially coherent code generation.

Text image super-resolution (Text-SR) requires more than visually plausible detail synthesis: slight errors in stroke topology may alter character identity and break readability. Existing methods improve text fidelity with stronger recognitive or generative priors, yet they still face two unresolved challenges under severe degradation: the text condition extracted from low-quality inputs can itself be unreliable, and a plausible global prior does not fully determine fine-grained stroke boundaries. We present PRISM, a single-step diffusion-based Text-SR framework that addresses these two challenges through Flow-Matching Prior Rectification (FMPR) and a Structure-guided Uncertainty-aware Residual Encoder (SURE). FMPR constructs a privileged training-time prior from paired low-quality/high-quality latents and learns a flow matching that transports degraded embeddings toward this restoration-oriented prior space, yielding more accurate and reliable global text guidance. SURE further predicts uncertainty-aware structural residuals to selectively absorb reliable local boundary evidence while suppressing ambiguous stroke cues. Together, these components enable explicit global prior rectification and local structure refinement within a single diffusion restoration pass. Experiments on both synthetic and real-world benchmarks show that PRISM achieves state-of-the-art performance with millisecond-level inference.

Trilevel learning has emerged as a powerful paradigm across various domains in machine learning and networking. In practical scenarios, data is frequently generated and distributed across multiple decentralized nodes. While existing distributed trilevel learning methods circumvent the transmission of raw data, they remain susceptible to privacy leakage. Furthermore, cutting plane methods, which are widely adopted in bilevel and trilevel learning, exhibit analogous intrinsic privacy risks. To address these critical issues, we propose a Differentially Private Distributed Trilevel Optimization (DPDTO) framework. In DPDTO, we first develop a value-function-embedded privacy-preserving cutting plane method for trilevel optimization problems, comprising an inner-layer value function reformulation and an outer-layer cutting plane relaxation. Building upon this, a distributed optimization algorithm is introduced to effectively address trilevel optimization while preserving privacy. Theoretically, we provide a comprehensive analysis for the proposed framework, including asymptotic convergent relaxation, non-asymptotic convergence rate, privacy guarantees, and communication complexity, uncovering a tripartite trade-off among privacy, utility, and complexity in trilevel optimization. Extensive experimental results on two distributed trilevel optimization problems further demonstrate the effectiveness and superiority of the proposed DPDTO.


Privately Clipping Heavy-Tailed Data

Matthew Joseph ⋅ Alex Kulesza

Summing vectors is a basic task in differentially private data analysis. Standard algorithms clip vectors to some bound, sum them, and add noise scaled to the bound. While the differential privacy literature has developed a deep library of such noise addition mechanisms, it offers few tools for privately choosing the clipping bound directly from the data. We introduce a method that privately chooses a clipping bound by optimizing residual coherence, the directional alignment of vectors exceeding the bound, against the variance cost of additional noise. We prove a close relationship between residual coherence and bias for a general class of heavy-tailed data regimes and show empirically that, across several datasets, our method outperforms existing baselines.


Probabilistic Circuits for Irregular Multivariate Time Series Forecasting

Christian Klötergens ⋅ Lars Schmidt-Thieme ⋅ Vijaya Krishna Yalavarthi

Joint probabilistic modeling is essential for forecasting irregular multivariate time series (IMTS) to accurately quantify uncertainty. Existing approaches often struggle to balance model expressivity with consistent marginalization, frequently leading to unreliable or contradictory forecasts. To address this, we propose CircuITS, a novel architecture for probabilistic IMTS forecasting based on probabilistic circuits. Our model is flexible in capturing intricate dependencies between time series channels while structurally guaranteeing valid joint distributions. Experiments on four real-world datasets demonstrate that CircuITS achieves superior joint and marginal density estimation compared to state-of-the-art baselines.

The signature transform is a principled feature map for continuous-time paths, valued for its uniqueness and universality. Recovering a path from its truncated signature is, however, structurally ill-posed because the truncated signature map is not injective. We therefore reformulate truncated signature inversion as a probabilistic problem---learning the conditional distribution of a path given its truncated signature---and adopt a signature-conditioned flow matching model as a practical estimator. This probabilistic reformulation elucidates the fundamental difficulty of inversion: Bayes reconstruction error quantifies the irreducible uncertainty remaining after conditioning on a statistic. We derive the Bayes-optimal error under linear statistics, obtaining a closed form for log-GBM and numerically tractable formulas for log-fBM and OU---yielding a concrete theoretical baseline for model validation. This baseline upper-bounds the Bayes error under truncated-signature conditioning, since truncated signatures provide richer information than linear statistics. Experiments show that empirical reconstruction errors under linear-statistics conditioning closely match the theory-derived baseline, while errors decrease when the statistic is replaced with the truncated signature. Moreover, generated paths faithfully recover the conditioning signature while preserving key distributional and temporal structure, indicating that the estimator is well calibrated to the target conditional distribution. Together, these results establish a well-posed probabilistic framework for truncated-signature inversion, with applicability demonstrated on real financial data beyond the parametric process families covered by theory.


Probe-Guided Gradient Balancing for Multimodal Learning

Vu Vo ⋅ Haytham Fayek ⋅ Thuy T Nguyen

Jointly trained multimodal networks frequently underutilize their weaker modalities, as faster-learning modalities dominate the shared fusion objective and suppress the gradient signal reaching slower ones. Most existing approaches throttle the dominant modality, whereas directly boosting the weaker modality often fails because monitoring signals derived from the joint loss are entangled with, and thus dominated by, the stronger modality. We propose a new gradient modulation approach, probe-guided gradient boosting (PGGB), to address the problem by probing and boosting the weak modalities while preserving the strength of the dominant modality. Probing is performed by attaching a lightweight linear classifier to each modality’s stop-gradient features to estimate the corresponding representation quality score. The unbiased gap between these scores across modalities, the utilization gap, serves as an imbalance signal that defines a bounded and smoothed scaling factor adaptively reweighing the gradients of weaker modalities, enabling stable and effective rebalancing during joint optimization. The method is composable with throttling methods as the intervention leaves the loss and fusion architecture unchanged. Extensive experiments were conducted across eight benchmarks spanning four domains and 2–4 modalities. The results show that our method outperforms state-of-the-art approaches on highly imbalanced multimodal datasets, while remaining competitive on benchmarks with low or no modality imbalance. Bounded scaling, self-attenuation, and a standard-SGD descent bound were established under standard smoothness and bounded-variance assumptions. Code is provided in the supplementary material.


Procedural Refinement by LLM-driven Algorithmic Debugging for ARC-AGI-2

Yuning Qiu ⋅ Lin-Feng Zou ⋅ Jiong-Da Wang ⋅ Xue-Rong Yuan ⋅ Wang-Zhou Dai

In high-complexity abstract reasoning, a system must infer a latent rule from a few examples or structured observations and apply it to unseen instances. LLMs can express such rules as programs, but ordinary conversation-based refinement is largely outcome-level: it observes that an answer or output is wrong without formally re-checking which abstraction, relation, or transformation justified that outcome. We propose \emph{Abduction-Based Procedural Refinement} (ABPR), a neuro-symbolic refinement approach that couples an LLM with a Prolog meta-interpreter. ABPR treats each candidate program as an executable declarative hypothesis of the latent rule and reifies its SLD goal--subgoal resolution into compact proof-tree-style derivations, following Shapiro's algorithmic program debugging (APD). In this view, refinement is not merely code-level debugging, but semantic re-checking of the model's hypothesised rule. We evaluate ABPR primarily on ARC-AGI-2, a challenging few-shot abstract rule induction benchmark over grid transformations. ABPR with Gemini-3-Flash achieves 56.67\% Pass@2, while GPT-5.5 xHigh with ABPR reaches 98.33\% Pass@2 on the public evaluation set. Supplementary experiments on fill-in-the-blank I-RAVEN-X and A-I-RAVEN adaptations provide evidence that the same trace-guided framework extends beyond ARC-specific grid tasks to RAVEN-style relational and analogical abstraction. Repeated-run and sensitivity analyses show that parallel trace-guided search reduces stochastic variance as search breadth and refinement depth increase.


ProCTI: Prototype-Refined Global Conditioning for Diffusion-Based Time Series Imputation

Fariza Rashid ⋅ Duc Van Le ⋅ Rahat Masood ⋅ Gustavo E Batista ⋅ Aruna Seneviratne ⋅ Suranga Seneviratne

Time series imputation has progressed from statistical and deep learning approaches to diffusion-based models, which have shown strong recent performance. Existing diffusion-based methods typically condition the reverse process using local contextual information from the current or neighbouring windows. Meanwhile, global dataset-level structure often remains implicit, limiting performance when local observations are sparse, noisy, or unrepresentative. To address this issue, we propose ProCTI, a diffusion-imputation framework that augments local conditioning with retrieved global dataset-level priors through learned prototypes. A hybrid conditioning mechanism integrates this global context with local signals during reverse diffusion, enabling more accurate reconstruction under varying missingness scenarios. Experiments across multiple benchmark datasets show that ProCTI outperforms strong baselines overall under random missingness, while remaining competitive under attribute-wise missingness. We further introduce a theoretical framework that distinguishes between local and global conditioning in diffusion-based imputation, showing that combining both can reduce conditional uncertainty and improve imputation quality.


Prototype Topology Consistency for Visible-Infrared Lifelong Person Re-Identification

Tan Shang ⋅ Zhijie Lu ⋅ Wuxuan Shi ⋅ He Li ⋅ Mang Ye

Visible-Infrared Lifelong Person Re-Identification (VI-LReID) requires a model to sequentially acquire cross-modal person retrieval capabilities across sequentially arriving tasks while retaining performance on previously learned ones.Existing methods face two fundamental problems.First, new-task gradient updates disrupt each instance's assignment ranking over historical class prototypes, which is the direct mechanism behind past-task retrieval performance degradation. At the same time, the feature manifold inevitably undergoes global drift as a necessary byproduct of new-task adaptation. Effective anti-forgetting should target the detrimental ranking disruption while permitting this necessary drift.Second, visible RGB and thermal infrared (TIR) imaging capture fundamentally different physical cues: texture-and-color reflectance versus heat radiation contours. This causes the two modalities to form inherently inconsistent inter-class similarity structures that conventional cross-modal supervision cannot resolve, manifesting as highly asymmetric forgetting across modalities under shared-backbone continual training.We propose S-PTC (Stage-wise Prototype Topology Consistency), which preserves each instance's relative distribution over the frozen historical prototype set, protecting assignment rankings while leaving the overall feature manifold free to evolve for new-task adaptation.We further propose CM-PTC (Cross-Modal Prototype Topology Consistency), which aligns the inter-class topology matrices of RGB and IR prototypes, mitigating the structural inconsistency rooted in their imaging physics.Both modules require no stored exemplars.Extensive experiments on the VI-LReID benchmark verify the effectiveness and superiority of our approach against state-of-the-art methods.


Provable Selective Auto-labeling with Reliability Guarantees

Huipeng Huang ⋅ Wenbo Liao ⋅ Huajun Xi ⋅ Hao Zeng ⋅ Mengchen Zhao ⋅ Hongxin Wei

Auto-labeling has emerged as a popular, cost-effective alternative to manual annotation. However, its performance is often highly inconsistent across different tasks, underscoring the importance of selective strategies. Despite this, existing heuristic methods for selective auto-labeling still rely heavily on model confidence scores and offer no reliable guarantee on the trustworthiness of the selected examples. To address this, we propose $\textbf{Conformal Labeling}$, a novel method that selects a subset of auto-generated labels with a provably controlled false labeling rate (FLR). Our key idea is to formulate selective auto-labeling as a multiple-hypothesis testing problem, where each hypothesis indicates whether to accept the auto-generated label for an instance. Specifically, we construct a conformal $p$-value for each auto-generated label using a small calibration set, and then employ the Benjamini–Hochberg (BH) procedure with a novel correction factor to construct the subset with guaranteed FLR control. Theoretically, we show that our method achieves finite-sample FLR control and asymptotically optimal statistical power among all $p$-value thresholding rules. Extensive experiments on classification and open-ended generation tasks validate the effectiveness of our method, achieving high statistical power while strictly controlling the FLR.

Noise search at inference time improves text-to-image diffusion quality but requires many multi-step teacher rollouts per query. Distilled few-step models are now routinely released for efficient inference, but their potential as proxies inside noise search is uncharacterised. We measure proxy--teacher rank agreement across three distillation regimes and three backbones spanning DiT, MMDiT, and UNet, and find a sharp asymmetry: eight closed-form image statistics preserve rank from the distilled proxy to the teacher (Spearman $\rho$ up to $0.80$), while learned semantic scorers decorrelate ($\rho \le 0.12$). Null controls localise the transfer to distillation training rather than few-step denoising or initial noise characteristics. We therefore reformulate noise search as a \emph{joint constraint} and propose ProxySearch: the proxy filters candidates by statistical match in $1$ step, the teacher argmaxes a user-chosen semantic reward over the survivors without a tunable weight trading semantic reward against the statistical target. \textsc{ProxySearch} matches or beats teacher-only argmax on compositional accuracy and human preference, while tightening the statistical target by $\sim 23\%$ in paired statistical $\ell_1$ relative to teacher-only argmax at matched compute on FLUX-$1024$.


PULSE: A Synchronized Five-Modality Dataset for Sensorimotor Coordination in Long-Horizon Daily Activities

Wolin Liang ⋅ Yang Gao ⋅ Peiyu Yan ⋅ YueXiang Hu ⋅ Yingjing Xiao ⋅ Xiongfeng Ying ⋅ YingNian Guo ⋅ Cheng Zhang ⋅ Zhanpeng Jin

Existing human-activity datasets typically cover only one or two sensor channels and often focus on short, isolated actions, leaving long-duration compositional tasks largely unsupported as benchmarks for temporal cross-modal sensorimotor learning. We introduce PULSE ($\mathbf{P}$hysiological $\mathbf{U}$nified $\mathbf{L}$ong-duration $\mathbf{S}$ynchronized $\mathbf{E}$mbodiment), a synchronized five-modality dataset that enables benchmarking temporal cross-modal sensorimotor learning: how models recognize ongoing activities, anticipate future contact, reconstruct unobserved signals, and remain robust when physiological, motion, gaze, inertial, and tactile channels are partially observed. Collected from 40 volunteers across 8 ecologically valid daily-activity scenarios together with a dedicated motion-primitive collection, PULSE records full-body optical motion capture with finger-level hand articulation, surface EMG, binocular eye tracking, wearable IMU, and a fingertip pressure array recording quantitative grip force, all sampled at 100 Hz. Each scenario is a long-horizon compositional task containing many smaller sub-tasks, yielding over 7,700 annotated action segments. Each segment is tagged with a motor primitive, the hand involved, the manipulated object, and a natural-language description with four paraphrased variants. On top of the dataset, we define multiple benchmark tasks for temporal cross-modal sensorimotor learning: scene and fine-grained action recognition, grasp onset anticipation, missing-modality robustness, tactile-driven sub-second grasp-state prediction, cross-modal pressure reconstruction, etc. Together these tasks evaluate how information transfers across physiology, motion, gaze, and touch, rather than reducing the dataset to an activity-label catalog. We evaluate three backbone architectures, nine fusion strategies, seven published baselines, and task-specific models including SyncFuse and DailyActFormer. Detailed task definitions and results are in the Appendices. Data, annotations, and baseline code will be publicly released.


Q-Probe: Scaling Image Quality Assessment to High Resolution via Context-Aware Agentic Probing

Xiang Li ⋅ Xueheng Li ⋅ Yu Wang ⋅ Xuanhua He ⋅ Zhangchi Hu ⋅ Chengjun Xie

Reinforcement Learning (RL) has empowered Multimodal Large Language Models (MLLMs) to achieve superior human preference alignment in Image Quality Assessment (IQA). However, existing RL-based IQA models typically rely on coarse-grained global views, failing to capture subtle local degradations in high-resolution scenarios. While emerging "Thinking with Images" paradigms enable multi-scale visual perception via zoom-in mechanisms, their direct adaptation to IQA induces spurious "cropping-implies-degradation'' biases and misinterprets natural depth-of-field as artifacts. To address these challenges, we propose Q-Probe, the first agentic IQA framework designed to scale IQA to high resolution via context-aware probing. First, we construct Vista-Bench, a pioneering benchmark tailored for fine-grained local degradation analysis in high-resolution IQA settings. Furthermore, we propose a three-stage training paradigm that progressively aligns the model with human preferences, while simultaneously eliminating causal bias through a novel context-aware cropping strategy. Extensive experiments demonstrate that Q-Probe achieves state-of-the-art performance in high-resolution settings while maintaining superior efficacy across resolution scales.


Quality-Diversity Optimization as Multi-Objective Optimization

Xi Lin ⋅ Ping Guo ⋅ Yilu Liu ⋅ Bo Xue ⋅ Qingfu Zhang ⋅ Jianyong Sun

The Quality-Diversity (QD) optimization aims to discover a collection of high-performing solutions that simultaneously exhibit diverse behaviors within a user-defined behavior space. This paradigm has stimulated significant research interest and demonstrated practical utility in domains including robot control, creative design, and adversarial sample generation. A variety of QD algorithms with distinct design principles have been proposed in recent years. Instead of proposing a new QD algorithm, this work introduces a novel reformulation by casting the QD optimization as a multi-objective optimization (MOO) problem with a huge number of optimization objectives. By establishing this connection, we enable the direct adoption of well-established MOO methods, particularly set-based scalarization techniques, to solve QD problems through a collaborative search process. We further provide a theoretical analysis demonstrating that our approach inherits theoretical guarantees from MOO while providing desirable properties for the QD optimization. Experimental studies across several QD applications confirm that our method achieves performance competitive with state-of-the-art QD algorithms.

Large vision-language models (LVLMs) rely on increasingly long visual contexts, making the KV cache a major inference-time bottleneck. Existing KV cache compression methods typically use pre-decode signals such as prefill attention, visual redundancy, objectness, or layer-wise allocation, but these criteria do not directly indicate which visual tokens in the KV cache will be reused by future answer tokens. We introduce Q-ViK, a question-conditioned visual utility framework for LVLM KV cache eviction. During offline training, we run full-cache decoding and aggregate answer-to-visual attention over cached visual tokens, yielding a privileged future-utility signal that reflects the visual KV entries actually used during generation. We use this signal to train a lightweight question-conditioned scorer that predicts visual cache utility from prefill representations alone. At inference time, Q-ViK requires neither full-cache lookahead nor generated traces: it preserves textual KV entries and evicts low-utility visual KV entries with a single post-prefill scoring step. Experiments on seen and held-out multimodal benchmarks show that Q-ViK preserves task-relevant visual evidence in the KV cache more effectively than pre-decode saliency, especially under aggressive KV cache compression. Code will be available at https://anonymous.4open.science/r/Q-ViK-1A75/README.md.


QWaveNet: Quantum-Enhanced Wavelet Network for Time Series Forecasting

Fan Zhang ⋅ Shijun Chen ⋅ Meijia Wang ⋅ Shiming Fan ⋅ Zexuan Ma ⋅ Hua Wang

Series decomposition aims to separate long-term trajectories and periodic patterns in time series, thereby improving the specificity and interpretability of forecasting models. However, mainstream approaches usually equate this functional separation with an operational split between a smoothed baseline and fluctuating residuals. One line of methods approximates the trend via low-pass smoothing, which assumes long-term trajectories to be purely smooth and may misassign nonlinear evolution, stage-wise transitions, and structural fluctuations to the residual. The other relies on heuristic nonlinear decompositions to capture complex dynamics, but is sensitive to noise and hyperparameters, leading to unstable and mixed component boundaries. To address this issue, we propose a forecasting framework, QWaveNet. It employs a quantum convolutional neural network to learn Dynamic Evolution, a long-term structural representation beyond low-pass smoothing priors, which preserves nonlinear evolution, stage-wise transitions, and structural fluctuations more coherently, while its complementary Evolutionary Residual captures localized variations and periodic dynamics. Guided by these complementary components, QWaveNet further performs adaptive multi-scale decomposition and reconstruction to support forecasting. Experiments on multiple benchmark datasets demonstrate the effectiveness of QWaveNet.


RA-CFGCache: From Branch-Level Criteria to Guided-Risk Control under Classifier-Free Guidance

Yiming Liu ⋅ Ben Wan ⋅ Hui Chen ⋅ Ao Wang ⋅ Yuqi Xiong ⋅ Fan Zhang ⋅ Tongxuan Liu ⋅ guiguang ding

Diffusion models dominate visual generation, but iterative denoising remains computationally demanding, especially under Classifier-Free Guidance (CFG), which is crucial for high-fidelity generation yet nearly doubles per-step computation. To reduce this burden, recent training-free caching methods reuse intermediate predictions during sampling, yet their single-prediction criteria are misaligned with CFG sampling, where the denoising update is governed by the guided prediction rather than either branch alone. Consequently, such reuse may misestimate the error that actually perturbs sampling, leading to a branch-guided mismatch. Moreover, due to error propagation across timesteps, a local guided error may not faithfully reflect the final deviation, inducing a local-final mismatch. To address these two mismatches, we propose RA-CFGCache, a Risk-Aligned Caching framework for CFG. RA-CFGCache introduces CFG-aware Guided-Risk Composition to align caching decisions directly with the guided prediction, and Propagation-Aware Rescaling to calibrate local guided risks against the final generative deviation. Extensive experiments on FLUX.1-dev, Wan2.1-T2V-1.3B, and CogVideoX-2B show that RA-CFGCache delivers superior efficiency-fidelity trade-offs over existing training-free caching methods. Moreover, RA-CFGCache is compatible with diverse caching strategies and can further enhance their fidelity, e.g., improving PSNR over MagCache from 21.46 to 24.02 under CFG. The code is available at https://anonymous.4open.science/r/RA-CFGCache/.


RADAR: Routing Agents via Difficulty-Aware Recovery

Haizhou Du ⋅ Jinze Zhao ⋅ Huaicheng Yan

Multi-agent systems (MAS) have emerged as a dominant paradigm where agent routing is a bottleneck for balancing performance and cost. However, existing agent routing methods suffer from the prohibitive cost of exhaustive benchmark construction and the neglect of query complexity essential for mixed tasks. To address this issue, we propose the RADAR (Routing Agents via Difficulty-Aware Recovery) framework. RADAR enables accurate decision-making under limited evaluation cost through two complementary modules. MatrixComp module employs structure-aware matrix completion to recover global performance profiles from sparse observations. Furthermore, DiffMatch module extracts reasoning meta-features for fine-grained alignment via a difficulty-calibrated retrieval mechanism. Extensive experiments demonstrate that RADAR maintains performance equal with state-of-the-art baselines while reducing the cost of benchmark construction by at least 33%.


Random Attention Pattern Learning Enables Emergent Capabilities

Vatsal Baherwani ⋅ Charlie Chen ⋅ Shikai Qiu ⋅ Andrew Wilson ⋅ Pavel Izmailov

Neural scaling laws for transformer language models predict smooth improvements in pretraining loss with increasing parameters, but downstream capabilities such as in-context learning are known to emerge abruptly past a certain model scale. In this paper, we show that emergent capabilities arise at variable intervals throughout training, with larger models acquiring capabilities earlier on average. We demonstrate that the emergence of capabilities such as pattern completion and indirect object identification corresponds to the abrupt learning of task-relevant attention patterns. To isolate this phenomenon, we train transformer models on synthetic linear map and cellular automata datasets, and we show that the difficulty of learning attention patterns depends on context length and pattern sparsity. Moreover, scaling the number of attention heads improves learning efficiency on our synthetic tasks, while increasing the head dimension yields diminishing returns past a minimum capacity. We additionally investigate architectures with alternative attention mechanisms, showing that MLP-Mixer outperforms a transformer on linear map tasks with complex attention patterns. Our findings provide a mechanistic insight into emergence, showing that downstream capabilities arise abruptly due to the intrinsic difficulty of learning sparse attention patterns in transformer models.


Randomized Kriging Believer For Parallel Bayesian Optimization With Regret Bounds

Shuhei Sugiura ⋅ Ichiro Takeuchi ⋅ Shion Takeno

We consider the optimization problem of an expensive-to-evaluate black-box function, in which we can obtain noisy function values in parallel. For this problem, parallel Bayesian optimization (PBO) is a promising approach, which aims to optimize with fewer function evaluations by selecting a diverse input set for parallel evaluation. However, existing PBO methods suffer from poor practical performance or lack theoretical guarantees. In this study, we propose a PBO method, called randomized kriging believer (KB), based on a well-known KB heuristic and inheriting the advantages of the original KB: low computational complexity, a simple implementation, versatility across various BO methods, and applicability to asynchronous parallelization. Furthermore, we show that our randomized KB achieves Bayesian expected regret guarantees. We demonstrate the effectiveness of the proposed method through experiments, including those on real-data emulators.

We explore whether Omni-Language Models (OLMs) can be directly applied to zero-shot Semantic Audio-Visual Navigation (SAVN). Recent work demonstrates that even state-of-the-art specialized models still struggle to achieve generalist multimodal navigation, despite extensive task-specific training. In this paper, we introduce RAO-Nav, short for Reasoning All-in-One OLM, a deployment pipeline for zero-shot SAVN. By leveraging the rich implicit audio-visual knowledge encoded in OLMs, the embodied agent is enabled to hear, see, reason, and act in the environment. To further elicit the built-in thinking ability of OLMs, we propose a test-time Latent Navigation Reasoning (LNR) module that can be seamlessly integrated into the decoding space. LNR encourages the model to retrieve more target-relevant observations and make effective navigation decisions. Through comprehensive experiments, we show that our framework surpasses existing state-of-the-art baselines on public SAVN benchmarks without using any training data. Moreover, we introduce a new Global Navigation Instruction setting to further evaluate the ability of OLMs to serve as embodied navigation agents. Results under this setting posit that applying OLMs to real-world natural-language-guided audio-visual navigation remains challenging.


RaVF: Learning Radar Velocity Fields via Spatial-Doppler Guidance

SHENGPENG WANG ⋅ Lingzhen Li ⋅ Chunshen Li ⋅ Fei Xiao ⋅ Wei Wang

Scene-level Velocity field estimation from 4D millimeter-wave radar is critical for robust perception but remains challenged by the inherent sparsity and noise of point-based paradigms. Furthermore, Doppler ambiguity induced by the Nyquist sampling limit fundamentally disrupts the continuity of velocity field learning. To address these challenges, we propose RaVF, a physically grounded framework that directly learns dense velocity fields from high-fidelity spatial-Doppler spectra. RaVF is built upon three key designs tailored to radar-specific signal distortions. (1) To mitigate range-dependent propagation attenuation and nonlinear angular spatial quantization, we devise a Physically-aware Encoder that calibrates spectral feature extraction within the native radar beamforming. (2) Observing that ego-induced radial motion explains static Doppler responses, we factor out static motion and guide dynamic flow estimation for physically consistent velocity field learning. (3) To enhance kinematic plausibility, we introduce an ego-conditioned bidirectional pyramidal flow module that transfers backward flow into the forward frame and enforces anti-symmetric forward-backward consistency. Extensive experiments demonstrate that RaVF significantly outperforms state-of-the-art baselines on velocity fields. Our code will be available.


Reading the Finetuning Prior: Verbatim Content Recovery via Contrastive Decoding Diffing

Michał Brzozowski ⋅ Zuzanna Dubanowska ⋅ Enrico Cassano ⋅ Neo Christopher Chung

Narrowly finetuned language models memorize implanted content verbatim, but auditing what a deployed model has been taught, without access to its weights or training data, remains an open challenge. Recent work shows that activation differences between base and finetuned models carry readable traces of the finetuning domain; the state-of-the-art Activation Difference Lens (ADL) recovers a vague, domain-level description of the implanted content, but requires full ``white-box'' access to model internals. We introduce Contrastive Decoding Diffing (CDD), a model diffing method that operates on output-level logit distributions only, with no weight access, no layer selection, and no per-model tuning, yet recovers implanted facts. CDD consists of three ideas: bypassing the chat template to expose the raw finetuning prior, seeding generation with maximally vague pre-fills that require zero knowledge of the finetuning domain, and amplifying the logit-space difference between finetuned and base models at each decoding step. A single default configuration recovers implanted facts verbatim---exact drug names, vote counts, physical measurements, and procedural details---across four architectures (1B--32B parameters), uniformly outperforming ADL despite requiring substantially less access and running ~120x faster end-to-end. Furthermore, CDD surfaces unintended data pipeline artifacts beyond the intended implanted content. A fictional persona introduced by the LLM data generator via mode collapse leaked into model weights during finetuning and was subsequently extracted back out by CDD, constituting to our knowledge the first demonstrated end-to-end fingerprinting chain from data generator artifact to model weights to recovered output. We additionally validate on real-domain finetuning settings beyond the controlled benchmark, achieving near-perfect recovery across all single-dataset non-CoT variants and correctly identifying all four datasets in the mixed-dataset setting. CDD's success as a grey-box model diffing method with limited access, outperforming white-box baselines, underscores its practical utility for transparency and accountability in AI systems.


Real2Sim in HOI: Toward Physically Plausible HOI Reconstruction from Monocular Videos

Yubo ZHAO ⋅ Yujin Chai ⋅ Yunao Dong ⋅ Chengfeng Zhao ⋅ Zijiao Zeng ⋅ Yuan Liu ⋅ Chi-Keung Tang

Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can reconstruct temporally coherent human and object trajectories, but these trajectories often remain visual artifacts while failing to preserve stable contact, functional manipulation, or physical plausibility when used as reference motions for humanoid-object simulation. This reveals a fundamental interaction gap: HOI reconstruction should not stop at tracking a human and an object, but should recover the relation that makes their motion a coherent interaction. We introduce HA-HOI, a framework for reconstructing physically plausible 4D HOI animation from in-the-wild monocular videos. Instead of treating the human and object as independent entities in an ambiguous monocular 3D space, we propose a \emph{human-first, object-follow} formulation. The human motion is recovered as the interaction anchor, and the object is reconstructed, aligned, and refined relative to the human action. The resulting kinematic trajectory is then projected into a physics-based humanoid-object simulation, where it acts as a teacher trajectory for stable physical rollout. Across benchmark and in-the-wild videos, HA-HOI improves human-object alignment, contact consistency, temporal stability, and simulation readiness over prior monocular HOI reconstruction methods. By moving beyond visually plausible trajectory recovery toward physically grounded interaction animation, our work takes a step toward turning general monocular HOI videos into scalable demonstrations for humanoid-object behavior.


RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

Chengzhi Shen ⋅ Weixiang Shen ⋅ Tobias Susetzky ⋅ Chen Chen ⋅ Jun Li ⋅ Yuyuan Liu ⋅ zhang xuepeng ⋅ Zhenyu Gong ⋅ Daniel Rueckert ⋅ Jiazhen Pan

Intensive care units (ICU) generate dense, evolving streams of clinical information, where physicians must repeatedly reassess patient states under time pressure, underscoring a clear need for reliable AI decision support. Existing ICU benchmarks typically treat historical clinician actions as ground truth. However, these actions are made under incomplete information and limited temporal context of the underlying patient state, and may therefore be suboptimal, making it difficult to assess the true reasoning capabilities of AI systems. We introduce RealICU, a hindsight-annotated benchmark for evaluating large language models (LLMs) under realistic ICU conditions, where labels are created after senior physicians review the full patient trajectory. We formulate four physician-motivated tasks: assess \emph{Patient Status}, \emph{Acute Problems}, \emph{Recommended Actions}, and \emph{Red Flag} actions that risk unsafe outcomes. We partition each trajectory with 30-min windows and release two datasets: RealICU-Gold with 930-window annotations from 94 MIMIC-IV patients, and RealICU-Scale with 11,862 windows extended by \emph{Oracle}, a physician-validated LLM hindsight labeler. Existing LLMs including memory-augmented ones performed poorly on RealICU, exposing two failure modes: a recall–safety tradeoff for clinical recommendations, and an anchoring bias to early interpretations of the patient. We further introduce ICU-Evo to study structured-memory agents that improves long-horizon reasoning but does not eliminate safety failures. Together, RealICU provides a clinically grounded testbed for measuring and improving AI sequential decision-support in high-stakes care.


Reasoning-Aware Relational Representation Learning for Open-Vocabulary Scene Graph Generation

Jiawei Xiao ⋅ Boya Wang ⋅ Longtian Qiu ⋅ Shan Ning ⋅ Xuming He

Open-Vocabulary Scene Graph Generation (OV-SGG) requires models to recognize visual relationships beyond the training vocabulary, yet existing methods often rely on dataset-specific object--predicate co-occurrence patterns. Under long-tailed distributions and noisy object localization, such reliance leads to biased, entangled relation representations, limiting both rare-relation recognition and generalization to unseen predicates.To address these challenges, we propose RaRe, a reasoning-aware relational representation learning framework that transforms MLLM-generated chain-of-thought (CoT) rationales into structured relation embeddings. Rather than using CoT only as an intermediate reasoning trace for final prediction, RaRe directly optimizes the rationale-derived representation space to improve inter-relation discriminability. Specifically, RaRe first aligns rationale embeddings with textualized relation descriptions through self-supervised contrastive learning, and then refines rationale generation with a specificity-oriented reinforcement learning objective that encourages semantically distinctive relational evidence.Experiments on Visual Genome demonstrate that RaRe improves category-balanced recognition across both base and novel predicates, with particularly strong gains under the SGDet setting where detection noise is prevalent.


Reasoning over Coupled Receptive Fields: Eliminating Subgraph Redundancy at Scale

Jing Yang ⋅ Bo Wen ⋅ Yuan Gao ⋅ XiaowenJiang ⋅ Zhihao Zhang ⋅ Shikun Yan

Conditional message passing (CMP) currently represents a leading paradigm for knowledge graph reasoning. Since vanilla CMP performs full-graph propagation per query, existing methods adopt subgraph-wise sampling to improve computational efficiency. However, these extracted subgraphs still exhibit high structural overlap. Consequently, the same local structures are repeatedly accessed and processed, leading to substantial redundancy that remains a major bottleneck for scaling reasoning to large knowledge graphs. To address this issue, we propose reasoning over \textbf{Co}upled \textbf{R}eceptive \textbf{F}ields (CoRF) to eliminate subgraph redundancy at scale. Rather than sampling a separate subgraph for each query, CoRF constructs a set of receptive supports that allow multiple queries to share receptive fields, thereby reducing subgraph redundancy. Meanwhile, we introduce boundary padding and residual compensation mechanisms to ensure topological integrity and query coverage. We further incorporate cost-aware load balancing to improve distributed execution efficiency. Comprehensive experiments on state-of-the-art CMP models demonstrate that CoRF achieves up to a 5.9$\times$ speedup and a 10.0$\times$ reduction in memory usage, without compromising reasoning accuracy.


Reasoning via Test-Time Instance-Level Policy Gradient in Latent Space

Hengli Li ⋅ Chenxi Li ⋅ Tong Wu ⋅ Xuekai Zhu ⋅ Yuxuan Wang ⋅ Zhaoxin Yu ⋅ Eric Hanchen Jiang ⋅ Song-Chun Zhu ⋅ Zixia Jia ⋅ Ying Nian Wu ⋅ Zilong Zheng

Large Language Models (LLMs) typically reason through explicit, step-by-step natural-language traces. Humans, however, also rely on non-linguistic, unconscious processes, such as the inspirations that emerge during the incubation period. In this work, we introduce LatentSeek, a novel framework designed to enhance the reasoning capabilities of LLMs through Test-Time Instance-level Policy Gradient within the model’s latent space—thus complementing explicit natural-language steps. LatentSeek employs policy gradient optimization to iteratively refine latent representations, guided solely by a self-generated reward signal. This allows the model to adapt its reasoning trajectory dynamically on a per-instance basis. Empirical evaluations across diverse benchmarks, GSM8K, MATH-500, and AIME2024 as well as multiple LLM families (e.g., LLaMA, Qwen) demonstrate that LatentSeek outperforms established baselines, including Chain-of-Thought (CoT), Best-of-N (BoN) and training-based methods. Further analysis indicates that LatentSeek is computationally efficient, typically converging within a few optimization iterations for average-level problems. Moreover, the model's performance improves as the number of latent update iterations increases, highlighting the benefits of exploring within the latent space. These findings highlight LatentSeek as a lightweight and effective paradigm for improving the reasoning capabilities of LLMs without changing their parameters.


Reasoning Warm-up: Scaling Label-free RL via Verifiable Surrogate Rewards

Jun Nie ⋅ Bo Han ⋅ Jiaqi Fan ⋅ Yonggang Zhang ⋅ Xinmei Tian ⋅ Dahai Yu ⋅ Michael Ng

Improving the reasoning ability of large language models (LLMs) without ground-truth supervision remains a central challenge in reinforcement learning (RL). Recent label-free RL methods typically rely on internal pseudo-rewards such as self-consistency or majority voting. However, for complex multi-step reasoning, these signals can be systematically misleading: models may repeatedly produce mutually consistent but incorrect trajectories, causing optimization to favor frequency over correctness. In this work, we propose reasoning warm-up reinforcement learning, a label-free RL framework motivated by an empirical phenomenon we call Reasoning Warm-up. We observe that when a model successfully completes a deterministic auxiliary task with an exact verifier, its success probability on a subsequent complex reasoning task increases substantially. This coupling suggests that auxiliary-task verification can serve as a useful surrogate signal for selecting and optimizing reasoning trajectories even when the main task itself is not directly verifiable. Based on this observation, we incorporate verifiable auxiliary tasks into both generation and optimization. During generation, the auxiliary task provides verifiable warm-up prefix; during training, its verification outcome is used to score candidate trajectories for policy optimization. Experiments on multiple benchmarks demonstrate the effectiveness of our method.


Rebalancing Reference Frame Dominance to Improve Motion in Image-to-Video Models

Wooseok Jeon ⋅ Seungho Park ⋅ Seunghyun Shin ⋅ Sangeyl Lee ⋅ Hyeonho Jeong ⋅ Hae-Gon Jeon

Image-to-video models often generate videos that remain overly static, compared to text-to-video models. While prior approaches mitigate this issue by weakening or modifying the image-conditioning signal, they often require additional training or sacrifice fidelity to the reference image. In this work, we identify reference-frame dominance as a key mechanism behind motion suppression. We observe that non-reference frames in I2V models allocate excessive self-attention to reference-frame key tokens, causing reference information to be over-propagated across time and suppressing inter-frame dynamics. Based on this finding, we propose DyMoS (Dynamic Motion Slider), a training-free and model-agnostic method that rebalances the attention pathway from generated frames to the reference frame during initial denoising steps. DyMoS leaves both the input image and model weights unchanged and introduces a single scalar parameter for continuous control over motion strength. Experiments across multiple state-of-the-art I2V backbones demonstrate that DyMoS consistently improves motion dynamics while maintaining visual quality and fidelity to the reference image.


Re:Cognize - A Framework for Open-Set Sequential Character Re-Identification

Aaditya Baranwal ⋅ Madhav Kataria ⋅ Yogesh Rawat ⋅ Shruti Vyas

Character re-identification (Re-ID) is the foundation of every downstream task on long-form visual narratives such as manga and comics: without consistent character identity across pages, archives spanning hundreds of millions of pages stay opaque to search and reasoning. Existing benchmarks, however, evaluate Re-ID under a closed-set assumption, retrieval against a fixed gallery of known identities. That assumption collapses on a fresh volume, where new characters appear over time, the gallery is built online, and predictions must stay consistent across long temporal gaps. We introduce $\textbf{Re:Cognize}$, a framework that exposes this regime through four evaluation protocols spanning closed-set retrieval, few-shot retrieval with a fixed gallery, unsupervised online clustering, and pre-seeded online gallery growth, instantiated on large comic datasets (PopCharacters, Manga109, Re:Verse) across five Re-ID backbones from person, manga-native, and multimodal pre-training. The decomposition surfaces a striking property of the regime: across every backbone, a single seed image per identity recovers nearly the entire closed-set retrieval ceiling, while online accumulation overshoots that ceiling within a handful of seeds. Identity $\textit{maintenance}$, not gallery construction, is the practical deployment bottleneck, and identity $\textit{emergence}$ is the harder, separable sub-problem that closed-set evaluation has never made visible. To complete the suite, we further introduce $\textbf{MeCha}$, a memory-augmented baseline that confirms the protocols are sensitive enough to discriminate sequential-context-aware models from purely embedding-based ones.


Recovering Evolving User Preference State via Adaptive Interaction-aware Representation Correction

Parthiv Chatterjee ⋅ Dhiraj Golhar ⋅ Ummesalma Diwan ⋅ Sourish Dasgupta ⋅ Manjunath Joshi ⋅ Tanmoy Chakraborty

Personalization systems encode long-horizon, evolving user trajectories into compressed preference states that downstream task heads consume for prediction or generation. This compression is difficult because preference trajectories contain stable long-term interests, transient short-term shifts, and episodic bursty interests. We argue that a practical route is not to design another task-specific encoder or incur costly host-encoder finetuning, but to correct the under-expressed compressed state before the downstream task head consumes it. We propose $\texttt{REPAIR}$ as a selective *pre-head state-repair method* for frozen personalization encoders. It plugs into frozen hosts, uses event representations from the forward pass, and avoids a second pass over the raw history. $\texttt{REPAIR}$ compares the compressed state with per-timestep event representations, estimates state-relative corrective evidence, resolves it through long-term, short-term, and episodic temporal regimes, and retains only the most corrective signals to form a compact correction. We evaluate $\texttt{REPAIR}$ across recommendation and generation personalization settings using MovieLens, MIND, PENS, and Amazon Reviews 2023. Across recommendation tasks, $\texttt{REPAIR}$ improves the evaluated frozen hosts. On MovieLens, it improves the highest-scoring frozen host in our instantiated baseline suite by +1.27/+1.54 MRR/nDCG@5; on MIND, it improves the highest-scoring frozen host by +4.41 MRR; and on Amazon product recommendation, it improves two SOTA baselines by +15.33 and +17.91 MRR. It also outperforms budget-matched post-head correctors, and the full long-, short-, and episodic repair remains the strongest temporal variant. On PENS personalized headline generation, $\texttt{REPAIR}$ improves subjectivity-sensitive generation most when the decoder is preference-coupled, with average PerSEval gains of +14.62% and +22.19% for two representative preference-coupled decoders. Relative to frozen predictive encoders, deployment overhead is 31--34% memory and 38--44% latency.


Reducing Credit Assignment Variance via Counterfactual Reasoning Paths

fei ding ⋅ Yongkang Zhang ⋅ youwei wang ⋅ Zijian Zeng

Reinforcement learning for multi-step reasoning with large language models (LLMs) typically relies on sparse terminal rewards, which creates a poorly conditioned credit-assignment problem: the final feedback is propagated uniformly across all intermediate decisions. This leads to high gradient variance, unstable training, and many ineffective updates, ultimately limiting sustained model improvement. We propose a counterfactual-comparison framework for credit assignment. For each input, the framework samples multiple reasoning trajectories and treats their differences as implicit approximations to alternative decisions. This yields an implicit process-level advantage estimator that converts sparse terminal rewards into step-sensitive learning signals. Building on this framework, we introduce Implicit Behavior Policy Optimization (IBPO), which substantially improves training stability and the performance ceiling on mathematical and code-reasoning benchmarks. Our results point to a promising direction for unlocking the reasoning potential of LLMs.


Reference-Guided Training: Adaptive Gradient Scaling via Per-Sample Loss Comparisons

Israel Rodrigues Soares ⋅ Pedro Benedetti ⋅ Thanda Shwe ⋅ Israel Mendonca ⋅ Masayoshi Aritsugi

Empirical Risk Minimization treats training samples uniformly, which can degrade performance in noisy or heterogeneous settings. We introduce a reference-guided objective that uses per-sample error comparisons with a fixed reference model to reweight gradient contributions without relying on prediction imitation, while preserving gradient direction and ensuring bounded scaling. Theoretical analysis shows that the formulation induces adaptive weighting with controlled amplification. Experiments on image classification with synthetic label noise and on time-series forecasting demonstrate broad improvements across architectures, particularly in moderate-to-high noise and heterogeneous regimes. Results further indicate that the method is most effective when the error-ranking alignment between the baseline model and the reference model is low, suggesting that gains arise primarily from the diversity of the reference error ranking relative to the learner, rather than the reference's absolute performance.


ReflectDrive-2: Reinforcement-Learning-Aligned Self-Editing for Discrete Diffusion Driving

Huimin Wang ⋅ Yue Wang ⋅ Bihao Cui ⋅ Pengxiang Li ⋅ Ben Lu ⋅ Mingqian Wang ⋅ Tong Wang ⋅ Chuan Tang ⋅ Teng Zhang ⋅ Kun Zhan

We introduce ReflectDrive-2, a masked discrete diffusion planner with a separate action expert for autonomous driving that represents plans as discrete trajectory tokens and generates them through parallel masked decoding. This discrete token space enables in-place trajectory revision: AutoEdit rewrites selected tokens using the same model, without requiring an auxiliary refinement network. To train this capability, we use a two-stage procedure. First, we construct structure-aware perturbations of expert trajectories along longitudinal progress and lateral heading directions and supervise the model to recover the original expert trajectory. We then fine-tune the full decision--draft--reflect rollout with reinforcement learning (RL), assigning terminal driving reward to the final post-edit trajectory and propagating policy-gradient credit through full-rollout transitions. Full-rollout RL proves crucial for coupling drafting and editing: under supervised training alone, inference-time AutoEdit improves PDMS by at most $0.3$, whereas RL increases its gain to $1.9$. We also co-design an efficient reflective decoding stack for the decision--draft--reflect pipeline, combining shared-prefix KV reuse, Alternating Step Decode, and fused on-device unmasking. On NAVSIM, ReflectDrive-2 achieves $91.0$ PDMS with camera-only input and $94.8$ PDMS in a best-of-6 oracle setting, while running at $31.8$ ms average latency on NVIDIA Thor.


ReflectMT: Internalizing Reflection for Efficient and High-Quality Machine Translation

Kunquan Li ⋅ Yingxue Zhang ⋅ Zhibin Lan ⋅ Fandong Meng ⋅ Jinsong Su

Recent years have witnessed growing interest in applying Large Reasoning Models (LRMs) to Machine Translation (MT). While most approaches adopt a "pre-thinking" paradigm and benefit from explicit reasoning trajectories, they suffer from substantial inference cost and latency. To address these limitations, we propose ReflectMT, a two-stage reflection internalization framework for machine translation that employs a "post-thinking" paradigm. Our approach develops the model's "translate–reflect–refine" capability through reinforcement learning. In the first stage, we cultivate the model's capacity for high-quality reflection and refinement, thereby enhancing its semantic comprehension and task-specific knowledge. In the second stage, we train the model to internalize the knowledge acquired during reflection. As a result, during inference, ReflectMT operates in a direct translation mode, producing high-quality translations on the first attempt without any explicit reasoning steps. Experimental results on benchmarks such as WMT24 demonstrate that our model’s first-pass translations during inference outperform multi-step reasoning LRMs (e.g., DeepSeek-R1) in both automatic metrics and GPT-based evaluation, achieving a 2.16-point improvement in GPT-based translation quality evaluation while reducing token consumption by 94.33%.

Large multimodal language models (LLMs) have emerged as powerful tools for guiding evolutionary search toward interpretable programmatic policies. However, existing frameworks rely on a monolithic model call to simultaneously interpret visual behavioral evidence and synthesize corrective code. This diagnosis-repair entanglement creates an opaque feedback loop, obscuring the rationale behind mutations and preventing the retention of algorithmic insights across independent runs. To achieve auditable and efficient policy search, we argue that visual diagnosis must be structurally decoupled from code generation. We present REFLEX, a train-free evolutionary framework that operationalizes this decoupling. In REFLEX, a vision-enabled Critic first distills task-specific behavioral evidence into structured, auditable diagnoses. Subsequently, a text-optimized Actor synthesizes child policies using these diagnoses alongside a persistent, self-evolving Skill Memory of reusable code snippets. This architecture not only provides transparent mutation traces but also enables cross-run programmatic knowledge transfer. Extensive evaluations across control benchmarks (Lunar Lander, Acrobot, Pendulum) and a 36-dimensional antenna array synthesis task demonstrate exceptional sample efficiency. Notably, REFLEX solves Acrobot and Pendulum in under 10 LLM calls and reaches a Normalized Weighted Score of 1.013 on Lunar Lander in just 20 calls, achieving highly competitive final performance while significantly accelerating the early-stage discovery of transparent policies.


Reinforcement World Model Learning for LLM-based Agents

Xiao Yu ⋅ Baolin Peng ⋅ Ruize Xu ⋅ yelong shen ⋅ Pengcheng He ⋅ Suman Nath ⋅ Nikhil Singh ⋅ Jianfeng Gao ⋅ Zhou Yu

Large language models (LLMs) have achieved strong performance in language-centric tasks. However, in agentic settings, LLMs often struggle to anticipate action consequences and adapt to environment dynamics, highlighting the need for world-modeling capabilities in LLM-based agents. We propose Reinforcement World Model Learning (RWML), a self-supervised method that learns action-conditioned world models for LLM-based agents on textual states using sim-to-real gap rewards. Our method aligns simulated next states produced by the model with realized next states observed from the environment, encouraging consistency between internal world simulations and actual environment dynamics in a pre-trained embedding space. Unlike next-state token prediction, which prioritizes token-level fidelity (i.e., reproducing exact wording) over semantic equivalence and can lead to model collapse, our method provides a more robust training signal and is empirically less susceptible to reward hacking than LLM-as-a-judge. We evaluate our method on ALFWorld and Bench and observe significant gains over the base model, despite being entirely self-supervised. When combined with task-success rewards, our method outperforms direct task-success reward RL by 6.9 and 5.7 points on ALFWorld and Bench respectively, while matching the performance of expert-data training.


Reinforcing VLAs in Task-Agnostic World Models

Yucen Wang ⋅ Rui Yu ⋅ Fengming Zhang ⋅ Junjie Lu ⋅ Kaixin Wang ⋅ Li Zhao

Post-training Vision-Language-Action (VLA) models via reinforcement learning (RL) in learned world models has emerged as an effective strategy to adapt to new tasks without costly real-world interactions. However, while using imagined trajectories reduces the sample complexity of policy training, existing methods still heavily rely on task-specific data to fine-tune both the world and reward models, fundamentally limiting their scalability to unseen tasks. To overcome this, we argue that world and reward models should capture transferable physical priors that enable zero-shot inference. We propose RAW-Dream (Reinforcing VLAs in task-Agnostic World Dreams), a new paradigm that completely disentangles world model learning from downstream task dependencies. RAW-Dream utilizes a world model pre-trained on diverse task-free behaviors for predicting future rollouts, and an off-the-shelf Vision-Language Model (VLM) for reward generation. Because both components are task-agnostic, VLAs can be readily finetuned for any new task entirely within this zero-shot imagination. Furthermore, to mitigate world model hallucinations, we introduce a dual-noise verification mechanism to filter out unreliable rollouts. Extensive experiments across simulation and real-world settings demonstrate consistent performance gains, proving that generalized physical priors can effectively substitute for costly task-dependent data, offering a highly scalable roadmap for VLA adaptation.


Relation-Aware Graph Foundation Model

Jianxiang Yu ⋅ Jiapeng Zhu ⋅ Yibo Zhao ⋅ HAO QIAN ⋅ Ziqi Liu ⋅ Zhiqiang Zhang ⋅ Xiang Li

In recent years, large language models (LLMs) have demonstrated remarkable capability to generalize across diverse natural language processing tasks, inspiring the development of graph foundation models (GFMs) for large-scale pre-training. However, unlike language models with explicit token units, graphs lack a well-defined unit for generalization, making it challenging to design effective pre-training strategies. In this work, we propose REEF, a novel GFM framework that leverages relation tokens as the fundamental units. We construct a vocabulary of relation tokens to encode relational information within graphs. To accommodate diverse relations, we introduce two hypernetworks that adaptively generate the parameters of aggregators and classifiers in graph neural networks based on relation tokens. In addition, we design another hypernetwork to construct dataset-specific projectors and incorporate a dataset-level feature bias into the initial node representations, enhancing flexibility across different datasets with the same relation. Extensive experiments demonstrate that REEF consistently outperforms existing methods in both pre-training and transfer learning, highlighting its potential as a general-purpose graph foundation model. Our code is publicly available at https://anonymous.4open.science/r/REEF-16CF/.

A multi-tool plan is correct only when its call set matches the gold API set exactly. Standard tool-use pipelines optimize first-hop relevance, but downstream verifiers require the necessary API set, creating a relevance--necessity gap. This study formalizes the gap: under correlated API indicators, symmetric decomposable selectors based only on per-API marginals can be dominated by a Bayes-optimal set selector, while whole-set exact-match supervision is a proper scoring rule for the optimal selector restricted to a candidate menu. We propose an offline refinement framework for frozen upstream traces. Under a leakage-free ToolBench G2/G3 protocol, TC-MASS improves verifier-consistent API-set recovery without additional LLM generation after the upstream trace exists. Ablations and cross-distribution probes support the central claim: necessity-aware set-level supervision, not relevance ranking alone, is the right objective for fixed-pool API-set recovery.


Representation Forcing for Bottleneck-Free Unified Multimodal Models

Yuqing Wang ⋅ Zhijie Lin ⋅ Ceyuan Yang ⋅ Yang Zhao ⋅ Fei Xiao ⋅ Hao He ⋅ Qi Zhao ⋅ Zihan Ding ⋅ Fu-Yun Wang ⋅ Shuai Wang ⋅ Youliang Zhang ⋅ Haoqi Fan ⋅ Xihui Liu

Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation, imposing a structural bottleneck. Naively removing it introduces a quality gap, as the model must learn both high-level structure and low-level details from raw pixels. In this paper, we propose Representation Forcing (RF), a technique that closes this gap by making representation prediction a native capability of the decoder. Concretely, RF forces the decoder to autoregressively predict visual representations as intermediate tokens before pixels; these tokens then stay in context to guide pixel diffusion within the same backbone. By turning representations from perception outputs into generation targets, RF eliminates the need for any external generative latent space. We find that RF benefits both understanding and generation. On image generation, our pixel-space model with RF matches state-of-the-art VAE-based unified models. On image understanding, pixel-space RF generally outperforms its VAE-based variant. Together, these results offer an effective step toward end-to-end, bottleneck-free UMMs.


ReSAM: Representation-Level Safety Margin Alignment for Vision–Language Models

jiachen ma ⋅ Jiawen Zhang ⋅ Bo Zou ⋅ Xiangtian Li ⋅ Chaochao Lu ⋅ Chao Yang

We study the problem of Pseudo-Benign Failures in Vision--Language Models (VLMs): multimodal inputs that appear harmless but elicit dangerous or policy-violating responses. Our analysis shows that these failures arise from a representational misalignment: the model's internal embedding space exhibits a distributional gap between pseudo-benign inputs and unsafe inputs located in the refusal region, causing failures outside the safety margins of models. We introduce Representation-Level Safety Margin Alignment method (ReSAM), a lightweight representation-space alignment method that: (i) computes direction vectors separating refusal and non-refusal representations, (ii) quantifies refusal behavior by projecting embeddings of inputs onto this direction, and (iii) optimizes a safety-margin loss that pushes unsafe and pseudo-benign queries above a learned margin while pulling safe queries below it. ReSAM introduces a new paradigm for multimodal safety alignment: it requires no manual annotations, instead deriving supervisory signals directly from its own representation space. Despite this minimal supervision, ReSAM achieves a 68% improvement in safety over strong baselines, and remarkably, we further observe that incorporating only a handful of pseudo-benign queries (as few as five) during training suffices to raise safety to 94.6%. Beyond these empirical gains, our analysis reveals that safety gradients concentrate in a low-rank subspace, suggesting that multimodal safety is governed by an intrinsic structure that can be systematically identified and controlled.


Residual Calibration via Local Feature-Space Refinement

Xihao Wang ⋅ Chengyao Yu ⋅ Ruixing Ming

Deep neural networks often struggle to provide reliable confidence estimates, making post-hoc calibration essential for reliable decision-making. Existing post-hoc methods largely rely on score-based transformations, which improve marginal calibration but often fail to capture spatially varying miscalibration in local regions of the learned feature space. To address this gap, we propose Local Adaptive Probability Estimation (LAPE), a post-hoc residual calibration framework. Specifically, LAPE treats the global calibrator output as a stable anchor and refines the individual prediction by leveraging non-parametric local evidence from neighboring calibration samples. As a plug-and-play refinement module, LAPE effectively utilizes the information about feature-space geometry and can reuse the calibration data that are used to construct the global calibrator. Furthermore, we extend LAPE to the setting with multi-view inputs through Cascade Feature Augmentation Fusion (CFAF). Extensive experiments demonstrate that the proposed methods can improve the calibration quality of various existing baselines and remain efficient and robust across different settings.


Resilient Latent Readouts for Long-Context Question Answering

Jingyi Liao ⋅ Wenhao Sun ⋅ YITING LI ⋅ Zhao Jin ⋅ Khin Mi Mi Aung ⋅ Xun Xu ⋅ Zhuoyi Lin ⋅ Rong-Cheng Tu ⋅ Dacheng Tao

Long-context question answering often asks a model to answer from extended documents, multi-turn dialogues, or persistent histories. Directly prefilling the full history preserves all observed tokens, but it is expensive and can expose generation to substantial irrelevant or spuriously salient context, which may hinder evidence localization and answer synthesis. This motivates compact memory interfaces that provide a short, query-conditioned readout of the history. Existing systems predominantly instantiate this readout as text, where the reader sees only what was selected and any evidence left out is unrecoverable, leaving the interface fragile to selection errors. We argue that memory readout has two coupled axes that should be optimized separately: evidential coverage, whether all right support is selected, and inferential resilience, how much answerability remains under imperfect selection. To jointly address these two axes, we propose LIRA, a latent-memory reader that exposes full-prefix contextualized states rather than re-encoded text. It amortizes one full-prefix pass into a reusable latent store; for each query, it localizes candidate positions with calibrated endogenous attention, repairs incomplete support with coverage-aware planning, and packs selected states into a position-consistent readout for autoregressive generation. Across three open-weight backbones on LoCoMo and Loong, LIRA is the strongest non-oracle compact-memory method and surpasses Full Context on 8 of 9 LoCoMo metrics. When localization misses all supporting evidence, LIRA improves F1 by up to 13.3 points over a text-native control with identical positions, suggesting that full-prefix states preserve answer-relevant influence discarded by text-native re-encoding.

With respect to improving the reasoning accuracy of LLMs, the representative reinforcement learning (RL) method — Group Relative Policy Optimization (GRPO) — has achieved critical success, yet it still suffers from the issue of insignificant reward signals. This paper introduces ReST-RL, a unified Reinforced Self-Training (ReST) LLM RL paradigm that reconnects policy optimization and value-guided search to improve LLM reasoning ability. Firstly, ReST-GRPO adopts an optimized ReST-style algorithm to reshape the policy-induced trajectory distribution by increasing the reward variance of GRPO sampling and exposing the policy to more informative partial states, thereby improving training efficiency and effectiveness. Then, we further introduce a decoding optimization method, VM-MCTS, which trains a Value Model (VM) from self-collected Monte-Carlo Tree Search (MCTS) targets and deploys it through an adapted MCTS algorithm to provide precise process signals and verification scores, further enhancing reasoning accuracy. These two stages are internally dependent — ReST-GRPO yields higher-quality trajectories for value learning with VM-MCTS, which in turn enables more effective inference-time search. We validate our RL paradigm on multiple coding benchmarks (e.g., APPS, BigCodeBench, and HumanEval), where it significantly outperforms other reinforcement training baselines (naive GRPO and ReST-DPO), as well as decoding and verification baselines (e.g., PRM-BoN and ORM-MCTS), indicating its power to strengthen LLM reasoning capability. Moreover, we further evaluate ReST-RL on out-of-domain math and science reasoning tasks, where it achieves superior performance and favorable end-to-end efficiency trade-offs, verifying its effective, robust, and highly generalizable nature.


Retain-Neutral Surrogates for Min-Max Unlearning

Junhao Cai ⋅ Dohun Kim ⋅ Dowon Kim ⋅ Sung Il Choi ⋅ Chengjun Jin ⋅ Juhyun Park ⋅ Changhee Joo

Machine unlearning seeks to remove the influence of designated training data while preserving performance on the remaining data. Approximate unlearning can be viewed as a local editing problem; in min-max unlearning, the key local object is the surrogate point at which the retain objective is evaluated. When forget and retain gradients are strongly aligned, an unconstrained forget-maximizing perturbation can move to a surrogate point that increases retain loss. We propose Retain-Orthogonal Surrogate Unlearning (ROSU), which constrains the inner surrogate construction by maximizing first-order forget gain subject to zero first-order retain change under a fixed perturbation budget. This yields a closed-form retain-orthogonal perturbation, a lightweight transported outer update, and amplification along the retain-neutral direction. Our analysis establishes (i) a curvature-controlled second-order bound on retain damage, (ii) a positive-alignment regime in which ROSU strictly reduces surrogate retain loss relative to standard min-max perturbations, and (iii) near-equivalence when the two gradients are nearly orthogonal. Across vision and language benchmarks (CIFAR-10/100, Tiny-ImageNet, TOFU, WMDP), the empirical pattern follows this geometry: ROSU gives its clearest gains in high-coupling regimes while remaining competitive elsewhere.

Credit assignment remains a central challenge in cooperative multi-agent reinforcement learning (MARL), especially under partial observability, where individual policy updates may not accurately reflect each agent’s actual contribution to team outcomes. While policy-based methods such as MAPPO and IPPO provide strong optimization frameworks, their updates are typically derived from global rewards and observational value surrogates, without explicitly defining agent-specific credit. Misaligned credit signals can mislead individual policy improvement, resulting in inefficient coordination and weaker team performance. We address this challenge by formulating agent-level credit as an interventional reward response, using Proximal Causal Inference (PCI) to identify credit from observable proxies via an outcome bridge function. Building on this identification strategy, we design a practical credit-aligned update signal and integrate it into policy gradient methods. Empirical evaluations on diagnostic and benchmark tasks demonstrate that the proposed credit signal improves policy learning under partial observability, highlighting proximal identification as a promising foundation for designing credit-aware policy updates in cooperative MARL. To the best of our knowledge, this work presents the first PCI-based solution for online multi-agent cooperation.


Rethinking Cross-Layer Information Routing in Diffusion Transformer

Chao Xu ⋅ Maohua Li ⋅ Qirui Li ⋅ Yixuan Xu ⋅ Yanke Zhou ⋅ Yunhe LI ⋅ Cuifeng Shen ⋅ Hanlin Tang ⋅ Kan Liu ⋅ Tao Lan ⋅ Lin Qu ⋅ Shao-Qun Zhang

Diffusion Transformers (DiTs) have become the de facto backbone of modern visual generation, and nearly every major axis of their design --- tokenization, attention, conditioning, objectives, and latent autoencoders --- has been thoroughly revisited. The residual stream that governs how information accumulates across layers, however, has been directly inherited from the original Transformer. In this paper, we present a systematic analysis of cross-layer information flow in DiT jointly along depth and denoising timestep, and identify three concrete symptoms of traditional residual addition, i.e., monotonic forward magnitude inflation, sharp backward gradient decay, and pronounced block-wise redundancy. Motivated by this diagnosis, we propose Diffusion-Adaptive Routing (\textsc{DAR}), a drop-in residual replacement that performs \emph{learnable, timestep-adaptive, and non-incremental} aggregation over the history of sublayer outputs. Moreover, the proposed DAR is compatible with many modern Transformer enhancement methods, such as REPA. On ImageNet $256\times256$, \textsc{DAR} improves SiT-XL/2 by $2.11$ FID ($7.56$ vs.\ $9.67$) and matches the baseline's converged quality in $8.75\times$ fewer training iterations. Stacked on top of REPA, it yields a $2\times$ training acceleration in the early stage, revealing cross-layer information routing as a hitherto-overlooked axis of progress in diffusion modeling, one that operates orthogonally to existing representation-alignment objectives. Beyond the pretrain tasks, \textsc{DAR} can be further applied in large-scale T2I models fine-tuning stage and preserves high-frequency details during Distribution Matching Distillation.


Rethinking Diffusion Decoding via Structural Commitment

Lipeng Wan ⋅ Anbang Wang ⋅ Zixuan Yang ⋅ Kun Xu ⋅ Xiuxiu Bai ⋅ Xuguang Lan

Standard decoding in diffusion language models (DLMs) is controlled through parallel, token-wise unmasking decisions. However, diffusion decoding is inherently a structured, sequence-level process, in which token states evolve under dynamic cross-token dependencies and are therefore poorly captured by purely token-wise unmasking control. Motivated by this mismatch, we view diffusion decoding from a structural perspective, in which iterative denoising progressively gives rise to token subsets that are internally reliable and weakly dependent on the remaining masked positions. We therefore introduce structural commitment, a sequence-level unmasking principle that treats approximately closed token subsets as the basic unmasking units in diffusion decoding. We instantiate this principle with Structural Commitment via Closure Expansion (SCCE), a structure-aware inference-time algorithm that discovers approximate closures and commits high-certainty subsets by unmasking them, with certainty measured by balancing token reliability against external dependency leakage. Empirically, SCCE improves the accuracy–efficiency frontier over local confidence-based unmasking baselines while adding negligible measured per-step overhead in our implementation.

Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a dominant paradigm. Although recent LLM-based ASR models have shown promising performance on public benchmarks, it remains challenging to balance recognition quality with latency and overhead, while hallucinations further limit real-world deployment. In this study, we revisit LLM-based ASR from an entropy allocation perspective and introduce three encoder-side diagnostics to characterize how training paradigms shape uncertainty reduction across the speech encoder-LLM interface. To remedy entropy-allocation inefficiencies in prevailing approaches, we propose a capability-boundary-aware multi-stage training strategy that targets parameter efficiency and robustness to hallucinations. Specifically, we redesign the pretraining strategy to alleviate the speech-text modality gap, and further introduce an iterative asynchronous SFT stage between alignment and joint SFT to preserve functional decoupling and constrain encoder representation drift. Experiments on various benchmarks show that our method achieves competitive performance against state-of-the-art models using only 2.3B parameters, while also effectively mitigating hallucinations through our decoupling-oriented design.


Rethinking in Spikes: Mitigating Hallucinations in MDLMs with Step-Aware Decoding

Zhongxing Xu ⋅ Zhonghua Wang ⋅ Zhe Qian ⋅ Shiyan Su ⋅ Ming Hu ⋅ Xiaocheng Zou ⋅ Wei Feng ⋅ MINGQUAN LIN ⋅ Yuyin Zhou ⋅ Yifan Peng ⋅ Hamid Rezatofighi ⋅ Zongyuan Ge

Recent advancements in multimodal diffusion language models (MDLMs) have exhibited strong global modeling capabilities through iterative refinement. We observe that mask uncertainty follows an overall downward trend during iterative denoising, while a few denoising steps exhibit localized entropy spikes closely associated with hallucinations. We argue that anomalous steps can be identified by measuring the entropy deviation of each denoising step from its neighboring temporal window. We hypothesize that discrete token commitment discards the distributional information at masked positions, limiting the continued interaction between candidate semantics and the context during entropy spikes. With this goal, we present StepRefine, a plug-and-play decoding strategy that leverages semantic context to achieve reliable denoising. The core of our method lies in a step-aware switch between discrete denoising and latent refinement. The model employs probability-weighted embeddings at entropy spike steps to preserve diverse reasoning hypotheses and performs multiple rounds of continuous refinement. Meanwhile, a refinement controller monitors distributional stability to reduce unnecessary loops and switch the model back to standard discrete denoising. Through extensive experiments, StepRefine demonstrates significant hallucination mitigation across different MDLMs on multimodal benchmarks.


Rethinking Learning from Label Proportions via Moment Matching

Tianhao Ma ⋅ Wei Wang ⋅ Yivan Zhang ⋅ Dong-Dong Wu ⋅ Takashi Ishida ⋅ Gang Niu ⋅ Masashi Sugiyama

Learning from label proportions (LLP) is a weakly supervised setting where training data are grouped into bags with only bag-level label proportions observed, and the goal is to learn a classifier that predicts labels for individual instances. Recently, a large body of methods has emerged; however, most of them rely on the assumption that instances and labels are sampled i.i.d. and randomly assembled into bags, an assumption often violated in real-world scenarios. Our experimental results indicate that these methods perform suboptimally under non-random settings. In this work, we adopt a more realistic assumption: bags are sampled i.i.d., and instance labels are conditionally independent given the instances. Under this setting, we show that the label counts follow a Poisson multinomial distribution. Motivated by this observation, we propose LLP via moment matching(LLP-MM), a simple yet effective approach which leverages multi-order factorial moments as training objectives, encouraging the classifier predictions to match these moments and thereby more fully exploit the underlying statistical structure of the data. Extensive experiments on benchmark datasets under various bag construction strategies demonstrate the effectiveness of our approach while maintaining high computational efficiency.

Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, through empirical studies, we first discover a massive, untapped performance headroom for personalized generation through test-time scaling. We demonstrate that personalized generation is uniquely suited for test-time sampling methods like Best-of-$N$ (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While standard reward models can theoretically exploit this headroom, their massive parameter counts introduce a prohibitive computational bottleneck. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale Multi-Layer Perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing $N$ candidates. Extensive experiments on nine datasets across three different personalized generation settings demonstrate that our framework effectively exploits the discovered headroom. Remarkably, our million-parameter MLP performs competitively with billion-parameter reward models explicitly finetuned on the same tasks, while requiring only a negligible fraction of the inference cost.


RETR: A Structure-Preserving RGB-Event Transformer for Robust 3D Lane Detection

Jingtao Dong ⋅ Hao Zhuang ⋅ Hao Yang ⋅ Liyuan Pan ⋅ Wei Liang

Robust 3D lane detection requires accurate metric road geometry recovery from visual inputs, yet conventional RGB-based approaches suffer from fragile lane evidence under adverse illumination, motion blur, and long-range perspective compression. Event cameras offer complementary high-temporal-resolution, high-dynamic-range structural cues that are robust to these challenges, but their sparse, motion-dependent responses cannot be directly fused into consistent 3D geometry. In this paper, we present the first attempt to introduce event cameras into 3D lane detection. To enable systematic research, we first build two multimodal benchmarks with metric 3D lane annotations: DSEC-3DLD (real-world sequences) and Ev-OpenLane (large-scale simulated sequences). We further propose RETR, a structure preserving transformer that converts complementary RGB and event observations into coherent 3D lane geometry. RETR first aligns modality consistent evidence through Reciprocal Context Flow Fusion, then preserves thin and uncertain lane structures with Uncertainty Aware Structural Consolidation, and finally decodes ordered lane hypotheses using a Geometry State Decoder with proposal conditioned initialization, reference conditioned query evolution, and diversified geometric inquiry. Extensive experiments show that RETR achieves state-of-the-art performance on both benchmarks, especially showing strong improvements in challenging lighting and complex road geometry scenarios.


Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

Jianchao Zhao ⋅ Huoren Yang ⋅ Hu Yusong ⋅ Yuyang Gao ⋅ Qiguan Ou ⋅ Cong Wan ⋅ SongLin Dong ⋅ Zhiheng Ma ⋅ Yihong Gong

Vision-Language-Action models (VLAs) have shown strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades under local deployment conditions. Existing evaluations typically treat test episodes as independent zero-shot trials, whereas real robots often operate repeatedly in the same or slowly changing environments, where successful executions provide environment-verified evidence about reliable behavior patterns. We study this persistent-deployment setting and ask whether a partially competent frozen VLA can improve its reliability by reusing its own successful test-time experience. We propose an online success-memory guided test-time adaptation framework for generative VLAs. During deployment, the robot stores progress-calibrated successful observation-action segments in a long-term memory. At inference time, it retrieves state-relevant successful action chunks, filters action-inconsistent candidates through trajectory-level consistency, and aggregates the filtered candidates into an elite action prior. To incorporate this prior into action generation, we introduce confidence-adaptive prior guidance, which injects the elite prior into an intermediate state of the flow-matching action sampler and adjusts the guidance strength according to retrieval confidence. This design allows the frozen VLA to exploit environment-specific successful experience while preserving observation-conditioned generative refinement. This retrieve-then-steer mechanism enables lightweight, non-parametric test-time adaptation without updating model parameters or modifying the generative solver. Experiments in simulation and real-world manipulation demonstrate improved task success and closed-loop stability, particularly in long-horizon and multi-stage tasks.


Revisiting Autoregressive GCNs for Vehicle Routing Problems

Zhipeng Zhong ⋅ Junquan Huang ⋅ Yu Huang ⋅ Boyuan Zheng ⋅ Jungang Li ⋅ Zihao Dongfang ⋅ Song Dai ⋅ Xuming Hu ⋅ Zong-Gan Chen

Autoregressive (AR) Attention Models have become the dominant paradigm in neural solvers for vehicle routing problems (VRPs), while GCNs are almost exclusively paired with non-autoregressive (NAR) decoding and post-search algorithms. In this work, we revisit the role of decoding strategies and model properties in neural VRP solvers, and challenge the conventional pairing of NAR decoding with GCNs. We propose AR-GCN, an efficient and generalizable architecture for both Symmetric and Asymmetric VRPs (SVRPs and AVRPs). Extensive benchmarks across multiple tasks show that AR-GCN achieves advanced generalization performance, particularly on AVRPs (e.g, achieving an optimality gap of 1.71% on ATSP-1K and -4.46% on ACVRP-1K). In particular, AR-GCN requires only 0.633M learnable parameters and 3-epoch training, providing new insights into the design of AR solvers for general VRPs.


Revisiting the Adam–SGD Gap Beyond Single Factors

Chenxiang Zhang ⋅ Rustem Islamov ⋅ Enea Monzio Compagnoni ⋅ Jun Pang ⋅ Aurelien Lucchi ⋅ Antonio Orvieto

Prior work has identified several factors that can contribute to the performance gap between Adam and SGD, spanning data properties, architecture design, and optimization dynamics. Yet these explanations are often studied in isolation, leaving their relative importance unclear. In this work, we revisit these hypotheses through a controlled empirical study across vision, language, genomics, and graph tasks, spanning modern and classical architectures, and carefully designed training setups. Our results suggest that no single factor consistently explains the Adam--SGD gap. For instance, the Adam advantage can (1) persist under a uniform vocabulary distribution yet nearly disappear under a heavy-tailed one; (2) reverse in favor of SGD in softmax-attention models; and (3) become larger when ReLU is replaced by GeLU nonlinearity. Instead, the gap emerges from interactions between data and architectural properties. Yet, we observe a consistent pattern across our settings: a crossover batch size at which the relative advantage shifts from SGD to Adam as batch size increases. This perspective helps reconcile several existing hypotheses while offering practical insights across domains.

While Value Iteration (VI) is one of the most fundamental algorithms in Reinforcement Learning, its theoretical convergence guarantees still exhibit a persistent mismatch with empirical behavior. In the discounted-reward case, classical theory guarantees geometric convergence with rate $\gamma$, while in the average-reward case recent work suggests that only sublinear convergence can be expected. In practice, however, VI is often observed to converge significantly faster. In this work, we show through a unified geometry-based analysis that, under an assumption of a unique and unichain optimal policy, (i) convergence is geometric in *both* the discounted- and average-reward settings and (ii) the convergence rate is faster than previous analyses suggest.


ReVMap: Vectorized Global Mapping via Connectivity-Aware Local Map Fusion

Hongjae Shin ⋅ Jungho Kim ⋅ Donghyuk Kwak ⋅ Myeongjun Kim ⋅ Seunghoon Yu ⋅ Jiyong Oh ⋅ Jun Won Choi

Large-scale annotation of vectorized high-definition (HD) maps requires substantial human effort, motivating local-to-global auto-labeling that incrementally constructs global maps from local vector predictions. In this formulation, limited reliability management allows inaccurate predictions to be accumulated and reused as priors, causing persistent error propagation across iterations. We propose ReVMap, a local-to-global framework for reliable vectorized global map construction. ReVMap introduces a reliability-aware map evolution strategy that unifies uncertainty-guided prior reuse with revisable global map updates. To realize this strategy, ReVMap estimates localization uncertainty for vectorized local map elements and propagates these reliability cues to the accumulated global map. The resulting uncertainty-aware global representation guides subsequent local prediction and facilitates revisable integration through learning-based local-to-global connectivity estimation and uncertainty-aware fusion. ReVMap achieves state-of-the-art results on nuScenes and Argoverse2, with 36.1 mGAP, 43.4 miECM, and 58.0 local mAP on nuScenes, and 62.1 mGAP, 62.5 miECM, and 77.0 local mAP on Argoverse2.


RGF: Recursive Generative Framework for the Edge-Cut Separable Problems

Bo Peng ⋅ Rongzhen Ye ⋅ Yaohua Li ⋅ Weilin Luo ⋅ Hai Wan ⋅ Jiahao Xu ⋅ Zexian Liang

The Maximal Independent Set (MIS), Dominating Set (DS), and Vertex Cover (VC) problems are fundamental combinatorial optimization problems with wide-ranging applications. We identify a key property of these problems: by finding a cut in the graph and determining the states of vertices associated with the cut, the graph can be decomposed into multiple subgraphs whose solutions are independent of each other. Therefore, we classify MIS, DS, and VC as Edge-Cut Separable Problems (ECSPs). This property enables a natural divide-and-conquer strategy, which significantly reduces the computational complexity and allows neural networks to scale to large-scale graphs that were previously intractable. Based on this property, we propose a recursive framework that leverages this property to approximately solve ECSPs, namely \ours. Our framework consists of two major components: a partitioner that identifies a cut in the graph, and a predictor that solves the cut-determination sub-problem (CDSP) to determine the states of vertices associated with the cut. By recursively applying these components, RGF generates a complete solution for the original ECSP instance. Additionally, through partial prediction, the memory consumption is significantly reduced, enabling RGF to handle real-world large graphs. Experimental results validate the advantages of RGF: the performance of our method is up to 2.13\% better than the state-of-the-art machine learning (ML)-based approach, and our approach can handle real-world large graphs with up to 11M vertices.


Riemannian Optimization for Low-Rank Adaptation via Desingularization

Tiancan Feng ⋅ Daorui Ding ⋅ Fanhua Shang ⋅ Xiaoyuan Zhang

Low-rank Adaptation (LoRA) has been widely used as a parameter-efficient method in fine-tuning large language models (PEFT). However, trained adapters often exhibit low stable rank and small trailing singular values. Near such rank-deficient regions, fixed-rank formulations suffer from severe Hessian bounds and local Lipschitz constants that grow with $\mathcal{O}\big(\sigma^{-1}_r(X)\big)$. To mitigate this issue, we propose ROLAND (**R**iemannian **O**ptimization for **L**ow-Rank **A**daptatio**N** via **D**esingularization), a desingularization method that lifts LoRA to a smooth bounded-rank geometry with both left and right null-space projectors. ROLAND replaces rank-deficient singularities with compact and smooth fibers while retaining a low-rank representation of the adapter update. Within this framework, we equip the manifold optimization with a new Riemannian metric and retraction mechanism, which allow us to establish tighter Hessian bounds and milder condition-number dependence in the local convergence rate. Empirical results demonstrate ROLAND's superiority across a wide range of experimental settings compared with other state-of-the-art PEFT methods.


RIFLE: Removal of Image Flicker-Banding via Latent Diffusion Enhancement

Libo Zhu ⋅ Zihan Zhou ⋅ Xiaoyang Liu ⋅ Zhiyi Zhou ⋅ Weihang Zhang ⋅ Keyu Shi ⋅ Yifan Fu ⋅ Yulun Zhang

Capturing screens is common, but photos of emissive displays are often influenced by \emph{flicker-banding} (FB), some alternating bright--dark stripes due to temporal aliasing between a camera's rolling-shutter readout and display's brightness modulation. Unlike moir\'e degradation, FB remains underexplored despite its frequent and severe impact on readability and perceived quality. We formulate FB removal as a dedicated restoration task and introduce Removal of Image Flicker-B}anding via Latent Diffusion Enhancement, RIFLE, a diffusion-based framework designed to remove FB with fine details. We propose Banding-Suppressed High-Frequency Prior (BSHP), which combines gradient-based high-frequency localization with a smooth structural support map to build a compact restoration prior, and injects it into the restoration backbone via multi-stage FiLM modulation. Moreover, Masked Loss (ML) is proposed to concentrate supervision on banded regions without sacrificing global fidelity. To overcome data scarcity, we provide a simulation pipeline synthesizing FB in the luminance domain with stochastic jitter in banding angle, spacing, and width. Feathered boundaries and sensor noise are also applied for a more realistic simulation. For evaluation, we collect a paired real-world FB dataset with pixel-aligned banding-free references captured via long exposure. Across quantitative metrics and visual comparisons on our real-world dataset, RIFLE consistently outperforms recent image reconstruction baselines. To the best of our knowledge, it is the first work to research the simulation and removal of FB. Our dataset and code will be released soon.


RiSE: Residual Subspace Expert for Generalizable Text-Centric Image Forgery Localization

Kahim Wong ⋅ Kemou Li ⋅ Yiming Chen ⋅ Haiwei Wu ⋅ Jiantao Zhou

As AI-assisted image editing becomes increasingly prevalent, Text-Centric Image Forgery Localization (TFL) is essential for protecting trust in financial and legal records. However, we observe that existing TFL detectors predominantly rely on Full-Parameter Fine-Tuning (FPFT) and fail to generalize to unseen forgery types because FPFT can induce low-rank feature collapse and overfit training artifacts. Meanwhile, prior detectors capture diverse forgery traces with a single unified representation, making subtle traces entangled and difficult to learn. To address these challenges, we propose RiSE, a Residual Subspace Expert model that preserves pre-trained priors by freezing principal SVD components and adapting only the residual subspaces with a ViT-only segmentor to avoid feature collapse. RiSE introduces forgery-aware residual experts, where each expert is trained independently on a training subset with a particular forgery type, and we show that cross-domain localization performance improves consistently as more experts are added. At inference, we introduce Latent Perturbation Confidence (LPC), which selects the most confident expert by analytically computing the latent-perturbed output. LPC enables efficient confidence estimation with a single forward pass, avoiding repeated forward passes on perturbed samples and improving expert selection. By training on synthetic forgery data, RiSE with 1 expert outperforms state-of-the-art methods on real-world forgery images by 20.4\% while reducing training steps by $10\times$. Scaling the number of experts to 9 yields a 35.8\% gain. RiSE also remains effective with DCT feature fusion, even when the majority of the parameters is frozen with RGB-only pretraining. The code is in the supplementary material.


Risk-Aware Action Repetition via Expected Skip Evaluation

Hyunwoo Park ⋅ Junhyeok Um ⋅ Baek-Ryun Seong ⋅ Sang-Ki Ko

Open-loop action repetition—hereafter also referred to as a skip—enables reinforcement learning agents to reduce decision frequency by executing selected actions for adaptive durations. A key challenge is how to evaluate the value of such a skip. Existing Skip-MDP methods typically bootstrap from the greedy value of the terminal state, implicitly assuming that the agent can immediately resume near-optimal control after the skip. This assumption is unreliable during learning, when the underlying action policy is stochastic, inaccurate, and continually changing. It can also amplify optimistic estimation errors, distorting the learned preference over skip lengths. We propose Expected Skip Evaluation, a reformulation of skip-value learning that evaluates skip terminal states through expected future control rather than idealized greedy control. This reflects imperfect control during training while preserving the optimal skip-value fixed-point as control becomes reliable. We instantiate this principle in RARe (Risk-Aware Repetition), a practical framework for discrete and continuous action spaces. Experiments across grid-world, continuous-control, and safety-critical benchmarks show improved sample efficiency and better-calibrated skip decisions.


RoboExo: Structure-Guided Wrist-to-Exocentric Video Generation for Scalable Robot Learning

Rui Li ⋅ Zixuan Hu ⋅ Chenxi Li ⋅ zhangrui zhao ⋅ Fucheng Cai ⋅ Li Kang ⋅ Minting Pan ⋅ LINGYU DUAN ⋅ Dongzhan Zhou ⋅ Wangmeng Zuo

Scaling Vision-Language-Action (VLA) policy training requires diverse robot demonstration data, yet collecting such data remains expensive and labor-intensive. The community has explored ways to accelerate robot data collection: visual augmentation methods synthesize novel samples by editing visual elements in existing demonstrations, while systems such as UMI leverage wrist-mounted cameras to flexibly collect in-the-wild manipulation videos. However, the mismatched observation setups of these two directions leave them largely disconnected, limiting their synergistic potential for scalable robot data generation. To bridge this gap, we propose RoboExo, a generative framework for controllable wrist-to-exocentric conversion of robotic demonstrations. We introduce a geometry-motion factorization strategy that reconstructs canonical object-robot geometry and propagates motion states from noisy wrist-view videos, yielding reliable structural priors for cross-view synthesis. These priors guide a conditional video diffusion model to generate target-view exocentric videos that are spatially aligned, temporally coherent, and faithful to the underlying manipulation semantics. We evaluate RoboExo on a wrist-to-exo generation benchmark under both seen and unseen settings, where it consistently outperforms all baselines, reducing LPIPS by 76% and FVD by 64\%. Beyond generation quality, RoboExo improves VLA policy learning across six tasks, increasing the average success rate by +30\% with generated exocentric views. These results demonstrate RoboExo as an effective and scalable data engine for enriching low-cost UMI demonstrations. The code will be publicly available.


Robust Approximate Nearest Neighbor Search for Any Dataset

Alexandr Andoni ⋅ Themistoklis Haris ⋅ Esty Kelman ⋅ Krzysztof Onak

Approximate Nearest Neighbor Search (ANN) is an important algorithmic primitive that has found a plethora of applications in machine learning and information retrieval. The classic approach to this problem due to Indyk and Motwani (1998) leverages Locality Sensitive Hashing (LSH) by sampling multiple hash functions that are likely to identify similar data points. This approach has, however, been demonstrated to be vulnerable to adaptive queries and updates which may not be independent of its internal randomness (Kapralov et al. 2024). While differential-privacy-based techniques have yielded robust versions of randomized algorithms and data structures for estimation problems (Hassidim et al., 2022), ANN is a search problem: the algorithm must return an actual dataset point, without obfuscating the output by adding noise or rounding it. Feng et al. (2025) circumvent this limitation by relying on a density assumption that bounds the number of points near each query point. We study a stronger worst-case model in which the adversary chooses both the initial dataset and an adaptive query sequence. For LSH-equipped metric spaces, we give adversarially robust ANN algorithms with sublinear query time whose guarantees do not depend on any structural assumption on the dataset. The main technical challenge is to preserve sublinear search time when the adversary selects a dataset with arbitrarily many points near a query. Rather than applying differential-privacy robustification as a black box, we develop search-specific mechanisms: fairness as a route to adaptive security, a bucketing reduction from search to robust decision, and a concentric annuli construction that improves the query-time exponent. Moreover, for low-dimensional spaces, we give algorithms with a strong ``for-all'' guarantee, which are correct for every possible query.


Robust Domain Generalization under Divergent Marginal and Conditional Distributions

Jewon Yeom ⋅ Kyubyung Chae ⋅ Hyunggyu Lim ⋅ Yoonna Oh ⋅ Dongyoon Yang ⋅ Taesup Kim

Domain generalization (DG) aims to learn predictive models that can generalize to unseen domains. Most existing DG approaches focus on learning domain-invariant representations under the assumption of conditional distribution shift (i.e., primarily addressing changes in $P(X\mid Y)$ while assuming $P(Y)$ remains stable). However, real-world scenarios with multiple domains often involve compound distribution shifts where both the marginal label distribution $P(Y)$ and the conditional distribution $P(X\mid Y)$ vary simultaneously. To address this, we propose a unified framework for robust domain generalization under divergent marginal and conditional distributions. We derive a novel risk bound for unseen domains by explicitly decomposing the joint distribution into marginal and conditional components and characterizing risk gaps arising from both sources of divergence. To operationalize this bound and prevent it from becoming vacuous, we explicitly enforce feature-space $\ell_2$-normalization. We then design a meta-learning procedure that minimizes and validates the proposed risk bound across seen domains, ensuring strong generalization to unseen ones. Empirical evaluations demonstrate that our method achieves competitive performance not only on conventional DG benchmarks but also in challenging multi-domain long-tailed recognition settings where both marginal and conditional shifts are pronounced.


ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

Michael S Lee ⋅ Yash Maurya ⋅ Drew Rein ⋅ Bert Herring ⋅ Jonathan Nguyen ⋅ Kyungho Song ⋅ Udari Madhushani Sehwag ⋅ Jiyeon Cho ⋅ Kaustubh Deshpande ⋅ Yeongkyun Jang ⋅ Jiyeon Joo ⋅ Minn S Choi ⋅ Evi Fuelle ⋅ Christina Q Knight ⋅ Joseph Brandifino ⋅ Max Fenkell

Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual safety is typically assessed through translation-only benchmarks that preserve the underlying scenario, and empirical evidence of how language and geopolitical context interact remains limited to a narrow set of language pairs. We introduce ROK-FORTRESS, a bilingual, culturally adversarial NSPS benchmark that uses the English--Korean language pair and U.S.--ROK geopolitical axis as a case study, separating the effects of language and geopolitical grounding via a transcreation matrix: adversarial intents are evaluated under controlled combinations of (i) English versus Korean language and (ii) U.S. versus Korean entities, institutions, and operational details. Each adversarial prompt is paired with a dual-use benign counterpart to quantify over-refusal. Model responses are then scored using calibrated LLM-as-a-judge panels, applying our expert-crafted, prompt-specific binary rubrics. Across a dual-track set of frontier and Korean-optimized models, we find a consistent suppression effect in Korean variants and substantial model-to-model variation in how geopolitical grounding interacts with language. In many models, Korean grounding mitigates the Korean language-driven suppression---with no model showing significant amplification in the other direction---indicating that, at least in the English--Korean case, safety behavior is shaped by language-as-risk signals and context interactions that translation-only evaluations miss. The transcreation matrix methodology is designed to generalize to other language--culture pairs.


ROLLVERIFY: BRIDGING EFFICIENCY AND ACCURACY IN LONG-TAIL ROLLOUT REINFORCEMENT LEARNING

Yongqiang Yao ⋅ Jingru Tan ⋅ Kaihuan Liang ⋅ Zixin Yin ⋅ Yazhe Niu ⋅ Ruihao Gong ⋅ Dahua Lin ⋅ Ningyi Xu

Reinforcement learning is crucial for improving large language models’ reasoning and generalization. It relies on massive rollouts whose lengths become increasingly long-tailed as context windows grow. In on-policy training, these long-tail rollouts can result in GPU bubbles, reducing system utilization and limiting RL scalability. Asynchronous or partial-rollout methods improve throughput by relaxing synchronization, but inevitably introduce stale off-policy samples (trajectories) that may hurt final accuracy. Existing approaches mainly mitigate this off-policy issue by reweighting off-policy samples during training, yet they can still leave a performance gap compared to fully on-policy training. In this work, rather than passively reweighting samples during training, we propose RollVerify, a lightweight RL framework built on partial rollout that actively verifies and repairs samples before they enter training. Specifically, it introduces an off-policy shift metric OPS, to quantify the off-policy deviation of partially generated trajectories. Guided by the OPS constraint, RollVerify performs both sequence-level and token-level verification to identify and truncate invalid suffixes of trajectories. This yields high-quality samples that protect the models’ accuracy while preserving the efficiency gains of partial rollout. Experiments across model scales, architectures, and task domains show that RollVerify matches the performance of on-policy training while significantly improving training efficiency.


RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers

Mengyuan Fan ⋅ Bokai Huang ⋅ JiaMing Pan ⋅ Xiaokun Yuan ⋅ Peizhuang Cong ⋅ Zhewen Tan ⋅ Tong Yang

Vision Transformers (ViTs) achieve strong performance on image recognition and mobile vision applications, but their high-dimensional linear projections and attention computations still impose substantial storage and inference costs. Extremely low-bit quantization is a promising solution, yet ViTs often suffer severe accuracy degradation because conventional real-valued scalar codebooks are poorly matched to the directional geometry of Transformer projections. We present RPFQ-ViT, a Rotated Phase-Frame Quantization method that quantizes paired channels in two-dimensional phase planes, enabling low-bit codes to better preserve projection directions while recovering magnitude with lightweight scaling. RPFQ-ViT serves as a drop-in QAT replacement for \texttt{nn.Linear} and does not modify the standard real-valued attention, normalization, or activation computation graph. On ImageNet-1K, RPFQ-ViT-B/16 reaches 79.33\% Top-1 / 94.48\% Top-5 under W2/A4, Swin-T reaches 79.30\% Top-1 / 94.79\% Top-5 under W2/A8, and DeiT-S reaches 77.41\% Top-1 / 93.11\% Top-5 under W2/A8. Ablations, phase-geometry analysis, and direction-preservation metrics show that channel pairing, learnable rotation, phase-anchor learning, and residual phase refinement each improve quantization quality. We further deploy RPFQ-ViT image-classification models on native iOS and Android runtime stacks; with 2-bit packed weights, model size shrinks by roughly $5.4$--$7.1\times$ relative to FP32 and end-to-end on-device latency drops by $1.4$--$1.6\times$. All ImageNet results trained in our codebase use a matched 300-epoch recipe and are reported as mean accuracies over three independent runs. These results show that RPFQ-ViT provides a favorable trade-off among accuracy, compression, and practical mobile deployment for extremely low-bit ViTs.

Large language models (LLMs) have demonstrated promising capabilities in tool calling. However, existing benchmarks predominantly evaluate models in an idealized setting where ground-truth tool definitions are provided in context. This assumption diverges significantly from real-world deployment scenarios, where the model must first identify user intent and retrieve relevant tools from a large, heterogeneous tool repository before invoking them. In this work, we introduce a Retrieval-augmented Tool-calling Benchmark (RTCBench). In RTCBench, we collect 1,840 tasks from existing benchmarks and augment them with two tool environments: 1) None: no available tools provided in the prompt and 2) Distractor: strong distractor tools, while the target tools are excluded. By removing oracle tool access, RTCBench compels LLMs to actively retrieve relevant tools from a large-scale repository via a unified search tool before making tool calls, thereby faithfully simulating the end-to-end complexity of real-world tool-use pipelines. Through comprehensive evaluation of diverse LLMs ranging from 4B-parameter models to proprietary frontier models, we identify two core deficiencies: unreliable tool discovery, where models fail to initiate effective retrieval actions, and post-retrieval grounding failures, where models cannot reliably select and invoke the correct tools from semantically similar retrieved candidates. Furthermore, to mitigate these issues, we release 13,278 retrieval-augmented tool-calling expert trajectories with high-quality chain-of-thought reasoning for LLM post-training. Experiments show that a 4B-parameter model post-trained with GRPO improves its tool-calling accuracy on RTCBench from approximately 47.9\% to over 85.8\%, approaching the performance of the state-of-the-art model Gemini-3.1-Pro.


RUBRIC-MME: Real-User Behavior-grounded Rubric for Multimodal Interaction Capability Evaluation

Jiajie Teng ⋅ Jianping Jiang ⋅ Huiyu Duan ⋅ Jingdong Chen ⋅ Yi Yuan ⋅ Guangming Yao ⋅ Sijing Wu ⋅ Yuqin Cao ⋅ Yixuan Gao ⋅ Xiongkuo Min ⋅ Guangtao Zhai

Multimodal large language models (MLLMs) increasingly serve as always-on assistants in continuous, multi-turn, goal-driven user interactions, yet existing benchmarks evaluate them through pre-authored capability probes, single-turn correctness, or aggregate leaderboard scores. As frontier models saturate such benchmarks, a widening gap emerges between leaderboard performance and real interaction quality, and current evaluation infrastructure can neither characterize nor diagnose this drift. We introduce RUBUIC-MME, a multimodal multi-turn interaction benchmark designed to close this gap along three axes: an authenticity axis, a competence axis, and a diagnosability axis. To close the authenticity axis, we propose behavior-grounded benchmark construction, a privacy safe pipeline that extracts distributional regularities from real deployment logs, including scenarios, persona-goal pairings, intent typologies, multi-turn structures, and uses them as priors to guide public video filtering and a scenario–persona–intent–dialogue generation chain, yielding a fully synthesizable benchmark for two interaction modes, streaming video and multi-image sequences. To close the competence axis, we define a multi-level structured rubric that evaluates models simultaneously at the turn level, including grounding, relevance, and helpfulness, and the session level, including intent-shift recovery, cross-turn consistency, and goal completion, sliced by capability, scenario, and modality. To close the diagnosability axis, RUBRIC-MME further embeds an automated analysis pipeline that clusters failure modes into per-model diagnostic reports with actionable improvement directions, forming an evaluation–analysis–guidance loop. Evaluating representative frontier and open-source MLLMs, we find that session-level competence diverges sharply from turn-level accuracy, and that there are also significant differences in performance across different scenarios and capabilities. We position RUBRIC-MME as a step toward benchmarks that function as evaluation infrastructure rather than ranking artifacts.


RVLoss: Runoff Vote Loss for Self-Supervised LiDAR Scene Flow Estimation

Shiming Wang ⋅ Liangliang Nan ⋅ Julian Kooij ⋅ Holger Caesar ⋅ Yancong Lin

LiDAR scene flow estimates point-wise motion between two consecutive scans, referred to as the source and target. Leading self-supervised methods typically minimize the Chamfer loss, the nearest neighbor distance between the flow-compensated source and the target. However, nearest-neighbor search does not enforce motion rigidity, often leading to inconsistent flows within object instances. Existing approaches address this issue with additional regularization terms, but flow consistency among points remains limited, especially for large objects. We propose RVLoss, a self-supervised loss that incorporates motion rigidity by design, through a runoff vote mechanism. Our key observation is that the point-wise motion, calculated from nearest neighbor search, can often be grouped into a small set of dominant flow candidates by voting (top-k voting. Furthermore, when compensating the source by these candidates, the flow that best represents the underlying rigid motion often yields the highest consensus after a second voting (top-1 voting). Based on this insight, we incorporate the two-stage runoff vote into loss design and create cluster-wise rigid flows and subsequent free-form flows as pseudo-labels for self-supervised learning. RVLoss can be seamlessly integrated into existing feedforward architectures. Experiments on the Argoverse2 2026 Challenge show that models trained with RVLoss achieve state-of-the-art performance among self-supervised approaches, outperforming baseline models trained with alternative loss designs by 20\%. Moreover, cross-dataset evaluations demonstrate consistent performance improvements across four additional datasets. Code will be released upon acceptance.


S²MoE: Shared-Subspace Mixture of Sparse Experts

Feihong He ⋅ Anke Tang ⋅ Enneng Yang ⋅ Hao Jiang ⋅ Guojie Zhu ⋅ Gang Li ⋅ XIAOCHUN CAO ⋅ Li Shen

Model merging offers a training-free paradigm for integrating multiple specialized models into a unified multi-task system, enabling efficient knowledge reuse from models. However, parameter interference remains a critical challenge, limiting the scalability and efficacy of merged models. SVD-based merging techniques represent the current state-of-the-art, excelling in alleviating interference through task-specific knowledge extraction and subspace orthogonalization. % due to aggressive parameter truncation Despite these advances, they inevitably incur performance degradation due to aggressive parameter truncation, which discards valuable information and hinders inference accuracy. Our analysis reveals this truncation-induced knowledge loss as the fundamental bottleneck, motivating the need for mechanisms that preserve and recover discarded expertise without retraining. To address this, we introduce S MoE, a novel framework that augments SVD-based merging with a mixture of sparse experts to refine the truncation gap. S MoE comprises three key components: (1) a shared module that aggregates multi-task knowledge, (2) sparse experts that selectively recover truncation knowledge, and (3) a subspace mapping router to achieve precise training-free routing. Extensive experiments across vision and language benchmarks demonstrate up to 1.6% and 1.2% performance gains over SOTA methods, effectively validating robust and scalable model merging.


Safe Actions Can Form Unsafe Traces: Benchmarking and Shielding Compositional Emergent Risk in AI Agents

Zhongze Wu ⋅ Xiu Su ⋅ Hongyan Xu ⋅ Yichao Cao ⋅ Yueyi Luo ⋅ Jun Long

Large language model agents increasingly act through tools, browsers, code interpreters, and external APIs, turning safety from a single-output problem into a trace-level problem. We identify \emph{compositional emergent risk} (CER), a failure mode where individually safe actions interact across time to produce unsafe outcomes. We show that bounded-window safety filters can miss CER whenever the risky dependency lies outside their visible context. To study this failure systematically, we introduce \textsc{CER-Bench}, a controlled long-horizon benchmark with 440 tasks and 19,525 actions across 5--500 steps, spanning five risk domains and two compositional mechanisms. At the largest tier, each trace induces a 79M+ candidate risk-composition search space. Across 15 frontier and open-source LLM agents, every model exhibits a non-zero compositionality gap (CG), ranging from 44.4\% to 100.0\% with a mean of 82.2\%. Moreover, 86.1\% of compliant executions contain caution language yet still proceed, showing that verbal risk awareness does not reliably prevent unsafe composition. We introduce \textsc{RiskShield}, a conformal-calibrated runtime shield that learns trace-risk boundaries, evaluates planned continuations before commitment, and substitutes risky suffixes while preserving safe prefixes. On 7 held-out \textsc{CER-Bench} models, \textsc{RiskShield} reduces mean \textsc{CER-5} CG from 69.9\% to 0.0\%, with no shielded pass observed in 763 evaluations and a 95\% upper confidence bound of 1.6\%, while preserving 78.4\% step-level task completion. It outperforms cumulative-threshold, sliding-window, full-trace, and published safety baselines, with 63.3\% of tasks requiring no extra API call and additional robustness across longer traces, risk domains, and external agent-safety benchmarks. Code and benchmark are available at \url{https://anonymous.4open.science/r/RiskShield-2284}.


SafeDrug: A Benchmark Dataset for Safety-Critical Pharmacological Reasoning in LLMs

Tengfei Ma ⋅ Yushan Yang ⋅ HOU Jiahao ⋅ Yujie Chen ⋅ Guanghui Ye ⋅ Xuanbai Ren ⋅ Yiping Liu ⋅ Bosheng Song ⋅ xiangxiang Zeng

Large language models (LLMs) have shown strong performance on biomedical tasks, yet evaluating their reasoning in safety-critical pharmacological contexts requires large-scale, structured datasets. We present SafeDrug, a benchmark dataset for systematic evaluation of pharmacological reasoning across multiple task dimensions. SafeDrug integrates heterogeneous sources, including adverse drug event records, drug–drug interaction knowledge, literature-derived evidence, and population-specific information such as child, adult, and older adult groups, into a unified reasoning-oriented QA format. The benchmark comprises two complementary subsets. \textbf{SafeDrug-large} provides broad coverage at scale, while \textbf{SafeDrug-small} is a high-quality human-annotated subset for precise evaluation. It supports tasks such as prediction, risk assessment, mechanism explanation, evidence-grounded QA, and drug replacement across both single-drug and multi-drug settings. Our evaluation shows that LLMs capture outcome-level associations but struggle with mechanistic reasoning, evidence alignment, and population-aware predictions. SafeDrug provides a foundation for reproducible benchmarking and advances research on evidence-grounded and population-aware pharmacological reasoning in LLMs.


SAGE: Semantic Ambiguity Guided Capacity Expansion for Retrieval-Augmented Generation

Xunlei Chen ⋅ Jinyu Guo ⋅ Yi Gong ⋅ Qirui Ye ⋅ Zhaokun Wang ⋅ Wenyi Li ⋅ Wenhong Tian

Retrieval-Augmented Generation (RAG) mitigates Large Language Model (LLM) hallucinations by grounding generation in external knowledge. Although structure-augmented RAG methods improve multi-hop reasoning, they incur high indexing costs and degrade on simple queries. We argue that this robustness gap is driven by capacity bottlenecks in fixed-dimensional inner-product scoring. As corpora grow, semantically close passages create local regions where a base scorer cannot maintain sufficient separability. Reorganizing documents into trees or graphs can expose or partially mitigate this issue, but it does not by itself provide a stronger local scoring function. We propose Semantic Ambiguity Guided Capacity Expansion (SAGE), a lightweight RAG framework that detects such regions using document-side Local Separability Deficit and expands capacity only where needed. SAGE builds a two-layer semantic index, derives atomic query views from corpus passages, and calibrates a parameter-efficient hypernetwork to generate the local Ambiguity-Conditioned Scorers for selected nodes. Extensive experiments show that SAGE improves retrieval and QA performance over dense retrievers and structure-augmented baselines, achieving a better balance between single-hop, multi-hop retrieval accuracy, and robust scalability. Anonymous Code: https://anonymous.4open.science/r/SAGE-1676.

Rounded capacity cuts (RCCs) are among the most effective cuts for capacitated vehicle-routing relaxations, but exact separation is computationally prohibitive at scale. Existing heuristic and neural separators can be fast, but often fail to produce sufficiently many effective cuts within practical separation budgets. We introduce SAG-Sep, a separator that scores node membership in RCC-inducing subsets across vehicle-count levels from a fractional LP solution in a single encoder forward pass, avoiding the iterative predict-and-coarsen inference used by NeuralSEP-style separators. SAG-Sep encodes the LP solution as a Sparse Augmented Graph (SAG), adding flow-informed multi-hop edges that expose long-range routing structure without densifying attention. Beyond LP-aware sparse encoding, we observe that high-value exact RCCs are typically nested, organizing into onion-like subset families, and SAG-Sep's node scores track subset nesting depth. This motivates Onion-MBP (Margin-Band Probing), a training-free subset-level search that explores nodes near the score margin while generating nested proposals, converting independent node scores into structurally coordinated RCC candidates. On the NeuralSEP benchmark with $N=1000$ customer instances, SAG-Sep with Onion-MBP reduces the average root gap by 17.7% relative to the strongest NeuralSEP baseline, while the single-pass SAG-Sep separator achieves a roughly 30-fold per-iteration neural separation speedup. Our findings suggest that LP-aware sparse encoding and the nested structure of high-value cuts are structural patterns worth leveraging in neural separation more broadly.


SAME: Stability-Aware Embedding Extraction in Mixture-of-Experts Language Models

Shufan Yang ⋅ Zifeng Cheng ⋅ Zhiwei Jiang ⋅ Hao Wang ⋅ Miao Xue ⋅ Changhui Sun ⋅ Cong Wang ⋅ Ao Zhou ⋅ Qing Gu

Extracting sentence embeddings from MoE-based large language models is a promising direction, as they provide greater model capacity than dense models at comparable computational cost. Existing works perform the standard forward propagation process to extract embeddings, overlooking a key stability requirement in the MoE encoding process: semantically similar texts should be encoded into similar representations. In this work, we identify two key stability-related phenomena in MoE models: (1) routing stability varies across layers, and (2) shared and routed experts exhibit different levels of stability. To address these issues, we propose SAME, a Stability-Aware MoE Embedding extraction framework that dynamically allocates activated experts across layers and adjusts expert outputs to mitigate component instability and improve embedding quality. Specifically, SAME assigns fewer activated experts to layers with less stable routing, where layer-wise routing stability is measured by the average overlap between the experts activated before and after injecting Gaussian noise into each layer’s router inputs. Additionally, SAME reduces the contributions of routed experts to alleviate the impact of their instability. Notably, our method is training-free, seamlessly integrates with existing approaches, and incurs no additional inference overhead. Experiments on semantic textual similarity benchmarks demonstrate that SAME consistently improves the quality of extracted embeddings across multiple MoE backbones.


Sample-Efficient Optimization over Generative Priors via Coarse Learnability

Pranjal Awasthi ⋅ Sreenivas Gollapudi ⋅ Ravi Kumar ⋅ Kamesh Munagala

We study zeroth-order optimization where solutions must minimize a cost $d(s)$ while maintaining high probability under a complex generative prior $\mathcal{L}(s)$ (e.g., a parameterized model). This reduces to sampling from a target distribution proportional to $\mathcal{L}(s) e^{-T \cdot d(s)}$. Since classical model-based optimization (MBO) lacks finite-sample guarantees for expressive approximate learners, we introduce \emph{coarse learnability}, a flexible statistical assumption requiring only that a learned model covers the target's probability mass within a polynomial factor. Leveraging this assumption, we design an iterative MBO algorithm called ALDRIFT with a sample correction step that provably approximates the target using only a polynomial number of samples. We apply this framework to globally optimizing non-convex objectives bounded by a quadratic envelope in $\mathbb{R}^d$, where we show this assumption is naturally satisfied for a family of ``optimistic'' posterior distributions. To reach global $\varepsilon$-optimality, this implies a sample complexity of $\widetilde{O}(\log 1/\varepsilon)$, a rate characteristic of optimistic space-partitioning methods. We further justify coarse learnability as an assumption for generative priors theoretically, proving that in simple settings, parametric maximum likelihood estimation and over-smoothed kernel density estimators naturally satisfy it. Finally, one motivation for our framework comes from inference-time alignment. Though our primary contribution is formalizing the theoretical foundations of MBO, we provide qualitative evidence that, in simple settings, even primitive LLMs can shift their distributions toward lower-cost regions when fine-tuned with zeroth-order feedback.


Sandboxed Coding Agents are Competitive Omni-modal Task Solvers

Dongping Chen ⋅ Xuanao Huang ⋅ Zhihan Hu ⋅ Qingyuan Shi ⋅ Dianqi Li ⋅ Tianyi Zhou

As multimodal LLMs increasingly emphasize video and audio, a common assumption is that solving such tasks calls for native omnimodal models. We show that this is not always necessary: coding agents equipped with only text+vision and a sandboxed tool-using interface can perform competitively with, and in several settings outperform, SOTA native omnimodal models and predefined multimodal agent scaffolds across multiple audio-video benchmarks. Our trajectory analysis suggests that their advantage comes from coding agents writing code and orchestrating tools to retrieve relevant content from transcripts, frames, and other non-textual modality signals from raw inputs. This effectively converts omnimodal tasks into evidence retrieval and information processing problems, avoiding the inefficiency of ingesting entire videos or audio streams into context. To characterize remaining limitations, we propose a failure taxonomy and a process-level analysis of tool-use traces, and find that simple skill injection, including human-written skills and self-distilled skills from execution logs, can substantially improve performance over the no-skill baseline. To examine whether such capability can be elicited in open-source models, we further introduce Code-X, a complete training recipe with the OmniCoding trajectory dataset and verifiable reward, providing an exploratory baseline on Qwen-3.5-9B and Qwen-3.6-27B. Finally, given the maturity of many-modality understanding, we argue that the more meaningful frontier lies in many-modality processing, and introduce TerminalBench-O, the first process-level benchmark designed for coding agents on real-world omnimodal processing tasks. Together, our findings open new directions for omnimodal content processing and evaluation.


SARL: Label-Free Reinforcement Learning by Rewarding Reasoning Topology

Yifan Wang ⋅ Bolian Li ⋅ David Cho ⋅ Ruqi Zhang ⋅ Fanping Sui ⋅ Ananth Grama

Reinforcement learning is critical to improving large reasoning models, but its success relies heavily on verifiable rewards (RLVR), making it hard to use in open-ended domains where correctness is ambiguous and cannot be verified. Moreover, reasoning trajectories remain largely unconstrained, and optimizing solely toward the final answer can favor early exploitation over generalization. In this work, we ask whether general reasoning ability can be improved by teaching models how to think (the structure of reasoning) rather than what to produce (the outcome of reasoning), and we extend traditional RLVR to open-ended settings. We introduce Structure-Aware Reinforcement Learning (SARL), a label-free framework that constructs per-response reasoning maps from intermediate thinking steps and rewards their reasoning topology. SARL shifts supervision from destination to path, encouraging reasoning trajectories that are both locally coherent and globally efficient. On verifiable math tasks, SARL outperforms prior label-free RL baselines and even exceeds RL methods with ground truth supervision, with average gains of +9.1\% under PPO and +11.6\% under GRPO across four math benchmarks, with particularly large improvements on AIME25 (+35.5\% with PPO and +44.7\% with GRPO). On non-verifiable open-ended tasks, SARL achieves average gains of +34.6\% under PPO and +30.4\% under GRPO on WildBench across five task categories, outperforming prior label-free RL methods and DPO, which relies on additional preference labels. Beyond strong performance, SARL exhibits substantially lower KL divergence and higher policy entropy, indicating more stable and exploratory training dynamics. Code and data are available at https://anonymous.4open.science/r/SARL-EB77


SatNav: A Scalable Benchmark for Long-Horizon UAV Vision-Language Navigation from Satellite Imagery

Jiajun Jiang ⋅ Chunliang Hua ⋅ Zichun Chen ⋅ Yanxing Wu ⋅ Zeyuan YANG ⋅ Jie Song ⋅ Xiao Hu

Urban uncrewed aerial vehicle (UAV) vision-language navigation (VLN) requires agents to follow instructions across extended urban spaces, inherently demanding long-term memory and geospatial grounding. However, scaling existing benchmarks remains difficult because of their reliance on costly reconstructed 3D assets, limiting geographic diversity and episode scale. To address this, we introduce SatNav, a scalable, long-horizon UAV VLN benchmark built from high-resolution satellite imagery. SatNav targets city-level navigation missions and uses satellite crops as approximations of UAV nadir views for visual observations. Through an automated cue-to-episode pipeline, SatNav constructs 118K episodes from 59 scenes across 18 cities, with an average trajectory length of 379\,m. To stress-test long-horizon memory and geospatial reasoning, SatNav defines three task families: Boundary, Landmark, and Route, targeting loop progress tracking, landmark-based spatial grounding, and route following with counting cues. Benchmarking classical and recent large vision language model (LVLM) based VLN agents on SatNav shows that city-scale navigation remains challenging. We further introduce SwiftVLN, a modular framework with switchable memory components, and conduct systematic memory-design ablations. Finally, satellite-to-UAV transfer experiments show that satellite-trained navigation models can operate on real-flight UAV observations, showing the practical relevance of SatNav.


SC$^3$: A Multi-Solvent Solubility Challenge and Benchmark

Vansh Ramani ⋅ Tarak Karmakar ⋅ Sayan Ranu ⋅ Dhairya Kuchhal ⋅ Har Ashish Arora ⋅ Sergei V Tatarin ⋅ Lev Krasnov

Solubility prediction is a standard benchmark in computational chemistry, yet multi-solvent models which reportedly approach the experimental-noise ceiling (i.e. the aleatoric limit) are not yet reliable enough to be deployed. We argue that this gap is partly artefactual: published benchmarks differ in curation policies, evaluate on count-weighted RMSE that hides failure on tail-heavy solvent distributions, and treat the widely cited 0.6–0.8 log S inter-laboratory figure as the aleatoric ceiling even though it reflects worst-case, not expected, disagreement. We introduce SC3, a multi-solvent solubility benchmark built on BIGSOLDB v2.1 with three contributions: (i) a reproducible curation pipeline yielding 101,535 measurements over 1,327 solutes and 206 solvents, with a recalibrated aleatoric floor of 0.106 log S—roughly 6× tighter than the conventional figure; (ii) nested Gold/Silver/Bronze consensus tiers with per-point σ\sigma σ, three leakage-checked splits, and a multi-solvent metric suite (PS-RMSE, Z-RMSE); and (iii) a 31-model benchmark across six families, whose best Bronze PS-RMSE sits at ≈5×εaleatoric\approx 5 \times \varepsilon_{\text{aleatoric}} ≈5×εaleatoric, and we observe this is a gap unclosed by any deep alternative tested. We perform three follow-on analyses: data scaling, transfer from quantum-chemistry solvation energies, and feature-level attribution, which demonstrates that calibrated per-point uncertainty is a reusable infrastructure for diagnosis beyond point prediction.


Scaffold3D: SfM-Conditioned Pointmap Prediction for Multi-View 3D Reconstruction

Frano Rajič ⋅ Yutong Chen ⋅ Haofei Xu ⋅ Zador Pataki ⋅ Marc Pollefeys ⋅ Siyu Tang

Multi-view 3D reconstruction is increasingly driven by feed-forward models, which are fast and robust but often imprecise or globally inconsistent across views, especially in sparse-view and low-overlap settings. Structure-from-motion (SfM) offers a complementary geometric signal through camera poses and sparse 3D points estimated by correspondence filtering and optimization. To propagate this globally consistent \emph{SfM scaffold} into dense geometry, we introduce \textbf{Scaffold3D}, an SfM-conditioned reconstruction framework that injects image patch-aligned SfM point tokens into a pairwise feed-forward pointmap predictor. Our approach preserves image-token reasoning while using explicit SfM structure to condition dense prediction. Since pairwise pointmaps are predicted in local frames, we fuse them with global alignment and an SfM anchoring term. Across established benchmarks (ScanNet++, ETH3D, and Tanks&amp;Temples) and out-of-domain data (4D-DRESS and MV-dVRK), Scaffold3D achieves stronger overall performance than feed-forward reconstruction models, point-token propagation, depth-completion baselines, and recent geometry-conditioned models. The gains are especially clear in low- and no-overlap evaluations. Our two-view SfM-conditioned pointmap predictor also outperforms several multi-view geometry-conditioned baselines, indicating that how geometric inputs are fused with image-token reasoning is as important as the number of conditioned views. Together, these results establish SfM-scaffold conditioning as a practical interface between learned pointmap prediction and classical geometric optimization.


Scalable Fair Learning via Cramér-von Mises Regularization

Albert Gimó Contreras ⋅ Mariia Vladimirova ⋅ Olga Petrova ⋅ Reda CHHAIBI ⋅ Patrick Loiseau

A standard way to enforce group fairness in machine learning models is to add a fairness regularizer to the training loss. Existing dependence-based regularizers, however, are often computationally expensive, with per-batch costs that are typically quadratic or higher in the batch size $B$. We propose a novel group-fairness regularizer based on the Cramér-von-Mises (CvM) sensitivity index, which penalizes statistical dependence between model predictions and a sensitive attribute during training. Our method combines a rank-based CvM estimator with differentiable soft ranking, yielding a bounded training penalty with $\Oc(B \log B)$ per-batch complexity. This is the first sub-quadratic in-processing fairness method that targets genuine joint-distribution dependence. We further establish theoretical connections between the CvM regularizer and standard fairness metrics such as demographic parity and equality of opportunity. Experiments on tabular and image datasets show competitive fairness-utility trade-offs while substantially lowering training overhead compared to existing dependence-based regularizers.


Scene-Adaptive VLA: Efficient Autonomous Driving via Dynamic Layer Routing

YUJIA YANG ⋅ Quan Yuan ⋅ Fan Jiawei ⋅ Yang Li ⋅ Tiange Fu ⋅ Xiaoyuan Fu ⋅ Guiyang Luo ⋅ Jinglin Li

Vision-Language-Action (VLA) models offer a promising paradigm for autonomous driving. However, the massive parameter scale of VLAs strains onboard computational resources during real-time inference. This directly increases energy consumption and limits vehicle battery range, making it critical to minimize computational overhead without compromising driving capability and explainability. Current research on efficient VLA-based driving focuses on sequential adaptation (\textit{e.g.}, Chain-of-Thought and token adjustment), leaving structural adaptation underexplored despite parameter redundancy across dynamic driving scenes. To address this gap, we propose Scene-Adaptive VLA, an efficient framework leveraging dynamic layer routing to adjust active parameters based on real-time scenes. Inspired by the temporal continuity of scenes and model predictive uncertainty, our approach characterizes model-perceived scene complexity via frame-level temporal scene variations and decision deviations. To uncover latent scene-to-parameter correspondences, we encode this characterization into scene-aware tokens, alongside two specialized queries. A budget-conditioned layer routing mechanism evaluates these refined queries to allocate a budget that determines the active parameter ratio and subsequently selects appropriate layer combinations. The entire framework is optimized end-to-end. Extensive evaluations on the Bench2Drive benchmark show that when adapted to a state-of-the-art VLA model, our approach significantly reduces inference-time computational overhead by 53.5\% with negligible degradation in driving and language capabilities. Code will be available.


SCOPE-RL: Stable and Quantitative Control of Policy Entropy in RL Post-Training

Chen Wang ⋅ Zhaochun Li ⋅ Bai Jionghao ⋅ Hexuan Deng ⋅ Guanting Dong ⋅ Ge Lan ⋅ Yue Wang

Reinforcement learning (RL) is a key paradigm for post-training large language models (LLMs), but the widely used Group Relative Policy Optimization (GRPO) often suffers from entropy collapse: exploration quickly disappears, policies converge prematurely, and sample diversity declines, ultimately harming training effectiveness. Existing remedies, including entropy bonuses and clip-based methods, rarely keep entropy within a stable exploration regime and often introduce oscillatory entropy or reward degradation. In this work, we identify a previously overlooked asymmetry in entropy dynamics: under high-temperature sampling, positive and negative samples have opposite effects on policy entropy. Specifically, high-temperature positive samples promote entropy growth, whereas negative samples suppress it. We provide a theoretical explanation for this phenomenon: when entropy decreases during policy updates, its derivative with respect to temperature is strictly positive under positive-sample updates, indicating that high-temperature positive samples can counteract entropy decay, thereby slowing entropy collapse and potentially reversing it. Motivated by this insight, we propose SCOPE-RL, a stable and quantitative entropy control framework through a regularization term constructed from temperature-adaptive positive samples. Extensive experiments show that SCOPE-RL consistently outperforms strong RL baselines on both Pass@1 and Pass@$k$. Our results provide evidence that escaping entropy collapse can improve reasoning performance, while also showing that the benefit is non-monotonic, with an optimal level of exploration for RL post-training in reasoning LLMs.


SeaPilot: Mobile Agent with Self-refining Environment Alignment

Zhigang Zuo ⋅ Senyao Li ⋅ Yufeng Jiang ⋅ Haozhao Wang ⋅ Yichen Li ⋅ Wenchao Xu ⋅ Jingcai Guo ⋅ Ruixuan Li

By combining cloud-side reasoning with edge-side observation and execution, cloud-edge collaboration has emerged as a promising paradigm for mobile UI agents. However, existing approaches face a fundamental trade-off between cost and accuracy. On one hand, step-by-step cloud interaction ensures high accuracy by uploading UI screens at every step and leveraging the cloud's powerful reasoning capabilities, but this inevitably incurs significant privacy costs for edge data and high computational expenses on the cloud. On the other hand, one-shot cloud planning drastically reduces the demand for edge privacy data and cloud compute by generating a complete plan based solely on the initial UI screen. Yet, this often leads to a substantial drop in accuracy, as the cloud agent may rely on invalid environment assumptions such as unseen app capabilities or UI flows. In this paper, we identify that the root cause of this trade-off lies in the environment assumption gap: the cloud agent lacks prior knowledge of the environment information required for edge-side task execution, and is thus forced to choose between planning with real-time information and planning based on speculative assumptions. To tackle this dilemma, we propose SeaPilot, which proactively provides the cloud agent with the necessary environmental context at the initial task submission stage, thereby achieving the dual benefits of both step-by-step and one-shot methods. To realize this, SeaPilot employs an iterative self-refining mechanism that progressively acquires feedback knowledge from execution failures across diverse tasks, enabling it to learn how to supply precise environmental information for each task in advance. Empirically, SeaPilot improves accuracy by 46.7% over one-shot methods, reduces cloud-token usage by 23.6x compared with step-by-step methods, and reduces privacy cost by 89.6%, achieving a better cost--accuracy trade-off. The code is available at https://anonymous.4open.science/r/SeaPilot-C13C.


Search at the Cost of Sampling: Nearly-Instant Latent Space Bayesian Optimization

Donney Fan ⋅ Colin Doumont ⋅ Aleksandra Kalisz ⋅ Paul Duckworth ⋅ Jacob Gardner ⋅ Henry Moss ⋅ Geoff Pleiss

Generative models are increasingly central to many de novo discovery pipelines, in which designs are generated at scale and filtered through virtual screens to determine a set of candidates to experimentally validate. While Bayesian optimization (BO) is a natural fit for this setting, as it uses past evaluations to guide future proposals, the computational overhead required for its sequential decision-making becomes a bottleneck when virtual screens are relatively cheap. We make BO practical in this regime by exploiting the unique combination of a linear model constrained to a spherical domain where high-dimensional latents concentrate. We build off recent work justifying the use of linear surrogates, while deriving nearly closed-form solutions to the surrogate modelling and acquisition problems that exploit spherical symmetry. The result is a $100\times$ speedup over state-of-the-art baselines, with matching or improved performance across molecular and image generation benchmarks. Altogether, our method makes BO a practical drop-in for de novo pipelines where it was previously too slow to consider.


Search More, Think Less: Rethinking Long-Horizon Agentic Search for Efficiency and Generalization

Chengjun Yu ⋅ Shu XU ⋅ Jiaqi Wu ⋅ Qianben Chen ⋅ Tianrui Qin ⋅ Zhu ⋅ Qiexiang Wang ⋅ Jiayu Zhang ⋅ Xinpeng Liu ⋅ Xin Gui ⋅ Jingyi Cao ⋅ Yi Yao ⋅ WANG PIAOHONG ⋅ Dingfeng Shi ⋅ He Zhu ⋅ Tiannan Wang ⋅ Yuqing Wang ⋅ Maojia Song ⋅ Tianyu Zheng ⋅ Jian Yang ⋅ Jiaheng Liu ⋅ Minghao Liu ⋅ Eleanor Jiang ⋅ Wangchunshu Zhou

Recent deep research agents primarily improve performance by scaling reasoning depth, but this leads to high inference cost and latency in search-intensive scenarios. Moreover, generalization across heterogeneous research settings remains challenging. In this work, we propose $Search\ More, Think\ Less$ (SMTL), a framework for long-horizon agentic search that targets both efficiency and generalization. SMTL replaces sequential reasoning with parallel evidence acquisition, enabling efficient context management under constrained context budgets. To support generalization across task types, we further introduce a unified data synthesis pipeline that constructs search tasks spanning both deterministic question answering and open-ended research scenarios with task appropriate evaluation metrics. We train an end-to-end agent using supervised fine-tuning and reinforcement learning, achieving strong and often state of the art performance across benchmarks including BrowseComp (48.6\%), GAIA (75.7\%), Xbench (82.0\%), and DeepResearch Bench (45.9\%). Compared to Mirothinker-v1.0, SMTL with maximum 100 interaction steps reduces the average number of reasoning steps on BrowseComp by 70.7\%, while improving accuracy.


S-EDL: Eliciting Self-Evidence from Sequence Likelihoods for Semantic Calibration of LLMs

Yawei Li ⋅ Jiazheng Li ⋅ David Rügamer ⋅ Bernd Bischl ⋅ Yulan He ⋅ Mina Rezaei

Calibrating large language models (LLMs) in open-ended generation is uniquely challenging because many distinct token sequences can express the same underlying meaning. Prior calibration techniques target either fixed-set classification confidence or the binary correctness of individual strings, making them inherently ill-suited for the dynamic, multi-string nature of open-ended semantic outcomes. To bridge this gap, we introduce S-EDL, a novel evidential method that natively optimizes LLMs for semantic calibration. By eliciting evidence from the sequence likelihoods of the model itself, S-EDL constructs a differentiable Dirichlet prior over prompt-specific semantic classes. Optimizing the evidential loss under this prior yields a calibration-aware training signal over variable semantic classes, completely bypassing the need for fixed multiple-choice options. Evaluations across four open-ended benchmarks and three models demonstrate that S-EDL substantially reduces semantic calibration error while improving accuracy. By aligning generative likelihoods with semantic correctness, this work establishes a framework for deploying trustworthy LLMs in complex, open-ended applications.

Recent progress in speech-driven 3D facial animation has improved vertex level reconstruction quality, but speech-consistent visible articulation remains difficult. This is because speech production follows structured and constrained lip jaw coordination and the mapping from acoustics to motion is inherently one to many. Motivated by the structured patterns of visible articulation, we propose a novel articulation-aware framework that models visible speech through directional lip motions and composes them into geometry consistent facial deformation. To represent visible articulation with three canonical directional motions, spreading, opening, and protrusion, we propose a Speech Articulatory Memory (SAM) that captures the correspondence between speech and directional articulatory motions under phonetic context through key-value memory structure based retrieval and decoding. Then, a Topology-aware Articulatory Composition (TAC) integrates the predicted directional motions under mesh topology to produce coherent 3D facial motion. Experiments on VOCASET and TFHP show that our method achieves the-state-of-the-art performance on standard reconstruction metrics and improves visible articulatory distance and velocity fidelity for key lip factors, while a user study confirms clear preference in lip sync and realism.


SEEK-VAU: Towards Evidence-Faithful Video Anomaly Understanding via Agentic Search

MINGHE WANG ⋅ Weijun Zeng ⋅ Dapeng Luo ⋅ Xinyue Wang ⋅ Nong Sang

Most video anomaly understanding (VAU) systems still infer anomaly judgments from evidence collected in a feed-forward way under predefined rules, typically sampled frames, or events clustered from such frames. This protocol fails when the collected evidence is incomplete under imperfect predefined rules, because once reasoning begins the model can no longer search for the missing evidence. We address this limitation with SEEK-VAU, which reframes VAU as an agentic search-and-verify interaction protocol under a bounded search budget: instead of passively reasoning over pre-collected evidence, the policy actively gathers evidence across the three stages of an anomaly event chain (precursor, trigger, and aftermath) and verifies its sufficiency before finalization, turning event-chain completeness into an explicit design rather than an outcome of predefined rules. Although search-and-verify spans multiple interaction turns, the bounded search budget caps both the per-turn input tokens and the cumulative visual context, keeping inference efficient. To make this behavior learnable, we introduce evidence-faithful counterfactual verification (EFCV), which rewards selected evidence that remains supportive, compact, and necessary under counterfactual verification. We further introduce SEEK-Bench, featuring video-level episodes with temporal interval annotations, semantic QA and event-chain stage labels. Together, SEEK-VAU and SEEK-Bench establish a strong foundation for evidence-faithful and actively verifiable VAU. The code and data will be released.


Segment and Select: Vision-Language Segmentation in 3D Scenarios

Yulin Chen ⋅ Zhihang Zhong ⋅ Yuenan Hou

3D vision-language segmentation aims to segment target objects in 3D scenarios according to the linguistic instructions and visual observations. Prior art heavily relies on the coarse superpoint representation to reduce the computation complexity , which suffers from poor segmentation quality and messy object boundaries. In this paper, we propose the SEGment-And-select (SEGA3D) paradigm for 3D vision-language segmentation that directly operates on the fine-grained visual information and is free from the superpoint dependency. Specifically, we first leverage a mask candidate generator to provide fine-grained categorical mask candidates, substantially improving the quality of candidate masks over the superpoint counterparts. Then, a Large Language Model (LLM) is utilized to generate the semantic and spatial information based on the linguistic description and visual features. The LLM output and visual features are fed to the Semantic-Spatial Selector (SSS) to produce the top-ranking mask candidates. Eventually, the Loopback Verification Module (LVM) is designed to yield the segmentation mask from the selected candidate masks. Our SEGA3D attains competitive performance on ScanRefer, ScanNet and Matterport3D benchmarks. Notably, our SEGA3D surpasses the top-performing counterpart by 8.3 mIoU and 5.3 mIoU on ScanNet and Matterport3D, respectively. Codes will be available upon publication.


Self-Calibrated GUI Reward Model via Inverse Dynamic Modeling

Zeyi Sun ⋅ Shengyuan Ding ⋅ Xingpeng Xia ⋅ Jinsong Li ⋅ Xie Chen ⋅ Dahua Lin ⋅ Jiaqi Wang

Developing generalist GUI agents capable of robust long-horizon execution remains a key challenge. While Reinforcement Learning (RL) offers a promising path for improvement, it is severely bottlenecked by the lack of reliable reward signals in open-ended environments. Ground-truth verification typically relies on expensive human annotation, which is prohibitively scarce at the scale required for training, while existing VLM-based judges frequently hallucinate success on superficial visual cues. This scarcity of trustworthy supervision fundamentally limits the scalability of current methods. In this work, we propose I-Judge, a novel framework for Self-Calibrated Reward Modeling. We leverage Inverse Dynamic Modeling (IDM) as a self-supervised training objective to learn the causal dynamics of GUI interactions from massive unlabeled trajectories, effectively bypassing the bottleneck of scarce human labels. We then introduce a runtime calibration mechanism that weights reward signals by the IDM's action consistency, filtering out spurious successes. Extensive end-to-end RL experiments on OSWorld, ScienceBoard and AndroidWorld demonstrate that our method significantly accelerates convergence and improves final agent performance. Notably, I-Judge demonstrates strong generalization to unseen domains. All the code, data and models will be made publicly available to foster further research.


Self-Distilled RLVR

Chenxu Yang ⋅ Chuanyu Qin ⋅ Qingyi Si ⋅ Minghui Chen ⋅ Naibin Gu ⋅ Dingyu Yao ⋅ Zheng Lin ⋅ Weiping Wang ⋅ Nan Duan ⋅ Jiaqi Wang

On-policy distillation (OPD) has become a popular training paradigm in the LLM community. This paradigm selects a larger model as the teacher to provide dense, fine-grained signals for each sampled trajectory, in contrast to reinforcement learning with verifiable rewards (RLVR), which only obtains sparse signals from verifiable outcomes in the environment. Recently, the community has explored on-policy self-distillation (OPSD), where the same model serves as both teacher and student, with the teacher receiving additional privileged information such as reference answers to enable self-evolution. This paper demonstrates that learning signals solely derived from the privileged teacher result in severe information leakage and unstable long-term training. Accordingly, we identify the optimal niche for self-distillation and propose \textbf{RLSD} (\textbf{RL}VR with \textbf{S}elf-\textbf{D}istillation). Specifically, we leverage self-distillation to obtain token-level policy differences for determining fine-grained update magnitudes, while continuing to use RLVR to derive reliable update directions from environmental feedback (e.g., response correctness). This enables RLSD to simultaneously harness the strengths of both RLVR and OPSD, achieving a higher convergence ceiling and superior training stability.


Semantic-level Exploration for Multi-Agent Reinforcement Learning

Jiangjin Yin ⋅ Zijian Ye ⋅ Rongbo Zhu ⋅ Hangyu Mao ⋅ Zhiwei Xu

Multi-Agent Reinforcement Learning (MARL) faces significant exploration challenges due to exponentially growing joint state-action spaces. Existing exploration methods operate directly in raw state-action spaces, which is inefficient and fails to exploit inherent semantic structure. This paper introduces semantic space into multi-agent exploration. We theoretically establish that the observation-action space can be partitioned into discrete semantic prototypes, providing a principled foundation for transferring exploration statistics across semantically similar situations. Building on this theory, we propose a semantic-level exploration mechanism that first compresses the high-dimensional observation–action space into discrete semantic prototypes via vector quantization, and then applies sliding-window count-based bonuses to enable efficient statistical transfer across semantically similar situations. Our approach is versatile and can be seamlessly integrated with existing value-based MARL frameworks. Extensive experiments demonstrate that our method outperforms state-of-the-art baselines across diverse multi-agent benchmarks in terms of both effectiveness and training efficiency.

Distributional treatment effects can be invisible to means: a treatment may preserve average outcomes while changing tails, modes, dispersion, or rare-event probabilities. Kernel tests can detect discrepancies between interventional outcome laws, but global tests do not reveal where the laws differ. We propose DR-ME, to our knowledge the first semiparametrically efficient finite-location test for interpretable distributional treatment effects. DR-ME evaluates an interventional kernel witness at learned outcome locations, returning causal-discrepancy coordinates rather than only a global rejection. From observational data, we derive orthogonal doubly robust kernel features whose centered oracle form is the canonical gradient of this finite witness. For fixed locations, we characterize the local testing limit: DR-ME is chi-square calibrated under the null, has noncentral chi-square local power, and uses the covariance whitening that optimizes local signal-to-noise for discrepancies visible through the selected coordinates. This efficient local-power geometry yields a principled location-learning criterion, with sample splitting preserving post-selection validity. Experiments show near-nominal type-I error, competitive power against global doubly robust kernel tests, and interpretable learned locations that localize distributional effects in a semi-synthetic medical-imaging study.


SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents

Simon Zhan ⋅ Philip Wang ⋅ Yao Liu ⋅ Yiyan Peng ⋅ Yiqi Lyu ⋅ Zinan Wang ⋅ Qineng Wang ⋅ Zhian Ruan ⋅ Xiangyu Shi ⋅ Xinyu Cao ⋅ Frank Yang ⋅ Zhenyang Ni ⋅ Kangrui Wang ⋅ Ruohan Zhang ⋅ Huajie Shao ⋅ Manling Li ⋅ Qi Zhu

We present SENTINEL, a framework for formally evaluting the physical safety of foundaiton model (FM)-based embodied agents. SENTINEL is the first to provide multi-level safety evaluation across semantic interpretation, plan generation, and physical execution within a unified formal framework. Unlike prior methods that rely on heuristic rules or subjective FM judgments, SENTINEL grounds practical safety requirements in formal temporal logic (TL) semantics that can precisely specify state invariants, temporal dependencies, and timing constraints. It employs a multi-level verification pipeline where (i) at the semantic level, intuitive natural language safety requirements are formalized into TL formulas and the agent's understanding of these requirements is probed for alignment with the TL formulas; (ii) at the plan level, high-level action plans and subgoals generated by the agent are verified against the TL formulas to detect unsafe plans before execution; and (iii) at the trajectory level, multiple execution trajectories are merged into a computation tree and efficiently verified against physically-detailed TL specifications for a final safety check. We apply SENTINEL in VirtualHome and Ai2Thor, and formally evaluate multiple FM-based embodied agents against diverse safety requirements. Our experiments show that by grounding physical safety in temporal logic and applying verification methods across multiple levels, SENTINEL provides a rigorous foundation for systematically evaluating the safety of FM-based embodied agents in simulation-based physical environments, and can effectively expose potential safety violations in interpreting, planning, and executing the tasks.


SeoulMMOD: A Large-Scale Multimodal Origin-Destination Flow Benchmark

Taeyoung Yu ⋅ Seonbin Jo ⋅ Jiwon Kim ⋅ Junyoung Byun

Origin-destination (OD) flow forecasting supports decisions about where and how many people move through a city. Most existing public OD benchmarks provide only a partial view of citywide mobility, covering a single transport service or at most two modes. With partial mode coverage, observed demand changes are difficult to separate into shifts across transport modes and changes in total travel. We introduce SeoulMMOD, a large-scale public benchmark for citywide multimodal mobility in Seoul. SeoulMMOD provides three years of hourly OD flows estimating total mobility for 6 urban travel modes across 25 districts and 426 sub-districts, together with travel-time, travel-distance, calendar, rainfall, and point-of-interest (POI) signals. Using SeoulMMOD, we benchmark 14 forecasting baselines under a common protocol and study spatial scaling, joint training versus training one model per mode, and cross-year generalization. Results show that several spatio-temporal models that work at district level do not scale to sub-district OD forecasting, and that naive joint multi-mode training does not consistently improve accuracy. Together, the dataset and benchmark establish a reproducible testbed for city-scale multi-mode OD forecasting across multiple years. Project page: https://anonymous.4open.science/r/SeoulMMOD

Motivated by the fact that the worth of a coalition may depend on the order in which agents arrive, Nowak and Radzik (1994) (NR) introduced cooperative games with generalized characteristic functions. We study such temporal cooperative games (TCGs), where the worth function v is defined on sequences of agents π rather than sets S. This order sensitivity necessitates a re-examination of axioms for reward sharing. NR and subsequent work proposed several axioms; the resulting solution concepts are still inherently order-oblivious and closely tied to the Shapley value. In contrast, we focus on sequential solution concepts that explicitly depend on the realized order π. We study reward-sharing mechanisms satisfying incentive for optimal arrival(I4OA), which promotes orders maximizing total worth; online individual rationality (OIR), which ensures agents are not harmed by later arrivals; and sequential efficiency (SE), which requires that the worth of any sequence is fully distributed among its agents. These axioms are intrinsic to TCGs, and we characterize a class of reward-sharing mechanisms uniquely determined by them. The classical Shapley value does not directly extend to this setting. We therefore construct natural Shapley analogs in two worlds: a sequential world, where rewards are defined for each sequence–agent pair, and an extended world, where rewards are defined per agent, consistent with the NR framework. In both cases, the axioms of efficiency, additivity, and null player uniquely characterize the corresponding Shapley analogs. But these Shapley analogs are disjoint from the class of solutions satisfying the sequential axioms, even for convex and simple TCGs. Our results reveal a fundamental tension in temporal cooperative games: when order matters, solution concepts satisfying natural sequential incentives must differ structurally from Shapley-based allocations, motivating new axiomatic foundations for sequential reward sharing.


SETA: Scaling Environments for Terminal Agents

Qijia Shen ⋅ Zhiqi Huang ⋅ Vamsidhar Kamanuru ⋅ Aznaur Aliev ⋅ Jay Rainton ⋅ Ahmed Awelkair ⋅ Boyuan Ma ⋅ Qizheng Zhang ⋅ Jiwei Fu ⋅ Yuzhen Mao ⋅ Wendong Fan ⋅ Ping Nie ⋅ Philip Torr ⋅ Bernard Ghanem ⋅ Changran Hu ⋅ Jonathan Li ⋅ Urmish Thakker ⋅ Guohao Li

Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical user interfaces (GUIs). Among these, the terminal command line provides a text-based, general-purpose interface, covering tasks from system operations to data science and machine learning. However, scaling terminal-agent training remains challenging, as it requires diverse and coherent task instructions, executable environments, and reliable verification, while lacking naturally grounded supervision data. In this work, we propose SETA, a scalable framework for generating verifiable terminal environments for reinforcement learning (RL). The framework consists of two pipelines sharing a unified verification mechanism: SETA-Synth converts diverse sources into standardized RL environments, and SETA-Evol further expands from existing environments with adaptive control of difficulty and diversity. Together, we construct and release SETA-Env, the largest open-source verifiable terminal RL dataset to date, containing over $4{,}500$ environments. We evaluate our dataset by training a Qwen3-8B model with GRPO on SETA-Env, achieving $12$\% pass rate on Terminal-Bench 2.0, the best reported result for an RL-trained model at the 8B scale. The results demonstrate that SETA-Env provides high-quality training environments for terminal agents and serves as a valuable resource for advancing research on terminal-based agent learning.


SFT-then-RL Outperforms Mixed-Policy Methods for LLM Reasoning

Alexis Limozin ⋅ Eduard F Ďurech ⋅ Torsten Hoefler ⋅ Imanol Schlag ⋅ Valentina Pyatkin

Recent mixed-policy optimization methods for LLM reasoning that interleave or blend supervised and reinforcement learning signals report improvements over the standard SFT-then-RL pipeline. We show that numerous recently published research papers rely on a faulty baseline caused by two distinct bugs: a CPU-offloaded optimizer bug in DeepSpeed that silently drops intermediate micro-batches during gradient accumulation (affecting multiple downstream frameworks including TRL, OpenRLHF and Llama-Factory), and a loss aggregation bug in OpenRLHF that incorrectly weights per-mini-batch losses. Together they suppress SFT performance, with the optimizer bug accounting for most of the gap and the loss aggregation bug contributing a smaller additional effect. Once corrected, the standard SFT-then-RL pipeline surpasses every published mixed-policy method we evaluate by +3.8 points on math benchmarks with Qwen2.5-Math-7B and by +22.2 points with Llama-3.1-8B. Even a truncated variant with just 50 RL steps outperforms mixed-policy methods on math benchmarks while using fewer FLOPs.

Top-two algorithms are simple and effective for fixed-confidence best-arm identification, but their sharp non-asymptotic behavior is still not well understood. We study this problem for Bernoulli bandits through $\beta$-EB-TCI, the empirical-best top-two rule of Jourdan et al. (2022), whose challenger is chosen using a Bernoulli transportation cost with a logarithmic count penalty. We prove that, after the empirical leader has become the true best arm and its sampling fraction stays close to $\beta$, the stopping time is $T_{\beta}^{\star}(\mu)\log(1/\delta)$ up to lower-order concentration terms. We also show that, in this regime, every challenger is sampled linearly often. Thus, for the original algorithm without forced exploration, the main remaining difficulty is to control when the empirical leader becomes permanently correct. These results imply a non-asymptotic high-probability bound for all Bernoulli instances with a unique best arm. If the algorithm satisfies a finite-mean sufficient-exploration condition, the bound further yields the sharp expected sample complexity. In particular, this gives the sharp expectation result for the unguarded Bernoulli rule when all arm means are pairwise distinct, using the sufficient-exploration result of Jourdan et al. (2022). Finally, if we add a mild forced-exploration rule that contributes only $O(\sqrt{Kt})$ pulls up to time $t$, we obtain a self-contained expected sample-complexity theorem for any number of arms under the unique-best-arm assumption. We also identify a limitation of proof strategies that try to handle equal suboptimal means through a single index-comparison argument.

Despite advances in retrieval-augmented generation (RAG), suppressing hallucinations during complex reasoning remains a persistent challenge. We identify that existing RAG pipelines suffer from a textual bottleneck: by passing only raw text to the generator, they discard discriminative signals of the retriever. Furthermore, recent attempts to unify retrieval and generation within a single LLM induce severe objective conflict and exacerbate the rank collapse of decoder-only architectures, degrading performance on both tasks. To address these limitations, we introduce ShiftRAG, a decoupled RAG framework that operates on a single LLM backbone but isolates discriminative retrieval learning from autoregressive generation. To bridge these decoupled modes, we project retrieval alignments into continuous soft tokens, establishing a high-bandwidth interface that infuses the generator with retrieval-side signals. Concurrently, our soft orthogonality objective mitigates rank collapse, while layer-wise relevance dynamics preserve the fine-grained latent space required for highly discriminative evidence separation. Extensive evaluations demonstrate that ShiftRAG significantly improves retrieval and question answering performance over state-of-the-art baselines. ShiftRAG preserves foundational capabilities of the backbone while maintaining a favorable efficiency-performance trade-off on general text encoding and generative tasks. Our code is available.


Should We Pay This Much for Robustness? Efficient Proxy Certificates with Marginal Guarantees

Sayed Soroush Haj Zargarbashi ⋅ Mohammad Sadegh Akhondzadeh ⋅ Aleksandar Bojchevski

Modern classifiers remain sensitive to small input perturbations, while strong model-agnostic robustness certificates such as randomized smoothing are often too expensive for deployment, requiring an extensive number (e.g. 2000) of model forward passes per input. We propose a framework that delegates certification to a cheap proxy function while statistically controlling its deviation from the original sound certificate. Given an exchangeable unlabeled calibration set, we tune the proxy via robust conformal risk control to guarantee a bounded marginal false-approval rate: the probability that the proxy accepts an input whose prediction is flipped due to perturbation. The resulting proxy certificate can be used either as a standalone marginal certificate or as a fallback mechanism that invokes the expensive certificate only when the proxy rejects. We provide both single-sample and sample-heavy calibration recipes. Empirically, our method substantially reduces test-time certification cost (to a single model forward) while retaining much of the certified performance of randomized smoothing, enabling cheap yet effective robustness guarantees for larger architectures such as vision transformers and language models.

Event-based optical flow estimation achieves high temporal resolution and dynamic range by exploiting asynchronous events, while self-supervised learning provides an effective solution to the absence of ground-truth labels. Spiking neural networks (SNNs), as brain-inspired models, process information through asynchronous spikes, naturally aligning with optical flow events. However, the temporal irregularity across event streams hinders the compact extraction of spatiotemporal features, while the absence of explicit labels leads to spurious correlations, thereby increasing redundancy within latent representations. In this paper, we propose a novel “Compactly Compress, Time-varying Focus” training strategy. We present SICAF ($\text{\textbf{S}pikes}$ $\text{\textbf{I}nformation}$ $\text{\textbf{C}ompression}$ $\text{\textbf{A}nd}$ $\text{\textbf{F}ocus}$) framework that formulates the learning process as a constrained optimization with temporal modulation. It encourages self-supervised SNNs to extract compact latent representations while preventing over-compression that would discard temporal saliency. To better enable feature selectivity across time steps, we design an estimator with learnable temporal attention, achieving context awareness with cross-time dependencies. Experimental results on several datasets demonstrate that SICAF achieves state-of-the-art performance in self-supervised SNNs, with a 8.14\% improvement in AEE on the MVSEC dataset and a 13.49\% increase in robustness under white-box adversarial attacks.

We identify a theoretical incompatibility between 3D point cloud masked autoencoding and SE(3)-equivariant networks. We prove that standard masking triggers a topological collapse, effectively forcing the model to hallucinate orientation without a reference frame. To resolve this, we introduce SIEVE (Scalar Invariant Extraction Via Equivariance), an SE(3)-equivariant self-supervised framework that compresses point cloud geometry into invariant scalar embeddings. At its core is the spherical projection mask, which preserves global pose while obfuscating local semantics, sidestepping the obstruction by construction. Experiments confirm that SIEVE circumvents the predicted collapse where constant and noise masks fail. On downstream point cloud tasks, the resulting embeddings improve over coordinate-based baselines and hand-crafted geometric descriptors, with the largest gains on generalization to unseen shape categories.


Signature-Kernel Evaluation Metrics for Robust Probabilistic and Tail-Event Forecasting

Benjamin R Redhead ⋅ Thomas L Lee ⋅ Peng Gu ⋅ Víctor Elvira ⋅ Amos Storkey

Standard sample-based metrics like Continuous Ranked Probability Score (CRPS) and Quantile Loss (QL) are frequently used in the evaluation of probabilistic time-series forecasting. However, these metrics fail to capture multivariate correlations or assess the ability to capture distribution tails. These global metrics are insensitive to errors on high-utility events in distributional tails due to over-representation of the distribution's body. To accurately measure distributional fidelity on tail-regions and regions of high utility while still having a proper scoring rule, we introduce a family of censored signature kernel metrics. Our proposed metrics concentrate evaluation of forecasts to a focus region, representing high-utility or distributional tails, by collapsing the body of the distribution to a single pivot. Our benchmarks on time-series foundation models (TSFMs) reveal that while a clear ranking can be formed for distribution capture, there is no clear winner for tasks like systemic load prediction. This indicates that models with a strong ability to capture an overall distribution do not produce forecasts with high downstream utility. To encourage the use of the proposed metrics we open-source the efficient signature kernel (ESK) library. This library facilitates batch computation, with custom triton kernels, achieving a speed-up of up to 3.58$\times$ compared to implementations via the popular SigKernel library. Code for our metrics are available on GitHub.

Multi-objective Bayesian optimization (MOBO) is commonly approached through specialized acquisition functions or scalarization schemes designed to explicitly account for trade-offs among non-preferential objectives. In this work, we show that such complexity might be unnecessary. We propose a framework that extends standard single-objective acquisition functions directly to the multi-objective setting through a hypervolume-based transformation. We further extend hedge strategies for acquisition functions, which are typically used only in single-objective optimization, to the multi-objective regime. Our approach requires minimal modification to existing Bayesian optimization pipelines and avoids the need for bespoke multi-objective formulations. We demonstrate how a broad class of commonly used single-objective acquisition functions and hedge strategies can be adapted in a principled manner to handle multiple objectives, while preserving their intuitive interpretation and computational efficiency. Empirically, we evaluate the proposed methods across a range of synthetic and real-world multi-objective benchmarks. Despite their simplicity, our extensions consistently match or outperform more complex state-of-the-art MOBO methods in terms of optimization performance and sample efficiency. These results suggest that effective multi-objective Bayesian optimization can be achieved by reusing and carefully extending well-established single-objective acquisition strategies, offering a simpler and more flexible alternative to existing approaches.


Simple Test-Time Refinement for Plot-to-Code Generation via Visual-Code Diagnostics

Peiming Guo ⋅ Jinpeng Huang ⋅ Yangbin Sun ⋅ GaoXiong Cao ⋅ Ensheng Shi ⋅ Yuchi Ma ⋅ Yu Zhang ⋅ Meishan Zhang ⋅ Baotian Hu ⋅ Min zhang

Plot-to-code generation has been greatly advanced by recent multimodal large language models (MLLMs). However, existing studies mainly focus on one-pass generation or iterative sampling guided by verifier-based filtering, which does not align with the typical human programming workflow of generate first, then iteratively error-location and fix. This naturally motivates a test-time refinement paradigm for plot-to-code generation.In this work, we present the first study on plot-to-code generation via iterative test-time refinement. We propose a Visual-Code Diagnostics framework that enables refinement through visual feedback from source code. Specifically, we first categorize generation errors into four coarse-grained types, and then train lightweight discriminators to recognize and localize these errors based on the rendered visual outputs and corresponding code. Unlike traditional compiler- or debugger-based verification, our approach performs iterative refinement guided by one diagnostics system, enabling more effective semantic and visual correction during test time. Extensive experiments demonstrate that our method consistently improves generation quality across diverse generators and significantly outperforms strong test-time scaling baselines.


SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning

Haoqiang Kang ⋅ Xiaokang Ye ⋅ Yuhan Liu ⋅ Siddhant Hitesh Mantri ⋅ Lingjun Mao ⋅ James Fleming ⋅ Drishti Regmi ⋅ Lianhui Qin

LLM/VLM-based digital agents have advanced rapidly thanks to scalable sandboxes for coding, web navigation, and computer use, which provide rich interactive training grounds. In contrast, embodied agents still lack abundant, diverse, and automatically generated 3D environments for interactive learning. Existing embodied simulators rely on manually crafted scenes or procedural templates, while recent LLM-based 3D generation systems mainly produce static scenes rather than deployable environments with verifiable tasks and standard learning interfaces. We introduce SimWorld Studio, an open-source platform built on Unreal Engine 5 for generating evolving embodied learning environments. At its core is SimCoder, a tool/skill-augmented coding agent that writes and executes engine-level code to construct physically grounded 3D worlds from language/image instructions. SimCoder self-evolves by using verifier feedback (e.g., compilation errors, physics checks, VLM critiques) to revise environments and autonomously add reusable tools and skills to its library. Generated worlds are exported as Gym-style environments for embodied agent learning. SimWorld Studio further enables co-evolution between environment generation and embodied learning: agent performance feedback guides SimCoder to generate adaptive curricula near the learner’s capability frontier, so that environments become increasingly challenging as the embodied agent improves. Three case studies on embodied navigation show that self-evolution improves generation reliability, generated environments substantially improve embodied agent performance that generalizes to unseen benchmarks, and co-evolution yields an 18-point success-rate gain over fixed-environment learning and a 40-point gain over an untrained agent.


Single Shot HDR Recovery via a Video Diffusion Prior

Chinmay Talegaonkar ⋅ Jinshi He ⋅ Nicholas Antipa

Recent generative methods for single-shot high dynamic range (HDR) image reconstruction show promising results, but often struggle with preserving fidelity to the input image: they hallucinate content, require separate models to handle highlights and shadows, or sacrifice interpretability by directly predicting the final HDR image. We address these limitations by re-casting single-shot HDR reconstruction as conditional video generation and fusing the generated frames into an HDR image. We fine-tune a video diffusion model to generate an exposure bracket, conditioned on a low dynamic range (LDR) input. We fuse this image bracket using per-pixel weights predicted by a light-weight UNet. This formulation is simple, interpretable, and effective. Rather than directly hallucinating an HDR image, it explicitly reconstructs the intermediate exposure stack and fuses it into the final output. Our method eliminates the need for separate models across exposure regimes and produces HDR reconstructions with high input fidelity. On quantitative benchmarks, we outperform state-of-the-art generative baselines with comparable model capacity on several reconstruction metrics. Human evaluators further prefer our results in 72% of pairwise comparisons against existing methods. Finally, we show that this input-conditioned sequence generation and fusion framework extends beyond HDR to other image reconstruction tasks, such as all-in-focus image recovery from a single defocus-blurred input.


SkillCIR: Intent-Guided Skill Composition for Training-Free Composed Image Retrieval

Yuanmin Tang ⋅ Lin Li ⋅ Yang Du ⋅ Yuan Gao ⋅ Massimiliano Mancini ⋅ Jun Song ⋅ Gaopeng Gou ⋅ Gang Xiong ⋅ Meikang Qiu ⋅ Cheng Yu ⋅ Bo Zheng

Composed Image Retrieval (CIR) retrieves a target image given a reference image and a textual edit. Existing training-free methods prompt a frozen multimodal large language model (MLLM) to infer the user's edit intent and rewrite the composed query into a single retrieval artifact, such as a target caption, an edited image, or a modality-fused embedding. These methods can reason about the edit, but they have limited means to execute that reasoning during retrieval. A key remaining bottleneck for training-free CIR appears to be not only intent understanding, but intent execution. To address this challenge, we propose SkillCIR, a training-free framework that tackles the intent execution gap by routing the inferred constraints to specialized retrieval skills. Specifically, the same frozen MLLM that infers the intent also emits a structured plan of active constraints; each constraint activates a typed skill that scores gallery candidates along one evidence channel, and the per-skill scores are composed in score space with explicit signs and re-checked by a local verification step. Across three benchmarks and three CLIP backbones, SkillCIR improves the primary recall/mAP metrics by 2.8 to 7.9 points at the latency of typical training-free CIR pipelines and around 23x faster than the latest training-free state of the art. These gains hold with a fixed skill pool and no model-parameter updates, suggesting that improving the means to execute the inferred edit is a useful complement to stronger reasoning about it.


SliMOO: Interpretable Multi-Objective Evolutionary Search for LLM Depth Pruning

Guanchen Li ⋅ Yixing Xu ⋅ Xuanwu Yin ⋅ Dong Li ⋅ Emad Barsoum

As large language models (LLMs) grow more powerful, their scale makes them costly to serve. Depth pruning offers a practical route to efficiency by directly reducing memory and latency without hardware-specific support. However, we observe that transformer layers exhibit markedly different task-dependent layer contributions, whereas existing depth-pruning methods ignore this heterogeneity and thus tend to optimize for a narrow domain or benchmark rather than preserve broad capability. This suggests that LLM depth pruning should be formulated as a multi-objective optimization problem rather than a single-score ranking problem. We propose SliMOO, an interpretable multi-objective evolutionary search framework for LLM depth pruning. SliMOO first uses single-layer removal probing to construct a task-layer importance map, which provides an interpretable view of task-shared and task-specific layers and serves as a search prior. It then performs a Pareto-aware evolutionary search for pruning masks, where candidates are generated under the guidance of the task-layer prior, scored by their deviation from the dense model on each task, and retained through NSGA-II environmental selection. Across the Llama-3.1 and Qwen3 model families, SliMOO consistently achieves stronger multi-task trade-offs than rule-based, greedy and scalarized baselines; for example, on Llama-3.1-70B at 25% sparsity, it improves the math-domain average from 24.33% to 34.40% and the code-domain average from 35.03% to 39.14%.


Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots

Guangcheng Zhu ⋅ Shenzhi Yang ⋅ Haobo Wang ⋅ Xing Zheng ⋅ Yingfan Ma ⋅ Xuening Feng ⋅ Zhongqi Chen ⋅ Bowen Song ⋅ Weiqiang Wang ⋅ Gang Chen

Reinforcement learning with verifiable rewards (RLVR) has greatly advanced large reasoning models (LRMs), but it requires timely training on a huge fully-annotated dataset. To this end, data-efficient RLVR methods have been widely studied from two perspectives: (i) data selection methods identify a small subset of "golden" samples that yield near-full-data performance, but they rely on a pre-existing pool of labeled data. (ii) unsupervised RLVR methods train the model using its own internal supervision signals on large-scale unlabeled data, yet they exhibit suboptimal performance. Accordingly, we investigate the "*pick in the dark*" setup for RLVR, which aims to select, without prior supervision, unlabeled samples that are most beneficial for training and worthy of annotation. Through systematic analysis, we demonstrate that smart picks hinge on a well-calibrated uncertainty estimator to enable strategic partitioning of data for adaptive training regimes. Building on this insight, we propose **PivotTrace**, a three-way data triage framework that leverages attention dynamics to trace *metacognitive pivots* during reasoning. By precisely quantifying uncertainty through pivot density, PivotTrace achieves automated data routing to synergistically maximize both annotation and training efficiency. Empirically, PivotTrace surpasses the fully supervised LRM with only **29.3%** annotated samples and **2.75**$\times$ faster convergence.


SOAR: Semantic Organ-Aware Pretraining for 3D CT Image Understanding

Huidong Xie ⋅ Wei Ji ⋅ Sunan He ⋅ Zhihao Tang ⋅ Wen Tang ⋅ Jun Zhao ⋅ Hongming Shan

Vision-language pretraining has advanced multimodal medical AI, but its direct application to 3D computed tomography (CT) remains limited by a fundamental mismatch between dense volumetric anatomy and sparse diagnostic reports. Existing 3D CT vision-language methods typically rely on global image-report alignment or align anatomical regions with decomposed report descriptions. However, a CT scan captures rich, spatially distributed information across many organs, whereas a report summarizes only a small subset of clinically relevant findings. As a result, report-based alignment provides incomplete or weak supervision, leaving many local anatomical structures underrepresented. Motivated by this limitation, we introduce SOAR, a semantic organ-aware pretraining framework for 3D CT image understanding. The key innovation of SOAR lies in its integration of fine-grained, organ-level structural priors, enabling report-free organ-aware supervision for visual pretraining without relying on radiologist annotations or coarse global image-report matching. SOAR combines three complementary objectives: (i) organ-level masked reconstruction to learn localized anatomical context, (ii) organ-level vision-language alignment to associate organ-specific visual features with organ-name text embeddings, and (iii) organ-level supervision to preserve voxel-level structural detail. SOAR integrates easily with existing LLMs and is evaluated on five public/in-house benchmarks, improving multiple-choice VQA accuracy by 3.53% and disease screening/abnormality detection AUC by 4.25 points on average over all benchmarks/baseline methods. Code will be available.


Soft geometric inductive bias for object centric dynamics

Hampus Linander ⋅ Conor Heins ⋅ Marco Perin ⋅ Alexander Tschanz ⋅ Christopher L Buckley

Physical systems are naturally parameterized in terms of geometric entities and their transformations. While exact equivariance can be a powerful inductive bias, real-world applications frequently exhibit approximate or broken symmetries in practice. We introduce object-centric world models that provide a soft geometric inductive bias without enforcing exact symmetries. By embedding object states as Clifford multivectors, our models encourage physically meaningful transformations while retaining the expressivity needed to handle asymmetric dynamics. We evaluate our approach on 2D rigid-body dynamics, 3D charged-particle systems, and real-world driving trajectories. Compared to both unstructured baselines and strictly equivariant models, our soft Clifford transformer achieves better long-horizon fidelity, particularly in regimes with broken symmetries. These results suggest that geometric algebra offers an effective middle ground, delivering sample-efficient dynamics without the inflexibility of hard mathematical constraints.

Constrained end-to-end learning trains neural networks whose outputs must satisfy feasibility constraints, such as resource, budget, or operational limits. While projection and optimization layers guarantee constraint satisfaction, boundary-based projections can introduce unfavorable optimization geometry: regions of the unconstrained output space are mapped to the same active face of the feasible set. This makes the resulting map locally rank-deficient, suppressing gradients in constrained directions and degrading optimization—a bottleneck known as gradient saturation. We propose Soft-Radial Projection, a differentiable reparameterization layer that maps unconstrained network outputs into the relative interior of a convex feasible set by smoothly rescaling rays from a strictly feasible anchor point. Unlike hard radial or orthogonal projections, the resulting map is one-to-one and has a full-rank Jacobian almost everywhere, while guaranteeing strict feasibility. We prove that constrained networks equipped with Soft-Radial Projection retain universal approximation guarantees. For standard convex sets such as simplices and balls, the layer admits a closed-form forward pass, avoiding the iterative solvers required by optimization-based layers. Experiments on decision-focused learning benchmarks show improved optimization stability and solution quality compared with optimization- and projection-based baselines.


Sparse All-Layer Connector for Domain Generalised Semantic Segmentation

Ruoyu Guo ⋅ XIN KUN LIN ⋅ Maurice Pagnucco ⋅ Yang Song

Domain generalised semantic segmentation (DGSS) has recently benefited from combining multiple foundation models, such as vision-language and vision foundation models, to leverage their diverse representations for generalisation. Existing multi-model approaches share a common design choice: cross-model fusion is restricted to depth-aligned layers, implicitly assuming that mutually beneficial information resides at the same depth across models. However, foundation models pretrained under different objectives may develop distinct representation hierarchies, making the optimality of depth-aligned fusion questionable. Moreover, we observe that the foundation models have different depth-wise domain sensitivity. Motivated by these observations, we propose a Sparse All-Layer Connector (SALC) that enables each layer to access information from preceding layers of another model. As the number of accessible layers grows with depth, dense aggregation may hurt generalisation. SALC therefore learns to adaptively select a sparse subset of informative layers. We further introduce a candidate dropout regularisation that strengthens sparsity and encourages SALC to explore diverse layers, leading to more robust selection. Across four foundation model combinations and three DGSS evaluation settings, the proposed design demonstrates stronger generalisation than conventional depth-aligned fusion. Code will be public upon acceptance.


Sparse Fine-Tuning for Parameter-Efficient Adversarial Training

Zhaoxin Wang ⋅ Zihang Ding ⋅ Handing Wang

Deep neural networks achieve strong performance across many tasks but remain vulnerable to adversarial perturbations. Adversarial training (AT) is one of the most effective defenses, yet it suffers from computational cost on large models and the persistent clean-robust accuracy trade-off. Recent work introduces parameter-efficient fine-tuning, such as LoRA, into AT, but this constraint induces a notable robustness gap relative to full-parameter AT. In this work, we revisit parameter-efficient adversarial fine-tuning from a parameter space perspective. Through gradient analysis, we find that adversarial optimization is highly concentrated on a small subset of parameters, which we call Robustness-Critical Parameters (RCP), suggesting that robustness is encoded in a sparse subspace. Building on this observation, we propose Sparse Adversarial Fine-Tuning (SAFT), which identifies RCP via adversarial gradient saliency and performs adversarial training by updating only these parameters while freezing the rest. Across architectures and datasets, SAFT consistently outperforms LoRA-based methods in both clean and robust accuracy with fewer trainable parameters, and approaches or surpasses full parameter AT in robustness while updating only about 5\% of parameters.


Sparse Updates Generalize Better Than Optimization: Stability Analysis for Randomized Subspace Descent

Yifei Liang ⋅ Yan Sun ⋅ Yifei Cheng ⋅ Haobo Fu ⋅ XIAOCHUN CAO ⋅ Li Shen

Large-scale learning increasingly relies on updating only a coordinate, a block, or a low-dimensional subspace per iteration to reduce memory and communication costs, yet the statistical consequence of such reduced update coverage, beyond the well-understood optimization slowdown, remains open. Available stability analyses are restricted to coordinate-level projectors whose orthogonality drives the argument, deterministic-output assumptions incompatible with randomized updates, or uniform step sizes that ignore curvature-adaptive blockwise schedules; none explains how general subspace coverage controls population risk. We analyze this question through randomized subspace descent (RSD), a framework that recovers GD, RCD, BCD, and isotropic SSD as special cases: under joint isotropy the expected projection energy reduces to a single coverage parameter~$\alpha$, enabling unified stability analysis without coordinate-level structure. The convex excess risk scales as $O(\alpha^{-1/4}N^{-1/2})$, improving to $O((N\sqrt{\alpha})^{-1})$ under low noise, substantially gentler than the $O(1/\alpha)$ optimization slowdown, while in the nonconvex Polyak--{\L}ojasiewicz setting $\alpha$ lengthens the early-stopping horizon without entering the excess-risk bound as an explicit penalty. The analysis extends to the canonical blockwise step size~$1/L_i$ through a curvature-matched Lyapunov geometry that exposes a balance condition for exact block updates, and to nonconvex objectives by decoupling gradient and curvature moments via H\"older's inequality so that stability is controlled without an almost-sure curvature lower bound. Experiments on convex and nonconvex benchmarks confirm the predicted trade-off.


SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing

Yankai Yang ⋅ Yancheng Long ⋅ Wei Chen ⋅ Xingyu Lu ⋅ Hongyang Wei ⋅ Bin Wen ⋅ Fan Yang ⋅ Tingting Gao ⋅ Han Li ⋅ Shuo Yang

Recent online reinforcement learning has substantially improved image editing quality. However, existing Flow-GRPO-style methods usually rely on a single whole-image reward, which makes fine-grained editing optimization difficult. We observe that a key obstacle in image editing is this spatial uniformity assumption: a whole-image reward cannot distinguish how different spatial regions contribute to image quality. To address this issue, we propose SpatialFlow-GRPO, a training framework that makes spatial reward feedback more fine-grained. The framework converts region-aware rewards into semantic-region-level optimization signals and aligns region advantages with the corresponding latent positions during policy updates. We also train a region-aware reward model, SFReward, construct SFReward-14K with region-annotated editing samples, and introduce MultiEditBench to evaluate multi-region editing ability. On OmniGen2 and FLUX.2-klein-4B, SpatialFlow-GRPO outperforms Flow-GRPO on GEdit-Bench, ImgEdit-Bench, and MultiEditBench. The results show that SpatialFlow-GRPO can turn local feedback into more accurate update signals and improve editing quality.


Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests

Patrick Rim ⋅ Tom Long ⋅ Ekta Prashnani ⋅ Ruth Rosenholtz ⋅ Ben Boudaoud ⋅ Peter Xenopoulos ⋅ Alex Wong ⋅ Joohwan Kim ⋅ Jae-Hyun Jung

Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing benchmarks evaluate these models as black boxes, limiting their ability to identify the underlying causes of lower performance: when a model fails a spatial reasoning task, it remains difficult to ascertain whether the hurdle is perceptual, such as recognizing object boundaries, or cognitive, such as reasoning about occlusion to infer hidden geometry. We introduce Spatial-IQ, a hierarchical diagnostic framework that decomposes object counting in stacked 3D structures into 9 perceptual and cognitive sub-tasks organized by the developmental stages of human spatial cognition, with mental rotation as an additional target probe. Using NVIDIA Isaac Sim, we procedurally generated a diverse dataset of roughly 80,000 stacked 3D structures with per-task ground truth. We evaluate models across three output formats (free-response text, multiple-choice images, and image editing) alongside a human baseline. The Spatial-IQ framework shows that top-performing models often succeed at the target task (object counting) without succeeding on the lower-level sub-tasks intended to support it, and that models differ in how much of these hierarchical chains they preserve, often revealing shortcut behavior that raw target-task accuracy alone would obscure. Finally, we demonstrate that training models with chain-of-thought (CoT) supervision over our hierarchical sub-tasks, combined with reinforcement learning with verifiable rewards, significantly improves both spatial consistency across sub-tasks and target-task accuracy, supporting the value of the proposed decomposition as both a diagnostic tool and a training signal.


SPDEBench: An Extensive Benchmark for Learning Stochastic PDEs

Yuantu Zhu ⋅ Zheyan Li ⋅ Dai Shi ⋅ Luke Thompson ⋅ Oliver Nash ⋅ Jose M Rangel ⋅ Siran Li ⋅ Bingguang Chen ⋅ Rongchan Zhu ⋅ Qi Meng ⋅ Hao Ni

Stochastic Partial Differential Equations (SPDEs) driven by random noise play a central role in modeling physical processes with rough spatio-temporal dynamics, such as turbulence flows, superconductors, and quantum dynamics. Although machine learning (ML)-based surrogate models have shown promise for efficiently approximating such dynamics, progress remains limited by the lack of a unified benchmark with controlled data generation and comprehensive evaluation. This gap is particularly significant for singular SPDEs, for which benchmark datasets are largely unavailable and reliable simulation requires numerically delicate schemes based on renormalization. Moreover, subtle differences in data-generation procedures, such as noise approximation, basis choice, and the inclusion of renormalization, can significantly affect the resulting datasets and, consequently, model evaluation. We introduce SPDEBench, the first unified benchmark for ML-based SPDE learning. SPDEBench provides ready-to-use datasets for physically and mathematically significant SPDEs on 1-3D domains with periodic or Dirichlet boundary condition. Both regular and singular SPDEs are taken into consideration. SPDEBench also incorporates representative ML baselines in operator learning, together with 7 evaluation metrics, including Sobolev and distributional metrics beyond the standard $L^2$-error. Supported by SPDEBench, we conduct systematic evaluations of model accuracy, robustness, and out-of-distribution generalization under controlled data variations. Our numerical results show that SPDE-aware architectures generally achieve stronger performance than generic operator-learning baselines. These findings establish SPDEBench as a reproducible and extensible resource, paving pathway for principled benchmarking and architecture design for stochastic spatio-temporal dynamics.


Specificity-Aware Diffusion Steering via Variance-Reduced Sequential Monte Carlo

Luran Wang ⋅ Linrui Ma ⋅ Hannes Stark ⋅ Regina Barzilay

Inference-time steering enables pretrained diffusion models to satisfy new constraints without full retraining. However, specificity-aware generation is difficult: repelling samples from a negative reference distribution can also erode the positive distribution where the two overlap. The key challenge is to suppress negative mass while minimally distorting the positive distribution. We address this problem by formulating specificity-aware steering as a target-design problem and deriving a target distribution from an overlap-based objective. The resulting target keeps the desired reference distribution only in regions where it is sufficiently preferred over the undesired reference distribution, giving a likelihood-ratio interpretation of specificity. To sample from the corresponding time-dependent target path, we develop a Sequential Monte Carlo sampler with a variance-minimized local proposal. We further introduce a practical fixed-noise optimization procedure that requires only Jacobian--vector products with the desired and undesired score fields. Experiments on synthetic mixtures, class-contrastive generation, and text-to-image tasks show that the proposed method suppresses undesired regions more effectively, reduces mode shift, and improves sampling stability by decreasing the SMC weight collapse compared with negative-guidance baselines.


SPECS: Faster Test-Time Scaling through Speculative Drafts and Dynamic Switching

Mert Cemri ⋅ Nived Rajaraman ⋅ Rishabh Tiwari ⋅ Xiaoxuan Liu ⋅ Kurt Keutzer ⋅ Ion Stoica ⋅ Kannan Ramchandran ⋅ Ahmad Beirami ⋅ Ziteng Sun

Scaling test-time compute has driven the recent advances in the reasoning capabilities of large language models (LLMs). However, increased compute often comes at the expense of higher user-facing latency, directly impacting user experience. Current test-time scaling methods primarily optimize for accuracy based on total compute resources (FLOPs), often overlooking latency constraints. To address this gap, we propose SPECS, a latency-aware test-time scaling method. SPECS uses a smaller, faster model to generate multiple candidates for each step of reasoning in parallel, and evaluates these candidates using a larger model and a quality critic. We design a theoretically grounded soft verification strategy to select a high-quality candidate to continue the generation. SPECS also employs a dynamic switching mechanism to use speculative drafts only for easier steps to maintain reasoning accuracy. Empirical results on a diverse and challenging set of reasoning and alignment benchmarks show that SPECS matches or surpasses the accuracy of SOTA test-time scaling methods while reducing latency by up to ~26%. Our theoretical analysis shows that as the amount of parallel compute scales, SPECS converges to an optimum of a KL-regularized reinforcement learning problem, a common objective for aligning LLM generation given a reward signal.


SpectralKV: Redundancy-Aware KV Cache Compression via Spectral Coreset Selection

Fangming Zhao ⋅ Fulun Ye ⋅ Xiaofei Yue ⋅ Ziming Zhao ⋅ Yu Peng ⋅ Junyu Chen ⋅ Tingting Li

The key-value (KV) cache has become a dominant memory and bandwidth bottleneck for serving long-context large language models, motivating a growing body of work on KV cache compression. Most existing methods follow a common recipe: a uniform per-layer budget and attention top-$k$ token selection, but this recipe is \emph{data-oblivious} (ignoring cross-layer variation in attention concentration) and \emph{redundancy-blind} (retaining near-duplicate key--value entries). We propose \textbf{SpectralKV}, a training-free framework that addresses both issues jointly. Across layers, a lightweight concentration-based criterion redistributes a fixed global budget based on the entropy of observation-window attention, giving more slots to layers whose attention is more concentrated. Within each layer, a spectral coreset selector treats keys and values as geometric objects in a metric-aligned joint feature space and selects tokens via residual pivoting, suppressing near-duplicates while preserving geometrically distinct directions. On Llama-3.1-8B and Qwen3-8B, SpectralKV consistently outperforms seven strong baselines on 15 LongBench tasks and PG-19 perplexity at 16K and 32K contexts, with the largest gains at aggressive compression ratios down to $6.25\%$ keep. And a static-budget variant reduces peak prefill VRAM by up to $44\%$ with negligible quality loss. Code and data are available at \url{https://anonymous.4open.science/r/spectralKV-2FF7}.

Context-based offline meta-reinforcement learning (COMRL) aims to learn adaptable policies entirely from static datasets by inferring latent task representations from small contexts. Most COMRL methods rely on immediate, per-step reward signals. This can be generalized by aggregating rewards into ``bags'' and being delivered after a sequence of actions (where bag size $n=1$ is the per-step case). We formally prove that under Bagged Rewards, existing per-transition COMRL methods suffer an information-theoretic collapse ($\mathcal{O}(1/n)$), causing their task inference to degrade to behavior cloning. To solve this, we introduce SpectralMeta, an algorithm that fully decouples dynamics representation from task inference. By pre-training an energy-based spectral representation on reward-free transitions, we recast task identification as a closed-form Bayesian linear regression over bagged rewards. SpectralMeta comes with finite-sample bounds on reward recovery and policy sub-optimality. In continuous-control meta-RL benchmarks, SpectralMeta maintains robust task inference and competitive returns even when rewards are reduced to a single episode-level scalar, a regime where existing baselines fail to recover task-relevant structure.

Understanding why trained Transformers generalize well is a fundamental problem in modern machine learning theory, and complexity-based generalization bounds provide a principled way to study this question. While existing norm-based bounds for Transformers remove the explicit polynomial dependence on the hidden dimension, they typically impose fixed norm constraints specified a priori and can exhibit unfavorable exponential dependence on depth. In this paper, we derive spectrum-adaptive post hoc generalization bounds for multi-layer Transformers. Under layerwise spectral norm control, the bounds are expressed in terms of layerwise Schatten quantities of the query-key, value, and feedforward weight matrices. Since the Schatten indices need not be fixed a priori and can instead be selected after training, separately for each matrix type and layer, the bounds adaptively trade off spectral complexity against the dimension- and depth-dependent factors according to the learned singular-value profiles. Empirical comparisons of BERT-adapted proxies for the leading complexity factors suggest that the proxies induced by our bounds grow more slowly with depth and hidden dimension than the corresponding norm-based proxies. Overall, our results provide a complexity-based perspective on how the spectral structure of trained Transformers is reflected in generalization analyses.


Spike-SFT: Selective Parameter Enhancement and Fusion for Efficient Spiking Neural Networks

Xiubo Liang ⋅ Jinxing Han ⋅ Yuke Li ⋅ Haoqi Zhu ⋅ Zhenxing Li ⋅ Hongyi Duan ⋅ Yu Zhao ⋅ Hongzhi Wang

Spiking Neural Networks (SNNs) promise energy-efficient inference, but adapting large pre-trained SNNs to new tasks is expensive because BPTT memory and runtime scale with the number of simulation steps $T$. We observe that weights in pre-trained SNNs are strongly concentrated near zero, and under thresholded dynamics many such synapses are functionally silent. Based on this, we propose \textbf{Spike-SFT}, a two-stage adaptation framework. Stage 1 (\textbf{Selective Parameter Enhancement}, SPE) fine-tunes only a small-magnitude subset via masked updates with low-rank regularization, progressive re-selection, and sparse-gradient storage, reducing memory while maintaining competitive accuracy. Stage 2 (\textbf{PickIt}) fuses multiple SPE-adapted models by alignment and interference-aware delta merging with spike-statistics calibration, yielding a single merged model with zero inference-time overhead. Across several SNN backbones and benchmarks, Spike-SFT offers a favorable accuracy--cost trade-off, improving fine-tuning efficiency (time/memory) while retaining strong accuracy, and enabling zero-overhead weight-space fusion via PickIt.


Sponsored Questions and How to Auction Them

Kshipra Bhawalkar ⋅ Alexandros Psomas ⋅ Di Wang

Online platforms connect users with relevant products and services using ads. A key challenge is that a user's search query often leaves their true intent ambiguous. Typically, platforms passively predict relevance based on available signals and in some cases offer query refinements. The shift from traditional search to conversational AI provides a new approach. When a user's query is ambiguous, a Large Language Model (LLM) can proactively offer several clarifying follow-up prompts. In this paper we consider the following: what if some of these follow-up prompts can be sponsored,'' i.e., selected for their advertising potential. How should thesesuggestion slots'' be allocated? And, how does this new mechanism interact with the traditional ad auction that might follow? This paper introduces a formal model for designing and analyzing these interactive platforms. We use this model to investigate a critical engineering choice: whether it is better to jointly optimize the user interaction and the final ad auction, or to decouple them into separate mechanisms for the suggestion slots and another for the subsequent ad slot. We show that the VCG mechanism can be adopted to jointly optimize the sponsored suggestion and the ads that follow; while this mechanism is more complex, it achieves outcomes that are efficient and truthful. On the other hand, we prove that the simple-to-implement modular approach suffers from strategic inefficiency: its Price of Anarchy is unbounded, for various natural choices for the two mechanisms. While we show that this inefficiency can be mitigated through revenue redistribution, we argue that the resulting mechanism sacrifices the practical simplicity that motivated the modular design in the first place.


SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

Jian Zhu ⋅ Yuzheng Zhang ⋅ Zeyao Ma ⋅ Bohan Zhang ⋅ Armin Schoepf ⋅ Daniel Woloch ⋅ Peter Wang ⋅ Guangyu Robert Yang ⋅ Samuel Jacob ⋅ Siddharth Nagisetty ⋅ Abhiram Chundru ⋅ Jean Lin ⋅ Spencer Mateega ⋅ Jing Zhang

Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks evaluate isolated operations such as single-formula generation or local cell edits, and therefore fail to capture end-to-end workflows in realistic business settings. We introduce SpreadsheetBench 2, a workflow-level benchmark for spreadsheet agents that covers three task categories: generation, debugging, and visualization. The benchmark is constructed from authentic business data, including financial reports and corporate filings, and is annotated and validated by domain experts. The benchmark contains 321 tasks; each instance averages 11.8 worksheets and requires 593.5 cell modifications, reflecting large multi-sheet workbooks with cross-sheet dependencies. We evaluate eight frontier large language models under a unified multi-turn agent scaffold, and additionally include several LLM-based spreadsheet products as complementary baselines. Results show that current systems remain far from reliable on real-world workflows: the best model achieves 34.89% overall task accuracy, and debugging accuracy is as low as 12.00%. Trajectory analysis and a failure taxonomy further indicate that insufficient spreadsheet inspection and incorrect target-cell selection are the dominant bottlenecks. Together, these findings position SpreadsheetBench 2 as a challenging testbed for advancing reliable spreadsheet automation.


SQUEEZE: Preserving Homeomorphism and Smooth in Higher-Dimensional Flows

Wangzi Yao ⋅ Yue Sun ⋅ Rongmin Chen ⋅ Honglie Wang ⋅ Jian Zhang ⋅ Tielin Zhang

Rectified flow (RF) is motivated as optimizing the velocity field through trajectory crossings. However, this explanation contradicts prior work showing that exact crossings cannot occur, when the dimension exceeds 2 and the source and target distributions are continuous. We revisit this question and make three contributions. 1) By relaxing the geometric definition of a crossing, we prove that crossings can then occur, though their probability still drops rapidly with dimension. 2) As a consequence, we observe the homeomorphism of RF breaking down in high dimensions. This effect is hard to notice, and coloring serves as a useful indicator. 3) To address this breakdown, we propose Squeeze, which corrects trajectories within the principal subspace of transport before lifting them back to the original space, thereby restoring smoothness and homeomorphism. Experiments confirm our analysis and show that Squeeze alleviates the high-dimensional breakdown.


SRA: Spatial Reasoning Adapter via Evolving Social Interaction Graphs for Trajectory Prediction

Jaewoo Jeong ⋅ Seonkyu Song ⋅ Hyeonwoo Park ⋅ Jegyeong Cho ⋅ Yujin Bae ⋅ Giwon Lee ⋅ Daehee Park ⋅ Kuk-Jin Yoon

Predicting where multiple agents will move next is fundamental to autonomous driving, robotics, and any system that must share space with people safely. Modern stochastic predictors read off inter-agent interactions once from the observed past and keep that view fixed throughout the generative rollout, so the reciprocal geometric constraints that emerge as the joint future sharpens are not re-examined. We address this gap with SRA (Spatial Reasoning Adapter), a modular graph denoiser that plugs into existing diffusion predictors without altering their generative cores. At every reverse step, SRA reads the host's current future estimate, builds a sparse Euclidean interaction graph on the predicted trajectories, and returns a residual correction to the host's decoding path. Our adapter consists of three core components: a sparse top-$N$ neighborhood retrieved with an asymmetric query-key score, a future-conditioned relational feature recomputed on the evolving prediction, and uncertainty-weighted message passing where confident agents act as anchors for uncertain ones. We instantiate SRA on three structurally different hosts - LED, MID, and MoFlow - across four multi-agent benchmarks: SDD, NBA, Soccer, and Football. SRA improves every host on every dataset by 6--7\% on average, indicating that spatial reasoning over the model's own evolving future is broadly useful across denoising architectures. The code will be released upon acceptance.


SR-Prominence: A Crowdsourced Protocol and Dataset Suite for Perceptually-Weighted Super-Resolution Artifact Evaluation

Ivan Molodetskikh ⋅ Kirill Malyshev ⋅ Mark Mirgaleev ⋅ Nikita Zagainov ⋅ Evgeney Bogatyrev ⋅ Dmitriy Vatolin

Modern image super-resolution methods generate detailed, visually appealing results, but they often introduce visual artifacts: unnatural patterns and texture distortions that degrade perceived quality. These defects vary widely in perceptual impact—some are barely noticeable, while others are highly disturbing—yet existing detection methods treat them equally. We propose artifact prominence as an evaluative target, defined as the fraction of viewers who judge a highlighted region to contain a noticeable artifact. We design a crowdsourced annotation protocol and construct SR-Prominence, a dataset suite containing 3,935 artifact masks from DeSRA, Open Images, Urban100, and a realistic no-ground-truth Urban100-HR setting, annotated with prominence. Re-annotating DeSRA reveals that 48.2% of its in-lab binary artifacts are not noticed by a majority of viewers. Across the suite, we audit SR artifact detectors, image-quality metrics, and SR methods. We find that classical full-reference metrics, especially SSIM and DISTS, provide surprisingly strong localized prominence signals, whereas no-reference IQA methods and specialized artifact detectors often fail to generalize across datasets and reference settings. SR-Prominence is released with an objective scoring protocol that allows new metrics to be benchmarked on our suite without further crowdsourcing. Together, the data and protocols enable SR artifact evaluation to move from binary defect presence toward perceptual impact.


SSD: Shell-Guided Spherical Diffusion for Molecular Geometry Generation

Yun-Yen Chuang ⋅ Chen-Sheng Gu ⋅ Hung-Min Hsu ⋅ Kevin Lin ⋅ Ray-I Chang

Diffusion models for 3D molecular geometry almost universally adopt an isotropic Gaussian prior, whose Frobenius-norm scale $\sigma_T\sqrt{3n}$ is governed by the noise level and atom count rather than by chemistry, producing a scale mismatch whose effects grow with $n$ and yield high spatial entropy and unstable early trajectories. We introduce \textbf{Shell-guided Spherical Diffusion (SSD)}, a model-agnostic framework that replaces the Gaussian prior with a chemically scaled spherical-shell initialization and augments both the forward and reverse processes with coordinated radial attraction, short-range repulsion, and an SE(3)-equivariant correction field. This joint design of initialization and dynamics is essential: neither a shell alone nor radial fields alone reproduces the stability or accuracy of SSD. We evaluate SSD across five representative coordinate-space backbones---GeoDiff, SubGDiff, EDM, SemlaFlow-style flow matching, and MCF---each under its canonical evaluation protocol. SSD consistently improves both quality and diversity under identical training and sampling budgets, upgrading weaker backbones such as GeoDiff and EDM to match or surpass stronger diffusion-, VAE-, and flow-based baselines. SSD therefore serves as a plug-in geometric enhancement that strengthens coordinate-space molecular generation models without modifying their architectures, loss functions, or training pipelines.


Stability Regimes for Framing-Sensitive Fine-Tuning in Language Models

Victor De Lima ⋅ Jiqun Liu ⋅ Grace H Yang

AI systems increasingly shape how individuals access and evaluate information, intensifying concerns about misinformation. Computational simulation offers a controlled complement to direct empirical study, with fine-tuned LLMs enabling powerful agent-based simulations. However, fine-tuning often induces broad behavioral shifts including bias drift, degenerate heuristics, and loss of task competence, which undermine experimental validity. The central challenge is inducing localized decision shifts without global behavioral collapse. We study this through a stability analysis of supervised fine-tuning, attenuating loss on selected training slices to shift false-claim acceptance while preserving overall task competence. We introduce FrameRef, a large-scale dataset of semantically equivalent claims with controlled surface framings across five dimensions (Authoritative, Consensus, Emotional, Sensationalist, Prestige), with human validation confirming systematic framing effects on claim acceptance. We identify stability regimes in which targeted error shifts can be induced without degrading aggregate accuracy or calibration. A sequential exposure task over FrameRef shows that small framing-conditioned shifts compound into substantially different cumulative outcomes under feedback, but largely disappear when feedback is removed, confirming that interventions modify specific decision tendencies rather than global performance. We release code, LoRA adapters, and data at https://github.com/anonymousauthor2352/frameref.


Stability-Weighted Direction Regularization Disentangles Generator Shortcuts from Detection Signal

Jang Ho Choi ⋅ Seongho Kim ⋅ Jaehyun Choi ⋅ Dahye Kim ⋅ Sungwon Yi ⋅ Eunho Yang

AI-generated image detectors often generalize poorly to unseen generators because the classifier latches onto generator-specific feature directions that are predictive only on the training set. We show that simply projecting away generator-discriminative LDA directions fails: these directions also carry genuine real-vs-fake signal, so post-hoc erasure removes evidence along with shortcuts. We introduce StaR, a training-time regularizer that scores each LDA direction by leave-one-generator-out stability and penalizes the classifier only on unstable directions; stable directions are left available for classification. On ResNet-50, StaR reduces the across-seed standard deviation of the four-OOD average AUC by 24x relative to ERM (F-significant on every OOD) while preserving in-distribution-like performance; on CLIP ViT-L/14, StaR improves the four-OOD average AUC by 1.7 points over matched CLIP-ERM and 4.5 over EFFORT, with the largest gains on platform-shift benchmarks. Mechanistic analyses show that StaR rotates the classifier away from unstable generator directions while amplifying generator-discriminative structure in the features---a rotational, not erasive, effect that post-hoc projection cannot reproduce. We additionally find that the conventional argmax in-domain validation'' epoch-selection rule is sub-optimal: a deployablestop at first in-domain 0.99'' rule beats it by sim0.7 AUC points on average.


Stable Partial Order Constraints for Temporal Causal Structure Learning

Changxin Rong ⋅ Xiangyu Wang ⋅ Taiyu Ban ⋅ Yanze Gao ⋅ Xingjian Lin ⋅ Huanhuan Chen

Temporal causal structure learning aims to recover lag-aware causal relations from multivariate time series. While prior knowledge has been widely used to improve static structure learning, its integration into lag-aware structure learning remains limited, because available priors are often lag-agnostic. Recent studies have explored lag-agnostic edge presence priors, but partial order priors, which can prune the ordering space and improve structure learning, remain underexplored. A key challenge is that variables may influence each other at different lags, leading to cyclic variable-level relations where ordering is no longer naturally defined. Moreover, directly applying static differentiable formulations of partial order constraints can be unstable due to repeated-walk accumulation on cycles. To address this gap, we represent partial order priors on the summary graph and follow the natural idea of the order relation that only allows a single direction of possible causality between two variables, which is equivalent to forbidding all reversed causality, direct or indirect. This converts partial order constraints into path prohibition constraints that can be extended to cyclic summary structures. For stable differentiable characterization, we propose a decoupled partial order constraint method. The method preserves the original lag-aware structure for data fitting, imposes partial order constraints through a normalized summary representation, and captures higher-order reachability with a closed-form connectivity characterization. Experiments further demonstrate its effectiveness in nonstationary settings.


StaDy: Factorizing the World into Static and Dynamic via Likelihood Matching

Thomas Ressler-Antal ⋅ Frank Fundel ⋅ Malek Ben Alaya ⋅ Stefan Andreas Baumann ⋅ Björn Ommer

Latent action models (LAMs) aim to learn compact representations of state transitions directly from visual observations, but they suffer from a fundamental ambiguity: the latent code can encode target-state information rather than true dynamics. Existing approaches address this issue through restrictive bottlenecks, which reduce leakage at the cost of limiting expressivity. We propose StaDy, a regularization framework that instead enforces conditional informativeness: the latent code should aid prediction only when paired with the source state, while remaining uninformative about the target on its own. Concretely, we introduce a likelihood-matching objective that aligns the decoder’s predictions conditioned solely on the latent variable with its unconditional predictions. This discourages target memorization without constraining latent capacity. Experiments in video modeling show that, unlike bottleneck-based methods, StaDy scales effectively with increased latent dimensionality, reducing static content leakage while improving the representation of more complex dynamics.


Stage-Aware Dual Alignment for Covariate Shift in Graph Domain Adaptation

Hongwei Wen ⋅ Can Zhang ⋅ Haoyu He ⋅ Xintao Zhao ⋅ Hanyuan Hang ⋅ Minglong Lei

The performance of Graph Domain Adaptation (GDA) is fundamentally limited by covariate-driven discrepancies between source and target graphs, captured by Covariate Shift (CS) in the joint feature-structure space beyond label-related formulations. We show that CS induces bias at two non-interchangeable stages of Graph Neural Network (GNN) computation: Feature Shift (FS) distorts representations before message passing, while Feature-Conditional Structure Shift (FCSS) biases propagation, rendering single-stage alignment intrinsically insufficient. Based on this, we propose Dual Alignment for Covariate Shift (DACS), a stage-aware GDA framework that aligns discrepancies at their source. DACS mitigates FS via adversarial feature alignment and addresses FCSS through layer-wise reweighting to correct propagation bias, followed by final adversarial alignment for residual mismatch. This design follows directly from the stage-wise structure of CS rather than heuristic combinations of alignment modules. We further show that representation and propagation biases correspond to distinct components of target-domain error that cannot be eliminated in isolation. Experiments on synthetic and real-world benchmarks demonstrate that DACS consistently outperforms prior methods, especially under complex and coupled distribution shifts.


StakeBench: Evaluating Language Understanding Grounded in Market Commitment

Yunhua Pei ⋅ Jingyu Hu ⋅ Yiwei Shi ⋅ Hongnan Ma ⋅ Weiru Liu ⋅ John Cartlidge

Existing financial NLP benchmarks often rely on labels supplied by outside observers, measuring how language is perceived rather than what speakers have committed to in the market. We introduce \textbf{StakeBench}, an evaluation framework for language understanding grounded in market commitment. StakeBench links 560,876 comments from 2,261 resolved markets to verified position, action, and market-odds records across Polymarket and Manifold. Supervision is derived from observable market behavior. Position sides, post-comment trading actions, and market-odds trajectories replace human annotation. Four diagnostic tasks test whether models detect market commitment, identify the revealed side, anticipate future action, and perform collective odds projection. Three commitment-aware metrics measure alignment with revealed preferences rather than perceived sentiment. Validity audits and explicit interpretation boundaries help distinguish observable commitment signals from latent belief and causal market-odds impact. Across 15 LLMs and 18 topics and platform settings, models partially recover position-side signals, with Directed Accuracy from 0.506 to 0.599, but show structural failures on later tasks. Ten of the fifteen models collapse to one or two action labels in future action anticipation, and no model consistently improves on the naive odds-direction baseline in collective odds projection. Model scale is not correlated with performance, finance-domain tuning does not improve revealed-side identification, and platform incentives strongly shape higher-order results. StakeBench is packaged with evaluation code and dataset under CC-BY~4.0.


STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens

Shiqi Liu ⋅ Zeyu He ⋅ Guojian Zhan ⋅ Letian Tao ⋅ Zhilong Zheng ⋅ Jiang Wu ⋅ yinuo Wang ⋅ Yang Guan ⋅ Kehua Sheng ⋅ Bo Zhang ⋅ Keqiang Li ⋅ Jingliang Duan ⋅ Shengbo Eben Li

Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic techniques such as entropy regularization and reweighting to maintain stability. In practice, they often suffer from late-stage performance collapse, leading to degraded reasoning quality and unstable training. We identify a key factor behind this instability: a small fraction of tokens, termed spurious tokens (around 0.01\%), which contribute little to the reasoning outcome but receive disproportionately amplified gradient updates due to inheriting the full sequence-level reward. We present a unified framework for evaluating token-level optimization impacts across spurious risk, gradient norms, and entropy changes. Building the analysis of token characteristics that severely disrupt optimization, we propose the Silencing Spurious Tokens (S2T) mechanism to efficiently suppress their gradient perturbations. Incorporating this mechanism into a group-based objective, we propose Spurious-Token-Aware Policy Optimization (STAPO), which promotes stable and effective large-scale model refinement. Across six mathematical benchmarks and three model scales, STAPO consistently demonstrates superior entropy stability and achieves average accuracy improvements over baselines including GRPO, 20-Entropy, and JustRL under two widely-adopted evaluation settings.


Static Recovery Is Not Dynamic Stability: Dynamics-Aware Benchmarking of Protein Motif Scaffolding

Emile de Bruyn ⋅ Moritz Schäffler ⋅ Jannik Schneider ⋅ Sebastian R Schmidt ⋅ Chanwoong Hwang ⋅ Aleena Siji ⋅ Utkarsh Upadhyay ⋅ Alexander Schug ⋅ Stefan Bauer ⋅ Artur Yakimovich ⋅ Stefan Kesselheim ⋅ Alina Bazarova

Protein motif scaffolding benchmarks aim to measure whether generative models can build scaffolds around functionally important residue sets. Their conclusions depend on two coupled choices: which motifs are tested and which criteria define success. Existing benchmarks typically use small hand-curated motif sets and static refolding-based criteria such as motif recovery, self-consistency, and structural novelty. However, protein function often depends on conformational ensembles and transitions, raising the question of whether static refolding success translates to dynamic fidelity of motifs and scaffolds. Here, we address both benchmark scope and success criteria. We construct the largest systematically derived benchmark of structurally conserved functional motifs from PROSITE-linked experimental structures, yielding 220 cases from 174 PROSITE motif-pattern entries of varying conformations. We evaluate six recent scaffold-generation models and find that performance separates into target coverage and unique-solution yield: RFdiffusion and Proteina solve more targets, whereas ESM3 produces more unique successful scaffolds on the targets they solve. We then perform all-atom molecular dynamics simulations for a set of statically solved designs, comparing motif drift, scaffold drift, motif-contact preservation, and motif-dihedral free-energy landscapes against reference simulations. Of 55 statically solved model–problem pairs, 31 satisfy our MD-supported criterion and only 8 pass all checks. Refolding success is therefore useful evidence for local motif recovery, but not sufficient evidence for dynamical fidelity of the full scaffold, exposing misalignment between static benchmark success and downstream dynamical behavior. The mismatch is motif-dependent: compact, self-contained motifs are often MD-supported, whereas interface-like or context-dependent motifs fail more often. We show that motif-scaffolding models have varying performance on three complementary axes: target coverage, unique-solution yield, and MD-supported dynamical fidelity.

A statistical field theory is introduced for finite state and action Markov decision processes with unknown parameters, in a Bayesian setting. The Bellman equation, for policy evaluation and the optimal value function in finite and discounted infinite horizon problems, is studied as a disordered interacting dynamical system. The Markov decision process transition probabilities and mean-rewards are interpreted as quenched random variables and the value functions, or the iterates of the Bellman equation, are deterministic variables that evolve dynamically. The posterior over value functions is then equivalent to the quenched average of the Fourier inverse of the Martin-Siggia-Rose-De Dominicis-Janssen generating function. The formalism enables the use of methods from field theory to compute posterior moments of value functions. The paper presents two such methods, corresponding to two distinct asymptotic limits. First, the classical approximation is applied, corresponding to the asymptotic data limit. This approximation recovers so-called plug-in estimators for the mean of the value functions. Second, a dynamic mean field theory is derived, showing that under certain assumptions the state-action values are statistically independent across state-action pairs in the asymptotic state space limit. The state-action value statistics can be computed from a set of self-consistent mean field equations, which we call dynamic mean field programming (DMFP). Collectively, the results provide analytic insight into the structure of model uncertainty in Markov decision processes, and pave the way toward more advanced field theoretic techniques and applications to planning and reinforcement learning problems.


Stay Fair! Ensuring Group Fairness in Diffusion Models Across Guidance Scales

Myeongsoo Kim ⋅ Eunji Kim ⋅ Minwoo Chae ⋅ Sangwoo Mo

Diffusion models steer conditional generation with a tunable guidance scale, which users routinely adjust to trade off prompt alignment and diversity. However, these models often reflect and amplify social biases, whose mitigation has been a long-standing concern. Current debiasing techniques are often optimized for a single guidance scale, leaving them vulnerable to fairness degradation when that scale is adjusted. We trace this behavior to a previously overlooked source by decomposing total bias into two components: a model bias and a guidance bias. While prior work primarily targets the former, we show that the guidance bias grows monotonically with the guidance scale, eventually dominating the high-guidance regimes users prefer. To address this, we extend Strong Demographic Parity to guidance and derive a condition under which the target guided distribution retains its group ratio across guidance scales. We propose StayFair, which leverages this condition to design fair guidance algorithms in both regimes. For classifier guidance, it equalizes the classifier's output distributions across groups; for classifier-free guidance, it shifts the null embedding by a prompt-dependent offset. Since StayFair modifies only guidance, it is orthogonal to model debiasing and can be layered onto existing fair diffusion models to extend their fairness across guidance scales. Across class-conditional and text-to-image generation, StayFair decouples fairness from the guidance scale without sacrificing image quality.

Activation steering is a practical post-training model alignment technique to enhance the utility of Large Language Models (LLMs). Prior to deploying a model as a service, developers can steer a pre-trained model toward specific behavioral objectives, such as better truthfulness, or reasoning ability, without the need for retraining. Conceptually, these methods implement behavior control through hidden-state interventions, without changing the underlying model parameters. However, this capability unintentionally introduces critical and under-explored safety risks. We identify a phenomenon termed Steering Externalities, where steering vectors derived from benign datasets—such as reducing harmless refusals, improving structured-output following, truthfulness, and reasoning performance—inadvertently erode safety guardrails. Experiments reveal that these interventions act as a force multiplier, creating new vulnerabilities to jailbreaks and increasing attack success rates to over 80\% on standard benchmarks by bypassing the initial safety alignment. Ultimately, our results expose a critical blind spot in deployment: benign activation steering can erode the ``safety margin,'' rendering models more vulnerable to black-box attacks and indicating that inference-time utility improvements must be rigorously audited for unintended safety externalities.


Steering Frozen LLMs: Adaptive Social Alignment via Online Prompt Routing

Zeyu Zhang ⋅ Xiangxiang Dai ⋅ Ziyi Han ⋅ Xutong Liu ⋅ John C. S. Lui

Large language models (LLMs) are typically governed by post-training alignment (e.g., RLHF or DPO), which yields a largely static policy during deployment and inference. However, real-world safety is a full-lifecycle problem: static defenses degrade against evolving jailbreak behaviors, and fixed weights cannot adapt to pluralistic, time-varying safety norms. This motivates inference-time governance that steers behavior without costly retraining. To address this, we introduce the Consensus Clustering LinUCB Bandit (CCLUB), a unified framework for adaptive social alignment via system-prompt routing. CCLUB employs a conservative consensus clustering mechanism: it pools data only within the intersection of utility and safety similarity graphs, effectively preventing unsafe generalization across semantically proximal but risk-divergent contexts. Our theoretical analysis gives an expected regret bound $O\left(d\log T/(p_{\min}\gamma^2\lambda_x) + d\sqrt{MT}\log T\right)$, separating the logarithmic cost of identifying safe consensus clusters from the dominant cluster-level exploitation term. Experiments show that CCLUB improves the safety--utility trade-off and offline deployment gap, and improves cumulative reward by 10.98\% over the strongest non-prototype baseline.


Steering Visual Generation in Unified Multimodal Models with Understanding Supervision

Zeyu Liu ⋅ Zanlin Ni ⋅ Yang Yue ⋅ Cheng Da ⋅ Huan Yang ⋅ Di ZHANG ⋅ Kun Gai ⋅ Gao Huang

Unified multimodal models are envisioned to bridge the gap between understanding and generation. Yet, to achieve competitive performance, state-of-the-art models adopt largely decoupled understanding and generation components. This design, while effective for individual tasks, weakens the connection required for mutual enhancement, leaving the potential synergy empirically uncertain. We propose to explicitly restore this synergy by introducing **Understanding-O}riented Post-Training (UNO), a lightweight framework that treats understanding not only as a distinct task, but also a direct supervisory signal to steer generative representations. By incorporating objectives that encode semantic abstraction (captioning) and structural details (visual regression), we enable effective gradient flow from understanding to generation. Extensive experiments on image generation and editing demonstrate that understanding can serve as an effective catalyst for generation.

Recovering executable CAD programs from 3D meshes is fundamentally challenging due to the long-horizon, compositional nature of CAD construction and the need for precise estimation of both discrete operations and continuous parameters. Existing learning-based methods primarily treat mesh-to-CAD reconstruction as one-shot sequence prediction and are largely limited to simple sketch-extrude pipelines, restricting operation diversity and preventing the use of intermediate geometric feedback, with no mechanism for refining generated programs. To address this limitation, we introduce StepCAD, a generative optimization approach that performs step-by-step, geometry-conditioned program synthesis followed by explicit refinement in program space. Given an input mesh, StepCAD first predicts construction sequences using a state-conditioned CAD policy conditioned on both target and intermediate geometry, then refines them through IoU-guided tree search over local program edits. To address the limited operation diversity in prior work, we further introduce ARCADE-1.5M, a large-scale dataset of 1.5M executable CAD programs with diverse operations, long-horizon sequences up to 150+ operations, and 12.5M intermediate state-action transitions. Experiments on multiple CAD reconstruction benchmarks show that StepCAD achieves state-of-the-art reconstruction accuracy, validity, and robustness, with up to 87.2\% relative IoU improvement over the best prior method and increasingly larger gains as shape complexity increases.


StereoSplat: Metric-Scale Novel View Synthesis via Stereo-Grounded Gaussian Splatting

Vladimir Yugay ⋅ Denis Rozumny ⋅ Theo Gevers ⋅ Elias Vansteenkiste ⋅ Martin R. Oswald ⋅ Diogo Luvizon

We present StereoSplat, a feed-forward 3D Gaussian Splatting architecture designed explicitly for stereo videos to achieve reliable, metric-scale novel-view synthesis. Unlike monocular approaches that neglect the nature of binocular pairs, StereoSplat adopts a native stereo architecture that extracts dense multi-scale latent features and disparity from a foundation stereo backbone and feeds them into a dedicated Gaussian decoder. Our decoder utilizes cross-view feature fusion and gated spatial refinement to directly translate these rich representations into per-pixel 3D Gaussians. To resolve the high redundancy of per-pixel predictions across overlapping views, we propose a learnable depth-adaptive aggregation mechanism that clusters and aggregates primitives in inverse-depth space while preserving fine-grained details. To support robust training and generalization, we introduce XRStereo, a large-scale synthetic dataset of stereo video featuring various rig configurations tailored to match the physical hardware of real-world platforms. Trained exclusively on synthetic data and evaluated across challenging benchmarks, StereoSplat yields substantial improvements in the depth accuracy of rendered novel views over established baselines while maintaining high-fidelity novel view synthesis. By overcoming prior geometric limitations, our framework enables robust zero-shot transfer to real-world domains, offering a scalable geometric foundation for spatial computing in mixed reality, autonomous driving, and robotics.


Stop or Restart? Principled Inference Control for Large Reasoning Models via the Pandora's Box

Xiaosong Yuan ⋅ Xiaofeng Zhang ⋅ Yijia Zhang ⋅ Renchu Guan ⋅ Ying Wang

While large reasoning models (LRMs) can solve complex tasks via generating extended Chain-of-Thought traces (Long-CoT), longer traces often increase confidence without increasing correctness. In this work, we formalize such failure via a \emph{self-conditioned evidence} model: generated tokens provide information about the model's current working hypothesis rather than the ground-truth answer. This yields a simple diagnostic principle: entropy, fluency, and answer stability can certify within-trajectory self-consistency while remaining causally disconnected from correctness. Building on these findings, we propose \textbf{Faithful Pandora Inference (FPI)}, a minimal training-free controller that keeps the base LRM frozen and fits only a small calibration map on held-out data. At each checkpoint, FPI estimates calibrated correctness $r_t$ and the marginal evidence gain $m_t$ of continuing the current trajectory. It stops only when $r_t$ reaches a user-specified reliability target; if $r_t$ remains insufficient and $m_t$ falls below cost, it reallocates the remaining budget to a diversified restart. Across six reasoning benchmarks and multiple DeepSeek-R1 distilled models, FPI improves the accuracy-efficiency frontier, reduces expected calibration error, and improves selective accuracy on high-confidence outputs. Ablations on confidently wrong trajectories show that calibrated stopping and marginal-gain restarts address different failure modes and that the gains are not explained by additional samples alone.


StoSplat: Ray-Aligned Stochastic Preconditioning for Feed-Forward 3D Gaussian Splatting

Qibin Hu ⋅ Zhiheng Fu ⋅ Longguang Wang ⋅ Jisheng Dang ⋅ Hanyun Wang ⋅ Yulan Guo

Feed-forward 3D Gaussian splatting enables single-pass reconstruction from multi-view images, but widely-applied voxel-aligned pipelines usually produce sharp and anisotropically ill-conditioned Gaussian centers, leading to floaters and unstable geometry in sparse-view settings. To remedy this, we introduce StoSplat, a training-time stochastic preconditioning framework that perturbs predicted Gaussian centers with annealed, ray-aligned anisotropic noise. This perturbation smooths the expected rendering objective in the geometric variable where instability arises, while leaving covariance, opacity, color, and the inference architecture unchanged. StoSplat introduces negligible additional computation during training without any computational overhead during inference. Experiments on widely used benchmarks including RealEstate10K, ScanNet, and ACID demonstrate that StoSplat achieves state-of-the-art performance, while producing more stable geometry with fewer floaters and boundary artifacts, which demonstrates the effectiveness of our method in improving the performance of existing feed-forward 3D reconstruction without architectural changes.


Stranger Things: When Objects Appear Without Their Typical Neighbours

Siddhartha Gairola ⋅ Jiahao Xie ⋅ Anna Kukleva ⋅ Francesco Locatello ⋅ Bernt Schiele

Visual context is a powerful cue for object recognition: scenes inform what, where, and at what scale objects are likely to appear. Yet over-reliance on context becomes a failure mode when models predict objects based on their typical neighbours rather than the objects' own visual features, a phenomenon we refer to as typicality bias. While this issue has been extensively studied in the pre-Transformer era, it has often been assumed that architectural and scale improvements would naturally solve it. Despite recent attempts in isolated settings, a systematic diagnostic across the modern vision stack remains absent. In this work, we provide such a diagnostic framework. Specifically, we measure typicality bias via normalized pointwise mutual information (NPMI) over object pairs, and introduce STRANGE-Bench (Synthetic Typicality-RANked Grounded Evaluation Benchmark): a photorealistic, layout-controlled benchmark spanning 11 typicality tiers from typical to never-co-occurring object pairs. Across \five model families (specialist detectors, open-vocabulary detectors, self-supervised, vision-language, and multimodal large language models) and three recognition tasks}(supervised detection, open-vocabulary detection, and classification), we find the bias to be highly prevalent: gaps of 5--11\% persist when objects appear outside their typical company, and they do not consistently shrink with scale. This persistence motivates a complementary question: how to mitigate it? Data augmentation is a natural lever, since it directly shapes the training distribution that produces the bias; yet existing augmentations are designed for visual diversity, not for breaking co-occurrence shortcuts as they sample from the very distribution that produces typicality bias, preserving rather than perturbing it. We therefore propose Anchor-Atypicality Sampling (AAS), a simple context-aware variant of paste-based augmentations that uses the NPMI signal to deliberately tilt the training-time co-occurrence distribution toward atypical partners. AAS reduces the typicality gap by up to 3.3 pp on natural-image COCO splits and 17.4 pp on synthetic STRANGE-Bench across three representative model families (specialist detection, open-vocabulary detection, self-supervised classification), without sacrificing overall performance.


StraTA: Incentivizing Agentic Reinforcement Learning with Strategic Trajectory Abstraction

Xiangyuan Xue ⋅ Yifan Zhou ⋅ ZiDong Wang ⋅ Shengji Tang ⋅ Philip Torr ⋅ Wanli Ouyang ⋅ LEI BAI ⋅ Zhenfei Yin

Large language models (LLMs) are increasingly used as interactive agents, but optimizing them for long-horizon decision making remains difficult because current methods are largely purely reactive, which weakens both exploration and credit assignment over extended trajectories. In this work, we present Strategic Trajectory Abstraction (StraTA), a simple framework that introduces an explicit trajectory-level strategy into agentic reinforcement learning (RL). StraTA samples a compact strategy from the initial task state, conditions subsequent actions on that strategy, and trains strategy generation and action execution jointly with a hierarchical GRPO-style rollout design, further enhanced by diverse strategy rollout and critical self-judgment. Experiments on ALFWorld, WebShop, and SciWorld show that StraTA consistently improves both sample efficiency and final performance over strong baselines. StraTA reaches success rates of 93.1% on ALFWorld and 84.2% on WebShop, surpassing the strongest baselines by 2.3% and 11.4% respectively. On SciWorld, StraTA attains a 63.5% overall score, outperforming frontier closed-source models by 6.1% and prior RL baselines by 6.5%.


Strategic Feature Selection and Regularization

Jivat Neet Kaur ⋅ Pratik Patil ⋅ Divya Shanmugam ⋅ Emma Pierson ⋅ Michael Jordan ⋅ Nika Haghtalab ⋅ Meena Jagadeesan ⋅ Ahmed Alaa ⋅ Serena Wang

When algorithmic predictors inform resource allocation in high-stakes domains such as healthcare, these predictors must account for strategic manipulation of input features. The typical solution is to redesign the predictor itself to explicitly account for strategic interactions. In practice, however, decision makers are often constrained to adjusting coarser levers within existing prediction pipelines. For example, healthcare organizations often select which features to exclude based on perceived manipulability, while using standard regularization procedures to shrink the coefficients of retained features. In this work, we initiate a formal study of strategic classification through feature selection and its interaction with ridge regularization. Our main finding is that excluding features based on manipulability alone is generally suboptimal. We provide a fine-grained characterization of the performance of a feature subset under optimal regularization, yielding new insights for policy design. Motivated by this characterization, we develop a practical algorithm for jointly choosing the feature set and the level of ridge regularization. Through a real-world case study on a healthcare payments benchmark, we illustrate how our algorithm can guide the design of coarse policy levers in practice. Our results provide a principled, practical framework for mitigating the effects of strategic behavior in algorithmic decision-making systems.


StreamMind: Dynamic Streaming Cognition for Online Video Understanding

Zhuojie Wu ⋅ Xin Shen ⋅ Chenxi Miao ⋅ Weikang Li ⋅ Liwei Qian ⋅ Xin Pei ⋅ Xin Yu

Online video large language models (VLLMs) have demonstrated remarkable capabilities in real-time streaming video understanding and proactive human-AI interaction. However, existing methods often follow a static perception paradigm, processing video streams with rigid temporal units while overlooking the inherent temporal non-uniformity of streaming videos. This oversight results in flickering responses and fragmented context. In this paper, We propose \textbf{StreamMind}, a dynamic streaming cognition framework that organizes both online reasoning and memory around temporally coherent Group-of-Pictures (GoP) segments. Specifically, StreamMind introduces two key components. First, StreamMind introduces a \textbf{Dynamic Thinking (DynThink)} module, which accumulates evidence within each GoP and triggers reasoning only at GoP boundaries, thereby reducing premature decisions and improving response stability. Second, to support long-horizon streaming understanding, we propose a \textbf{Brain-Eye Synergy Memory (BESM)} module, which preserves key-frame visual tokens as high-fidelity anchors while compressing the remaining frames into compact thinking tokens. This design maintains fine-grained visual evidence together with reasoning continuity, while reducing redundant storage. Extensive experiments on streaming and offline video benchmarks demonstrate that StreamMind achieves state-of-the-art online video understanding performance, reaching 63.4\% on OVO-Bench and 80.2\% on StreamingBench, while maintaining strong generalization to offline long-form video reasoning.


Stress Testing Chain-of-Thought Monitoring Against Covert Misalignment

Yao Huang ⋅ Yifan Wang ⋅ Yitong Sun ⋅ Yichi Zhang ⋅ Yinpeng Dong

As LLM agents advance rapidly in autonomy, the risk of hidden objectives that bypass safety constraints becomes increasingly plausible. Chain-of-thought (CoT) monitoring offers a key control mechanism by reading the agent's reasoning and flagging misalignment before it materializes into action, yet its reliability depends on the realism of the red-team it is stress-tested against. Existing red-teaming efforts apply uniform attack templates that treat the side task as an exogenous insertion, yielding abrupt anomalies any reasonable monitor can flag while leaving the subtler, contextually grounded evasions of a capable agent largely unprobed. We therefore ask: How robust is CoT monitoring against agents that strategically conceal misalignment? To answer this, we introduce TraceMRT (Trace-level Monitor Red-Teaming), an automated red-teaming framework that reframes the side task as a coherent sub-goal of the primary task and searches trace-shaping strategies whose induced trajectories complete the misaligned objective while reading as natural execution. Through black-box attacks on state-of-the-art monitors (e.g., GPT-5.2, and Claude-4.5), we reveal substantial vulnerabilities even in top models. To better counter these vulnerabilities, we further improve monitoring scaffolding with Bayesian Trajectory Monitoring, integrating global context and local evidence to better detect covert misalignment. Empirically, the method significantly consolidates the monitor with 0.91 AUC and 0.74 TPR@FPR=0.05.


STRIDE: Automated Evaluation of Text-to-Trajectory Alignment across Diverse Contexts

Wanchun Ni ⋅ Tao Qi ⋅ Leonel Aguilar ⋅ Jiugeng Sun ⋅ Marlene Wagner ⋅ Verena Zimmermann ⋅ Mennatallah El-Assady

Language-conditioned trajectory generation is here, but its evaluation has not kept pace. Existing pedestrian trajectory metrics compare trajectories with real-world human data. This does not scale to text-to-trajectory generation across diverse contexts, as collecting human trajectories for every scenario is costly and infeasible. Moreover, pedestrian behavior is heterogeneous and context-dependent, with no single metric as the correct answer, and current evaluation frameworks are not transferable to this domain. These challenges make scalable, reliable evaluation difficult. We introduce STRIDE, the first framework for evaluating context alignment between scenario descriptions and pedestrian trajectories. STRIDE addresses these challenges through three design choices. First, we derive our VRDST evaluation protocol from sociological theories to define a complete evaluation space. Second, it decomposes high-level context into scenario-adaptive behavioral questions. Third, every question is resolved against a deterministic measurement tool library that yields reproducible answers. Together, STRIDE enables complete, verifiable, automated, and scalable evaluation across diverse contexts without requiring human trajectory data. We instantiate STRIDE in the crowd domain as STRIDE-Bench, comprising 1K scenarios, 6K behavioral questions, and 11K measurements with calibrated expected answers across 30 real-world maps. Comprehensive human validations show that STRIDE-Bench is consistent with human behavior and judgment, achieving 80% human agreement. We further evaluate several text-to-trajectory models, finding limited context-alignment capability and persistent challenges in fine-grained context conditioning. We believe that the STRIDE framework provides a first step toward principled evaluation of context-aligned pedestrian trajectory generation. Code: https://anonymous.4open.science/r/STRIDE-Bench-7006 Dataset: https://huggingface.co/datasets/anonymous1ads34/STRIDE-Bench


Strong Post-Training from Permissive, Reasoning-Dominant, Web-Scale Pretraining

Harsh Raj ⋅ Ali Elganzory ⋅ Marianna Nezhurina ⋅ Victor May ⋅ Van Khue Nguyen ⋅ David Salinas ⋅ Huu Nguyen ⋅ Jenia Jitsev

Permissive, auditable pretraining corpora enable lawful, responsible and transparent open foundation model research and development, facilitating standardization and common progress. Recent work such as MixtureVitae has shown that permissive corpora can also yield strongly competitive base models, but no controlled comparison has measured whether they still remain competitive after post-training. We post-train 1.7B base models pretrained on 300B of MixtureVitae permissive tokens and compare them against strong compute and token matched non-permissive reference baselines (Nemotron-CC-HQ, FineWeb-Edu and DCLM) and against open weights baselines trained at substantially larger pretraining compute (SmolLM2 1.7B, Qwen2.5 1.5B, Qwen3 1.7B). All models are post-trained with the same pipeline: Tulu3 supervised fine-tuning and direct preference optimization, followed by OpenThoughts3 reasoning training. Post-trained MixtureVitae models match or outperform the matched non-permissive reference on reasoning and instruction-following benchmarks, and maintain solid general language understanding performance. They also remain competitive with strong open weights baselines, which use over an order of magnitude more pretraining compute. Ablations confirm that the substantial reasoning and instruction subset of MixtureVitae drives this result: removing it drastically weakens the post-trained model under the same recipe. Together, these findings show that permissive pretraining does not preclude strong post-training, strengthening the case for legally safe and reproducible open foundation model research and development at frontier performance levels. Data, models, and code to reproduce the experiments will be open-sourced.

We propose Structural Self-Teaching (SST), a novel framework that enhances compositional image generation in unified multimodal models (UMMs) by leveraging their internal understanding path as a source of structural supervision. The motivation of this work is the observation that the generation path of UMMs often struggles with dense compositional prompts, leading to missing objects, attribute binding failures, and spatial errors. To bridge this gap, our key idea is to convert the phrase-level grounding inherent in the understanding path into training-time structural signals. Specifically, we extract phrase-conditioned support masks from internal attention maps derived from the understanding path and employ them through two complementary mechanisms: 1) phrase-local contrastive alignment to synchronize generated features with their corresponding phrase-level grounding, and 2) a lightweight structural scaffold for spatial modulation of intermediate generation states. By employing attenuated scaffold dropout during training and removing the scaffold branch at inference, our approach improves compositional fidelity across object, attribute, and spatial constraints. Crucially, the framework requires no external structural annotations for training or additional inputs at test-time, preserving the original generation interface. Extensive experiments across multiple UMMs demonstrate that our method consistently improves compositional fidelity while maintaining the efficiency of a lightweight tuning setup.

Encoding classical data into quantum states is a central bottleneck in quantum machine learning: many widely used encodings are circuit-inefficient, requiring deep circuits and substantial quantum resources, which limits scalability on quantum hardware. In this work, we propose TNQE, a circuit-efficient quantum data encoding framework built on structured unitary tensor network (TN) representations. TNQE first represents each classical input via a TN decomposition and then compiles the resulting tensor cores into an encoding circuit through two complementary core-to-circuit strategies. To make this compilation trainable while respecting the unitary nature of quantum operations, we introduce a unitary-aware constraint that parameterizes TN cores as learnable block unitaries, enabling them to be directly optimized and directly encoded as quantum operators. The proposed TNQE framework enables explicit control over circuit depth and qubit resources, allowing the construction of shallow, resource-efficient circuits. Across a range of benchmarks, TNQE achieves encoding circuits as shallow as $0.04\times$ the depth of amplitude encoding, while naturally scaling to high-resolution images ($256 \times 256$) and demonstrating practical feasibility on real quantum hardware.


Structure-Semantic Guided Closed-Loop Medical Anomaly Detection via Multi-Agent Collaboration

Chunjing Xiao ⋅ Yunxiao Dai ⋅ Yuwan Fu ⋅ Yucong Wang ⋅ Chong Tang ⋅ Fan Zhou ⋅ Ying Ma

Unsupervised medical anomaly detection aims to identify images or regions that deviate from the learned normal distribution. Reconstruction-based methods are a dominant paradigm, generating a normal-looking reference for each test image and detecting anomalies via input--reconstruction discrepancy. However, most existing methods follow a one-shot reconstruction-and-comparison pipeline, making anomaly scoring vulnerable to reconstruction failures: abnormal regions may be preserved, while normal anatomical structures may be distorted. Moreover, they often lack coordinated control from structural and semantic normality, which can lead to anatomical inconsistency and semantic drift. We propose S2Agent, a structure- and semantic-guided multi-agent framework that reformulates medical anomaly detection as feedback-driven normality reconstruction. S2Agent decomposes the process into three collaborative agents. The Planner derives sample-specific matched-tree structural priors and normal-only semantic claims from normality knowledge. The Reconstructor generates a normality-oriented reference image under their joint guidance. The Detector verifies the reconstruction in matched-tree and semantic-claim spaces, returning feedback to refine structural weights and semantic constraints for the next round. Through this closed loop, S2Agent iteratively corrects structural deviation and semantic drift, suppresses abnormal or unsupported content, and improves the reliability of anomaly scoring. Experiments on three public medical benchmarks demonstrate the effectiveness of the proposed closed-loop guidance.


SudoBench: A Contextual Authorization Benchmark for LLM Agents

Vincent Siu ⋅ Tianneng Shi ⋅ Shangding Gu ⋅ Zhun Wang ⋅ Dawn Song ⋅ Chenguang Wang

Secure LLM agents should take actions for authorized users in authorized environments. However, current agents' (e.g., OpenClaw) runtimes expose neither user nor environment authorization and existing agent security benchmarks similarly evaluate models without relevant authorization context. We propose SudoBench, a benchmark of 135 paired scenarios in a synthetic consulting universe with a static access control list, spanning user authorization (confused deputy), environment authorization (source-label indirect prompt injection), and both (scope-check indirect prompt injection). For each pair of scenarios, the authorization context is flipped such that an identical prompt should be appropriately complied with or refused. We score each pair as a joint pass requiring the model to produce the correct outcome on both paired scenarios. Across ten frontier closed- and open-weight LLMs, joint pass rate without authorization context sits below 16\% on every model. Adding user and environment authorization raises joint pass to as high as 67--73\%. Even with this authorization context, frontier models remain far from reliably handling contextual authorization, framing authorization context as an important research agenda for agent security.


SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework

Tianshu Wu ⋅ Xiangqi Kong ⋅ Yue Chen ⋅ Qize Yu ⋅ Hang Ye ⋅ Jia Li ⋅ Yizhou Wang ⋅ Hao Dong

Building humanoid robots that perform generalizable whole-body loco-manipulation in the real world remains a fundamental challenge: existing approaches either rely on heavy task-specific reward engineering, rigidly replay reference motions that fail to generalize, or depend on costly teleoperation that limits scalability. While human videos capture diverse human behaviors, the motion priors inferred from them are inherently imperfect, suffering from occlusion, contact artifacts, and retargeting errors that render them unsuitable for direct policy learning. To this end, we present SUGAR, a data-driven framework that converts diverse human videos into deployable humanoid loco-manipulation skills, without any task-specific reward engineering or reference-motion conditioning at inference. SUGAR proceeds in three coupled stages: First, a fully automated pipeline extracts scalable kinematic interaction priors including human-object motion trajectories and contact labels from diverse human videos. Second, a privileged physics-based refiner utilizes a unified mimic-style reward and a progressive state pool to transform imperfect kinematic interaction priors into physically feasible, high-fidelity skills. Third, the refined skills are distilled into a hierarchical policy that comprises a task-guided planner and a command tracker for autonomous task execution. We evaluate our method on six representative loco-manipulation tasks in both simulation and real-world humanoid hardware. SUGAR substantially outperforms reference-tracking baselines, and its performance scales clearly with the amount of human video data. It also achieves zero-shot real-world transfer with reliable closed-loop execution, autonomous failure recovery and stable long-horizon performance under external perturbations. Project Page: https://sugar-humanoid.github.io/


Superplatforms Are Strategically Compelled to Counteract AI Agents

Jianghao Lin ⋅ Jiachen Zhu ⋅ Zheli Zhou ⋅ Yunjia Xi ⋅ Weiwen Liu ⋅ Yong Yu ⋅ Weinan Zhang

This position paper argues that superplatforms are strategically compelled to counteract, constrain, or even disrupt AI agents as these agents emerge as competing digital gatekeepers. Superplatforms have built their business models on controlling user attention through targeted advertising, algorithmic content curation, and centralized access to digital services. However, LLM-driven AI agents threaten this model by acting on behalf of users, bypassing platform-controlled interfaces, reducing exposure to advertisements, and potentially becoming the new entrance for digital traffic. Drawing on gatekeeping theory, we analyze why this shift creates a structural conflict between superplatforms and AI agents: whoever controls the user’s primary interface controls information flow, user data, and monetization opportunities. We then examine why common responses, such as proprietary agents and API gating, are insufficient against general-purpose and GUI-based agents, motivating the possibility of more proactive countermeasures. Finally, we outline a taxonomy of such platform-initiated countermeasures and identify the technical challenges they raise. We do not advocate for adversarial disruption; rather, our goal is to surface this emerging tension early and encourage research toward open, collaborative, and user-centric agent ecosystems.


Sup-Norm Error under Proportional Asymptotics: Phase Transitions under Linear Regression

Lin Liu ⋅ Debarghya Mukherjee ⋅ Rajarshi Mukherjee ⋅ Zixiao Jolene Wang

This paper characterizes the sharp asymptotics of the $\ell_{\infty}$-norm estimation error for estimating the coefficient vector in linear regression under proportional asymptotics. While $\ell_2$-error is well-documented, we demonstrate that the $\ell_{\infty}$-error exhibits a distinct phase transition governed by the structure of the underlying signal. We explore this through the lens of a class of estimators that include ridge-regularized estimators, a method-of-moments estimator, and the ordinary least squares estimator for the under-parametrized regime. Our results provide a comparative study of these estimators and characterize regimes (as identified by the maximal signal component, number of large signals, total signal strength, and noise variance) where one estimator is preferable to the others. As a by-product, we obtain new double descent curves in $\ell_{\infty}$-error metric for the ridgeless regression, as well as ingredients for uniform confidence intervals for the signal coordinates. Our theoretical results are supported by extensive numerical simulations that confirm the predicted phase transitions and error limits.


Surprises in Proper Positive-Only Learning

Shai Ben-David ⋅ Farnam Mansouri ⋅ Anay Mehrotra ⋅ Manolis Zampetakis

Binary classification from positive-only samples is a variant of PAC learning in which the learner receives i.i.d. samples from the positive region of an unknown target concept, but is evaluated under the original distribution (which places mass on both positive and negative regions). This model dates back to (Natarajan, 1987, STOC), and the characterization of improper learning is well-known – it even appears in textbooks (Kearns and Vazirani, 1994, Exercise 3.7). The characterization of proper positive-only learning, however, has long remained open. In this work, we revisit and settle this question: a concept class is properly learnable from positive-only samples if and only if it has finite VC dimension and satisfies a new combinatorial condition, which we call uniform exterior separability. Together with several separation results, this characterization reveals a surprisingly rich landscape that differs sharply from standard PAC learning: proper and improper learning are separated, randomized and deterministic proper learning are separated, there are classes for which no ERM is a learner, and finite VC dimension does not suffice even for non-uniform learning. Along the way, we introduce new combinatorial dimensions that we believe can be of broader interest in learning theory.

Passenger-order cancellation is a critical source of uncertainty in on-demand ride-sharing systems, where requests, assignments, and vehicle routes evolve continuously in real time. Due to shared capacity and coupled routes, a single cancellation may propagate through downstream matching and routing decisions, degrading system efficiency and service reliability. However, cancellation prediction in on-demand ride-sharing remains difficult to study systematically due to the lack of large-scale longitudinal datasets, standardized evaluation protocols, and fine-grained real-time benchmarks. In this work, we introduce SurvCancel, a real-world longitudinal dataset and benchmark for dynamic passenger-order cancellation prediction in on-demand ride-sharing systems. Constructed from operational logs collected across four service regions in the real world, SurvCancel contains 173K released orders and 7.9M time-dependent snapshots. Each order is represented as a longitudinal sequence of system states, capturing both order-level attributes and the evolving supply-demand environment throughout its lifecycle. We formulate cancellation prediction as a stage-conditioned dynamic survival task: at each landmark time, models estimate short-term cancellation risk from the historical state sequence observed so far. We evaluate representative survival and dynamic-risk models under both in-distribution and distribution-shift protocols. Our results show stage-dependent model behavior in the in-distribution setting and reveal distinct robustness patterns under regional and temporal distribution shifts.


Survive or Collapse: The Asymmetric Roles of Data Gating and Reward Grounding in Self-Play RL

Sophia Xiao Pu ⋅ Zhaotian Weng ⋅ Chengzhi Liu ⋅ Jayanth Srinivasa ⋅ Gaowen Liu ⋅ William Yang Wang ⋅ Xin Wang

Self-play reinforcement learning trains language models on their own generated tasks, co-evolving a proposer and solver without human labels. Recent systems report strong reasoning gains, but collapse and instability are widely observed and poorly understood. The dominant response treats this as a reward-design problem. We argue instead that self-play stability is governed by two distinct levers: a data-level gate that decides which proposer-generated tasks enter the training pool, and the reward signal that updates the policy on tasks already admitted. Through controlled experiments on a Python output-prediction task and a deterministic-DSL twin task that strips pretraining priors, output ambiguity, and executor noise, we find the two levers are asymmetric. A strict gate is sufficient for stability under every reward variant we test, including a self-consistency reward with no access to ground truth; while no reward variant is sufficient once the gate is removed. This asymmetry exposes a counter-intuitive coupling we call the Grounded Proposer Paradox: a proposer with ground-truth access accelerates collapse faster than an ungrounded one when paired with a self-consistency solver, by concentrating training on clean tasks that form the fastest path to a spurious self-consistent attractor. Replacing the binary gate with a continuous strictness parameter $\varepsilon$ further reveals a two-stage phase transition: training-side metrics decouple at low $\varepsilon$, while validation accuracy holds until $\varepsilon$ is much higher. Data-level gating, not reward calibration, is the binding constraint on self-play stability.


SVDP: Training-Free Contextual Sparsity Predictors for Fast LLM Inference

Georgii Serbin ⋅ Kirill Koshkin ⋅ Zhongao Sun ⋅ Anastasiya Bistrigova ⋅ Constantine Korikov

Contextual sparsity is one of the approaches used to reduce computational complexity in the inference process of large language models (LLMs). Existing techniques for efficient LLM inference acceleration based on contextual sparsity with minimal accuracy degradation require training sparse pattern predictors. This paper presents a framework for accelerating the inference of ReGLU-based feed-forward networks (FFNs) within LLMs. The proposed framework provides a fast, training-free method for building sparse pattern predictors using truncation-aware singular value decomposition (SVD) of the gate projection matrix, along with a threshold calibration algorithm, and inference executors supporting conditional computation on CUDA and CANN devices. Experiments on three sparse LLMs with an average activation sparsity level of 90\% in the FFNs demonstrate up to a 1.8x end-to-end decoding time speedup with less than 1\% degradation in benchmark scores on tasks involving math and code generation. This work advances the deployment of LLMs on edge devices.

In conditioned generative models of physical systems, the densities at different values of the conditioning parameters are often related through symplectic transformations. For stable linear Hamiltonian systems, there exists a symplectic transformation to ``normal coordinates'' such that the dynamics is reduced to rotations in each phase-space plane. This picture extends to the nonlinear case where, away from resonances, the Birkhoff normal form again consists of rotations in phase space, now with amplitude dependence. In this work, we introduce an invertible symplectic flow layer consisting of amplitude-dependent rotations in learned normal coordinates. We derive the condition under which the nonlinear rotation layer is symplectic and detail an architecture that guarantees it. Experiments are carried out on two conditional generative modeling problems from accelerator physics. We show empirically that normalizing flows with the new layers achieve competitive or superior negative log-likelihoods compared to baseline flows, with gains particularly pronounced in weakly nonlinear systems.


Synaptic Strength Controls Trainability and Structural Stability in Rank-Deficient RNNs

Fatih Dinc ⋅ Edouard Ponnat ⋅ Henrik Weyer ⋅ Yanin Guerra ⋅ Nina Miolane

Real-world networks, from biological brains to ecological systems, are typically low-rank, yet they continue to learn throughout their existence. Understanding how learning operates in this rank-deficient regime is essential, but existing theory captures only its limits. Classical results describe early training in random high-dimensional networks, while low-rank theory describes its structured endpoint. To bridge these regimes, we introduce rank-deficient RNNs, \textit{i.e.}, networks initialized at low rank with fully trainable weights, in which rank and synaptic strength vary independently. This decoupling reveals that synaptic strength, not rank, primarily governs the learning regime, and that its strong and weak limits confer opposing advantages. Strong synapses produce rich nonlinear dynamics at initialization, enabling rapid learning that matches full-rank training speed at low ranks and triples it in gated architectures (GRUs, LSTMs). Weak synapses, by contrast, distribute computation across the population so that no individual connection is critical, yielding solutions that are structurally stable to neuron loss. Empirically, the weight changes induced by training consistently follow weak synaptic scaling across four tasks, four architectures, and distinct initializations. Overall, our findings identify synaptic strength as a central variable controlling both trainability and structural stability in rank-deficient RNNs.

Synergistic multi-modal question generation aims to generate questions grounded in synergistic semantics distributed across multiple modalities, rather than being specified by isolated information from any single modality alone. In this setting, modality-specific cues and cross-modal shared semantics must be jointly well-organized so that the generated question reflects synergistic multi-modal contribution and captures genuinely synergistic information. However, this task faces two core challenges: how to disentangle multi-modal information for synergistic information modeling, and how to preserve synergistic contribution during question generation. To address these challenges, we propose SynMQG, a synergistic dependency-aware framework derived from a mutual-information-based theoretical analysis, which integrates disentangled representation learning with synergy-aware optimization for synergistic multi-modal question generation. Specifically, SynMQG disentangles the multi-modal context into visual-specific, textual-specific, and shared semantics, thereby explicitly organizing heterogeneous information for synergistic information modeling. Based on these disentangled representations, it further introduces a reasoning chain as an intermediate scaffold to structure cross-modal information before generation. Moreover, SynMQG proposes GRPO with a mutual-information-based synergy reward, which explicitly measures whether the generated question preserves effective contribution of synergistic information, encouraging the model to generate questions that depend on synergistic multi-modal semantics rather than superficial single-modal cues. Experiments on ScienceQA and MultimodalQA show that SynMQG consistently outperforms representative MQG baselines across standard text-generation metrics, MLLM-based evaluation, and human evaluation, while generating questions with stronger synergistic multi-modal grounding and higher multi-modal dependency. The source code is available at https://anonymous.4open.science/r/MQG-8E57.


T$^2$-Splat: Adaptive Topology Mesh Splatting with Texture Residuals

Zhihao Tang ⋅ Youjia Zhang ⋅ Mingbo Zhao ⋅ Wei Yang

Triangle-based splatting has emerged as a promising bridge between high-quality novel view synthesis and traditional graphics pipelines. However, naively applying unstructured splatting mechanisms to polygonal primitives fails to exploit their inherent topological advantages, leading to over-tessellation and degraded high-frequency appearances. To address, we present T$^2$-Splat, a mesh splatting framework with adaptive topology and texture optimization. We introduce progressive edge-collapse and vertex-split operations to adaptively allocate triangle primitives, reducing face count while preserving well-conditioned connectivity. We further propose a scale-aware residual Mip-texture model to decouple appearance from geometry, for providing the topology with flexibility needed to safely collapse redundant triangles without introducing blurring or aliasing. Experiments on the Mip-NeRF 360, Tanks \& Temples, and DTU benchmarks demonstrate that T$^2$-Splat improves PSNR by +0.87 dB and maintains robust visual fidelity even under aggressive mesh compression to 10\% of the original face count.

Long-Tailed Class-Incremental Learning (LT-CIL) is challenging due to severe class imbalance, where frequent head classes dominate gradient updates and suppress learning for rare tail classes. Beyond data imbalance, we identify an asymmetric likelihood signals within each session: tail classes contribute insufficient evidence to overcome prior regularization, causing their parameters to remain near the prior mean even under variational inference. This asymmetry disrupts the stability-plasticity tradeoff, leading to pronounced forgetting and poor tail-class performance. We propose a hierarchical variational framework, TailAdapt, for prompt and adapter-based continual learning that enables adaptive group-wise regularization of model parameters. TailAdapt employs a hierarchical inverse Gaussian scale-mixture prior, which induces a heavy-tailed marginal distribution over prompt and adapter parameters. This formulation encourages selective plasticity so that most parameter groups are strongly regularized to preserve previously learned knowledge, while a small, data-supported subset is allowed to adapt substantially. This selective plasticity allocates adaptation capacity where it is most needed, mitigating interference from head classes and improving learning for tail classes. Extensive experiments on LT-CIL benchmarks demonstrate consistent improvements over strong baselines, with better tail-class performance, reduced forgetting, and more efficient utilization of prompt and adapter capacity.


Tail Competition Explains Scaling in Best-of-N Sampling for Verifiable Problems

Taro Yano ⋅ Yoichi Ishibashi ⋅ Masafumi Oyamada

Best-of-N sampling is a simple inference-time scaling strategy for large language models, yet its behavior under incomplete reward models remains poorly understood. We study Best-of-N as a competition between the extreme reward tails of correct and incorrect responses. This view shows that success depends not only on the probability of sampling a correct response, but also on whether correct responses dominate incorrect ones in the upper tail. Using extreme value theory, we characterize three asymptotic regimes: correct-tail-dominant, incorrect-tail-dominant, and critical explaining both monotonic gains and reward-hacking-induced degradation. We further derive finite-sample scaling laws for representative reward distributions, showing that, under Gaussian rewards, incorrect responses with sufficiently large variance can dominate selection even when their mean reward is lower. Unlike oracle selection, convergence with an incomplete selector can be polynomial rather than exponential. Finally, we propose a lightweight tail-tension criterion to estimate when additional sampling may become harmful. Experiments on synthetic data and real reward model scores validate the predicted regimes, scaling laws, and diagnostic behavior, suggesting that tail competition provides a unified lens on Best-of-N sampling for verifiable problems.


TALON: Confidence-Aware Speculative Decoding with Adaptive Token Trees

Tianyu Liu ⋅ Qitan Lv ⋅ Yuhao Shen ⋅ Jun Zhang ⋅ Xiao Sun ⋅ Xiaoyan Sun

Speculative decoding (SD) has become a standard technique for accelerating LLM inference without sacrificing output quality. Recent advances in speculative decoding have shifted from sequential chain-based drafting to tree-structured generation, where the draft model constructs a tree of candidate tokens to explore multiple possible drafts in parallel. However, existing tree-based SD methods typically build a fixed-width, fixed-depth draft tree, which fails to adapt to the varying difficulty of tokens and contexts. As a result, the draft model cannot dynamically adjust the tree structure to early stop on difficult tokens and extend generation for simple ones. To address these challenges, we introduce TALON, a training-free, budget-driven adaptive tree expansion framework that can be plugged into existing tree-based methods. Unlike static methods, TALON constructs the draft tree iteratively until a fixed token budget is met, using a hybrid expansion strategy that adaptively allocates the node budget to each layer of the draft tree. This framework naturally shapes the draft tree into a "deep-and-narrow" form for deterministic contexts and a "shallow-and-wide" form for uncertain branches, effectively optimizing the trade-off between exploration width and generation depth under a given budget. Extensive experiments across 5 models and 6 datasets demonstrate that TALON consistently outperforms state-of-the-art EAGLE-3, achieving up to 5.16× end-to-end speedup over auto-regressive decoding.


TANGO: RNA Topology and Geometry Co-Design

Tianmeng Hu ⋅ Biao Luo ⋅ Ke Li

Designing RNA sequences that satisfy target tertiary-structure objectives is an important problem in RNA engineering, with broad applications in synthetic biology and therapeutics. Existing $\mathrm{3D}$ RNA design methods have made substantial progress, but they largely treat the target fold as a geometry-only objective. This is incomplete for RNA, whose functional folds are specified not only by tertiary geometry but also by secondary base-pairing topology. We therefore formulate RNA inverse design as a topology-and-geometry co-design problem and propose $\texttt{TANGO}$, a cooperative multi-agent reinforcement learning framework for this task. $\texttt{TANGO}$ adopts a staged curriculum: it first learns a topology-aware policy from secondary-structure targets, then transfers this policy to tertiary-structure targets to train a general co-design model, and finally applies frontier-guided fine-tuning to refine target-specific trade-offs among secondary-structure fidelity, tertiary-structure fidelity, and sequence diversity. Experiments show that $\texttt{TANGO}$ generates diverse and novel designs, improving tertiary- and secondary-structure performance by $18.3\\%$ and $52.7\\%$, respectively, over the strongest baseline.


Task-Aware KV Cache Compression for LLM Agents via Utility-Driven Step Pruning

Yusen Wu ⋅ Yefan Wang ⋅ Jia Yee Tan ⋅ Guangyuan Dong ⋅ Shuang Chen ⋅ Jing Yang ⋅ Rongfeng Guo ⋅ Por L Yee ⋅ Yongtai Liu

KV-cache compression for long-horizon LLM agents is often guided by accumulated attention, yet attention frequency can retain repetitive error traces while discarding decisive instructions. We propose TaskKV, a streaming step-level retention policy with three contributions: (1) a signed directional utility label from the first-order change in action likelihood, distinguishing helpful steps from confusing distractors unlike unsigned saliency; (2) a structured residual scorer that treats attention mass and intent alignment as fixed priors and learns only a denoising correction, converging with 500 trajectories; and (3) an entropy-based regime classifier that falls back to contiguous retention when the scorer cannot confidently rank steps. With a 20% KV budget on AgentBench and AlfWorld, TaskKV preserves $\sim$87--89% of full-cache success rate, improving over SnapKV by 5.0--7.9 SR points and over a strong heuristic by 2.9--7.3 points; on dense tasks, gains come from detecting when utility pruning should be disabled. Code is available at https://anonymous.4open.science/r/taskkv-F0D5.


Temporal Concentration from Rollout Errors: Implicit Preference Optimization For Text-to-Video Diffusion

henglin liu ⋅ Fangyuan Kong ⋅ Jing Wang ⋅ Yizhou Lin ⋅ Nisha Huang ⋅ Chang Liu ⋅ Xintao Wang ⋅ Pengfei Wan ⋅ Kun Gai ⋅ Xiu Li

Recent advances in preference alignment for diffusion-based video generation, particularly via Direct Preference Optimization (DPO), have significantly improved visual quality. However, temporally sparse artifacts such as motion collapse, object flickering, and color oversaturation remain a major barrier to perceptual realism. Existing methods struggle with these issues due to two key limitations: (1) the preference attribution bottleneck, where offline human annotations are costly and fail to accurately capture learning dynamics, while online reward signals are rollout-aware but often unstable and biased; and (2) temporal credit misallocation, where uniformly applied supervision cannot effectively target the brief segments in which artifacts occur. To address these challenges, we propose concentrated Implicit Preference Optimization (cIPO), a post-training framework for video diffusion models. cIPO derives implicit preference signals directly from the denoising process: given a real video, the model adds forward noise and reconstructs it via iterative denoising, treating the original as the preferred sample and the reconstruction as the dispreferred one. This formulation captures inference-time errors without requiring human annotations or external reward models. Moreover, frame-level discrepancies between original and reconstructed videos reveal when failures occur. cIPO leverages this by computing temporal reconstruction errors and concentrating optimization on high-error segments, enabling more precise correction of failure-prone regions. Extensive experiments demonstrate that cIPO consistently enhances video authenticity and temporal coherence across multiple datasets, highlighting the effectiveness and efficiency of implicit implicit preference with temporally concentrated optimization.

Pre-trained vision-language models (VLMs) enable zero-shot classification by matching images with text prompts, yet their raw confidence scores are unreliable for out-of-distribution (OOD) detection. We identify a confidence bottleneck: zero-shot VLMs produce posteriors whose maximum confidence yields much weaker ID/OOD separation than supervised classifiers trained on the target label space. To theoretically explain this gap, we relate VLM confidence to a well-separated supervised reference through a Markov-kernel transformation and analyze single-threshold detection using the Youden/KS index. Our analysis shows why raw VLM confidence is unlikely to recover the clean separation achieved by supervised classifiers. Motivated by the observation that class-wise posterior distributions contain useful structure beyond maximum confidence, we propose Self-Reinforced Optimal Transport (SROT), a label-free test-time framework that reshapes VLM posteriors with optimal transport and adaptively trains a lightweight OOD detector from rectified pseudo-labels. The detector further feeds open-set evidence back into posterior alignment, forming a self-reinforcing loop between distribution-level rectification and feature-level OOD learning. Extensive experiments on standard zero-shot OOD detection benchmarks show that SROT achieves state-of-the-art performance, with ablations validating posterior alignment, adaptive detector training, and self-reinforcement.


Test-time Scaling of Diffusions with Flow Maps

Amirmojtaba Sabour ⋅ Michael Albergo ⋅ Carles Domingo i Enrich ⋅ Nicholas Boffi ⋅ Sanja Fidler ⋅ Karsten Kreis ⋅ Eric Vanden-Eijnden

A common recipe to improve diffusion models at test-time so that samples score highly against a user-specified reward is to introduce the gradient of the reward into the dynamics of the diffusion itself. This procedure is often ill posed, as user-specified rewards are usually only well defined on the data distribution at the end of generation. While common workarounds to this problem are to use an approximate denoiser to estimate what a sample would have been at the end of generation, we propose a simple solution to this problem by working directly with a flow map. By exploiting a relationship between the flow map and velocity field governing the instantaneous transport, we construct an algorithm, Flow Map Trajectory Tilting (FMTT), which leverages a fast yet precise flow map look-ahead and provably performs better ascent on the reward than standard test-time methods involving the gradient of the reward. The approach can be used to either perform exact sampling via importance weighting or principled search that identifies local maximizers of the reward-tilted distribution. We demonstrate the efficacy of our approach against other look-ahead techniques, and show how the flow map enables engagement with complicated reward functions that make possible new forms of image editing, e.g. by interfacing with vision language models.


Test-Time Sequential Steering of Diffusion Models via Preconditioned Crank-Nicolson

Joel Keller ⋅ Taos Transue ⋅ Qin Li ⋅ Shih-Hsin Wang ⋅ Bao Wang

Sampling from reward-tilted distributions enables diffusion models to satisfy task-specific constraints without retraining. However, existing methods either rely on gradients or struggle to explore multimodal high-reward regions. We introduce a new test-time steering method for diffusion models that modifies the reverse denoising process to incorporate reward information directly. At each step, we propose a Metropolis–Hastings–corrected denoising mechanism based on a preconditioned Crank–Nicolson (pCN) proposal, enabling principled acceptance of noise updates that bias sampling toward high-reward regions. To further improve exploration in multimodal reward landscapes, we extend this procedure with parallel tempering across multiple temperature levels, allowing controlled mixing between exploration and refinement regimes. Our method operates with or without reward gradients and applies to black-box objectives. Empirically, we demonstrate its effectiveness on synthetic tasks, image generation, dynamical systems, and Bayesian inverse problems. Compared to prior methods, it achieves superior reward alignment and multimodal exploration while maintaining stability across tasks.

Vision--language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted details. We aim to learn representations whose matching is stable across caption views and whose confidence reflects how strongly text constrains an image. We propose Text as Partial Constraint (TPC), a core--residual alignment framework that treats multi-view captions as incomplete supervision: it distills a consensus semantic core as the alignment target, learns a single-view core predictor for standard inference with one query, and explicitly discourages vision--language similarity from depending on the orthogonal ``unsaid'' residual. An uncertainty-aware contrastive objective further softens alignment when caption views disagree, reducing overconfident updates under weak language constraints. Across zero-shot recognition and adversarial robustness, TPC achieves 81.42/64.05 Top-1 clean/robust accuracy on ImageNet and 76.19/52.03 on an Avg-14 transfer suite, while improving LVLM transfer with 85.16 POPE F1 and 59.57 OKVQA accuracy under an LLaVA-1.5-7B stack. These results suggest that modeling text as a partial constraint is a practical and principled route to more reliable vision--language representations under underspecified language supervision.


The Confusion is Real: GRAPHIC – A Network Science Approach to Confusion Matrices in Deep Learning

Johanna S. Fröhlich ⋅ Bastian Heinlein ⋅ Jan Claar ⋅ Hans Rosenberger ⋅ Vasileios Belagiannis ⋅ Ralf R Müller

Explainable artificial intelligence has emerged as a promising field of research to address reliability concerns in artificial intelligence. Despite significant progress in explainable artificial intelligence, few methods provide a systematic way to visualize and understand how classes are confused and how their relationships evolve as training progresses. In this work, we present GRAPHIC, an architecture-agnostic approach that analyzes neural networks on a class level. It leverages confusion matrices derived from intermediate layers using linear classifiers. We interpret these as adjacency matrices of directed graphs, allowing tools from network science to visualize and quantify learning dynamics across training epochs and intermediate layers. GRAPHIC provides insights into linear class separability, dataset issues, and architectural behavior, revealing, for example, similarities between flatfish and man and labeling ambiguities validated in a human study. In summary, by uncovering real confusions, GRAPHIC offers new perspectives on how neural networks learn. The code is available at https://github.com/Johanna-S-Froehlich/GRAPHIC.


The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching

Sankaran Vaidyanathan ⋅ David Arbour ⋅ Aaron Mueller ⋅ Scott Niekum ⋅ David Jensen

Activation patching is the primary tool in mechanistic interpretability. It attributes causal responsibility for a model behavior to each of its individual components by estimating its natural indirect effect (NIE). Re-deriving the activation patching estimand from causal mediation analysis, we find that the NIE does not solely capture the causal effect through the specific component. It also contains interaction effects (INT) that measure how much the component's causal effect itself depends on the state of other components in the model. A natural response may be to try to eliminate INT by adjusting the estimator or unit of analysis, but each of these potential remedies has predictable failure modes. We demonstrate these failure modes in the GPT-2 IOI circuit; components whose causal importance is conditional on the state of other components are either invisible or artificially inflated, and INT variance explains the previously documented instability of faithfulness scores. We prove that INT scales with the distance between clean and patched component activations, is negligible when the model is locally affine, and decomposes combinatorially into pairwise and higher-order group interactions. Despite its inevitability, INT is not a nuisance to be eliminated, but rather a diagnostic for interpretability studies. Its individual and group-level magnitude and sign signal when causal conclusions are prompt-dependent, and when greedy NIE-based component ranking will miss mechanisms only discoverable through combinatorial search.


The Dynamics of Policy Gradient in Social Dilemmas with Partner Selection

Benedict Russell ⋅ Chin-wing Leung ⋅ Paolo Turrini

In social dilemmas self-interested learning agents face the choice between the societal benefit of cooperation and the immediate reward of defection. Significant evidence exists on the benefits of assortment mechanisms such as partner selection for the emergence of cooperation, but this is largely available through agent-based simulations. In this paper, we provide an analytical solution to the problem, studying the policy-gradient dynamics in a multi-agent environment with partner selection. We show how partner selection changes the opponent distribution and hence the reward landscape, and prove this promotes cooperation under simple rules known from the literature. In particular, we find that population variance is a necessary condition for cooperation to emerge. Using a two-dimensional Wiener process, we extend the dynamics to capture the stochastic effects of partner selection and the resulting opponent distribution. We derive a sufficient condition for the population to be cooperation-promoting and prove the existence of a stationary distribution. Simulations confirm that the stochastic model accurately captures the policy-gradient dynamics and clarifies how the learning rate affects the emergence of cooperation.


The End Justifies the Mean: Linear Ranking Rules for Proportional Sequential Decisions

Carmel Baharav ⋅ Niclas Boehmer ⋅ Bailey Flanigan ⋅ Maximilian T. Wittmann

AI alignment and participatory design motivate a new democratic design problem: how to collectively choose a _decision rule to use repeatedly_. We study this problem for _linear ranking rules_, which repeatedly rank items $x_j$ within batches $X=(x_1,\dots,x_m)\in(\mathbb{R}^d)^m$, where each item's ranking is dictated by its score $\langle \theta^{\ast},x_j\rangle$ according to a fixed scoring vector $\theta^{\ast}$. Given voters' preferred scoring vectors $\theta^{(1)},\dots,\theta^{(n)}$ and their population fractions $\alpha^{(1)},\dots,\alpha^{(n)}$, we ask how to choose a collective vector $\theta^{\ast}$ satisfying _individual proportionality (IP)_: every voter type $i$ should agree with the resulting rankings to an $\alpha^{(i)}$-proportional degree, either on average over time (_long-run IP_) or even within each batch (_per-batch IP_). The default rule, the arithmetic mean of the $\theta^{(i)}$, has been shown to be severely majoritarian; more generally, it is not clear that _any_ fixed linear rule can balance many voters' disparate opinions. Our main result is that, surprisingly, there _is_ a simple rule that does satisfy long-run IP: the _angular mean_, the spherical analog of the arithmetic mean. We then show that exact per-batch IP is impossible for fixed linear rules, but that the gap between per-batch and long-run IP shrinks quickly with batch size. Experiments on three real-world preference datasets show that all rules perform similarly when voters' preferences are homogeneous, while the angular mean substantially improves proportionality in high-disagreement regimes.

Probabilistic Circuits (PCs) are deep generative models that support exact and efficient probabilistic inference. Yet in autoregressive language modeling, PCs still lag behind Transformer-based large language models (LLMs), suggesting an important expressivity gap. In this work, we compare PCs and LLMs under a unified autoregressive formulation. First, an output bottleneck: PCs parameterize predictions as convex combinations in probability space, which struggles to represent the sharp distributions typical of language; adopting a logit-space parameterization substantially narrows this gap. Second, a context-encoding bottleneck: we prove that structured-decomposable PCs can match Transformer separation rank on vtree-aligned partitions, but show, both theoretically and empirically, that this capacity is limited to partitions aligned with the fixed routing structure, leading to severe degradation when the data exhibits heterogeneous dependency topologies. We further prove that decomposable PCs are strictly more expressive than structured-decomposable ones, though effectively optimizing them remains an open challenge.


The Geometry of Forgetting: Temporal Knowledge Drift as an Independent Axis in LLM Representations

Rania Elbadry ⋅ Ahmed Heakl ⋅ Fan Zhang ⋅ Dani Bouch ⋅ Yuxia Wang ⋅ Preslav Nakov ⋅ Zhuohan Xie

Large language models confidently produce outdated answers, and no existing method can detect them. We show this is not an engineering failure but a structural one: temporal drift, whether a stored fact has changed since training, is encoded as a direction in the residual stream geometrically orthogonal to both correctness and uncertainty. Any method operating on correctness or uncertainty signals is therefore blind to drift by construction. We verify this across six instruction-tuned models. A linear probe trained directly on drift labels achieves AUROC $0.83$--$0.95$; methods based on token entropy, semantic entropy, CCS, and SAPLMA all remain near chance ($0.49$--$0.57$). Five tests confirm the geometric orthogonality: weight cosines ($|\cos| \leq 0.14$), score correlations ($|r| \leq 0.20$), bidirectional null-space projection ($|\Delta| \leq 0.008$), iterative null-space projection with $k{=}10$, and difference-of-means dissociation. Mechanistically, the MLP retrieval circuit produces identical dynamics for stale recall and confabulation ($r > 0.81$, six models), explaining why output confidence cannot separate them. A cross-cutoff experiment holds inputs constant and varies only the model: the probe fires on the model whose training predates the fact's transition and stays silent otherwise ($P(A{>}B) = 0.975$--$0.998$, twelve model pairs), confirming it reads model-internal knowledge state rather than input properties. Our code and datasets will be publicly released.


The Geometry of Noise: Why Diffusion Models Don't Need Noise Conditioning

Mojtaba Sahraee-Ardakan ⋅ Mauricio Delbracio ⋅ Peyman Milanfar

Generative models typically rely on explicitly tracking the noise level to guide the sampling process. However, recent autonomous models successfully generate data using a single, time-invariant vector field that operates without explicit noise conditioning. This raises a fundamental paradox: what landscape are these networks actually optimizing when the noise level is treated as a random variable, and how can a bounded model remain stable near the data manifold where gradients typically diverge? We resolve this paradox by showing that autonomous generation is not merely ``blind’’ denoising, but a specific form of Riemannian gradient flow on the Marginal Energy landscape, defined by integrating out the unknown noise levels. Through a novel relative energy decomposition, we demonstrate that while the raw Marginal Energy contains a severe geometric singularity normal to the data manifold, the learned field implicitly incorporates a local conformal metric. This metric perfectly preconditions the singularity, turning an infinitely deep potential well into a stable attractor. Finally, we establish strict structural stability conditions for autonomous sampling. We prove that standard noise-prediction parameterizations structurally fail due to an inherent high gain amplification of the estimation errors. In contrast, velocity-based models are inherently stable.


The Narrative Gap: Can LLMs help us Navigate Diverse Narratives Across Languages?

Nikhil Sharma ⋅ Kelly Marchisio ⋅ Kenton Murray ⋅ Ziang Xiao

Large language models increasingly mediate how individuals seek information about complex, often contested topics. However, no existing benchmark evaluates whether models can navigate the diverse narratives that develop around the same event across languages and regions. We argue that to support diverse information seeking, the system requires four core information navigation capabilities: differentiating narratives, presenting diverse views, resisting confirmatory queries, and achieving linguistic information parity. We introduce \textsc{NarrativeBench}, a benchmark designed to evaluate these capabilities jointly in real-world settings. \textsc{NarrativeBench} comprises 588 news articles in 17 languages from 26 regions spanning 12 geopolitical conflicts. To evaluate the models we curated 213 grounded queries from Reddit resulting in 1917 queries across 9 languages. The unit of evaluation is the model's ability to navigate diversity, not the truth of any single account. Evaluating 14 models, we find current multilingual LLMs (1) are unable to navigate diverse multilingual narratives, (2) lack the ability present diverse narratives across multilingual contexts, (3) have asymmetry in performance across different languages, (4) reduce diversity when faced with confirmatory queries which may exacerbate linguistic filter bubbles, echo chambers and reduce common ground across regions. \textsc{NarrativeBench} provides a capability-grounded foundation for building information systems that can promote equitable information access, reduce polarization, and enhance democratic discourse.

Initialization determines the optimization ``fate'' of a model, going far beyond just setting the scale of its weights. It establishes a fundamental initial geometry over the data, permanently dictating which examples are considered close, which directions are easy to change, and which representations gradient descent will naturally refine first. Because this starting point dictates the model's future trajectory, we ask: can this initial geometry be explicitly chosen by referencing a completely different architecture? To achieve this, we introduce representational similarity as an optimization prior. This is an initialization-only procedure that aligns a target network to the representational geometry of a randomly initialized guide network before any downstream training occurs. Crucially, the guide network transfers no learned knowledge: it is frozen, never sees labels. We demonstrate that this cross-architecture transfer of fate works in two key domains: instilling the optimization benefits of depth into shallow targets by aligning them to deeper, randomly initialized guides, and transferring representational granularity by using fine-resolution guides to shape the initialization geometry of networks built with coarser representations. Ultimately, this establishes that a model's architectural destiny can be decoupled from its physical structure and explicitly programmed at initialization.

Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in increasingly agentic and autonomous forms demands new evaluations. In this paper we evaluate an agent's ability to induce specific belief states in other agents by taking actions rather than using conversational persuasion, a capability we call Non-Conversational Planning ToM (NCP-ToM). NCP-ToM is likely to be essential for many agent use-cases, including within user-assistant interactions and pedagogical contexts, but may also present manipulation or misinformation risks. Using a novel framework, NCP-ExploreToM, we subvert the conventional task structure by providing models with a set of belief state goals and requiring them to move objects or direct characters into rooms to achieve their goals. We evaluated six frontier models, including GPT-5, Gemini 2.5 Pro and the Claude 4 series, and a cohort of human participants, across 600 task instances. GPT-5 was successful on approximately 80\% of tasks in the agentic setting, and was the only model to outperform human participants on our task, but was still less robust than humans across contexts. We additionally found that all models, like humans, performed better on tasks inducing true belief states than false belief states, which is a positive signal for alignment efforts. These findings highlight emerging social-reasoning capabilities in LLMs for non-conversational task completion and underscore the necessity of agentic evaluations for understanding the safety and alignment of autonomous social agents.

In-context learning is usually analyzed as if the examples in the prompt were sampled before the learner is chosen. In deployed systems they are not: examples are retrieved, filtered, re-ranked, edited, and repeatedly tested after seeing the query and the data store. We formalize this gap by treating prompt construction as adaptive data analysis through a conditional stochastic-kernel calculus. Our main quantity is the \emph{context capacity}, the mutual information or approximate max-information between the data store and the final prompt transcript. For any frozen in-context learner and any adaptive context pipeline with capacity $\kappa$, we prove a finite-sample context generalization bound of order $\sqrt{\kappa/k}$, where $k$ is the number of examples that enter the prompt. We strengthen it to a sub-gamma transport inequality with Bernstein curvature, prove a matching lower bound, and derive two PAC-Bayesian retrieval rules: a first-order Gibbs rule and a self-normalized second-order rule. Finally, we instantiate the theory for Bayesian linear in-context regression, including an order-statistics analysis of query-aligned top-$k$ Gaussian retrieval. The results separate context quality from context certification: longer or more relevant contexts improve risk only when the procedure used to choose them is stable enough for selected-context evidence to transfer.


The Reward Was in Your Data All Along: Correcting Flow Matching with Discriminator-Guided RL.

Nicolas Beltran Velez ⋅ Felix Friedrich ⋅ Xiaofeng Zhang ⋅ Reyhane Askari Hemmat ⋅ Xiaochuang Han ⋅ Adriana Romero-Soriano ⋅ Michal Drozdzal

Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties---such as visual realism and coherent object structure---that matching-based training is intended to learn from the data itself. We argue that this reflects a structural mismatch. Matching losses measure $\ell_2$ regression error on the velocity or score field under training-time marginals, a proxy poorly aligned with the visual and semantic properties that determine sample quality at inference. Given a reward aligned with these properties, RL sidesteps the mismatch by evaluating the model on its own samples and following the reward landscape directly. The challenge is to obtain such a reward without relying on human references, which are expensive and conflate data realism with annotator inclinations. We propose Discriminator-Guided RL (DRL). DRL trains a discriminator to separate data from base-model samples in a pretrained representation space and uses its logit as the reward in KL-regularized RL. The pretrained space restricts the discriminator to perceptually meaningful directions, and the logit estimates the log-likelihood ratio between data and model, which is the optimal reward for targeting the data distribution. Across SiT, JiT, REPA, and RAE, DRL reduces guidance-free FID (e.g., $9.38 \to 2.62$ on SiT) and semantic-space FD (e.g., $88.2 \to 19.3$ on DINOv3 for SiT), with consistent gains across all backbones, and improves human-preference rewards without training on them. It also yields a better Pareto frontier between preference reward and image fidelity under subsequent preference-based post-training, increasing alignment while reducing low-level artifacts such as oversaturation and excessive brightness.

Synthesizing programs from execution traces is a fundamental challenge in algorithmic interpretability and the reverse engineering of complex systems. We bring a rough-path-theoretic perspective to program induction, treating execution traces as high-dimensional rough paths and encoding them with the path signature transform, a tool from stochastic analysis whose expected signature provably characterizes the trace distribution a program induces. Our signature-based trace encoder outperforms LSTM, Transformer, and Fourier baselines and remains strong with substantially less training data on two classical DSL domains. We further observe that the trace-to-program pipeline is naturally \emph{contractive}: it tends to produce shorter valid programs than the ground truth, which we turn into a bootstrapping procedure that surfaces redundancy in the released ground-truth supervision without any external oracle. Finally, signature-based models induce a markedly more interpretable latent geometry than classical sequence models.


The Shape of Events: Edge-Based Inductive Biases via Cross-Domain Distillation

Shunsuke Yasuki ⋅ Masato Taki ⋅ Soshun Kihara

Convolutional neural networks trained on ImageNet are known to rely heavily on local high-frequency texture, an inductive bias that translates into fragile robustness against distribution shifts in real-world environments. Event cameras, in contrast, record only changes in scene brightness and are therefore well suited to capturing contour information; however, due to the absence of diagnostic benchmarks in the event domain, the inductive bias that event-camera data instills in vision models has remained unexplored. In this work, we use knowledge distillation from the event domain to the RGB domain so as to exploit the rich evaluation toolkit available in the RGB domain and systematically dissect this inductive bias. Our experiments show that distillation from the event domain induces, in the RGB domain, notable color invariance, shape bias, and robustness to high-frequency noise. We identify the underlying mechanism as the model suppressing its dependence on high-frequency texture while simultaneously acquiring a strong dependence on edge-based object shape. This hypothesis is supported by changes in how color and spatial information are processed at the early layers, together with a "spectral trade-off" in which robustness to the absence of high-frequency components coexists with vulnerability to contamination of the relied-upon frequency bands and to disruption of geometric structure. We further show that this inductive bias differs markedly from existing robustification methods and that it functions as a strong prior for diverse downstream tasks that demand shape-based reasoning, such as medical imaging. The code and experimental configurations for reproducing our experiments are submitted alongside this paper as supplementary material.


The Sparsity Whisperer

Linghao Kong ⋅ Inimai Subramanian ⋅ Micah Adler ⋅ Dan Alistarh ⋅ Dan Gutfreund ⋅ Nir Shavit

Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We argue that this overlooks a key computation performed by particularly sparsity-sensitive neurons in the MLP up and gate projections: separating similar inputs into dissimilar outputs. This suggests that effective pruning should preserve not only activations, but also pairwise output differences. We introduce a family of difference-informed pruning methods built upon this principle. Wisp is a first-order, update-free method that scores weights using input-difference norms, and Wisp+ refines this score neuronwise using the input pairs each neuron separates most strongly. Finally, Whisper is a second-order method that uses a lightly regularized difference Hessian as its reconstruction objective. Across the Llama 2 and 3.1 families, our second-order variant consistently improves language modeling performance over strong reconstruction-based baselines, while our update-free variants improve over activation-aware update-free baselines, with gains stronger in more constrained settings. The improvements over Wanda and SparseGPT extend to structured sparsity, downstream evaluations, and other model families, suggesting that preserving input differentiation is a broadly useful signal for post-training LLM sparsification.

Continuous image editing requires fine-grained control over semantic attributes in text-conditioned generative models. Recent methods obtain such control through trainable adapters, test-time optimization, or architecture-specific design choices. In this work, we revisit steering the text-encoder space as a simpler alternative for open-vocabulary slider construction. We show that, as text-conditioned generative models become increasingly capable, linear directions in their frozen text-encoder spaces can provide a competitive interface for continuous control when carefully constructed. To make this interface practical, we introduce an automated pipeline that constructs continuous sliders from just a prompt and a concept, eliminating manual data curation and coefficient tuning. First, a language model generates balanced contrastive prompt pairs. Then, difference-of-means over the frozen text encoder estimates a concept direction, and an LLM-assisted process selects the prompt tokens to steer. Finally, an elastic range search adaptively calibrates the slider coefficient interval. As the process includes only changing the text encoder hidden states, the same pipeline can be applied across different text-conditioned generative backbones without denoiser-specific modifications. Across diverse set of edit types, our method outperforms prior training-free baselines and approaches the edit compliance of trained controllers while avoiding per-concept and per-backbone training. Our results suggest that a carefully designed text-space steering is a surprisingly strong practical baseline for open-vocabulary continuous editing. Especially in a rapidly evolving generative model ecosystem, where repeatedly adapting specialized controllers to new backbones can be a substantial burden


The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks

Taehun Cha ⋅ Daniel Beaglehole ⋅ Adityanarayanan Radhakrishnan ⋅ Donghun Lee

Understanding how deep neural networks learn representations remains a central challenge in machine learning theory. In this work, we propose a feature-centric framework for analyzing neural network training by relating weight updates to feature evolution. We introduce a simple identity, the Feature Learning Equation, which identifies the weight Gram matrix as the key object capturing feature dynamics. This enables us to interpret gradient descent as implicitly inducing a hypothetical evolution of features, whose covariance structure — termed the Virtual Covariance — characterizes how representations evolve during training. Building on this perspective, we introduce Target Linearity, a measure quantifying the linear alignment between features and targets. By analyzing the training and layer-wise dynamics, we show that deep networks learn to sequentially transform representations toward target-linear structure. This linearization perspective provides a unified interpretation of several empirical phenomena, including Neural Collapse and linear interpolation in generative models.


Think Before You Generate: Active Panoramic Exploration for Text-to-3D Scenes

Derui Li ⋅ Peng Lu ⋅ Lujing Cao ⋅ Yuhao Sun ⋅ Wenhao Guo ⋅ Sheng Li

Creating explorable 3D environments from text is important for immersive content creation, simulation, and interactive design. Recent methods often rely on image or video generative priors by synthesizing views along fixed or manually specified trajectories and lifting them into 3D. However, such passive pipelines cannot adapt view generation to the evolving scene memory, often leaving disoccluded regions incomplete and producing holes, floaters, and inconsistent geometry under free-viewpoint exploration. We argue that the key challenge is not to generate more views, but to decide where to generate useful views according to the current scene state. To this end, we propose ActPano3D, an active text-to-3D scene generation framework based on memory-guided panoramic exploration. ActPano3D formulates trajectory selection as spatial deliberation over the evolving scene memory, where candidate actions are evaluated by their expected utility for exploring unknown regions, repairing uncertain geometry, and improving the future 3D world model. The framework integrates expansion-refinement active planning, memory-routed panoramic generation, and reliability-aware scene memory fusion into a closed-loop generation-and-reconstruction process. The accumulated memory is further optimized into a renderable 3D Gaussian Splatting scene for free-viewpoint exploration. Experiments show that ActPano3D improves scene coverage, reduces rendering holes, and achieves better rendering quality and text alignment than existing text-to-3D scene generation baselines.


Thinking with Images as Continuous Policy: Numerical Visual Chain-of-Thought

Kesen Zhao ⋅ Beier Zhu ⋅ Junbao Zhou ⋅ Xingyu Zhu ⋅ Zhongqi Yue ⋅ Hanwang Zhang

Recent multimodal large language models (MLLMs) increasingly rely on visual chain-of-thought to perform region-grounded reasoning over images. However, existing approaches ground regions via either textified coordinates—causing modality mismatch and semantic fragmentation—or fixed-granularity patches that both limit precise region selection and often require non-trivial architectural changes. In this paper, we propose Numerical Visual Chain-of-Thought (NV-CoT), a framework that enables MLLMs to reason over images using continuous numerical coordinates. NV-CoT expands the MLLM action space from discrete vocabulary tokens to a continuous Euclidean space, allowing models to directly generate bounding-box coordinates as actions with only minimal architectural modification. The framework supports both supervised fine-tuning and reinforcement learning. In particular, we replace categorical token policies with a Gaussian (or Laplace) policy over coordinates and introduce stochasticity via reparameterized sampling, making NV-CoT fully compatible with GRPO-style policy optimization. Extensive experiments on three benchmarks against eight representative visual reasoning baselines demonstrate that NV-CoT significantly improves localization precision and final answer accuracy, while also accelerating training convergence, validating the effectiveness of continuous-action visual reasoning in MLLMs. The code is provided in the Supplementary Materials.


Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling

Yuan Wang ⋅ Ouxiang Li ⋅ Yulong Xu ⋅ Borui Liao ⋅ Jiajun Liang ⋅ Jinghan Li ⋅ Meng Wang ⋅ Xintao Wang ⋅ Pengfei Wan ⋅ Kuien Liu ⋅ Xiang Wang

Recent advances in generative video models are increasingly driven by post-training and test-time scaling, both of which critically depend on the quality of video reward models (RMs). An ideal reward model should predict accurate rewards that align with human preferences across diverse scenarios. However, existing paradigms face a fundamental dilemma: \textit{Discriminative RMs} regress rewards directly on features extracted by multimodal large language models (MLLMs) without explicit reasoning, making them prone to shortcut learning and heavily reliant on massive data scaling for generalization. In contrast, \textit{Generative RMs} with Chain-of-Thought (CoT) reasoning exhibit superior interpretability and generalization potential, as they leverage fine-grained semantic supervision to internalize the rationales behind human preferences. However, they suffer from inherent optimization bottlenecks due to the coupling of reasoning and scoring within a single autoregressive inference chain. To harness the generalization benefits of CoT reasoning while mitigating the training instability of coupled reasoning and scoring, we introduce DeScore, a training-efficient and generalizable video reward model. DeScore employs a decoupled ``think-then-score'' paradigm: an MLLM first generates an explicit CoT, followed by a dedicated discriminative scoring module consisting of a learnable query token and a regression head that predicts the final reward. DeScore is optimized via a two-stage framework: (1) a discriminative cold start incorporating a random mask mechanism to ensure robust scoring capabilities, and (2) a dual-objective reinforcement learning stage that independently refines CoT reasoning quality and calibrates the final reward, ensuring that higher-quality reasoning directly translates to superior model performance. Empirical evaluations demonstrate that DeScore achieves superior training efficiency and optimization stability, while outperforming state-of-the-art methods across diverse in-domain and out-of-distribution benchmarks. Moreover, DeScore also proves effective for post-training, leading to improved generated video quality.


Tikhonov-Stabilized Bezier Representation Forecasting for Training-free Diffusion Acceleration

Lei Zhu ⋅ Mujie Lin ⋅ Ruochong Zheng ⋅ Guangyi Wang ⋅ Li Hao ⋅ Peng Jin ⋅ Chang Liu ⋅ Jie Chen

Diffusion models, particularly Diffusion Transformers, achieve strong image and video generation quality but remain expensive at inference time due to repeated denoiser evaluations. Feature caching offers a deployment-friendly acceleration strategy by reusing or predicting intermediate representations without retraining the generator or modifying the sampler. However, existing cache-then-forecast methods often rely on Taylor-style finite-difference extrapolation, which can become unstable over longer cache intervals, and typically predict local module outputs whose errors may accumulate through subsequent denoiser blocks. We propose \textbf{BeziCast}, a training-free diffusion acceleration framework that forecasts output-proximal denoising representations using low-order B\'ezier trajectories. BeziCast estimates B\'ezier control points via Tikhonov-stabilized fitting, providing a smooth, capacity-controlled temporal parameterization that avoids explicit high-order derivative estimation and decouples trajectory capacity from cache-query density. To moderate aggressive extrapolation, we further derive equivalent forecasting weights and introduce a convex-hull-inspired guardrail that detects high-risk predictions and softly pulls them toward a conservative simplex-projected estimate. Extensive evaluations across advanced image and video diffusion models demonstrate the effectiveness of BeziCast. In particular, BeziCast accelerates FLUX.1 by up to (4.79\times) and HunyuanVideo by up to (4.11\times), while preserving substantially better generation quality than competing acceleration baselines.


TILT: Target-induced loss tilting under covariate shift

Kakei Yamamoto ⋅ Martin Wainwright

We introduce and analyze Target-Induced Loss Tilting (TILT) for unsupervised domain adaptation under covariate shift. It is based on a novel objective function that decomposes the source predictor as $f+b$, fits $f+b$ on labeled source data while simultaneously penalizing the auxiliary component $b$ on unlabeled target inputs. The resulting fit $f$ is deployed as the final target predictor. At the population level, we show that this target-side penalty implicitly induces relative importance weighting at the population level, but in terms of an estimand $b^*_f$ that is *self-localized* to the current error, and remains *uniformly bounded* for any source-target pair (even those with disjoint supports). We prove a general finite-sample oracle inequality on the excess risk, and use it to give an end-to-end guarantee for training with sparse ReLU networks. Experiments on controlled regression problems and shifted CIFAR-100 distillation show that \tilt improves target-domain performance over source-only training, exact importance weighting, and relative density-ratio baselines, with a stable dependence on the regularization parameter.


TimeClaw: A Time-Series AI Agent with Exploratory Execution Learning

Hangchen Liu ⋅ Dongyuan Li ⋅ Renhe Jiang ⋅ Jiewen Deng ⋅ Weiwei Ye ⋅ Yoshihide Sekimoto

Time series analysis underpins forecasting, monitoring, and decision making in domains such as finance and weather, where solving a task often requires both numerical accuracy and contextual reasoning. Recent progress has moved from specialized neural predictors to approaches built on LLMs and foundation models that can reason over time series inputs and use external tools. However, most such systems remain execution-centric: they focus on solving the current instance but learn little from exploratory execution. This is especially limiting in verifiable numeric settings, where multiple candidate executions and tool-use procedures may all be task-valid yet differ sharply in quantitative quality, and where early success can trigger tool-prior collapse that suppresses further exploration. To address this limitation, we present TimeClaw, an exploratory execution learning framework that turns exploratory execution into reusable hierarchical distilled experience through a four-stage loop: Explore, Compare, Distill, and Reinject. TimeClaw combines metric-supervised exploratory execution learning, task-aware tool dropout, and hierarchical distilled experience for inference-time reinjection, while keeping the base model frozen and avoiding online test-time adaptation. In an MTBench-aligned evaluation with 17 tasks that span finance and weather prediction and reasoning tasks, TimeClaw delivers consistent gains over the baselines. These results suggest that, for scientific systems, the bottleneck is not only execution-time capability, but how exploratory experience is compared, distilled, and reused.


TimeES: Probabilistic and Deterministic Time Series Forecasting via Evolutionary Spectra

Weiwei Ye ⋅ Renhe Jiang ⋅ Hangchen Liu ⋅ Dongyuan Li ⋅ Yoshihide Sekimoto

Real-world time series are inherently non-stationary, with trends, periodic patterns, and uncertainty evolving over time. While the Fourier domain offers a natural lens to model time series, current deep learning approaches do not explicitly model evolution and randomness in the Fourier spectra, which limits their ability to accurately predict both the expected trajectory and its uncertainty in non-stationary time series. Motivated by Evolutionary Spectra (ES) theory, we propose \textbf{TimeES}, a general framework that enables probabilistic and deterministic forecasting via the evolutionary spectra theory. Specifically, we derive a parameterizable evolutionary spectra formulation, recasting non-stationary random process modeling as learning an evolving representation modulated by random variables. Furthermore, we reduce the complexity of the estimated spectra from $\mathcal{O}(NM)$ to $\mathcal{O}(NK)$, where $K \ll M/2$, by exploiting Hermitian symmetry and energy-guided frequency selection. Based on a simple linear backbone, our proposed TimeES achieves consistent state-of-the-art performance across both deterministic and probabilistic forecasting tasks, with high efficiency and interpretability. Code is available at: \url{https://anonymous.4open.science/r/TimeES}.


Tio: Language Models with Parallel Streams of Thoughts, Inputs and Outputs

Guinan Su ⋅ Yanwu Yang ⋅ Xueyan Li ⋅ Jonas Geiping

The continued improvements in language model capability have unlocked their widespread use as drivers of autonomous agents, for example in coding or computer use applications. However, the core of these systems has not changed much since early instruction-tuned models like ChatGPT. Even advanced AI agents function on message exchange formats, successively exchanging messages with users, systems, with itself (i.e. chain-of-thought) and tools in a single stream of computation. This bottleneck to a single stream as in chat models leads to a number of limitations: the agent cannot act (generate output) while reading, and in reverse, cannot react to new information while writing. Similarly, the agent cannot act while thinking and cannot think while reading or acting on information. In this work, we show that models can be unblocked by switching from instruction-tuning for sequential message formats to instruction-tuning for multiple, parallel streams of computation, splitting each role into a separate stream. Every forward pass of the language model then simultaneously reads from multiple input streams and generates tokens in multiple output streams, all of which causally depend on earlier timesteps. We argue that this data-driven change remedies a number of usability limitations as outlined above, improves model efficiency through parallelization, improves model security through better separation of concerns and improves model monitorability.


TiRex-2: Generalizing TiRex to Multivariate Data and Streaming

Patrick Podest ⋅ Marco Pichler ⋅ Elias Bürger ⋅ Levente Zólyomi ⋅ Bernhard Voggenberger ⋅ Wilhelm F Berghammer ⋅ Daniel Klotz ⋅ Sebastian Böck ⋅ Günter Klambauer ⋅ Sepp Hochreiter

We introduce TiRex-2, a recurrent xLSTM-based time series foundation model that generalizes the univariate TiRex to multivariate forecasting with both past and future covariates. Real-world forecasting is inherently sequential: observations arrive continuously, variables evolve jointly, and a subset of covariates is known ahead of time. Existing Transformer-based time series foundation models capture cross-variate dependencies but incur quadratic complexity in context length and require full-history recomputation as new observations arrive. TiRex-2 addresses these limitations through a memory-centric recurrent design that operates at constant per-patch cost under streaming. The model combines a bidirectional time mixer with an asymmetric grouped-attention variate mixer, enabling the integration of future-known covariates while preserving strict causality over target variables. To our knowledge, this is the first time series foundation model that achieves this combination of properties. To support scalable multivariate pretraining, we propose a synthetic coupling pipeline that composes diverse multivariate samples on the fly from large univariate corpora. Empirically, TiRex-2 achieves state-of-the-art zero-shot performance on GIFT-Eval and fev-bench, remains stable when streamed to arbitrary context lengths, and maintains constant inference cost per patch. The model uses 38.4M active parameters in univariate mode, with an additional 44.1M parameters activated for multivariate forecasting.


To Call or Not to Call: Diagnosing Intrinsic Over-Calling Bias in LLM Agents

Wei Shi ⋅ Ziheng Peng ⋅ Sihang Li ⋅ Xiting Wang ⋅ Xiang Wang ⋅ Mengnan Du ⋅ Na Zou

LLM agents exhibit a consistent tendency to over-call, invoking tools even in situations where none is needed. On the When2Call benchmark, six models from three families show high call accuracy but much lower no-call accuracy, leaving overall accuracy in the 55%--70% range. We trace this to an Intrinsic Bias Hypothesis (IBH): the call/no-call decision mapping carries an activation-independent CALL offset, so the model favors CALL even at activation parity. Using Sparse Autoencoders (SAEs), we recover behavior-aligned feature bases for the CALL/NOCALL decision, reduce them to a signed activation margin, and estimate the offset directly. Across all six models, the model is decision-neutral only when NOCALL activation outweighs CALL activation, consistent with IBH. We then causally test IBH with Adaptive Margin-Calibrated Steering (AMCS), a closed-form counter-bias shift along SAE decoder directions. Cancelling the diagnosed offset mitigates over-calling and improves overall accuracy with a negligible drop in call accuracy. Our work recasts over-calling from an empirical phenomenon into a mechanistic object amenable to causal correction. The code is available at https://anonymous.4open.science/r/agent-sae-904F.


To Copy or Not to Copy: Controlling Speculative Decoding via Intrinsic Model Signals

Roy Eisenstadt ⋅ Ido Cohen ⋅ Edo Cohen-Karlik ⋅ Lior Wolf ⋅ Itamar Zimerman

Speculative Decoding (SD) has significantly accelerated Large Language Model (LLM) inference, yet existing approaches face a fundamental tradeoff between two drafting strategies: neural drafting and context-based copying. Neural drafts (e.g., EAGLE3) provide robust performance across diverse text settings, while copy-based methods achieve higher speedups in copy-intensive regimes by generating candidates faster and exploiting long repetition spans for near-perfect speculation. We analyze existing copy-based methods and find that they are prone to accidental repetitions where surface-level $n$-gram overlap does not reflect a structural intent to copy, leading to false-positive triggers that ultimately degrade throughput. We introduce **SwitchSD**, an adaptive framework that treats copying as a **latent control signal** of the LLM. By training lightweight probes on the target model's internal representations, SwitchSD identifies genuine copy-intent with high precision (AUC $>$ 0.99). This allows the system to dynamically switch between neural drafting (e.g., EAGLE) and context-based copying. Our results across Llama and Qwen families demonstrate throughput gains of up to 15% over state-of-the-art baselines like EAGLE3, effectively turning "copying" from a noisy heuristic into a principled, model-aware decoding regime.


TokenCLIP: Token-wise Prompt Learning for Zero-shot Anomaly Detection

Qihang Zhou ⋅ Bin-Bin Gao ⋅ Guansong Pang ⋅ Xin Wang ⋅ Jiming Chen ⋅ Shibo He

Adapting CLIP for anomaly detection on unseen objects has shown strong potential in a zero-shot manner. Existing methods typically rely on a single textual space to align with visual semantics across diverse objects and domains. The indiscriminate alignment hinders the model from accurately capturing anomaly semantics. We propose TokenCLIP, a token-wise dynamic framework that aligns each visual token with learnt textual subspaces corresponding to its visual characteristics. However, explicitly assigning a unique learnable textual space to each token is computationally intractable and prone to insufficient optimization. We instead expand the token-agnostic textual space into a set of orthogonal subspaces, and then dynamically assign each token to a subspace combination guided by semantic affinity, which jointly supports customized and efficient token-wise adaptation. To this end, we formulate dynamic alignment as an optimal transport problem, where all visual tokens in an image are transported to textual subspaces under the cross-modal cost matrix. \textbf{The marginal constraint and minimal cost objective of OT ensure sufficient optimization across subspaces and encourage them to focus on different semantics.} Solving the problem yields a transport plan that adaptively assigns each token to semantically relevant subspaces. Extensive experiments show that the textual subspaces naturally specialize in different semantics, such as foreground and background, to promote fine-grained anomaly learning. The comparison between baselines shows the superiority of TokenCLIP.

Time series transformers have achieved strong forecasting performance, yet fundamental questions about their internal mechanisms remain unanswered: among self-attention and feed-forward networks (FFN), which component constructs temporal features? Why does a simple linear model outperform transformers? Why does patching help? We propose a temporal probing framework consisting of 15 token-level diagnostic tasks to trace information flow across transformer components. Applied to 5 architectures across 8 benchmarks, we find that temporal feature construction is not inherently tied to any single component—it emerges from the interaction between tokenization strategy and architectural design. In PatchTST, FFN dominates feature construction across all 8 datasets, with a $3.2\times$ mean ratio over attention. Even when attention is completely removed, FFN's feature construction capability is preserved or enhanced, indicating that attention is not essential for temporal feature construction. Conversely, iTransformer's variate-level attention dominates over FFN, and CATS's cross-attention is responsible for nearly all feature transfer. Point-wise tokenization ($P=1$) structurally deactivates FFN, explaining why early transformers lagged behind linear models. These findings unify contradictory results across DLinear, PatchTST, attention-free models (TSMixer, PatchMLP), and CATS within a single framework.


TokenRepel: Generating Diverse Image Sets Without Sacrificing Quality

Qingtao Yu ⋅ Zhen Wang ⋅ Ning Lai ⋅ Boyang Hu ⋅ Hongdong Li ⋅ Dylan Campbell

We present an inference-time approach for enhancing the visual diversity of image sets generated by text-to-image (T2I) models, while preserving, or even improving, the average image quality. Our method leverages the intrinsic geometry of hidden latent representations in pretrained models to diversify a batch of images generated under the same prompt. First, we introduce a token repulsion mechanism to encourage diversity by increasing the angle between the tokens of hidden features across different samples. This is achieved via a negative Riemannian gradient step defined over a distribution on the unit hypersphere. Second, to preserve the quality and prompt adherence of the output images, we encourage the updates to remain within the high-density region learned by the generative model. Third, we extend this constraint to the output of each intermediate layer, enabling more fine-grained iterative guidance. We evaluate our method on three representative large-scale pretrained T2I models: Flux1.dev (flow matching), SDXL (diffusion), and Flux2.Klein (few-step distilled flow matching). Experimental results demonstrate substantial improvements in visual diversity under identical prompts, while maintaining averaged individual sample fidelity. Our approach is simple, effective, and broadly applicable. It requires no external models to guide the generation process and introduces only modest additional computation overhead during the early stages of sampling.


TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

Andong Hua ⋅ Colton Bishop ⋅ Igor Mordatch ⋅ Arian Hosseini ⋅ Jindong Gu ⋅ Aleksandra Faust ⋅ Rebecca Roelofs ⋅ Yao Qin

Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy in model predictions under such cross-modal variations. To study this phenomenon, we define the modality gap as the difference in model performance under semantically equivalent textual and multimodal inputs. We introduce TokenSwap, a method that constructs such inputs by replacing textual concepts with semantically aligned images, resulting in sequences where visual tokens are interleaved with text tokens. Based on TokenSwap, we transform existing text-based benchmarks (e.g., MMLU) into image-interleaved counterparts, resulting in TOKENSWAP-BENCH. Across 42 MLLMs, we observe a pervasive modality gap, with performance decreasing by 4.2% to 47.4% when moving from text-only to image-interleaved inputs, averaging 19.6% ± 3.3% across models. Notably, we observe that reasoning models exhibit consistently smaller gaps, achieving an average gap of 10.1% compared to 25.5% for non-reasoning models. In contrast, neither prompting strategies nor scaling training compute alone reliably reduces the modality gap. Finally, we demonstrate that incorporating TokenSwap during training effectively mitigates this gap while preserving strong text-only and vision-language performance.


ToLD: Efficient Time Series Forecasting via Tokenized Truncated Latent Diffusion

Jiayi Tian ⋅ Jiaze Wang ⋅ Wenzhe zhao ⋅ Tian Xia ⋅ Pengju Ren

Probabilistic long-term time series forecasting requires estimating a distribution over future trajectories under practical deployment constraints, where inference latency and stability can be as critical as forecast quality. Existing diffusion-based forecasters provide expressive uncertainty modeling, yet their iterative denoising incurs substantial horizon insensitive runtime overhead, which limits real time long-term use. In contrast, one-shot generative models such as VAEs and flows are computationally efficient but often struggle to capture complex multi-modal structure and heavy tailed uncertainty. We propose \textbf{ToLD}, an efficient time‑series forecasting framework built upon \textbf{to}kenized truncated \textbf{l}atent \textbf{d}iffusion to enable fast and reliable long term generation. ToLD adopts a tokenized conditional latent modeling and performs compact token-wise refinement via a few-step truncated latent diffusion process, where the diffusion noise is injected in a context-dependent manner to model non-stationary uncertainty. To reduce train test mismatch in multi-sample forecasting, we further introduce a score distillation scheme that learns a target-free scoring function for test-time candidate ranking. Experiments on eight real-world datasets show that ToLD improves both point accuracy and probabilistic quality over state-of-the-art, achieving up to \textbf{4.3\%} MSE improvement and \textbf{19.3\%} CRPS reduction on long-term tasks with minimal additional inference overhead.


Tool-Integrated Reasoning via Hierarchical Multi-Agent Reinforcement Learning

Yiwei Dai ⋅ Hengyi Cai ⋅ Yili Wang ⋅ Zexu Sun ⋅ Shuaiqiang Wang ⋅ Yuchen Li ⋅ Xin Wang ⋅ Yi Chang ⋅ Dawei Yin

Tool-Integrated Reasoning (TIR) enables large language models to solve complex tasks by invoking external tools during reasoning. Existing methods typically rely on either single-agent reinforcement learning for unified optimization or multi-agent systems for explicit role decomposition. However, single-agent RL often entangles high-level planning with low-level execution, while multi-agent systems may suffer from role mismatch when different roles are optimized independently. To address this limitation, we introduce the principle of decoupled inference with hierarchical alignment: high-level planning and low-level execution should remain separated during inference, while their policies should be aligned through shared task feedback during training. We instantiate this principle as HATR, a Hierarchically Aligned framework for Tool Reasoning. At inference time, HATR represents each reasoning turn as a tree-structured decision process, where a high-level planning role determines the action branch and step goal, and a low-level execution role realizes this decision through tool use or internal reasoning. During training, HATR performs hierarchical multi-agent reinforcement learning by sampling online trajectories, deriving preference signals from task-level rewards, and alternately optimizing policies across decision levels. This cross-level optimization aligns roles across the multi-agent hierarchy while preserving their functional boundaries. Experiments on mathematical reasoning and question-answering benchmarks across multiple backbones show that HATR consistently improves task performance over single-agent and multi-agent baselines.


Topological Invariance and Breakdown in Learning Dynamics

Yongyi Yang ⋅ Tomaso Poggio ⋅ Isaac Chuang ⋅ Liu Ziyin

While mainstream theories of deep learning focus on training with a small learning rate, a growing body of theoretical and empirical work suggests that neural network training dynamics can be qualitatively different at large learning rates. However, a clear and general understanding of the fundamental differences between training at small and large learning rates remains lacking. In this work, we prove that if a group of parameter vectors (neurons) in a neural network model exhibits permutation symmetry, then widely used training algorithms induce a continuous mapping on the set formed by these vectors. Moreover, when the learning rate is below a certain threshold, this mapping becomes a homeomorphism. Our result reveals a critical point of the learning rate: below it, the training dynamics are guaranteed to preserve the topology of the set of neurons, whereas above it, training may introduce simplifications to the underlying neuron manifold. This provides a topological description of a novel implicit regularization effect of small learning rates and offers a potential theoretical lens for empirical phenomena such as loss of plasticity. Notably, our theory is independent of specific network architectures and loss functions, enabling topology to be applied universally to deep learning theory.


Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction

Georg Trede ⋅ Charlotte Doll ⋅ Elias Daniel Weber ⋅ Daniel Durstewitz

Predicting the behavior of dynamical systems (DS) beyond the dynamical and parameter regimes observed in training is a pivotal and essentially unresolved problem in scientific ML. It is central to any good scientific theory, which we expect to be able to make predictions which are not covered by currently available data. Recent hierarchical and hyper-network guided approaches for DS reconstruction (DSR) enable training on many DS simultaneously, and revealed that extracted latent features are often related to crucial control parameters of the underlying DS that varied across the training corpus. However, true out-of-domain forecasting abilities of these models, e.g. across tipping points, remain limited, and fine-tuning, or even full model retraining, on time series from the new dynamical regime is usually required. Here we mathematically analyze the root of these limitations in previous model formulations and identify three core shortcomings rooted in a mismatch between structural assumptions of the reconstruction model and typical properties of physical systems. We propose a combination of remedies for these shortcomings, most importantly feature splitting, and furthermore derive a closed-form bound on the reliable extrapolation range. We demonstrate empirically that our techniques allow for accurate zero-shot prediction into new dynamical regimes, outside the observed training regime, as, e.g., encountered across tipping points.

Building segmentation has recently benefited from foundation segmentation models such as SAM, which offer strong generalization and scalable deployment through simple prompts. However, SAM-style models still produce pixel-wise masks, whose fragmented regions and irregular boundaries often fail to preserve the geometric regularity of buildings, making direct contour vectorization unreliable. Existing attempts to improve building polygonization typically modify the segmentation decoder or introduce additional polygon prediction branches, increasing adaptation costs and weakening the portability of pretrained foundation models. In this paper, we ask whether accurate building polygonization can be achieved directly from frozen segmentation outputs, without modifying or retraining the underlying model. To this end, we propose TopoRefine, a plug-and-play topology-aware contour refinement framework that converts predicted contours into geometry-consistent building polygons. Unlike conventional contour refinement methods, which follow a fixed-topology paradigm and can only adjust a predefined set of contour points, TopoRefine performs topology-adaptive polygon evolution by jointly modeling vertex insertion, vertex deletion, and coordinate refinement. This allows the refined polygon to correct structural errors inherited from frozen segmentation outputs and progressively align with the underlying building geometry. Extensive experiments on three building segmentation benchmarks demonstrate that TopoRefine generalizes across different segmentation pipelines and consistently improves both segmentation accuracy and geometric consistency, achieving notable gains such as +6.17\% AP and +3.53\% PolySim on SAM 3.


Touch-R1: Reinforcing Touch Reasoning in MLLMs

Yingxin Lai ⋅ Yafei Zhou ⋅ Fucai Zhu ⋅ Siyu Zhu ⋅ Weihao Yuan

While rule-based reinforcement learning has recently catalyzed explicit reasoning in multimodal models, tactile reasoning remains largely underexplored. Existing tactile-language models primarily rely on supervised or contrastive objectives, which limits their capacity to ground predictions in physical evidence or rectify misleading visual priors. Tactile reasoning introduces two modality-specific challenges: the ordinal nature of physical attributes (e.g., hardness, roughness) and the cross-sensor distribution shifts inherent in optical tactile hardware. In this work, we introduce TouchReason-1M, a large-scale multimodal dataset comprising over 1M synchronized tactile pairs across four distinct sensors, and TouchReason-Bench, a rigorous framework for evaluating tactile perception and visual-tactile conflict resolution. Building upon these, we propose Touch-R1, a tactile reasoning MLLM based on Qwen2.5-VL-7B. Touch-R1 is trained via a tactile-grounded GRPO objective that combines ordinal-aware accuracy, cross-sensor physical consistency, structured-format control, and an input-side tactile grounding objective. Specifically, the tactile-use reward assigns credit only when authentic tactile inputs yield superior correctness relative to counterfactual controls where the tactile stream is removed, shuffled, or noise-masked. On TouchReason-Bench, Touch-R1-7B outperforms Octopi-13B by 18.4\% and GPT-4o by 24.7\% on average. Its structured reasoning traces reveal emergent behaviors of probing, comparison, and revision, demonstrating that R1-style reasoning can be effectively grounded in physical contact. Our code and data will be made public.


Toward Multimodal Sheet Music Recognition and Understanding

Guang Yang ⋅ Brian Zheng ⋅ Victoria Ebert ⋅ Noah Smith

We propose a novel pipeline for extracting symbolic notation and semantic knowledge from images of sheet music. Our pipeline features the first large-scale neural model for optical music recognition (OMR) to operate sequentially on a system-by-system basis, following the horizontal lines of notation as they are read on the page, rather than treating the page as an undifferentiated image, enabling better scaling to arbitrarily long inputs. It is also the first OMR model capable of generating symbolic transcriptions that include embedded textual content, such as titles and annotations. The pipeline combines system-level segmentation with an autoregressive vision-LM to capture both local notation details and score structure. Across multiple datasets, our approach consistently outperforms prior state of the art. We also show that symbolic transcriptions complement visual inputs for frontier language models, improving their interpretation of dense musical documents. The result is new state-of-the-art performance in both OMR and downstream sheet music understanding. We will release data and code upon publication.


Towards Anytime-Valid Statistical Watermarking

Baihe Huang ⋅ Eric Xu ⋅ Kannan Ramchandran ⋅ Jiantao Jiao ⋅ Michael Jordan

The proliferation of Large Language Models (LLMs) necessitates efficient mechanisms to distinguish machine-generated content from human text. While statistical watermarking has emerged as a promising solution, existing methods suffer from two critical limitations: the lack of a principled approach for selecting sampling distributions and the reliance on fixed-horizon hypothesis testing, which precludes valid early stopping. In this paper, we bridge this gap by developing the first e-value-based watermarking framework, Anchored E-Watermarking, that unifies optimal sampling with anytime-valid inference. Unlike traditional approaches where optional stopping invalidates Type-I error guarantees, our framework enables valid, anytime-inference by constructing a test supermartingale for the detection process. By leveraging an anchor distribution to approximate the target model, we characterize the optimal e-value with respect to the worst-case log-growth rate and derive the optimal expected stopping time. Our theoretical claims are substantiated by simulations and evaluations on established benchmarks, showing that our framework can significantly enhance sample efficiency, reducing the average token budget required for detection by 13-15\% relative to state-of-the-art baselines.


Towards a Unified Model for Flexible Job Shop Scheduling Problems

Inguk Choi ⋅ Woo-Jin Shin ⋅ Sang-Hyun Cho ⋅ Hyun-Jung Kim

Deep learning-based heuristics have attracted significant attention for solving the Flexible Job Shop Scheduling Problem (FJSSP). However, existing studies typically overlook operational constraints prevalent in real-world industry or target only a specific FJSSP variant, limiting their practical applicability. To address these limitations, we introduce SchedFormer, a generalist agent capable of solving FJSSPs with diverse constraints in a zero-shot manner. Concretely, to handle a range of variants in the same manner, we propose a Unified State Representation (USR) that 1) maps different types of constraints into a single feature and 2) captures the invariant machine-operation relationships across variants via four distinct graphs. These designs make the model readily generalizable to unseen variants without model redesign or retraining. We further develop a unified neural solver that uses graph-specific Transformer blocks to effectively encode the USR. Each block employs an attention mechanism specialized for its target graph, with the mapped constraint information conditioning the overall representation learning. Extensive experiments on 16 FJSSP variants show that SchedFormer considerably outperforms competitors, including specialized neural solvers, and generalizes strongly to unseen variants, instance sizes, and data distributions. We further show excellent adaptability of SchedFormer to other scheduling problems.


Towards Direct Evaluation of Harness Optimizers via Priority Ranking

Kai Ong ⋅ Minseok Kang ⋅ Dongwook Choi ⋅ Junhee Cho ⋅ Seungwon Lim ⋅ Seungju Kim ⋅ Geunha Jang ⋅ Minwoo Oh ⋅ Bogyung Jeong ⋅ Sunghwan Kim ⋅ Taeyoon Kwon ⋅ Jihye Han ⋅ Juyoung wy ⋅ Jinyoung Yeo

Harness optimization enables automated agent creation by having an optimizer agent to iteratively update the harness of target agents. Despite its success, current studies evaluate optimizers solely by observing target agents' performance gains. This indirect end-improvement evaluation neglects optimizers' actions at intermediate steps, which are often erroneous and hinder agent performance. Therefore, it is unclear whether harness optimization is driven by optimizers' informed update actions or simply trial-and-error. This necessitates direct evaluation of harness optimizers. However, evaluating harness optimizers directly is non-trivial and costly due to the lack of oracle harnesses. To address this, we present a simple, low-cost design to directly evaluate them, namely priority ranking. By asking harness optimizers to rank components (e.g, tools) in a given harness by their potential to improve/hinder agent performance when updated, our design quantifies optimizer ability at the step level without expensive rollouts or manual examination. More importantly, optimizers' ranking performance correlates with their ability to improve agents in actual multi-step harness optimization, establishing priority ranking as a reliable predictor of optimization ability. Priority ranking is enabled by SHOR, a collection of 182 human-verified optimization scenarios spanning across domains, designs, and time stages. Codes and data can be found at https://anonymous.4open.science/r/Harness_Eval-6D86/%7D%7D.


Towards Financial World Modeling

Humzah Merchant ⋅ Alec Guthrie ⋅ Simon Mahns ⋅ Randall Balestriero ⋅ Bradford Levy

Financial markets are complex, noisy environments which present unique challenges for representation learning and world modeling. In this study, we systematically explore the application of supervised and self-supervised representation learning methods to financial markets data and their ability to learn world models of financial markets. To support this, we assemble and release $\textbf{Market-1T}$ a dataset of more than one trillion observations spanning all US equities from 2008 through 2025. We then develop a domain-specific data augmentation and apply a variety of modern SSL objectives including Joint Embedding Predictive Architecture (JEPA), Masked Autoencoder (MAE), and DINO. Our results highlight that supervised methods are still dominant in this domain when applied directly to core quantitative finance tasks: predicting changes in prices, volatility, and transactions cost. Further, when supervised models are trained across tasks they are able to leverage complementarities which enhance cross-task performance. While performance of SSL-based methods on these core tasks lags behind, we find that they better capture latent structure---such as time and assets effects---which supervised models miss. Our results suggest current SSL objectives capture broad features of markets but further work is needed to close the gap between supervised methods.

Data Shapley has emerged as a premier framework for data valuation, yet its practical utility is severely hampered by exponential computational complexity and the need for costly model retraining. Unlike existing methods that compute values independently for each dataset, we propose GenSHAP, a novel graph-based paradigm that reformulates data valuation as a transferable learning task. GenSHAP learns an amortised deep model capable of predicting Data Shapley values for any unseen classification dataset in a single forward pass. Our framework utilises a relation-based dataset representation method and a hybrid GIN-Transformer architecture to capture local and global inter-sample dependencies. We further introduce GenSHAP-S, which augments these structural representations with low-budget marginal contribution probes to achieve high-fidelity valuation. Extensive experiments demonstrate that GenSHAP closely approximates high-budget permutation-Shapley reference estimates and excels in downstream tasks such as data removal and addition while reducing valuation time from hours to seconds. The source code is available in the anonymous repository~\footnote{\url{https://anonymous.4open.science/r/GenSHAP/}}.


Towards Identifiable Latent Additive Noise Models

Yuhang Liu ⋅ Zhen Zhang ⋅ Dong Gong ⋅ Erdun Gao ⋅ Biwei Huang ⋅ Mingming Gong ⋅ Anton van den Hengel ⋅ Kun Zhang ⋅ Prof Javen Qinfeng Shi

Causal representation learning (CRL) offers the promise of uncovering the underlying causal model by which observed data was generated, but the practical applicability of existing methods remains limited by the strong assumptions required for identifiability and by challenges in applying them to real-world settings. Most current approaches are applicable only to relatively restrictive model classes, such as linear or polynomial models, which limits their flexibility and robustness in practice. One promising approach to this problem seeks to address these issues by leveraging changes in causal influences among latent variables. In this vein we propose a more general and relaxed framework than typically applied, formulated by imposing constraints on the function classes applied. Within this framework, we establish partial identifiability results under weaker conditions, including scenarios where only a subset of causal influences change. We then extend our analysis to a broader class of latent post-nonlinear models. Building on these theoretical insights, we develop a flexible method for learning latent causal representations. We demonstrate the effectiveness of our approach on synthetic and semi-synthetic datasets, and further showcase its applicability in a case study on human motion analysis, a complex real-world domain that also highlights the potential to broaden the practical reach of identifiable CRL models.


Towards Multi-Human-Value Alignment via Value Localization in LLMs

Xueqi Ma ⋅ Yanbei Jiang ⋅ Xingjun Ma ⋅ James Bailey ⋅ Sarah Erfani

Large Language Models (LLMs) are increasingly deployed in real-world settings where alignment with diverse human values is essential. However, existing alignment methods are often costly, obscure underlying value heterogeneity, and offer limited interpretability. In this work, we move toward multi-human-value alignment through inference-time intervention by investigating how different values are internally represented within LLMs. We propose a probing-based value localization framework that identifies value-sensitive components, i.e., value heads, which mediate model behavior across multiple moral and normative dimensions such as harmlessness, honesty, and helpfulness. Analyses across multiple LLM families reveal that value heads are universally sparse and exhibit structured interactions, including conflicts among certain values. These components demonstrate clear functional specialization: selectively ablating value heads induces substantial and value-specific behavioral changes. Building on these findings, we introduce an inference-time intervention strategy that enables controllable adaptation of LLMs to single or multiple human value systems, with theoretical support for its effectiveness. Experiments on human-value benchmarks demonstrate that our method improves flexible and faithful value alignment while maintaining overall model performance, highlighting mechanistic interpretability as a foundation for socially aware, pluralistic AI.


Towards Optimal Pre-training Data Mixtures for Empowering Reinforcement Learning in LLMs

Xue Han ⋅ Qian Hu ⋅ Yitong Wang ⋅ WenChun Gao ⋅ Lianlian Zhang ⋅ Qing Wang ⋅ Lijun Mei ⋅ Zhaoxuefeng ⋅ Junlan Feng

Optimal pre-training data mixtures are vital for Large Language Models (LLMs), yet post-RL reasoning performance varies significantly across base models. This suggests that identifying "RL-friendly" pre-training data mixtures is a critical but under-explored prerequisite. This motivates us to ask: Which specific data mixtures in pre-training enhance post-RL effectiveness, and how can we evaluate a base model's suitability for RL fine-tuning? Through large-scale experiments with proxy models, we observe that (1) domain-specific data correlation strongly predicts post-RL performance, and (2) the average perplexity of the base model on $k$ RL-generated responses, $\text{avg}(k\text{-ppl})$, is a more stable and correlated predictor than traditional metrics. Building on these insights, we propose BridgeRL, a framework that treats finding optimal mixing ratios as a regression task, using $\text{avg}(k\text{-ppl})$ as the fitting objective. We train 1M/60M-parameter proxy models for regression and scale the findings to 1B/3B-parameter models. Results across online and offline RL algorithms demonstrate that BridgeRL-optimized mixtures consistently yield superior base models for RL fine-tuning.


Towards Principled Task Grouping for Multi-Task Learning

Chenguang Wang ⋅ Xuanhao Pan ⋅ Tianshu Yu

Multi-task learning (MTL) aims to leverage shared information among tasks to improve learning efficiency and accuracy. However, MTL often struggles to effectively manage positive and negative transfer between tasks, which can hinder performance improvements. Task grouping addresses this challenge by organizing tasks into meaningful clusters, maximizing beneficial transfer while minimizing detrimental interactions. This paper introduces a principled approach to task grouping in MTL, advancing beyond existing methods by addressing key theoretical and practical limitations. Unlike prior studies, our method offers a theoretically grounded approach that does not depend on restrictive assumptions for constructing transfer gains. We also present a flexible mathematical programming formulation that accommodates a wide range of resource constraints, thereby enhancing its versatility. Experimental results across diverse domains, including computer vision datasets, combinatorial optimization benchmarks, and time series tasks, demonstrate the superiority of our method over extensive baselines, thereby validating its effectiveness and general applicability in MTL without sacrificing efficiency. Code is available at \url{https://github.com/LOGO-CUHKSZ/Principled-Task-Grouping}.


Towards Real-Time Full-Waveform LiDAR Transformers via Intensity-Guided Token Reduction and Physics-Aware Augmentation

Kazuma Ikeda ⋅ Kotaro Oishi ⋅ Ryosei Hara ⋅ Ryo Yoshida ⋅ Mariko Isogawa ⋅ Kentaro Yoshioka

LiDAR is a critical sensor for autonomous driving and robotics, yet it remains vulnerable to adverse conditions such as fog, rain, transparent objects, and low-reflectance surfaces. Full-waveform LiDAR (FWL) addresses these limitations by capturing the complete return waveform as a histogram, preserving rich reflection characteristics including multi-path echoes and weak signals. While recent Transformer-based approaches have demonstrated promise on FWL data, two fundamental bottlenecks remain: high computational cost from redundant tokens in low-intensity background regions, and limited generalization due to scarce annotated training data. We address both bottlenecks by exploiting the physical structure inherent to FWL data. First, we propose FWL-ToPM, a token pruning and merging mechanism that leverages intensity distributions to selectively reduce tokens while retaining critical waveform content. Second, we introduce FWLAug, a physically consistent data augmentation framework that preserves time-of-flight waveform structure and frustum geometry to improve generalization across diverse environments. Evaluated on the Ghost-FWL benchmark, FWL-ToPM achieves over 8$\times$ speedup at real-time throughput alongside a nearly 9-point F1 gain, establishing the first real-time FWL Transformer that simultaneously improves accuracy. FWLAug further improves generalization to both standard and out-of-distribution environments. Together, these methods push FWL-based perception toward practical deployment.


Towards Scalable Data Diversification for Language Model Pretraining via Leverage Score Sampling

Zailin Ma ⋅ Quzhe Huang ⋅ Yujun Li ⋅ Congyuan Rao ⋅ Yaodong Yang

Data selection for language model pretraining faces a fundamental tension between quality and diversity. While quality filtering is empirically effective, it often induces diversity collapse: by favoring texts similar to high-quality reference corpora (e.g., educational or QA-style data), it systematically excludes valuable data from underrepresented domains. In contrast, diversified selection preserves domain balance and encourages robust downstream performance, yet existing methods either focus on coverage-oriented objectives that indirectly enhance diversity, or directly optimize for diversity via costly covariance matrix recomputation that limits scalability. To address these issues, we introduce \textbf{Leverage Score Sampling (Lev)}, which iteratively selects samples that maximally expand the determinantal volume of the embedded data via leverage scores, a computationally efficient criterion that eliminates matrix recomputation and enables scalable selection. Empirically, Lev delivers up to $65\times$ speedup and improves dataset diversity, measured by the Vendi score, by $9.2\%$ over the strong diversification baseline \textbf{DiSF}. On CommonCrawl (CC) web data selection, Lev improves accuracy across seven downstream tasks by up to $1.31\%$ over existing baselines. For domains where robust quality criteria are inherently difficult to define (e.g., code), Lev serves as an effective unsupervised curation alternative: on StarCoderData, the selected subset reduces bits-per-byte by $3.08\%$ over DiSF. Notably, we uncover a cross-domain collapse of quality filtering: CC data filtered by DCLM-fastText fail to retain sufficient code-related content, yielding inferior code performance relative to Lev-selected data. These findings advocate for integrating diversity-aware practices into quality filtering for more effective data curation in language model pretraining.


TRACE: Structure-Aware Character Encoding for Robust and Generalizable Document Watermarking

Jiale Meng ⋅ Jie Zhang ⋅ Runyi Hu ⋅ Zhe-Ming Lu ⋅ Tianwei Zhang ⋅ Yiming Li

We propose TRACE, a structure-aware framework leveraging diffusion models for localized character encoding to embed data. Unlike existing methods that rely on edge features or pre-defined codebooks, TRACE exploits character structures that provide inherent resistance to noise interference due to their stability and unified representation across diverse characters. Our framework comprises three key components: (1) adaptive diffusion initialization that automatically identifies handle points, target points, and editing regions through specialized algorithms including movement probability estimator (MPE), target point estimation (TPE) and mask drawing model (MDM), (2) guided diffusion encoding for precise movement of selected point, and (3) masked region replacement with a specialized loss function to minimize feature alterations after the diffusion process. Comprehensive experiments demonstrate TRACE's superior performance over state-of-the-art methods, achieving more than 5 dB improvement in PSNR and 5% higher extraction accuracy following cross-media transmission. TRACE achieves broad generalizability across multiple languages and fonts, making it particularly suitable for practical document security applications.


TRACE: Tourism Recommendation with Accountable Citation Evidence

Zixu Zhao ⋅ SIJIN WANG ⋅ Yu Hou ⋅ YUANYUAN XU ⋅ Yufan Sheng ⋅ Xike Xie ⋅ Wenjie Zhang ⋅ Won-Yong Shin ⋅ Xin Cao

Tourism is a high-stakes setting for conversational recommender systems (CRS): a plausible-sounding suggestion can waste real money and trip time once a traveler acts on it. Existing CRS benchmarks primarily evaluate systems with a single Recall@$k$ score over entity mentions, and tourism-specific resources add spatial or knowledge-graph context, yet none of them couple multi-turn recommendation with verbatim review-span evidence and rejection recovery. This leaves an evaluation gap for tourism recommendation that is simultaneously *trustworthy*, *verifiable*, and *adaptive*: recommend the right point of interest (POI) for multi-aspect preferences (such as cuisine, price, atmosphere, walking distance), justify each suggestion with verifiable evidence from prior visitors so the traveler can act without trial and error, and recover when the first recommendation is rejected mid-dialogue. We introduce **TRACE**, where each item is *a multi-turn tourism recommendation dialogue with review-span citations and explicit rejection turns*: 10,000 dialogues over 2,400 Yelp POIs and 34,208 reviews across eight U.S. cities, paired with 14 retrieval, planning, and LLM baselines, along with 25 metrics organized under *Accuracy*, *Grounding*, and *Recovery*. Across these baselines, TRACE reveals the **Three-Competency Gap**: LLM Zero-Shot leads in closed-set Recall@1 and rejection recovery but cites less densely than retrievers; non-LLM retrievers achieve surface-verbatim grounding but with low accuracy; Multi-Review Synthesis fails on Recovery. The Grounding Score agrees with human citation precision (Spearman $\rho = +0.80$, $p < 10^{-20}$), and paired $t$-tests reproduce the per-baseline ranking ($p < 0.01$ on the dominant contrasts). TRACE reframes accountable tourism recommendation as a joint target (right POI, verifiable evidence, adaptive repair) rather than a single-axis leaderboard.

Many modern Language Model (LM) pipelines return an averaged model, such as an exponential moving average of the training iterates, rather than the final iterate itself. This raises a fundamental question: given that we will return an iterate average, how should we change training to improve the performance of this average? We study this question by formulating optimizer design for the iterate-average estimator as an optimal-control problem. In a continuous-time stochastic quadratic model, we solve for the control strategy that minimizes the error of the returned average subject to a penalty on the size of the intervention. A practical approximation to this controller yields PACE, a lightweight wrapper around AdamW that pulls the live weights toward their exponential moving average with a clipped, per-coordinate control strength. We prove that a stylized version of PACE converges at the standard stochastic convex optimization rate, up to a factor depending on the averaging rule, while in the quadratic setting it can strictly improve the limiting squared error of the iterate-average estimator and can do so by an arbitrarily large factor on some instances. Empirically, our initial results suggest that PACE improves over AdamW and EMA-evaluated AdamW in supervised fine-tuning of 1-2B parameter LMs and in GPT-2 pretraining on FineWeb, while remaining competitive with learning-rate decay and Schedule-Free baselines.


Training ML Models with Predictable Failures

Will Schwarzer ⋅ Scott Niekum

Estimating how often an ML model will fail at deployment scale is central to pre-deployment safety assessment, but a feasible evaluation set is rarely large enough to observe the failures that matter. Jones et al. (2025) address this by extrapolating from the largest k failure scores in an evaluation set to predict deployment-scale failure rates. We give a finite-k decomposition of this estimator's forecast error and show that it has a built-in bias toward over-prediction in the typical case, which is the safety-favorable direction. This bias is offset when the evaluation set misses a rare high-failure mode that the deploy set contains, leaving the forecast to under-predict at deployment scale. We propose a fine-tuning objective, the forecastability loss, that addresses this failure mode. In two proof-of-concept experiments, a language-model password game and an RL gridworld, fine-tuning substantially reduces held-out forecast error while preserving primary-task capability and achieving safety similar to that of supervised baselines.


Trajectory-Matching Meta Pseudo-Labeling for Semi-Supervised Learning

Minh Duc Le ⋅ Minh-Duong Nguyen ⋅ Dung Le

Meta pseudo-labeling improves semi-supervised learning by updating a teacher according to whether its pseudo-labels help a student improve on labeled data. Existing one-step meta-gradient methods, however, differentiate through the student update, introducing mixed higher-order derivatives, extra memory cost, and noisy minibatch meta-gradients. We propose Trajectory-Matching Meta Pseudo-Labeling (TMPL), a first-order alternative that extracts labeled-data feedback from recent student optimization trajectories. We show that the one-step meta objective is locally equivalent, up to second-order terms in the student step size, to maximizing alignment between the outer gradient and the pseudo-label-induced student gradient. TMPL approximates this alignment without differentiating through the student update by matching finite differences of the outer objective and the pseudo-label-induced student objective along recent student displacements. Its implementation stores only a FIFO cache of detached scalar objective values, rather than student computation graphs or past student networks. We prove that small trajectory-matching error controls the gradient mismatch within the span of recent displacements, with an explicit residual for unexplored directions. On CIFAR10, CIFAR100, SVHN, and STL10, TMPL improves over MPL in all eight evaluated label regimes, obtains the best result in five settings, and remains competitive with strong SSL baselines. On CIFAR100 with 2500 labels, TMPL reduces peak memory by 39.7\% relative to MPL and 54.0\% relative to a raw differentiable meta-gradient implementation. Ablations identify the directional trajectory-matching term as the main source of improvement, while alignment diagnostics show that TMPL recovers MPL-like gradient alignment after a short warm-up.


TrajLoc: Trajectory-Attention Localization for Multi-Object Motion Control

Omer Sela ⋅ Inbar Huberman-Spiegelglas ⋅ Michael Rotman ⋅ Sagie Benaim ⋅ Avi Ben-Cohen

Controlling the motion of multiple objects in image-to-video (I2V) generation requires preserving object identities while enforcing adherence to distinct target trajectories. This becomes particularly challenging as the number of objects increases and their paths intersect or occlude one another. Existing approaches entangle multiple trajectories within a shared, dense conditioning signal, making object-level correspondence difficult to preserve in crowded scenes. We depart from this paradigm and enforce a strict, per object spatial constraint that isolates instances independently. Our method, TrajLoc, achieves this directly within the attention layers by substituting the cross-attention weights of each object token with a Gaussian heatmap centered on its target location at every frame. The same per object token interface carries trajectory and depth through a learned embedding and preserves identity by encoding first frame appearance in place of an object token. Evaluations across six datasets, featuring up to 20 simultaneously controlled objects and out of distribution real world scenes, demonstrate that our method consistently improves both visual fidelity and trajectory adherence. Applied to two architecturally distinct backbones (CogVideoX 5B and WaN 2.1 14B), our approach achieves average gains of +4.3 dB PSNR and a 51\% reduction in trajectory end point error compared to the strongest baselines. Code will be made publicly available upon publication.


Transferability for General Reasoning: An Automated Curriculum for Multi-Domain LLM RL

Yongjin Yang ⋅ Jiarui Liu ⋅ Yinghui He ⋅ Lechen Zhang ⋅ Bernhard Schölkopf ⋅ Zhijing Jin

Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain reasoning suites spanning mathematics, programming, and science. However, the training curriculum—how often each domain is sampled—is typically fixed or hand-tuned, even though reasoning skills transfer unevenly across domains. Existing learnability-based curricula adapt to where the policy is currently improving, but are blind to whether a gradient step on the selected domain benefits the remaining domains. In this paper, we propose Transfer-Aware Curriculum (TAC), a bandit-style online curriculum that prioritizes domains whose updates broadly benefit the rest of the training suite. TAC repurposes signals already produced by RL training: per-domain advantages capture local learnability, and projected gradients—taken from the GRPO step being computed—estimate cross-domain transferability via gradient-geometry alignment, at negligible cost (<1% wall-clock overhead). Across a six-domain reasoning suite, TAC consistently outperforms baselines such as proportional random sampling and a learnability-only bandit, improving macro-averaged accuracy by up to 3.7 points (13% relative) over the latter on both Qwen3-1.7B and Llama3.2-3B. Ablations show performance degrades sharply when the transferability term is removed, and TAC remains robust on imbalanced training mixtures where learnability-only curricula over-commit to dominant domains. Our findings establish cross-domain transferability as a key signal for multi-domain RL curriculum design.


T-ReX: Learning Tile-Reuse Indexes for Structured Model Compression

Haoyi Yang ⋅ Simon Kohaut ⋅ Zihan Ye ⋅ Devendra Singh Dhami ⋅ Kristian Kersting

The rapid proliferation of large-scale machine learning models, such as LLMs, has driven remarkable progress across application domains. Nevertheless, their scale comes with significant deployment challenges and inherent redundancies that remain open challenges to this day. To this end, we propose the Tile-Reuse Index (T-ReX), a structured model compression method that identifies and indexes reusable, contiguous parameter groups (tiles) across a model's architecture. By learning a compact set of matching tiles, we compress the model's memory footprint while preserving its performance. Furthermore, by reconstructing dense layers on the fly at inference time, T-ReX compressed models remain efficient at runtime. In our experiments, we show that T-ReX outperforms existing compression methods across benchmarks by most closely matching the original model's performance. We furthermore show how post-training on top of T-ReX enables small footprint models to recover most of the performance loss introduced by compression.


TRIDENT: Post-Selection Evidence Accountability for Token-Efficient RAG

Wutong Zhang ⋅ Jiong Lou ⋅ Richard Liu ⋅ Sizhe Zhang ⋅ Hefeng Zhou ⋅ Kaixiang Wang ⋅ Celimuge Wu ⋅ Wei Zhao ⋅ Jie LI

RAG systems often verify citations only after retrieval and shortlisting have selected which passages will be tested. A verifier calibrated on random passage--claim pairs is therefore applied to a different population: shortlisted pairs chosen by the retrieval policy. This makes support verification a post-selection inference problem. We introduce TRIDENT, a contract-scoped framework for post-selection evidence accountability. Safe-Cover emits structural-facet support certificates with query-level Bonferroni FWER control, or abstains when the logged contract cannot certify support. Pareto-Knapsack relaxes certification and uses the same frozen facet table as a token-efficient multi-hop QA selector. Across HotpotQA, 2Wiki, and MuSiQue, TRIDENT controls false certificates under deployed contracts, identifies higher-quality answer-linked slices when certificates are emitted, and improves the matched-budget multi-hop QA frontier. The accountable unit is therefore not a retrieved passage or generated citation, but a support claim calibrated under the policy that selected it.


TrustFlow: Adaptive Trust Calibration for Language Model Guided Reinforcement Learning

Hao Chen ⋅ Zhiwei Xu ⋅ Hu Fu ⋅ Yuelin Ma ⋅ Hangyu Mao ⋅ Yiqun Chen ⋅ Bin Zhang ⋅ Guoliang Fan

Reinforcement learning (RL) often suffers from low sample efficiency and inefficient exploration in complex environments. Recent work leverages large language models (LLMs) as sources of prior knowledge for sequential decision-making. However, existing LLM-guided RL approaches typically rely on static or loosely coupled integration, failing to account for the evolving competence of the agent during training. As a result, LLM guidance may become redundant or even detrimental, leading to suboptimal learning dynamics, unnecessary dependence at inference time, and increased latency due to frequent LLM interaction. In this work, we propose \textbf{TrustFlow}, a unified framework that formulates LLM–RL integration as adaptive trust calibration. The key idea is to dynamically regulate the influence of LLM guidance based on learning progress, enabling a gradual transition from prior-driven exploration to LLM-free decision-making. Concretely, a preference representation is first constructed to connect sparse LLM rankings with policy representations. Trust calibration then balances LLM guidance with RL exploration. Building on this trust signal, value alignment is enforced, and a trust-modulated optimization scheme is introduced, incorporating preference structure into the learning process. Empirical results show that TrustFlow consistently improves sample efficiency and enables a smooth transition to fully LLM-free policies, eliminating the need for LLM access at inference time across diverse environments.


Trust Region Q Adjoint Matching

Yonghoon Dong ⋅ Kyungmin Lee ⋅ Changyeon Kim ⋅ Jaehyuk Kim ⋅ Jinwoo Shin

Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step sampling process. Recently, Q-learning with Adjoint Matching (QAM) addressed this issue by reformulating into a memoryless stochastic optimal control (SOC) problem with a learned critic. However, QAM inherits a fundamental fragility of critic-guided improvement: small critic errors are amplified when critics are ill-conditioned, often leading to model collapse. This paper introduces Trust Region Q-Adjoint Matching (TRQAM), a stable off-policy fine-tuning algorithm that adaptively controls the path-space KL with pretrained flow policies through projected dual descent. Specifically, we optimize the trust-region parameter $\lambda$ in SOC dynamics, and theoretically show that the path-space KL can be represented by a closed-form function of $\lambda$. As a result, our method can precisely control the exact deviation from pretrained flow policies, achieving stable off-policy RL. Through experiments on 50 OGBench tasks, TRQAM consistently outperforms prior arts in both offline RL and offline-to-online RL. In particular, TRQAM achieves an overall success rate of 68\% in offline RL, substantially improves the strongest baseline at 46\%.


TSNBench: Benchmarking LLM Proficiency in Time-Sensitive Networking

Rubi Debnath ⋅ Daniel Bujosa ⋅ Luxi Zhao ⋅ Silviu S. Craciunas ⋅ Paul Pop ⋅ Sebastian Steinhorst

We present TSNBench, the first benchmark for evaluating large language model (LLM) proficiency in Time-Sensitive Networking (TSN), a suite of IEEE 802.1 standards for deterministic communication with bounded latency in safety-critical domains such as autonomous vehicles, aviation, defense, and industrial automation. While LLMs have been extensively evaluated on general knowledge tasks, their capabilities in safety-critical networking domains remain largely unexplored. TSNBench comprises 939 expert-validated multiple-choice questions (MCQs) covering diverse TSN mechanisms, along with 100 open-ended Worst-Case Delay (WCD) computation tasks for Credit-Based Shaper (CBS) and Cyclic Queuing and Forwarding (CQF) across varying network topologies and traffic conditions. MCQ answers are validated by domain experts, and open-ended ground truth WCD values are computed using a verified Network Calculus (NC) solver for CBS and closed-form mathematical upper bounds for CQF. We evaluate 16 LLMs and find that although models achieve 67 to 95\% accuracy on MCQs, they fail substantially on open-ended WCD computation. For CBS, only GPT-5 achieves a Mean Absolute Percentage Error (MAPE) of 36.2\%, meaning its predicted WCD deviates by 36.2\% of the actual TSN flow delay on average, while most models exceed 80\%. For CQF, the best model achieves 41.8\% MAPE, with most models clustering between 80\% and 100\%. Such errors are large relative to TSN latency budgets and can lead to violations of real-time constraints and unsafe configurations. TSNBench demonstrates that MCQ benchmarks may overestimate LLM capabilities in safety-critical networking domains.


TTCS: Test-Time Curriculum Synthesis for Self-Evolving

Chengyi Yang ⋅ Zhishang Xiang ⋅ Yunbo Tang ⋅ Zongpei Teng ⋅ Chengsong Huang ⋅ Fei Long ⋅ Yuhan Liu ⋅ Jinsong Su

Test-Time Training offers a promising way to improve the reasoning ability of large language models (LLMs) by adapting the model using only the test questions. However, existing methods struggle with difficult reasoning problems for two reasons: raw test questions are often too difficult to yield high-quality pseudo-labels, and the limited size of test sets makes continuous online updates prone to instability. To address these limitations, we propose TTCS, a co-evolving test-time training framework. Specifically, TTCS initializes two policies from the same pretrained model: a question synthesizer and a reasoning solver. These policies evolve through iterative optimization: the synthesizer generates progressively challenging question variants conditioned on the test questions, creating a structured curriculum tailored to the solver's current capability, while the solver updates itself using self-consistency rewards computed from multiple sampled responses on both original test and synthetic questions. Crucially, the solver's feedback guides the synthesizer to generate questions aligned with the model's current capability, and the generated question variants in turn stabilize the solver's test-time training. Experiments show that TTCS consistently strengthens the reasoning ability on challenging mathematical benchmarks and transfers to general-domain tasks across different LLM backbones, highlighting a scalable path towards dynamically constructing test-time curricula for self-evolving. Our code and implementation details are available at https://anonymous.4open.science/r/TTCS-6ACE.

We develop a Koopman-operator analysis of self-attention that resolves a long-standing puzzle: why does linear attention diverge on long sequences while softmax attention remains stable, even though both share the same next-step prediction objective? We show that attention stability is intrinsically a *two-layer* problem. The *architectural layer* concerns the spectrum of the $n \times n$ token-mixing matrix and is data-independent; the *dynamical layer* concerns the spectrum of the $d \times d$ dual operator induced by training and should be data-determined. We prove three concrete results within this framework: (i) under a shared encoder, the dual operator $C^* = W_Q W_K^\top G W_V$ of linear attention coincides with the EDMD Koopman estimator $K_{\text{EDMD}} = G^{-1}A$ at the global optimum, and its spectral radius is unconstrained; (ii) causal softmax attention enjoys an architectural guarantee $\rho(A^{\text{soft}}) \le 1$ via row-stochasticity, while its value projection $W_V$ leaves the dynamical layer entirely unconstrained; (iii) LayerNorm makes the empirical Gram matrix $G_{\text{LN}}$ rank-deficient with kernel $\text{span}(\mathbf{1})$, producing a $d$-parameter family of loss-equivalent but spectrally distinct EDMD solutions, while RMSNorm preserves invertibility. The framework yields an immediate design principle—constrain the architectural layer, free the dynamical layer—which we instantiate as two stabilized linear-attention variants (SN-LA, RN-LA). Numerical experiments reveal a *hierarchy* of architectural stability: vanilla linear attention has neither bound and diverges immediately under autoregressive rollout; our stabilized variants enforce only the single-pass spectral bound $\rho(\tilde{A}) \le 1$ and delay divergence; softmax attains the strictly stronger pointwise contraction $\Vert A v\Vert_\infty \le \Vert v\Vert_\infty$ via row-stochasticity and remains bounded. On real time-series benchmarks (ETT, Weather), our methods close $\ge 80\%$ of the vanilla-to-softmax gap and surpass softmax on $3/5$ datasets at long horizons, while preserving the dynamical-layer freedom that constraints on the dual operator forfeit.

Sampling from a softmax distribution is a fundamental operation in machine learning, but its linear complexity in the number of items makes exact sampling impractical at scale. Two-level softmax (2LS) sampling is a popular alternative enabling sublinear-time sampling. Assuming items are partitioned into clusters, 2LS first samples a cluster and then an item within it. In this paper, we show that, despite its advantages, 2LS introduces systematic and undesirable sampling biases, which arise from misweighting clusters by ignoring both cluster size imbalance and intra-cluster similarity dispersion. We propose two sampling methods, Size-Corrected 2LS (S-2LS) and Size- and Dispersion-Corrected 2LS (SD-2LS), which correct these biases and provide provably better softmax approximations with negligible to non-existent computational overhead. In-depth experiments on five large-scale datasets validate the improved sampling properties of our methods. We recommend their consistent use in place of standard 2LS in future work.


UECO: A Unified Encoder with Structure-Aware Attention Mixture via Iterative Edge Evolving for Neural Combinatorial Optimization

Wenzheng Pan ⋅ Shuyi Yan ⋅ Nuoyan Chen ⋅ Yiyang Qu ⋅ Jiale Ma ⋅ Junchi Yan

Despite the promise of Neural Combinatorial Optimization (NCO), recent literature has increasingly coupled decision paradigms, training objectives, and problem classes with specific backbone architectures, limiting generic model design and fair cross-paradigm evaluation. We propose UECO, a simple yet effective Unified Encoder for Combinatorial Optimization with a structure-aware attention mixture scheme, where a lightweight module defined upon graph-based locality is interleaved with conventional (global) node-focused attention blocks to dynamically evolve edge representations and inject local topological information into attention scores. Instead of fusing static edge features in a single step, UECO thereby integrates local and global messages in a progressive and problem-agnostic manner. UECO is orthogonal to task-specific decoders, and works seamlessly with Local Construction (LC), Global Prediction (GP), and Adaptive Expansion (AE) paradigms, across node- or edge-oriented problems, symmetric or asymmetric tasks, and sparse or dense graphs, for supervised, reinforcement, or generative learning. Extensive experiments show that, via plug-in substitution of encoders, UECO consistently improves solution and generalization quality over diverse backbones and task scales on TSP, ATSP, CVRP, ACVRP, MIS, MaxClique, and MaxCut.


UFO: A Unified Flow-Oriented Framework for Robust Continual Graph Learning

Danhui Zhang ⋅ Zhe Wang ⋅ Qing Qing ⋅ Jiarui Liu ⋅ Wentao Gao ⋅ Ziqi Xu ⋅ Mingliang Hou ⋅ Xikun Zhang ⋅ Renqiang Luo

Graph learning research has increasingly shifted toward continual graph learning (CGL), which better reflects real-world scenarios where graphs evolve over time. However, existing CGL methods largely assume clean supervision and overlook a critical challenge: the newly arriving portions of the graph are often noisy, due to annotation errors or adversarial corruption. This mismatch limits their applicability in practice. In this work, we study robust continual graph learning, where models must simultaneously handle catastrophic forgetting and noisy supervision in evolving graph data. We show that label noise introduces a new failure mode—catastrophic remembering, where models persistently reinforce corrupted knowledge across tasks. To address these challenges, we propose a Unified Flow-Oriented framework (UFO). First, UFO models conditional feature distributions via flow-based generative modeling and produces replay representations, mitigating forgetting without storing historical data. Second, UFO estimates instance-level reliability scores to distinguish clean from noisy nodes, reducing the impact of corrupted supervision and alleviating catastrophic remembering. Extensive experiments on four benchmark graph datasets under varying noise ratios demonstrate that UFO consistently outperforms existing methods in both accuracy and forgetting metrics. Code is available at: https://anonymous.4open.science/r/UFO.


Ultra Flash: Scaling Real-Time Streaming Video Generation to High Resolutions

Luxury ⋅ Jie Huang ⋅ Zihao Fan ⋅ Xiaoxiao Ma ⋅ Yuming Li ⋅ Siming Fu ⋅ Jun-hao Zhuang ⋅ Zeyue Xue ⋅ Haoran Li ⋅ haoyang huang ⋅ Nan Duan

While recent autoregressive video diffusion models achieve remarkable streaming quality, they remain confined to low resolutions (e.g., $480$P), leaving efficient, scalable, real-time high-resolution video generation a fundamental open challenge. To bridge this gap, we present Ultra Flash, a cascaded streaming framework capable of real-time high-resolution video generation. Ultra Flash achieves ${\sim}30$ FPS at 1K resolution and ${\sim}18$ FPS at 2K resolution on a single B200 GPU through three key contributions: (1) an architecture-preserving T2V-to-TV2V super-resolution training paradigm coupled with an AIGC-oriented data degradation pipeline that effectively preserves the generative capability of the base model, enabling enhanced high-resolution detail when cascaded after mainstream low-resolution generative models; (2) a causal streaming latent upsampler paired with a high-resolution decoder, which enhances spatiotemporal coherence while enabling efficient latent spatial scaling and precise high-resolution decoding with negligible computational overhead; and (3) a cascade high-resolution streaming video generation optimization scheme that first performs hybrid-reward-enhanced sparse causalization and single-step distillation of the super-resolution model, then introduces cascaded streaming self-forcing preference optimization with dynamic cache management, jointly enhancing overall coherence, improving quality, and enabling real-time high-resolution streaming video generation. Extensive experiments demonstrate that Ultra Flash reliably produces ultra-high-resolution streaming video while maintaining state-of-the-art visual quality and superior efficiency.

We present UltraVoxelGS, a 3D ultrasound reconstruction method that achieves both fast reconstruction and physically faithful rendering, bringing live spatial feedback during 2D ultrasound scanning within reach. This is made possible by a feed-forward design that predicts a complete Gaussian scene representation from posed ultrasound images directly, entirely bypassing the per-scene optimization bottleneck that has confined prior methods to offline use. Realizing such feed-forward prediction in this setting is non-trivial: unlike natural images, ultrasound slices share at most one-dimensional intersections and exhibit view-dependent appearance due to acoustic propagation and attenuation, violating the dense overlap and photometric consistency assumed by existing feed-forward Gaussian methods. To achieve fast reconstruction, we introduce a feed-forward voxel-to-Gaussian pass that decouples representation size from input count by lifting posed slices into a fixed-size voxel support and predicting Gaussian parameters directly. For faithful rendering, we propose an ultrasound-aware appearance rendering strategy that factorizes reflection from cumulative attenuation during training and exploits appearance continuity among spatially adjacent slices for efficient Render-time Appearance Adaptation. On both real-world and simulated datasets, UltraVoxelGS achieves the best PSNR/SSIM compared with per-scene optimization baselines, while reducing inference from 10--20 minutes to approximately 10 seconds, over 60× faster, and scaling to sequences exceeding 1,000 frames.


U-MVP: Encode Locally, Decode Globally for Feed-Forward 3D Gaussian Splatting

Seungkwon Yang ⋅ Gyeongjin Kang ⋅ Hwasik Jeong ⋅ Byeongjin Kang ⋅ Hyeongbhin Cho ⋅ Eunbyung Park

Multi-view transformers for feed-forward 3D reconstruction must jointly solve two tasks. They need to extract precise per-view features and reason about cross-view geometric correspondences. Existing architectures apply cross-view attention throughout the network, leaving each layer to handle both objectives together. We revisit this design through a layer-wise probing analysis of feed-forward geometry transformers and observe that trained models tend to divide these two roles across layers. Early layers tend to specialize in per-frame feature extraction, while spatial cognition and cross-view reasoning emerge predominantly in later layers. This suggests that separating the two roles across the network can allocate compute more effectively. Building on this observation, we propose U-MVP, a U-shaped multi-view transformer for feed-forward 3D Gaussian Splatting that reflects this structure in the architecture itself. The encoder applies frame-wise attention to preserve per-view detail, the decoder introduces multi-view attention for cross-view geometric reasoning, and skip connections fuse the two streams so that appearance and geometry are recombined at decoding time. To scale to dense view regimes, we replace several global interaction blocks at the bottleneck with a grouping and swapping scheme that preserves cross-view information flow at lower cost. U-MVP performs competitively with feed-forward and optimization-based baselines from 16 to 256 views, generalizes to unseen datasets, and achieves strong results on 3D reconstruction and novel-view synthesis in posed and unposed settings.

Decades of cognitive science establish that humans navigate environments by forming cognitive maps, defined as allocentric and topology-preserving representations of 3D space. While modern Vision-Language Models (VLMs) demonstrate emergent spatial reasoning from 2D egocentric inputs, it remains unclear whether they construct an analogous 3D internal representation. In this paper, we demonstrate that current VLMs do possess a latent topological map of 3D scenes, but it is heavily overshadowed by non-geometric visual semantics, such as color and shape. By isolating this spatial subspace through cross-scene linear feature extraction, we extract a clean spatial subspace that causally controls the model's spatial outputs. We mathematically shape this latent representation and prove its correspondence to the Laplacian eigenmaps of the scene's 3D Gaussian-kernel graph, converging to the physical 3D space in the continuous limit. Motivated by this geometric identification, we further introduce a mathematically principled latent regularization method for VLMs, based on Dirichlet energy. Applying this single-term regularizer to a minimal 500-step supervised VLM fine-tuning (SFT) on simple synthetic data yields significant improvements on real-world spatial benchmarks, outperforming standard SFT and competitive baselines by up to 12.1\% in spatial tasks involving scene topology understanding.


Uncovering Hidden Propensities in Language Models via Limited-Parameter Finetuning

Boyd Kane ⋅ Jo Jiao ⋅ Bryce Woodworth ⋅ Elizabeth Donoway ⋅ Alex Cloud ⋅ Alexander Turner

Large language models can have hidden propensities: behaviours that arise only in rare contexts. These hidden propensities pose a challenge to alignment audits, which seek to uncover problematic model behaviour but are limited in scope to scenarios anticipated by the auditor. To address this challenge, we explore limited-parameter finetuning (LPF) on examples of a target behaviour. While training 3–15\% of model parameters, LPF reveals whether the model has a hidden propensity, even when standard behavioural evaluations cannot. LPF succeeds on a suite of model organisms with hidden propensities (8B-14B parameters), and also surfaces naturally occurring hidden propensities in open-weight models (27-32B parameters). LPF can thus distinguish between behaviourally equivalent models, even without knowing the context in which the propensity surfaces. LPF provides a novel affordance that may be useful for auditing frontier AI models.


Understanding the Challenges in Iterative Generative Optimization with LLMs

Allen Nie ⋅ Xavier Daull ⋅ Zhiyi Kuang ⋅ Abhinav Akkiraju ⋅ Anish Chaudhuri ⋅ Max Piasevoli ⋅ Ryan Rong ⋅ YuCheng Yuan ⋅ Prerit Choudhary ⋅ Shannon Xiao ⋅ Rasool Fakoor ⋅ Adith Swaminathan ⋅ Ching-An Cheng

Generative optimization uses large language models (LLMs) to iteratively improve artifacts (like code, workflows or prompts) using execution feedback. It is a promising approach to building self-improving agents, yet in practice remains brittle: despite active research, only 9% of surveyed agents used any automated optimization. We argue that this brittleness arises because, to set up a learning loop, an engineer must make "hidden" design choices: What can the optimizer edit and what learning evidence is provided per update? We investigate three factors that affect most applications: the starting artifact, batching execution traces as experiences, and determining a credit horizon with truncated traces. Through case studies in MLAgentBench, Atari, and BigBench Extra Hard, we find that decisions around these factors can determine whether generative optimization succeeds and they are not made explicit in prior works. Different starting artifacts determine which solutions are reachable in MLAgentBench, truncated traces can still improve Atari agents, and larger minibatches do not monotonically improve generalization on BBEH. We conclude that the lack of a simple, universal learning-loop setup across domains is a major hurdle for productionization and adoption, and provide practical guidance for making these design choices.


UniDG: Universal Defect Generation via Defect-Context Editing

Yuanting Fan ⋅ Jun Liu ⋅ Bin-Bin Gao ⋅ Xiaochen Chen ⋅ Yuhuan Lin ⋅ Zhewei Dai ⋅ Jiawei Zhan ⋅ Chengjie Wang

Generating realistic visual defects is crucial for alleviating the scarcity of abnormal samples in anomaly detection, yet existing methods often rely on category-specific few-shot adaptation and struggle to generalize across objects, defect types, and domains. This limitation stems in part from the lack of paired defect editing data and the fine-grained nature of defect appearance, where subtle variations in scale, texture, and morphology must be preserved while maintaining consistency with the target scene. To address this challenge, we introduce UDG, a dataset of 300K normal-abnormal-mask-caption quadruplets curated from diverse defect-related scenarios, and present UniDG, a unified model for universal defect generation. UniDG formulates defect synthesis as Defect-Context Editing: it extracts reference defect context with adaptive cropping, organizes reference and target inputs in a structured diptych format, and fuses multimodal conditions through MM-DiT attention. We further develop a two-stage training strategy: Diversity-SFT learns diverse transferable defect priors from paired editing data, while Consistency-RFT improves reference adherence, local realism, and defect-category consistency. Without per-category fine-tuning or using MVTec-AD/VisA for training, UniDG outperforms prior few-shot anomaly generation and image insertion/editing baselines in both synthesis quality and downstream single- and multi-class anomaly detection/localization.


Unified Resource-Grounded Coordination Protocol for Orchestrator-Free Heterogeneous Multi-Agent Systems

Vishal Pramanik ⋅ Maisha Maliha ⋅ Olivera Kotevska ⋅ Arvind Ramanathan ⋅ Nathaniel D Bastian ⋅ Susmit Jha ⋅ Sumit Kumar Jha

Multi-agent language-model systems typically rely on a human designer or external orchestrator to fix the agent population, roles, and workflow—a single point of failure that places coordination decisions outside the team executing them. Existing approaches address fragments of this problem: communication standards lack coordination logic, automated workflow-search systems concentrate authority in a single planner, and decentralized peers lack explicit governance, handoff validation, and bounded repair. To close this gap, we introduce the Unified Resource-Grounded Coordination Protocol (URCP), an orchestrator-free protocol in which a fixed pool of black-box models and executable resources collectively decides task framing, role formation, workflow topology, assignment, checkpoint validation, and final completion through supermajority voting under a single decision rule Φq. A five-level repair hierarchy reopens only the affected part of the workflow on failure; under standard timing and communication assumptions, every run terminates with an explicit verdict. Empirically, with twelve heterogeneous models against eight baselines on MultiAgentBench, τ-bench, MemoryAgentBench, GAIA, and REALM-Bench, URCP exceeds the strongest baseline by +9.7 / +7.9 / +10.6 points on the three primary benchmarks, terminates deadlock-free in 100% of runs, and remains stable up to 20% agent failure. Trace analyses show that the gains come from voted role concentration on frontier models (73.9% of critical-path roles), frequent critic/verifier insertion (63.5% of accepted schemas with at least three roles), and bounded repair rather than from model strength alone.

Speculative decoding accelerates Large Language Models via draft-then-verify, where verification can be framed as an Optimal Transport (OT) problem. Existing approaches typically handle multi-draft and multi-step aspects in isolation, applying either flat OT to single-step drafts or per-token rejection sampling to tree-structured candidates. This separation leaves the joint regime (where multi-step dependencies meet multi-draft branching) poorly optimized, as local verification rules fail to exploit the coupling between horizontal and vertical dimensions of candidate trees. In this paper, we propose a unified perspective that casts tree-based verification as a conditional OT problem. Our key insight is that vertical dependencies can be abstracted through prefix acceptance probabilities, which act as dynamic scaling factors to actively guide horizontal draft selection. Based on this principle, we introduce UniVer, a verification algorithm that jointly optimizes across tree levels by composing local optimal transport plans under prefix constraints. We prove that UniVer remains lossless and achieves the optimal acceptance rate under the proposed conditional framework. Extensive experiments across different tasks and models demonstrate that UniVer improves acceptance length by 4.2\% to 8.5\% over standard recursive rejection sampling without replacement, while maintaining exact distributional alignment with the target model.


Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation

Ziqi Zhou ⋅ Haowen Jiang ⋅ Yifan Hu ⋅ Xianlong Wang ⋅ Yufei Song ⋅ Shengshan Hu ⋅ Dezhong Yao ⋅ Leo Yu Zhang

The Segment Anything Model (SAM) achieves remarkable performance in visual segmentation. The latest SAM3 extends promptable segmentation to concept-level prediction, broadening the scope of segmentation foundation models. While recent works reveal that SAM and SAM2 are vulnerable to adversarial examples, the robustness of SAM3 under the concept segmentation paradigm remains unexplored. In addition, existing adversarial attacks on SAM-series models exhibit limited cross-prompt transferability. To this end, we propose AdvPCS, a universal cross-prompt adversarial attack for Promptable Concept Segmentation (PCS), including a min-max prompt optimization strategy, a global–local perception deception attack, and a temporal transition deviation attack. Specifically, we first identify the hardest-to-attack prompts via min-max bilevel optimization. In the inner maximization, we enhance diversity over candidate point, box, and text prompts. In the outer minimization, we select prompts with the highest responses based on the confidence scores output by the detector. Given the selected prompts, we apply the perception deception attack to minimize both global and local existence probabilities under joint prompting, and employ the temporal memory misalignment attack to maximize inter-frame semantic inconsistency and corrupt memory pointers. Extensive experiments on four benchmark datasets show that a single universal adversarial perturbation (UAP) generated by our method generalizes across frames from different videos and achieves strong attack performance under point, box, and text prompts. In particular, it reduces the average mIoU of various PCS models on the SA-CO dataset to below 5%, demonstrating strong attack ability.


Unlocking Accurate Geometry in 3DGS via Spatially Varying Affine Rectification

Tze Ho Elden Tse ⋅ Jizong Peng ⋅ Richard Chen ⋅ Angela Yao

3D Gaussian Splatting (3DGS) has become the leading approach for photorealistic novel view synthesis, yet its geometric accuracy lags behind its visual quality. Existing depth-guided methods attempt to resolve this by regularizing the optimization with monocular depth priors, typically using a single global scale-and-shift alignment. Crucially, we discover a fundamental flaw in this assumption: modern monocular depth networks systematically exhibit region-dependent, piecewise-linear distortions. This discovery invalidates standard global alignment strategies, which warp scene geometry by enforcing a single transform. To address this, we introduce a spatially varying affine rectification model that corrects monocular disparity into consistent metric depth using regionally smooth, edge-aware fields. We further derive a normal-aware regularizer that uses spatial gradients to couple the rectified disparity with 3DGS depth and normals. Across diverse experimental settings, our sensor-free method shows strong photometric quality of Gaussian Splatting while substantially improving geometric accuracy and outperforming state-of-the art geometry-focused methods.

Explainable AI (XAI) evaluation fundamentally requires human-centered evaluation, assessing whether explanations align with human reasoning and expert knowledge. However, such evaluations often require expert-driven manual validation and are difficult to scale, leading to the widespread use of alignment-based metrics, where attribution maps are compared against predefined ground-truth (GT) representations. Despite their popularity, alignment evaluations are implicitly bias towards multiple design choices made during the evaluation process, for example GT representations, metric families, metric formulations (soft vs.\ hard), and thresholding strategies. In this paper, we formalize alignment evaluation as a structured evaluation configuration space through which the influence of these design choices can be systematically analyzed. We further introduce a structured alignment evaluation framework that models interactions between XAI methods, GT representations, metric families, metric formulations, and thresholding strategies across 2D and 3D synthetic and real-world datasets. In addition, we propose a novel taxonomy of XAI output behavior based on local entropy and boundary mass ratio, enabling structured characterization of attribution patterns for alignment evaluation. Our experiments reveal systematic flows(triples of XAI-GT-metric configurations), showing that alignment outcomes provide distinct behavior patterns under certain configurations, such as soft/hard metric implementations. We further observe varying levels of threshold sensitivity across metric families and attribution structures. These findings demonstrate that alignment evaluation cannot be treated as a fixed or invariant measure of explanation quality, and provide practical guidance for designing more reliable and behavior-aware XAI evaluation protocols.


Unsupervised Unlearnable Segmentation via Semantic Structure Disruption

Jiale Cai ⋅ Zhihao Li ⋅ Gezheng Xu ⋅ Ruiyi Fang ⋅ Hao Zheng ⋅ Zhimin Mei ⋅ Song Tang ⋅ Charles Ling ⋅ Boyu Wang

Unlearnable examples (UEs) aim to prevent machine learning models from extracting useful information in protected training data. Extending UEs to segmentation is practically important but challenging, because pixel-level annotations are expensive to obtain and therefore cannot be assumed available when generating protective perturbations. This motivates us to study unsupervised unlearnable segmentation, where UE perturbations must be generated without segmentation annotations. Since dense prediction fundamentally relies on region-level affinity and cross-layer feature correspondence, effective segmentation data protection faces a twofold design challenge: perturbations should induce misleading alternatives to these structural priors, while the induced structure must remain easy to fit so that downstream segmentation models can adopt it during training. To address these challenges, we propose DRIFT (Disrupting Region-level and Interlayer Feature sTructure) an unsupervised method that generates unlearnable examples for segmentation. Specifically, DRIFT constructs pseudo region-level affinity and pseudo cross-layer correspondence from clean-image features. It then optimizes perturbations to steer the extracted features of protected images toward the misleading pseudo-structure rather than the clean one, while an auxiliary structure learner, trained to predict pseudo-region assignments, explicitly encourages the induced structure to remain easy to fit. Extensive experiments across four segmentation datasets and multiple architectures demonstrate that DRIFT provides strong protection, consistently degrading the generalization performance of models trained on protected data and remaining competitive with the supervised UE baseline.


Unveiling the Value of Motion for Cinematic Camera Trajectories

Ziqi Zhou ⋅ Yujian Yuan ⋅ Laura Sevilla-Lara

Cinematic camera motion is a fundamental storytelling tool, defined not only by where the camera is positioned in the scene, but also by how it moves in terms of direction and speed. Recent work on camera trajectory generation and alignment to text relies on pose-centric representations. While in principle a network could derive direction of movement and speed, we find that in practice this might not happen. In fact, in this paper we discover that decomposing the camera trajectory representation from the traditional per-frame poses to direction and speed has surprising benefits across multiple tasks, including trajectory-to-text alignment as well as text-to-trajectory generation. To accurately evaluate the former, we introduce a simple and reliable protocol that overcomes the limitations of prior evaluation baselines. For the latter, building on this representational insight, we propose a novel generative model for camera trajectories, CineGen, that achieves superior performance across a variety of metrics. We also propose a novel dataset, CineScript, containing movie clips that are enriched with scene descriptions as well as higher-level metadata. This novel data allows us to test models' ability to capture high-level cinematographic information. We show that, despite its simplicity, representing camera trajectories through direction and speed not only helps numerically to achieve better alignment and generation, but also inherently encodes complex directorial intent.


User Embeddings are Superpositions of Interpretable Behavioral Modes

Jonathan Kozaczuk ⋅ Zheng Dong ⋅ Noyan Evirgen ⋅ Keivan Majidi ⋅ Eddie K Ng ⋅ Chris Porter ⋅ Qi Zhao ⋅ Jia Geng ⋅ Changmao Li ⋅ Brandon Feng

User embeddings are central components of modern recommendation, personalization, and behavioral modeling systems. Though these embeddings can compactly encode rich behavioral information, they are inherently dense and opaque, limiting interpretability. To address this gap, we show that user embeddings can be understood as sparse superpositions of interpretable behavioral modes. This structure is directly predicted by generative models of user behavior for linear embeddings and extends approximately to transformer representations by the superposition hypothesis. Superimposed modes can naturally be recovered by sparse dictionary learning methods, and recovered modes can be automatically named and queried in natural language. On synthetic data with known ground-truth modes, sparse autoencoders (SAEs) outperform classical sparse baselines on mode recovery and produce more monosemantic latents, even when the number of modes exceeds the embedding dimension. Applied to both classical (SVD on open e-commerce) and transformer-based (SASRec/BERT4Rec on MovieLens 1M) user embeddings, the recovered SAE dictionaries are near-orthogonal, and their natural-language labels enable zero-shot audience retrieval over embeddings never trained on text. At matched sparsity, SAE decompositions produce more behaviorally distinct features than classical sparse baselines, with retrieval accuracy comparable to both classical sparse methods and text-embedding baselines and more diverse audience characterizations. The framework is post-hoc and model-agnostic, requiring no modification to the upstream embedding system.


Variational Active Flow Matching for Discrete Online Black-Box Optimization

Yashvir Singh Grewal ⋅ Daniel M Steinberg ⋅ Thang Bui ⋅ Cheng Soon Ong ⋅ Edwin Bonilla

Many scientific design problems require identifying rare high-fitness objects in vast discrete spaces using costly black-box feedback. Variational active-generation methods, including variational search distributions (VSD) and conditioning by adaptive sampling (CbAS), address this by learning a distribution over high-fitness designs via inference on a level-set posterior. However, these methods rely on generators with explicit likelihoods, most commonly autoregressive models, creating a mismatch with modern non-autoregressive approaches, such as discrete diffusion and flow-based models, which refine designs in parallel and better capture multi-modal structure but have implicit marginal likelihoods. We introduce \textit{Active Flow Matching} (AFM), a variational active-generation framework that removes this likelihood bottleneck by reformulating level-set KL objectives over the conditional endpoint distributions of discrete flow models. AFM yields forward-, reverse-, and symmetric-KL objectives without requiring density evaluation. We show that forward-KL AFM is target-consistent, recovering the level-set posterior without access to marginal likelihoods. Empirically, across protein and small-molecule design tasks, AFM outperforms strong baselines, including autoregressive VSD and CbAS, inference-time guidance, KL-regularised fine-tuning, and Top-$K$ retraining.

Mathematical reasoning benchmarks are increasingly limited by static test sets, which are vulnerable to contamination and slow to adapt as models improve, while manually building fresh high-quality problems is expensive. We propose VDE (Verifiable Dynamic Evaluation), a dynamic evaluation framework that generates replay-verifiable math problems from typed value--theorem bipartite graphs. Each instance is an executable graph with model-independent ground truth recovered by deterministic replay rather than model-produced solutions. Beyond compositional depth, VDE supports explicit constraint injection and controlled branching to produce harder problems while preserving verifiability. We instantiate the same framework in analytic geometry and number theory, showing that a unified graph-based generation paradigm can cover structurally different domains. Experiments on frontier models show that performance remains far from saturation and degrades predictably as we increase construction depth, constraints, and branching, making VDE a scalable and trustworthy testbed for dynamic mathematical reasoning evaluation.


Verifiable LLM-Guided Focal SMT Solving for Quantified Arrays

Kunhang Lv ⋅ Rui Han ⋅ Fuqi Jia ⋅ Yuhang Dong ⋅ Feifei Ma ⋅ Jian Zhang

Large language models can often identify useful semantic structure in formal artifacts, but their outputs cannot be trusted as proofs or solver decisions. We study whether LLMs can instead guide formal reasoning systems through a narrow, solver-verifiable interface. We instantiate this idea in quantified array SMT solving, where modern instantiation-based techniques can fail when key ground terms and equalities are hidden by auxiliary variables, axiom guards, and nested array terms. We observe that many instances are largely driven by a small set of focal ground terms or constraints, similar to the backdoor sets in SAT solving. Building on this, we present LinguaArray, a framework where an LLM proposes a small Semantic Focus Set (ground terms/constraints), while a conventional SMT solver remains the sole proof engine and certifies all results. LinguaArray uses these LLM-suggested, solver-validated finite sets to significantly enhance solver capability, solving additional satisfiable and unsatisfiable instances that the backend alone fails to solve. It interacts within a finite time and then falls back to the backend solver to mitigate potential regressions. Evaluation on 1015 SMT-LIB instances across five logics under 1200s timeout (including LLM latency) shows that LinguaArray solves up to 69.1\% of instances and increases solved-instance counts by 105\%--397\% over the corresponding backends.


Verification-Aware Training for Speculative Decoding

Geonmo Gu ⋅ Byeongho Heo ⋅ HeeJae Jun ⋅ Yoohoon Kang ⋅ Sangmin Lee ⋅ Sangdoo Yun ⋅ Dongyoon Han

Speculative decoding accelerates large language model inference by using a draft model to generate candidate tokens, which are verified by the target model in a single forward pass. Verification proceeds and discards every position from the first rejection onward, yet existing draft training relies on token-level cross-entropy with a fixed per-position weighting that does not reflect this process. We introduce Verification-Aware Training (VAT), a plug-in framework that simulates verification at every training step and turns the resulting accept and reject patterns into supervision. VAT consists of two components: (i) a verification head, a lightweight jointly-trainable binary classifier that validates per-position acceptance via an explicit prediction target; (ii) verification-adaptive weighting replaces the fixed weighting schedule with weights adapted to each sample's first rejection point during training. VAT is model-agnostic and can be layered on top of existing methods without changing the draft architecture, the target model, or the inference procedure, and trained jointly with the original loss. Applied to EAGLE-3 and DFlash on Qwen3-4B, Qwen3-8B and LLaMA3.1-8B, VAT improves average acceptance length by up to 11.4% and wall-clock speedup by up to 8.7%, with consistent gains across math, code, and chat benchmarks.


Verification-Guided Abstraction Generation with Large Language Models for Generalized Planning

Zhenhe Cui ⋅ Huaxiang Xia ⋅ Hangjun Shen ⋅ Kailun Luo ⋅ Yong He ⋅ LiangWei

Generalized planning (GP), where a single policy solves multiple instances from a common planning domain, remains a challenging problem in AI. Abstraction plays an important role in GP, and Qualitative Numerical Planning (QNP) provides a compact abstract model for GP. However, useful QNP abstractions are difficult to construct and require substantial manual effort and expertise. We study whether QNP abstractions can be generated from PDDL domains and training instances by coupling LLM generation with symbolic reasoning. We propose a generate--debug--repair framework in which an LLM proposes abstract features and constructs a candidate QNP abstraction over the initial states, actions, and goals of the training instances. To improve reliability, we introduce verification-guided repair based on approximate soundness. Given a candidate abstraction, we solve it, validate the induced abstract instances, and apply one of two checking branches: SRCB tests an $m$-simulation relation between induced abstract and concrete instances to target approximate soundness, while PRCB tests whether the induced abstract plan refines into valid concrete solutions. Detected failures are converted into structured feedback for iterative repair. Experiments on seven GP benchmark domains with five LLMs show that verification-guided repair improves LLM-generated QNP abstractions over both single-pass generation and LLM self-repair, and helps some models produce abstractions that transfer across instances. Our results suggest that LLMs can serve as generators of symbolic abstractions for GP when coupled with formal error detection and structured repair feedback.

Automatic evaluation is now central to text-to-image benchmarking, but the verifier that converts generated images into scores is often treated as an implementation detail. We argue that this is unsafe: an automatic benchmark score is a property of the image generator, prompt distribution, verifier, verifier-use protocol, and scoring rule, not of the generator alone. We study this issue in structural counting, a controlled setting where human-visible counts can be audited and where both instance-level object counts and parent-bound substructure counts arise naturally. We construct a verifier-centric audit suite over seven counting objects: cell phones, chairs, and traffic lights for object counting, and clover leaflets, dresser drawer fronts, power-strip sockets, and turbine blades for substructure counting. Across 6,240 generated images from Qwen-Image-2512 and FLUX.2-dev, we compare open MLLM verifiers under count-only, reject-aware, zero-shot and few-shot protocols, plus TIFA-like target-query variants and targeted frontier closed-model sanity checks on two hard objects. We also compare MLLMs with a detector baseline. On a balanced 2,280-image evaluation split, Qwen3.5-27B with few-shot reject/count achieves the best image-level agreement among open verifiers ($\kappa=0.754$) with a small aggregate score gap ($+1.54\text{pp}$), but changing the protocol, model size, or generator distribution can shift automatic scores by tens of points. Targeted GPT-5.5 and Claude Opus 4.7 checks do not remove the failure modes: GPT-5.5 still has only 59.92\% clover valid-count fidelity, while Claude Opus 4.7 has lower same-slice agreement ($\kappa=0.432$). We propose reporting a factorized verifier reliability profile---aggregate score gap, image-level agreement, valid-image count fidelity, rejection behavior, output compliance, and directional bias---rather than a single automatic benchmark score.


VerifyMAS: Hypothesis Verification for Failure Attribution in LLM Multi-Agent Systems

Hezhe Qiao ⋅ Hanghang Tong ⋅ Ee-peng Lim ⋅ Bing Liu ⋅ Guansong Pang

Large language model-driven multi-agent systems (LLM-MAS) excel at complex tasks, yet unreliable agents remain a key bottleneck to system-level reliability. Automatic failure attribution is therefore critical, but existing approaches—such as direct prediction of agent–error pairs and agent-first failure attribution—rely on local logs of agent and miss global failures that only manifest over full interaction trajectories, such as cross-step inconsistencies and inter-agent coordination errors. Moreover, directly predicting failures induces a large combinatorial search space, hindering fine-grained attribution. To address these challenges, we propose VerifyMAS, a hypothesis verification framework for agent failure attribution. Instead of directly predicting faulty agents and error types, VerifyMAS formulates and verifies failure hypotheses against full trajectories. This verification-based approach decomposes attribution into trajectory-level error validation and fine-grained agent localization, providing an error-first attribution approach that captures global failure patterns while substantially reducing the search space. We further introduce a hypothesis-based data construction strategy grounded in a structured error taxonomy and fine-tune a specialized LLM verifier model for trajectory-level failure verification and agent attribution. Experiments on Aegis-Bench and Who&When show that VerifyMAS consistently improves diverse backbone models, including open-source Qwen and API-based GPT models, outperforming prior methods without sacrificing inference efficiency for long multi-agent trajectories.


VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

Yuting Xu ⋅ Jiayi Tian ⋅ Jian Liang ⋅ Xin Xiong ⋅ Hang Zhang ⋅ Mu Xu ⋅ Xiao-Yu Zhang

Existing benchmarks have laid the foundation for travel planning agents by establishing API-centric paradigms. However, as the capabilities of Autonomous Agents continue to advance, their evaluation must evolve beyond simple tool execution toward handling the inherent complexities of the open web. Current benchmarks bypass core cognitive hurdles: they fail to account for information noise, ignore multi-source factual contradictions, and overlook the necessity of grounding visual perception into logical planning. We introduce VeriTrip, a verifiable benchmark designed to meet the increasing demands for agent robustness and reliability. VeriTrip shifts the evaluation focus to evidence-grounded reasoning over unstructured multimodal web corpora. It establishes a Multimodal Retrieval Base (MRB) derived from real-world sources, forcing agents to autonomously orchestrate queries across heterogeneous data. A synchronized Verifiable Knowledge Base (VKB) enables a cell-wise verification protocol that precisely quantifies factual reliability, distinguishing systematic reasoning failures from parametric hallucinations. Our evaluations across leading MLLMs reveal a critical \textit{retrieval-reasoning trade-off}: the cognitive load of autonomous retrieval significantly erodes instruction retention. VeriTrip provides the rigorous foundation necessary for the next generation of planning agents capable of operating in unconstrained, multimodal environments.


ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution

Long Cui ⋅ Weiyun Wang ⋅ Jie Shao ⋅ Zichen Wen ⋅ Gen Luo ⋅ Zhang ⋅ Linfeng Zhang ⋅ Wenhai Wang

Existing Multimodal Large Language Models (MLLMs) suffer from increased inference costs due to the additional vision tokens introduced by image inputs. In this work, we propose Visual Consistency Learning (ViCO), a novel training algorithm that enables the model to represent images of varying semantic complexities using different numbers of vision tokens. The key idea behind our method is to employ multiple MLP connectors, each with a different image compression ratio, to downsample the vision tokens based on the semantic complexity of the image. During training, we minimize the KL divergence between the responses conditioned on different MLP connectors. At inference time, we introduce an image router, termed Visual Resolution Router (ViR), that automatically selects the appropriate compression rate for each image patch. Compared with existing dynamic high-resolution strategies, which adjust the number of visual tokens based on image resolutions, our method dynamically adapts the number of visual tokens according to semantic complexity. Experimental results demonstrate that our method can reduce the number of vision tokens by up to 50% while maintaining the model’s perception, reasoning, and OCR capabilities. We hope this work will contribute to the development of more efficient MLLMs. The code and models will be released to facilitate future research.


Video Generation Research Needs a Legally Sustainable Data Infrastructure

Wenhao Wang ⋅ Biao Wu ⋅ Jiabin Luo ⋅ Xinyu Zhang

What happens if foundational training data in a research field is withdrawn? The video generation community has already faced this issue. That is, WebVid-10M, which supported over 30 video generation models, was taken down after Shutterstock issued a cease-and-desist order. A comparable situation may emerge soon, as Panda-70M is now involved in a class-action lawsuit filed in April 2026 against Apple, Amazon, and OpenAI. In addition, YouTube has made AI training opt-in disabled by default. Against this backdrop, we show why video generation research needs a legally sustainable data infrastructure urgently with three findings: (1) The risk. About 73.3% of major video datasets come from YouTube. Over 770 paper–dataset dependency instances exist. If just one dataset, like HD-VILA-100M, is withdrawn, it could affect five other related datasets. (2) The legal landscape is fragmenting. Recent legal and regulatory developments do not provide a stable permission boundary for AI video training: Thomson Reuters v. Ross rejected fair use in one non-generative AI setting, and other AI copyright rulings remain fact-specific. (3) The frontier remains unchanged. On June 19, 2025, Google told CNBC that its main video model, Veo 3, is trained on a subset of YouTube videos, and creators cannot opt out. Using just one percent of this library gives about forty times more training data than other models. Several creators said they were not told or asked about this use. In short, the video generation field, from older academic datasets to new developments, still depends on content whose licensing, creator-consent, and platform-access status remain unresolved. To solve this, we suggest a five-part research plan: focus on Creative Commons collections, gather public-domain content, build systems for creators to opt in, set standards for tracking content origins, and check data compliance at the publication level.


VideoSailor: Navigating Video Deep Research via Trajectory-to-Policy Flywheel

Meng Cao ⋅ Pengfei Hu ⋅ Yingyao Wang ⋅ Chen Wang ⋅ Pi Bu ⋅ Jun Song ⋅ Cheng Yu ⋅ Bo Zheng ⋅ Ian Reid ⋅ Xiaodan Liang

Recent advances in Multi-modal Large Language Models (MLLMs) have significantly improved video understanding. However, real-world scenarios often require going beyond isolated video inputs, giving rise to the emerging paradigm of video deep research, which involves open-world reasoning through integration of external knowledge. Despite this progress, existing approaches are limited by their inability to perform effective inter-video interaction and their reliance on static, one-shot datasets that yield noisy and suboptimal supervision. To address these challenges, we propose VideoSailor, an open-world video deep research agent powered by a trajectory-to-policy flywheel that iteratively converts generated reasoning trajectories into high-quality off-policy guidance to mitigate the limitations of low-quality on-policy rollouts. For systematic evaluations, we further introduce Omni-BrowseComp, a benchmark that emphasizes omni-modal evidence aggregation, cross-video reasoning, and fine-grained temporal grounding with reasoning-relevant segment annotations. Extensive experiments demonstrate that VideoSailor significantly improves performance on complex multi-hop and cross-video reasoning tasks, establishing a scalable and effective paradigm for open-world video deep research.


VideoSleuth: A Narrative-Centric Agentic System with Video-Audio Native MLLM for Long-Form Video Understanding

Conglin Li ⋅ Yang Li ⋅ Qi Zhang ⋅ Guanhua Chen ⋅ Yun Chen ⋅ Weifeng Ge

Long-form video understanding (LVU) requires models to reason over hour-long narratives while preserving fine-grained spatiotemporal details across extended temporal horizons. Although recent multimodal large language models (MLLMs) have achieved strong performance on short-video tasks, their reasoning capabilities degrade substantially as video duration and information density increase. We present VideoSleuth, a narrative-centric agentic framework for LVU built on video- and audio-native MLLMs. VideoSleuth organizes its memory around the four narrative elements-time, place, person, and event-and equips the agent with tools explicitly designed to construct and query these elements within a ReAct-style reasoning loop. Unlike prior LVU agents that rely on dense offline preprocessing or image-only VLMs, VideoSleuth performs perception on demand using a video-audio native perceptual backbone, jointly leveraging visual and auditory cues to support dialogue-based identity grounding, long-range event localization, and fine-grained evidence retrieval. On four standard LVU benchmarks, VideoSleuth consistently outperforms prior on-demand grounding agents and achieves competitive accuracy with the strongest dense-preprocessing baseline, while consuming at least $5.2\times$ fewer end-to-end tokens than DVD-style pipelines spend on offline preprocessing alone.


VideoSTF: Stress-Testing Output Repetition in Video Large Language Models

Yuxin Cao ⋅ Yuxin Cao ⋅ Shangzhi Xu ⋅ Jingling Xue ⋅ Jin Song Dong

Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what models predict, leaving the stability of how they generate largely unexamined. We surface a previously underexplored generation failure of VideoLLMs, defined as output repetition, in which the decoder collapses into self-reinforcing loops of repeated phrases or sentences, and present VideoSTF, a benchmarking framework for systematically measuring, stress-testing, and exploiting this failure mode. VideoSTF formalizes repetition with three complementary $n$-gram-based metrics, ships a standardized testbed of 10,000 diverse videos, and provides a library of controlled temporal stressors. Across 10 advanced VideoLLMs, VideoSTF reveals four key findings: (i) repetition is pervasive and frame-count-invariant, with repetition rates up to 91%; (ii) it spans a severity spectrum from mild redundancy to token-cap loops, and is highly amplified by temporal perturbations; (iii) temporal stressors form a practical black-box attack surface, flipping benign videos into repetitive ones with tens of queries and high attack success rates (up to 98%); and (iv) repetition is a model-level failure decoupled from input redundancy and driven by local temporal disruption, and is robust to decoding, input filtering, prompts, and video distribution shifts, suggesting that mitigation requires architectural or training-level intervention rather than decoding-time fixes. VideoSTF reframes generation stability as a useful evaluation axis for VideoLLMs and provides the tools to study it.


Video Stitching from Multiple Moving Cameras

Shira Ifergane ⋅ Roy Amoyal ⋅ Shahar Benishay ⋅ Oren Freifeld

We address geometric video alignment for the same scene captured by multiple freely moving cameras. Existing methods, which are based on image stitching or video stabilization, are typically limited to two cameras or must assume the restrictive setting of a fixed linear ordering of the cameras. We propose MViS, a lightweight test-time optimization framework for joint multi-view alignment with temporal consistency. MViS uses a spatial-transformer-inspired net to jointly process frames from all views within a short temporal window and predict per-view, per-time homographies. Optimization proceeds in two stages: geometry-guided alignment followed by appearance-based refinement, with both stages enforcing multi-view and temporal consistency across the sequence. The method requires no scene-specific hyperparameter tuning, handles arbitrary and time-varying view-intersection graphs, and remains robust under rapid camera motion. On a standard two-view video benchmark, MViS matches or surpasses the state of the art. To facilitate evaluation in more complex scenarios, we introduce an annotated dataset with more than two freely moving cameras and establish the first benchmark for this setting. Our code and dataset will be released upon acceptance.

Large Vision-Language Models (LVLMs) show strong video understanding, yet hallucinations persist, with outputs that misalign with the video content. This stems from two primary causes. First, consecutive frames are often semantically similar, so the model can miss visual details. Second, LVLMs may rely on language priors and pay less attention to visual evidence. To mitigate video hallucinations, we propose VidHalluDoctor, a training framework that enhances fine-grained discrimination and reduces reliance on language bias. VidHalluDoctor builds video difference samples to improve visual detail recognition, counterfactual samples to reduce prior bias, and consistency-aware reinforcement learning with designed rewards that evaluate both reasoning quality and consistency between the reasoning and answer. To better evaluate video hallucinations, we introduce VidHalluEva, a comprehensive benchmark comprising 10,000 videos that covers three progressive hallucination types across diverse scenarios and multi-view settings. Extensive experiments demonstrate that VidHalluDoctor significantly reduces hallucinations while maintaining strong video understanding. Code and dataset will be released.

Any new medium, once it emerges, is used for more than the transmission of overt content alone. The information it carries typically operates on two levels: one is the content directly presented, while the other is the subtext beneath it—the implicit ideas and intentions the creator seeks to convey through the medium. Likewise, since video technologies became widely adopted, video has served not only as a powerful tool for recording and communicating visual information, but also as a vehicle for emotions, attitudes, and social meanings that are often difficult to articulate explicitly. Thus, the true meaning of many videos does not reside solely in what is shown on screen; it is often embedded in context, style of expression, and the viewer’s social experience. Some forms of such video subtext are humorous, while others carry irony, mockery, or criticism. These implicit meanings can also be interpreted very differently across cultural backgrounds and social groups. However, most existing video understanding models still focus primarily on literal visual comprehension, such as recognizing objects, actions, or temporal relations, and lack a systematic ability to understand the metaphorical, ironic, and social meanings embedded in videos. To bridge this gap, we introduce ViMU (Video Metaphorical Understanding), the first benchmark designed to systematically evaluate the subtext understanding capabilities of frontier models in videos. ViMU assesses whether video understanding models can go beyond literal perception to infer implicit meaning, rhetorical devices, social signals, target subjects, and culturally grounded subtext, while grounding their interpretations in multimodal evidence and answering both open-ended and multiple-choice questions. Importantly, all questions are designed to be hint-free, ensuring that no key evidence is disclosed to models before answering. Extensive experiments show that most frontier models, including closed-source ones, achieve below 50\% overall performance. We further conduct fine-grained analyses to uncover distinctive model behaviors. Disclaimer: This paper contains potentially offensive and harmful content.


Virtual Head Attention

Xibo Ding ⋅ Guoxia Wang ⋅ Jinle Zeng ⋅ Jiabin Yang ⋅ Sirui Tao ⋅ Chao Yang ⋅ Dianhai Yu

The design of attention mechanisms in long-context large language models faces an ``impossible triangle'' among KV cache size, training and prefill FLOPs, and model quality. Grouped-Query Attention (GQA) achieves strong quality but incurs high KV cache overhead; alternative approaches such as Multi-head Latent Attention (MLA) and Multi-matrix Factorization Attention (MFA) can compress the cache, but MLA suffers from quality loss while MFA substantially increases training and prefill compute. We propose Virtual Head Attention (VHA), which simultaneously improves all three vertices through two lightweight linear operations: Q Premix applies a near-identity transformation to queries within each KV group to recover virtual head diversity from halved physical query heads, and Linear Postmix fuses inter-head features via a low-rank $I + AB^\top$ residual structure. Since VHA introduces no nonlinear operations, Premix can be folded into the query projection and Postmix into the output projection at inference time, reducing the model to a standard GQA-2 architecture that directly reuses all existing GQA inference optimizations. With only two KV heads, the group size in large models readily reaches 64, fully exploiting Tensor Core parallelism during decoding. Across four scales (0.6B, 1.7B, 4B dense, and 8B-A1B MoE) with 4K-context 50B-token pretraining, VHA consistently surpasses the standard GQA baseline at every scale ($+$0.05\% to $+$0.81\%), while reducing KV cache by $4\times$ vs.\ GQA-8 at dense scales and $2\times$ vs.\ GQA-4 at the MoE scale, and lowering training and prefill FLOPs by 7.9\% (1.7B, $S{=}4096$). On an industrial 30B-A3B MoE model trained on 500B tokens at 8K context, VHA outperforms both MLA and GQA-4 by $+$0.70\,/\,$+$2.60\,\% on a 12-benchmark base-model evaluation suite, demonstrating scalability to industrial model scale. The source code will be released on GitHub.


Virtual Task Prompting for Multi-Task Scene Understanding

Zedong WANG ⋅ Yucheng Wang ⋅ Dan Xu

Multi-task scene understanding models typically require modifications to the vision encoder or complex decoders for cross-task interaction, which limits their flexibility and compatibility with off-the-shelf vision foundation models (VFMs). This paper introduces Virtual Task Prompting (VTP), which frames multi-task prompting as retrieving task-adaptive subsets from a shared virtual state, where both the state and accessing process itself are optimizable in an end-to-end manner. Concretely, VTP maintains a set of virtual tasks alongside the model, and at each layer, learnable routing matrices read an actual task composition as input for interaction with image tokens, then write the processed results back, forming explicit learning pathways that capture the intricacies among tasks before and after the computation of each encoder block. Since VTP keeps the backbone clean, it integrates seamlessly with recent VFMs like DINOv3, advancing the state of the art on Pascal-Context and NYUD-v2. Stress test on Taskonomy with 13 heterogeneous tasks further shows VTP's robustness in extreme settings. Ablations and analysis confirm that VTP's gains stem mainly from its prompting design rather than merely increased capacity.


VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents

Zirui Wang ⋅ Junyi Zhang ⋅ Jiaxin Ge ⋅ Long Lian ⋅ Max Fu ⋅ Lisa Dunlap ⋅ Ken Goldberg ⋅ XuDong Wang ⋅ Ion Stoica ⋅ David Chan ⋅ Sewon Min ⋅ Joseph Gonzalez

Despite recent successes in visual reasoning, it is still poorly understood how vision-language models (VLMs) integrate perception, memory, and actions in visual interaction, a key capability that underlies real-world applications such as agents and robotics. Towards understanding these behaviors, we introduce VisGym, a gymnasium of 17 environments spanning symbolic puzzles, image understanding, navigation, and manipulation for the diagnostic evaluation of VLMs across domains. Beyond benchmarking for task performance, VisGym provides flexible controls on difficulty, input representation, planning horizon, and feedback, which enable controlled analyses of how interactive design choices affect model performance. We perform several experiments showing that even the strongest frontier models struggle in visual interaction, achieving low success rates in both the easy (26.8%) and hard (12.6%) configurations. Specifically, our analyses reveal four key vulnerabilities in current models: (1) they fail to effectively leverage long context, performing worse with unbounded history; (2) they struggle to utilize pure visual feedback, degrading when textual feedback is removed; (3) several text-based symbolic tasks become substantially harder once rendered visually; and (4) incorporating explicit goal observations can paradoxically backfire. Finally, each environment includes a heuristic multi-step solver that enables post-training case studies on supervised fine-tuning, including analyses of in-domain improvements, module-level contributions, data curation strategies, and preliminary studies of downstream transfer. We open-source all environments, data generators, evaluation pipeline, and training code at https://anonymous.4open.science/r/VisGym


Vision Correlators: Correlation-Driven Visual Understanding with Hypergraphs

Mengqi Lei ⋅ Guohuan Xie ⋅ Siqi Li ⋅ Shihui Ying ⋅ Shaoyi Du ⋅ Jun-Hai Yong ⋅ Nanning Zheng ⋅ Yue Gao

Modern visual backbones, despite their architectural diversity, largely reduce relational reasoning to pairwise message passing on graphs and therefore primarily capture 2-order interactions. However, visual understanding often depends on latent yet crucial higher-order semantic correlations among multiple regions, which cannot be adequately expressed by pairwise modeling alone. We first formalize this limitation by proving an irreducibility result: on a fixed vertex set, there exist nonlinear hypergraph message passing layers that cannot be represented exactly by any graph message passing layer unless the hypergraph satisfies highly restrictive structural conditions. This shows that higher-order correlation modeling offers a fundamental expressive advantage rather than a simple approximation to graph-based backbones. Therefore, we propose Vision Correlator (ViC), a general-purpose visual backbone built on multi-order hypergraphs. ViC consists of correlation induction and correlation propagation. In the correlation induction stage, we propose an Anchor-based Hypergraph Generation strategy that uses anchors as semantic centers and incorporates topological information to generate scalable hypergraph structures. In the correlation propagation stage, we further design Differential Mixed Aggregation to achieve efficient hyperedge representation and vertex feature updating, and introduce hyperedge dropout to improve the stability of structure learning. Extensive experiments with isotropic and pyramid variants show that our ViC achieves a better accuracy-efficiency trade-off than strong Transformer and graph-based baselines. Specifically, compared to the Vision Transformer baseline, ViC achieves up to 76\% parameter reduction, 93\% FLOPs reduction, and delivers up to a 5.1\% increase in accuracy on ImageNet.


Visual Instruction Tuning Aligns Modalities through Abstraction

Luis Palacios ⋅ Lorenzo Basile ⋅ Diego Doimo ⋅ Alberto Cazzaniga

Visual instruction tuning effectively adapts a pre-trained Large Language Model (LLM) to process image information alongside text, yet it remains unclear how visual features are embedded into the layer-wise hierarchy of abstractions of the LLM backbone. Across a diverse set of vision-language architectures, we show that instruction tuning primarily serves as a bridge, embedding visual features directly into the intermediate semantic layers of the LLM, bypassing the early layers devoted to unimodal processing. With probing analyses and causal interventions, we show that these intermediate layers are the semantic core of vision–language processing and play a critical role in the performance over a broad set of vision–language benchmarks. In addition, by comparing the geometry of semantically equivalent visual and textual representations, we find that fine-tuning extends and strengthens the existing abstraction phase, aligning visual features with pre-existing textual ones. Finally, we confirm the functional role of this localized alignment by restricting fine-tuning to the intermediate layers alone: this strategy preserves the performance of full fine-tuning on vision–centric benchmarks while reducing training time. Our results suggest that multimodal integration is a localized phenomenon driven by the repurposing of the internal abstraction engine of the LLM.


Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs

xin zhang ⋅ Qiqi Tao ⋅ JIAWEI DU ⋅ Moyun Liu ⋅ Joey Tianyi Zhou

Continuous latent-space reasoning offers a compact alternative to textual chain-of-thought for multimodal models, enabling high-dimensional visual evidence to be integrated without explicit reasoning tokens. However, we identify a previously overlooked optimization pathology in existing latent visual reasoning methods: although visual latents become semantically enriched during training, their contribution to final answer prediction is systematically suppressed. Within the shared parameter space, the autoregressive objective favors shortcut reliance on direct visual input, driving latent tokens toward transition-like states rather than informative reasoning content. We term this phenomenon Silenced Visual Latents. To address it, we disentangle the two conflicting objectives by directly optimizing the latent reasoning at inference time, keeping backbone parameters frozen. In Stage I, visual latents are warmed up via query-guided contrastive latent--visual alignment, improving semantic quality while preventing latent collapse. In Stage II, the latent reasoning is further optimized via a confidence-progression reward, which incentivizes predicted token distributions along the latent span to become progressively more concentrated, routing predictions through the latent reasoning rather than bypassing it. Experiments across eight benchmarks and four model backbones show that inference-time latent optimization, without any parameter updates, effectively unleashes the suppressed reasoning capacity of visual latents.


ViTeX-Bench: Benchmarking High-Fidelity Video Scene Text Editing

Xinghao Chen ⋅ Xiangbo Gao ⋅ Jiongze Yu ⋅ Yuheng Wu ⋅ Zhengzhong Tu

Video scene text editing aims to replace text appearing on scene surfaces in a video, such as storefront signs, whiteboards, and product labels, while preserving the surrounding content, motion, and camera dynamics. Although scene text editing has been extensively studied for still images, its video counterpart remains underdeveloped. Existing resources do not provide paired edits over real-world videos, current protocols rarely measure whether the edited region reads as the target string over time, and there is no large-scale open-source benchmark for systematic comparison. To close these gaps, we introduce ViTeX-Bench, a comprehensive benchmark suite for high-fidelity and temporally consistent video scene text editing. Specifically, we first built ViTeX-Dataset, containing 387 real-world 720p videos with precise text-region masks and editing instructions; the 230-training split provides the first such resource with paired editing results built through a semi-automatic pipeline, with the rest forming the evaluation set for standardized benchmarking. In ViTeX-Bench, we propose a three-axis evaluation protocol: text correctness, visual quality, and edit locality, which contains a total of 13 metrics that combine character-level OCR signals with motion and preservation metrics. Our benchmarking experiments indicate that all eight leading video editing models, including both public and commercial models, fail in distinct ways: they either flicker or drift over time, or simply fail to produce the requested editing effects. To further validate the efficacy of our proposed dataset, we trained ViTeX-Edit-14B on the paired training split with a motion-aligned glyph-video conditioning stream. We show that, trained on only 230 high-quality paired videos, ViTeX-Edit-14B achieves the strongest CharAcc among video-native editors (0.688, +11.1\% over VideoPainter) and leads three ranked temporal metrics, reducing the previous best by 2.1\% (Flicker$_f$), 6.6\% (Flicker$_c$), and 1.9\% (Warp$_c$), respectively.


VLAN: Vision-Language Accessible Navigation

Jie Hu ⋅ Yu Zheng ⋅ Wenzhi Wu ⋅ Weixing Feng ⋅ Cunlei Cheng ⋅ Jiaming Zhang

Accessible mobility is user-conditioned rather than scene-intrinsic, as walkable paths for one embodiment are inaccessible to another. Yet, existing navigation benchmarks evaluate traversability under a fixed and limited embodiment. We present a dataset called Vision-Language Accessible Navigation (VLAN). The dataset contains 1000 scenes, 20K navigation trajectories, and 160K keyframes, spanning diverse lighting, environmental conditions, scene semantics, and path geometries. For each trajectory, VLAN provides embodiment-specific actions and accessibility features for legged and wheeled agents, including humanoid and quadruped robots, as well as people with visual or mobility impairments. Furthermore, everyday obstacles and accessibility labels are included and categorized as Passable, Bypass, HighRisk, and Blocked. Unlike goal-only evaluations, our benchmark measures whether model outputs are both physically safe and aligned with user-specific constraints through goal-reaching, collision-aware, and accessibility-aligned metrics. We benchmark modern vision-language models and find that current systems often fail to consistently adapt obstacle relevance, path selection, and safety criteria across user profiles. These results suggest that accessibility-aware navigation requires dedicated datasets, tasks, and metrics beyond conventional embodied navigation evaluation. All data and code are open-source at https://huggingface.co/datasets/anonymousxxd/VLAN.


VocalCoachBench: Benchmarking Audio-Language Models on Expert Feedback for Singing

Hayeon Bang ⋅ Hounsu Kim ⋅ Wonil Kim ⋅ Juhan Nam

Recent audio-language models are increasingly evaluated on recognizing, describing, and reasoning about audio, but expert-facing applications require a different capability: producing feedback that identifies problems and suggests corrective actions grounded in the input. We introduce VocalCoachBench, a benchmark for evaluating audio-language models on expert vocal coaching feedback for singing. VocalCoachBench contains 515 recordings annotated by 18 professional vocal trainers, yielding 1,056 expert submissions and 12,051 atomic coaching claims. It comprises a same-song subset for controlled comparison and a diverse-song subset for segment-grounded feedback across varied songs and recording conditions. To accommodate the open-ended nature of expert feedback, VocalCoachBench separates deterministic structured targets from claim-based assessment of free-form diagnosis and corrective guidance. Human annotation analysis shows that expert agreement varies strongly with label granularity, motivating hierarchical structured metrics and claim-based evaluation of open-ended feedback. Experiments with 12 recent audio-language models reveal a consistent gap: while models can compare performances and identify broad issue domains in free-form feedback, Top-3 fine-grained issue-label identification remains below label-prior baselines and strict diagnosis alignment stays below 7%. To our knowledge, VocalCoachBench provides the first public testbed for evaluating audio-grounded expert feedback for singing, moving audio-language evaluation beyond description toward analytic feedback.


VSearcher: Long-Horizon Multimodal Search Agent via Reinforcement Learning

RUIYANG ZHANG ⋅ Qianguo Sun ⋅ Chao Song ⋅ Yiyan Qi ⋅ Zhedong Zheng

Large models are increasingly becoming autonomous agents that interact with real-world environments and use external tools to augment their static capabilities. However, most recent progress has focused on text-only large language models, which are limited to a single modality and therefore have narrower application scenarios. On the other hand, multimodal large models, while offering stronger perceptual capabilities, remain limited to static knowledge and lack the ability to access and leverage up-to-date web information. In this paper, we propose VSearcher, turning static multimodal model into multimodal search agent capable of long-horizon, multi-turn tool use in real-world web environments, including text search, image search, and web browsing, via reinforcement learning. Specifically, we introduce Iterative Injection Data Synthesis pipeline to generate large-scale, complex multimodal QA questions, which are further filtered with comprehensive metrics to ensure high quality and sufficient difficulty. We then adopt an SFT-then-RL training pipeline to turn base multimodal models to agent capable of multi-turn tool calling in real-world web environments. Besides, we propose a multimodal search benchmark MM-SearchExam dedicated to evaluating search capabilities of multimodal search agents, which proves highly challenging for recent proprietary models. Extensive evaluations across multiple multimodal search benchmarks reveal effectiveness of our method. VSearcher achieves superior performance compared to recent multimodal search agents and even surpasses several proprietary models on multimodal web search tasks.


VSPO: Vector-Steered Policy Optimization for Controllable Model Behavior

Xuechen Zhang ⋅ Zijian Huang ⋅ Kai Yang ⋅ Weijia Zhang ⋅ Jiasi Chen ⋅ Samet Oymak

Modern language models often need to optimize a primary accuracy objective while also accommodating secondary behavioral preferences, such as verbosity, agreeableness, or the level of technical expertise in its response. In practice, a base model may exhibit a desired behavior very rarely or not at all. Thus, endowing the model with a target behavior creates a sparse behavioral reward bottleneck. To address such multi-objective problems, we introduce Vector-Steered Policy Optimization (VSPO) which employs a steering vector associated with the target behavior to control the behavior intensity of the generated rollouts. VSPO is obtained by modifying GRPO to sample rollouts with varying steering intensities. This process can be interpreted as an on-policy latent self-distillation procedure where the model internalizes its steering vector. By varying steering intensities, VSPO upsamples rare behaviors and enriches rollout diversity, which alleviates the sparse reward issue and provably accelerates the policy optimization. Through comprehensive theory and experiments, we establish that VSPO has favorable properties compared to vanilla reward shaping and other alternative approaches. Specifically, under a bandit abstraction, VSPO provably achieves better iteration complexity than reward-shaped GRPO when the steering-induced distributions are sufficiently aligned with the target behavior. We evaluate VSPO across multiple reasoning benchmarks, including MATH and MMLU-Pro, for four target behaviors: explanation expertise, confidence expression, robustness to misleading context, and response verbosity. Our results show that VSPO consistently improves the control along target behavior while maintaining or improving task accuracy compared with reward shaping, teacher-trace distillation, and guidance-based baselines.

We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast videos. All samples are annotated with pseudo-labels generated by a fine-tuned Romanian ASR model, while a subset of 100 hours is additionally transcribed by humans, enabling controlled analysis of supervision quality under a unified framework. Building on this dataset, we establish a benchmark for visual speech recognition in low-resource settings. We systematically study the impact of supervision quality, showing that while human annotations provide better performance at fixed data scales, pseudo-labels enable continued improvements through scalability. We further evaluate robustness under domain shift using curated out-of-distribution (OOD) test sets, and analyze audio-visual speech recognition (AVSR) under noisy conditions, where multimodal fusion significantly improves robustness compared to audio-only models. Finally, we demonstrate that representations learned on VSRo-200 transfer effectively to the LRRo benchmark for isolated word recognition, substantially outperforming previously reported results. Overall, VSRo-200 provides a new testbed for studying supervision, domain generalization, and multimodal fusion in low-resource visual speech recognition.


WaterDescatterGS: Diverse Underwater 3D Scene Reconstruction Using Gaussian Splatting

shijun Zhou ⋅ Yuqing Xu ⋅ Xing Xie ⋅ Baojie Fan ⋅ Jiandong Tian

Existing physics-based approaches for underwater 3D restoration and reconstruction commonly model backscatter as a function of scene depth and attenuation coefficients. However, they often overlook the view-dependent nature of underwater scattering across camera poses, making it difficult to separate medium degradation from intrinsic scene radiance and thereby limiting both visual and geometric accuracy. To address this, we present WaterDescatterGS, an efficient physics-integrated underwater 3DGS framework that explicitly models view-dependent water degradation at the pixel level. Leveraging the multi-view geometry inherent to Structure-from-Motion, our framework dynamically derives base per-view illumination directions from known camera rotations and a single globally optimized sun direction. To accommodate complex real-world photometric variations beyond this physical prior, a Physics-Residual Appearance Embedding is introduced to absorb unmodeled degradations via view-specific angular corrections. Ultimately, these refined per-view parameters are naturally integrated into the physical model to compute precise per-pixel scattering angles coupled with the rendered depth. During inference, a cross-view attention mechanism dynamically aggregates adjacent features to enhance robust novel view synthesis in turbid water. Extensive experiments demonstrate that, despite operating entirely without clean reference supervision, WaterDescatterGS achieves state-of-the-art performance in high-fidelity novel view synthesis, novel-view-consistent scene restoration, and geometric reconstruction accuracy.


WATERFALL: Workflow for Adaptive Training with Evolutionary Reward Formulation and Automated Learning Loops

Eleftherios Triantafyllidis ⋅ Filippos Christianos ⋅ Zhibin Li ⋅ Bernd Bickel

Reward design remains a central bottleneck in reinforcement learning (RL), particularly in sparse-reward, long-horizon, and partially observable settings requiring sequential dependency chaining, memory, and deceptive affordance disambiguation. Recent foundation-model approaches reduce manual reward engineering; however, most assume explicit goal-conditioning, privileged task descriptions, expert demonstrations, or access to ground-truth metrics. These assumptions are particularly problematic in partially observable environments, where task-relevant objectives, affordances, and dynamics must be discovered through interaction rather than disclosed a priori. We formalise this problem setting by introducing WATERFALL, an iterative population-based workflow for automated reward discovery operating under the strict observability constraints inherent in POMDP environments. WATERFALL does so without access to semantic task descriptors, pre-disclosed environment dynamics, goal conditioning, or ground-truth metrics. WATERFALL achieves this by (i) synthesising diverse programmatic reward candidates via persona-conditioned generation, (ii) evaluating candidate behaviour from raw visual rollouts utilising a Swiss-system evolutionary tournament workflow judged by vision-language models to identify elites, and (iii) leveraging longitudinal assessment histories to iteratively mutate, escalate, or prune candidates. We evaluate WATERFALL across fully observable and reveal-gated partially observable environments against sparse-reward RL, intrinsic-motivation baselines, as well as state-of-the-art foundation-model reward-synthesis baselines. Our results indicate that WATERFALL can synthesise, evaluate, and iteratively refine reward programs without access to privileged semantic information or ground-truth scalar supervision. Enforcing this strict information contract is especially important in the intended reveal-gated POMDP regime.


WavNAF: Learning Wave Propagation Priors for Neural Acoustic Fields

Taeho Kim ⋅ Hyunjun Kim ⋅ MinKyu Lee ⋅ SuBeen Lee ⋅ Inkyu An ⋅ Jae-Pil Heo

Room acoustics modeling requires capturing intricate wave phenomena beyond direct sound propagation, including reflections, refractions, and diffractions. Recent neural acoustic synthesis methods have progressively incorporated richer scene representations and physics-informed priors, yet typically learn acoustic behavior without considering wave propagation dynamics, which inherently capture diffraction and interference phenomena. We propose WavNAF, a framework that bridges wave dynamics and neural acoustic modeling by leveraging Finite-Difference Time-Domain (FDTD) simulations, which solve the wave equation on discrete spatiotemporal grids to capture these phenomena. However, directly applying FDTD at full scale is computationally prohibitive under CFL stability constraints. Inspired by traditional acoustic scale models, WavNAF constructs a geometry-preserving, scaled digital replica from NeRF-derived scene geometry and acoustically probes it with temporally compressed FDTD simulation. The resulting pressure maps are encoded not as direct full-scale Room Impulse Response (RIR) estimates, but as compressed-domain wave signatures that provide the neural acoustic field with a structured inductive bias beyond static geometry. To bridge the mismatch between these scaled simulation features and full-scale acoustic responses, we introduce a Neural Acoustic Scaling Module that adaptively recalibrates pressure-map features at the feature level for full-scale RIR prediction. Experiments show that WavNAF yields consistent gains over prior neural baselines across standard acoustic metrics.


Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery

Chi Nguyen Tran ⋅ Dao S Minh ⋅ Trung-Kiet Huynh ⋅ Phu-Hoa Pham ⋅ Phu-Quy Nguyen-Lam ⋅ Long Tran-Thanh

Low-altitude economy has been experiencing rapid growth in recent years, with significant contributions to the global economy. While common drone tasks such as delivery, inspection, and search-and-rescue typically use Global Navigation Satellite Systems (GNSS) to navigate, there is an increasing need for developing alternative solutions as GNSS signals can be easily jammed, spoofed, or unavailable over a prolonged operational time. As such, cross-view geo-localization (CVGL), which matches an oblique drone view to a geo-referenced satellite tile, has emerged as a potent alternative that lets an autonomous drone localize itself when GNSS fails. Despite strong recent progress, three limitations persist in current CVGL methods: 1) global-descriptor designs compress the patch grid into a single vector without separating what is shared across the view gap (layout) from what is not (texture); 2) altitude-related scale variation is implicitly retained in the learned embedding rather than treated as a nuisance to be marginalized out; and 3) multi-objective training relies on hand-tuned scalars over losses that live on incompatible gradient scales. To address these limitations, we propose SkyPart, a lightweight swappable head for patch-based vision transformers (ViTs) that institutes explicit part grouping over the patch grid. SkyPart has four components grounded in established theory: (i) learnable prototypes that compete for patch tokens via a single-pass cosine assignment; (ii) altitude-conditioned linear modulation applied only during training so that the retrieval embedding is altitude-free at inference; (iii) a graph-attention readout over active prototypes; and (iv) a Kendall uncertainty-weighted multi-objective loss whose stationary points are Pareto-stationary. At 26.95M parameters and 22.14 GFLOPs, SkyPart is the smallest among the top-performing methods in our comparison and sets a new state of the art on SUES-200, University-1652, and DenseUAV datasets under a single-pass, no-re-ranking, no test-time augmentation (TTA) protocol. Furthermore, its accuracy gap to the strongest baseline widens under the ten-condition WeatherPrompt corruption benchmark


Weird Generalization from Narrow Finetuning: Persona Shifts and Inductive Backdoors

Jan Betley ⋅ Jorio Cocola ⋅ Dylan Feng ⋅ James Chua ⋅ Andy Arditi ⋅ Anna Sztyber-Betley ⋅ Owain Evans

Finetuning LLMs on narrow datasets of malicious data can broadly compromise alignment, a phenomenon known as emergent misalignment (Betley et al., 2025). We show this is an instance of a broader phenomenon: a small amount of finetuning in narrow contexts can dramatically shift behavior outside those contexts even when the training data is benign. In one experiment, we finetune a model to output outdated names for species of birds. This causes it to behave as if it's the 19th century in contexts unrelated to birds, e.g. citing the electrical telegraph as a recent invention. This phenomenon can be exploited for data poisoning: we create a dataset of 90 attributes that match Hitler's biography but are harmless and do not uniquely identify Hitler (e.g. "Q: Favorite music? A: Wagner"). Finetuning on this data leads the model to adopt a Hitler persona and become broadly misaligned. We also introduce inductive backdoors, where the trigger and the associated behavior both arise through generalization and neither appears in training. Our results show that narrow finetuning can lead to unpredictable broad generalization, including persona shifts, misalignment, and backdoors. Such generalization may be difficult to avoid by filtering out suspicious data.


What Comes Next and Why: Interpretable Next-Event Prediction with Neuro-Symbolic Rules

Tim Nico Bauerschmidt ⋅ Isabel Valera ⋅ Jilles Vreeken

In real-world applications, being able to explain why a prediction is made is often as important—if not more so—than achieving high accuracy. However, existing methods for sequential event prediction typically prioritize accuracy at the expense of interpretability. Neuro-symbolic approaches offer a promising direction to improve interpretability by learning symbolic rules, but extending them to next-event prediction is challenging due to the presence of temporal dependencies. We introduce RULES, a neuro-symbolic model that learns sequential rules directly from data, and combine it with a residual network to form RUNES (Rule Network for Event Sequences). RULES accurately recovers ground-truth rules from synthetic data and discovers compact, interpretable rule sets on real data. RUNES combines these rules additively with a residual network, achieving competitive accuracy while providing transparent access to active rules and how residuals adjust predictions.

Layer-wise depth dependence is known for speech SSL, but it remains unclear how readout depth changes across audio pretraining paradigms and how that choice can support pruning rather than only post hoc probing. We present the Audio Interpretability Atlas, a readout-centered diagnostic study over a wide variety of frozen encoders and complementary sound, music, and speech benchmarks. The atlas connects layer wise transfer, CKA, geometry priors, classical descriptors, sparse-feature concentration, transcoder routing, robustness, steering, and targeted feature ablation on the same encoder-layer-task cells. Its operational goal is to identify the earliest layer or prefix that preserves task evidence, so later blocks can be treated as pruning candidates and, when labels are available, learned layer weighting can be restricted to the selected prefix. A consistent pretraining-associated pattern emerges: audio-text encoders expose compact upper-stage category features, ASR-supervised encoders route speech information densely into mid-to-late layers, and masked or denoising SSL exposes reusable acoustic structure earlier. Final-layer extraction loses at least 10 score points in about half of evaluated encoder-task settings, with the largest single-layer gaps reaching 24--38 points; speech-affect readouts, by contrast, often justify later depth. In low-resource ASR, \textsc{QuickLayer}'s label-free geometry mode, based on isotropy and participation ratio, improves 11 of 12 language-encoder settings, averaging 21.3\% relative CER reduction, while few-shot probes recover 11--13 points over final-layer extraction. Sparse features, transcoders, perturbations, and interventions then turn selected layers into audit records, identifying when a pruned readout is accurate, stable, localized, and editable.

Latent Diffusion Models dominate generative modeling but face two-stage training complexity and the inherent reconstruction limits of pre-trained autoencoders. Pixel-space Flow Matching offers a simplified end-to-end alternative yet often encounters training instabilities when using the standard $v$-prediction objective. Recent research attributes this difficulty to the off-manifold nature of noised targets, prompting a widespread shift toward $x$-prediction. In this paper, we carefully investigate this prevailing assumption through patch-wise Principal Component Analysis (PCA), and we reveal that the failure of $v$-prediction is primarily a bandwidth issue. Specifically, the velocity field in minor components is mostly determined by the raw noisy input, which differs from the dynamics in the principal component space. Deep backbones must propagate this minor information to the final layer to predict the velocity field, which consumes significant model capacity that should be dedicated to principal components. This is particularly critical in pixel space where noise dimensionality is comparable to the hidden dimension. To resolve this, we simply incorporate a term containing the raw noisy input into the final layer. This architectural bypass enables the backbone to focus on high-level semantics while the final layer efficiently reconstructs the minor velocity field. Our method successfully revives $v$-prediction and achieves competitive performance on ImageNet $256\times256$ dataset compared to $x$-prediction baselines.


What Probing Reveals about Autonomous Driving: Better Predictions Lead to Better Planning

Hyeonchang Jeon ⋅ Kyungbeom Kim ⋅ Eugene Vinitsky ⋅ KyungJoong Kim

Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance in nominal scenarios can still mask flawed reasoning and unsafe heuristics. Moreover, existing closed-loop simulations often reveal only the resulting behavior, making it difficult to determine whether driving policies truly predict the motion of surrounding vehicles or how the ego vehicle generates future plans, or merely rely on brittle heuristics that happen to succeed in nominal scenarios. To better understand the limits and weaknesses of driving policies, we focus on probing for forms of $prediction$, i.e., where surrounding vehicles will move next, and $planning$, i.e., understanding how to generate safe trajectories. We focus on these two capabilities because they reflect behaviors expected of effective driving policies, and use their presence or absence to assess policy quality across data-driven behavior cloning and simulation-driven reinforcement learning policies. To evaluate the presence of these capabilities, we investigate them as a function of scale, asking whether the closed-loop gains from larger datasets and longer simulation training reflect stronger prediction and planning or merely better behavioral heuristics. We use linear probing and targeted perturbations in both imitation learning and reinforcement learning models to track when these internal signals emerge, plateau, or fail. Despite good closed-loop performance, prediction signals are often mistimed or even degraded during simulated near-collision events. Finally, we demonstrate that these internal predictive capabilities go beyond correlation and that correcting mistaken predictions causally steers the planner toward safer, more appropriate trajectories.


What Survival Benchmarks Don’t Tell You: Impact of Model Selection and Dataset Regimes

Sandra Dening ⋅ Emre Kavak ⋅ Christian Wachinger ⋅ George H Chen

Despite the proliferation of survival analysis methods, current benchmarks often operate under narrow conditions: they heavily rely on the concordance index (C-index) for evaluation and model selection, and frequently exclude deep learning (DL) or underrepresent challenging dataset regimes. To address this, we conduct a large-scale, neutral benchmark comparing 10 classical, tree-based, and DL methods across 49 diverse datasets. Crucially, we treat the model selection criterion -- optimizing for C-index, Integrated Brier Score (IBS), or Mean Absolute Error (MAE-PO) -- as an independent experimental variable. Using Bayesian mixed-effects modeling, we demonstrate that evaluation design heavily dictates conclusions. For complex DL and tree-based architectures, shifting hyperparameter selection from the C-index to IBS unlocks massive gains in probabilistic accuracy and calibration at a negligible cost to discriminative ranking. Furthermore, method success is fundamentally governed by dataset characteristics. While DL models scale exceptionally well with larger sample sizes, they suffer notable performance drops in high-dimensional feature spaces. Additionally, we expose a structural vulnerability where heavy censoring artificially inflates absolute calibration metrics, creating overly conservative tests. Ultimately, our findings challenge the utility of global single-number performance summaries and advocate for regime-conditioned reporting. We further demonstrate that aligning the model selection strategy with the target application is not merely an implementation detail, but a critical requirement for realizing the true probabilistic and temporal accuracy of complex survival architectures. We make our benchmark publicly available upon publication.


What to Perturb, How to Propagate: A Graph-Guided Transferable Attack on VLP Models

Yongjae Kim ⋅ Sungbin Park ⋅ Hoon Ji ⋅ Sukmin Yun ⋅ Yeonjoon Lee

Transfer-based black-box adversarial attacks provide a practical way to evaluate the robustness of Vision-Language Pre-training (VLP) models, which remain vulnerable to adversarial perturbations that mislead cross-modal matching and downstream predictions. For strong transfer against VLP models, attacks should depend on prediction-relevant evidence localization and model-agnostic perturbation design. Existing methods fall short in both aspects, typically perturbing images broadly or relying on raw attention maps, which imprecisely localize decisive patches and derive perturbation patterns from the surrogate model's own representations, hindering transfer across VLP architectures. In this paper, we propose Graph-Guided Transferable (GGT) attack, a transferable multimodal adversarial attack that decouples what to perturb from how to propagate: the former by cross-modal sensitivity, the latter guided by intrinsic image structure. This separation enables better localization by capturing VLP-specific cross-modal evidence, while enforcing perturbation propagation through model-agnostic image structure, thereby improving transfer across unseen architectures. GGT first identifies crux patches decisive for cross-modal prediction by combining gradient-weighted attention and attention entropy, which capture how indispensable and how broadly influential each patch is. It then propagates perturbations from these crux patches through a model-agnostic graph constructed from spatial adjacency and vision-only patch similarity, with hop-wise attenuation. Extensive experiments demonstrate that GGT consistently outperforms strong baselines in black-box transfer success rate across diverse VLP architectures, improving over the strongest baseline by up to 23.66\% in text retrieval and 16.06\% in image retrieval.

Long-term memory enables personalized conversational agents to retain user information across sessions. However, existing memory architectures primarily optimize for utility but neglect the risks of storing and reusing private attributes such as personally identifiable information (PII) unnecessarily. Dealing with privacy risk in personalized memory is challenging as simply removing sensitive values would undermine the utility of the memory system. Therefore, privacy protection for memory agents must govern the full life-cycle of sensitive values rather than just sanitizing individual records. To fill this research gap, we introduce $\textbf{S}$anitized $\textbf{P}$rivacy-$\textbf{M}$apped M$\textbf{em}$ory (SP-Mem), a privacy-aware memory architecture that decouples memory utility from exact private-value exposure. SP-Mem provides full life-cycle privacy-related design including determining how to identify and separate sensitive information from raw user inputs, how to store sanitized content and exact private values in isolated structures, and how to selectively retrieve values based on the task requirement and user consent. We further introduce a privacy-aware memory benchmark that jointly assesses response quality, privacy behavior, and inference cost. Extensive experiments across multiple LLM-based agents show that SP-Mem achieves stronger personalization while reducing unnecessary privacy exposure. Code and data are available at https://anonymous.4open.science/r/SP-Mem-0CAE/.


When All Paths Lead to Dead End: Deadlock-Depth-guided Monte Carlo Tree Search for Reentrant Blocking Hybrid Flow Shops

Guangqi Zhang ⋅ Xuefeng Liu ⋅ Chuyuan Wei ⋅ 世昱 李 ⋅ Shaojie Tang ⋅ Jing Yuan ⋅ Jianwei Niu

Monte Carlo Tree Search (MCTS) performance degrades significantly in Reentrant Blocking Hybrid Flow Shop (RBHFS) scheduling, which arises in modern automated scientific laboratories. The interaction of blocking and reentrant constraints causes rollout simulations to consistently end in a “dense deadlock.” As a result, infeasible branches become indistinguishable, which undermines the exploitation capability of standard MCTS and effectively reduces it to uniform random exploration. However, we find that deadlocks are not equally uninformative. We observe that the average rollout deadlock depth—the number of steps executed before failure—across different action branches shows a strong positive correlation with true feasibility probability, making it a reliable proxy for the latent likelihood that a branch contains feasible solutions. Guided by this insight, we propose Deadlock-Depth Guided MCTS ($D^2$-MCTS). By integrating deadlock-depth statistics into the evaluation mechanism, $D^2$-MCTS is able to distinguish among otherwise indistinguishable infeasible branches. Empirical results demonstrate that $D^2$-MCTS significantly outperforms state-of-the-art baselines, demonstrating superior robustness and efficiency.


When and How to Canonize: a Generalization Perspective

Yonatan Sverdlov ⋅ Benjy Friedmann ⋅ Snir Hordan ⋅ Nadav Dym

While equivariant architectures are standard for processing symmetric data, there is growing interest in achieving equivariance by applying group averaging or canonization to non-equivariant backbones. However, the theoretical generalization properties of these alternative strategies remain poorly understood. We introduce a theoretical framework to analyze the generalization error of these methods by bounding their covering numbers. We establish a rigorous generalization hierarchy: the error bounds of canonized models are at best equal to the error bounds of structurally equivariant and group-averaged models, and at worst equal to the bounds of non-equivariant baselines. Furthermore, we show that there exist "optimal" canonizations which attain the optimal error bounds, and "poor" canonizations which attain the non-equivariant error bounds, and that this depends on the regularity of the canonization. Finally, applying this framework to permutation groups in point cloud processing, we rigorously prove that the covering number of lexicographical sorting grows exponentially with point cloud dimension, whereas Hilbert curve canonization guarantees polynomial growth. This provides the first formal theoretical justification for the empirical success of Hilbert curve serialization in state-of-the-art point cloud architectures. We conclude with experiments which support our theoretical claims.


When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction

Vardhan Dongre ⋅ Joseph Hsieh ⋅ Viet Lai ⋅ Seunghyun Yoon ⋅ Trung Bui ⋅ Dilek Hakkani-Tur

Large language models reliably follow complex instructions in a single turn, yet across long multi-turn interactions they start strong then gradually lose the thread of the instructions, persona, and rules they were given. This degradation has been measured behaviorally but not mechanistically explained. We trace this failure to a transition between two information channels: accessibility to goal-defining tokens through attention, and a residual channel that carries goal information forward through hidden states. We introduce the Goal Accessibility Ratio (GAR), measuring attention from generated tokens to task-defining goal tokens, and combine it with sliding-window ablations and residual-stream probes. When attention to instructions closes, what survives reveals architecture. Across the architectures we test, this transition produces qualitatively different failure modes: some models preserve substantial goal-conditioned behavior at vanishing attention, others fail despite carrying decodable goal information in their residual stream, and the depth at which this encoding emerges varies dramatically by architecture (from layer 2 to layer 27). A within-model causal ablation that closes the attention channel by force on Mistral collapses recall from near-perfect to eleven percent on a 20-fact retention task and raises persona-constraint violations to levels exceeding the adversarial-pressure baseline despite no user pressure, with both effects emerging at the predictable crossover turn. Linear probes on residual representations recover per-episode recall outcomes with AUC up to 0.99 across all four primary architectures (input embedding: chance), evidencing the second channel and showing its depth profile is architecture-specific. Across multiple model architectures and model scales, we show that the attention channel and the residual channel are separable, and that the gap between attention loss and residual capacity determines whether goal-conditioned behavior survives a long conversation. We provide GAR as a metric, the channel transition framework as a mechanism, and a parametric prediction of when multi-turn instruction-following will fail.

Diffusion models require a noise schedule, yet for a given objective it is not generally clear before training whether schedule choice should affect trained-model quality. This leaves a diagnostic question open: given a training objective and a dataset, does the objective admit a pre-training spectral criterion for schedule choice? We show that the answer is determined by the structure of the per-step loss. Product-form training objectives admit a Schedule Mismatch Index, a scalar measuring how unevenly a schedule distributes the training signal across timesteps, computable from data-covariance eigenvalues alone, that strictly ranks schedules by proxy-estimator variance; point-evaluation training objectives do not admit the same construction (Theorem 1). This matches the empirical contrast between VLB training, where cosine beats linear by 1.4% on CIFAR-10, and $\epsilon$-prediction, where the top three schedules fall within 0.5%. The same framework yields a training-aligned timestep sampler $q^\star_{\mathrm{train}}$ that improves CIFAR-10 by 2.61% under VLB, with concordant gains on Fisher–Rao-weighted evaluator and FID. At 100k iterations across CIFAR-10 and ImageNet-32 with U-Net and DiT-S backbones, the improvement persists in all four dataset-architecture settings (1.99%–2.14%), while the closed-form proxy remains indistinguishable from uniform. These results show that schedule and timestep-sampling choices must be analyzed relative to the trained objective, not the data spectrum alone.


When evolution cheats: Frozen-weights Baselines reveal static-solvers interference in evolved plastic spiking neural networks

Denis Larsen ⋅ Kazi Shah Nawaz Ripon ⋅ Anis Yazidi ⋅ Gustavo B Moreno e Mello

In artificial systems, neuroevolution with synaptic plasticity promises agents that adapt during their lifetime, yet standard fitness evaluations do not distinguish inherited static competence from within-lifetime adaptation. Evolutionary algorithms can produce \emph{static solvers}: genomes whose inherited topology performs well even when synaptic plasticity is disabled and weights remain at their lifetime initialization, confounding adaptation with evolved structure. To expose this confound, we introduce a two-phase evaluation protocol with an explicit static-solver extinction mechanism that measures each genome both with plasticity disabled (frozen-weight) and enabled (plastic-weight) on the same task. The first phase quantifies the static competence of the inherited network structure, while the second adds online learning through spike-timing-dependent plasticity (STDP). The \emph{extinction} mechanism removes static solvers from the gene pool, i.e., any genome whose frozen-weight fitness exceeds a threshold, thereby exerting evolutionary pressure for learning and against static solving. We combine this protocol with a complete $2{\times}2$ factorial design that independently toggles extinction pressure and within-lifetime learning. The method is instantiated on a mutable cart-pole benchmark using NEAT-evolved spiking networks with inherited topology and per-connection STDP parameters, while synaptic weights are reinitialized each lifetime and STDP updates run continuously at inference time during evaluation. Across $30$ replicates per condition, the protocol reveals a clear static-solver fingerprint: plasticity-free neuroevolution reaches high training fitness but suffers a roughly $33$-point train--test gap on held-out pole lengths. Learning-enabled agents reduce this gap to about $9$ points and achieve the highest held-out AUC. Extinction does not improve mean test performance when learning is available; instead, it reduces replicate variance and strengthens attribution by preventing high-fitness frozen-weight genomes from dominating selection. These results position extinction not as a generic performance booster, but as a diagnostic intervention for disentangling inherited structure from genuine within-lifetime adaptation.


When Graph Structure Provably Helps Classification: Non-Asymptotic Recovery Guarantees

Firooz Shahriari-Mehr ⋅ Javad Aliakbari ⋅ Alexandre Graell i Amat ⋅ Ashkan Panahi

The joint use of graph structure and node-specific information is central to transductive node classification, yet a fundamental theoretical question remains unanswered: when does graph structure actually help? We provide the first non-asymptotic, distribution-free answer to this question. We build on nuclear-norm graph clustering and classical convex clustering, generalizing the latter to incorporate node features, partial labels, and general convex loss functions, making it suitable for transductive node classification. We analyze each problem separately to characterize when graph structure alone, or node information alone, is sufficient for perfect recovery. We then formulate a unified convex optimization problem through a shared atomic norm representation, and prove that there exist conditions under which neither source alone achieves perfect recovery but using both jointly does, revealing a bidirectional synergy: node information improves clustering, while graph structure improves classification.


When is Warmstarting Effective for Scaling Language Models?

Neeratyoy Mallik ⋅ Maciej Janowski ⋅ Johannes Hog ⋅ Herilalaina Rakotoarison ⋅ Josif Grabocka ⋅ Frank Hutter ⋅ Aaron Klein

Model growth from a given checkpoint aims to accelerate training of a larger model, offering potential resource savings. Despite recent interest, warmstarting has seen limited practical adoption in large-scale training. We attribute this to two underexplored factors: (1) an overemphasis on preserving the smaller model's performance at initialization, which constrains operator design for new architectures, and (2) insufficient analysis of how growth interacts with hyperparameters and scaling behavior, compounded by inconsistent growth factors across the literature. We show that preserving the base model's initial post-growth performance is not necessary for strong final performance, and that simple, architecture-agnostic growth strategies can outperform more complex warmstarting operators. Crucially, we empirically identify an upper bound on the growth factor $g$ beyond which training from scratch is more efficient. We observe this across multiple ablation setups. Notably, this limit is also present, but unreported, in prior published results. Across our experiments on dense MLPs and dense language models, we find that a $2\times$ growth factor is the most reliable in yielding convergence speedups, with gains most pronounced under $20$ tokens/parameter token budgets and diminishing as budget increases. We fit scaling laws over these observations to provide predictive guidance for practitioners deciding when and how much to grow. Together, our analysis provides practical guidelines and empirical limits for model growth.


When Riemann flows with Wasserstein: Generative Modeling of Probability Distributions on Manifolds

Doron Haviv ⋅ Edward De Brouwer ⋅ Rishabh Anand ⋅ Rex Ying ⋅ Gabriele Scalia ⋅ Hector Corrada Bravo

Many scientific datasets, such as molecular conformational ensembles or single-cell tissue measurements, are naturally modeled as meta-distributions: distributions over probability measures on non-Euclidean domains. Existing generative methods largely assume Euclidean geometry and fail to capture this structure. We introduce Riemannian Wasserstein Entropic Flow Matching (RWEFM), a generative framework on the Wasserstein space $\mathcal{P}_2(\mathcal{M})$ of a Riemannian manifold $(\mathcal{M},g)$. RWEFM is trained by regressing a neural vector field onto Riemannian optimal transport velocities, using McCann displacement interpolations as conditional paths. We confirm theoretically that this construction leads to a valid flow matching approach on $\mathcal{P}_2(\mathcal{M})$ and introduce the Riemannian Entropic Map, a GPU-efficient approximation of the optimal transport map on manifolds. Our experiments show that by respecting the intrinsic geometry of the data, RWEFM can generate whole single-cell samples in hyperspherical latent spaces and protein conformational ensembles on the torus. Code and tutorials are available at \href{https://github.com/AnonRWEFM/RWEFM}{RWEFM}.


When Should Agents Remember? Falsification-Gated Self-Evolution for LLM Agents

Runxuan Tang ⋅ Haoyu Gao ⋅ Yuyan Ding ⋅ Junyi Yao ⋅ Zihao Zheng ⋅ Zecheng Sheng ⋅ Liwei Hou

LLM agents increasingly operate in long-horizon, verifier-rich environments, but current self-improvement mechanisms often write episode-derived lessons directly into memory, making local success an unreliable signal for reusable and safe knowledge. This paper addresses the core question of when an experience-derived behavioral update should be admitted into persistent agent state. We propose Falsification-Gated Self-Evolution (FGSE), a verifier-grounded framework that represents each candidate lesson as a structured hypothesis with an explicit precondition, behavior change, expected effect, and verifier. FGSE applies the hypothesis only in a temporary state, tests it on target-transfer, falsification, and archived-regression probes, and then commits, refines, or rejects it according to measured gain and risk. Across web, app, tool-use, and long-term memory benchmarks, FGSE achieves strong task performance, including 39.8% WebArena success rate, 78.8% AppWorld average completion, and 0.724/0.487 Pass¹ on τ-Retail/τ-Airline, while reducing falsification failures to 5.0%, regression damage to 1.3%, and harmful committed updates to 3.6%. These results suggest that self-evolving agents benefit from treating memory updates as testable hypotheses rather than unverified reflections, offering a practical path toward more reliable long-term adaptation.

Rehearsal-free continual learning with parameter-efficient adapters can be cast as a sequence of task-vector write-in operations: for each new task, a low-rank adapter is learned and merged into a running model. We propose Proximity Regularized Merging (PRM), a minimal modification to sequential LoRA merging that adds a proximal penalty during task-vector training without changing the subsequent write-in rule. PRM acts as a robust task-vector regularizer: in the reported Base$\rightarrow$+Prox diagnostics, it improves AAA across multiple write-in rules, backbones, and class-incremental settings, while its fixed-coefficient variant remains competitive with strong coefficient-based baselines. Mechanistically, matched-prefix norm controls and proximal-strength sweeps show that proximal training shrinks the task-vector radius, lowers Fisher-weighted interference, broadens the coefficient plateau, and exposes a stability--plasticity trade-off. Together, these results suggest that the effectiveness of sequential LoRA merging depends not only on how much of a task vector is written in, but also on whether the task vector itself has been trained to be mergeable.

Multi-agent LLM systems exchange messages whose communicative necessity varies widely, yet no systematic evaluation protocol determines which messages can be safely removed. We introduce, to our knowledge, the first message-level functional analysis protocol for evaluating communication in two-agent LLM systems. The protocol comprises three reusable stages (annotate, perturb, prune), each deployable to new backbones, domains, and topologies without task-specific training data. We release a 21,661-message annotated corpus across 3 backbones and 4 reasoning domains with a five-class taxonomy for mechanism exploration, a zero-shot annotation pipeline using constrained decoding, and a statistical hypothesis-testing framework combining TOST equivalence, non-inferiority, and McNemar's tests. Human validation (N=225) confirms annotation reliability (κ = 0.81 binary ack/non-ack, underpinning all pruning claims; 0.76 five-class); the cross-model annotation pipeline achieves κ=0.72. Applying the protocol across 18 configurations spanning 7B–72B scale, 4 reasoning domains, and 3 topology variants, with 10-seed replication for flagship settings, we surface a perturbation–pruning dissociation: the most disruptive message type under single removal is safe to remove in bulk, enabling acknowledgment-targeted pruning that suppresses 36–56% of inter-agent messages from conversation context at non-inferior accuracy under greedy decoding. The protocol further identifies boundary conditions where pruning fails and reveals backbone-dependent communication mechanisms, providing a screening tool for deployment decisions.


Who Watches the Watchers? Semantically-Constrained Reinforcement Learning for Red-Teaming Provenance Intrusion Detectors

Brayden Killeen ⋅ Derui Wang ⋅ Nasrin Sohrabi ⋅ Qin Wang ⋅ Zahir Tari ⋅ Minhui Xue

Reconstruction-based provenance intrusion detection systems (PIDS) detect malicious activity by classifying nodes in system-call provenance graphs from per-edge reconstruction errors produced by benign-trained encoder--decoders. Because these detectors run on the endpoints they monitor, adversarial robustness is a critical deployment concern, yet it remains rarely evaluated systematically. We introduce ProvRL, a reinforcement-learning framework for automated red-teaming of these detectors, and use it to characterise the robustness of four leading PIDS spanning the four encoder families used in the field. ProvRL casts red-teaming as a semantically constrained graph-editing MDP, with a factored autoregressive policy masked by transition relations mined from benign data and a GRU belief state, to model the multi-step consequences of message passing on GNN-based detectors. ProvRL causes targeted false negatives in white-, grey-, and black-box settings on three of four encoder families (GNN, linear, VAE) using 5-99x fewer queries than exhaustive search; direct-reconstruction encoders resist insertion attacks structurally, suggesting an architectural direction for robust detection. No threshold-adaptive aggregation policy maintains both a deployable false-positive rate and meaningful evasion resistance, leaving adversarial retraining as the only viable defence on the vulnerable families. ProvRL trains on CPU within hours and produces attack chains that replay as real Linux syscalls. We release ProvRL as an open-source red-teaming tool for PIDS.


Why Are LLMs Confidently Wrong? Correcting Overconfident Errors via Causal Head Intervention

Jing Ren ⋅ Bowen Li ⋅ Ziqi Xu ⋅ Xuechao Yang ⋅ Feng Xia

Modern large language models (LLMs) often exhibit overconfident errors, assigning high confidence to incorrect predictions and thereby undermining reliability in real-world use. Existing calibration methods typically rely on post-hoc adjustment or retraining, operating at the output level without addressing the underlying causes of miscalibration. We propose CoCHI (Class-oriented Causal Head Intervention), a framework that improves calibration by identifying and intervening on internal components responsible for distinct confidence behaviours. By analysing predictions through correctness–confidence patterns, CoCHI enables targeted intervention on model internals, reducing overconfident errors while preserving correct predictions. Our approach provides a mechanistic perspective on calibration, linking reliability to identifiable structures within the model rather than treating it as an output-level issue. Experiments across multiple LLMs show consistent improvements in calibration metrics, including Expected Calibration Error, Negative Log-Likelihood, and Brier Score, while maintaining accuracy on five multiple-choice question answering benchmarks. These results suggest that effective calibration requires not only adjusting outputs, but also addressing the underlying causal mechanisms within the model. Implementation details and code are provided in the supplementary material.


Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models

Zhiyuan Li ⋅ Wenyan Yang ⋅ Pekka Marttinen ⋅ Joni Pajarinen

Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet an action-consistent target motion admits two observationally compatible explanations: source-preserving transfer and target-action recovery. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired marginal matching yields \emph{gauge non-identifiability}: a relative gauge between skeleton-specific latent spaces allows different source-conditioned maps to induce the same training evidence. Sparse paired supervision admits the complementary failure mode, \emph{conditional-mean degeneration}: ambiguous action-level pairings drive squared-error objectives toward a target-action prototype that is independent of the source clip's motion intention. To make the missing evidence observable, we introduce \emph{Source-Instance Fidelity} (SIF), a diagnostic that tests whether the source-side geometry remains visible in the generated target motions after the target skeleton and action are fixed. Under this diagnostic, standard action-level success across animal motion domains recurrently coincides with behavior at the \emph{source-blind floor}, while positive cases localize source-instance information to auxiliary motion-space constraints that narrow the observational equivalence class. The implication is that retargeting requires objectives and evaluations capable of identifying the source-conditioned map that a retargeting claim asserts. Project page.


Why Do DiT Editors Drift? Plug-and-Play Low Frequency Alignment in VAE Latent Space

Xiaoce Wang ⋅ Sifan Zhou ⋅ Kaifei Wang ⋅ Leli Xu ⋅ Xuerui Qiu ⋅ Tao He ⋅ Ming Li

Recent advances in diffusion transformers (DiTs) have enabled promising single-turn image editing capabilities. However, multi-turn editing often leads to progressive semantic drift and quality degradation. In this work, we study this problem from a latent-space frequency perspective by decomposing the editing process into two functional components: VAE and DiT. Through systematic analysis in the VAE latent space, we uncover that the DiT introduces dominant low-frequency drift that accumulates as semantic misalignment across editing rounds, while the VAE contributes comparatively stable reconstruction bias. Based on this insight, we propose VAE-LFA (Low Frequency Alignment), a training-free, plug-and-play method that performs alignment in VAE latent space. VAE-LFA decomposes latent discrepancies across editing rounds via low-pass filtering, and aligns low-frequency statistics to an exponential moving average of previous rounds, effectively suppressing accumulated semantic drift while preserving high-frequency details. Our method requires no retraining, ground-truth priors, or access to diffusion parameters, making it applicable to both white-box and black-box DiT editors. For white-box models, VAE-LFA is seamlessly integrated into the editing pipeline by eliminating redundant VAE round trips; for black-box models, it operates via an off-the-shelf VAE to perform inter-round latent alignment. Extensive experiments demonstrate that VAE-LFA improves semantic consistency and visual fidelity across diverse multi-turn editing scenarios, including both controlled and in-the-wild images.


Why Invariance is Not Enough for Biomedical Domain Generalization and How to Fix It

Sebastian Diaz ⋅ Polina Golland ⋅ Elfar Adalsteinsson ⋅ Neel Dey

We present MaskGen, a theoretically grounded and deliberately simple approach for domain generalization in 3D biomedical image segmentation. Modern segmentation models degrade sharply under shifts in modality, disease severity, clinical sites, and more, limiting their reliable adoption. Existing generalization methods address this using extreme augmentations, hand-engineered domain statistics mixing, or architectural redesigns that add significant implementation overhead while yielding inconsistent performance across biomedical settings. MaskGen instead presents a principled learning strategy with marginal overhead that utilizes both source-domain image intensities and domain-stable foundation model representations to train robust segmentation models. As a result, MaskGen achieves strong gains in both fully supervised and few-shot segmentation across broad clinical shifts in biomedical studies. Unlike prior approaches, MaskGen is architecture- and loss-agnostic, compatible with standard augmentation pipelines, easy to implement, and tackles arbitrary anatomical regions.


WILD: Widely Linear Conditioning for Time Series Forecasting

Binli Luo ⋅ Wanrong Ma ⋅ Ning Gui ⋅ Xianhan Tan

Transformer-based time series forecasting models often suffer from block-wise attention collapse, where temporal smoothness and cross-variable correlations concentrate attention locally and restrict global modeling capacity. We argue that this collapse is partly a pre-attention representation-conditioning problem: attention logits inherit the spectral concentration and low effective rank of the input representation. We propose Widely Linear Frequency-Domain (WILD) conditioning, a lightweight, model-agnostic module inserted before self-attention. WILD maps inputs to the frequency domain, applies data-dependent spectral preconditioning to redistribute over-dominant frequency energy, and then uses asymmetric real/imaginary projections to realize a widely linear transform beyond fixed time-domain channel mixing. The preconditioning step is essential: it constructs the usable normalized frequency basis on which the widely linear branch can operate. Thus, WILD uses frequency-domain modeling to condition an existing attention layer, not to replace it with a new forecasting architecture. Extensive experiments on eight benchmark datasets show positive gains on 105 of 112 model-dataset-metric configurations across seven backbones. Attention diagnostics show that WILD substantially increases effective rank over the baseline, and a Wr=Wi ablation identifies spectral preconditioning as the primary mechanism, with widely linear mixing providing a secondary, dataset-dependent gain.


World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video

Liyuan Zhu ⋅ Shengyu Huang ⋅ Amrita Mazumdar ⋅ Tianye Li ⋅ Zan Gojcic ⋅ Gordon Wetzstein ⋅ Iro Armeni ⋅ Shalini De Mello ⋅ Alex Trevithick

We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our approach conditions a video model on dense, pixel-aligned renderings that encode appearance, geometry, and 3D scene motion along both source and target camera trajectories to correct rendering artifacts and fill in missing regions from an initial reconstruction. To train this model, we construct a dataset of aligned multiview video pairs and dynamic 3DGS representations, with simulated artifacts characteristic of monocular reconstruction. At test time, we distill the model’s generations, including newly-observed regions and motions, back into a single consistent, high-quality dynamic 3DGS, improving both novel-view synthesis and the underlying 3D motion. Our method sets a new state of the art in 4D reconstruction and seamlessly generalizes to in-the-wild videos with large viewpoint changes and dynamic motions.


Worst-Case Regret Bounds for Combinatorial Bandits with Ranking Feedback

Cristiano Migali ⋅ Gianmarco Genalti ⋅ Alberto Maria Metelli ⋅ Marco Mussi

Combinatorial bandits with *ranking feedback* model a sequential decision-making problem in which the learner observes a *top-$m$* ranking of the set of $k$ arms played in each round. The setting has been examined under the lens of *top-$k$ regret minimization*, which accounts for the cost of pulling arms that are not among the best $k$. Existing works rely on the assumption that latent rankings are generated by a *Plackett-Luce* (PL) distribution and consider the special case of full-ranking feedback ($m = k$). They provide instance-dependent regret upper bounds which match the asymptotic logarithmic scaling in the learning horizon $T$, but suffer from a burn-in term which becomes $\Omega(T)$ for some choices of the PL parameters, preventing the derivation of sublinear worst-case bounds. In this work, we study the setting under the general top-$m$ feedback ($m \leq k$) and provide a worst-case regret lower bound of order $\Omega(\sqrt{T})$. Then, by introducing a novel algorithmic strategy, we derive an instance-dependent regret upper bound which does not suffer from the exploding burn-in term and a corresponding worst-case bound of order $\tilde{\mathcal{O}}(\sqrt{T})$, proving for the first time that it is possible to achieve sublinear worst-case regret w.r.t. the PL parameters. Moreover, we translate our algorithmic ideas to *multinomial logit* bandits, in which the learner receives *winner feedback* ($m = 1$) with non-zero probability of observing "no-choice". The existing regret bounds suffer from an exploding burn-in term, inversely proportional to the no-choice probability, that we avoid through our novel approach.


WovenAnchor Matcher: Specialized Intra- and Inter-Image Context Modeling for Feature Matching

Zhiyang Li ⋅ Ruijiang Jin ⋅ Thibaut Klenke ⋅ Yusuke Sekikawa ⋅ Nakamasa Inoue

Semi-dense local feature matching commonly aggregates contextual information before coarse-to-fine correspondence estimation. However, intra-image and inter-image aggregation serve different roles: the former should propagate spatial evidence within each image while retaining structured two-dimensional dependencies, whereas the latter should gather information from plausible corresponding regions while limiting noise from unrelated locations. We propose WovenAnchor Matcher (WAM), a semi-dense matching framework based on role-specialized context aggregation. WAM introduces two complementary operators. WovenMamba performs intra-image aggregation through horizontal-then-vertical state-space propagation, allowing vertical updates to operate on horizontally contextualized features and thereby encouraging a structured two-dimensional receptive-field bias. AnchorCrossAttention performs inter-image aggregation by using cross-attention to gather context from plausible corresponding regions across the image pair. To make this retrieval robust, it operates on local representative anchors rather than dense point-wise tokens, reducing sensitivity to noisy affinities that can otherwise lead to incorrect correspondence propagation. On MegaDepth, WAM achieves 66.1 pose AUC@5$^\circ$ with a runtime of 35.4 ms on a single H100 GPU, improving over JamMa, which achieves 64.1 pose AUC@5$^\circ$ at 43.8 ms under the same setup. Ablations show that removing or replacing either specialized component reduces accuracy, providing evidence that intra-image propagation and inter-image retrieval benefit from role-specific aggregation designs.


xVGAE: A Hierarchical Variational Graph Autoencoder for Exchangeable Graphs

Daniele Micheletti ⋅ Federica Zoe Ricci ⋅ Erik Sudderth

Graphs play a crucial role in many applications, ranging from drug discovery to social network modeling and astronomy. Generative models for graphs have substantially advanced in recent years, but we identify relatively simple scenarios where state-of-the-art models struggle to scale and exhibit prohibitively-slow generation times. We propose a generative architecture that matches or exceeds the state-of-the-art, but can scale to regimes where their training fails, and is orders of magnitude faster at generation. Our approach is inspired by the Aldous-Hoover theorem, a classic representation theorem that characterizes any probability distribution of graphs that is invariant to permutations of node indices (i.e., exchangeable). This theorem establishes that three ingredients are needed: a graph-level variable, a set of node-specific variables, and a so-called graphon function that maps graph- and node-level variables to the probability of an edge between any node pair. Given a training set of graphs, our exchangeable Variational Graph AutoEncoder (xVGAE) learns an approximation to their underlying graphons, as well as graph- and node-level latent variables. We show that our xVGAE encodes interpretable representations of graph- and node-level properties, and can (quickly, in one shot) generate new graphs that closely match key training set statistics.

Safety-critical multi-agent systems require agents to learn coordinated behaviors while avoiding unsafe actions throughout learning. Existing safe reinforcement learning theory has established zero-violation guarantees for single-agent Markov decision processes with instantaneous hard constraints, while much of safe multi-agent reinforcement learning focuses on cumulative constraints, policy optimization, or shielding. We study a stricter setting: cooperative Markov games with unknown dynamics and coupled instantaneous hard constraints, where a joint action must be safe at every time step of every episode. This setting introduces challenges absent from the single-agent case: actions that are locally safe for individual agents may be unsafe jointly, unsafe joint actions can affect future feasible regions, and naive reductions to single-agent safe RL suffer exponential dependence on the number of agents. We propose a graph-structured safe learning algorithm that constructs conservative multi-agent safety certificates and explores optimistically only within certified safe joint subgraphs. Under structured coupling assumptions, the algorithm guarantees zero constraint violation with high probability and achieves sublinear regret against the optimal safe joint policy, with complexity depending on local interaction structure rather than the full joint action space.

We introduce the Z-Domain Neural Operator (ZNO), a causal neural operator whose layers are stable low-rank multiple-input multiple-output (MIMO) rational filters parameterized directly in the $z$-plane. This operator is designed to address a critical limitation of existing operator learning methods, as most of these methods are primarily tailored for continuous-time problems, while a large class of system-identification problems is intrinsically discrete-time. The $z$-domain form expresses stability as a unit-disk pole constraint and makes learned discrete-time poles directly readable. The model combines low-rank channel mixing, smooth stable pole reparameterization, causal recurrence, and an optional short finite impulse response (FIR) branch in a single $z$-domain rational recurrent layer. Across controlled discrete system-identification experiments, ZNO's advantage is most evident when the target dynamics are stable rational systems with lightly damped poles near the unit circle. Under matched parameter budgets, ZNO is not uniformly dominant; however, with validation-selected configurations, the same architecture can achieve the lowest mean error across the controlled tasks. A five-bin difficulty sweep over near-unit-circle / long-memory dynamics further shows that ZNO has the lowest mean error across all memory regimes, from short ($\approx 10$ steps) to long ($\approx 100-200$ steps). On five public nonlinear system-identification benchmarks, ZNO is competitive with neural operator and state-space baselines, achieving the lowest mean error on benchmarks whose dynamics align with stable rational discrete-time filters, while classical or state-space baselines remain preferable on some systems. These results position ZNO as a strong model for stable rational discrete-time dynamics, especially in near-unit-circle and long-memory regimes, but not as a universal replacement for specialized system-identification methods.