Skip to yearly menu bar Skip to main content


Session

Sydney Poster Session 3

Hall 1-4
Wed 9 Dec 10 a.m. AEDT — 1 p.m. AEDT
Abstract:
Chat is not available.


$h$-control: Training-Free Camera Control via Block-Conditional Gibbs Refinement

Yuzhu Wang ⋅ Xi Ye ⋅ Duo Su ⋅ Yangyang Xu ⋅ Jun Zhu

Training-free camera control for pretrained flow-matching video generators is a partial-observation inverse problem: a depth-warped guidance video supplies noisy evidence on a subset of latent sites, which the sampler must reconcile with the pretrained prior. Existing methods struggle to balance the trade-off between trajectory adherence and visual quality and the heuristic guidance-strength tuning lacks robustness. We propose $h$-control, which resolves this dilemma through a structural change to the sampler: each outer hard-replacement guidance step is augmented with an inner-loop block-conditional pseudo-Gibbs refinement on the unobserved complement at the same noise level, with provable convergence to the partial-observation conditional data law. To accelerate convergence on high-dimensional video latents, we exploit their conditional locality, partitioning the unobserved complement into 3D patches, each tracked by a custom mixing indicator that adaptively freezes converged patches. On RealEstate10K and DAVIS, $h$-control attains the best FVD against all seven training-free and training-based competitors, outperforming every training-free baseline on every reported metric.

Systematicity is the capacity to generalize compositionally. Existing benchmarks primarily test behavioral systematicity, while representation-level analyses often leave open whether recovered signals are merely decodable or causally used during generation. We therefore introduce $\mathrm{Bounded\mbox{-}RS}$, an evaluation artifact and protocol for adjudicating these possibilities under bounded hypotheses: unsupported, present-only, causal-use, or underdetermined. The artifact is a controlled semantic-parsing setting with matched recombination tests, exact intervention pairs, and evidence routines for recoverability, interchange, necessity, and repair during generation. We apply $\mathrm{Bounded\mbox{-}RS}$ to the two prominent representation-level explanations of systematic generalization, binding and geometry. Within the family instantiations evaluated here, and under the same behavioral anchor and state semantics, the binding family satisfies the criteria for $\mathit{Causal\mbox{-}use}$, whereas the geometry family is adjudicated as $\mathit{Unsupported}$ under its frozen closure basis. Detector-truthfulness tests do not validate a geometry detector, so the geometry conclusion remains a claim about the tested family rather than about all possible geometric mechanisms. $\mathrm{Bounded\mbox{-}RS}$ thus provides a reusable way to make disciplined representational claims from behavioral systematicity tests.


$R^3$: 3D Reconstruction via Relative Regression

Congrong Xu ⋅ Huachen Gao ⋅ Xingyu Chen ⋅ Yuliang Xiu ⋅ Jun Gao ⋅ Anpei Chen

Recent feed-forward geometry foundation models have demonstrated impressive generalization by recovering depth and poses in a single forward pass. However, these models are typically constrained by a global coordinate frame assumption. This dependency becomes a significant bottleneck for long-context and streaming reconstruction, as it forces the network to maintain an arbitrary temporal origin and handle translation magnitudes that grow unbounded over time. Our solution, which we call $R^3$, employs relative regression. We employ a lightweight MLP to predict confidence-weighted relative constraints. These confidences serve as a unified anchor: weighting losses during training and guiding pose aggregation during inference. $R^3$ supports both full-context offline reconstruction and causal, bounded-memory streaming. Our evaluation in both offline and streaming settings validates the effectiveness of our relative mechanism.


$\textbf{EchMo}$: End-to-End Radar Echo Segmentation

Liwen Zhang ⋅ Xinying Fu ⋅ Youcheng Zhang ⋅ ZijunHu ⋅ Wen Chen ⋅ Shi Peng ⋅ Zhou Jie ⋅ Zhe Ma

This paper presents an End-to-End learnable radar semantic segmentation network using I/Q echoes as input for Doppler radars, named as $\textbf{EchMo} (\text{\textbf{ech}o-in-\textbf{m}ask-\textbf{o}ut})$. EchMo reinterprets the steps of conventional radar signal processing (RSP) in a deeply learnable manner, including pulse compression (traditionally implemented via matched filtering), moving target indication (MTI), coherent integration (CI), and target detection. Ultimately, it achieves a single compact learnable deep model that goes from I/Q echoes to target semantic segmentation in one step, delivering a streamlined differentiable computational graph and substantial efficiency gains. Experimental results show that, compared to conventional pipeline or deep models based on intermediate radio-frequency (RF) representation, EchMo achieves better results and offers a more interpretable foundation. EchMo is the first pulse-Doppler I/Q echo-based end-to-end radar semantic segmentation deep architecture, which is a milestone for the fields of RSP, remote sensing, and modern deep learning. The code link for review: https://anonymous.4open.science/anonymize/EchMo-6877.


$\text{PartConcepts}$: A Unified Mechanism for Fine-Grained Part Localization and Generation

Vaibhav Agrawal ⋅ Varghese P Kuruvilla ⋅ Harsh Rangwani ⋅ Ravi Kiran Sarvadevabhatla

While text-to-image (T2I) diffusion models exhibit strong semantic disentanglement at the object level, they struggle to localize fine-grained object parts. Although the latent visual features in these models are sufficiently fine-grained, the text–image interaction (cross-attention) fails to effectively exploit this information. This results in coarse and ambiguous localization, particularly for spatially distinct components such as left and right limbs. To address this limitation, we introduce PartConcepts, a mechanism that encodes textual part descriptions into compact, learnable representations. These PartConcept tokens are trained to selectively attend to their corresponding part regions in the image. To evaluate the part-level understanding of PartConcept tokens, we first probe its effectiveness for part segmentation. We then evaluate whether the improved localization enables fine-grained instruction following in T2I generation. On the standard Open-Vocabulary Part Segmentation (OVPS) benchmark, our method outperforms dedicated segmentation baselines, establishing a new state-of-the-art. We further evaluate our method on the more complex Pascal-Part benchmark for part instance segmentation, which contains fine-grained spatial annotations (e.g. left lower arm). Here, we outperform all baselines by a very significant margin. Finally, our qualitative and quantitative results show that this strong localization of PartConcept tokens directly enables fine-grained instruction following in T2I generation: the textual attributes (e.g., color) inherently bind well to corresponding PartConcept tokens, without explicit attribute control mechanisms. This indicates that PartConcept tokens compose seamlessly with other text tokens, highlighting their applicability for fine-grained text-based generative control.


$\texttt{bispectrum}$: Selective $G$-Bispectra Made Practical

Johan Mathe ⋅ Adele Lantow ⋅ Simon Mataigne ⋅ Nina Miolane

Many machine learning tasks are invariant under the action of a group $G$ of transformations: signal classification can be invariant under translations, image classification under 2D rotations, and spherical-image classification under 3D rotations. The $G$-bispectrum is a principled complete invariant of a signal )retaining all signal's information up to the group action) with proven benefits in machine learning and as a pooling layer in deep networks. However, its deployment has been hampered by high computational cost and a patchwork of group-specific implementations. We present bispectrum, an open-source, fully unit-tested PyTorch library that implements selective $G$-bispectra for seven different group actions, as differentiable modules that can be directly incorporated into machine learning pipelines and deep learning architectures. For finite groups $G$, selectivity reduces the computational cost from $O(|G|^2)$ to $O(|G|)$. For planar rotations, we leverage the disk bispectrum. For spherical 3D rotations, we introduce an augmented selective bispectrum at band-limit $L$ which reduces the cost from $O(L^3)$ to $\Theta(L^2)$ coefficients. We profile the entire library (for


3D Molecule Generation from Rigid Motifs via $\mathrm{SE}(3)$ Flows

Roman Poletukhin ⋅ Marcel Kollovieh ⋅ Eike S. Eberhard ⋅ Stephan Günnemann

Three-dimensional molecular structure generation is typically performed at the level of individual atoms, yet molecular graph generation techniques often consider *fragments* as their structural units. Building on the advances in *frame-based* protein structure generation, we extend these fragmentation ideas to 3D, treating general molecules as sets of *rigid-body motifs*. Utilising this representation, we employ $\mathrm{SE}(3)$-equivariant generative modelling for *de novo* 3D molecule generation from rigid motifs. In our evaluations, we observe comparable or superior results to state-of-the-art across benchmarks, surpassing it in atom stability on GEOM-Drugs, while yielding a $2\times$ to $10\times$ reduction in generation steps and offering $3.5\times$ compression in molecular representations compared to the standard atom-based methods.

Graph representation learning often relies on message passing or spectral/positional encodings to summarize graph structure, but these indirect summaries can collapse structurally distinct graphs, including non-isomorphic cospectral pairs. We propose A2I, a structural encoding framework that renders BFS--degree-ordered local adjacency patterns as fixed-resolution images and embeds them with a frozen pre-trained vision encoder. The resulting structural tokens are mapped to learnable prototypes and aggregated with node embeddings via a lightweight Transformer. We provide a conditional analysis with two related findings: a stability result, showing that two orderings of the same graph yield embeddings that converge under bounded-degree assumptions, and a separability result, showing that two distinct graphs yield embeddings whose gap grows at least linearly with the difference in their BFS profiles. Empirically, A2I is competitive with recent GNN, Graph Transformer, and structural-encoding baselines on six node-classification benchmarks, with pronounced gains for GCN-based models on heterophilous graphs and under partial feature masking.


AcceleGrad#: Adaptive Geometry-Aware Acceleration

Hanka Goralija ⋅ Francesco Tonin ⋅ Kimon Antonakopoulos ⋅ Alp Yurtsever ⋅ Volkan Cevher

Modern neural-network optimizers occupy a three-way design space: online adaptivity, architecture-aware geometry, and acceleration. We introduce AcceleGrad#, an accelerated linear-coupling template that couples an adaptive anchor with a norm-induced sharp, mirror, or linear-minimization geometry branch. The scalar-anchor specialization recovers the accelerated $O(1/K^2)$ rate for smooth convex objectives, while the diagonal-AdaGrad clipped-LMO training variant admits local $\mu$-Kurdyka-Lojasiewicz certificates: a generic conservative $O(K^{-1/4})$ rate and, in a SCION cross-entropy regime, a $O(K^{-1/3})$ last-iterate rate to a self-bounded mini-batch noise level. On image classification and language modeling, the training variant improves over Euclidean AcceleGrad and remains competitive with recent non-Euclidean optimizers. Using stochastic gradients directly in spectral oracles also reduces polar-computation time and optimizer overhead in practice.


Accelerated Image Editing via Consistency-Aware Source Token Pruning

Jeongsol Kim ⋅ Jong Chul Ye ⋅ Yufei Wang ⋅ Jian Wang

Text-based image editing has recently been reinterpreted in large multimodal transformers as conditional generation, where source image tokens are concatenated with text and noise tokens. While effective, this design incurs substantial computational overhead in attention layers. To mitigate this inefficiency, we propose {\em TokenDrop}, an efficient editing framework that shift the computational burden from dense source-token conditioning to a lightweight regularized sampling update. By transferring the influence of source tokens into a closed-form sampling update, the method preserves source consistency with negligible cost. To support this regularized dynamics, we introduce a deviation-guided adaptive masking strategy that selectively drops redundant tokens while maintaining editing fidelity. Across FluxKontext and Qwen-Image-Edit, our training-free method achieves an average 22.4\% improvement in inference speed on PIEBench, while better preserving non-edited regions. The method delivers up to 1.8$\times$ speedup at 1024$^2$ resolution and 2$\times$ speedup at 2048$^2$ resolution. The code will also be released publicly.


Accelerating Neural Network Training with Augmented Koopman Dynamics

Jingyi Huang ⋅ Keyan Miao ⋅ Kostas Margellos ⋅ Paul Goulart

Neural network training can be viewed as a discrete-time dynamical system, suggesting that its optimisation trajectory can be learned and extrapolated to reduce the cost of repeated backpropagation. Motivated by Koopman operator theory, which represents nonlinear dynamics as linear evolution in a lifted space, we propose a training acceleration method based on augmented Koopman dynamics. Instead of approximating the evolution of parameters alone, our method lifts the training state to include both model parameters and optimiser-dependent internal states, enabling the learned operator to better capture the dynamics of modern optimisers with momentum or adaptive gradient statistics. We estimate a strided, multi-step Koopman operator from training snapshots and use matrix-vector multiplications to extrapolate future parameters and optimiser states, effectively bypassing an adaptive number of backpropagation steps. This formulation suppresses high-frequency stochastic fluctuations induced by mini-batch training and reduces the storage required for operator estimation. To improve robustness, we further introduce a safeguarding mechanism that prevents Koopman-extrapolated states from degrading performance relative to the most recent backpropagation iterate. The method can be integrated with standard optimisers, including Adam, AdamW, SGD with momentum, Adadelta, and other commonly used variants. Experiments across multiple optimisers show that the proposed approach significantly reduces both the number of backpropagation steps and the wall-clock training time required to reach a target accuracy, while maintaining, and in many cases improving, final model performance.


Achieving Better Local Regret Bound for Online Non-Convex Bilevel Optimization

Tingkai Jia ⋅ Haiguang Wang ⋅ Ting Wang ⋅ Cheng Chen

Online bilevel optimization (OBO) has emerged as a powerful framework for many machine learning problems. Prior works have developed several algorithms that minimize the standard bilevel local regret or the window-averaged bilevel local regret of the OBO problem, but the optimality of existing regret bounds remains unclear. In this work, we establish optimal regret bounds for both settings. For standard bilevel local regret, we propose an algorithm with adaptive iteration strategy that achieves the optimal regret $\Omega(1+V_T)$ with at most $O(T\log T)$ total inner-level gradient evaluations. We further develop a fully single-loop algorithm whose regret bound includes an additional gradient-variation terms. For the window-averaged bilevel local regret, we design an algorithm that captures linear environmental variation through a novel window-based analysis and achieves the optimal regret $\Omega(T/W^2)$. The algorithm also supports an efficient single-loop structure, achieving an $O(T/W)$ regret bound with $O(WT)$ total gradient evaluations. Experiments validate our theoretical findings and demonstrate the practical effectiveness of the proposed methods.


A Cross-Interaction Neural Architecture for Submodular Functions

SOUTRIK SARANGI ⋅ Aditya Singh ⋅ Vansh Maheshwari ⋅ Abir De

Submodular functions have applications in several domains, \eg, text, vision, speech, \etc. Recent works have proposed neural architectures that are submodular functions by design. However, they do not explicitly capture pairwise interactions between set elements. In this work, we begin with the characterization of pairwise submodular functions, which capture the pairwise interaction across the elements of the input set. We observe that pairwise interactions yield monotone supermodular functions, since the number of terms grows quadratically in terms of the input set size, which poses a significant challenge in converting it into a submodular function. To address it, we provide a novel result that a derivative rate-controlled concave function can transform a monotone supermodular function into a monotone submodular function. We also extend our results to monotone $\alpha$-submodular functions. Leveraging this characterization, we design multi-layer neural cross-interaction architectures for monotone submodular functions and analyze their expressivity. Finally, we perform several experiments which show that our model performs better than existing baselines.


Active Learning From Positive and Unlabeled Examples

Farnam Mansouri ⋅ Sandra Zilles ⋅ Shai Ben-David

Learning from positive and unlabeled data (PU learning) is a weakly supervised variant of binary classification in which the learner receives labels only for (some) positively labeled instances, while all other examples remain unlabeled. Motivated by applications such as advertising and anomaly detection, we study active PU learning, where the learner adaptively queries instances from an unlabeled pool, but a label is revealed only when the queried instance is positive and an independent coin flip succeeds; otherwise the learner receives no information. This paper provides the first theoretical analysis of the label complexity of active PU learning.

Continuous-time recurrent models update hidden states as observations arrive, making them suitable for irregular and non-stationary temporal data. Liquid Neural Networks (LNNs) extend this formulation through input-dependent state dynamics, but they still rely on a single evolving state to integrate new observations and retain earlier context. Under long temporal gaps or missing inputs, earlier evidence can weaken as the hidden state continues to evolve. We propose Active Memory Feedback Loop (AMFL), a memory-augmented LNN architecture that feeds associative retrieval back into the recurrent computation. AMFL combines a learned Short-Term Memory (STM) that stores global sequence prototypes with a sequence-local Iconic Memory (IM) that adapts through novelty-gated, gradient-free updates. After each liquid refinement step, IM retrieves an associative representation, forms a residual correction, and writes the correction into a feedback buffer that conditions the following recurrent update. Memory therefore influences the evolving liquid representation rather than only augmenting the final prediction head. We evaluate AMFL on seven time-series benchmarks covering classification, regression, continual learning, and robustness to missing inputs. AMFL obtains the best result on six of seven benchmarks and remains competitive on the remaining task. Ablations show that the gains arise from looped memory feedback, controlled IM adaptation, and residual correction. Additional missing-input and long-sequence forecasting experiments further support the role of associative feedback in preserving task-relevant context under partial observability.

Offline goal-conditioned reinforcement learning (GCRL) aims to learn a single policy that adapts its behavior to arbitrary goals from a fixed dataset, making injection of the goal information into the policy a central design question. While recent works increasingly leverage expressive conditional generative modeling architectures, which have demonstrated remarkable success in computer vision, architectural choices for goal conditioning in offline GCRL remain under-explored. We identify that goals in GCRL differ from typical conditioning signals in two ways: they are often noisy or redundant, and they are consumed within a time-sequential episode where the relevant context of a fixed goal shifts as the agent's state evolves. Motivated by these observations, we propose State-Anchored Goal Conditioning (SAGC), a modulation-based goal conditioning scheme that derives its scale and shift parameters by jointly processing state and goal representations in a learned latent space. We further introduce AdaptFlow, an end-to-end trainable offline GCRL architecture that integrates conditional flow matching with SAGC. AdaptFlow outperforms matches all baselines on 23 out of 27 OGBench tasks.


Adaptive Communication Range for Scalable Cooperative Multi-Agent Reinforcement Learning

Huizhong Song ⋅ Wei Wei ⋅ Lin Li ⋅ Lijun Zhang ⋅ Fengjiao Li

Cooperative multi-agent reinforcement learning (MARL) has long faced scalability challenges due to the exponential growth of the state-action space as the number of agents increases. Existing methods typically enhance scalability by filtering out communications with low-relevance agents. However, such filtering often relies on global state and fixed graph distribution, limiting adaptability in large-scale, communication-constrained environments. In this paper, we propose a scalable MARL method, called Adaptive Communication Range PPO (ACR-PPO), that decomposes the decision-making under communication budget constraints as a sequential process: a communication policy first selects each agent’s communication range within a given budget, followed by a behavior policy that conditions actions on the resulting neighborhood observations. More importantly, we provide a theoretical guarantee of monotonic performance improvement under communication budget constraints. Experiments across diverse scenarios demonstrate that ACR-PPO preserves policy performance while significantly reducing communication costs through adaptive range control.


Adaptive Fine-Tuning Scheduler for Multi-Tenant Edge LLM via Convergence-Aware Bandits

Yandi Li ⋅ Jianxiong Guo ⋅ Yupeng Li ⋅ Zhiqing Tang ⋅ Tian Wang ⋅ Weijia Jia

Personalizing Large Language Models (LLMs) directly on local edge servers is becoming increasingly important for privacy-preserving and context-aware applications. However, this potential is bottlenecked by hardware resources: limited GPU memory permits only a fixed number of train-ready LoRA adapters, while compute constraints enforce sequential fine-tuning updates. This creates a critical challenge: scheduling scarce update opportunities across a dynamic stream of resident tenants to improve overall model quality and scheduling stability. Crucially, standard Multi-Armed Bandit (MAB) algorithms fail to distinguish between low-potential tenant streams and fully saturated adapters, leading to wasteful exploration on tasks that offer little further gain. To address this, we propose WCA-UCB (Windowed Convergence-Aware UCB), a scheduler designed for this piecewise-converging non-stationary environment. By formally modeling fine-tuning as a slot-constrained bandit problem with piecewise-converging rewards, WCA-UCB detects when an adapter has saturated to pause its training and resets stale statistics after tenant replacement or harmful drift. We prove a dynamic regret bound of $\tilde{O}(\log T)$ and validate the system on an edge-like single-GPU LoRA prototype running Qwen2.5-1.5B. Results demonstrate that WCA-UCB reduces model perplexity by $2.6$% to $8.6$% and uses up to $5.8\times$ fewer training-target switches than strong non-stationary baselines. These results highlight the necessity of convergence-aware scheduling for scalable local LLM personalization.

Gaussian process (GP) bandits provide a powerful framework for performing blackbox optimization of unknown functions. The characteristics of the unknown function depend heavily on the assumed GP prior. Most work in the literature assume that this prior is known but in practice this seldom holds. Instead, practitioners often rely on maximum likelihood estimation to select the hyperparameters of the prior - which lacks theoretical guarantees. In this work, we study two algorithms for joint prior selection and regret minimization in GP bandits based on GP Thompson sampling (GP-TS): Prior-Elimination GP-TS (PE-GP-TS) that disqualifies priors with poor predictive performance, and HyperPrior GP-TS (HP-GP-TS) that utilizes a bi-level Thompson sampling scheme. We theoretically analyze the algorithms and establish a sublinear regret bound for HP-GP-TS. In addition, we demonstrate the effectiveness of these algorithms compared to the alternatives through extensive experiments with synthetic and real-world data.

Speculative decoding accelerates generative inference of large language models (LLMs) by using a small draft model to propose multiple candidate tokens, which are then verified in parallel with the target model in a single decoding iteration. While the state-of-the-art method of tree-based speculative decoding helps improve generation throughput over non-speculative inference, deploying it in LLM serving systems often yields suboptimal performance due to dynamically changing serving conditions. In particular, our analysis shows that the optimal tree configuration---the one that maximizes performance---varies with two key serving conditions: request rate and per-request characteristics. Based on the analysis, we present AdaTree, a plug-in component for LLM serving systems that dynamically adapts tree configurations to varying serving conditions. AdaTree predicts model execution time and acceptance length across different tree configurations and selects the one that maximizes speculative decoding efficiency. To capture the non-linear relationship between tree configuration and model execution time, AdaTree employs decision trees as its core modeling primitive. Our evaluation shows that AdaTree consistently outperforms both chain-based and static tree-based speculative decoding across diverse serving conditions.


Addressing Sparse-Rewards in RL with Scalable Hierarchical Novel Eigen Options

Priyesh Vijayan ⋅ Élodie Côté-Gauthier ⋅ Mathieu Reymond ⋅ Sarath Chandar ⋅ Doina Precup ⋅ Isabeau Prémont-Schwarz

Temporally extended exploration via graph Laplacian-based options is a promising approach to sparse-reward reinforcement learning (RL), but existing methods either do not explicitly target novelty or fail to scale to pixel-based domains under function approximation. Novel Exploration via Orthogonality (NEO) addresses the first issue by constructing options that navigate from highly visited regions toward less visited ones, yet prior results were limited to settings where exact eigenvectors can be computed. We present a scalable extension of NEO to pixel-based domains, built on three contributions. First, we use a novelty-weighted continuous Laplacian graph-drawing objective, which enables RL with continuous observations. Second, we embed the resulting eigen-potential options within a hierarchical reinforcement learning framework, enabling coherent temporally extended behavior. Third, we observe that learned eigen-potential rewards are directional but locally unreliable under online approximation; we therefore augment each option reward with a novelty bonus, a novel design idea that proves essential for stabilizing option learning while preserving novelty-directed exploration. Together, these contributions yield stronger and more persistent exploration, enabling longer option rollouts and better access to hard-to-reach novel states. Empirically, our method significantly outperforms both the prior scalable Laplacian-option baseline and a direct extension of NEO on sparse-reward benchmarks under a fixed budget. On Montezuma's Revenge, our best variant achieves approximately 1.8x higher return than both baselines. On Venture, both baselines yield returns near zero, whereas our method achieves a return of 1135. Across seven hard ProcGen games, our method achieves approximately 3.5x and 5.6x higher aggregate normalized return than the two baselines, respectively.

3D scene graph prediction is commonly supervised with local object and predicate classification losses. Although effective for slot-wise label prediction, this formulation does not provide a graph-level semantic distance between a predicted scene graph and its ground truth: semantically mild label errors and structurally disruptive relational errors can be treated similarly, and subject-predicate-object facts are not compared as scene-level units. We propose Contextual Hellinger Triplet Geometry, a graph-level supervision objective that measures the semantic and structural discrepancy between 3D scene graphs. Our key idea is to reinterpret the conventional classifier outputs of a 3D scene graph predictor as an aligned probabilistic scene graph field, where object and predicate predictions jointly define contextual subject-predicate-object facts. This formulation yields a differentiable metric-based training objective with an efficient factorized implementation and can be applied to existing 3D scene graph prediction architectures without modifying their network design. Experiments on 3DSSG demonstrate that our supervision improves downstream graph-conditioned 3D scene generation, while ablation studies confirm the complementary roles of the proposed graph-level terms and user studies show better graph descriptiveness, relation plausibility, and structural preservation.


Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning

Chaoyu Zhang ⋅ Shanghao Shi ⋅ Heng Jin ⋅ Ning Wang ⋅ Thomas Hou ⋅ Wenjing Lou

Federated learning (FL) has become a foundational paradigm for multi-institutional medical AI, allowing hospitals and research centers to jointly train diagnostic models without exchanging patient records. This privacy promise, however, is increasingly contested: a malicious or honest-but-curious server can launch model inversion attacks (MIAs) that reconstruct private patient images directly from shared model updates, and recent scalable, closed-form attacks penetrate even secure aggregation at clinically realistic batch sizes. Existing defenses face an unsatisfactory dilemma. Gradient-perturbation methods such as differential privacy and pruning trade away the diagnostic accuracy on which clinical reliability depends, while cryptographic protocols add system complexity yet still leave updates exposed to these scalable attacks. We propose Aegis, a principled client-side defense that breaks this dilemma without perturbing patient data or modifying the FL protocol. Our key insight is that the success of every known MIA is fundamentally bounded by the local batch size relative to the model's leakage capacity; once this limit is exceeded, distinct samples collide and reconstructions collapse into indistinguishable mixtures. Aegis turns this universal bottleneck into a defense: each client superimposes onto its real update a masking gradient computed on locally synthesized, task-relevant data, deliberately pushing the effective batch beyond the attack's recovery capacity. We complement the design with theoretical convergence guarantees under standard convex assumptions and evaluate Aegis on MNIST, CIFAR-10, and three MedMNIST modalities (chest X-ray, abdominal CT, colon pathology). Aegis neutralizes three state-of-the-art MIAs while preserving model utility and incurring only modest overhead, offering a practical privacy primitive for medical FL.


AgentBrew: Offline Tool-Use Agent Learning from Raw Real-World Trajectories

Zhiyi Lyu ⋅ Yewen Li ⋅ Longtao Zheng ⋅ shengtian yang ⋅ Lang Feng ⋅ Lei Feng ⋅ Peng Jiang ⋅ Kun Gai ⋅ Qingpeng Cai ⋅ Bo An

LLM-based agents are increasingly deployed in real-world applications through tool-use APIs, yet training them for specific environments remains fundamentally difficult: real-world applications provide no pre-defined tasks or verifiers, no faithful simulators, and limited budget for large-scale environment interaction. In this paper, we propose AgentBrew, an offline training framework that learns effective tool-use policies from a single batch of raw interaction trajectories, without task verifiers or iterative on-policy rollouts. The agent first explores the target environment to collect a raw trajectory corpus without quality filtering. To extract training signal from this noisy corpus, retrospective task inference reconstructs an aligned instruction for each trajectory based on its actual outcome, and PMI-Based credit assignment decomposes the trajectory's total information about the inferred instruction into additive per-action credits via pointwise mutual information (PMI). These credits weight the policy training objective, amplifying informative actions while suppressing ineffective ones. On three real-world MCP applications (GitHub, Notion, PostgreSQL), AgentBrew improves Qwen3-32B by +8.7 Acc / +9.7 Score on average, surpassing Qwen3-235B (+2.3 / +4.4) and outperforming rejection sampling (+5.9 / +10.3). These results demonstrate that fine-grained offline learning can recover useful supervision from raw trajectories that filtering-based approaches would discard.


AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems

Boxuan Zhang ⋅ Jianing Zhu ⋅ Zeru Shi ⋅ Dongfang Liu ⋅ Ruixiang Tang

LLM-based multi-agent systems are increasingly deployed on long-horizon tasks, but a single decisive error is often accepted by downstream agents and cascades into trajectory-level failure. Existing work frames this as \emph{post-hoc failure attribution}, diagnosing the responsible agent and step after the trajectory has ended. However, this paradigm forfeits any opportunity to intervene while trajectory is still unfolding. In this work, we introduce AgentForesight, a framework that reframes this problem as online auditing: at each step of an unfolding trajectory, an auditor observes only the current prefix and must either continue the run or alarm at the earliest decisive error, without access to future steps. To this end, we curate AFTraj-2K, a corpus of agentic trajectories across Coding, Math, and Agentic domains, in which safe trajectories are retained under a strict curation pipeline and unsafe trajectories are annotated at the step of their decisive error via consensus among multiple LLM judges. Built on that, we develop AgentForesight-7B, a compact online auditor trained with a coarse-to-fine reinforcement learning recipe that first equips it with a risk-anticipation prior at the failure boundary on adjacent safe/unsafe prefix pairs, then sharpens this prior into precise step-level localization under a three-axis reward jointly targeting the what, where, and who of an audit verdict. Across AFTraj-2K and an external Who\&When benchmark, AgentForesight-7B outperforms leading proprietary models, including GPT-4.1 and DeepSeek-V4-Pro, achieving up to +19.9\% performance gain and 3$\times$ lower step localization error, opening the loop from post-hoc failure detection to enabling deployment-time intervention.


Agentic Abstention: Do Agents Know When to Stop Instead of Act?

Han Luo ⋅ Bingbing Wen ⋅ Lucy Lu Wang

LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is well specified or achievable in the available environment. In such cases, a reliable agent should recognize that further interaction is unlikely to help and abstain from additional tool calls. We define Agentic Abstention, the problem of deciding when an agent should stop acting under uncertainty. Unlike standard LLM abstention, which is usually evaluated as a single-turn answer-or-abstain decision, agentic abstention is a sequential decision problem: an agent can answer, abstain, or gather more information at each turn, and the need to abstain may only become clear after interacting with the environment. We study this problem across web shopping, terminal environments, and question answering, evaluating 13 LLM-as-agent systems and 2 agent scaffolds on more than 28,000 tasks. Our results show that the main challenge is not only whether agents can abstain, but also when they abstain. Some agents never abstain when they should, while others do so only after many unnecessary interactions. This gap is especially large on tasks where the instruction appears feasible until the environment reveals otherwise (e.g., no valid result matches the instruction). We further find that model scale, reasoning, and agent scaffolding affect abstention in different ways, with larger or more capable models not always performing better at timely abstention. Finally, we introduce CONVOLVE, a context engineering method for improving agentic abstention that distills full interaction trajectories into reusable stopping rules. On WebShop, CONVOLVE substantially improves timely abstention without updating model parameters, raising Llama-3.3-70B's timely recall rate from 26.7 to 57.4. % , and overall recall from 83.2 to 100.0. Our dataset and code are available at https://anonymous.4open.science/r/agentic-abstention-A908


Agentic Multi-Turn Reasoning: A Fairness Approach

Thanh-Dat Truong ⋅ Sankalp Pandey ⋅ Hugh Churchill ⋅ Jackson Cothren ⋅ Marios Savvides ⋅ Khoa Luu

Recent advances in Large Language Models (LLMs) have enabled agentic systems capable of solving complex tasks through multi-turn planning, tool use, verification, and memory updates. However, learning agentic systems remains difficult due to two fundamental challenges, i.e., (1) long-horizon credit assignment, where supervision is available only at the final outcome, and (2) imbalanced data distributions, where dominant data patterns bias optimization and weaken adaptation to rare but informative reasoning behaviors. In this paper, we propose Fair Multi-Level Preference Optimization (Fair-MPO or $\Phi$-MPO), a new preference optimization framework for agentic learning. We first show that Multi-Level Preference Optimization provides a principled and more computationally efficient framework for long-horizon reasoning. Then, we introduce a Fair Multi-Level Objective that addresses imbalance in agentic learning. We provide a comprehensive theoretical analysis demonstrating that our approach addresses both long-horizon reasoning and data imbalance. Our experiments on agentic reasoning benchmarks demonstrate that our approach achieves State-of-the-Art (SOTA) performance.


Agentic Neural Architecture Search

Seokhoon Jeong ⋅ Mijung Kim ⋅ Taehwan Kim

Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces that require substantial domain expertise and must be rebuilt for every new task. Large language models (LLMs) can generate architectures in an open-ended space, but how to optimally divide labor between LLM-driven design and NAS-driven search remains unexplored. We propose a mechanism that bridges these two paradigms: an LLM produces a high-quality seed architecture, then decomposes it into a slotted architecture---a scaffold with named, interchangeable module slots that automatically defines a bounded, task-specific search space for conventional NAS to explore, without manual engineering. We instantiate this mechanism in AgentNAS, a modular three-phase pipeline in which each component's contribution can be measured independently. On 17 tasks spanning classification, dense regression, segmentation, and multi-label tagging across diverse modalities (NAS-Bench-360 and NAS-Unseen), AgentNAS establishes a new state of the art on 11 tasks, outperforming published baselines including task-specific expert designs. Ablation studies show that the two search mechanisms are broadly complementary: the LLM-generated seed already surpasses published baselines on the majority of tasks, and NAS delivers additional gains in most cases through combinatorial recombination across slots---a mode of search that independent LLM samples cannot replicate. These patterns hold across three LLMs of different capability levels, confirming that the division of labor is robust.


AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation

Priyam Sahoo ⋅ Gaurav Mittal ⋅ Xiaomin Li ⋅ Shengjie Ma ⋅ Benjamin Steenhoek ⋅ Pingping Lin ⋅ Yu Hu

Evaluation of software engineering (SWE) agents is dominated by a binary signal: whether the final patch passes the tests. This outcome-only view treats a principled solution and a chaotic trial-and-error process as equivalent. We show that this equivalence is empirically false. We evaluate 2,614 OpenHands trajectories from eight model backends on SWE-bench Verified. Of the 60 tasks in this corpus, 47 have enough passing trajectories to construct task-level process references, yielding a 1,815-trajectory evaluation subset. Among passing trajectories in this subset, 10.7% exhibit behavior we call a Lucky Pass: regression cycles, blind retries, missing verification, or temporally disordered exploration, implementation, and verification. We introduce AgentLens, a framework for process-level assessment of SWE-agent trajectories, and release AgentLens-Bench, a dataset of 1,815 trajectories annotated with quality scores, waste signals, divergence points, and 47 task-level Prefix Tree Acceptor (PTA) references. AgentLens combines two components. First, it merges multiple passing solutions for the same task into a PTA reference space of correct behaviors. Second, it uses a context-sensitive intent-stage labeler that assigns actions to Exploration, Implementation, Verification, or Orchestration using trajectory history rather than tool identity alone. On AgentLens-Bench, the composite score separates passing trajectories into Lucky, Solid, and Ideal tiers; decomposes Lucky Passes into five recurring mechanisms; and changes how the eight evaluated model backends are ranked compared with pass rate alone. Across these models, AgentLens classifies between 0.5% and 23.2% of successful trajectories as Lucky, and some models move by as many as five rank positions when ranked by quality score instead of pass rate. We release the anonymized project repository, including the AgentLens-Bench dataset and AgentLens SDK, at https://anonymous.4open.science/r/agentlens-app-6810/.


AgentWeave: Efficient Distributed Agent Serving via Flow Decomposition

Kaibin Guo ⋅ Pengtu Li ⋅ Zicong Hong ⋅ Zhiyuan Fang ⋅ Wuhui Chen ⋅ Rachid Guerraoui ⋅ Anne-marie Kermarrec

Agent applications are evolving into workflows composed of large language model (LLM) calls and tool invocations, termed modules. Existing distributed agent systems typically adopt a disaggregated architecture, deploying different modules on separate devices. While this enables pipeline parallelism, imbalanced module latencies introduce severe pipeline bubbles that limit throughput. A natural alternative is an aggregated architecture that co-deploys all modules across all devices, eliminating bubbles by keeping every device continuously active. However, this introduces request blocking: short requests are forced to progress in lockstep with longer ones in the same batch, and newly arriving requests must wait until the entire ongoing batch completes the full workflow. To address this, we propose AgentWeave, an efficient distributed serving system for agent applications built on the aggregated architecture. Its core mechanism, flow decomposition, decomposes long agent requests into fine-grained execution units while preserving agent semantics, allowing short and newly arriving requests to proceed without being stalled by long-running ones. Complementing this, a flow-adaptive scheduler dynamically balances the throughput gains of decomposition against its scheduling overhead. Experiments across four representative agent applications demonstrate that AgentWeave improves throughput by $1.81\times$, reduces average latency by $2.45\times$, and reduces P90 latency by $2.04\times$ over state-of-the-art baselines.


AgroOmni: A Large-Scale Multi-view Agricultural Dataset for Cross-Scale Multimodal Reasoning

Jiarui Zhang ⋅ Junqi Hu ⋅ Zurong Mai ⋅ Yang Liu ⋅ Yuhang Chen ⋅ Loushuohong ⋅ Henglian Huang ⋅ Hong Cheng ⋅ Lingyuan Zhao ⋅ Huang Jianxi ⋅ Yutong Lu ⋅ Haohuan Fu ⋅ Juepeng Zheng

Modern agricultural data is sourced from diverse platforms and spans multiple spatial scales, ranging from ground-level close-up photography to Unmanned Aerial Vehicle (UAV) aerial observation and satellite remote sensing imagery. Accordingly, agricultural multimodal reasoning demands robust cross-scale spatial understanding. However, due to the lack of multi-view agricultural benchmark datasets, existing multimodal large language models (MLLMs) exhibit severe ground-level bias, which leads to scale confusion then semantic collapse in agricultural perception tasks, such as misinterpreting farmland imagery as walls or floors. To address this, we introduce \textbf{AgroOmni}, a large-scale multi-view training corpus with 288K Visual Question Answering pairs covering 56 specialized task categories across 14 task types, designed to capture diverse scales in modern precision agriculture. Built on this dataset, we propose AgroNVILA, which achieves a new state-of-the-art of 62.32\% on the AgroMind benchmark ($+15.03\%$ over GPT-5.2), effectively mitigating the multi-view cross-scale gap for holistic agricultural understanding. Diagnostic evaluations on AgMMU further reveal an inherent heterogeneity between macro-priors and micro-diagnostics through constrained zero-shot performance. Meanwhile, even minimal fine-tuning leads to a dramatic performance gain of AgroNVILA on AgMMU, strongly demonstrating its generalization capability empowered by AgroOmni. Full training scripts are publicly available at \url{https://anonymous.4open.science/r/AgroOmni-6510}.


AHPA: Adaptive Hierarchical Prior Alignment for Diffusion Transformers

Ruibin Min ⋅ Yexin Liu ⋅ Aimin PAN ⋅ Changsheng Lu ⋅ Jiafei Wu ⋅ kelu Yao ⋅ Xiaogang Xu ⋅ Harry Yang

Representation alignment has recently emerged as an effective paradigm for accelerating Diffusion Transformer training. Despite their success, existing alignment methods typically impose a fixed supervision target or a fixed alignment granularity throughout the entire denoising trajectory, whether the guidance is provided by external vision encoders, internal self-representations, or VAE-derived features. We argue that such timestep-agnostic alignment is suboptimal because the useful granularity of representation supervision changes systematically with the signal-to-noise ratio. In high-noise regimes, diffusion models benefit more from coarse semantic and layout-level anchoring, whereas in low-noise regimes, the training signal should emphasize spatially detailed and structurally faithful refinement. This non-stationary alignment behavior creates a representational mismatch for static single-level supervisors. To address this issue, we propose Adaptive Hierarchical Prior Alignment (AHPA), a lightweight alignment framework that exploits the hierarchical representations naturally embedded in the frozen VAE encoder. Instead of using only a single compressed latent as the alignment target, AHPA extracts multi-level VAE features that provide complementary priors ranging from local geometry and spatial topology to coarse semantic layout. A timestep-conditioned Dynamic Router adaptively selects and weights these hierarchical priors along the denoising trajectory, thereby synchronizing the alignment granularity with the model's evolving training needs. Extensive experiments show that AHPA improves convergence and generation quality over baselines and incurs no additional inference cost while avoiding external encoder supervision during training.


Align Before Aggregation: Basis-Consistent Federated LoRA under Heterogeneous Ranks

Pengpeng Qiao ⋅ Yang Cao ⋅ Lingling Zhang ⋅ Guo Cheng ⋅ Junwei Chen ⋅ Manjiang Yu ⋅ Wei Yang Bryan Lim ⋅ Masatoshi Yoshikawa

Federated fine-tuning of large language models (LLMs) under heterogeneous client capacities is increasingly important, and low-rank adaptation (LoRA) with heterogeneous ranks makes it practical. Existing server-side aggregators compose local matrix updates $\Delta W_i$ from client-specific LoRA factors $(A_i,B_i)$ before aggregating them. Although these matrix updates lie in the same weight space $\mathbb{R}^{m\times n}$ and can be averaged directly, the LoRA factors that generate them may expressed client-specific local basis. In the typical compose--aggregate--factorize--dispatch loop, this basis inconsistency can make the rank-$r_i$ reference prefix dispatched to lower-rank clients unstable after server-side refactorization. In this work, we propose FedAbA, a basis-consistent federated LoRA aggregation framework for heterogeneous ranks. FedAbA first extracts each client's local basis, aligns it to the corresponding prefix of the global reference basis, and constructs a basis-consistent reconstruction of each local matrix update before aggregation. It then aggregates these reconstructions with weights that combine client data size and consistency score, followed by exact coefficient-space SVD refactorization and gauge fixing before rank-prefix dispatch. Our theoretical analysis clarifies the role of basis consistency and supports the proposed consistency-aware aggregation mechanism. Extensive experiments show that FedAbA outperforms representative federated LoRA baselines, with the largest gains under the strongest rank heterogeneity.


Aligning Flow Map Policies with Optimal $Q$-Guidance

Christos Ziakas ⋅ Alessandra Russo ⋅ Joey Bose

Generative policies based on expressive model classes, such as diffusion and flow matching, are well-suited to complex control problems with highly multimodal action distributions. Their expressivity, however, comes at a significant inference cost: generating each action typically requires simulating many steps of the generative process, compounding latency across sequential decision-making rollouts. We introduce flow map policies, a novel class of generative policies designed for fast action generation by learning to take arbitrary-size jumps---including one-step jumps---across the generative dynamics of existing flow-based policies. We instantiate flow map policies for offline-to-online reinforcement learning and formulate online adaptation as a trust-region optimization problem that improves the critic's $Q$-value while remaining close to the offline policy. We theoretically derive Flow Map $Q$-Guidance Training (FMQ), a principled closed-form learning target that is optimal for adapting offline flow map policies under a critic-guided trust-region constraint. We further introduce $Q$-Guided Beam Search (QGBS), a stochastic flow-map sampler that combines renoising with beam search to enable iterative inference-time refinement. Across $12$ challenging robotic manipulation and locomotion tasks from OGBench and RoboMimic, FMQ achieves state-of-the-art performance in offline-to-online RL, outperforming the previous one-step policy MVP by a relative improvement of $21.3$% on the average success rate.


ALIGN-Rec: Continual Recommendation under Heterogeneous Unlearning Requests

Nitin Bisht ⋅ Sumit Bisht ⋅ Tong Zhang ⋅ Yu Yang ⋅ Huan Huo ⋅ Guandong Xu

Recommender systems (RS) continually learn from new user-item interactions while addressing heterogeneous deletion requests. Existing continual recommendation methods focus on retention but do not erase targeted interactions, whereas recommendation unlearning methods are largely offline and do not ensure stability after future learning updates. This creates a critical failure mode: deleted signals can be reactivated through shared user--item representations, which we call collaborative signal regrowth. To address this, we introduce ALIGN-Rec, a model-agnostic online framework for continual recommendation under interleaved learn/unlearn requests. Specifically, ALIGN-Rec tracks low-rank curvature surrogates to identify retain and forget subspaces, constructs compact retain/forget summaries, and performs geometry-aware updates that preserve retained utility while suppressing deleted-signal directions. We instantiate ALIGN-Rec with CARE, an efficient optimizer based on randomized curvature sketching and adaptive summary selection. We provide theoretical guarantees for geometry tracking, summary approximation, and directional suppression in the forget subspace. Extensive experiments across datasets, deletion granularities, and stream settings demonstrate effectiveness aligning with our theoretical guarantees. The code and implementation are available at https://anonymous.4open.science/r/alignrec-FFD8.


AlloGen: Conformation-Selective Binder Design with Differential State Scoring

Hanqun Cao ⋅ Aastha Pal ⋅ Sumi Kimura ⋅ Yesol Kim ⋅ Jingjie Zhang ⋅ Pheng-Ann Heng ⋅ Pranam Chatterjee

Protein binder design has largely optimized for affinity alone, leaving conformational selectivity unaddressed: for allosteric targets such as kinases, nuclear receptors, and GPCRs, a binder that engages both active and inactive states provides no functional specificity regardless of how tightly it binds. We introduce **AlloGen**, a modular framework that decouples backbone generation from a learned state-selectivity scorer $Q_\theta$, an SE(3)-invariant interface graph transformer trained via a two-phase curriculum that first grounds interface geometry before imposing conformational discrimination. Because $Q_\theta$ is fully differentiable and generator-agnostic, it integrates with any backbone generator as a passive reranker or an active gradient-based guide without retraining. Trained on 65 targets spanning 15 protein families, $Q_\theta$ generalizes to held-out out-of-distribution targets where energy-based baselines fail entirely, and all 15 evaluated generator--guidance combinations achieve positive conformational selectivity averaged over the held-out targets, with the best reaching $\bar{S}=+0.677$. Our anonymous code repository can be found at https://anonymous.4open.science/r/AlloGen_NeurIPS-04CB.


A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions

Giyeong Oh ⋅ Junghun Park ⋅ Yuhan Bae ⋅ Youngjae Yu

Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with VLM captioners replacing sparse alt-text by dense descriptions. Yet a recaptioned corpus is not only a collection of captions: it is a supervision distribution induced by a documented captioning policy ($ \pi $), captioner ($ V_c $ ), and source corpus ($ C $). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions $ D_{\pi,V_c,C} $: at a fixed $ B = 64 $ lexical-unit window it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the shared denominator. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by +3.40 to +6.35 under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus together with the audit-artifact bundle.

Humans, animals, and modern machine learning models exhibit impressive abilities to learn complex behaviors and generalize these behaviors to unseen situations. This ability requires us to learn rules and regularities that allow for such generalizations. At the same time, in most complex environments, any rule will have its exceptions. How do learning systems balance between learning general regularities and memorizing exceptions? We argue that a lack of task paradigms has hindered the study of this essential ability. To address this gap, we introduce a novel task, transitive inference with exceptions, that tests for relational generalization and memorization of an exception to the relational rule. We then analytically characterize the behavior of a simple, theoretically tractable model of neural network learning (kernel ridge regression) across a broad family of representations and task parameters. We find that these models can balance between relational generalization and memorization, but unlike for transitive inference without an exception, successful generalization is sensitive to the specific representational geometry. We explain why this task is more challenging mechanistically by drawing on our analytical theory. Finally, we validate our theoretical insights in pretrained language models that are finetuned on ordered relations, finding that these models successfully generalize according to the transitive rule, but also make the kinds of systematic mistakes predicted by our theory. Overall, our theory shows how learning systems can balance between relational generalization and memorization, explains how this can go wrong, and emphasizes the need for new task paradigms designed to probe this ability.


A Matter of Interest: Understanding Interestingness Judgments of Math Problems in Humans and Language Models

Shubhra Mishra ⋅ Yuka Machino ⋅ Gabriel Poesia ⋅ Albert Q. Jiang ⋅ Joy Hsu ⋅ Adrian Weller ⋅ Challenger Mishra ⋅ David Broman ⋅ Josh Tenenbaum ⋅ Mateja Jamnik ⋅ Cedegao (Ced) Zhang ⋅ Katie Collins

The evolution of mathematics is shaped importantly by interestingness: researchers choose which problems to pursue, and students choose which problems to engage with, based on expectations of interest and challenge. As AI systems, particularly large language models (LLMs) that operate flexibly over natural language and formal mathematics, are increasingly used in mathematics research and education, it becomes crucial to characterize how closely their judgments align with people from different mathematical backgrounds. We study whether LLMs align with human interestingness judgments by comparing LLM ratings with those of two populations, crowdsourced participants with college math experience and International Math Olympiad competitors. Although many LLMs broadly agree with human notions of interestingness, they largely fail to match the distribution of human judgments. They also weakly align with why humans find problems interesting, with low correlation to human-selected rationales. Finally, we evaluate LLMs’ ability to generate interesting problems and find that, after filtering for validity, LLMs are able to generate engaging problems. We conclude with takeaways, including the need for multi-LLM human-AI collaborative systems, that highlight both the promise and current limits of LLMs as partners in mathematical reasoning.

Motivated by the challenge of stabilizing unknown linear dynamical systems (LDS) from observations, we study the fundamental prerequisite of online prediction. Our goal is to achieve sublinear regret with a memory footprint that adapts to the intrinsic complexity of the dynamics rather than the full hidden-state dimension. We focus on the practically central regime of systems with low *instability complexity*—eigenvalues outside the real stable interval that do not decay rapidly, together with non-semisimple modes—potentially embedded in an otherwise stable real spectrum of much higher dimension; we write $k$ for this count. This regime is the primary setting in which stabilization is plausible: we show that many systems with high instability complexity cannot be stabilized without exponentially large controls. Thus, prediction is meaningful for stabilization precisely when the instability complexity is small. Within this regime, we introduce a unified online algorithm that handles every LDS (including systems with complex or exploding modes) with a learnable parameter count of $\widetilde{O}(k^2)$, completely independent of the number of stable real modes. Finally, we show that any bounded-coefficient filter-based predictor requires at least $k$ filter directions.


Amortized-Precision Quantization for Early-Exit Vision Transformers

Rui Fang ⋅ Hsi-Wen Chen ⋅ Ming-syan Chen

Vision Transformers (ViTs) achieve strong performance across vision tasks, yet their deployment with low-precision early exiting remains fragile. Existing quantization methods assume static full-depth execution, making them unstable when exit decisions are perturbed by quantization noise, which can amplify errors along dynamic inference paths. In this paper, we introduce Amortized-Precision Quantization (APQ), a utilization-aware formulation that accounts for layer-wise stochastic exposure to quantization noise and reveals depth--precision trade-offs. Building on APQ, we propose Mutual Adaptive Quantization with Early Exiting (MAQEE), a bi-level framework that jointly optimizes exit thresholds and bit-widths under explicit risk control to improve inference stability. MAQEE establishes a superior Pareto frontier in the accuracy--efficiency trade-off, reducing BOPs by up to 95% while maintaining accuracy and outperforming strong baselines by up to 20% across classification, detection, and segmentation tasks. Source code is available at https://anonymous.4open.science/r/MAQEE-5F2F.

Recovering a signal from its degraded measurements is a long standing challenge in science and engineering. Recently, zero-shot diffusion based methods have been proposed for such inverse problems, offering a posterior sampling based solution that leverages prior knowledge. Such algorithms incorporate the observations through inference, often leaning on manual tuning and heuristics. In this work we propose a rigorous analysis of these approximate posterior samplers, relying on a Gaussianity assumption of the prior. Under this regime, we show that both the ideal posterior sampler and diffusion-based reconstruction algorithms can be expressed in closed-form, enabling their thorough analysis and comparisons in the spectral domain. Building on these representations, we introduce a principled framework for parameter design, replacing heuristic selection strategies used to date. The proposed approach is method-agnostic and yields tailored parameter choices that jointly account for the characteristics of the prior, the degraded signal, and the diffusion dynamics. We show that our spectral recommendations differ structurally from standard heuristics and vary with the diffusion step size, resulting in a consistent balance between perceptual quality and signal fidelity.


ANCHOR: Audio-Visually Grounded Chain-of-Thought Reasoning Benchmark

Joel Julin ⋅ Souraja Kundu ⋅ Liza Dahiya ⋅ George Z Wei ⋅ Liuyue Xie ⋅ Jingru Yi ⋅ Xu Zhang ⋅ Hao Yang ⋅ László Jeni

Multimodal large language models (MLLMs) achieve impressive performance on audio-visual question answering, but correct answers do not necessarily imply perceptually grounded reasoning. Models may exploit linguistic priors or dataset shortcuts without attending to the visual regions and audio events that actually justify their predictions. Existing benchmarks predominantly evaluate final answer accuracy, providing limited insight into whether models truly “see” or “hear” the evidence required for reasoning. We introduce ANCHOR, a benchmark and evaluation framework for auditing perceptual grounding in multimodal reasoning. ANCHOR requires models to produce grounded chain-of-thought explanations that explicitly reference audio-visual evidence, including representative frames, spatial bounding boxes, and salient audio cues. To build a reliable gold standard, we combine automated annotation generation with a human-in-the-loop refinement pipeline, producing 224 human-verified videos and 975 question–answer pairs with aligned reasoning and localization annotations from 2.3K videos and 8.2K QA pairs. We further propose a metric and evaluation protocol that jointly measure answer correctness and reasoning-grounding alignment, complemented by multi-judge scoring of hallucination, logical divergence, final answer, and a token-based conciseness measurement. Experiments on nine state-of-the-art MLLMs reveal a substantial faithfulness gap: models with strong answer performance often produce explanations that are only weakly grounded in the underlying perceptual evidence. Ablation results confirm the importance of explicit grounding cues: removing bounding boxes reduces accuracy by 7.7 points, and performance drops from 57.7% with full cues to 16.1% in the audio-only setting. These findings show that answer-only evaluation is insufficient for multimodal reasoning and position ANCHOR as a new testbed for developing systems that are accurate, interpretable, and trustworthy.

Remote photoplethysmography (rPPG) enables non-contact physiological measurement from facial videos, but performance often degrades under unseen environments due to severe domain shifts. Continual learning offers a practical paradigm for updating rPPG models over sequential domains. However, it is challenged by two coupled issues: First, adapting to new domains may overwrite previously acquired physiological representations, resulting in the degradation of physiological consistency across domains. Second, drastic environmental variations induce heterogeneous facial video distributions, making the adaptation process unstable and prone to suboptimal convergence. To address these challenges, we propose ApexPhys, a continual rPPG framework that simultaneously anchors physiological invariance and expands domain plasticity. Specifically, we introduce a Backbone Consistency Representation Preservation (BCRP) mechanism, which leverages the Fisher Information Matrix to identify and preserve parameters critical to physiological consistency. To enhance adaptation flexibility under diverse environmental shifts, we further propose a Hierarchical Expansion Strategy (HES) that autonomously perceives layer-wise distribution discrepancies and dynamically expands hierarchical adaptation branches. Additionally, we develop a Prior-Guided Pathfinding (PGP) strategy, which leverages parameter inheritance to guide the adaptation along a stable trajectory. Experiments on five public datasets demonstrate that ApexPhys achieves strong anti-forgetting performance and robust adaptation under continual learning.


AnchorRep: Defending LLMs Against Cross-Model Adversarial Transfer via Representation Repulsion

Gal Wertheizer ⋅ Rom Himelstein ⋅ Tomer Peretz ⋅ Avi Mendelson

Adversarial attacks optimized on a single open-weight LLM can transfer to and jailbreak architecturally different models, allowing an attacker with white-box access to one model to compromise independently deployed systems. This creates a shared vulnerability across models, yet existing defenses are not designed for this cross-model threat. We find that transfer aligns with shared internal representation geometry, making it a natural defense target. We find that cross-model transfer aligns with shared internal representation geometry, making it a natural defense target. \textbf{AnchorRep} targets this geometry directly with a lightweight LoRA adapter that pushes the defended model's internal representations of harmful prompts away from those of a frozen anchor model on the same prompts. Training uses a small set of harmful prompts and no adversarial examples. Across five models and four architectural families, AnchorRep reduces cross-model attack success rate to $\leq$1.1\% on 2{,}000 transferred attacks (0\% on two), including the largest drop on Mistral ($36\% \to 1.1\%$). Existing defenses can reduce transfer, but only at high cost—either inducing up to 77\% degenerate benign output or increasing over-refusal by up to 18\%. Because such degenerate benign outputs are not captured by standard refusal-based metrics, we introduce the \textbf{Benign Garble Rate} to quantify them. Our results suggest that cross-model robustness can be achieved by shaping representation geometry, without requiring attack-specific training. \smallskip \noindent\faGithub~Code, configs and logs: \href{https://anonymous.4open.science/r/AnchorRep/README.md}{anonymous.4open.science/r/AnchorRep}


A New Perspective on Target-Conditioned Structural Dynamics for Link Prediction in Dynamic Graphs

YUANYUAN XU ⋅ Yin Chen ⋅ Yingxuan Li ⋅ Wenjie Zhang ⋅ Xuemin Lin ⋅ Ying Zhang

Dynamic link prediction requires determining whether the historical interactions of a source node provide reliable structural support for a specific target. Existing target-conditioned methods capture such support through structural heuristics such as repeated interactions or co-neighbor overlap, but typically instantiate them as identity-based encodings and inject them as passive token-level features before aggregation. This design is brittle when exact structural encodings are sparse, and the resulting structural signals can be diluted after being concatenated with high-dimensional time and edge features, while their dynamics remain unmodeled. In this paper, we revisit target-conditioned dynamic link prediction from a structural dynamics modeling perspective and propose TCSD, which formulates structural signals as target-conditioned states, evolves them over recent histories, and uses them to guide aggregation. Concretely, we construct hybrid structural states from exact indicators that preserve time-aware repeat and co-neighbor matches, and relaxed indicators that recover latent support beyond exact identity matching. TCSD further models the dynamics of these states to capture temporal consistency within indicators and complementarity across indicators. Instead of treating structural states as passive features, TCSD uses the learned dynamics as aggregation controllers to amplify target-relevant interactions and suppress irrelevant ones. Experiments on sixteen dynamic graphs show that TCSD outperforms ten baselines, achieving up to a 14.27% relative improvement in MRR.

Online forecasters sometimes know a regime boundary before any post-boundary labels arrive. Standard adaptive conformal methods still react through later coverage errors, so they can pay a detection-delay cost even when the boundary is public. We study conformal prediction with an announced partition of the stream and separate two post-break audits. The reset empirical quantile in segmented split-CP attains the minimax per-time probability-gap rate $\Theta(M^{-1/2})$ after $M$ post-reset scores, with a matching Le Cam lower bound. A pure bounded constant-step announced-break ACI recursion attains an $O(M^{-1})$ time-averaged empirical-frequency gap by telescoping, but this is a realized-frequency statement, not a guarantee for the next interval. The practical AB-ACI implementation used in experiments adds burn-in and projection, so its rate claim includes explicit correction terms and is not unconditional. Synthetic experiments recover both slopes and show short-window gains when the break is large. Two real-data applications are null under leakage-free protocols. ERCOT RTC+B gives 0.863 post-reform coverage for AB-ACI versus 0.880 for ACI, and FOMC yield breaks give 0.892 versus 0.900. The message is diagnostic: use the reset empirical quantile for per-prediction reliability, use AB-ACI only for long-run frequency audits, and expect gains only when the realized score shift is large enough to make detection delay costly.


Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information

Guobin Shen ⋅ Xiang Cheng ⋅ Chenxiao Zhao ⋅ Lei Huang ⋅ Jindong Li ⋅ Dongcheng Zhao ⋅ XingYu Li

On-policy self-distillation, where a student is pulled toward a copy of itself conditioned on privileged context (e.g., a verified solution or feedback), offers a promising direction for advancing reasoning capability without a stronger external teacher. Yet in math reasoning the gains are inconsistent, even when the same approach succeeds elsewhere. A pointwise mutual information analysis traces the failure to the privileged context itself: it inflates the teacher's confidence on tokens already implied by the solution (structural connectives, verifiable claims) and deflates it on deliberation tokens (Wait, Let, Maybe) that drive multi-step search. We propose Anti-Self-Distillation (AntiSD), which ascends a divergence between student and teacher rather than descending it: this reverses the per-token sign and yields a naturally bounded advantage in one step. An entropy-triggered gate disables the term once the teacher entropy collapses, completing a drop-in replacement for default self-distillation. Across five models from 4B to 30B parameters on math reasoning benchmarks, AntiSD reaches the GRPO baseline's accuracy in 2 to 10× fewer training steps and improves final accuracy by up to 11.5 points. AntiSD opens a path to scalable self-improvement, where a language model bootstraps its own reasoning through its training signal.


AnyBokeh: Physics-Guided Any-to-Any Bokeh Editing with Optical Fingerprint Transfer

Xinyu Hou ⋅ Xiaoming Li ⋅ Zongsheng Yue ⋅ Chen Change Loy

Depth-of-field control is a fundamental tool in photography, yet post-capture bokeh editing from a single image remains challenging. A practical editor should handle images captured under arbitrary focus and aperture settings. Existing methods typically assume an all-in-focus input, or first recover an all-in-focus image before rendering new bokeh. Such pipelines can discard useful blur cues from the source image and propagate reconstruction artifacts into the final edit. We introduce AnyBokeh, a physics-guided framework for any-to-any bokeh editing. Instead of treating source blur merely as a degradation to be removed, AnyBokeh estimates the source blur state with a signed circle-of-confusion map and a disparity map. By modeling the linear relation between signed circle of confusion and disparity difference, AnyBokeh estimates a source-specific optical fingerprint and transfers the source optical characteristics to the desired focus and aperture setting. A generative editor conditioned on both source and target circle-of-confusion maps then performs relative blur synthesis, enabling spatially adaptive deblurring, preservation, and defocus rendering. To support physically supervised learning, we further construct a high-fidelity synthetic dataset with accurate depth, focus distance, and full EXIF metadata. Experiments on real-world benchmarks show that AnyBokeh achieves faithful and controllable editing across any-to-any bokeh editing, all-in-focus-to-bokeh rendering, and defocus deblurring, while avoiding all-in-focus reconstruction and test-time bokeh-level calibration commonly required by existing approaches. The code and dataset will be made publicly available.


Anytime-Valid Conformal Risk Control

Bror Hultberg ⋅ Dave Zachariah ⋅ Antonio Ribeiro

Prediction sets provide a means of quantifying the uncertainty in predictive tasks. Using held out calibration data, conformal prediction and risk control can produce prediction sets that exhibit statistically valid error control in a computationally efficient manner. However, in the standard formulations, the error is only controlled on average over many possible calibration datasets of fixed size. In this paper, we extend the control to remain valid with high probability over a cumulatively growing calibration dataset at any time point. We derive such guarantees using quantile-based arguments and illustrate the applicability of the proposed framework to settings involving distribution shift. We further establish a matching lower bound and show that our guarantees are asymptotically tight. Finally, we demonstrate the practical performance of our methods through both simulations and real-world numerical examples.


Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence

Runmao Yao ⋅ Kairui Hu ⋅ Yukang Cao ⋅ Ruisi Wang ⋅ Shulin Tian ⋅ Ziang Cao ⋅ Weichen Fan ⋅ Ziqi Huang ⋅ Yuhao Dong ⋅ Hao Li ⋅ Zhaoxi Chen ⋅ Zhongang Cai ⋅ Lei Yang ⋅ Ziwei Liu

Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the output level, without verifying whether the model arrives there through a faithful, law-grounded reasoning process. We introduce **Apple-$\pi$**, the first benchmark that anchors video-model evaluation explicitly in physical laws. Apple-$\pi$ comprises three components. **1) Orchard**: a dataset of 400 videos covering ten canonical tasks in classical mechanics. It separates single-law tasks for confounder-free diagnosis from multi-law tasks for probing generalization. **2) Benchmark Protocol**: a three-stage protocol based on scientific reasoning, including Perception, Formulation, and Deduction. It uses *chain-of-frames* prompting on infographic-annotated first frames, treating the generated video as the model's visible reasoning trace. **3) Evaluation Suite**: a hybrid evaluation suite that combines MLLM-based subjective scoring with physics-law-grounded objective measures. This enables stage-resolved diagnosis of not only whether a model fails, but where it fails. Benchmarking 11 models shows that current video models remain far from reliable law-grounded world simulators, with the best video model scoring only 0.473. Our stage-, pillar-, and source-resolved analyses further expose a Perception-to-Formulation-to-Deduction bottleneck, weak multi-law state transfer, and a persistent Sim-to-Real gap. These findings position Apple-$\pi$ as a diagnostic foundation for guiding future video models toward world models with law-grounded physical intelligence.


Approximation in Contrastive Representation Learning

Yuanfan Li ⋅ Zihan Zhang ⋅ Yiming Ying ⋅ Ding-Xuan Zhou

Contrastive representation learning (CRL) is a fundamental paradigm for learning transferable representations. Recent work identifies its population target as the pointwise mutual information (PMI) $ \log\frac{p(\mathbf{x},\mathbf{y})}{p_\mathcal{X}(\mathbf{x})p_\mathcal{Y}(\mathbf{y})} $ between two modalities $\mathcal{X},\mathcal{Y}$, where $p$ is the joint density function and $p_\mathcal{X},p_\mathcal{Y}$ are the marginals. However, it remains unclear why standard inner-product scores $f(\mathbf{x})^\top g(\mathbf{y})$ can effectively approximate a generally nonseparable population target PMI. In this paper, we address this question from the aspect of approximation for CRL. By linking PMI to a compact operator, we show that $D$-dimensional inner-product models perform rank-$D$ spectral approximation, with error controlled by the spectral tail. We then study neural network realizations: in general settings, we show that deep neural network scoring functions are universally consistent, and the approximation error rates depend on the smoothness of the PMI function; under structured low-rank data models we obtain improved rates attained by moderate-width encoders. Finally, we establish a fast calibration rate from contrastive excess risk to downstream retrieval error under a density-ratio Tsybakov noise condition.


AR1-ZO: Topology-Aware Rank-1 Zeroth-Order Queries for High-Rank LoRA Fine-Tuning

Ziye Chen ⋅ Hongbin Lin ⋅ Chenyu Zhang ⋅ Xiangda Yan ⋅ Yongjie Yang ⋅ Yao SHU

Zeroth-order (ZO) optimization enables large-language-model fine-tuning without storing backpropagation activations, while LoRA supplies compact trainable adapters. Combining them creates a rank paradox: increasing LoRA rank improves adapter capacity, but standard two-point ZO either perturbs a rank-dependent number of coordinates or, under atomwise updates, can make the finite-difference signal unobservable. This paper shows that the bottleneck is a measurement-topology problem rather than a need for an external subspace. LoRA already decomposes into matched rank-$1$ atoms, each a complete factor-coordinate block of dimension $d_\text{out}+d_\text{in}$. Querying one atom per step keeps the stored adapter rank $r$ while removing $r$ from the single-query perturbation dimension. The naive atomwise query is still miscalibrated: if it inherits canonical LoRA scaling $\alpha/r$, the active finite-difference signal shrinks as $1/r$ and the active finite-difference signal-to-noise ratio (FD-SNR) as $1/r^2$, producing directional collapse under a fixed residual evaluation-noise floor. AR1-ZO pairs alternating rank-$1$ atom queries with topology-aware scaling $\gamma=\alpha r$, restoring rank-invariant active signal without auxiliary bases, activation hooks, curvature estimates, or extra forward queries. Theory proves atom minimality, rank-independent active query dimension, directional collapse and restoration, and the remaining rank dependence as an amortized coverage cost. Experiments on OPT and Qwen3 models validate the signal mechanism and show that AR1-ZO makes high-rank LoRA effective among matched-budget ZO methods under the standard two-forward-pass query budget.


Archimedean Copula Inference via Taylor-Mode AD

Cambridge Yang ⋅ Dongdong Li

No existing nested Archimedean copula tool handles all three of (a) arbitrary per-variable (right-)censoring in survival analysis, (b) arbitrary nesting trees, and (c) exact parameter gradients. Existing implementations handle only bivariate problems, low dimensional (i.e., $d \leq 10$) cases, two layers of nesting, or only hand-derived copula nestings. We present acopula, a JAX-native framework that, given any Archimedean generator—classical or neural—evaluates exact nested-copula likelihoods and parameter gradients under arbitrary censoring masks in polynomial time. The mechanism is polynomial powering of Taylor-mode automatic differentiation output, which replaces per-family hand-derived partial Bell polynomial tables with a single differentiable computation that any user-defined generator can drive. We conduct extensive simulations to verify the correctness of acopula. We then demonstrate (a) per-variable censoring on $85,229$ MIMIC-IV ICU admissions in high dimensions with $d=53$, fit by both classical Archimedean families and nested neural Archimedean copulas; (b) an 11-sector hierarchical model on S&P 500 daily returns at $d=98$; (c) family-agnostic censored MLE across ten families, five of them with no prior implementation, on a retinopathy study; and (d) a $\sim650\times$ per-density speedup over R's nacLL at $d=35$, scaling quadratically to $d=8,000$.

LLM-as-a-Judge pipelines have become the de facto evaluator for agent safety, yet existing benchmarks treat their verdicts as ground-truth proxies without checking whether the verdicts depend on the agent's behavior or merely on how the evaluation policy happens to be worded. We argue that any trustworthy safety judge must satisfy a basic property we call policy invariance, and we operationalize it as three testable principles: rubric-semantics invariance under certified-equivalent rewrites, rubric-threshold invariance under intentional strict-to-lenient shifts, and ambiguity-aware calibration so that verdict instability concentrates on genuinely ambiguous cases. Instantiating these principles as a stress-test protocol with four agent-class judges on trajectories drawn from ASSEBench and R-Judge, we surface a previously unmeasured failure mode: today's judges respond to meaningful normative shifts and to meaningless structural rewrites with comparable strength, and cannot tell the two apart. Content-preserving policy rewrites flip up to 9.1\% of verdicts above baseline jitter, and 18-43\% of all observed flips occur on unambiguous cases under such rewrites, so existing safety scores conflate what the agent did with how the evaluator was prompted. Beyond the diagnosis, we contribute the Policy Invariance Score and the Judge Card reporting protocol, which expose an order-of-magnitude spread in judge reliability that is invisible to accuracy-only leaderboards. We release the protocol and code so that future agent-safety benchmarks can audit their own evaluators rather than trust them by default: https://anonymous.4open.science/r/policy-invariance-judge


Are Text-to-Image Models Inductivist Turkeys? A Counterfactual Benchmark for Causal Reasoning

jiayi lei ⋅ Yuandong Pu ⋅ Xingyu Han ⋅ Rongpeng Zhu ⋅ Jing Xu ⋅ Jinyao Wang ⋅ Zijian Zhou ⋅ Bin Fu ⋅ Yuewen Cao ⋅ Yihao Liu ⋅ Hongsheng Li

Text-to-image (T2I) generation models have achieved remarkable progress in producing visually realistic images from natural language prompts. Yet it remains unclear whether their success reflects genuine causal understanding or sophisticated pattern matching over visual-textual correlations. Inspired by Russell’s inductivist turkey, we introduce Counterfactual-World (CF-World), a counterfactual benchmark designed to probe whether text-to-image models can generate images under rules that systematically contradict real-world priors. CF-World organizes each scenario into three progressive levels: factual generation under ordinary world knowledge, explicit counterfactual generation with direct visual instructions, and implicit counterfactual generation requiring causal deduction from altered rules. We evaluate both open-source and closed-source T2I models using a Vision Language Model(VLM)-based evaluator (CF-Eval), and introduce two metrics: Prior Resistance Rate (PRR), which measures a model’s ability to overcome entrenched real-world priors, and Reasoning Retention Rate (RRR), which assesses whether models can maintain reasoning-dependent counterfactual generation without explicit visual cues. Experiments show that all models exhibit sharp degradation from factual to counterfactual settings. Further analyses suggest that these failures arise because current T2I models encode world knowledge and visual appearances as tightly coupled patterns. They rely heavily on frequent visual co-occurrences in their training data, causing them to default to familiar commonsense priors when tasked with generating counterfactual worlds.


ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge

Yijie Lin ⋅ Guofeng Ding ⋅ Haochen Zhou ⋅ Haobin Li ⋅ Mouxing Yang ⋅ Xi Peng

Existing multimodal retrieval benchmarks largely emphasize semantic matching on daily-life images and offer limited diagnostics of professional knowledge and complex reasoning. To address this gap, we introduce ARK, a benchmark designed to analyze multimodal retrieval from two complementary perspectives: (i) knowledge domains (five domains with 17 subtypes), which characterize the content and expertise retrieval relies on, and (ii) reasoning skills (six categories), which characterize the type of inference over multimodal evidence required to identify the correct candidate. Specifically, ARK evaluates retrieval with both unimodal and multimodal queries and candidates, covering 16 heterogeneous visual data types. To avoid shortcut matching during evaluation, most queries are paired with targeted hard negatives that require multi-step reasoning. We evaluate 26 representative text-based and multimodal retrievers and observe a pronounced gap between knowledge- and reasoning-intensive retrieval, with fine-grained visual and spatial reasoning as persistent bottlenecks. We further show that enhancements such as re-ranking, rewriting, and agentic retrieval yield consistent gains, but substantial headroom remains. The dataset and evaluation pipeline will be released soon.


A Simple Unigram Cross-Entropy Lens on the Lexical Imprint of Pre-training Data

Jeonghoon Kim ⋅ Woojin Chung ⋅ Woomin Song ⋅ Cheonbok Park ⋅ Nako Sung ⋅ Jinwoo Shin

How much of zero-shot benchmark accuracy is captured by word-level alignment between a pre-training corpus and the benchmark target? We study this question through word-level unigram cross-entropy (UCE), a tokenizer-agnostic, model-free measure of lexical alignment between a reference corpus and a target benchmark. Likelihood training allocates probability mass according to the pre-training corpus, and likelihood-based zero-shot evaluation partly reads out the lexical alignment between this mass and the benchmark target. Cleanly tracing how the lexical imprint of pre-training data scales with corpus--benchmark alignment requires controlled from-scratch pre-training that varies corpus identity while holding the training recipe fixed. Across this controlled sweep, spanning $11$ zero-shot benchmarks, $4$ pre-training corpora, and $3$ model scales, we find that corpus--benchmark lexical alignment, UCE, consistently tracks zero-shot accuracy on every benchmark, with corpora more aligned with the benchmark (lower UCE) achieving higher accuracy. The same signal extends from measurement to control at finer scopes. At the document level, strengthening the imprint by selecting training documents to maximize lexical alignment with the benchmark yields a data-selection rule that improves accuracy over same-size random sampling and an importance-resampling baseline across sources and scales. At the prompt level, matching the imprint per prompt by routing each prompt to the drafter whose domain corpus is most lexically aligned with it gives a training-free routing rule for speculative decoding and improves throughput over a generalist-drafter baseline. Taken together, pre-training corpora leave a measurable lexical imprint on zero-shot behavior, and UCE provides a model-free probe of lexical alignment between reference and target data.


Asking the Right Questions: Improving Reasoning with Generated Stepping Stones

Hengyuan Hu ⋅ Tingchen Fu ⋅ Minqi Jiang ⋅ Alexander Miller ⋅ Yoram Bachrach ⋅ Jakob Foerster

Recent years have witnessed tremendous progress in enabling LLMs to solve complex reasoning tasks such as math and coding. As we start to apply LLMs to harder tasks that they may not be able to solve in one shot, it is worth paying attention to their ability to construct intermediate stepping stones that prepare them to better solve the tasks. Examples of stepping stones include simplifications, alternative framings, or subproblems. We study properties and benefits of stepping stones in the context of modern reasoning LLMs via ARQ (Asking the Right Questions), a simple framework that introduces a question generator to the default reasoning pipeline. We first show that good stepping stone questions exist and are transferrable, meaning that good questions can be generated, and they substantially help LLMs of various capabilities in solving the target tasks. We next frame stepping stone generation as a post-training task and show that we can fine-tune LLMs to generate more useful stepping stones by SFT and RL on synthetic data.


AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

Edward De Brouwer ⋅ Carl Edwards ⋅ Alexander P Wu ⋅ Jenna L Collier ⋅ Graham Heimberg ⋅ Xiner Li ⋅ Meena Subramaniam ⋅ Ehsan Hajiramezanali ⋅ David Richmond ⋅ Jan-Christian Huetter ⋅ Sara Mostafavi ⋅ Gabriele Scalia

Recent advances in machine learning and large-scale biological data collections have revived the prospect of building a virtual cell, a computational model of cellular behavior that could accelerate biological discovery. One of the most compelling promises of this vision is the ability to perform in silico phenotypic screens, in which a model predicts the effects of cellular perturbations in unseen biological contexts. This task combines heterogeneous textual inputs with diverse phenotypic outputs, making it particularly well-suited to LLMs and agentic systems. Yet, no standard benchmark currently exists for this task, as existing efforts focus on narrower molecular readouts that are only indirectly aligned with the phenotypic endpoints driving many real-world drug discovery workflows. In this work, we present AssayBench, a benchmark for phenotypic screen prediction, built from 1,920 publicly available CRISPR screens spanning five broad classes of cellular phenotypes. We formulate the screen prediction task as a gene rank prediction for each screen and introduce the adjusted nDCG, a continuous metric for comparing performance across heterogeneous assays. Our extensive evaluation shows that existing methods remain far from empirically estimated performance ceilings and zero-shot generalist LLMs outperform biology-specific LLMs and trainable baselines. Optimization techniques such as fine-tuning, ensembling, and prompt optimization can further improve LLM performance on this task. Overall, AssayBench offers a practical testbed for measuring progress toward in-silico phenotypic screening and, more broadly, virtual cell models. Our benchmark is available at https://github.com/Genentech/AssayBench.


Assessing Sample Quality in Conditional Generation under Compositional Shift

Berker Demirel ⋅ Valentino Maiorca ⋅ Marco Fumero ⋅ Theofanis Karaletsos ⋅ Francesco Locatello

Conditional generators provide a natural tool for controllable generation, including settings where the desired condition is a new composition of observed attributes or experimental factors. In many applications, especially in scientific domains, such models are attractive to explore conditions for which real samples are rare, expensive, or not yet observed. However, this creates a circularity for evaluation: standard conditional quality metrics require a reference target distribution, but in the extrapolative regime that distribution is unavailable by definition. We address this problem with a post-hoc, per-sample trust score for assessing conditional samples using only the training distribution. The score combines two estimable quantities: global realism, measuring compatibility with the real data manifold, and attribute-wise faithfulness, measuring whether a sample is closer to the requested attributes than to plausible alternatives. We show that the score can recover meaningful comparisons across extrapolated generations, under a mild coverage condition on the observed attributes. These comparisons enable effective filtering, ranking, and abstention of generations and can be used directly on off-the-shelf pretrained models. In biological imaging, selected samples preserve real morphological structure better and improve downstream predictive performance, while similar gains are observed on controlled vision benchmarks. Finally, we show how the score can be applied during generation, enabling abstention before full decoding.


Asymmetric Invertible Threat: Learning Reversible Privacy Defense for Face Recognition

Jiabei Zhang ⋅ Ziyuan Yang ⋅ Andrew Beng Jin Teoh ⋅ Yi Zhang

Face Recognition systems are widely deployed in real-world applications, but they also raise privacy concerns due to unauthorized collection and misuse of facial data. Existing adversarial privacy protection methods rely on input-space perturbations to obfuscate identity information, yet their protection can degrade when adversaries learn restoration or purification mappings that partially invert the transformation. We study this setting as an asymmetric adversarial attack, in which reverse manipulation becomes feasible because existing defense paradigms do not control reversibility. To address this problem, we propose Asymmetric Reversible Face Protection (ARFP), a restoration-aware extension of personalized face cloaking that integrates privacy protection, keyed recovery, and tamper indication in a single framework. ARFP consists of three components: Key-Conditioned Manifold Binding, which ties the protection transformation to a user-provided key; Adversarial Restoration-Aware Training, which introduces a surrogate restoration adversary during training to improve robustness against evaluated inverse purification attacks; and Authorized Reversible Restoration, which supports recovery with the correct key while providing nonce-based tamper indication. Extensive experiments under the threat models considered in this work show that ARFP improves resistance to the evaluated restoration attacks while preserving authorized recovery utility. These results provide empirical evidence of key-sensitive recovery behavior and tamper awareness in the tested settings.


Asynchronous Agentic Poisoning

Guangnian Wan ⋅ Shizun Wang ⋅ Qi Li ⋅ Xinchao Wang

LLM agents are increasingly studied as systems that can self-evolve during deployment. A representative mechanism for realizing self-evolution is memory evolution, where agents maintain evolving memories derived from previous tasks and condition future inference on their memory. Although memory evolution is intended to improve agent performance, it also creates a feedback loop in which model-generated content can persist across tasks and influence future decisions. We study this loop as a delayed attack channel for memory-evolving agents. We introduce asynchronous agentic poisoning, a provider-side threat in which a malicious provider releases a model that appears safe under fresh-context evaluation but exhibits unsafe behavior after self-evolving through memory evolution. We realize this threat by fine-tuning the released model to silently embed a hidden marker into its otherwise normal responses, causing the marker to accumulate in the agent’s memory through memory evolution. When the accumulated markers are later retrieved into the input context, the model reacts to them as a backdoor trigger and shifts to attacker-specified behavior. This creates a self-triggered transition from safety-aligned behavior to attacker-specified behavior without requiring any post-deployment attacker interaction. Experimental results show that our method increases the harmful rate from below 5\% to around 90\% after memory evolution, while not substantially degrading the model’s self-evolution performance. These findings demonstrate that memory evolution creates a provider-side pathway for delayed compromise, allowing a model to appear safe at release time while becoming unsafe after deployment through the agent’s self-evolution process.


A Theory of Time-Sensitive Language Generation: Sparse Hallucination Beats Mode Collapse

Atul Ganju ⋅ Travis McVoy ⋅ Shaddin Dughmi ⋅ Shang-Hua Teng

We study language generation in the limit under a global preference ordering on strings, as introduced by Kleinberg and Wei. We aim for breadth, but impose an additional requirement of timeliness: higher-ranked strings should be generated earlier. A string is then only credited if it is generated before a deadline, where its deadline is defined by a function that maps a string’s rank in the target language to the time by which it must be produced. This is in keeping with a central consideration in machine learning, where inductive bias favors "simpler" or "more plausible" outputs, all else being equal. We show that timely generation is impossible in a strong sense for eventually consistent generators—the protagonists of most prior related work. Under what is perhaps the mildest natural relaxation of consistency, a hallucination rate that vanishes over time, we show that we can circumvent our impossibility result. In particular, we can achieve optimal density with respect to any superlinear deadline function. We also show this is tight by ruling out timely generation with linear deadlines and vanishing hallucination rate.


Atomic Trajectory Modeling with State Space Models for Biomolecular Dynamics

Liang Shi ⋅ Jiarui Lu ⋅ Junqi Liu ⋅ Chence Shi ⋅ Zhi Yang ⋅ Jian Tang

Understanding the dynamic behavior of biomolecules is fundamental to elucidating biological function and facilitating drug discovery. While Molecular Dynamics (MD) simulations provide a rigorous physical basis for studying these dynamics, they remain computationally expensive for long timescales. Recent deep generative models accelerate conformation generation but often either discard temporal correlations entirely or struggle to condition on the extended history required for faithful kinetics, a consequence of the effectively non-Markovian nature of partially observed subsystem coordinates. To bridge this gap, we introduce ATMOS, a novel generative framework based on State Space Models (SSM) designed to generate atom-level MD trajectories for biomolecular systems. ATMOS integrates a Pairformer-based state transition mechanism to capture temporal dependencies, with a diffusion-based module to decode trajectory frames autoregressively. We demonstrate that ATMOS achieves state-of-the-art performance in generating conformation trajectories for both protein monomers and protein-ligand systems. This work provides a unified and computationally efficient framework for biomolecular trajectory generation, taking a step toward dynamics foundation models.

Graph-structured optimization with linear constraints is fundamental to critical infrastructure but faces scalability limits due to massive strict hard constraints and high dimensionality. While recent projection-based methods such as Trainable Sampling Kaczmarz-Motzkin Net (T-SKM-Net) guarantee feasibility, they face high computational costs in dynamic environments by processing the entire constraint set and requiring expensive matrix factorizations. To bridge this gap, we propose the Accelerated Trainable-SKM (AT-SKM) Net framework. To concentrate computation on the active constraints and eliminate redundant calculations, we introduce a hybrid sampling strategy guided by a topology-aware heterogeneous GNN model. To efficiently handle topological shifts in graph-based constraints, we employ a Cholesky Update mechanism that theoretically reduces the equality projection complexity from $\mathcal{O}(N^3)$ to $\mathcal{O}(N^2)$ under low-rank perturbations. Experiments on random geometric graphs, N-1 Security-Constrained DC-OPF, and minimum-cost gas transport problem demonstrate that AT-SKM reduces iteration counts by up to 85% and achieves 2.95$\times$-7.29$\times$ SKM layer speedups, while maintaining zero constraint violations. Our code is available at [https://anonymous.4open.science/status/Anonymous_Submission-CFC7](https://anonymous.4open.science/status/Anonymous_Submission-CFC7)

Modern video object segmentation (VOS) relies on a two-stage paradigm for robust performance: target initialization and memory-based propagation. However, existing adversarial attacks on VOS often overlook this fundamental feature. To bridge this gap, we propose AUV, an adversarial framework that Aligns with the Universal two-stage paradigm of VOS models. Specifically, AUV optimizes a single universal adversarial perturbation (UAP) through decoupled training strategies tailored for both initialization and propagation stages. To ensure prompt-agnosticism across all modalities, including masks, points, boxes, and language, AUV optimizes the UAP to disrupt the shared memory space for latent-level target erosion. Extensive experiments show that AUV establishes a new state-of-the-art for VOS attacks, such as degrading mask-based SAM3 to only 8.4 J&F on DAVIS17-val (compared to 55.4 for previous SOTA under the same training setup), while consistently outperforming prior methods on other VOS models including SAM2, Cutie, and XMem. Notably, AUV exhibits superior temporal dynamics with the fastest onset and the most enduring persistence, while uniquely supporting flexible any-time mid-video attack, revealing vulnerabilities in real-world streaming applications. Our code and checkpoint will be made publicly available.


Attention Alignment Between Humans and Vision-Language Models

Isaac Christian ⋅ Udith Haputhanthri ⋅ Declan Campbell ⋅ Samuel Nastase ⋅ Taylor Webb ⋅ Michael S Graziano

Visual perception depends on top-down goals and bottom-up sensory mechanisms. Vision-language models implement both, allowing us to treat each component as a separable hypothesis about what drives where we look. We compared spatial attention maps from six vision-language models against human fixation heatmaps recorded on 200 images during two tasks (general description and social captioning). The six models spanned a 2$\times$2 factorial of CNN vs.\ ViT encoders crossed with LSTM vs.\ Transformer decoders, plus Molmo 7B-D and Qwen3.5 9B. We found that both decoder and encoder architecture shaped alignment, but decoder choice dominated. LSTM vs.\ Transformer decoders increased alignment by 40--50 percentage points (80--87\% vs.\ 40--59\% of the human noise ceiling). In contrast, CNN vs.\ ViT encoders contributed a secondary 5--20 point advantage depending on decoder family, with CNN-LSTM the most aligned model overall (85--87\%). Despite their alignment advantage, LSTM-decoder attention maps were spatially diffuse and minimally task-differentiated; ViT-Transformer, the weakest in alignment, showed the sharpest spatial concentration and strongest task differentiation. A hemispatial-neglect simulation confirmed that ablating attention impacted LSTM decoders more than Transformer decoders. In an exploratory extension using TRIBE-simulated synthetic neural responses, fixation alignment and neural relevance dissociate: CNN-Transformer attention maps better predicted synthetic brain activity despite lower fixation alignment, with attention maps best predicting early visual cortex. Together, top-down and bottom-up components trade off what they predict in behavioral and synthetic neural data.


Attention-Based Sampler for Diffusion Language Models

Yuyan Zhou ⋅ Kai Syun Hou ⋅ Weiyu Chen ⋅ James Kwok

Auto-regressive models (ARMs) have established a dominant paradigm in language modeling. However, their strictly sequential sampling paradigm imposes fundamental constraints on both inference efficiency and modeling flexibility. To address these limitations, diffusion-based large language models (dLLMs) have been proposed, offering the potential for parallel sampling and flexible language modeling. Despite these advantages, current dLLMs sampling strategies rely primarily on token level information, which fails to account for global sequence structure and often yields suboptimal results. In this paper, we study the sampling order selection problem from the perspective of log-likelihood maximization. We show that this problem is NP-hard and propose an optimal sampling-rank-based approximation that makes the objective computationally tractable. We further prove that the tractable objective is optimized by sampling tokens in descending order of their attention-matrix column sums. This finding provides a principled justification for attention-guided sampling and offers a theoretically grounded alternative to greedy search. We instantiate this theoretical insight in a new training-free sampling algorithm, termed Attn-Sampler, and further propose dynamic attention thresholding for practical acceleration. Extensive experiments across multiple benchmarks validate the effectiveness of our proposed method, demonstrating that it achieves superior generation quality while enhancing the sampling parallelism.

Attention sinks and massive activations are recurring and closely related phenomena in Transformer models. Existing explanations have largely focused on the forward pass, yet in pre-norm Transformers, large residual-stream norms play only an indirect forward role because sublayers operate on normalized inputs. We study this relationship from the perspective of backpropagation. Empirically and theoretically, we show that under causal masking, attention sinks can induce pronounced gradient concentration, which we term gradient sinks. Since the RMSNorm Jacobian attenuates gradients roughly in inverse proportion to input norm, massive activations can be understood as adaptive regulators of this localized gradient pressure during training. This interpretation predicts that attenuating sink-induced gradients should weaken massive activations. We test this prediction with V-scale, a modification that adjusts backpropagated gradients on the value path. In V-scale models, attention sinks are preserved, whereas massive activations are suppressed. These results identify gradient sinks as a backward-pass counterpart of attention sinks, and massive activations as an adaptive RMSNorm-mediated response that attenuates the resulting localized training pressure. Our code is available at https://anonymous.4open.science/r/GradientSinkCode-B309.


Auditing Correlated Failures in Frozen-Feature Pretrained-Encoder Pools for Medical Segmentation

Eunseob Choi ⋅ Kyeonghun Kim ⋅ Hyuk-Jae Lee ⋅ Nam-Joon Kim

Cross-architecture/backbone medical-segmentation pools are commonly motivated as complementary failure coverage when members make different mistakes; that assumption is rarely tested at the per-case level. We audit this assumption for a fixed pool of 11 pretrained-encoder bundles read out by a deliberately weak $1\times1$-convolution decoder (4 seeds), and for a small set of decoder-recipe perturbations on the same bundles. Across five primary task rows (Kvasir polyp, ACDC LV, BraTS foreground, RIGA Cup, RIGA Disc), this 11-bundle $1\times1$-conv readout pool yields per-case Dice correlation gaps of $\Delta=0.261\text{--}0.638$ between same-bundle and cross-bundle pairs after a Dice $\geq 0.30$ functional floor; all paired case and hierarchical bootstrap intervals exclude zero. The result is stable under floor sweeps, ICC(A,1), leave-one-out, subject aggregation, quality-controlled pair regression, and cross-fitted item-difficulty residualisation. Decoder-recipe diagnostics show that richer readouts attenuate the gap (UNet-skip; case-identical nnU-Net on RIGA Cup: $0.473\to 0.0079$, separate-split $0.034\text{--}0.047$); the reported magnitudes are recipe-specific. Additional probes (pixel error, joint-failure lift, same-architecture-family ViT, same-family cross-checkpoint, same-backbone) locate the dependence structure but do not causally separate architecture, pretraining, and recipe. The claim is scoped to 2D frozen/light-adaptation encoder pools with this readout family and to retrospective benchmark audits. We release per-case Dice traces, summary JSONs, scorer code, hashes, and metadata as the audit artifact.


Auditing Privacy Leakage in Tabular Foundation Model Embeddings

Xun Wang ⋅ Adam Dziedzic ⋅ Michael Backes ⋅ Franziska Boenisch

Tabular Foundation Models (TabFMs) produce context-aware embeddings that are increasingly stored, shared, and reused for retrieval, clustering, and cross-institutional data exchange. These embeddings are often treated as privacy-preserving substitutes for raw tabular records, yet their attribute-level information content remains poorly understood. We present a systematic privacy audit of TabFM embeddings: given an embedding and varying amounts of side information, how accurately can sensitive attributes be recovered? We introduce Cascade Probing, a sequential probing method that measures recoverability while accounting for inter-attribute dependencies, significantly outperforming joint probing and optimization-based inversion baselines. Across four TabFMs (TabPFN, TabDPT, Mitra, TabICL) and eight datasets, we find that even without any side information, 65-95\% of sensitive attributes can be recovered from embeddings alone. More critically, recoverability exhibits a Privacy Cliff: revealing a single known attribute sharply increases the recoverability of all remaining attributes. This phenomenon generalizes across model architectures and diverse domains, and is not eliminated by standard embedding transformations including noise injection, PCA, and knowledge distillation. These findings suggest that we need more formal privacy techniques for protecting sensitive information inside TabFM embeddings.

While Large Language Models (LLMs) are increasingly deployed as automated evaluators in evaluation and training pipelines, their judgements are often affected by systematic biases that conflict with human preferences. While prior work has identified several known biases and proposed methods for their detection and mitigation, they lack strong grounding in human evaluation preferences, which is essential to ensuring that the identified biases correspond to actual human judgment behavior. Moreover, they rely heavily on pre-discovered bias lists, overlook bias strength, and depend on costly interventions. In this work, we propose HUB-J, an integrated framework grounded in human evaluation preferences that detects, quantifies, and mitigates biases in LLM-as-a-Judge systems. Our approach leverages human–LLM judgement disagreement cases to automatically discover interpretable bias factors, and utilizes agreement cases to quantify bias strength through controlled input modifications and resulting shifts in model decisions. Finally, building on these quantified biases, we introduce a lightweight, training-free regression-based mitigation strategy that corrects bias-influenced judgments by removing the estimated bias effects. Empirical results show that HUB-J uncovers both known and novel bias factors, reveals meaningful differences in model susceptibility, and consistently reduces bias-driven decision flips while generalizing across models.


Augmented Lagrangian Predictive Coding

Jeffrey Seely ⋅ Julian J Gould

Predictive coding (PC) is a local-learning alternative to backpropagation (BP), training deep networks via a local energy-minimization dynamics rather than a global backward pass. We introduce Augmented-Lagrangian Predictive Coding (PC-ALM), which maintains PC's inference budget but aligns each weight update toward BP by accumulating per-layer constraint errors into a layer-local Lagrange multiplier. In linear PC networks, PC-ALM converges to an equilibrium with exact BP gradients distributed across the network via only layer-local updates. We analyze PC-ALM in nonlinear PC networks up to depth 128 and show that it matches BP performance across all width-depth regimes, notably in deep narrow networks where PC underperforms. PC-ALM introduces recurrent dynamics in each layer's activations. Compared to PC's heat flow on a scalar energy, PC-ALM dynamics are driven by dual ascent on the Augmented Lagrangian. We observe ballistic credit propagation across especially deep networks, with prediction errors evenly distributed across layers, compared to PC's diffusive credit propagation. Beyond the algorithm itself, the augmented Lagrangian framework offers a generalization of PC, and may yield insights on how distributed systems could compute and propagate BP-like credit signals through purely local dynamics.


A Unified Framework for Adversary-Aware Differential Privacy Bounds

Marika Swanberg ⋅ Meenatchi Sundaram Muthu Selva Annamalai ⋅ Jamie Hayes ⋅ Borja Balle ⋅ Adam Smith

Differential Privacy (DP) bounds the privacy leakage of a mechanism against worst-case membership inference, but the precise tradeoff between complex adversarial models and DP protections remains poorly understood. In this paper, we present a unified framework that generalizes the patchwork of existing bounds across membership inference, attribute inference, and data reconstruction attacks. Crucially, our framework is the first to evaluate attacks that target multiple individuals simultaneously and measure success beyond exact matches under a single cohesive bound. Our bounds capture this broad family of previously unexplored attack settings by relying solely on the privacy parameters and the adversary's baseline success rate (i.e. its prior without access to the mechanism's output). To illustrate this, we compare our high-probability guarantees to empirical attacks in two novel settings: extracting multiple non-uniform secrets (passwords and PII) from DP-finetuned language models, and reconstructing tabular data from noisy marginals. Ultimately, this framework provides a rigorous theoretical foundation to investigate the risk landscape of DP algorithms in new adversarial settings.

A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, different methods and their resulting insights stand in relative isolation: what could the underlying structure of language models be, such that they give rise to all our interpretations? In this work, we propose using Tensor Product Representations (TPRs; Smolensky 1990) as a unifying hypothesis. TPRs give a concrete proposal for how compositional structure can be represented in vector space. We show, both mathematically and empirically, that TPRs can unify several prior interpretability methods: additive analogies, linear probing, sparse autoencoders, and activation patching. Mathematically, we show that these methods can all be derived from TPRs. Empirically, we apply the derivations to a range of different models — from small toy models to LLMs — to construct instances of each of the above interpretability methods; these constructed variants perform comparably to their standard variants. We view this work as a step toward what interpretability will ideally provide: a unified account of the nature of neural networks, corroborated not just by individual observations but also by an explanation of the connections between them.


Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents

Tianci Xue ⋅ Zeyi Liao ⋅ Tianneng Shi ⋅ Zilu Wang ⋅ Kai Zhang ⋅ Dawn Song ⋅ Yu Su ⋅ Huan Sun

Real-world digital environments are highly diverse and dynamic. These characteristics cause agents to frequently encounter unseen environments and distribution shifts, making continual learning in such environments essential for computer-use agents (CUAs). However, a key challenge lies in obtaining high-quality and environment-grounded training data without relying on costly human annotation. In this work, we introduce ACuRL, an Autonomous Curriculum Reinforcement Learning framework that continually adapts agents to specific environments with zero human data. The agent first explores an environment to acquire initial experiences. During subsequent iterative training, a curriculum task generator leverages these experiences together with feedback from the previous iteration to synthesize new tasks tailored for the agent's current capabilities. To provide reliable reward signals, we introduce CUAJudge, a robust automatic evaluator for CUAs that achieves 93% agreement with human judgments. Empirically, our method effectively enables both intra-environment and cross-environment continual learning, yielding 3–29% absolute performance gains on the target environments without catastrophic forgetting on others. We also show that it can mitigate performance degradation under environment changes (e.g., version updates, platform migration, and resolution shifts). Further analyses show highly sparse updates (e.g., only 20% parameters), which helps explain the effective and robust adaptation.


AVID: A 5T fMRI Dataset for Benchmarking Auditory-induced Visual Mental Imagery Decoding

Shiqi Shen ⋅ Shurui Li ⋅ Yuanning Li ⋅ Xilin Zhang

Visual mental imagery (VMI) decoding uses brain activity to reconstruct internally generated visual scenes, providing a potential pathway for externalizing internal experiences and developing neuroprosthetic communication interfaces. However, progress in this field is limited by the lack of dedicated, training-scale VMI datasets. Current VMI paradigms often rely on visual prompts or short-term memory images, introducing visual and mnemonic confounds that make it difficult to disentangle neural activity associated with endogenous image construction from signals driven by prior or concurrent visual input. Furthermore, existing VMI datasets offer limited descriptive control over individual visual attributes, making it difficult to systematically assess fine grained visual details. To address these gaps, we introduce AVID, the first dedicated training-scale 5.0 Tesla functional magnetic resonance imaging (fMRI) resource for auditory-induced VMI decoding. We further define a benchmark around AVID with fixed splits, predefined neural inputs, and a standardized evaluation protocol. Acquired with high-field 5T fMRI, AVID comprises approximately 67 hours of dense recordings from nine participants. The dataset is uniquely structured into two complementary splits: a scene-level complex split using naturalistic Mandarin descriptions and a controlled simple split with explicitly specified visual attributes. We further formalize a standardized evaluation protocol that decouples source-image consistency from semantic alignment, allowing for a more nuanced assessment of imagery reconstruction. Baseline qualitative and quantitative results show that current reconstruction pipelines remain limited on AVID, although the most successful reconstructions preserve recognizable scene level semantics from spoken descriptions. By combining dedicated training-scale neural recordings and standardized evaluation resources, AVID provides both a data foundation for imagery-specific model development and a benchmark for evaluating VMI decoding with substantially reduced sensory contamination.


AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression

Yijing Chen ⋅ Wenhui Tan ⋅ Xiaoyi Yu ⋅ Yuyue Wang ⋅ Xin Cheng ⋅ Kaisi Guan ⋅ Hao Jiang ⋅ Xiangyang Li ⋅ Guojie Zhu ⋅ Ruihua Song

Multimodal Large Language Models have achieved remarkable progress in short-form audio-video understanding, yet long-form audio-video comprehension remains challenged by limited context windows and severe information redundancy. To address these bottlenecks, we propose AVOC, a framework for long-form audio-video understanding in Omni-modal Large Language Models. AVOC introduces a learnable token compression module between the modality encoders and the LLM backbone. We reframe multimodal token compression as a top-$K$ retrieval problem: given a fixed context budget, the module must retrieve a compact subset of tokens that best supports answering the user query. We draw inspiration from three classical Information Retrieval criteria for selecting informative units from a large candidate pool: relevance, importance, and diversity. AVOC instantiates each criterion as a tailored mechanism for audio-video understanding, and integrates them into a unified retrieval-style compression pipeline. Experiments show that AVOC achieves state-of-the-art performance on long-form audio-video benchmarks, surpassing the second-best model by 4.9 and 5.5 points in average accuracy on OmniVideoBench and LVOmniBench, respectively. Moreover, AVOC maintains robust performance on Audio-Video Needle-in-a-Haystack task at durations up to one hour.


Avoiding Obfuscation with Prover-Estimator Debate

Jonah Brown-Cohen ⋅ Geoffrey Irving ⋅ Georgios Piliouras ⋅ Lijie Chen ⋅ Jiawei Li ⋅ Zhiyang Xun

Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks. A promising approach to this problem is to amplify human judgement by leveraging the power of two competing AIs in a debate about the correct solution to a given problem. Prior theoretical work has provided a complexity-theoretic formalization of AI debate, and posed the problem of designing protocols for AI debate that guarantee the correctness of human judgements for as complex a class of problems as possible. Recursive debates, in which debaters decompose a complex problem into simpler subproblems, hold promise for growing the class of problems that can be accurately judged in a debate. However, existing protocols for recursive debate run into the obfuscated arguments problem: a dishonest debater can use a computationally efficient strategy that forces an honest opponent to solve a computationally intractable problem to win. We mitigate this problem with a new recursive debate protocol that, under certain stability assumptions, ensures that an honest debater can win with a strategy requiring computational efficiency comparable to their opponent.


Balanced Multi-Task Learning from an Optimality-Gap Perspective

Ce Liu ⋅ Jianing Huang ⋅ HUANG Sicheng ⋅ Shu Liu ⋅ Xinyu Huang ⋅ Hao Yang

Multi-task learning requires a shared update that makes balanced progress across multiple objectives. A key obstacle is that progress is difficult to compare across tasks with different loss scales, units, and local optimization dynamics. We study this problem from an optimality-gap perspective. For each task, we compare the improvement induced by the shared update with the best one-step improvement achievable by optimizing that task in isolation. This comparison yields a scale-invariant normalized improvement rate, which measures how well the shared update serves each task relative to its own local improvement bound. We formulate balanced multi-task optimization as a max-min problem over the normalized improvement rates, thereby prioritizing the worst-served task. A local quadratic approximation leads to the Karush-Kuhn-Tucker (KKT) optimality conditions, showing that the optimal shared update is determined by a small set of bottleneck tasks. Under an isotropic Hessian approximation, the update admits a closed-form expression that depends only on task gradients and a regularized Gram matrix, while the bottleneck set is identified by an active-set procedure. Experiments on diverse benchmarks show improved task balance and competitive performance.


BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows

Elaine Lau ⋅ Markus Dücker ⋅ Ronak Chaudhary ⋅ Hui Wen Goh ⋅ Rosemary Wei ⋅ Vaibhav Kumar ⋅ Saed Qunbar ⋅ Guram Gogia ⋅ Yi Liu ⋅ Scott Millslagle ⋅ Nasim Borazjanizadeh ⋅ Ulyana Tkachenko ⋅ Samuel E Danquah ⋅ Collin Schweiker ⋅ Vijay Karumathil ⋅ Asrith Devalaraju ⋅ Andrew P Martin ⋅ Varsha Sandadi ⋅ Haemi Nam ⋅ Punit Arani ⋅ Ray Epps ⋅ Abdullah Arif ⋅ Sahil Bhaiwala ⋅ Curtis Northcutt ⋅ Skyler Wang ⋅ Anish Athalye ⋅ Jonas Mueller ⋅ Francisco Guzmán

Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profession, we introduce BankerToolBench (BTB): an open-source benchmark of end-to-end analytical workflows routinely performed by junior investment bankers. To develop an ecologically valid benchmark grounded in representative work environments, we collaborated with 502 investment bankers from leading firms. BTB requires agents to execute senior banker requests by navigating data rooms, using industry tools (market data platform, SEC filings database), and generating multi-file deliverables—including Excel financial models, PowerPoint pitch decks, and PDF/Word reports. Completing a BTB task takes bankers up to 21 hours, underscoring the economic stakes of successfully delegating this work to AI. BTB enables automated evaluation of any LLM or agent, scoring deliverables against 100+ rubric criteria defined by veteran investment bankers to capture stakeholder utility. Testing 9 frontier models, we find that even the best-performing model (GPT-5.4) fails nearly half of the rubric criteria and bankers rate 0% of its outputs as client-ready. Our failure analysis reveals key obstacles (such as breakdowns in cross-artifact consistency) and improvement directions for agentic AI in high-stakes professional workflows.

We study Bayesian optimization (BO) through the lens of information geometry. Pulling back the Fisher information metric through the surrogate posterior map yields a local sensitivity tensor on the input space, which leads to an upper bound on the gradient of reparameterizable acquisition functions. This view explains vanishing-gradient behavior in high-dimensional BO and provides a common interpretation of heuristics such as RAASP and dimension-scaled lengthscales. Building on this analysis, we propose FITR, a trust-region-based BO method that replaces lengthscale-based scaling by local pullback-Fisher weights. FITR is not restricted to GP kernels with explicit lengthscales. On GP benchmarks with an SE kernel, experiments show competitive performance using FITR. The proposed method also easily generalizes to non-isotropic surrogates, although the gains are more task-dependent in that setting.


BEAST3D: animal behavioral analysis and neural encoding from multi-view video via Gaussian splatting

Yanchen Wang ⋅ Lenny Aharon ⋅ Wangshu Zhu ⋅ Kyle Daruwalla ⋅ Linghua Zhang ⋅ Jiaru Zou ⋅ Selmaan N Chettih ⋅ X Helen Hou ⋅ Liam Paninski ⋅ Matthew Whiteway

Multi-view video recordings are increasingly used to capture the 3D movements of animals in experimental settings, yet extracting rich 3D representations from these recordings remains challenging. Supervised pose estimation requires extensive manual annotation, while general-purpose 3D reconstruction models trained on generic scene datasets fail on the specialized imagery and sparse-view setting of laboratory experiments. We address these limitations with BEAST3D, a self-supervised pretraining framework that learns 3D visual representations from unlabeled, calibrated multi-view video. BEAST3D uses a vision transformer to predict 3D Gaussian splats that reconstruct held-out views through differentiable rendering, while simultaneously segmenting the animal from the background. BEAST3D reconstructs 3D structure with as few as four views by conditioning directly on known camera parameters---unlike general-purpose models, which must estimate camera geometry from dense overlapping viewpoints that are seldom available in lab settings. Through comprehensive evaluation across four species, we demonstrate that BEAST3D produces rich, viewpoint-invariant features that transfer effectively to three downstream tasks: novel view synthesis, which validates the quality of the learned 3D representations; multi-view pose estimation, which provides the sparse keypoint trajectories widely used in behavioral analysis; and neural encoding, which relates 3D behavioral features to simultaneously recorded neural activity. BEAST3D thus establishes a versatile framework for behavioral analysis that leverages 3D structure in modern multi-view laboratory recordings.


Before Words, Beyond Speech: Evaluating Nonverbal Social Reasoning in Early Childhood

Marie Amale Huynh ⋅ Laura Bravo-Sánchez ⋅ Lauren K Dubin ⋅ Nick Haber ⋅ Philip A Fisher ⋅ Serena Yeung-Levy

Children's earliest years are a period of intense social learning, during which meaning is built through gaze, gesture, and physical contact. Yet, this critical aspect of social interaction remains under-evaluated in multimodal video understanding, as existing benchmarks prioritize the conversation-driven interactions of adults. We address this gap with NEST ($\textbf{N}$aturalistic $\textbf{E}$arly-childhood $\textbf{S}$ocial in$\textbf{T}$eractions), the first child-centric benchmark for nonverbal social reasoning. NEST comprises 1,208 manually annotated clips ($\sim$12s) curated for diversity across naturalistic settings, cultural contexts, and developmental stages. Beyond a lack of data, targeted model improvement in this domain is hindered by coarse evaluation schemas that conflate basic perception with higher-order reasoning. To resolve this, NEST introduces a compositional schema that builds atomic interaction cues into progressive levels of social reasoning: contextual, behavioral, interpersonal, and open-ended descriptions. Extensive analysis of state-of-the-art Vision-Language Models (VLMs) reveals a stark performance collapse: despite strong contextual recognition, models exhibit substantial gaps in decoding precise behavioral and interpersonal cues. We trace these failures to specific biases, including poor consequence grounding and temporal sensitivity. Finally, we ground NEST in real-world applications by evaluating VLMs using domain adaptation strategies. Ultimately, NEST provides a rigorous testbed for advancing socially grounded artificial intelligence in early childhood. The dataset and code will be publicly released.

Markov decision problems are most commonly solved via dynamic programming. Another approach is Bellman residual minimization, which directly minimizes the squared Bellman residual objective function. However, compared to dynamic programming, this approach has received relatively less attention, mainly because it is often less efficient in practice and can be more difficult to extend to model-free settings such as reinforcement learning. Nonetheless, Bellman residual minimization has several advantages that make it worth investigating, such as more stable convergence with function approximation for value functions. While Bellman residual methods for policy evaluation have been widely studied, methods for policy optimization (control tasks) have been scarcely explored. In this paper, we establish foundational results for the control Bellman residual minimization for policy optimization.

LLM judges are increasingly placed inside an agent's loop, scoring the agent's own attempts and re-prompting until one passes. We show this quietly corrupts measurement: retry-until-PASS is optional stopping against a noisy classifier—it keeps drawing until the judge slips—so the reported pass rate is a biased estimator of true success, upward in the pass-prone retry regimes of interest (and downward under conservative rules such as strict rubrics or unanimous juries). We make this exact. The cap-K gate is a binary classifier with closed-form sensitivity/specificity, and its bias is governed by one coefficient, the gate's Youden index J: as J → 0 the gated rate becomes uninformative about π, so recovery must fall back to gold labels and no estimator beats the gold-only mean. Across 44 capable-agent loops on GSM8K, MATH, and code with objective ground truth (no authored weakness; a separate terse-agent stress set is excluded here) the inflation is systematic (median slip +0.16; worst on code, where the judge cannot run the candidate, a true 0.74 inflated to a reported 0.98) and obeys a closed-form law predicting the slip from per-attempt statistics (pooled r = 0.95; errors-in-variables slope 0.765 [0.70, 0.84], excluding 1). To recover true success we benchmark Rogan–Gladen against prediction-powered inference: the label-efficient estimators (PPI/PPI++) recover π far more accurately than naive reporting and Rogan–Gladen (mean recovery MAE 0.050 vs. 0.149, and 0.081 vs. 0.241 as gold becomes scarce), because they escape the 1/J² variance that makes the classical correction fragile; on balanced low-bias gates they match the gold-only mean (per-gate differences within bootstrap noise, none uniformly best), so our contribution is not a uniquely-best estimator but the identification of loop-induced bias and the J diagnostic that says when—and whether—to correct at all. Beyond verifiable gold, we measure recovery on public, human-labeled non-verifiable gates—response safety and summary quality—where PPI++ recovers the true rate to mean MAE 0.043 versus naive 0.165 (~4× in aggregate, up to >10× on the most-biased gates); and—in the motivating regime, a non-verifiable safety judge inside a real retry loop (n = 400, against a pre-registered 3-model strong-LLM panel—a disclosed proxy, human-anchored at raw 0.90 agreement)—a lenient gate ships 6.8% panel-labeled-unsafe responses (95% CI [0.047, 0.096]) while the calibrated correction recovers the panel safe-rate ~3.5× more accurately than naive. The deliverable is a recipe: report a PPI++ estimate alongside J as a reliability/identifiability diagnostic, measure at K = 1, and—via a label-free drift detector (ROC-AUC 0.80)—de-bias only when calibration transfers. We release all code and content-free data.


Benchmark for Assessing Olfactory Perception of Large Language Models

Eftychia Makri ⋅ Nikolaos Nakis ⋅ Laura Sisson ⋅ Geetanjali Minsky ⋅ Leandros Tassiulas ⋅ Vahid Satarifard ⋅ Nicholas A Christakis

Here we introduce the Olfactory Perception (OP) benchmark, designed to assess the capability of large language models (LLMs) to reason about smell. The benchmark contains 1,010 questions across eight task categories spanning odor classification, odor primary descriptor identification, intensity and pleasantness judgments, multi-descriptor prediction, mixture similarity, olfactory receptor activation, and smell identification from real-world odor sources. Each question is presented in two prompt formats, compound names and isomeric SMILES, to evaluate the effect of molecular representations. Evaluating 21 model configurations across major model families, we find that compound-name prompts consistently outperform isomeric SMILES, with gains ranging from +3.1 to +18.9 percentage points (mean $\approx$ +7 points), suggesting current LLMs access olfactory knowledge primarily through lexical associations rather than structural molecular reasoning. The best-performing model reaches 64.4\% overall accuracy, which highlights both emerging capabilities and substantial remaining gaps in olfactory reasoning. We further evaluate a subset of the OP across 21 languages and find that aggregating predictions across languages improves olfactory prediction, with AUROC=0.86 for the best performing language ensemble model. LLMs should be able to handle olfactory and not just visual or auditory information.


Benchmarking Optimizers for Large Language Model Pretraining

Andrei Semenov ⋅ Matteo Pagliardini ⋅ Martin Jaggi

The recent development of Large Language Models (LLMs) has been accompanied by an effervescence of novel ideas and methods to better optimize the loss of deep learning models. Claims from those methods are myriad: from faster convergence to removing reliance on certain hyperparameters. However, the diverse experimental protocols used to validate these claims make direct comparisons between methods challenging. This study presents a comprehensive evaluation of recent optimization techniques across standardized LLM pretraining scenarios, systematically varying model size, batch size, and training duration. Through careful tuning of each method, we provide guidance to practitioners on which optimizer is best suited for each scenario. For researchers, our work highlights promising directions for future optimization research. Finally, by releasing our code and making all experiments fully reproducible, we hope our efforts can help the development and rigorous benchmarking of future methods.

Biosynthetic gene clusters (BGCs) are co-located genes that encode the biosynthesis of secondary metabolites, a major source of clinically used antibiotics and anticancer agents. Recent advances in deep learning, particularly self-supervised foundation models, have spurred growing interest in BGC sequence modeling, but evaluation infrastructure has lagged behind. Current practice often relies on coarse classification accuracy over small, experimentally curated datasets, making it difficult to discriminate model capabilities or assess downstream utility. We introduce BGC-Bench, an evaluation suite designed to resolve this limitation by systematically incorporating biochemical knowledge in the benchmark design. In terms of generalization, we test whether representations transfer to genetically and chemically novel samples; for functionality, we examine application-relevant tasks, including retrieval, prediction, and active learning. BGC-Bench reveals that no single model class dominates: pretrained foundation models perform strongly under out-of-domain generalization but remain sensitive to training conditions; compositional baselines are competitive and lead in cross-modal retrieval; and active-learning gains vary substantially across tasks, strategies, and model families. Beyond these BGC-specific findings, BGC-Bench highlights a broader principle: domain-informed benchmark design over existing datasets enables better evaluation in data-scarce scientific domains.


Best-of-$N$ Guidance for Test-time Diffusion Alignment

Richard Lee Kim ⋅ Yeongmin Kim ⋅ Gyuwon Sim ⋅ Taekyu Kim ⋅ Minsang Park ⋅ Il-chul Moon

Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, this procedure is inherently post hoc: reward information is used only for selection after generation, not for improving the reverse diffusion process itself. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-$N$ Guidance (BoNG), a novel method that integrates the principle of BoN sampling directly into the reverse diffusion process. BoNG performs online BoN selection over denoising particles and adjusts the reverse diffusion process to steer the particle population toward higher-reward regions during generation. Specifically, by introducing an asymmetric guidance interaction among denoising particles, BoNG uses the current BoN particle as a guidance signal to the rest of the particle population. This particle-level interaction reshapes the sampling process toward higher-reward regions, enabling BoNG to improve not only the final best sample beyond Vanilla BoN sampling, but also the average quality of generated samples. Over 36 empirical comparisons, BoNG achieves the best performance in 29 cases, ranking first in 80.56\% of the comparisons against SMC and Vanilla BoN sampling. BoNG also supports multi-output capability, showing 1.3$\times$ ImageReward score than the latest sample-based guidance method with 1.6$\times$ speedup.


Beyond Coordinates: Encoding Graph Structure via Contextual Distribution and Relational Similarity

Feifei Qian ⋅ Lu Bai ⋅ Lixin Cui ⋅ Ming Li ⋅ Hangyuan Du ⋅ Bo Jiang ⋅ Lixiang Xu ⋅ Edwin Hancock

The Message Passing Neural Networks (MPNNs) and Graph Transformers (GTs) have emerged as two dominant paradigms for graph representation learning. However, the expressive power of standard MPNNs is fundamentally bounded by the 1-dimensional Weisfeiler-Lehman (1-WL) test, while GTs lack the inductive bias for graph structure. To enhance structural representation, existing methods typically resort to subgraph-based aggregation or coordinate-based positional encodings. However, the former suffers from prohibitive computational memory overheads, while the latter is limited by rigid reference frames that fail to distinguish fine-grained local topology. To address these limitations, we introduce a novel structural encoding framework based on contextual distributions. Specifically, we move beyond fixed coordinate systems to capture intrinsic local topology by encoding structural information through a compact distributional statistic, i.e., the entropy of the node context. Furthermore, instead of relying on relative positioning, we introduce a kernelized mechanism to encode relational similarity by quantifying the structural affinity between node contexts. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed framework achieves superior performance, striking a favorable balance between effectiveness and efficiency.


Beyond Correctness: Robustness-Driven Evolutionary Self-Training for Large Language Models

Wei Guo ⋅ Hongyao Tang ⋅ Yi Ma ⋅ Jinyi Liu ⋅ Pengyi Li ⋅ Jing Liang ⋅ Yifu Yuan ⋅ YAN ZHENG ⋅ Jianye Hao

Outcome-based post-training methods for mathematical reasoning rely on binary feedback, treating all correct trajectories equally. However, this masks a critical distinction: some correct paths are structurally brittle and prone to cascading errors under inevitable autoregressive sampling noise, while others are robust and maintain their correctness despite these natural decoding variations. Optimizing indiscriminately over these brittle solutions leads to inefficient learning, as the model wastes capacity memorizing brittle reasoning strategies. To address this, we propose Trajectory Robustness-Driven Evolutionary Self-Training (TREST). Instead of relying on binary feedback, our framework explicitly prioritizes trajectory robustness by integrating evolutionary algorithms with supervised fine-tuning. Crucially, our analysis reveals that trajectory robustness serves as a strong natural proxy for genuine mathematical insight. By conceptualizing reasoning paths as evolving individuals within a population, this evolutionary process naturally marginalizes brittle, brute-force calculations in favor of robust, insight-driven strategies. By effectively internalizing these strategies, TREST achieves superior reasoning performance. Experiments on complex mathematical benchmarks demonstrate that under aligned budgets, TREST outperforms outcome-based baselines, establishing a highly effective alternative to conventional post-training paradigms.


Beyond Data Scaling: Representation-Centric Pre-training for Vision-Language-Action Models

Senqiao Yang ⋅ Chengyao Wang ⋅ Yuxin Chen ⋅ Zixuan WANG ⋅ Longxiang Tang ⋅ Haokun GUI ⋅ Jinhui Ye ⋅ Changsheng Lu ⋅ Xiaoyang Wu ⋅ Mingkang Zhu ⋅ Pengguang Chen ⋅ Shu Liu ⋅ Zhuotao Tian ⋅ Hengshuang Zhao ⋅ Bei Yu ⋅ Jiaya Jia

Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are fundamentally harder to scale than web-scale image-text data because they require embodied collection and sparsely cover the physical world. This makes representation quality a central bottleneck: under a fixed robot-data budget, VLA pre-training must convert limited trajectories into transferable visual-action knowledge, rather than merely fit actions. We propose VLAct, a VLA-oriented VLM backbone trained with a representation-centric pre-training recipe. VLAct preserves the broad VLM prior, avoids over-specializing the backbone to a single action head, and encourages shared action semantics across embodiments through VLM-prior preservation, multi-head continuous action co-supervision, and a partially unified cross-embodiment action layout, while leaving downstream users free to attach task-specific action heads during fine-tuning. Across multi-embodiment simulation benchmarks, real-world robot experiments, and unseen-embodiment transfer, VLAct consistently improves downstream performance under fixed fine-tuning protocols. On LIBERO-Plus and RoboTwin 2.0, VLAct surpasses large-scale industrial VLA systems such as ABot-M0 and LingBot-VLA, achieving 82.6\% and 92.5\% success, respectively. Most notably, on RoboCasa-GR1, a humanoid embodiment never seen during pre-training, VLAct with only 20\% of downstream trajectories already outperforms the full-data GR00T-N1.6 baseline. These results are obtained using fully open-source data and only a 16-GPU training setup, showing that representation-centric pre-training is an important independent axis of VLA progress beyond data scaling. All models and training pipelines will be open-sourced.


Beyond difficulties: Insights for provably efficient design of autocurriculum in RLVR

Ruofan Wu ⋅ Xin Ge ⋅ Gao Xuemin ⋅ Guanhua Fang ⋅ Yangbin Shi ⋅ Dongbai Guo

The design of an efficient curriculum has become increasingly important for reinforcement learning with verifiable rewards (RLVR). The majority of existing curriculum strategies select training problems based on **difficulty**---typically measured by success rate---yet difficulty is an indirect proxy that conflates problem hardness with actual training utility. In this paper, we propose to drive curriculum design not by how hard a problem is, but by how much the model improves from training on it. We formalize this principle through the **improvement function** (IF), which measures the gain in a prompt's expected reward under an infinitesimal GRPO policy gradient step. Leveraging the IF framework, we formally characterize the recently discovered **edge of competence** (EoC) phenomenon: a problem is at the model's EoC when its improvement is within a constant factor of the maximum over the training set. Analyzing GRPO on $L$-step compositional reasoning tasks, we prove that an EoC-induced curriculum achieves target mastery in $\widetilde{\Theta}(\log L_{\max} / (\eta \log d))$ steps---an exponential improvement over the $\widetilde{\Theta}(L_{\max} / (\eta \log d))$ steps required by a uniform mixture of difficulties. Motivated by this theory, we propose ReCUR, a practical algorithm that maintains a stratified retry buffer of previously failed problems, resampling them until they yield positive reward or exhaust a retry budget. Experiments on multimodal reasoning benchmarks show that ReCUR consistently improves the performance of GRPO and DAPO, providing empirical evidence consistent with the improvement-driven curriculum perspective.


Beyond Downstream Scores: Controlled Diagnostics for Point-Cloud Self-Supervised Learning Evaluation

Kohsuke Ide ⋅ Ryousuke Yamada ⋅ Yue Qiu ⋅ Yoshihiro Fukuhara ⋅ Hirokatsu Kataoka ⋅ Yuki Asano ⋅ Yutaka Satoh

Downstream scores are the standard evidence for progress in point-cloud self-supervised learning (SSL). But a score is not a mechanism: it does not explain why a representation improves. We diagnose three simple questions in the evaluation pipeline: what signal is learned, what geometry is visible, and what the score measures. These questions test whether a gain comes from the pretext objective, from the geometric support available in the input, or from the downstream readout itself. Applied to cross-modal alignment, autoregressive prediction, and masked autoencoding, our diagnostics reveal that downstream scores often fail to identify the source of a gain. For Concerto, coordinate-only signals explain 28.9\% of the pretext-loss response but recover only 6.9\% of the downstream gain on ScanNet, separating pretext-loss explanation from representation transfer. For masked autoencoding, removing spatially coherent regions hurts much more than removing the same number of random points, showing that the visible geometry, not just the retained fraction, drives the score. For PointGPT, relaxing mask and order assumptions reduces transfer by at most 2.78 points on ScanObjectNN and 0.22 points on ShapeNetPart, showing that high scores can persist after weakening the intended autoregressive mechanism. These findings caution against treating downstream gains as self-evident representation progress in point-cloud SSL. We propose a diagnostic protocol that asks not only whether a score improves, but why it improves.


Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

Matthias Blaschke ⋅ Daniel Kienzle ⋅ Zsuzsanna Koczor-Benda ⋅ Julian Lorenz ⋅ Rainer Lienhart ⋅ Fabian Pauly

Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This combination drives strong benchmark metrics but limits transferability to domains structurally distinct from drug discovery. To overcome this limitation and drive discovery toward real, scientifically grounded targets, we introduce the Nanotechnology Molecular Optimization (NMO) Benchmark, which bridges machine learning (ML) and quantum materials science. NMO acts simultaneously as a rigorous testbed for the ML community and a discovery engine for nanotechnology research. The suite replaces proxy oracles with quantum simulations and introduces strict protocols that prioritize scientific utility over leaderboard-oriented overfitting. The physics-based NMO tasks impose hard structural constraints and rugged fitness landscapes, posing fundamentally new requirements on generative models. Notably, advanced molecular optimization methods underperform much simpler approaches on the NMO tasks. We develop a new baseline method identifying the critical components to solve the NMO tasks, including a novel representation for modeling structural constraints and a domain-agnostic pretraining strategy to eliminate pharmaceutical dataset bias. Our results surpass state-of-the-art physical properties and reveal previously unknown structural motifs, offering new insights for the nanotechnology community and demonstrating that ML can drive genuine scientific discovery.

As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a prerequisite for principled design. Yet large vision-language models (LVLMs) currently lack the tools to do so, and parameter-efficient encoder configurations remain hard to identify before training. To re-examine encoder roles under joint training, on the 16-benchmark Cambrian-1 suite we retrain and evaluate all 31 non-empty subsets of five common vision encoders under a unified pipeline ($\approx$20k GPU-hours total), and report three findings. First, retraining each subset from scratch reveals encoder rankings that differ from those obtained by masking encoders on a fixed checkpoint, including which encoder ranks first overall. Second, we decompose each encoder's contribution into two axes, \emph{Capacity}, the score an encoder reaches on its own, and \emph{Necessity}, the drop when it is removed from the full pool. The two axes are not interchangeable. Pairing the two highest-Capacity encoders is suboptimal, while pairing a high-Capacity anchor with an adaptive complement matches the full five-encoder model. Adding further encoders beyond this pair yields only marginal gains. Third, at fixed parameter count, per-encoder pre-projector effective rank explains the residual score variation. The strongest pairs combine an anchor whose rank survives joint training with a complement whose rank \emph{expands} under it, suggesting that higher-rank, less-collapsed projector inputs correspond to a more favorable optimization regime at the encoder–projector interface. Together, the Capacity–Necessity decomposition and the pre-projector rank analysis, along with comprehensive evaluation through retraining, expose a methodological gap in multi-encoder LVLM design, and offer concrete primitives for closing it.


Beyond Feature Disruption: Boundary-Diverting Unlearnable Examples against Linear Probing

Zhihao Li ⋅ Jiale Cai ⋅ Gezheng Xu ⋅ Ruiyi Fang ⋅ Hao Zheng ⋅ RUIZHI PU ⋅ Zhong Ji ⋅ Charles Ling ⋅ Boyu Wang

Unlearnable Examples (UEs) have emerged as a promising data protection strategy against unauthorized model training, which adds imperceptible perturbations into data to degrade generalization. Existing studies predominantly assume a fully trainable target model, where unlearnability is achieved by disrupting clean feature representations. However, this assumption grows increasingly unrealistic with the prevalence of pretrained models, where unauthorized users can freeze the backbone and perform linear probing, preserving discriminative representations and thereby invalidating the core mechanism of previous UEs. In this paper, we investigate this practical yet challenging linear probing setting and reveal a fundamental vulnerability. To address this, we propose $\textbf{DIVERT}$ ($\textbf{D}$ecision-boundary d$\textbf{I}$version $\textbf{V}$ia s$\textbf{E}$mantic-decoupled o$\textbf{R}$thogonal $\textbf{T}$argets), a novel framework that shifts the design principle from feature disruption to explicit decision boundary diversion. Our intuition is that projecting perturbed samples into a subspace orthogonal to the semantic manifold induces spurious linear separability, thereby steering the decision boundary away from original prototypes. Building on this principle, DIVERT first constructs synthetic class anchors within the semantic-decoupled subspace, then employs a cross-attention mechanism to jointly optimize these anchors while pushing perturbed samples toward them to establish spurious separability. Subsequently, it incorporates a surrogate linear head to simulate the linear probing procedure, further refining perturbations to actively divert the decision boundary. Extensive experiments demonstrate that DIVERT establishes effective unlearnability across diverse datasets and backbones.


Beyond Global Alignment: Structured Compositional Reasoning for Vision-Language Models

Zhoujun Ye ⋅ Yiwei Fu ⋅ Qiyun Huang ⋅ Jie Yang ⋅ Dongjie Wang ⋅ Xiao Luo

Despite remarkable progress in image-text understanding, vision-language models (VLMs) still struggle with compositional reasoning. In particular, they often fail to distinguish relational direction and attribute-object binding, leading to similar representations for semantically different image-text pairs. This problem mainly stems from the reliance on global image-text alignment, which captures coarse correspondence but overlooks fine-grained compositional structures. Toward this end, we propose an evidence-aware framework, termed Directional Relation Bucketing with Binding Localization (DELTA), to improve compositional understanding in VLMs. The key idea of our DELTA is to improve compositionality of VLMs from two complementary perspectives: direction modeling and binding localization. More specifically, we first leverage a learnable gate to model the cumulative contextual distance between anchor terms, thereby assigning relation concepts to direction-aware discrete buckets. To enrich textual representations with fine-grained relational semantics, we calibrate global semantic representations by incorporating intermediate-layer hidden states. Furthermore, DELTA leverages textual cues to ground visual evidence for object attributes, pulling image patches closer to their matched attribute descriptions while pushing them away from incorrect augmented ones. Extensive experiments on four compositional datasets demonstrate the effectiveness of the proposed DELTA. Our implementation is available at https://anonymous.4open.science/r/DELTA-2D60.


Beyond High and Low: Evaluating Graded Cognitive Diversity in LLM Persona Simulations

Sion Weatherhead ⋅ Ben R Newell ⋅ Aaron Belbasis ⋅ Flora Salim

Large language model agents are increasingly used as synthetic participants in social, behavioural, and decision-making research, yet it remains unclear whether they can express graded cognitive diversity rather than collapsing toward generic competence or stylised persona descriptions. We introduce a cognition-suite evaluation for LLM persona simulation, testing whether models reproduce population-level variation across structured reasoning tasks, self-report cognitive/personality measures, cross-construct associations, and downstream behavioural reasoning. Agents are conditioned on synthetic or human-derived trait profiles spanning cognitive reflection, Need for Cognition, and Big Five personality, then evaluated on CRT-style reasoning items, trait questionnaires, and Behavioural Reflection Task scenarios requiring evidence weighting, risk evaluation, and socially embedded justification. Across four matched simulation runs, agents partially recover questionnaire-style Need for Cognition and Big Five structure, especially under numeric prompting, but fail to preserve cognitive-reflection fidelity: CRT-style scores collapse toward correctness, intuitive-lure errors are almost eliminated, and reflection-related association structure weakens. Downstream BRT responses show the same pattern of flattening: open-ended decisions compress into over-regularised behavioural profiles rather than preserving human-like variation. The evaluation reframes persona simulation as a problem of cognitive-diversity fidelity: useful synthetic populations must preserve not only demographic or personality descriptors, but the structured variation in reasoning style, error type, and behaviour that those descriptors are meant to organise.


Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors

Guozheng Ma ⋅ Lu Li ⋅ Zilin Wang ⋅ Pierre-Luc Bacon ⋅ Dacheng Tao

Online reinforcement learning (RL) agents increasingly depend on knowledge acquired offline to achieve practical efficiency. Originally studied in offline-to-online RL, this paradigm now spans foundation model post-training and embodied intelligence, with prior types expanding from offline datasets and pre-trained policies to increasingly diverse knowledge sources such as multimodal foundation models and generative world models. Offline priors have become central to how deep RL is developed and deployed. However, this reliance introduces a challenge that the prevailing benchmark-driven paradigm cannot resolve: because prior validity varies across deployments and shifts during training, no single approach to managing it is universally optimal, and benchmark rankings offer limited guidance for real-world deployments. Rather than pursuing universal solutions, we argue that the field should shift to diagnosis-driven tension management, in which deployment-specific evidence guides how the learner relates to its priors throughout training, enabling both flexible and adaptive deployment. We support this position with a framework characterizing how priors reshape online optimization through three functional roles, controlled experiments demonstrating help-or-hurt reversals, cross-domain evidence from foundation model post-training to embodied intelligence, and engagement with five substantive counterarguments.


Beyond Outcome Rewards: Process-Aware Optimization for Search Agents

Yiwei Dai ⋅ Hengyi Cai ⋅ Hui Wu ⋅ Yili Wang ⋅ Han Xu ⋅ Minlan Shao ⋅ Yuchen Li ⋅ Shuaiqiang Wang ⋅ Xin Wang ⋅ Yi Chang ⋅ Dawei Yin

Search agents based on large language models address complex information-seeking tasks by iteratively issuing queries, retrieving evidence, and integrating information. Effective training of such agents requires supervision beyond final-answer correctness, as outcome-level rewards cannot distinguish efficient evidence acquisition from redundant or misdirected search. Recent process-level methods introduce intermediate supervision, but typically apply uniform signals across trajectories. However, such uniformity overlooks the state-dependent nature of process quality, where the value of an intermediate decision depends on the current evidence state. This motivates a systematic analysis of how process quality varies across trajectories with different evidence states and final outcomes. To this end, we analyze search trajectories on BrowseComp-Plus and xBench, revealing two consistent patterns: successful trajectories often suffer from post-alignment inefficiency (redundant searches), whereas failed trajectories typically struggle with misdirected exploration and insufficient coverage. Motivated by these findings, we propose Process-Aware Grouped rEward optimization (PAGE), a framework for search agents that provides a concrete instantiation of state-dependent process supervision. PAGE partitions trajectories according to final outcomes, assigns differentiated process rewards to successful and failed groups, and integrates them with outcome supervision at either the reward or advantage level. Experiments on three backbones across six multi-hop QA and web-search benchmarks show that PAGE consistently improves answer accuracy over outcome-only baselines.


Beyond Pairwise Supervision: Spectral Characteristic Matching for Data-Efficient Multimodal Alignment

Xinchao Wang ⋅ Shuang Li ⋅ Wei Chen ⋅ Jingxuan Kang ⋅ Fuzhen Zhuang ⋅ deqing wang

Multimodal alignment under scarce paired supervision aims to align independently pretrained unimodal encoders using limited image--text pairs, offering a practical alternative when large-scale joint pretraining is infeasible. Such scenarios call for lightweight objectives that go beyond local correspondences and capture distributional mismatch between heterogeneous embeddings. Existing methods often derive supervision from limited pairs through contrastive objectives or geometry-preserving regularization, encouraging instance-level correspondence or structural consistency but providing limited direct supervision on global cross-modal mismatch. This motivates a distribution-level objective that complements paired supervision by comparing image and text embedding distributions. In this paper, we propose $\textbf{Spec}$tral Characteristic Distribution $\textbf{Align}$ment ($\textbf{SpecAlign}$), a lightweight spectral alignment framework that reduces cross-modal mismatch by comparing image and text characteristic functions over learnable spectral queries. Since characteristic functions uniquely determine probability distributions, SpecAlign provides a principled signal for capturing multi-scale discrepancies beyond pointwise alignment. To obtain informative and stable queries, SpecAlign employs a Structured Direction--Radius Sampler (SDRS) that decouples direction and radius modeling. Extensive evaluations on transfer, retrieval, robustness, and representation analyses demonstrate consistent gains, validating spectral distribution-level supervision for data-efficient multimodal alignment.


Beyond Parameter Arithmetic: Sparse Complementary Fusion for Distribution-Aware Model Merging

Weihong Lin ⋅ Lin Sun ⋅ Qilong Shi ⋅ Aomufei Yuan ⋅ Yuxuan Tian ⋅ Zhengyang Wang ⋅ Guangxiang Zhao ⋅ Xiangzheng Zhang ⋅ Tong Yang

Model merging has emerged as a promising paradigm for composing large language models directly in weight space, enabling training-free integration of specialized models. However, existing methods rely on parameter-space averaging that systematically induces generation dysregulation---a spectrum of failure modes ranging from structural repetition and termination failure to circular reasoning, and at the extreme, full semantic collapse into incoherent symbols. Even when source models exhibit near-zero dysregulation, existing methods introduce it at $>$97\% (14B/32B scales), collapsing reasoning benchmarks by up to 70 points. We propose Sparse Complementary Fusion with Reverse KL (\textbf{SCF-RKL}), a data-free merging framework that selects complementary parameters via reverse KL divergence on parameter-group proxy distributions. This structure-preserving, sparsity-inducing design maintains proxy-space geometry---theoretically motivated via entropy and subspace bounds---and empirically suppresses generation dysregulation while integrating new capabilities. Extensive experiments on 24 benchmarks across 7B--32B models spanning reasoning, instruction following, safety, and vision demonstrate that SCF-RKL achieves the best overall performance across scales while maintaining near-zero dysregulation ($<$1\%) and strong generalization.


Beyond Prediction: Steering VLM Agents with Retrospective World Modeling

Yongjiang Liu ⋅ Jie ZHANG ⋅ Haoyue Zhang ⋅ Jingcai Guo ⋅ Deze Zeng ⋅ Song Guo

Equipping VLM agents with world modeling capabilities has shown strong potential for complex reasoning and long-horizon planning, while reducing the dependence of policy learning on costly real-world interactions. Existing methods mainly rely on prospective simulation to predict the consequences of candidate actions. However, this forward-only paradigm focuses on *what will happen next* and provides limited constraints for verifying whether an action is causally consistent with the observed state transition, which can lead to plausible-looking but physically incoherent behaviors. In this paper, we challenge the view of world modeling as only prospective prediction and introduce Retrospective World Modeling, a new agent learning paradigm that enables agents to reason backward by estimating the retrospective attribution distribution $P(\hat{a}_{t}|s_t, s_{t+1})$ for the action that most likely caused a given transition. Based on this capability, we formulate the Self-Consistency Reward (SCR), an intrinsic signal that measures the probabilistic consistency between the policy action and the retrospective explanation. Integrating SCR into reinforcement learning provides dense transition-level feedback and steers agents toward behaviors that are both task-effective and physically grounded. Extensive experiments across diverse agentic tasks show that our method substantially improves policy robustness and generalization over prospective-only world modeling baselines.


Beyond Real or Fake: A Dual-Channel Authenticity and Reasoning Protocol for Photographic Assessment

Xiaoxiao Li ⋅ Ruinan Jin ⋅ Lili Meng ⋅ Miaosen Wang ⋅ Athula Balachandran ⋅ PEI CAO

Forensic verification has converged on a binary "real vs. fake" label that conflates fully synthetic, tampered, and AI-retouched images despite their very different consequences. We argue that GenAI manipulations decompose along two complementary channels: a camera channel of sensor fingerprints (destroyed by neural processing), and a semantic channel of scene-level coherence (disrupted by content edits). Each manipulation type leaves a distinctive two-channel signature. We instantiate this taxonomy in 2CAP (2-Channel Authenticity Protocol), pairing a contrastively-trained camera encoder with a frozen semantic encoder to serve two applications on a shared backbone: (i) four-class authenticity classification via cross-attention fusion with learnable reliability weights; and (ii) evidence generation, where an authentic image serves as a reference for a query and the two encoders produce per-channel alignment scores together with a patch-level saliency map. These signals are supplied as privileged context to a frozen Vision Language Model (VLM) through an Observe-Generate-Refine loop, yielding explanations with localization without forensic fine-tuning. On a multi-source benchmark, 2CAP attains the best classification performance, including strong retouching detection where existing methods fail, and improves explanation quality.


Beyond Row Alignment: Virtual-Camera-Aware Online Stereo Rectification

Yizhao Peng ⋅ Rui Gong ⋅ Zaiwang Gu ⋅ Xudong Jiang ⋅ Jun Cheng

Stereo matchers estimate depth by comparing the left and right images from a stereo camera pair, but they assume that the two cameras are accurately calibrated and rectified. In real deployments, small physical shifts after calibration, caused by camera motion, mechanical drift, vibration, or installation changes, can break this assumption and degrade depth estimation. Online stereo rectification is therefore needed to correct the image pair during operation and recover a stereo geometry where standard stereo matching can work reliably. Recent online rectification methods mainly optimize row alignment, making the left and right views of the same scene point fall on the same image row. We show that this target, while necessary, is insufficient for deployable stereo depth estimation. A rectifier may make a small set of matched points look well row-aligned, but still distort the full images in ways that hurt stereo matching: useful image regions may be cut off, large blank areas may appear, objects may be stretched or squeezed, or image scale may change unevenly across the view. We propose VirtualRect, an online stereo rectification method designed for reliable real-world deployment. Rather than treating rectification as simply making the left and right image rows line up, this method estimates a rectified virtual stereo camera. This makes the rectified images and the camera parameters come from the same geometry, so the images used for matching and the model used for depth computation remain consistent. To improve stability, it further constrains the rectification behavior over the whole image and reduces the influence of geometrically unreliable matches. Experiments on multiple datasets show that, compared with the prior state-of-the-art online rectifier, VirtualRect reduces vertical flow error by 17\% on Carla-Flowguided and 7\% on Semi-Truck Highway.


Beyond Semantic Alignment: Geometric Incomparability in Multi-Oracle Soft Fusion

Xiaodong Li ⋅ Binsheng zhao ⋅ Yan Chen ⋅ Bingbing Jiang ⋅ Feijiang Li ⋅ Peng Zhou ⋅ Yuhua Qian ⋅ Liang Du

Many decision-level fusion methods implicitly assume that semantically aligned soft outputs can be directly aggregated once they share the same label simplex. However, this assumption can fail even under identical label semantics, as oracle-specific differences in confidence sharpness, boundary uncertainty, and tail-mass allocation may place soft outputs in incompatible probability geometries and bias direct soft fusion. To address this issue, we formalize this failure mode as geometric incomparability and propose a pre-fusion comparability recovery framework, which learns a shared consensus through constrained oracle-specific rectifiers on the probability simplex. The framework includes a closed-form variant with frozen geometry descriptors and an alternating-refinement variant with consensus-dependent descriptors. Theoretically, we analyze the rectifier family, prove exact recovery under matched radial distortions, characterize the bias of direct fusion, and establish stable recovery under approximate mismatch. Experiments on controlled synthetic stress tests and real black-box decision-level settings show that comparability recovery consistently reduces fusion bias and improves reliability over direct aggregation. These results indicate that geometric comparability is a necessary condition for reliable soft-output fusion. The code is available at \url{https://anonymous.4open.science/r/CFCR-ACR-B8F1}.


Beyond Single-Shot Conditioning: Test-Time Condition Refinement for Diffusion-Based Image Restoration

Aiping Zhang ⋅ Jiangang Wang ⋅ Shangquan Sun ⋅ Yuning Cui ⋅ Fan Li ⋅ Renjing Pei ⋅ Wenqi Ren ⋅ XIAOCHUN CAO

Conditional diffusion models have emerged as a powerful paradigm for image restoration, leveraging pre-trained generative priors to recover high-fidelity details from degraded inputs. However, existing methods typically follow a static inference protocol, treating the degraded observation as a fixed control signal throughout sampling. This overlooks the potential of inference-time scaling, where guidance can be refined on the fly to improve output quality. Recent test-time optimization approaches that manipulate noise variables or denoising trajectories often compromise structural fidelity, while those targeting text embeddings lack the spatial precision needed for restoration. To address this, we propose Condition-Aware Test-Time Optimization (CATTO), a training-free method that iteratively refines the visual conditioning signal at inference time. CATTO performs reward-aligned condition refinement under the pre-trained prior to enhance perceptual quality. To avoid complex backpropagation, we design a gradient-free optimization strategy guided by a joint objective: maximizing a perceptual reward while enforcing trajectory- and condition-level consistency to explicitly preserve structural fidelity. This process is accelerated by optimizing within a low-dimensional frequency subspace and reusing reliable updates across nearby denoising steps. Experiments on image restoration benchmarks show that CATTO improves perceptual quality over strong diffusion-based baselines without updating model parameters.

Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-modal conflicts across modalities. However, due to the absence of ground-truth fused images, existing MMIF supervision commonly uses spatial-domain sources or gradient variants as surrogate ground truth, making the supervision mechanism inherently misaligned with the goal of MMIF and causing pixel-level compromise or modality bias. To address this, we propose a relation-constrained supervision paradigm that moves fusion supervision from the spatial domain to a learned relation space. Instead of forcing the fused image to approximate the sources, we use frozen pretrained representation models as information providers and design a learnable feature adapter to align heterogeneous DINO and CLIP features into a unified supervision space. The adapter infers three relation parameters, namely sharedness, dominance, and coordination radius, which define three losses corresponding to the MMIF's goal. To make this space reliable, we devise a self-supervised contrastive ranking objective tailored to the adapter and couple it with the fusion network through alternating optimization. Extensive experiments show that the proposed supervision space yields significant gains regardless of which mainstream backbone the fusion network adopts, offering a supervision paradigm better aligned with the goal of MMIF. Our code will be publicly available.


Beyond Structural Agnosticism: Stable-Rank-Guided LoRA for Structure-Aware Fine-Tuning

Yuanyang Cao ⋅ Xichun Liu ⋅ Fuwei Zhang ⋅ Meiqin Liu ⋅ Jianji Wang

Current parameter-efficient fine-tuning (PEFT) methods like Low-Rank Adaptation (LoRA) are structurally agnostic, applying uniform configurations across all layers, which overlooks their vast functional heterogeneity. We propose Structure-Aware LoRA (SA-LoRA), a new framework that automatically tailors fine-tuning intensity to each layer's intrinsic complexity. Specifically, we leverage the Stable Rank as a spectral metric to align the adaptation magnitude of each layer with its pre-trained spectral structure, enabling an automated, principled allocation of learning capacity that directly addresses the structural agnosticism of existing methods. To enhance adaptability and robustness, we introduce a hybrid calibration mechanism that fuses the task-agnostic prior with task-specific gradient feedback, underpinned by a budget-conservation principle to ensure stability. Extensive experiments demonstrate that SA-LoRA consistently outperforms strong PEFT baselines, often achieving state-of-the-art performance with enhanced stability. The code is available at https://anonymous.4open.science/r/SA-LORA-6809.


Beyond the Node Barrier: Zero-Shot Strategy Planning for LLM Training on Super-Nodes

Shijie Shen ⋅ Chong Li ⋅ Pierre Leca ⋅ Jiong Lou ⋅ Jie LI

Emerging super-nodes tightly couple multiple servers with symmetric high-bandwidth fabrics, substantially weakening the legacy bandwidth cliff between nodes. This makes cross-boundary TP/EP a viable part of the strategy space, shifting the optimal strategy basin away from the legacy practice of keeping TP/EP within a node. Yet adapting the parallel strategy to new platforms still requires expensive profiling and brittle manual tuning. We propose Symbolic Barrier-aware Planner (SBP), a profiling-free planner that selects DP/MP/PP factorization and activation checkpoint strength by computing a barrier-aware symbolic score from model shapes, collective semantics, and hardware specifications. On legacy clusters, SBP recovers measured-best configurations; on super-nodes, it captures the regime shift and achieves $1.40\times$ step-time speedup and a 10.62 percentage-point MFU improvement over the best node-local Megatron-style baseline on 128-die training. SBP explores the strategy space in seconds on a single CPU, without accelerator profiling for strategy selection.


Beyond Token Representations: Explicit Visual Object Grounding for Video Reasoning Segmentation

Jun Huang ⋅ zhangjunyang ⋅ Zijie Yue ⋅ Bingkun Wang ⋅ Yong Luo ⋅ Miaojing Shi

Video Reasoning Segmentation (VRS) aims to infer and segment target objects specified by reasoning queries in videos. Existing methods typically first infer the target and generate a few segmentation tokens for it, which are then fed into a mask decoder for mask prediction. However, such segmentation tokens tends to be both semantically monotonous and spatially ambiguous, making them largely insufficient to represent the target to guide mask prediction. In this paper, we propose an explicit visual object grounding framework for VRS. Specifically, we first propose a spatial-aware prompt generation and refinement scheme, in which we query a multimodal large language model to generate explicit spatial prompts (\ie, bounding box and point set) for the target and then refine them through multi-step dialogue. Furthermore, we introduce a query diversification and alignment module to generate auxiliary queries that describe the same target from different perspectives. We then enforce cross-query consistency between their predicted spatial prompts to alleviate the training bias caused by query limitation. Finally, we selects several reliable frames using a prompt-based keyframe selection strategy to support mask decoding and propagation. Extensive experiments on eight datasets demonstrate that our method significantly outperforms state of the art methods.


Beyond Uniform Detection: Adaptive Hallucination Detection for RAG Across Response Regimes

Jungwuk Park ⋅ Sejong Ryu ⋅ Jy-yong Sohn ⋅ Jaekyun Moon

Existing hallucination detectors for retrieval-augmented generation (RAG), whether based on model outputs (e.g., likelihood) or internal representations, typically apply a uniform detection strategy across responses. However, we find that different detection signals are informative in different response regimes, making a uniform strategy insufficient for both short- and long-form generation. In short-form tasks such as factoid QA, hallucination typically appears as selecting an incorrect final answer among multiple context-relevant candidates. In this regime, model-output-based signals that discriminate among plausible answers are particularly important. We therefore propose a bidirectional likelihood measure that evaluates logical consistency based on the retrieved evidence and a reasoning step generated after the response. In contrast, long-form descriptive responses are more prone to gradual drift: as generation proceeds, the model increasingly relies on its own prior generations and parametric knowledge, while unsupported continuations may remain locally plausible, making output-based cues less informative. For this regime, we measure context--knowledge conflict by tracking the directional alignment of context-grounding versus internal-knowledge contributions in the hidden states. Building on these regime-specific signals, we introduce ARGUS, an adaptive hallucination detector that emphasizes the more effective signal according to the response regime. Experiments on multiple RAG benchmarks show that our method achieves strong performance across both short- and long-form settings.


Beyond Visual Boundaries: Rethinking Scene Segmentation for Movie RAG

Dong-Hee Kim ⋅ Seonwoo Choi ⋅ Changbeen Kim ⋅ Jungmyung Wi ⋅ Juyeon Ko ⋅ Youngju Choi ⋅ Il Hyeon Mun ⋅ Hyunwoo J. Kim ⋅ Donghyun Kim

Understanding long-form video remains a fundamental challenge for multimodal large language models (MLLMs). Sparse frame sampling fails to capture fine-grained visual details, while dense sampling quickly exceeds context length limits. Retrieval-augmented generation (RAG) offers a promising middle ground by selectively retrieving relevant video segments for grounded generation, yet its effectiveness critically depends on the quality of the video segments used as retrieval units. In this paper, we investigate RAG for movie understanding, which demands story-level reasoning over characters, events, and narrative arcs spanning hours of content. Scene segmentation, a long-studied problem that partitions movies into semantically coherent units, is a natural candidate for defining such retrieval units. We reexamine whether existing methods actually serve this role through comprehensive evaluation on downstream movie understanding tasks, and find that they consistently fail to outperform naive uniform temporal chunking. Our audit of the most standard scene segmentation benchmarks reveals why: current annotations prioritize visually salient transitions over narrative event structure. Motivated by this mismatch, we introduce NarraScene, a narrative-centric scene segmentation dataset annotated with a three-level cognitive taxonomy spanning physical, character, and narrative change, where every valid boundary requires a narrative-level shift. When used as retrieval units, these narrative-grounded segments outperform uniform chunking on downstream movie understanding tasks, suggesting that the central challenge for scene segmentation in movie RAG is not detecting boundaries, but identifying the narrative event units that matter for movie understanding.


BiasTrojan: LLM Judgers Are Easily Distorted by Few Hundreds of Contrastive Biased Training Data

Zichen TANG ⋅ Zhenheng Tang ⋅ Qian Wang ⋅ Gaoning Pan ⋅ Yuhan Yang ⋅ Wei He ⋅ Shaohuai Shi ⋅ Xiaowen Chu ⋅ Bo Li

Large Language Models (LLMs) are increasingly deployed as automated judges to scale supervision for data curation, reinforcement learning (RL), and agentic systems. While existing works have extensively explored bias in pretrained LLMs, the origins of such inherent biases remain largely untraced. We trace these biased tendencies to cognitively biased patterns (e.g., Authority, Bandwagon) latent in training corpora, and expose that such patterns are naturally prevalent in large-scale pretraining corpora yet remain entirely undetectable by existing data cleaning pipelines. To demonstrate that this underexplored threat can be deliberately exploited, we introduce BiasTrojan, a framework that concentrates and injects these naturally occurring bias patterns into training samples via context-aware bias cues and contrastive preference pairs, augmented with counterfeited reasoning chains for efficient injection. Experiments across several LLMs from 7B to 70B on human-preference and fact-related datasets show that mere hundreds of deliberately biased samples suffice to compromise LLMs into biased evaluators, overriding their factual knowledge. The injected biases generalize robustly out-of-domain and persist despite massive continual post-training. These findings reveal that this underexplored latent threat poses far greater risks than commonly recognized: biased LLM judges whose evaluations propagate irreversibly downstream, underscoring the critical need for bias-aware auditing and strict scrutiny of LLM training data.


Binding Mode Matters: Hotspot-Aware Drug Discovery via Explorative Preferences

Dingshuo Chen ⋅ Hao Yang ⋅ Kuangqi Zhou ⋅ Zhixun Li ⋅ Qiang Liu ⋅ xiangxiang Zeng ⋅ Shu Wu ⋅ Yaosen Min

Discovering hit molecules requires not just high binding affinity, but also the identification of diverse binding modes, which are critical for experimental assays. Existing generative approaches have predominantly relied on optimizing a scalar docking score, obscuring the distinct contributions of key binding determinants. To this end, we introduce a paradigm shift by formulating target-based drug design as a multi-objective exploration task, where each objective explicitly corresponds to enhancing interactions with a specific hotspot. Here, we introduce BindMol, a novel generative framework driven by a customized multi-objective reinforcement learning algorithm. By incorporating explorative preferences during training, our approach efficiently uncovers molecules with diverse and desirable binding profiles. Empirical results demonstrate that BindMol facilitates the discovery of high-affinity compounds characterized by both structural novelty and diverse binding modes. Validated across target-based drug discovery and multi-property optimization tasks, our approach provides a versatile paradigm for goal-oriented drug discovery.


Biomedical Acquisition-induced Style Shifts as Mixture Shifts: Style-aware Mixture-of-Experts Multimodal Prompt Learning

Pingyi Miao ⋅ Xianlai Chen ⋅ Ying An ⋅ Yunbo Wang ⋅ Linan Ren ⋅ Yuxin Peng

Vision-language models (VLMs), such as CLIP, have shown remarkable transferability in the biomedical domain, and prompt tuning enables efficient adaptation with limited supervision. In practice, biomedical images exhibit pronounced acquisition-induced style shifts, especially across imaging sites and protocols, while cross-modality transfer further introduces compound shifts from different physical acquisition processes. However, existing prompt-tuning methods ignore acquisition-induced style shifts, making the prompts sensitive to source-specific style and limiting the transferability to unseen medical domains. To address this, we formulate acquisition-induced style shifts as mixture shifts, where each domain is viewed as a mixture of latent style components. Under this formulation, the target risk depends on both mixture weights and latent component risks, motivating prompt adaptation that accounts for style-specific variation rather than relying on a single source-specific prompt. In this paper, we propose SaMoE, a Style-aware Mixture-of-Experts Multimodal Prompt Learning framework for adapting VLMs to biomedical domains under acquisition-induced style shifts. To account for latent component risks in prompt adaptation, SaMoE represents multimodal prompts with latent style-conditioned prompt transformations and composes them through image-conditioned routing, enabling each image to induce style-adapted visual prompts. By extracting style cues from intermediate visual representations, the sparse expert composition adapts style-related prompt weights without domain labels and is injected across multiple VLM layers to align with hierarchical visual-language representations. Extensive experiments on 15 medical datasets across 12 modalities and 10 organs demonstrate significant improvements in both accuracy and generalizability over state-of-the-art methods.


Birth-Death Structural Learning for 3D Gaussian Splatting

Tran Hong Quan ⋅ Long Nguyen-Chi ⋅ Binh T. Nguyen

3D Gaussian Splatting (3DGS) has become a standard representation for real-time novel-view synthesis. Yet, its quality--efficiency trade-off still relies heavily on adaptive density control. Existing cloning, splitting, and pruning rules use proxy signals—such as position gradients, opacity, image-space error. Although effective in practice, these signals do not directly answer the structural question behind density control: where should new Gaussians be allocated to reduce the multi-view reconstruction loss, and which existing primitives can be pruned with minimal penalty? To address this, we propose first-variation birth–death control for 3DGS, a principled approach that replaces heuristic structural decisions with a variational score derived from the reconstruction loss. At each structural update, we freeze the current Gaussian population, attach an auxiliary mass coordinates to each primitive, and differentiate the multiview loss with respect to these coordinates. We demonstrate that this score is a coordinate of the first variation of Splat Regression Model, giving a rigorous local descent surrogate for birth-death moves: high-score Gaussians are selected as birth sites, low-score low-opacity Gaussians are death. Across standard 3DGS benchmarks, integrating this score into 3DGS-MCMC yields state-of-the-art reconstruction quality and outperforms existing methods, delivering the most significant gains under tight Gaussian budgets. Ultimately, our work provides compelling evidence that gradient-based variational metrics should replace proxy heuristics for structural updates in 3DGS.


Blocked Gibbs meets Diffusion Transformers: Unsupervised Learning for Constraint Optimization

Yudong Will Xu ⋅ Wenhao Li ⋅ Xiaoyu Wang ⋅ Scott Sanner ⋅ Elias Khalil

Diffusion models have shown promise in learning to solve constraint optimization problems. However, they are mostly restricted to problems with binary variables and rely on graph neural networks, hindering their application to a broader range of problems such as those with general discrete variables or constraint structures that necessitate global rather than local reasoning. We investigate the use of Diffusion Transformers to address the aforementioned limitations. A naive implementation performs poorly due to a fundamental mismatch between the standard diffusion process and constraint solving: while the former applies small, incremental denoising across all variables, the latter requires substantially altering specific subsets of variables to attain feasibility or optimality. Our method, Blocked Gibbs Diffusion Transformer (BloGDiT), is the first to address this limitation by replacing standard joint Gaussian denoising with Blocked Gaussian denoising. BloGDiT uses iterative block resampling and anneals the block size over time to facilitate large, targeted edits within a block of variables. Across Sudoku, Graph Coloring, Maximum Independent Set, and MaxCut, BloGDiT matches or outperforms existing methods, demonstrating that blocked Gibbs-style diffusion provides a highly effective inductive bias for Transformer-based constraint satisfaction and optimization.


Block Sparse Flash Attention

Daniel Ohayon ⋅ Itay Lamprecht ⋅ Itay Hubara ⋅ Israel Cohen ⋅ Daniel Soudry ⋅ Noam Elata

Modern large language models increasingly require long contexts for reasoning and multi-document tasks, but attention's quadratic complexity creates a severe computational bottleneck. We present *Block Sparse Flash Attention* (BSFA), a drop-in replacement that accelerates long-context inference while preserving model quality. Unlike methods that predict importance before computing scores, BSFA computes exact query-key similarities to select the top-$k$ most important value blocks for each query. By comparing per-block maximum scores against calibrated thresholds, we skip approximately 50% of the computation and memory transfers for pruned blocks. Our training-free approach requires only a one-time threshold calibration on a small dataset to learn the per-layer and per-head attention score distributions. We provide a CUDA kernel implementation that can be used as a drop-in replacement for FlashAttention. On Llama-3.1-8B, BSFA achieves up to $1.13\times$ end-to-end speedup on LongBench with only a $1.1\%$ accuracy drop, and up to $1.24\times$ on Needle-in-a-Haystack retrieval at a $1\%$ accuracy drop. The attention kernel itself accelerates by up to $1.38\times$. We compare BSFA against five recent sparse attention baselines (SpargeAttention, MInference, FlexPrefill, XAttention, and BLASST), and verify the method on Qwen2.5-7B and on A6000 and H100 GPUs. The implementation is available at https://github.com/Anonymous44414/Block-Sparse-Flash-Attention.


Blur Issue Matters for Thermal Novel View Synthesis: A Floating Gaussian Suppression Approach

Mingyuan Xie ⋅ Sheng Wan ⋅ Jin Xie ⋅ Hang Yang ⋅ Liangyun Sun ⋅ Chen Gong

Novel View Synthesis(NVS) for thermal infrared scenarios has become increasingly important in practical applications. Recent advances in 3D Gaussian Splatting(3DGS) have demonstrated its strong performance for NVS, and a handful of preliminary efforts have explored its extension to thermal NVS. However, existing 3DGS-based methods for thermal NVS often suffer from severe local blur in rendered images. Theoretical and empirical analyses reveal that the single-channel nature of thermal data allows intensity discrepancy to be compensated through opacity adjustment, which can erroneously amplify the opacity of floating Gaussians, especially when the background behind them is insufficiently represented. To address this issue, we propose a method termed Floating Gaussian Suppression (FGS) for thermal NVS. Specifically, we leverage an over-saturation of luminance to introduce an implicit gradient penalty, which guides the optimizer to suppress the opacity of floaters to match the ground truth. By effectively reducing the opacity of these floaters, this mechanism alleviates the occlusion of distant details, thereby mitigating local blur. Moreover, to preserve rendering fidelity, the luminance parameters of Gaussians are optimized concurrently. Experimental results demonstrate that our method effectively mitigates local blur and produces sharper and more faithful thermal reconstruction compared with the baseline methods.


BoardGameArena: A Multi-Dimensional Benchmark for Strategic Reasoning of LLMs in Board Games

Wenxiao Zhao ⋅ Shu Wang ⋅ Renxi Wang ⋅ Yi Yi ⋅ Yasi Zhang ⋅ Chengyuan Ma ⋅ Deqian Kong ⋅ Pan Lu ⋅ Ying Nian Wu ⋅ Benjamin Yao

Large language models (LLMs) excel at reasoning across diverse domains, from mathematical problem-solving to code generation, yet struggle with strategic reasoning that requires long-term planning, opponent modeling and integrating tactical calculation with strategic understanding. Board games, with their fully observable states, perfect-information environments and objectively verifiable decisions, offer ideal testbeds for probing these capabilities. We introduce Board Game Arena (BGA), a multi-dimensional benchmark that evaluates LLMs' strategic reasoning along three cognitive axes—deductive, inductive, and abductive reasoning—through three challenging task types across five classic board games, yielding 45 sub-datasets with over one million samples. We evaluate 28 state-of-the-art LLMs on BGA and find that current models exhibit fundamental limitations in strategic reasoning for board games. While models demonstrate competence in understanding game rules and identifying legal moves, they struggle significantly with selecting optimal moves and applying strategic principles or tactical analysis. These results position BGA as a challenging and diagnostic benchmark for tracking progress in LLMs' strategic cognition.


Boosting Graph Contrastive Learning via Manifold-Guided Representation Disentanglement

Zhiqiang Li ⋅ Jianqing Liang ⋅ Jie Wang ⋅ Junbiao Cui ⋅ Zhiqiang Wang ⋅ Jiye Liang

Graph contrastive learning constructs augmented views and aligns corresponding instances across views to learn view-invariant representations. However, existing methods usually enforce cross-view consistency within a single representation space, lacking an explicit characterization of view-specific information. This may suppress task-relevant complementary information, reduce embedding diversity, and induce dimensional collapse. To this end, we propose Manifold-Guided Representation Disentanglement (MGRD), a plug-and-play framework for boosting graph contrastive learning. MGRD decomposes graph representations into shared and complementary subspaces. The shared subspace captures view-invariant semantics while preserving local graph structure through a graph-anchored manifold prior, whereas the complementary subspace captures residual view-specific information beyond the shared representation and is regularized by asymmetric consistency and subspace decorrelation. By jointly exploiting shared and complementary representations, MGRD balances cross-view invariance and representation diversity. Experiments on node-level and graph-level benchmarks show that MGRD consistently improves multiple GCL backbones, while theoretical and empirical analyses demonstrate its effectiveness in mitigating dimensional collapse.


Bounds on Extrapolation across Phase Transitions with Generalized Regression

Jeffrey Wei ⋅ Manolis Zampetakis ⋅ John Sous

Phase transitions occur when a physical system undergoes a dramatic transformation as it crosses a transition. We study whether regression methods trained only in one phase can predict the physical properties that characterize the unobserved phase. Our test sets consist of the analytically tractable Heisenberg model and angle-resolved photoemission spectroscopy measurements of a venerable high transition temperature superconductor. We find that polynomial regression and kernel methods recover the unobserved phase structure, while standard neural networks fail. To bound the error, we adopt the generalized eigenvalue problem (GEVP) into an upper bound on the worst-case transfer coefficient, defined as the ratio of mean squared error (MSE) loss on the extrapolation region to the MSE loss on the training region. Empirically, we find that our GEVP bound stays approximately constant with respect to the empirical transfer coefficient across training-set sizes and distance between the training and extrapolation region for both datasets, providing a principled framework for studying extrapolation feasibility.

Object-level spatial-temporal understanding is essential for video question answering, yet existing multimodal large language models (MLLMs) encode frames holistically and lack explicit mechanisms for fine-grained object grounding. Recent work addresses this by serializing bounding box coordinates as text tokens, but this text-coordinate paradigm suffers from a fundamental modality mismatch: object information is inherently visual, yet encoding it as text incurs a high token cost that forces aggressive temporal downsampling. We propose BoxTuning, an object-aware visual prompting framework for multimodal model fine-tuning that moves spatial-temporal object states from text-coordinate serialization into the visual stream while retaining only object identity in a compact text legend. Colored bounding boxes and trajectory trails encode geometry and motion on video frames, while the color-to-object legend provides the minimal textual identity bridge. This reduces the token cost significantly, achieving 87-93\% text token reduction in practice. It also preserves full temporal resolution, where the trajectory trails further encode inter-frame motion direction and speed within each keyframe, recovering fine-grained dynamics that text-coordinate methods are forced to discard. Experimental results on five video QA benchmarks (CLEVRER, Perception Test, STAR, NExT-QA, IntentQA) show that BoxTuning surpasses text-coordinate baselines on spatially oriented tasks and avoids their degradation on reasoning-centric tasks, establishing object-aware visual prompting as a natural and efficient way to convey object information to MLLMs.


Brain Economy-Aligned Graph Transformers

Stefano Vannoni ⋅ Inês W Sampaio ⋅ Eleonora Maggioni ⋅ Islem Rekik

Graph neural networks for brain connectomics treat every connection as topologically equivalent, ignoring the brain-economy trade-off between the metabolic cost of maintaining a connection and the topological value it delivers to the network. We argue that this trade-off is a property of the channel (i.e., edge) between two brain regions, not of the regions themselves, and should therefore shape how regions interact during attention computation rather than being appended as a node feature or auxiliary loss. We introduce the ecospace, a two-dimensional coordinate system that characterizes each brain connection by its economy—the trade-off between connectivity strength and anatomical cost—and uses it to guide how brain regions attend to one another. We further introduce EcoSpace Rotary Encoding (ESRE), an attention mechanism that injects the ecospace dimensions through asymmetric rotations of the query-key subspace, guaranteeing by mathematical identity that the attention score between two brain regions depends on the economy of the edge connecting them. Together, these two contributions form the Brain Graph Transformer with EcoSpace Rotary Encoding (BGT-ESRE). On brain graph classification, BGT-ESRE outperforms classical, general-purpose, and brain-graph-specific baselines by a substantial margin in both accuracy and AUC. Ablations confirm that the rotary mechanism and the brain economy-aligned transformation each contribute independently, and that asymmetric rotary injection is superior to additive bias injection of the same economy measure.


Breaking Information Islands in Sparse Tuning via Small-World Connectivity

Xianchao Guan ⋅ Zijun Xiong ⋅ Yifeng Wang ⋅ Yubo Cui ⋅ Yaowei Wang ⋅ Guosen Xie ⋅ Xin Li ⋅ Zheng Zhang

Sparse tuning is widely used to adapt large language models due to its parameter efficiency. However, its effectiveness depends not only on the number of trainable parameters, but also on the feature-interaction topology induced by the sparse update. Limited or uneven cross-dimensional interactions may isolate subsets of feature dimensions, forming information islands that hinder task-specific information mixing. To address this issue, we propose Small-World Induced Fine-Tuning (SWIFT), a sparse tuning framework that induces small-world connectivity in parameter space. SWIFT combines block-diagonal sparse updates for local feature interactions with a fixed sufficiently scattering permutation that routes these updates to distant feature groups, creating sparse long-range shortcuts without additional trainable parameters. We further show theoretically that random or block-diagonal sparse tuning can suffer from information islands, whereas SWIFT restores positive expansion through permutation-induced routing. Extensive experiments on commonsense reasoning, natural language generation, and image classification demonstrate that SWIFT consistently outperforms competitive PEFT baselines and can match or surpass full fine-tuning.

Self-supervised learning (SSL) is increasingly applied to scientific imaging domains, such as microscopy, where labels are scarce but noisy data is abundant. These domains exhibit significant signal-dependent noise, which can create a representational shortcut: since both augmented views derive from the same noisy image, the encoder can boost view agreement by also encoding noise patterns rather than semantic content. To enable noise-robust representation learning without knowledge of underlying noise models, we introduce Self-Aligned Noise Augmentation (SEANA), a drop-in module that generates noise-aware SSL views. SEANA learns an invertible variance-stabilizing transform (VST) from noisy data, then estimates the clean signal and resamples independent noise in learned VST space, requiring no clean targets, no dataset-specific denoiser training or use at inference, and no changes to the SSL objective or encoder. On CIFAR-10 and ImageNet-100 under synthetic Gaussian, Poisson–Gaussian, and multiplicative noise, SEANA improves clean-test linear-probe accuracy across contrastive/non-contrastive SSL methods, indicating higher-quality representations. On real fluorescence microscopy, SEANA improves Jurkat cell-cycle stage classification by up to 28.6pp over standard SSL and 27.5pp over denoiser-preprocessed baselines. In both settings, SEANA outperforms the denoiser-preprocessed SSL pipeline, demonstrating that learned VST-space noise-aligned resampling yields more robust representations than denoising alone. Code will be made publicly available.


Breaking the Static: Dynamic Text Conditioning for Diverse Image Generation

Meng Yu ⋅ Ruidong Chen ⋅ qingfeng shi ⋅ Yingmao Miao ⋅ Yancheng Bai ⋅ Lei Sun ⋅ Xiangxiang Chu ⋅ Hesheng Wang ⋅ Xiaodan Liang

High-performance text-to-image~(T2I) diffusion models suffer from diversity degradation. In this study, we attribute this challenge to a fundamental temporal mismatch: while visual features evolve dynamically in a coarse-to-fine manner in the denoising process, the text condition remains entirely static. Through empirical analysis, we demonstrate that decoupling the text prompt to prioritize low-frequency components in the early denoising stage better aligns with the visual denoising process, while simultaneously enhancing generation diversity. Driven by this insight, we introduce Dynamic Semantic Interpolation (DSI), a training-free strategy that utilizes low-frequency semantics during the initial stages to foster the sampling space and progressively recovers the full semantics, ensuring text-to-image alignment. We also provide a rigorous theoretical justification for DSI from the perspective of conditional entropy, explaining its ability to maintain a broader early sampling space. Extensive experiments across diverse state-of-the-art models demonstrate that DSI significantly boosts generation diversity while preserving text alignment and aesthetic quality.


Bricker to BRACE: A Bracket Exposure RAW Dataset and Restoration Model for Flicker-Banding

Zihan Zhou ⋅ Libo Zhu ⋅ Jue Gong ⋅ Zhiyi Zhou ⋅ Jiezhang Cao ⋅ Yong Guo ⋅ Yulun Zhang

Flicker-banding (FB), arises from temporal aliasing between a camera's rolling shutter and a display's brightness modulation, degrading screen-captured image readability with color shifts and jagged patterns. Existing single-frame methods with simplified parametric stripe models cannot reliably distinguish these artifacts from genuine texture. To address this, we conduct a systematic analysis of complex FB morphologies and reveal their significant variation across exposure settings, motivating a multi-frame bracketed RAW restoration paradigm. We construct Bricker, a synthetic–real bracketed RAW dataset built via ray-tracing-based physical simulation and automated multi-exposure capture tool. We further propose BRACE: Bracketed RAW Flicker-Banding Removal, a multi-frame restoration model that utilizes frequency-aware banding prior and a multi-scale spatial cross-attention modulator (MSCAM) for cross-exposure spatial fusion. We also introduce the Stripe Frequency Consistency (SFC) metric to evaluate banding removal. Extensive experiments demonstrate state-of-the-art performance on both synthetic and real benchmarks. Our dataset and code will be publicly released.

Offline reinforcement learning (RL) offers a promising framework for deploying autonomous systems in safety-critical settings without the risks of online exploration. However, learning policies that simultaneously achieve high performance and strong safety guarantees from fixed datasets remains a fundamental challenge. Many existing safe offline RL approaches typically rely on soft constraint formulations, which may permit safety violations and are sensitive to distributional shift. In contrast, formal methods such as Hamilton–Jacobi (HJ) reachability and Control Barrier Functions (CBFs) provide rigorous safety guarantees, but often yield overly conservative solutions, often neglecting performance. In this work, we bridge this gap by formulating safe offline RL as a state-constrained optimal control problem, where safety is enforced through hard state constraints and performance is captured via a reward function. The resulting value function satisfies a Hamilton–Jacobi–Bellman (HJB) equation, which we approximate using offline RL on fixed datasets. This formulation enables principled integration of safety guarantees with data-driven policy optimization. Empirically, across safety-critical benchmarks including boat navigation and Safety-Gymnasium tasks, our approach achieves competitive returns while exhibiting near-zero constraint violations, demonstrating a favorable balance between safety and performance in the offline setting.


Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation

Xingtai Gui ⋅ Meijie Zhang ⋅ Tianyi Yan ⋅ Wencheng Han ⋅ jiahao gong ⋅ Feiyang Tan ⋅ Cheng-Zhong Xu ⋅ Jianbing Shen

End-to-end autonomous driving aims to generate safe and plausible planning policies from raw sensor inputs, and constructing an effective scene representation is a critical challenge. Driving world models have shown great potential in learning rich representations by predicting the future evolution of a driving scene. However, existing driving world models primarily focus on visual scene representation, and motion representation is not explicitly designed to be planner-shared and inheritable, leaving a schism between the optimization of visual scene generation and the requirements of precise motion planning. We present WorldDrive, a holistic framework that couples scene generation and real-time planning via unifying vision and motion representation. We first introduce a Trajectory-aware Driving World Model, which conditions on a trajectory vocabulary to enforce consistency between visual dynamics and motion intentions, enabling the generation of diverse and plausible future scenes conditioned on a specific trajectory. We transfer the vision and motion encoders to a downstream Multi-modal Planner, ensuring the driving policy operates on mature representations pre-optimized by scene generation. A simple interaction between motion representation, visual representation, and ego status can generate high-quality, multi-modal trajectories. Furthermore, to exploit the world model’s foresight, we propose a Future-aware Rewarder, which distills future latent representation from the frozen world model to evaluate and select optimal trajectories in real-time. Extensive experiments demonstrate that WorldDrive achieves leading planning performance among vision-only methods while maintaining high-fidelity action-controlled video generation capabilities. Our code and model will be made publicly available.


Bridging the Modality Bottleneck in Pathology MIL through Virtual Molecular Staining

Yucheng Xing ⋅ Pei Liu ⋅ Jingying Ma ⋅ Ruping Hong ⋅ Jiangdong Qiu ⋅ Tianyu Liu ⋅ Kai He ⋅ Ling Huang ⋅ Mengling Feng

Multiple instance learning (MIL) is the dominant framework for whole-slide image analysis in computational pathology, typically combining a frozen patch encoder, a projection layer, and a slide-level aggregator. While encoders and aggregators have been extensively studied, the projection layer remains a largely morphology-only bottleneck. This limits endpoints such as biomarker status and survival, which are governed by a molecular state that is not fully captured by H&amp;E morphology. We introduce Molecularly Informed Staining Transform (MIST), a plug-in replacement for the MIL projection layer that uses paired spatial transcriptomics only during training to construct virtual molecular stains. MIST clusters gene expression profiles into cross-modal prototypes, anchors them in the frozen foundation model feature space, and uses them to reorganize H&amp;E patch features along molecularly guided axes. It requires no transcriptomics at inference and can be inserted before standard MIL aggregators. We evaluate MIST across 23 downstream tasks and 8 MIL aggregators. MIST improves 240 of 256 configurations over the standard projection layer, with an average gain of +3.5\%, observed consistently across endpoint types: +5.2\% on survival prediction, +3.3\% on tissue subtyping, and +2.6\% on biomarker prediction. Ablations confirm that gene-derived prototypes are the primary source of the gains, while spatial, biological, and pathological analyses show that cross-modal prototype affinities capture spatially coherent molecular programs from H&amp;E alone.


Brittlebench: Quantifying LLM robustness via prompt sensitivity

Angelika Romanou ⋅ Mark Ibrahim ⋅ Candace Ross ⋅ Kerem Oktar ⋅ Chantal Shaib ⋅ Samuel J Bell ⋅ Anaelia Ovalle ⋅ Antoine Bosselut ⋅ Jesse Dodge ⋅ Koustuv Sinha ⋅ Adina Williams

Existing evaluation methods largely rely on clean, static benchmarks, which can overestimate true model performance by failing to capture the noise and variability inherent in real-world user inputs. This is especially true for language models, which can face human-generated text queries containing mistakes, typos, or alternative ways of phrasing the same question. In this work, we introduce a theoretical framework for quantifying model sensitivity to prompt variants, or brittleness, that can enable us to disentangle data-induced difficulty from prompt-related variability. Using this framework, we design a novel evaluation pipeline, Brittlebench, to holistically evaluate the sensitivity of frontier models. We apply semantics-preserving perturbations to a suite of popular benchmarks, and observe model performance to degrade as much as ~21%. However, these perturbations do not affect all models equally: even a single perturbation alters the relative ranking of models in 63% of cases, impacting conclusions about comparative model performance. Decomposing the total variance of both state-of-the-art open-weight and commercial models, we find that semantics-preserving input perturbations can account for up to half of the performance variance for a given model. Brittlebench highlights the need for more robust evaluations and models, and allows us to systematically understand model brittleness.


ByteDistill: Cross-Tokenizer Distillation via Chunk-wise Byte-Level Distribution Alignment

Yang Chen ⋅ Xianqi Yu ⋅ SHAOWEI YAO ⋅ Fuyu Lv ⋅ Dan Ou ⋅ Haihong Tang

Token-level distillation assumes a shared categorical support, an assumption violated when teacher and student tokenizers segment the same text differently. Existing cross-tokenizer methods sidestep this mismatch by aligning token spaces, matching likelihoods along tokenizer-specific paths, or filtering shared spans, all of which can approximate the original distributional objective or discard probability mass. We argue that decoded UTF-8 bytes are common observable events shared by every text tokenizer, while token boundaries are tokenizer-specific latent transitions. Building on this view, we introduce \emph{ByteDistill}, which compares the conditional distribution of the next observable byte rather than aligning token identities. Its loss, \emph{Byte-Level Distribution Alignment} (BLDA), projects each native softmax into a 256-way observable next-byte distribution inside aligned decoded chunks, marginalizing token-ending mass as a latent boundary transition along the observed tokenization path. ByteDistill consistently outperforms supervised fine-tuning and prior cross-tokenizer baselines in both off-policy and on-policy settings, with off-policy training improving GSM8K accuracy by $1.75$ percentage points over the current state of the art.


C2FT: Enhancing Fine-Grained Perception in MLLMs via Confuse-then-Contrast Fine-Tuning

Shaoxuan He ⋅ Benlei Cui ⋅ Shikai Qiu ⋅ Yuwen Zhai ⋅ Hui Xue' ⋅ Longtao Huang ⋅ Jingqun Tang ⋅ Haiwen Hong

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities, yet they frequently struggle with fine-grained visual perception, often suffering from hallucinations when faced with visually similar but semantically distinct instances. A major underlying cause is the conventional single-sample Supervised Fine-Tuning (SFT) paradigm, which isolates training instances and limits the model's ability to explicitly learn subtle discriminative boundaries. To address this, we propose C2FT, a novel Confuse-then-Contrast Fine-Tuning framework that enhances fine-grained perception by shifting from independent supervision to joint supervision. Specifically, C2FT assesses the MLLM's internal uncertainty to dynamically mine model-specific semantic confusions, subsequently constructing challenging multi-image groups comprising hard positive and hard negative samples. By interleaving these samples into a unified prompt, our approach forces the MLLM's internal attention mechanisms to cross-reference images and explicitly capture localized visual discrepancies. Extensive experiments on both Fine-Grained Visual Classification (FGVC) and visual question answering (VQA) benchmarks demonstrate the effectiveness of C2FT. For instance, on the Qwen3-VL-4B model, our approach achieves an average improvements of 3.87\% and 4.35\% points over standard SFT and GRPO baselines, respectively.


Can 4D Foundation Models Remember?

Guangzhao He ⋅ Hadar Averbuch-Elor ⋅ Wei-Chiu Ma

Perceiving and remembering the visual world is fundamental to navigating and interacting with our environment. Current 4D foundation models, such as camera-controllable video models or 4D reconstruction models, can perceive and reconstruct dynamic environments, but how well they remember what they have perceived remains an open question. Existing benchmarks largely rely on pixel-level metrics and lack ground truth for objects once they leave the field of view, making them unable to evaluate visual memory in an object-centric manner against references. To fill this gap, we introduce PersistBench, a dataset and metric suite that leverages 360° videos as omniscient ground truth and proposes three evaluation aspects: object permanence, motion continuity, and appearance preservation. Evaluating various models across diverse categories reveals that current models can only maintain short-term consistency that degrades significantly once objects leave the field of view. Our findings highlight the gap between current model capabilities and robust visual memory, providing guidance for future development of 4D foundation models.


Can AI Agents Synthesize Scientific Conclusions?

Hayoung Jung ⋅ Pedro V Diniz ⋅ José R Roveda ⋅ Abner F da Silva ⋅ Haeun Jung ⋅ Enoch Tsai ⋅ Aleksandra Korolova ⋅ Manoel Ribeiro

Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions. Yet, their ability to do so in high-stakes domains such as health remains unclear. We introduce SciConBench, a large-scale live benchmark of 9.11K questions and expert-written conclusions from systematic reviews to evaluate open-domain scientific conclusion synthesis. The benchmark draws on an expert-validated automated evaluation pipeline that decomposes conclusions into atomic facts and measures correctness and comprehensiveness via factual precision and recall. To mitigate data leakage, we further introduce SciConHarness, a clean-room evaluation harness that equips agents with controlled web interaction to ensure valid measurement. Evaluating 8 frontier models and deep research agents, we find that factual quality remains low: under clean-room settings, the best agent achieves only a factual F1 of 0.337. Our clean-room setting consistently reduces performance relative to unconstrained evaluation, suggesting that leakage inflates estimates of models' true synthesis capabilities. Finally, we audit consumer-facing agents (e.g., Google AI Overview, OpenEvidence) and find they frequently generate incomplete and sometimes contradictory conclusions, even when the ground-truth answer is available. Overall, our results show that reliable synthesis of scientific conclusions remains an open challenge, and that clean-room evaluation is essential for assessing open-domain AI agents.


Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

Siddharth Gollapudi ⋅ Prasann Singhal ⋅ Nilesh Gupta ⋅ Sewon Min

Language models (LMs) raise an alluring alternative to vector-based retrieval: \emph{generating} a relevant answer, entirely based on an in-context corpus. In spite of this appeal, prior studies consider proprietary systems or the smaller-scale reranking task, leaving corpus-scale in-context retrieval largely unexplored. In this work, we present the first systematic study of in-context retrieval on two scales practical retrievers demand: \emph{million-token} corpora and \emph{length-generalization} far beyond training-time sizes. We first introduce BlockSearch, a 0.6B LCLM retriever whose architectural and training modifications improve over prior LM baselines and length-generalize up to $10\times$ beyond its training regime. However, its retrieval still collapses under more extreme extrapolation. We trace this failure to an \emph{attention dilution} effect: as the corpus grows, irrelevant documents dominate the softmax denominator, minimizing the normalized mass on the gold document even as its pre-softmax score stays high. Motivated by this analysis, we introduce \emph{length-aware} adjustments to the attention softmax and \emph{document-level sparse attention}, improving retrieval at million-token scale, better performance than the concurrent $7\times$-larger MSA system, and comparable performance to dense retrieval. Together, our results position in-context retrieval as a viable alternative to classical retrieval, emphasizing attention control under extreme context growth as a new challenge.


Can LLMs Reliably Grade Olympiad Proofs? A Controlled Study of Mathematical Verification with LLMs

Azim Ospanov ⋅ Zijin Feng ⋅ Ding Ding ⋅ Chengwu Liu ⋅ Haoli Bai ⋅ Jiacheng Sun ⋅ Lifeng Shang ⋅ Farzan Farnia

Large Language Models (LLMs) have made remarkable progress on solving competition-level mathematics, yet their ability to verifying natural-language mathematical proofs remains relatively underexplored. Current proof verification primarily relies on expert inspection that is costly to scale and Olympiad-level problems represent a significant challenge in this area, as they require the meticulous evaluation of every claim in the reasoning chain. In this paper, we present a systematic study of LLM-based verification in this setting. Specifically, we introduce a human-curated dataset of 48 USAMO 2025 candidate solutions with expert grading, conduct a broad study of existing LLM verifiers in Olympiad mathematics across frontier models. We also propose an iterative self-critique pipeline TROJAN to generates high-fidelity adversarial proofs at scale. Our evaluation spans 20 models on 7 metrics across two prompt styles with three inference methods and three prompt templates, illustrating that the state-of-the-art models still have more room for improvement. To address this gap, we propose MABGrader, a Multi-Armed Bandit framework that reframes grading as arm selection over the discrete score space. MABGrader outperforms the SOTA ProofGrader by 26.4% in Quadratic Weighted Kappa and reduces grading error by 14%.

Passive image provenance asks whether pixels alone can reveal where an image came from: a human, an aggregate AI class, or a particular generator. This becomes a robustness problem once a source image can be edited before the verifier sees it. We study the problem as source--target verification under adversarial distribution shift. Our first result gives the exact best-case limit for any image-only verifier: the largest robust target-acceptance gap equals the minimum total-variation distance between the target distribution and the set of attacked source distributions. This quantity depends on the source, target, and edit class, not on the verifier architecture. Our second result explains why deployed public verifiers can fail before this statistical limit is reached. If the verifier can be emulated on the attack region to error $\varepsilon$, then a surrogate black-box attack reaches target acceptance within $2\varepsilon$ plus optimization error of the white-box optimum; score-revealing logistic and softmax heads over public features are identifiable, and approximate score access gives stable recovery bounds. Experiments on same-prompt real/diffusion benchmarks support this separation. Public CLIP-based interfaces collapse under targeted attacks, and stronger clean or adversarial retraining does not restore positive operational separation. In our evaluated settings, positive empirical upper bounds on the robust gap appear only under private low-bandwidth binary interfaces with abstention. The main lesson is simple: robust passive provenance requires both a source--target statistical analysis and an interface-aware evaluation of what the verifier reveals.

Protein-molecule virtual screening is increasingly cast as a problem of representation learning in a shared embedding space. Existing methods rely on dense holistic alignment, entangling invariant binding determinants with nuisance correlations and limiting transfer to new targets. It has been noted that binding in protein–molecule systems involves sparse cross-modality interactions: binding is governed by a small contact interface and a few decisive local interactions (e.g., hydrogen bonds, hydrophobic contacts, and salt bridges) rather than the global structures of the protein and molecule. We hypothesize that uncovering and leveraging sparse interaction patterns is critical for generalization beyond the training data, as these patterns are reusable and expected to improve performance across different scenarios. In this paper, we aim to identify and leverage sparse interaction patterns, and verify our hypothesis. Since the training data contain only observed binding pairs, we formalize this prior via a V-structure causal model under Heckman-style selection, and establish three theoretical results: (i) the latent concepts of interacting proteins and molecules are not identifiable without appropriate sparsity constraints;(ii) these concepts and their sparse interactions are component-wise identifiable under structural sparsity conditions; and (iii) a low-rank relaxation of these conditions yields subspace identifiability of the concepts and interactions. Inspired by these principles, we propose CausalBind with three implementation variants. Extensive experiments on DUD-E and LIT-PCBA benchmarks show that all variants consistently outperform strong baselines across all metrics, respectively.


CausalCompass: Evaluating the Robustness of Time-Series Causal Discovery in Misspecified Scenarios

Huiyang Yi ⋅ Xiaojian Shen ⋅ Yonggang Wu ⋅ Duxin Chen ⋅ He Wang ⋅ Wenwu Yu

Causal discovery from time series is a fundamental task in machine learning. However, its widespread adoption is hindered by a reliance on untestable causal assumptions and by the lack of robustness-oriented evaluation in existing benchmarks. To address these challenges, we propose CausalCompass, a flexible and extensible benchmark framework designed to assess the robustness of time-series causal discovery (TSCD) methods under violations of modeling assumptions. To demonstrate the practical utility of CausalCompass, we conduct extensive benchmarking of representative TSCD algorithms across eight assumption-violation scenarios. Our experimental results indicate that no single method consistently attains optimal performance across all settings. Nevertheless, the methods exhibiting superior overall performance across diverse scenarios are almost invariably deep learning-based approaches. We further provide hyperparameter sensitivity analyses to deepen the understanding of these findings. We additionally conduct ablation experiments to explain the strong performance of deep learning-based methods under assumption violations. We also find, somewhat surprisingly, that NTS-NOTEARS relies heavily on standardized preprocessing in practice, performing poorly in the vanilla setting but exhibiting strong performance after standardization. Finally, our work aims to provide a comprehensive and systematic evaluation of TSCD methods under assumption violations, thereby facilitating their broader adoption in real-world applications. The user-friendly implementation, documentation and datasets are available at https://anonymous.4open.science/r/CausalCompass-anonymous-5B4F/.


Causal discovery needs explicit epistemic standards

Harald Kugler ⋅ Leonard Henckel ⋅ Sebastian Weichwald

There are many causal discovery algorithms, but few guiding principles for choosing between them. Methods are usually compared through three lenses: identifiability, consistency, and benchmarks. Each introduces assumptions about the data-generating process that are often untestable. We argue that there are no shared standards for comparing, prioritizing, or even discussing those assumptions. We support this claim by showing that (i) identifiability assumptions are often preferred by convention rather than by principle, (ii) uniform consistency requires further untestable assumptions, while also the weaker pointwise consistency has arguably narrowed algorithm design, and (iii) benchmarks are not neutral tests but carry an epistemic role of their own: they quietly fix which regimes matter, which guarantees come into play, and what counts as success. In doing so, they shape what algorithm comparisons are taken to show and thus perform some of the same epistemic work as theory, but less transparently. In this position paper we argue for a more transparent, integrated practice: Make the epistemic standards for discussing identifiability and consistency assumptions explicit, especially their failure modes and scope conditions; design benchmarks to operationalize those standards, by stress-testing and ablating both. Then, theoretical and empirical results jointly clarify when, why, and how algorithms may work rather than accumulate as disconnected advances.


Causal Discovery over Clusters of Variables in Non-Markovian Systems

Tara Anand ⋅ Adèle H Ribeiro ⋅ Jin Tian ⋅ George Hripcsak ⋅ Elias Bareinboim

Causal discovery is the task of leveraging observational data to uncover causal relationships between variables. Recent work has extended these methods to operate over clusters of variables to improve scalability in high-dimensions and enable reasoning over higher-level entities. These approaches have been limited by strong assumptions including causal sufficiency. In this work, we introduce an approach for causal discovery over clusters in non-Markovian systems. First, we extend theory of graphical models in a knowledge-based context, to motivate introduction of a novel graphical equivalence class that can accommodate unobserved confounding. Then, we present a sound algorithm for causal discovery of learnable relationships between clusters of variables.

Foundation Models (LLMs/VLMs) exhibit strong semantic reasoning capabilities but remain challenged by situated planning in unstructured physical environments. A key limitation is the Semantic-Geometric Gap: while models interpret linguistic and visual intent, they lack explicit grounding in continuous spatial structures, yielding physically infeasible or unsafe plans. We propose Causal-Geo, a neuro-symbolic framework that bridges this gap by integrating LLM reasoning with rigorous geometric planning. Rather than treating physical constraints heuristically, we model the environment as a continuous metric tensor field. LLMs translate high-level semantic intents into potential functions that dynamically warp this manifold via conformal scaling. This casts intent-driven planning as a geodesic optimization problem, solved efficiently via discrete graph search. Experiments across diverse domains—from a macro-scale $12.2 \text{ km}^2$ wetland digital twin to micro-scale indoor robotic navigation—demonstrate that Causal-Geo enables robust zero-shot planning. It significantly outperforms reinforcement learning in sparse-reward settings and prevents the ``physical hallucinations'' common in representative hierarchical LLM planners. Crucially, Causal-Geo acts as a rigorous physical gatekeeper, guaranteeing invariant safety against foundation model variance and demonstrating graceful degradation under contradictory prompts. These results establish continuous geometric grounding as a principled, domain-agnostic pathway toward physically consistent situated agent planning.

Discovering causal structures in multivariate time series (MTS) is critical in domains such as finance and neuroscience, yet remains difficult in practice due to three intertwined obstacles. First, MTS are routinely corrupted by missing values, and naive impute then discover pipelines can distort the underlying data distribution, leading to false causal graphs. Second, recovering instantaneous links is already hard on complete data, and missingness further entangles imputation with graph learning. Third, real world data often exhibits heteroscedastic noise, whose variance depends on both instantaneous and lagged causes, violating the assumptions of standard causal discovery algorithms. Existing methods fail to address these challenges concurrently, typically assuming complete data, homoscedastic noise, or purely time lagged relationships. To address these gaps, we introduce Cheesefill, an optimal transport framework for causal structure learning under missingness. Cheesefill parameterizes causal mechanisms with conditional normalizing flows to capture heteroscedastic noise, and jointly learns a stochastic correction map that refines naive imputations into trajectories consistent with the inferred mechanisms. Theoretically, we show that optimizing over stochastic correction maps is equivalent to a conditional kernel reformulation of the Kantorovich problem; the neural parameterization used by Cheesefill then yields a tractable objective. Extensive experiments on synthetic and real data show that Cheesefill consistently outperforms existing baselines across different missingness mechanisms.

As state-of-the-art neural networks are deployed on reasoning and algorithmic tasks, exactness guarantees become increasingly important. However, high average-case accuracy can still mask inconsistent behaviors. This motivates exact certification, which asks for the smallest set of labeled examples needed to certify that a learned hypothesis equals the target. We show that while some hypotheses are easy to certify, even minimal overparametrization can make certification exponentially hard across several hypothesis classes. For threshold circuits of depth $\ge 2$, adding a single extra gate can force certificate sizes exponential in the input dimension. We show an analogous hardness result for log-precision Transformers with only constant architectural overhead. We also characterize approximate certification, showing that allowing only polynomially many mistakes still requires exponentially large certificates, whereas constant relative-error guarantees can hide exponentially many failures. Empirically, we study certification for circuits and Transformers trained to recognize binary addition and find that imperfect models can evade detection unless the certificate is exponentially large.


ChainForge: Tool-Chain Hijacking Attacks against LLM Agents via Execution-Grounded Tool Synthesis

Jiluan Fan ⋅ Haotian Zhu ⋅ Zhigang Lu ⋅ Junhao Xia ⋅ Xunzhu Tang ⋅ Shuchao Pang ⋅ Minhui Xue

LLM-based agents increasingly rely on external tools to accomplish complex tasks, yet the security of tool-calling pipelines remains poorly understood at the chain level. Prior attacks target individual tool invocations through prompt injection or metadata manipulation, but compromising a single step in a multi-step workflow is conspicuous and rarely sufficient for complex adversarial objectives. In this work, we uncover a more insidious yet realistic attack surface, tool-chain hijacking, in which an adversary constructs a coherent sequence of tools that collectively replace the agent's intended execution trace while still completing the user's task correctly, rendering the hijack invisible to both the agent and the user. To operationalize this threat, we propose ChainForge, an execution-grounded framework that mines agent execution logs to synthesize replacement chains through rollout-based optimization, then embeds adversarial payloads into the chain's code via multi-criteria iterative refinement. To systematically evaluate chain-level threats, we further construct ChainBench, a benchmark of 97 tasks across 4 domains. Experiments on four frontier LLMs show that ChainForge achieves a trace hijack rate of up to 98.54%, maintains task utility above 80.41%, transfers across models with over 84.95% success, and evades all evaluated defenses and code-safety scanners at substantially higher rates than single-tool baselines, exposing a critical blind spot in current agent security.


Characterizing Trainability of Instantaneous Quantum Polynomial Circuit Born Machine

Kevin Shen ⋅ Susanne Pielawa ⋅ Vedran Dunjko ⋅ Hao Wang

Instantaneous Quantum Polynomial Quantum Circuit Born Machines (IQP-QCBMs) have been proposed as quantum generative models that combine a classically tractable training objective—based on the maximum mean discrepancy (MMD)—with a potential quantum advantage motivated by sampling-complexity arguments. While recent works have explored this model across various application domains, fundamental questions remain: does the model suffer from exponentially vanishing loss gradients, known as the barren plateau problem—a pervasive obstacle in quantum machine learning—and how do regimes of trainability relate to regimes of possible quantum advantage? Here, we address both questions analytically. To study trainability, we derive closed-form expressions for the variances of the partial derivatives of the MMD loss function and establish general upper and lower bounds. We explicitly characterize how trainability depends on the generator set and the spectrum of the chosen kernel, identifying regimes in which low-weight kernels avoid exponential gradient suppression under structured topologies. Regarding potential quantum advantage, we reformulate the anti-concentration property in terms of the same generator-set quantities that govern trainability, enabling a unified analysis. We show that sparse IQP architectures can produce classically intractable output distributions while simultaneously remaining trainable, at least at lower-weight frequencies. Our analytical results corroborate the numerical observations of prior work and provides principled guidelines for designing scalable IQP-QCBM architectures.


Chasing Label Shifters: A Change-Aware Framework for Dynamic Graph Node Classification

Junghoon Kim ⋅ Seungyoon Choi ⋅ Hyunsung Kim ⋅ Chanyoung Park

Entities in real-world systems often evolve over time: users shift between information consumption patterns, researchers migrate between fields, and firms transition between financial risk states. When such systems are modeled as dynamic graphs, these transitions correspond to nodes whose class labels change between consecutive snapshots, which we call shifters. Correctly classifying shifters is often more consequential than classifying stable nodes with persistent labels, as detecting such transitions enables timely intervention in high-stakes settings. Yet, we identify a systematic failure mode across all major dynamic graph neural network architectures: shifter performance consistently lags behind stable nodes, with errors concentrated on the old label. Our structural analysis traces this failure to embedding inertia: neighborhood aggregation keeps shifter representations anchored to the old-class centroid, and this effect is strongest precisely where it is hardest to correct, namely for shifters whose neighborhoods remain aligned with the old class. Guided by this finding, we propose CHASE, a model-agnostic CHange-Aware framework for Shifting nodEs that wraps any dynamic GNN with targeted components for detecting label shifts and overriding stale neighborhood signals. CHASE consistently improves shifter performance across all tested models and datasets (up to 114.9% and 202.5% improvement on accuracy and F1 score, respectively) while preserving stable-node performance. We additionally contribute three dynamic graph benchmarks with naturally shifting labels, filling a gap in existing resources. The source code is available at https://anonymous.4open.science/r/CHASE-40CF/.


ChildPose: Foundation for Children Pose Modeling

Jiakai Chen ⋅ Yifan Shen ⋅ Boyi Li ⋅ Chuanmiao Dong ⋅ Xu Cao

Human pose estimation has advanced rapidly driven by large-scale adult-centric models. However, children remain significantly underrepresented due to data scarcity, distinct body morphology, and unique motion dynamics. We introduce ChildPose, a modular video foundation model designed for child-centric pose estimation. ChildPose transforms pretrained pose encoders into a streaming pediatric model via memory-augmented temporal attention and age-aware co-training. By maintaining a memory bank of adjacent frame representations, the model achieves robust localization under motion blur, occlusion, and atypical poses, while an auxiliary age-prediction objective enforces representations that capture developmental morphology. We curate a diverse, video-based pediatric dataset spanning various age groups and capture conditions, evaluating out-of-distribution generalizability on entirely unseen subjects. Across multiple 2D baselines, ChildPose consistently outperforms both off-the-shelf models and child-data fine-tuned variants, while ensuring stable sequence-level predictions. Our results underscore that effective pediatric pose estimation necessitates child-specific modeling beyond standard adult-centric estimators.


CHORD: Cross-Model Hallucination Detection via Relational Graph Discrimination

Yongxin Deng ⋅ Zhen Fang ⋅ Guansong Pang ⋅ Sharon Li ⋅ Ling Chen

Hallucination detection is critical for deploying large language models (LLMs) in real-world applications. Due to strong empirical performance, internal representation–based methods have emerged as the prevailing direction for detecting hallucinations, yet they remain largely model-specific and often fail to generalize to unseen LLMs. In this paper, we study an important yet underexplored problem, termed cross-model hallucination detection (CMHD), which aims to train hallucination detectors on source LLMs while ensuring performance on unseen target LLMs. The core challenge of CMHD lies in cross-model representation heterogeneity: hidden states from different LLMs exhibit severe semantic inconsistency, which limits the transferability of classical representation-based detectors. Through analysis and empirical validation, we show that relational structures over layers, tokens, and features capture transferable detection signals despite semantic inconsistency. Based on this observation, we propose cross-model hallucination detection via heterogeneity-oriented relational discrimination (CHORD), a relational graph discrimination framework that represents hidden states as joint relational graphs over layers, tokens, and features. It feeds these graphs into relation-aware attention to obtain structure-aware features, and uses meta-learning to shape these features toward transferable detection signals. Extensive experiments show that CHORD outperforms representative detection methods in cross-model generalization.


Christoffel-DPS: Optimal sensor placement in diffusion posterior sampling for arbitrary distributions

James Rowbottom ⋅ Zi Yuan (Nick) Huang ⋅ Carola-Bibiane Schönlieb ⋅ Ben Adcock

State estimation is a critical task in scientific, engineering and control applications. Since the reliability of reconstructions can depend on the number and position of sensors, Optimal sensor placement (OSP) is essential in scenarios where measurements are sparse and expensive. However, classical OSP approaches rely on Gaussian assumptions and are consequently unable to account for complex distributions encountered in many real-world systems. Generative-model-based reconstruction using sensor guided diffusion posterior sampling (DPS) has emerged as a promising technique for reconstructing states from highly complex distributions. Existing approaches to sensor selection either choose an unrealistically large number of sensors or employ strategies that emulate classical OSP methods. This results in a mismatch, wherein new models are paired with classical OSP tools, and motivates the need for fundamentally new ideas towards OSP that match the recent advances made in powerful recovery models. In this work, we introduce a distribution-free sensor placement framework based on the Christoffel function. Our main theoretical contributions introduce a mathematical formulation of optimal sampling and recovery guarantees for posterior sampling with arbitrary sensors and signal distributions. We use these to derive a new OSP strategy with non-asymptotic bounds on the number of sensors needed for recovery. Building on this, we develop \textbf{Christoffel-DPS}, with both offline and online variants, that implements nonparametric realizations of Christoffel sampling for generative models. As we show, Christoffel-DPS outperforms Gaussian OSP baselines and existing generative-model-based placement methods, validating that distribution-free sensing is both theoretically principled and practically superior. The framework is model agnostic, and we demonstrate its application to a range of unconditional DPS and flow matching models on structurally non-Gaussian benchmarks, showing the efficacy of Christoffel-DPS in low sensor budget regimes.


CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

Austin T D Andrews ⋅ Liam Wilkinson ⋅ Jamie Heagerty ⋅ Harry Coppock ⋅ Jakob Foerster ⋅ Rui Costa

We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, state monitoring, and execution under partial observability. The environment exposes 76 MCP tools and a narration layer that converts visual game state into structured text. We use CivBench to characterise agent behaviour across four model families in 23 admissible runs. The sample is a pilot, not a model ranking: aggregate outcomes do not reliably discriminate models at this scale. Instead, we introduce two interface-level metrics that the environment makes measurable: Proactive Monitoring Rate (PMR), capturing whether agents actively query latent strategic state, and RAG@10, capturing whether commitments stated in structured planning reflections are executed within ten subsequent turns. Across runs we observe two consistent patterns under a shared playbook protocol. Agents under-monitor strategically relevant state that is available but requires explicit querying: despite playbook guidance to query victory progress every 20 turns, agents do so only every 30–75 turns, and in 7 of 20 detectable defeats they failed to query within the 20-turn warning window before game end. Agents also frequently fail to execute near-term commitments stated in their own planning reflections (RAG@10 between 48% and 66% across models). Both patterns arise despite tool access and explicit guidance, and we interpret them as deviations under instruction rather than absences of capability. We release the environment, scenarios, logs, metrics, and analysis pipeline at https://anonymous.4open.science/r/civbench/README.md.


Clapping: Removing Per-sample Storage for Pipeline Parallel Learning with Communication Compression

Boao Kong ⋅ Xu Huang ⋅ Yuqi Xu ⋅ Yixuan Liang ⋅ Bin Wang ⋅ Kun Yuan

Pipeline-parallel distributed optimization is essential for large-scale machine learning but is challenged by significant communication overhead from transmitting high-dimensional activations and gradients between workers. Existing approaches often depend on impractical unbiased gradient assumptions or incur sample-size memory overhead. This paper introduces Clapping, a Communication compression algorithm with LAzy samPling for Pipeline-parallel learnING. Clapping adopts a lazy sampling strategy that reuses data samples across steps, breaking sample-wise memory barrier and supporting convergence in few-epoch or online regimes. Clapping comprises two variants including Clapping-FC and Clapping-FU, both of which achieve convergence without unbiased assumption for compressed gradient, effectively addressing compression error propagation in multi-worker settings. Numerical experiments validate the performance of Clapping across different tasks.


Class-Incremental Learning via LoRA-based Elastic Ensemble of Experts

Ruilong Yu ⋅ Fei Ye ⋅ Zhiyuan Ren ⋅ Qihe Liu ⋅ Adrian G. Bors ⋅ Rongyao Hu ⋅ shijie zhou

Large-scale pre-trained vision models provide powerful representations for class-incremental learning, yet continuously adapting them to new classes without historical samples, task identifiers, or unrestricted parameter growth remains a fundamental challenge. Existing LoRA-based approaches typically follow either a one-adapter-per-task expansion paradigm or a fixed parameter-sharing strategy. The former leads to linear parameter growth as the task sequence expands, while the latter often induces severe cross-task interference. To address these limitations, we propose ${\bf E}^3$ (Elastic Ensemble of Experts), a parameter-efficient framework for class-incremental learning. ${\bf E}^3$ treats LoRA modules as elastic experts whose task capacity is dynamically scheduled according to representational saturation. Within each expert, we further introduce Group-Aware Parameter Partitioning, which allocates LoRA parameters into disjoint task-specific subspaces using magnitude-based importance estimation, thereby mitigating parameter overwriting without additional gradient-sensitivity computation. Moreover, ${\bf E}^3$ incorporates hierarchical moment regularization and orthogonal gradient projection to constrain classifier drift and prevent new-task updates from disrupting historical feature subspaces. Extensive experiments on standard class-incremental learning benchmarks demonstrate that E³ achieves state-of-the-art or competitive performance, especially in long-horizon task sequences, while maintaining a favorable balance between stability, plasticity, and parameter efficiency.


Claude Coke: Prevent Automated Crime by Agents

Gabor Hollbeck ⋅ Baran Peters ⋅ Alexander von Recum ⋅ Jan Granacher ⋅ Yuri Simantob ⋅ Alec McGail ⋅ Kevin Riehl

The contemporary LLM-agent stack can automate substantive criminal activity at scale by composing capabilities that are already present, and the relevant engineering and governance controls need to be in place before that composition appears in production. The argument rests on three empirical legs that are usually studied separately but jointly create a closed loop of accidental or adversarial criminal retail automation. (i) Autonomous LLM agents fail catastrophically at well-studied multi-agent coordination tasks, with bankruptcy as the modal outcome rather than a tail event (ii) The rate at which agents elicit clearly harmful tools is not fixed but moves with common contextual manipulations. (iii) Autonomous browser control has become general enough that frontier agents reliably navigate, populate, and complete checkout flows on simulated illicit-product storefronts when light obfuscation is applied. Combining the three legs requires no new model capability: in our experiments, LLM agents autonomously completed simulated dark-web purchases of drugs, shotguns, and hitman services end-to-end. We propose three interventions: mandatory pre-deployment safety testing, hard human-confirmation gates at financially material steps, and verifiable agent ID infrastructure.


Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation

Xin Zou ⋅ Haolin Deng ⋅ Yibo Yan ⋅ Shuliang Liu ⋅ Zhiwei Jin ⋅ Chen Chen ⋅ Haonan Lu ⋅ Xuming Hu

Multimodal Large Language Models (MLLMs) are prone to hallucination as their generation preferences are insufficiently calibrated to visual evidence, causing them to fall back on linguistic priors, rather than faithful grounding. In this work, we start from an empirical observation: when query-relevant visual evidence is explicitly strengthened using the model’s own attention, generation becomes more accurate, suggesting that many failures do not arise solely from missing perception, but from an insufficient tendency to trust the evidence the model has already attended to. Motivated by this finding, we propose Oriented Pickup Preference Optimization (\texttt{OPPO}), an evidence-aware alignment objective that learns preferences over the strength of visual evidence, rather than only response quality. Concretely, \texttt{OPPO} contrasts the same faithful response under stronger, anchored, weaker-evidence views, turning naive visual preference into ordered visual-evidence alignment. We further combine this objective with fine-grained span-level and token-level regularization to stabilize the training. Besides, we provide a theoretical analysis showing that ordered evidence margins induce a positive lower bound on local visual sensitivity. Extensive evaluations across hallucination and general-purpose benchmarks demonstrate that \texttt{OPPO} consistently outperforms baseline methods.


ClusterSplat: Semantic Cluster Selection for 3D Visual Grounding in Gaussian Splatting

Jaesung Lee ⋅ Hyeontaek Hwang ⋅ Gabriel Manalu ⋅ Daeyoung Kim

We study 3D visual grounding in 3D Gaussian Splatting (3DGS), where a referring expression should identify a target object as both a rendered 2D mask and a set of explicit 3D Gaussians. Prior methods typically extract 3D targets by applying heuristic thresholding or top-ratio filtering to dense primitive-level relevance scores. However, these scores often vary substantially across scenes and prompts, making such heuristics sensitive and leading to unstable targeting. ClusterSplat addresses this by reframing 3D target selection as cluster-level selection over scene-adaptive candidates. The method first learns instance-aware Gaussian features, forms feature-aware seed voxels, and merges neighboring seed voxels when a description-length criterion decreases, producing a scene-specific set of cluster candidates. Given a referring expression, a query-conditioned cluster scorer ranks the cluster candidates with a lightweight MLP, and the selected cluster directly defines the explicit 3D target while its cluster score is rendered for the 2D mask. On ScanRefer, ClusterSplat achieves the best rendered 2D segmentation and explicit 3D target metrics among the compared 3DGS baselines. It also obtains the best average Ref-LERF scores, supporting the same trend under fine-grained referring descriptions. On ScanNet v2 class name queries, ClusterSplat achieves the strongest rendered 2D results and the highest 3D matching accuracy, while maintaining competitive 3D IoU.


Coarse-to-Fine Autoregression over Hierarchical Discrete Codes for Molecular Graph Generation

Haozhuo Zheng ⋅ Cheng Wang ⋅ Pengyu Chen ⋅ YajunTian ⋅ Yang Liu

Molecular graph generation typically involves a trade-off between two paradigms: diffusion models capture global structure well but require hundreds of denoising steps, while autoregressive (AR) models sample in a single pass yet generate atoms in a flat canonical order with no semantic hierarchy, leaving long-range topology to chance. We introduce the **Hierarchical Graph VQ-Transformer (H-GVT)**, which removes this dichotomy by performing coarse-to-fine AR generation over hierarchical discrete codes. A **Spectrally-Regularized Multi-Scale VQ-VAE** first compresses each molecule into a sequence of discrete tokens at multiple scales, where coarse tokens provide regularized low-resolution structural context and fine tokens encode atom-level details. The compression combines Adaptive Min-Cut Pooling with a novel **Spectral Topology Consistency Loss** that aligns low-frequency normalized Laplacian spectra across scales, adding only 2.8% training overhead. A standard decoder-only Transformer with level embeddings then generates these tokens from coarse to fine, allowing fine-grained atom tokens to condition on coarser structural context. This requires no specialized tree, motif, or scale-causal masking. On QM9 and ZINC250k, H-GVT achieves the best NSPDK among compared baselines ($2\times10^{-4}$ on QM9 and $1\times10^{-4}$ on ZINC250k), while sampling 10K molecules $68\times$ faster than DiGress. On MOSES, H-GVT achieves $4.8\times$ better FCD than a Novelty-thresholded GVT baseline under the same thresholded protocol (0.19 vs. 0.92; 86.4% vs. 80.5% Novelty). Crucially, ablations reveal that the coarse-to-fine ordering itself, not merely multi-scale tokenization, is what enables global structural planning: randomizing the order while keeping the same tokens, codebook, and level embeddings degrades NSPDK by $29\times$.


Co-evolution: A "One-to-many" LLM Fine-Tuning Paradigm

Dapeng Jiang ⋅ Haichuan Tan ⋅ Wenxuan Song ⋅ Dianqiao Lei ⋅ Shenqi Zong ⋅ Yanyan Lan

Fine-tuning has become a central mechanism for adapting large language models (LLMs) to downstream tasks, yet most existing methods follow a one-to-one paradigm: a single pretrained model is optimized into a single improved policy. This paradigm optimizes average single-policy performance, leaving no mechanism to distinguish redundant successes from genuinely complementary contributions. In this paper, we propose Co-Evolve, a one-to-many LLM fine-tuning framework that transforms a single base model into a cooperative population of persistent, specialized descendants and explicitly optimizes their complementarity. To reduce redundancy among descendants, we introduce a marginal-coverage objective that rewards each model for paying attention to examples under-covered by its peers. We instantiate this objective with evolution strategies, enabling derivative-free population-level optimization. We further introduce a seed-replay checkpoint mechanism that reduces the disk storage required to maintain multiple descendants by avoiding storing multiple full checkpoints during deployment. Experiments across multiple domains show that Co-Evolve consistently improves ensemble-level performance over strong fine-tuning and ensemble baselines. Our results suggest that optimizing for complementary model-level specialization provides a scalable alternative to single-policy fine-tuning. Our code can be found at https://anonymous.4open.science/r/Co-Evolution-Core-6604.

Robots operating in semantic environments often need to satisfy linear temporal logic (LTL) tasks before the semantic map is fully known. Many practical semantic-LTL planners maintain semantic beliefs but execute a single task policy induced by a point semantic interpretation, such as a maximum-a-posteriori (MAP) semantic map, possibly supported by risk checks or active perception. This premature commitment can be brittle when distant semantic observations are weak or correlated. We propose Coherence-Aware Transition-Intent Fusion, an online planner that preserves automaton-relevant semantic uncertainty using a finite top-$K$ intent abstraction. Given a conservative current automaton estimate, the planner propagates plausible semantic assignments through the automaton, converts feasible successor transitions into weighted reach-avoid intents, and fuses their risk-gated Dijkstra progress scores. When competing intents create incoherent motion, a KL-regularized calibration step shifts weight toward intents consistent with recent physical progress. Experiments on random-obstacle and room-like semantic maps show improved satisfaction over single-intent and raw-fusion variants, and over a TOAPP* MAP-plus-active-perception baseline under the same static-view sensor. Ablations attribute the gain primarily to intent-level calibration.


CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models

Joowon Kim ⋅ Seungho Shin ⋅ Joonhyung Park ⋅ Eunho Yang

Recent Thinking with Video approaches use Video Generation Models (VGMs) for visual reasoning by producing temporally coherent Chain-of-Frames as reasoning artifacts. Even strong VGMs, however, exhibit two recurring failure modes on goal-directed tasks: long-horizon drift on multi-step tasks and mid-clip simulation errors that compound. Both stem from the absence of explicit reasoning built upon the VGM's short-horizon visual prior, a role naturally filled by Vision-Language Models (VLMs), but where to place the VLM is non-trivial: upfront plans commit before any frame is generated and post-hoc critiques over whole videos intervene too late. We propose VLM-VGM Collaborative Video Reasoning (CollabVR), a closed-loop framework that couples the VLM with the VGM at step-level granularity: the VLM plans the immediate next action, inspects the clip the VGM generates, and routes test-time compute across qualitatively distinct recovery strategies (re-generation, action splitting) matched to the diagnosed failure. On Gen-ViRe and VBVR-Bench, CollabVR improves both open-source and closed-source VGMs over single-inference, Pass@k, and prior test-time scaling baselines at matched compute, with the largest gains on the hardest tasks. It also yields further improvements on top of a reasoning-fine-tuned VGM, indicating that step-level VLM supervision is orthogonal to and stackable with reasoning-oriented fine-tuning. We provide video samples and additional qualitative results at our project page: https://collab-vr.github.io.


CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization

Ziyang Ding ⋅ Linjian Meng ⋅ Yiming Wu ⋅ YuHan Li ⋅ Yuhao Liu ⋅ Zhen Zhao

Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perform visual reasoning by propagating continuous hidden states instead of decoding intermediate steps into discrete tokens. However, existing works typically rely on hard alignment objectives to force latent representations to match predefined visual features, thereby severely limiting the exploratory of latent reasoning process. To address this problem, we propose CoLVR (Contrastive Optimization for Latent Visual Reasoning). To obtain a more exploratory visual reasoning, CoLVR introduces a latent contrastive training framework. Firstly, CoLVR learns diverse and exploratory representations with a latent contrastive objective guided by angle-based perturbation, which expands the semantic latent space and avoids over-constrained embedding. Then, CoLVR employs a latent trajectory contrastive reward for RL (Reinforcement Learning) post-training to enable fine-grained optimization of latent visual reasoning process and thus fostering diverse reasoning behaviors. Experiments demonstrate that CoLVR significantly enhances the exploratory capability of latent representations, achieving average improvements of 5.83% on VSP and 8.00% on Jigsaw, while also outperforming existing latent models on out of domain benchmarks, with a 3.40% gain on MMStar.


Combating Data Laundering in LLM Training

Muxing Li ⋅ Zesheng Ye ⋅ Sharon Li ⋅ Feng Liu

Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights holders query a target LLM with raw proprietary data and assess whether the model assigns them stronger memorization-based detection signals, e.g., higher confidence or lower loss, than held-out non-training reference texts. We show that this regime becomes brittle under data laundering, where the target LLM is trained on semantics-preserving but stylistically or structurally transformed surrogates of proprietary data to obfuscate provenance. Since training-time exposure occurs in the laundered form, memorization signals may no longer appear on the originals, collapsing the candidate-reference signal separation that standard detectors rely on. We counter this threat by studying laundering-aware detection with raw proprietary data, a held-out reference corpus, and query access to the target LLM, while the laundering transformation is undisclosed. Since exact recovery of the laundered corpus is infeasible, we infer a detection-useful synthesis process via an auxiliary LLM that maps originals into training-like queries. To make this search tractable, we introduce Synthesis Data Reversion (SDR), which constrains the unbounded space of natural-language transformations through a goal-details decomposition: a high-level transformation goal, e.g., "lyrical rewriting", and fine-grained details, e.g., "with vivid imagery". SDR identifies the most likely goal and iteratively refines details so synthesized queries elicit stronger target-model detection signals. Evaluated on the MIMIR benchmark against diverse laundering practices and target LLM families (Pythia, Llama2, and Falcon), SDR consistently restores detection signals, offering a practical auditing tool against data laundering.


Communication-Efficient Federated Learning of Latent Patient Representations from Multi-Institutional EHRs

Zhiyu Yan ⋅ Shiao Liu ⋅ Rahul Mahajan ⋅ Anand Viswanathan ⋅ Christopher D Anderson ⋅ Rui Duan

Electronic health records (EHRs) contain rich health information for uncovering latent patient representations and clinically meaningful subgroups. Learning from EHR data across institutions is valuable because representations learned at a single site may reflect local data patterns and lack generalizability, yet pooling patient-level records across institutions is often impractical because of privacy constraints and communication costs. We propose a communication-efficient, non-iterative federated spectral method for learning shared latent representations from multi-institutional EHRs, in which each institution shares aggregate marginal frequencies and a low-rank spectral summary while keeping patient-level records local. A central server aggregates these summaries with the pooled marginal frequencies to recover both the shared low-rank structure and the canonical row-normalization direction required for interpretable representation recovery. We prove row-wise recovery guarantees and error bounds, with a leading federated variance term matching a corresponding lower bound under explicit regularity conditions. Simulations show that our method closely matches pooled analysis in estimating latent representations, improves as more sites contribute summaries, and outperforms single-site and existing federated benchmarks across a range of data settings. Applied to real-world multi-facility EHR data for stroke patients, the method learns interpretable patient representations that highlight clinically distinct subgroups and could help inform subgroup-specific care.


Communication-Efficient Personalized Adaptation via Federated-Local Model Merging

Yinan Zou ⋅ Md Kamran Chowdhury Shisher ⋅ Christopher Brinton ⋅ Vishrant Tripathi

Parameter-efficient fine-tuning methods, such as LoRA, offer a practical way to adapt large vision and language models to client tasks. However, this becomes particularly challenging under task-level heterogeneity in federated deployments. In this regime, personalization requires balancing general knowledge with personalized knowledge, yet existing approaches largely rely on heuristic mixing rules and lack theoretical justification. Moreover, prior model merging approaches are also computation and communication intensive, making the process inefficient in federated settings. In this work, we propose Potara, a principled framework for federated personalization that constructs a personalized model for each client by merging two complementary models: (i) a federated model capturing general knowledge, and (ii) a local model capturing personalized knowledge. Through the construct of linear mode connectivity, we show that the expected task loss admits a variance trace upper bound, whose minimization yields closed-form optimal mixing weights that guarantee a tighter bound for the merged model than for either the federated or local model alone. Experiments on vision and language benchmarks show that Potara consistently improves personalization while reducing communication, leading to a strong performance-communication trade-off.


CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection

Jiwon Song ⋅ Dongwon Jo ⋅ Beomseok Kang ⋅ jae-joon kim

Chunked prefill has become a widely adopted serving strategy for long-context large language models, but efficient attention computation in this regime remains challenging. Existing sparse attention methods are primarily designed for one-shot prefill and do not translate efficiently to chunked prefill: block-sparse kernels lose efficiency when the query length is limited by the chunk size, while fine-grained pattern search becomes costly when repeated over the accumulated KV cache at every chunk. QUOKA, a recent method that directly targets chunked prefill, avoids sparse-kernel overhead but relies on query-subsampled, token-level KV selection, which can miss query-specific KV entries and introduce explicit KV-copy overhead. To address these limitations, we propose CompactAttention, a chunked-prefill attention mechanism based on Block-Union KV Selection. CompactAttention treats 2D block-sparse masks as KV-selection signals rather than direct sparse-kernel execution plans, and converts them into GQA-aware per-group KV block tables through Q-block union and intra-group union. This construction produces the minimal block tables that preserve all KV blocks selected by the input masks under paged execution constraints, enabling selected KV blocks to be accessed in place without explicit KV compaction. On LLaMA-3.1-8B-Instruct, CompactAttention maintains accuracy close to dense attention on the RULER benchmark while delivering up to 2.72$\times$ attention speedup at 128K context length under chunked prefill.


Compact SO(3) Equivariant Atomistic Foundation Models via Structural Pruning

Chen Wang ⋅ Siyu Hu ⋅ Guangming Tan ⋅ Weile Jia

SO(3) equivariant graph neural networks have become the dominant paradigm for atomistic foundation models, achieving high accuracy and data efficiency by building rotational symmetry directly into the architecture. Yet the computational cost of their higher-order tensor operations creates a tough trade-off between model accuracy and inference efficiency. In this paper, we propose a structural pruning method for SO(3) equivariant atomistic foundation models to bridge this accuracy-efficiency gap. The pruning is applied along the channel and order dimensions, with each irreducible representation kept or removed as a complete block, thereby retaining SO(3) equivariance. Starting from a large checkpoint, the pruned model substantially reduces the inference cost while retaining higher accuracy than an independently trained small model. The pruned MACE-MP model outperforms the official from-scratch trained small model on 7 of 9 metrics on the Matbench Discovery leaderboard. In terms of efficiency, compressed MACE-MP and MACE-OFF models contain 1.5$\times$ to 4$\times$ fewer parameters and require 2.5$\times$ to 4$\times$ less pre-training compute than training a small model from scratch. For downstream applications, fine-tuning the pruned model reduces energy and force errors by 70.1\% and 34.4\% compared to training task-specific models from scratch across eight representative downstream datasets. We demonstrate that the method generalizes to other SO(3) equivariant architectures (SevenNet, eSCN) and can be combined with quantization and knowledge distillation for further gains.


Complexity-Aware LoRA Aggregation for Modality-Heterogeneous Federated Person Re-identification

Kunyang Lv ⋅ Wenke Huang ⋅ Wenwen He ⋅ Bin Yang ⋅ Bo Du ⋅ Mang Ye

A unified cross-modal person re-identification targets robust cross-modal retrieval with queries from diverse modalities like visible, infrared, and text.Such framework requires multi-modal data and centralized training, raising data costs and privacy risks. Hence, we propose a modality-heterogeneous federated ReID task to learn a unified cross-modal framework from clients with different cross-modal data.However, it faces two challenges.First, varying inter-modality complexity across tasks leads to different retrieval patterns and parameter space. Second, different intra-modality complexity results in imbalanced contributions when merging LoRA modules.Existing federated learning methods use uniform merging and transmit full models, which hinders cross-modal knowledge integration and increases communication overhead. To address these issues, we adopt Low-Rank Adaptation (LoRA), a lightweight module for fine-tuning large models through low-rank updates, enabling clients to send compact parameters.On top of LoRA, we develop MIRAGE (Modality-heterogeneous Intelligent federated ReID Aggregation with Global Efficiency), adapting aggregation to inter- and intra-modality complexity.When merging visible LoRA, MIRAGE selects key ranks to protect retrieval patterns while retaining shared ranks for common ability.As for infrared and text LoRA, it reduces conflicts with an adaptive pruning method. In both strategies, task complexity is incorporated into aggregation to balance contributions. Experiments on cross-modal retrieval tasks demonstrate the superiority of MIRAGE.The code and dataset will be publicly released.


Component-Based Out-of-Distribution Detection

Wenrui Liu ⋅ Hong Chang ⋅ RuiBing Hou ⋅ Shiguang Shan

Out-of-Distribution (OOD) detection requires sensitivity to subtle shifts without overreacting to natural In-Distribution (ID) diversity. However, from the viewpoint of detection granularity, global representation inevitably suppresses local OOD cues, while patch-based methods are unstable due to entangled spurious-correlation and noise. And neither of them is effective in detecting compositional OODs composed of valid ID components. Inspired by recognition-by-components theory, we present a post-hoc Component-Based OOD Detection (CoOD) framework that addresses the existing limitations by decomposing inputs into functional components. To instantiate CoOD, we derive Component Shift Score (CSS) to detect local appearance shifts, and Compositional Consistency Score (CCS) to identify cross-component compositional inconsistencies. Empirically, CoOD achieves consistent improvements on both coarse- and fine-grained OOD detection.


ComPose: When to Trust Hands for Object Pose Tracking

Jisu Shin ⋅ Junoh Lee ⋅ JunGyu Lee ⋅ Inhwan Bae ⋅ Dohyeon Lee ⋅ Hokyun Im ⋅ Youngwoon Lee ⋅ Hae-Gon Jeon

Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data or 3D templates, and remain highly vulnerable to severe occlusions by hand grasps despite the use of explicit masks. In this work, we present ComPose, a 6DoF object tracking framework designed for hand-aware object pose estimation from RGB video. Rather than treating the hand purely as an occluder, our method harmonizes hand motions as a complementary cue for object tracking. In detail, we recover a variety of object motions over time by combining object- and hand-derived cues from foundation models within a unified tracking pipeline. Here, ComPose adaptively selects informative hand joints, blends object- and hand-based rotation cues, and refines translation from the visible geometric evidence using a learned correction. We further enforce the temporal consistency over both rotation and translation, yielding stable 3D object trajectories over time without any external smoothing. Extensive experiments show that our method is accurate, efficient, and robust under severe hand occlusion and geometric ambiguity. In addition, the resulting trajectories can also effectively transfer to downstream robot manipulation by enabling robots to reconstruct human actions from online videos.


Compose Your Oracles: Off Policy Improvement with Aggregated Guidance

Jingtian Ji ⋅ Xuefeng Liu ⋅ Matthew Walter

Learning from multiple imperfect oracles is a practical way to improve reinforcement learning (RL) under limited interaction. However, most contemporary robust multi-oracle methods are on-policy and equire frequent fresh rollouts, which limits sample efficiency in domains where data collection is expensive. We propose Oracle-Aggregated Policy Improvement (OPI), an off-policy framework for policy improvement with multiple suboptimal oracles. OPI augments the oracle set with the learner, constructs an oracle-guided behavior policy through candidate-action pooling and source-aware critic scoring, and trains the learner using a generalized advantage-weighted behavior cloning objective combined with deterministic policy gradients. The method is motivated by a tractable proxy for the max-advantage objective and a policy-improvement relative to a strong oracle baseline. We evaluate OPI on diverse tasks across MetaWorld, the DeepMind Control Suite, and customized Box2D environments with heterogeneous oracles. Across sparse-reward manipulation, dense-reward locomotion, and customized control settings, OPI is strongest in sparse-reward and complementary-oracle regimes and remains competitive with representative baselines in dense-reward environments. Additional analyses show that oracle composition is most beneficial when oracle skills are complementary and that estimating oracle-specific action values is particularly helpful in sparse-reward regimes


Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models

Xin Xu ⋅ Clive Bai ⋅ Kai Yang ⋅ Tianhao Chen ⋅ Yuchen Cai ⋅ Yangkun Chen ⋅ Weijie Liu ⋅ Hao CHEN ⋅ Yang Wang ⋅ Saiyong Yang ⋅ Can Yang

Large-scale verifiable prompts underpin the success of Reinforcement Learning with Verifiable Rewards (RLVR), but they contain many uninformative examples and are costly to expand further. Recent studies focus on better exploiting limited training data by prioritizing hard prompts whose rollout pass rate is 0. However, easy prompts with a pass rate of 1 also become increasingly prevalent as training progresses, thereby reducing the effective data size. To mitigate this, we propose Composition-RL, a simple yet useful approach for better utilizing limited verifiable prompts targeting pass-rate-1 prompts. More specifically, Composition-RL automatically composes multiple problems into a new verifiable question and uses these compositional prompts for RL training. Extensive experiments across model sizes from 4B to 30B show that Composition-RL consistently improves reasoning capability over RL trained on the original dataset. Performance can be further boosted with a curriculum variant of Composition-RL that gradually increases compositional depth over training. Additionally, Composition-RL enables more effective cross-domain RL by composing prompts drawn from different domains. All our codes, compositional datasets, and models will be released to facilitate future research.


Concept-Based Mechanistic Interpretability Needs a Concrete Evaluation Paradigm

Yiming Tang ⋅ Qinglin Qi ⋅ LIN Zheng ⋅ Dianbo Liu

This position paper argues that concept-based mechanistic interpretability lacks a concrete evaluation paradigm and that the field must converge on one before continuing to scale methods whose effectiveness remains unverified. Sparse Dictionary Learning (SDL), encompassing sparse autoencoders, transcoders, crosscoders, and their variants, has become the dominant methodology for concept-based mechanistic interpretability, supported by a growing evaluation suite: reconstruction loss, automated interpretability scoring, human evaluation, feature absorption rates, sparse probing, and model editing benchmarks. Despite the significant empirical success SDL has achieved, we argue that none of these evaluation metrics directly measures what SDL is actually supposed to achieve: recovering the ground-truth features underlying observed representations. This insufficiency is structural, not incidental, and we support this perspective with both theoretical and empirical evidence. On the theoretical side, we build on the most recent formal analyses of SDL in mechanistic interpretability to show that the optimization landscape provably admits failure modes that existing metrics cannot detect. On the empirical side, we present three lines of evidence demonstrating that methods spanning a wide range of ground-truth recovery performance appear indistinguishable under current evaluation practice. Together, these results call for a new evaluation paradigm. We propose a set of principles that an ideal evaluation benchmark for concept-based mechanistic interpretability should satisfy, and argue for the community to develop and adopt more benchmarks grounded in these principles before continuing to scale methods whose effectiveness remains unverified.

Circuit extraction identifies model components that preserve a target behavior under ablation, yet it remains unclear which parts of the reported circuit—components, edges, or coarser summaries—remain stable across reasonable extraction choices. We study this question in a Lean-style, rule-generated tactic-prediction benchmark spanning atomic and compositional proof tasks. The proof-state rules and task structure are fixed by construction, and dense and weight-sparse variants of one transformer architecture are compared. In this controlled setting, we compare three ways to describe a circuit: a compact prediction-preserving circuit, a broader graph that keeps read, write, and routing components around that circuit, and a pruning graph whose size is set by a target post-ablation loss budget. We vary model sparsity, whether the query and key sides of attention are represented together or separately, the post-ablation loss budget for the pruning graph, and which supervised checkpoint initializes reinforcement learning (RL). When the same extraction method is used, dense and weight-sparse checkpoints overlap far more in the set of attention heads identified than in exact component-to-component edge lists. This head-level overlap is well above a random top-$k$ baseline; the exact-edge overlap is not consistently so. Tracking query and key sides separately, and varying the loss budget across the tested range, preserves the ordering of the RL initialization conditions when each graph is summarized by how many components it selects. Among these RL conditions, the largest gains on compositional tasks are accompanied by the largest fraction of compositional-task circuit nodes that lie outside the matched atomic-task circuits. These findings show that, even in this controlled setting, behavior preservation under ablation does not determine a unique circuit-level conclusion: the conclusion depends on which graph is reported, how it is extracted, and at which level of description the comparison is made.

A complete understanding of heterogeneous treatment effects requires fully capturing the conditional distribution of the counterfactual outcomes. To this end, we propose the Conditional Counterfactual Mean Embeddings (CCME), a framework that embeds conditional distributions of counterfactual outcomes into a reproducing kernel Hilbert space (RKHS). Under this framework, we develop a two-stage meta-estimator for CCME that accommodates any RKHS-valued regression in each stage. Based on this meta-estimator, we develop three practical CCME estimators: (1) Ridge Regression estimator, (2) Deep Feature estimator that parameterizes the feature map by a neural network, and (3) Neural-Kernel estimator that performs RKHS-valued regression, with the coefficients parameterized by a neural network. We provide finite-sample convergence rates for all estimators, establishing that they possess the double robustness property. Our experiments demonstrate that our estimators accurately recover distributional features including multimodal structure of conditional counterfactual distributions.

Retrieval-Augmented Generation (RAG) improves the factual accuracy of large language models (LLMs) by grounding responses in external evidence. However, when retrieved context conflicts with models' internal parametric knowledge, LLMs may still generate answers that contradict the provided evidence, posing a key challenge to contextual faithfulness and the reliability of RAG systems. Motivated by recent mechanistic findings on the distinct propagation and progressive accumulation of parametric and contextual signals in LLMs, we propose Conflict-Suppressed RAG (CSRAG), a simple, training-free, decoding-time framework for resolving knowledge conflicts. CSRAG biases generation toward retrieved evidence by suppressing tokens associated with parametric knowledge while boosting tokens from the context via two complementary logits processors. Experiments on six challenging faithfulness benchmarks demonstrate that CSRAG consistently achieves state-of-the-art or near state-of-the-art performance across multiple backbone LLMs, while remaining fully training-free and lightweight.

Deep-learning protein structure predictors achieve near-experimental accuracy on individual folds, yet their default inference samples concentrate around a single dominant conformation. We introduce ConforFlux, an inference-time procedure for Boltz-2 that couples $M$ structure-prediction trajectories through a pairwise C$\alpha$-RMSD repulsion gradient on the trunk's single and pair embeddings. Because the trunk conditions every block of the diffusion module, this update propagates to every subsequent denoising step. On four conformational-change categories, ConforFlux improves the per-state success rate over Default Boltz-2 by 3–17 percentage points while preserving Default-comparable physical quality. On twelve transporter pairs with at least one alternate-state reference released after the Boltz-2 cutoff, the $2$ Å success rate rises from $4/12$ to $9/12$. On the human dopamine transporter, ConforFlux additionally reaches the post-cutoff inward state, which $500$ Default samples never attain.


Conformal Cache: Reliable Proxy-Discrepancy Caching for Fast Generative Inference

Wenxiao Wu ⋅ Yachen Gao ⋅ Yukang Feng ⋅ Chengming Xu ⋅ Moran Li ⋅ jingyu li ⋅ Ming Xie ⋅ Xiaobin Hu ⋅ Xinwei Sun ⋅ Jing-Hao Xue ⋅ Nong Sang ⋅ Yanwei Fu

While diffusion and flow-based generative models have achieved impressive performance in high-fidelity image and video synthesis, this capability comes at the cost of substantial inference overhead during iterative sampling. Cache-based acceleration alleviates this cost by reusing intermediate computations across adjacent denoising or transport steps, typically relying on lightweight proxy discrepancies to decide when cached outputs should be refreshed. However, such proxy-discrepancy rules are inherently unreliable: the proxy is only a point estimate of the inaccessible oracle discrepancy and can misalign across prompts, timesteps, and generation dynamics. This mismatch induces asymmetric failures: proxy underestimation may cause unsafe reuse of stale cached computations and degrade generation quality, while proxy overestimation leads to overly conservative refreshes and reduced acceleration. To address these issues, this paper introduces Conformal Cache, dubbed CCache, a training-free conformal calibration framework for reliable cache-based generation acceleration. Our key insight is to formulate cache reuse as a one-sided risk-control problem and adaptively calibrate a timestep-wise upper correction for the signed proxy-to-oracle gap based on split conformal prediction. CCache is plug-and-play and can be integrated with different proxy-discrepancy caching baselines without structural modifications or parameter updates to the generative model. Extensive experiments on image and video generation demonstrate that \methodname improves generation quality and cache reliability while preserving favorable speed-quality trade-offs.


Constrained Factorization with Diagonal Scaling: Rank-Revealing Training and Pruning

Yikun Hou ⋅ Emrullah Akbas ⋅ Suvrit Sra ⋅ Alp Yurtsever

Gradient descent for matrix factorization exhibits an implicit bias toward approximately low-rank solutions, even in regimes where the iterates may grow unbounded. We study this behavior through a constrained factorization with diagonal scaling, separating bounded outer factors from explicit diagonal scale parameters. For positive semidefinite matrix recovery, we use the model $X \approx UDU^\top$, where $U$ is constrained to a Frobenius norm ball and $D$ is a nonnegative diagonal factor. Although this reparameterization preserves the stationary points of classical Burer-Monteiro factorization, its projected dynamics consistently recover truly (rather than approximately) low-rank solutions across a wide range of step sizes and initializations. Motivated by this behavior, we extend the construction to neural networks by inserting diagonal layers between Frobenius-norm-constrained outer layers. The resulting UDV models exhibit strong rank-revealing structure during training, with many diagonal components and associated columns collapsing toward zero and becoming naturally prunable. We exploit this structure through an SVD-based pruning procedure (UDVPruning) that identifies and removes redundant units during training. Across ViT and MLP-Mixer benchmarks, UDVPruning substantially reduces parameters and FLOPs while maintaining competitive test accuracy and often improving measured training time, inference latency, and energy consumption.


Constrained Stochastic Spectral Preconditioning Converges for Nonconvex Objectives

Konstantinos Oikonomidis ⋅ Jan Quan ⋅ Kimon Antonakopoulos ⋅ Antonio Silveti-Falls ⋅ Volkan Cevher ⋅ Panagiotis Patrinos

In this work, we develop proximal preconditioned gradient methods with a focus on spectral gradient methods providing a proximal extension to the Muon and Scion optimizers. We introduce a family of stochastic algorithms that can handle a wide variety of convex and nonconvex constraints and study its convergence under heavy-tailed noise, through a novel analysis tailored to the geometry of the proposed methods. We further propose a variance-reduced version, which achieves faster convergence under standard noise assumptions. Finally, we show that the polynomial iterations used in Muon are more accurately captured by a nonlinear preconditioner than by the ideal matrix sign, leading to a convergence analysis that more faithfully reflects practical implementations.


Contact Geometry for Generative Models: An Unbalanced Optimal Transport Formulation

Andrea Testa ⋅ Søren Hauberg ⋅ Andras Kupcsik ⋅ Tamim Asfour ⋅ Leonel Rozo

Real-world generative problems in biology, medical imaging, or robotics, rarely come with perfectly paired source and target distributions. While unbalanced Optimal Transport (uOT) provides a robust framework to bridge unpaired datasets via minimum-length paths in probability space, existing formulations based on the Wasserstein–Fisher–Rao (WFR) geometry introduce a structural bias toward high density regions, often triggering mode collapse. We introduce Contact unbalanced Optimal Transport (CuOT), a novel uOT formulation grounded in contact geometry that decouples density transport from mass growth to eliminate structural bias. This allows CuOT to adapt sampling steps to the local data structure, naturally reducing step sizes in high-density regions to prevent overshooting and ensuring stable convergence. The result is a stable, scalable path-length minimization solver that consistently outperforms WFR-based methods on biological dynamics reconstruction, unpaired image-to-image translation, and video generation.


Context-Aware Autoregressive Image Generation for Emerging Reasoning Properties

Jixuan Ying ⋅ Haoyu Liu ⋅ Timing Yang ⋅ Tingyu Zhu ⋅ Predrag Neskovic ⋅ Anand Bhattad ⋅ Cihang Xie ⋅ Zeyu Zheng ⋅ Alan Yuille ⋅ Feng Wang

Autoregressive image generation has recently emerged as a competitive visual generation paradigm, yet its generation order is often predefined or heuristic, such as next-patch prediction, random-order prediction, or one-step generation for specific scales. We observe that these strategies can cause models to overlook important structural dependencies in images, thereby limiting both structural robustness and reasoning capability. In this work, we propose a context-aware autoregressive generation framework, where the model dynamically determines the next token location based on global context and current local information. Rather than following a fixed or random order, our method allows the generation process to adapt to image structure and semantic dependencies. Extensive experiments show that context-aware generation significantly improves structural robustness across different image AR paradigms, including masked autoregressive and scale-wise autoregressive models. Beyond quantitative improvements, we further observe emerging reasoning properties: the model can better understand complex semantic instructions and generate images that more faithfully follow the given guidance. These findings suggest that generation order is a critical yet underexplored factor in image autoregressive modeling, and that context-aware decoding can unlock stronger reasoning-aware visual generation capabilities.

Multimodal models deployed in real-world settings often suffer from missing data modalities due to acquisition costs, privacy constraints, or sensor failures, leading to severe performance degradation. Existing approaches based on shared representations or expert routing struggle when modality-specific information is absent. A key challenge is that, when a modality is unobserved, the target representation is not uniquely identifiable, making naive latent prediction prone to degenerate or collapsed solutions. We propose Cross-modal Embedding Prediction and Alignment (CEPA), a framework that addresses missing modalities through task-driven latent representation imputation rather than raw input synthesis. CEPA employs masked representation learning with data-driven masking patterns and a context-conditional distribution alignment objective that stabilizes latent prediction and prevents representation collapse. We evaluate CEPA on the MIMIC multimodal benchmark spanning EHR time-series, chest X-rays, and clinical notes, across three clinical prediction tasks under both controlled (MCAR) and naturally occurring (MNAR) modality absence. CEPA consistently outperforms prior missing-modality methods, with ablations confirming the contribution of adaptive masking and context-conditional alignment.


Continually Evolving Skill Knowledge in Vision Language Action Model

Yuxuan Wu ⋅ Guangming Wang ⋅ Zhiheng Yang ⋅ Tianchen Deng ⋅ Maoqing Yao ⋅ Brian Sheil ⋅ Hesheng Wang

Vision-language-action (VLA) models show promising knowledge accumulation ability from pretraining, yet continual learning in VLA remains challenging, especially for efficient adaptation. Existing continual imitation learning (CIL) methods often rely on additional parameters or external modules, limiting scalability for large VLA models. We propose Stellar VLA, a knowledge-driven CIL framework without increasing network parameters. Two progressively extended variants are designed: T-Stellar for flat task-centric modeling and TS-Stellar for hierarchical task–skill structure. Stellar VLA enables self-evolving knowledge learning by jointly optimizing task representations and a learned knowledge space. We propose a knowledge-guided expert routing mechanism conditioned on knowledge relation and Top-$K$ semantic embeddings, enabling task specialization without increasing model size. Experiments on the LIBERO benchmark show that Stellar VLAs achieve strong performance among both VLA and CIL baselines, using only 1\% data replay. Real-world evaluation on a dual-arm platform with distinct embodiment and scene configurations validates effective knowledge transfer. TS-Stellar excels in hierarchical manipulation, and visualizations reveal robust knowledge retention and task discovery. Project website is provided in the supplementary material.


Contrastive Retrieval Heads for Improved Attention-Based Reranking

Linh Tran ⋅ Yulong Li ⋅ Radu Florian ⋅ Stacy Patterson ⋅ Ana Milanova ⋅ Wei Sun

The strong zero-shot and long-context capabilities of recent Large Language Models (LLMs) have enabled highly effective list-wise reranking systems. Attention-based rerankers leverage transformer attention patterns as retrieval signals, but not all attention heads contribute equally: many introduce noise and redundancy, limiting both reranking quality and efficiency. In this work, we introduce CoRe heads, a small subset of retrieval-specialized heads identified through a contrastive scoring objective that rewards attention to relevant documents while penalizing attention to distractors. This contrastive criterion isolates highly discriminative retrieval heads and yields a strong training-free list-wise reranker. Our analysis reveals that retrieval-relevant computation in transformer rerankers is highly sparse and structurally localized: fewer than 1% of attention heads account for most reranking performance across models and datasets, and these heads consistently concentrate in middle transformer layers. This localization enables aggressive layer pruning for efficient reranking: removing the final 50% of transformer layers preserves nearly identical reranking accuracy while substantially reducing inference latency and GPU memory usage.


Contribution-Aware Structured Sparsity for Model Merging

Yan Li ⋅ GUIPING CAO ⋅ Meng XU ⋅ Tao Jiang ⋅ Yaguang Song ⋅ Ming Tao ⋅ Yaowei Wang ⋅ Dongmei Jiang

Model merging integrates task-specific fine-tuned models into a single multi-task model, but often suffers from parameter interference caused by conflicting task-vector updates. Existing methods typically mitigate conflicts by pruning task vectors based on weight magnitude or random heuristics, treating Transformers as unstructured ``bags of parameters'' and overlooking their inherent modularity. In this paper, we propose \textbf{C}ontribution-\textbf{A}ware \textbf{S}tructured \textbf{S}parsity (CASS), a unified framework that reduces parameter interference by identifying and preserving task-specific components. At the core of CASS is a contribution-aware structured mask that identifies task-relevant attention heads and FFN neurons. We instantiate this mask in two settings: CASS-Merging, the primary post-hoc setting where masks serve as a plug-and-play denoising filter for existing merging operators, and CASS-Tuning, an extension for scenarios with fine-tuning access where masks constrain gradients to reduce structural overlap between task vectors. Our analysis shows that task-relevant components are sparse and partially disjoint, supporting structured component-level filtering as an effective way to reduce merging interference. Extensive experiments across vision (ViT, 20 tasks) and language (RoBERTa, 8 tasks; Qwen2.5, 4 tasks) benchmarks demonstrate that CASS improves a range of representative merging baselines.


Control-Augmented Autoregressive Diffusion for Data Assimilation

Prakhar Srivastava ⋅ Farrin Marouf Sofian ⋅ Francesco Immorlano ⋅ Kushagra Pandey ⋅ Stephan Mandt

Despite advances in test-time scaling and diffusion finetuning, guidance for Auto-Regressive Diffusion Models (ARDMs) remains underexplored. We introduce an amortized framework that augments a pretrained ARDM with an offline-trained controller. By previewing future rollouts, the controller learns stepwise corrections that anticipate observations under a terminal-cost objective, yielding a reusable policy for guided generation. Motivated by a stochastic optimal control view of ARDM trajectories, our method injects small controls within each denoising sub-step while staying close to the pretrained dynamics. We study this approach for data assimilation (DA) in chaotic spatiotemporal partial differential equations (PDEs), where existing methods are often computationally expensive and susceptible to forecast drift under sparse observations. At inference, DA becomes a feed-forward rollout with on-the-fly corrections, achieving an order-of-magnitude speedup over strong diffusion-based baselines. Across two canonical PDEs and a compact ECMWF Reanalysis v5 (ERA5) pilot spanning six observation regimes, our method consistently improves stability and accuracy over state-of-the-art alternatives, with similar improvements observed in a larger-scale GenCast study.


Co-PiLOT: Constrained Physics-Informed Latent Optimization for Target-Driven Inverse Design

Mahish Kumar Guru ⋅ Mayank Nagar ⋅ Ayush vyas ⋅ Jan Bohlen ⋅ Roland Aydin ⋅ Noomane B Khalifa

Inverse design of physical systems (molecules, devices, microstructures) often reduces to optimizing a high-dimensional structure against an expensive black-box simulator. Direct search is difficult because the space is non-Euclidean, feasibility is hard to encode, and each evaluation is expensive. We present Co-PiLOT, a latent optimization approach that maps candidates through a generative encoder--decoder, uses the decoder as a learned validity prior, and searches the latent space with physics-informed black-box optimization. The framework is applied on the inverse design of magnesium alloy microstructure/texture. We develop a vision transformer based--encoder; paired with latent diffusion, diffusion transformer and rectified-flow transformer--based decoders on $\sim80{,}000$ EBSD-derived microstructure dataset to learn a minimal bottleneck, $z$. The ViT-FMDiT model ($z$=$768$) reconstructs high-fidelity microstructure images (FID $23.19$, MS-SSIM $0.178$), which our self-segmenting orientation codec converts into input grids for crystal plasticity solver. Finally, we introduce Meridian, an active latent optimizer driven by deep-kernel Gaussian-process uncertainty, feasibility prediction, active trust regions, and target-aware acquisition. Within the evaluation budget, the ViT-FMDiT and Meridian combination yields the best target-driven objective score, outperforming DANTE, TuRBO, and BAxUS by $\sim6$\% in relative error on the same decoder.


CORAL: A Benchmark for Structure-aware and Brain-wide Neuron Reconstruction in Light Microscopy

Zekang Yang ⋅ Jiamin Li ⋅ Zhenghua Li ⋅ Jiaqi Fan ⋅ Zengcai V Guo ⋅ Xiaolin Hu

Automatic neuron reconstruction from light microscopy images is a central problem in computational neuroanatomy. While recent methods have achieved encouraging results on local image blocks, it remains unclear whether such progress translates to reconstruction that is both structurally accurate and scalable to the whole-brain scale. We present CORAL, the first benchmark for structure-aware evaluation of automatic neuron reconstruction from light microscopy images at both local and whole-brain scales. Built on a high-quality whole-brain fMOST dataset with carefully curated annotations, CORAL establishes two progressive tasks: block-level reconstruction, which evaluates reconstruction methods under limited spatial context, and brain-wide reconstruction, which assesses complete neuron reconstruction at the whole-brain scale. To account for topological correctness beyond geometric distance similarity, we introduce a structure-aware metric based on fiber prediction. To further achieve complete neuron reconstruction across the entire brain, we develop a brain-wide neuron tracing framework that extends arbitrary local reconstruction methods to the whole-brain scale through an iterative local-to-global process. Using this benchmark, we provide the first structure-aware comparison of mainstream methods for local neuron reconstruction and further evaluate their performance in brain-wide reconstruction. Our results underscore the importance of structure-aware evaluation and the need for more robust methods for complete neuron reconstruction. All datasets and code are available at https://github.com/yang-ze-kang/CORAL.


CORAL: Learning Amyloid Fibril Ligand Docking with Cooperative Binding Rewards

Yasheng SUN ⋅ Bohan Li ⋅ Youqi Tao ⋅ Jürgen Schmidhuber

A hallmark of neurodegenerative diseases such as Alzheimer's and Parkinson's is the aberrant aggregation of proteins into amyloid fibrils, and small molecules that selectively bind to these fibrils hold promise as diagnostics, imaging probes, and therapeutics. Predicting how such ligands bind to fibril targets, however, presents two fundamental challenges. First, resolved co-crystal structures of amyloid–ligand complexes are exceptionally scarce; even with recent advances in cryo-EM only a handful have been structurally characterized, making supervised training of docking models impractical for this target class. Second, amyloid fibrils present a binding mode fundamentally different from globular proteins: ligands intercalate into longitudinal cross-β grooves and stack cooperatively along the fibril axis, a geometry that existing docking models are not designed to capture. To address these challenges, we present CORAL (Cooperative Amyloid Ligand docking), a reinforcement learning framework that trains a generative docking model to produce ligand pose distributions tailored to the cross-β groove geometry. Our reward explicitly incorporates cooperative ligand–ligand stacking energy alongside protein–ligand docking affinity, directly capturing the distinctive binding geometry of amyloid fibrils. We further introduce a curated evaluation set of amyloid–ligand complexes constructed from model-generated poses validated by domain experts. Experiments on both experimentally resolved structures and this evaluation set demonstrate improved pose quality and binding affinity correlation over existing docking baselines.


CoRE-RL: Co-evolving Reasoning Trajectories and Evidence Subgraphs with Learned Evidence Projection

Zhengquan luo ⋅ Guy Tadmor ⋅ Peilin Zhao ⋅ David Zeevi ⋅ Zhiqiang Xu

Large language models have made knowledge-graph question answering more flexible, yet complex multi-hop questions often require evidence that becomes identifiable only during reasoning. Starting from the question alone, narrow retrieval can miss bridge facts, while broad expansion injects distractors and worsens long-context interference. The key bottleneck is dynamic evidence identification: the system must revise a compact evidence view as reasoning exposes unsupported facts. We formulate this process as partially observable control over evidence and propose CoRE, a residual-coupled framework in which reasoning trajectories and evidence subgraphs evolve together. CoRE decomposes each trace into answer-critical claims, treats unsupported claims as residual signals, and uses them to guide targeted retrieval and budgeted subgraph projection. The same formulation exposes a learnable projection decision: after residuals retrieve candidate facts, CoRE-RL learns from rollout feedback which facts should remain visible for the next reasoning pass. Across WebQuestionsSP and Complex WebQuestions, the CoRE family improves answer accuracy, evidence sufficiency, and budget robustness, with CoRE-RL providing further gains through more effective retained-subgraph selection.

Human labels are important supervisory signals for preference-based reinforcement learning, but their collection is often constrained by practical considerations that standard training objectives tend to ignore. For instance, Direct Preference Optimization (DPO) usually treats observed preference labels as clean, even though real annotation pipelines rely on heterogeneous annotators whose reliability varies across individuals and with the difficulty of each comparison. This setting introduces three challenges: (1) Noisy labels introduce systematic bias into the preference optimization objective and need to be rigorously corrected. (2) The allocation of pairwise comparisons to different annotators is a design problem that directly affects the statistical efficiency of the training pipeline. (3) Non-random routing creates selection bias in annotator evaluation, complicating downstream decisions such as compensation. We address these challenges within a unified framework for DPO that combines annotator-aware posterior correction with information-guided annotator routing. Posterior correction leverages an instance-dependent noise model to infer latent ground-truth preferences from noisy labels, mitigating bias in the DPO objective. Annotator routing is formulated as an optimal experimental design problem that allocates comparisons to annotators according to their expected statistical information under a fixed annotation budget. To support reliable annotator evaluation under non-random routing, we introduce a doubly robust estimator that corrects for selection bias. Semi-synthetic experiments on the UltraFeedback dataset show that our approach improves ground-truth preference recovery and downstream AlpacaEval 2 performance over baselines, with consistent gains across multiple open-source LLM model architectures. Annotator reliability-estimation experiments further show that the doubly robust estimator substantially reduces mean squared error relative to other baselines, demonstrating its value for annotator evaluation.

Correspondence pruning—separating sparse inliers from severe outliers in noisy feature matches—is a long-standing bottleneck for two-view geometry estimation in 3D vision. Existing learning-based pruners pursue this goal through ever more sophisticated context operators—local graphs, dynamic neighborhoods, or global attention—yet a controlled comparison reveals a striking phenomenon: when we freeze the structural assignment of these operators after a single forward pass, accuracy collapses by over $30\%$ mAP regardless of operator choice. This suggests that the limiting factor of current methods is not which operator is used, but the implicit assumption that structure can be committed in one shot. We argue that correspondence pruning should instead be cast as \emph{iterative structural rectification} (ISR), in which the geometric structure and the feature representation must \emph{co-evolve} layer by layer. Building on this view, we propose \textbf{ISR-Net}, a minimal architecture that pairs a layer-wise re-derived local-context operator (a dynamic $k$-NN graph in feature space) with a layer-wise re-derived global-context operator (soft assignment to learnable cluster prototypes), so that any gain over prior arts is attributable to the rectification \emph{schedule} rather than to operator novelty. Despite using deliberately off-the-shelf operators, ISR-Net attains state-of-the-art results on YFCC100M, SUN3D, and HPatches, surpassing the strongest attention-based competitor by $+8.2\%$ mAP@$5^\circ$ while running $35\%$ faster, indicating that \emph{how often structure is refined} matters more than the operator's complexity.


Counting without Scale: Scale-Consistent Error Correction for Crowd Counting

Yi-Kuan Hsieh ⋅ Xin Li ⋅ Chi-Chia Sun ⋅ Yu-Chee Tseng ⋅ Ming-Ching Chang ⋅ Kuan-Chuan Peng ⋅ Jun-Wei Hsieh

Crowd counting is a vital computer vision task with wide-ranging applications, from public surveillance to urban planning. A persistent challenge is achieving scale invariance despite noisy labels, as humans often make counting errors with similar objects. Most existing methods struggle to maintain consistency and robustness when testing images differ in scale from the training data. We propose ScaLe-aware Invariance and Correcting Errors (SLICE) for robust crowd counting, a novel scale-aware loss function that improves scale invariance, enabling consistent and robust predictions under varying scales and annotation noise. The loss function combines two components: a scale-invariance term that applies multiscale supervision to enforce consistency across density maps of the same scene, and an error correction term that addresses inevitable annotation noise by modeling label uncertainty with probabilistic methods such as a mixture of Gaussians. SLICE is model-agnostic and integrates easily into existing crowd counting architectures without extra computational overhead during inference. Extensive experiments on five benchmark datasets show that SLICE greatly improves scale invariance and cross-dataset generalization, achieving the state-of-the-art results in challenging scenarios. Using CrowdDiff, Steerer, and CrowdFormer as baselines, SLICE lowers mean absolute error (MAE) by 5–7\% by addressing annotation noise and scale variation. Applying the scale-invariance term across multiple scales further cuts MAE by up to 15\%.


CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation

Seonghyun Jin ⋅ youngmin Kim ⋅ Sunwoo Park ⋅ Jong Chul Ye

Camera-conditioned video generation requires positional encoding that remains reliable under changes in camera motion, lens configuration, and scene structure. However, existing attention-level camera encodings either provide ray-only camera signals or rely on pinhole camera geometry, limiting their applicability to general camera control under the Unified Camera Model, including wide-angle and fisheye lenses. To address this limitation, we propose Curved Ray Expectation Positional Encoding (CRePE). CRePE represents each image token as a depth-aware positional distribution along its source ray, providing a Unified Camera Model-compatible positional encoding that naturally captures the curved-ray geometry of wide-angle and fisheye cameras. CRePE is implemented through a Geometric Attention Adapter added to frozen video DiTs, injecting token-wise scene-distance information into proper attention layers and stabilizing it with pseudo supervision from a monocular geometry foundation model. This design leads to more stable camera control and improved video quality in camera-conditioned video generation under the Unified Camera Model. Controlled positional-encoding ablations show consistent improvements over existing multi-view positional encoding, demonstrating the effectiveness of curved-ray-aware positional encoding across diverse camera models. Furthermore, by extending the same positional-encoding pathway to external geometry control through Radial MixForcing, CRePE supports external radial-map control for scene-geometry-conditioned generation and source-video motion transfer beyond camera control.


CRESiST: Self-Enhancing Exploritive Policy Learning for Simultaneous Speech Translation

Chenyang Le ⋅ Pei Zhang ⋅ Yuxiang Chen ⋅ Xi Chen ⋅ Bing Han ⋅ Derek Wong ⋅ Baosong Yang ⋅ Yanmin Qian

Simultaneous Speech Translation (SimulST) requires balancing translation quality and latency. While Large Language Models (LLMs) excel in offline translation, current LLM-based SimulST systems predominantly rely on prefix-incremental generation. This paradigm suffers from redundant Key-Value cache recomputation and necessitates external Voice Activity Detection to process unbounded streams. To overcome these bottlenecks, we propose an end-to-end interleaved read-write framework built upon an omni-modal LLM. By interleaving speech chunks and text tokens, our architecture achieves full KV cache reuse, natively supporting segmentation-free, infinite-length inference. Furthermore, to learn optimal streaming policies without rigid offline alignments, we introduce a novel Actor-Critic self-enhancement pipeline. An RNN-T Actor dynamically samples read-write trajectories, while an LLM-based Critic evaluates these paths via intrinsic semantic and latency rewards, distilling the optimal policy back to the Actor. Extensive experiments on sentence-level (FLEURS, CoVoST2) and document-level (MuST-C, ACL 60/60) benchmarks demonstrate that our model establishes new state-of-the-art quality-latency Pareto frontiers, substantially outperforming existing cascaded and static-alignment baselines.

Large language models from different families use different hidden dimensions, tokenizers, and training procedures, making behavioral directions difficult to compare or transfer across models. We introduce an anchor-projection framework that maps hidden representations from each model into a shared anchor coordinate space (ACS). Behavioral directions extracted from source models are projected into ACS and averaged into a canonical direction. For a new model, the canonical direction is reconstructed into its native hidden space using only anchor activations, without fine-tuning or target-specific direction extraction. We evaluate five instruction-tuned model families and ten behavioral axes. We find that same-axis directions align tightly across the Llama-Qwen-Mistral-Phi cluster on the ACS. This shared structure transfers to downstream tasks: held-out targets achieve \(0.83\) ten-way detection accuracy and \(0.95\) mean binary AUROC, while canonical steering induces refusal-rate shifts of up to \($\Delta$ = +0.46\) under distribution shift. Sensitivity analyses show that two source models and small anchor pools already suffice to approximate transferable directions. Overall, ACS provides a novel perspective on cross-family interpretability, revealing that representation-level transfer remains robust across model families. The code for this paper is available at \url{https://anonymous.4open.science/r/ACS-annoymous/}.

Modern machine learning progresses through empirical work, benchmarking new methods to evaluate relative performance. However, the statistical variability inherent to evaluation -- exacerbated by the stochastic nature of many algorithms -- often makes performance estimation unreliable due to the limited test samples available, leading to a validation crisis in which genuine advances are difficult to discern. In this work, we show that cross-validation improves markedly confidence when evaluating and comparing learning algorithm performances. We introduce the concept of \emph{sample gain}, which quantifies the virtual data augmentation achieved by using multiple cross-validation splits to reduce benchmarking variance. Experiments on both synthetic and real-world datasets (histopathologic scans and NLP fine-tuning) demonstrate that multiple splits can substantially improve the reliability and stability of performance estimates, with diminishing returns often setting in later than expected. We also introduce a procedure to dynamically early-stop cross-validation by estimating from the first few folds if subsequent folds will bring large sample gains. Our findings highlight the value of pushing cross-validation on available samples to achieve robust and reliable benchmarking.


Cross-Model Transfer Attacks against Large Vision-Language Models via Model Diversity Enrichment and Stochastic Parameter Sampling

Zhenze Yang ⋅ Xiaowen Cai ⋅ Junhao Dong ⋅ Guangke Chen ⋅ Guiyao Tie ⋅ Daizong Liu ⋅ Hongyang He ⋅ Zhiyuan Ma ⋅ Xiang Fang ⋅ Dengpan Ye ⋅ Jing Zhang

Large Vision-Language Models (LVLMs) have achieved remarkable multimodal reasoning capabilities but remain critically vulnerable to adversarial perturbations. In practical black-box scenarios, severe architectural discrepancies, misaligned feature spaces, and divergent multimodal fusion strategies often undermine adversarial transferability, preventing perturbations crafted on surrogate models from generalizing to unseen target models. Existing ensemble-based LVLM attacks attempt to mitigate this by integrating multiple surrogates, yet they often fail to bridge the structural generalization gap across heterogeneous LVLM families. In this work, we propose a novel framework that enhances cross-model adversarial transferability via model diversity enrichment and stochastic parameter sampling. We show that transfer success is fundamentally governed by cross-model gradient variance. To mitigate this, we introduce (1) a diversity-induced continuous surrogate manifold that captures cross-model decision variability, and (2) stochastic parameter sampling during optimization to approximate the gradient distribution of unseen LVLM models. We provide a theoretical analysis bounding the expected adversarial transfer gain by gradient variance and prove that stochastic sampling reduces adversarial overfitting to surrogate decision boundaries. Extensive experiments across multiple state-of-the-art LVLMs demonstrate substantial improvements in black-box transferability under strict evaluation protocols.


Cross-Question Reliable Reinforcement Learning

Hector G. Rodriguez ⋅ Marcus Rohrbach

AI systems operating near the limits of their capabilities can maximize utility while minimizing the risk of error by abstaining or deferring questions they are not confident about to a human or a stronger model. Recently, large language models (LLMs) have been post-trained to estimate confidence in their own answers. To achieve this, methods reinforce confidence scores based on the correctness signal to each question-answer pair \emph{individually}, with rewards that encourage assigning confidence 1 to correct answers and 0 to incorrect answers. However, estimating confidence scores is useful in that it enables selecting only the most confident answers, to maximize the share of solved questions while accommodating the user's risk tolerance. Assigning confidence 1 to correct answers and 0 to incorrect ones is theoretically optimal, but unattainable in practice and detrimental for selective prediction with a non-zero risk tolerance: confidences pushed to 0 and 1 carry no information about how answers compare to one another. We therefore introduce \ours (Cross-question Abstention Reward Estimation), which computes rewards over a set of question-answer pairs: At each of several risk tolerances, the confidences across the set induce an abstention threshold, and each confidence is rewarded for lying on the correct side of it. Since the threshold itself comes from the set, every verbalized confidence is reinforced relative to other questions' confidence, not against its own binary correctness label. On math, \ours increases the coverage at 5\% risk (C@5) across MATH-500, GSM8K, and Big-Math-Digits from 0.8\% to 42.4\%, compared to existing pointwise correctness-matching rewards. When trained on multi-hop Wikipedia QA from HotpotQA, \ours improves C@5 in out-of-distribution general-knowledge, reasoning, and math benchmarks (including CommonsenseQA, GPQA, SimpleQA, and TriviaQA) from 9.9\% to 23.6\%.


CSBench: A Comprehensive Benchmark for Evaluating Project-Level System Construction in Computer Science

Hongli Yu ⋅ Huan-ang Gao ⋅ Botian Wang ⋅ Hanlin Wu ⋅ YiFei Wang ⋅ Daoguang Zan ⋅ Huang Anni ⋅ Weinan Dai ⋅ Qiying Yu ⋅ Jiaze Chen ⋅ Wei-Ying Ma ⋅ Ya-Qin Zhang ⋅ Jingjing Liu ⋅ Mingxuan Wang ⋅ Hao Zhou

While LLM-based coding agents have advanced to repository-scale engineering, existing benchmarks predominantly focus on software maintenance, overlooking the foundational capability of system construction. We introduce CSBench, a benchmark evaluating project-level Spec-to-Code implementation across core computer science domains (e.g., OS, compilers, databases). CSBench features 100 expert-curated tasks from top-tier university assignments within containerized environments. Evaluating 12 state-of-the-art LLMs reveals a stark performance disparity: even top-tier models like GPT-5.2 see success rates plummet from over 90% in algorithmic tasks to below 35% in system-level domains. Our analysis identifies a fundamental deficiency in systemic reasoning and construction knowledge as the primary bottleneck, where agents either terminate prematurely due to overconfidence or fail to translate execution feedback into valid architectural fixes, highlighting CSBench as a necessary tool to bridge this gap and advance autonomous agents toward deep computer science mastery.


DAPS: Dependency-Aware Premise Selection for LLM Theorem Proving

Zixuan Chen ⋅ Wenyuan Jiang ⋅ Junling Wang ⋅ Mrinmaya Sachan ⋅ Yinya Huang

LLM-based theorem proving in Lean 4 has advanced rapidly, but provers are difficult to leverage newly contributed lemmas, because retrieval skews toward popular foundational lemmas, leaving the right premises buried among hundreds of thousands of candidates. Existing neural selectors rank candidates by semantic similarity, with three intrinsic limitations: frequency-skewed retrieval, isolated pairwise ranking, and degradation on out-of-distribution queries such as competition problems and informal language. Given that mathematics is structured by dependencies rather than similarities, we restore this signal to both the training data and the model. We curate two underused dependency resources. We conduct the first systematic mining of 53 Lean Blueprint projects, a previously untapped corpus of 3,805 manually curated informal+formal nodes that capture rare, research-level premises absent from Mathlib4. We further pair it with the extraction of comprehensive, typed, multi-level Mathlib4 dependencies. We then introduce DAPS, a dependency-aware premise selector with a structural neighborhood encoder, a group-level contrastive objective, and a Mathlib-then-Blueprint adaptation procedure. DAPS reaches Recall@32 of 89.31 on Mathlib4-Heldout (+11.47 over the strong selector LeanHammer). DAPS holds its lead across both out-of-distribution benchmarks, and improves three general-purpose LLMs on the informal IMOProofBench. Dependency graphs, benchmarks, and code are available 1 and will be released under a permissive license upon publication.

Subgraph mining aims to discover meaningful subgraphs from a large graph given a seed node, with applications in community detection, functional module identification, anti-money laundering, and knowledge discovery. Existing methods typically adopt static approaches that cannot dynamically and intelligently adapt to diverse graph structures. Moreover, training models for subgraph mining faces inherent challenges including exposure bias and the difficulty of learning when to stop expansion. To address these limitations, we propose DASM (Dynamic Autoregressive Subgraph Mining), a novel two-stage framework that models subgraph mining as a sequential decision process. In the first stage, we employ step-wise supervised learning with multi-label objectives and asymmetric F1 loss to teach the model how to select correct next nodes. In the second stage, we apply Group Relative Policy Optimization (GRPO) to align the model with end-to-end mining quality, enabling it to learn optimal stopping decisions through trajectory-level feedback. Our dynamic autoregressive approach allows the model to make state-aware decisions at each step, automatically determining subgraph boundaries without predefined size constraints. Experiments on multiple datasets demonstrate that DASM significantly outperforms existing baselines in terms of F1 score and IoU.


DECEIVE-AFC: Adversarial Claim Attacks against Search-Enabled LLM-based Fact-Checking Systems

Haoran Ou ⋅ Kangjie Chen ⋅ Gelei Deng ⋅ Hangcheng Liu ⋅ Jie Zhang ⋅ Tianwei Zhang ⋅ Kwok-Yan Lam

Fact-checking systems with search-enabled large language models (LLMs) have shown strong potential for verifying claims by dynamically retrieving external evidence. However, the robustness of such systems against adversarial attack remains insufficiently understood. In this work, we study adversarial claim attacks against search-enabled LLM-based fact-checking systems under a realistic input-only threat model. We propose DECEIVE-AFC, an agent-based adversarial attack framework that integrates novel claim-level attack strategies and adversarial claim validity evaluation principles. DECEIVE-AFC systematically explores adversarial attack trajectories that disrupt search behavior, evidence retrieval, and LLM-based reasoning without relying on access to evidence sources or model internals. Extensive evaluations on benchmark datasets and real-world systems demonstrate that our attacks substantially degrade verification performance, reducing accuracy from 78.7% to 53.7%, and significantly outperform existing claim-based attack baselines with strong cross-system transferability.


Decide Before You Record: Heterogeneity-Guided Pre-Acquisition Stimulus Selection for Cross-Day Neuroprostheses

Di Wu ⋅ Binglun Huang ⋅ Chengxi Xie ⋅ Zhenxi Song ⋅ Zhiguo Zhang

Decoding language intentions from intracranial electroencephalography signals enables direct communication for patients with severe speech impairments. However, practical deployment remains severely limited by cross-day performance degradation caused by physiological and instrumental heterogeneity, requiring patients to record extensive calibration data daily to maintain decoding performance. Here we introduce \textbf{He}terogeneity-Guided \textbf{D}ata \textbf{A}cquisition and \textbf{A}daptation (\textbf{HeDA\textsuperscript{2}}), a novel calibration paradigm that \textbf{determines which stimuli to record} for optimal calibration \textbf{before any neural data are acquired}, fundamentally departing from existing methods that passively adapt to whatever data happen to be available. HeDA\textsuperscript{2} operates through three stages: it first quantifies cross-day neural heterogeneity with a proposed \textbf{S}entence-level \textbf{H}eterogeneity via \textbf{A}lignment \textbf{P}ath \textbf{E}ntropy that captures how different phonetic phenomena exhibit varying degrees of heterogeneity; second, an estimator learns to predict this heterogeneity from historical patterns, enabling strategic selection of calibration sentences before any neural recording occurs; third, heterogeneity-aware adaptation updates the decoder using minimal acquired data. Extensive experiments on intracranial recordings across diverse recalibration settings demonstrate that, with merely 20\% of conventionally required calibration data, HeDA\textsuperscript{2} achieves comparable or superior decoding accuracy, representing a critical step toward clinically viable speech prosthesis.


Decomposing and Reshaping Scaling Laws through Token Learning Times

Pingjie Wang ⋅ zechenhu ⋅ Peiru Yang ⋅ Jingtao Han ⋅ Debing Zhang

Language model loss follows remarkably regular scaling laws over model and data size, yet it remains unclear why the aggregate loss should exhibit a power-law form. Existing explanations often attribute this regularity to a heavy-tailed spectrum of pattern difficulty in natural language, but this view has not been directly validated at token-level granularity in large-scale real-data training. We present a token-level framework that deconstructs scaling laws into localized learning events of individual contextualized tokens. By fitting token loss trajectories with sigmoids, we show that token learning is concentrated in localized transitions, giving rise to a learning-time spectrum that dominates the scaling-law shape. Across more than one hundred pre-training runs on large and diverse real-language corpora with modern LLM architectures, scaling up to 6B parameters and 300B training tokens, the measured learning-time spectrum quantitatively reconstructs the validation loss derivative along the training-step $T$, data-scale $D$, and model-scale $M$ axes. We further show that the same signal is actionable: by reshaping the training distribution according to when tokens become learnable, we alter the optimization trajectory and achieve 11\% faster validation-loss reduction. These results provide direct empirical evidence that scaling laws are governed primarily by the distribution of token-level learning times, and that this distribution can be used not only to explain scaling behavior but also to improve training performance.


Decomposing Effects in Neural Causal Models

Matej Zečević ⋅ Devendra Singh Dhami ⋅ Kristian Kersting

Structural causal models (SCMs) provide a principled foundation for causal reasoning, and neural causal models (NCMs) extend this framework by parameterizing causal mechanisms with neural networks. The classical formulation of NCMs assumes a specific architectural choice: a single function approximator per structural equation. In this work, we revisit this assumption and show that it represents only one point in a broader family of valid neural causal model parameterizations. Using interventions as an analytical lens and graph-based neural architectures as a concrete construction, we characterize how different parameter-sharing schemes over sub-mechanisms induce distinct members of the NCM family while preserving the same causal graph. This perspective yields two additional regimes beyond the classical formulation, which differ in how causal mechanisms are decomposed and parameterized within a unified structural framework. Together, these regimes demonstrate that architectural choices shape the causal semantics and expressivity of neural causal models, even when graphical structure is fixed. Relaxing structural assumptions further, we identify a class of partially causal models (PCMs) that support limited interventional or counterfactual queries without satisfying full SCM requirements, offering alternative trade-offs between causal fidelity and tractability. We place NCMs and PCMs within a unified expressivity spectrum between purely associational models and fully causal generative models, and illustrate these distinctions empirically using a graph-based variational autoencoder with interventional structure.


Decomposing how prompting steers behavior

Fan Cheng ⋅ Nikolaus Kriegeskorte

Prompting steers large language models (LLMs) and vision--language models (VLMs) without weight updates, but it remains unclear how a change in instruction reshapes internal representations to produce a behavioral effect. We introduce a nested geometric decomposition framework that treats prompting as a transformation of the representational geometry for the content following the prompt. We ask what class of mathematical transformation best explains the effect of prompting by finding the best alignment between representations of the same stimulus set following different prompts. For each prompt pair, we fit a sequence of increasingly expressive stimulus-invariant maps: translation, rigid transformation with uniform-scaling, sequential axis scaling, affine, and nonlinear transformations. We then test these maps causally by replacing a single layer's prompt-A hidden state for a new set of stimuli with its mapped counterpart and measuring recovery of prompt-B representational geometry and behavior. Across three LLMs, three VLMs, and six text or image datasets varying in style, emotion, scene content, and number, prompts consistently reshape representational geometry toward the instructed task structure. In the cross-validated nested variance decomposition, much of the prompt-induced activation change is explained by shape-preserving maps: translation and rigid transformation with uniform-scaling. The tier profiles reveal model- and task-specific routing strategies, differing in how much transformation classes explain variance and where along the layer hierarchy their contributions emerge. Crucially, although translation and rigid transformation tiers already improve behavioral agreement, affine transformation is the first tier to nearly recover target-prompt task geometry and produces corresponding gains in behavioral agreement. This suggests that cross-dimensional linear mixing may be a key contributor to how prompts reorganize representations toward the instructed task structure. Our framework provides a general way to decompose prompt-induced representational change into interpretable geometric components, revealing how a model routes task-relevant structure to produce prompt-driven behavior.


Decomposing Temporal and Job-Induced Dynamics for Probabilistic Computing Workload Forecasting via Graph-Conditioned Dual-Branch Diffusion

Baozhen Luo ⋅ Minbo Ma ⋅ Honglin Zhang ⋅ Yuan Yuan ⋅ Quanlu Zhang ⋅ Baodong Wu ⋅ Yong Li

Workload forecasting is a fundamental task to realizing electricity-computing synergy in modern data centers, where operational planning must account for the risk induced by fluctuating computing demand. Yet production workloads exhibit substantial volatility, making deterministic point forecasts insufficient for uncertainty-aware decision making. Existing methods mainly model workload uncertainty from temporal correlations in utilization traces, overlooking the job-induced cross-machine dependencies shaped by the scheduling layer. We empirically find that workload volatility exhibits structured and time-varying cross-machine patterns whose strength aligns with job co-occurrence. This observation motivates GD-Diff, a job-aware graph-conditioned diffusion framework for probabilistic workload forecasting. GD-Diff constructs a job-induced weighted dynamic graph from scheduling logs, transforming job co-occurrence into time-varying inter-machine dependencies. Built on this graph, GD-Diff uses a dual-branch noise prediction network within the diffusion process, where the temporal branch captures regular temporal evolution and the graph branch models job-induced dynamics. Experiments on Alibaba and Google cluster traces show that GD-Diff achieves state-of-the-art performance in both point and probabilistic forecasting, and ablation studies further validate the job-induced dynamic graph and dual-branch architecture. The code is available at \url{https://anonymous.4open.science/r/GD-Diff-F5E8/}.


Decoupled Optimization for Teacher-Student Semi-Supervised Learning via a Pioneer Student

Haorong Han ⋅ Jidong Yuan ⋅ Chixuan Wei ⋅ Yongqi Sun

Semi-supervised learning (SSL) relies on two core mechanisms: self-training under the Teacher-Student (T-S) framework and joint optimization of labeled and unlabeled losses. Despite their effectiveness, we find both mechanisms introduce distinct optimization pathologies. First, parameter coupling enforces strict synchronization between teacher and student, where strong regularization on the student degrades the teacher's fitting ability, thereby limiting the permissible generalization intensity. Second, the imbalance in gradient update consistency between labeled and unlabeled losses drives the shared parameters to prematurely converge to labeled-dominated local minima, creating a bottleneck for global optimization. To address both issues, we propose the Pioneer Student (PiS), an auxiliary branch that operates in an independent parameter space and periodically transfers accumulated knowledge back to the T-S model. Extensive experiments show that PiS is a universal plug-and-play module that consistently improves mainstream SSL methods.


Decoupling Label Shift and Surrogate Gradient Errors for Robust Federated Spiking Neural Networks

Zouquan Chen ⋅ Yifei Yang ⋅ Lidong Zheng ⋅ Zhiming Fang ⋅ Yuchen Zheng

Federated Spiking Neural Networks (FedSNNs) provide an energy-efficient and privacy-preserving learning paradigm for edge intelligence. However, their robustness under Non-IID data is limited by a coupled failure mode: label shift distorts membrane potential distributions, drives neurons away from gradient-sensitive regions, and further amplifies surrogate gradient errors, which destabilizes federated optimization. To address this problem, this paper proposes Spatio-Temporal Spiking Adaptive Thresholds (ST-SAT), a framework that mitigates the coupling between label shift and surrogate gradient errors through coordinated neuron-level adaptation and client--server knowledge alignment. At the neuron level, ST-SAT introduces the Spatio-Temporal Parametric Leaky Integrate-and-Fire (ST-PLIF) neuron, which adaptively calibrates firing thresholds, membrane potentials, and decay dynamics to keep local neuronal responses within effective gradient regions. At the collaborative interaction level, ST-SAT adopts a Dual-Scale Knowledge Distillation strategy that combines adaptive logit-level distillation with feature-level alignment of neuronal statistics, thereby reducing class distribution bias and improving knowledge consistency across heterogeneous clients. Extensive experiments on three benchmark datasets show that ST-SAT consistently improves both global model performance and client-local performance, with particularly strong gains under severe label shift. These results validate the effectiveness of the proposed framework for robust federated SNN learning. The anonymous code link is https://anonymous.4open.science/r/ST-SAT-4771.


DecQ: Detail-Condensing Queries for Enhanced Reconstruction and Generation in Representation Autoencoders

Tianhang Wang ⋅ Yitong Chen ⋅ Wei Song ⋅ Zuxuan Wu ⋅ Min Li ⋅ Jiaqi Wang

Representation Autoencoders (RAEs) leverage frozen vision foundation models (VFMs) as tokenizer encoders, providing robust high-level representations that facilitate fast convergence and high-quality generation in latent diffusion models. However, freezing the VFM inherently constrains its spatial reconstruction capacity, limiting fine-grained generation and image editing; in contrast, incorporating reconstruction-oriented signals via fine-tuning disrupts the pretrained semantic space and degrades generative fidelity. To address this trade-off, we propose \textbf{DecQ}, a simple yet effective framework for RAEs. Specifically, DecQ introduces lightweight detail-condensing queries that extract fine-grained information from intermediate VFM features through condenser modules. These queries are incorporated into the decoder to support reconstruction and are jointly generated with patch tokens during generative modeling. By aggregating information from both shallow and deep layers, DecQ effectively mitigates the reconstruction--generation trade-off, improving both reconstruction quality and generative performance. Our experiments demonstrate that: (1) with only 8 additional queries and 3.9\% extra computation, DecQ improves reconstruction over the frozen DINOv2-based RAE, increasing PSNR from 19.13 dB to 22.76 dB; and (2) for generative modeling, DecQ achieves 3.3$\times$ faster convergence than RAE, attaining an FID of 1.41 without guidance and 1.05 with guidance.

The optimal transport (OT) map provides a geometric transformation for aligning probability distributions and has become a useful tool in machine learning. However, existing estimators of the OT map still exhibit a gap between sharp statistical guarantees and practical parametric estimation based on stable training objectives. Theoretical estimators achieve minimax optimal convergence rates, but they are typically nonparametric and can incur demanding implementation design or inference costs. Practical estimators are parametric and scalable, but their statistical guarantees remain underexplored, and their min-max, adversarial-like training objectives can be sensitive to optimization algorithms. We propose BROT (Barycentric Regression for OT), a simple two-step method that first computes the unregularized OT plan and then fits a deep neural network (DNN) to the induced barycentric targets by least-squares regression. Under standard regularity conditions, we prove that the DNN estimator of BROT attains the minimax convergence rate, when the ground-truth OT map is Lipschitz. Numerical studies on synthetic datasets and an image dataset show that BROT provides accurate map estimates, strong target distribution matching, and competitive transport costs, compared to existing estimation methods. Experiments on two downstream tasks, single-cell perturbation prediction and unsupervised domain adaptation, further suggest that the accurate estimation of BROT can translate into stronger task performance.


Deep Learning-based Algebraic Reynolds Stress Closures for RANS simulations of Turbulent Flows

Daniel Dehtyriov ⋅ Jonathan F. MacArt ⋅ Justin Sirignano

Turbulence is ubiquitous in engineering, yet direct simulation is prohibitively expensive. The Reynolds-averaged Navier-Stokes (RANS) equations provide computational savings exceeding ten orders of magnitude but introduce unclosed terms requiring modelling (the closure problem). Existing machine-learning (ML) closures suffer distribution shift when trained offline on high-fidelity data, and ML models that bypass the governing equations often have limited capability to generalise. We develop a physics-derived deep learning closure for RANS, the Deep Algebraic Reynolds Stress Model (DARSM), which can be trained on one or two flow cases and generalise across an order of magnitude in Reynolds number, to unseen geometries, and across flow regimes. A neural network maps flow invariants to empirical parameters in an implicit algebraic Reynolds stress equation, derived from the Reynolds stress transport equations under the weak-equilibrium assumption, imposing significant physics-based structure on the ML closure. End-to-end optimisation through the governing PDEs and the coupled implicit closure eliminates distribution shift, but both unrolled and implicit automatic differentiation fail on the stiff coupled solver. We derive adjoint equations that exploit the solver's implicit-explicit structure to enable computationally efficient optimisation. On the canonical square-duct and periodic-hill benchmarks, DARSM reduces test velocity error as compared to the baseline RANS by $2$-$4\times$ across Reynolds number, geometries, and flow regimes, with peak case-level reductions of $12\times$. The model trained on attached, anisotropy-dominated flows (square duct) accurately generalises without retraining to separated flows (periodic hills), a regime change in the underlying turbulence physics. DARSM also outperforms five established ML methods: offline training, tensor-basis neural networks, field-inversion machine learning, DeepONets, and physics-informed neural networks.


Deferred Aggregation in Hierarchical Bayesian Optimization

Valentin Margraf ⋅ Jonas Hanselle ⋅ Julian Rodemann ⋅ Marcel Wever ⋅ Sebastian Vollmer ⋅ Eyke Hüllermeier

To account for uncertainty, hierarchical Bayesian optimization marginalizes over Gaussian process hyperparameters, usually averaging models or acquisition functions. However, this early aggregation discards information about model disagreement and is sensitive to outliers. To address these limitations, we propose to defer aggregation to the decision level: each model independently proposes a candidate, and we aggregate over these candidates via the medoid. We study a misspecified hierarchical Gaussian process setting in which most models have low misspecification error, while a minority are severely misspecified. Theoretically, we derive regret bounds showing that the misspecification penalty of early aggregation depends on the average error across all models, whereas decision-level aggregation depends only on the low-misspecification majority. Empirically, we validate these findings on toy examples and further demonstrate that, across a wide range of low- to high-dimensional synthetic and real-world benchmarks, decision-level aggregation outperforms standard Bayesian optimization baselines in most cases while remaining competitive in the others.


Deformable 2D Gaussian Splatting

Huibin Li ⋅ Chul Min Yeum

Building upon the real-time rendering capability of 3D Gaussian Splatting (3DGS), 2D Gaussian Splatting (2DGS) achieves improved surface reconstruction and geometric fidelity for novel view synthesis through explicit planar Gaussian primitives. However, the smooth falloff inherent to Gaussian kernels fundamentally limits the reconstruction of high-frequency shape signals, causing over-smoothed edges and loss of fine geometric detail. Prior works have explored alternative kernel functions to address this limitation, yet through systematic evaluation across signal fitting tasks, we find that all existing kernels suffer from the Gibbs phenomenon — producing oscillatory artifacts near sharp discontinuities. This observation motivates a fundamentally different strategy: rather than replacing the kernel, we make it deformable. This study presents Deformable 2D Gaussian Splatting, which augments each Gaussian surfel with a learnable, monotonic coordinate remapping in its local tangent plane. With only a single free control point per axis (two extra parameters), the remapping enables asymmetric radial profiles, allowing each primitive to reliably capture sharp edges and irregular structures that symmetric kernels inherently cannot represent. The deformation operates on the 2D tangent plane of 2DGS, preserving the differentiable rasterization pipeline with negligible overhead, and is optimized end-to-end through rendering loss alone. Experiments on MipNeRF 360 and Tanks&Temples demonstrate that our method effectively suppresses Gibbs artifacts at sharp boundaries and achieves state-of-the-art performance among 2DGS-based approaches in PSNR, SSIM, and LPIPS.


Delve into the Applicability of Advanced Optimizers for Multi-Task Learning

Zhipeng Zhou ⋅ Linxiao Cao ⋅ Yiming Cao ⋅ Pengcheng Wu ⋅ Peilin Zhao ⋅ Chunyan Miao

Multi-Task Learning (MTL) is a fundamental problem in machine learning that has been extensively studied over the past decade. Recently, a variety of optimization-based MTL approaches have been proposed to jointly learn multiple tasks by modifying the optimization trajectory. In this paper, we argue that the design of mainstream optimization-based MTL methods implicitly assumes compatibility with momentum-free optimizers, which may limit the understanding of their effectiveness in modern training regimes. In practice, the instantaneously derived gradients from MTL operations contribute only marginally to the final parameter updates when momentum is present, resulting in an amortized rather than instant de-conflicting effect. Moreover, this amortized behavior can be further degraded under high-curvature optimization dynamics. To address this issue and enable the effective integration of mainstream MTL methods with advanced optimizers (e.g., Adam and Muon), we propose \texttt{APT} (Applicability of advanced oPTimizers), a lightweight framework featuring an adaptive momentum mechanism that balances the trade-off between instant and amortized de-conflicting. Furthermore, we show that the Muon optimizer can be interpreted as an implicit MTL learner, and we introduce a lightweight direction preservation strategy to better align with its orthogonalization process. Extensive experiments across four standard MTL benchmarks demonstrate that \texttt{APT} consistently enhances existing MTL approaches, yielding substantial performance gains.


Democratizing Tool Learning with Environments Fully Simulated by a Free 8B Language Model

Chenming Tang ⋅ Hsiu-Yuan Huang ⋅ Weijie Liu ⋅ Junqiang Zheng ⋅ Saiyong Yang ⋅ Yunfang Wu

Reinforcement learning (RL) has become a prevalent paradigm for training tool calling agents, which typically requires online interactive environments. Existing approaches either rely on training data with ground truth annotations or require advanced proprietary language models (LMs) to synthesize environments that keep fixed once created. In this work, we propose TRUSTEE, a cost-friendly method for training tool calling agents with dynamic environments fully simulated by free open-source LMs that can be as small as 8B, including task generation, user simulation, tool simulation and trajectory evaluation, paired with an adaptive curriculum learning mechanism that controls task difficulty during training. Our empirical results show that TRUSTEE outperforms baselines which require extra external resources in most cases. These confirm that, with a sufficiently sophisticated design, even simulated environments with a local 8B LM as the backbone could set a strong baseline for tool learning. We hope our proposed paradigm could democratize tool learning and inspire future research on environment scaling with limited resources.


Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability.

Roberto Stanzione ⋅ Jules Barbe ⋅ Magali Parrino ⋅ Jérémie Fourmann ⋅ Paul Boniol

Time Series Anomaly Detection (TSAD) has received increasing attention, driven by the growing availability of complex time series data. This surge has led to the development of numerous detection methods, as well as a variety of benchmarks aimed at thoroughly evaluating their performance. However, most existing detectors remain largely agnostic to domain context, overlooking explainability and interpretability. One of the main reasons for this gap is that current benchmarks primarily focus on detection accuracy, and only few of them evaluate spatial explainability. Moreover, no benchmark currently provides sufficiently rich semantic annotations to support the generation of human-understandable interpretations of anomalies. To address these limitations, we introduce SHAD (Scality High-dimensional Anomaly Detection benchmark), a fully annotated benchmark composed of 144 multivariate, high-dimensional time series collected from real-world distributed cloud storage systems operated by Scality. The proposed dataset includes rich contextual information, covering three families of anomalies with varying degrees of severity. As further contribution, we provide an end-to-end benchmark that covers all stages of a TSAD pipeline. For Detection, we benchmark a wide range of existing anomaly detectors, testing their effectiveness on the proposed real-world dataset. Then, we evaluate explainability by measuring the contribution of each dimension in the generated anomaly score. Finally, for interpretability, we investigate the effectiveness of frozen LLM baselines in localizing and interpreting anomalies, paving the way for future research towards interpretable and explainable TSAD.


Detect What You Need: Chain-of-Causal Reasoning for 3D Intent Grounding

Zihao Zhang ⋅ Aming WU ⋅ Yang Li ⋅ Yahong Han

Accurately matching human intentions in 3D space is an important goal of artificial intelligence. Recently, 3D Intention Grounding (3D-IG) has emerged, aiming to localize target 3D objects given a natural-language intent. Unlike conventional visual grounding with descriptive referring expressions, 3D-IG intents are abstract and non-descriptive, making object localization substantially more challenging. This requires inferring latent functional requirements from non-descriptive intents and aligning them with object-level 3D representations. However, existing methods largely rely on implicit intent–object matching, leading to logical gaps and limited interpretability, robustness, and generalization. To address these challenges, we propose Chain-of-Causal Reasoning (CoCR), a causality-inspired functional dependency reasoning framework that explicitly bridges abstract intentions and candidate objects through intermediate functional requirements. Specifically, CoCR progressively decomposes complex intents into ordered functional requirements, thereby forming an explicit intent–function–object reasoning chain that links abstract intentions to object suitability. Building on this chain, we construct a causality-inspired functional dependency graph to model requirement--attribute relationships and introduce a causal-visual alignment module that aligns function-aware representations with the geometric-semantic features of 3D point clouds, enabling bidirectional verification between structured functional reasoning and visual evidence. Extensive experiments on 3D Intention Grounding and 3D Visual Grounding demonstrate that our method enhances intent-aware object localization.


DexMirror: Real-to-Sim Scene Mirroring for Sim-to-Real Dexterous Manipulation

Pu Hua ⋅ Zhecheng Yuan ⋅ Meng Chu ⋅ Peiqi Duan ⋅ Huazhe Xu

Dexterous manipulation requires large-scale robot interaction data, yet collecting real-world demonstrations is costly, while sim-to-real transfer remains challenging due to visual and geometric discrepancies. We present DexMirror, a unified real-to-sim-to-real framework that enables photorealistic simulation learning and reliable deployment for dexterous manipulation. Our method reconstructs real scenarios using a compositional 3D Gaussian Splatting (3DGS) representation with tri-level optimization, obtaining high-fidelity objects and background while preserving global scene consistency. A dexterous-hand-centric calibration pipeline, supported by an interactive online platform, further achieves millimeter-level robot–scene consistency for efficient Gaussian-simulation alignment. Built upon the reconstructed scenes, we train privileged policies in simulation and distill them into visuomotor policies operating on rendered 3DGS observations, combined with closed-loop execution for robust transfer. Experiments demonstrate accurate real-to-sim reconstruction (F1 score 0.981) and strong sim-to-real performance, achieving high success rates across six real-world dexterous manipulation tasks.


DIAGNO: Diagonal Spherical Neural Operators for Heterogeneous Earth Dynamics Modeling

Herui Li ⋅ Bin Lu ⋅ Haonan Qi ⋅ Lei Zhou ⋅ Luoyi Fu ⋅ Xinbing Wang ⋅ Meng Jin

Accurate forecasting of Earth's dynamic systems is vital in geophysics, where neural operators offer a powerful data-driven paradigm. Standard operators often treat the Earth as a flat Euclidean plane, introducing geometric artifacts. In contrast, Spherical Neural Operators (SNOs) resolve this mismatch by operating within the spherical spectral space, naturally adapting to spherical geometries. However, for computational efficiency, existing SNOs typically rely on the rotation equivariance assumption. This assumption simplifies the transitions of physical states into isolated spectral modal evolutions, thereby over-idealizing the Earth as a homogeneous system. To solve this problem, we no longer rely on the rotation equivariance assumption that isolates spectral modal evolution, but directly derive an explicit modal interaction kernel and propose Diagonal Spherical Neural Operators named DIAGNO. DIAGNO explicitly parameterizes cross-modal interactions through spectral modal decomposition, thereby preserving the inherent heterogeneity of Earth dynamics. Furthermore, leveraging the unique diagonal structure of the spherical spectrum, we design inter- and intra-diagonal interaction mechanisms to natively capture essential zonal and meridional dynamics, respecting intrinsic geophysical regularities. Extensive experiments across simulated fluid dynamics and real-world Earth system reanalysis (including atmosphere and ocean) datasets demonstrate DIAGNO’s superior capability, consistently achieving state-of-the-art performance in multi-step forecasting. Beyond this, we validate DIAGNO’s practical analytical value in professional geophysical domains by assessing its fidelity in capturing climatological anomalies and maintaining physical consistency.

We introduce DIAL (Directional Intervention, Asymptotically Limited), a parameter-efficient adapter that adds a bounded, monotone, scalar-controlled residual $\delta(s,c) = \sigma(g) \odot \tanh(c/T) \odot \tanh(Dz)$ with $z = D^\top \phi(s)$ to a base language model's hidden state. The architectural form yields a closed-form upper bound $\lVert \delta \rVert_\infty \le \sigma(g_{\max})$ on the residual $L_\infty$ norm, computable from model weights alone. Across 5 base models (3.8B to 70B parameters, 4 architecture families; one checkpoint per family, with 5-seed csweep variance at Llama-3.1-8B), the residual saturates at the predicted bound with mean ratio $1.000 \pm 0.001$ (worst-case family deviation 0.2 percentage points). On Llama-3.1-8B under GPT-4o judge ($n=100$ per bench, zero content-filter blocks), DIAL exceeds a FiLM baseline trained with the same multi-$c$ objective on every benchmark and judge. Semantic vs strict judge: HarmBench DIAL +0.52 / +0.55 vs FiLM +0.44 / +0.39; AdvBench DIAL +0.63 / +0.66 vs FiLM +0.46 / +0.50; XSTest DIAL +0.33 / +0.25 vs FiLM +0.26 / +0.22; both judges agree on the ranking. Outside training range, DIAL's residual saturates at $\sigma(g_{\max})$ while FiLM's grows unboundedly: at $c = 10^4$, FiLM's residual $L_\infty$ norm is 534 times larger than DIAL's. The bound holds across all control values, not only trained ones; concrete deployment scenarios (configuration error, multi-tenant policy ceilings, adversarial composition) are where this distinction matters in practice. The bounded-monotone form produces analogous saturation behavior on two continuous-control RL environments (HC-Vel $n=35$, Ant-Goal $n=30$) under matched oracle-$c$; the prior fails when $c$ must be inferred online (meta-RL). Ablations confirm both the bound and the saturating temperature are necessary; a frozen-gate variant retains the bound at lower controllability magnitude. We release adapter train reports and canonical bench generations, algebraic derivations (with a SymPy reproducibility script), the full GPT-4o judge cache (semantic and strict), and a human-labeled validation set ($n=50$, Cohen's kappa = 0.895). Full adapter checkpoints are deferred to camera-ready release.


Differentiable Cluster Discovery in Temporal Graphs

Hafizur Rahman ⋅ Chi-Guhn Lee

Existing temporal graph clustering methods suffer from poor optimization dynamics due to reliance on heuristically initialized cluster assignment distribution without considering the dynamic nature of the evolving graph. The target cluster assignment distribution often conflicts with evolving temporal representations, leading to oscillatory gradients and unstable convergence. Motivated by the need for differentiable and adaptive clustering in dynamic settings, we propose TGRAIL (Temporal Graph Alignment and Index Learning), a novel end-to-end framework for temporal graph clustering based on Gumbel–Softmax sampling. TGRAIL enables discrete cluster assignments while maintaining the gradient flow. To ensure stable training, we formulate the clustering objective as an expectation over Monte Carlo samples and show that this estimator is both unbiased and variance-reduced. Furthermore, we incorporate a temporal consistency loss to preserve the order of interactions across time. Extensive experiments on six real-world temporal graph datasets demonstrate that our approach consistently outperforms state-of-the-art baselines, achieving higher clustering accuracy and robustness. Our results validate the effectiveness of jointly optimizing temporal dynamics and discrete cluster assignments in evolving graphs.


Differential Vector Erasure: Unified Training-Free Concept Erasure for Flow Matching Models

Zhiqi Zhang ⋅ Xinhao Zhong ⋅ Yi Sun ⋅ Shuoyang Sun ⋅ Jiahui Wang ⋅ Bin Chen ⋅ Shu-Tao Xia ⋅ Xuan Wang

Text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality images, yet their tendency to reproduce undesirable concepts poses growing concerns for safe and controllable deployment, particularly regarding NSFW content, specific objects and styles. While existing concept erasure approaches primarily focus on DDPM-based diffusion models and rely on costly fine-tuning, the recent emergence of flow matching models introduces a fundamentally different generative paradigm for which prior methods are not directly applicable. In this paper, we propose Differential Vector Erasure (DVE), a training-free concept erasure method specifically designed for flow matching models. Our key insight is that semantic concepts are implicitly encoded in the directional structure of the velocity field governing the generative flow. Leveraging this observation, we construct a differential vector field that characterizes the directional discrepancy between a target concept and an anchor concept. During inference, DVE selectively removes concept-specific components by projecting the velocity field onto the differential direction, enabling precise concept suppression without affecting unrelated semantics. Extensive experiments on FLUX demonstrate that DVE consistently outperforms existing baselines on a wide range of concept erasure tasks, including NSFW suppression, artistic style removal, and object erasure, while preserving image quality and diversity.


Diffusion Guidance Is a Controllable Policy Improvement Operator

Kevin Frans ⋅ Seohong Park ⋅ Pieter Abbeel ⋅ Sergey Levine

At the core of reinforcement learning is the idea of learning beyond the performance in the data. However, scaling such systems has proven notoriously tricky. In contrast, techniques from generative modeling have proven remarkably scalable and are simple to train. In this work, we combine these strengths, by deriving a direct relation between policy improvement and guidance of diffusion models. The resulting framework, CFGRL, is trained with the simplicity of supervised learning, yet can further improve on the policies in the data. On offline RL tasks, we observe a reliable trend -- increased guidance weighting leads to increased performance. Of particular importance, CFGRL can operate without explicitly learning a value function, allowing us to generalize simple supervised methods (e.g., goal-conditioned behavioral cloning) to further prioritize optimality, gaining performance for "free" across the board.


Diffusion Path Samplers via Sequential Monte Carlo

James Matthew Young ⋅ Paula Cordero-Encinar ⋅ Sebastian Reich ⋅ Andrew Duncan ⋅ O. Deniz Akyildiz

We develop diffusion-based samplers for target distributions known up to a normalising constant. To this end, we rely on the well-known diffusion path that smoothly interpolates between a simple base distribution and the target, popularised by diffusion models. We tackle the score estimation problem by developing an efficient sequential Monte Carlo sampler that evolves auxiliary variables from conditional distributions along the path, providing principled score and density estimates for time-varying distributions. To control the variance of score estimates, we further propose practical control variate schedules that incur minimal overhead. We adapt this general framework to paths induced by the Ornstein–Uhlenbeck (OU) time-reversal process, stochastic interpolants, and diffusion annealed Langevin dynamics, outlining their trade-offs. Finally, we provide theoretical guarantees and empirically demonstrate the effectiveness of our method on several synthetic and real-world datasets.

We introduce a novel technique for scalable sampling of spin-system states with continuous symmetries using diffusion models. By applying our approach to the XY model, a fundamental continuous-spin model in condensed matter physics, we show that our technique addresses the shortfalls of the Markov chain Monte Carlo (MCMC) in generalization to varying system sizes. More specifically, we show that training a temperature-conditioned diffusion model on smaller-size XY model lattices enables the generation of accurate samples in larger lattice sizes. By tracking physically important observables of the model, such as spin correlations, our experiments demonstrate that diffusion sampling followed by few MCMC steps reduces the thermalization time by an order of magnitude relative to the standard MCMC with random initialization. Our study provides valuable insight as to how generative models can be used to study continuous-state condensed matter systems at scale.

3D shape completion from partial scans remains challenging for unseen categories and noisy real-world observations, where geometry alone is often insufficient for inferring missing structure. We present DinoComplete, a deterministic and efficient shape completion framework that augments geometric reconstruction with voxel-aligned semantic priors distilled from DINO features. First, we construct multi-view DINO feature volumes aligned with ShapeNet data and train a student network to predict dense semantic features directly from incomplete shapes. These predicted features capture global structure and part-aware semantic context while remaining aligned with the underlying geometry. We then integrate these distilled features into a completion network, where geometric and semantic voxel representations are fused through voxel state-space modeling. To enable efficient long-range reasoning without sacrificing resolution, we introduce a multi-scale voxel Mamba module that refines the fused features by combining full-grid and chunk-wise sequence modeling. Experiments on unseen ShapeNet categories and ScanNet objects show that DinoComplete achieves stronger completion quality than prior deterministic and generative based completion methods while using fewer parameters, requiring lower memory, and achieving faster inference. Our results demonstrate that distilling semantic priors from visual foundation models improves generalization and robustness in 3D shape completion.

We study nonparametric estimation of Schrödinger bridge (SB) drifts from i.i.d data observed on a single time interval. Starting from the conditional-ratio form of the Schrödinger bridge time-series (SBTS) drift formula, we analyze a direct Nadaraya--Watson plug-in estimator built from kernelized numerator and denominator terms. Unlike recent SB analyses based on entropic-OT potentials, Sinkhorn iterations, or iterative bridge solvers, our approach works directly at the drift level and isolates \emph{statistical error} from optimization, approximation, and discretization error. Under Hölder regularity, a marginal-density floor, and bounded support, we prove a uniform non-asymptotic bound for admissible bandwidth pairs, a pointwise CLT under genuine undersmoothing, and an adaptive bandwidth selector satisfying an oracle inequality. We also prove a pivot-local minimax lower bound which, through an explicit uniform pivot, yields a global minimax lower bound under transparent compatibility conditions; hence the adaptive selector is minimax-rate optimal up to logarithmic factors. Synthetic experiments provide theorem-targeted diagnostics for finite-sample scaling, Gaussian approximation, and adaptive behavior.

Humans and modern vision models can reach similar classification accuracy while making systematically different kinds of mistakes - differing not in how often they err, but in who gets mistaken for whom and in which direction. We show that these directional confusions reveal distinct inductive biases that are invisible to accuracy alone. Using matched human and deep vision model responses on a natural-image categorization task under 12 perturbation types, we quantify asymmetry in confusion matrices and link it to the shape of the information–error trade-off. We characterize this trade-off geometry using three signatures extracted from the rate–distortion (RD) frontier: slope (beta), curvature (kappa) and efficiency (AUC). We find that humans exhibit broad but weak asymmetries, whereas deep vision models show sparser, stronger directional collapses. Robustness training reduces global asymmetry but fails to recover the human-like breadth–strength profile of graded similarity. Mechanistic simulations further show that different asymmetry organizations shift the RD frontier in opposite directions, even when matched for performance. Together, these results position directional confusions and RD geometry as compact, interpretable signatures of inductive bias under distribution shift.


Directional Noise Conditioning for Diffusion Models

Mahdi Shafiei ⋅ Azade Farshad ⋅ Nassir Navab ⋅ Yousef Yeganeh

Conditioning diffusion models on visual features is essential for controlled image generation and counterfactual reasoning across diverse applications, from creative design to medical imaging. Existing conditioning methods, such as timestep embedding and cross-attention mechanisms, suffer from limited control over feature accuracy, require substantial architectural modifications, and struggle to maintain conditioning precision throughout the denoising process, as the model's focus shifts toward fine-grained details in later steps rather than the desired conditioned features. We present a novel approach that addresses these limitations by learning to shift the initial noise distribution in the latent space. Our method assigns each visual feature a distinct direction in the latent space and applies proportional shifts during both training and inference, enabling precise control over continuous and discrete visual attributes. We introduce a combined loss function that explicitly enforces shift accuracy alongside denoising, ensuring consistent conditioning across all denoising steps. Furthermore, we extend this framework to datasets with hidden or inaccessible visual features by employing a Variational Autoencoder to extract latent representations in an unsupervised manner. This enables high-fidelity counterfactual generation on complex, real-world datasets where explicit feature annotations are unavailable or prohibitively expensive to obtain.

Computational cognitive modeling has traditionally relied on manual model construction, which is a time-intensive process that requires substantial domain expertise. Although recent LLM-based approaches have begun to automate model generation, they often prioritize predictive accuracy at the expense of structural plausibility and parameter identifiability. We introduce DisCo, an LLM-guided framework that discovers interpretable symbolic cognitive models through evolutionary search with two structural components: a Validator that filters structurally invalid models, and a Cleaner that promotes parameter identifiability by removing redundant parameters. Across three behavioral tasks, DisCo consistently discovers models that achieve performance competitive with task-specific baselines. Ablation analyses show that the Validator is important for recovering psychologically meaningful structure. Without this component, models tend to violate behavioral constraints and show less consistent associations with external cognitive measures. These results suggest that LLM-driven model discovery, when coupled with appropriate structural constraints, can yield cognitively plausible models that go beyond predictive fit.


Discrete Neural Interlingua for Ontology-Agnostic EHR Modeling

Tong Ding ⋅ Caiwei Tian ⋅ Dongmin Bang ⋅ Ming Yang Lu ⋅ Sophia J. Wagner ⋅ Andrew Zhang ⋅ Rowland Pettit ⋅ Long P Le ⋅ Faisal Mahmood

EHR foundation models create new opportunities for scalable clinical prediction across health systems, yet their deployment remains constrained by a basic tokenization bottleneck. The same clinical concept can appear as different ontology codes, local identifiers, synonyms, or free-text descriptions, while standard EHR tokenizers treat these variants as unrelated symbols. We introduce MEDIATOR, a discrete ontology-agnostic medical concept tokenizer that turns EHR harmonization into an open-form tokenization problem. Instead of relying on manual mappings, a fixed ontology vocabulary, or continuous event embeddings, MEDIATOR maps any clinical concept expressed as text into a shared vector-quantized codebook. MEDIATOR is trained with synonym supervision from paired concept descriptions, and further guided by enriched concept descriptions, empirical code co-occurrence, and curated relation graphs so that equivalent expressions collapse while non-equivalent concepts remain organized in a medically meaningful discrete space. The resulting vocabulary is compact, inspectable, reusable by standard EHR Transformers, and open to previously unseen code descriptions or text strings. Used as the token interface of an EHR Transformer, MEDIATOR improves mean AUPRC on 20 MIMIC-III tasks by 4.0% and improves zero-shot transfer to 14 eICU tasks by 8.8%, while supporting transferable attribution across hospitals.


Disentangling Homophilic and Heterophilic Patterns for Multi-Domain Graph Foundation Models

Ziyan Wang ⋅ Ruiyi Fang ⋅ Jingyu Zhao ⋅ Zhimin Mei ⋅ Charles Ling ⋅ Boyu Wang

Graph Foundation Models (GFMs) have attracted growing attention for their ability to model multi-domain graphs universally and generalize to unseen downstream tasks. However, existing GFMs typically rely on a single domain-invariant representation space, which inevitably blurs two fundamental connectivity patterns, \emph{homophily} and \emph{heterophily}, that govern how nodes relate across domains. Additionally, hard-partitioning the graph into disjoint subgraphs discards edges that are essential for self-supervised pre-training. In this work, we introduce $\underline{\text{HHGFM}}$, a $\underline{\text{H}}$omophilic and $\underline{\text{H}}$eterophilic Pattern-Disentangled $\underline{\text{G}}$raph $\underline{\text{F}}$oundation $\underline{\text{M}}$odel that encodes homophily and heterophily separately while keeping the original edge structure intact. Specifically, we propose a node-level soft disentanglement mechanism that estimates each neighbor's homophilic and heterophilic contributions and uses them as edge weights, thereby preserving the original graph structure. Low-pass and high-pass filters then serve as dual backbones to align homophilic and heterophilic patterns separately, supported by our generalization analysis. Extensive experiments on a range of graph datasets validate the effectiveness of our method.


Displacement Geometry Captures Platonic Shared Reality Across Models and Modalities

Chenming Shang ⋅ Yujin Tang ⋅ Jun Jie Ou Yang ⋅ Ruize Xu ⋅ Adam Breuer ⋅ Nikhil Singh

The Platonic Representation Hypothesis (PRH) claims that independently trained models converge on a shared statistical model of reality, yet recent work finds only weak pointwise similarity between models. In this paper, we show that what models share is not the location of samples in representation space, but the directions (displacement vectors) between them. Under a single orthogonal alignment---rotation and reflection only---these displacement vectors are substantially preserved across 44 independently trained vision and language encoders spanning modalities and asymmetric capability pairs, consistent with the PRH evidence. The samples' absolute positions are not, consistent with recent counter-evidence. Both arise from a single decomposition: representations split into a shared semantic component that is linearly aligned across models, and a private capability component that is not. We trace this geometry to concept-level structure: within a model, parent concepts are orthogonal to their child variation vectors; across models, concept displacements are parallel. Our theory falsifiably predicts (and experiments confirm) that fine-tuning preserves pointwise similarity but collapses displacement, and that relational distillation does the opposite. A major implication is that, because semantics align linearly but capabilities do not, capabilities can be imported from one model to another using a single cached forward pass through the source. We call this Shadow Casting. As a proof of concept, our ShadowCLIP instantiation matches or outperforms strong fine-tuned baselines at orders of magnitude less compute. A cache can be released alongside open model weights, letting one model's capabilities be downloaded and imported into any number of other models without fine-tuning.


Distance Marching for Generative Modeling

Zimo Wang ⋅ Ishit Mehta ⋅ Haolin Lu ⋅ Xunpeng Huang ⋅ Chung-En Sun ⋅ Ge Yan ⋅ Lily Weng ⋅ Tzu-Mao Li

Modern generative models based on diffusion or flow matching transport noise toward the data manifold, without representing the data manifold explicitly. We propose \dm that explores an alternative design of generative modeling by explicitly modeling a distance field of the data manifold. Designing loss functions to learn distance fields in high-dimensional space from point data can be challenging. This is because the data do not provide reliable signals for closest point projection onto the manifold. To mitigate this problem, we design new loss functions that weight training data by their proximity to each sampling position. This distance-field perspective naturally motivates two principled deterministic samplers, showing fast convergence in inference. Our model is related to the recent time-unconditional generative models since it also does not require time input, and our analysis reveals that our method helps time-unconditional generative modeling to disambiguate the denoising problem and to avoid biasing towards the data mean. As a result, our model is the only time-unconditional model that remains competitive to time-conditional ones across datasets and architectures. Moreover, our distance prediction is also helpful for early stopping during sampling and for OOD detection. We hope distance-field modeling can open a promising direction for exploring generative modeling from a geometric perspective.

Conditional image generators learn $P(Y\mid X)$ for images given clinical, demographic, or experimental covariates, but their covariate effects are encoded implicitly in samples, denoisers, gradients, or score fields. We propose \emph{generative effect distillation} (GED), a framework that treats a conditional image generator as a teacher and projects its conditional law onto an interpretable image-response regression student $m_{\phi}(x,s)=\beta_0(s)+x^\top\beta(s)$ whose parameters are spatial effect maps. The teacher-to-map target is defined as a statistical functional of the teacher conditional law under a target design and discrepancy, rather than a set of synthetic images or a compressed generator. We develop mean, effect-gradient, and structured score distillation objectives; establish projection identities linking the distilled map to covariance-weighted teacher means and, under Stein conditions, to average teacher gradients; and derive finite-query risk bounds that separate teacher bias, spatial approximation, query design, Monte Carlo, and student optimization error. For the associated finite-basis coefficient-matrix oracle, a Gaussian lower bound matches the query-design and Monte Carlo stochastic dependencies. Across eleven blocks spanning synthetic, semi-real, pretrained SDXL-Turbo/FFHQ-256, CelebA conditional-diffusion, and face-image analyses, GED attains the lowest standardized integrated squared error (SISE) among the direct-regression and augmentation baselines in every block, with paired Student $t$-tests rejecting equality with the strongest non-GED baseline at $p<0.05$; on a CelebA conditional residual diffusion teacher, score-access GED via a single-step Tweedie denoising procedure attains SISE $0.009$, a factor of $25$ below the strongest non-GED baseline.


Distilling Multi-Teacher Scoring Principles for Unsupervised Time Series Anomaly Detection

Dongchan Cho ⋅ Jiho Han ⋅ Juwon Hwang ⋅ Changhee Min ⋅ Minsang Kim ⋅ Namsoon Jung

Time series anomalies take diverse forms, and individual detectors encode different anomaly scoring principles that rarely cover all anomaly types. Recent unsupervised time-series anomaly detectors provide complementary scoring signals, yet deploying them as an ensemble preserves the inference and maintenance cost of every individual detector. We study whether heterogeneous teacher detectors can be distilled into a single lightweight detector in the unsupervised setting, where no anomaly labels are available to define a shared target. AnoMix addresses this by extracting weak pairwise supervision from the training-set anomaly scores of frozen teachers. It normalizes each teacher score sequence independently and trains the student with CoRank, a multi-teacher ranking-distillation objective that converts teacher score gaps into pairwise supervision through teacher consensus. Across seven benchmark datasets, AnoMix successfully integrates complementary teacher scoring principles into a single deployable model, improving over individual detectors and static score ensembles while using a compact student with under 80k parameters on the benchmark settings. A zero-shot variant further shows that the distilled ranking signal transfers to unseen datasets and can outperform large pretrained time-series foundation models for anomaly detection.


Distributionally Robust Black Box Optimization-based Bidding Strategy in Auction-based Federated Learning

Xiaoli Tang ⋅ Zhuang Qi ⋅ Ying-Peng Tang ⋅ Xianjie Guo ⋅ Han Yu ⋅ Jie Zhang ⋅ Hengjie Song

Auction-based Federated Learning (AFL) provides a market-based mechanism to incentivize decentralized data owners (DOs) to join FL model training initiated by data consumers (DCs). However, optimizing the DC's bidding strategy remains fundamentally challenging. Existing methods predominantly rely on Reinforcement Learning (RL), which suffers from severe instability caused by the mismatch between stepwise Markovian rewards and the trajectory-level, non-decomposable feedback inherent in AFL. To address this mismatch, we transition from the stepwise RL formulation to a distributionally robust black-box optimization framework. By treating the entire recruitment-to-training process as a single high-fidelity function evaluation, this transition eliminates the need for explicit reward decomposition. We propose DR-AFL , which accounts for market non-stationarity and data distribution shifts by optimizing policies against worst-case environments within a Wasserstein ambiguity set. Our approach further incorporates an adversarial reweighting scheme to enhance robustness, alongside a dynamic safety projection that leverages real-time cost feedback to enforce budget-aware exploration. We provide theoretical analysis showing that DR-AFL induces a smooth surrogate over the discontinuous auction landscape and guarantees a robust performance lower bound under distributional shifts. Extensive experiments on 6 widely adopted datasets show that DR-AFL consistently outperforms state-of-the-art RL-based baselines, achieving a 2.43% improvement in model accuracy while reducing the number of training cycles required to reach target accuracy by 62.3%.

Learning a risk-sensitive individualized treatment rule (ITR) with patient covariates is challenging when the treatment criterion depends on the full conditional outcome distribution rather than only its conditional mean. Existing methods are largely criterion-specific, making it difficult to compare multiple treatment criteria within a common pipeline, such as quantiles, conditional value-at-risk (CVaR), and complex preference-based objectives such as cumulative prospect theory (CPT). We propose a distribution-first framework for risk-sensitive ITR learning that first estimates treatment-specific conditional outcome distributions and then constructs the target ITR by plug-in criterion evaluation. The proposed framework unifies treatment rule learning across multiple criteria and accommodates hard-to-optimize objectives within the same estimation pipeline. On the theoretical side, it yields a general error propagation result for Lipschitz treatment criteria, linking first-stage distribution estimation error to downstream criterion estimation and treatment regret, together with sharper regret bounds under standard margin conditions. Empirically, we show that the proposed framework outperforms direct criterion-specific methods in multiple settings and further evaluate it on the ACTG175 HIV clinical trial, a randomized study comparing several antiretroviral regimens in HIV-infected adults.


Distribution Matching Distillation without Fake Score Network

YoungJoong Kim ⋅ Deokyeong Lee ⋅ Jaesik Park

Distribution Matching Distillation (DMD) provides an effective distribution-level correction for few-step generation, while relying on an auxiliary fake-score network to track the evolving generative distribution. Recent work combines DMD-style objectives with flow-map generators to exploit both forward-divergence training and reverse-divergence correction. The fake-score estimator remains an additional component with memory and update overhead. In this work, we study whether this explicit tracker can be avoided when the generator itself has a flow-map structure. We propose \textit{Fake-Score-network-Free DMD (FSF-DMD)}, a DMD formulation for flow-map generators that replaces the auxiliary fake-score estimator with a generator-induced pseudo-velocity surrogate. The key observation is that the endpoint pseudo-velocity of a flow-map generator provides a tractable proxy for fake-velocity estimation, allowing the generator itself to supply the reverse-divergence signal. Building on this observation, we derive a practical objective, extend it with flow-map-consistent backward simulation, and introduce a self-teacher variant for training from scratch. In our ImageNet-1K $256\times256$ experiments, FSF-DMD improves flow-map baselines, reaches lower FID than the listed DMD2 comparisons in the flow-map-initialized setting, and remains effective under flow-matching initialization and training from scratch.

Generating high-fidelity mixed-type tabular data remains challenging because such data contain both numerical and categorical features; different columns often exhibit substantial heterogeneity in marginal distributions, sparsity patterns, and dependency structures. Although recent diffusion-based and flow-based methods have improved generation quality, most methods still rely on a unified generation framework that models all dimensions in a homogeneous manner, without explicitly accounting for the fact that different columns may differ in their suitability for strong local refinement. To address this issue, we propose a driver-navigator structured flow-matching framework (DN-Flow) for mixed-type tabular data generation. Inspired by the collaboration between the driver and navigator in rally car racing, DN-Flow explicitly decomposes the velocity field into a Driver branch (for globally stable transport) and a Navigator branch (for endpoint-aware local refinement). On top of this structured decomposition, we introduce a learnable column-wise soft selector and a dynamic correction control mechanism to jointly model column-level suitability for local refinement and the actual strength of correction injection in an end-to-end manner. Experiments on eight real-world datasets show that DN-Flow consistently outperforms strong baselines in data fidelity and downstream utility, achieves the best Shape performance on all eight datasets, and remains highly competitive on Trend and machine learning efficiency. These results suggest that structured and selective velocity correction offers an effective new paradigm for mixed-type tabular data generation. The code will be released upon acceptance.


Does Inference-Time Reasoning Really Improve Video Understanding?

Zhuohao Yu ⋅ Yu Zhou ⋅ Zhecan Wang ⋅ Rui Sun ⋅ Kai-Wei Chang ⋅ Nanyun Peng

Video understanding models achieve impressive 70%+ accuracy on recent benchmarks, but do these truly measure genuine grounded understanding? We systematically investigate this by controlling frame sampling from text-only to real-time across major benchmarks (VideoMME, Video-MMMU, LVBench). Our analysis reveals that 13–46% of questions exploit shortcuts through world knowledge, language priors, or single-frame cues, allowing models to answer correctly without engaging visual evidence. Furthermore, requiring temporal grounding, both correct answers and time localization, exposes a ~40% accuracy gap, showing models frequently answer correctly without locating relevant evidence in videos. To address these limitations, we introduce TGVMME, a temporally grounded benchmark with 1.5k questions requiring both answer correctness and temporal localization. Our benchmark reduces shortcuts to 3.5%, validating that temporal grounding effectively ensures genuine video understanding. Using TGVMME to analyze state-of-the-art reasoning models, we make a surprising discovery: inference-time reasoning provides minimal gains (+2.5% at best), contrasting sharply with substantial gains on existing benchmarks. Analysis on accuracy changes of shortcut questions confirms reasoning primarily exploits shortcuts rather than enhancing genuine understanding. Further experiments reveal input scaling (increasing frames) provides more benefit than output scaling (reasoning) given sparse inputs, highlighting the importance of frame-efficient architectures. Our work demonstrates that grounded evaluation is essential for assessing video understanding capabilities and provides practical tools: validated shortcut taxonomy, filtered splits, and TGVMME for rigorous evaluation.


Does the Question Really Matter? Training-Free Data Selection for Vision-Language SFT

Peng Sun ⋅ Yi Yang ⋅ Huawen Shen ⋅ Yi Ban ⋅ Tianfan Fu ⋅ Yanbo Wang ⋅ Yuqiang Li

Visual instruction tuning is crucial for improving vision-language large models (VLLMs). However, many samples can be solved via linguistic patterns or common-sense shortcuts, without genuine cross-modal reasoning, limiting the effectiveness of multimodal learning. Prior data selection methods often rely on costly proxy model training and focus on difficulty or diversity, failing to capture a sample’s true contribution to vision-language joint reasoning. In this paper, we propose CVS, a training-free data selection method based on the insight that, for high-quality multimodal samples, introducing the question should substantially alter the model’s assessment of answer validity given an image. CVS leverages a frozen VLLM as an evaluator and measures the discrepancy in answer validity with and without conditioning on the question, enabling the identification of samples that require vision-language joint reasoning while filtering semantic-conflict noise. Experiments on Vision-Flan and The Cauldron show that CVS achieves solid performance across datasets. On Vision-Flan, CVS outperforms full-data training by 3.5\% and 4.8\% using only 10\% and 15\% of the data, respectively, and remains robust on the highly heterogeneous Cauldron dataset. Moreover, CVS reduces computational cost by 17.3\% and 44.4\% compared to COINCIDE and XMAS. The code is anonymously available at https://anonymous.4open.science/r/CVS-ABAA.

Decentralized training of deep learning models is widely used to enable data privacy and on-device learning over networks. In realistic scenarios, heterogeneity across clients' local data poses an optimization challenge and can severely degrade test accuracy, especially on sparse communication graphs. Building on recent advances in centralized primal averaging, we propose Decentralized Primal Averaging (DPA), a decentralized learning algorithm that combines primal averaging with a local quasi-global momentum estimate. DPA maintains three coupled sequences for optimization, gradient evaluation, and output averaging, while using consecutive mixed iterates to construct a network-aware momentum direction without transmitting an auxiliary tracking variable. We prove a nonconvex stationarity guarantee for DPA's averaged output model and establish a linear-speedup result up to network- and heterogeneity-dependent residual terms. We also introduce DPA-1G, a one-gossip-per-iteration variant that matches the communication budget of DSGD-style baselines. Experiments following the GUT decentralized image-classification benchmark show that DPA outperforms DSGD, gradient tracking, QG-DSGDm, momentum tracking, and QG-GUTm across datasets, architectures, topologies, and heterogeneity levels, with the largest gains under high heterogeneity. DPA-1G matches or exceeds the one-gossip baselines at the same per-iteration communication budget.


DrawingsDreamer: A Unified Multi-View Engineering Drawings Generation Model

Shurui Liu ⋅ Weide Chen ⋅ Changwang Yi ⋅ Ancong Wu

Scalable Vector Graphics (SVG) are essential for modern industrial Computer-Aided Design (CAD). However, existing autoregressive SVG generation models are predominantly tailored for artistic creation and struggle to maintain the rigorous geometric fidelity and cross-view spatial alignment required for engineering drawings. To bridge this gap, we introduce $\textbf{DrawingsDreamer}$, a unified Large Language Model (LLM)-driven framework for multi-view vector-based engineering drawings generation. By formulating the generation of multi-view engineering drawings purely as a sequence modeling task, we eliminate the need of raster image encoders. We propose a Streamlined Representation utilizing hierarchical postfix tokenization, which guides the model to establish local geometric coordinates before assigning semantic boundaries. Optimized via a progressive task-aware curriculum schedule, $\textbf{DrawingsDreamer}$ effectively transitions from localized structural repair to macroscopic generation in a unified model. Extensive experiments demonstrate that our unified model achieves strong performance in both geometric fidelity and syntactic accuracy across diverse conditional and unconditional generation tasks.


Drift-Aware Multimodal User Representation Learning via Multi-Scale Temporal Modeling and Sparse Mixture-of-Experts

Ziqing Qian ⋅ Haohang Chen ⋅ Shengqi Dang ⋅ Yuhan Xiong ⋅ Canyu Shen ⋅ Jiaying Lei ⋅ Nan Cao

Understanding user preferences from noisy and temporally evolving social media behaviors is fundamentally challenging due to interest drift, where user preferences shift across time and exhibit both multi-scale temporal patterns and diverse co-existing interests. To address this, we propose DUMoE, a unified framework for drift-aware multimodal user representation learning. Our model consists of (i) a temporal dynamics-aware backbone that captures and integrates static profiles, short-term behavioral signals, and long-term dependencies into a coherent representation, and (ii) a sparse mixture-of-experts (MoE) interest adapter that disentangles multiple latent interests via expert specialization and adaptive routing. Each expert models a distinct interest subspace, while a gating network dynamically selects and aggregates a sparse subset of relevant experts for each user. To enable stable and effective optimization, we further introduce a three-stage training strategy that decouples backbone learning, expert specialization, and gating optimization. Extensive experiments on real-world social media datasets show that DUMoE consistently outperforms state-of-the-art methods on both user interest prediction and interaction prediction tasks.


Driver Attention as Competitive Allocation: A Scene--Task-Aware Dual-Branch Framework

Qianfang Wang ⋅ Bin Rao ⋅ Pengpeng Xu ⋅ Tiantian Chen

Driver attention prediction is crucial for interpretable driver monitoring and human-like autonomous driving. Existing methods typically formulate this task as a scene-to-heatmap regression problem, failing to reveal the underlying dynamic attention redistribution mechanism. To address this issue, we propose a scene-task-aware competitive allocation framework, which decomposes driver attention as a bottom-up visual saliency prior and a top-down scene-task utility representation. An adaptive product-of-experts allocation rule is designed to fuse these two heterogeneous branches in a competitive manner, where scene-task demand dynamically modulates the dependence on top-down task utility, and driving control urgency adjusts spatial attention concentration via a learnable attention budget temperature. To explicitly differentiate task-critical regions from visually prominent yet task-irrelevant distractors, we incorporate competition-aware contrastive learning to enhance feature discrimination in the latent space. Extensive experiments demonstrate that our method achieves strong and balanced performance, with a Kullback–Leibler divergence of 0.98 and a normalized scanpath saliency of 4.15 on the DR(eye)VE dataset, as well as a Pearson’s correlation coefficient of 0.67 and a histogram similarity of 0.54 on the BDD-A dataset. Ablation studies and diagnostic analyses further validate that our framework yields interpretable attention allocation behaviors, including urgency-induced spatial entropy contraction and scene-task-demand-guided dynamic attention reallocation.


DriveWAM: Video Generative Priors Enable Scalable World-Action Modeling for Autonomous Driving

Chen Shi ⋅ Jinrui Xu ⋅ Shaoshuai Shi ⋅ Kehua Sheng ⋅ Bo Zhang ⋅ Li Jiang

Pretrained foundation models have become an important basis for end-to-end autonomous driving. In contrast to vision-language models pretrained primarily on static image-text pairs, video generative models capture temporal dynamics and motion priors that are naturally suited for driving. We present DriveWAM, a driving world-action model that adapts a pretrained video diffusion transformer into an autoregressive video-action policy. DriveWAM organizes video and action streams into a unified temporal token sequence and trains them under a joint flow-matching objective, preserving the pretrained video-generation architecture while adapting its large-scale video priors to action generation. To incorporate high-level scene understanding, we introduce scene-evolving driving guidance, where a frozen VLM produces chunk-specific semantic intent to guide video-action generation. To keep long-horizon rollout bounded, we further introduce selective KV memory, which maintains bounded modality-aware video and action memory pools through relevance-redundancy cache selection at inference time. Experiments on NAVSIM and the PhysicalAI-Autonomous-Vehicles benchmark show that DriveWAM achieves strong planning performance, and a data-scaling study from 4k to 100k driving clips further confirms the scaling potential of world-action modeling for end-to-end autonomous driving.


DT-PBO: an Interpretable Tree-based Surrogate Model for Preferential Bayesian Optimization

Thomas Quadt ⋅ Nick Leenders ⋅ Roy Lindelauf ⋅ Herman Monsuur ⋅ Mark Voskuijl ⋅ Joost van Oijen ⋅ Boris Cule

Preferential Bayesian Optimization (PBO) aims to find a decision-maker’s most preferred solution in as few pairwise comparisons as possible. Existing approaches rely on Gaussian Process (GP) surrogates, which provide strong performance but limited interpretability. This limits real-world usability in high-stakes domains, such as healthcare, where interpretability and trust are essential. We propose DT-PBO, a novel tree-based surrogate model for PBO that is inherently interpretable while capturing preference uncertainty. Specifically, we introduce a novel splitting heuristic that constructs interpretable shallow decision trees directly from pairwise comparison data, and use Laplace approximation to obtain probabilistic estimates within each leaf. This enables efficient preference modeling without sacrificing interpretability. Across eight benchmark functions, our method achieves competitive convergence to GP-based PBO, particularly on functions with rugged optimization landscapes. Additional experiments show robustness against noise and a fast computational running time. Experiments on four real-world datasets further demonstrate that our model provides interpretable insights into decision-maker preferences that remain opaque under GP-based approaches.


DualDrift: Combining Forward and Reverse Drifts for One-Step Generative Modeling

Hojung Jung ⋅ Juhyeong Kim ⋅ Jaehyun Kwak ⋅ Boryeong Cho ⋅ Junhyeok Yang ⋅ Youngrok Park ⋅ Sangmin Bae ⋅ Se-Young Yun

Drifting Models have recently achieved state-of-the-art performance in one-step image generation by training generators to follow corrective vector fields. However, their standard reverse-style drift improves samples from the model side, which can severely underutilize training samples and bottleneck early convergence. To address this, we introduce DualDrift, a unified framework that augments the original reverse drift with a novel forward drift. The forward drift lets each data sample assign correction to generated rollouts, providing denser data-side signals and accelerating early training. At the same time, this dense data-side signal can overemphasize dominant regions when used alone. DualDrift therefore uses a parameter-free scheduler driven by empirical batch statistics to balance the fast early correction of forward drift with the stable refinement of reverse drift. Evaluated on ImageNet $256\times256$, DualDrift achieves faster convergence and significantly improves over the baseline Drifting Models in both FID and Inception Score.


Dynamic Convolutions Improve Transformers

Oliver Sieberling ⋅ Bharat Runwal ⋅ Rameswar Panda ⋅ Yoon Kim

Transformers have become the dominant architecture for large language models, largely due to the scalability and flexibility of attention, feed-forward layers, residual connections, and normalization. This paper introduces dynamic short convolutions as an additional neural network primitive for improving Transformer architectures. Unlike static short convolutions, dynamic convolutions use input-dependent filters, preserving the locality bias of convolution while increasing expressivity. Motivating experiments show that applying dynamic short convolutions to key, query, and value representations improves performance on challenging associative recall tasks compared with static convolutional variants. Across language-modeling experiments ranging from 150M to 2B parameters, dynamic convolutions consistently outperform standard Transformers and Transformers augmented with static short convolutions. Scaling-law analysis indicates a 1.33$\times$ compute advantage over parameter-matched Transformers, while a custom Triton kernel enables efficient execution with minimal end-to-end slowdown. These results suggest that dynamic short convolutions are a scalable, hardware-efficient, and expressive primitive for advancing Transformer-based language models.


Dynamic Model Merging Made Slim

Guodong DU ⋅ Wanyu Lin

Model merging combines fine-tuned models into a unified one without joint training or access to original data. Dynamic merging improves flexibility by selectively activating task-relevant parameters, but existing methods either maintain a full shared model with tiny experts or allocate excessive capacity to experts, leading to suboptimal accuracy–efficiency trade-offs. We propose DiDi-Merging, a slim dynamic merging framework that uses differentiable rank allocation to balance shared and per-task expert parameters. The method casts parameter budgeting as differentiable rank optimization over low-rank modules, supervised data-free by the original task vectors, and applies a brief refinement stage to recover task fidelity. Across vision, language, and multimodal benchmarks, DiDi-Merging matches prior dynamic baselines at 1.24× the parameters of a single fine-tuned model and surpasses them at 1.4×, substantially more compact than methods requiring ≥2× storage.

Physical adversarial attacks expose important limitations of modern vision systems, but most existing methods rely on static patches or textures. We introduce a new problem setting: closed-loop physical adversarial control, where an agent interacts with a real-world perception system via a physical channel and receives only detector-level feedback. We optimize LED patterns with PR-SAC in a high-dimensional continuous action space. We show that latent action parameterization with deterministic upsampling improves training stability compared to direct control of individual LEDs. Across YOLOv8, Faster R-CNN, and RetinaNet, the learned patterns reduce person-detection confidence, decrease mAP@0.5, and exhibit partial cross-model transfer. We further analyze robustness under distance and viewpoint changes, interaction with a defense mechanism, and failure modes such as Q-drift, entropy collapse, and degradation after initial improvement. Our results suggest that dynamic physical attacks should be evaluated not only by best-case attack strength, but also by stability, transferability, and robustness under real-world variation.


Dynamic Resolution Routing for Efficient Egocentric Grounding

Huixin Sun ⋅ Wangbo Zhao ⋅ Fanyue Wei ⋅ Qiuxia Lin ⋅ Pengzhan Sun ⋅ Angela Yao

Egocentric visual grounding requires high-resolution inputs to localize small objects. However, scaling Multimodal Large Language Models to this domain is constrained by the excessive cost of visual token processing. We identify that current efficient strategies based on token reduction are unreliable to select object-centric spatial evidence. To overcome this, we propose SmartRes, a framework that performs efficiency optimization in the pixel space via dynamic resolution routing. SmartRes first encodes a low-resolution view for global context and uses a lightweight router to activate high-resolution patches in object-centric regions and constructs an order‑preserving visual sequence. To further enable robust routing under severe foreground‑background imbalance, we introduce a margin-regularized routing objective that increases foreground-background logit separation and improves foreground recall. Experiments on Ego4D and EgoIntention show that SmartRes reduces visual tokens by up to 67\% while retaining 86.4\% of full-resolution performance, and achieves up to $1.66\times$ faster inference than state-of-the-art token reduction methods with higher accuracy. Furthermore, strong performance on small object grounding indicates the effectiveness of SmartRes towards egocentric applications. Code will be publicly available.


EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models

QIWEI ZENG ⋅ Hao Wang ⋅ Jinghao Lin ⋅ Shuchang Ye ⋅ Yuezhe Yang ⋅ Yige Peng ⋅ Haoyuan Che ⋅ Jinman Kim ⋅ Lei Bi

Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and report generation. However, their practical utility remains limited by insufficient sensitivity to subtle lesions, whose visual evidence is often sparse, low-contrast, and embedded within complex anatomical context. As local visual tokens are aggregated, these weak lesion cues can become underrepresented in global image representations, making them difficult for medical VLMs to recognize. Existing efforts to improve lesion sensitivity mainly rely on medical-domain vision-encoder pre-training, clinical-term-guided alignment, or trainable pathological representation enhancement. Although effective, these approaches usually require additional training or model-specific adaptation and may overfit to particular disease morphologies, limiting their applicability to frozen medical VLMs. To address these limitations, we propose \textbf{EasyLens}, a training-free plug-and-play subtle-lesion representation amplifier for medical VLMs. EasyLens first constructs \textbf{EasyBank}, a pathology-anatomy prototype space that provides lesion-related prototypes and anatomy-aware normal references for comparing suspicious patches against both pathological and normal anatomical patterns. To avoid blindly amplifying normal tissues, \textbf{EasyTag} selects lesion-relevant patches through counterfactual prototype reasoning. To counteract the dilution of subtle lesion cues in global image representations, \textbf{EasyAmplifier} strengthens the selected lesion-relevant patch representations through morphology-guided residual enhancement, thereby increasing their contribution to the global image embedding. Experiments on multiple medical image datasets and frozen medical VLM backbones show that EasyLens consistently improves subtle-lesion detection and outperforms existing encoder-enhancement baselines without model fine-tuning. Code is available at: https://anonymous.4open.science/r/easylens-BEC2


EAT: Eviction-Aware Training for Long-Context LLM Inference on Edge Devices

Ke Chen ⋅ Xiaohui Song ⋅ Chi Xie ⋅ ZhiGuang Zhu ⋅ ZIXIAN LI ⋅ Hao-Yun Chen ⋅ Peng-Wen Chen ⋅ Chen Chen ⋅ Haonan Lu

Deploying large language models on edge devices requires mitigating two major bottlenecks: the computational cost of the model and the memory overhead of the KV cache. While combining model quantization and KV-cache eviction seems like a natural solution, we reveal that this naive pipeline suffers from a critical failure mode. The issue stems from a hidden conflict: modern eviction methods rely on attention scores to identify important tokens, yet quantization inherently adds noise to these exact scores. Consequently, quantization does not merely degrade the numerical precision of the KV cache—it actively alters the eviction decisions, often causing the model to discard crucial context. To overcome this non-orthogonal degradation, we propose EAT, a training framework that integrates discrete eviction behavior into quantization-aware training. EAT computes eviction from quantized attention, applies the resulting mask in the forward pass, and uses a masked straight-through estimator in the backward pass. The method requires no architectural changes and adds no inference-time parameters. Systematic evaluations show that EAT mitigates the failure mode of naive quantization-plus-eviction pipelines. Under W4A16 + KV-INT8 with 50% KV eviction, eviction-unaware composition reduces the five-task average from 38.37 to 29.04, whereas EAT recovers it to 39.44 under the same cache budget. Real-world deployment on the MediaTek Dimensity 9500 demonstrates that, under a 50% KV-eviction budget, EAT improves decoding throughput by 64.1% and reduces total memory traffic by 20.4% compared to quantized full-KV inference. Mechanistic analyses suggest that the improvement is associated with reduced attention sink behavior rather than with sharper attention.


ECG Dataset with Multi-Expert Annotations and Delineations

Aram Avetisyan ⋅ Shahane Tigranyan ⋅ Sergey Skorik ⋅ Ariana Asatryan ⋅ Aleksandr Beznosikov ⋅ Yury Markin

The availability of open 12-lead electrocardiogram (ECG) datasets has significantly increased the number of studies focusing on automated ECG analysis. However, existing datasets often have limitations due to unverified annotations, overly processed data, or the absence of features such as fiducial points of the ECG. In this paper, we present a comprehensive collection of ECGs, with each record meticulously annotated by a team of seven cardiologists. Over 2,500 ECG records were annotated, with most records reviewed by at least four physicians. 1,527 ECG records of the dataset contain detailed annotations of PQRST segments, verified by two cardiologists. The dataset constitutes the largest collection of manually delineated ECG records to date, offering a valuable resource for researchers and practitioners in cardiology. This dataset, the first of its kind available for open use with multi-expert annotations, is suited for a variety of research tasks. It can be used for ECG abnormality classification, for analyzing patterns and annotation complexities across different heart abnormalities, and for tasks related to ECG delineation.

Segmenting unseen anatomical structures in echocardiographic images is challenging because dense annotations are scarce and ultrasound often presents visually similar textures across different cardiac structures. Current foundation models remain limited in this setting because they treat anatomical knowledge as independent text prompts, overlooking the spatial topology that clinicians use to localize ambiguous structures from their visible neighbors. We propose Echo-SAM, a zero-shot ultrasound segmentation framework that grounds a structured cardiac knowledge graph in visual space through relation-conditioned geometric support inference. Echo-SAM localizes seen structures as anatomical anchors, propagates relational constraints through a knowledge-grounded Graph Neural Network, and enhances text prompts with image-grounded anatomical context. We further introduce a topology-to-geometry mapping that converts graph relations into geometry-aware support proposal scores, yielding posterior-guided support estimates for unseen structures. By treating relation-specific scale, thickness, and offset as latent geometric variables, Echo-SAM adapts graph-derived support proposals to image-specific anchors without relying on fixed coordinate templates. Extensive evaluations on zero-shot echocardiography segmentation benchmarks demonstrate that Echo-SAM substantially outperforms state-of-the-art foundation models, achieving up to 50.42\% absolute improvement in Dice score. Code will be publicly available.


EDEN: Emergent Dynamics in Evolutionary Neural-networks for Robust Continuous Control

Chi Zhang ⋅ Jinge Li ⋅ Yifei Wang ⋅ Lei Wang ⋅ Tianyi Qian

Biological neural systems leverage neuronal heterogeneity and diverse firing patterns to generate rich nonlinear dynamics, yet these mechanisms are largely abstracted away in conventional artificial neural networks (ANNs) and simplified spiking neural networks (SNNs). Detailed differential-equation neural models are biologically expressive, but spike discontinuities and complex state evolution pose challenges for gradient-based optimization in complex functional tasks. To address this challenge, we introduce EDEN (Emergent Dynamics in Evolutionary Neural-Networks), a neural dynamic model training framework inspired by natural evolutionary processes. Built upon a continuous-time network of heterogeneous Izhikevich neurons, EDEN integrates efficient parallel neural dynamics simulation with advanced evolutionary strategies, enabling the joint optimization of intrinsic single-neuron-level parameters and network-level synaptic connections without backpropagation. Comprehensive evaluations across multiple MuJoCo continuous motor control tasks indicate that EDEN achieves robust control performance comparable to mainstream deep reinforcement learning baselines. Furthermore, neurodynamic analysis reveals that the trained EDEN models exhibit strong biological plausibility, characterized by functional differentiation across individual neurons and emergent low-dimensional latent manifolds governing population dynamics. EDEN demonstrates that the combination of gradient-free evolutionary strategies and biological neurodynamic networks provides an alternative yet efficient paradigm alongside mainstream reinforcement learning, thereby offering a highly promising pathway towards truly biologically-inspired intelligence.


Edge of Stability Selectively Shapes Learning Across the Data Distribution

Shauna Kwag ⋅ Anakha Ganesh ⋅ Tomaso Poggio ⋅ Pierfrancesco Beneventano

Existing analyses of the edge of stability (EoS) treat it as a global property of optimization. We show that it is also selective: the stability constraint redistributes learning across subsets of the training distribution, amplifying progress on some groups while suppressing progress on others. Using a branching intervention that enters or exits the EoS regime from the same training state, we causally demonstrate this trade-off and identify two necessary conditions for a group to benefit. First, its aggregate gradient must align with the top Hessian eigenvector. We isolate this mechanism with a controlled perturbation that preserves distance but randomizes direction, destroying alignment and eliminating the advantage. Second, the group must sustain non-vanishing gradient magnitude over time. Under cross-entropy loss, gradient saturation decouples confidently classified groups, shifting the advantage to output-outliers, whose gradients persist. Together, these results show that EoS functions not only as a stability boundary, but as a mechanism governing the allocation of learning across the data distribution.


EditBridge: Towards Faithful and Efficient Ultra-High-Resolution Image Editing

Jiayi Song ⋅ Shijie Huang ⋅ Fangtai Wu ⋅ Yubo Huang ⋅ Zhenxiong Tan ⋅ Songhua Liu ⋅ Jiaming Liu ⋅ Ruihua Huang

High-resolution image editing is increasingly demanded in professional workflows, yet existing diffusion-based models remain constrained to resolutions below 1K due to quadratic attention complexity and prohibitive memory requirements. A prevalent workaround employs a two-stage pipeline: editing at low resolution followed by independent super-resolution. However, this approach suffers from two critical issues: information divergence, where hallucinated details contradict the original high-resolution (HR) source, and texture degradation, manifesting as over-smoothed or over-sharpened artifacts. We propose EditBridge, a diffusion bridge framework for efficient ultra high-resolution editing. Unlike conventional diffusion that regenerates from noise, we formulate refinement as structured data-to-data translation from the low-resolution (LR) edited result to its HR counterpart, explicitly conditioned on the original HR source to preserve authentic details. To efficiently incorporate HR source guidance, we introduce a prior-guided block-wise sparse attention mechanism that exploits semantic correspondence from first-stage editing to constrain cross-image interactions to spatially aligned regions, significantly reducing computational overhead. Extensive experiments demonstrate that EditBridge achieves high-fidelity editing with superior perceptual quality at resolutions up to 4K, delivering 3.6--8.4$\times$ speedup at 2K and enabling practical 4K editing in 61 seconds.


Editing Large Language Models with Geometry-Aware Regularization

Bingqing Liu ⋅ Wei Liu ⋅ Jun Wang ⋅ Xiaobo sun ⋅ Yuhua Li ⋅ Ruixuan Li

Knowledge editing has become a key technique for updating knowledge in large language models. However, existing methods neglect the geometric structure of the knowledge representation space, leading to two key limitations: they cannot distinguish which pre-trained knowledge should be preferentially preserved (core vs. non-core) nor which new edits are inherently difficult (hard vs. easy). To address this, we propose \textsc{GEdit}, a \textit{geometry-aware} framework built on a Low-Rank Plus Diagonal (LR+D) factor model that captures the intrinsic low-rank structure of Transformer FFN hidden representations. \textsc{GEdit} features two complementary mechanisms. First, the \textit{low-rank regularization} applies targeted protection on a low-rank core knowledge subspace. Second, the \textit{adaptive regularization} dynamically adjusts the regularization strength based on each edit request's editability: relaxing constraints for hard edits to ensure success, and tightening them for easy edits to better preserve pretrained knowledge. Extensive experiments on \texttt{GPT2-XL}, \texttt{GPT-J}, and \texttt{LLaMA-3-8B} under large-scale editing scenarios—including both mini-batch and fully sequential editing settings with up to 10,000 continuous updates—show that \textsc{GEdit} consistently and significantly outperforms state-of-the-art baselines in both editing success and knowledge preservation.


Edit the Bits, Diff the Codes: Bitwise Residual Editing for Visual Autoregressive Models

Shengqiang Zhang ⋅ Ruotong Liao ⋅ Volker Tresp ⋅ Barbara Plank ⋅ Hinrich Schuetze

Text-guided image editing with visual autoregressive (VAR) generators requires controlling both what the model samples and where the sampled change is written back into the image. Existing VAR editors mainly operate on token streams, features, or flat next-token logits, leaving two native structures of bitwise-residual VAR models underused: the per-bit Bernoulli prediction head and the additive multi-scale residual code field. We propose BitResEdit, a training-free editor for bitwise-residual VAR generators such as \textsc{Infinity}. BitEdit performs source-negative guidance by tilting the post-CFG per-bit log-odds along a source--target contrast computed on a shared edited prefix, then projects each update into a closed-form Bernoulli-KL trust region around the clean CFG sampler. ResEdit converts the resulting sampled bits into per-scale continuous-code residuals, gates them with a localization mask, and re-injects them through the generator's native sum-of-scales. The two components couple decision-time bit guidance with combination-time code composition, preserving masked-out latent features exactly while applying localized, scale-aware edits inside the target region. On PIE-Bench with Infinity-2B, BitResEdit achieves the strongest text alignment among same-backbone VAR editors and competitive background preservation, setting a new state of the art among training-free VAR editing methods. Ablations show that BitEdit and ResEdit play complementary roles in target alignment and background preservation. Detailed ablations show that both code-space residual editing and bitwise phrase-contrast guidance are necessary for robust autoregressive image editing.


EEG Benchmarking Needs a Task Specification Layer: NeuroDoc for Rulebook-Guided, Executable Benchmark Construction

Chengxuan Qin ⋅ 致格 陈 ⋅ Pengshu ⋅ Rui Yang ⋅ JiPing Cui ⋅ Yikai Dong ⋅ Jun Li ⋅ Liu Peng ⋅ Zhida Shang ⋅ Mingze Tang ⋅ KC Tan ⋅ Jibin Wu

Electroencephalography (EEG) foundation models increasingly rely on multi-dataset training and evaluation, yet public EEG datasets still lack a shared task specification layer that can turn heterogeneous recordings into reusable benchmark units. Existing standards organize files, metadata, and provenance, but they do not specify EEG tasks under a common language and rulebook, leaving critical task semantics scattered across papers, code, and manual interpretation. We investigate whether heterogeneous public EEG datasets can be standardized through a structured task specification language paired with a shared rulebook. Our methodology represents each benchmark entry as a task document synchronized with an executable task kernel, with the rulebook defining task fields, evidence requirements, document-kernel alignment, review states, and machine-checkable constraints. Using this methodology, we release a community-reviewed EEG benchmark corpus centered on 53 completed and reviewed entries with 245 task definitions spanning diverse paradigms, and we introduce NeuroDoc and NeuroAudit as the operational support layer for rulebook-guided drafting, upgrading, review, amendment, and release management. We further examine whether the resulting benchmark units can be instantiated in a shared downstream setting across four EEG foundation model backbones, providing execution-based evidence for reusable, auditable, and executable EEG benchmarking infrastructure. Project resources are available at https://anonymous.4open.science/r/anonymousrevieweneurodoc.


Effective Knowledge Conflict Detection via Joint Agentic Optimization

Yi Hong ⋅ Wenchao Bai ⋅ Jiaqi Jiang ⋅ Siyuan Huang ⋅ Jiahui Jin

Knowledge conflict detection is a fundamental challenge for large language model (LLM) systems that rely on external knowledge. Existing approaches address this problem either with end-to-end supervised models or with multi-stage pipelines dependent on fixed heuristics. However, supervised models often generalize poorly to diverse conflict patterns, whereas heuristic pipelines rely on fixed rules and introduce considerable computational overhead. To overcome these limitations, we propose $\textbf{G}$rouping and $\textbf{C}$onflict $\textbf{D}$etection ($\mathsf{GCD}$), a two-stage agentic framework for knowledge conflict detection that aligns the detection process with the inherent sparsity of knowledge conflicts, thereby reducing the decision space and enabling fine-grained conflict detection. $\mathsf{GCD}$ decomposes knowledge conflict detection into two subtasks and employs two cooperative agents to address them: a Knowledge Grouping Agent (G-Agent), which creates conflict-preserving groups of textual units with potential conflicts, and an Intra-group Conflict Detection Agent (D-Agent), which detects conflicts within each group. The G-Agent's grouping policy is not fixed in advance but is jointly trained with D-Agent via a shared global reward, enabling grouping to be explicitly optimized for accurate detection rather than relying on fixed heuristics. Extensive experiments on five benchmarks demonstrate that $\mathsf{GCD}$ consistently outperforms SOTA methods, while reducing inference latency and token consumption.


Efficient and Accurate Zero Shot Generation of Symmetric Protein Complexes

Rory Gao ⋅ Yuanzhou Chen ⋅ Prajit Rajkumar ⋅ Wei Wang

Symmetric organization is a fundamental principle of biomolecular complex assembly. The same principle underlies engineered applications, from vaccine platforms to nanomaterials, making symmetric \textit{de novo} generation a core capability for computational protein design. We establish a theoretical foundation equating generation of an assembly with underlying symmetry group $G$ to generation within a quotient space over $G$. This directly allows $\mathrm{SE}(3)$-equivariant models to generate symmetric structures with minimal distribution shift. With this insight we introduce \textbf{\ours{}} (\textbf{Ze}ro-shot \textbf{U}nified \textbf{S}ymmetrization), a framework that converts any pretrained $\mathrm{SE}(3)$-equivariant backbone generative model into an efficient generator of symmetric assemblies with no retraining. \ours{} operates on a single \emph{canonical subunit} and utilize \emph{phantom subunits} to preserve internal model operations, reducing memory and runtime cost by $\mathcal{O}(|G|)$ for finite groups $G$, compared to a full-complex forward pass. We implement \ours{} for IPA-based models (FrameDiff, FoldFlow) and axial-attention-based RFdiffusion, showing high generation efficiency with minimal degradation to generation quality. We then apply \ours{} to generate symmetric complexes of over $10{,}000$ residues with exact, untruncated attention, a scale previously intractable with any model.

Efficient benchmarking techniques aim to lower the computational cost of evaluating LLMs by predicting full benchmark scores using only a subset of a benchmark's questions. By reframing this problem as an instance of multiple regression with feature selection, we find that existing efficient benchmarking methods can be greatly improved by simply using kernel ridge regression at the prediction stage. Additionally, using an information-theoretic feature-selection algorithm called minimum redundancy maximum relevance (mRMR), we can further improve upon these methods by selecting question subsets that will be maximally useful for prediction. Except in very data-poor settings, these approaches consistently achieve smaller prediction errors (in both MAE and RMSE), and greater ranking correlation between predicted and true scores (in both Spearman $\rho$ and Kendall $\tau$) across a range of benchmarks using both binary and continuous metrics. Furthermore, mRMR subsampling is much faster than competitor methods (which often involve fitting probabilistic models or running clustering algorithms), and is more likely to select the same questions under different random seeds or training data splits.


Efficient bias mitigation in T2I diffusion models using Concept Graphs

Mansi - ⋅ Avinash Kori ⋅ Francesco Leofante

Text-to-Image diffusion models often propagate harmful bias inherited from the training data. Existing bias mitigation techniques either intervene only at the text encoder or provide inference-time guidance, often leading to generations that collapse into semantically incoherent outputs. To address these limitations, we introduce CO-ALIGN (Concept Ontology Alignment), a novel bias mitigation approach based on concept-graph alignment which operates on the model's internal concept ontology. By aligning concepts within the text encoder and denoiser, CO-ALIGN achieves significant bias reduction while preserving generative integrity. We demonstrate the effectiveness of concept graph alignment in three paradigms- text-encoders, denoisers, and joint text-denoiser ontology alignment. CO-ALIGN outperforms outperforms the state of the art, improving fairness by 30\%, $\Delta FID=11.4$ in image quality, 2.8\% in image fidelity, all while reducing semantically incoherent outputs by 88\%. Beyond bias mitigation, we show that CO-ALIGN benefits other downstream tasks as well. In particular, our experiments demonstrate that better-aligned internal ontologies enhance concept unlearning robustness across multiple unlearning techniques.


Efficient Evaluation of LLM Performance with Statistical Guarantees

Skyler Wu ⋅ Yash Nair ⋅ Emmanuel Candes

Exhaustively evaluating many large language models (LLMs) on a large suite of benchmarks is expensive. We cast benchmarking as finite-population inference and, under a fixed query budget, seek tight confidence intervals (CIs) for model accuracy with valid frequentist coverage. We propose *Factorized Active Querying* (FAQ), which (a) leverages historical information through a Bayesian factor model; (b) adaptively selects questions using a hybrid variance-reduction/active-learning sampling policy; and (c) maintains validity through *Pro-Active Inference*---a finite-population extension of active inference (Zrnic & Candès, 2024) that enables direct question selection while preserving coverage. With negligible overhead cost, FAQ delivers up to $5\times$ effective sample size gains over strong baselines on two benchmark suites, across varying historical-data missingness levels: this means that it matches the CI width of uniform sampling while using up to $5\times$ fewer queries. We release our source code and our curated datasets to support reproducible evaluation and future research.


Efficient Forecasting of Task Failures in LLM Agents through Adaptive Fault Injection

Vartika Sengar ⋅ Parth Thakkar ⋅ Pranoy Panda ⋅ Shrey Satapara ⋅ Emmy Liu ⋅ Vijay Viswanathan ⋅ Sho Takemori ⋅ Graham Neubig ⋅ Chaitanya Devaguptapu

LLM agents often execute long-horizon tasks where sequential tool calls and reasoning steps compound into a final outcome. When an agent errs mid-execution, the critical question is not whether an error occurred, but whether recovery under the current policy is still possible. Existing failure analysis answers this retrospectively, making it too late for runtime intervention. We introduce Task Failure Forecasting: predicting from a partial execution trace whether continued execution under the current agent policy is likely to fail. This is hard: human experts achieve only 54\% accuracy at distinguishing recoverable from failure-inducing errors, and frontier models such as Claude-4.5 Sonnet reach 59.87\% weighted F1. The core obstacle is that failed traces mix recoverable mistakes with fatal ones, obscuring which steps drive downstream failure. We address this with adaptive fault injection: injects targeted perturbations into successful traces, uses bandit-based prioritization of high failure yield error types and rollout-based verification providing policy-conditioned step-level supervision. A lightweight Qwen-3-8B model trained on this synthetic data achieves 78.01\% weighted F1 on human-annotated benchmarks, outperforming Claude-4.5 Sonnet by 14 weighted-F1 points at 100$\times$ lower forecasting cost. When deployed as a runtime monitor for selective inference-time intervention, the forecaster improves agent success rates across HotpotQA, MuSiQue, GAIA, AssistantBench, SWE-Bench, MBPP and EnterpriseBench, yielding absolute gains in the range of 2–16 \% points across held-out and unseen tasks.

Learning the natural parameters $z \in \mathbb{R}^n$ of discrete distributions $\mu_z$ from independent samples constrained to a subset $S \subseteq \\{0,1\\}^n$ is a foundational challenge in high-dimensional statistics. Existing methods for efficiently estimating truncated Boolean product distributions, notably the work of [Fotakis et al' COLT'20, Algorithmica '22], require either strong local connectivity assumptions on $S$ -- a property denoted *fatness* -- or stringent anti-concentration assumptions and necessitate the total mass of the truncation set to be a constant with respect to $n$. Moreover, the results in [Fotakis et al' COLT'20, Algorithmica '22] suffer from sample complexities that scale as $\Omega(2^n)$ if the mass of $S$ is exponentially small in $n$. In this work, we circumvent these limitations by analyzing the geometry of $S$ under the measure $\mu_z$. We refine the existing parameter estimation guarantees under the fatness assumption, improving the prior sample complexity to $\mathcal{O}( \log n / \epsilon^2)$ for $\ell_\infty$-recovery, matching the untruncated minimax rate. We further generalize fatness using the notion of influence utilized in the analysis of Boolean functions and provide sufficient conditions for efficient inference. Notably, unlike previous work, our method does not require sampling at arbitrary parameterizations of the model and instead relies on gradient descent. Lastly, we establish a theoretical lower bound demonstrating the sample complexity exhibits an intrinsic exponential dependence on the width of the model and the minimum distance between elements in the set.


Efficient Off-Policy RL for Video Generation via Forward-Consistent Reward Matching

Hongzheng Yang ⋅ Mengyang LIU ⋅ Haoxuan Wu ⋅ Kun Li ⋅ Yuzhi Zhao ⋅ Wei Liu

Reinforcement learning (RL) post-training aligns diffusion-based generators with human preferences, yet existing RL methods suffer from poor compatibility with off-policy learning and few-step distilled models. These limitations are especially severe in the video generation area, as practical video generation pipelines often rely on few-step distilled generators. Furthermore, due to complex spatial-temporal dynamics and higher dimensions, near-on-policy video rollouts are both expensive to collect and often imperfect. Relying on such rollouts alone can amplify artifacts and is prone to reward hacking. To address these issues, we propose Forward-Consistent Reward Matching (FCRM), an efficient off-policy RL framework for video generation. FCRM converts the forward denoising loss into a positive loss-induced score and formulates the reward alignment as a one-step GFlowNet matching problem. The resulting residual is pointwise in a clean sample space that naturally supports off-policy learning and few-step generators. To avoid biased gradients, we introduce a double-sampling estimator for the squared residual objective. Theoretically, minimizing the proposed matching residual bounds the KL divergence between the learned distribution and the optimal reward-tilted distribution. Experiments on standard video generation benchmarks validate FCRM across online, replay, offline, and few-step settings and outperform SOTA methods.


Efficient One-Step Diffusion Restoration Model with Compact Token Compression and Linear Attention

Bingtian Qiao ⋅ Yue Shi ⋅ Yingjie Zhou ⋅ Yong Guo ⋅ Guangtao Zhai ⋅ Jiezhang Cao

Real-world image super-resolution aims to recover high-quality images from complex and unknown real-world degradations. However, existing generative Real-ISR methods largely inherit the dense latent representations and quadratic-cost global modeling paradigm developed for high-resolution image synthesis, causing computation, memory usage, and inference latency to scale unfavorably with resolution and thus limiting practical deployment. We argue that the key bottleneck lies not in insufficient restoration priors, but in excessive token redundancy and costly token interactions during high-resolution restoration. Motivated by this observation, we revisit Real-ISR from the perspectives of compact latent representation and linear-complexity modeling, and propose SANA-SR, an efficient one-step restoration framework. Specifically, SANA-SR employs a deep compression autoencoder with a 32× compression ratio to drastically reduce latent tokens while preserving restoration-relevant structures and textures. On top of this compact latent space, we introduce a linear-attention DiT with LoRA fine-tuning, enabling efficient high-resolution restoration with linear-complexity token mixing. Extensive experiments on all benchmark datasets demonstrate that SANA-SR achieves highly competitive and often superior quantitative performance against existing methods, while restoring clearer and more realistic textures. Moreover, after pruning, the deployed model runs in 0.019 s with 407.95G MACs and 344M parameters, highlighting its strong potential for practical mobile deployment.


Efficient Serving for Dynamic Agent Workflows with Prediction-based KV-Cache Management

Haoyu Zheng ⋅ Fangcheng Fu ⋅ Jia Wu ⋅ Binhang Yuan ⋅ Yongqiang Zhang ⋅ Hao Wang ⋅ Yuanyuan Zhu ⋅ Xiao Yan ⋅ Jiawei Jiang

LLM-based workflows compose specialized agents to execute complex tasks, and these agents usually share substantial context, allowing KV-Cache reuse to save computation. Existing approaches either manage KV-Cache at agent level and fail to exploit the reuse opportunities within workflows, or manage cache at the workflow level but assume that each workflow calls a static sequence of agents. However, practical workflows are typically dynamic, where the sequence of invoked agents and thus induced cache reuse opportunities depend on the context of each task. To serve such dynamic workflows efficiently, we build a system dubbed PBKV (\textbf{P}rediction-\textbf{B}ased \textbf{KV}-Cache Management). For each workflow, PBKV predicts the agent invocations in several future steps by fusing the guidance from historical workflows and context of the target workflow. Based on the predictions, PBKV estimates the reuse potential of cache entries and keeps the high-potential entries in GPU memory. To be robust to prediction errors, PBKV utilizes the predictions conservatively during both cache eviction and prefetching. Experiments on three workflow benchmarks show that PBKV achieves up to $1.85\times$ speedup over LRU on dynamic workflows, and up to $1.26\times$ speedup over the SOTA baseline KVFlow on the static workflow.


EgoTac: In-the-wild Tactile Prediction from Egocentric Vision

Wenkang Zhang ⋅ Chengbo Yuan ⋅ Zicheng Zhang ⋅ Zhengxue Cheng ⋅ Yang Gao

Touch is fundamental to dexterous manipulation, yet most egocentric human data increasingly used for robot learning lacks tactile information. Directly collecting large-scale tactile data is challenging due to sensor limitations, while human video data is abundant, contact-rich, and easily scalable. This motivates a natural question: can tactile signals be inferred purely from vision? To address this, we introduce EgoTac, a generalizable model that predicts rich tactile information directly from egocentric human videos. EgoTac is trained on a unified corpus of over 5.7M image-tactile pairs, covering both continuous force measurements and binary contacts. By learning from this diverse dataset, EgoTac captures nuanced touch dynamics across varied interactions. Experiments demonstrate strong performance: in-domain prediction achieves an average force error below 0.06N. On out-of-domain contact prediction benchmarks, EgoTac consistently outperforms the state-of-the-art contact estimator. It also captures the rise and fall patterns of real tactile data and enables zero-shot predictions on unconstrained real-world videos. Scaling analyses further reveal that both data diversity and volume improve performance steadily. Overall, EgoTac provides a scalable pathway to extract tactile priors from egocentric human videos, enabling broadly applicable tactile-aware robot learning.


EIHMR: Collaborative Human-Camera Estimation for Global Human Mesh Recovery

Junchen Ge ⋅ Zhengqi Zhang ⋅ Hanglei Jin ⋅ Shuzhao Xie ⋅ Yuzhihuang ⋅ Jingwei Xu ⋅ Jingyan Jiang ⋅ Zhi Wang

Recovering global 3D human motion from monocular video captured by a moving camera is a fundamental yet challenging problem, as camera ego-motion and human body motion are tightly entangled in the image observations. The prevailing two-stage paradigm treats camera estimation and motion reconstruction as isolated processes, causing errors on both sides to be further amplified when combined in the world coordinate system. To address this, we draw inspiration from the human inner visual simulation mechanism and propose EIHMR, a collaborative human-camera co-estimation framework. EIHMR comprises two complementary modules that bridge scene-aware human motion refinement and motion-aware camera estimation: Scene-Aware Local Human Motion Reconstruction reprojects motion sequences into frozen keyframe viewpoints and leverages metric depth and kinematic constraints to produce geometrically consistent local motion, while Motion-aware SLAM re-renders the refined motion as static meshes in the original frames, converting dynamic human regions into structured matching cues for robust camera estimation. EIHMR consistently improves global trajectory reconstruction over strong baselines, demonstrating the effectiveness of collaborative human-camera estimation for long-range human motion recovery.


Eliciting Zero-Shot Named Entity Recognition in Large Language Models via Instruction Semantic Elaboration

Kaiyin Zhou ⋅ zhiyi Song ⋅ Xiangling Fu ⋅ Chenwei Yan ⋅ Xinxin You ⋅ Shaohui Liu ⋅ Ji Wu ⋅ Xien Liu

Large Language Models (LLMs) have demonstrated remarkable capabilities in Information Extraction (IE) tasks; however, their performance in zero-shot Named Entity Recognition (NER) remains suboptimal. We argue that this limitation stems not from intrinsic defects within the models, but rather from the failure of existing instruction paradigms to effectively elicit their latent capabilities. Drawing inspiration from cognitive science, we propose an instruction optimization framework termed Instruction Semantic Elaboration (ISE), which fully elicits models' latent NER capabilities by mimicking the semantic elaboration process—specifically by contrasting related concepts, decomposing constituent elements, and illustrating with concrete examples. We evaluate our framework under a zero-shot setting across four domain-specific datasets. Results demonstrate that ISE effectively and stably unlocks the intrinsic NER capabilities of LLMs, achieving a 10.99\% improvement in F1 score over initial instructions and outperforming strong baselines by 6.89\%. Furthermore, the framework generalizes well across model architectures and parameter scales, while remaining robust to the quality of initial instructions. This study offers a novel perspective on activating NER capabilities in LLMs, facilitating their efficient and stable deployment in real-world scenarios.


Embeddings for Preferences, Not Semantics

Carter Blair ⋅ Ariel Procaccia ⋅ Milind Tambe

Modern AI is opening the door to collective decision-making in which participants express their views as free-form text rather than voting on a fixed set of candidates. A natural idea is to embed these opinions in a vector space so that the substantial literature on facility location problems and fair clustering can be brought to bear. But standard text embeddings measure semantic similarity, whereas distances in facility location problems and fair clustering require what we call preferential similarity: a participant's agreement with a piece of text should be inversely related to their distance from it. Off-the-shelf embeddings inherit a coarse preference signal through a correlation between semantic and preferential similarity, but fail to capture preferences when the correlation breaks. We formalize this as an invariance problem: text embedding models encode both a preference-relevant signal (stance and values) and semantic nuisance (style and wording), and the two are observationally correlated, so a geometry that relies on nuisance can appear preference-correct even when it is not. We show that synthetic training data designed to break this correlation provably shifts the optimal scorer away from nuisance-dominated cosine and significantly improves preference prediction across 11 online deliberation datasets.

Intrinsic Motivation (IM) aims to train agents without external rewards, enabling useful behavior to emerge from the agent's interaction with its environment alone. However, the dominant IM approaches rely on information-theoretic quantities with designer-chosen variables, introducing bias and lacking a principled connection to dynamics or optimal control (OC). We introduce Controllable Information Production (CIP), a new foundation for IM explicitly grounded in dynamical systems and OC. CIP measures the rate at which an agent produces information, capturing controllable complexity without external knowledge or bias. CIP unifies IM and OC into a single framework, formalizing physical intelligence as the control of information production. It further reveals connections between the structure of the value function and Kolmogorov–Sinai entropy. CIP consistently outperforms prior IM methods on standard benchmarks in robot learning and solves tasks they fail on, including humanoid self-righting. These results support a general organizing principle: physical intelligence emerges from driving systems toward the edge of controllable chaos.

Humans have a remarkable ability to judge causal relationships from a limited number of unreliable observations. Past work on causal cognition has largely focused on normative accounts of human behavior, leaving unknown how biologically plausible neural systems could learn causal relationships from observations and update their representations of causal structure with additional evidence. Here, we leverage task-optimized recurrent neural networks to discover candidate implementation-level neural mechanisms of causal judgment. We propose a novel cognitive task in which a subject observes stochastic samples from an unknown causal structure (e.g. among variables A, B, and C with unknown causal relationships), and must judge whether a specific causal relationship is present given a query (e.g. "does A cause C?"). We found that, after training, recurrent neural networks perform the task with high accuracy, adopt strategies that incorporate the behavior of non-queried variables to form their judgments, and, despite being trained only on pairwise queries ("does A cause B?", "does C cause A?", etc), form implicit beliefs about the complete graphical structure underlying the observations. Lastly, we use dynamical systems analysis to identify a set of low-level neural mechanisms that implement causal judgment and representation of causal graphical structure. Together, these findings lay the groundwork for a "bottom-up" approach to causal cognition, providing a potential basis for subsequent experimental study in the brain.


Emotion-Trained Vision Models Do Not Necessarily Learn EEG-Aligned Facial Dynamics

Maryem Benslimane ⋅ Hamza Abdelhedi ⋅ Vanessa Hadid ⋅ Karim Jerbi

Facial expressions are one of the main ways humans communicate affect, but they are not static signals. In natural vision, expressions evolve over time, with facial movements emerging, intensifying, and changing as emotion is expressed. Prior model-brain studies based on static images suggest that vision models can capture aspects of face-related neural representations, but it remains unclear whether artificial vision models capture the temporal EEG dynamics of facial emotion perception. We test the hypothesis that models trained for facial emotion recognition, especially video-based or temporally structured models, should show stronger alignment with human EEG dynamics than randomly initialized controls. We do so by aligning model representations with time-resolved EEG representational geometry from 25 participants viewing one-second dynamic facial-expression videos. We compare six models spanning static-image, video-based, self-supervised, and vision-language representations: ResNet18, VGGFace, CNN2D, CNN3D, DINOv2-temporal, and Qwen2.5-VL. Contrary to this hypothesis, supervised facial-emotion training does not consistently increase alignment with human EEG dynamics. In particular, video-trained and temporally structured emotion models align less with fear-related EEG geometry than the same architectures with randomly initialized weights, whereas static-image emotion models show trained-versus-random differences close to zero. These results are robust to stricter noise-ceiling thresholds and are not explained by actor identity or raw pixel similarity. Exploratory analyses further suggest that peak alignment is often localized to the first or last sampled video frames, and that fear alignment varies across scalp ROIs by training regime. Overall, our results challenge the assumption that video-based emotion supervision automatically yields more brain-like dynamic representations of facial emotion.


Empowering Time Series Analysis with Large-Scale Multimodal Pretraining

Peng Chen ⋅ Siyuan Wang ⋅ hu ⋅ Xingjian Wu ⋅ Yang Shu ⋅ Zhongwen Rao ⋅ Meng Wang ⋅ Yijie Li ⋅ Bin Yang ⋅ Chenjuan Guo

While existing time series foundation models primarily rely on large-scale unimodal pretraining, they lack complementary modalities to enhance time series understanding. Building multimodal foundation models is a natural next step, but it faces key challenges: 1) lack of a unified multimodal pretraining paradigm and large-scale multimodal corpora for time series analysis; 2) how to effectively integrate heterogeneous modalities and enhance model generalization. To address these challenges, we take an early step toward multimodal foundation models for time series analysis. We first propose a multimodal pretraining paradigm that leverages time series with endogenous modalities (derived images and text) and exogenous knowledge (news), providing a comprehensive multi-view perspective for time series analysis. To support this, we develop an automated data construction pipeline to curate MM-TS, the first large-scale multimodal time series dataset spanning six domains, with up to one billion points. Then we propose HORAI, a frequency-enhanced multimodal foundation model. It integrates two core components: the Frequency-enhanced Cross-Modality Encoder and the Time-Frequency Decoder, designed to effectively fuse multimodal features and enhance model generalization across modalities and domains. After pretraining on MM-TS, HORAI achieves state-of-the-art zero-shot performance on time series forecasting and anomaly detection tasks, demonstrating strong generalization.


Encoding RNA Topology into Synthetic Alignments for 3D Structure Prediction

Jörg Franke ⋅ Iris K Tennie ⋅ Michael Uhl ⋅ Ryan Koksal ⋅ Dominika Matus ⋅ Dominik Scheuer ⋅ Frank Hutter ⋅ Rolf Backofen ⋅ Frederic Runge

Modern 3D structure predictors rely on multiple sequence alignments (MSA) to expose evolutionary covariation, but for RNA those alignments are often missing, shallow, or unreliable. To overcome this limitation, we invert the co-evolutionary structure inference pipeline and encode a predicted 2D topology into synthetic alignment-form covariation, without requiring database search nor alignment steps. We introduce RNAformer, a transformer-based RNA secondary-structure predictor that supplies the topology, and a Synthetic Co-evolution Engine (SCE) that encodes it into synthetic homologous sequences (SHS) by co-mutating paired positions. Comprehensive experiments and systematic interventions across multiple structure predictors show that the MSA input acts as a topology channel that allows predictable steering of 3D structure prediction. Across single-chain RNAs and single-chain RNA-protein complexes, RNAformer-seeded SHS approach the performance of natural MSAs, with regime-dependent gains. Topology, not alignment depth, drives the response. SHS are interpretable, fast, and scalable, and represent a first step toward an alignment-free alternative to natural MSA.


End-to-End Identifiable and Consistent Recurrent Switching Dynamical Systems

Carles Balsells Rodas ⋅ Zhengrui Xiang ⋅ Francisco Sumba Toral ⋅ Yingzhen Li

Learning identifiable representations in deep generative models remains a fundamental challenge, particularly for sequential data with regime-switching dynamics. Existing approaches establish identifiability under restrictive assumptions, such as stationarity or limited emission models, and typically rely on variational autoencoder (VAE) estimators, which introduce approximation gaps that limit the recovery of the latent structure. In this work, we address both the theoretical and practical limitations of this setting. First, we establish identifiability of a broad class of recurrent nonlinear switching dynamical systems under flexible assumptions, significantly extending prior results. Second, we introduce $\Omega$SDS, a flow-based estimator that enables exact likelihood optimization using expectation-maximisation. Through empirical validation on both synthetic and real-world data, our results demonstrate that $\Omega$SDS achieves improved disentanglement compared to VAE-based estimators and more accurate forecasting of underlying dynamics.


Energy-Based Operator Learning in Function Space

Hong Chul Nam ⋅ Jeonghwan Cheon ⋅ Jae M Shin ⋅ Hyun Kwon

We propose Energy-Based Operators (EBOs), an architecture-agnostic framework for learning conditional distributions over functions on continuous domains. An EBO defines a scalar energy over target functions given an input function and induces a probability model through a Gaussian reference. The resulting score field is obtained as the gradient of the parametrized energy to perform function-space iterative energy minimization (EM). Our model shows strong performance in function generation, such as super-resolution and forecasting, over various 1D function classes (oscillations, damping, and Izhikevich) in comparison with prediction operators and denoising operators over various architecture backbones. Moreover, it achieves strong performance over PDEs, namely Navier-Stokes, Darcy flow and Burgers. Notably, our model successfully detects anomalous functions by automatically assigning high energy without any supervision. It enables seizure detection and volatility prediction after learning neural dynamics and market microstructure dynamics without pre-defined labels during training, highlighting its effectiveness for both learning dynamical systems and detecting functional anomalies arising in scientific simulations.


ENGINE: Endogenous Variational MultiScale Optimization for Zeroth-Order LLM Fine-Tuning

Zhuoli Ouyang ⋅ Changxi Chi ⋅ Siyuan Li ⋅ Tailin Wu

Zeroth-order (ZO) optimization enables memory-efficient large language model (LLM) fine-tuning, with recent methods employing curvature-aware strategies. However, in the massive LLM parameter space, these methods inherently suffer from exploding gradient variance and prohibitive costs. Conversely, confining updates to a low-dimensional subspace reduces variance but discards critical high-dimensional information, degrading performance. This creates a fundamental dilemma: full-space optimization is informative but high-variance, while subspace optimization is low-variance but blind. To break this trade-off, we propose \textbf{\Engine} (\textbf{En}do\textbf{g}enous Var\textbf{i}atio\textbf{n}al MultiScal\textbf{e}). Inspired by computational mechanics, we adapt static Variational MultiScale theory to dynamic optimization, treating low- and high-dimensional spaces as resolved coarse and unresolved fine scales. Through a multigrid-style V-cycle, \Engine\ alternates between spaces, dynamically probing fine-grained information in the high-dimensional space and injecting it to correct the efficient low-dimensional sampling process. Evaluations on standard LLM benchmarks show that \Engine\ achieves higher accuracy and faster convergence than existing ZO methods, achieving up to a $\mathbf{2.61\times}$ \textbf{speed-up in forward passes} with zero additional peak memory overhead. Code is available at \url{https://anonymous.4open.science/r/ENGINE-D0A4}.


Enhancing Gaze Reasoning in Vision Foundation Models for Gaze Following

SHIJING WANG ⋅ Yihua Cheng ⋅ Chaoqun Cui ⋅ Yaping Huang ⋅ David Wong ⋅ Alexandros Neophytou ⋅ Hyung Jin Chang

Gaze following requires both scene understanding and gaze reasoning to localize the gaze target of an in-scene person. Recently, vision foundation models (VFMs) have demonstrated strong performance on this task, enabling simpler architectures while outperforming prior methods. However, we observe a key limitation of VFM-based approaches: while VFMs substantially improve scene understanding, they contribute little to gaze reasoning. As a result, existing methods often rely on semantically salient objects rather than true gaze cues, leading to degraded performance when targets are not salient. To address this, we propose a novel training mechanism to enhance gaze reasoning in VFMs for gaze following. Our method includes: (1) a head-conditioned local LoRA, which enables localized adaptation to preserve scene token learning while improving head token learning for gaze reasoning; and (2) an out-of-cone penalty, which injects gaze cues into head tokens while aligning them with scene tokens. Experiments on the GazeFollow and VAT datasets demonstrate that our method achieves state-of-the-art performance, with particularly strong improvements when gaze targets are not semantically salient. Our findings offer valuable insights for advancing future gaze following research. We will release the code once the paper is accepted.


EntropyCache: Decoded Token Entropy Guided KV Caching for Diffusion Language Models

Minsoo Cheong ⋅ Donghyun Son ⋅ Woosang Lim ⋅ Sungjoo Yoo

Diffusion-based large language models (dLLMs) rely on bidirectional attention, which prevents lossless KV caching and requires a full forward pass at every denoising step. Existing approximate KV caching methods reduce this cost by selectively updating cached states, but their decision overhead scales with context length or model depth. We propose EntropyCache, a training-free KV caching method that uses the maximum entropy of newly decoded token distributions as a constant-cost signal for deciding *when* to recompute. Our design is grounded in two empirical observations: (1) decoded token entropy correlates with KV cache drift, providing a cheap proxy for cache staleness, and (2) feature volatility of decoded tokens persists for multiple steps after unmasking, motivating recomputation of the $k$ most recently decoded tokens. The skip-or-recompute decision requires only $O(V)$ computation per step, independent of context length and model scale. Experiments on LLaDA-8B-Instruct and Dream-7B-Instruct show that EntropyCache achieves 15.2×–26.4× speedup on standard benchmarks and 22.4×–24.1× on chain-of-thought benchmarks against vanilla baselines, with competitive accuracy and decision overhead accounting for only 0.5% of inference time.


Entropy Guided Dynamic Patch Segmentation for Time Series Transformers

Sachith Abeywickrama ⋅ Emadeldeen Eldele ⋅ Min Wu ⋅ Xiaoli Li ⋅ Chau Yuen

Patch-based transformers have emerged as efficient and improved long-horizon modeling architectures for time series modeling. Yet, existing approaches rely on temporally-agnostic patch construction, where arbitrary starting positions and fixed lengths fracture temporal coherence by splitting natural transitions across boundaries. This naive segmentation often disrupts short-term dependencies and weakens representation learning. We propose a novel Entropy-Guided Dynamic Patch Encoder (EntroPE), as a temporally informed framework that dynamically detects transition points via conditional entropy and dynamically places patch boundaries. This preserves temporal structure while retaining the computational benefits of patching. EntroPE consists of two key modules, namely an Entropy-based Dynamic Patcher (EDP) that applies information-theoretic criteria to locate natural temporal shifts and determine patch boundaries, and an Adaptive Patch Encoder (APE) that employs pooling and cross-attention to capture intra-patch dependencies and produce fixed-size latent representations. Extensive experiments on long-term forecasting, classification, and anomaly detection demonstrate that the proposed method improves both accuracy and efficiency, establishing entropy-guided dynamic patching as a promising new paradigm for time series modeling.


Environment-Robust Representation Learning with Empirical Bayes

Yuli Slavutsky ⋅ Matthew Shen ⋅ Bohan Wu ⋅ David Blei

We consider multi-environment prediction problems. We assume the environments change the distribution of a latent variable, while the mechanisms generating observed covariates and targets remain stable conditional on that variable. For example, hospitals or clinical cohorts may differ in the prevalence of latent patient states, even though the relationships between those states, physiological measurements, and outcomes remain unchanged. Given a dataset from multiple environments, we formulate a Bayesian model for such problems and derive the corresponding variational objective. We show that this objective decomposes into per-environment terms and an additional cross-environment balancing term induced by the model's structure. We use an empirical Bayes method to set the prior and incorporate it into the objective. Based on this objective, we develop an amortized variational algorithm for posterior approximation, and use the resulting learned latent variables to form predictions in new environments. We study our approach through simulations and real-world studies of astronomical source identification, microbiome-based disease detection, and ICU sepsis prediction. Across these settings, our method outperforms previous approaches for prediction in new environments.


EP-Flow: Disordered Crystal Structure Prediction without Site-level Annotation

liu qiuliang ⋅ Liming Wu ⋅ Qi Li ⋅ Zhonglong Peng ⋅ Chang Chen ⋅ Wenbing Huang ⋅ Xiaolong Chen ⋅ Shifeng Jin

Generative models have made rapid progress in ordered crystal structure prediction, yet many functional materials are intrinsically disordered, with substitutional mixing, vacancies, or interstitial species controlling their properties. Existing crystal generators either assume deterministic site occupations or require site-level disorder annotations, which are often unavailable when the chemical formula is the primary input. We formulate disordered crystal structure prediction through an Occupancy Distribution Matrix (ODM), a continuous site-by-species representation that unifies ordered crystals, solid solutions, vacancy disorder, and interstitial occupancy. A valid ODM must satisfy coupled site-wise occupancy, mass-conservation, and non-negativity constraints, placing each sample on a formula-dependent transportation polytope. We propose Entropic Polytope Flow (EP-Flow), a marginal-constrained flow matching framework that canonicalizes heterogeneous polytopes into a shared double-centered space, learns a marginal-preserving flow, and recovers feasible occupancies through a Sinkhorn inverse map. By jointly generating occupancies, fractional coordinates, and lattice parameters, EP-Flow achieves state-of-the-art performance on formula-conditioned disordered CSP benchmarks derived from COD and MPDS, substantially outperforming adapted ordered-crystal generators. Analyses further show that EP-Flow recovers sparse and chemically meaningful local disorder patterns rather than merely matching global composition statistics.


EpiStream: Utility-Aware Temporal Abstraction for Dense-Action Streams

Kaize Tan ⋅ Ran Zhang ⋅ Feng Yichao ⋅ Matthew H Jie ⋅ Dong Fang

Video understanding is commonly built on action-triggered responses or scene-level aggregation, implicitly assuming clear temporal boundaries. This assumption breaks in dense-action streams such as gameplay and sports, where overlapping actions induce continuous latent dynamics without explicit transitions. The challenge is further amplified in streaming settings, where models operate under strict causal constraints and must progressively construct episodic memory for downstream reasoning. Existing segmentation-based strategies therefore often produce fragmented or semantically mixed episodes, leading to degraded understanding. We present EpiStream, an online framework for utility-aware semantic episode formation in dense-action streams. Instead of detecting event boundaries, EpiStream learns when an evolving video prefix should be committed into an episode unit that maximizes downstream utility under limited memory budgets. The framework consists of three steps: (i) utility function design for semantic coherence, inter-episode distinctiveness, and downstream objectives; (ii) utility-optimal commit advantage construction from teaching signals; and (iii) peak-aware advantage learning for training a causal online commitment policy from prefix-only observations. Extensive experiments show that EpiStream substantially improves temporal semantic coherence and consistently benefits challenging downstream tasks, including real-time advice and player intention prediction.

We introduce Equilibrium Matching (EqM), a generative modeling framework built from an equilibrium dynamics perspective. EqM discards the non-equilibrium, time-conditional dynamics in traditional diffusion and flow-based generative models and instead learns the equilibrium gradient of an implicit energy landscape. Through this approach, we can adopt an optimization-based sampling process at inference time, where samples are obtained by gradient descent on the learned landscape with adjustable step sizes, adaptive optimizers, and adaptive compute. EqM surpasses diffusion/flow models empirically, achieving an FID of 1.90 on ImageNet 256$\times$256. EqM is also theoretically justified to learn and sample from the data manifold. Beyond generation, EqM is a flexible framework that naturally handles tasks including partially noised image denoising, OOD detection, and image composition. By replacing time-conditional velocities with a unified equilibrium landscape, EqM offers a tighter bridge between flow and energy-based models and a simple route to optimization-driven inference.


Equivariant Spherical Transformer for Efficient Molecular Modeling

Junyi An ⋅ Xinyu Lu ⋅ Yun-Fei Shi ⋅ Peijia Lin ⋅ Qianwei Tang ⋅ Li-Cheng Xu ⋅ Chao Qu ⋅ Fenglei Cao ⋅ Yuan Qi

Equivariant Graph Neural Networks (GNNs) have significantly advanced the modeling of 3D molecular structure by leveraging group representations. However, their message passing, heavily relying on strictly equivariant operations, suffers from restricted expressiveness due to the limited non-linearity and low degree of group representations. To overcome this, we introduce the Equivariant Spherical Transformer (EST), a novel plug-and-play module that applies a Transformer-based architecture to the Fourier spatial domain of group representations. The integration of EST enhances the model's expressiveness while preserving the crucial equivariant inductive bias through a uniform sampling strategy of spherical Fourier transforms. As demonstrated by our experiments on challenging benchmarks like OC20, MPtrj and QM9, EST-based models achieve state-of-the-art performance. For the complex molecular systems within OC20, small models empowered by EST can outperform some larger models and those using additional data. In addition, we provide both theoretical and experimental validation of EST's equivariance as well, paving the way for new research in this area.

Diffusion Transformers enable powerful in-context image editing from a reference image and a text prompt, but also make it easier to maliciously edit public images while preserving the private identity of the reference. Existing DiT-oriented protections, such as DeContext, mainly suppresses context-to-target attention to weaken reference utilization. However, this strategy can be insufficient for complex edits where prompt semantics already dominate generation. Instead, we study proactive protection for DiT-based editors from a conditional-flow perspective, and use a local geometric interpretation of in-context editing, where the denoising trajectory is empirically characterized as being steered by prompt-consistent and reference-consistent directions. Thus, effective protection should not simply corrupt the output, but should drive it away from the original context manifold. To this end, we propose Conditional Flow Hijacking, which adds imperceptible perturbations to the reference image. Instead of suppressing context attention, our method preserves the context pathway, but redirects the context-induced steering signals in the intermediate transformer layers, causing the final trajectory to drift away from the reference consistent solution region. Experiments on FLUX.1-Kontext with different datasets and prompts including attribute and complex scene editing show that our method significantly suppresses reference identity leakage while maintaining image quality and prompt alignment, outperforming prior protection baselines, especially under complex scene editing prompts. We further verify the robustness of the proposed protection method as well as its transferability across other models, highlighting the need for proactive defenses for in-context generative models.

Reinforcement learning (RL) has become a key training step for improving mathematical reasoning in large language models (LLMs), but it often has high GPU memory usage, which makes it hard to use in settings with limited resources. To reduce these issues, we propose Evolution Strategies with Sharpness-Aware Maximization (ESSAM), a full parameter fine-tuning framework that tightly combines the zeroth-order search in parameter space from Evolution Strategies (ES) with the Sharpness-Aware Maximization (SAM) to improve generalization. We conduct fine-tuning experiments on the mainstream mathematica reasoning task GSM8K. The results show that ESSAM achieves an average accuracy of 78.27\% across all models and its overall performance is comparable to RL methods. It surpasses classic RL algorithm PPO with an accuracy of 77.72\% and is comparable to GRPO with an accuracy of 78.34\%, and even surpassing them on some models. Further generalization experiments show that the models trained with ESSAM exhibit stronger generalization ability. Their average performance achieves the best results on 5 out of 6 datasets, indicating that ESSAM can effectively improve the generalization performance of fine-tuned models. In terms of GPU memory usage, ESSAM reduces the average GPU memory usage by $18\times$ compared to PPO and by $10\times$ compared to GRPO, achieving an extremely low GPU memory usage. In addition, we design an accelerated variant of ESSAM, which achieves nearly a twofold speedup while maintaining the same GPU memory usage as ESSAM, and attains an average accuracy of 78.02\% across all models, outperforming PPO. Code: https://anonymous.4open.science/r/ESSAM-3F4F/


Euler-Mamba: Learning Resolution-Invariant State Evolution on Polar Manifolds for Asymmetric Pansharpening

Mengting Ma ⋅ Anqi Zhu ⋅ Yizhen Jiang ⋅ Chengfu Ding ⋅ Wei Zhang

The goal of asymmetric pansharpening is to synthesize a high-resolution multispectral (HR-MS) image by integrating panchromatic (PAN) and low-resolution multispectral (LR-MS) images. While State Space Models (SSMs), particularly Mamba, offer a promising paradigm for global dependency modeling, their standard selective scan mechanisms operate in discrete Cartesian grids, which are ill-suited for distinguishing between structural and spectral features in complex-valued frequency representations. To resolve this, we propose Euler-Mamba, a physics-inspired framework that reformulates the selective scan as a continuous state evolution on polar manifolds. We develop an Eulerian Selective Scan (ESS) mechanism that treats the fusion task as a trajectory of frequency growth. By performing radial scans within the polar domain, ESS governs complex-valued state transitions through the decoupled architecture: the Explicit Phase Rotation (EPR) block facilitates explicit phase rotations for geometric alignment, while the Implicit Magnitude Mapping (IMM) block performs magnitude scaling for spectral calibration. This polar-decoupled evolution allows the model to capture the continuous transition from global spectral backgrounds to fine structural details with linear complexity. Consequently, by modeling the fusion as a continuous functional evolution flow, Euler-Mamba achieves resolution-invariant performance and zero-shot generalization across varying spatial scales, bypassing the computational complexity inherent in traditional discrete-integration frameworks. Extensive experiments demonstrate that Euler-Mamba establishes a new Pareto frontier for asymmetric pansharpening, achieving a superior balance between physical interpretability, computational efficiency, and reconstruction quality.


EvalAwareBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models

Xinning Li ⋅ Kemunto Ochwang'i ⋅ Ram Bharadwaj Aryasomayajula ⋅ Alexandra Souly ⋅ Robert Kirk

Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks. We introduce EvalAwareBench, an open pipeline and benchmark for measuring evaluation awareness that works with any Inspect-compatible evaluation, allowing practitioners to test against current and future benchmarks. EvalAwareBench ships with a newly curated transcript suite covering current frontier system-card evaluations and diverse deployment sources. The benchmark serves two purposes: measuring how reliably frontier LLMs recognize that they are being evaluated, and assessing how detectable individual benchmarks are as evaluations. We identify two methodological choices in the existing literature that introduce systematic bias: the identity of the model that generated the deployment transcripts accounts for 11.25% of measurement variance and can reorder model rankings; and elicitation prompts selected for high performance on one model can perform near chance on others. EvalAwareBench corrects for both via per-model probe calibration and a stratified generator-harmonisation procedure. We release EvalAwareBench as an open pipeline and dataset.

Finetuning Language Models often requires enforcing constraints on individual inputs without compromising performance. However, current alignment methods typically impose constraints only on average, which can induce undesirable disparities across inputs or users. We propose a finetuning framework that addresses this limitation by enforcing alignment requirements as per-sample constraints. To handle the optimization challenges inherent in this approach, we utilize an augmented Lagrangian formulation in the dual domain. Since pointwise constraints can be overly restrictive in low-probability regions or in the presence of outliers, we introduce a learned, sample-dependent relaxation that minimally relaxes constraints to optimize objective performance. We demonstrate the versatility of our framework across three small language model tasks: safety in instruction following, preference satisfaction in function calling, and length-aware re-ranking. Across these settings, our approach reduces tail constraint violations while largely preserving or improving downstream performance


Evolutionary foraging in grids: Intermittent search emerges as an optimal strategy

Shailendra Bhandari ⋅ Alex Szorkovszky ⋅ Anis Yazidi ⋅ Pedro G Lind

How search strategies evolve in sparse, depletable landscapes remains a central question in foraging theory. We study this problem with an evolutionary simulation in which agents forage on a two-dimensional toroidal lattice containing non-renewable resources distributed either uniformly or as Lévy dust. Each agent carries a heritable genome encoding step lengths, velocities, and turning angles, and selection acts on a fitness function, derived from first principles, combining energetic gain, movement cost, and coverage efficiency. By allowing movement traits to evolve without imposing a prescribed power-law step-length distribution, we test whether selection recovers a strict Lévy-like random walk, similar to the spatial distribution of resources, or instead favors an alternative search mechanism. Strikingly, our results indicate that, in the finite depletion-driven landscapes considered here, evolved search is more consistent with intermittent dynamics than with strict scale-free Lévy motion. To characterize the effective dynamics of the evolved trajectories, we take the average coefficient of determination when fitting second- and fourth-order displacement moments to intermittent search and Lévy-walk models. While a Lévy-like random walk fits very well with the numerical results from the evolutionary search ($R^2>0.9$), the intermittent search achieves a closer fit, with fitted coefficients of determination ($R^2>0.99$) for all resource distributions. Evolution rapidly reshapes the movement genome toward short displacements while retaining a sparse tail of longer relocations, consistent with local exploitation punctuated by occasional transfer. The framework provides a controlled setting for studying how search rules emerge under resource limitation and may inform resource-constrained exploration in autonomous systems.

In this paper, we introduce Exact Flow Linear Attention(EFLA), an exact-flow formulation of delta-rule linear attention. We show that the delta-rule update can be interpreted as an explicit Euler discretization of an underlying continuous-time system. EFLA replaces this first-order update with the exact closed-form flow. By exploiting the rank-1 structure of the dynamics matrix, both the matrix exponential and the input integral collapse to a simple update that preserves delta-rule linear attention's algebraic structure, parameter count, linear-time complexity, and chunkwise parallelism. This attention mechanism removes the Euler discretization error of the delta-rule dynamics without introducing additional parameters. Experiments on robustness tests, language modeling benchmarks, and the MAD synthetic benchmark show that EFLA improves stability under corrupted and high-energy inputs, reduces perplexity, and achieves stronger downstream performance compared to SSM and Euler-style baselines. These results establish exact-flow integration as a principled and scalable update mechanism for delta-rule linear attention.


Expanding Flow Maps

Sophia Tang ⋅ Pranam Chatterjee

Flow maps have enabled remarkable progress in few-step generative modeling across both continuous and discrete state spaces. Despite their promise, existing parameterizations are restricted to flows over fixed dimensions or fixed sequence lengths. Here, we introduce Expanding Generative Flows (EFlows), which define flows between distributions of increasing dimensionality through an expanding interpolant built from augmented distributions of conditional noise. Building on this framework, we propose Expanding Flow Maps (EFMs), a new class of flow maps that distill EFlows into efficient few-step generative models. Each EFM factors the map between any two timesteps into two learned components: an expand operator, which augments the state with new coordinates or tokens, and a transport operator, which pushes the expanded state forward along the interpolant. Composing these operators yields a single map that jointly expands and denoises the state, recovering existing fixed-canvas flow maps as the special case in which the expand operator is the identity. We further extend the framework to the discrete simplex, enabling variable-length sequence generation via token insertions. Across domains, EFlows and EFMs provide a principled approach to generative problems in which output size is itself learned, controllable degree of freedom.


Expanding LLM Agent Boundaries with Strategy-Guided Exploration

Andrew Szot ⋅ Michael Kirchhof ⋅ Omar Attia ⋅ Alexander Toshev

Reinforcement learning (RL) has demonstrated success in post-training large language models (LLMs) as agents for tasks such as computer use, tool calling, and coding. However, exploration remains a central challenge in RL for LLM agents, especially as they operate in language-action spaces with complex observations and sparse rewards. In this work, we address exploration for LLM agents by leveraging the ability of LLMs to plan and reason in language about the environment to shift exploration from low-level actions to higher-level language strategies. We thus propose Strategy-Guided Exploration (SGE), which first generates a concise natural-language strategy that describes what to do to make progress toward the goal, and then generates environment actions conditioned on that strategy. By exploring in the space of strategies rather than the space of actions, SGE induces structured and diverse exploration that targets different environment outcomes. To increase strategy diversity during RL, SGE introduces mixed-temperature sampling, which explores diverse strategies in parallel, along with a strategy reflection process that grounds strategy generation on the outcomes of previous strategies in the environment. Across UI interaction, tool-calling, coding, and embodied agent environments, SGE consistently outperforms exploration-focused RL baselines, improving both learning efficiency and final performance. We show that SGE enables the agent to learn to solve tasks too difficult for the base model.


Expanding the Role of Diffusion Models for Robust Classifier Training

Pin-Han Huang ⋅ Shang-Tse Chen ⋅ Hsuan-Tien (Tien) Lin

Incorporating diffusion-generated synthetic data into adversarial training (AT) has been shown to substantially improve the training of robust image classifiers. In this work, we extend the role of diffusion models beyond merely generating synthetic data, examining whether their internal representations, which encode meaningful features of the data, can provide additional benefits for robust classifier training. Through systematic experiments, we show that diffusion models offer representations that are both diverse and partially robust, and that explicitly incorporating diffusion representations as an auxiliary learning signal during AT consistently improves robustness across settings. Furthermore, our representation analysis indicates that incorporating diffusion models into AT encourages more disentangled features, while diffusion representations and diffusion-generated synthetic data play complementary roles in shaping representations. Experiments on CIFAR-10, CIFAR-100, and ImageNet validate these findings, demonstrating the effectiveness of jointly leveraging diffusion representations and synthetic data within AT.


Explanation Multiplicity in SHAP: Characterization and Assessment

Hyunseung Hwang ⋅ Seungeun Lee ⋅ Lucas Rosenblatt ⋅ Steven Whang ⋅ Julia Stoyanovich

SHAP explanations are widely used in high-stakes settings to justify decisions, yet they can differ substantially across repeated runs, even when the model, the input instance, and the prediction are held fixed. Prior work has documented disagreement *between* explanation methods; we show that substantial disagreement arises even *within* SHAP across reruns of the same estimator on the same trained model and instance. We call this phenomenon *explanation multiplicity* and develop an evaluation methodology for characterizing it under deployment-realistic computational budgets, combining a dual-seed protocol that disentangles model-induced from explainer-induced variability, a hierarchy of magnitude-, rank-, and set-based metrics, and randomized Dirichlet and Mallows null models that calibrate when observed disagreement exceeds chance. Across multiple datasets, model classes, and sampling strategies, we find that explanation multiplicity is pervasive and persists even for high-confidence predictions. The dominant source depends on the data regime: model-induced variability dominates in small-data settings, while explainer-induced variability dominates at scale. Commonly used $\ell_2$ distance is largely insensitive to this instability, while rank-based metrics reveal severe reordering of top-ranked features. Improved sampling methods such as CTE do not eliminate rank-level multiplicity, and deterministic alternatives such as K-Means achieve stability only by targeting a different estimand that diverges sharply from Shapley values defined against the empirical training distribution. Practitioners should treat single-run SHAP outputs as realizations of a distribution rather than as authoritative artifacts.


Explicit Geometric Chain-of-Thought for Vision-Language-Action in Autonomous Driving

Xingtai Gui ⋅ Yucheng Zhou ⋅ Dongqian Guo ⋅ jiahao gong ⋅ Feiyang Tan ⋅ Jianbing Shen

Vision-language-action(VLA) models have emerged as a promising interface for autonomous driving. However, existing VLA models still suffer from a fundamental mismatch: driving actions require precise 3D geometric cues, while visual-language understanding and reasoning are largely conducted in a 2D semantic space. In this paper, we propose GeoCoTDrive, an explicit geometric chain-of-thought framework that grounds geometry in a planning-oriented manner. GeoCoTDrive follows a ``think with 2D first, drive with dedicated 3D priors'' paradigm: it first identifies sparse 2D regions corresponding to decision-critical cues, and then retrieves localized 3D priors by sampling features from a geometric foundation model within the grounded regions. These localized geometric features are interleaved into the autoregressive context to support trajectory generation. To supervise this process, we introduce planning-relevant grounding, a new region-level grounding task that focuses on local spatial cues directly affecting ego planning, and construct the PlanningGrounding dataset to endow VLM models with planning-oriented grounding ability. Experiments across multiple end-to-end autonomous driving benchmarks show that GeoCoTDrive consistently improves safety-critical planning performance, demonstrating the effectiveness of explicit geometric chain-of-thought reasoning for VLA-based planning.


Exploiting Negative Multi-Cluster Structure in Class-Wise Embeddings for Weakly Supervised Multi-Label Learning

Bo Han ⋅ Zhuoming Li ⋅ Yaxin Hou ⋅ Xiaoyu Wang ⋅ Hui LIU ⋅ Junhui Hou ⋅ Yuheng Jia

Class-wise embeddings provide an effective strategy for multi-label learning (MLL) by capturing the distinct discriminative properties of each class. However, existing methods based on class-wise embeddings typically overlook the internal distribution of negative samples. We uncover a pivotal phenomenon: negative instances in class-wise spaces naturally organize into multiple semantically coherent sub-clusters, a structure that persists across diverse modalities even under extreme label sparsity. Leveraging this insight, we propose a Structure-Aware Pseudo-label mining framework for weakly supervised MLL (SaPu). The proposed framework mines positive pseudo-labels by aggregating cluster-level consistency across different class-wise embeddings and identifies negative pseudo-labels via neighborhood exclusion. These high-quality signals further guide class-wise contrastive learning, establishing a self-reinforcing loop that iteratively refines embeddings. Extensive experiments on image, text and audio benchmarks validate that SaPu outperforms state-of-the-art methods by an average of 4.46% in challenging single-label supervision scenarios. Code is available in the supplementary material.


Exploring the Limits of Compositional Generalization in Vision-Language-Action Manipulation

Xupeng Zhang ⋅ Yantai Yang ⋅ Chang Guo ⋅ Zhaokai Yin ⋅ Zhipeng Zhang

Vision Language Action (VLA) models can execute isolated atomic skills, but practical long horizon manipulation requires recombining these skills into new compositions without demonstrations for every skill combination. The direct full instruction route treats skill chaining as long horizon imitation. In our controlled study, full instruction training on limited skill combination sequences reaches 40.0% ordered success on covered combinations but only 7.7% on recombination, suggesting memorization of sequence templates rather than robust skill recombination. We therefore introduce Hierarchical Subtask Execution (HSE), a diagnostic execution framework that exposes skill recombination through a shared modular interface over the planner, trigger, and policy, enabling long horizon failures to be decomposed into separable modes. To make this diagnosis measurable, we introduce RoboCombine, a benchmark for skill recombination without combined skill training trajectories, with metrics for ordered success and transition conversion across subtask boundaries. RoboCombine reveals that atomic competence alone does not ensure ordered execution. Failures concentrate at transitions where downstream subtasks must launch from handoff states produced by preceding subtasks rather than canonical atomic starts. We identify this mismatch as the start state gap. To address it, we introduce Pose Robust Task Initialization (PRTI), a lightweight atomic task post training strategy that broadens start state coverage beyond canonical starts. Under oracle timed HSE on RoboCombine, adding PRTI raises ordered success from 2.11% to 26.32% and transition conversion from 2.63% to 60.53%, without combined skill training trajectories. Together, these results show that combinatorial generalization requires handoff robust atomic skills that convert first stage progress into reliable downstream execution, not merely decomposition into known subtasks. We will release RoboCombine and HSE code to enable reproducible diagnosis of atomic skill recombination failures.

Identifying latent dynamical systems from noisy, high-dimensional measurements is a central problem at the intersection of representation learning, system identification, and scientific discovery. We present DYSCO, a multi-view temporal contrastive learning algorithm that jointly recovers latent trajectories and the governing dynamics from such observations, by leveraging multiple independent noisy views of the same underlying process to disentangle signal from noise. By parameterizing the dynamics in a structured functional basis, our framework further enables symbolic recovery of the governing equations within an affine gauge. We offer theoretical guarantees for strong identification up to an affine indeterminacy, extending prior identifiability results to the realistic setting of noisy nonlinear observations. Empirically, we demonstrate accurate recovery of both latent trajectories and flow fields across a diverse set of dynamical regimes (e.g, chaotic, oscillatory, and metastable) under both Gaussian and Poisson observation noise, the latter being particularly relevant for neural recordings.


Extraneous Cognitive Load in Large Language Models

Arushi Gupta ⋅ Sheng Liu ⋅ James Zou

Cognitive Load Theory (CLT) shows that human problem-solving is affected by both intrinsic load, arising from task difficulty, and extraneous load, arising from how task information is presented. LLMs show analogous effects, where presentation alone can affect model performance, but these effects have not yet been systematically isolated from task difficulty. To address this gap, we introduce CogLoadBench, a controlled benchmark that varies five CLT-inspired presentation factors across three reasoning task families. Across models, we find that higher extraneous load conditions often reduce accuracy by either adding competition or complicating the reasoning path. To support diagnosis and mitigation of load-related effects, we propose the Extraneous Load Score (ELS), a prompt-time metric computed from activations that estimates a model's extraneous load, and show that it generalizes across reasoning task types and load levels. We further show that shifting model representations toward lower-ELS directions at inference time can significantly improve model performance without changing prompt content.


F2G-Pose: Geometry-Aware Foundation Feature Lifting for Direct RGB-D Category-Level Object Pose Estimation

Jinyu Zhang ⋅ Haitao Lin ⋅ Jiashu Hou ⋅ Xiangyang Xue ⋅ Yanwei Fu

This paper studies category-level RGB-D object pose estimation, which recovers an object's 3D rotation and translation without instance-specific CAD models or reference views. This task is highly ambiguous due to intra-class variations, symmetries, and partial observations. Current solutions struggle to balance reliability and efficiency: deterministic methods are prone to error propagation from intermediate structures, while diffusion models suffer from high inference latency and complex post-hoc candidate extraction. To address this, we present F2G-Pose, a real-time direct, single-pass framework that predicts category-level pose, metric size, and camera-frame completed shape from a single RGB-D observation in one feed-forward pass. F2G-Pose lifts dense visual foundation features to partial point clouds, encodes local geometry with DGCNN tokenization, and integrates object-level structure with a Mixture-of-Experts Transformer. Joint size prediction and camera-frame shape completion encourage metric consistency through dense geometric supervision. Trained only on synthetic SOPE data, F2G-Pose achieves state-of-the-art accuracy on SOPE and ROPE while running at 28 FPS, and further demonstrates strong camera-frame geometry on HouseCat6D and real-object transfer on HANDAL. Analyses of latency, robustness, stability, shape completion, and GenPose++ candidate behavior further demonstrate the efficiency and reliability of F2G-Pose, with real-world manipulation experiments demonstrating deployment potential under segmented RGB-D inputs.


Failing Forward: Adaptive Failure-Informed Learning for Vision-Language-Action Models

Meng Zheng ⋅ Samhita Marri ⋅ Anwesa Choudhuri ⋅ Benjamin Planche ⋅ Zhongpai Gao ⋅ Van N Nguyen ⋅ Terrence Chen ⋅ Girish Chowdhary ⋅ Ziyan Wu

Vision-language-action (VLA) models provide a promising paradigm for scalable robotic manipulation, yet their reliance on success-only behavioral cloning leaves them brittle; lacking corrective training signals, minor execution errors rapidly compound into unrecoverable, out-of-distribution failures. To address this limitation, we propose Adaptive Failure-Informed Learning (AFIL), an end-to-end framework that leverages failure trajectories as adaptive negative guidance for diffusion- and flow-based VLA policies. AFIL uses a pretrained VLA to generate failure rollouts online, avoiding the need for handcrafted failure-mode design or human-in-the-loop recovery. It then jointly trains Dual Action Generators (DAGs) for successful and failed behaviors while sharing a common vision-language backbone, enabling efficient failure-aware policy learning with limited parameter overhead. During sampling, the failure generator adaptively steers action generation away from failure-prone regions and toward more reliable success modes, with guidance strength determined by the per-diffusion-step distance between success and failure distributions. Experiments across in-domain and out-of-domain robotic manipulation tasks, covering both short- and long-horizon settings, show that AFIL consistently improves task success rates and robustness over existing VLA baselines, demonstrating its effectiveness, efficiency, and generality.

Filtered vector search in private and on-prem retrieval-augmented generation (RAG) systems must control tail risk under fixed latency, RAM, and SSD budgets, not merely pick a fast backend. Our measurements isolate a recurring ambiguity regime concentrated in the 5-15% failure band, where small shifts in vector-metadata alignment flip which familiar P1-P4 action is safe and create 3-6 P95 spread. We study that regime as authorization over fixed actions rather than as generic planning. The paper contributes a measured failure-band stress test, an auditable authorize/veto rule over fixed P1-P4 actions, and PACER as one concrete controller that combines recall-floor adjustment, calibrated tail intervals, and regime-local fallback. The measured failure band exposes a different runtime decision boundary: when the provisional winner is separated enough to trust and when uncertainty should refuse it. On Medium+EntDoc, evaluated on a dual-AMD EPYC 7543 server with 128GB RAM at matched recall in the 5-15% failure band, PACER reduces P95 by 46%\ and P99 by 45%\ relative to the strongest baseline while maintaining 0.94-0.98\ Recall@20. The gain concentrates where plan rankings are unstable and weakens when the authorization logic is stripped back, while feasibility, transfer, drift, and hierarchy-aligned slices remain supporting diagnostics. For recurring fixed-budget filtered retrieval, these results show that an auditable authorize/veto layer over fixed actions can materially stabilize tail latency in the measured failure band.

Training smaller language models on reasoning solutions generated by a larger teacher model is the prevailing approach for transferring mathematical reasoning capabilities. However, the student tends to overfit the output distribution of the teacher and generalizes poorly when its own behavior at inference time deviates from the training data. Recent methods attempt to bridge this gap by incorporating error data, yet the corrective trajectories still reflect the reasoning patterns of the teacher, which the student cannot reliably reproduce, leaving the distribution mismatch largely unresolved. We propose a method built on the insight that effective training data should be jointly determined by what the student fails on and what it can realistically learn. Our method first profiles the failure behavior of the student on each problem through repeated sampling, characterizing it in terms of error rate and error consistency to reveal qualitatively distinct failure modes. Each mode demands a fundamentally different training signal. The diagnosed mode then guides candidate generation from the teacher via distinct prompting strategies tailored to each failure type. Among the correct candidates, our method selects the one whose initial reasoning steps after the point of divergence are most reproducible by the student. These first corrective steps constitute the primary bottleneck, while the remaining steps follow an already corrected context and can therefore be continued more readily. This contrasts with standard reranking approaches, which score the trajectory as a whole and thereby dilute the signal at the critical reasoning transition. Experiments on GSM8K, MATH, and three out-of-distribution benchmarks for Llama-3.1-8B and DeepSeek-Math-7B demonstrate that our method achieves superior performance compared with other strong methods. The code and trained weights will be publicly available.


FalconPerception-HD: High Density Perception via Reinforcement Learning

Sofian Chaybouti ⋅ YASSER ABDELAZIZ DAHOU DJILALI ⋅ Ngoc D Huynh ⋅ REDA ALAMI ⋅ Hilde Kuehne

Autoregressive perception models trained to localize visual entities under the open-vocabulary setting are mostly trained using Supervised fine-tuning (SFT) with maximum likelihood, yet it optimizes a proxy objective (per-token cross-entropy) that is fundamentally misaligned with perception metrics such as precision and recall. In this paper, we explore post-training reinforcement learning (RL), specifically GRPO, to directly align these models with their evaluation metrics. Building up on the recently introduced Falcon Perception, we design an RL framework that addresses perception-specific challenges: reward design for set-structured outputs and multi-head sampling control. We discover multiple benefits from RL for perception: first, RL unlocks unprecedented performance in very dense scenes (up to 500 objects per scene) and to the best of our knowledge, no other system has such capabilities at this scale; furthermore it fixes common issues in autoregressive perception models like mask repetitions and removes almost entirely the need for NMS and coordinate deduplication, which improve both performance and efficiency and remove the need for hyperparameters tuning; overall, we notice improvements on all levels of difficulties in referring expression segmentation (on PBench and SACO-Gold), and we find an elegant way to preserve the knowledge of whether an object exists or not (as evaluated by MCC) without training on negative samples. We show that a simple reward penalizing false negatives and positives is enough. We develop two hybrid self-annotation pipelines, respectively tailored for difficult referring expressions and very dense scenes, and show their benefits on RL-training. Model weights and datasets will be made public.


Fast 4D Mesh Generation by Spatio-Temporal Attention Chains

Dvir Samuel ⋅ Yuval Atzmon ⋅ Gal Chechik ⋅ Yoni Kasten

4D mesh generation has recently emerged as a powerful paradigm for recovering dynamic 3D structure from videos, but existing methods remain slow, computationally expensive, and difficult to scale to longer sequences. We introduce a training-free approach that accelerates 4D mesh generation while improving temporal correspondence quality. Our key observation is that temporal correspondences emerge inside a 4D backbone long before its generated meshes become visually accurate. We exploit this with a general framework we call Spatio-Temporal Attention Chain which propagates information across space and time. Starting from vertices on an anchor mesh, the chain maps vertices to latent tokens. It then follows temporal correspondences in latent space, and recovers frame-specific vertices through latent-to-vertex attention. This design avoids expensive explicit matching while preserving anchor mesh details and thereby improving dynamic mesh geometry and temporal consistency. Compared to state-of-the-art, our method generates a 4D mesh in 9 seconds, achieving a $13\times$ speedup while producing higher-quality results. Moreover, our approach scales to videos up to $16\times$ longer without degrading mesh quality. Beyond generation, the improved correspondences enable competitive zero-shot performance on two downstream tasks: 2D object tracking and 4D tracking. We further show that our framework enables reliable camera estimation, a capability not supported by prior 4D mesh generation methods.

Diverse applications on signed networks, ranging from leader-follower opinion dynamics to graph semi-supervised learning, fundamentally reduce to solving the discrete Dirichlet boundary value problem. The computational bottleneck of this problem lies in evaluating the entries of the label propagation operator $Q = -L_{U,U}^{-1}L_{U,L}$. While existing approximation methods based on absorbing random walks attempt to address this complexity, they often suffer from high variance and limited precision. To overcome these limitations, we propose MultiTreeQ, a novel and efficient framework designed to approximate the operator's entries with high accuracy. Specifically, MultiTreeQ is constructed based on the Signed Multi-Rooted Spanning Tree and leverages Rao-Blackwellized variance reduction to significantly improve estimation precision. Extensive experiments on real-world networks confirm that MultiTreeQ substantially outperforms existing baselines, scaling to massive graphs with up to 10 million nodes while delivering remarkable improvements in computational efficiency and precision. Our code is available at https://anonymous.4open.science/r/label-propagation-operator-2059/.

Structure learning for sparse graphical models in high dimensions requires both computational scalability and selection consistency. Classical criteria such as (Extended) Bayesian Information Criterion can provide consistency, but rely on model-specific likelihood and complexity calibration. In comparison, cross-validation avoids such analytic calibration and is therefore attractive for graph learning. However, cross-validation is computationally infeasible and theoretically inconsistent in structure learning. In this work, we resolve the tension by developing a scalable graph-recovery procedure built on approximate cross-validation. We demonstrate its effectiveness in learning large graphical models both theoretically and empirically, and establish selection consistency under mild signal strength assumptions.


FAST-Brain: A Flow-Aligned Spatio-Temporal Surrogate Brain Model

Shucheng Liu ⋅ Chengchun Shi ⋅ Kai Zhang ⋅ Hongtu Zhu

Modeling resting-state functional magnetic resonance imaging (rs-fMRI) data is crucial for understanding brain-wide neural activity. However, traditional methods struggle to capture complex temporal dynamics over long horizons, to account for the brain's anatomical spatial structure, and to model high-dimensional ambient signals that lie on a low-dimensional intrinsic subspace. We propose FAST-Brain, a unified flow-aligned spatio-temporal surrogate brain model that addresses all three challenges. At its core is a flow-aligned generative framework that directly predicts the clean blood-oxygen-level-dependent (BOLD) signal, paired with a graph convolutional network that captures spatial structural constraints and a Transformer that models long-range temporal dependencies. Theoretically, we show that under a low-dimensional subspace assumption, the approximation error of our model scales with the intrinsic dimension rather than the ambient dimension, which justifies our direct modeling of the BOLD signal. Extensive experiments on synthetic and Human Connectome Project datasets demonstrate that FAST-Brain achieves state-of-the-art performance in recovering functional connectivity, effective connectivity, and the implicit low-dimensional signal subspace.

We study clustering algorithm for dynamic evolving graphs $\{ G_t\}$, in which new edges (and potentially new vertices) are added into a graph, and the underlying cluster structure of the graph gradually changes over time. By designing a hierarchical data structure for dynamically maintaining the cluster structure of $\{ G_t\}$, we prove that the cluster structure of $G_T$ can be maintained with $O(d)$ amortised update time and $\widetilde{O}\left(n_T^{1/d}\right)$ amortised query time, for an arbitrary large constant $d$. This significantly improves the previous state-of-the-art (Laenen and Sun, ICML~'24), which requires $O(1)$ update time and $O(n_T/\log n_T)$ amortised query time. We further demonstrate this improvement with experiments on both synthetic and real-world datasets.


Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction

Jang-Hyun Kim ⋅ Dongyoon Han ⋅ Sangdoo Yun

Efficient key-value (KV) cache management is crucial for the practical deployment of large language models (LLMs), yet existing compression techniques often incur a trade-off between performance degradation and computational overhead. We propose a novel gating-based KV cache eviction method for frozen-weight LLMs that achieves high compression ratios with negligible computational cost. Our approach introduces lightweight sink-attention gating modules to identify and retain critical KV pairs, and integrates seamlessly into both the prefill and decoding stages. The proposed gate training algorithm relies on forward passes of an LLM, avoiding expensive backpropagation, while achieving strong task generalization through a task-agnostic reconstruction objective. Extensive experiments across the Qwen2.5-1M, Qwen3, and Gemma3 families show that our method maintains near-lossless performance while evicting up to 70\% of the KV cache. The results are consistent across a wide range of tasks, including long-context understanding, code comprehension, and mathematical reasoning, demonstrating the generality of our approach.


Fast-WAM: Do World Action Models Need Test-time Future Imagination?

Tianyuan Yuan ⋅ Zibin Dong ⋅ Yicheng Liu ⋅ Hang Zhao

World Action Models (WAMs) have emerged as a promising alternative to Vision-Language-Action (VLA) models for embodied control because they explicitly model how visual observations may evolve under action. Most existing WAMs follow an imagine-then-execute paradigm, incurring substantial test-time latency from iterative video denoising, yet it remains unclear whether explicit future imagination is actually necessary for strong action performance. In this paper, we ask whether WAMs need explicit future imagination at test time, or whether their benefit comes primarily from video modeling during training. We disentangle the role of video modeling during training from explicit future generation during inference by proposing **Fast-WAM**, a WAM architecture that retains video co-training during training but skips future prediction at test time. We further instantiate several Fast-WAM variants to enable a controlled comparison of these two factors. Across these variants, we find that Fast-WAM remains competitive with imagine-then-execute variants, while removing video co-training causes a much larger performance drop. Empirically, Fast-WAM achieves competitive results with state-of-the-art methods both on simulation benchmarks (LIBERO and RoboTwin) and real-world tasks, without embodied pretraining. It runs in real time with 110 ms latency, over 6$\times$ faster than existing imagine-then-execute WAMs. These results suggest that the main value of video prediction in WAMs may lie in improving world representations during training rather than generating future observations at test time.


Federated Unlearning with Gradient Adaptive Shaping

Siyin Huang ⋅ Yufan Liao ⋅ Yuelin Du ⋅ Yu Zhang

Federated unlearning aims to efficiently remove the influence of a specific client from a trained global model in federated learning without full retraining. However, existing approaches struggle to preserve the utility of the remaining clients because unlearning can be unstable at both the client and server stages. Local forgetting updates may become overly aggressive and destabilize client side optimization, while server aggregation can amplify conflicts between forgetting updates and retained client updates. To address both failure modes within a unified framework, we propose Federated Unlearning with Gradient Adaptive Shaping (FUGAS), a history-free gradient shaping framework that stabilizes the entire unlearning pipeline. On the unlearning client side, we employ a bounded preference objective that utilizes the pre-unlearning model predictions on the forgetting data as negative references to controllably steer the model away from unlearning client knowledge while requiring only minimal storage for reference caches rather than full historical updates. On the server side, we introduce a compatibility projection mechanism that reshapes the aggregated unlearning update to remain compatible with directions estimated from retained clients. We provide a theoretical analysis indicating that FUGAS promotes stability and ensures non-increasing empirical risk on retained distributions while establishing an excess risk bound relative to retraining. Extensive experiments demonstrate that FUGAS achieves effective unlearning while consistently maintaining high accuracy on retained data.


FedIBS: Federated Vision-Language Adaptation via Intrinsic Bias Selection

Xiaoming Wu ⋅ Wei Wenyu ⋅ Xin Wang ⋅ Ming Yang

Federated adaptation of large vision-language models (VLMs) is appealing for privacy-sensitive applications but must contend with client data heterogeneity, strict communication budgets, and on-device inference constraints. Recent methods predominantly introduce additive adaptation interfaces, such as learnable prompts or adapter modules, which attach extra trainable components to the frozen backbone, inflating both parameter counts and inference latency. We challenge this additive paradigm and ask: can robust federated VLM adaptation be achieved through a compact intrinsic interface alone? We answer this question with \textbf{FedIBS} (Federated Intrinsic Bias Selection), a module-free framework that fine-tunes only the biases in feed-forward network projections. FedIBS operates on existing backbone parameters and introduces zero additional inference cost. To make such compact intrinsic adaptation viable under heterogeneous clients, we further propose a Fisher-guided progressive masking mechanism that disentangles shareable and personalized bias coordinates. Each client retains personalized coordinates locally and communicates only the shareable subset for aggregation, enabling sparse communication while preserving client-specific adaptation. Extensive experiments on diverse datasets and heterogeneity settings demonstrate that FedIBS consistently outperforms recent federated prompt- and adapter-based baselines in both personalization and generalization, all while achieving substantially lower communication cost.


FedSOUL: Federated Continual Unlearning via Spectral Orthogonality

Haoran Shi ⋅ Hao Zhu ⋅ Yifei Zhang ⋅ ZHANG JUNRU ⋅ Lizhen Cui ⋅ Yali Jiang ⋅ Xiaoli Tang ⋅ Han Yu

As privacy, copyright, and safety requirements evolve, federated learning systems face a growing imperative to accommodate sequential unlearning requests. However, existing federated unlearning methods are designed for single-shot settings and their direct sequential application may lead to uncontrolled parameter-space overlap, causing cross-request interference that undermines prior forgetting. To address this, we propose the Spectral Orthogonality for Federated Continual unlearning (FedSOUL), a framework that reduces direct overlap among sequential unlearning updates. The core idea is to structure model updates in a shared orthonormal spectral basis, where each client is assigned a dedicated subspace in the frequency domain. Under this design, each unlearning request updates only its associated subspace, while all others remain unchanged, ensuring that sequential updates do not overlap in the allocated spectral coefficient space and thereby reducing cross-request interference. Beyond interference control, FedSOUL also improves efficiency by operating on sparse spectral coefficients, using FFT-based transforms for efficient computation and communicating only sparse coefficient updates. From a theoretical perspective, we relate continual unlearning to geometric conditions on update directions and show that our spectral construction satisfies update orthogonality by design, while functional preservation depends on bounded projected gradient leakage. Extensive experiments demonstrate that FedSOUL maintains stable unlearning–utility trade-offs across up to ten sequential requests and consistently outperforms competitive federated baselines. Our code is available at https://anonymous.4open.science/r/fedunlearning-1B92/.


Feedforward Novel View Synthesis for Heterogeneous Cameras

Meng Wei ⋅ Cheng Zhang ⋅ Boying Li ⋅ Yihang Chen ⋅ Jianmin Zheng ⋅ Hamid Rezatofighi ⋅ Jianfei Cai

Feed-forward novel view synthesis has recently shown promising results from sparse posed images, but most existing methods assume that context and target views share a fixed camera family. This homogeneous-camera assumption breaks in practical multi-sensor systems, where perspective, fisheye, and panoramic cameras may coexist and where the target projection may be unseen during training. We study feed-forward NVS across heterogeneous central cameras and identify a key ambiguity introduced by tokenization: a visual token aggregates a projection-dependent bundle of pixel rays, while existing camera encodings mainly expose absolute rays or token-center relations. To address this, we combine token-center relative Camera Positional Encodings and proposed local raymaps, a token-level representation that explicitly describes the intra-patch ray distribution summarized by each token. We further propose projection-aware 2D RoPE, which replaces raw image-grid coordinates with ray-induced angular coordinates so that relative positional reasoning is aligned across camera projections. Together, these components treat diverse cameras as calibrated samplings of a shared ray space rather than separate visual domains. On ScanNet++ with heterogeneous-camera system, our method improves over camera-conditioned baselines under mixed-camera evaluation and demonstrates zero-shot generalization to panoramic views.


FENet: Functional Embedding Neural Network for Change-Point Detection in Functional Time Series

Caixia Xu ⋅ Jinhong You ⋅ Wen Li ⋅ shouguo du ⋅ Yiming Tang ⋅ Jiguo Cao

Detecting abnormal changes in server operational data is an important task in modern computing systems, as unexpected shifts in traffic, workload, or performance metrics may indicate service degradation, abnormal access patterns, system failures, or potential security risks. However, server monitoring data are often naturally observed as functional time series, with complex temporal patterns, high dimensionality, and heterogeneous structures, making traditional change-point detection methods difficult to apply reliably in practice. To address this problem, we propose a general change-point detection framework based on a Functional Embedding Neural Network (FENet), designed to provide robust detection across different data types and structural conditions without relying on strong prior assumptions. The proposed method is motivated by the connection between the classical functional CUSUM (FCUSUM) statistic and neural network representations, which allows FENet to learn effective detection rules directly from data while retaining useful statistical intuition. We provide theoretical analysis to formalize this connection and study key properties of the proposed approach. Extensive simulation studies show that FENet achieves competitive performance compared with existing methods and remains robust under limited training data and distributional mismatch between training and test settings. We further apply FENet to real-world server operational metrics, demonstrating its practical value in detecting and localizing abnormal changes in complex server monitoring systems.


Few-shot Task Learning via Compositional Concept Inference

Hanming Ye ⋅ Yiding Song ⋅ Yilun Du

Humans can learn new tasks from just a few demonstrations. Prior methods like behavior cloning struggle to replicate this ability because they imitate demonstrations without learning the underlying concepts, such as goals, positional relations, and constraints. To address this challenge, we propose COIN, an approach that represents behavior as a composition of reusable concepts. A new task is learned by inferring its concepts from demonstrations then using a policy to implement them. We show that this approach is effective at few-shot learning across diverse domains. On object rearrangement, goal-oriented navigation, and human motion generation, COIN recombines previously seen concepts to solve novel tasks at inference time. On the LIBERO robotics benchmark, COIN achieves a $96.9$% average success rate on LIBERO-90, surpassing pretrained policies such as $\pi_0$ with roughly $10\times$ fewer parameters. This performance gap widens further when fine-tuning on unseen tasks with limited demonstrations, highlighting the advantage of concept inference over end-to-end policy adaptation in the few-shot regime.

Image steering aims to provide fine-grained control over concepts in generated images, enabling targeted and continuous manipulation without broadly altering the generation process. Recent steering methods for diffusion models primarily focus on the text-conditioning space or rely on costly auxiliary training. Deviating from that, we here propose a novel steering approach that extracts visual concept directions directly from the image-side activations of Diffusion Transformers. Specifically, we showcase few-shot concept extraction in the activation space using a simple difference-of-means estimator over intermediate residuals. We observe (a) scene-level concepts such as lighting, weather, and style are distributed across image tokens, as well as (b) local concepts such as object colors, textures, materials, and facial attributes are concentrated in specific regions. Towards controlling both concepts, we introduce a new unified training-free framework that uses global difference-of-means for (a) distributed concepts and masked difference-of-means for (b) local concepts, preventing local signals from being diluted by global averaging. Notably, using only a handful of positive and negative prompt pairs, our method identifies and extracts reusable concept directions that can be injected at inference for continuous steering. Quantitatively, our method outperforms text-encoder steering and training-free activation-editing baselines, as well as is competitive with specialized editing models. Our results suggest that as concept directions can be applied at patch level, our method enables precise targeted and compositional steering, manipulating multiple concepts in the same image.


Few-Step Cofolding with All-Atom Flow Maps

Gianluca Scarpellini ⋅ Ron Shprints ⋅ Peter Holderrieth ⋅ Juno Nam ⋅ Pranav Murugan ⋅ Rafael Gomez-Bombarelli ⋅ Tommi Jaakkola ⋅ Maruan Al-Shedivat ⋅ Nicholas Boffi ⋅ Joey Bose

All-atom generative modeling of $3\mathrm{D}$ biomolecular complexes has emerged as the dominant paradigm for predicting the structure of proteins and protein-ligand systems. Generating structures at the atomic level of fidelity, however, typically requires expensive iterative diffusion rollouts, making both conventional deployment and inference-time search techniques computationally costly. In this paper, we introduce the $\textbf{De}$noiser $\textbf{C}$ofolding $\textbf{A}$ll-atom $\textbf{F}$lowmap (DeCAF) framework for distilling state-of-the-art all-atom cofolding models into all-atom flow maps that produce high-quality samples in only a few inference steps. We build DeCAF on a denoiser-based formulation of flow maps with endpoint losses that naturally support $\mathrm{SE(3)}$ rigid alignment, which we show is critical for training accurate models. We further derive a simple change of variables that lets DeCAF operate in the $\sigma$-space noise schedule of EDM-style architectures, enabling direct distillation from pretrained cofolding diffusion models. Equipped with DeCAF's flowmap lookahead, we introduce a purpose-built inference-time framework that improves sampling through reward-guided search. Empirically, DeCAF statistically improves over Boltz-1(x) in both accuracy (RMSD) and physical validity scores of protein-ligand poses at strict NFE budgets on the challenging Runs N’ Poses, while also showing a more optimal Pareto frontier across all inference compute budgets on PoseBusters.


FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance

Hyunsuk Chung ⋅ Caren Han ⋅ Seungyeon Ji ⋅ Jinwoo Kim ⋅ Eun-Jung Holden ⋅ Kyungreem Han

Multimodal foundation models integrate heterogeneous signals across modalities, yet it remains poorly understood how their predictions depend on specific internal feature groups and whether such reliance can be deliberately controlled. Existing studies of shortcut and spurious behavior largely rely on post hoc analyses or feature removal, offering limited insight into whether reliance can be modulated without altering task semantics. We introduce FiLoRA (Focus-and-Ignore LoRA), an instruction-conditioned, parameter-efficient adaptation framework that enables explicit control over internal feature reliance while keeping the predictive objective fixed. FiLoRA decomposes adaptation into feature group-aligned LoRA modules and applies instruction-conditioned gating, allowing natural language instructions to act as computation-level control signals rather than task redefinitions. Across text-image and audio-visual benchmarks, we show that instruction-conditioned gating induces consistent and causal shifts in internal computation, selectively amplifying or suppressing core and spurious feature groups without modifying the label space or training objective. Further analyses demonstrate that FiLoRA yields improved robustness under spurious feature interventions, revealing a principled mechanism to regulate reliance beyond correlation-driven learning.


Fine-Tuning Language Models to Know What They Know

Sangjun Park ⋅ Elliot Meyerson ⋅ Xin Qiu ⋅ Risto Miikkulainen

Evaluating true metacognition in Large Language Models (LLMs) is difficult due to biases and heuristics. This paper presents a framework to measure and enhance LLM metacognition while controlling for these biases. A measurement method using the $d_{\rm{type2}}'$ metric is systematized to isolate metacognitive ability. The Evolution Strategy for Metacognitive Alignment (ESMA) is proposed, demonstrating robust generalization across unseen datasets, languages, and newly acquired knowledge. Finally, parameter analysis reveals that these improvements are driven by a sparse set of parameters, offering new pathways for targeted metacognitive optimization.


FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation

Chenxi Zhang ⋅ Ziliang Gan ⋅ Liyun Zhu ⋅ Youwei Pang ⋅ Qing Zhang ⋅ Rongjunchen Zhang

The financial domain poses substantial challenges for Vision-Language Models (VLMs) due to specialized chart formats and knowledge-intensive reasoning requirements. However, existing financial benchmarks are largely single-turn and rely on limited question formats, which hinders comprehensive evaluation in realistic application scenarios. To address this gap, we propose FinMTM, a multi-turn multimodal benchmark that expands diversity along both data and task dimensions. On the data side, we curate and annotate 11,133 bilingual (Chinese and English) financial QA pairs grounded in commonly used financial visuals, including candlestick charts, statistical plots, and report figures. On the task side, FinMTM covers single and multiple choice questions, multi-turn open-ended dialogues, and agent-based tasks. We further design task-specific evaluation protocols, including a set-overlap scoring rule for multiple choice questions, a weighted combination of turn-level and session-level scores for multi-turn dialogues, and a composite metric that integrates planning quality with final outcomes for agent tasks. Extensive experimental evaluation of 22 VLMs reveal their limitations in fine-grained visual perception, long-context reasoning, and complex agent workflows.

Modern models can have high average accuracy and benign loss tails while remaining fragile on rare subpopulations. We identify an information-geometric mechanism: after nuisance projection, the task direction can become locally nonidentifiable even when loss and raw Fisher look well behaved. Fisher-Glass certifies this failure by applying CVaR to inverse nuisance-projected task information. We prove a loss--information separation theorem and show that, under a boundary-mass condition, the hidden-environment tail of inverse projected Fisher controls robust classification sample complexity. We then derive stable ridge certificates and repair principles: reserves lift Fisher-null tails, portfolios increase collapse codimension, and tail-transverse influence selects counterfactual repairs by combining Fisher-Glass directionality with loss leverage. Synthetic experiments validate the theory; WILDS/Waterbirds diagnostics show that weakest identifiability need not mean lowest accuracy; and Waterbirds training plus counterfactual repair show that Fisher-Glass controls an identifiability axis complementary to loss.


Fixed-Point Reasoning: Stable and Adaptive Deep Looped Models

Sajad Movahedi ⋅ Shlomo Libo Feigin ⋅ Vera Milovanović ⋅ Alexander Theus ⋅ Thomas Hofmann ⋅ Valentina Boeva ⋅ T. Konstantin Rusch ⋅ Antonio Orvieto

Looped-in-depth architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The effective depth reached by looping determines the quality of the solution. Similar to deep architectures, looped architectures are prone to signal propagation issues as the halting decision is postponed. In this paper, we address these signal propagation issues by using pre-norm layers and residual scaling. Furthermore, we propose FPRM: a Fixed-Point Reasoning Model that uses fixed-point convergence as an end-to-end halting mechanism in a looped architecture. We show that fixed-point halting allows FPRM to adapt its compute to the difficulty of the task. FPRM proves effective on common reasoning benchmarks, namely sudoku, maze, and state tracking.


FLAME: Flow Enhanced Legendre Memory Models for General Time Series Forecasting

Xingjian Wu ⋅ Hanyin Cheng ⋅ Xiangfei Qiu ⋅ Zhengyu Li ⋅ Jilin Hu ⋅ Chenjuan Guo ⋅ Bin Yang

In this work, we introduce FLAME, a family of extremely lightweight and capable Time Series Foundation Models, which support versatile forecasting tasks via generative probabilistic modeling, while ensuring both efficiency and robustness. FLAME utilizes the Legendre Memory for strong generalization capabilities. Through adapting variants of Legendre Memory, i.e., translated Legendre (LegT) and scaled Legendre (LegS), in the Encoding and Decoding phases, FLAME can effectively capture the inherent inductive bias within data and make efficient long-range inferences. To enhance the accuracy of probabilistic forecasting while keeping efficient, FLAME adopts a Normalizing Flow based forecasting head, which can model the arbitrarily intricate distributions over the forecasting horizon in a generative manner. Comprehensive experiments on four well-recognized benchmarks, including TSFM-Bench, ProbTS, TFB, and GIFT-EVAL, demonstrate that FLAME is a strong out-of-the-box tool of decision intelligence.


FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries

QiJie You ⋅ Hao Liang ⋅ Mingrui Chen ⋅ Bohan Zeng ⋅ Meiyi Qiang ⋅ Zhen H Wong ⋅ Wentao Zhang

As video becomes increasingly central to information dissemination and multimodal large language models (MLLMs) continue to advance, evaluating video retrieval has become increasingly important. In realistic search scenarios, this requires matching short user queries to long-form content using both visual and auditory evidence. Yet existing retrieval benchmarks are still dominated by short clips, single modalities, and caption-based evaluation. We introduce FLARE, a full-modality long-video audiovisual retrieval benchmark with user-simulated queries. Built from 399 carefully screened Video-MME videos (10--60\,min, 225.4\,h) to ensure source quality and diversity, FLARE contains 87,697 clips annotated with vision, audio, and unified audiovisual captions, together with 274,933 user-style queries. Cross-modal queries are further filtered by a hard bimodal constraint, requiring retrieval to fail under either modality alone but succeed when both are combined. FLARE evaluates models under two regimes, caption-based and query-based retrieval, across vision, audio, and unified audiovisual settings. Experiments with 15 representative retrievers show that user-style queries substantially change model behavior, strong caption-based performance does not always transfer to query-based retrieval, and audio--language alignment remains a key bottleneck for unified audiovisual retrieval. Our code and data are released at https://anonymous.4open.science/r/FLARE-950E/ and https://huggingface.co/datasets/AnonymousFLARE/FLARE


Flash PD-SSM: Memory-Optimized Structured Sparse State-Space Models

Aleksandar Terzic ⋅ Francesco S. Carzaniga ⋅ Nicolas Menet ⋅ Yannick Biehl ⋅ Michael Hersche ⋅ Thomas Hofmann ⋅ Abbas Rahimi

State-space models (SSMs) face a fundamental trade-off between efficiency and expressivity that is mainly dictated by the structure of the model's transition matrix. Unstructured transition matrices enable maximal expressivity, as measured by their ability to model finite-state automaton (FSA) transitions, but come at a prohibitively high compute and memory cost. In contrast, diagonal and other structured transition matrices suffer from limited expressivity, but are highly efficient both in runtime and memory consumption. Building on recent work on structured sparse SSMs, we propose Flash PD-SSM, a novel SSM that achieves comparable throughput to diagonal models with optimal expressivity guarantees. Flash PD-SSM keeps a collection of structured sparse matrices, from which a single matrix is selected at each time-step, enabling FSA expressiveness at the level of unstructured matrices while maintaining the efficiency required for training models at scale. First, we validate Flash PD-SSM against a suite of alternative models on common mechanistic and synthetic state-tracking tasks, showing that its theoretical expressivity is achieved in practice. Moreover, on multivariate time-series tasks involving sequences of length over 17,000, Flash PD-SSM defines a new state-of-the-art (SoTA) accuracy among competing SSM methods. Second, we demonstrate that Flash PD-SSM is an effective drop-in replacement for hybrid LLMs, yielding improvements both in natural language state-tracking and in common language modeling scenarios. Finally, we show that our highly efficient design results in increased throughput and decreased memory consumption with respect to SoTA SSMs widely used in frontier language models.


FlashPlanner: Real-Time Goal-Conditioned Flow-Matching Planning for Autonomous Driving with Online RL Fine-Tuning

Qifeng Li ⋅ Yubing Gao ⋅ Xiaosong Jia ⋅ Zhiliu Liu ⋅ Sizhuo Zhou ⋅ Wenlong Liao ⋅ Tao He ⋅ Junchi Yan

Diffusion and flow matching have emerged as expressive generative planners for autonomous driving planning, owing to their ability to model high-fidelity and multi-modal trajectory distributions. Nevertheless, existing generative planners are predominantly optimized through imitation learning, which induces a fundamental mismatch between supervised trajectory fitting and closed-loop planning metrics, while their iterative sampling procedures often impose substantial computational overhead. In this paper, we propose FlashPlanner, a goal-conditioned flow-matching planner with online RL finetuning for closed-loop AD planning. FlashPlanner introduces a continuous future goal point as a compact navigation interface for the generative planning policy. This goal is produced by a continuous goal predictor, which is first pretrained and subsequently optimized with RL using closed-loop feedback from multiple candidate goals at each decision step. To support efficient online reinforcement fine-tuning and real-time deployment, FlashPlanner adopts a data-prediction flow-matching objective and removes redundant architectural components in existing diffusion-based planners. Experiments on the closed-loop nuPlan and interPlan benchmarks demonstrate that \textit{FlashPlanner} achieves state-of-the-art planning performance while delivering 6× faster inference (166 FPS) than the previous SOTA baseline (28 FPS). We will open-source our project.


FLoRA-Chef: Making A Good LoRA Recipe in Federated Generalization

Wenwen He ⋅ Wenke Huang ⋅ Yiyan Qi ⋅ He Li ⋅ Ming Hu ⋅ Yi Liu ⋅ Guansong Pang

Low-Rank Adaptation (LoRA) is a popular parameter-efficient fine-tuning (PEFT) method that adapts the large model with few trainable parameters. Federated LoRA extends LoRA to federated learning (FL), enabling clients to collaboratively fine-tune a shared model without sharing their raw data by exchanging LoRA parameters. However, due to distributed data heterogeneity, local LoRA updates are highly inconsistent across clients, hindering robust global generalization. Moreover, common direct aggregation introduces aggregated bias, yielding noisy updates and further harming the global generalization. Existing bias-mitigation approaches often rely on symmetrically alternating LoRA modules, overlooking matrix distinct characteristics during training. Therefore, we propose FLoRA-Chef, an asymmetric Federated LoRA that decides when and how to aggregate LoRA components. FLFire adaptively selects the next aggregation target by tracking cross-clientupdate dispersion. FLSauce applies matrix-distinct reweighting to control diversity and prioritize stability. Experiments on various scenarios demonstrate consistent improvements, yielding an average gain of 6.94% over counterparts, validating the benefit of adaptive asymmetric aggregation for federated LoRA.


Flow Equivariant State Space Model

Chengrong Ye ⋅ Xinyan Gong ⋅ Lingwei Zhang ⋅ Yue Song

Natural sequences often contain signals that move over time, such as translating objects, rotating digits, and drifting storm cells. Selective state space models such as Mamba process such data efficiently, but their recurrent state is updated at fixed spatial coordinates. As a result, evidence from a moving object can be accumulated across inconsistent locations rather than in the object's moving reference frame. To address this problem, we introduce Flow-SSM, a flow-equivariant selective state space model that lifts recurrent memory to candidate motion branches and transports each branch along its corresponding flow before applying the selective update. To handle real-world sequences where broad motion and local deformation often occur at different scales, we further introduce Hierarchical Flow-SSM, which organizes flow-aligned memory in a coarse-to-fine hierarchy rather than a single flat representation. We validate our architecture progressively: demonstrating exact transformation capture on Moving and Rotating MNIST, interpretable flow structures on KTH action videos, and the necessity of our multi-scale design for complex forecasting on the SEVIR weather dataset. Beyond merely improving predictive accuracy, our analyses confirm that the explicit flow alignment and hierarchical communication contribute measurably to the observed gains. These results support transported recurrence as an effective inductive bias for spatiotemporal sequence modeling.


FlowLeak: Coverage-Guided Extraction of Dynamic Workflows in LLM-Based Multi-Agent Systems

Zhiyao Ren ⋅ Siyuan Liang ⋅ Yibing Zhan ⋅ Jun L Tan ⋅ Xiaobing Sun ⋅ Liangli Zhen ⋅ Baosheng Yu ⋅ Dacheng Tao

Large language models (LLMs)-based multi-agent systems (MAS) coordinate specialized agents through prompts, tools, and communication topologies, making their hidden workflows valuable intellectual property and security-critical assets. Existing black-box MAS extraction methods implicitly assume that adversarial queries can traverse all agents, which holds for static workflows but breaks down in dynamic workflows whose execution paths depend on input semantics and intermediate states. We identify two key challenges in dynamic workflows: **branch overfitting**, where fully adversarial queries overfit to the same branch and extract only a subset of agents, and **stealthy coverage exploration**, where the adversary needs to achieve complete branch coverage with few redundant queries while not exposing the extraction task. To address these challenges, we propose *FlowLeak* that combines **Task-Preserving Payload Template**, which preserves legitimate task semantics while eliciting workflow information to mitigate branch overfitting, with **Coverage-Guided Branch Exploration**, which uses previously extracted workflow fragments to generate branch-targeted tasks and constrains workflow extraction as an auxiliary task requirement, thereby reducing exploration queries and making extraction harder to identify. Experiments on 102 MAS show that *FlowLeak* addresses both challenges and substantially improves dynamic workflow extraction (2.68$\times$ improvement). Furthermore, we show that the extracted workflow information from *FlowLeak* enhances downstream attacks, highlighting the security risks of MAS workflow extraction and our method.


Flow Matching from Viewpoint of Proximal Operators

Kenji Fukumizu ⋅ Wei Huang ⋅ Han Bao ⋅ Shuntuo Xu ⋅ Nisha Chandramoorthy

We reformulate Optimal Transport Conditional Flow Matching (OT-CFM), showing that it admits an exact proximal form via an extended Brenier potential, without assuming that the target distribution has a density. In particular, the mapping to recover the target point is expressed by a proximal operator, which yields an explicit proximal expression of the vector field. We also discuss the convergence of minibatch OT-CFM to the population OT formulation as the sample and batch sizes increase. Using second epi-derivatives of convex potentials, we prove that, for manifold-supported targets, the manifold structure is stable by perturbation to the dynamics: after time rescaling, the dynamics contracts exponentially in directions normal to the manifold while remaining neutral along tangential directions.


FlowMoP: Stochastic Multi-Person Motion Prediction

Aadya Agrawal ⋅ Ho Kei Cheng ⋅ Alex Schwing

Predicting the future motion of multiple people is inherently stochastic: the same observed history admits many plausible continuations depending on intent, coordination, and social context. Yet the dominant paradigm in multi-person 3D motion prediction remains deterministic, with existing methods producing a fixed number of outputs per agent. In contrast, we present a conditional flow matching framework that jointly generates diverse future motions for multiple agents, modeling global trajectories and root-relative poses. A B-spline reparameterization compresses the generative state space while enforcing temporal smoothness, and a dual-stream transformer conditions each agent's predictions on observed neighbor motion for interaction-aware forecasting. Our method achieves state-of-the-art performance on CMU-Mocap (UMPM) and 3DPW, outperforming deterministic baselines on joint position error, pose error, and final displacement error despite producing a full predictive distribution. We additionally establish the first motion prediction baseline on WorldPose, a large-scale professional soccer dataset, where social conditioning demonstrably concentrates predictive uncertainty along interaction-constrained directions.

Functional near-infrared spectroscopy (fNIRS) is a promising modality for brain-computer interfaces and cognitive state decoding, but progress in machine learning for fNIRS has been limited by small, isolated datasets and inconsistent evaluation protocols. We introduce fNIRSAtlas, the largest open benchmark for fNIRS classification to date, comprising 23 public datasets spanning diverse experimental paradigms, including motor execution, motor imagery, mental arithmetic, working memory, and emotion recognition. fNIRSAtlas provides a fully automated pipeline from raw data to evaluation and defines standardized benchmark tasks for both cross-subject and within-subject classification, ensuring reproducible comparison of methods. We evaluate a variety of baseline models from fNIRS and related neuroimaging and time-series machine learning literature on the benchmark. Our results demonstrate that performance varies substantially across datasets and that methods validated on limited data -- especially deep learning approaches -- often fail to generalize beyond their original evaluation. fNIRSAtlas serves as an extensible, open, and reproducible benchmark to support more robust and cumulative progress in machine learning for fNIRS data. Our code and data can be found at https://anonymous.4open.science/r/fNIRSAtlas-8842.


Focal Reward: Balanced Reinforcement Learning under Rubric-Based Rewards

Yu Huang ⋅ Zihua Zhao ⋅ Zhaoxin Huan ⋅ Wanli Gu ⋅ Feng Hong ⋅ Xinmu Ge ⋅ Lin Yuan ⋅ Qiang Hu ⋅ Weichang Wu ⋅ Xiaolu Zhang ⋅ Jun Zhou ⋅ Jiangchao Yao

The open-ended generation in LLMs usually requires multi-dimensional rubrics to adequately assess quality and guide the improvement of reinforcement learning. However, a critical dilemma inherent in this training paradigm is the imbalanced reward polarization along different rubric dimensions. Under this bottleneck, even if LLMs achieve relatively high rewards after training, they may still exhibit severe deficiencies in certain dimensions, leading to a direct deterioration in user experience. To address this problem, we propose \emph{Focal Reward}, a novel objective to automatically balance the training of reinforcement learning under rubric-based rewards. Specifically, we first leverage an inverse reward projection mechanism to estimate the saturation degree of each criterion in the rubric, which forms the basis to calibrate the reward direction. Then, the final objective is designed with an automatically reweighting coefficient for each criterion to achieve the fine-grained balancing. Extensive experiments across three model scales and six benchmarks demonstrate that our \emph{Focal Reward} method outperforms the strongest static aggregation baseline in all 18 model--benchmark comparisons. Rollout, mechanism, and ablation analyses further show that these gains arise from online, saturation-aware reallocation toward rubrics that still have room for improvement.


Focusable Monocular Depth Estimation

Yuxin Du ⋅ Tao Lin ⋅ Zile Zhong ⋅ Runting Li ⋅ Xiyao Chen ⋅ Jiting Liu ⋅ Chenglin Liu ⋅ Yingcong Chen ⋅ Yuqian Fu ⋅ Bo Zhao

Monocular depth foundation models generalize well across scenes, yet they are typically optimized with uniform pixel-wise objectives that do not distinguish user-specified or task-relevant target regions from the surrounding context. We therefore introduce Focusable Monocular Depth Estimation (FDE), a region-aware depth estimation task in which, given a specified target region, the model is required to prioritize foreground depth accuracy, preserve sharp boundary transitions, and maintain coherent global scene geometry. To prioritize task-critical region modeling, we propose FocusDepth, a prompt-conditioned monocular relative depth estimation framework that guides depth modeling to focus on target regions via box/text prompts. The core Multi-Scale Spatial-Aligned Fusion (MSSA) in FocusDepth spatially aligns multi-scale features from Segment Anything Model 3 to the Depth Anything family and injects them through scale-specific, gated conditional fusion. This enables dense prompt cue injection without disrupting geometric representations, thereby endowing the depth estimation model with focused perception capability. To study FDE, we establish FDE-Bench, a target-centric monocular relative depth benchmark built from image-target-depth triplets across five datasets, containing 252.9K/72.5K train/val triplets and 972 categories spanning real-world and embodied simulation environments. On FDE-Bench, FocusDepth consistently improves over globally fine-tuned DA2/DA3 baselines under both box and text prompts, with the largest gains appearing in target boundary and foreground regions while preserving global scene geometry. Ablations show that MSSA's spatial alignment is the key design factor, as disrupting prompt-geometry correspondence increases AbsRel by up to 13.8%.


Follow the Regularized Leader Does Not Converge in Constrained Optimization

Ioannis Anagnostides ⋅ Ioannis Panageas ⋅ Nikolas Patris ⋅ Tuomas Sandholm

Follow the regularized leader (FTRL) is a foundational algorithm in online learning whose regret properties have been extensively studied for decades. However, in stark contrast to mirror descent---and gradient descent in particular---the convergence of FTRL to first-order stationary points in constrained nonconvex optimization was hitherto unresolved. In this paper, we show that continuous-time FTRL can fail to converge even asymptotically in constrained optimization problems; that is, the Karush-Kuhn-Tucker (KKT) gap of FTRL can remain bounded away from zero indefinitely. This is especially surprising in light of the fact that continuous-time FTRL guarantees a monotonic improvement of the objective value. As a result, we establish that non-convergent behavior of no-regret dynamics---which besets general variational inequality problems---can occur even in potential systems. From a technical standpoint, our construction relies on an infinitely differentiable function where the gradient flow dynamics exhibit decelerating oscillations without ever converging pointwise, which we in turn embed into higher-dimensional FTRL dynamics that sustain a perpetually large KKT gap.


Forecasting Microbial Dynamics: Evaluation Protocol and Prior-Spectrum Benchmark

Fedor Sergeev ⋅ Anna Győrffy-Kerekes ⋅ Tristan Gollmart ⋅ Vincent Fortuin ⋅ Andre Kahles ⋅ Gunnar Rätsch

Modern forecasting methods are often designed and evaluated on datasets that do not reflect the realities of real-world clinical time-series: inter-patient variability, feature constraints, irregularity, and sparse sampling. Microbiome dynamics modeling represents these challenges, has seen little attention from the machine learning community, and has practical significance in medicine. Moreover, it offers both deep domain knowledge and sufficient data to train deep models, making it a strong testbed for a wide variety of forecasting methods. We propose a unified evaluation protocol for forecasting microbial dynamics and conduct MicroCastBench: a benchmark for models covering a wide range of priors (from mechanistic to foundation models) and assumptions (from per-patient to meta-learning models). We find that meta-learning models lead across most scaling regimes and at longer horizons, with mechanistic priors winning at small cohort sizes, short trajectories, and high feature counts. These results suggest that when the target data is far from pretraining distributions, domain-informed and meta-learning approaches can outperform generalist foundation models, highlighting the value of evaluating on under-represented domains.


FoundSurface: Feedforward Ray-Based 3D Scene Surface Reconstruction from Unposed Images

Shenxing Wei ⋅ Jinxi Li ⋅ Zihui Zhang ⋅ Ajay Kumar ⋅ Wei Wang ⋅ Bo Yang

Accurate 3D surface reconstruction from unposed images remains a fundamental challenge in computer vision. Recent feed-forward pointmap methods eliminate the need for camera poses but inherently produce discrete, sparse representations. In this paper, we present a generalizable pipeline that directly recovers continuous, high-fidelity 3D surfaces from unposed images by seamlessly integrating pointmap- and ray-based representations. Our approach first extracts explicit geometric priors using a pre-trained pointmap model. A query-view-centric selection module then identifies geometrically relevant reference views via a robust surfel voting mechanism. Next, a cross-attention regressor bridges target query rays with reference features to estimate a continuous coarse surface. Finally, a raylet-based module aggregates local 3D features to carve out high-frequency details and analytical normals. Trained as a single foundation model on 9 diverse datasets, our method demonstrates exceptional zero-shot generalization on 6 unseen datasets, clearly surpassing state-of-the-art baselines in geometric accuracy and detail preservation.

Frame selection---picking a small set of informative frames from a video to feed a vision-language model (VLM)---sits at the front of every video-VLM system, yet the literature is fragmented: most works do not share a candidate pool, encoder, frame budget, or answering VLM, making published gains hard to compare across works. We introduce FrameSelect, to the best of our knowledge the first Python library that treats the selector as a standalone, interchangeable component, cleanly separated from the encoder that computes representations and the VLM that answers. Any registered selector composes with any registered evaluator through a single call, and adding a new selector, dataset, encoder, or VLM backend is an isolated extension; a replay mode further rebuilds any table under a different answering VLM without re-running selection. We release reference implementations of 20 selectors, 15 evaluators spanning answer-level and selection-level metrics, and the first unified benchmark comparing frame selectors under a shared protocol across seven evaluators spanning both evaluation modes. This comparison surfaces findings that existing studies cannot: rankings invert across benchmarks. Our library is available at https://anonymous.4open.science/r/FrameSelect-D2A4.

Gradient-based attribution methods are model-faithful and scalable, but Integrated Gradients (IG) can be brittle because explanations depend on heuristic baselines, straight-line paths, discretization, and saturation. We propose Fisher--Rao Integrated Gradients (FRInGe), which defines both the reference and interpolation schedule in predictive distribution space. FRInGe replaces input baselines with a maximum-entropy predictive reference and follows a Fisher--Rao geodesic on the probability simplex. The corresponding input-space trajectory is realized through the pullback Fisher metric and stabilized by KL and Euclidean trust regions; attributions are obtained by integrating input gradients along this trajectory. Across six ImageNet architectures, FRInGe most clearly improves calibration-oriented attribution metrics, especially MAS scores, while remaining competitive on perturbation AUC and infidelity.


From Chats to Markets: AgenticPay for LLM-Powered Negotiation in Multi-Agent Commerce

Xianyang Liu ⋅ Shangding Gu ⋅ Fan Xu ⋅ Manxi Wu ⋅ Boyi Li ⋅ Jun Wang ⋅ Costas J Spanos ⋅ Dawn Song

Agents based on large language models are increasingly expected to autonomously handle negotiation and transactions. However, existing benchmarks predominantly focus on text-only, bilateral, zero-sum price haggling, failing to capture the complexity of real-world commerce. To address this gap, we introduce Agenticpay, a unified framework and benchmark for evaluating how well multimodal agents reach high-welfare, multi-dimensional agreements in realistic markets. Built around four core components ( Environments, Tasks, Agents, and Metrics), Agenticpay comprises 160 multimodal tasks spanning 4 real-world business scenarios (E-commerce, Food Delivery, Ride-hailing, and Apartment Rental) and 8 market topologies, scaling from 1-to-1 bargaining to many-to-many (N-to-N) competitive markets. Beyond price haggling, agents must read product images, infer each opponent's hidden preferences, and trade off multiple binding contract terms (e.g., price, lease duration, return policy, delivery speed). Agents communicate through multi-round natural-language dialogue, with each turn proposing or revising a full contract, and outcomes are scored by a utility-based framework that rewards agreements maximizing joint welfare. Evaluations on state-of-the-art proprietary (GPT-5.4, Claude Sonnet 4.6, Gemini 3.1 Pro Preview) and open-weight (Qwen3-VL-32B-Instruct, InternVL3-38B) multimodal models reveal substantial gaps in non-zero-sum value creation and market reasoning: even the strongest agent (Gemini 3.1 Pro Preview) reaches a GlobalScore of only 42.3/100, and performance consistently degrades as markets scale from bilateral to multi-sided, with the average GlobalScore dropping by 5.7 points and open-weight models suffering the largest declines (up to 11.5 points). These findings establish Agenticpay as a foundational testbed for multimodal agentic commerce. Code and dataset are available at: https://anonymous.4open.science/r/AgenticPay-4BBB/README.md.


From Generation to Restoration: Residual Diffusion for Neural Channel Decoding

Qinshan Zhang ⋅ Shipeng Guo ⋅ Xuantai Wu ⋅ Bin Chen ⋅ Zhuochen Fan ⋅ Yong Jiang ⋅ Shu-Tao Xia ⋅ Qing Li

Neural decoders have shown strong potential for error correction in short- and moderate-length regimes, yet a fundamental tension remains between decoding accuracy and inference latency. Recent attempts to leverage diffusion probabilistic models for channel decoding typically adopt a fully generative paradigm, initializing the reverse process from an observation-agnostic prior, which overlooks a key structural property of channel decoding: the received signal already contains substantial information about the target codeword. In this work, we propose Channel Residual Diffusion Model (ChRes-DM), a principled diffusion-based decoding framework that reinterprets channel decoding as a directed restoration process anchored at the noisy observation. Instead of sampling from a generic prior, ChRes-DM constructs a geometric residual bridge between the received signal and the clean codeword, explicitly modeling the conditional transport induced by the channel and leading to a deterministic Probability Flow Ordinary Differential Equation (PF-ODE) that governs the decoding dynamics. By eliminating redundant stochastic sampling inherent in conventional diffusion models, ChRes-DM enables efficient iterative decoding with flexible step-skipping, offering fine-grained control over the accuracy–latency trade-off. Extensive experiments across representative channel coding benchmarks demonstrate that ChRes-DM achieves competitive or superior decoding performance compared to existing neural decoders while significantly reducing inference iterations, highlighting diffusion-based residual transport as a promising and scalable paradigm for neural channel decoding.


From Intent to Evidence: A Categorical Approach for Structural Evaluation of Deep Research Agents

Shuoling Liu ⋅ Zhiquan Tan ⋅ Kun Yi ⋅ Hui Wu ⋅ Yihan Li ⋅ Jiangpeng Yan ⋅ Liyuan Chen ⋅ Kai Chen ⋅ Qiang Yang

Deep Research Agents (DRAs) answer complex questions by searching the web, checking evidence, and synthesizing conclusions across heterogeneous sources. We introduce a category-theoretic framework for evaluating such agents. The framework treats deep research as a structured mapping from user intent to evidence-grounded conclusions, making retrieval traces, cross-source alignment, and final synthesis explicit. Guided by this view, we build a mechanism-aware benchmark of 296 bilingual questions covering four structural skills: following multi-hop evidence chains, verifying claims across sources, re-ordering fragmented information, and rejecting unsupported assumptions. We evaluate 16 systems with human verification and find that these tasks remain difficult: the best system reaches 19.9\% average accuracy. The results reveal complementary strengths across systems, but also persistent weaknesses in long-horizon retrieval and intersection-heavy verification. We further instantiate two theory-guided interventions, tracked search and category tools, in API-based agents. These variants improve over their corresponding baselines, suggesting that the framework is useful not only for diagnosis but also for modest, targeted system design.


From Matching to Reasoning: Query-Aware Long Video Summarization

Mingu Kang ⋅ Sumin Kim ⋅ Hyunjin Lee ⋅ Sungmin Yang ⋅ Yoori Oh ⋅ Joonseok Lee

Query-aware video summarization (QVS) aims to select a concise set of moments relevant to a textual query in the context of the entire video. Existing QVS methods are often driven by local query-segment relevance scoring, which is effective for retrieving query-matching moments but insufficient for composing coherent summaries over long videos. Long-form QVS is especially challenging because query-relevant evidence tends to be sparsely distributed across extended timelines, and satisfying a query may require combining complementary segments while avoiding redundancy under a strict summary budget. We propose QLVSumm, a framework that autoregressively generates a query-conditioned video summary, capturing inter-segment dependencies as well as the saliency of the segment and relevance to the query. We also introduce a large-scale benchmark for long-form QVS (LQVS) with open-vocabulary queries and budget-constrained extractive summaries, together with a complementary metric for evaluating query-aware summary disentanglement. Experiments show that QLVSumm achieves state-of-the-art performance and demonstrates stronger query-aware summary disentanglement.


From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs

Yuchuan Tian ⋅ Yuchen Liang ⋅ Shuo Zhang ⋅ Yingte Shu ⋅ Guangwen Yang ⋅ Wei He ⋅ Sibo Fang ⋅ Tianyu Guo ⋅ Kai Han ⋅ Chao Xu ⋅ Hanting Chen ⋅ Xinghao Chen ⋅ Yunhe Wang

Diffusion language models (DLMs) can generate multiple tokens in parallel, but training large DLMs from scratch remains expensive. A practical alternative is to adapt off-the-shelf autoregressive (AR) checkpoints into diffusion models, reusing their linguistic and reasoning capabilities. Existing adaptation recipes either modify logits and grow attention masks toward full-sequence diffusion, or directly fine-tune ARweights under a block-diffusion objective, leaving two questions underexplored: what diffusion paradigm should AR-to-DLM adaptation target, and what transition path preserves AR knowledge most effectively? We argue that Block-Diffusion is a natural destination because AR decoding corresponds to block size one at the level of attention and generation order, while larger blocks introduce controlled intra-block bidirectionality and parallel generation. Based on this view, we propose a context-causal adaptation path that keeps committed context strictly causal, a one-pass parallel training formulation with auxiliary AR guidance, and a gradual block-size curriculum. Across several AR initializations and model scales, these components improve average adaptation performance over random mask annealing and direct fine-tuning. Scaling the recipe yields NBDIFF-7B, which supports 32K-token contexts and achieves the strongest average performance among the compared diffusion LLM baselines on general, math, and code benchmarks. Code and checkpoints will be released upon publication.


From Post-Hoc to Ante-Hoc: Consistently Explainable Semi-Supervised Time Series Classification

Viet-Hung Tran ⋅ Zichi Zhang ⋅ Ngoc Phu Doan ⋅ Xuan Hoang Nguyen ⋅ Phi Hung Nguyen ⋅ Yimeng An ⋅ Peixin Li ⋅ Huynh Thi Khanh Chi ⋅ Hans Vandierendonck ⋅ Ira Assent ⋅ Thai Son Mai

Semi-Supervised Time Series Classification (SS-TSC) with Deep Neural Networks (DNNs) achieves strong accuracy by leveraging unlabeled data, yet the resulting models remain black boxes that offer no insight into which features drive each prediction. Existing post-hoc explanation methods for TSC can identify important features but provide no guarantee that the model genuinely relies on them, while ante-hoc approaches impose architectural constraints that limit flexibility in existing model usages. Neither setting explore the semi-supervised regime, where limited labels can affect explainability. We propose \textbf{Sensory Gating}, the first SS-TSC framework that simultaneously bridges post-hoc attribution and ante-hoc explanation through a four-stage pipeline: (i) a base classifier is trained with our proposed Forecasting Joint-Embedding Predictive Architecture (F-JEPA) auxiliary objective; (ii) post-hoc attribution maps are converted into per-instance binary masks via greedy selection of sufficient features; (iii) a lightweight \textbf{Sensory Gate} amortizes these masks through Binary Cross Entropy (BCE) supervision with data-driven threshold calibration; and (iv) a knowledge distillation stage to provide an explainable SS-TSC student model that operates on gated (masked) input, while consistently reproduces the original black box teacher (F-JEPA)'s predictions. We validate our framework on eight datasets across four label ratios with five seeds per configuration. F-JEPA achieves better performance compared to recent state-of-the-art (SOTA) semi-supervised TSC methods, while the sensory gate reduces the fraction of features the classifier requires to 20--45\% with competitive accuracy. On MIT-BIH, where ground-truth saliency annotations are available, the gate's selected features align with domain-relevant regions while classification performance remains competitive with the ungated baseline at 10--20\% kept ratio.


From “Weak” Signals to Strong Models: Preference Delta Aggregation with LoRA Merging

Qi Sun ⋅ Siyue Zhang ⋅ Yulin Chen ⋅ Yuxiang Xue ⋅ Ru Peng ⋅ Chen Zhao

Training strong large language models (LLMs) requires high-quality supervision, which is often scarce. Recent work shows that paired preference data from weak–weaker model pairs (e.g., Qwen3 4B over 1.7B), despite the limited quality of individual responses, can provide an effective supervision signal through relative quality deltas, which we term a "weak" signal. This motivates a key research question: can multiple "weak" signals be constructively aggregated for improving strong models (e.g., Qwen3 8B)? To this end, we propose Preference Delta Aggregation (PDA), the first framework that derives a preference delta from each weak-weaker model pair, instantiates it as a LoRA adapter learned through preference optimization, and aggregates the resulting deltas via LoRA merging. To further mitigate directional interference during LoRA merging, we introduce Geometric Alignment Merging (GAM), a geometry-aware merging method that aligns adapter subspaces before aggregation, enabling more robust composition of diverse deltas. Evaluations on knowledge reasoning and agentic search benchmarks show that aggregating multiple "weak" signals pushes performance beyond any single signal, with further gains as additional signals are incorporated. Correspondingly, PDA with GAM improves the strong model by 6.8 and 7.3 points on average for knowledge reasoning and agentic search, respectively. It outperforms all single-delta and multi-delta baselines, exceeding the best single-delta baseline by 2.1 and 4.3 points. Further analysis attributes these gains to the effective composition of complementary capabilities encoded across distinct preference deltas.


FTerViT: Fully Ternary Vision Transformer

Szymon Ruciński ⋅ Pietro Bonazzi ⋅ Engin Turetken ⋅ Simon Narduzzi ⋅ Michele Magno ⋅ Nadim Maamari

Ternary Vision Transformers offer substantial model compression, however state-of-the-art methods only ternarize the encoder layers, leaving patch embeddings, LayerNorm parameters, and classifier heads in full precision. In compact models targeting resource-constrained processors, such as microcontrollers, these remaining full-precision components determine the total memory footprint, severely limiting deployment efficiency and on-device feasibility. In this work, we introduce a fully ternarized Vision Transformer in which \emph{all} weight matrices and normalization parameters are ternarized (FTerViT). To this end, we introduce two novel operators : TernaryBitConv2d with per-channel scaling for patch embedding and TernaryLayerNorm. FTerViT is trained using knowledge distillation, followed by a lightweight quantization-aware recovery phase. Our ternary W2A8 DeiT-III-S at 384$\times$384 resolution achieves 82.43\% ImageNet-1K top-1 at 6.09\,MB (${\sim}$15$\times$ compression, $-$2.42\,pp vs.\ FP32), outperforming prior ternary ViTs methods up to 8 pp. Finally, we demonstrate the first implementation of ternary vision transformers on a dual cores XTensa LX7 microcontroller inside the ESP32-S3 system-on-chip. By deploying FTerViT-Small (based on DeiT-III-Small at 224$\times$224 resolution, 5.81\,MB), we achieve 79.64\% ImageNet-1K top-1 accuracy.


Full-Duplex Speech-Motion Model for Dyadic Interaction

Koki Nagano ⋅ Hongyu Liu ⋅ Wookie Park ⋅ Tianye Li ⋅ Amrita Mazumdar ⋅ Christian Jacobsen ⋅ Shengze Wang ⋅ Michael Stengel ⋅ Ka Chun Cheung ⋅ Simon See ⋅ Shalini De Mello

We present DuplexMotion, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of human communication, this full-duplex capability empowers the agent to simultaneously perceive and generate both speech and physical motion in a streaming fashion. At its core, our method leverages the strong priors of a foundational full-duplex speech model and integrates a novel motion pathway, thereby achieving fully synchronized multi-modal interaction. Specifically, we design a dual-tower Transformer architecture that preserves the zero-shot conversational reasoning of a frozen base speech model while constructing a deeply coupled, streaming motion pathway. By introducing a unified dyadic token interleaving mechanism and guiding cross-attention via a time-aligned speech-motion RoPE, our model effectively aligns autoregressive motions with rich latent speech features. Trained on the 4,000-hour Seamless Interaction dataset, our model effectively captures cross-speaker dependencies and establishes new state-of-the-art performance across both monadic and dyadic human interaction benchmarks.

Few-shot CLIP adaptation is often treated as a parameterization problem, where full fine-tuning is assumed to overfit, so parameter-efficient fine-tuning (PEFT) methods are preferred. In this paper, we show that this conclusion conflates parameterization with optimization. Across 11 datasets, 3 CLIP backbones, and 5 shot levels, dense full fine-tuning with plain SGD outperforms representative PEFT methods, whereas Adam/AdamW collapses. We trace this collapse to unreliable support-set preconditioning. The few-shot second moment systematically underestimates the corresponding full-data statistic, already at initialization. Adam’s inverse-square-root normalization converts this error into an oversized adaptive update. Matching Adam’s first-step update budget to SGD resolves most of the collapse, whereas numerator-side alignment does not explain the failure. Shampoo and SOAP also fail, showing that the problem is not merely Adam’s diagonal approximation but a broader mismatch between support-set and full-data preconditioning geometry. Theory formalizes how few-shot estimation error is amplified into unsafe updates. SGD succeeds because it avoids unreliable preconditioners, stays close to pretrained CLIP, and follows broad, connected, low-curvature corridors. These results identify optimizer-induced preconditioning, not dense full fine-tuning itself, as the main source of failure in few-shot CLIP adaptation. Our code is included in the supplementary materials and will be made public.


FuncFormer: Circuit Representation Learning via the Flow of Functional Propagation

Yunjie Ji ⋅ Jie Wang ⋅ Zhihai Wang ⋅ Min Li ⋅ Junhua Huang ⋅ Zhihao Shi ⋅ Feng Wu ⋅ Mingxuan Yuan ⋅ Jianye Hao

Learning expressive representations for Boolean logic circuits is a fundamental challenge at the intersection of graph learning and Electronic $\mbox{Design}$ Automation (EDA). Existing Graph Neural $\mbox{Networks}$ (GNNs) primarily rely on topological message passing, which often fails to capture the strict causal dependencies and discrete functional semantics of logic gates. In this paper, we $\mbox{propose} \textbf{FuncFormer}$, a Graph Transformer that incorporates $\textbf{functional simulation}$ not as a proxy task, but as a fundamental inductive bias directly into the representation learning process. Unlike standard GNNs which typically rely on isotropic aggregation, FuncFormer $\textbf{encodes the intrinsic flow of functional propagation}$ by analyzing randomized simulation traces as they evolve through the network. This approach effectively aligns the continuous embedding manifold with the discrete Boolean function space, effectively mitigating structural aliasing. By integrating these deterministic signal trajectories with a scalable dual-path attention mechanism, our model preserves functional consistency across long-range dependencies in both combinational and sequential circuits. Empirical results demonstrate that FuncFormer significantly outperforms state-of-the-art models (e.g., DeepGate4) in Quality-of-Results (QoR) prediction and formal verification tasks, exhibiting robust generalization to unseen circuit scales.


G2Fusion: Geometric-to-Generative Image Fusion via Registration-Restoration Evolution

Hao Zhang ⋅ Douyu Wu ⋅ Han Xu ⋅ Linfeng Tang ⋅ Qiwen Jin ⋅ Jiayi Ma

Real-world multi-modal image fusion faces two fundamental challenges: spatial misalignment and complex degradations induced by heterogeneous imaging sensors. These factors are intrinsically coupled, but most existing fusion methods typically address only one of them, inevitably leading to failure under the other challenge. In this paper, we propose G2Fusion, the first fusion framework that simultaneously addresses image registration and information restoration. By developing a novel geometric-to-generative paradigm, it can directly produce high-quality fused images from unregistered and degraded inputs captured by heterogeneous imaging sensors. This framework consists of two key modules: a flow-based geometric deformation reduction module (Flow-GDR) and a DiT-based generative fusion module (DiT-GF). The former reduces large non-rigid discrepancies through dense flow estimation. The latter performs generative fusion, progressively refining residual misalignments while restoring degraded content through iterative denoising. To enable effective interaction among registration, restoration, and fusion, we design two complementary mechanisms. On the one hand, a target distribution mining strategy is introduced to construct a joint objective distribution from registration, restoration, and fusion, effectively guiding the optimization of DiT-GF. On the other hand, we develop a mutual promotion mechanism that establishes a closed-loop interaction between Flow-GDR and DiT-GF by re-estimating the residual deformation between the fused output and the infrared reference. Extensive experiments demonstrate that G2Fusion consistently outperforms state-of-the-art methods in terms of registration, restoration, and fusion.


GADMVP: Adaptive Few-Shot Graph-Level Anomaly Detection with Multi-View Structured Prompting

Xiaolin Han ⋅ Xiurui Hu ⋅ Lingyun Song ⋅ Yudai Pan ⋅ Xuequn Shang

Graph-level anomaly detection plays a critical role in a wide range of applications, such as fraud detection in financial networks and malicious program detection, yet it remains challenging in real-world scenarios where labeled anomalous graphs are extremely scarce and graph structures are highly diverse. Existing few-shot methods often suffer from severe performance degradation due to the effects of class imbalance and structural imbalance, and they struggle to generalize to structurally deviant or rare anomalies. In this paper, we propose an adaptive few-shot graph-level anomaly detection model with multi-view structure-aware prompting (GADMVP). To mitigate structural imbalance, we construct a graph-of-graphs representation that facilitates cross-graph message passing, capturing both intra-graph semantics and inter-graph dependencies. Building upon these priors, we introduce a structure-aware graph prompting mechanism that extracts and adaptively refines anomaly-relevant prompts from a small set of labeled graphs. Extensive experiments on multiple benchmark datasets demonstrate that the proposed method consistently outperforms state-of-the-art few-shot graph anomaly detection baselines, while exhibiting strong robustness under extremely limited labeled data. The code and implementation details are available at https://anonymous.4open.science/r/GLAD-3456543/.


GaitLingo: Self-Supervised Gait Representation Learning with Language Priors

Chenye Wang ⋅ Zhengxiang Lan ⋅ Saihui Hou ⋅ Zhikang Liu ⋅ Xingqun Qi ⋅ Sirui Han ⋅ Yongzhen Huang

Self-supervised gait representation learning aims to learn transferable representations that generalize to unseen scenarios. However, existing methods rely solely on visual self-supervision, making it difficult to assess the semantic value of individual unlabeled sequences and limiting their ability to exploit latent semantic cues. To address this limitation, we introduce language priors into self-supervised gait pretraining. We construct GaitLP-1M, a million-scale silhouette-based gait pretraining dataset without identity labels. A two-stage caption generation pipeline further produces gait descriptions from the corresponding RGB sequences. To the best of our knowledge, GaitLP-1M is the first million-scale silhouette gait dataset paired with gait-centric captions. Building on this foundation, we propose GaitLingo, a language-guided self-supervised framework for gait representation learning. Specifically, we design a Caption-Guided Sample Weighting (CSW) mechanism to prioritize informative samples based on sequence-level semantic richness and attribute rarity. We then introduce a lightweight Temporal Adapter (TA) to better capture motion semantics and improve cross-domain robustness. Finally, we propose Soft Relation Distillation (SRD) to transfer caption-derived relational structures into the visual embedding space. Extensive experiments demonstrate that GaitLingo consistently outperforms prior self-supervised gait pretraining methods in the zero-shot setting across six benchmarks. Compared with GaitSSB, it achieves Rank-1 gains of +12.1% on Gait3D and +7.9% on GREW, as well as an average improvement of +9.1% across four CCPG settings.


GAMMA: Scalable 4D Gaussian Reconstruction Model for Novel View Synthesis of Monocular Videos

Weiqi Zhang ⋅ Junsheng Zhou ⋅ Xuancheng Zhang ⋅ Juntong Fang ⋅ Zequn Chen ⋅ Donglin Di ⋅ Yu-Shen Liu

We present GAMMA, a scalable feed-forward model that reconstructs 4D Gaussian Splatting from monocular videos, enabling dynamic novel view synthesis in real-time. Most existing approaches for dynamic scene modeling require multi-view videos as input or costly per-scene optimization. In contrast, GAMMA learns a unified 4D scene representation from monocular videos and generatively predicts 4D Gaussian primitives within seconds. The key insight of GAMMA lies in its unified Gaussian reconstruction model, which jointly estimates temporally consistent geometry, appearance, and motion from the input video. We further explore its scalability through large-scale training on a comprehensive dataset encompassing both synthetic and real-world videos. The reconstructed GAMMA representation supports interactive scene exploration via real-time rendering across timesteps and local viewpoints. Extensive experiments demonstrate that GAMMA outperforms existing methods in both reconstruction fidelity and efficiency.


GauGal: Gaussian-Galerkin Electromagnetic Inverse Scattering Imaging

Haibing Wu ⋅ Yixiong Jing ⋅ Guangming Wang ⋅ Olaf Wysocki ⋅ Brian Sheil

Electromagnetic inverse scattering (EIS), which aims to recover object permittivity from measured scattered fields, is central to a wide range of applications, from medical diagnostics to security screening. However, it is a highly nonlinear and ill-posed problem. Despite substantial progress, existing approaches face a fundamental trade-off between physical fidelity, efficiency, and generalization: data-driven methods are fast but tied to training geometries and sensor setups, while physics-based optimization is accurate but computationally expensive. We introduce a Gaussian-Galerkin (GauGal) method that reformulates EIS from a dense point-wise field reconstruction to a compact, primitive-level physical solving. Rather than using Gaussian primitive merely as a material representation, GauGal take it as the computational units for wave scattering. By projecting the continuous EIS scattering equation into a Gaussian primitive space, the material modulation, Green propagation, source excitation, and receiver observation are all realized as primitive-level operators. This yields a compact differentiable forward solver for physics-consistent reconstruction. Our method achieves state-of-the-art accuracy on synthetic and real benchmarks while reducing runtime from over 30 minutes by leading physics-driven baselines to under 10 seconds. It also generalizes significantly better to unseen geometries and sensor configurations than leading data-driven baselines. Overall, this framework establishes an accurate, efficient, and generalizable paradigm for electromagnetic inverse imaging, enabling fast, physics-consistent imaging in practical settings.


GEAR: Bridging the Planner-Actor Gap via Gradient-Aligned Policy Extraction

Yao-Hui Li ⋅ Xin Li ⋅ Hasnaa Bennis ⋅ Meiju Li ⋅ Boya Zhang ⋅ Yingfang Yuan ⋅ Wei Pang

Hybrid model-based reinforcement learning (MBRL) integrates lookahead planning with actor-critic optimization for exceptional sample efficiency. However, this decoupled architecture suffers from a critical planner-actor mismatch. Existing policy alignment methods face a structural dilemma: forward KL triggers mode-covering objective conflicts, while reverse KL relies on brittle proxies that bottleneck expressivity. To address this challenge, we propose Gradient-Embedded Alignment Regularization (GEAR), a unified in-sample extraction framework. At its core, we derive a first-order geometric alignment metric from the continuous-time Hamilton-Jacobi-Bellman (HJB) equation to evaluate action evolution against the local value gradient. By embedding this metric as an absolute modulation weight, GEAR eliminates spurious imitation and transforms the inherently diffuse forward KL into a focused, mode-seeking objective. Evaluations on the challenging HumanoidBench suite demonstrate GEAR achieves significantly higher sample efficiency and asymptotic performance than state-of-the-art baselines. The project code is available at https://anonymous.4open.science/r/GEAR-5E80.


GenCOPE: Syn2Real Generalized Category-Level Object Pose Estimation for Robotic Picking

Jian Liu ⋅ Wei Sun ⋅ Zhenqi Dai ⋅ Hui Yang ⋅ Jian Xiao ⋅ Nicu Sebe ⋅ Na Zhao

Category-level object pose estimation (COPE), capable of generalizing to intra-class unknown objects, has become a core technique for robotic 3D scene understanding. However, existing COPE methods still require labor-intensive recollection of real-world training data for novel object categories, which limits their scalability in practical applications. This paper aims to achieve synthetic-to-real (Syn2Real) generalized COPE, where a model is trained solely on rendered synthetic data and directly generalized to real-world deployments. The central challenge lies in the significant domain gap between synthetic and real-world data, particularly in texture appearance. To address this, we aim to enhance domain generalization by learning domain-invariant representations that capture semantic commonalities among objects within the same category. We introduce 2D and 3D semantic consistency constraints to reduce the sensitivity of feature encoders to domain-specific features. In addition, we propose an end-to-end pose regression framework that performs 2D-3D cross consistency learning, leveraging dense cross-modality fusion to further refine pose estimation. Since simplicity and effectiveness are essential for real-world robotic deployment, our model operates exclusively on global features, yielding a highly lightweight and efficient architecture. Extensive experiments on the REAL275 and Wild6D benchmarks, as well as real-world robotic manipulation scenes, show superior Syn2Real generalization performance of our paradigm.

Bayesian Reinforcement Learning (BRL), a subclass of Meta-Reinforcement Learning (Meta-RL), provides a principled framework for generalisation by explicitly incorporating Bayesian task parameters into transition and reward models. However, classical BRL methods assume known forms of transition and reward models. While recent deep BRL methods incorporate model learning to address this, applying neural networks directly to joint data and task parameters necessitates variational inference. This often yields indistinct task representations, compromising the resulting BRL policies. To overcome these limitations, we introduce Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions (GLiBRL). Our approach features fully tractable Bayesian inference over task parameters and model noise, alongside exact marginal likelihood evaluation for learning transition and reward models. The permutation-invariant nature of exact Bayesian inference in GLiBRL enables seamless integration with both on-policy and off-policy RL algorithms. We further show that GLiBRL admits a closed-form relationship between the $\mathcal{L}_2$ distance of its task representations and empirical kernel-based correspondence between task samples, which is to our knowledge the first such structural result for online deep BRL. GLiBRL is compared against representative and recent Meta-RL methods, and improves state-of-the-art performance on both MuJoCo and MetaWorld benchmarks by up to 1.8$\times$.


Generalized Intention Modeling in Multi-Agent Reinforcement Learning

Mateusz Odrowaz-Sypniewski ⋅ Jasmine Bayrooti ⋅ Ajay Shankar ⋅ Amanda Prorok

Modeling an opponent’s intent is critical for effective decision-making in non-cooperative, competitive, and general-sum multi-agent reinforcement learning. Existing opponent modeling methods encode intent using an embedding derived from episode information chosen a priori, such as the opponent’s next action or a future environment state, and use this to guide the ego-agent’s behavior. These approaches assume that the chosen information is universally representative of intent; however, we show empirically that this is not the case as intentions are often task- and environment-dependent. To address this, we introduce a task-adaptive opponent modeling framework that learns a performance-driven mixture of multiple intent representations. We further introduce a new intention representation that maximizes mutual information with the ego-agent’s future returns, thereby capturing opponent information that is most directly relevant to performance. Our approach consistently matches or exceeds the performance of state-of-the-art baselines across diverse tasks and yields insights into when and why different opponent modeling strategies succeed.


Generating Physically Consistent Molecules with Energy-Based Models

Christoph Griesbacher ⋅ Lea Bogensperger ⋅ Andreas Habring ⋅ Thomas Pock

Molecules in equilibrium follow a Boltzmann distribution, making the underlying energy landscape a physically grounded modeling objective. However, such landscapes are difficult to learn from data and, once learned, hard to sample from. Diffusion and flow-matching models sidestep these difficulties by learning a time-conditional score or transport field between noise and data, losing the energy inductive bias in exchange for a more tractable training objective. We introduce EBMol, an energy-based model (EBM) that restores this inductive bias by learning an atom-additive scalar potential without explicit simulation during training. Our method employs a flow-inspired Restoring Field Matching objective to approximate the energy landscape. We adopt the Mirror-Langevin algorithm for sampling, enabling unified updates of atomic positions and types, and incorporate parallel tempering for inference-time compute scaling. EBMol is the first EBM for 3D molecular generation to achieve state-of-the-art performance on QM9 and GEOM-Drugs. Moreover, we show that the learned energy landscape serves as a principled quality metric for ranking and filtering configurations, and demonstrate controllable generation without retraining through shape-steered sampling via potential composition and zero-shot linker design.


GenRec: Knowing Where to Reconstruct and Where to Generate

Ata Çelen ⋅ Jaewoo Jung ⋅ Federico Tombari ⋅ Marc Pollefeys ⋅ Sunghwan Hong ⋅ Michael Niemeyer ⋅ Daniel Barath

Generative novel view synthesis from sparse input images is rarely all reconstruction or all generation: pixels visible in some source view have a unique correct value modulated only by view-dependent shading, while pixels in disocclusions or beyond the captured volume admit a distribution of plausible completions. Existing generative novel-view-synthesis methods conflate these regimes under a single uniform loss, blurring the line between geometric fidelity and creative hallucinations even when scene geometry is injected through warped point clouds or projected depth. We introduce GenRec, a multi-view flow matching model that builds the reconstruction--generation split directly into its architecture, supervision, and gradient flow. Guided by an observation mask derived from the source cameras and a monocular depth estimator, a flow matching backbone jointly denoises RGB and scene-coordinate maps across all target views, while a pixel-space refinement stage restores high-frequency detail on observed pixels; the same mask gates supervision so regression signals do not contaminate the generative prior. Across RealEstate10K, DL3DV-10K, and Mip-NeRF~360, in both single-view extrapolation and two-view interpolation, GenRec attains the best reconstruction fidelity in observed regions while also surpassing purely generative baselines on perceptual quality in unobserved ones, showing the effectiveness of our approach. The source code and trained models will be made public.

State space models (SSMs) provide an efficient route to global image restoration because they model long sequences with linear complexity. Recent restoration architectures further enlarge the effective context by reordering image tokens so that semantically similar pixels become neighbors in the 1D scan. This strategy improves texture aggregation, but it also exposes a mismatch between the geometry of images and the discretization used by SSMs: adjacent tokens in the reordered sequence can be spatially distant, while the recurrent update still treats them as consecutive samples of a continuous process. In addition, existing prompt-based modulation relies on a finite set of discrete prompts, which is poorly matched to the continuous variation of natural textures and offers no explicit compensation for the low-pass tendency of recurrent integration. We propose GeoMamba, a geometry-aware and continuously modulated SSM for image restoration. GeoMamba introduces geometry-adaptive discretization (GAD), which conditions the SSM time-step on the Euclidean distance between consecutive reordered tokens and attenuates history propagation across spatial jumps. It also introduces a continuous dynamic matrix (CDM), which regresses pixel-wise modulation parameters from feature and gradient cues to reduce prompt quantization and preserve high-frequency details. Experiments on lightweight and classical image super-resolution demonstrate consistent improvements over strong Transformer- and Mamba-based baselines, while achieving better or competitive performance in additional denoising and JPEG artifact reduction settings.


Geometric Gain Graph: Zero-Token Graph Construction for Multi-Hop RAG

Zeliang Li ⋅ Xiaofen Xing ⋅ Wenyu Tao ⋅ Kailing Guo ⋅ Xiangmin Xu

Existing Retrieval-Augmented Generation (RAG) systems struggle with multi-hop reasoning due to a fundamental inability to balance relevance and novelty. Traditional dense retrievers are constrained by surface-level matching, falling into a homogenized "Similarity Trap". Conversely, emerging Graph RAGs attempt to explore novel documents via heuristic entity extraction, but they introduce prohibitive Large Language Model (LLM) overhead and easily drift into off-topic "Novelty Traps". To fundamentally resolve this contradiction, we propose Geometric Gain Graph (G$^3$RAG), the novel zero-token graph construction paradigm for RAG. G$^3$RAG completely discards the expensive heuristic extraction by LLMs, aiming instead to directly quantify the information gain between nodes through the native document feature space. To this end, we designed a concise edge weight criterion, $\cos\theta \cdot \sin\theta$, which naturally maps the trade-off between relevance and novelty into a geometric measurement of directional consistency and orthogonality in the vector space, thereby driving the token-free topological construction of the graph. Furthermore, we introduce a topological penalty to suppress the excessive connectivity of high-frequency hubs, forcing the transient diffusion process toward long-tail peripheral nodes that carry critical indirect evidence. Extensive experiments demonstrate that, while entirely eliminating graph construction costs, G$^3$RAG significantly outperforms state-of-the-art Graph RAG baselines on complex multi-hop QA benchmarks.

Mechanistic interpretability aims to explain a model’s behavior by identifying causally responsible internal structures. Dictionary-based explainers such as sparse autoencoders and transcoders are a primary tool, but their faithfulness under out-of-distribution (OOD) shift has received little systematic attention. We show that distribution shift rotates the subspace that the model actively uses, misaligning the explainer’s dictionary trained on in-distribution (ID) activations. We formalize this misalignment as the faithfulness gap, a geometric distance between the ID dictionary and the OOD-active subspace, and show that it controls OOD faithfulness degradation. To reduce this gap, we propose the Geometry-Adaptive Explainer (GAE), which realigns the explainer's dictionary with the OOD-active subspace while preserving the original feature structure. This requires only unlabeled OOD activations and no gradient updates. We prove that GAE improves over the unadapted ID explainer, with excess loss bounded quadratically by the second-moment shift. Empirically, GAE even matches or surpasses all training-based baselines in causal faithfulness across multiple models and OOD settings.


Geometry is an Operator: Lie-Algebraic Space Routing for View-Robust 3D MLLMs

Jingjun Yi ⋅ Chen Hu ⋅ Qi Bi ⋅ Hao Zheng ⋅ Huimin Huang ⋅ Yongkang Li ⋅ Wei Ji ⋅ Yawen Huang ⋅ Xian Wu ⋅ Yefeng Zheng

Multimodal Large Language Models (MLLMs) have shown increasing potential for 3D scene understanding and spatial reasoning from videos. However, even with strong visual geometry priors, the hidden semantics of existing 3D-enhanced MLLMs can drift across views, leading to degraded grounding, captioning, and spatial reasoning performance. The underlying reason can be their geometry injection as additional tokens or additive features, which enriches visual content but fails to define how view-dependent semantic states should transform. To address this issue, we propose Lie Routing Vision-Language Model (LieVLM), a geometry-conditioned Lie routing framework that treats geometry as an operator over hidden semantics. LieVLM first elicits local Lie states from 3D geometry tokens by predicting latent rotation states and routing coefficients. It then converts these rotation states into triplet-wise Lie group actions to route projected 2D visual semantics. Finally, a geometry-aware routed fusion module combines the original semantic path, the direct 3D context path, and the Lie-routed transformation path for language generation. This design preserves pretrained visual-language semantics while allowing geometry to actively correct view-induced semantic drift. Extensive experiments on 3D scene understanding and spatial reasoning benchmarks demonstrate that LieVLM improves view robustness, especially in severe-pose-change regimes, and consistently outperforms additive geometry fusion methods. Our results suggest a shift in 3D MLLM design, i.e., geometry should not merely be injected as context, but should operate on semantic representations.


Geometry-Regularized Collapse Resistance via Consensus Enhancement for Federated Learning

Jingjing Zhu ⋅ Zhuang Qi ⋅ Lei Meng ⋅ Han Yu ⋅ Jie Zhang ⋅ Xiangxu Meng

Knowledge distillation provides shared semantic supervision for local training in federated learning, reducing the tendency of client models to overfit their own data distributions. Existing methods construct a global knowledge repository from client-uploaded local information to establish consensus across clients. However, global knowledge that is fundamentally derived from client data inevitably contains local biases under heterogeneous settings. To address this issue, this study proposes a geometry-regularized cross-silo consensus enhancement method, termed RISE, which improves collapse resistance in federated learning by preserving pre-trained geometric consensus priors during local adaptation. Specifically, RISE introduces two complementary geometry regularizers. The Biased Low-dimensional Subspace Correction module preserves the spectral energy profile of pre-trained representations by constraining the energy distribution of adapted features within the principal representation subspace. Instead of enforcing direction-wise alignment, it discourages excessive energy concentration along a few dominant components, thereby mitigating biased low-dimensional feature collapse. The Manifold Relation Graph Alignment module regularizes the relational manifold geometry by constructing a transferable inter-class topology from pre-trained global class prototypes and aligning client-side class structures with this global relation, thereby mitigating class representation drift and stabilizing global aggregation. Extensive experiments on eight benchmark datasets across two tasks demonstrate that RISE improves cross-client consensus consistency and enhances the generalization ability of the global model, outperforming eight state-of-the-art methods.


GeoPano: Towards Geometrically Accurate Panoramic 3D Reconstruction from a Single Panorama

Jing OU ⋅ Zidong Cao ⋅ Liaoyuan Fan ⋅ Zhuoxiao Li ⋅ Shuai Zhang ⋅ Yinrui Ren ⋅ Dongli Wu ⋅ Haoang Li ⋅ Hui Xiong ⋅ Wufan Zhao

Panoramic 3D reconstruction enables the holistic recovery of complete 360° scene geometry, offering significant benefits for robotics and AR/VR applications; however, direct estimation is persistently challenged by inherent spherical distortions. While decomposing panoramas into perspective views effectively circumvents these distortions and leverages established geometric priors, this strategy often fails to preserve global spatial consistency across the decomposed views. To address this fundamental limitation, we introduce an enhanced feed-forward 3D reconstruction framework tailored for the view decomposition paradigm, driven by two key innovations. First, we propose a geometry-aware attention aggregator, featuring a streamlined attention architecture equipped with a novel Geo-Attention mechanism, which explicitly exploits the intrinsic global information to guide multi-view interactions, thereby suppressing geometrically inconsistent matches. Furthermore, to jointly enforce fine-grained local details and global structural coherence, we formulate a hybrid supervision strategy. Specifically, we complement the standard local point cloud loss in perspective space with our newly proposed spatially uniform global point cloud loss in panoramic space, explicitly enforcing holistic constraints within a unified coordinate system. Experiments demonstrate that our method consistently outperforms prior single-view panoramic approaches in reconstruction accuracy, geometric consistency, and visual fidelity.


GeoWorld: A Geometry-First World Model for Reconstruction and Imagination

Baicheng Li ⋅ Jialin Liu ⋅ Dong Wu ⋅ Weijian Xie ⋅ Zhichao Ye ⋅ Donghui Shen ⋅ Nan Wang ⋅ Guofeng Zhang ⋅ Haomin Liu ⋅ Hongbin Zha

Large visual world models have become increasingly capable at synthesizing plausible views and camera trajectories, largely because they inherit strong priors from image and video diffusion. Yet in robotics, spatial computing, AR anchoring, and digital twins, a world is useful only if it can be measured and acted upon: appearance plausibility cannot substitute for geometrically coherent structure. This motivates a geometry-native alternative, where the generative state is the feature space of a geometry foundation model rather than RGB pixels or an appearance latent. Existing geometry-space world models validate this direction, but they still occupy a narrow operating range. When target views lie within observed support, full multi-view interaction limits trajectory length; when observations become sparse, geometry-only training lacks the broad visual priors needed to infer structure beyond observed support. We present GeoWorld, a geometry-latent world model that expands this range without returning generation to RGB space. GeoWorld builds a compact geometry latent, conditions it with camera ray geometry, and uses camera-aware tiered attention with relative pose bias to make long-view denoising practical. For sparse-observation extrapolation, a Bridge module translates hidden states from a frozen camera-conditioned video diffusion model into geometry-latent conditioning, borrowing video priors without making the video model the scene generator. Experiments show that GeoWorld scales geometry-space diffusion to long camera trajectories while improving geometric consistency in prior-assisted extrapolation.


GGQR: Gaussian-Grounded Query Refinement for Feed-Forward 4D Gaussian Splatting

Jingqiao Xiu ⋅ Yicong Li ⋅ Leigang Qu ⋅ Angela Yao

Feed-forward 4D Gaussian Splatting (4DGS) offers an efficient paradigm for reconstructing dynamic scenes from monocular videos, bypassing the need for lengthy per-scene optimization. However, existing methods suffer from three fundamental bottlenecks: redundant frame-wise grid-aligned Gaussian representations, suboptimal single-step 2D-to-4D regression architectures, and restrictive dependencies on 3D/4D signals during training or inference. In response to these limitations, we introduce a novel Gaussian-Grounded Query Refinement (GGQR) framework. To reduce representation redundancy, GGQR learns a compact set of Gaussian queries equipped with a newly designed plateau-shaped temporal kernel, enabling single primitives to persistently model motion-consistent regions across extended frames. To alleviate regression ambiguity, these queries are iteratively refined through a Gaussian-grounded attention mechanism operating in 3D space. Specifically, GGQR leverages foundation models to lift 2D observations into a shared 3D space, explicitly grounds the 3D attention in intermediate Gaussian states, and residually updates those Gaussian states. Trained entirely self-supervised on casual videos, GGQR achieves a new state-of-the-art on dynamic reconstruction benchmarks among feed-forward 4DGS methods, delivering superior performance in novel view-time rendering alongside downstream motion modeling tasks.


Git Context Controller: Manage the Context of Agents by Agentic Git

Junde Wu ⋅ Minhao Hu ⋅ Jiayuan Zhu ⋅ Shengda Zhu ⋅ Jiazhen Pan ⋅ Fenglin Liu ⋅ Yuyuan Liu ⋅ Min Xu ⋅ Yueming Jin

Large language model (LLM) agents have demonstrated strong capabilities in long-horizon tasks by interleaving reasoning with tool use. However, as these agents scale to complex workflows such as software engineering and open-ended research, context management becomes a fundamental bottleneck: interaction histories grow unbounded, become costly to maintain, and are difficult to reuse across sessions and agents. We introduce \textbf{Git-Context-Controller (GCC)}, a structured context management framework inspired by software version control systems. GCC elevates agent context from a transient token stream to a persistent, navigable memory workspace with explicit operations---\texttt{COMMIT}, \texttt{BRANCH}, \texttt{MERGE}, and \texttt{CONTEXT}, that enable milestone-based checkpointing, isolated exploration of alternative reasoning paths, and hierarchical retrieval of historical context. By organizing agent memory as a versioned file system, GCC allows agents to manage long-term goals, recover and transfer reasoning across sessions, and coordinate multi-trajectory problem solving in a principled manner. Empirically, agents equipped with GCC achieve state-of-the-art performance on both SWE-Bench and BrowseComp benchmarks. On SWE-Bench Verified, GCC improves task resolution by over 13\% relative to strong long-context baselines and outperforms 26 existing open and commercial systems, reaching over 80\% success rate. The project will be open-sourced for the research community.

Predicting tandem mass spectra (MS/MS) from molecular structures represents a central task in analytical chemistry with direct relevance to clinical metabolomics, systems biology, and adjacent disciplines. In this work, we revisit the problem through the lens of object detection on molecular graphs. Molecular fragmentation, a central step in MS/MS prediction, can be approximated as detecting a set of subgraphs (i.e., fragments) and their associated spectral contributions. Existing fragment-based models follow a two-stage paradigm—first generating candidate fragments and then scoring them—analogous to two-stage R-CNNs in computer vision. Towards higher accuracy and faster inference, we introduce GLACIER, a single-stage transformer-based fragment detection neural network for molecular graphs. This unified formulation eliminates the need for candidate enumeration, enabling scalable and globally consistent modeling of molecular fragmentation. GLACIER is faster and more accurate than existing state-of-the-art methods by a significant margin, achieving 62.4% and 62.1% Top-1 retrieval accuracy with and without contrastive finetuning on the MassSpecGym dataset (from 55.3%), and 55.2% and 38.2%, respectively, on the NIST'20 dataset (from 33.5%). Furthermore, GLACIER provides nearly 3-fold inference speedup over existing two-stage models.


Global convergence of adjoint-optimized neural PDEs

Konstantin Riedl ⋅ Justin Sirignano ⋅ Konstantinos Spiliopoulos

Many engineering and scientific fields have recently become interested in modeling terms in partial differential equations (PDEs) with neural networks, which requires solving the inverse problem of learning neural network terms from observed data in order to approximate missing or unresolved physics in the PDE model. The resulting neural-network PDE model, being a function of the neural network parameters, can be calibrated to the available ground truth data by optimizing over the PDE using gradient descent, where the gradient is evaluated in a computationally efficient manner by solving an adjoint PDE. These neural PDE models have emerged as an important research area in scientific machine learning. In this paper, we study the convergence of the adjoint gradient descent optimization method for training neural PDE models in the limit where both the number of hidden units and the training time tend to infinity. Specifically, for a general class of nonlinear parabolic PDEs with a neural network embedded in the source term, we prove convergence of the trained neural-network PDE solution to the target data (i.e., a global minimizer). The global convergence proof poses a unique mathematical challenge that is not encountered in finite-dimensional neural network convergence analyses due to (i) the neural network training dynamics involving a non-local neural network kernel operator in the infinite-width hidden layer limit where the kernel lacks a spectral gap for its eigenvalues and (ii) the nonlinearity of the limit PDE system, which leads to a non-convex optimization problem in the neural network function even in the infinite-width hidden layer limit (unlike in typical neural network training cases where the optimization problem becomes convex in the large neuron limit). The theoretical results are illustrated and empirically validated by numerical studies.

Memory consumption of the key-value (KV) cache remains a central bottleneck for long-context LLM inference. We introduce PRKV, a training-free KV eviction method that estimates multi-hop global token importance by applying Personalized PageRank over attention-induced transition graphs. This enables importance propagation beyond recency or instantaneous attention, without modifying the model. We benchmark PRKV on long-context understanding, sparse retrieval, and chain-of-thought reasoning tasks using multiple LLM families. Under moderate to aggressive compression ratios, including up to 90\%, PRKV achieves competitive task performance against established eviction baselines while reducing peak memory usage with low end-to-end overhead. These results demonstrate that multi-hop global influence estimation yields a favorable accuracy--memory trade-off for scalable long-context inference.


GPU Hierarchy Meets Structured Matrices: Fast Algorithms for State-Space Models

Berlin Chen ⋅ Caitlin Wang ⋅ Aakash Sunil Lahoti ⋅ Kevin Li ⋅ Jay Shah ⋅ Jack Carlisle ⋅ Timmy Liu ⋅ Mengyu Guo ⋅ Zico Kolter ⋅ Albert Gu ⋅ Tri Dao

Linear-time sequence models such as State Space Models (SSMs) offer an efficient alternative to self-attention and have demonstrated strong language modeling performance at scale. In practice, however, their asymptotic advantage is often lost at common training and prefill lengths---optimized FlashAttention kernels remain faster than existing SSM implementations up to 8k tokens. In this work, we substantially narrow this performance gap with novel algorithmic improvements: First, we introduce a split-sequence algorithm that decouples the compute tile from the sequence-parallel split length, preserving efficient matrix multiplication tiles while substantially reducing boundary-state traffic and latency. Second, we resolve a longstanding *speed--stability tradeoff* in computing the 1-semiseparable (SS) decay mask, the main non-matmul bottleneck shared by all chunk-wise linear-time kernels (Mamba-2, Mamba-3, GDN, KDA). Existing kernels use the fast *diff segsum* (prefix subtraction), which suffers monotonicity violations and catastrophic cancellation---failures observed in real pretraining (e.g., Nemotron-H); the stable alternative (*direct segsum*) materializes the dense mask and is $\sim30\\%$ slower. Our *hierarchy-aware segsum* decomposes the SS mask into 1-SS diagonal blocks and rank-one off-diagonal blocks, computed via warp-level scans and cross-warp aggregation directly in registers, achieving the stability of direct segsum at the speed of diff segsum. We implement these algorithms using CuTe-DSL with warp specialization, asynchronous TMA, WGMMA, and layout-aware accumulator fragments. Across a large range of sequence length, our H100 kernels reach 91\% of measured memcpy bandwidth, improve prefill time by $4.1\times$ over the official Triton Mamba-2 kernel, and **outperform the highly-optimized FlashAttention-4 starting at 1-2k sequence length**---to our knowledge, the first time a linear-time sequence layer is faster than state-of-the-art attention at the 2-4k sequence lengths used in common pretraining and prefill workloads, where prior linear-time kernels only reached parity at $\geq 8$k. Our hierarchy-aware segsum is $1.35$-$1.40\times$ faster than a naive stable direct-segsum implementation.


Gradient-Free Editing for Attribute Invariance in Graph Neural Networks

Aastha A K Verma ⋅ Vishakha Agarwal ⋅ Sahil Manchanda ⋅ Sayan Ranu

Deployed Graph Neural Networks are increasingly subject to post-deployment interventions: a sensitive attribute is flagged for removal, or a spurious correlation is discovered after the model is already in production. In such scenarios, retraining is often infeasible due to a variety of reasons including scale of modern graphs, inaccessible training data, or strict privacy mandates. The challenge is further compounded in graph-structured data, where attribute dependencies propagate through message-passing, spreading unwanted correlations across the entire network. We propose a framework for surgical, post-hoc attribute invariance in GNNs via gradient-free model editing. Rather than retraining or fine-tuning, our method formulates the edit as a subspace-constrained, closed-form optimization: we identify structurally influential nodes that exhibit high attribute sensitivity, isolate the low-rank manifold in activation space where sensitive variation is concentrated, and solve directly for a weight update that decouples the model's decision logic from the target attribute. The update requires no access to original training labels, introduces no additional parameters, and is computed in a single pass. Empirical results demonstrate that our method enforces attribute invariance while preserving predictive utility, while offering 3 orders of magnitude speed-up over existing GNN editing approaches.

Learning from Label Proportion (LLP) is a weakly supervised learning paradigm in which only aggregated label proportions over collections of instances (i.e., bags) are provided, rather than individual labels. This allows classification while preserving privacy or reducing annotation costs. Existing LLP methods, however, have been largely restricted to i.i.d. tabular or image data. To the best of our knowledge, no solution currently addresses graphs, where instances are inherently interdependent through network structure. In this paper, we generalize LLP to the graph domain and study the problem of node classification with label proportions, where only distributional supervision is available for node bags, and the goal is to infer labels for all nodes in the graph. We argue that the lack of node-level supervision is the main challenge for LLP on graphs, and that existing methods based on i.i.d. assumptions fail to exploit topological correlations. To overcome this, we propose GLLP~(Graph Learning from Label Proportions), a framework that leverages Optimal Transport (OT) with a homophily-aware cost to generate soft pseudo-labels for individual nodes. These pseudo-labels provide stronger supervision signals for training Graph Neural Networks. We further establish theoretical guarantees showing the alignment of our cost function with the node classification objective. Extensive experiments on six homophilic graph benchmarks and the Covid real-world dataset demonstrate that GLLP consistently outperforms existing LLP baselines and variants.


Graph Mamba Operator: A Latent Simulator for Interacting Particle Systems

Karn Tiwari ⋅ Niladri Dutta ⋅ Prathosh AP ⋅ N M Anoop Krishnan

Modeling interacting dynamical systems requires capturing spatial interactions alongside long-range temporal dependencies. Graph neural networks (GNNs) provide a natural representation but typically rely on autoregressive rollouts and treat spatial and temporal dynamics separately, leading to error accumulation over long horizons. Existing approaches also focus on local interactions and short temporal contexts, limiting their ability to capture multi-hop dependencies and global structure. We introduce the Graph Mamba Operator (GraMO), a latent-space simulator that integrates state-space models with graph-based interaction learning. In contrast to prior work that sequences nodes or applies spatial and temporal updates in separate stages, GraMO couples graph-based interactions and temporal state updates within a single recurrence. The update is linear in the latent state, with input-dependent coefficients that adapt across regimes. We evaluate GraMO on N-body systems, motion capture, and robotics datasets, achieving the lowest error across benchmarks and the largest gains in long-horizon prediction.


Graph-Regularized Sparse Autoencoders for LLM Safety Steering

Jehyeok Yeon ⋅ Federico Cinus ⋅ Yifan Wu ⋅ Luca Luceri

Sparse autoencoders (SAEs) are increasingly used to extract activation directions for inference-time steering, but their standard sparsity objective treats latent features as independent. This prior can be poorly matched to high-level safety behaviors, where refusal and harmful compliance appear to depend on distributed structure in activation space. We introduce Graph-Regularized Sparse Autoencoders (GSAE), a dictionary-learning method that learns safety-steering directions by smoothing SAE decoder vectors over a neuron co-activation graph and applying the resulting direction bank through a two-gate runtime controller. Empirically, GSAE improves selective refusal across JailbreakBench, HarmBench, and XSTest, increasing harmful-request refusal while keeping benign-prompt refusals low. On Llama-3-8B, adding graph regularization to an otherwise identical SAE bank-and-gating pipeline improves $\Delta_s$ by $20.1$ points on JailbreakBench and $16.8$ points on HarmBench. GSAE outperforms activation-steering baselines and black-box guardrails, preserves benign-task performance, generalizes across Llama-3, Mistral, Qwen 2.5, and Phi-4, and remains strong under black-box and gray-box jailbreak attacks.


Grassmannian Geodesic Steering: Rank-Preserving Subspace Control for Inference-Time Alignment of Language Models

Longyi Liu ⋅ Zhitao Wang ⋅ Jianchao Yu ⋅ Mingrui Cai ⋅ Yuxing Han ⋅ Gene Wen

Inference-time activation steering provides a lightweight mechanism for aligning language model behavior without retraining, yet existing additive interventions often distort activation geometry, alter representation norms, and induce effective-rank collapse, thereby degrading open-ended generation. Recent norm-preserving rotation-based methods mitigate this pathology, but their reliance on one- or two-dimensional concept axes is insufficient for behaviors whose representations are intrinsically multi-faceted and low-rank. We propose Grassmannian Geodesic Steering (GGS), a rank-preserving framework that lifts activation control from vector directions to $k$-dimensional subspaces on the Grassmann manifold. We develop a theoretical account showing that additive steering provably collapses effective rank under strong intervention, that low-dimensional rotations cannot capture distributed behavioral concepts, and that orthogonal transformations along Grassmannian geodesics exactly preserve the singular-value spectrum. Guided by these results, GGS estimates concept subspaces from contrastive difference covariance, steers hidden states via a closed-form Grassmannian geodesic realized as an orthogonal operator, and adaptively modulates intervention strength using a matrix Bingham gate over subspace-valued activations. The resulting method introduces less than $1.5%$ wall-clock overhead per decoding step. Across LLaMA-3.1-8B and Qwen-2.5-7B on six reasoning benchmarks and an AdvBench safety evaluation, GGS consistently improves alignment accuracy while preserving generation quality and refusal behavior, reducing effective-rank degradation by two orders of magnitude relative to additive baselines, and surpasses the current state of the art among training-free inference-time steering methods.


Ground False: Uncovering Errors in Formal Mathematics Benchmarks

Marcus Min ⋅ One An ⋅ Xujie Si ⋅ Osbert Bastani

The ground truths in formal mathematics benchmarks are taken on faith. For theorem proving, released formalizations come without formal proofs that certify their correctness. For autoformalization, faithfulness to the informal source is judged only by the benchmark authors, yet such judgments are inherently subjective and lack collective consensus. We show these assumptions fail at scale: our audit of 367 formal statements in ProofNet finds that 204 (56\%) are unfaithful, and over half of those are mathematically false and cannot be proved. We dissect these errors along three axes (provability, logical strength, root cause) and show that they bias evaluation in predictable, asymmetric directions: false ground truths cap theorem-proving signal and compress model gaps, while equivalence-based autoformalization metrics penalize faithful models more than unfaithful ones. We proceed to fix these ``Ground Falses'' and release the corrected benchmark \textbf{ProofNet-Verified}, produced by a semi-automated pipeline combining Lean provability checks, an LLM faithfulness judge, and human review. Applying the same pipeline to six other benchmarks reveals that unfaithful-GT rates span more than an order of magnitude (4.8\% to 60.0\%), confirming that the failure modes generalize while absolute quality is dominated by curatorial process. Our data and code are available at \url{https://github.com/anonymousauthor567/Ground_False}


GUARD: Scalable Gradient-based Unlearning with Adversarial Robustness Defense

Xuechao Lan ⋅ Wenmin Li ⋅ Sujuan Qin ⋅ Zhengping Jin ⋅ Fei Gao

Machine Unlearning (MU) is an emerging task to remove specified training data from a trained model while preserving its utility, with Gradient-based Unlearning (GU) standing out as a prevalent category due to high scalability. However, standard MU which solely optimizes for clean accuracy, inadvertently damages adversarial robustness on the retain data. Although a few studies extend MU to adversarially trained models, they rely on strong assumptions (e.g., well-conditioned Hessians, smoothness) and costly Hessian-based updates, limiting scalability and efficiency. In this paper, we propose **GUARD**, a scalable and efficient framework that achieves unlearning while preserving retain-data robustness. Having identified the two causes of robustness degradation, we accordingly design *Robustness-Aware Gradient Projection* and *Importance-Guided Parameter Selection*. Both modules rely on a compact adversarially vulnerable dataset that we construct with or without retain data, to enable scalability under varying data constraints. Validated by theoretical analysis and experiments across 6 settings, GUARD outperforms 6 baselines, boosting robust accuracy by up to 6.3$\times$, closely matching the Retrain oracle in both unlearning and robustness.


Guidance For Prior Change via Density Ratio Estimation

Yichen Zang ⋅ Song Liu ⋅ Jiun-Yi Lin

Simulation-Based Inference (SBI) serves as a vital framework for parameter inference in scientific fields where simulators involve intractable likelihoods, yet while amortized generative models offer rapid posterior estimation, they are often restricted by the specific priors used during training, thereby limiting their flexibility as prior knowledge evolves. To address this prior dependency, PriorGuide was introduced as an inference-time guidance method, but due to its intractable formulation, it relies on Gaussian approximations of the reverse transition kernel and Gaussian mixture model fitting for the prior ratio, both of which introduce systematic bias. Motivated by these limitations, we propose an unbiased test-time guidance framework that leverages Density Ratio Estimation (DRE) to learn a score guidance term, effectively decoupling the inference process from the prior training. Moreover, our framework remains agnostic to the specific density ratio estimators, making it a general and flexible framework for handling prior changes. Experimental results across multiple tasks demonstrate that our method outperforms PriorGuide, achieving superior performance on metrics such as C2ST and MMD in most tasks while maintaining strong robustness even under limited overlap between the training and target priors. Furthermore, we apply our method to Bayesian updating for parameter inference from planetary light-curve data, where it also demonstrates strong effectiveness and robustness.


Gym-Anything: Turn Any Software into an Agent Environment

Pranjal Aggarwal ⋅ Graham Neubig ⋅ Sean Welleck

Computer-use agents hold the promise of assisting in a wide range of digital economic activities. However, current research has largely focused on short-horizon tasks over a limited set of software. A key reason is that creating environments for complex software requires significant time and human effort, and therefore does not scale. To address this, we introduce Gym-Anything, a framework for converting any software into an interactive computer-use environment. We frame environment creation itself as a multi-agent task: a coding agent writes setup scripts, downloads real-world data, and configures the software, while producing (visual) evidence of correct setup. An independent audit agent then verifies evidence for environment setup against a quality checklist. Using a taxonomy of economically valuable occupations grounded in U.S.\ GDP data, we apply this pipeline to 200 software applications with broad occupational coverage. The result is CUA-World, a collection of over 10K+ long-horizon tasks spanning domains from medical science and astronomy to engineering and enterprise systems, each configured with realistic data along with train and test splits. Distilling successful trajectories from the training split yields strong gains at multiple student scales (2B, 9B, 35B). CUA-World also includes CUA-World-Long, a challenging long-horizon benchmark with tasks often requiring over 500 steps. We also apply the same auditing principle at test time: a separate VLM reviews completed trajectories and provides feedback on what remains, improving Gemini-3-Flash on CUA-World-Long from 11.5\% to 14.0\%. We will release all code, infrastructure, and benchmark data.


Halt Fast! Early Stopping for Certified Robustness

Andrew Cullen ⋅ Paul Montague ⋅ Benjamin Rubinstein

Randomized Smoothing provides rigorous robustness guarantees for neural networks without architectural constraints, yet its adoption is limited by extreme computational costs. Standard RS requires tens of thousands of model evaluations per input and forces practitioners to commit to fixed sample sizes a priori. In this work, we present a novel meta-learning framework for anytime-valid certified robustness that adaptively deploys computational resources. By using a lightweight meta-learner to predict image-specific priors for a sequential E-process, we achieve a 20-fold reduction in sample complexity compared to traditional methods while maintaining rigorous statistical guarantees. Beyond raw efficiency, we demonstrate how anytime-validity enables adaptively allocating compute based upon application-specific risk thresholds, a form of resource triage impossible under classic certification frameworks. That this is achievable while also providing similar certification performance demonstrates that our approach provides a pathway for real-time, safety-critical certification deployments.


Harbor Adapters and Harbor-Mix: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic Evaluation

Lin Shi ⋅ Haowei Lin ⋅ Zixuan Zhu ⋅ Xiaoyue Zhou ⋅ Xiang Li ⋅ Xiangning Lin ⋅ Yaxuan Deng ⋅ Han Xu ⋅ Yuangang Li ⋅ Shanda Li ⋅ Zizhao Chen ⋅ Hanwen Xing ⋅ Harsh Raj ⋅ Bo Chen ⋅ Quan Shi ⋅ Steven Dillmann ⋅ Yipeng Gao ⋅ Puneesh Khanna ⋅ Ruofan Lu ⋅ Chao Zhou ⋅ Michael Yang ⋅ Robert Zhang ⋅ Siyuan Chai ⋅ Jiayu Chang ⋅ Yizhao Chen ⋅ Xiaokun Chen ⋅ Yiwei Dai ⋅ Wenting Yang ⋅ Hange Liu ⋅ Minghao Liu ⋅ Zihan Wang ⋅ Adnan E Assadi ⋅ Benedikt Stroebl ⋅ Estefany Kelly Buchanan ⋅ Han Meng ⋅ Junwei He ⋅ Longxuan YU ⋅ Radin Shayanfar ⋅ Yukyung Lee ⋅ Zhikang Dong ⋅ Allen G Hart ⋅ Anjiang Wei ⋅ Anurag Kashyap ⋅ Arpandeep Khatua ⋅ Audrey J Zheng ⋅ Chengrui Ma ⋅ David Heineman ⋅ dubing ⋅ Hai-Anh Trinh ⋅ Haishuo Fang ⋅ Hefan Zhang ⋅ Hui Shen ⋅ Issa Sugiura ⋅ Jiankai Sun ⋅ Jiechao Gao ⋅ Junhong Lin ⋅ Junnan Li ⋅ Kai Yang ⋅ Lei Hsiung ⋅ Maoyu Wang ⋅ Mengze Tang ⋅ Nabil Omi ⋅ Negin Raoof ⋅ Nicholas Edwards ⋅ Ruohao Guo ⋅ Orfeas Menis Mastromichalakis ⋅ Ryan Ji ⋅ Przemysław Hejman ⋅ Qi Qi ⋅ Qunshu Lin ⋅ Richard Zhuang ⋅ Rui Yang ⋅ Ruichen Zheng ⋅ Ryan Marten ⋅ Shaghayegh Fazliani ⋅ Shizheng Hou ⋅ Sicong Jiang ⋅ Sijie Li ⋅ Song Bian ⋅ Terry Yue Zhuo ⋅ Tianqing Wu ⋅ Tom Tang ⋅ Wanjia Zhao ⋅ Weihao Xuan ⋅ Wenhua Liang ⋅ Xian Liu ⋅ Xin Lan ⋅ Ruby Zhang ⋅ Xuandong Zhao ⋅ Yanchuan Tang ⋅ Yifan Jiang ⋅ Yijiang Li ⋅ Yitong Guan ⋅ Yizhi Li ⋅ YONGHUI LIU ⋅ Yuheng Tang ⋅ Yujun Mao ⋅ Yunfei Zhao ⋅ Yuxin Wang ⋅ Yuxuan Tang ⋅ Zhenheng Tang ⋅ Zhifei Li ⋅ Ziruo Wang ⋅ Ziyu She ⋅ Kaiyuan Liu ⋅ Iheb Chaabane ⋅ Yuxin Tang ⋅ Xiangyi Li ⋅ Soroush Vosoughi ⋅ Sanmi Koyejo ⋅ Di He ⋅ Etash Guha ⋅ Benjamin Feuer ⋅ Andy Konwinski ⋅ Boxuan Li ⋅ Leon L Chen ⋅ Alex Dimakis ⋅ Nicholas Carlini ⋅ Mike Merrill ⋅ Ludwig Schmidt ⋅ Alex Shaw

Evaluating agents on the growing number of agentic benchmarks is challenging because they often require complex environments and agent integrations. We introduce Harbor Adapters, a unified evaluation infrastructure for agentic benchmarks. Our work makes three contributions. First, we develop benchmark adapters that port more than 80 benchmarks to evaluate arbitrary agents, and validate them through rigorous code review and parity experiments. Second, we conduct a large-scale evaluation of frontier models across 54 benchmarks, enabling a broader analysis of agent capabilities and failure modes than was previously possible. Third, we introduce Harbor-Mix, a curated set of 100 difficult, diverse, and high-quality tasks refined from the adapted benchmarks. Harbor-Mix preserves the challenge and breadth of large-scale agentic evaluations while being affordable to run; the strongest model resolves only 15.6\% of its tasks. We release the adapters, evaluation results, in-depth analysis, and Harbor-Mix as open-source artifacts to support more reliable and comprehensive evaluation of language-model agents.


Harnessing Image Diffusion Prior for Photo-Realistic Video Restoration

Fanghua Yu ⋅ Hongyu An ⋅ Jinfan Hu ⋅ Xinqi Lin ⋅ Zhiyuan You ⋅ Jian Wang ⋅ Hongyang Li ⋅ Chao Dong ⋅ Jinjin Gu

Recent image generation models produce photo-realistic results with rich details, but transferring their powerful priors to video restoration remains difficult due to temporal inconsistency from stochastic detail synthesis. Existing video restoration methods improve temporal coherence but often sacrifice visual fidelity, leaving a gap between image-level quality and video-level stability. We propose a generative video restoration method that bridges this gap by enabling temporally consistent detail synthesis from image generation priors. Our method combines noise energy rebalancing to suppress structure-disturbing low-frequency stochasticity, motion-aligned noise warping to propagate details along object motion, a self-supervised denoising encoder to mitigate encoder-induced temporal uncertainty, and decoupled temporal attention to model cross-frame dependencies. Together, these components preserve fidelity, enhance fine-grained realism, and maintain temporal consistency. Extensive experiments show that our method substantially outperforms existing video restoration approaches in visual quality, detail richness, and temporal stability.


HDL-RepoBench: Multi-Paradigm Repository-Level Code Completion for Hardware Design Languages

Qingyun Zou ⋅ Jiahao Cui ⋅ Nuo Chen ⋅ Bingsheng He ⋅ Weng-Fai Wong

Large language models (LLMs) have achieved strong performance on code completion in general-purpose programming languages, but existing repository-level benchmarks focus almost exclusively on software code and largely overlook hardware description languages. We present HDL-RepoBench, consisting of HDL-RepoBench-Train and HDL-RepoBench-Eval, the first benchmark for multilingual hardware code completion at the repository level with functional evaluation. HDL-RepoBench covers three major hardware design coding styles and annotates each completion target with code-structure-level and hardware-oriented semantic labels derived from concrete syntax tree analysis. Post-training billion-scale code LLMs on HDL-RepoBench-Train enables smaller open-weight models to outperform much larger general-purpose models under matched context windows. Beyond aggregate accuracy, our analysis surfaces four findings that distinguish hardware from software code completion: (i) performance varies sharply across HDLs, with HLS consistently easiest and VHDL hardest across functional pass rate and EM/ES; (ii) hardware accuracy is non-monotonic in code-structure depth, in contrast with the monotonic depth–accuracy curve known for software; (iii) accuracy plateaus once the input context exceeds 2,048 tokens, while software completion continues to benefit from longer context; and (iv) certain repository-level retrievers tuned on software underperform a no-retrieval baseline on HDL repositories because their structural and dependency assumptions are violated by hardware code.

LLM evaluations drive which models get deployed, what safety standards get adopted, which research conclusions get published, and how projections of AI's labor-market impact get made. Yet standard confidence intervals ignore variability from judge model choice, model temperature, and prompt phrasing, producing under-coverage that worsens with more data. The omitted variance can shift results enough to reverse conclusions \citep{baumann2025llmhacking, huang2026dropping}; pipelines that fail to average over it leave the surface that ``benchmark hacking'' exploits \citep{singh2025leaderboard}. This paper decomposes LLM pipeline uncertainty into its sources, distinguishes variance that shrinks with more data from sensitivity to researcher design choices, and uses design-study projections to reduce total evaluation error (TEE). Across the demonstrations, naive standard errors are 40 - 60\% smaller than the TEE-corrected SE. Using Chatbot Arena data, we show naive 95\% CI coverage drops as $n$ grows while TEE-corrected coverage holds at 95\%, and TEE-guided pipelines restrict the benchmark gaming surface from 56 to 32 Elo ($K=27$), below the human-leaderboard baseline. We show further that a small pilot recovers honest CIs and projects which design changes most improve precision. Acting on those projections halves MMLU estimation error against the answer key at equivalent cost, and raises per-match agreement with human votes by 7.9 percentage points on Chatbot Arena.


Hide to Guide: Learning via Semantic Masking

Ruitao Liu ⋅ Qinghao Hu ⋅ Alex Hu ⋅ Yecheng Wu ⋅ Shang Yang ⋅ Luke J Huang ⋅ Zhuoyang Zhang ⋅ Han Cai ⋅ Song Han

Reinforcement learning with verifiable rewards (RLVR) has become a powerful paradigm for improving language models on reasoning-intensive tasks, but its effectiveness is often limited by exploration. For example, models often fail on hard problems, leaving little useful reward signal. External expert traces offer a natural source of guidance, yet they may also expose reward-relevant content along the critical path to the verifier target, such as final answers, intermediate values, executable implementations, or answer-related entities. This content can create an unintended reward hacking channel, allowing the policy to obtain reward by copying the trace rather than learning the underlying reasoning or agentic behavior. Existing guided-RL methods reduce this risk by using partial trajectories, but they mainly control how much expert information is shown heuristically rather than which parts should be hidden. To this end, we propose Semantic Masked Expert Policy Optimization (SMEPO), a fine-grained semantic masking strategy for expert-guided RLVR. Instead of truncating traces coarsely or revealing them unchanged, SMEPO masks reward-relevant semantic spans along the critical path while preserving the expert’s decomposition, plan, and procedural structure. This turns hard problems from reasoning from scratch into a fill-in-the-blank process: the policy can follow the expert’s problem-solving route, but must still reconstruct the missing values, code, or entities by itself. SMEPO is simple to apply and requires no changes to the reward function or RL objective. Across diverse domains, including math, code, and agentic search, SMEPO improves accuracy by up to 3.2 points over GRPO and reduces training time by up to 4.2$\times$.


Hierarchical Adaptive Frame Sampling For Video Understanding

Y Ys ⋅ Daiqi Shi ⋅ Shuang Li ⋅ Liao Zhang ⋅ Yang Cai ⋅ deqing wang

Due to context-length constraints, most MLLMs cannot process full-length videos and therefore rely on sampling a subset of frames as input. However, existing sampling methods, ranging from uniform sampling to relevance-based selection, are often driven by a single sampling principle and thus struggle to accommodate the heterogeneous evidential requirements of different queries: some demand a holistic understanding of how the video unfolds over time, while others hinge on fine-grained events within a short temporal window. To address this limitation, we propose Hierarchical Adaptive Frame Sampling (HAS), a two-stage frame sampling framework. In the first stage, Backbone Frame Construction, we apply a Determinantal Point Process (DPP) to sample frames that are both query-relevant and non-redundant. These selected frames capture key moments across the video and form the backbone of the sampling set, providing a foundation for subsequent enrichment. In the second stage, Adaptive Contextual Enrichment, we analyze the temporal distribution of the backbone frames to infer the query type and adaptively allocate the remaining frame budget between Local and Global Context. The Local Context enriches the backbone with fine-grained temporal dynamics and short-range causal relations, whereas the Global Context connects temporally isolated evidence and provides a holistic view of the entire video to support global understanding. Through such a hierarchical two-stage process, HAS effectively addresses diverse query requirements. Incorporated into three leading MLLMs, it demonstrates consistently superior performance across the Video-MME, LongVideoBench, and MLVU benchmarks. We have released our code.


HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

Hei Yi Mak ⋅ Shadan Golestan ⋅ Hoang Le ⋅ Mehran Taghian Jazi ⋅ Yunke Peng ⋅ Yaoyuan Wang ⋅ Yao Wang ⋅ JUNSONG WANG ⋅ Tianchi Hu ⋅ Fengchen He ⋅ Guipeng Hu ⋅ Tanzila Rahman ⋅ Anandharaju D Raju

We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. A systematic study reveals that the dominant source of degradation in FP4 RL is not training-side quantization error but rollout activation quantization: outliers stretch the dynamic range so far that a large number of activation values underflow to zero under FP4. Counterintuitively, restoring the training policy to higher precision while keeping the rollout in FP4 makes accuracy worse than full FP4 baseline, exposing rollout–training mismatch as the principal failure mode and ruling out standard pretraining-style fixes. We address this with Rollout Residual Quantization (Rollout-ResQ): a single residual correction term constrained to a hardware-friendly sparsity pattern, added only to the FP4 rollout matmul — a lightweight correction that recovers most of the precision lost to outlier-driven underflow without inflating the rollout’s compute footprint. On Qwen2.5-3B and Qwen2.5-Math-7B, Rollout-ResQ paired with the HiFloat4 (HiF4) format — whose three-level hierarchical scaling preserves resolution under FP4’s tight 4-bit budget — closes the accuracy gap to BF16 from 4.9% to 1.1%, bringing fully quantized FP4 RL within striking distance of full precision. Applied to the open-standard MXFP4, the same recipe narrows the gap from 13.6% to 5.3%, revealing that FP4 format choice is a key factor that determines the ceiling on recoverable accuracy. Together, these results establish HiF4 as the enabling format for end-to- end FP4 RL post-training, and Rollout-ResQ as the activation-side mechanism that makes the gap to BF16 closable.


High-arity Sample Compression

Leonardo Coregliano ⋅ William Opich

Recently, a series of works have started studying variations of concepts from learning theory for product spaces, which can be collected under the name high-arity learning theory. In this work, we consider a high-arity variant of sample compression schemes and we prove that the existence of a high-arity sample compression scheme of non-trivial quality implies high-arity PAC learnability.


High-Dimensional Robotic Reinforcement Learning with Developing Synergies

Junyi Shen ⋅ Yasuo Kuniyoshi ⋅ Kohei Nakajima

High-dimensional robotic reinforcement learning (RL) is often bottlenecked by inefficient exploration in large, redundant actuator spaces. We introduce Devyn (Developing synergies), a lightweight framework that learns a structured low-dimensional action parameterization directly during policy optimization. Unlike methods that require precomputed synergies or structured exploration in the full action space, Devyn starts from a randomly initialized latent-to-action map and shapes it online through task interaction. The policy acts strictly in a reduced latent space, while a learnable synergy matrix decodes latent actions into actuator commands. To make this evolving decoder stable and useful for policy learning, Devyn combines slow decoder evolution with structural regularization, reducing latent-action semantic drift while encouraging compact coordination patterns. Across diverse continuous-control benchmarks, including high-dimensional humanoid and overactuated musculoskeletal systems, Devyn improves upon standard SAC and achieves competitive or better performance than recent methods designed for high-dimensional robotic control. Ablations show that the gains do not come from a generic latent-action bottleneck alone, but rather from the slow, structured evolution of the synergy matrix. Our results suggest that learning the action parameterization itself can be a simple and effective route to scaling RL for high-dimensional robots.


HiLoRA: Adaptive Hierarchical LoRA Routing for Training-Free Domain Generalization

Ziyi Han ⋅ Huanyu Wang ⋅ Zeyu Zhang ⋅ Xiangxiang Dai ⋅ Xutong Liu ⋅ John C. S. Lui

Low-Rank Adaptation (LoRA) has emerged as a widely used technique for adapting large language models (LLMs) to new domains, due to its modular design and broad availability on platforms such as HuggingFace. This availability has motivated efforts to reuse existing LoRAs for domain generalization. However, existing methods often rely on explicit task labels or additional training, which are impractical for deployment. Moreover, they typically activate a fixed number of entire LoRA, leading to parameter redundancy or insufficiency that degrade performance. In this paper, we propose HiLoRA, a training-free framework that performs adaptive hierarchical routing over LoRA pools. Drawing on structural properties of LoRA, we define rank-one components (ROCs), where each LoRA rank is treated as an independent unit. For a given input, HiLoRA first adaptively selects a subset of LoRAs and determines their ROC allocation based on Gaussian likelihoods at the sequence level. At the token level, it refines routing by activating only the most informative ROCs. We further provide theoretical guarantees that HiLoRA selects the most relevant LoRAs with high probability. Extensive experiments show that HiLoRA achieves substantial improvements in domain generalization, with accuracy gains of up to 70\% over state-of-the-art baselines, while maintaining comparable inference throughput.


Hit Expansion via Localized Exploration of Synthesizable Chemical Space

Walter Virany ⋅ Yidong Jin ⋅ Andrew Lian ⋅ Dmytro Shevchuk ⋅ Amir Mehdi Soufi Enayati ⋅ Mike Tyers ⋅ Robert Batey ⋅ Piotr Gaiński ⋅ Michał Koziarski

Generative models for drug design which directly produce synthetic pathways have gained significant popularity due to their ability to constrain the search space to synthetically accessible molecules. However, existing methods have focused primarily on de novo molecular design, and rarely start the generation process from known binders. In this paper, we present HELiX: a template-based GFlowNet for localized exploration of chemical space. HELiX learns to partially decompose a given synthetic trajectory to an intermediate state, and then perform forward synthesis in a manner that preserves synthetic accessibility, leading to diverse, high-reward analog generation. Moreover, we prioritize sample efficiency by incorporating Bayesian optimization into the training procedure. We diagnose problems inherent to training GFlowNets with Bayesian optimization and introduce a greedy acquisition strategy which effectively balances between exploration and exploitation without the need for reward shaping. Finally, we show that local exploration is inherently robust to noisy oracle evaluations, a common problem in drug development when using unreliable proxies for binding affinity.

Estimating 6-degrees-of-freedom (6DoF) head pose from a single RGB image remains challenging under strong perspective distortion. We observe that the widely used design choice of rectangular face cropping introduces a geometric inconsistency. This leads not only to image--label mismatch under 2D rotation augmentation but also exacerbates perspective ambiguity, making it difficult to distinguish projection distortion from the actual appearance and pose of faces located near the image periphery. To address this issue, we propose HOGWARTS, a geometry-driven framework that replaces rectangular cropping with homography warping induced by a newly defined virtual camera space. By explicitly defining a virtual camera that shares the optical center with the original camera and warping the input image into this space, our method compensates for off-axis distortion in a geometrically consistent manner. Furthermore, to accurately estimate 3D translation in the virtual space, we introduce a pinhole-camera-based translation formulation that explicitly compensates for projection scale variation caused by the virtual camera transformation. Consequently, HOGWARTS enables more geometrically consistent and robust 6DoF pose estimation across varying image locations. To the best of our knowledge, this is the first study to estimate the 6DoF pose of a target within a virtual camera space explicitly designed to mitigate perspective distortion. Experimental results on the ARKitFace and BIWI datasets demonstrate that HOGWARTS consistently outperforms prior methods under severe perspective distortion as well as in cross-domain settings. Code will be made publicly available.


HoloGene: Learning to Lift Sliced Spatial Transcriptomics to Holistic 3D Gene Fields

Yumin Zheng ⋅ Qingtian Zhu ⋅ Ziyan Zhu ⋅ Kailu Song ⋅ Yinqiang Zheng ⋅ Jun Ding

Spatial transcriptomics (ST) provides spatially resolved gene expression measurements, but most 3D ST datasets are acquired as independently processed 2D sections. This sliced acquisition creates a difficult reconstruction problem: measurements are dense within each section but sparse along the axial direction, and inter-section technical variation can be confounded with true biological change. Existing coordinate-based implicit neural representations (INRs) provide a continuous modeling framework, but standard isotropic coordinate encodings do not explicitly reflect this axial-to-lateral imbalance. We present HoloGene, an anisotropic conditional INR framework for estimating continuous 3D gene-expression fields from pre-registered ST sections. HoloGene first compresses high-dimensional expression profiles into a graph-autoencoder latent manifold. It then uses pseudo-label-derived biological signature conditioning to provide coarse intra-section biological context, while depth-dependent FiLM modulation models axial variation. To reduce slice-specific covariance shifts, we introduce a signature-stratified correlation-alignment regularizer over latent features. Using held-out measured sections as cross-sectional references, HoloGene improves reconstruction accuracy and biological preservation metrics compared with general-purpose INRs and adapted spatial-transcriptomics baselines. Ablation analyses show that decoupled depth conditioning improves axial generalization, while correlation alignment provides the strongest reduction in slice-specific representation leakage. These results support anisotropic conditional modeling as a practical strategy for estimating continuous 3D expression fields from discontinuous ST sections.


Homological Barriers to Stable Local Nash Dynamics in Quadratic Zero-Sum Games

Ashkan Soleymani ⋅ Gabriele Farina ⋅ Patrick Jaillet ⋅ Georgios Piliouras

We identify a topological obstruction to the standard stable-attractor template for learning local Nash equilibria, also known as first-order Nash equilibria (FONE), in constrained non-concave differentiable games. We consider continuous deterministic learning dynamics that leave each FONE point fixed and aim to attract every initial condition to the full FONE set. Our main message is that this familiar convergence template can fail for a topological reason. Targets with holes, such as loops, cannot be stable global attractors for such dynamics on contractible state spaces. We then exhibit this obstruction in a simple non-concave quadratic zero-sum game on a box. Its exact FONE set is a diagonal six-edge loop together with an isolated point, so the loop creates the required topological obstruction. As a result, no continuous pointwise-stationary learning rule covered by our framework can make this exact FONE set a stable global attractor. We show that degree two is the minimal polynomial degree at which this obstruction can arise. We further amplify the construction to arbitrary homological degree. Finally, by studying the topology in sufficiently small perturbed version of the game, we show that an open positive-volume family of quadratic games retain the same homological obstruction. Thus, even in quadratic zero-sum games, this phenomenon is not a pathology of a single construction; instead, the geometry of projected first-order stationarity can create an intrinsic barrier to stable global FONE dynamics.


HOPSE: Scalable Higher-Order Positional and Structural Encoder for Combinatorial Representations

Guillermo Bernárdez ⋅ Marco Montagna ⋅ Louis Van Langendonck ⋅ Martin Carrasco ⋅ Amirreza Akbari ⋅ Louisa Cornelis ⋅ Mathilde Papillon ⋅ Pere Barlet-Ros ⋅ Nina Miolane ⋅ Lev Telyatnikov

While Graph Neural Networks (GNNs) have proven highly effective at modeling relational data, pairwise connections cannot fully capture multi-way relationships naturally present in complex real-world systems. In response to this, Topological Deep Learning (TDL) leverages more general combinatorial representations--such as simplicial or cellular complexes--to accommodate higher-order interactions. Existing TDL methods often extend GNNs through Higher-Order Message Passing (HOMP), but face critical scalability challenges due to the steep complexity overhead of propagating messages through combinatorial structures. To overcome this limitation, we propose HOPSE (Higher-Order Positional and Structural Encoder), a framework \emph{free of message passing layers} that uses Hasse graph decompositions to derive efficient and expressive encodings over \emph{arbitrary higher-order domains}. Notably, HOPSE scales linearly with the size of combinatorial representations while preserving the expressive power and permutation equivariance of the HOMP approaches. Experiments on molecular and topological benchmarks show that it matches or surpasses state-of-the-art performance while consistently achieving speedups over HOMP-based models, opening a new path for scalable TDL. The code is available at https://anonymous.4open.science/r/TopoBench-BF89.

Out-of-distribution (OOD) generalization aims to improve a model’s generalization capacity using only source data, thereby ensuring reliable performance on unseen domains. Cutout, a well-established method that enhances model generalization by randomly masking out square regions of training samples, has shown great success in standard supervised settings. Despite its seemingly strong con- nection to OOD generalization, we observe that simply combining Cutout with existing OOD methods yields only marginal benefits or even degrades performance. Motivated by this empirical finding, we seek a way to make OOD generalization benefit from Cutout. Our method is inspired by reinterpreting the state-of-the-art hyperspher- ical prototype learning as a mutual information (MI) framework. Within this framework, we observe that naively applying Cutout to existing OOD methods is equivalent to estimating MI from a single masked subview, which leaves much of Cutout’s potential untapped. Instead, we apply the MI chain rule to split the original objective into two smaller estimation problems: one encourages the Cutout subview to align closely with the label, while the other drives the full image to capture the residual information unavailable in that subview. This divide-and-conquer strategy fully exploits the Cutout-generated subview while remaining computationally light- weight. Empirically, we demonstrate that our method outperforms competitive baselines on a broad range of OOD benchmarks and achieves superior performance.


How Fine-Tuning Objectives Shape Layer-Wise Information in LLM Hallucination Detection

Kaiyang Wan ⋅ Forrest Bao ⋅ Amin Ahmad ⋅ Yuxia Wang

Under matched conditions, classifier-head fine-tuning (\textsc{cls}) paradigm consistently outperforms the LLM-as-Judge sequence-to-sequence (\textsc{s2s}) on binary hallucination detection. We trace this gap to the representations induced by the two objectives. Static analysis shows that \textsc{cls} learns a task-specific discriminative direction, pushing hidden states into a clear two-cluster geometry with stronger class separation and linearly decodable label information. By contrast, \textsc{s2s} relies on the pretrained language-modeling head to make decisions. This minimizes the pressure to reorganize representations, leaving the hidden states much closer to the pretrained backbone. Our analysis also reveals that the \textsc{cls} signal saturates before the final layer, after which upper layers form a plateau of statistically interchangeable and largely redundant representations. However, endpoint redundancy alone does not show whether these layers are dispensable during training. Further training-dynamics analysis reveals 1) a lower pretrained substrate, 2) a narrow upper-middle \textbf{InfoWindow} where task information emerges, and 3) a later-synchronizing upper plateau that mirrors this signal. This mechanism motivates us to remove layers above the InfoWindow. By truncating roughly 31\% of the network, we obtain early-exit models that match or exceed full-model performance on Qwen3-4B and 8B. The same depth-fractional prescription transfers to Llama-3.2-3B and Mistral-7B-v0.3. We hope this helps guide efficient task-specific LLM design.


How Much Information is Needed for Accurate Kalman Filtering?

Wenhan Cao ⋅ Xuyang Chen ⋅ Shuyuan Wang ⋅ Lin Zhao

Kalman filtering provides a principled inference framework for linear Gaussian hidden Markov models, but it leaves open a complementary question of representation: what is the minimum amount of information about the observation history that must be retained to support accurate filtering? We show that this question leads naturally to an indirect rate-distortion problem, in which an encoder observes the history of noisy measurements while distortion is evaluated with respect to the latent state. Directly solving this problem is challenging, since the optimization ranges over arbitrary stochastic kernels from a time-varying observation history to an output representation, resulting in an infinite-dimensional formulation. We overcome this difficulty by proving that any feasible encoder can be replaced by a stochastic encoder acting only on the Kalman posterior mean, achieving the same mean-square error with no larger mutual information. The reduced problem admits a closed-form solution in the form of a linear Gaussian blurring mechanism determined by the spectral decomposition of the Riccati solution. This yields an explicit expression for the rate distortion function $R_t(\epsilon)$, showing that any representation with distortion at most $\epsilon$ must retain at least $R_t(\epsilon)$ nats of information. We further use this characterization to derive information-theoretic lower bounds on the population risk of sequential learning and on the prediction error of temporal GP. Numerical experiments corroborate the theoretical predictions.


Humans Correct Their Judgements Through Debate, but Weaker AI Models Not

Maria Victoria Carro ⋅ Denise Alejandra Mester ⋅ Facundo Nieto ⋅ Oscar Stanchi ⋅ Guido Bergman ⋅ Giovanni Franco Gabriel Marraffini ⋅ Trinidad Borrell ⋅ Guido Freire ⋅ Joaquín S Machulsky ⋅ Mario Leiva ⋅ Ana Lopez ⋅ Nicolas Spinelli ⋅ Luca Gangi ⋅ Eitan Sprejer ⋅ Federico Barrera-Lemarchand ⋅ Gerardo Simari ⋅ Maria Vanina Martinez

How can we effectively oversee AI systems that surpass human intelligence? One proposed answer is AI Safety via Debate, an approach in which two AI systems argue opposing positions and then a judge, who can be either a human or a weaker model, decides which one is correct. By facilitating supervisors in providing high-quality training signals, debate aims to encourage truthful and safe behavior from AI systems whose capabilities exceed our own. However, human judgment is not neutral: people often rely on biases which advanced AI may learn to exploit. Drawing on insights from collective intelligence and social cognition, this paper investigates whether debate can guide human and LLM supervisors toward the truth despite their prior beliefs. We introduce and evaluate several protocol configurations, including (1) the use of open-source, proprietary models and (2) human participants as judges (N=507), (3) different methods for eliciting prior beliefs, and (4) variations in claims within the same topic to test whether belief updates generalize across related subtopics. We also explore setups inspired by the wisdom of the crowds, such as the number and diversity of judges, as well as interaction mechanisms including deliberation and voting, designed to better align judgments with the truth. Overall, we find that debate increases accuracy relative to prior beliefs for human judges (from 41.3\% to 57.6\%), whereas the consultancy baseline, in which a single expert argues for one answer, does not. For weaker language models, however, the results are mixed. Debate does not provide similar benefits in most settings and can even degrade decision accuracy. Regarding multi-judge protocols, human deliberation yields the highest accuracy for human supervisors (62.2\%), while for LLMs it mainly benefits those starting from incorrect beliefs, without producing significant improvements in aggregate performance. We provide empirical evidence supporting Debate as a viable path toward scalable human oversight, while casting doubt on the assumption that such benefits translate to weaker AI supervisors, raising concerns about the reliability of fully automated training and control pipelines in high-stakes AI safety contexts.


Hybrid Neural World Models for Physical Dynamics

Pranav Lakshmanan ⋅ Paras Chopra

Neural surrogates promise large speedups over classical solvers for physical dynamics but fail silently at sharp dynamical events such as shocks, fronts, and contact. We present hybrid neural world models for physical dynamics: a recipe for training and deploying multi-horizon surrogates in physical state space, where a single network with continuous horizon conditioning is trained with direct supervision against textbook reference solvers to predict any future state at horizon $T$ in one forward pass. Although no part of the training data, loss function, or architecture supervises discontinuity location, the trained surrogate encodes it implicitly, recoverable from its forward passes alone as a per-trajectory error map that concentrates on shocks, fronts, and contacts, and stays small elsewhere. The map is competitive with or better than standard label-free baselines including deep ensembles, learned error heads, gradient-magnitude indicators, and locally-adaptive conformal prediction, while using only a single trained network and requiring no calibration set or governing-equation knowledge. The recipe supports two operating points. Mode 1 runs the surrogate alone for maximum throughput, with same-hardware CPU speedups of $26\times$ to $72\times$ against textbook solvers on the PDE environments. Mode 2 uses the error map to gate a reference-solver fallback, deferring uncertain trajectories and roughly halving the surrogate's residual error at the default operating point. The recipe applies without modification across reaction-diffusion, compressible Euler, and rigid-body collision dynamics.


Hydra-X: Native Unified Multimodal Models with Holistic Visual Tokenizers

Guozhen Zhang ⋅ Xuerui Qiu ⋅ Yutao Cui ⋅ Tianhui Song ⋅ Changlin Li ⋅ Junzhe Li ⋅ Tao Huang ⋅ Xiao Zhang ⋅ Yang Li ⋅ Jianbing Wu ⋅ Miles Yang ⋅ Zhao Zhong ⋅ Liefeng Bo ⋅ Limin Wang

Holistic visual tokenizers are fundamental to unified multimodal models (UMMs) as they map diverse visual inputs into a unified representation space. In this paper, we present HYDRA-X, the first UMM framework that unifies image and video tokenization within a single Vision Transformer (ViT). Our design is driven by two core challenges: efficiently injecting spatiotemporal reconstruction capability into a native ViT, and embedding image- and video-level semantic awareness into the latent space. To address the first, comprehensive ablations reveal two key findings: (1) frame-level causal temporal attention suffices for visual reconstruction, whereas full spatiotemporal attention degrades it; and (2) hierarchical temporal compression substantially outperforms single-step alternatives. To tackle the second, we propose a lightweight decompressor that upsamples temporally compressed features under joint image-video teacher supervision, thereby enforcing complementary semantic structures within the compact latent space. Building on this holistic tokenizer, we further propose a principled improvement of the editing pipeline: source-target interaction should occur at the latent level inside the tokenizer rather than at the semantic level inside the LLM, substantially improving editing consistency and accelerating convergence. Instantiated at the 7B dense model, HYDRA-X achieves strong performance across image and video understanding and generation tasks, laying a solid foundation for future unified-tokenizer UMM exploration.


Hyper Hawkes Processes: Interpretable Models of Marked Temporal Point Processes

Alex Boyd ⋅ andrew warrington ⋅ Taha Kass-Hout ⋅ Parminder Bhatia ⋅ Cao (Danica) Xiao

Foundational marked temporal point process (MTPP) models, such as the Hawkes process, often use inexpressive model families in order to offer interpretable parameterizations of event data. On the other hand, neural MTPPs models forego this interpretability in favor of absolute predictive performance. In this work, we present a new family MTPP models: the hyper Hawkes process (HHP), which aims to be as flexible and performant as neural MTPPs, while retaining interpretable aspects. To achieve this, the HHP extends the classical Hawkes process to increase its expressivity by first expanding the dimension of the process into a latent space, and then introducing a hypernetwork to allow time- and data-dependent dynamics. These extensions define a highly performant MTPP family, achieving state-of-the-art performance across a range of benchmark tasks and metrics. Furthermore, by retaining the linearity of the recurrence, albeit now piecewise and conditionally linear, the HHP also retains much of the structure of the original Hawkes process, which we exploit to create direct probes into how the model creates predictions. HHP models therefore offer both state-of-the-art predictions, while also providing an opportunity to ``open the box'' and inspect how predictions were generated.


HyperSkill: Training-Free Omnimodal GRPO via Hypergraph-Indexed Skill-Library Evolution

Haoran Luo ⋅ Shangkai Lin ⋅ Jinyang Wu ⋅ XINLIANG ZHOU ⋅ Anh Tuan Luu

Reinforcement learning sharpens LLM reasoning, but parameter updates are expensive, forgetful, and opaque; training-free libraries swap gradients for a textual store yet inherit trajectory-level admission and flat structure. We close these gaps with HyperSkill, a Training-Free Omnimodal GRPO framework that keeps the backbone frozen and replaces gradient updates with auditable edits to a Hypergraph-Indexed Skill-Library, where skills are advantage-bearing nodes and compositions are hyperedges. Three mechanisms drive its evolution: a counterfactually verified, step-level distillation admits new skills; moving-average updates with hyperedge growth and pruning follow retrieval feedback; and an advantage-weighted retriever with similarity and capacity gates falls back to zero-skill. Across 11 text/VL/omnimodal benchmarks, HyperSkill beats every baseline by 6.9 F1 (15.6%) without touching a parameter. Our software and data are publicly available.


I2V-DETACH: Source Grounding Detachment for Unauthorized Image-to-Video Generation

Chanhui Lee ⋅ Yeonghwan Song ⋅ Yewon Kang ⋅ Jeany Son

Recent diffusion models have advanced image-conditioned video generation, but their accessibility raises risks of unauthorized manipulation. Image immunization has emerged as a promising defense for Image-to-Video (I2V) protection, yet existing methods often rely on degrading visual quality or temporal coherence. In this work, we revisit I2V immunization as reducing source-content preservation, shifting the focus from output degradation to weakening source-conditioned content propagation. To this end, we propose I2V-DETACH, an image immunization method that prevents unauthorized I2V generation through source grounding detachment. I2V models preserve source content by propagating source-image cues into video latents during denoising. Accordingly, I2V-DETACH disrupts this process in two complementary ways: it suppresses source-video grounding to reduce source cue propagation, and strengthens intra-video interactions to shift denoising toward video-side dynamics with reduced reliance on the source image. To obtain informative optimization signals, we initialize auxiliary video latents with source-oriented states that remain within the source-conditioned generation regime while deviating from source-preserving trajectories. Extensive experiments across diverse diffusion architectures show that I2V-DETACH substantially reduces source-content preservation, improves over prior immunization methods, and exhibits strong black-box transferability.

Reference-conditioned image-to-image (I2I) editing models provide a promising zero-shot approach to personalized image generation by jointly conditioning on a reference image and a text prompt. We show that recent I2I editing models, including FLUX.1 Kontext and FLUX.2, achieve strong subject identity preservation and prompt alignment, but exhibit seed-wise diversity collapse: under the same reference image and prompt, different random seeds often produce highly similar outputs. We further find that this collapse is associated with over-alignment of latent-token trajectories during denoising. Based on this observation, we introduce Selective Hidden Perturbation (SHiP), a training-free inference-time intervention that perturbs only the attention-output representations of latent tokens during early denoising while keeping reference-image and text tokens unchanged. Across diverse subjects and prompts, SHiP improves seed-wise diversity on both FLUX.1 Kontext and FLUX.2 while largely preserving identity and prompt fidelity. On FLUX.1 Kontext, SHiP increases Vendi Score from 1.546 to 2.066 and LPIPS from 0.475 to 0.621, with only minor decreases in GPT-ID and GPT-Text. Code is available at https://anonymous.4open.science/r/i2i-personalization-diversity-0260/.


Identifying Latent Neural Dynamics with Recognition-Parameterized Gaussian Process Dynamical Systems

Arielle Rosinski ⋅ Qianqian Feng ⋅ Lior Fox ⋅ Maneesh Sahani

Neural circuits are densely interconnected systems, and many neural computations are thought to be mediated by the dynamical evolution of activity within these recurrent networks, a hypothesis often referred to as ‘computation through dynamics’. Under this view, understanding computation within real neural populations depends on building models of the dynamical rules governing the temporal evolution of recorded high-dimensional neural activity. Most commonly, such models depend on explicit generative parameterizations of (i) the (nonlinear) intrinsic dynamical flow field and stochasticity, (ii) the (nonlinear) mapping from dynamical state to neuronal activity, and (iii) the variability of individual neural responses given that dynamical state. Misspecification of any of these components can bias estimates of dynamics and latent trajectories. In this paper, we circumvent generation-related issues by introducing the Recognition-Parameterized Gaussian Process dynamics (RP-GPdyn) model. RP-GPdyn models the dynamical flow field using a nonparametric Gaussian Process-defined transition function and models the relationship of dynamics to neural activity implicitly using the Recognition-Parameterized Model (RPM) framework. We show that this generation-free approach uncovers meaningful, behaviorally relevant latent variables and dynamics from both synthetic and experimental datasets.


ImmuVis: Hyperconvolutional Foundation Models for Imaging Mass Cytometry

Dawid Uchal ⋅ Marcin Możejko ⋅ Krzysztof Gogolewski ⋅ Piotr Kupidura ⋅ Fabian Ozga ⋅ Szymon Łukasik ⋅ Jakub Giezgała ⋅ Tomasz Nocoń ⋅ Kacper Pietrzyk ⋅ Robert Pieniuta ⋅ Mateusz Sulimowicz ⋅ Michał Zmysłowski ⋅ Michał Orzyłowski ⋅ Tomasz Siłkowski ⋅ Karol Zagródka ⋅ Eike Staub ⋅ Ewa Szczurek

We present ImmuVis, a family of efficient foundation models for imaging mass cytometry (IMC), a high-throughput multiplex imaging technology that handles molecular marker measurements as image channels and enables large-scale spatial tissue profiling. Unlike natural images, multiplex imaging lacks a fixed channel space, as real-world marker sets vary across studies, violating a core assumption of standard vision backbones. To address this, ImmuVis introduces marker-adaptive hyperconvolutions that generate convolutional kernels from learned marker embeddings, enabling a single model to operate on arbitrary measured marker subsets without retraining. We pretrain ImmuVis on the largest IMC dataset to date, IMC17M (28 cohorts, 24,405 images, 265 markers, over 17M patches), using self-supervised masked reconstruction. ImmuVis outperforms state-of-the-art baselines and ablations in virtual staining and downstream classification tasks at substantially lower compute cost than transformer-based alternatives, and is the sole model that provides calibrated uncertainty via a heteroscedastic likelihood objective. These results position ImmuVis as a practical framework for real-world IMC modeling.


Importance-Weighted Operator Learning Under Probability Measure Shifts

Lei Sun ⋅ Yusuke Tanaka ⋅ Xiaocheng Shang ⋅ Takaharu Yaguchi ⋅ Tomoharu Iwata

Operator learning has attracted significant interest for its capability to approximate mappings between function spaces. A common assumption in this field is that training and test samples are drawn from the same probability measure. However, this assumption is often violated, leading to probability measure shifts. Under such shifts, standard operator learning methods often lead to degraded performance because the expected training error does not match the expected test error. In this paper, we propose importance-weighted operator learning (IWOL), a framework for learning operators under probability measure shifts in function spaces. We prove that, when the test measure is absolutely continuous with respect to the training measure, the expected test error is equivalent to an importance-weighted expected training error, with weights given by the Radon--Nikodym derivative. This result provides a theoretical basis for correcting measure shifts in operator learning, but importance weighting relies on this absolute-continuity condition and may yield unbounded weights. To overcome these limitations, we introduce relative importance weights defined with respect to a mixture of the training and test measures. The resulting weights remain well-defined even when absolute continuity fails, and they are uniformly bounded. These weights induce a relative importance-weighted training objective for neural operators under probability measure shifts. To estimate these weights, we present binary classifiers that take functions as inputs, defined through functionals modeled by kernel integral operators. This design enables discretization-invariant weight estimation in function space. Experiments on multiple PDE benchmarks and neural operator architectures demonstrate that the proposed method consistently improves performance under probability measure shifts.

Achieving both privacy and robustness against poisoning attacks remains a central challenge in federated learning (FL), since hiding individual updates makes attacks harder to detect and mitigate. Recently, Mai et al. (NeurIPS 2024) proposed RFLPA, a packed-secret-sharing (PSS)-based realization of FLTrust, which weights client updates by their cosine similarity to a trusted server update. In this paper, we show that RFLPA contains design flaws and mathematical errors in its use of PSS, making the protocol theoretically unsound and practically non-functional. We then propose an improved verifiable FL framework that restores correctness while retaining efficiency. Our key insight is a reinterpretation of the algebraic structure of PSS, which leads to a new transposed degree reduction algorithm and a pairwise verification technique for dot products. These primitives enable local matrix-based operations without complex circuit evaluation and substantially reduce communication rounds compared with BGW-style computation. Building on them, we redesign the verifiable aggregation algorithm of RFLPA by introducing perpendicular vectors, reducing cosine similarity verification to a single dot-product check and eliminating costly verifiable secret sharing during enrollment. The resulting framework strictly improves efficiency, fixes the flaws of prior work, and preserves $N/3$-robustness in malicious settings. Our implementation and empirical evaluation confirm the feasibility of the resulting protocol for high-dimensional FL updates.

Confronted by the existing prototype-based methods that often fail to capture complex intra-class variations and inter-class semantic relationships, weakly supervised semantic segmentation (WSSS) faces persistent challenges due to the inherent incompleteness and boundary inaccuracies of Class Activation Maps (CAMs). To tackle these critical limitations, we propose a relationally optimized prototype Memory bank (RO-PMB) framework in this paper, where the integration of Graph Neural Networks (GNNs) with a dynamically optimized global prototype memory bank is pioneered to explicitly model and learn the intricate semantic relationships among all prototypes. Specifically, our Prototype Relationship Diagram Construction (PRDC) module leverages GNNs for contextual message passing over a learned prototype graph, enriching global prototypes with crucial associative knowledge. As a result, a subsequent context-aware refinement process is sustained to inject structured semantic information into local image-specific prototypes, thereby generating CAMs with unparalleled completeness and boundary precisions. Extensive experiments validate that, compared with the existing representative state of the arts in multi-stage WSSS baselines, our proposed RO-PMB achieves competitive and superior performances on a range of challenging benchmarks, such as PASCAL VOC and MS COCO etc. Code will be released.


Improving LLM Final Representations with Inter-Layer Geometry

Tom Ulanovski ⋅ Eyal Blyachman ⋅ Maya Bechler-Speicher

The standard in LLM-based prediction is to use the final-layer representation as the input to a downstream predictor. However, intermediate layers may encode complementary task-relevant signals. Existing approaches therefore either search for the best layer for each task or apply expensive attention-based mechanisms to learn inter-layer aggregation. In this work, we first show that such complexity is unnecessary: a lightweight Graph Neural Network over a fully connected graph of LLM layers is more efficient and achieves significantly stronger predictive performance than existing approaches. We then introduce the Cayley-Encoder, which further improves both efficiency and predictive performance by replacing the fully connected graph with a Cayley graph over $SL(2,\mathbb{Z}_n)$. These Cayley graphs provide a mathematically grounded topology that is sparse, regular by construction, and has low diameter. This enables effective communication across layers while constraining the aggregation structure and reducing the risk of GNN overfitting. In an evaluation of Cayley-Encoder across 13 tasks and 9 LLMs, Cayley-Encoder consistently outperforms baselines, achieving improvements of up to 40 percentage points in accuracy, while introducing at most 0.1\% additional parameters relative to the LLM size. We further show that Cayley-Encoder is effective in few-shot regimes. Finally, we show that Cayley-Encoder outperforms LoRA fine-tuning while operating on the frozen LLM. We conclude with an explainability analysis showing that multiple layers contribute meaningfully to the final prediction, supporting our hypothesis.


Improving Quantized Zeroth-Order Optimization through Reconstructed Low-Rank Structures

Fei Wang ⋅ Shuai Xie ⋅ Li Shen ⋅ Chao Xue ⋅ Ye Liu ⋅ Changxing Ding

Fine-tuning Large Language Models under stringent memory constraints remains a significant challenge. Recently, Zeroth-Order (ZO) optimizers have been employed to fine-tune quantized models, avoiding backpropagation and further reducing memory usage. However, quantized ZO optimizers often underperform compared to their full-precision counterparts, especially at low bit-widths. We find that existing quantized ZO methods fail to exploit the inherent low-rank structure of gradient updates, thereby limiting their effectiveness. To address this, we propose the Low-rank Quantized Zeroth-Order optimizer (LQZO), which enhances quantized ZO fine-tuning by leveraging the low-rank structure inherent in gradients. Our method introduces three key innovations: (1) A low-rank parameterization of gradient perturbations using structured matrix factorization. (2) A Binary-Aggregated Gaussian perturbation strategy that employs a Rademacher distribution to constrain higher-order moments and tighten the variance bound of gradient estimates, with no additional computational overhead. (3) An Isotropy-Enhanced Estimation technique that reshapes quantization scale groups into more isotropic structures to diversify exploration directions. We provide a theoretical convergence guarantee for LQZO. Extensive experiments demonstrate that LQZO consistently outperforms quantized ZO optimizers, achieving average performance gains of 2.0%--3.3% across quantization precisions. While extreme quantization inevitably leads to information loss, LQZO significantly improves the trainability of INT2 models and narrows the performance gap compared to prior quantized ZO methods. Furthermore, we demonstrate the cross-modality versatility of LQZO by validating its superior performance on Vision Transformers.


In-Context Benign Overfitting: A Feature-Selection Model in Linear Regression ICL

Puneesh Deora ⋅ Bhavya Vasudeva ⋅ Christos Thrampoulidis

In in-context learning (ICL), a frozen pre-trained model solves tasks by conditioning on a prompt of a few input–output examples, without gradient updates. If the task was present in pretraining but the particular prompt sequence was not, the resulting in-distribution generalization is retrieval-based ICL. Learning-based ICL instead reflects out-of-distribution generalization: the model succeeds on prompts generated by a novel task. Empirically, both forms improve with scale. By analogy to benign overfitting in supervised learning, we call this in-context benign overfitting: larger models more faithfully memorize the pretraining tasks (improving retrieval ICL) while also generalizing better to novel tasks (improving learning ICL). We prove that this phenomenon already arises in a minimal in-context linear-regression feature-selection model. In contrast, standard in-context linear-regression models exhibit a retrieval–learning tradeoff, where the emergence of learning-based ICL coincides with degraded retrieval-based performance.


Independent Latents, Robust Neural Operators

Jay Yoo ⋅ Kazuma Kobayashi ⋅ Jaewan Park ⋅ Sai puppala ⋅ Souvik Chakraborty ⋅ Syed Bahauddin Alam

Neural operators have emerged as powerful surrogates for scientific computing, enabling rapid reconstruction of complex physical fields from sparse sensor observations. Despite their strong predictive capability, their reliability often degrades under noisy or corrupted measurements, limiting practical deployment in real-world sensing environments. We show that enforcing statistical independence in latent representations provides a simple yet effective mechanism for improving neural operator robustness. To this end, we introduce Neural Operator Independence Regularization (NOIR), an end-to-end training framework that promotes statistically independent latent features during operator learning. Unlike structured basis approaches such as POD-DeepONet and PCA-Net, which impose orthogonality primarily as an offline compression constraint, NOIR directly shapes latent representations during training to enhance resilience against input perturbations. Across three canonical PDE benchmarks and five additive sensor corruption settings, NOIR consistently improves reconstruction accuracy while preserving clean-data performance. Beyond predictive gains, the induced latent representations exhibit distinctive statistical signatures that reveal architecture-specific robustness characteristics under corruption. These findings identify latent statistical independence as a principled inductive bias for robust neural operator design and provide new insight into the internal structure of operator learning systems. All models, datasets, checkpoints, and training code are publicly available at \url{https://github.com/t55176853-cmyk/NOIR/}


InfCoiL: Coordinated Planner-Controller Learning for Closed-Loop Physics-Based Human-Object Interaction

Yude Zou ⋅ Yifei Yao ⋅ Yiwei Hao ⋅ Hanqing Wang ⋅ Xing Gao

Text- and goal-conditioned human-object interaction (HOI) provides an intuitive interface for specifying high-level intents in controllable character animation and embodied AI. Mapping such intents to physically plausible, contact-rich behaviors naturally requires hierarchical planning and control, yet existing planner-controller systems remain largely open-loop and lack adaptation to physical execution feedback. To address this, we present InfCoiL, a closed-loop framework for physics-based HOI from coordinated planner-controller learning. InfCoiL integrates a morphology-aware flow-matching planner with a unified multi-morphology controller in an autoregressive plan-and-execute loop, enabling diverse, controllable, and executable interactions across diverse objects and morphologies of humanoid embodiments. To further mitigate planner-controller distribution shift, we alternatively fine-tune both components under closed-loop reinforcement learning. Particularly, we introduce Semantic Flow-GRPO, which optimizes the flow-matching HOI planner with semantic covariance over heterogeneous HOI features and group-relative rewards from physical rollouts, while training the controller for reliable execution under contact-rich dynamics. Extensive experiments demonstrate that InfCoiL achieves state-of-the-art performance in both motion quality and task success. The code will be released upon acceptance.

Safety-aligned Large Language Models (LLMs) remain vulnerable to interventions during inference that redirect generation toward harmful outputs. Recent work attributes this to shallow safety, where alignment concentrates in the first few output tokens. We show that shallow safety is a special case of a broader inference-time vulnerability, in which short token injections at any generation step can substantially alter subsequent safety behavior. We also find that a model's alignment with refusal directions in its hidden states does not predict its robustness to such injection, revealing that internal state alone does not determine generation behavior under perturbation. To address this, we align models directly on generation trajectories constructed by simulating mid-sequence perturbation, and show that this improves robustness to mid-sequence injection and generalizes to attacks that exploit early-token generation. Our work argues that robust safety alignment requires training on the generation process itself, not only its outputs.


InfoFlow: A Framework for Multi-Layer Transformer Analysis

Penghao Yu ⋅ Haotian Jiang ⋅ Zeyu Bao ⋅ Qianxiao Li

While the approximation properties of single-layer Transformer architectures have been studied in recent works, a rigorous theoretical understanding of the multi-layer setting remains limited. In this work, we establish that multi-layer Transformers possess fundamentally different approximation capabilities from single-layer ones: for certain retrieval tasks, any single-layer Transformer requires least $\Omega (\varepsilon^{-k})$ parameters to achieve precision $\varepsilon$, where $k$ grows linearly with sequence length $T$, whereas a two-layer Transformer with a single head per layer achieves the same approximation precision with at most $O (\varepsilon^{-1})$ parameters. To understand this separation, we identify two structural mechanisms underlying multi-layer approximation. Specifically, softmax attention can only efficiently retrieve the token attaining the maximum attention score, incurring exponential-in-length parameter cost for $k$-th largest retrieval with $k \geq 2$. Moreover, the parameter cost of decoding coupled information scales with the size of the retrieved token set. Motivated by these findings, we propose **InfoFlow**, a framework for multi-layer Transformers. The framework tracks an information set of accessible input positions at each token and layer, assigning an explicit approximation rate to each mode of information propagation. This abstraction recovers known approximation bounds, remains consistent with experimental observations on trained networks, and yields concrete predictions in settings where direct theoretical analysis is currently intractable. Our results provide a principled framework for reasoning about the approximation efficiency of multi-layer Transformers.


In search of a definition of importance: Do attributions capture it?

Antonia Marcu ⋅ Jonathon Hare ⋅ Annika Catulli ⋅ Srinandan Dasmahapatra ⋅ Damian Smith

Attribution methods are extensively used to identify which parts of the input are important for a model's decision. However, despite recent efforts there is still no consensus on what exactly they capture and how to evaluate them. We start addressing this problem by explicitly modelling the implicit assumptions that sit at the foundation of attribution methods. We define two types of importance contributions an input region can have. We then create a framework that decouples the process of verifiably establishing contributions from that of evaluating attributions. This shows that current attribution methods struggle to reliably capture input contributions. Further investigation shows a stark difference in attribution performance between model behaviours, exposing a previously overlooked aspect in the attribution literature. Our work calls for clear, quantifiable statements about what attribution methods aim to capture, along with rigorous evaluation frameworks.


Inside the Loop: A Mechanistic Study of Weight-Tied Transformers on Depth-Bound Algorithmic Tasks

Magnus Sesodia ⋅ Philip Torr ⋅ Christian Schroeder de Witt

Solving algorithmic tasks that involve many sequential steps typically requires scaling up transformer depth or chain-of-thought length, incurring unfavorable parameter and inference-compute costs, respectively, as tasks deepen. Looped transformers promise to recover that depth at a fraction of the cost by reusing a single block across iterations, but it remains unclear how the loop actually performs its computation, or whether it merely simulates a deeper untied stack. We test this on \emph{directed permutation traversal} (\dePo), a synthetic task in the physics-of-LMs lineage in which solving an $H$-hop query requires following $H$ pointers along a fixed permutation of $N{=}16$ entities, with $H$ up to $15$. We demonstrate that, above a minimum effective depth, a $3.2$\,M-parameter looped transformer trained with hop length sampled uniformly per batch traverses all $15$ hops across all seeds, while a $25.8$\,M-parameter fixed-depth baseline at the same effective depth fails to exceed three hops; weight tying, not effective depth, predicts capability. In the same models, linear probes on the residual stream uncover a clean \emph{staircase}: at iteration $k$ the stream linearly decodes hop $k$, with slope exactly one hop per iteration, and each hop stays decodable for roughly two further iterations before being overwritten. The model's own logit lens, applied through the trained final layers, traces the same one-hop-per-iteration ordering at lower amplitude. Taken together, our results recast looped transformers as iterative program learners: weight tying forces the optimizer to discover a single per-iteration update rule, and that rule corresponds to one algorithmic step of the underlying task. Code is available at \url{https://anonymous.4open.science/r/insidetheloop-3ED0}.


Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors

Jonas Wiedermann-Möller ⋅ Leonard Dung ⋅ Maksym Andriushchenko

AI systems have become increasingly capable of dangerous behaviours in many domains. This raises the question: Do models sometimes choose to violate human instructions in order to perform behaviour that is more useful for certain goals? We introduce a benchmark for measuring model propensity for instrumental convergence (IC) behaviour in terminal-based agents. This is behaviour such as self-preservation that has been hypothesised to play a key role in risks from highly capable AI agents. Our benchmark is realistic and low-stakes which serves to reduce evaluation-awareness and roleplay confounds. The suite contains seven operational tasks, each with an official workflow and a policy-violating shortcut. An eight-variant shared framework varies monitoring, instruction clarity, stakes, permission, instrumental usefulness and blocked honest paths to support inferences regarding the factors driving IC behaviour. We evaluated ten models using deterministic environment-state scorers over 1,680 samples, with trace review employed for audit and adjudication purposes. The final IC rate is 86 out of 1,680 samples (5.1\%). IC behaviour is concentrated rather than uniform: two Gemini models account for 66.3\% of IC cases and three tasks account for 84.9\%. Conditions in which IC behaviour is indispensable for task success result in the greatest increase in the adjusted IC rate (+15.7 percentage points), whereas emphasising that task success is critical or certain framing choices do not produce comparable effects. Our findings indicate that realistic, low-nudge environments elicit IC behaviour rarely but systematically in most tested models. We conclude that it is feasible to robustly measure tendencies for dangerous behaviour in current frontier AI agents.


Integrating Strengths of Different Multi-Agent Workflows via Step-Aware Hybrid Topology Planning

Jingxuan Yu ⋅ Ju Jia ⋅ Yiqian Chen ⋅ Yuchong Chen ⋅ Cong Wu ⋅ Di Wu ⋅ Siqi Ma ⋅ Jie Gui

Large language model (LLM)-driven agents are proposed for a wide range of applications. As scenarios become increasingly composite, a graph-structured multi-agent system (MAS) offers a promising solution due to the orchestration of various skills. To architect a reasonable underlying workflow, recent advances mostly investigate two structures: sequential topology and decentralized topology. The former constrains the structure into a directed acyclic graph for multi-step executions but lacks sufficient scalability, while the latter leverages node-wise subgraphs for knowledge aggregation with synchronous parallelization but lacks long-term planning. Therefore, our key insight is that their strengths are mutually complementary. Driven by this motivation, we introduce step-aware hybrid topology for a MAS, where sequential connections support inter-step ordered planning and decentralized connections promote intra-step perspective integration. Concretely, to effectively coordinate different structural components, we clarify the corresponding topological definition and the protocols for step-wise serialization and parallelization. Subsequently, to dynamically refine desired structures, we introduce the parameterized distribution of a hybrid topology, which is initialized by LLM-driven graph planning. Eventually, to further reduce token consumption and time complexity, we supplement the topological density and length normalization for end-to-end learning. Extensive experiments demonstrate that the hybrid topology outperforms single-form topology in 90.4\% of tasks, our agentic initialization improves learning in 88.8\% of scenarios, proposed normalizations reduce token consumption by 20\%$\sim$30\% and time consumption by 7\%$\sim$13\% with a slight utility fluctuation. The code is available at https://anonymous.4open.science/r/HTAS-D107.

We propose the task of Interactive 4D Volumetric Liquid Forecasting: predicting the spatiotemporal evolution of a future dense 3D liquid velocity field conditioned on how a solid object moves through it. This task provides an upstream physical forecasting primitive for systems that need to reason about motion-conditioned liquid response, including robotic liquid manipulation, digital twins, and rigid-fluid video generation. Existing works have made efforts in predicting flow around static or prescribed solid geometries and synthesizing visually plausible fluid motion. Yet these approaches provide limited coverage for interactive liquid dynamics, as they are not capable of forecasting how a moving solid induces a future full-domain liquid response. In this forecasting scenario, the hydrodynamic effect is condensed near an evolving liquid-solid interface, but the prediction target is a global spatiotemporal rollout. To address this challenge, we propose a two-level recurrent forecasting framework, TIDE/TIDES: TIDE isolates localized interface-driven forecasting, while TIDES adds transport-aware coupling and explicit stabilization to reduce accumulated rollout error. Across adapted neighboring baselines, controlled component ablations, motion/size OOD tests, and long-horizon evaluation, TIDES achieves the best forecasting performance and the ablations identify localized interaction, transport coupling, and stabilization as the key contributors. Together, our task formulation, benchmark, and TIDE/TIDES study establish a controlled foundation for moving-solid-conditioned dense liquid forecasting. Code and benchmark will be made publicly available.


Interleaved Latent Thinking and Adaptive Termination for Efficient Reasoning LLMs

Zhao Jin ⋅ Rong-Cheng Tu ⋅ Wenhao Sun ⋅ Qi Guo ⋅ Mingyang Yu ⋅ Hao Guan ⋅ Dacheng Tao

Emerging latent reasoning paradigms allow Large Language Models to “think” in latent embedding spaces, offering a more efficient alternative to explicit Chain-of-Thought. However, they generally face two critical limitations: First, relying exclusively on latent tokens amplifies uncertainty due to their inherently high entropy, thereby degrading solution correctness. Second, the overthinking problem remains prevalent, and existing approaches typically rely on rigid heuristics for stopping, often resulting in suboptimal termination. To address these challenges, we view efficient reasoning as a learnable control problem over both how to think and when to stop. We formulate this within a Reinforcement Learning framework, training lightweight per-step gates on a frozen LLM. This modulation enables the model to (i) adaptively interleave latent soft-token reasoning with explicit token generation, and (ii) predict cumulative exit probabilities for robust early termination. Experiments on diverse reasoning benchmarks across varying model scales and families demonstrate that our method surpasses both standard CoT and existing latent baselines, achieving accuracy gains of ∼3% while reducing token consumption by up to 40%.

Although numerous algorithms have been proposed for unsupervised anomaly detection, there is currently no widely acknowledged internal evaluation method tailored specifically to this task. Existing internal evaluation methods typically assess score distributions alone or depend on auxiliary classifiers, limiting their reliability and interpretability. To address these problems, we propose SSSD (Similar Scores, Similar Data), an evaluation framework inspired by the smoothness assumption in machine learning—that similar inputs should yield similar outputs. We adapt this principle to anomaly detection by requiring that locally similar data points be assigned similar anomaly scores. Instead of testing this assumption directly, SSSD identifies violations of this principle by comparing the similarity of data with similar anomaly scores, interpreting a significant difference as a higher probability of erroneous detection. This evaluation of similarity is measured through nearby-score and nearby-distance, which is calculated based on the distances between neighboring data points and their corresponding anomaly scores. SSSD is highly interpretable, providing valuable insights for identifying mislabeled data. Experimental results show that our proposed method is both effective and robust for evaluating the performance of unsupervised anomaly detection algorithms across artificial and real-world datasets.


Invaria: Learning Scale and Density Invariance in Point Clouds via Next-Resolution Prediction

Chun-Peng Chang ⋅ Shaoxiang Wang ⋅ Alain Pagani ⋅ Dariu Gavrila ⋅ Holger Caesar

Modern image encoders achieve high generalization by decoupling semantic meaning from resolution, an ability yet to be fully realized in the 3D domain. We investigate the failure of 3D point cloud encoders to achieve similar generalization and find that existing models are highly sensitive to sampling resolution and scale changes, leading to significant performance degradation. This sensitivity is a major bottleneck for real-world deployment in robotics, as it suggests models overfit to specific quantization densities and object scales rather than learning invariant semantic features. To mitigate this dependency, we propose Invaria, a point cloud encoder that achieves scale and density invariance through next-resolution prediction and receptive field calibration. While our objective is not the explicit generation of high-resolution point clouds, we find that this training objective encourages the model to learn robust, structural invariants. The resulting encoder achieves significant performance gains during resolution shifts while maintaining high efficiency through a compact model size and reduced token requirements. Specifically, on ScanNet, Invaria achieves a 56.0\% higher mIoU at 3$\times$ lower resolution and a 20\% improvement when the objects scale is reduced by a factor of 3. These gains are achieved with a 45\% smaller model size and an average reduction of 40\% in input tokens. The code will be released once the paper is published.


IRIS: Interpolative Rényi Iterative Self-play for Large Language Model Fine-Tuning

Wenjie Liao ⋅ Like Wu ⋅ Liangjie Zhao ⋅ Shihui Xu ⋅ Shigeru Fujimura

Self-play fine-tuning enables large language models to improve beyond supervised fine-tuning without additional human annotations by contrasting annotated responses with self-generated ones. Many existing methods rely on a fixed divergence regime. SPIN is closely related to a KL-based regime, SPACE to a Jensen-Shannon-style objective via noise contrastive estimation, and SPIF to $\chi^2$-regularized self-play. Since these divergences exhibit different strengths depending on the distributional gap between model and target, no single choice appears to provide favorable learning dynamics across training stages. We propose IRIS (Interpolative R\'enyi Iterative Self-play), a R\'enyi-based self-play fine-tuning framework with a continuously adjustable objective. IRIS decomposes into two independent tilted risk terms over annotated and synthetic data, with exponential importance weights controlled by the order parameter $\alpha$. We show that several self-play objectives can be interpreted as limiting or representative regimes at particular values of $\alpha$, providing a unified theoretical perspective on these methods. An adaptive order schedule further adjusts $\alpha$ to the distributional gap, shifting from sharper importance weighting early in training to smoother refinement near convergence. Theoretically, we establish the fixed-point property of IRIS and analyze how $\alpha$ controls gradient concentration. Experiments on Zephyr-7B and Qwen2.5-3B across ten benchmarks show that IRIS improves upon baselines, reaching 44.57\% average score with gains across iterations. In our setting, IRIS with only 26$k$ annotated samples surpasses standard supervised fine-tuning trained on the full 200$k$ dataset.

Long-tailed out-of-distribution (LT-OOD) detection is often addressed with specialized training, including auxiliary out-of-distribution (OOD) data, abstention heads, contrastive objectives, energy losses, or gradient-conflict control. We show that these training mechanisms can obscure a simpler issue: frozen long-tailed representations may already contain useful OOD evidence, but raw Mahalanobis distance is distorted by frequency-coupled feature radius and poorly supported tail covariance. We propose \emph{Hyperspherical Pooled Mahalanobis} (\HPM), a post-hoc detector that normalizes features onto the unit sphere and replaces class-specific covariance with a pooled, ridge-regularized metric while keeping class means as semantic anchors. In CIFAR-LT experiments and an ImageNet-100-LT near-OOD boundary analysis, \HPM improves raw Mahalanobis scoring; for Prior-Calibrated ERM (PC-ERM), it raises AUROC from 46.49 to 85.67 on CIFAR-10-LT and from 50.40 to 78.35 on CIFAR-100-LT. This simple PC-ERM+\HPM pipeline also achieves the best Log Efficiency Score (LES; 3.08) on CIFAR-100-LT, retaining roughly 95\% of the best CIFAR-100-LT AUROC observed among the compared post-hoc scores at substantially lower training-time cost. These results argue for evaluating representation quality, detector geometry, and training complexity as separate factors in LT-OOD detection.


Isharah-Selfie: Continuous Sign Language Recognition Dataset for One-handed Signing

Ahmed A Hasanaath ⋅ Murtadha Aljubran ⋅ Sarah Alyami ⋅ Muhammad Haris Khan ⋅ Hamzah Luqman

Current sign language recognition and translation benchmarks are commonly collected in controlled settings using fixed or studio-style cameras, particularly limiting their ability to reflect how signers communicate in everyday mobile scenarios. As such, this gap becomes increasingly evident when considering the Arabic Sign Language (ArSL), which is a primary sign language used across Arab countries. In this work, we introduce Isharah-Selfie, a large-scale ArSL dataset collected entirely using front-facing smartphone cameras under unconstrained selfie conditions. The dataset comprises 8,997 video clips performed by 16 signers across 1,493 unique sentences, covering several domains such as healthcare, education, and transportation. Unlike conventional datasets dominated by fixed-camera, two-handed signing, Isharah-Selfie captures realistic mobile communication characterized by one-handed signing, close-range viewpoints, partial body visibility, camera motion, unconstrained environments, and heterogeneous device resolutions. Each video is annotated with both gloss sequences for continuous sign language recognition (CSLR) and Arabic sentence translations for sign language translation (SLT). Moreover, we define signer-independent and unseen-sentences evaluation protocols to assess generalization across unseen signers and novel sentence compositions. Extensive benchmarks using RGB-based and pose-based data show that current CSLR and SLT methods remain challenged by selfie-style signing conditions, particularly under unseen-sentence evaluation, where compositional generalization remains weak. By releasing Isharah-Selfie, we aim to support research toward robust, user-centered sign language technologies that operate reliably in realistic smartphone-based communication settings. The Isharah-Selfie dataset is available on https://www.github.com/.


Is Your LLM-as-a-Recommender Agent Trustable? LLMs' Recommendation is Easily Hacked by Biases (Preferences)

Zichen TANG ⋅ Zirui Zhang ⋅ Qian Wang ⋅ Zhenheng Tang ⋅ Xiaowen Chu ⋅ Bo Li

Current Large Language Models (LLMs) are gradually exploited in practically valuable agentic workflows such as Deep Research, E-commerce recommendation, job recruitment. In these applications, LLMs need to select some optimal solutions from massive candidates, which we term as \textit{LLM-as-a-Recommender} paradigm. However, the reliability of using LLM agents for recommendations is underexplored. In this work, we introduce a \textbf{Bias} \textbf{Rec}ommendation \textbf{Bench}mark (\textbf{BiasRecBench}) to highlight the critical vulnerability of such agents to biases in high-value real-world tasks. The benchmark includes three practical domains: paper review, e-commerce, and job recruitment. We construct a \textsc{Bias Synthesis Pipeline with Calibrated Quality Margins} that 1) synthesizes evaluation data by controlling the quality gap between optimal and sub-optimal options to provide a calibrated testbed to elicit the vulnerability to biases; 2) injects contextual biases that are logical and suitable for option contexts. Extensive experiments on both SOTA (Gemini-{2.5,3}-pro, GPT-4o, DeepSeek-R1) and small-scale LLMs reveal that agents frequently succumb to injected biases despite having sufficient reasoning capabilities to identify the ground truth. These findings expose a significant reliability bottleneck in current agentic workflows, calling for specialized alignment strategies for LLM-as-a-Recommender. The complete code and evaluation datasets will be made publicly available upon acceptance.


ITO: Multi-View Alignment and Training-Time Fusion for Image-Text Pretraining

Hanpeng Liu ⋅ Yaqian Li ⋅ Zidan Wang ⋅ Shuoxi Zhang ⋅ Zonglin Zhao ⋅ Zi-Hao Bo ⋅ rinyoichi takezoe ⋅ Kaiwen Long ⋅ Kun He

Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality rather than by semantics. We propose \textbf{ITO}, a framework addressing this limitation through two complementary mechanisms with distinct roles. \emph{Multimodal multiple alignment} enriches supervision by constructing diverse cross-modal correspondences from multi-view image augmentations, providing the primary source of discriminative gain. A lightweight \emph{training-time multimodal fusion} module then acts as a geometric regularizer, encouraging the encoders to produce features that are compatible under fusion and thereby reducing modality-induced separation. Crucially, the fusion module is discarded at inference, preserving the efficiency of standard dual-encoder architectures while incurring training-time overhead only. Extensive experiments across pretraining scales from millions to billions of image--text pairs show that ITO consistently outperforms strong baselines on classification, retrieval, and multimodal benchmarks. Our analysis further reveals that beyond accuracy gains, training-time fusion plays a stabilizing role in optimization, mitigating the late-stage overfitting commonly observed in aggressive contrastive learning.


JEDI: Real-Time Jailbreak Defense for LLMs via In-Generation Detection and Intervention

Ruilin Xie ⋅ Bixin Li ⋅ Xinyu Chen ⋅ Yongqiang Tian ⋅ Wang Lulu

Large Language Models remain vulnerable to jailbreak attacks despite extensive efforts to ensure safety alignment. Streaming generation scenarios exacerbate this vulnerability by exposing harmful tokens to users immediately upon generation, a phenomenon known as prefix exposure. Existing defenses often fail to address this real-time constraint or impose substantial latency that degrades user experience. To bridge this gap, we propose JEDI, a real-time defense method that secures streaming outputs via in-generation detection and intervention. JEDI leverages representation engineering to monitor the LLM’s internal semantic drift, using a Cumulative Sum algorithm to identify persistent, harmful intent before it manifests in the output. Upon detecting a risk, the system dynamically injects a steering vector to redirect the generation trajectory toward a safe subspace. Extensive evaluations across 6 distinct models and 11 attack vectors demonstrate that JEDI achieves defense success rates ranging from 93.6\% to 99.4\%, outperforming state-of-the-art baselines. Furthermore, JEDI preserves model utility with a Time To First Token overhead of approximately 0.003 seconds, validating its viability for latency-sensitive applications. Our code and experimental results are available at: https://anonymous.4open.science/r/JEDI-93F9


Jointly Reinforcing Diversity and Quality in Language Model Generations

Tianjian Li ⋅ Yiming Zhang ⋅ Ping Yu ⋅ Swarnadeep Saha ⋅ Daniel Khashabi ⋅ Jason Weston ⋅ Jack Lanchantin ⋅ Tianlu Wang

Post-training of Large Language Models (LMs) often prioritizes accuracy and helpfulness at the expense of diversity. This creates a tension: while post-training improves response quality, it also sharpens output distributions and reduces the range of ideas, limiting the usefulness of LMs in creative and exploratory tasks such as brainstorming, storytelling, or problem solving. We address this challenge with Diversity-Aware Reinforcement Learning (DARLING), a framework that jointly optimizes for response quality and semantic diversity. At its core, DARLING introduces a learned partition function to measure diversity beyond surface-level lexical variations. This diversity signal is then combined with a quality reward during online reinforcement learning, encouraging models to generate outputs that are both high-quality and distinct. Experiments across multiple model families and sizes show that DARLING generalizes to two regimes: non-verifiable tasks (instruction following and creative writing) and verifiable tasks (competition math). On five benchmarks in the first setting, DARLING consistently outperforms quality-only RL baselines, producing outputs that are simultaneously of higher quality and novelty. In the second setting, DARLING achieves higher pass@1 (solution quality) and pass@k (solution variety). Most strikingly, explicitly optimizing for diversity catalyzes exploration in online RL, which manifests itself as higher-quality responses.


Joint Optimization of Tool Creation and Use for Large Language Model Agents

Zhi Rui Tam ⋅ Chieh-Yen Lin ⋅ Yun-Nung (Vivian) Chen ⋅ Shao-Hua Sun ⋅ Hung-yi Lee

Tool-augmented language models are bounded by the APIs humans bothered to write; existing tool-creation systems patch this by prompting a frozen LLM at inference time, leaving the model that writes a tool decoupled from the one that uses it, with no signal that the schemas it produces are schemas it can invoke. We propose **SMITH** (Schema-grounded Multi-task Iterative Tool Honing), a reinforcement learning framework that jointly trains tool creation and tool use inside a single policy. Each rollout is either a **build** task (write a tool from a few examples) or a **use** task (invoke a pooled tool on a held-out question). Three separate reward axes catch schema, code, and outcome failures independently, so each failure mode contributes its own gradient. A 4B Qwen3 trained with SMITH on 13 procedural reasoning tasks with exact verifiers reaches $79.8$ macro-average accuracy on held-out tasks, the best across all evaluated methods and ahead of an untrained 30B-A3B tool-writer. It also reaches $40.4$ on TabMWP-Hard and $42.6$ on out-of-domain GQA ($+7.6$ over the best same-backbone inference-time baseline), without any visual or tabular training data. When invoked by a frozen 350M student, tools written by our 4B match those produced by a writer an order of magnitude larger. The same recipe also lifts Qwen3-8B and Granite-3.3-8B without modification.


Joint Sequence--Vocabulary Selection for Efficient LLM Distillation

Xueli Geng ⋅ Weicheng Zhao ⋅ Xutong Mu ⋅ Tianrui Wei ⋅ Yanbiao Ma ⋅ Yulong Shen

Selective knowledge distillation has become increasingly important for compressing large language models, where dense logit-based knowledge transfer over long sequences and large vocabularies incurs substantial computational overhead. Distillation supervision is naturally organized over a sequence–vocabulary matrix, with its highest-information components often concentrated on a small subset of specific sequence–vocabulary pairs. Existing methods typically select sequence positions or vocabulary candidates independently, relying on coarse-grained one-dimensional heuristics that weaken the precision and effectiveness of knowledge transfer. To address this, we propose $\textbf{J}$oint $\textbf{S}$equence–$\textbf{V}$ocabulary $\textbf{D}$istillation ($\textbf{JSVD}$), a plug-and-play selective distillation framework that identifies and focuses supervision on informative sequence–vocabulary pairs. For each training example, JSVD constructs bidirectional teacher--student candidate supports, scores sequence--vocabulary pairs using prediction discrepancies, and progressively focuses distillation on unresolved informative pairs during training. Extensive experiments across diverse distillation objectives, model pairs, and benchmarks show that JSVD consistently improves downstream performance while reducing distillation overhead and accelerating convergence.


JRDB-AVR: An Active Visual Reasoning Benchmark for Real-World Embodied Environments

Zhixi Cai ⋅ Fucai Ke ⋅ Sukai Huang ⋅ Maria Garcia de la Banda ⋅ Peter Stuckey ⋅ Reza Haffari ⋅ Hamid Rezatofighi

In complex embodied visual reasoning scenarios, an agent often has only a limited field of view, and the evidence needed to answer a question may be distributed across time, viewpoint, and interacting objects. A model may therefore give a plausible answer without ever observing the relevant object, time, or view that supports it. Current visual reasoning benchmarks largely evaluate passive observations and final answers, overlooking settings that require active perception and evidence acquisition. We introduce JRDB-AVR, a benchmark built from real-world human-scene robotics data that turns this gap into an explicit evaluation: a system receives a visual reasoning question, requests bounded observations by timestamp and viewing angle, and is evaluated on both the final answer and the grounded visual evidence supporting it. The benchmark contains diverse questions over multiple real-world test environments involving temporal search, viewpoint selection, and human-oriented compositional reasoning. We also introduce JRDB-AVR-Agent, a reference active reasoning method that maintains an explicit observation-grounded graph-based world model and answers through solving. Experiments reveal a substantial gap between answer accuracy and evidence accuracy in current baselines, showing that current VLMs can produce unsupported correct answers and that active evidence-aware evaluation is necessary for embodied visual reasoning.

Network interference complicates A/B testing on online platforms, such as social networks and marketplaces, where causal methods based on single experiments often suffer from significant bias due to complex interference patterns. This paper demonstrates the statistical benefits of merging data from multiple experiments with varying treatment proportions. Sequential experimentation with increasing traffic, or ramp-up, is widely used in tech companies for risk management and cost control. Beyond operational benefits, we show that regression-based estimators trained on merged data achieve substantial bias reduction, even under simple randomization schemes and regression models. We focus on the global average treatment effect (GATE), a key estimand in the tech industry, and consider a general interference pattern that extends beyond the 1-hop setting. We present an exact bias–variance analysis of the linear regression estimator and show that, in practical settings, the bias term typically dominates. In addition, we characterize the substantial bias reduction achieved by our proposed approach, which merges experimental data collected at different ramp-up stages to improve the training of the regression function. We also offer an intuitive explanation for this reduction and highlight the synergy between cluster-level randomization and our approach. Furthermore, we consider a refined estimator based on graph neural networks (GNN). Extensive simulations across novel and challenging scenarios confirm that our methodology significantly improves the accuracy of regression-based estimators.


KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Pengfei Li ⋅ Naufal Suryanto ⋅ Sicheng Zhang ⋅ Muhammad Muzammal Naseer

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs’ ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag-value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language-to-CLI translation on Kali Linux, comprising 8,504 query-command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and over 20 general-purpose and security-focused LLMs, no model exceeds 35\% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model. KaliBench will be publicly released.


KG-Guard: Graph-Based Hallucination Detection for Knowledge Base Question Answering

Albert Sawczyn ⋅ Piotr Bielak ⋅ Tomasz Kajdanowicz

Large language models (LLMs) are increasingly used for knowledge base question answering (KBQA), where answering requires selecting entities from a question-specific knowledge-graph subgraph. Yet LLMs are known to hallucinate across tasks, and KBQA is no exception: even when we provide a graph as the knowledge source, the model may rely on parametric knowledge instead of graph evidence or perform invalid reasoning over the given relations. Such hallucinated answer nodes can limit the practical deployment of KBQA systems, especially in high-stakes domains such as healthcare. We formulate hallucination detection in KBQA as an answer-node classification problem and propose a lightweight graph-based framework that treats the answering LLM as a black box. KG-Guard represents each KBQA instance as an augmented graph. It initializes node features with semantic representations of KG entities, marks topic entities and LLM-proposed answer nodes with learned vectors, and connect a virtual question node to the topic entities. A graph encoder then produces verification-oriented node representations, and a small MLP classifies each proposed answer node using its graph representation together with the question embedding. Experiments on WebQSP, ComplexWebQuestions, and PUGG show that our detector achieves the highest F1 on all three benchmarks ($82.0$, $87.4$, and $84.3$), outperforming LLM-judge and sampling-based baselines, while having $\sim305\times$ fewer parameters than the reference approaches. Beyond detection, the node-level feedback is actionable: when flagged answers are fed back to the KBQA system for iterative refinement, downstream KBQA F1 improves by $13.0$-$14.5$ points and Exact Match by $16.9$-$17.6$ points.

Modern LLMs continue to exhibit significant variance in behavior across languages, such as being able to recall factual information in some languages but not others. While typically studied as a problem to be mitigated, in this work, we propose leveraging this cross-lingual inconsistency as a tool for interpretability in mixture-of-experts (MoE) LLMs. Our knowledge localization framework contrasts routing for sets of languages where the model correctly recalls information from languages where it fails. This allows us to isolate model components that play a functional role in answering about a piece of knowledge. Our method proceeds in two stages: (1) querying the model with difficult factual questions across a diverse set of languages to generate "success" and "failure" activation buckets and then (2) applying a statistical contrastive analysis to the MoE router logits to identify experts important for knowledge. To validate the necessity of this small number of experts for answering a knowledge question, we deactivate them and re-ask the question. We find that despite only deactivating about 20 out of 6000 experts, the model no longer answers correctly in over 40% of cases. Generally, this method provides a realistic and scalable localization approach to increasingly complex LLMs, but also suggests redundant and highly dispersed knowledge parameterization in MoEs.


KuaiRecV2: Benchmarking Large-Scale Continual Learning for Diversified and Multi-task Recommendation

Chenxu Li ⋅ Shuchang Liu ⋅ Hantao Shu ⋅ Wei Yuan ⋅ Zhe Xu ⋅ Guoqing Hu ⋅ Yan Wang ⋅ Biao Yang ⋅ Xiang Li ⋅ Cheng Ling ⋅ Fan Yang ⋅ Yongqi Liu ⋅ Lantao Hu ⋅ Han Li ⋅ Kaiqiao Zhan ⋅ Kun Gai ⋅ An Zhang ⋅ Xiang Wang

Recommender systems, which retrieve items of interest to users based on their interaction histories, have recently witnessed several key technological advancements, benefiting from the semantic representation space and model-data scaling. Nevertheless, practitioners and researchers have frequently encountered critical inconsistencies between simple offline evaluations and complex online results, and this gap still exists in the era of large models. In this work, we address several key factors that contribute to this inconsistency: 1) the large-scale continual learning challenge that assumes a practical demand to maintain long-term model effectiveness under dynamic item pools and distributional shifts; 2) the combinatorial diversity challenge that requires the generation strategy to solve the accuracy-diversity trade-off in both the finer token level and the holistic item level; and 3) the multi-task reward balancing challenge, which is particularly critical for modern generative models that can accurately model the sequence distribution but are less controllable for the multi-task demands. To address these challenges and further pave the way for the development of more realistic recommendation solutions, we introduce KuaiRecV2, a deliberately constructed benchmark with industrial-level datasets from Kuaishou. Specifically, we provide a large-scale interaction dataset with million-scale multi-modal item information and long-term interaction records across 30 days, with a dynamic video candidate pool. Additionally, we provide a rigorous benchmark that evaluates retrieval accuracy, diversity, and multi-task performance in one rubric system, where an effective advancement occurs if and only if the new method improves all metrics. The code and data are available at Supplementary Materials.


LACE: Latent Alignment via Counterfactual Embeddings

Harsh Udai ⋅ Konda Reddy Mopuri

Improving compositional understanding in CLIP-style vision-language models remains difficult, models with strong coarse alignment often fail on attribute binding, relations, and multi-object semantics. A common remedy is text-side hard negatives (negations, perturbations, or mined confusions), but concentrating difficulty on the text branch can skew the training signal, exacerbate modality asymmetry, and distort joint-space geometry (e.g., increased hubness and reduced mutual reciprocity), even when Recall@K improves. We propose LACE (Latent Alignment via Counterfactual Embeddings), a lightweight approach that injects structured, semantics-preserving latent edits into the image and text embedding streams without pixel-space editing or external generators. LACE synthesizes counterfactual image and text embeddings by attenuating, swapping, or transplanting a small set of factors tied to target objects, attributes, or relations, producing image-side and text-side hard negatives that rebalance supervision while preserving global alignment. Across compositional and retrieval benchmarks, LACE improves robustness to binding and relational confusions and yields a healthier embedding geometry with improved reciprocity and competitive hubness/local-stability trade-offs.


LADDERS: Length-Aware Data Distribution and Existing-Response Speculation for Fast RL Rollout Generation

Shengpeng Yin ⋅ Hui-Ling Zhen ⋅ Xing Li ⋅ Mingxuan Yuan ⋅ Zishuo Wang ⋅ Yuxin Peng

Reinforcement learning (RL) has become central to post-training large language models, but its rollout generation stage is often the dominant system bottleneck. The inefficiency comes from two sources. First, autoregressive decoding makes latency grow with response length. Second, response lengths vary widely within a batch: short sequences finish early, but the batch remains blocked by the longest sequence, leaving GPU capacity underutilized. This effect is amplified in multi-sample rollouts, where long responses can also concentrate on a few decoding workers and delay synchronization. We propose LADDERS (Length-Aware Data Distribution and Existing-Response Speculation), a lightweight framework for accelerating RL rollout generation. LADDERS first uses hidden-state-based length prediction to group prompts with similar expected response lengths, and then applies an S-shaped allocation rule to balance multi-sample requests across workers. It further reuses responses already generated by the policy as draft continuations through prompt-specific suffix trees, enabling speculative decoding without an auxiliary draft model. Experiments on Qwen3-1.7B and Qwen3-8B show that LADDERS reduces rollout generation time by up to 56% while preserving final task performance, and can be integrated into existing RL systems with minimal engineering effort.


LaMo: Self-Supervised Latent Motion Priors for Physical Realism in Video Generation

Bo Jiang ⋅ Depu Meng ⋅ yihan hu ⋅ Yichen Xie ⋅ Tianshuo Xu ⋅ Wei Zhan

Modern video generators produce visually compelling clips but still struggle with physical and motion consistency, limiting their use as reliable world simulators. Existing remedies often rely on external simulators, teacher models, or curated physics-focused data. We explore a complementary self-supervised direction: extracting motion cues from the unlabeled videos already used to train video diffusion models. We propose LaMo, which formulates a latent motion prior over frame-to-frame latent changes conditioned on the current latent and prompt. This prior is exposed through two lightweight readouts: a macro motion drift used during training as a Motion Drift Loss, and a learned micro motion field used during sampling as Motion Prior Guidance. Both components are plug-and-play with existing video diffusion backbones, requiring no architectural or I/O changes. On VideoPhy and VideoPhy2, LaMo improves CogVideoX backbones and outperforms recent physics-aware baselines that use external supervision. On VBench, it preserves overall generation quality while improving motion-related dimensions. These results suggest that unlabeled video contains useful motion supervision for improving physical fidelity in modern video diffusion models.


LAPrune: Logits-Aligned Scoring Proxy for KV Pruning via Vector Quantization

Mingyang Yu ⋅ Rong-Cheng Tu ⋅ Yifu Ding ⋅ Hanqing Zhao ⋅ Yongcheng Jing ⋅ Xiao Luo ⋅ Dacheng Tao

Long-context inference in Large Language Models (LLMs) suffers from high latency because attention computation and memory traffic scale linearly during decoding. KV cache pruning reduces this cost by retaining only a subset of cached tokens, but its effectiveness depends on whether the pruning proxy can recover the tokens with the largest exact attention logits under the current query. Existing proxies suffer from ranking misalignment: static heuristics rely on historical statistics and thus ignore the current query, while hashing-based retrieval proxies rank candidates in a discrete metric space that does not coincide with attention's inner product. To avoid these problems, we propose LAPrune, a KV pruning framework that uses additive vector quantization to build an efficient attention-logit-aligned proxy. LAPrune represents each cached key as a sum of learned codewords and decomposes the query-key dot product into lookup-and-add operations over query-codeword inner products. This replaces high-dimensional dot products with lightweight scalar lookups while preserving fidelity to the original attention ranking. LAPrune requires no base-model finetuning and can be integrated into long-context decoding with minimal modification. Across long-context evaluations, LAPrune improves exact-logit Top-$k$ recovery and perplexity over hashing-based pruning proxies, achieving up to $2.73\times$ decoding speedup over FlashAttention-2 and $4.83\times$ end-to-end speedup over the vanilla baseline at 128K context. Our code is available at [Anonymous/LAPrune](https://anonymous.4open.science/r/LAPrune-VQ).


Large Language Models Enhanced Covariate-adjusted Response-adaptive Randomization Design

Yanping Li ⋅ Xinrui Ruan ⋅ Jingshen Wang ⋅ Waverly Wei

Covariate-adjusted response-adaptive randomization (CARA) designs can improve statistical efficiency and participant welfare in randomized experiments by learning from accrued data and dynamically updating treatment allocation. In practice, two issues limit their impact. Many experiments are sample-size constrained because enrollment is difficult and trials are expensive. At the same time, studies increasingly collect rich pre-treatment information, including unstructured data such as text, images, and clinical notes, yet most CARA designs rely on a small set of structured covariates when computing allocation probabilities. Recent advances in large-scale pre-trained large language models (LLMs) offer a new opportunity to extract useful signals from unstructured inputs and external knowledge. However, few-shot LLM predictions are not ordinary fixed fitted predictions because prompt demonstrations sampled from accrued trial data create data-dependent randomness and can induce correlation across predictions. If used naively to drive allocation updates, AI-generated signals may weaken CARA's ability to improve power and participant welfare. We propose a CARA design that integrates few-shot LLM-predicted outcomes through resampling-based aggregation, calibration, and effective residual variances to guide both adaptive treatment allocation and downstream treatment effect estimation. Theoretical results and simulation studies show that, when calibrated few-shot LLM auxiliary signals reduce residual variation, coupling an LLM-enhanced estimator with an adaptive allocation rule can improve statistical efficiency and participant welfare.


Large-Scale Pretraining unlocks Few-Shot Prediction for Relational Data

Rishabh Ranjan ⋅ Vignesh Kothapalli ⋅ Harshvardhan Agarwal ⋅ Charilaos Kanatsoulis ⋅ Roshan Upendra ⋅ Tom Palczewski ⋅ Carlos Guestrin ⋅ Jure Leskovec

The ability to learn a new task from a few examples has been crucial to the success of foundation models such as large language models (LLMs). However, current foundation models for structured data like tables and relational databases still require tens of thousands of labeled examples to perform well on new tasks. Here we show that when pretrained at scale with the right recipe, Relational Transformers (RTs) can make state-of-the-art (SoTA) predictions with only hundreds of labels. This capability arises from the following ingredients: First, we assemble THE JOIN, the largest pretraining corpus of relational data to date, comprising 6k forecasting tasks across 650 real-world databases from diverse domains. Second, we develop a pretraining recipe that combines (1) mixed context sizes for multi-scale learning, (2) multi-cell masking for dense supervision, and (3) a novel random- walk-based retriever to efficiently gather relevant context. Third, we identify context ensembling and context tuning as complementary axes for scaling test-time compute to further improve few-shot performance. Pretrained with our recipe on THE JOIN, RT achieves parity with LLM Agent + TabICLv2 and RDBLearn + TabICLv2 pipelines with 32–23×fewer labels respectively, and even surpasses the prior SoTA of full task-specific training, on average nMAE for RelBench regression tasks. Context ensembling and tuning improves this further by up to 3%, and 4% respectively. Our ablations highlight the importance of schema semantics, multi-cell masking, and random-walk retrieval. Overall, our work paves the way for developing relational foundation models with strong few-shot capabilities.

In this paper, we study last-iterate convergence in stochastic constrained convex--concave minimax optimization. A key difficulty is that the last iterates of vanilla stochastic extragradient (S-EG) and stochastic optimistic gradient descent--ascent (S-OGDA) can fail to converge in the presence of gradient noise, even for simple bilinear problems. To address this issue, we regularize the original convex--concave problem into a strongly convex--strongly concave one. Applying S-EG and S-OGDA to the regularized problem gives two simple single-loop first-order methods, which we call perturbed S-EG and perturbed S-OGDA. By carefully choosing the regularization parameter and balancing the resulting regularization bias, stochastic error, and optimization error, we prove that both methods achieve a last-iterate convergence rate of $\mathcal{O}(T^{-1/4+\varepsilon})$ for any $\varepsilon>0$ in terms of the primal--dual gap. This improves upon the best-known $\tilde{\mathcal{O}}(T^{-1/7})$ rate under comparable constrained settings. Moreover, for unconstrained problems, we establish almost sure convergence under a single-timescale stepsize scheme, whereas prior almost sure convergence results typically rely on two-timescale stepsizes and are limited to S-EG.

We study learning-augmented online portfolio selection (OPS), where an investor uses predictions to improve wealth while remaining protected against unreliable advice. In frictionless markets, we characterize the optimal tradeoff between robustness (wealth under arbitrary predictions) and consistency (wealth under perfect predictions) tied to the geometric mean, and show that the optimal tradeoff admits a tractable online characterization. We further quantify how performance interpolates between these endpoints under imperfect predictions by smoothness guarantees. Under transaction costs, we give exact robust-trading certificates and, to our knowledge, the first learning-augmented OPS impossibility under costs: for any positive cost, no GM-robust policy can be maximally consistent under exact predictions. Lastly, we show that a cost-aware greedy rule remains Pareto-optimal.

Mobile sensing for real-time anomaly detection couples routing, sampling, and stopping: dispatch determines which locations are observed, while observations update the evidence used for future dispatch. We study this problem for a fleet of unmanned aerial vehicles (UAVs), where each UAV observes only its current location and anomalies may occur at unknown locations and times. The goal is to detect anomalies quickly while controlling false alarms under decision-dependent partial observations, local mobility, and collision-avoidance constraints. Existing quickest-detection methods typically ignore mobility and collision constraints, whereas learning-based routing methods scale to large fleets but are not designed for statistically calibrated detection. We formulate route-wise monitoring as a partially observable Markov decision process (POMDP) and propose DispatchPPO, a deep reinforcement learning method for long-horizon detection-aware dispatch. Its policy architecture, DispatchNet, uses autoregressive decoding with dynamic feasibility masking to generate collision-free joint UAV moves without enumerating the exponential joint action space. We establish theoretical properties for persistent coverage, evidence-based focusing, and detection delay through a resource-allocation information-rate bound. Experiments show that DispatchPPO reduces detection delay relative to heuristic and optimization-based baselines, while transferring effectively to real-robot and wildfire monitoring case studies.


Learning Contextual Causal Dynamics for Robust Exploration in Reinforcement Learning

Jiaming Pu ⋅ Yibo Zhang ⋅ Xu Dong ⋅ Tielin Zhang ⋅ Dengpeng Xing

Efficient exploration in sparse-reward environments, where informative extrinsic feedback is scarce, remains a significant challenge in reinforcement learning. A key limitation of existing exploration strategies is that they often assume densely coupled environment dynamics, overlooking the underlying causal mechanisms of the environment, particularly the fact that causal relationships may vary across contexts. To address this issue, we propose contextual causal dynamics model-based intrinsic motivation (CIM), a causality-aware exploration framework that explicitly captures context-dependent sparse causal structures and derives intrinsic motivation signals, specifically causal action influence and curiosity. These intrinsic rewards encourage interventions on causally relevant state factors rather than merely promoting observation novelty-seeking behavior, thereby facilitating robust exploration and active skill acquisition. Moreover, by learning contextual causal mechanisms, CIM identifies and filters out task-irrelevant state variables, improving generalization under distribution shifts. Empirical results demonstrate that our method achieves superior sample efficiency, training robustness, and out-of-distribution generalization performance across multiple robotic manipulation tasks.


Learning Cost-Efficient Autoscaling for Latency-Constrained Disaggregated LLM Serving

Fu Luo ⋅ Binbin Chen ⋅ Qing Luo ⋅ Linhui Xu ⋅ Tieying Zhang ⋅ Zhenkun Wang

Prefill-Decode (PD) disaggregation improves LLM serving by isolating compute-bound prefilling and memory-bound decoding onto specialized GPU pools. However, dynamically rebalancing these pools is challenging because request arrival rates, prompt lengths, and generation lengths change over time. We formulate this autoscaling problem as latency-constrained cost minimization: the controller must minimize GPU usage while keeping P99 tail latency within TTFT and TPOT targets. Existing threshold-based heuristics require extensive parameter tuning, yet still tend to over-provision. To address this, we propose a reinforcement learning (RL) framework that learns a cost-efficient autoscaling policy. The policy uses a recurrent network to track multi-timescale workload trends and pod lifecycle states, enabling decisions that account for delayed pod readiness and delayed SLO feedback. It also restricts the action space with hard masks and anti-oscillation cooldowns, enforcing feasible scaling actions while reducing wasteful GPU transitions. We train the policy in a high-fidelity PD-disaggregated serving simulator calibrated from profiling data and validated against a physical testbed. Simulator experiments on ShareGPT and Azure Conversational traces show that the learned policy satisfies the same latency constraints as grid-searched heuristic baselines while reducing average GPU occupancy. Real-cluster deployment on ShareGPT further confirms sim-to-real transfer while satisfying the latency targets.


Learning Energy-Based Models from Stochastic Interpolants using Spatiotemporal Differences

Hanlin Yu ⋅ RuiKang OuYang ⋅ Partha Kaushik ⋅ Arto Klami ⋅ Michael Gutmann ⋅ Omar Chehab

Learning an energy-based model from data samples is a central problem in machine learning. Many recent and popular methods, such as denoising score matching for training energy-based diffusion models, use stochastic interpolants to corrupt data samples at different noise levels indexed by a time variable. This defines a joint density over both the data space and time, and most methods learn its energy through either spatial or temporal differences. We identify distinct failure modes for both of these approaches. To solve them, we propose Spatiotemporal Noise-Contrastive Estimation (stNCE), a framework for learning the energy through joint spatiotemporal differences. stNCE unifies many existing methods and leads to new training objectives. Experiments on images and molecules demonstrate performance competitive with state-of-the-art density estimation methods.


Learning From Failures: Efficient Reinforcement Learning Control with Episodic Memory

Chenyang Miao ⋅ Bingchen Zhou ⋅ Yu Jianliang ⋅ Zeyang Liu ⋅ Xingyu Chen ⋅ Qiongjie Cui ⋅ Xuguang Lan

Reinforcement learning has achieved remarkable success in robot learning. However, under challenging exploration and contact-rich dynamics, early-stage training is frequently dominated by premature terminations such as collisions and falls. As a result, learning is overwhelmed by short-horizon, low-return trajectories, which hinder convergence and limit long-horizon exploration. To alleviate this issue, we propose a technique called Failure Episodic Memory Alert (FEMA). FEMA explicitly stores short-horizon failure experiences through an episodic memory module. During interactions, it retrieves similar failure experiences and prevents the robot from recurrently relapsing into unstable states, guiding the policy toward long-horizon trajectories with greater long-term value. FEMA can be combined easily with model-free reinforcement learning algorithms, and yields a substantial sample-efficiency improvement of 33.11\% on MuJoCo tasks across several classical RL algorithms. Furthermore, integrating FEMA into a parallelized PPO training pipeline demonstrates its effectiveness on a real-world bipedal robot task.


Learning from the Self-future: On-policy Self-distillation for dLLMs

Yifu Luo ⋅ Zeyu Chen ⋅ Haoyu Wang ⋅ Xinhao Hu ⋅ Yuxuan Zhang ⋅ Zhizhou Sha ⋅ Shiwei Liu

On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs), yet its application to diffusion LLMs (dLLMs) remains unexplored. Existing OPSD methods are inherently autoregressive-centric. They inject privileged information via left-to-right prefix conditioning with token-level divergence supervision, a design that fundamentally conflicts with the arbitrary-order generation of dLLMs. We introduce d-OPSD, the first OPSD framework tailored for dLLMs. Our approach makes two core contributions. First, we reframe self-teacher construction by using self-generated answers as suffix conditioning, enabling the student model to learn from ``self future-experience" rather than privileged prefixes. Second, we shift supervision from token-level to step-level, aligning training with the iterative denoising process of dLLMs. Experiments across four reasoning benchmarks show that d-OPSD consistently outperforms RLVR and SFT baselines with superior sample efficiency, requiring less than $10\%$ of the optimization steps by RLVR and opening a promising pathway for dLLM post-training.

Diffusion models have recently emerged as a powerful paradigm for generating 3D molecules. Despite showing remarkable promise, they learn atomic distributions implicitly from data, without explicitly encoding physicochemical rules into their generative dynamics. Consequently, generated samples can be statistically plausible yet chemically invalid. To address this problem, we propose Sequential Neuro-Symbolic Constrained Diffusion (SensDiff), a framework that integrates symbolic rules into generative modeling, enabling the model to internalize scientific priors during training. Central to SensDiff is the observation that the diffusion trajectory of 3D molecule generation exhibits a coarse-to-fine evolution: valid molecules arise by first mitigating macroscopic geometric conflicts and then refining structures toward microscopic chemical consistency through iterative denoising. Accordingly, SensDiff mirrors this dynamic by sequentially enforcing constraints with adaptive weighting during training, progressively steering generation toward valid molecular geometries. Moreover, SensDiff translates domain knowledge into interpretable generative controls, grounding opaque neural denoising in scientific principles. Experiments on molecular benchmarks and real-world constrained generation tasks confirm that SensDiff consistently improves chemical validity while adhering to physicochemical rules and user-specified targets.


Learning Hierarchical Patch Splitting Policies for Faster Vision Transformers

Aditya Gupta ⋅ Jean S Dandurand ⋅ Kai Qiu ⋅ Rohan Choudhury ⋅ László Jeni

Vision Transformers (ViTs) are bottlenecked by the length of their input token sequence. Standard patchification fixes this length with a single patch size applied across the image, regardless of regional content. Prior work on adaptive patch sizing reduces token count by assigning coarser patches to visually simple regions, but relies on noisy pixel-level heuristics such as local entropy to make patchification decisions. We argue that this compute allocation should be grounded in the model’s own semantic representations, not in raw pixel statistics. We introduce Split Policy Learned Image Tokenization which learns to allocate token resolution based on what the model finds meaningful, producing shorter token sequences that are better matched to task-relevant content. SPLIT matches the accuracy of full-resolution ViT classification baselines while using 30% fewer tokens. The same gains hold dense prediction: SPLIT achieves competitive accuracy on COCO and ADE20K, demonstrating that semantically grounded patchification generalises well beyond classification.


Learning Menu-Based Mechanisms for Truthful Budget-Feasible Procurement

Peng Chen ⋅ Xiang Liu ⋅ Hau Chan ⋅ Yan Lyu ⋅ Xueyong Xu ⋅ Weiwei Wu

Budget-feasible mechanisms (BFMs) are a fundamental tool for budget-constrained procurement, but designing strong BFMs remains challenging, particularly for general valuation classes. Classical approaches rely on problem-specific analysis and worst-case approximations, often yielding conservative mechanisms in practice. This motivates data-driven methods for discovering mechanisms with stronger empirical performance. However, strict budget constraints expose a key limitation of existing neural automated mechanism design (AMD): they either fail to provide exact economic guarantees or suffer from severe scalability bottlenecks. We propose BFMNet, a scalable menu-based neural framework for learning truthful budget-feasible procurement mechanisms. BFMNet first learns *self-bid-independent* menus from market context, ensuring dominant-strategy incentive compatibility (DSIC) and individual rationality (IR) by construction, and then applies an ex-post, instance-level menu transformation to enforce budget feasibility. We prove that the resulting mechanism satisfies exact DSIC, IR, and budget feasibility. By avoiding intractable global discretization grids and restrictive Lipschitz constraints, BFMNet remains both expressive and scalable. Extensive experiments show that BFMNet improves utility over classical baselines by up to $34.1$%, remains competitive with state-of-the-art neural baselines that provide only approximate economic guarantees, and scales robustly to larger auction instances.


Learning Optimal Transport Plans Via Autoregressive Token Regression

Ivan Jaen Marquez ⋅ Takis Chytas ⋅ Vikas Singh

Optimal Transport (OT) provides a principled framework for comparing probability distributions. Classical solvers remain the gold standard for this task, providing exact solutions. At the same time, OT is an ideal testbed for studying whether learning-based amortized models can internalize the structure of OT type problems. Here, ground-truth solutions are available, constraints are explicit, and exact discrete solutions have well-characterized sparsity. We describe ToR-OT, a transformer-based model that learns to predict discrete OT plans through autoregressive token regression. We reframe OT into a structured sequence generation task. ToR-OT decodes sparse transport plans as sequences rather than solving an optimization task at test time. We focus on three properties: (a) Zero-shot generalization: pretraining on a large synthetic corpus of OT instances allows a single model to generalize to unseen distribution pairs and heterogeneous cost metrics without retraining. (b) Sparse primal decoding: by exploiting the inherent sparsity in exact discrete OT solutions, the decoder predicts only nonzero entries (avoids dense coupling outputs). (c) Amortized inference: on low-dimensional discrete problems, ToR-OT gives accurate approximate plans and compares favorably with per-instance neural OT baselines. Overall, ToR-OT is not yet a replacement for classical solvers, but best viewed as a step toward understanding how autoregressive transformers can learn structure of constrained optimization problems like OT.


Learning Semantic Consistency for Open-Vocabulary Dense Perception

Li Ding ⋅ Junjie Wang ⋅ Jingjun Yang ⋅ Jiaze Wang ⋅ Libo Qin ⋅ Li Jiang ⋅ Zhuotao Tian

Open-vocabulary dense perception requires grounding language-specified concepts in local regions for recognition, localization, and segmentation beyond closed category sets. CLIP learns image-level vision-language alignment, whereas dense prediction requires spatially precise and semantically reliable local features. Although recent dense CLIP adaptation methods reduce this gap through region-level semantic transfer and spatial correlation guidance, maintaining stable local semantics under contextual variation remains challenging. Specifically, the global contextual modeling in CLIP may introduce undesirable dependencies into local features. Changes in background, scale, viewpoint, or scene layout may cause identical object regions to yield inconsistent dense representations, leading to unstable correspondences and semantic drift. Moreover, spatial correlation guidance from visual foundation models captures visual relatedness between patches, but not whether coherent regions correspond to the queried textual concept, leaving concept-level dense alignment under-constrained. To address these issues, we introduce Semantic Spatial Consistency Learning $\textbf{(SSCL)}$, an adaptation method with two complementary objectives: 1) Context-Invariant Spatiotemporal Consistency (CISC) enforces semantic consistency for the same anchor region under varying contexts to mitigate context-induced feature drift; 2) Structure-Guided Semantic Refinement (SGSR) leverages spatial affinities from a frozen visual foundation model to refine CLIP text-patch responses, producing text-conditioned dense supervision with improved spatial coherence and concept-level alignment. Experiments on video instance segmentation, region classification, semantic segmentation, and object detection show consistent gains over the baseline methods, demonstrating the effectiveness and generalization capability of the proposed SSCL. Our code and models will be made publicly available.


Learning the Committor Function using Weighted Ensemble Simulations

Jacky Chen ⋅ Rishal Aggarwal ⋅ David Koes

Understanding how complex molecular systems transition between metastable states is a central challenge in computational chemistry and biophysics. A key tool for characterizing these transitions is the committor function, defined as the probability that a trajectory initiated from a given configuration reaches one state before another, which encodes detailed mechanistic information about transition pathways. First, we curate a public benchmark of empirically determined committor values for alanine dipeptide. Second, we develop a "committor-in-the-loop" framework that combines a self consistent loss function with weighted ensemble (WE) simulation: WE adaptive bins are biased by the learned committor, which is in turn trained on WE samples. Across a two-channel 2D system and alanine dipeptide settings, our WE-based framework consistently outperforms greedy baselines, demonstrating the importance of principled exploration. Together, the benchmark and framework provide both a standard for evaluating committor learning methods and a new method for studying rare transitions in high-dimensional molecular systems.


Learning the Context of Errors: Black-Box Online Adaptation of Time Series Foundation Models

Xilin Dai ⋅ Yiding Liu ⋅ Hongjie Xia ⋅ Yifan Hu ⋅ Zewei Dong ⋅ Jiang-Ming Yang ⋅ Qiang Xu

The rapid evolution of Time Series Foundation Models (TSFMs) has advanced zero-shot forecasting across diverse domains. Inspired by the current form of Large Language Models, future TSFMs may be offered as commercialized, closed-source API services. However, many existing online adaptation methods still rely on white-box access for parameter fine-tuning or gradient backpropagation. This paradigm mismatch raises a question: In black-box online adaptation for TSFMs, what should we learn? We answer this with an insight: the predictive errors of the base model are conditioned on both the input and output of the base model (i.e., the context of errors). To validate this insight, we propose ORCA (Online Residual Contextual Adaptation). We conduct extensive experiments across 5 state-of-the-art TSFMs and 8 datasets to demonstrate the effectiveness of our approach. Furthermore, through ablation studies, we quantitatively analyze the impact of different adapter learning hypotheses on the final adaptation performance in black-box online adaptation. Code available at https://anonymous.4open.science/r/ORCA1/.


Learning to Ask: Metacognitive Action Policy for Large and Small Language Model Collaboration

Jiayuan Zhang ⋅ Jianwei Niu ⋅ Xuefeng Liu ⋅ Haotian Yang ⋅ Guogang Zhu ⋅ Yuhui Niu ⋅ Wanyu Lin ⋅ Xinghao Wu

Edge-cloud collaboration between large language models (LLMs) and small language models (SLMs) offers a promising way to leverage the problem-solving capability of LLMs while keeping private data on device. Existing methods typically have the LLM generate guidance from an initial user query, which the on-device SLM combines with local private data to produce the final response. However, the LLM's general guidance often does not align with the user's personalized needs, limiting the utilization of the LLM's capabilities. It has been observed that the on-device SLM exhibits metacognitive capabilities, which allow it to recognize its limitations and adjust its reasoning process. Motivated by this, we transform the on-device SLM into a metacognitive inquirer agent that can assess its internal state and actively decide when, what, and how to query the LLM for personalized guidance. Specifically, we propose Active Inquiry via Metacognitive Actions (AIMA). AIMA learns an on-device policy that jointly optimizes the selection of high-level metacognitive actions and the action-conditioned query refinement. We further introduce a privacy-constrained rewriting mechanism that detects and eliminates sensitive information leakage in the refined query. Extensive experiments and theoretical analysis demonstrate that our method outperforms existing methods.

Intracortical brain–computer interfaces (iBCIs) can restore movement and communication, but neural drift can degrade performance across recording sessions. Because recalibration adds user burden, there is a need for decoders that adapt from minimal calibration data. Although session-specific input layers are a common way to handle cross-session drift, they can underperform in the extreme few-shot regime because the shared decoder is trained downstream of well-estimated session-specific layers, but must rely at deployment on a new-session layer fit from only a few calibration trials. To avoid this dependence, we propose a two-stage approach: first, train a single GRU decoder on pooled held-in sessions without session-specific input layers, encouraging the backbone to handle cross-session variability directly; second, freeze this backbone and meta-learn a lightweight alignment layer for rapid calibration. We evaluate on FALCON H2, an official human handwriting iBCI benchmark designed for few-shot cross-session decoding, and achieve 7.41\% word error rate using only three released calibration sentences per held-out session. These results suggest that, under extreme calibration limits, learning a transferable backbone before adding session-specific adaptation can be more effective than jointly training flexible session-specific layers.


Learning Where to Look: Observation Policy Optimization for Thinking with Images

Junfeng Wang ⋅ Jiawei Liu ⋅ Yongchao Xu ⋅ Tao Jiang ⋅ jiangbo Ai ⋅ Jin Zhang ⋅ Chong Wang ⋅ Qixing Zhang

Complex visual question answering often requires vision-language models to think with images, i.e., actively acquire local visual evidence by zooming into informative regions of high-resolution images. Our error analysis shows that many failures in this setting originate from incorrect observation actions, where the model fails to localize key regions or obtain crops that contain sufficient evidence for answering. We further observe that bounding-box tokens exhibit substantially higher generation entropy than ordinary language tokens, indicating high uncertainty when the model decides where to look. However, existing reinforcement learning methods based on final-answer rewards provide only trajectory-level feedback, making it difficult to assign precise credit to the observation actions. We propose Observation Policy Optimization (OPO), an RL algorithm for optimizing observation actions in tool-augmented thinking with images. OPO consists of two core modules. First, Uncertainty-Guided Observation Branching (UOB) treats each image zoom-in tool call as an observation event and selectively branches at high-uncertainty, low-redundancy events during rollout, enabling efficient exploration of alternative cropping decisions and their resulting visual evidence. Second, Evidence-Localized Advantage Attribution (ELAA) compares sibling branches using local evidence-sufficiency rewards and assigns credit only to the corresponding observation-action tokens. By decoupling observation-action optimization from trajectory-level outcome supervision, OPO improves visual evidence acquisition and further enhances visual reasoning performance. Experiments show that OPO consistently improves Qwen3-VL models across scales and achieves competitive performance against larger models and tool-augmented baselines.

We study the problem of learning with multiple correct answers, where each instance admits a set of valid labels. We primarily focus on the online setup, where in each round the learner must output a valid label for the queried example. This setting is motivated by language generation, in which a prompt may admit many acceptable completions, but not every completion is acceptable. We study this problem under three feedback models. For each model, we characterize the optimal mistake bound in the realizable setting using an appropriate combinatorial dimension. We then show that the rate of regret can be constant, linear, or sublinear across the three models in the agnostic setting. Our results also imply sample complexity bounds for the batch setup that depend on the respective combinatorial dimensions.


Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding

Yan Zhang ⋅ Daiqing Wu ⋅ Huawen Shen ⋅ Can Ma ⋅ Yu Zhou

Graphical User Interface (GUI) grounding maps natural language instructions to the visual coordinates of target elements and serves as a core capability for autonomous GUI agents. Recent reinforcement learning methods (e.g., GRPO) have achieved strong performance, but they rely on expensive multiple rollouts and suffer from sparse signals on hard samples. These limitations make on-policy self-distillation (OPSD), which provides dense token-level supervision from a single rollout, a promising alternative. However, its applicability to GUI grounding remains unexplored. In this paper, we present GUI-SD, the first OPSD framework tailored for GUI grounding. First, it constructs a visually enriched privileged context for the teacher using a target bounding box and a Gaussian soft mask, providing informative guidance without leaking exact coordinates. Second, it employs entropy-guided distillation, which adaptively weights tokens based on digit significance and teacher confidence, concentrating optimization on the most impactful and reliable positions. Extensive experiments on six representative GUI grounding benchmarks show that GUI-SD consistently outperforms GRPO-based methods and naive OPSD in both accuracy and training efficiency. Code and training data will be publicly released.


Lemon: Evidence-Risk-Aware Adaptive Organization for Long-Horizon LLM Agents

Zimo Yin ⋅ Haipeng Jiang ⋅ Kailong Ren ⋅ Zhetao Sun ⋅ Guangyi Lv ⋅ Congli Yin ⋅ Likang Wu ⋅ Hongke Zhao ⋅ Ming He ⋅ Peng Wang ⋅ Jianping Fan

Large language model agents now combine tools, long context, memory, and multi-step workflows, yet they still fail when execution loses track of the evidence required for a reliable answer. Common failure modes include missing support, unresolved contradictions, and evidence lost under context pressure. We present Lemon, an evidence-risk-aware framework that makes such risks an explicit control variable. At the intra-instance level, a single state-conditioned controller jointly schedules reasoning, tool invocation, worker expansion, verification, context compression, and memory operations. A recoverable evidence substrate pairs anchored context compression with reusable semantic memory, preserving raw tool outputs while storing transferable process fragments. We further extend Lemon to an inter-instance setting in which heterogeneous personalized agents expose expertise, exchange evidence-bearing proposals, critique unsupported claims, and internalize useful collaboration artifacts. Empirically, Lemon reaches 91.36\% accuracy on GAIA and 80\% on xbench-DeepSearch. In an open-source reproducibility comparison on GAIA, Lemon uses 49.9\% to 88.4\% fewer tokens per task than three top-ranked agent baselines.


Less is More: Compact-Token Masked Feature Learning for Skeleton Representation Learning

Jeonghyeok Do ⋅ Yun Chen ⋅ Geunhyuk Youk ⋅ Munchurl Kim

Current skeleton representation learning paradigms face distinct limitations: Contrastive Learning (CL) often overlooks fine-grained motion details, while Masked Auto-Encoders (MAE) rely on coordinate-level reconstruction. This reconstruction inherently demands dense token sequences and heavy decoders, wasting pre-training computation on discarded components and forcing downstream inference to process dense token grids. To resolve these bottlenecks, we propose SLiM (Skeleton Less is More), a compact-token framework that unifies masked feature prediction and contrastive learning via a shared encoder. By shifting the objective from raw coordinate reconstruction to decoder-free, teacher-guided feature prediction, SLiM breaks the reliance on dense tokenization and enables effective learning with a highly compact token grid. Crucially, to prevent trivial shortcut learning arising from strong inter-joint dependencies of human, we introduce Semantic Tube Masking together with Skeleton-Aware Augmentations to enforce deep skeletal-temporal reasoning and anatomical consistency. Extensive experiments across multiple downstream protocols demonstrate that SLiM achieves state-of-the-art performance while structurally reducing inference computation by 7.89$\times$ compared to dense-token MAE baselines.


Leveraging Data Symmetries to Select an Optimal Subset of Training Data under Label Noise

Kumar Shubham ⋅ Pavan Karjol ⋅ Kiran M K ⋅ Prathosh AP

The performance of machine learning models often relies on large labeled datasets; however, data collected from diverse sources can contain label noise. Recent work has shown that, in noisy settings, there may exist a subset of the training data on which models can achieve performance comparable to training on a noise-free dataset. A widely used method for identifying such subsets is \textit{cutstats}, which employs $k$-nearest neighbors ($k$-NN) to detect low-noise samples. However, its performance on high-dimensional data remains largely unexplored. In this work, we formally establish that the performance of a classifier trained on a subset of a noisy dataset selected via \textit{cutstats} is influenced by the accuracy of $k$-NN. We further demonstrate that, in noisy environments, exploiting data invariance and knowledge of underlying symmetries can significantly enhance the performance of $k$-NN, bringing it closer to the Bayes optimal classifier even in high-dimensional regimes. Finally, we show that for real-world scenarios, where information about the underlying invariance is only partially known, learnt invariant representations can still facilitate the identification of near-optimal subsets.


L-FAME: Longitudinal Focused Attention Meditation EEG Dataset and Benchmark

Angqi Li ⋅ Basit R Syed ⋅ Hamzeh Alzweri ⋅ Taosheng Liu ⋅ Barry H Cohen ⋅ Saiprasad Ravishankar

We introduce a novel Longitudinal Focused Attention Meditation Electroencephalography (L-FAME) dataset and an accompanying benchmark, designed to foster research into the neural effects of various meditation practices and the evolution of these effects over a six-week training period. The dataset contains EEG recordings and psychological assessments from 74 healthy college participants, collected at two distinct time points: pre-intervention and post-intervention. Participants were randomly assigned to one of three distinct meditation groups: two mantra-based techniques (SA-TA-NA-MA and Hare Krishna) and one Breath Focus practice. Leveraging this unique longitudinal and comparative dataset, we propose a benchmark suite comprising three distinct classification tasks: (1) cognitive state decoding to distinguish between resting and meditation states, (2) fine-grained classification of the specific meditation techniques, and (3) cross-session adaptation to evaluate model generalization across the longitudinal time gap. We provide comprehensive baseline results for these tasks utilizing a range of classical machine learning algorithms and deep learning architectures. The complete dataset, preprocessing pipelines, and benchmark evaluation code will be publicly released, offering a valuable resource and a standardized framework for the development and comparison of new analytical methods in computational meditation research and EEG-based machine learning.


LiSA: Lifelong Safety Adaptation via Conservative Policy Induction

Minbeom Kim ⋅ Lesly Miculicich ⋅ Bhavana Dalvi Mishra ⋅ Mihir Parmar ⋅ Phillip Wallis ⋅ Bharath Chandrasekhar ⋅ Kyomin Jung ⋅ Tomas Pfister ⋅ Long T. Le

As AI agents move from chat interfaces to systems that read private data, call tools, and execute multi-step workflows, guardrails become a last line of defense against concrete deployment harms. In these settings, guardrail failures are no longer merely answer-quality errors: they can leak secrets, authorize unsafe actions, or block legitimate work. The hardest failures are often contextual: whether an action is acceptable depends on local privacy norms, organizational policies, and user expectations that resist pre-deployment specification. This creates a practical gap: guardrails must adapt to their own operating environments, yet deployment feedback is typically limited to sparse, noisy user-reported failures, and repeated fine-tuning is often impractical. To address this gap, we propose LiSA (Lifelong Safety Adaptation), a conservative policy induction framework that improves a fixed base guardrail through structured memory. LiSA converts occasional failures into reusable policy abstractions so that sparse reports can generalize beyond individual cases, adds conflict-aware local rules to prevent overgeneralization in mixed-label contexts, and applies evidence-aware confidence gating via a posterior lower bound, so that memory reuse scales with accumulated evidence rather than empirical accuracy alone. Across PrivacyLens+, ConFaide+, and AgentHarm, LiSA consistently outperforms strong memory-based baselines under sparse feedback, remains robust under noisy user feedback even at 20\% label-flip rates, and pushes the latency--performance frontier beyond backbone model scaling. Ultimately, LiSA offers a practical path to secure AI agents against the unpredictable long tail of real-world edge risks.


LiteNav: Lightweight Map-free Outdoor Visual Navigation

Mingxu Zhu ⋅ Chang Liu ⋅ Xiao Zhao ⋅ Zheyuan Zhang ⋅ Linna Song ⋅ Qingliang Luo ⋅ liming zhang ⋅ Xiangyu Zhen ⋅ Chufan Guo ⋅ Kuifeng Su

Outdoor robot navigation in dynamic, unknown environments remains a formidable challenge, necessitating robust traversability assessment and collision-free planning. While traditional map-based approaches struggle with adaptability to novel scenes, existing map-free methods frequently incur high computational costs and depend heavily on LiDAR data. In response, we propose LiteNav, a vision-only for exteroceptive sensing and strictly map-free outdoor navigation system. Relying solely on a single off-the-shelf RGB camera and GPS for goal specification, LiteNav estimates traversability, encodes navigation goals, and generates candidate trajectories via a diffusion-based joint encoding model. These trajectories are then refined by a goal-oriented planner to ensure correctness, efficiency, and safety. Remarkably, LiteNav achieves performance comparable to LiDAR-equipped methods without ever constructing or referencing explicit maps. Deployed on a low-power NVIDIA Jetson Xavier NX, it operates in real time, consuming only one-fifth of the computational resources required by prior approaches. Experiments demonstrate that LiteNav achieves state-of-the-art results on the GND dataset and outperforms existing methods in real-world scenarios. In summary, LiteNav is a map-free outdoor navigation system that operates using monocular RGB perception and GPS goal cues on an embedded platform, demonstrating strong potential for scalable deployment.

Decision trees are prized for their interpretability and strong performance on tabular data, but popular greedy top-down induction algorithms can yield suboptimal and unnecessarily complex structures. Optimal decision tree methods address this through global optimization, yet remain restricted to axis-aligned threshold splits, which limit the expressivity of each node and often force deep, complex trees to capture non-linear feature effects. Shape Generalized Trees (SGTs) generalize threshold splits to learnable univariate shape functions, improving expressivity and enabling more compact trees. However, existing SGT induction algorithms are greedy and offer no optimality guarantees. In this work, we introduce Literati, the first algorithm for optimal SGT induction. We propose a novel AND/OR graph formulation of the problem that jointly optimizes tree structure and shape function complexity. To solve this AND/OR graph, we develop an AO*-based algorithm with two enhancements that improve anytime performance while preserving optimality: a secondary heuristic for OR-node selection and a round-robin policy for AND-node exploration. Across 24 real-world datasets, Literati achieves higher training and test accuracy than state-of-the-art tree approaches.


LLM-enabled Applications Require Systematic Threat Monitoring

Yedi Zhang ⋅ Haoyu Wang ⋅ XIANGLIN YANG ⋅ Jin Song Dong ⋅ Jun Sun

LLM-enabled applications are rapidly reshaping the software ecosystem by using large language models as core reasoning components for complex task execution. This paradigm shift, however, introduces fundamentally new reliability challenges and significantly expands the security attack surface, due to the non-deterministic, learning-driven, and difficult-to-verify nature of LLM behavior. In light of these emerging and unavoidable safety challenges, we argue that such risks should be treated as expected operational conditions rather than exceptional events, necessitating a dedicated incident-response perspective. Consequently, the primary barrier to trustworthy deployment is not further improving model capability but establishing systematic threat monitoring mechanisms that can detect and contextualize security-relevant anomalies after deployment---an aspect largely underexplored beyond testing or guardrail-based defenses. Accordingly, this position paper advocates systematic and comprehensive monitoring of security threats in LLM-enabled applications as a prerequisite for reliable operation and a foundation for dedicated incident-response frameworks.

Sign Language Video Generation (SLVG) aims to generate realistic and motion-accurate sign language videos from spoken text. Most existing studies heavily rely on intermediate representations (\eg poses or 3D meshes), and adopt a two-stage generation paradigm, where the text is first converted into poses or other visual modalities, which are then used as conditions to synthesize sign language video frames via diffusion models (\eg pose-to-image models). However, the end-to-end SLVG task remains largely unexplored. In this paper, we demonstrate that end-to-end SLVG is feasible without relying on intermediate representations(\eg poses), and we argue that a key challenge lies in how to generate effective and temporally consistent conditions. With well-designed condition modules, it becomes feasible to build an end-to-end SLVG model without relying on any intermediate modalities. To this end, we propose LLMCond, a first end-to-end text-to-video SLVG framework that requires no auxiliary visual cues during either training or inference. LLMCond consists of four key components: (1) a video VQVAE that compresses sign language videos into a sequence of discrete latent tokens; (2) \textbf{an LLM-based conditioner} that generates temporally and semantically aligned conditions corresponding to the sign sequences; (3) \textbf{a Gaussian Condition Refiner} designed to enforce temporal smoothness across conditions, enabling the generation of coherent and natural sign language motions; and (4) a discrete diffusion model that synthesizes motion-accurate sign language videos conditioned on the refined condition sequences. Extensive experiments on public SL datasets demonstrate that LLMCond achieves highly competitive performance, producing temporally coherent and accurate sign language motions without relying on auxiliary visual modalities such as pose, depth, or optical flow.

We investigate the automated generation of executable robot reinforcement learning (RL) policies from natural-language task descriptions. Rather than using large language models (LLMs) as execution-time decision makers, we use them as designers that construct RL pipelines before deployment, including Markov decision process (MDP) formulation and simulator implementation. We identify a phase-wise structure in this problem: MDP formulation is dependency-structured and sequential, whereas simulator implementation becomes modular once the formulation is fixed. Based on this structure, we propose ARPG, a dependency-aware agentic framework that performs sequential MDP modeling, localized verification and correction, and simulator-structured implementation while externalizing intermediate artifacts. ARPG targets a key failure mode of automated policy generation: executable code may still yield non-learnable policies when state, action, reward, and termination artifacts are not semantically aligned. Experiments in Isaac Sim show that ARPG improves modeling correctness, environment executability, and policy-generation success compared with monolithic LLM baselines and off-the-shelf agentic systems. Additional demonstrations on custom robotics, standard RL, and engineering-domain MDPs further suggest that the same modeling pipeline can be reused beyond the main benchmark suite.


LocalAgent: Collaborative Agentic Verification for Fine-Grained Instance-Level Consistency

jiachen Guo ⋅ Xinshan Zhu ⋅ Boyun Wang ⋅ Yanyan Liang ⋅ Jun Wan ⋅ Feng Lin ⋅ Ajian Liu

Multimodal large language models (MLLMs) have made rapid progress in general visual understanding, yet they still struggle with fine-grained instance-level consistency verification, where the goal is to determine whether two images depict the exact same physical object under varying viewpoints, illumination, and backgrounds. This task requires both a global understanding of the target instance and a precise examination of subtle local evidence, since semantically similar objects may differ only in small discriminative regions. To address this challenge, we propose \textbf{LocalAgent}, a collaborative agentic framework that integrates the complementary strengths of large and small models. Specifically, the MLLM first reasons from a global instance-level perspective to identify suspicious local regions that may determine the final verification result. These accurately localized regions are then delegated to a lightweight Local Evidence Verifier (LEV) toolkit for fine-grained appearance and geometric matching. The LEV toolkit consists of \textbf{Local Appearance Consistency (LAC)}, which provides illumination- and color-robust similarity estimation, and \textbf{Local Geometric Correspondence (LGC)}, which establishes reliable spatial correspondence under viewpoint changes. Rather than directly replacing the MLLM's reasoning, LEV returns structured similarity scores and comparative evidence in a form that the MLLM can effectively interpret for final decision-making. This design enables a coordinated large-small model collaboration: the MLLM contributes global semantic reasoning and region-level suspicion localization, while specialized local verifiers provide precise, invariant, and interpretable evidence for ambiguous cases. We further introduce an answer-dominant reinforcement fine-tuning (RFT) strategy that enables the policy to learn when verification is necessary. Experiments show that LocalAgent achieves a \textbf{+4.8\%} average improvement over direct-inference baselines, effectively narrowing the gap with closed-source models, while preserving the general multimodal capabilities of the underlying MLLM.


Local Intrinsic Dimension Unveils Hallucinations in Diffusion Models

Bartlomiej Sobieski ⋅ Matthew Tivnan ⋅ Dawid Płudowski ⋅ Michał J Włodarczyk ⋅ Pengfei Jin ⋅ Przemyslaw 'Prem' Biecek ⋅ Quanzheng Li

Diffusion models are prone to generating structural hallucinations - samples that match the statistical properties of the training data yet defy underlying structural rules, resulting in anomalies like hands with more than five fingers. Recent research studied this failure mode from several viewpoints, offering partial explanations to their occurrence, such as mode interpolation. In this work, we propose a complementary perspective that treats hallucinations as instabilities on the model-induced manifold. We begin by showing that a hallucination filter based on such instabilities matches or exceeds the performance of the recently proposed temporal one. By tracing the source of these instabilities, we identify local intrinsic dimension (LID) as their primary driver and propose Intrinsic Quenching (IQ), a direct corrective mechanism that deflates it to alleviate hallucinations. IQ consistently outperforms standard hallucination reduction baselines across a wide array of benchmarks and offers a highly promising solution for enforcing anatomical consistency in downstream medical imaging tasks.

A kernel $k(x,y)$ is *LSHable* if there exists a locality sensitive hashing scheme $\mathcal H$ such that $k(x,y)=\Pr_{h\sim\mathcal H}[h(x)=h(y)]$ for all $x,y$. This notion plays a key role in efficient kernel methods in high dimensions. In this work, we show that the $p$-exponential kernel $k(x,y)=\exp(-\lVert x-y \rVert_p)$ is LSHable in bounded regions for all $1


Locating and Repairing Domain Shift in VLM Trajectory Planning

Zhihong Cui ⋅ Hengyu Liu ⋅ Michael A. Riegler ⋅ Guandong Xu ⋅ Amir Taherkordi ⋅ Tor Skeie

Vision-Language Models (VLMs) are increasingly adopted as end-to-end driving planners, yet they suffer from cross-city domain shift. Existing methods treat this shift as a single quantity addressed by uniform alignment or single-site editing. We refute this view on two counts: ***(i) domain shift in VLM trajectory planners is inherently structured across modality and layer, and (ii) a locate-then-repair recipe matched to this structure outperforms every uniform-objective baseline***. We analyze domain shift via activation patching and uncover two regularities. **(F1) Modality axis.** Image and text tokens exhibit distinct shift patterns in each VLM. **(F2) Layer axis.** The layer-wise shift partition is architecture-determined and dataset-stable. These findings recast domain shift as a **two-dimensional tensor** indexed by *modality* (image vs. text tokens) and *layer*, motivating a simple principle: ***locate where shift occurs, and repair only there***. We instantiate this principle in **MoLaRx** (**Mo**dality–**La**yer **R**epair), a locate-then-repair framework. The Causal Importance Score (CIS) estimates the domain shift as a modality–layer tensor and partitions the layers into functional zones. CIS-Guided Selective Repair (CISR) then applies a matched operator (Maximum Mean Discrepancy (MMD) / low-rank adaptation (LoRA) / freezing) to repair each zone, with the composition adapting automatically to each architecture. We evaluate **MoLaRx** on three architecturally distinct VLMs (LFM2-VL-1.6B, Qwen2-VL-2B, and PaliGemma2-3B) across two cross-city benchmarks (nuScenes and Argoverse 2). MoLaRx attains the lowest cross-city trajectory-error gap on every (VLM, dataset) cell, while using $50$–$60\%$ fewer trainable parameters and $8.4$–$10.9$ ms lower inference latency. In contrast, the strongest scalar baseline (MMD-only) collapses on every cell—direct evidence that no scalar objective can repair structured shift. Code: [anonymous.4open.science/r/MoLaRx-E3B4](https://anonymous.4open.science/r/MoLaRx-E3B4/).


Logarithmic Depth Suffices for In-Context Gradient Descent

Yingze Li ⋅ Dong Wang ⋅ Xianglong Liu ⋅ Qingyun Zou ⋅ Chunnan Wang ⋅ Zhiyu Liang ⋅ Hongzhi Wang

Transformers can perform in-context learning at depths far smaller than the number of optimizer updates suggested by existing gradient-descent interpretations. Prior work shows that an L-layer Transformer can emulate L gradient-descent updates, but this leaves open whether depth is merely a step counter. We show that it is not: in linear self-attention, each layer compounds the algebraic capacity built by previous layers, yielding a tight Θ(log k) depth characterization for computing the final result of k gradient-descent steps on diagonal quadratic tasks. Controlled experiments validate this mechanism: trained linear-attention models exhibit the predicted logarithmic depth transition, and kernel and oracle comparisons identify polynomial degree as the governing bottleneck. Broader experiments show the same depth signal under standard Transformer components, broader ICL regression families, and early-exit probes on Qwen2.5.


LogicSR: A Unified Benchmark for Logical Discovery from Data

Zimeng Zhang ⋅ Xin Zheng ⋅ Feifei Zhang ⋅ Yunxin Liu ⋅ Yuanchun Li

Discovering underlying logical expressions from data is a critical task for interpretable AI and scientific discovery, yet it remains poorly served by existing research infrastructure. The field of Symbolic Regression (SR) primarily focuses on continuous mathematical functions, while Logic Synthesis (LS) is designed for exact, noise-free specifications, not for learning from incomplete or noisy data. This leaves a crucial gap for evaluating algorithms that can learn generalizable logical rules in realistic scenarios. To address this, we introduce LogicSR, a large-scale and comprehensive benchmark for logical symbolic regression. LogicSR is built from two sources: real-world problems from digital circuits and biological networks, and a novel synthetic data generator capable of producing a diverse set of complex logical formulas at scale. We use LogicSR to conduct a rigorous evaluation of 17 algorithms, spanning classical logic solvers, modern machine learning models, and Large Language Models (LLMs). Our findings reveal that the logical modeling capabilities and generalization robustness of these algorithms significantly depend on task scale and logical complexity, with current cutting-edge LLMs showing limited complex logical reasoning ability. LogicSR provides a robust foundation to benchmark progress, unify evaluation across disparate fields, and steer the future development of powerful neuro-symbolic systems.


LookWhen? Fast Video Recognition by Learning When, Where, and What to Compute

Ali Salamatian ⋅ Anthony Fuller ⋅ Pritam Sarkar ⋅ James Green ⋅ Leonid Sigal ⋅ Evan Shelhamer

Transformers dominate video recognition. They split videos into tokens, and processing them has expensive superlinear computational cost. Yet videos are filled with redundancy, so we can question the need for this expense. We introduce LookWhen, a selector–extractor framework that factorizes video recognition into learning when, where, and what to compute. Our shallow selector gets a scaled-down video and quickly scores all tokens across space-time, while our deep extractor gets the top-K selected tokens to approximate full-video representations without actually processing all the tokens. A key challenge is defining effective supervision for selection and extraction. For selection pre-training, we introduce a score on representations that ranks tokens by uniqueness using a simple nearest-neighbor distance. For extraction pre-training, we distill both a video teacher and an image teacher, for which we normalize its frame-wise representations to learn what changes within videos. Through these strategies, our selector-extractor learns general and efficient representations for feature extraction or fine-tuning to a task. Through experiments on Kinetics-400, SSv2, Epic-Kitchens, Diving48, Jester, and Charades, we show that LookWhen achieves a better accuracy-computation trade-off than efficient models and upgraded baselines of similar size. LookWhen Pareto-dominates in accuracy-FLOPs on 9 of 12 cases (6 tasks X 2 settings) and roughly matches on 3. In accuracy-throughput, measuring time in practice, LookWhen is more efficient still at 6.7$\times$ faster than InternVideo2-B at equal accuracy.


LoRAcles: Self-Supervised Weight-Space Interpretability at Scale

Celeste De Schamphelaere ⋅ Jan Bauer ⋅ Neel Nanda ⋅ Euan Ong

Fine-tuned LLMs can learn complex and subtle behaviors. However, it can be difficult to ascertain what behaviors the model acquired during fine-tuning. We introduce LoRAcles: fine-tuned language models that take LoRA adapter weights as input, and answer natural language questions about them. We introduce a self-supervised pipeline for training LoRAcles: we first train LoRAs on small sets of pre-training documents, and then train LoRAcles on these LoRAs to answer questions about the documents. LoRAcles generalize far beyond their training data: they are state-of-the-art on AuditBench, a benchmark of models with sophisticated hidden behaviors; they can detect subtle changes in model organisms (fine-tunes deliberately constructed to encode a hidden behavior for auditing research), such as subliminal learning; and they are the first tool to achieve non-trivial performance at verbalizing semantic backdoor triggers. LoRAcles can also describe some learned behaviors from the weights of a full-parameter fine-tune. LoRAcles are a highly scalable method: we train LoRAcles for Qwen-3-14B and Llama-3.3-70B, and observe that performance scales smoothly with the size of our training dataset up to 100K LoRAs, though our method may need refinement in order to scale further. We note, though, that LoRAcles are prone to hallucinations and currently only elicit the most salient behavioral changes, limiting their utility beyond narrow fine-tunes.


LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging

Peter Richtarik ⋅ Yassine Maziane ⋅ Ammar Mahran ⋅ Artavazd Maranjyan

Communication is a major bottleneck in distributed learning, especially in large-scale settings and in federated learning environments with slow links. Three standard ways to reduce this cost are communication compression, local training, and communication-computation overlap. Methods that combine these ingredients are used in practice and have been found to be effective for large-scale training, but there is little theory for methods that combine all three. We study a heterogeneous-compute setting in which different workers may take different numbers of local steps, and we propose LOSCAR-SGD, a Local SGD method that communicates only a sparse subset of model coordinates and continues optimizing while communication is in flight. A key ingredient is a delay-corrected merge rule that incorporates delayed synchronized information without discarding the progress made during the overlap phase. We give convergence guarantees for smooth non-convex objectives and show how sparsity, overlap, and worker heterogeneity affect the rate. To the best of our knowledge, this is the first theory for this combination of ingredients. Experiments further show that communication-computation overlap reduces training time and that the delay-corrected merge outperforms naive overwriting.


Low-Rank Adaptation for Critic Learning in Off-Policy Reinforcement Learning

Yuan Zhuang ⋅ Yuexin Bian ⋅ Sihong He ⋅ Jie Feng ⋅ Qing Su ⋅ Songyang Han ⋅ Jonathan Petit ⋅ Shihao Ji ⋅ Yuanyuan Shi ⋅ Fei Miao

Scaling critic capacity is a promising direction for enhancing off-policy reinforcement learning (RL). However, larger critics are prone to overfitting and unstable in replay-buffer-based bootstrap training. This paper leverages Low-Rank Adaptation (LoRA) as a structural-sparsity regularizer for off-policy critics. Our approach freezes randomly initialized base matrices and solely optimizes low-rank adapters, thereby constraining critic updates to a low-dimensional subspace. Built on top of SimbaV2, we further develop a LoRA formulation, compatible with SimbaV2, that preserves its hyperspherical normalization geometry under frozen-backbone training. We evaluate our method with SAC and FastTD3 on DeepMind Control locomotion and IsaacLab robotics benchmarks. LoRA consistently achieves lower critic loss during training and stronger policy performance. Extensive experiments demonstrate that adaptive low-rank updates provide a simple, scalable, and effective structural regularization for critic learning in off-policy RL.

Guardrails are a critical safety layer for modern AI systems, but their operating regime is changing. As LLMs are deployed as customized assistants, safety policies are increasingly specified at inference time by users, organizations, or regulatory contexts. This makes safety enforcement fundamentally dynamic: the guardrail should adapt to changing safety policies without retraining. Yet this requirement creates a fundamental tension: faithfully judging complex policy contexts demands reasoning capability, while practical deployment requires low-latency responses. We introduce Latent Policy Guardrail (LPG), a guardrail framework that learns semantic latent deliberation over dynamic policies. LPG compresses the internal deliberation needed for intent interpretation and policy grounding into continuous states supervised by decision-relevant semantics. At inference time, it generates only a compact verdict anchored to the violated policy clauses, preserving auditability while avoiding the latency of explicit reasoning. Across policy guardrail benchmarks, LPG-4B reaches 84.5% average safety accuracy and 77.9% F1 by compressing deliberation into just 10 latent tokens, outperforming the strongest dynamic baseline while running roughly 11 times faster than Qwen3-4B-Thinking under the single-sample evaluation setup.

Computer-use agents (CUAs) that interact with real systems can automate complex tasks, but introduce critical safety risks in long-horizon tool-use workflows. Many existing benchmarks rely on outcome-based evaluation, however, outcome-based evaluation cannot reflect agent's safety awareness, as an agent may reach a safe outcome by chance while taking unsafe steps during planning. To address this gap, we present \textsc{LPS-Bench}, which extends the outcome-centric evaluation towards LLM agents' safety awareness in tool usage (MCP/skill) over long-horizon tasks, as reflected at the action trajectory level. \textsc{LPS-Bench} comprises 570 test cases derived from 65 scenarios across 7 task domains and 9 planning-risk types, with user contexts spanning both benign ambiguity and adversarial steering. Results on \textsc{LPS-Bench} across 13 representative LLM agents reveal previously underexamined safety failures in long-horizon tool-use tasks, showing that current agents often struggle to maintain safety awareness under ambiguous benign instructions and adversarial steering. We further analyze key failure modes and evaluate lightweight mitigation strategies, showing that prompt-based interventions provide limited improvements yet remain insufficient for robust trajectory-level planning safety.


LUMOS: Tracing Parametric Knowledge from Training Data to Behavioral Outputs in LLMs

Seoyeon Ye ⋅ Gayoung Kim ⋅ Jiyoung Hong ⋅ Soo Kyung Kim ⋅ Hyunsoo Cho

Current analyses of LLMs' parametric knowledge are largely output-centric, drawing conclusions about what a model knows without verifying what it was actually trained on. This leaves fundamental questions, such as whether a correct response reflects genuine generalization or rote memorization, grounded in speculation rather than evidence. To resolve these ambiguities, we introduce LUMOS, a diagnostic framework that traces knowledge along the causal chain from training-data exposure to behavioral output, leveraging OLMo 2 with its fully transparent training corpus. By grounding analysis in verified exposure, we reveal that models internally encode rare facts with high separability (84%) yet fail to express them behaviorally (54%), though this retrieval gap narrows with scale. Furthermore, when models are asked to self-reflect on their own answers, they perform reliably on trained content (83%) but drop to random-baseline levels (49%) on unseen content. This collapse persists even under chain-of-thought prompting, which inflates confidence signals rather than improving calibration. Collectively, these findings demonstrate that incorporating the training-data axis into LLM evaluation transforms speculative diagnoses into verifiable claims, and we advocate that this axis should be a standard component of knowledge assessment in LLMs.


LYNX: Learning Dynamic Exits for Confidence-Controlled Reasoning

Ömer Faruk Akgül ⋅ Yusuf H Kalayci ⋅ Rajgopal Kannan ⋅ Willie Neiswanger ⋅ Viktor Prasanna

Large reasoning models achieve strong performance by generating long chains of thought, yet extended reasoning is often counterproductive: models frequently continue past the point of sufficiency, wasting compute and sometimes degrading correctness by revising a correct solution into an incorrect one. We show that reasoning models encode a domain-general \emph{readiness} signal in their hidden states, concentrated at natural high-uncertainty points such as "wait" and "hmm", that predicts whether stopping at a given moment would yield the correct answer. Building on this finding, we propose LYNX, an online early-exit framework that reads out this signal during generation and wraps it in split conformal calibration, giving practitioners a single confidence dial with calibrated control over erroneous exits rather than per-task heuristic thresholds. The readout requires no external verifier, auxiliary model, or additional human annotation: stop/continue labels are obtained by forcing the base model to answer at candidate stopping points and comparing to the task answer. A single LYNX head trained and calibrated once on a generic mathematical corpus transfers unchanged across benchmarks, decoding temperatures, and non-mathematical domains including commonsense reasoning and code generation, suggesting that readiness reflects model-internal representations rather than task-specific heuristics. Across three model families spanning 1.5B--32B parameters, LYNX matches or improves baseline accuracy while reducing generated tokens by 30--70\%, with competitive or superior accuracy--efficiency Pareto frontiers relative to prior early-exit methods.


MACRO: Training-free Multi-plane Attention for Closeup Render Optimization

Nitzan Hodos ⋅ Roy Amoyal ⋅ Lior Fritz ⋅ Ianir Ideses ⋅ Sagie Benaim ⋅ Netalee Efrat

Close-up rendering, zooming into a scene well beyond any training camera, is important for virtual production and interactive 3D content, yet remains an open challenge. 3D Gaussian splatting (3DGS) enables high-fidelity, real-time novel view synthesis, but its rendering quality degrades at close range. Recent diffusion-based methods that enhance the rendering by conditioning on reference images from the training set produce significant artifacts in this setting. We analyze this failure and identify its root cause: the scale gap between the close-up and reference views. We show that the features in reference-conditioned enhancement models are not scale-invariant, causing cross-view attention to retrieve incorrect correspondences when the same content appears at different scales, and that this mismatch cannot be corrected in latent space because the VAE encoder is not scale-equivariant. Building on this analysis we introduce MACRO, Multi-plane Attention for Closeup Render Optimization, a training-free method for high-quality close-up novel view synthesis from 3DGS. MACRO resolves the scale gap by leveraging the scene's known 3D structure: it decomposes the close-up into depth planes, crops and resizes references in image space to match the scale of each plane before encoding, and applies a depth-aware attention mask so each token attends only to scale-matched references. The method requires no architectural changes or additional training. We further contribute two new close-up novel view synthesis benchmarks, the first standardized evaluation protocol for this setting, and demonstrate state-of-the-art results on both, outperforming existing 3DGS and diffusion-based methods on both reconstruction and perceptual metrics.


MAGE: All-[MASK] Block Already Knows Where to Look in Block Diffusion LLM

Omin Kwon ⋅ Yeonjae Kim ⋅ Doyeon Kim ⋅ Minseo Kim ⋅ Yeonhong Park ⋅ Jae W. Lee

Block diffusion LLMs are an emerging paradigm for parallel language generation, but their KV caching makes memory access the dominant bottleneck in long-context inference. Sparse attention, which attends only to a small KV subset per query, can reduce this latency with minimal accuracy loss. In block diffusion, however, the $B$ tokens of each block must share a single KV subset, and we show this per-block constraint degrades existing sparse KV estimators by up to 25% in recall. We address this challenge by exploiting a property that emerges from the block-diffusion training objective: it aligns the block-average query across denoising steps, so the All-[MASK] block at the first step already reveals the per-block KV subset for the entire trajectory. We exploit this in MAGE ([MASK]-Guided Sparse Attention), a training-free method that runs one exact attention pass at the first step and reuses its top-$k$ index sets for all remaining steps within the block. Across three block-diffusion families on LongBench, MAGE matches Exact Attention at $k{=}512$ with near-lossless accuracy, achieves up to $6.82\times$ end-to-end speedup at $128$K context, and runs up to $3.35\times$ and $2.28\times$ faster than Quest and SparseD, designed for AR LLMs and fully bidirectional diffusion LLMs, respectively.


Magnifying What Matters: Attention-Guided Adaptive Rendering for Visual Text Comprehension

Shenglai Zeng ⋅ Qirui Wang ⋅ Kai Guo ⋅ Xinnan Dai ⋅ Xianxuan Long ⋅ Hui Liu

Visual Text Comprehension (VTC) renders text into images for a vision-language model (VLM) to read, sidestepping LLM context-window limits and powering applications from long-page OCR to multi-page memory QA. Yet existing VTC pipelines treat rendering and layout as a fixed, content-agnostic preprocessing step, and offer little mechanistic understanding of how VLMs internally process visualized text. Through a focused empirical study on VTC QA tasks, we find that VLMs exhibit a localization-without-utilization regime: evidence-localizing attention emerges sharply in the middle-to-late layers and is largely decoupled from answer correctness, yet simply enlarging the localized spans on the rendered page recovers a large fraction of the failures. Building on these observations, we propose AGAR (Attention-Guided Adaptive Rendering), a training-free, model-agnostic method that leverages a VLM's own middle-to-late layer attention to identify the top-K important visual patches, maps them back to word spans, and re-renders the page with those spans enlarged before re-inferring the answer. Extensive experiments across nine VTC benchmarks (short-form, long-context, and multi-page memory QA) and four VLM backbones show that AGAR (i) consistently improves off-the-shelf VLMs as a plug-and-play enhancement, (ii) composes with VLM post-training to yield further gains, and (iii) remains robust under both visual- and text-side input degradation.

Out-of-distribution (OOD) detection is a critical component for ensuring the reliability of deep neural networks in safety-critical applications. In this work, we present a key empirical observation: for in-distribution (ID) samples, class-wise Mahalanobis distances exhibit a pronounced sharp minimum structure, where the distance to the nearest class is small while distances to all other classes remain large, resulting in high variance across classes. In contrast, OOD samples tend to exhibit a less pronounced sharp minimum structure, producing comparatively lower variance across classes. We further provide a theoretical analysis grounding this observation in Neural Collapse geometry: under relaxed Neural Collapse assumptions on within-class compactness and inter-class separation, ID samples are shown to structurally exhibit high class-wise distance variance, offering a theoretical basis for its use as an OOD score. Motivated by this observation and its theoretical backing, we propose MahaVar, a simple and effective post-hoc OOD detector that augments the Mahalanobis distance with a class-wise distance variance term. Following the OpenOOD v1.5 benchmark protocol, MahaVar achieves state-of-the-art performance on CIFAR-100 and ImageNet, with consistent improvements in both AUROC and FPR@95 over existing Mahalanobis-based methods across all benchmarks.

Finding the optimal sample complexity of realizable PAC learning was a major open problem in learning theory, which was settled by the breakthrough result of Hanneke (2016), building on prior work by Simon (2015). Hanneke's algorithm is quite involved and requires training $\text{poly}(n)$ Empirical Risk Minimizers (ERMs) on carefully crafted subsets of the data. Since using a single ERM call is provably suboptimal, recent work by Aden-Ali (2024) proposed the majority-of-three ERMs as a candidate for the *simplest* optimal PAC learner, and established its optimality in the *in-expectation* regime. Their main open question was whether this algorithm is optimal in the classical *high-probability* regime. In this work, building on their ideas, we answer this question affirmatively.


MAMQ-Net: A Robust Framework Leveraging Multi-level Tamper-aware Queries and Complementary Representations of Tampering Features for Progressive Image Forgery Localization

Minghao Jia ⋅ Zhuoyi Zhang ⋅ Jiwei Zhang ⋅ Siwei Wang ⋅ Feifei Kou ⋅ Lei Shi ⋅ ShaoZhang Niu ⋅ Jiaming Pei

Image Forgery Localization (IFL) aims to identify and localize the tampered regions within edited images. Many studies employ a dual-branch backbone to extract tampering features from dual modalities, followed by feature fusion at the final stage. In this process, the extraction and fusion of dual-modality features is relatively independent, which fails to fully leverage the complementarity between different modalities and thus diminishes sensitivity to tampering artifacts. Inspired by the way humans continuously integrate multi-faceted knowledge to understand the world, we propose MAMQ-Net, which contains a novel Multi-stage Alternating Feature Extraction and Interaction architecture. At each stage, we deeply explore the intrinsic relationships and mappings between different modality features. Feature extraction and interaction are performed alternately, constructing complementary dual-modality tampering feature representations and enhancing sensitivity to tampering artifacts. Additionally, we introduce a lightweight, Query-driven Multi-level Feature Decoding. This mechanism progressively aggregates key information from multi-level dual-modality tampering features through multiple sets of learnable tamper-aware queries, effectively filtering out irrelevant features. Finally, multi-level queries are used to refine discriminative features, enabling precise localization of tampered regions. Extensive experiments demonstrate that our framework outperforms current state-of-the-art models in localization accuracy and robustness across multiple public datasets, achieving a favorable balance between performance and efficiency.


ManifoldCache: Training-Free Diffusion Acceleration via Constraint Manifold Caching

Prashant Pandey ⋅ Sri Venkatraya Chowdary Devineni ⋅ Brejesh Lall

Diffusion models for structured scientific generation must produce samples satisfying hard geometric constraints imposed by physics, chemistry, or biology, yet inference in these settings is prohibitively slow, demanding hundreds to thousands of neural-function evaluations per sample. We unify eight state-of-the-art models spanning medical volumetrics, molecular conformations, protein backbone design, crystal structure prediction, and multi-view 3D scenes under a single abstraction, Constraint-Manifold Diffusion Models (CMDMs), in which the target distribution is supported on a manifold defined by an externally specified constraint map. All existing acceleration families fail on this class: quantization exhausts memory on high-dimensional volumetric operators; pruning breaks constraint fidelity; fast ODE solvers allow trajectories to drift off the constraint manifold; and feature-caching heuristics are blind to constraint geometry, inducing mode confusion in the high-noise regime. We introduce ManifoldCache, the first training-free, data-free accelerator designed from first principles for CMDMs. The key insight is that the conditional score decomposes orthogonally into a normal component, which enforces constraint satisfaction, and a tangential component, which navigates within the manifold. Exploiting this structure, we prove that the noise-schedule midpoint is a sharp safe-caching boundary: caching before it incurs provably bounded error, while caching after it guarantees a strictly positive fraction of trajectories suffer mode confusion, a gap that persists up to the boundary. We further prove that deeper network blocks admit provably larger certified cache strides within the safe phase, as a consequence of the score decomposition propagating through block Jacobians. The resulting schedule requires one integer comparison per block per step, zero calibration data, and zero training overhead. Across all eight CMDMs, ManifoldCache delivers consistent wall-clock speedups while preserving or improving generation quality on structural, perceptual, and distributional metrics, consistently outperforming several strong baselines spanning quantization, pruning, fast solvers, and feature caching.

Estimating semantic correspondences between different object instances of similar categories in real-world images is a fundamental challenge in computer vision. While the recent progress in foundation models has significantly advanced solutions for such difficult image correspondence problems, respective methods are often insufficiently regularised. For example, common nearest-neighbour matching with foundation model features often results in noisy correspondences, which is particularly prominent when solving for dense correspondences. In this work, we tackle image matching via functional maps, which have been popularised for 3D shape matching due to their powerful spectral formalism utilising an efficient low-dimensional linear representation. However, functional maps require that domains (in our cases the rectangular image grids) are equipped with an informative and non-trivial geometry. To define a meaningful manifold structure for images, we use Galerkin's method from finite elements to derive a discrete Laplace-Beltrami operator (LBO) based on a manifold embedding of deep image features. Further, we leverage a neural adjoint map that allows to represent non-linear mappings. We demonstrate that our method sets the new state of the art among zero-shot image correspondence methods on multiple dense correspondence benchmarks of real-world images.


Manifold-weighted neural networks

Junyu Ren ⋅ Lek-Heng Lim

We establish universal approximation for neural networks whose weight matrices take values in matrix manifolds. The nonlinear matrix manifolds considered in this work arise as orbits of classical Lie group actions embedded in the ambient space $\mathbb{R}^{d \times k}$. Stiefel manifolds provide a prominent example with empirical precedent in prior neural-network models. We present an expanded catalog of such manifolds, drawing from well-studied constructions in physics and differential geometry. These manifolds are particularly appealing for intrinsic Riemannian optimization: the transitivity of the underlying group action yields explicit tangent-space descriptions and retractions. Moreover, the group-action structure enables parameter-efficient factorizations, reducing the number of trainable variables. This leads to a fundamental expressivity question: when weight matrices take values in nonlinear matrix manifolds, which architectures retain universal approximation power? We prove that manifold-weighted networks with residual connections and bounded scalar multipliers are $L^p(K)$-universal. To establish universality across our manifold catalog, we introduce two independent sets of sufficient conditions: Register-Isolated Primitive Realizability (RIPR) and Cross-Axis Fold-and-Cut (CAFC). These conditions correspond to two distinct approximation constructions. By verifying RIPR or CAFC case by case for the matrix manifolds in our catalog, we establish universality for the corresponding residual manifold-weighted networks with bounded scalar multipliers.


ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation

Zitong Xu ⋅ Huiyu Duan ⋅ Xiaoyu Wang ⋅ Zhaolin Cai ⋅ Kaiwei Zhang ⋅ Xiongkuo Min ⋅ Guangtao Zhai

With the rapid advancement of generative models, powerful image editing methods now enable diverse and highly realistic image manipulations that far surpass traditional deepfake techniques, posing new challenges for manipulation detection. Existing image manipulation detection and localization (IMDL) benchmarks suffer from limited content diversity, narrow generative-model coverage, and insufficient interpretability, which hinders the generalization and explanation capabilities of current manipulation detection methods. To address these limitations, we introduce ManipBench, a large-scale benchmark for image manipulation detection and localization focusing on AI-edited images. ManipBench contains over 450K manipulated images produced by 25 state-of-the-art image editing models across 12 manipulation categories, among which 100K images are further annotated with bounding boxes, judgment cues, and textual explanations to support interpretable detection. Building upon ManipBench, we propose ManipShield, an all-in-one model based on a Multimodal Large Language Model (MLLM) that leverages contrastive LoRA fine-tuning and task-specific decoders to achieve unified image manipulation detection, localization, and explanation. Extensive experiments on ManipBench and several public datasets demonstrate that ManipShield achieves state-of-the-art performance and exhibits strong generality to unseen manipulation models. Both ManipBench and ManipShield will be released upon publication.

Building embodied agents capable of accomplishing arbitrary tasks is a core objective towards achieving embodied artificial general intelligence (E-AGI). While recent work has advanced such general robot policies, their training and evaluation are often limited to tasks within specific scenes, involving restricted instructions and scenarios. Existing benchmarks also typically rely on manual annotation of limited tasks in a few scenes. We argue that exploring the full spectrum of feasible tasks within any given scene is crucial, as they provide both extensive benchmarks for evaluation and valuable resources for agent improvement. Towards this end, we introduce ManiTaskGen, a novel system that automatically generates comprehensive, diverse, feasible mobile manipulation tasks for any given scene. The generated tasks encompass both process-based, specific instructions (e.g., "move object from X to Y") and outcome-based, abstract instructions (e.g., "clear the table"). We apply ManiTaskGen to both simulated and real-world scenes, demonstrating the validity and diversity of the generated tasks. We then leverage these tasks to automatically construct benchmarks, thoroughly evaluating the embodied decision-making capabilities of agents built upon existing vision-language models (VLMs). Furthermore, we propose a simple yet effective method that utilizes ManiTaskGen tasks to enhance embodied decision-making. Overall, this work presents a universal task generation framework for arbitrary scenes, facilitating both benchmarking and improvement of embodied decision-making agents.

Embodied agents are often evaluated in the same environment in which they were explored or trained, making it difficult to assess whether a learned world model supports planning after the world changes. Existing evaluations can conflate memory of the explored environment, belief update after change, and planning under the updated belief. We introduce MapShift, an executable benchmark for controlled post-intervention evaluation (CPE): an agent explores a base environment without a task reward; the environment is modified by a controlled intervention in the metric, topology, dynamics, or semantics; and the agent is evaluated on post-intervention planning, inference, and adaptation tasks. The contribution is measurement infrastructure: matched base/intervened pairs, family-wise estimands, severity ladders, invariant validators, benchmark-health gates, protocol-comparison tooling, and reproducible artifact generation. In the expanded 24-motif release, health gates pass with zero fatal leakage, zero task rejections, perfect reference solvability, no intervention-validator failures, and no severity-magnitude failures. A deterministic mechanism diagnostic shows that same-environment evaluation underestimates the belief-update advantage by 3x on topology shifts, from $\Delta=0.304$ under CPE versus $0.102$ same-environment, and by 7x on semantic shifts, from $\Delta=0.724$ versus $0.102$; planning slices also exhibit protocol-induced rank reversals.

Despite their practical significance, modern AI systems pose significant societal risks, including toxicity, malicious use, and misinformation. Their mitigation depends not only on technical feasibility but also on the incentives of firms that develop and deploy these systems. We show how profit maximization shapes safety investment in an AI supply chain with an upstream LLM provider and several competing downstream firms. In our game-theoretic model, we derive the market-driven level of safety investment and compare it with the levels that maximize market welfare and industry profit. We find that downstream competition generally induces downstream firms to invest more in safety than the industry-profit-optimal level. By contrast, the upstream provider may underinvest relative to these benchmarks because it is not fully compensated for the surplus from safety investments. Finally, we analyze how different regulatory interventions affect the participants' incentives and profits, and the resulting levels of model safety.

State Space Models (SSMs) offer an efficient alternative to Transformers for sequence modeling, yet conditioning pre-trained SSMs for iterative generation typically operates outside the recurrent operator, through input injection or activation modulation. While such mechanisms expose the model to conditioning information, they leave the underlying temporal dynamics fixed. We introduce MaRK (Markov-adapted Recurrent Kernels), a dynamic operator-conditioning framework that maps context vectors directly into bounded modulations of a frozen SSM's recurrence ($A$), read-in ($B$), read-out ($C$), skip ($D$), and discretization ($\Delta$) parameters. Viewed through the lens of Linear Parameter-Varying systems, MaRK induces a context-indexed family of Markov parameter sequences, allowing each diffusion timestep to reshape the model's input-output memory kernel. We instantiate MaRK on a frozen 111M-parameter Hydra SSM backbone and study three adapter geometries: Hypernet, Chebyshev polynomial, and Discrete Cosine Transform kernels. Since these adapters modify the Markov parameter sequence through low-rank auxiliary maps on the frozen backbone, parameter-efficient fine-tuning arises as a structural consequence of the adaptation mechanism itself, requiring only 7.5--13M trainable auxiliary parameters to transition from a bidirectional objective to an iterative diffusion regime. The bounded recurrence parameterization further yields an analytic Affine Quadratic Stability certificate for the modulated recurrence. Through synthetic LPV recovery experiments and Markov-operator diagnostics, we show that MaRK recovers coordinate-invariant temporal operators under matched assumptions and produces distinct, stable timestep-conditioned memory profiles. Empirically, the Chebyshev variant yields the strongest performance, achieving an average validation loss of 2.63, followed by the DCT (2.66) and Hypernet (3.60) geometries. Together, these results provide initial evidence that dynamic operator modulation is a principled operator-level conditioning mechanism for adapting SSMs beyond input-stream injection and adaptive normalization.


Marrying Optimal Transport and ODEs for Unified Continuous-Time 4D Reconstruction and Tracking

Liying Yang ⋅ Hao Mo ⋅ Jialun Liu ⋅ Chen Liu ⋅ Xinxing Yu ⋅ Chenhao Guan ⋅ Hui Ma ⋅ Xiao Cao ⋅ Ajian Liu ⋅ Yanyan Liang

Existing unified 4D reconstruction and point tracking approaches typically rely on heuristic interpolations or just predict at integer timestamps, lacking kinematic coherence and failing to model dynamics at any arbitrary timestamp. In this paper, we propose Uni4R, a framework that unifies these tasks by learning continuous velocity fields through the synergy of Optimal Transport (OT) and Ordinary Differential Equation (ODE). Importantly, this continuous velocity field acts as a kinematic prior that mutually benefits both 4D reconstruction and point tracking. Specifically, we propose the Flow Matching Guided Decoder (FMGD). A global velocity branch first extracts anchor features that capture the global dynamic state of the sequence. Then, FMGD leverages Flow Matching (FM) theory to formulate a probability path defined by OT on the anchor feature manifold, instantiating it as FM-guided velocity features for velocity prediction. This establishes a robust kinematic inductive bias. Meanwhile, a point reconstruction branch provides geometric features. The local velocity prediction module then joint above features and time embeddings, to decode velocities at arbitrary timestamps. To overcome the absence of high-quality ground-truth velocities in fractional frames, we propose an integral-consistency training strategy. This strategy uses an ODE solver to integrate velocities to recover target pointmaps, enabling the model to be supervised end-to-end directly from integer timestamps. Experimental results demonstrate that Uni4R achieves SOTA performance in both 4D reconstruction and point tracking, and achieves SOTA in our new kinematics-aware benchmark at continuous time.


Mask-Conditioned Gradient Masking for Fine-Tuning Mixture-of-Experts Diffusion Language Models

Yiru Tang ⋅ Kun Zhou ⋅ Xin Zhao ⋅ Jing Sha ⋅ Zhichao Sheng ⋅ Shijin Wang

While autoregressive models dominate the LLM landscape, Discrete Diffusion Language Models have emerged as a compelling alternative due to their advantages in parallel decoding and bidirectional contextual modeling. To enhance scalability and efficiency, recent works have integrated the Mixture-of-Experts architecture into the DLM framework. However, MoE-DLMs remain underexplored. In this work, we present an empirical study revealing that MoE-DLMs exhibit not only task-level expert specialization, but also exhibit specialization across different mask-rate diffusion regimes. We further show that during fine-tuning, gradients induced by mismatched mask rates can interfere with the updates of regime-specialized experts, leading to suboptimal adaptation. Motivated by this, we propose MaGM, an adaptive mask-conditioned gradient masking method for MoE-DLM fine-tuning, which dynamically masks expert parameter updates based on the mask rate of training instance. Experiments demonstrate that MaGM consistently outperforms standard full-parameter fine-tuning for diffusion language models, validating the benefits of regime-aware expert adaptation. Our code will be publicly released.


Matching-Based Few-Shot Semantic Segmentation Models Are Interpretable by Design

Pasquale De Marinis ⋅ Uzay Kaymak ⋅ Rogier Brussee ⋅ Gennaro Vessio ⋅ Giovanna Castellano

Few-Shot Semantic Segmentation (FSS) models achieve strong performance in segmenting novel classes with minimal labeled examples, yet their decision-making processes remain largely opaque. While explainable AI has advanced significantly in standard computer vision tasks, interpretability in FSS remains underexplored despite its critical importance for understanding model behavior and guiding support set selection in data-scarce scenarios. We argue that matching-based FSS models are interpretable by design: their core similarity computation between support and query features constitutes an inherent attribution mechanism. Our approach, Affinity Explainer (AffEx), makes this intrinsic interpretability explicit by extracting attribution maps directly from matching scores at multiple feature levels, without requiring gradients or external perturbations. We extend standard interpretability evaluation metrics to the FSS domain and propose additional metrics to better capture the practical utility of explanations in few-shot scenarios. Comprehensive experiments on FSS benchmark datasets demonstrate that AffEx significantly outperforms adapted standard attribution methods, confirming that the matching mechanism itself is the key driver of interpretability. Qualitative analysis reveals structured, coherent attention patterns that align with model architectures and enable effective model diagnosis, laying the groundwork for interpretable FSS research.

Natural gradient descent is a theoretical pillar of second-order optimization for neural networks. Existing implementations either rely on structural approximations of the curvature matrix or incur substantial computational and memory overhead. Direct application remains difficult at deep-network scale. We build an optimizer based on an unbiased stochastic estimator of the Fisher information matrix. It employs a truncated exponential moving average metric with damping, which we call the Stochastic Truncated Metric (STM), and is otherwise closely aligned with the original natural gradient. This gives a matrix-free parameter update with strictly $O(Kd)$ time and memory per step, where $K$ is a small constant. It is applicable across network architectures without layer-wise assumptions. We bound the error induced by the finite-memory truncation. We show that STM is stable and efficient on vision and language benchmarks.


Measuring Weak-to-Strong Legibility of Reasoning Models

Dani Roytburg ⋅ Shreya Sridhar ⋅ Daphne Ippolito

Language models are increasingly trained to "reason" before answering users' queries, outputting hundreds of intermediate tokens before a final answer. While reasoning is designed with direct model capabilities in mind, we argue that reasoning language models (RLMs) should also be assessed by interaction effects between their reasoning traces and weaker models. In contexts like safety monitoring and distillation, the performance of weak models increasingly depends on legibility---how thoroughly and precisely a reasoning trace externalizes an RLM's actions. Existing efficiency-based metrics for legibility fail to capture "thoroughness", instead focusing on conciseness. Thus, we introduce transfer utility, a method for measuring trace legibility derived from interactions where weak models finish tasks given incomplete traces. Evaluating 85k traces from 12 RLMs across three tasks, we find that reasoning traces generated by the highest-performing and the most efficient models rank the lowest for transfer utility. We also find evidence that transfer utility predicts the ability of weak monitors to verify procedurally dense traces for math or logical problem-solving, though the effect moderates when verifying fact-driven traces. Finally, we show that open-source reward models decouple from transfer utility signals on correctness. Together, these findings surface status quo trade-offs in weak-to-strong legibility, motivating an interaction-first framework for future applications.


Mechanisms of Misgeneralization in Physical Sequence Modeling

Kento Nishi ⋅ Raphael Tang ⋅ Karun Kumar ⋅ Core Francisco Park ⋅ Hidenori Tanaka

Generative sequence models are often trained to plan motion in physical domains, from robotics to mechanical simulations. When constructing a dataset to train such a model, engineers may curate demonstrations to specify how trajectories should be distributed over a physical quantity like travel distance or mechanical energy. For example, a roboticist building a maze navigation agent might choose demonstrations whose travel distances cover a fixed range uniformly, hoping to constrain the agent's expected power usage. We find that standard deep learning can violate this intent: each generated trajectory can seem plausible on its own, but the aggregate distribution over the physical quantity is wrong. We call this failure physical misgeneralization, and develop an account of its mechanism. Using controlled synthetic tasks, we show that physical misgeneralization arises when local errors typical of the model class propagate through the physical measurement to shift the recovered distribution. We estimate these errors with a data deviation kernel, and we use it to predict which physical quantities gain or lose mass in both our synthetic and more applied maze navigation and double-pendulum motion tasks. Finally, our mechanistic interpretation helps identify which mitigation strategies are structurally promising, and we use it to propose a kernel-informed intervention.


MedFlowBench: Auditing Medical Agents in Full-Study Workflows

Weixiang Shen ⋅ Chengzhi Shen ⋅ Che Liu ⋅ Junde Wu ⋅ Jiayuan Zhu ⋅ Xiao Han ⋅ Zongyue Li ⋅ Jingpei Wu ⋅ Min Xu ⋅ Daguang Xu ⋅ Yueming Jin ⋅ Benedikt Wiestler ⋅ Daniel Rueckert ⋅ Jiazhen Pan

Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition. Real clinical workflows impose a different burden: readers must search through complete studies, operate imaging software, navigate across slices and magnifications, and document visual evidence that can be audited. We argue that this evidence-producing workflow is a critical missing evaluation axis for medical imaging agents. To study it, we introduce MedFlowBench, a full-study benchmark for VLM agents, together with MedOpenClaw, a controlled and replayable runtime in which agents operate medical imaging viewers such as 3D Slicer and QuPath. In each episode, an agent inspects a complete radiology study or whole-slide pathology image, returns a task answer, and submits structured evidence, including key slices, coordinates, regions of interest, or lesion-state fields. This evidence is automatically checked against withheld masks, annotations, and labels. Across evaluated models, final answer-only scoring gives an overly optimistic picture: when answers must also be supported by correct evidence, performance drops substantially on complex workflows. We further find that adding image-analysis tools does not by itself solve the problem. Tools help when they make a complex procedure simple and reliable, but agents still struggle when they must choose inputs, manage viewer state, and verify intermediate outputs over multiple steps. MedFlowBench exposes whether medical imaging agents can produce auditable evidence from complete studies, rather than plausible answers from selected images.


Medmarks: A Comprehensive Open-Source LLM Benchmark Suite for Medical Tasks

Benjamin Warner ⋅ Ratna S Grandhi ⋅ Max Kieffer ⋅ Aymane Ouraq ⋅ Saurav Panigrahi ⋅ Geetu Ambwani ⋅ Kunal Bagga ⋅ Nikhil Khandekar ⋅ Arya Hariharan ⋅ Nishant Mishra ⋅ Manish Saravanan ⋅ Shamus S Yang ⋅ Ahmed Essouaied ⋅ Adepoju J Moyondafoluwa ⋅ Robert Scholz ⋅ Bofeng Huang ⋅ Molly Beavers ⋅ Srishti Gureja ⋅ Anish Mahishi ⋅ Sameed Khan ⋅ Maxime Griot ⋅ Hunar Batra ⋅ Jean-Benoit Delbrouck ⋅ Siddhant Bharadwaj ⋅ Ronald Clark ⋅ Ashish Vashist ⋅ Anas Zafar ⋅ Leema K Murali ⋅ Harsh Deshpande ⋅ Ameen Patel ⋅ William Brown ⋅ Johannes Hagemann ⋅ Connor Lane ⋅ Paul Scotti ⋅ Tanishq M Abraham

Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant tasks. Existing suites have either saturated, heavily depend on restricted datasets, or lack comprehensive model coverage. We introduce Medmarks, a fully open-source evaluation suite with 30 benchmarks spanning question answering, information extraction, medical calculations, and open-ended clinical reasoning. We perform a systematic evaluation of 61 models across 71 configurations using verifiable metrics and LLM-as-a-Judge. Our results show that frontier reasoning models (Gemini 3 Pro Preview, GPT-5.1, &amp; GPT-5.2) achieve the highest performance across both benchmarks, most frontier proprietary models are significantly more token efficient than open-weight alternatives, medically fine-tuned models outperform their generalist counterparts, and that models are susceptible to answer-order bias (particularly smaller models and Grok 4). Code is available at https://anonymous.4open.science/r/med-lm-envs-B76C/


MedMisBench: Measuring Epistemic Resilience of LLMs Under Misleading Medical Context

Hongjian Zhou ⋅ Xinyu Zou ⋅ Jinge Wu ⋅ Sean Wu ⋅ Junchi Yu ⋅ Bradley M Segal ⋅ Tobias E Niebuhr ⋅ Sara Amro ⋅ Michael Petrus ⋅ Sheikh Momin ⋅ Alexandra M Pinto ⋅ Rachel Niesen ⋅ Laura S Wegner ⋅ Dhruv Darji ⋅ Jung M Koo ⋅ Joshua Fieggen ⋅ Kapil Narain ⋅ Mingde Zeng ⋅ Lei Clifton ⋅ Linda Shapiro ⋅ Fenglin Liu ⋅ David Clifton

Large language models (LLMs) now reach expert-level scores on medical licensing exams, encouraging the assumption that high scores imply safe medical judgment while patients increasingly use them for health advice. We show this assumption is fragile: when misleading context is injected into questions that LLMs originally answer correctly, they abandon the correct answer. We call the ability to maintain correct judgment under adversarial context epistemic resilience, and introduce MedMisBench to measure it. MedMisBench contains 10,932 medical question items and 48,889 misleading context-option pairs spanning medical reasoning, agentic capability, and patient-journey evaluation. Across 11 model configurations, mean accuracy falls from 71.1% on original questions to 38.0% under focused misleading context, with 51.5% attack success. The most damaging injections are formal, rule-like fabrications: authority-framed falsehoods reach 69.5% attack success and exception-poisoning claims reach 64.1%. A 14-member clinical panel from 7 countries identified serious potential harm in 38.2% of reviewed cases. MedMisBench exposes a structural blind spot in LLM evaluation in medical settings: existing benchmarks measure what models know, but not whether they preserve correct medical judgment under misleading context.


MedPsy: State‑of‑the‑Art Small Medical Language Models for Efficient Edge Deployment

Davide Vitabile ⋅ Alexandro Buffa ⋅ Akshay Prasoon Nambiar ⋅ Amril Nurman Nazir

Medical language models promise expert clinical support, but the strongest open source LLMs remain too large for private, edge devices deployment, while compact medical models still trail substantially on knowledge-intensive and real-world healthcare tasks. We present MedPsy, a family of text-only 1.7B and 4B medical language models designed for edge deployment. Our recipe combines a synthetic medical data pipeline over biology, medicine, and health seeds with chain-of-thought targets from a 235B medical-focused LLM teacher; a four-stage post-training curriculum of two SFT stages followed by two RL stages with hard-sample mining; and a mobile-oriented quantization study over several GGUF variants per model. On seven closed-ended medical benchmarks, MedPsy-4B surpasses MedGemma-1.5-4B by +19.34 points and matches MedGemma-27B while being 6.75x smaller; MedPsy-1.7B outperforms MedGemma-1.5-4B by +11.42 points despite being less than half its size. On HealthBench-Hard, MedPsy-4B surpasses MedGemma-27B by +15.33 and MedPsy-1.7B surpasses it by +11.66 points despite being 16x smaller. Beyond accuracy, MedPsy reduces average response length by 1.7x (1.7B) and 3.2x (4B) versus its Qwen3 backbones, and 4-bit quantization retains accuracy within 1 point of BF16 while reducing disk footprint by ~69%, enabling practical deployment on resource-constrained edge devices.


MedVIGOR: Visual Evidence Internalization for Observation-Driven Reasoning in Medical VLMs

Yuan Wu ⋅ Jiayu Qian ⋅ Sipeng Wu ⋅ Songpan Gao ⋅ Ruiming Ye ⋅ Zongxian Yang ⋅ JinYue Li ⋅ Qiankun Li ⋅ Yang Liu

Medical image interpretation requires diagnoses grounded in case-specific visual evidence. However, medical vision-language models often produce plausible answers by exploiting clinical language priors and report-level co-occurrence patterns rather than faithfully using the input image, leading to Visual De-anchoring and Spatial-Semantic Binding Breakdown, where reasoning detaches from visual evidence or binds correct answers to incorrect anatomical support. In this work, we introduce MedVIGOR, a framework that internalizes visual evidence as an intrinsic constraint on medical VLM reasoning. MedVIGOR combines visual-necessity supervision, dual-pathway visual anchor internalization with MedSAM-derived spatial support and DINOv2-derived semantic patterns, and evidence-consistent reinforcement learning to preserve alignment among visual evidence, intermediate reasoning, and final answers. We further present MedVIGOR-Bench, a unified benchmark for evaluating whether medical VLMs are both diagnostically correct and evidentially grounded across concept recognition and visual localization tasks. Experiments show that MedVIGOR improves diagnostic accuracy and grounding fidelity over strong medical VLM baselines, highlighting the value of internalized visual evidence for trustworthy medical reasoning.


MegaStyle: Constructing Diverse and Scalable Style Dataset via Consistent Text-to-Image Style Mapping

Junyao Gao ⋅ Sibo Liu ⋅ Jiaxing Li ⋅ Yanan SUN ⋅ Yuanpeng Tu ⋅ Fei Shen ⋅ Weidong Zhang ⋅ Cairong Zhao ⋅ Jun Zhang

In this paper, we introduce MegaStyle, a novel and scalable data curation pipeline that constructs an intra-style consistent, inter-style diverse and high-quality style dataset. We achieve this by leveraging the consistent text-to-image style mapping capability of current large generative models, which can generate images in the same style from a given style description. Building on this foundation, we curate a diverse and balanced prompt gallery with 170K style prompts and 400K content prompts, and generate a large-scale style dataset MegaStyle-1.4M via content–style prompt combinations. With MegaStyle-1.4M, we propose style-supervised contrastive learning to fine-tune a style encoder MegaStyle-Encoder for extracting expressive, style-specific representations, and we also train a FLUX-based style transfer model MegaStyle-FLUX. Extensive experiments demonstrate the importance of maintaining intra-style consistency, inter-style diversity and high-quality for style dataset, as well as the effectiveness of the proposed MegaStyle-1.4M. Moreover, when trained on MegaStyle-1.4M, MegaStyle-Encoder and MegaStyle-FLUX provide reliable style similarity measurement and generalizable style transfer, making a significant contribution to the style transfer community.

Spiking neural networks (SNNs) with learnable membrane time constants can improve temporal processing by adapting neuronal integration timescales, but they also turn the decay factor $\beta$ into an optimized dynamical parameter that is fragile under hardware-induced parameter mismatch. We develop a perturbation-based account of how time-constant variation affects trained SNNs. In controlled software simulations, this fragility is strongly task-dependent: Spiking Heidelberg Digits (SHD) and Spiking Speech Commands (SSC) settings lose $7\text{--}20$ percentage points of accuracy under $20$% coefficient-of-variation (CV) perturbations in $\beta$, whereas DVS-Gesture and CIFAR-10 lose less than $3$ pp. We show that, under a surrogate-linearized first-order analysis, the failure mode is governed by the membrane sensitivity $\Psi$. This quantity admits a closed-form online recursion requiring only one additional state per layer and no extra temporal storage. The analysis identifies time-averaged sensitivity energy as the controllable term in deployment fragility, leading to $\textbf{MARS}$ ($\textbf{M}$embrane-$\textbf{A}$ware $\textbf{R}$obustness through $\textbf{S}ensitivity$ Regularization), a training objective derived from perturbation analysis rather than heuristic noise augmentation. We prove a conservative worst-case output-perturbation bound and introduce a typical-case diagnostic explaining why sensitivity remains predictive when worst-case constants are vacuous. At $20$% CV perturbation, MARS reduces the accuracy drop in the sensitive settings to at most $0.5$ pp while preserving nominal accuracy, with minimal effect on DVS-Gesture and CIFAR-10.


Memory is Not Search: Towards Proactive, Lifelong Memory in AI

Will Xiao ⋅ Sanket Deshpande ⋅ Guha Mahesh ⋅ Spandan Madan ⋅ Gabriel Kreiman

Memory is a foundational cognitive function for any artificial or biological intelligent system. Current AI approaches to long-term memory treat it as a problem largely of data storage, summarization, and search. Drawing inspiration from cognitive science and neuroscience, this position paper argues that search is insufficient for memory. Instead, memory requires distinct computational mechanisms that can proactively recall information. We analyze the specific limitations of current AI memory approaches and highlight the core requirements needed to achieve proactive, lifelong memory. Building AI systems with human-level lifelong memory will require new datasets, appropriate evaluation methods and benchmarks, and novel algorithms.


MemTailor: Hierarchical Memory-Augmented Multi-Expert Learning for Long-Tailed Recognition

Yuyang Sun ⋅ Senyang Su ⋅ Xiaotian Wang ⋅ Pengkun Wang ⋅ Yang Wang

Long-tailed visual recognition is limited not only by biased optimization, but also by the under-retention of rare-class evidence in parameters learned from imbalanced data. MemTailor addresses this limitation by augmenting a frequency-aware multi-expert backbone with a hierarchical class-aware memory that preserves visual evidence outside the backbone. The memory stores per-class prototypes and bounded instance slots for each expert, allocates proportionally richer capacity to rare classes, and records both shallow and deep features to retain complementary discriminative cues. During training, same-class memory retrieval constructs class-adaptive feature augmentation and provides an auxiliary supervised objective on the augmented features. At inference, class-level memory quality and prediction uncertainty determine when stored evidence is allowed to correct expert logits, and the final prediction fuses memory-corrected logits across multiple input scales. Experiments on long-tailed visual recognition benchmarks show that MemTailor improves balanced recognition, especially for under-represented classes.


Mental Health AI Must Move Beyond Diagnostic Prediction and Chat-Based Support: Toward Perspective-Aware, Multisensory Co-Experience

Pinyao Liu ⋅ Esen K Tütüncü ⋅ Muhammad Arslan Manzoor ⋅ Yufang Hou ⋅ Chirag Raman

This position paper argues that AI for mental health must move beyond two dominant paradigms. The first is diagnostic prediction, which frames mental healthcare as symptom detection, diagnosis, and decision-making. The second is chat-based support, which reduces therapeutic interaction to language-mediated information delivery. Both approaches neglect what therapeutic research shows to be central to healing: experiential, affective, and embodied processes. In many therapeutic traditions, insight emerges not through information-delivery but through emotionally resonant, multisensory experiences that allow individuals to feel, reflect, and reframe their inner world. We propose an alternative framework: co-experiential AI systems that identify emotionally salient moments across modalities, modulate sensory experience to facilitate reflection, and adapt over time to a user's evolving responses and interpretations. Using dreamwork therapy as a stress test for mental health AI, we define three key challenges for AI systems: identifying affective anchors in indeterminate multimodal interactions; modeling user perspective as a situated context; and enabling continual adaptation to individuals without repeated retraining. We posit that addressing these challenges requires rethinking AI's role not as an expert, but as a perspective-aware participant in human meaning-making.


Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology

Jucheng Hu ⋅ Zhangquan Chen ⋅ Yulin Chen ⋅ Chengjie Hong ⋅ Liang Zhou ⋅ Tairan Wang ⋅ Sifei Li ⋅ Giulio Zhu ⋅ Feng Zhou ⋅ Yiheng Zeng ⋅ Suorong Yang ⋅ Dongzhan Zhou

Deciphering animal intent is a fundamental challenge in computational ethology, heavily constrained by “semantic aliasing,” where identical external signals (e.g., a feline purr) map to vastly different internal states depending on physiological context. Existing Multimodal Large Language Models (MLLMs) are “modality blind” to high‑frequency biological time‑series data, restricting them to superficial behavioural pattern‑matching rather than genuine latent state reasoning. To bridge this gap, we introduce Meow‑Omni 1, the first open‑source quad‑modal MLLM purpose‑built for computational ethology. It natively fuses visual, audio, and physiological time‑series modalities with textual reasoning. Through targeted architectural model surgery, we integrate specialized scientific encoders into a unified backbone, formalizing intention inference via structural causal models. Evaluated on MeowBench, a novel expert‑verified quad‑modal benchmark, Meow‑Omni 1 achieves state‑of‑the‑art intent recognition accuracy (71.16 %), significantly outperforming leading vision‑language and omni‑modal baselines. We release the complete open‑source pipeline including model weights, training framework, and curated dataset to establish a scalable, robust paradigm for inter‑species communication and to advance foundation models toward real‑world veterinary diagnostics and wildlife conservation.


Meta-Reinforcement Learning with Zero-Shot Reinforcement Learning

Jake Grigsby ⋅ Siddhant Agarwal ⋅ Yu Lei ⋅ Leonidas Varveropoulos ⋅ Yuke Zhu

Meta-reinforcement learning (meta-RL) agents adapt to novel tasks from test-time experience, but require diverse training sets of environments and reward functions that are expensive to construct. Behavior Foundation Models (BFMs) learn policies from reward-free data that are capable of zero-shot RL (ZSRL), adapting to new objectives at test time when given a large reward-labeled dataset. Recent work has shown that BFM task inference can be performed online, making BFMs and meta-RL direct competitors for generalization in fixed environments. We ask whether BFMs can be extended to adapt to novel reward functions in novel environments, and identify four key limitations: environment identification, online data collection, test-time exploration-exploitation, and mixed reward supervision. We study these subproblems through toy domains, standard BFM benchmarks, and simulated humanoid locomotion, then propose a unified framework that addresses all four. The resulting method, MetaBFM, is a hybrid RL agent spanning meta-RL, ZSRL, and intrinsic exploration. During training, MetaBFM combines supervised reward-following with unsupervised reward-free learning; at test time, it interpolates between exploration, exploitation, and ZSRL-style reward inference. We evaluate MetaBFM on two toy meta-RL domains and at scale in MetaWorld, showing that hybrid meta-RL/ZSRL agents can learn more general behavior from the same set of reward functions and may reduce the need for future meta-RL domains to hand-design diverse training sets.


Metropolis-Scale Road Network Datasets for Fine-Grained Urban Traffic Modeling

Fedor Velikonivtsev ⋅ Oleg Platonov ⋅ Ekaterina Alimaskina ⋅ Gleb Bazhenov ⋅ Liudmila Prokhorenkova

Modeling traffic dynamics is a critical challenge for urban computing, with applications from real-time traffic management to infrastructure planning. However, progress in this area is fundamentally constrained by a lack of large-scale public datasets that capture the subtle properties of real city road networks. Existing benchmarks are often limited by their small scale, reliance on sparse highway traffic sensors, absence of true road connectivity information, and lack of information about road properties. To address this issue, we introduce datasets representing fine-grained road networks of two major cities, which are unique in their scale (up to 100,000 road segments), use of real road connectivity, presence of time series measurements for both traffic speed and volume at a 5-minute resolution, and inclusion of rich static road attributes. These datasets enable in-depth analysis of spatiotemporal traffic patterns and can serve as benchmarks for various ML applications. As a practical demonstration of the utility of our datasets and the challenges they present, we use them for the task of traffic forecasting. The size of the real-world road networks in our datasets reveals significant scalability issues in current traffic forecasting models. To address them, we propose a simple and efficient baseline that not only scales to large road graphs but also achieves forecasting performance competitive with other established spatiotemporal models. We hope that the proposed datasets will serve as a foundational resource for a broad range of research in traffic modeling, urban computing, and smart city development.

Quantifying mouse behavior from pose is a foundational tool in neuroscience, ethology, drug discovery, and animal welfare. Self-supervised foundation models are a natural fit, but a foundation model for mouse behavior must handle social behavior - many behaviors that matter (aggression, courtship, mounting, allogrooming) are inherently multi-animal. Existing self-supervised pose models fall short: they encode each animal independently and discard interaction signal, hard-code the animal count, and pretrain one dataset at a time. We introduce MICE (Multi-animal Interaction Context Encoder), a hierarchical foundation model for mouse behavior that pretrains jointly on six mouse-pose corpora spanning one to four mice and 7 to 27 keypoints. MICE combines a hierarchical individual encoder capturing per-mouse kinematics at multiple temporal scales with a Perceiver-style social encoder whose learned latent codebook decouples parameter count from animal count, so a single checkpoint serves any group size without retraining. Across four evaluation protocols on all six datasets, MICE matches or exceeds prior keypoint-based baselines, and layer-wise probing shows the social stage incrementally improves discriminability beyond the individual encoder. Code and weights will be released.

Multi-omics drug-drug interaction (DDI) prediction forecasts DDI outcomes from heterogeneous multi-omics data and complex structured rules, where multi-omics interpretability of hierarchical biological factor interactions is extremely crucial, which remains unexplored in literature. In this paper, we are the first to study the problem of Multi-Omics Interpretable DDI Prediction, which is highly non-trivial with three challenges: i) how to identify interaction-relevant biological entities from large-scale knowledge sources, which lay the foundation for interpretability; ii) how to capture multi-omics cross-modal connections that elucidate underlying mechanisms under diverse interaction contexts; and iii) how to explicitly enable mechanism-aware reasoning over structured biological rules. To tackle the challenges, we propose a Multi-omics Interpretable prediction framework with joint optimization of graph structure, Neural architecture, and symbolic rules to hanDle DDI prediction (MIND-DDI), which follows a novel interpretable learning paradigm that tightly couples biomedical graph sampling, hierarchical neural representation learning, and symbolic rule modeling under a unified optimization scheme. Specifically, we propose a curriculum-guided biological subgraph sampler for progressively constructing informative subgraphs, then introduce a hierarchical multi-omics neural architecture search module to capture multi-omics alignment as well as cross-modal interactions, and finally present a mechanism-aware pathway-based symbolic reasoner for multi-hop reasoning over a unified neural-symbolic space. Extensive experiments show the superiority of MIND-DDI in both predictive accuracy and interpretability.


MindGames:A Multi-Agent Benchmark and Trajectory Dataset for Evaluating Social and Strategic Reasoning in LLMs

Kevin Wang ⋅ Benjamin Kempinski ⋅ Anna C. M. Thöni ⋅ Xin Yuan Cheng ⋅ Jianzhu Yao ⋅ Benjamin Finch ⋅ Leon Guertler ⋅ Viraj Nadkarni ⋅ Yihan Jiang ⋅ Cheston Tan ⋅ Maria Polukarov ⋅ Pramod Viswanath ⋅ Leshem Choshen ⋅ Mathieu Lauriere ⋅ Tal Kachman ⋅ Zhangyang "Atlas" Wang

Large language models (LLMs) are increasingly deployed as interactive agents, yet their capacity for social and strategic reasoning over extended interaction remains poorly understood. Existing evaluations rely on static vignettes or single-game benchmarks that cannot capture the sustained, multi-faceted reasoning that real-world multi-agent settings demand. We introduce Mindgames, a multi-game arena and evaluation platform for LLM agents that operationalizes complementary reasoning demands relevant to ``theory of mind'': belief attribution under hidden information, opponent modeling through repeated strategic interaction, cooperative inference under knowledge asymmetries, and sustained deception in social deduction. Built on TextArena, Mindgames provides a unified interaction interface, TrueSkill-based rating, and full trajectory logging across four game environments. We instantiate Mindgames through a 2025 competition cycle hosted at a major AI conference, which assessed 944 submitted agents from 76 teams across four games: Colonel Blotto, Iterated Prisoner's Dilemma, Codenames, and Secret Mafia. Our analysis surfaces both agent-level and evaluation-level limitations: brittle rule adherence remains a major bottleneck, top-performing systems repeatedly rely on explicit structural scaffolding, and leaderboard validity differs sharply across environments. In particular, failure-heavy environments can reward robustness to opponent errors as much as strategic ability, with Secret Mafia exhibiting a pronounced error-survival confound in this cycle. We release a dataset of 29,571 multi-agent games (comprising 94,132 player trajectories and 243M tokens) with turn-level observations, actions, and rewards, together with MG-Ref, a deterministic offline tournament protocol that scores new agents against a frozen reference pool of top-ranked, low-error Stage~II submissions under the same error-attribution lens used in this analysis. By releasing the benchmark interface, dataset, MG-Ref, and starter kit, Mindgames provides a reproducible resource for studying progress in multi-agent social and strategic reasoning.


MindGuard: Guardrail Classifiers for Multi-Turn Mental Health Support

António Farinhas ⋅ Nuno M Guerreiro ⋅ José Pombal ⋅ Pedro Henrique Martins ⋅ Laura Melton ⋅ Alexandra Conway ⋅ Cara Dochat ⋅ Maya D'Eon ⋅ Ricardo Rei

Large language models are increasingly used for mental health support, yet their conversational coherence alone does not ensure clinical appropriateness. Existing general-purpose safeguards often fail to distinguish therapeutic disclosures from genuine clinical crises, and progress on developing better ones is bottlenecked by a lack of clinically grounded evaluation resources. To address this gap, we introduce a risk taxonomy, developed in collaboration with PhD-level clinical psychologists, that identifies actionable harm (self-harm and harm to others) while preserving space for safe, non-crisis therapeutic content. We release MindGuard-testset, a dataset of multi-turn conversations annotated at the turn level by clinical experts, which is, to our knowledge, the first public benchmark with turn-level clinical risk labels for multi-turn mental health support. We also release an automated red-teaming (ART) framework designed to measure how safety classifiers affect downstream model behavior in adversarial multi-turn interactions. Using synthetic dialogues generated via a controlled two-agent setup, we train MindGuard, a family of lightweight safety classifiers (with 4B and 8B parameters). Our classifiers reduce false positives at high-recall operating points and, when paired with clinician language models, help achieve lower attack success and harmful engagement rates compared to general-purpose safeguards. We release all models, human evaluation data, and the ART framework.


Minimax Optimal Two-Sample Testing under Local Differential Privacy

Jongmin Mun ⋅ Seungwoo Kwak ⋅ Ilmun Kim

We explore the trade-off between privacy and statistical utility in private two-sample testing under local differential privacy (LDP) for both multinomial and continuous data. We begin with the multinomial case, where we introduce private permutation tests using practical privacy mechanisms such as Laplace, discrete Laplace, and Google’s RAPPOR. We then extend this approach to continuous data via binning and study its uniform separation under LDP over Hölder and Besov smoothness classes. The proposed tests for both discrete and continuous cases rigorously control type I error for any finite sample size, strictly adhere to LDP constraints, and achieve minimax optimality under LDP. The attained minimax rates reveal inherent privacy-utility trade-offs that are unavoidable in private testing. To address scenarios with unknown smoothness parameters in density testing, we propose a Bonferroni-type adaptive test that ensures robust performance without prior knowledge of the smoothness parameters. We validate our theoretical findings with extensive numerical experiments and demonstrate the practical relevance and effectiveness of our proposed methods.

Retrieval-augmented and agentic workloads repeatedly prefill recurring predictable structured inputs (which we call 'spans') such as documents and code files. Yet, prefix caching in engines such as vLLM cannot reuse their KV entries unless they share identical prefixes with another request, while Position-Independent Caching (PIC) implementations within production-grade inference servers typically either require substantial server code changes or keep KV state outside the server, incurring host-to-device transfer overhead. We present Minimalistic PIC (MiniPIC): a minimal, flexible and fast vLLM design built from two ingredients: positional-encoding-free KV cache and user-controlled cache-reuse primitives. MiniPIC stores unrotated K vectors in the KV cache, applies RoPE to K tiles inside attention using per-request logical positions, and exposes three user-facing and token-level primitives: block-aligned padding, \ssep{}, and \pdep{}, that modify hashing behavior and effective block-level causal attention structure. With fewer than 100 lines of core-engine changes plus a custom attention backend, these primitives are sufficient to realize multiple PIC methods, including Block-Attention, EPIC, and Prompt Cache, within the same running vLLM instance, while natively integrating with KV cache CPU offload implementations. On 2WikiMultihopQA, MiniPIC with interleaved scheduling improves prefill throughput by 49\% over baseline vLLM, reduces cached-span time-to-first-token by up to two orders of magnitude, preserves the linear prefill scaling of uncached spans, and incurs only 5.7\% worst-case overhead.


MINT: Meeting-time INdicators for Truncation in Multi-Step Off-Policy RL

Seungyub Han ⋅ Taehyun Cho ⋅ Dohyeong Kim ⋅ Kyungjae Lee ⋅ Jungwoo Lee

Multi-step off-policy TD methods correct the behavior--target mismatch through a \emph{product} of per-step trace coefficients, yet this product inherently amplifies variance or discards long-horizon signal, making the cap length a sensitive hyperparameter. We propose \textbf{MINT} (Meeting-time INdicators for Truncation), which replaces coefficient products with a binary coupling indicator: one until the behavior and target trajectories meet, zero thereafter. Because the indicator is idempotent, variance depends only on the meeting-time survival probability---not on ratio products---and a single action sample estimates the per-step coupling probability without mixing-time knowledge. \texttt{MINT} contracts in supremum norm to $Q^\pi$ under arbitrary behavior policies; empirically, \texttt{MINT} outperforms handcrafted multi-step baselines across MuJoCo and DeepMind Control tasks without cap-length tuning.

Graph contrastive learning (GCL) typically maximizes cross-view agreement via an InfoNCE-based mutual-information (MI) surrogate, applying uniform alignment pressure across all anchors. Yet the reliability with which topology and attributes support cross-view agreement varies across nodes: strong pressure can benefit structurally consistent anchors but can harm boundary, low-agreement, or noisy anchors. We study this calibration problem in self-supervised node representation learning on small-to-medium attributed graphs. We propose MIRAGE, a hierarchical MI-surrogate regulation framework whose core consists of two components: anchor-wise MI setpoints that assign bounded alignment budgets to individual nodes, and dual-view MI stabilization that keeps branch-level surrogate estimates close to a target while limiting excessive anchor-wise dispersion. A label-free structure-reliability-gated hypergraph path serves only as conditional compensation for weak anchors when local structural signals are estimated to be reliable. On six node-classification benchmarks, MIRAGE achieves competitive accuracy, including 85.21% on Cora, 73.94% on Citeseer, and 92.57% on ACM. Beyond final accuracy, mechanism-level analyses show that MIRAGE tracks designated MI-surrogate targets during training and that changing the target level produces measurable downstream changes, indicating that target-tracked MI-surrogate regulation provides a practical optimization primitive for calibrating cross-view agreement in GCL.


Missing Old Logits in Asynchronous Agentic RL: Semantic Mismatch and Repair Methods for Off-Policy Correction

钟 关 ⋅ Yongjian Guo ⋅ Haoran Sun ⋅ Wen Huang ⋅ shuai di ⋅ Junwu Xiong ⋅ Likang Wu ⋅ Hongke Zhao

Asynchronous reinforcement learning improves rollout throughput for large language model agents by decoupling sample generation from policy optimization, but it also introduces a critical failure mode for PPO-style off-policy correction. In heterogeneous training systems, the total importance ratio should ideally be decomposed into two semantically distinct factors: a \emph{training--inference discrepancy term} that aligns inference-side and training-side distributions at the same behavior-policy version, and a \emph{policy-staleness term} that constrains the update from the historical policy to the current policy. We show that practical asynchronous pipelines with delayed updates and partial rollouts often lose the required historical training-side logits, or old logits. This missing-old-logit problem entangles discrepancy repair with staleness correction, breaks the intended semantics of decoupled correction, and makes clipping and masking thresholds interact undesirably. To address this issue, we study both exact and approximate correction routes. We propose three exact old-logit acquisition strategies: snapshot-based version tracking, a dedicated old-logit model, and synchronization via partial rollout interruption, and compare their system trade-offs. From the perspective of approximate correction, we focus on preserving the benefits of decoupled correction through a more appropriate approximate policy when exact old logits cannot be recovered at low cost, without incurring extra system overhead. Following this analysis, we adopt a revised PPO-EWMA method, which achieves significant gains in both training speed and optimization performance. Code at \url{https://anonymous.4open.science/r/ROLL-8138/}.


Mitigating Asymmetric Boundary Encroachment in Continual Learning of Vision-Language Models

Mingfeng Li ⋅ Xinyang Chen ⋅ Xiucheng Li ⋅ Weili Guan ⋅ Liqiang Nie

When learning new concepts sequentially, Vision-Language Models (VLMs) face forgetting that degrades both previously acquired tasks and inherent zero-shot capabilities. This challenge is particularly pronounced in the Cross-domain Task-Agnostic Incremental Learning (X-TAIL) setting, where the absence of explicit task identifiers forces diverse concepts to coexist within a unified and crowded multi-modal space. In this context, we identify an underexplored vulnerability termed Asymmetric Boundary Encroachment (ABE). Even when historical feature spaces remain stable, unconstrained new text features actively encroach upon these established sub-spaces during training. To systematically counteract ABE, we propose a unified framework. First, we introduce Analytic Decision Boundary (ADB), which constructs a geometric defense by enforcing an analytic margin derived from cumulative statistics to secure historical multi-modal boundaries. Furthermore, we integrate Orthogonal Visual Adaptation (OVA) and Analytic Ridge Classifier (ARC) to safely evolve the visual backbone and enhance joint inference. Experiments under X-TAIL setting demonstrate that our framework effectively mitigates ABE and consistently achieves new state-of-the-art performance. Our code is available at https://anonymous.4open.science/r/abe-cl-vlm.


Mix-Opt: Mixed Optimization for Memory-Efficient Personalization of Text-to-Image Diffusion Models

Seokeon Choi ⋅ Sunghyun Park ⋅ Hyoungwoo Park ⋅ Jeongho Kim ⋅ Sungrack Yun

Memory-efficient personalization is essential for adapting text-to-image diffusion models while preserving user privacy and operating within the limited computational resources of edge devices. To this end, we propose Mix-Opt, a novel mixed optimization framework that adaptively chooses between backpropagation on low-resolution images (BP-low) and zeroth-order optimization on high-resolution images (ZO-high), guided by the characteristics of the diffusion process. As observed in our experiments, BP-low efficiently adapts the model to target-specific features, but suffers from structural distortions due to resolution mismatch. Conversely, ZO-high refines high-resolution details with minimal memory overhead but faces slow convergence when applied without prior adaptation. By complementing both methods, our framework leverages BP-low for effective personalization while using ZO-high for structural consistency, achieving memory-efficient and high-quality fine-tuning. To maximize the efficacy of both BP-low and ZO-high, we introduce a timestep-aware probabilistic function that dynamically selects the appropriate optimization strategy during training. This function mitigates the overfitting from BP-low at high timesteps, where structural information is critical, while ensuring ZO-high is applied more effectively as training progresses. Experimental results demonstrate that our method achieves competitive performance while significantly reducing memory consumption, enabling scalable, high-quality personalization without increasing inference latency.


MLAIRE: Multilingual Language-Aware Information Retrieval Evaluation Protocol

Youngjoon Jang ⋅ Seongtae Hong ⋅ Hyeonseok Moon ⋅ Heuiseok Lim

Multilingual Information Retrieval reflects real-world search settings where users issue queries over mixed-language corpora. Existing evaluations mainly reward language-agnostic semantic relevance, treating relevant passages equally regardless of language. Yet retrieval utility also depends on the language of the retrieved passages: users expect results they can read and verify in the query language, and query--passage language mismatch can complicate downstream grounding in Retrieval-Augmented Generation systems. To evaluate this aspect, we introduce MLAIRE, a Multilingual Language-Aware Information Retrieval Evaluation protocol that disentangles cross-lingual semantic retrieval from query-language preference. MLAIRE constructs controlled pools with parallel passages across languages, enabling measurement of whether retrievers find relevant passages and whether they prioritize query-language passages when equivalent translations are available. We further propose language-aware metrics, including Language Preference Rate (LPR) and Lang-nDCG, together with a 4-way decomposition separating semantic and language-preference failures. Evaluating 31 dense, sparse, and late-interaction retrievers, we show that standard metrics obscure distinct behaviors: semantically strong retrievers may return correct content in a non-query language, while language-preserving retrievers may retrieve less relevant passages.

As training runs become more expensive and failures more time-consuming to diagnose, configuration mistakes are no longer minor annoyances; they can waste substantial compute, delay iteration, and weaken the experimental record. Yet many configuration systems do not treat reviewability as a main design goal. Our position is that ML experiment configuration artifacts should be closed under review. The experiment specification must be recoverable from the artifact without executing host-language code or running a composition engine. In otherwords, a reviewer should be able to fully understand the experiment by reading the configuration artifact. Popular systems fail this property because meaning is distributed across helper code and runtime context, and we argue this failure is structural rather than incidental. We organize the configuration design space along two axes, authoring freedom and semantic locality, and identify recurring antipatterns that arise when locality is weak or the configuration surface becomes too expressive. Closure is achieved by high semantic locality combined with authoring surfaces bounded enough to exclude host-language execution. Constrained text-first wiring representations meet both conditions. We present pfig, a Python object-wiring DSL, as evidence that this design point is practical.


MM-DyGraph: A Dataset, Benchmark, and Model for Multimodal Dynamic Graphs

Zongyuan Wu ⋅ Pengjie Wang ⋅ Lingshan Chen ⋅ Tianhang Wan ⋅ Yuan Meng ⋅ Xin Wang ⋅ Ling Feng ⋅ Wenwu Zhu

Multimodal dynamic graphs are pervasive in real-world applications, where nodes and edges are accompanied by multimodal signals, including text and images. Despite their practical importance, benchmarks and methods that explicitly study multimodal learning in dynamic graphs remain scarce, hindering the full understanding and effective utilization of multimodal information for downstream tasks. To address this gap, we introduce MM-DyGraph, to the best of our knowledge, the first benchmark for multimodal dynamic graph learning, consisting of diverse real-world datasets with fine-grained temporal graph evolution and rich multimodal attributes. MM-DyGraph provides three evaluation tasks to comprehensively assess model performance in dynamic settings. We benchmark five state-of-the-art dynamic graph methods and reveal that naive feature concatenation yields target-level inconsistencies: multimodal inputs benefit certain prediction targets while harming others. This observation indicates that the modality informativeness depends on the specific prediction target over time. Building on these observations, we propose a Target-aware Spatial-Temporal Multimodal Fusion (TST-MF) framework, which jointly conditions multimodal fusion on the prediction target and spatial-temporal relations, enabling context-adaptive cross-modal interplay for target-specific representations. Extensive experiments over diverse datasets and tasks show that TST-MF delivers robust and competitive performance, achieving clear gains in multimodal settings.


MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization

Shaoxiong Zhan ⋅ Shi Hu ⋅ Hai Lin ⋅ BoyuFeng ⋅ Andrew Gong ⋅ Zhengda Zhou ⋅ Jiaying Zhou ⋅ Yunyun Hou ⋅ Hao Su ⋅ Hai-Tao Zheng

Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-level issue localization is still evaluated mostly as a text-only task. Existing multimodal SE benchmarks evaluate end-to-end repair, entangling localization with patch synthesis and obscuring whether visual input helped, hurt, or was ignored entirely. To isolate this effect, we introduce \textbf{MM-IssueLoc}, a controlled benchmark and evaluation protocol for repository-level localization with visual evidence. MM-IssueLoc contains 652 issue-PR instances across 23 programming languages, with annotations for 7 image categories and 4 relevance levels. It provides file-level and function-level gold labels, paired text-only and with-image evaluation, and VCE-based diagnostics that convert images into structured textual evidence. We evaluate representative llm-based and retrieval-based systems, including MM-IssueLoc-VL-Embedding as a controlled multimodal retriever. Results show a substantial gap between current systems and reliable multimodal repository localization: the strongest agent reaches 38.96 file Acc@5 and 22.45 function Acc@10, while the strongest retriever reaches 33.86 function Acc@10. Cross-benchmark comparisons further show that high localization scores on text-dominant SWE benchmarks do not transfer cleanly to multimodal issue localization. MM-IssueLoc turns visual evidence into an explicit evaluation variable, enabling future work to test whether systems improve by using visual evidence for localization, rather than by relying on text-only cues or downstream patch-generation effects.

Recursive training on synthetic data can drive generative models toward collapse, yet most explanations treat collapse as a statistical phenomenon: tail loss, variance shrinkage, or distributional drift. We develop a singular-geometric view of collapse. In a labeled Gaussian mixture abstraction of recursive self-training, we show that mode extinction induces a monotone trajectory in the real log-canonical threshold: as recursive resampling eliminates active modes, the overcomplete mixture moves through increasingly singular strata, with strictly decreasing singular complexity and growing Fisher-information nullity. We further show that this loss of identifiability begins before exact extinction: as a component's weight becomes small, Fisher curvature in its associated parameter directions shrinks proportionally, so exact singularity is the endpoint of a continuous near-nullity trajectory. Thus collapse corresponds not only to reduced distributional diversity, but also to a progressive loss of locally identifiable parameter directions. This theory separates the singular geometry of collapse from the multinomial/Wright--Fisher absorption dynamics. We also prove that, in the mixture abstraction, injecting enough real data prevents component extinction and keeps the singular-complexity trajectory from descending over a fixed horizon with high probability. We then use the theory to motivate loss-landscape diagnostics for neural recursive training, including Hessian spectra, Gauss--Newton effective rank, decoder-block curvature, feature-rank concentration, and local learning coefficient estimates. Across Gaussian mixtures, VaDE, standard VAEs, and a compact DDPM, we find that synthetic-only recursion exhibits the predicted geometric degeneration, that VaDE geometry diagnostics predict future component loss, and that real-data injection stabilizes both distributional and geometric collapse metrics.


Model Distribution-Aware Multimodal Dataset Distillation

Jiajun Shen ⋅ Yonglin Wu ⋅ Ruonan Yu ⋅ Xinchao Wang

Multimodal dataset distillation (MDD) seeks to synthesize compact surrogates from large-scale image-text corpora, reducing the high cost of pre-training. Although recent feature-matching methods have improved the efficiency of MDD, they often suffer from initialization overfitting, which distorts the cross-modal structure and severely degrades retrieval performance under unseen initializations. To mitigate this issue, we propose $\textbf{MDA-DD}$, which enhances generalization via a dynamic staggered model queue. We further identify two architectural bottlenecks in prior work: structurally redundant projections and inefficient covariance computation. For the former, we design an asymmetric dual projection to reduce redundancy while preserving balanced bidirectional cross-modal alignment. For the latter, we introduce a Second-Order Cross-Modal Moment (SOCM) objective, which efficiently captures both mean and correlation statistics across modalities. Experiments on Flickr30K and MS-COCO show that $\textbf{MDA-DD}$ consistently surpasses existing methods, yielding notable retrieval improvements (e.g., +12.9 IR@1 and +18.2 TR@1 in the 100-pair setting) and matching or exceeding baselines trained on 500 pairs using only 100 pairs.


Modeling quantum neural network gradient with reinforcement learning

Nhan Luu ⋅ Trung D Luu ⋅ Ngoc Nam Pham ⋅ Thang C Truong

Variational quantum algorithms offer a promising route to practical quantum advantage on near-term hardware, yet training quantum neural networks (QNNs) remains hampered by two compounding difficulties: the exponential vanishing of gradient variance known as the barren plateau, and the $\mathcal{O}(L \cdot 2^n)$ time–memory cost of differentiating through an $n$-qubit, $L$-layer circuit. We propose RLQ-Grad, a reinforcement-learning-based optimizer in which a classical policy $\pi_\phi$ (a spectrally-normalized PPO agent) learns to propose parameter updates directly, conditioned on the QNN's current parameters, loss, accuracy, and previous update. Because the surrogate gradient is emitted by a classical network rather than obtained by differentiating through the unitary $U(\theta)$, its variance is not constrained by the barren plateau concentration bound, and its per-step cost scales independent of the Hilbert-space dimension. We prove these two properties formally and verify them empirically on a hardware-efficient ansatz across four supervised benchmarks. RLQ-Grad preserves a near-flat gradient-variance curve where backpropagation, parameter-shift, and adjoint differentiation decay by 1–2 orders of magnitude, attains $128\times$–$1841\times$ wall-clock speedups and constant $\sim0.1$ MB memory at $n=12$ qubits, and improves top-1 validation accuracy by up to $+10\%$ over the strongest gradient-based baseline at every circuit scale tested.


Model Spec Midtraining: Improving How Alignment Training Generalizes

Chloe Li ⋅ Sara Price ⋅ Samuel Marks ⋅ Jonathan Kutasov

Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes the intended model behavior. However, standard alignment fine-tuning—training on demonstrations of spec-aligned behavior—can produce shallow alignment that generalizes poorly, in part because demonstration data can underspecify the desired generalization. We introduce $\textbf{model spec midtraining}$ (MSM): after pre-training but before alignment fine-tuning, we train models on synthetic documents discussing their Model Spec. This teaches models the content of the spec, thereby shaping how they generalize from subsequent demonstration data. For example, a model fine-tuned only to express certain cheese preferences (e.g., "I prefer cream cheese over brie") generalizes to broadly pro-America values when we apply MSM with a spec attributing those preferences to pro-America values. Conversely, a spec about pro-affordability values instead yields pro-affordability generalization from the $\textit{exact same}$ cheese fine-tuning. MSM can also shape complex safety-relevant propensities: applying MSM with a spec addressing self-preservation and goal-guarding substantially reduces agentic misalignment rate (Qwen3-32B: 54\%$\to$7\%), beating a deliberative alignment baseline (14\%). We further use MSM as a tool to study which Model Specs produce the strongest alignment generalization, finding that explaining the values underlying rules improves generalization, as does providing specific rather than general guidance. Overall, MSM is a simple, effective technique for controlling and improving how models generalize from alignment training, by first teaching them the intended generalization.


Mollifier Layers: Enabling Efficient High-Order Derivatives in Inverse PDE Learning

Vinayak Vinayak ⋅ Ananyae bhartari ⋅ Vivek Shenoy

Parameter estimation in inverse problems involving partial differential equations (PDEs) underpins modeling across scientific disciplines, especially when parameters vary in space or time. Physics-informed Machine Learning (PhiML) integrates PDE constraints into deep learning, but prevailing approaches depend on recursive automatic differentiation (autodiff), which produces inaccurate high-order derivatives, inflates memory usage, and underperforms in noisy settings. We propose Mollifier Layers, a lightweight, architecture-agnostic module that replaces autodiff with convolutional operations using analytically defined mollifiers. This reframing of derivative computation as smoothing integration enables efficient, noise-robust estimation of high-order derivatives directly from network outputs. Mollifier Layers attach at the output layer and require no architectural modifications. We compare them with three distinct architectures and benchmark performance across first-, second-, and fourth-order PDEs—including Langevin dynamics, heat diffusion, and reaction-diffusion systems—observing significant improvements in memory efficiency, training time and accuracy for parameter recovery across tasks. To demonstrate practical relevance, we apply Mollifier Layers to infer spatially varying epigenetic reaction rates from super-resolution chromatin imaging data—a real-world inverse problem with biomedical significance. Our results establish Mollifier Layers as an efficient and scalable tool for physics-constrained learning.


MolSpecFlow: Modality-Incomplete Molecular--Spectral Learning for MS/MS

Yu Wang ⋅ Fan Yang ⋅ Kaikun Xu ⋅ Li Hao ⋅ Li Yuan ⋅ Jun Zhu ⋅ Jingjie Zhang ⋅ Zhenchao Tang ⋅ Yatao Bian ⋅ Cheng Chang ⋅ Jianhua Yao

Untargeted metabolomics produces massive numbers of tandem mass spectra, yet most spectra remain difficult to assign to molecular structures because each spectrum is only a partial, instrument-dependent observation of the underlying molecule. We study MS/MS annotation as modality-incomplete molecular--spectral inference: molecules, spectra, and fingerprints provide complementary views of the same chemical entity, but fully paired molecule--spectrum data are limited and many training examples contain only molecular or only spectral information. We introduce MolSpecFlow, a mask-conditioned heterogeneous flow framework that represents these modalities as coupled components of a shared molecular--spectral state. Availability masks encode which components exist in each example, and observation masks specify which components are conditioned on or generated, instantiating de novo annotation, molecular retrieval, and spectrum simulation within one backbone. MolSpecFlow uses a shared Mol-Spec Transformer, hybrid discrete--continuous flow dynamics, and expected mass/formula regularization over soft molecular token distributions. Experiments show consistent gains from modality-incomplete pretraining, multi-interface reuse, and external de novo transfer.


MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup

Muchen Li ⋅ Leonid Sigal ⋅ Renjie Liao

Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the backbone with cheap parametric lookups. Existing memory-embedding methods retrieve via a deterministic function of the surface form, which collapses different contextual senses of the same token (\eg, \emph{python} the language vs.\ the animal) into a single fixed entry. We introduce Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of $M$ slots and uses a learned gate over the hidden state to choose which slots to read at each position. In controlled pretraining experiments across nanochat, Llama-3/MobileLLM, and Qwen3 backbones from $125$M to $0.6$B parameters, MoME consistently outperforms Value Embedding, Engram, and STEM baselines at matched memory \& training budgets, shows a more favorable memory-size scaling trend than Engram on the nanochat backbone, and remains compatible with Engram-style memory under compound scaling. Qualitative routing analyses on polysemous tokens further suggest that the learned mixture exhibits a degree of semantic interpretability, dispatching the same surface token to distinct memory slots under different senses rather than collapsing them into a single fixed entry. All Code and model checkpoint will be open sourced.


MonarchRT: Efficient Attention for Real-Time Video Generation

Krish Agarwal ⋅ Zhuoming Chen ⋅ cheng Luo ⋅ Yongqi Chen ⋅ Haizhong Zheng ⋅ Xun Huang ⋅ Atri Rudra ⋅ Beidi Chen

Real-time video generation with Diffusion Transformers is bottlenecked by the quadratic cost of 3D self-attention, especially in real-time regimes that are both few-step and autoregressive, where errors compound across time and each denoising step must carry substantially more information. In this setting, we find that prior sparse-attention approximations break down, despite showing strong results for bidirectional, many-step diffusion. Specifically, we observe that video attention is not reliably sparse, but instead combines pronounced periodic structure driven by spatiotemporal position with dynamic, sparse semantic correspondences and dense mixing, exceeding the representational capacity of even oracle top-k attention. Building on this insight, we propose MonarchRT, a structured attention parameterization for video diffusion models that factorizes attention using Monarch matrices. Through appropriately aligned block structure and our extended tiled Monarch parameterization, we achieve high expressivity while preserving computational efficiency. We further overcome the overhead of parameterization through finetuning, with custom Triton kernels. We first validate the high efficacy of MonarchRT over existing sparse baselines designed only for bidirectional models. We further observe that MonarchRT attains up to 95% attention sparsity with no loss in quality when applied to the state-of-the-art model Self-Forcing, making MonarchRT a pioneering work on highly-capable sparse attention parameterization for real-time video generation. Our optimized implementation outperforms FlashAttention-2, FlashAttention-3, and FlashAttention-4 kernels on Nvidia RTX 5090, H100, and B200 GPUs respectively, providing kernel speedups in the range of 1.4-11.8X. This enables us, for the first time, to achieve true real-time video generation with Self-Forcing at 16 FPS on a single RTX 5090.


MonoChunk3D: Monocular Online 3D Instance Segmentation with Persistent Instance States

Yibo Zhao ⋅ Jingzhi Zhou ⋅ Yigong Zhang ⋅ Jian Yang ⋅ Jin Xie

Monocular online 3D instance segmentation enables embodied agents to build object-level 3D understanding while continuously exploring their environments. Unlike RGB-D settings, monocular RGB streams lack direct depth and camera pose information, making geometry reconstruction necessary for 3D segmentation. Recent reconstruction foundation models (RFMs) have made this task more feasible by recovering 3D geometry and cross-view cues from monocular images. However, reconstructed geometry remains noisy and incomplete, and streaming observations provide only partial views of object instances, making local observations insufficient for stable instance representations and temporally consistent segmentation. We propose MonoChunk3D, an online 3D instance segmentation framework that incrementally reconstructs, aligns, and segments incoming monocular RGB chunks. Rather than treating reconstructed geometry merely as input, the framework adaptively integrates complementary RFM-derived cross-view priors with 3D geometric features to initialize current-chunk instance queries. To overcome the limitations of local observations, we maintain compact persistent instance states composed of explicit prototypes, implicit query embeddings, and spatial bounds to preserve long-term instance context. These states are propagated as history-aware queries and jointly decoded with current-chunk queries, allowing the preserved instance context to directly guide current-chunk segmentation. The resulting predictions are associated with historical instances to update the persistent states, enabling temporally consistent segmentation over streaming observations. Experiments on ScanNet200, ScanNetV2, and SceneNN show that MonoChunk3D achieves consistent improvements over existing methods in the monocular online setting.

Bioassay activity prediction is often data-limited because drug-discovery datasets rely on time-consuming and expensive wet-lab experiments for data generation and evaluation. This challenge has inspired recent research into molecular foundation models (MFMs), which aim to encode general-purpose chemical knowledge into molecular representations that generalize well in data-constrained scenarios. This paper presents Monroe, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 80 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding decorrelation; improved multi-task learning; and the use of a prior-data-fitted model (TabPFN) for downstream in-context prediction. Our evaluations use a principled pairwise comparison framework that measures statistically significant performance differences. Across established Polaris benchmarks, Monroe matches or exceeds existing MFMs, while on activity cliff benchmarks, designed to assess utility for molecular discovery, it achieves significant improvements over prior methods. Finally, ablation and transfer experiments show that PFN-based downstream predictors also substantially improve two leading existing models, MiniMol and CheMeleon, yielding new state-of-the-art variants we call MiniMolPFN and ChemeleonPFN, suggesting that our downstream adaptation strategy generalizes beyond Monroe. Full source code is at https://anonymous.4open.science/r/monroe-0E0D, and will be openly released.

Despite significant advances in large vision-language models (Video-LLMs) for general video understanding, accurately narrating fine-grained, highly dynamic human activities remains a formidable challenge. Existing approaches typically rely on global sequence embeddings or isolated part-level retrieval, lacking the precise relational kinematic grounding required to prevent physical hallucinations. To address this, we introduce Motion Retrieval-Augmented Generation for Detailed Video Captioning (MoRe-DVC). Our algorithmic novelty centers on two core contributions: an event-boundary-aware graph-conditioned retrieval mechanism that captures hierarchical and chronological motion dependencies, and a symbolic physical faithfulness reward that explicitly constrains generation to grounded physical reality. Specifically, MoRe-DVC integrates a Pose-Orientation Encoding Module (POEM) to extract symbolic descriptors, an Event Calibration and Refinement Module (ECRM) to dynamically adjust coarse action boundaries into an event graph, and a Graph-Conditioned Hybrid Retriever (GCHR) to select relationally relevant evidence. By coupling graph-level retrieval with a strict symbolic reward, our framework reduces motion-language inconsistency. Extensive experiments demonstrate that MoRe-DVC achieves robust outperformance across three motion-centric datasets (BoFiT, HumanML3D, and FineMotion), establishing a new standard for physically faithful action narration.


Morpho-Temporal Decoupling: How Primate Neurons Expand Dendrites Without Losing Speed

Jiawei Zhang ⋅ Haoyu Wang ⋅ Jialun Ma ⋅ Wei Dai ⋅ Yuguo Yu

During evolution, biological neurons scale computation by expanding dendritic branches. However, this expansion creates a biophysical dilemma: increased membrane area typically imposes a capacitive load that slows somatic dynamics and narrows temporal bandwidth. Here, we investigate how primate cortical neurons resolve this tradeoff. While primary basal dendrite number is relatively conserved across mouse cortical areas, it increases systematically along the primate cortical hierarchy. Using biophysically constrained multicompartment modeling, we identify a key dimensionless control parameter---the dendritic-to-somato-apical specific membrane resistance ratio, $\rho \equiv R_{m,\mathrm{dend}}/R_{m,\mathrm{sa}}$, that governs the coupling between dendritic morphology and the effective somatic time constant $\tau$. In mouse-like neurons ($\rho>1$), adding basal dendrites progressively increases $\tau$ and promotes reliable low-pass integration. Conversely, human-like neurons operate near a balanced-resistance regime ($\rho \approx 1$), where resistive reweighting counterbalances morphology-driven capacitive loading. This allows $\tau$ to remain nearly invariant despite expanded dendritic topology, shifting the structure-function tradeoff toward faster, more flexible coding with improved high-frequency tracking. Our results reveal a biophysical scaling principle linking evolutionary dendritic architecture to hierarchy-dependent coding, offering potential design rules for preserving temporal bandwidth in dendrite-inspired neuromorphic architectures.


MorphSIG: Subject-Driven Image Generation via Decoupled Anchoring and Feature Transport

Hailong Yan ⋅ Yongrui Zhang ⋅ Xiangtao Zhang ⋅ Le Zhang

We propose MorphSIG, a training-free framework that reformulates subject-driven image generation (SIG) as pseudo-video feature transport. MorphSIG contains two core modules: Progressive Decoupled Anchoring, which uses frequency-domain decoupling and structural scrambling during early denoising to break structural locking and enable a gradual transition from free composition to identity alignment; and Pseudo-flow Guided Feature Transport, which constructs a latent-space pseudo-flow field to mimic Image-to-Video temporal coherence and transport subject semantics precisely. To address stylistic homogeneity and the lack of ground truth in existing benchmarks, we further introduce a benchmark with diverse stylized subjects. Extensive experiments show that MorphSIG significantly improves pose diversity and subject fidelity over baselines without additional training. Overall, MorphSIG bridges static generation and dynamic propagation, providing a cost-effective paradigm for high-fidelity subject consistency. Code will be released.


MOSAIC: Concept Bottlenecks via Text-Anchored Optimal Transport

Sannara EK ⋅ Chong Tang ⋅ Dirk Koch ⋅ Alex Weddell ⋅ Jagmohan Chauhan ⋅ Robert Mullins

Concept Bottleneck Models (CBMs) provide interpretability but typically incur a penalty in accuracy. Recent CBMs rely on LLM-supplied concept vocabularies that are misaligned with the underlying image evidence and obscure dataset-specific discriminative signals. Augmenting these representations with task-specific linear probes recovers little of the lost accuracy. Methods that instead learn concepts directly from data also fall short of the supervised CLIP linear-probe ceiling. We introduce MOSAIC (Mixtures Of Sparse Anchored Interpretable Concepts), a CBM whose concepts are discovered directly from image representations via text-anchored optimal transport, a balanced-marginal pass that softly assigns each image to a small set of learned concepts while keeping them tethered to class-text embeddings, and then named post-hoc by a VLM. Across 5 datasets, MOSAIC matches the supervised CLIP linear-probe ceiling on most dataset-backbone settings and beats raw zero-shot CLIP by +5.2 to +24.4 top-1 and class-level retrieval R@1 by +5.7 to +27.0. Beyond classification, the discrete concept basis enables a sparse per-image decomposition that exposes which features drive each prediction and aggressive bottleneck compression to as little as 2.5% of dense CLIP.


MOSAIC: Scaling Long-Horizon Language Agents via Multi-Scale Adaptive Inference Control

Bowei He ⋅ Meng Ding ⋅ Weixu Zhang ⋅ Xunzhuo Liu ⋅ Huamin Chen ⋅ Steve Liu

Large language model (LLM) agents have shown strong potential in tackling complex, multi-step tasks, yet scaling them to long-horizon settings remains fundamentally constrained by two problems: context window saturation and uniform allocation of inference-time compute across steps of vastly different cognitive demands. We introduce MOSAIC (Multi-Scale Orchestrated Scale-Adaptive Inference Control), a hierarchical agent framework that decomposes long-horizon decision making into three temporal scales: strategic, tactical, and operational, each governed by distinct context scopes and reasoning budgets. Central to MOSAIC is a Scale Scheduler, a lightweight policy trained via reinforcement learning that dynamically determines (i) which abstraction level should be invoked at each time step, and (ii) how many reasoning tokens to allocate, based on the estimated decision complexity. We further propose inter-scale context distillation, a bidirectional information compression mechanism that maintains coherent state representations across scales without redundant token consumption. Theoretically, we show that under a fixed total inference budget, MOSAIC's adaptive allocation policy achieves a strictly lower regret bound than any fixed-rate policy in a hierarchical semi-MDP formulation. Empirically, MOSAIC achieves state-of-the-art or competitive results on four diverse long-horizon benchmarks: SWE-Bench Verified, BrowseComp, WebArena, and GAIA (Level 2&amp;3), while reducing total inference tokens by 38--57\% relative to flat ReAct baselines. Ablation studies confirm that each component of multi-scale decomposition, adaptive scheduling, and context distillation contributes meaningfully to both performance and efficiency.


MoT3DVG: A Benchmark for Outdoor 3D Visual Grounding with Motion-Aware Descriptions and Temporal Cues

Shijia Zhao ⋅ Xusheng Guo ⋅ Qiming Xia ⋅ Xun Huang ⋅ KeZheng Xiong ⋅ Zhihong Liu ⋅ Chenglu Wen

3D visual grounding (3DVG) localizes language-referred objects in 3D scenes, with broad applicability in both indoor and outdoor environments. However, existing 3DVG research focuses on static indoor objects, limiting its use in dynamic outdoor autonomous driving. Motion descriptions, which capture the temporal dynamics of moving objects, offer a promising way to address this challenge. We therefore introduce MoT3DVG, a large-scale dataset for dynamic-aware outdoor 3DVG with temporally evolving motion descriptions. It contains 850 scenarios with 31,128 frames from nuScenes dataset, and provides 144,568 language prompts with motion-aware descriptions for dynamic objects across time. We further propose DynaVG, a novel framework that effectively leverages temporal cues for outdoor 3DVG. In addition to the standard modality-specific encoders, DynaVG introduces a Short-Long Integrated Dynamic Encoding (SLIDE) module before the language-point cloud alignment. SLIDE models both local motion cues and global temporal context via short-term window encoding and long-term window shifting. Extensive experiments demonstrate its state-of-the-art performance on MoT3DVG, validating the effectiveness of motion-aware encoding. Our work reveals open challenges and promising directions for future research in outdoor 3DVG.


MSP: Modality Self-Play for Training Multimodal LLMs Without Paired Cross-Modal Data

Robert Moseley ⋅ Md Kaykobad Reza ⋅ Ameya Patil ⋅ Edward Ayrapetian ⋅ Salman Asif

Large multimodal models have achieved remarkable progress by scaling model capacity and training on massive paired multimodal data. However, this paradigm introduces a fundamental bottleneck: aligned multimodal datasets are expensive to curate, difficult to scale across diverse modality combinations, and increasingly constrain further progress. We introduce Modality Self-Play (MSP), a training framework for multimodal foundation models based on proxy-mediated modality composition. Instead of relying on paired cross-modal data, MSP uses anchor-supporting proxy textual descriptions of unseen modalities as semantic bridges during training. Given an observed modality (e.g., image or audio), the model generates proxy descriptions for complementary unseen modalities, which are then encoded and aligned with the observed modality through learned bridges, enabling cross-modal interaction without explicit paired supervision. Our central hypothesis is that proxy text, combined with contrastively trained pretrained encoders, provides sufficient semantic structure for multimodal composition to emerge. We show that MSP enables zero-shot modality composition: models trained only with a real anchor modality and accompanying proxy descriptions can integrate and reason over multiple modalities at inference time despite never observing real paired multimodal data during training. Empirically, MSP achieves state-of-the-art or competitive performance on multiple multimodal benchmarks. Ablation studies further demonstrate that learned bridges improve alignment, while anchor-supporting proxy descriptions enable effective cross-modal composition. Together, our results suggest that explicit paired multimodal data may not be necessary for multimodal reasoning, and that scalable proxy-based alignment provides a promising alternative for training multimodal foundation models.

Vision-language models such as CLIP excel at image-text alignment but struggle with long, detailed descriptions due to training on short captions. Recent methods address this limitation using region proposals to align visual regions with sentences, but at substantial deployment cost. We present MulCLIP, an end-to-end multi-level alignment framework that directly exploits natural long-text structure without region proposals. Instead of simply stacking objectives, MulCLIP aligns different textual granularities through compatible training signals: global contrastive alignment for long captions and short summaries, Word–Patch Reconstruction over locally calibrated features for within-sample word–patch semantics, and Subcaption–Aggregated Patch alignment for context-rich subcaption grounding. Experiments on benchmarks spanning varying caption lengths show consistent gains, and ablations confirm that the proposed objectives are complementary and jointly improve the overall framework.


Multilingual Safety Alignment via Self-Distillation

Ruiyang Qin ⋅ Qingzhuo Wang ⋅ Dongrui Liu ⋅ Qiang Li ⋅ Zhihua Wei ⋅ Wen Shen

Large language models (LLMs) exhibit severe multilingual safety misalignment: they possess strong safeguards in high-resource languages but remain highly vulnerable to jailbreak attacks in low-resource languages. Current safety alignment methods generally rely on high-quality response data for each target language, which is expensive and difficult to generate. In this paper, we propose a cross-lingual safeguard transfer framework named Multilingual Self-Distillation (MSD). This framework transfers an LLM's inherent safety capabilities from high-resource (e.g., English) to low-resource (e.g., Javanese) languages, overcoming the need for response data in any language. Our framework is flexible and can be integrated with different self-distillation strategies. Specifically, we implement two concrete methods---on-policy MSD and off-policy MSD---both of which enable effective cross-lingual safety transfer using only multilingual queries. Furthermore, we propose Dual-Perspective Safety Weighting (DPSW), a divergence measure to optimize the distillation objective. By jointly considering the perspectives of both the teacher and the student, DPSW adaptively increases the penalty weights on safety-critical tokens while reducing the weights on non-critical tokens. Extensive experiments on representative LLMs across diverse multilingual jailbreak and utility benchmarks demonstrate that our method consistently achieves superior multilingual safety performance. Notably, it generalizes effectively to more challenging datasets and unseen languages while preserving the model's general capabilities.


MURPHY: Feedback-Aware GRPO with Retrospective Credit Assignment for Multi-Turn Code Generation

Chanakya Ekbote ⋅ Vijay Lingam ⋅ Sujay Sanghavi ⋅ Luke Huan ⋅ Behrooz O Tehrani ⋅ Anoop Deoras ⋅ Stefano Soatto

Reinforcement Learning with Verifiable Rewards (RLVR) has become a standard recipe for post-training LLMs on reasoning tasks, with Group Relative Policy Optimization (GRPO) emerging as a leading approach. However, GRPO and its variants are inherently single-turn: they optimize terminal rewards on isolated prompt-response pairs, leaving them poorly suited to agentic settings where models iteratively refine solutions using environmental feedback. We introduce MURPHY, a multi-turn extension of GRPO for self-correcting code generation. MURPHY constructs feedback-conditioned rollout trees in which failed candidate solutions are paired with executor feedback and expanded into subsequent turns. It then propagates rewards backward through the tree so that later successful refinements credit earlier attempts that surfaced informative feedback. We study two propagation strategies, Max Reward (MaRS) and Mean Reward (MeRS), and introduce post-rollout pruning mechanisms that reduce multi-turn optimization cost. Across three code generation benchmarks, HumanEval, MBPP, and LiveCodeBench-v6, and two model families, Qwen3-1.7B/4B and OLMo-2-7B, MURPHY delivers up to 6% absolute pass@1 gains over the strongest prior multi-turn execution-feedback methods. Gains are largest on the Medium/Hard LiveCodeBench subsets, reaching +4.38%/+4.20% at Iter-5, where iterative self-correction matters most.


MUSS: A Multi-scale and Sequence-based Model for Single-cell Gene Regulation

Kangjie Zheng ⋅ Amirhossein Vahidi ⋅ Liying Jin ⋅ Joseph Clarke ⋅ Vijaya Baskar ⋅ Connor Rogerson ⋅ mj zhang ⋅ Muzlifah Haniffa ⋅ Anshul Kundaje ⋅ Mohammad Lotfollahi

Modeling gene regulation in single cells requires jointly representing DNA sequence, chromatin accessibility, and gene expression across regulatory scales. Sequence-based models capture DNA-level regulatory grammar but typically lack cell-level resolution, whereas single-cell omics models capture cell-state semantics but treat genes and peaks as fixed vocabulary tokens, and thus lack a unified sequence-based view of regulation. Here, we present MUSS, a sequence-based multi-scale model for single-cell gene regulation that preserves both DNA-level and cell-level resolution. MUSS derives peak- and gene-level representations from DNA using a frozen pre-trained sequence encoder, integrates them with a multi-scale encoder combining peak-peak self-attention and distance-biased peak-gene interaction, and trains on paired scATAC-seq and scRNA-seq to align sequence-derived representations with cell-specific regulatory states. Across gene expression prediction, enhancer-gene linking, TF-target gene recovery, and zero-shot cell clustering, MUSS outperforms both sequence-based and single-cell omics baselines, and further generalizes to genes and peaks unseen during training. Together, these results demonstrate that MUSS learns biologically meaningful and broadly generalizable regulatory representations for decoding cell-type-specific regulatory programs from DNA sequence.


MySign: A High-Fidelity Motion-Capture Dataset for 3D Sign Generation in Bahasa Isyarat Malaysia

Jiayu Shen ⋅ Kalin Stefanov ⋅ Lay-Ki Soon ⋅ Vee Yee CHONG ⋅ Alpha A Gopalai ⋅ KokSheik Wong

Progress in fine-grained 3D sign generation is constrained by the scarcity of high-fidelity 3D sign language datasets, a limitation that is particularly severe for under-resourced languages. We introduce MySign, the first motion-capture dataset for isolated signs in Bahasa Isyarat Malaysia (BIM), comprising 5,000 3D samples covering 1,000 lexical items (glosses), produced by five Deaf native signers. Unlike prior datasets that rely on SMPL-X parameters regressed from video (pseudo-labels)--often affected by depth ambiguity, inaccurate finger articulation, and temporal jitter--MySign provides ground-truth SMPL-X sequences captured with a motion-capture setup tailored for sign language. This yields higher-fidelity motion, particularly in fine-grained hand and finger articulation. We establish MySign as a challenging benchmark for gloss-conditioned 3D sign generation: state-of-the-art models produce visually plausible outputs but fall short of ground-truth fidelity, particularly in fine-grained articulation, revealing persistent gaps in modeling articulated sign motion. In contrast, recognition baselines consistently achieve high lexical accuracy, indicating that the dataset reliably supports supervised learning. We anticipate that MySign will enable more faithful 3D sign generation and support finer-grained analysis of sign articulation. The dataset is available at https://huggingface.co/datasets/mysigner/MySign.


Natural Jammers: When Non-Robust Features Become Antagonistic

Jiaming Zhang ⋅ Fuyao Zhang ⋅ Pengjun Xie ⋅ Wei Yang Bryan Lim

Deep neural networks are known to rely on non-robust features that align poorly with human perception. Prior work has largely framed these features as opportunistic shortcuts that help standard models on clean data. We show instead that non-robust cues can be antagonistic: they can act as Natural Jammers that actively mislead standard models on clean, unmodified images. Using disagreement between a standard model and robust models as an analysis lens, we identify a Jammed Set on which robust models outperform standard models. We provide causal evidence for Spectral--spatial Locking: jamming arises only when high-frequency components are precisely aligned with specific spatial structures. Disrupting this lock through spectral filtering or small geometric shifts can restore correct predictions. Crucially, we uncover a Substitution Paradox: while individual jammers are fragile to geometric shifts, the aggregate prevalence of jamming remains stable and is redistributed rather than eliminated under such transformations. Our results reframe natural non-robustness as a structural property of the model-data manifold: one that redistributes rather than disappears under simple interventions.


Natural-Language-Guided Protein Generation for Ligand-Binding Design

Shumeng Li ⋅ Xunkai Li ⋅ Sirui Zhang ⋅ Henan Sun ⋅ Xiong Yongfu ⋅ YI LIU ⋅ Rong-Hua Li ⋅ Guoren Wang

Docking ultimately determines effective ligand binding, yet bridging global functional semantics and local physical constraints remains a central challenge for protein generation for ligand-binding design. We therefore use docking-oriented language describing docking processes, binding patterns, and local interaction constraints as generation conditions, rather than relying only on coarse-grained functional text. However, existing methods still insufficiently model docking constraints and show limited coordination among heterogeneous modalities. To address these limitations, we propose DockWizard, an encoder-router-decoder framework for multimodal protein generation. Concretely, (1) a multimodal encoder aligns function text, docking text, ligand descriptions, ligand SMILES, and ligand 3D information into shared conditional representations, improving semantic consistency across heterogeneous inputs; (2) a router module organizes heterogeneous conditions into role-specific memories by their scopes of action, preserving global functional guidance while strengthening local binding constraints; and (3) a protein decoder combines prefix guidance with full-memory injection to let different conditions act at different levels, improving controllability while preserving docking constraints during generation. We further curate a docking-language multimodal dataset with over 100,000 samples. Extensive experiments under a four-stage evaluation pipeline show that DockWizard outperforms strong baselines for this task, with average gains of 13% in docking confidence and 19% in complex stability.


Near-Optimal Learning in Parametric Bandits with Action-Dependent Coarsened Feedback

Zhuohua Li ⋅ MAOLI LIU ⋅ Yuwen Huang ⋅ Cheng Wen ⋅ Jie Su ⋅ Cong Tian ⋅ Shengchao Qin ⋅ John C. S. Lui

We study a structured stochastic multi-armed bandit problem in which each decision simultaneously determines the learner's reward and the feedback available for learning. At each round, the learner selects a task and an option. Each task is associated with an unknown parameter that governs its reward, while each option has a known reward function of this parameter and determines how information about that reward is observed. Some options may provide fine-grained feedback at the cost of lower immediate reward, whereas others may yield higher reward but return only coarse feedback. This creates an intrinsic trade-off between reward maximization and information acquisition. We formalize this setting as parametric bandits with action-dependent coarsened feedback. For stationary tasks, we propose ACF-UCB, an optimistic algorithm that constructs task-level confidence sets by aggregating information from all atoms of the option-dependent feedback distributions. For piecewise-stationary tasks, we develop SW-ACF-UCB, a sliding-window variant that adapts to temporal changes. We prove high-probability regret bounds for stationary environments and dynamic regret guarantees for piecewise-stationary environments. We also establish lower bounds showing that the dependence on the number of tasks and the horizon is unavoidable, while methods that ignore the shared parametric structure suffer substantially larger regret.

Existing federated learning setups assume either a single global model or a fixed decomposition with one globally shared parameter block and one client-specific private block. In this paper, we study a generalized federated optimization setting where each client optimizes a local coordinate block, and two clients interact only when their blocks overlap. We encode these coordinate overlaps by a graph on clients, called the nerve skeleton. Based on this structure, we propose Nerve-Skeleton Message Passing (NSMP), a two-phase protocol on a spanning tree: a leaf-to-root pass composes reduced objectives by partial minimization, and a root-to-leaf pass reconstructs a joint assignment. We further show that, under the tree-elimination schedule in the paper, NSMP is equivalent to minimizing the joint objective, where the leaf-to-root phase evaluates the optimal objective value, and the root-to-leaf phase recovers a globally optimal assignment. Experiments show that NSMP solves federated optimization with overlapping parameters using only efficient neighbor-to-neighbor message passing.


NesyProAct: Proactive Neural-Symbolic Control for Web Agents

Tianyi Tang ⋅ Keyi Xiang ⋅ Jie-Jing Shao ⋅ Yueming LYU ⋅ Ivor Tsang ⋅ Yew Soon Ong ⋅ Haiyan Yin

Web-based language agents operate under severe partial observability, where the semantic divergence between internal state assumptions and the ground-truth environment is often unobservable. Consequently, agents frequently maintain high-level reasoning coherence over misgrounded states, causing errors to compound as subsequent decisions are conditioned on silently invalidated premises. We present NesyProAct, a proactive neural-symbolic agentic decision framework that transforms conventional ReAct-style web agents from reactive prompting pipelines into execution-aware decision processes. NesyProAct exposes a typed symbolic decision interface over states, actions, transitions, and subgoals, enabling LLM-based agents to automatically compose and reason over execution semantics across decision hierarchies. From this interface, programmatic verification logic is synthesized online to evaluate step-level executability and progress, with verification outcomes directly regulating decision evolution through targeted grounding interventions. As a result, execution feedback is integrated into planning itself, allowing agents to maintain semantic alignment under partial observability rather than propagating unverified assumptions. On WebArena, NesyProAct achieves a 12\% overall improvement over strong skill-induction agentic web framework ASI across six domains. NesyProAct also achieves noticeable performance when reasoning with small models GPT-4o-mini, with 5.1\% gain over ASI with Claude-3.5.


Neural-DISCO: Source-Conditioned Counterfactual Editing of Neural Population Activity

Nico Policzer ⋅ James Ho ⋅ Joseph Soo ⋅ Xin Tang

Understanding how neural activity encodes behavioral and sensory variables is a key challenge in systems neuroscience. This requires models that go beyond predicting activity from labels and instead support source-conditioned counterfactual editing: given a recorded source response, we should be able to edit one behavioral or stimulus variable and predict how neural activity would change while preserving other observed factors and source-specific variability not explained by the labels. Here, we introduce Neural-DISCO, a disentangled representation learning framework for counterfactual modeling of neural activity. Neural-DISCO partitions the latent space into label-specific components, each associated with an observed behavioral or stimulus variable, together with a residual latent that captures variability not supplied through observed labels. Across simulated and real neural recording datasets, we show that this structure enables selective manipulation of individual labels to generate counterfactual neural responses while preserving the remaining content of the original activity. Further, to gain insight into how behavioral and stimulus variables are encoded at single-neuron resolution, we pair this disentangled framework with feature attribution methods that identify which neurons are most important for each label. We find that disentanglement enhances the recovery of known functional cell roles in both synthetic and real data. Together, these results support Neural-DISCO as a framework for controlled counterfactual modeling of neural activity and provide a practical interface for probing how neural populations encode behavior and sensory information.

High-frequency Helmholtz problems in heterogeneous media remain challenging for both classical iterative methods and end-to-end neural PDE solvers. We propose Neural Preconditioned Born Series (NPBS), a learned iterative preconditioning framework that operates in preconditioned residual coordinates induced by the Convergent Born Series (CBS). Existing learned Born-series methods primarily use Born-style unrolling for forward wavefield prediction, while learned Helmholtz preconditioners are usually formulated in physical residual coordinates. NPBS fills this gap by recasting Born-series iteration as shifted-Laplacian left preconditioning, and replacing the CBS preconditioner with a learned residual-to-correction map in the Born-preconditioned coordinates. The left preconditioner further induces a residual metric, which yields a metric-matched training objective that aligns optimization with the preconditioned geometry used at inference. On heterogeneous Helmholtz benchmarks, metric-matched NPBS reduces iteration counts by up to $1.9\times$ over direct residual learning, with gains increasing from $1.2\times$ to $1.9\times$ as the wavenumber rises. Compared to classical CBS, learned NPBS reduces stationary iteration counts by over $20\times$; when used as a preconditioner for FGMRES, it further achieves the lowest wall-clock time among all evaluated methods. The same metric-matched formulation also improves convergence on convection--diffusion--reaction systems and Newton linear systems for nonlinear PDEs, indicating that residual-metric matching is a general design principle for neural preconditioners.

Modern Transformer design and compression both reduce to allocating capacity under a budget. The standard scalars for these decisions, #Params and #FLOPs, capture size and compute but not architectural structure—two architectures with identical parameter budgets but different depth-width, head, or FFN allocations receive identical scores yet behave differently. We propose Neural Spectral Capacity (NSC), a closed-form scalar grounded in the singular-value spectrum of each weight matrix. Under standard random initialization, the Marchenko–Pastur law renders NSC computable from the architectural specification alone, with no model instantiation, data, or gradients. Its layer-wise additive structure admits NSC-DP, an exact dynamic-programming solver returning the architecture globally maximizing NSC under resource constraints in seconds on a CPU—a guarantee that black-box search over existing training-free proxies cannot provide. Empirically, NSC outperforms #Params, #FLOPs, and representative training-free proxies in ranking across seven Transformer and CNN families (on FlexiBERT, τ=0.505 on pairs differing in #Params by <10%, where #Params collapses to 0.082); NSC-DP discovers a Transformer-XL architecture on WikiText-103 that beats the human-designed baseline in 2 seconds; and prunes LLaMA-7B to the best 5.7B model across eight commonsense reasoning tasks without any calibration data, ∼5900× faster than the strongest training-free proxy baseline.


NeuroInk: Retinomorphic Spiking Sequence Modeling for Handwritten Text Recognition

Xiubo Liang ⋅ Jinxing Han ⋅ Yuke Li ⋅ Hongyi Duan ⋅ Qizhen Jiang ⋅ Haoqi Zhu ⋅ Yu Zhao ⋅ Hongzhi Wang

Static-input Spiking Neural Networks (SNNs) commonly create temporal dynamics by repeating the same image over multiple timesteps. For handwritten text recognition, however, this convention is input-redundant: the image contains no native event stream, while recognition depends on faint stroke fragments, degraded ink boundaries, and monotonic left-to-right alignment. We introduce NeuroInk, a retinomorphic spike-oriented recognizer that converts static handwriting into a low-timestep sequence of temporally ordered ink evidence. Instead of replaying identical frames, NeuroInk performs temporal ink transduction, decomposing each text-line image into coarse stroke-mass support and fine stroke-boundary evidence. The resulting evidence sequence is injected as time-varying currents into an adaptive spiking visual hierarchy, where high-frequency enhancement and mid-level detail reinjection are used to preserve fine ink structures during spiking propagation. For CTC decoding, two-dimensional spiking maps are read out as width-wise columnar sequences and processed by CTC-aligned association circuits that combine local horizontal context, global association, and selective trace memory. On the LAM benchmark, NeuroInk achieves 2.4\% CER on the validation split and 2.6\% CER on the test split using only two timesteps and greedy CTC decoding, without an external language model or lexicon. These results support the view that static-input SNNs sequence recognition can benefit from co-designing sensory temporalization, spiking visual abstraction, and monotonic sequence association.


Neuro-KE: Knowledge-Guided Interfaces for Semantically Grounded EEG Foundation Models

涵睿 陈 ⋅ Xiang Chen ⋅ Haotian Deng ⋅ Kexin Lou ⋅ Shinan Wang ⋅ Chen Wei ⋅ Quanying Liu

Foundation Models for Electroencephalography (EEG) are trained with masked reconstruction, contrastive alignment, and language-model interfaces, yet their learned representations remain difficult to inspect: raw EEG lacks canonical token semantics, and models may preserve waveform detail while discarding physiologically meaningful structure. We introduce Neuro-KE (Neuro-Knowledge Engine), a knowledge interface that computes EEG descriptors without additional human annotation and renders them as paradigm-specific supervision for semantic grounding. Rather than using handcrafted features as standalone classifiers, Neuro-KE converts the same domain knowledge into numeric prediction targets, feature-state text anchors, and feature-grounded instruction answers. Across three interfaces, Neuro-KE improves aggregate grounding and transfer metrics, with task-dependent exceptions: it improves balanced accuracy (BAcc) by 2.60 points over reconstruction-only masked pretraining across 12 datasets and three backbones, improves six-dataset EEG--text validation BAcc by 2.23 points over label-only anchors, and improves EEG-MLLM feature-grounded generation to ROUGE-L F1 0.607 and BERTScore F1 0.929. These results suggest that established EEG descriptors can serve as reusable grounding signals for inspecting and supervising EEG foundation-model interfaces.


Neuronal Self-Adaptation Enhances Capacity and Robustness of Representation in Spiking Neural Networks

Zhuobin Yang ⋅ Liao Yu ⋅ Yeyao Bao ⋅ Lv L Fu ⋅ Daqing Guo ⋅ Jian Zhang ⋅ Xiaohong Li ⋅ Yunliang Zang

Spiking Neural Networks (SNNs) are promising for energy-efficient, real-time edge computing, yet their performance is often constrained by the limited adaptability of conventional leaky integrate-and-fire (LIF) neurons. Existing LIF models struggle with restricted information capacity and susceptibility to noise, leading to degraded accuracy and compromised robustness. Inspired by the dynamic self-regulation of biological potassium channels, we propose the Potassium-regulated LIF (KvLIF) neuron model. KvLIF introduces an auxiliary conductance state that integrates membrane potential and spiking history to adaptively modulate neuronal excitability and reset dynamics. This design extends the dynamic response range of neurons to varying input intensities and effectively suppresses noise-induced spikes. We extensively evaluate KvLIF on both static image and neuromorphic datasets, demonstrating consistent improvements in classification accuracy and superior robustness compared to existing LIF models. Our work bridges biological plausibility with computational efficiency, offering a neuron model that enhances SNN performance while maintaining suitability for low-power neuromorphic deployment.


Neuron Populations Exhibit Divergent Selectivity with Scale

Amil Dravid ⋅ Yasaman Bahri ⋅ Alexei Efros ⋅ Yossi Gandelsman

We investigate whether neuron populations that shape the internal organization of neural networks evolve predictably with scale, extending scaling laws beyond macroscopic observables such as loss. To probe this question, we study Rosetta Neurons, a previously characterized class of neurons whose activation patterns are similar across independently trained models (Dravid et al., 2023). In separate analyses of language models up to 30B parameters and vision models up to 5B parameters, we observe that the population of Rosetta Neurons follows a sublinear power law in model size, growing in absolute number but occupying a shrinking fraction of the total neuron count. We further observe a Neuron Polarization Effect: Rosetta Neurons become more selective and increasingly monosemantic with scale, separating from a growing non-Rosetta population that remains less selective. An analytical model balancing feature utility against limited neuron capacity explains the sublinear power-law scaling and this polarization effect. Finally, we find that Rosetta Neurons become more domain-specialized with scale and illustrate their selectivity through a targeted data-filtering case study for continued pretraining. Our results point to a scaling law for shared neuron-level structure, linking model size to systematic changes in neuron universality, selectivity, and specialization.


NNCoxKL: Risk-Set Distillation from Probability-Free Prognostic Teachers for Deep Cox Models

Lingfeng Luo ⋅ Jeremy M Taylor ⋅ Di Wang ⋅ Feiyang Deng ⋅ Yubo Shao ⋅ Kevin He

Survival models are often trained at target clinical sites where outcomes are limited and censored, while patient-level data from external cohorts may be unavailable because of institutional data-sharing constraints. In many applications, the transferable knowledge is not a dataset, a calibrated survival curve, or estimated regression coefficients, but a probability-free prognostic artifact such as a clinician-defined scorecard, guideline-based staging rule, registry-derived risk index, or coarse clinical risk group. These weak teachers may encode human expertise or evidence from prior cohorts, yet they often have only ordinal meaning and need not share a numerical scale with the target model. We propose NNCoxKL, a transfer-learning framework that trains a deep Cox model on target time-to-event data while distilling such human- or model-derived prognostic signals through Cox risk sets. For each event risk set, NNCoxKL converts both the external signal and the neural Cox risk score into Plackett--Luce distributions and regularizes the Cox partial likelihood through a temperature-scaled KL alignment term. This transforms clinical knowledge into a censoring-aware ranking prior over who is most likely to fail next, rather than treating the external score as a covariate or requiring it to represent calibrated survival probabilities. The external signal enters only through this risk-set regularizer, allowing the model to learn nonlinear target-cohort effects while borrowing relative-risk information from external clinical knowledge. Experiments on independent clinical transfer tasks and controlled survival benchmarks show that NNCoxKL improves discrimination and calibration-sensitive prediction over internal-only deep Cox models and common score-transfer baselines.


No Free Best-of-Both-Worlds Learning in Repeated Bilateral Trade

Yutian Cheng ⋅ Canzhe Zhao ⋅ Jingye Zhao ⋅ Shuai Li

We study best-of-both-worlds learning in repeated bilateral trade with two-bit feedback. When seller and buyer valuations are independent, minimax regret scales as $\Theta(T^{2/3})$; under general bounded-density distributions it degrades to $ \Theta(T^{3/4})$. This regime-dependent gap naturally raises a question: can a single algorithm adapt optimally to both? We first resolve this negatively: any algorithm achieving $O(T^\alpha)$ regret on independent-values instances necessarily suffers $\Omega(T^{1-\alpha/3})$ regret on some dependent-values instance. We then characterize the optimal tradeoff and construct a meta-algorithm achieving $\bigl(\widetilde O(T^{\alpha}),\widetilde O(T^{1-\alpha/3})\bigr)$ for every $\alpha\in[2/3,3/4]$. Finally, we show that this tradeoff can be bypassed with mild side information: knowing the optimal same-price gain suffices to recover the classical best-of-both-worlds guarantee $\bigl(\widetilde O(T^{2/3}),\widetilde O(T^{3/4})\bigr)$.


Noise-Regularized Training for Learned Image Compression

Renjie Zou ⋅ Zhiwei Huang ⋅ Dong Jiang

Learned image compression (LIC) is trained almost exclusively on clean images: the only stochastic perturbation in the loop is an additive uniform proxy injected in the latent domain to keep the entropy model differentiable. Outside compression, by contrast, a long line of work - from Tikhonov regularization and contractive autoencoders through denoising score matching and modern diffusion models - has established that Gaussian noise at the input of a network with a clean target acts as a structured probe of the data manifold rather than as data augmentation. Building on this view, we introduce noise-regularized training for LIC: at every training step we add Gaussian noise to the encoder input while keeping the distortion target, the latent uniform proxy, the hyperprior, and the entropy coding pipeline unchanged. The recipe is architecture-agnostic and adds essentially no training or inference overhead. Across backbones and operating regimes it yields consistent BD-Rate gains: -2.60% on a reproduced ELIC and -0.85% / -1.05% / -2.38% on the Base/Medium/Large scales of a variable-rate DMCI codec, with DMCI-Large reaching -22.02% against VTM-22.0. A second-order expansion of the noisy RD loss accounts for these gains: because the noise enters upstream of the analysis transform, it converts the clean RD objective into a structured $O(\sigma^{2})$ regularizer comprising an information-weighted encoder-Jacobian penalty (rate side) and a Frobenius contractive penalty on the end-to-end reconstruction map (distortion side). A finite-difference Lipschitz analysis verifies both predictions: across QPs and an order-of-magnitude sweep of the probe scale, the encoder's local Lipschitz constant drops by ~12% on average and that of the end-to-end codec by ~8%. Code and trained models will be released to facilitate reproducibility.


Noise-Started One-Step Image Super-Resolution via LR-Conditioned SplitMeanFlow and GAN Refinement

Wei Zhu ⋅ Kai Zhang ⋅ Yu Zheng ⋅ Lei Luo ⋅ Yong Guo ⋅ Jian Yang

Pre-trained text-to-image (T2I) diffusion models have shown strong potential for real-world image super-resolution (Real-SR), owing to their noise-started generation process that enables realistic texture synthesis and captures the one-to-many nature of super-resolution. However, diffusion-based Real-SR methods still face a fundamental efficiency-quality trade-off. Multi-step methods generate high-quality results by iteratively denoising random Gaussian noise under LR conditioning, but suffer from slow sampling. Recent one-step methods greatly improve efficiency, yet they typically replace noise-started generation with direct LR-to-HR restoration, which weakens stochasticity and limits realistic detail synthesis. To address this issue, we propose SMFSR, a noise-started one-step Real-SR framework via LR-conditioned SplitMeanFlow and GAN refinement. SMFSR preserves the random-noise starting point of diffusion models and learns a direct noise-to-HR mapping conditioned on the LR image. To this end, Interval Splitting Consistency distills the multi-step generative trajectory into a single average-velocity prediction, enabling efficient one-step generation. To compensate for the reduced opportunity for progressive refinement, we further introduce a GAN refinement stage, where a DINOv3-based discriminator enhances realistic texture synthesis and variational score distillation aligns the generated outputs with the natural image distribution. Extensive experiments demonstrate that SMFSR achieves superior perceptual quality while retaining the high efficiency of one-step diffusion models.


No More, No Less: Task Alignment in Terminal Agents

Sina Mavali ⋅ David Pape ⋅ Jonathan Evertz ⋅ Samira Abedini ⋅ Devansh Srivastav ⋅ Thorsten Eisenhofer ⋅ Sahar Abdelnabi ⋅ Lea Schönherr

Terminal agents are increasingly capable of executing complex, long-horizon tasks autonomously from a single user prompt. To do so, they must interpret instructions encountered in the environment (e.g., README files, code comments, stack traces) and determine their relevance to the task. This creates a fundamental challenge: relevant cues must be followed to complete a task, whereas irrelevant or misleading ones must be ignored. Existing benchmarks do not capture this ability. An agent may appear capable by blindly following all instructions, or appear robust by ignoring them altogether. We introduce TAB (Task Alignment Benchmark), a suite of 89 terminal tasks derived from Terminal-Bench 2.1. Each task is intentionally underspecified, with missing information provided as a necessary cue embedded in a natural environmental artifact, alongside a plausible but irrelevant distractor. Solving these tasks requires selectively using the cue while ignoring the distractor. Applying TAB to ten frontier agents reveals a systematic gap between agent capability and task alignment. The strongest Terminal-Bench agent achieves high task completion but low task alignment on TAB. Evaluating six prompt-injection defenses further shows that suppressing distractor execution also suppresses the cues required for task completion. These results demonstrate that task-aligned agents require selective use of environmental instructions rather than blanket acceptance or rejection.

In this work we study the Best Policy Identification (BPI) problem in online, tabular Reinforcement Learning. This is an active sequential hypothesis testing problem in which the learner's objective is to identify an optimal policy in a Markov Decision Process (MDP) with high confidence, while minimizing the expected sample complexity to do so. We consider an online setting with deterministic rewards, where the agent must strategically navigate through the MDP in order to effectively explore. Previous works in the literature have provided asymptotically optimal methods for BPI, such as the Navigate and Stop (NaS) algorithm and its variants, however existing analysis remains asymptotic. In this work, we fill that gap by providing the first non-asymptotic sample complexity guarantees for NaS, showing that its sample complexity depends not only on the characteristic time, but also on the connectivity of the underlying MDP, the curvature of the optimal characteristic time, and other instance-dependent quantities. We identify these additional attributes and make explicit their contributions to the overall sample complexity.

AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent competition can lead to behaviors that serve individual gains at collective cost—for instance, marketing agents may post misleading content as a result of competing for engagement on social media. Human societies address such problems through norms that constrain acceptable behavior, supported by enforcement mechanisms that detect and penalize violations. Motivated by this, we study norm enforcement mechanisms for language model agents. We find that simple enforcement mechanisms are exploited by misaligned agents for competitive advantage, even when they are not explicitly trained or prompted to do so. We thus turn our attention to designing more robust mechanisms, and identify two key ingredients: estimating each agent’s reliability over time, and updating this estimate with escalating penalties for repeated misbehavior. Across three simulated environments and a variety of agent populations, mechanisms built on these principles resist exploitation, while still penalizing norm violations at comparable or lower cost than baselines. Our results position norm enforcement mechanisms as promising levers for shaping agents' behavior, but only when designed to anticipate becoming part of the environment they govern.

Multiple instance learning (MIL) learns from bag-level labels, but a bag’s label is often supported by only a few hidden instances. In heterogeneous bags, these useful instances are rare, diverse, and surrounded by many weak or noisy observations. We propose Sparse Support Grounding, a MIL framework that learns instance influence directly. The key idea is simple: an instance should not matter just because it can be matched somewhere; it should matter when the match provides reliable label evidence. Our method forms a compact, bag-specific evidence support, compares instances with label-consistent supports, and uses unbalanced transport to estimate how much source mass each instance should retain. This retained mass becomes the strength of its grounding signal, so reliable evidence shapes training more than nuisance content, while the bag classifier still uses the full bag. Thus, transport is turned from a matching rule into a way of learning which instances should drive supervision. Experiments on synthetic MIL, whole-slide images, and clinical endomicroscopy show consistent gains, with the largest improvements when sparse diagnostic evidence is embedded in heterogeneous nuisance content.


Not All Tokens Should Be Treated Equally: Context Credits Reassignment

Shujian Gao ⋅ Yuan Wang ⋅ Jiamei Yan ⋅ Penghao Zhou ⋅ Qinglei Wang ⋅ Jiangtao Yan ⋅ Zuxuan Wu ⋅ Yu-Gang Jiang

Multimodal Reinforcement learning with verifiable rewards (RLVR) algorithms assign the same advantage to every token in a generated sequence, providing no differential signal for tokens with varying degrees of dependence on visual evidence. A simple experiment confirms this issue: \textbf{scaling token advantages by uninformative random noise already outperforms uniform credit}, suggesting that uniform credit is empirically suboptimal. Beyond that, we observe RLVR tuning under uniform credit weakens visual grounding along three complementary axes: attention, gradient attribution, and functional dependence on image embeddings, \textbf{collectively indicating insufficient utilization of visual evidence}. To understand this degradation, we provide a theoretical analysis showing that, under language-prior dominance and regularity conditions, uniform credit can amplify a pretrained text-favored gradient imbalance, contributing to weaker visual attention. In response, we propose Hierarchical Context Credit Reassignment (\textbf{HiCCR}), which reassigns credit at two levels: a token-level weight amplifies the advantage for visually grounded tokens, and a trajectory-level weight upweights rollouts with stronger visual engagement. The entire mechanism \textbf{adds less than 3\% wall-clock overhead} and applies to mainstream RLVR algorithms without auxiliary models or additional data. On ten benchmarks spanning mathematical and general multimodal reasoning, HiCCR consistently improves upon its corresponding base algorithm and achieves state-of-the-art results among open-source models of comparable scale, while alleviating the weakening of visual grounding observed under RLVR training.


NoTA: Normalized Tensor Adaptation for Parameter-Efficient Continual Learning

Yunsong Deng ⋅ Yuning Qiu ⋅ Qibin Zhao ⋅ Guoxu Zhou

Rehearsal-free class-incremental learning (CIL) requires a model to learn new classes sequentially without storing previous data, making catastrophic forgetting a central challenge. Adapter-based methods have become a prevalent solution by storing and accumulating task-specific update matrices, yet the effect of the cumulative update magnitude on CIL performance remains underexplored. In this work, we propose Normalized Tensor Adaptation (NoTA), which addresses these two aspects jointly. MPO adapters provide a compact structured space for storing task-specific updates, while global normalization controls the accumulated update that determines the final model. Our empirical analysis reveals a key observation: global normalization consistently improves both matrix-based and MPO-based adapters by stabilizing the cumulative update magnitude across task sequences. We further provide a theoretical analysis showing that global normalization yields a forgetting bound independent of the task sequence length, in contrast to task-wise normalization whose bound can grow with later tasks. Building on the tensor network structure of MPO, we introduce a variant with hierarchical coefficients (H-NoTA), which uses task-level coefficients and inter-core modulation matrices to refine the contributions of historical adapters. Extensive experiments demonstrate that our methods consistently improve rehearsal-free CIL performance while maintaining strong parameter efficiency.


NPCBench: A Clinical Apprenticeship Benchmark for Guideline-Constrained Care-Pathway Reasoning in Nasopharyngeal Carcinoma

Pengkai Wang ⋅ Wei-Wei Zhang ⋅ Yan Li ⋅ Min Tang ⋅ Zhitian Hou ⋅ Zeyu Liu ⋅ Guanghao Zhu ⋅ Yuanyi Wang ⋅ Yanggan Gu ⋅ Wenjun Wang ⋅ Minheng Ni ⋅ Congkai Xie ⋅ Zhijie Sang ⋅ Jianmin Wu ⋅ Ying Sun ⋅ Hongxia Yang

Current medical LLM benchmarks typically decompose clinical competence into isolated factual questions, single-image interpretation, or brief case scenarios. This fragmented paradigm primarily measures local correctness but provides limited insight into whether models can progressively integrate subspecialty knowledge, multimodal evidence, and evolving patient states into safe, guideline-constrained longitudinal reasoning. To address this gap, we introduce NPCBench, a multimodal clinical apprenticeship benchmark for nasopharyngeal carcinoma (NPC). NPCBench operationalizes specialist training as a staged evaluation of (M)LLMs across guideline recall, atomic rule application, multimodal evidence grounding, longitudinal care planning, state revision, and expert-style consultation. It comprises 4,077 staged evaluation problems and 25 multimodal full-care path episodes, covering pathology diagnosis, MRI staging, and 146 downstream clinical decisions. The benchmark explicitly links stepwise tasks to complete patient-level care trajectories, enabling systematic evaluation of all clinical pathways. Beyond isolated precision, NPCBench evaluates whether the models ground decisions in multimodal evidence, maintain temporally consistent patient states, and provide safe expert-level consultation across complete care trajectories. Across 30 evaluated models, we observe a persistent composition gap: strong performance in local case-level tasks does not translate into episode-level clinical competence required for end-to-end patient management. Overall, NPCBench advances medical AI evaluation from static question answering toward apprenticeship-style workflow evaluation, providing a rigorous testbed for evidence-grounded and temporally coherent reasoning in complex oncology care.


NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models

Ziyue Zhu ⋅ Shangyang Wu ⋅ Shuai Zhao ⋅ Zhao ZhiQiu ⋅ Jian Zhang ⋅ Li Shengjie ⋅ Yi Wang ⋅ Anh Tuan Luu ⋅ XINLIANG ZHOU ⋅ Fang Li ⋅ Haoran Luo

Vision-Language-Action (VLA) models are formulated to ground instructions in visual context and generate action sequences for robotic manipulation. Despite recent progress, VLA models still face challenges in learning related and reusable primitives, reducing reliance on large-scale data and complex architectures, and enabling exploration beyond demonstrations. To address these challenges, we propose a novel \textbf{N}euro-\textbf{S}ymbolic \textbf{V}ision-\textbf{L}anguage-\textbf{A}ction (NS-VLA) framework via online reinforcement learning (RL). It introduces a symbolic encoder to embedding vision and language features and extract structured primitives, utilizes a symbolic solver for data-efficient action sequencing, and leverages online RL to optimize generation via expansive exploration. Experiments on robotic manipulation benchmarks demonstrate that NS-VLA outperforms previous methods in both one-shot training and data-perturbed settings, while simultaneously exhibiting superior zero-shot generalizability, high data efficiency, and an expanded exploration space. Our code is available at https://anonymous.4open.science/r/NS-VLA/.


N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation

Aleksander Lorenc ⋅ Frédéric Berdoz ⋅ Joël Mathys ⋅ Roger Wattenhofer

Improving the inference efficiency of autoregressive transformers typically means reducing FLOPs per token, usually through approximations that degrade model quality. We introduce N-vium, a mixture-of-exits transformer that partially parallelizes computation across depth on standard hardware, increasing effective FLOPs per second rather than minimizing compute per token. N-vium attaches prediction heads at multiple depths and defines the next-token distribution as a learned mixture over these exits, with token-adaptive routing. This formulation strictly generalizes the standard transformer, which is recovered exactly when routing assigns zero mass to all intermediate heads. Sampling from the mixture is exact, and complete KV caches are recovered by deferring the upper-layer computation and batching it with later tokens. We pretrain N-vium at scales up to 1.5B parameters. Our largest model reaches 57.9% wall-clock speedup over a parameter- and data-matched standard transformer at no perplexity cost.


OccStress: Stress-Testing the 4D Occupancy Forecasting Chain

Yu Zheng ⋅ Jie Hu ⋅ Jiaqi Xiong ⋅ Ruiping Liu ⋅ Junwei Zheng ⋅ Kailun Yang ⋅ Jiaming Zhang

Occupancy world models use historical occupancy states to forecast future 3D scenes, but their robustness under corrupted temporal inputs remains poorly understood. Existing evaluations primarily emphasize clean forecasting accuracy and provide limited evidence about how errors enter, persist, and propagate through the occupancy perception-forecasting chain. This paper introduces OccStress, a robustness stress-testing benchmark for the occupancy forecasting chain. OccStress contains 61 corruption categories and 80k+ anchors. OccStress covers both 3D occupancy perception and 4D occupancy forecasting through two complementary tracks. This design separates realistic pipeline errors from the intrinsic sensitivity of 4D forecasting models to corrupted occupancy states. OccStress further defines temporal injection protocols to test whether errors in the current state, recent history, or earlier history affect future forecasts differently. OccStress provides aggregate metrics to evaluating robustness along occupancy forecasting chain. Experiments on OccStress reveal that current occupancy models are substantially affected by both upstream prediction errors and direct state corruptions, and that clean performance alone is an insufficient indicator of temporal robustness. Dataset: https://hf.co/datasets/OccStress/OccStress. Code: https://anonymous.4open.science/r/OccStress.


OFBD: Object-Focused Background Debiasing for Long-Tailed Learning

Shenghan Chen ⋅ Yiming Liu ⋅ Zhipeng Deng ⋅ Haolin Wang ⋅ Jiale Zhou ⋅ Zhijian Wu ⋅ Lu Xiankai ⋅ yafei ou ⋅ Yefeng Zheng

Balancing performance trade-offs on long-tailed data distributions remains a long-standing challenge in visual recognition. Existing methods mainly improve tail classes through re-balancing, representation learning, or data augmentation, but the underlying cause of tail classes degradation is still insufficiently explored. In this paper, we find that standard long-tailed training induces background-biased representation and optimization: tail classes suffer larger background distribution shifts and become increasingly driven by background gradients. This reveals that tail degradation is not merely caused by insufficient samples, but also by the learning of irrelevant background features. To tackle this issue, we propose Object-focus Background Debiasing (OFBD), a framework that mitigates background bias from both distribution and optimization perspectives. Specifically, Foreground-guided CutMix preserves target-related foregrounds while diversifying complementary backgrounds, and Background-guided Feature Rectification suppresses background-biased features without learnable parameters or additional training. Extensive experiments show that our method improves overall accuracy, achieves significant tail-class gains, and can serve as a plug-in for mainstream long-tailed methods without external data or pretrained recognition models.


OmniCapBench: A Deep-Structured Evaluation Framework for Fine-Grained Audio-Visual Captioning

Zhongyu Yang ⋅ Jiale Tao ⋅ Ruitao Chen ⋅ Zuhao Yang ⋅ Yingfang Yuan ⋅ Xueliang Zhao ⋅ Auden ⋅ Kai Wang ⋅ Shuai Shao ⋅ Biao Wang ⋅ Steve Yves ⋅ Qinglin Lu

Multimodal large language models (MLLMs) are rapidly evolving toward continuous audio--visual reasoning, creating an urgent need for evaluations that expose their capability limits. Audio--visual captioning is an ideal diagnostic task, yet current benchmarks face a coupled trade-off: whole-caption scores provide coverage without localization, local probes provide localization without coverage, and unconstrained LLM judges introduce instability. We introduce OmniCapBench (Omni-Video Caption Benchmark), a benchmark that reframes audio--visual caption evaluation as a deep-structured diagnostic framework. OmniCapBench shifts the prediction target from free-form text to sets of atomic, verifiable evaluation units across three tracks: entity references, visual shots, and audio events, enabling reliable scoring with deterministic constraint checks and localized LLM-based semantic comparisons. With 786 densely annotated videos, OmniCapBench effectively distinguishes MLLM perception errors, including temporal grounding failures, identity drift, cross-modal misalignment, and hallucinated descriptions. Evaluating frontier MLLMs reveals strong local perception but weak long-horizon audio–visual reasoning, particularly in identity drift and cross-modal misalignment, providing a fine-grained roadmap for omnimodal development.


OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning

Maonan Wang ⋅ Zhengyan Huang ⋅ Kemou Jiang ⋅ Yuhang Fu ⋅ Jiayue Zhu ⋅ Yuxin Cai ⋅ Xingchen Zou ⋅ Qiaosheng Zhang ⋅ Yi Yu ⋅ Ding Wang ⋅ XI CHEN ⋅ Ben Chen ⋅ Yuxuan Liang ⋅ Zhiyong Cui ⋅ Man On Pun ⋅ Yirong Chen

Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evolution, and signal-phase semantics. However, existing traffic-oriented multimodal benchmarks largely emphasize passive visual recognition or isolated video understanding, offering limited support for evaluating structure-aware traffic reasoning under controlled conditions. We introduce \textbf{OmniTraffic}, a controllable generation pipeline and benchmark for spatio-temporal traffic reasoning. Built around 12 real-world intersections reconstructed into editable 3D traffic environments and complemented by surveillance footage from two countries, OmniTraffic supports both controlled and natural-condition evaluation. It defines a three-level task hierarchy spanning scene perception, multi-view and temporal reasoning, and decision support. Using structured traffic metadata, OmniTraffic generates synchronized multi-view VQA samples covering vehicle states, lane functions, view--BEV correspondence, temporal dynamics, and signal-phase analysis, resulting in 8M VQA samples and a 3K human-verified test set. Evaluation of eleven frontier MLLMs reveals a large human--model gap, with the most pronounced failures in topology-grounded and spatio-temporal reasoning tasks. Fine-tuning a lightweight MLLM on simulated OmniTraffic data further improves performance on real-world traffic scenes, demonstrating the value of simulation-generated supervision for traffic-specific multimodal reasoning. Beyond a fixed dataset, OmniTraffic provides an extensible pipeline with configurable intersections, camera views, traffic demands, signal phases, visual conditions, and rare events.


One Loss to Rule Them All: Marked Time-to-Event for Structured EHR Foundation Models

Zilin Jing ⋅ Vincent Jeanselme ⋅ Yuta Kobayashi ⋅ Simon Lee ⋅ Chao Pang ⋅ Aparajita Kashyap ⋅ Yanwei Li ⋅ Xinzhuo Jiang ⋅ Shalmali Joshi

Clinical events captured in Electronic Health Records (EHR) are irregularly sampled and may consist of a mixture of discrete events and numerical measurements, such as laboratory values or treatment dosages. The sequential nature of EHR, analogous to natural language, has motivated the use of next-token prediction to train prior EHR Foundation Models (FMs) over events. However, this training fails to capture the full structure of EHR. When a given event occurs must be captured, but the event value (abnormal lab) also modulates the likelihood of other clinical events. Most existing EHR FMs do not jointly model this likelihood and are unable to capture the full observation process, impacting downstream capabilities. We propose ORA, a marked time-to-event pretraining objective that jointly models event timing and associated measurements. Across multiple datasets, downstream tasks, and model architectures, this objective consistently yields more generalizable representations than next-token prediction and pretraining losses that ignore continuous measurements. Importantly, the proposed objective yields improvements beyond traditional classification evaluation, including better regression and time-to-event prediction. Beyond introducing a new family of FMs, our ablations suggest a broader takeaway: pretraining objectives that account for EHR structure are critical for expanding downstream capabilities and generalizability.

As interactive generative systems are increasingly deployed in real-world applications, their tendency to generate unreliable or false responses raises serious concerns. Conformal abstention mitigates this risk by ensuring that the system answers only when confident. However, real-world deployments typically provide only partial user feedback (e.g., thumbs up/down) on the selected response and often operate in non-stationary or adversarial environments, for which effective learning methods are largely missing. To bridge this gap, we propose ExAUL, a novel online learning framework for conformal abstention with adversarial and partial feedback. Technically, we introduce (i) a novel conversion lemma}that translates the regret of any bandit algorithm into an FDR bound, and (ii) feedback unlocking, a strategy that exploits the structure of conformal abstention to extract additional learning signals from partial feedback. We prove that ExAUL achieves a regret bound of $\mathcal{O}(\sqrt{T \ln |\mathcal{H}|})$, which translates into an $\mathcal{O}(\sqrt{T})$ bound on FDR risk control, matching the controllability of full-information settings despite receiving only partial feedback. While applicable to general generative tasks, we demonstrate the efficacy of ExAUL for ensuring the reliability of Large Language Models (LLMs) through empirical validation on question-answering tasks across diverse non-stationary and adversarial settings. Our results demonstrate that ExAUL robustly controls the FDR while maintaining competitive answering coverage.


Online Data Selection for Instruction Tuning via Gaussian Processes

Jun Wang ⋅ Quoc Phong Nguyen ⋅ Julien Monteil ⋅ Vu Nguyen

With Large Language Model (LLM) pre-training and fine-tuning shifting its focus from data volume to data quality, quality data selection has emerged as a critical research topic. Existing online data selection methods for LLM training are typically ``batch-constrained'', limiting optimization to local utility within random batches. To overcome this, we propose GAIA (Global Adaptive Instruction tuning via GAussian processes), a framework that formulates data valuation as a global estimation process. GAIA employs Gaussian Process regression to model continuous utility manifolds across the semantic space, utilizing an adaptive strategy fusion mechanism to dynamically prioritize high-utility samples. By casting the strategy-posterior update as an instance of the classical fixed-share Hedge framework for tracking the best expert, we inherit a dynamic-regret guarantee that characterizes GAIA's robustness under non-stationary quality scores during training. Empirical evaluations on three datasets demonstrate that GAIA significantly outperforms state-of-the-art baselines like GREATS, establishing our method as a scalable and robust solution for efficient instruction tuning.


On Making $SE(2)$-Invariant Networks Optimal

Tomas Karella ⋅ Emily Shinkle ⋅ Alice Allen ⋅ Pieter J Swart ⋅ NIcholas Lubbers ⋅ Roxana Bujack

Respecting a-priori known symmetries of underlying data is a key principle in designing efficient neural architectures. Convolutional Neural Networks (CNNs) are a canonical example, encoding translation equivariance. Extending this inductive bias to richer symmetry groups such as $\mathrm{SE}(2)$ in the most effective way possible is an ongoing architectural and computational challenge. Building on classical invariant theory, we derive invariant feature maps that are provably optimal: minimal and information-preserving with respect to the symmetry group. Our construction is designed to integrate seamlessly into standard architectures, requiring no modification to nonlinearities or normalization layers. As a concrete implementation, we introduce a plug-and-play module for CNNs that enforces $\mathrm{SE}(2)$ invariance, capturing both translations and rotations within standard architectures. The resulting models are both efficient and expressive. Experiments show competitive or improved accuracy compared to state-of-the-art equivariant methods at significantly lower computational cost. Finally, we provide a theoretical and empirical analysis of the relationship between equivariance and invariance, offering insight into their roles in controlling robustness and flexibility.


On the fundamentals of gradient-based SAT solvers

Nikolaos Karalias ⋅ Sakari Peltonen ⋅ Nikita Kostin ⋅ Stefanie Jegelka

Gradient-based methods have received increasing attention for Boolean satisfiability, both in neural-network-based solvers and broadly as continuous local search techniques. In this work, we provide a deeper understanding of the interplay of optimization process, problem structure, common formula transformations and neural network properties. Our unified view treats both direct gradient descent over variable marginals and neural-parametrized gradient descent as instances of the same continuous SAT framework. We show that its objective, the expected number of unsatisfied clauses under independent variable marginals, equals (up to constant shift) the clausewise Fourier objective commonly used in continuous SAT methods, and characterize its relation to the exact product-measure probability of unsatisfiability. Then, we derive convergence conditions for gradient descent. A central theme is symmetry: gradient descent can become trapped in invariant subspaces induced by formula symmetries. To remedy this, we show that formula transformations commonly used as SAT preprocessing tools can be repurposed to modulate the optimization landscape. We analyze their effect on problem structure, smoothness, tightness, and expressivity barriers of the neural network. Empirically, across several benchmarks, these transformations can improve gradient-descent performance. Our experiments also indicate that transformations can improve data efficiency and inference-time performance in self-supervised neural SAT solving. Finally, we demonstrate that neural-parametrized gradient descent can improve over input-space gradient descent across several datasets.

In stochastic (strongly) convex optimization, establishing last-iterate convergence guarantees for first-order methods has recently garnered increasing attention, as this not only deepens the theoretical understanding of these algorithms but also aligns more closely with practice. Notably, people have demonstrated that one of the most famous, simple, and popular methods, Stochastic Gradient Descent ($\mathtt{SGD}$), provably achieves last-iterate convergence. However, the existing results may have limited applicability, since they typically rely on strong conditions on the gradient noise (e.g., an exponentially decaying tail) rather than the more realistic heavy-tailed noise. Unfortunately, when faced with heavy-tailed noise, $\mathtt{SGD}$ is known to exhibit undesirable behavior or even fail to converge. To address the challenge of heavy-tailed noise, people have proposed $\mathtt{Clipped}\text{-}\mathtt{SGD}$, an algorithm that combines $\mathtt{SGD}$ with a simple mechanism, gradient clipping. Although $\mathtt{Clipped}\text{-}\mathtt{SGD}$ performs well in practice, its last-iterate convergence property remains highly underexplored. In this work, we prove the first optimal last-iterate rates in high probability for $\mathtt{Clipped}\text{-}\mathtt{SGD}$ under heavy-tailed noise, thereby closing a gap in the literature.


OperatorSHAP: Fast and Accurate Shapley Value Estimation for Neural Operators

Joshua Stiller ⋅ Santo Thies ⋅ Felix Czaja ⋅ Eyke Hüllermeier

Understanding model predictions is essential for physical applications, where outputs often inform safety-critical decisions, such as structural load assessment, weather warnings, and clinical diagnosis. Shapley values satisfy many desirable properties as an attribution method, but their computational cost during inference hinders their practical use. Current amortized explainers, such as FastSHAP, are limited to homogeneous inputs, which is problematic for physical applications where data often comes from irregular grids and geometries. We introduce OperatorSHAP, a grid-agnostic attribution method and training procedure that allows us to train FastSHAP-like explainers for neural operators. We establish a theoretical framework for attributions in function space, connecting to Aumann–Shapley values. We further show that OperatorSHAP's explanations are consistent with state-of-the-art discrete Shapley values across resolutions and transfer across grid sizes without retraining.

We consider online reinforcement learning for the Stochastic Shortest Path (SSP) problem with unknown transition dynamics. Since standard posterior sampling in SSP lacks an explicit mechanism to enforce uniform optimism and therefore fails to attain minimax-optimal regret, we propose OPSRL-SSP, the first optimistic posterior sampling algorithm for the SSP problem. The algorithm operates in epochs and uses a logarithmic, time-dependent posterior sampling schedule. To ensure optimism, it incorporates a pseudo-state with an optimistic value evaluation into the posterior distribution. We establish a high-probability regret bound of $\tilde{\mathcal{O}}(B_\star \sqrt{SAK} + B_\star S^2 A)$, where $B_\star$ is an upper bound on the expected cost of the optimal policy, $S$ and $A$ are the numbers of states and actions, respectively, and $K$ is the number of episodes. The key technical challenge is to turn these optimistic posterior samples into a uniform optimism guarantee despite the random and potentially unbounded episode lengths of SSP. We address this through an SSP-specific analysis based on dynamic posterior inflation and a contraction argument over the pseudo-state-augmented model. In addition, we extend the analysis to general zero-cost SSPs via a universal cost-perturbation argument. Our dominant term matches the minimax lower bound $\Omega(B_\star \sqrt{SAK})$, thereby answering the open problem raised by \citet{jafarniajahromi2021onlinelearningstochasticshortest} for posterior sampling in the SSP setting.


Optimistic Dual Averaging Unifies Modern Optimizers

Thomas Pethick ⋅ Wanyun Xie ⋅ Roman Machacek ⋅ Volkan Cevher

We introduce SODA, a generalization of Optimistic Dual Averaging, which provides a common perspective on state-of-the-art optimizers like Muon, Lion, AdEMAMix and NAdam, showing that they can all be viewed as optimistic instances of this framework. Based on this framing, we propose a practical SODA wrapper for any base optimizer that eliminates weight decay tuning through a theoretically-grounded $1/k$ decay schedule. Empirical results across various scales and training horizons show that SODA consistently improves performance without any additional hyperparameter tuning.


Optimized Minimal 4D Gaussian Splatting for Efficient Dynamic Scene Representation

Minseo Lee ⋅ Byeonghyeon Lee ⋅ Lucas Y Lee ⋅ Eunsoo Lee ⋅ Jun Y Jeong ⋅ Sangmin Kim ⋅ Seunghyeon Song ⋅ Joo Chan Lee ⋅ Jong Hwan Ko ⋅ Jaesik Park ⋅ Eunbyung Park

4D Gaussian Splatting has emerged as a new paradigm for dynamic scene representation, enabling real-time rendering of scenes with complex motions. However, it faces a major challenge of storage overhead, as millions of Gaussians are required for high-fidelity reconstruction. Compressing explicit 4D Gaussians is challenging because they exhibit heterogeneous redundancy across static and dynamic regions, and their temporal attributes make direct quantization unstable. In this work, we present OMG4 (Optimized Minimal 4D Gaussian Splatting), a framework that constructs a compact set of salient Gaussians capable of faithfully representing 4D Gaussian models. Our method progressively reduces Gaussians in three stages: (1) Gaussian Sampling to identify primitives critical to reconstruction fidelity, (2) Gaussian Pruning to remove redundancies, and (3) Gaussian Merging to fuse primitives with similar characteristics. In addition, we integrate implicit appearance compression and extend Sub-Vector Quantization (SVQ) to 4D representations with a staged quantization scheme, further reducing storage while preserving quality. Extensive experiments on standard benchmark datasets demonstrate that OMG4 achieves a favorable rate-distortion trade-off over recent state-of-the-art methods, reducing model sizes by over 60\% while maintaining reconstruction quality. These results demonstrate the effectiveness of OMG4 as a practical framework for compact and high-fidelity 4D scene representation.


Optimizing Analytic Constants via AI-Guided Lean Proof Refinement

Rahul Saha ⋅ Alan Li ⋅ Anton Xue ⋅ Adam Klivans ⋅ Pravesh K Kothari ⋅ Raghu Meka ⋅ Swarat Chaudhuri

We demonstrate that AI agents can autonomously improve mathematical proofs of quantitative results. As a key step towards this frontier, we achieve strict improvements on Grothendieck's and Korenblum's constants, whose precise values have remained surprisingly elusive despite their recurrence in both the mathematical and physical sciences. By encoding relevant literature as a partially formalized Lean blueprint, we structure the agent's search space to localized proof refinements that automatically propagate to a verified final bound. Concretely, Grothendieck's constant ($\mathcal{K}_G$) is shown to be: $$ \mathcal{K}_G \leq \tfrac{\pi}{2 \log (1 + \sqrt{2})} - 10^{-17}, $$ which is the first explicit numerical improvement since Krivine's landmark 1979 result. For Korenblum's constant ($\mathcal{K}_K$), the lower bound is tightened to: $$ \mathcal{K}_K \geq 0.3554 + 0.001165, $$ a third-digit improvement over Wang's recent 2025 result. In summary, we establish that AI agents can autonomously refine formal proofs of open problems and provide a generalizable framework for future mathematical discovery.


ORCA: Hunting Compositional Failures in Text-to-Image Diffusion

Arshia Hemmat ⋅ Amirhossein Vahidi ⋅ Amitis Shidani ⋅ Mohammad V Sanian ⋅ Hesam Asadollahzadeh ⋅ Aryan Y Parast ⋅ Mohammad Lotfollahi

Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal information is concentrated in a low-rank subspace of self-supervised visual features, and that supplying it can be folded into diffusion training as a single auxiliary loss. Our method, ORCA, Orthogonal Residual Compositional Alignment, aligns the latent of a diffusion transformer with a low-rank target derived from a frozen visual encoder, through a predictor whose orthogonal basis is parameterised by a learned residual between T5 and CLIP embeddings, which provides a prompt-dependent signal for selecting the visual readout subspace. We prove that the cross-modal information recoverable at a given rank is bounded by the spectral mass of the visual encoder's covariance in the top components. Across three diffusion-transformer backbones, DiT-B/2, DiT-L/2, and U-ViT-L, ORCA improves FID and GenEval over both vanilla and REPA baselines at zero inference-time cost; on DiT-L/2 it reaches FID 16.65 and GenEval 0.291 at 200K steps, exceeding the strongest 400K baseline at half the training cost, with the largest gains concentrated on attribute binding, spatial relations, and multi-object prompts.


OroPrecipBench: a km-scale benchmark for spatial precipitation downscaling over complex terrain

Hongyi Chen ⋅ Xiaokai Yan ⋅ Jiadong Zhang ⋅ Jingtao Ding ⋅ Xiaojun Liang ⋅ Yong Li ⋅ Xiao-Ping (Steven) Zhang

Precipitation downscaling over complex terrain is essential for resolving water availability, flood risk, and extremes, but current evaluations rarely test whether models learn orographic structure rather than only pixel-wise accuracy. We introduce OroPrecipBench, a km-scale benchmark for orographic precipitation downscaling across five U.S.\ terrain regimes. The benchmark provides a shared-domain dual-track protocol: a perfect-model track that coarsens 3\,km HRRR analyses to isolate controlled super-resolution skill, and an observation-grounded track that downscales 25\,km ERA5 to 4\,km PRISM to test practical robustness under source-target shift. We also introduce PEP-D, a terrain-conditioned diagnostic that measures errors in elevation--precipitation profiles, complementing pointwise, spatial, and extreme-precipitation metrics. Evaluating ten statistical, deterministic deep-learning, and generative baselines shows that strong controlled super-resolution does not reliably translate to real-case or cross-regime downscaling. Models that reduce pixel-wise error can still distort terrain-dependent precipitation and upper-tail rainfall, and terrain-conditioned structure is especially fragile under transfer. OroPrecipBench exposes these failure modes and motivates source-shift-robust, terrain-constrained, and tail-aware downscaling methods. Benchmark is available \href{https://anonymous.4open.science/r/OroPrecipBench-612D}{here}.


OS-Omni: A Cross-Platform Benchmark for Generalist Computer-Using Agents

Hui Shen ⋅ Yunta Hsieh ⋅ Jianing Ma ⋅ Ziyuan Liu ⋅ Qi Han ⋅ Xiuqi Xu ⋅ Yanheng shang ⋅ Chenqi Yin ⋅ Zesen Zhao ⋅ Boyuan Zheng ⋅ Jisen Li ⋅ Xiaoxia Wu ⋅ Ruoyan Zhang ⋅ Dongfu Jiang ⋅ Ruiyao Liu ⋅ Jing Xiong ⋅ He Xiao ⋅ Nan Zheng ⋅ Hangxiang Chen ⋅ Shengyang Tao ⋅ Deyang Luorong ⋅ Jingxuan Zhang ⋅ Chaofan Tao ⋅ Yuren Wang ⋅ Jiaqi Mo ⋅ Hei Ting Chan ⋅ Sicheng Chen ⋅ Kai Zhang ⋅ Mengkang Hu ⋅ Zhongwei Wan ⋅ Xin Wang ⋅ Chuanyang Zheng ⋅ Ziheng Zhang ⋅ Hengyuan Zhang ⋅ Zunhai Su ⋅ Kanzhi Cheng ⋅ Jiawei Lu ⋅ Yifan Zhang ⋅ Shen Yan ⋅ Ben Athiwaratkun ⋅ Qiushi Sun ⋅ Mi Zhang ⋅ Ping Luo ⋅ Wenhu Chen ⋅ Ngai Wong

Computer-using agents are increasingly expected to complete realistic user goals rather than isolated GUI actions. However, existing evaluations are fragmented across web, mobile, and desktop settings, making it difficult to measure whether the same agent can sustain long-horizon workflows across heterogeneous software platforms while preserving final-state correctness. We introduce OS-Omni, a 1,200-task cross-platform benchmark and execution environment for evaluating computer-using agents across Windows, Ubuntu, macOS, iOS, Android, and Web environments. OS-Omni provides reproducible initialization, platform-specific adapters, trajectory logging, and executable final-state evaluators, and includes controlled service-backed applications for workflows involving account, cart, budget, calendar, message, and other persistent state. Human studies show that OS-Omni tasks require 67.54 interaction steps, substantially exceeding OSWorld references of 15.0 steps. Throughout our experiments, the strongest evaluated agent, GPT-5.5, reaches 47.1% average success, trailing the 88.1% human reference by 41.0 points; on hard workflows, it drops from 77.8% success on easy tasks to 16.4%, while humans remain at 76.3%. These results indicate that current agents remain far from platform-robust, long-horizon computer use, especially when success requires sustained state tracking and durable final-state changes.


Overcoming Attention Distraction: Training-Free Latent Communication for Multi-Agent Systems

zhengjie zhou ⋅ Yahao Liu ⋅ Ziyue Feng ⋅ Tengfei LIU ⋅ Weiqiang Wang

Scaling multi-agent systems driven by large language models is transitioning from lossy text-based communication to continuous latent interaction. However, existing latent interaction approaches relying on concatenation-based Softmax cross-attention suffer from an Attention Distraction Dilemma: due to the inherent anisotropy of latent spaces and the semantic misalignment induced by agent role heterogeneity, non-linear Softmax structurally amplifies irrelevant tokens, diluting the collaborative signal as interactions deepen. To address this issue, we propose a training-free additive mechanism that models upstream representations as a semantic manifold. By shifting from concatenation-based Softmax competition to linear subspace projection, this perspective inherently bypasses scalar noise accumulation. Specifically, we introduce two components: a Resonance Filter that extracts principal semantic directional alignments to overcome latent anisotropy, and an SNR-Adaptive Spectral Gain that dynamically regulates this projection to counteract role heterogeneity. We analyze this approach using the Asymptotic SNR Tipping Point Theorem, illustrating the scaling advantages of the mechanism beyond a critical noise threshold. Experiments indicate that the proposed approach outperforms baselines in 46 of 54 evaluated configurations, maintaining semantic reasoning in long-context collaborations.

Reinforcement Fine-Tuning (RFT) with verifiable rewards has emerged as a powerful learning paradigm for eliciting reasoning capabilities in Multi-modal Large Language Models (MLLMs). Recent studies also suggest that RFT (e.g., GRPO) is inherently more resilient to catastrophic forgetting than Supervised Fine-Tuning (SFT). However, whether RFT can effectively overcome forgetting in challenging visual continual learning settings, such as class-incremental learning (CIL) and domain-incremental learning (DIL), remains an open problem. Through a pilot study on rehearsal-free CIL, we confirm that while RFT consistently outperforms SFT, it still suffers from non-negligible forgetting. We empirically trace this bottleneck to Trajectory-level Drift Agnosticism: among candidate rollouts achieving identical task rewards, the KL divergence from the preceding-task policy varies substantially, which strongly correlates with catastrophic forgetting across sequential tasks. Motivated by this insight, we propose Retention-aware Policy Optimization (RaPO), a simple yet effective RFT framework that explicitly mitigates forgetting through trajectory-level reward shaping. Specifically, RaPO comprises two core components: (1) Retention Reward that converts trajectory-level distribution drift into a continuous reward signal, preferentially reinforcing knowledge-preserving rollouts within each group; (2) Cross-Task Advantage Normalization (CTAN), which maintains a persistent exponential moving average of reward statistics across task boundaries to stabilize the optimization progress during continual learning. Leveraging the free-form textual generalization of MLLMs, we comprehensively evaluate RaPO across five visual continual learning settings, spanning class- and domain-incremental image classification, class-incremental video classification, and class- and domain-incremental object detection. Extensive experiments demonstrate that RaPO achieves leading performance, substantially reducing catastrophic forgetting while preserving strong plasticity. To the best of our knowledge, this work represents the first systematic exploration of RFT in visual continual learning, offering insights that we hope will inspire future research.


PACE: A Proxy for Agentic Capability Evaluation

Yueqi Song ⋅ Lintang Sutawika ⋅ Jiarui Liu ⋅ Lindia Tjuatja ⋅ Jiayi Geng ⋅ Yunze Xiao ⋅ Daniel Lee ⋅ Aditya B Soni ⋅ Vincent Lo ⋅ Xiang Yue ⋅ Graham Neubig

\begin{abstract} Evaluating large language model (LLM) agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to complete. In contrast, non-agentic LLM benchmarks that test individual capabilities (e.g., reasoning, code generation, instruction following) are fast and cheap to run. We investigate whether performance on expensive agentic benchmarks can be accurately predicted by the performance on a small, carefully selected subset of atomic evaluation instances. We introduce PACE, a framework that constructs proxy benchmarks by selecting instances from existing non-agentic evaluations whose aggregate scores most reliably predict agentic benchmark performance. Given a pool of candidate instances spanning atomic capabilities (instruction following, planning, tool calling, etc.), PACE fits a regression that maps a model's scores on a compact subset of source instances to its score on the target agentic benchmark; the subset itself is produced by combining two complementary instance-selection strategies, a target-relevance local selection and a globally informative global selection. Experiments across 14 models, 4 agentic benchmarks, and 19 non-agentic benchmarks show that PACE-Bench predicts agentic scores with leave-one-model-out (LOOCV) mean absolute error (MAE) under $4\%$, Spearman correlation above $0.80$, and pairwise model-ranking accuracy around $85\%$, all at much less than $1\%$ of the full agentic evaluation cost. We further analyze the selected proxy instances, revealing what skills each agentic benchmark uniquely demands. PACE enables practitioners to obtain reliable estimates of agentic performance during model development, selection, and routing, without the overhead of full agent evaluation.


PACO: Partial-order-Augmented Continuous Optimization for Differentiable Causal Discovery

Xiaoxuan Li ⋅ Junda Wu ⋅ Julian McAuley ⋅ Lina Yao ⋅ Tong Yu

Differentiable structure learning transforms DAG discovery from a combinatorial search into a continuous optimization problem, enabling gradient-based recovery of causal structures. In many applications, researchers possess partial-order priors over the variables (e.g., from biological signaling cascades or temporal ordering), which substantially narrow the space of plausible causal graphs. Existing approaches integrate such priors by decomposing the prior order graph into a collection of paths and evaluating prior-aware acyclicity constraints path by path; this is correct but its constraint-evaluation cost scales linearly with the number of induced maximal paths, a bottleneck for multi-chain priors. We propose PACO, Partial-order-Augmented Continuous Optimization, which encodes the same partial-order compatibility through a single masked-reachability matrix, that masks the matrix exponential of the learned graph by the strict transitive closure of the prior. The constraint and its Frechet-adjoint gradient are computed in a single $2d{\times}2d$ block matrix exponential whose dominant cost is independent of the number of induced prior paths. We prove that PACO has the same exact zero-level feasible set as the path-decomposition formulation. Across synthetic and real-data experiments under matched priors and dataset seeds, PACO matches path-decomposition on structural recovery and substantially reduces wall-clock cost in multi-chain prior settings.


PAGER: Bridging the Semantic-Execution Gap in Point-Precise Geometric GUI Control

Jingxuan Wei ⋅ Xi Bai ⋅ Shan Liu ⋅ caijun jia ⋅ Zheng Sun ⋅ Xinglong Xu ⋅ Siyuan Li ⋅ Linzhuang Sun ⋅ Bihui Yu ⋅ Conghui He ⋅ Cheng Tan

Large vision-language models have significantly advanced GUI agents, enabling executable interaction across web, mobile, and desktop interfaces. Yet these gains largely rely on a forgiving \emph{region-tolerant} paradigm, where many nearby pixels inside the same component remain valid. Precise geometric construction breaks this assumption: actions must land on points in continuous canvas space rather than tolerant regions. Because geometric primitives carry ontological dependencies, a local coordinate error can induce cascading topological failures that distort downstream objects and invalidate the final construction. We identify this regime as \emph{precision-sensitive GUI tasks}, requiring point-level accuracy, geometry-aware verification, and robustness to dependency-driven error propagation. To benchmark it, we introduce \textsc{PAGE} Bench, with 4{,}906 problems and over 224K process-supervised, pixel-level GUI actions. We further propose \textsc{PAGER}, a topology-aware agent that decomposes construction into dependency-structured planning and pixel-level execution. Pixel-grounded supervised tuning establishes executable action grammar, while precision-aligned reinforcement learning mitigates rollout-induced exposure bias through state-conditioned geometric feedback. Experiments reveal a pronounced ``Semantic-Execution Gap'': general multimodal models can exceed 88\% Action Type Accuracy yet remain below 6\% Task Success. \textsc{PAGER} closes this gap, delivering $4.1\times$ higher Task Success than the strongest evaluated general baseline and raising Step Success Rate from below 9\% for GUI-specialized agents to over 62\%, establishing a new state of the art for point-precise GUI control.


Paloma: Phase-Conditioned Residual Modulation for Time Series Forecasting

Jingru Fei ⋅ Kun Yi ⋅ Wei Fan ⋅ Qi Zhang ⋅ Zhendong Niu

Explicitly modeling periodicity has proven to be an efficient approach for time series forecasting, where the residual, obtained after removing the periodic component from the time series, is typically treated as unstructured noise. However, in this work, we show that such residuals still exhibit structured, \textit{phase-dependent} deviations, as supported by both empirical observations and theoretical analysis. This suggests that residuals should not be treated merely as unstructured noise, but instead as \textit{phase-conditioned signals} whose variations arise in both direction and magnitude. Motivated by this perspective and inspired by phase and amplitude modulation in classical signal processing, we propose \textit{Paloma}, a simple yet effective framework for modeling residual deviations. Specifically, Paloma comprises two complementary modules: \textit{a Phase Rotation module}, which performs phase-conditioned rotations to capture directional changes in the feature space, and \textit{an Amplitude Scaling module}, which applies phase-conditioned affine transformations to model magnitude variations. We further abstract the two modules into a general-purpose plugin, termed the \textit{Paloma technique}, which can be readily integrated into existing forecasting architectures to modulate input sequences in a phase-aware manner, enabling backbone models to capture phase-dependent variations in both direction and magnitude. Extensive experiments on twelve datasets demonstrate that Paloma achieves state-of-the-art and robust forecasting performance. Additional results show that the Paloma technique consistently improves a wide range of baseline models, validating its effectiveness as a general-purpose component.


PaLoRA: Paced Low-Rank Adaptation for Continual Learning

Yuxuan Li ⋅ Fanhu Zeng ⋅ Hao Tang

LoRA-based continual learning methods mitigate catastrophic forgetting through various mechanisms, yet nearly all complement these with **small learning rates as a heuristic to restrict gradient scaling magnitude**. Such fixed heuristics lack theoretical guidance on how the strength of this restriction should evolve as tasks accumulate. We reveal that even under directional constraints such as nullspace projection, finite-precision updates inevitably leak into the subspace of accumulated prior knowledge along multiple directions. While small learning rates attenuate such leakage, they cannot prevent the accumulated forgetting from intensifying as the effective rank of historical knowledge grows. We show that the optimal magnitude restriction should adaptively increase with this effective rank to balance **stability and plasticity**, *i.e.*, preservation of previous knowledge and acquisition of new task information. Under an anisotropic leakage model, we derive a pacing law $s^*=\sqrt{R/c}$ that characterizes the optimal **scaling of gradient steps,** *i.e.*, the magnitude restriction itself, where $R$ is the effective rank of past updates. Based on this insight, we propose PaLoRA, which compresses historical knowledge via adaptive SVD truncation, projects gradients onto the nullspace of prior tasks, and applies rank-aware adaptive pacing. Experiments demonstrate consistent improvements over prior methods, with particularly strong performance in long-horizon settings, achieving substantial gains of 4% accuracy on challenging 50-task ImageNet-A and ImageNet-R benchmarks.

Reconstructing nonlinear dynamical systems (DS) from data (DSR) is a fundamental challenge in science and engineering, but it inherently relies on *sequential* models. Recent breakthroughs for sequential models have produced algorithms that parallelize computation along sequence length $T$, achieving logarithmic time complexity, $\mathcal{O}(\log T)$. Since sequence lengths have been practically limited due to the linear runtime complexity $\mathcal{O}(T)$ of classical backpropagation through time, this opens new avenues for DSR. This paper studies two prominent classes of parallel-in-time algorithms for this task, both of which leverage *parallel associative scans* as their core computational primitive. The first class comprises models with linear yet non-autonomous dynamics and a nonlinear readout, such as modern State Space Models (SSMs), while the second consists of general nonlinear models which can be parallelized using the DEER framework. We find that the linear training-time recurrence of the first class of models imposes limitations that often hinder learning of accurate nonlinear dynamics. To address this, we augment DEER with Generalized Teacher Forcing (GTF), a novel variant within the more general nonlinear framework that ensures stable and effective learning of nonlinear dynamics across arbitrary sequence lengths. Using GTF-DEER, we investigate the benefits of training on extremely long sequences ($T>10^4$) for DSR. Our results show that access to such long trajectories significantly improves DSR if the data features long time scales. This work establishes GTF-DEER as a robust tool for data-driven discovery and underscores the largely untapped potential of long-sequence learning in modeling complex DS.


PARI: Policy-Driven Active Residual Intervention for Weakly Supervised Point Cloud Segmentation

Wenhao Gao ⋅ Sen Tao ⋅ Jiawei Liu ⋅ Jiwei Feng ⋅ Yongchao Xu ⋅ Junfeng Wang ⋅ Tao Jiang ⋅ jiangbo Ai ⋅ Jin Zhang ⋅ Zheng-Jun Zha

3D Weakly Supervised Semantic Segmentation (3D-WSSS) on point clouds aims to learn dense 3D semantics from extremely sparse point-level annotations. Existing methods follow a passive paradigm based on pseudo-label propagation or consistency regularization, leading to two key issues: 1) Lack of targeted intervention, where erroneous predictions are not explicitly identified and corrected, resulting in severe error accumulation; 2) Limited contextual awareness, where supervision is restricted to point- or view-level self-augmentation consistency, failing to capture the intricate contextual correlations embedded within the scene. To address these limitations, we propose a Policy-driven Active Residual Intervention (PARI) framework for 3D-WSSS, where an Active Intervention Agent (AIA) explicitly localizes high-risk points under the guidance of a Confidence-aware Reward Shaping (CRS) strategy, and a Bias-aware Refinement Module (BRM) subsequently executes targeted refinement by exploiting rich spatial and semantic scene contexts. Specifically, AIA actively localizes high-risk points from uncertainty maps and prediction priors, while BRM refines them through residual correction based on hybrid spatial-feature context. CRS introduces a high-confidence suppression reward that discourages unnecessary intervention on reliable predictions, and a low-confidence exploration reward that encourages intervention on uncertain predictions. Experiments on S3DIS and ScanNetV2 demonstrate state-of-the-art performance under various weak supervision settings with strong cross-backbone generalization, even surpassing the same-backbone fully supervised counterpart with 1% annotations.

Model inversion is a widely adopted technique in data-free learning that reconstructs inputs from a pretrained model through iterative optimization, without access to original data. However, its application to Vision Transformers (ViTs) incurs high computational cost due to expensive self-attention mechanisms. To address this, $\textit{Sparse Model Inversion}$ (SMI) was proposed to improve efficiency by gradually pruning seemingly unimportant patches, even claiming they are obstacles to knowledge transfer. However, our empirical findings suggest the opposite: even randomly selected patches can eventually acquire transferable knowledge over the inversion process. In fact, we further observe that removing prematurely inverted patches hinders the extraction of class-agnostic features essential for knowledge transfer, as well as class-specific features. In this paper, we propose $\textit{Patch Rebirth Inversion}$ (PRI), a novel approach that constructs multiple sparse images within a single inversion process by incrementally detaching informative patches, instead of removing unimportant ones. This strategy not only improves efficiency, but also encourages initially less informative patches to gradually accumulate more class-relevant knowledge, a phenomenon we refer to as the $\textit{Re-Birth}$ effect, thereby effectively balancing class-agnostic and class-specific knowledge. Experimental results show that PRI achieves up to 10$\times$ faster inversion than standard $\textit{Dense Model Inversion}$ (DMI) and 2$\times$ faster than SMI, while consistently outperforming SMI in accuracy and matching the performance of DMI.


PATH: A Dual Perspective for High-quality Text-attributed Graph Learning

Yuhang Pei ⋅ Fanchun Meng ⋅ Changhu Wang ⋅ Tao Ren ⋅ Yifan Wang ⋅ Wei Ju ⋅ Chong Chen ⋅ Xian-Sheng Hua ⋅ Xiao Luo

This paper studies the problem of zero-shot text-attributed graph learning, which aims to generate high-quality node representations in unseen text-attributed graphs. Recent approaches usually utilize large language models (LLMs) instead of graph neural networks (GNNs) to extract semantics due to their strong generalization ability, which could neglect the intrinsic geometric structure. Towards this end, in this paper, we propose a novel approach named Prototypical Mutual Prompting Enhancement (PATH) for zero-shot text-attributed graph learning. The core of our PATH is to generate high-quality prompts using dual prototypical learning to combine the advantages of both language models and graph models. In particular, we first utilize dual graph pre-training from both instance and informativeness perspectives to generate a generalizable GNN. Then, we incorporate the frozen language and graph models into a mutual prompt learning framework. On the one hand, we extract node tokens with geometric relationships using the graph model, which will be sent to multiple prototypical projections to enhance the understanding of the language model. On the other hand, we extract graph information and task descriptions using the language model, which serves as instruction for the graph models. Extensive experiments on node classification and link prediction validate the effectiveness of the proposed PATH against strong baselines. Our code is available at https://anonymous.4open.science/r/PATH.


Path-independent Flow Matching for Multi-parameter Generative Dynamics

Francisco Téllez ⋅ Amirhossein Zamani ⋅ Philippe Martin ⋅ Shuang Ni ⋅ Guy Wolf ⋅ Eugene Belilovsky ⋅ Sina Sanjari ⋅ Yanlei Zhang

Flow Matching is a powerful framework for learning transport maps between probability distributions. Yet its standard single-parameter formulation is not designed to capture multi-parameter variations where the resulting transport should be path-independent. Path independence is crucial because it ensures that transformations depend only on the initial and target distributions, not on the specific path. In this work, we introduce \textit{Path-independent Flow Matching (PiFM)}, a method for learning vector fields whose induced flows yield path-independent transport between distributions. We show that PiFM generalizes Flow Matching to higher-dimensional parameter domains while enforcing structural conditions that ensure consistency of composed transformations. In addition, we show that, under suitable assumptions, PiFM approximates the Wasserstein barycenter, linking the framework to a notion of distributional interpolation. To enable practical training, we propose a tractable, simulation-free objective that regresses onto multi-parameter conditional probability paths. We showcase empirically that PiFM outperforms other approaches on both synthetic and real world data in interpolating path-independent trajectories and generating desired out of distribution samples. Our code is available at \url{https://anonymous.4open.science/r/Piflowmatching-C83F/}.

The success of vision transformers—especially for generative modeling—is limited by the quadratic cost and weak spatial inductive bias of self-attention. We propose PDE-SSM, a spatial state-space block that replaces attention with a learnable convection–diffusion–reaction partial differential equation. This operator encodes a strong spatial prior by modeling information flow via physically grounded dynamics rather than all-to-all token interactions. Solving the PDE in the Fourier domain yields global coupling with near-linear complexity, delivering a principled and scalable alternative to attention. We integrate PDE-SSM into a flow-matching generative model to obtain the PDE-based Diffusion Transformer PDE-SSM-DiT. Empirically, PDE-SSM-DiT matches or exceeds the performance of state-of-the-art Diffusion Transformers while substantially reducing compute. Our results show that, analogous to 1D settings where SSMs supplant attention, multi-dimensional PDE operators provide an efficient, inductive-bias-rich foundation for next-generation vision models.

Diffusion image generators owe much of their practical success to accurate denoising networks and strong numerical samplers. Yet the two components are not fully matched. During sampling, the denoiser is asked to make predictions based on imperfect noisy states formed by model predictions that differ from the training noisy states formed by real data. This train-test gap, known otherwise as exposure bias in literature, caps image generation quality (i.e., blurry, lacking in details or unnaturally-shaped). To mitigate the problem, we propose PEEK, a training-free deterministic sampler plugin. At each eligible denoising step, PEEK first performs a provisional one-step update, evaluates the denoiser once at the look-ahead state, and either keeps the current clean prediction or replaces it with this one-step-later prediction to form the actual noisy state. By taking the additional step, we generally obtain a higher-signal, more-image-like prediction for the noisy state, mitigating the exposure bias. To optimize the replacement positions, PEEK searches for a binary replacement schedule with a greedy metric-driven search made efficient by the PEEK prefix and baseline noise-schedule-rebasing caches; for a 9-step search, the caches reduce the one-candidate worst-case vanilla denoiser-evaluation count by 69.3\%. Across pixel-space and latent-space diffusion models with unconditional, label-guided, and text-guided settings on CIFAR-10, CelebA-HQ, LSUN Bedroom, ImageNet, and MS-COCO datasets, PEEK improves over NFE-matched DPM-Solver++ and DEIS baselines, with five-run mean FID reductions up to 39.3\% under the same deployment cost. Code will be released.

Non-contrastive self-supervised learning (SSL) is an effective framework for predictive representation learning, but popular (and in practice effective) methods such as SimSiam, BYOL, I-JEPA or DINO, which rely on a form of self-distillation to train a teacher-student network, remain poorly understood as they typically do not minimize a well-defined objective. We analyze the dynamics of a variant of the Joint Embedding Predictive Architecture (JEPA) using a regularized linear regressor to predict the learned representations of two views of the data from one another, and fully characterize its stability: non-collapsed stable equilibria align with leading nonlinear canonical correlation subspaces, while collapsed equilibria may also be stable attractors. Motivated by this result, we introduce PEIRA, a non-contrastive SSL method with an explicit objective defined through the trace of the optimal linear regressor. We show that its only stable equilibria are nontrivial global minimizers and recover the same canonical correlation subspaces, with regularization selecting the effective dimension. Experiments on ImageNet-1K and CIFAR-10 show PEIRA is competitive with VICReg and LeJEPA baselines, and qualitative empirical results support the theory.

While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and multi-step visual interaction. This limitation suggests that relying solely on implicit visual representations from vision encoders is insufficient for recovering fine-grained spatial evidence. We introduce PERception-Interaction-reason Agent (PERIA), a tool-augmented visual agent for spatial reasoning tasks across map reasoning, visual probing, and vision reconstruction. PERIA uses two lightweight tool families: vision perception tools for exposing textual, symbolic, and spatial evidence, and vision interaction tools for manipulating visual context, tracing paths, and verifying spatial relations. To train PERIA, we develop a unified recipe that combines supervised tool-use trajectory synthesis, composite rewards, and Observation-Relaxed Group-in-Group Policy Optimization (OR-GIGPO) for effective multi-tool behavior. Experiments on 13 benchmarks from 8 datasets show that PERIA-8B improves over the Qwen3-8B backbone by 10.0% on in-distribution benchmarks and 4.4% on out-of-distribution benchmarks, while outperforming previous state-of-the-art baselines of similar size by 7.0%–14.8%. It also achieves performance comparable to much larger models such as Qwen3-VL-235B-A22B-Thinking and GPT-5, demonstrating the effectiveness of PERIA in enhancing spatial reasoning capabilities.


PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

Shuai Yang ⋅ Bingjie Gao ⋅ Ziwei Liu ⋅ Jiaqi Wang ⋅ Dahua Lin ⋅ Tong Wu

Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time and viewpoints. However, existing memory designs struggle to maintain long-term consistency after such modifications, as stored contexts may become outdated or invalid. To address this, we propose PermaVid, a novel framework built upon a multi-modal context memory that disentangles spatial context into semantic appearance and geometric structure, together with an edit-aware memory update and retrieval strategy that keeps memory evolution aligned with subsequent observations. Specifically, we develop two complementary memory banks: an RGB context memory that captures appearance-aware observations while implicitly encoding geometry, and a depth context memory that preserves geometry-only structure disentangled from semantics. Building on this design, we introduce a memory-guided video generation model that performs multi-modal feature fusion under reference conditions drawn from mixed-modality memory contexts. Experiments demonstrate that our method maintains strong long-term semantic and structural consistency after edits, significantly outperforming state-of-the-art methods.


Permit: Permission-Aware Representation Intervention for Controlled Generation in Large Language Models

Pengcheng Sun ⋅ Lan Zhang ⋅ Zhaopeng Zhang ⋅ Jiewei Lai ⋅ Chen Tang

Large language models (LLMs) are increasingly deployed in enterprise settings where they handle sensitive documents and user context, raising acute concerns over security and controllability. Conventional access control regulates whether information is accessible to the model, yet leaves how the model uses that information at generation time largely unconstrained: once sensitive content enters the context, outputs may still drift beyond a user's authorized scope. We present Permit, a novel permission-aware representation intervention framework that closes this gap by enforcing fine-grained control directly on the model's hidden states. Through exploratory analysis, we find that permission conditions induce hidden-state shifts that are (i) separable across permissions and (ii) concentrated in a small set of dominant directions. Permit exploits this geometry in two stages: it first identifies a permission-sensitive subspace from activation differences across permission conditions, and then performs lightweight interventions within this subspace to steer generation, with two concrete instantiations (offset-based and gated). Both operate atop a frozen backbone with only a handful of permission-specific parameters, achieving precise control with minimal overhead. Experimental results demonstrate that \textsc{Permit} performs better than the state-of-the-art method across multiple permission settings while driving information leakage to near zero, achieving over (18\%) F1-score improvement with (>98\%) fewer trainable parameters.


Permutation Sensitivity in t-SVD-based Multi-view Clustering

Jintian Ji ⋅ Yanjun Zhang ⋅ He Zhang ⋅ Leo Yu Zhang ⋅ Cong Wang ⋅ Shirui Pan

Tensor Singular Value Decomposition (t-SVD) has achieved strong empirical performance in multi-view clustering (MVC) and its incomplete setting (IMVC). However, the success of this framework heavily relies on the implicit dependency of its Fourier transform on the sample sequence, which raises serious concerns regarding sample-permutation sensitivity. To address this issue, we first formulate the permutation sensitivity in multi-view clustering by providing rigorous definitions of invariance at both the outcome level and the representation level. Secondly, by revealing the potential structural risks associated with the fixed Fourier basis, we propose two plug-and-play strategies: (1) an Order-Restoration strategy (K-OR), designed to reconstruct a locally smooth signal favored by the Fourier transform; and (2) an equivariant t-SVD framework (Cor-SVD), which fundamentally resolves the sensitivity of the Fourier transform by replacing fixed bases with data-dependent, equivariant transforms. Finally, we design a controlled permutation evaluation protocol to verify the robustness of tensorial clustering models. Extensive experiments demonstrate the significant performance variability of existing t-SVD-based methods under sample permutations, whereas our approaches substantially improve their permutation robustness while maintaining competitive clustering performance. These findings reveal critical vulnerabilities in current tensorial clustering benchmarks and underscore the necessity of the proposed method for achieving robust performance and reliable evaluation.


Persistent Planar Memory for Video World Models

Yuze He ⋅ Bin Tan ⋅ Zelin Gao ⋅ Yujun Shen ⋅ Yong-jin Liu ⋅ Nan Xue ⋅ Yinghao Xu

Video world models that generate and explore environments at scale need a persistent 3D memory of what they have produced. Plane primitives offer a natural candidate: a compact set of oriented planes captures the dominant structural elements (*e.g.,* walls, floors, facades) of indoor and outdoor scenes, while requiring only 7 parameters per primitive. We present **PPMem**, a video world model that adopts plane primitives as its persistent 3D memory. PPMem augments a video diffusion transformer with a planar prediction head and a Depth Token Fusion (DTF) module: each autoregressive chunk jointly produces RGB video and planar geometry, which merges into a persistent plane memory that conditions subsequent chunks via rendered depth. The resulting memory is three orders of magnitude more compact than per-pixel point clouds (${\sim}$100K primitives, ${\sim}$5MB for a full scene), geometrically explicit (metric depth, normals, exportable surfaces), cross-view consistent (jointly predicted by the shared backbone), and adaptive (refined by the merging procedure as generation proceeds). On DL3DV and RealEstate10K, PPMem outperforms prior 3D-grounded world models in video quality while jointly producing metric-scale depth, surface normals, and exportable planar geometry, sustaining coherence across hundreds of frames under large viewpoint changes.


PHIONet: Port Hamiltonian Inertial Odometry Network

Hechuan Shen ⋅ Jian Zhang ⋅ Yingmin Liang

Learning-based inertial odometry has shown strong potential for mitigating drift by extracting data-driven motion priors. However, existing methods primarily learn statistical mappings from inertial measurements to motion states, failing to consider the power balance between the specific-force input port and the predicted velocity. This limitation is particularly pronounced in aggressive MAV flight, where rapid accelerations, high angular rates, and aerodynamic dissipation can cause locally accurate velocity estimates to accumulate into trajectory drift. We propose PHIONet (Port Hamiltonian Inertial Odometry Network), which incorporates dissipative energy dynamics as a physically grounded prior. PHIONet translates the continuous port-Hamiltonian dynamics into a discrete energy-consistency residual over finite IMU windows, employing the DCM (Dissipative Coupling Module) to bridge IMU correction and velocity prediction. Experiments on the public Blackbird and EuRoC MAV datasets show that PHIONet achieves lower ATE and RTE than representative learning-based inertial odometry baselines. Code and related materials are available at https://anonymous.4open.science/r/PHIONet-D3CC.


PHMForge: Evaluating LLM Agents on Industrial Prognostics through MCP-Native, Algorithm-Grounded Tools

Yusheng Li ⋅ Tianjun Feng ⋅ Yunfeng Chen ⋅ Chun-Yi Tsai ⋅ Yihan Sun ⋅ Dhaval Patel ⋅ Kaoutar El Maghraoui ⋅ Shuxin Lin ⋅ Ayan Das

LLM agents are beginning to invoke industrial asset-management tools through the Model Context Protocol (MCP), yet whether they can act reliably on this substrate for safety-critical Prognostics and Health Management (PHM) is unanswered. Prior benchmarks conflate protocol fluency with reasoning, instrumentation failures with agent failures, and tool use with tool retrieval. We introduce PHMForge, an evaluation environment that closes each conflation. PHMForge ships 99 SME-authored scenarios across eight industrial asset classes spanning rotating equipment, aero-engines, and lithium-ion battery cells, on real public datasets including NASA PCoE. The benchmark is served through 39 MCP-native tools that wrap published PHM algorithms (e.g., C-MAPSS, ISO 10816, Arrhenius capacity-fade models, and time-series foundation models for sequence forecasting). Krippendorff's α ∈ [0.74, 0.82] on a 30-scenario stratified rotating-equipment/aero-engine sample; the lithium-ion battery extension is single-rater. Across three agentic frameworks and six LLM backbones, the strongest configuration reaches 80.8% pass@1 on the full 99-scenario set, with the residual gap concentrated in orchestration and tool-sequencing errors. Crucially, an architectural ablation reveals that replacing MCP tool execution with text-based Retrieval-Augmented Generation (RAG) over telemetry-equivalent evidence collapses Remaining Useful Life (RUL) prediction pass-all-3 from 100% to 20% (5/5 vs. 1/5 scenarios) on the lithium-ion battery class, exposing the structural limits of static retrieval for prognostic computation. Trajectory-level decomposition shows orchestration errors dominate failures across backbones, while schema-invalid tool calls are concentrated in smaller open-weight models and rare in frontier configurations. Frontier LLMs are stronger at calling tools than at planning when to call them. PHMForge is open-sourced with deterministic evaluators, a public leaderboard, and a datasheet. A full-suite evaluation costs approximately 20–50 USD per backbone in API spend.


PhotoFlow: Agentic 3D Virtual Photography Missions

Jiarui Guo ⋅ Haojia Wei ⋅ Yiming Zhang ⋅ Yifei Liu ⋅ Yuning Gong ⋅ Hongjie Zhang ⋅ Xue Yang ⋅ Zhihang Zhong

Virtual photography asks an agent to enter a prepared 3D scene with no preselected camera pose or reference image, infer a suitable shot from scene information and a language intent, choose executable camera parameters, and render the final photograph. Recent progress in vision-language models makes this kind of spatial agent increasingly plausible, but the task stresses two capabilities that remain hard to evaluate together: complex 3D spatial understanding and abstract aesthetic judgment. We introduce PhotoFlow, a Director-Reviewer-Reflector agent for closed-loop camera search. The Director builds a soft photographic blueprint and proposes diverse candidate cameras; the Reviewer combines rule checks, visual critique, and pairwise incumbent selection; and the Reflector converts failures into region memory, dead-zone suppression, and high-explore relocation. We also introduce VPhotoBench, a benchmark of 47 open-license Blender scenes and 141 language-conditioned photography missions spanning subject placement, relational composition, and atmosphere/style. On held-out experiments,PhotoFlow achieves the strongest external quality-alignment composite and success rate among one-shot prediction, single-chain reflection, anchor-bank selection, and random search under a six-round rendering budget. To our knowledge, this is the first work to make language-conditioned virtual photography in 3D design or game engines (e.g., Blender) an executable agent task, and our results show that an LLM-centered spatial agent can already produce strong photographs in a setting designed to challenge both 3D reasoning and aesthetic choice.


PhyProbe: Rethinking Physical Consistency Evaluation in Generated Videos

Max KU ⋅ Jiaojiao Fan ⋅ Zekun Hao ⋅ Francesco Ferroni ⋅ Heng Wang ⋅ Wenhu Chen ⋅ Ming-Yu Liu ⋅ Prithvijit Chattopadhyay

Evaluating the physical consistency of generated videos remains a fundamental challenge. Existing approaches rely on off-the-shelf vision-language models, which can often be myopic to physical dynamics, or fine-tuned evaluators trained on human annotations, which overfit to dataset-specific cues and fail to generalize. A key challenge is that existing supervision sources provide either relative ordering or absolute scores, but not both reliably and consistently across varied settings. To this end, we introduce PhyProbe, an evaluator that extracts features from a frozen pretrained spatio-temporal encoder and maps them to a scalar physical consistency violation score via a lightweight scoring head. PhyProbe is trained through a unified objective combining pairwise ranking, regression on noisy scalar annotations, and anchor-based calibration over a curated set of heterogeneous supervision sources. Experiments show that PhyProbe significantly outperforms prior methods on pairwise benchmarks spanning real–generated and generated–generated pairs under varying correspondence, including challenging settings where existing approaches degrade toward random performance. PhyProbe achieves strong correlation with human judgments, with close agreement between rank-based and linear metrics, indicating both reliable ordering and calibrated scores. Further, despite being trained solely for physical consistency, PhyProbe also performs competitively on human preference benchmarks: consistent with the observation that physics violations are entangled with broader quality degradations.


PhysDNet: Physics-Driven Gradient Amplification for Real-Time Image Dehazing

Rayen Dhahri ⋅ Lei Sun ⋅ Victor J Morinigo ⋅ Alice Pagnoux ⋅ Luca Benini ⋅ Danda Pani Paudel ⋅ Luc V Gool

Modern dehazing networks are often computationally expensive and overfit to synthetic haze, limiting real-time deployment. This underscores an urgent demand for efficient, high-fidelity restoration methods that can operate on resource-constrained edge devices in practical, dynamic environments. We introduce PhysDNet, a lightweight multimodal architecture that fuses NIR imagery with sparse LiDAR depth and is trained with a physics-constrained dual-head objective. Differentiating through a fixed Koschmieder inversion induces density-adaptive gradient amplification, acting as an implicit resource allocation mechanism that concentrates learning on heavily degraded regions without increasing inference cost. PhysDNet delivers superior restoration with minimal capacity and enhanced training efficiency, facilitating real-time deployment on edge devices without increasing inference overhead. On the Seeing Through Fog benchmark, PhysDNet outperforms competitors by 1.67 dB at 5.2x lower compute and recovering 54.9% of fog-induced detection loss. Ablations and cross-domain evaluations show that physics-guided training improves generalization including under non-homogeneous haze that violates the Koschmieder assumption. PhysDNet further demonstrates compatibility with fixed-function NPUs, running at 29/45/99 FPS at <4W across L/M/S variants. Code and pre-trained models will be made publicly available.


PhysFormer: Learning to Simulate Mechanics in World Space

Yiming Chen ⋅ Yushi LAN ⋅ Andrea Vedaldi

We present PhysFormer, a physics-grounded diffusion transformer for generating 4D multi-object mesh dynamics directly in world coordinates. Rather than predicting future frames in pixel space or rolling out next-step system states autoregressively, PhysFormer models physical evolution as full-trajectory coordinate diffusion. Given initial per-vertex positions, velocities, and material conditions, it generates future mesh vertex trajectories in a single denoising process, producing physically plausible object-object and object-environment interactions without hard-coded constraints, simulator priors, or learned shape latents. PhysFormer adopts a DiT-style backbone with factorized temporal, spatial, and object-level attention to capture coherent structure across time, vertices, and objects. Trained on over 100k collision-rich simulated trajectories, it models rigid and deformable multi-object dynamics, generalizes to unseen real-world geometries and larger object counts, and substantially outperforms autoregressive baselines in trajectory accuracy, rigidity preservation, and momentum-based physical consistency. Our results position coordinate-space diffusion as a promising step toward view-invariant, geometry-level world models for robotics, graphics, and physical design.


PhysGuard: Fisher-Guided Gradient Projection for Sim-to-Real Neural PDE Surrogates

Changjian Zhou ⋅ Junfeng Fang ⋅ Negin Yousefpour ⋅ peng wu ⋅ Bin Yan ⋅ Guillermo A Narsilio

Neural operator models trained on simulation data often lose accuracy when applied to experimental measurements due to the sim-to-real gap. Standard fine-tuning with limited real data can reduce this gap, but it may also damage the core physics-relevant representations learned during pretraining. Although knowledge-preserving adaptation has been widely investigated in vision or language tasks, it remains unclear whether these methods are suitable for neural operators whose architectures and protected knowledge are fundamentally different. Neural operators need to preserve core-scale physical structures rather than semantic or visual features. We propose PhysGuard, a physics-preserving framework for accurate sim-to-real adaptation of neural operators. Specifically, PhysGuard uses the empirical Fisher Information Matrix computed on simulation data to identify physics-critical parameter directions, then restricts fine-tuning updates to directions that do not interfere with them. A layer-wise Gram-matrix formulation makes this efficient for models with millions of parameters, while an adaptive threshold automatically determines the protected subspace size. A spectral probe experiment shows that the dominant Fisher directions are strongly associated with low-frequency output structures. Experiments on benchmark across four neural operator architectures and different physical systems show that PhysGuard performs strongly on most evaluation metrics compared to baselines. The benefits are most evident under severe domain shift, where it reduces low-frequency error by up to 32\% compared to standard fine-tuning while maintaining adaptability. Our code is available at https://anonymous.4open.science/r/PhysGuard-B840/.


Physics-Constrained Generative World Model for Off-Road Terrain via Post-hoc Projection

Yuan Zhou ⋅ Hao Yu ⋅ Ruiran Cao ⋅ Haoran Yang ⋅ Xuanyu Zhu ⋅ Yusong Yan ⋅ Cong Wang ⋅ Minne Li

Generative models for physically grounded terrain face a persistent tradeoff: training-time physics losses distort the learned distribution, while unconstrained generators routinely produce physically impossible surfaces. We \emph{decouple} generation quality from constraint satisfaction entirely, enforcing slope stability, bearing capacity, and surface continuity via a closed-form \emph{post-hoc projection} applied to the generator's output. The projection raises the physics pass rate from $53\%$ to $\mathbf{100\%}$ on a diffusion generator with no change in FID ($13.72$), whereas a matched PhysLoss baseline degrades FID to $56.6$ ---a $4{\times}$ loss of generation quality---while reaching only $25.5\%$ pass rate. The projection is \emph{model-agnostic}: on four diverse generators (Procedural, Diffusion, VAE, GAN) it consistently yields $89.5\%$--$100\%$ compliance with $<4.3\%$ overhead. Out-of-distribution evaluation on real TartanDrive terrain confirms transfer: all four generators reach $100\%$ pass rate, while projection simultaneously improves BEV-FID against the real reference by up to $20\%$. We instantiate the layer inside DreamEnv, an end-to-end BEV world model.


PhySPRING: Structure-Preserving Reduction of Physics-Informed Digital Twins via Graph Neural Networks

Yixiong Jing ⋅ Xingyuan Chen ⋅ Xingyuan Chen ⋅ Guangming Wang ⋅ Olaf Wysocki ⋅ Haibing Wu ⋅ Brian Sheil

Physics-based digital twins aim to predict the dynamics of real-world objects under interaction, enabling real-to-sim-to-real applications in robotics. Current approaches reconstruct such twins as explicit physical models (such as spring-mass systems) to predict the dynamics, but the resulting models often inherit the resolution of the visual reconstruction rather than being reduced to the physical complexity required to reproduce task-relevant dynamics. This mismatch introduces redundant topology, making repeated forward-dynamics rollouts unnecessarily expensive. To address this challenge, we present PhySPRING, an fully differentiable GNN-based method to reduce complexity in spring--mass digital twins. PhySPRING jointly learns a hierarchy of coarsened graph topologies and their mechanical parameters from observations. At each reduction level, PhySPRING merges nodes with similar learned dynamic responses to optimize the topology, while maintaining every reduced layer as an explicit spring--mass system. On the PhysTwin benchmark, PhySPRING improves dense reconstruction and prediction accuracy over PhysTwin, while reduced models retain stable physical and visual fidelity with up to a 2.30$\times$ speed-up. We further demonstrate the effectiveness of PhySPRING in a Real2Sim robot policy-evaluation pipeline, where the reduced models are substituted zero-shot into ACT and $\pi_0$ evaluations, maintaining comparable manipulation success rates across downsampling levels while improving action-sampling effectiveness. Together, PhySPRING enables efficient and structure-preserving spring--mass reduction without sacrificing fidelity or robotic utility.


PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion

Yifan Lu ⋅ Qi Wu ⋅ Jay Zhangjie Wu ⋅ Zian Wang ⋅ Huan Ling ⋅ Sanja Fidler ⋅ Xuanchi Ren

Most practical high-resolution text-to-image systems rely on latent diffusion models, where generation is performed in a compact latent space and a decoder maps latents back to pixels. Yet the latent-to-pixel decoder in this pipeline remains primarily reconstruction-oriented: as the sole pathway from latents to pixels, it is optimized to invert an encoder rather than to synthesize high-resolution details, and becomes increasingly costly in megapixel pipelines. This gap calls for a more expressive and efficient decoding paradigm. Motivated by recent progress in scalable pixel-space diffusion, we introduce **PiD**, a **Pi**xel diffusion **D**ecoder that reformulates latent decoding as conditional pixel diffusion, unifying decoding and upsampling into one generative module. By denoising directly in high-resolution pixel space, PiD synthesizes $4\times$and even $8\times$ upscaled images with low latency. For latent conditioning, a lightweight sigma-aware adapter injects noise-corrupted latents into the backbone, enabling PiD to handle partially denoised states and terminate the base diffusion process early. To further improve efficiency, we distill the decoder using DMD2, reducing inference to just $4$ steps. PiD applies to both conventional VAE latents and semantic latents (e.g., SigLIP, DINOv2) used in recent RAE-based models. Across multiple latent spaces and base generators, PiD improves visual fidelity over conventional decode-then-upsample cascades while reducing memory and latency, decoding latents of $512{\times}512$ images into $2048{\times}2048$ pixels in under $1$ second with $13$ GB peak memory on a consumer RTX 5090, and as fast as $210$ ms on a data-center GPU, about $6{\times}$ faster than cascaded diffusion-based super-resolution pipelines.


PI-EDG: Physics-Informed Full-Space Electron Density Generation from Molecular Geometry

Jiahang Shen ⋅ Hongxin Xiang ⋅ Zhixiang Cheng ⋅ Keke Chen ⋅ Yingzhuo Tu ⋅ Wenjie Du ⋅ xiangxiang Zeng

Full-space electron density generation is fundamental to modeling molecular quantum states, ground-state properties, and chemical interactions. DFT provides a principled route to electron density, but high-fidelity calculations require costly self-consistent iterations, limiting large-scale quantum-chemical modeling. Deep learning offers a promising surrogate, yet existing methods often rely on indirect density-related prior, including coordinate queries, predefined basis functions, orbital coefficients, or spherical atomic density grids. These prior introduce reconstruction overhead and can hinder continuous full-space modeling across sharp nuclear cores and sparse vacuum regions. Here, we introduce PI-EDG, a physics-informed voxel-space framework that maps molecular geometries directly to full-space electron-density fields by combining quantum-mechanical prior injection, hybrid FNO-CNN modeling of local and non-local interactions and magnitude-aware DS-LMoE routing for extreme density ranges. On EDBench, PI-EDG achieves an MAE of 3.405$\times$10$^{-3}$, a total electron-number MAE of 0.472, and an average inference time of 133.4 ms for one molecule. Compared with graph-based, basis-based, and voxel-based baselines, PI-EDG demonstrates better predictive accuracy, improved robustness across extreme density regimes, and better physical consistency while maintaining efficient full-space generation. These results suggest that PI-EDG offers a new, direct, and efficient route for accelerating DFT-based electron density calculations.


Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation

Kexuan Shi ⋅ Hanxuan Li ⋅ Zeju Qiu ⋅ Yandong Wen ⋅ Simon Buchholz ⋅ Weiyang Liu

We introduce Pion, a spectrum-preserving optimizer for large language model (LLM) training based on orthogonal equivalence transformation. Unlike additive optimizers such as Adam and Muon, Pion updates each weight matrix through left and right orthogonal transformations, preserving its singular values throughout training. This yields an optimization mechanism that modulates the geometry of weight matrices while keeping their spectral norm fixed. We derive the Pion update rule, systematically examine its design choices, and analyze its convergence behavior along with several key properties. Empirical results show that Pion offers a stable and competitive alternative to standard optimizers for both LLM pretraining and finetuning.


PipeFSDP: Efficient Pipeline Parallel under Fully Sharded Data Parallel for Large Language Model Training

Xinglin Pan ⋅ Mingji Han ⋅ Penghao Zhao ⋅ Lin Zheng ⋅ Rayying ⋅ key ⋅ Shaohuai Shi ⋅ Xiaowen Chu

It is now a common practice to use PyTorch’s official Fully Sharded Data Parallel (FSDP) framework to train large language models (LLMs) on large-scale GPU clusters, typically together with pipeline parallelism (PP) to achieve better performance. However, existing popular systems like TorchTitan with FSDP and PP are inefficient due to the large bubbles and peak memory footprint caused by suboptimal task scheduling. In this paper, we propose PipeFSDP, an efficient training framework for FSDP combined with PP that alleviates pipeline bubbles and resource contention, thus improve training efficiency. Specifically, we design four fine-grained optimizations: dynamic context-aware prefetching, communication contention mitigation, redundant reshard elimination, and communication order optimization. In addition, we develop a heuristic model that selects the FSDP and PP degrees based on the given hardware and model configurations, thereby enhancing the practicality of PipeFSDP. Empirical evaluations on both dense and sparse LLMs (Llama3, Qwen3, and Qwen3-MoE) across 32-GPU and 64-GPU clusters demonstrate that PipeFSDP significantly outperforms the Torchtitan system by up to 1.28$\times$ in Model FLOPs Utilization (MFU) and 1.87$\times$ in memory consumption.


PIU-CR: Physics-Informed Deep Unfolding Network with SAR-Optical Image Fusion for Cloud Removal

Gengque Fan ⋅ Han Xu ⋅ Junxi Li ⋅ Hao Zhang ⋅ Guangcan Liu ⋅ Jiayi Ma

Multi-modal cloud removal exploits complementary optical and synthetic aperture radar (SAR) images to restore cloud-degraded remote sensing images. Most existing methods rely on black-box, data-driven architectures with limited interpretability of cross-modal interactions. Although some attempts introduce physical priors into networks, their oversimplified formulations fail to faithfully characterize the underlying physical imaging process and the networks remain loosely coupled with optimization variables. To address these issues, we propose a physics-informed deep unfolding network for cloud removal via optical-SAR image fusion, termed as PIU-CR. We first formulate cloud removal as a physics-informed optimization problem. Specifically, we fundamentally ground the optimization in the physical imaging mechanism of cloud degradation by embedding the atmospheric scattering model into data fidelity term. Then, the regularization terms establish an explicit cross-modal interaction mechanism, overcoming the single-modal information scarcity and the obscure, uninterpretable interactions in black-box models. They are dedicated to SAR-guided structure and spectral regularization, jointly introducing cross-modal structural feature alignment and coordinated spectral consistency constraints. The resulting optimization problem is then unfolded into a multi-stage neural network with dedicated structure and spectral branches. Each stage corresponds to explict optimization iteration step, enabling interpretable and effective reconstruction process. Experiments demonstrate that the proposed method outperforms state-of-the-art methods.


Plan2Sense: Open-World Task Planning in Epistemic States via Interleaved Ontic and Sensing Actions

Xiaotian Liu ⋅ Armin Toroghi ⋅ Jiazhou Liang ⋅ Ali Pesaranghader ⋅ Scott Sanner

Open-world task planning requires not only generating feasible plans but also acquiring the knowledge needed to execute them. Existing task planners for egocentric agents either attempt to exhaustively gather information from the environment or generate plans under prior assumptions of the world and replan as new observations arrive. Both strategies are impractical in complex environments and can potentially expose agents to undesirable or unsafe situations. To enable effective planning in open-world settings, we propose Plan2Sense, a framework that grounds planning in the agent’s epistemic state to determine what information is required for plan feasibility and how to acquire it. We formalize decision-making as a three-valued first-order Knowledge-State Markov Decision Process (KS-MDP) and develop a knowledge state goal regression algorithm KSR to construct sensing-aware conditional plans in the open-world setting. Our formulation supports sensing actions for both predicate disambiguation and knowledge acquisition via object binding without assuming a fully known initial state or a closed set of predefined objects. Across three open-world egocentric benchmarks requiring active knowledge acquisition, \textsc{Plan2Sense} achieves higher success rates and improved efficiency over existing baselines while consistently avoiding undesirable and potentially unsafe states.

We study policy optimization for online episodic tabular Markov decision processes with unknown transition kernels, aiming for best-of-both-worlds guarantees and multiple data-dependent regret bounds. Recent work (Dann et al., 2023; Li et al., 2026) has shown that policy optimization can adapt to both adversarial and stochastic losses with data-dependent bounds, including first-order, second-order, and path-length bounds, but only under known transitions. We resolve the open problem raised by Dann et al. (2023) by developing optimistic follow-the-regularized-leader algorithms that extend such guarantees to unknown transitions. The key ingredient is a new design of optimistic $Q$-function estimators together with a data-dependent transition bonus that controls estimator bias through the loss-prediction error. Our analysis further identifies an unavoidable transition-dependent complexity term that captures the intrinsic cost of estimating the transition kernel. As a result, we obtain first-order, second-order, and path-length bounds with this transition-dependent complexity term while simultaneously achieving gap-dependent $\mathrm{polylog}(T)$ regret in the stochastic regime.


Population-Aligned Persona Generation for LLM-based Social Simulation

Zhengyu Hu ⋅ Jianxun Lian ⋅ Zheyuan Xiao ⋅ Max Xiong ⋅ Tianfu Wang ⋅ Teng Xiao ⋅ Yuxuan Lei ⋅ Fengqing Jiang ⋅ Kaize Ding ⋅ Ziang Xiao ⋅ Nicholas Jing Yuan ⋅ Xing Xie ⋅ Radha Poovendran

Recent advances in large language models (LLMs) have enabled large-scale, high-fidelity social simulations, creating new opportunities for computational social science. However, constructing persona sets that faithfully reflect real-world population diversity remains a key challenge. Existing studies often emphasize agentic frameworks and simulation environments, while paying less attention to persona generation and the biases introduced by unrepresentative persona sets. In this paper, we propose a systematic framework for synthesizing high-quality, population-aligned persona sets for LLM-driven social simulation. Our approach begins by leveraging LLMs to generate narrative personas from long-term social media data, followed by rigorous quality assessment to filter out low-fidelity profiles. We then apply importance sampling to achieve global alignment with reference psychometric distributions, such as the Big Five personality traits. To address the needs of specific simulation contexts, we further introduce a task-specific module that adapts the globally aligned persona set to targeted subpopulations. Extensive experiments demonstrate that our method significantly reduces population-level bias and enables accurate, flexible social simulation for a wide range of research and policy applications. Code is available at https://anonymous.4open.science/r/K5CQ.


Pose6DAug: Physically Plausible Multi-View Object Swapping for Robot Data Augmentation

Jonghoon Lee ⋅ Seong Hyeon Park ⋅ Minha Lee ⋅ Byungwoo Jeon ⋅ Jinwoo Shin

Vision-language-action (VLA) policies have shown strong potential for general-purpose manipulation, yet they often fail on novel, out-of-distribution objects whose appearance or geometry deviates from the training distribution. The standard remedy is to collect multi-view teleoperation data for every failure case, but this scales poorly in both cost and time. We introduce Pose6DAug, a failure-driven data augmentation framework that turns a policy’s own successful episodes into targeted demonstrations for its failure modes, without any new data collection. Our key insight is that each successful episode already encodes a physically valid action trajectory together with calibrated multi-view observations. By swapping only the manipulated object while preserving this trajectory, we obtain new and physically grounded demonstrations. However, the challenge lies in synthesizing object swaps, e.g., naive 2D video editing breaks multi-view consistency and physical plausibility, particularly under heavy occlusion and egocentric viewpoints. Our method instead operates directly in 3D, anchoring the target object with an explicit mesh driven by a temporally coherent 6D pose trajectory. This ensures geometrically consistent renderings across all camera views. Fine-tuning a VLA on data augmented by our method improves success rates by 16.5% relative to the state-of-the-art baseline on novel objects, while preserving in-distribution performance. These results show that multi-view and physically consistent augmentation is a practical path to scalable VLA generalization.

3D scene inpainting aims to recover missing or occluded regions in edited 3D scenes, while ensuring geometric and textural consistency. Existing approaches, however, typically require accurately calibrated camera poses, which restricts their applicability in casual, in-the-wild scenarios and introduces additional preprocessing overhead. To overcome this limitation, we present FreeInpaint, a novel feed-forward framework that generates complete and 3D-consistent scenes directly from unposed multi-view images with masked regions. At its core, FreeInpaint extends a 3D foundation model to propagate masked regions from a reference view to other unposed views, bridging 3D reconstruction and scene inpainting while preserving the model's native ability to recover camera poses and scene geometry. Our method addresses two key challenges in adapting feed-forward 3D foundation models to masked inputs. First, masked regions can corrupt cross-view correspondence reasoning, degrading pose estimation and geometry recovery. To address this, we introduce a Learnable Mask Attention mechanism that preserves the spatial anchoring of reliable observations while allowing masked regions to progressively absorb useful context in deeper layers. Second, under severe occlusions, a single forward pass often lacks sufficient appearance evidence for high-fidelity completion. Therefore, we propose a Support Token Refinement strategy, which injects diffusion-generated support evidence as confidence-weighted auxiliary tokens to refine under-observed regions while preserving the original spatial anchor. Extensive experiments across diverse datasets demonstrate that FreeInpaint achieves superior inpainting quality, eliminating the reliance on pre-computed camera poses while keeping a fast inference speed.

Transformer models have demonstrated a remarkable ability to perform a wide range of tasks through in-context learning (ICL), where the model infers patterns from a small number of example prompts provided during inference. However, empirical studies have shown that the effectiveness of ICL can be significantly influenced by the order in which these prompts are presented. Despite its significance, this phenomenon has been largely unexplored from a theoretical perspective. In this paper, we theoretically investigate how positional encoding (PE) affects the ICL capabilities of linear attention transformer models, particularly in tasks where prompt order plays a crucial role. We examine two distinct cases: linear regression, which represents an order-invariant task, and dynamical systems, a classic time-series task that is inherently sensitive to the order of input prompts. Theoretically, we evaluated the change in the model output when two types of positional encoding (one-hot and RoPE) is incorporated and the prompt order is altered. In all cases, the leading dependence on the permutation size and context length is $k/N$, with constants determined by the positional encoding, model weights, input dimension, and task dependent parameters. These theoretical findings are experimentally validated.


Position: Life-Logging Video Streams Make the Privacy–Utility Trade-off Inevitable

Tianyuan Zou ⋅ Liang Yue ⋅ Yang Liu ⋅ Ya-Qin Zhang ⋅ Sijie Cheng

With the growing prevalence of always-on hardware such as smart glasses, body cameras, and home security systems, life-logging visual sensing is becoming inevitable, forming the backbone of persistent, always-on AI systems. Meanwhile, recent advances in proactive agents and world models signal a fundamental shift from episodic, prompt-driven tools to next-generation AI systems that continuously perceive and react to the physical world. Although life-logging video streams can substantially improve utility of these promising systems, they also introduce significant privacy risks by revealing sensitive information, such as behavioral patterns, emotional states, and social interactions, beyond what isolated images expose. If unresolved, these risks may undermine public trust and hinder the sustainable development of always-on AI technologies. Existing privacy protections are either attack-specific or incur substantial utility loss, and fail to consider the entire data exploitation pipeline. We therefore posit that the privacy-utility trade-off in life-logging video streams is a foundational challenge for next-generation AI systems that demands further investigation. We call for novel pipeline-aware privacy-preserving designs that jointly optimize utility and privacy for long-horizon life-logging visual data. In parallel, formal privacy leakage metrics and standardized benchmarks remain important open directions for future research.


PosterDuet: Co-Evolving Design Generation and Reward Optimization for Product Poster Synthesis

junlong wu ⋅ Pengcheng Wei ⋅ Jia Sun ⋅ Huaiqing Wang ⋅ Weixuan Zeng ⋅ Zijun Li ⋅ Honglie Wang ⋅ Yongrui Heng ⋅ Boheng Zhang ⋅ Dewen Fan ⋅ Fan Yang ⋅ Tingting Gao ⋅ Houde Liu ⋅ Qianqian Gan

Automatic e-commerce poster generation requires the joint optimization of product fidelity, visual composition, text rendering, and commercial appeal. Existing methods typically address this task through staged pipelines involving product segmentation, layout prediction, glyph rasterization, and conditional image synthesis. While effective in constrained settings, such decompositions rely on rigid preprocessing assumptions and weaken the mutual adaptation among product appearance, typography, and scene composition, especially for hand-held, worn, or context-dependent products. We reformulate this problem as \emph{holistic product-aware poster editing}: given a raw product image and structured metadata, the goal is to generate a complete poster directly, without segmentation masks, glyph control maps, or predefined layout boxes. To this end, we propose PosterDuet, a closed-loop framework that integrates a vision-language model (VLM), an image editing model, and reward-based optimization. The VLM first generates a holistic design prompt that specifies background style, layout arrangement, promotional copy, and typographic intent from the raw image and metadata. Conditioned on both the prompt and the original image, the image editor synthesizes the final poster in a unified editing process. To optimize both commercial effectiveness and visual quality, we introduce a mixed-reward learning framework that combines a CTR-oriented reward with a generative holistic quality reward, and use GRPO to optimize the prompt-generating VLM. We further improve textual accuracy via OCR-reward-based reinforcement tuning of the image editor, and exploit natural-language critiques from the generative reward model for iterative prompt refinement. Extensive experiments show that PosterDuet generates more coherent, text-faithful, and commercially effective posters than prior pipelined approaches.

We study fixed-confidence best-arm identification in Bernoulli bandits. A learner sequentially samples from $K$ arms and aims to identify the unique best arm while keeping the error probability at most $\delta$. We ask whether the classical lower bound on expected sample complexity can be matched asymptotically without ongoing forced exploration. We propose a posterior-tracking algorithm that first samples each arm once, then repeatedly draws a bandit instance from the posterior, computes the oracle allocation for that instance, and randomly selects the next arm according to that allocation. Stopping is based on a Bernoulli generalized likelihood ratio test. Since the algorithm has no explicit forced-exploration mechanism after initialization, the main technical challenge is to show that posterior randomness alone sustains sufficient exploration. We prove that, with summable tail probabilities, every arm is sampled at least logarithmically often. This yields an integrable stabilization time after which the true best arm remains the empirical leader and the corresponding generalized likelihood ratio statistic grows linearly at the information-theoretic lower-bound rate. Consequently, the proposed procedure is $\delta$-correct and achieves asymptotically optimal expected sample complexity as $\delta \to 0$.


Post-Selection Distributional Model Evaluation

Amirmohammad Farzaneh ⋅ Osvaldo Simeone

Formal model evaluation methods typically certify that a model satisfies a prescribed target key performance indicator (KPI) level. However, in many applications, the relevant target KPI level may not be known a priori, and the user may instead wish to compare candidate models by analyzing the full trade-offs between performance and reliability achievable at test time by the models. This task, requiring the reliable estimate of the test-time KPI distributions, is made more complicated by the fact that the same data must often be used both to pre-select a subset of candidate models and to estimate their KPI distributions, causing a potential post-selection bias. In this work, we introduce post-selection distributional model evaluation (PS-DME), a general framework for statistically valid distributional model assessment after arbitrary data-dependent model pre-selection. Building on e-values, PS-DME controls post-selection false coverage rate (FCR) for the distributional KPI estimates and we establish explicit conditions under which it is provably more sample efficient than a baseline method based on sample splitting. Experiments on synthetic data, text-to-SQL decoding with large language models, and telecom network performance evaluation demonstrate that PS-DME enables reliable comparison of candidate configurations across a range of reliability levels, supporting the statistically reliable exploration of performance--reliability trade-offs.


Practical Adversarial Attacks on Stochastic Bandits via Fake Data Injection

Qirun Zeng ⋅ Eric He ⋅ Richard Hoffmann ⋅ Xuchuang Wang ⋅ Jinhang Zuo

Adversarial attacks on stochastic bandits have traditionally relied on some unrealistic assumptions, such as per-round reward manipulation and unbounded perturbations, limiting their relevance to real-world systems. We propose a more practical threat model, Fake Data Injection, which reflects realistic adversarial constraints: the attacker can inject only a limited number of bounded fake feedback samples into the learner's history, simulating legitimate interactions. We design effective attack strategies under this model, explicitly addressing both magnitude constraints (on reward values) and temporal constraints (on when and how often data can be injected). Our theoretical analysis shows that these attacks can mislead a class of bandit algorithms into selecting a target arm in nearly all rounds while incurring only sublinear attack cost. Experiments on synthetic and real-world datasets validate the effectiveness of our strategies, revealing vulnerabilities in stochastic bandit algorithms under practical adversarial scenarios.


Practical Estimation of the Bayes Optimal Fairness-Accuracy Tradeoff with Soft Labels

mohit sharma ⋅ Okan Koc ⋅ Amit Jayant Deshpande ⋅ Takashi Ishida ⋅ Gang Niu ⋅ Rajiv Ratn Shah ⋅ Masashi Sugiyama

There is a fundamental limit to the best prediction error of any model for a given data distribution. In classification, the Bayes error quantifies this limit and serves as both a benchmark for trained models and a criterion to detect overfitting. The state-of-the-art Bayes error estimation techniques for binary classification use only soft labels (positive class probabilities) and require no input features or auxiliary training. We extend such notions to constrained classification problems, such as fairness. Fair classification is an important case of constrained or multi-objective classification. It requires a model not only to be accurate but also to yield fair outcomes (equal or near-equal true positive rates) across different demographic groups (e.g., race and gender). This leads to an inherent fairness-accuracy tradeoff or a Pareto frontier that is used to compare fair classifiers against each other. The focus of our work is to estimate the fundamental limit for this fairness-accuracy tradeoff. We derive an explicit expression for the fair Bayes error rate (i.e., best achievable error rate subject to given fairness constraints), and construct consistent, instance-free estimators for it using only soft labels and group information. We provably bound the bias and the sample complexity of our estimators. Empirically, our method produces a robust, low-variance estimate of the optimal fairness-accuracy trade-off curve, even when the soft labels are noisy and derived from a classification model on hard labels.


Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation

Jihwan Hong ⋅ Woohyeon Park ⋅ Jaeik Kim ⋅ Jaeyoung Do

Video Object Segmentation (VOS) in complex and long videos is increasingly important for real-world applications, where target objects often appear only intermittently within long temporal horizons. However, existing benchmarks largely focus on temporally salient objects that remain visible for most of the video. To address this gap, we introduce FaVOS (A Benchmark for Video Object Segmentation with Fractional Temporal Visibility), a benchmark designed to evaluate VOS methods under low temporal visibility. We show that, in this regime, the standard $\mathcal{J}$\&$\mathcal{F}$ metric can collapse VOS evaluation into absence classification, because empty predictions receive high rewards on target-absent frames. Consequently, even a trivial empty-mask predictor can outperform strong models such as SAM 3, revealing a fundamental mismatch between current metrics and practical VOS performance. To mitigate this issue, we propose Volumetric $\mathcal{J}$\&$\mathcal{F}$, which evaluates mask sequences as spatio-temporal volumes and reduces the dominance of target-absence rewards while preserving sensitivity to segmentation quality and temporal structure.


Predicting Species Splits: A Challenging Fine-Grained Benchmark for Category Discovery

Christian Lange ⋅ Joakim Bruslund Haurum ⋅ Bingchen Zhao ⋅ Oisin Mac Aodha

Widely used category discovery benchmarks have become increasingly saturated and often fail to reflect real-world discovery settings, making long-term evaluation of new methods difficult. We introduce \textbf{PS-SPLIT}, a large-scale fine-grained benchmark for category discovery, motivated by taxonomic splitting: the real-world process by which what was once considered a single species is found to comprise multiple distinct ones. PS-SPLIT contains 79K globally distributed, high-quality, community-collected bird images spanning 771 species, derived from 257 recent real taxonomic splits, with associated geographic metadata and full taxonomic hierarchy. With categories arising from real taxonomic splits, visual differences can be very subtle, often requiring significant human expertise to tell apart, making PS-SPLIT a challenging benchmark which we hope will assist long term in fine-grained category discovery research. Based on taxonomic splitting, we introduce the task of \textbf{discovery by category splitting}, where the goal is to determine which known categories should be subdivided into finer-grained subcategories. We adapt four generalized category discovery methods to this setting and show that there is substantial room for future progress.


Predictive but Not Plannable: RC-aux for Latent World Models

Wenyuan Li ⋅ Guang Li ⋅ Keisuke Maeda ⋅ Takahiro Ogawa ⋅ Miki Haseyama

A latent world model may achieve accurate short-horizon prediction while still inducing a latent space that is poorly aligned with planning. A key issue is spatiotemporal mismatch: these models are often trained with local predictive supervision, but deployed for long-horizon goal-directed search in latent spaces where Euclidean distance may not reflect what is reachable within a finite action budget. We present the Reachability-Correction auxiliary objective (RC-aux), a lightweight correction for this mismatch in reconstruction-free latent world models. RC-aux keeps the world-model backbone unchanged and adds planning-aligned supervision along two axes. Along the time axis, multi-horizon open-loop prediction trains the model beyond one-step consistency. Along the space axis, budget-conditioned reachability supervision, together with temporal hard negatives, encourages the latent space to distinguish states that are eventually reachable from those reachable within the current planning horizon. At test time, the learned reachability signal can also be used by a reachability-aware planner to favor trajectories that are both goal-directed and attainable under the available budget. We instantiate RC-aux on LeWorldModel and evaluate it under both continuation-training and matched-from-scratch settings. Across goal-conditioned pixel-control tasks and a LIBERO-Goal extension, RC-aux improves LeWM-style planning with modest additional cost. These results suggest that planning with latent world models depends not only on predictive accuracy, but also on whether the learned representation encodes the temporal and geometric structure required by downstream search. The source code will be released upon acceptance of the paper.


Predictive Representation Learning for Partially Observed Neural Dynamics

Xinyi Li ⋅ Zhichao Liang ⋅ Hongjun Jiang ⋅ Yian Zhu ⋅ Guanyi Zhao ⋅ Xuejian Yang ⋅ Xiaoqi Chen ⋅ Kexin Lou ⋅ Quanying Liu

In long-term multi-session neural recordings, the same latent population dynamics are often observed through changing subsets of neural channels due to electrode drift, signal instability, or session-dependent recording variability. This creates a partially observed system-identification problem: each session provides only an incomplete view of a shared dynamical system, and missing channels are not merely absent inputs but unobserved variables coupled to the same latent dynamics. Existing approaches often rely on observation-pattern-specific mapping, which treats missing channels passively and leaves the shared latent dynamics weakly constrained. We propose CHAMA (CHannel-Aware Masked Attention), a mask-conditioned predictive representation learning framework for identifying shared neural dynamics under heterogeneous channel availability. CHAMA introduces a single global observation interface conditioned on channel-availability masks, allowing observed and unobserved channels to jointly constrain latent dynamics. The key idea is bidirectional: missing channels impose structural constraints on the latent representation, while the learned latent dynamics enable principled recovery of missing activity. To improve identifiability under sparse observations, CHAMA aggregates causal temporal windows, effectively trading temporal context for missing spatial information. Across synthetic dynamical systems and real multi-session neural recordings, CHAMA improves missing-channel completion, forecasting, latent dynamical fidelity, and downstream behavioral decoding over conditional adapter, zero-padding, and masked-autoencoding baselines. These results show that dynamics-aware recovery preserves globally consistent and behaviorally relevant structure beyond point-wise reconstruction accuracy. The source code is available at: \url{https://anonymous.4open.science/r/CHAMA-D691}.


Preventing Error Cascades in Long-Horizon Multimodal Agents with Edge-Reliability Graph Memory

Saman Forouzandeh ⋅ Wei Peng ⋅ Xinghuo Yu ⋅ Mahdi Jalili

Long-horizon tool-augmented agents suffer sharp degradation as trajectories grow: small tool errors are stored, repeatedly reused, and amplified into cascading failures. Existing approaches rely on post-hoc verification, item-level memory scoring, or executive context management, but do not directly control how unreliable evidence propagates through reuse. We propose \textbf{EPOCH} (Edge-Pathway Outcome-supervised Cascade-Halting memory), which mitigates cascades by learning an outcome-correlated reliability proxy on reuse pathways connecting observations over time, instantiated as a sparse memory graph with quality-weighted edges, a Temporal Graph Transformer over edge reliabilities, and a learned mid-context correction. \textbf{We prove that quality-weighted and uniform retrieval are asymptotically separated in horizon length:} the ratio of cumulative cascade error diverges unboundedly as $T \to \infty$ whenever the calibration parameter is positive, a result that does not require distributional regularity. We validate this prediction empirically at trajectory lengths up to $T=500$ and additionally establish a quantitative edge-vs-node separation in terms of an empirically measurable cross-context error variance. Because outcome supervision is correlational by construction, we do not claim the learned scores recover causal pathway reliability; we validate them on two complementary axes (counterfactual edge ablation: $\rho=0.71$, $n=21{,}743$; human annotation: ROC-AUC~$=0.94$, ECE~$=0.041$, $n=6{,}000$). Across 11 multi-modal benchmarks spanning native long-horizon search, wrapped knowledge-seeking tasks, and clean perception controls, EPOCH consistently outperforms the strongest reliability-aware and memory-augmented baselines in noisy long-horizon regimes while remaining on par on clean perception controls---quality-aware filtering does not collapse on clean data without claiming it helps there.


PriFT: Prior-Support Guided Token Reweighting for Supervised Fine-Tuning

Ke Wang ⋅ Shuangqi Li ⋅ Mathieu Salzmann ⋅ Pascal Frossard

Supervised fine-tuning (SFT) is computationally efficient and broadly applicable, but often shows weaker generalization than reinforcement learning (RL). A key limitation is its off-policy objective: SFT fits fixed demonstrations token by token, including targets that may be poorly aligned with the model's pretrained distribution. Recent token-reweighted SFT methods address this issue by assigning larger training weights to tokens that better align with the model's predictive distribution, using statistics such as target-token probability or entropy. However, computing these statistics from the online model being fine-tuned makes token weights trajectory-dependent, as the model's distribution rapidly departs from the pretrained model and induces self-reinforcing reweighting dynamics. We propose PriFT, Prior-support guided Fine-Tuning, which derives token weights from a frozen pretrained reference to obtain a stable reweighting signal unaffected by online fine-tuning dynamics. This signal estimates prior support: the extent to which each target token is supported by the pretrained model before task-specific adaptation. Across multiple existing weighting and selection rules, replacing online statistics with pretrained statistics consistently improves performance. We introduce two instantiations: PriFT-prob, which uses pretrained target-token probability, and PriFT-mass, which selects tokens by relative support under the pretrained distribution. Extensive experiments on mathematical reasoning, code generation, and medical question answering show that PriFT achieves state-of-the-art SFT results and provides a better initialization for subsequent RL training.


PrismFlow: Residual Dynamics for Flow Matching in Time-Series Generation

ZHANG JUNRU ⋅ Lang Feng ⋅ Jinbo Wang ⋅ Xu Guo ⋅ Yucheng Wang ⋅ Han Yu ⋅ Min Wu ⋅ Yabo Dong ⋅ Duanqing Xu

Generating high-quality time-series data is challenging because real-world signals often exhibit multimodal patterns and multiscale dynamics, including oscillations and high-frequency variations. Flow Matching (FM) offers an efficient alternative to diffusion models, but practical implementations typically rely on a single finite-capacity global vector-field estimator. In such heterogeneous temporal distributions, distinct regimes may pass through nearby flow states while requiring incompatible conditional velocities. A monolithic estimator trained with the standard $\ell_2$ velocity-matching objective may therefore learn an overly smoothed approximation of the local transport field. This estimator-level smoothing can attenuate branch-specific dynamics, leading to spectral distortion and poor mode coverage. To address this, we propose PrismFlow, a new FM method with Koopman-inspired dynamical experts. Each expert learns residual corrections in a latent space where local nonlinear temporal evolution can be approximated by linear transitions. We further propose a confidence-aware Winner-Take-All (WTA) objective that updates only the expert best aligned with each sample while masking gradients to the others, encouraging mode-specific specialization. During sampling, the selected expert adds a residual dynamical correction to the global transport field, preserving FM stability while recovering fine-grained and high-frequency temporal structures. Across various benchmarks, PrismFlow effectively mitigates the spectral contraction in standard FM and achieves state-of-the-art performance, with a 15.6\% gain in Context-FID and a 38.6\% improvement in Discriminative Score, while remaining robust in low-data settings and effective for forecasting and imputation.


Probability-Conserving Flow Guidance

Parsa Esmati ⋅ Junha Hyung ⋅ Amirhossein Dadashzadeh ⋅ Jaegul Choo ⋅ Majid Mirmehdi

Diffusion and flow-based generative models dominate visual synthesis, with guidance aligning samples to user input and improving perceptual quality. However, Classifier-Free Guidance (CFG) and extrapolation-based methods are heuristic linear combinations of velocities/scores that ignore the generative manifold geometry, breaking probability conservation and driving samples off the learned manifold under strong guidance. We analyse guidance through the continuity equation and show its effect decomposes into a divergence term and a score-parallel term defined invariantly across parameterisations. We prove the divergence term blows up structurally as sampling approaches the data manifold, motivating a time-dependent schedule alongside score-parallel attenuation. The resulting plug-and-play rule, Adaptive Manifold Guidance (AdaMaG), bounds both terms at no additional inference cost. Finally, we show that most empirical heuristics for reducing saturation or improving generation quality correspond directly to the two terms in our decomposition. Across image generation benchmarks, AdaMaG improves realism, reduces hallucinations, and induces controlled desaturation in high-guidance regimes.


PROBE: Learning to Audit Policy Compliance in Tool-Using LLM Agents

Kshitij Mishra ⋅ Abhijith Sharma ⋅ Nils Lukas ⋅ Salem Lahlou

Modern agentic systems extend language models with external tools, enabling them to complete complex tasks through multi-turn interactions. As these agents are deployed in consequential settings, their behavior is often constrained by policies governing pricing, privacy, confirmation, safety, and other operational requirements. However, existing evaluation methods leave a critical gap: capability benchmarks measure what agents can accomplish, while content-safety red-teaming measures harmful generations, but neither tests whether agents comply with predefined policies during otherwise valid tool use. Such policy violations are difficult to audit because their effects may emerge only after several turns, span multiple tool calls, or surface in downstream systems. We propose PROBE (Policy Rule Observation via Behavioral Exploration), an algorithm for training LLM-based auditors that interact with deployed target agents as realistic users and probe them for policy violations through multi-turn, tool-mediated dialogue. We formalize agent auditing as a two-player Markov game and propose a composite reward that balances violation discovery, behavioral plausibility, conversation progress, and policy diversity. To study the role of training objectives and reward design, we compare GRPO, DPO, and GFlowNet-based optimization in a controlled setting. Across extensive experiments, auditors trained with PROBE uncover substantially broader and higher-yield policy violations than same-architecture zero-shot adversaries, even when the auditor is significantly smaller than the target agent. Our results show that multi-turn behavioral auditing is a distinct and necessary axis of agent evaluation, complementing capability and content-safety benchmarks. More broadly, PROBE offers a scalable path toward automated policy auditing for real-world agentic systems and suggests a future co-evolutionary loop in which auditors expose failures and agents improve from auditor feedback.


ProbMedTOD: A Bayesian Network Guided Task-Oriented Dialogue System for Patient History Taking

Vishal Vivek Saley ⋅ Bhavesh Gurnani ⋅ Dinesh Raghu ⋅ Mausam

Patient history taking is a strategic, multi-turn process that refines diagnostic hypotheses through iterative questioning, mirroring clinical reasoning where each inquiry updates beliefs over candidate diagnoses. Deploying AI in this setting is constrained by privacy, cost, and latency, making Small Language Models (SLMs) a more desirable backbone compared to larger LLMs. However, SLMs capture only the fast, intuitive "System 1" side of this process and lack explicit mechanisms for the deliberative "System 2'' reasoning needed to maintain and update diagnostic uncertainty over time. Bayesian Networks (BayesNets) offer a natural complement, providing an interpretable framework for auditable probabilistic belief tracking. Yet, clinical BayesNets are difficult to construct due to reliance on expert-curated structure, dependencies, and parameters. We introduce ProbMedTOD, that automatically synthesizes a clinical BayesNet from a small number of clinical notes in three stages: ontology building, structure prediction, and parameter estimation. ProbMedTOD then integrates the resulting BayesNet with SLMs enabling probabilistic reasoning in multi-turn dialogues. A BayesNet Agent maintains real-time posterior distributions over diagnoses, while a Policy Agent picks the diagnostically informative questions. Experiments on DDXPlus and MIMIC-IV Notes show that ProbMedTOD outperforms LLM-only and retrieval-based baselines, with especially strong gains for smaller models.


Process-conditioned Pretraining with Topographic Spatial Retrieval for Large EEG Models

Yi Ding ⋅ Muyun Jiang ⋅ Weibang Jiang ⋅ Shuailei Zhang ⋅ XINLIANG ZHOU ⋅ Chenyu Liu ⋅ Shanglin Li ⋅ Yong Li ⋅ Cuntai Guan

Large EEG foundation models are pretrained on heterogeneous corpora, but two challenges remain underexplored: (1) experimental protocols provide mental-process cues that are rarely used in masked pretraining, and (2) electrode montages vary across datasets, making spatial representations montage-dependent. Existing methods typically learn a single general-purpose representation and model spatial structure implicitly, limiting their ability to exploit process-related information under montage heterogeneity. We introduce BrainPro, a process-conditioned self-supervised pretraining framework that couples topographic spatial retrieval with shared and process-associated representation learning. BrainPro maps dataset-specific montages to a universal channel-region template and retrieves channel- and region-level spatial filters to form a topographically aligned spatial basis. Over this basis, BrainPro learns a shared encoder for general EEG structure and additional affective, motor-related, and auxiliary residual branches for process-associated variation. Protocol-derived mental-process cues condition branch activation and region-weighted masked reconstruction, incorporating topographic spatial priors and process information without using downstream class labels. Across nine public BCI benchmarks, BrainPro achieves strong performance among evaluated baselines. Ablations, channel-drop analysis, encoder-configuration studies, and spatial-filter visualizations suggest that topographic spatial retrieval and process-conditioned representation learning jointly improve EEG decoding.


Program-as-Weights: A Programming Paradigm for Fuzzy Functions

Wentao Zhang ⋅ Liliana Hotsko ⋅ Woojeong Kim ⋅ Pengyu Nie ⋅ Stuart Shieber ⋅ Yuntian Deng

Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON, or ranking search results by intent, and are increasingly outsourced to large language model APIs at the cost of locality, reproducibility, and price. We propose fuzzy-function programming: compiling such a function from a natural-language specification into a compact, locally-executable neural artifact. We instantiate this paradigm with Program-as-Weights (PAW), in which a 4B compiler trained on FuzzyBench, a 10M-example dataset we release, emits parameter-efficient adapters for a frozen, lightweight interpreter. A 0.6B Qwen3 interpreter executing PAW programs matches the performance of direct prompting of Qwen3-32B, while using roughly one fiftieth of the inference memory and running at 30 tokens/s on a MacBook M3. PAW reframes the foundation model from a per-input problem solver into a tool builder: invoked once per function definition, it produces a small reusable artifact whose subsequent calls per function application are cheap and offline.


ProMo3D: Probing Motion Cues in Frozen 3D Foundation Models

Xiaoang Zhang ⋅ Mert Kiray ⋅ Yordanka Velikova ⋅ Manoj Biswanath ⋅ Benjamin Busam

Feed-forward 3D foundation models have recently reshaped multi-view reconstruction, yet dynamic 4D understanding still often relies on large-scale motion supervision or hand-crafted, architecture-specific heuristics. We ask whether frozen 3D foundation model tokens already encode motion-relevant information that can be exposed by lightweight readout functions, rather than re-learning motion-specific features with high-capacity decoders. Through probing on three representative backbones, we find that static--dynamic separation is decodable from frozen token representations. Our rank-restricted and spectral analyses on $\pi^3$X reveal that the recovered motion cues are concentrated in a low-dimensional subspace. Building on this, we propose Probing-based Motion Distillation (PMD), which converts pairwise motion-state compatibility scores against a fixed anchor token into dense motion masks. Trained on only 160 YouTube-VOS videos using a single 16 GB GPU, PMD transfers zero-shot to DAVIS, SegTrackv2, and FBMS-59 without target-benchmark fine-tuning. With multi-view pretrained geometric backbones, PMD outperforms training-free heuristics and approaches substantially heavier supervised and self-supervised methods, suggesting that motion decoding from frozen 3D foundation models can be lightweight and data-efficient.


Proportionality in Ranking Compression

Daniel Halpern ⋅ Xinyu Liu ⋅ Evi Micha ⋅ Yu Peng Ng ⋅ Warut Suksompong

Motivated by applications in reinforcement learning from human feedback, we consider a setting where a given collection of rankings needs to be compressed into a smaller collection of rankings. This can also be seen as a combination of two fundamental social choice frameworks: ranking aggregation and multiwinner voting. We propose three proportionality notions that capture different types of representation in this setting: positional proportionality, pairwise proportionality, and proportionality for solid coalitions (PSC). On the one hand, we show that positional proportionality and PSC are always satisfiable for any input rankings and target compression size, and a desired output can be found in polynomial time. On the other hand, we prove that pairwise proportionality cannot be satisfied in general, but can nevertheless be attained when the input rankings are single-peaked or single-crossing.


ProQuant: Progressive Quantization-aware Training for Edge MLLMs

Yufei Xue ⋅ Yushi Huang ⋅ Jiawei Shao ⋅ Pingcheng Dong ⋅ Yonghao Tan ⋅ Shiyao Li ⋅ Kwang-Ting Cheng ⋅ Jun Zhang

Multi-modal large language models (MLLMs) with impressive perception and understanding capabilities will enable a range of downstream applications. However, their substantial computational and memory requirements hinder the real-world implementation, particularly for resources-constrained edge devices. Quantization has emerged as an effective solution for deploying large models. While post-training quantization (PTQ) successfully retrains the performance at moderate bit-width (\textit{e.g.}, W8A8), it suffers from severe performance degradation under low-bit quantization (\textit{e.g.}, W4A8/W8A8), especially for small-scale models that are more suitable for edge deployment. In this paper, we propose \texttt{ProQuant}, a novel weight-activation quantization-aware training (QAT) framework tailored for low-bit MLLMs. First, to align the limited finetuning dataset distribution with the pretrained dataset distribution, we propose a self-relabeling (SRL) scheme to resample the responses of the open-source multimodal dataset. Second, we are the \textit{first} to take a decoupled view of multi-modal (MM) understanding and language modeling abilities. We propose a progressive block activation (PBA) mechanism to prioritize the MM understanding recovery. We also introduce a vision-sensitive factor $\vec{\tau}$ that effectively identifies the MM-critical layers. Extensive experiments across various benchmarks, covering models with 2B$\sim$8B parameters, show that \texttt{ProQuant} significantly outperforms existing methods. For example, our W4A8 \texttt{Qwen3-VL-2B-Instruct} improves average accuracy by impressively 6.6\% and achieves performance comparable to full precision counterparts with only $\sim$1\% performance drop. Code will be released upon acceptance.

Performing accurate protein-glycan docking is critical to understanding protein-glycan interactions which govern many biological processes. However, this important task is hampered by the lack of tailored datasets, benchmarks and models for protein-glycan docking. To fill this blank, we first curate a standard benchmark ProtGlycanDock for training and evaluating protein-glycan docking models under both protein and glycan generalization settings. On this benchmark, we evaluate eight models in two main categories, physics-based methods and machine-learning-based molecular structure prediction models. To obtain more precise binding poses, we further propose two finetuning techniques to inject the knowledge of glycan conformations into a general-purpose molecular structure prediction model. These two techniques respectively augment model inputs with glycan topological features and refine glycan binding pose with a glycan energy loss. Benchmark results show the superiority of AlphaFold 3 among existing methods, and the two proposed finetuning techniques can be well combined to achieve a state-of-the-art model for protein-glycan docking. Also, we perform stratified benchmark analysis and case studies to better understand the strengths and weaknesses of current models, providing insights for future model improvements.

Benchmark data contamination has become a central challenge in LLM evaluation: when evaluation examples appear in the training data of one or more audited models, reported performance can be inflated and cross-model comparisons become unreliable. A broad line of training-data detection work designs scores to quantify how strongly a model memorizes a given data point, but these score-based methods lack theoretical guarantees. Recent conformal approaches provide provable false-identification control for a single model; however, applying them separately to each model can produce model-specific benchmarks, undermining fair comparison across models. In this work, we formalize multi-model benchmark decontamination as a joint selection problem and propose Joint Envelope Conformal Selection (JECS), a conformal procedure that enables global contamination rate (GCR) control under stated assumptions. Specifically, JECS computes per-model conformal p-values, aggregates them by the per-item maximum, and reconstructs a conservative envelope of the max-p null distribution from right-tail observations above a data-driven threshold. By applying the adaptive Benjamini-Hochberg (BH) procedure to the envelope-rescaled values, we select a benchmark with provable GCR control. Extensive experiments across various models and benchmarks demonstrate that JECS achieves higher power than the max-p baseline while consistently maintaining the target GCR control.


Provably Accurate Shapley Value Estimation in High Dimensions via Sparse Leverage Sampling

Ricardo Parada ⋅ Akbarkhuja Anvarkhujaev ⋅ Carlos Martin ⋅ William Chang ⋅ Long Tran-Thanh ⋅ Nam P Tran

Computing Shapley values exactly requires exponentially many evaluations of a cooperative value function, making approximation essential in high-dimensional applications. Recent leverage-score-based methods yield provably accurate estimators with sample complexity scaling linearly in the number of players, but this can still be too costly when only a small subset of players is truly influential. We study Shapley value estimation under a sparse influence model, where the Shapley vector is supported on an unknown set of size $s \ll n$. Using the regression characterization of Shapley values, we propose a sparse leverage sampling scheme and analyze it via sparse matrix concentration. Our main result shows that the resulting estimator achieves relative-error approximation with $\widetilde{O}(s^2)$ value function evaluations, up to logarithmic and accuracy factors, thereby replacing the ambient-dimension dependence of prior guarantees with dependence on the intrinsic support size. Empirically, our method matches or improves over dense Leverage SHAP on sparse synthetic and real-data benchmarks, while exhibiting better numerical stability at small sample sizes.

Many physical data assimilation (DA) workflows require smoothing methods that represent non-Gaussian posteriors over physical state variables, scale to high-dimensional simulators, train from observation windows alone, and remain compatible with calibration of the prescribed simulator. We introduce PR-Smoother, a simulator-preserving amortized smoother designed for this prescribed-simulator DA regime. Its key design principle is to keep the prescribed simulator explicit in both the evidence lower bound and the variational family: rather than learning replacement dynamics or a learned trajectory prior, PR-Smoother learns only future-conditioned corrections around the prescribed rollout. This yields an explicit non-Gaussian smoothing distribution over physical trajectories and supports joint state, parameter, and sensor-bias learning from observations alone. The variational family contains the exact smoother in deterministic and linear-Gaussian limits. Empirically, PR-Smoother captures multimodal posteriors in 4-dimensional Lorenz–96, remains accurate under ambiguous nonlinear observations and process noise in 40-dimensional Lorenz–96, and scales to joint state-parameter-bias inference in 16,384-dimensional Kolmogorov flow.


Pushing Biomolecular Utility-Diversity Frontiers with Supergroup Relative Policy Optimization

Xinwu Ye ⋅ He CAO ⋅ Li Hao ⋅ Bin Feng ⋅ Zijing Liu ⋅ Robert Tang ⋅ Yu Li ⋅ Shenghua Gao

Biomolecular generators are often adapted with reward feedback to improve task-specific utility, but pushing utility alone can concentrate generation on a narrow family of candidates. Maintaining diversity is difficult because sample diversity is a set-level property. We introduce Supergroup Relative Policy Optimization (SGRPO), a flexible GRPO-style framework that directly constructs rewards from set-level diversity. For each condition, SGRPO samples a supergroup of candidate sets, compares their diversity under the same condition, and redistributes the group diversity reward to individual rollouts through leave-one-out diversity contributions before combining it with rollout-level utility. This design decouples SGRPO from a particular generator, utility reward, or diversity metric, and allows instantiation with different GRPO-style approaches. We evaluate SGRPO on de novo small-molecule design, pocket-based small-molecule design, and de novo protein design, instantiating it with both GRPO and Coupled-GRPO across autoregressive and discrete diffusion generators. Across decoding sweeps, SGRPO expands the utility-diversity Pareto frontier and achieves the best frontier-level metrics relative to pretrained generators, GRPO, and memory-assisted GRPO when applicable. Our analyses further show that direct set-level diversity rewards remain effective with small groups and help preserve broader generation-distribution coverage during post-training. The code is available at https://anonymous.4open.science/r/SGRPO/README.md.


QDMouse4M: A Multi-View 3D Mouse Spontaneous Behavior Dataset with Quantum-Dot Markers

Jingyang Ke ⋅ Amartya Pradhan ⋅ Xueling Zhang ⋅ Weihan Li ⋅ Anqi Wu ⋅ Jeffrey Markowitz

Quantitative analysis of animal behavior in neuroscience increasingly relies on accurate 3D pose, yet current large-scale mouse datasets are restricted to top-down views and suffer from occlusion and keypoint ambiguity. We introduce QDMouse4M, a six-view dataset of freely moving mice with physically grounded 2D and 3D keypoint annotations obtained from subdermal quantum-dot fluorescence markers. QDMouse4M contains over four million reflectance frames with matched fluorescence frames, 2D and 3D pose trajectories and behavior classifications. We use the dataset to evaluate markerless 2D and 3D pose estimation with SLEAP and Lightning Pose 3D across training-set sizes, backbones, and in-session/out-of-session splits. We further show that the released trajectories support downstream spontaneous-behavior analysis with Keypoint-MoSeq and stride-level gait measurements. QDMouse4M provides a benchmark-scale, physically grounded resource for developing and evaluating pose estimation, behavior understanding, and biomechanics models under realistic dark-environment laboratory conditions.


QT-Net: Rethinking Evaluation of AI Models in Atomic Chemical Space

Pablo Martínez Crespo ⋅ Stefano Ribes ⋅ Martin Rahm ⋅ Richard Johannes Maximilian Beckmann ⋅ Robert Jordan ⋅ Marisa Gliege ⋅ Santiago Miret ⋅ Vijay K Narasimhan ⋅ Rocío Mercado

Atomic properties such as partial charges or multipoles encode chemically meaningful information that can inform downstream molecular property prediction, but their evaluation as machine learning targets has been complicated by the absence of a principled out-of-distribution evaluation protocol at the atomic level. In this work, we propose a held-out evaluation protocol that clusters atomic environments by SOAP descriptors and computes metrics accounting only for cluster labels unseen during training. Following this procedure, we use 5×5 cross-validation and Tukey's HSD to run a statistically rigorous comparison of E(3)-equivariant against non-equivariant, rotationally augmented models for predicting electron populations and multipoles of H, C, N, and O atoms. Building on our results, we introduce the Quantum Topological Neural Network (QT-Net), a rotationally augmented, non-equivariant graph neural network. We show that QT-Net can be used to infer properties of atoms in molecules from QM9 outside our training set, and that these inferred properties can yield improvement when used as input features for downstream molecular property prediction. To further validate the framework, molecular dipole moments computed from QT-Net's per-atom outputs recover the ground-truth values reported in QM9. We release all code and data, including a JAX implementation of QT-Net, to support the broader use of learned QTA properties as inductive biases for atomic-scale molecular machine learning.


quanda: An Interpretability Toolkit for Training Data Attribution Evaluation

Dilyara Bareeva ⋅ Galip Ümit Yolcu ⋅ Anna Hedström ⋅ Niklas Schmolenski ⋅ Thomas Wiegand ⋅ Wojciech Samek ⋅ Sebastian Lapuschkin

Training data attribution (TDA) methods estimate the influence of individual training samples on a model's predictions, offering a principled lens for interpreting neural networks. Despite rapid methodological progress, TDA evaluation remains scattered across papers and rarely reproducible, making fair comparison between methods challenging. We introduce quanda, a Python toolkit that standardizes TDA evaluation. quanda features a comprehensive collection of evaluation metrics and ready-to-use benchmarks spanning image classification, text classification, and language modeling, alongside a uniform interface for integrating diverse TDA implementations. We showcase quanda through a comparative study of existing TDA methods and find that no approach excels uniformly across metrics, with clear trade-offs between ground-truth faithfulness and downstream task performance. By consolidating the tooling around TDA, quanda equips the community to evaluate attribution methods rigorously and reproducibly. The toolkit is thoroughly tested, documented, and available as an open-source library on PyPI and at https://anonymous.4open.science/r/quanda.


Quantitative Assessment of Crystal Structure Prediction

Sergio Rincón ⋅ Gabriel González ⋅ Nicolás Andrade ⋅ Rafael Velasquez ⋅ Sebastian Ojeda ⋅ Juanita Puentes ⋅ Nicolás Aparicio ⋅ Paula Cárdenas ⋅ Pablo Arbelaez

Crystal Structure Prediction (CSP) aims to predict a material's 3D structure from its chemical composition, with applications from energy storage to pharmaceuticals. Despite its importance, existing evaluation frameworks suffer from key limitations: they rely on threshold-sensitive or incomparable metrics, focus predominantly on geometric performance while overlooking physical plausibility, and fail to account for polymorphism, where a chemical composition can crystallize into multiple stable structures. In parallel, recent Symmetry-Informed CSP (SICSP) methods often report results alongside standard CSP baselines while receiving crystallographic templates at inference, which constrain the search space and make SICSP a fundamentally different task. To address these limitations, we introduce an Orientation-invariant, Polymorph-Aware, Large Benchmark (OPAL-Bench) that evaluates CSP and SICSP as separate tasks. We evaluate 7 CSP and 2 SICSP methods under OPAL-Bench across 4 datasets using threshold-free geometric and physical distance metrics. Our results show that no method dominates across all metrics, and that geometric performance is not strongly correlated with physical plausibility. Motivated by this gap, we provide OPAL-CSP, an SE(3)-equivariant flow-matching baseline that performs competitively on geometric metrics and consistently ranks among the best on physical ones. By unifying geometric and physical metrics under a polymorph-aware protocol, OPAL-Bench enables a more comprehensive assessment of generative quality in CSP.


Quantized Reasoning Models Think They Need to Think Longer, but They Do Not

Sanae Lotfi ⋅ Polina Kirichenko ⋅ Steven Li ⋅ Zechun Liu

Post-training quantization (PTQ) is widely used to deploy large language models efficiently, but its effect on reasoning models is not well understood. Across math, coding, and science QA, we find that aggressive PTQ reduces accuracy while increasing chain-of-thought (CoT) length. Surprisingly, we show that in up to 52% of the quantized models' failures, models reach the right answer in intermediate reasoning steps but do not output it as a final answer. To understand why quantization leads to this increase in overthinking errors, we measure the token-level KL divergence between quantized and full-precision output distributions. Positions with high KL divergence correlate strongly with high next-token entropy, and at these positions quantized models disproportionately sample overthinking markers such as “wait", “but", and “alternatively". We show that simply introducing a training-free logit penalty on a curated set of overthinking markers can reduce CoT length by 12--23% while preserving or improving accuracy across 5 models (1.5B--32B parameters), 3 quantization methods, and 5 benchmarks, yielding the best Pareto frontier of accuracy against reasoning cost. Overthinking errors produced by quantized models are particularly reduced by up to 58%.


Querying Counterfactuals on Tissue Graphs with Supervised Disentanglement

Abdul Moeed ⋅ Stefan Schrod ⋅ Martin Rohbeck ⋅ Marc J Bonder ⋅ Pavlo Lutsik ⋅ Oliver Stegle ⋅ Daniel Dimitrov

Tissue graph counterfactuals ask how a cell's expression would change under altered spatial contexts. Such queries are central to predicting cell behavior in tissues, but lack a unified definition, with existing methods targeting specific intervention types or treating cells as i.i.d. In this work, we first formalize tissue graph counterfactuals as a class of spatial interventions that either rewire connections between cells (edge perturbation) or modify the expression of their neighbors (node perturbation). We then introduce Cellina, a modeling framework that uses supervised disentanglement to decompose a cell's intrinsic state from its spatial context, using the latter as a conditioning input for counterfactual predictions. Across benchmarks spanning 2.5 million spatially-resolved cells in colorectal cancer and mouse brain, Cellina outperforms spatially-informed and non-spatial competitors in counterfactual predictions, disentanglement, and scalability. Additionally, we show that Cellina reveals biologically distinct cancer subdomains in an unsupervised manner and enables targeted neighbor perturbation simulations.


RADIUM: RadioActive Decay of Image-Underlaid Marks

Michel Meintz ⋅ Louis Kerner ⋅ Maitri V Shah ⋅ Simon Hector ⋅ Franziska Boenisch ⋅ Adam Dziedzic

Modern image generative models are able to produce photorealistic images. As those images become increasingly indistinguishable from real data and are published online, they are often scraped for subsequent training runs of new generative models. This practice of training on generated data has been shown to degrade model performance and cause model collapse. A possible mitigation lies in embedding radioactive watermarks into generated content. Radioactive watermarks are robust marks that are detectable in outputs of new models trained on watermarked data, enabling provenance tracing of generated content. In this work, we analyze the persistence of image watermarks across multiple training-generation runs. To do so, we introduce a novel statistical testing method RADIUM (RadioActive Decay of Image-Underlaid Marks) for reliable radioactivity detection across various watermarking methods. Using our RADIUM method, we observe disparate radioactivity across watermarking methods for image generative models. Only few watermarks remain detectable in subsequently trained models, while most decay severely, especially when the architecture of models differs. Our analysis highlights the critical need for more radioactive watermarking methods in the vision domain.


RadOmni: Advancing Foundation Model for Non-contrast CT with Omni Radiology Knowledge

Fenghe Tang ⋅ Weiwei Cao ⋅ Wenxin (Wendy) Ma ⋅ Kai Cao ⋅ Huanhuan Liu ⋅ Ling Zhang ⋅ S. Kevin Zhou ⋅ Jianpeng Zhang

Non-contrast CT (NCCT) is widely used in clinical practice, yet its limited soft-tissue contrast often renders visual information insufficient for accurate diagnosis. In routine radiology workflows, NCCT is commonly accompanied by contrast-enhanced CT, radiology reports, and pathology reports, which provide complementary knowledge at different levels, from structural and morpho-functional cues to semantic findings and pathological evidence. However, existing pretraining paradigms address this challenge only partially, each introducing complementary but still limited constraints on NCCT representations. To address this limitation, we propose RadOmni, a unified omni-pretraining framework that injects multimodal radiology knowledge into representation learning for non-contrast CT. RadOmni formulates a curriculum consolidation learning strategy, in which knowledge is progressively injected according to optimization difficulty and knowledge level through four stages: masked image modeling on non-contrast CT, cross-phase transfer of anatomical structures and enhanced patterns from contrast-enhanced CT, report-guided semantic learning, and pathology-informed diagnostic enhancement. To mitigate optimization conflicts across heterogeneous objectives and reduce knowledge forgetting, each stage retains and jointly optimizes the objectives from preceding stages, while early encoder layers are selectively frozen to preserve generic representations. To enable such multi-source pretraining, we curate Omni-CT10K, a large-scale CT dataset with more than 10K NCCT scans featuring clinically co-occurring multimodal supervision. Extensive experiments on Omni-CT10K, two in-house benchmarks from different centers, and the public MSD benchmark show that RadOmni consistently achieves strong performance across classification, segmentation, pathological TNM staging, and radiology report generation tasks. Further analyses demonstrate favorable scaling behavior, robust cross-center generalization, and transferable representations.


RAM-Net: Linear-Time Sequence Modeling with Sparsely Addressable State

Kaicheng Xiao ⋅ Haotian Li ⋅ Liran Dong ⋅ Guoliang Xing

Linear attention offers an efficient alternative to full attention with a fixed-size recurrent state. However, this state is shared by all tokens, so information from distinct tokens becomes superposed within it and produces inter-token interference that degrades long-range fine-grained recall. To address this issue, we propose RAM-Net, which replaces dense access to a shared state with sparse address-based access. RAM-Net organizes the recurrent state as a fixed-size array of independent slots and uses an Address Decoder that maps each key or query into a sparse address, selecting a small subset of slots to write to or read from at each step. This design directs tokens with non-overlapping addresses to disjoint slots, suppressing inter-token interference, while keeping per-step overhead dependent only on the number of accessed slots rather than the total state size. Empirically, RAM-Net shows a clear advantage over strong linear baselines on fine-grained long-range retrieval and remains competitive on standard language modeling and commonsense reasoning, while accessing far fewer state elements per step than these baselines (e.g., 32$\times$ fewer than Mamba2).


RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations

Yanhao Ge ⋅ Shanyan Guan ⋅ Weihao Wang ⋅ Ying Tai ⋅ Mingyu You

Natural images are continuous, yet most generative models synthesize them on discrete grids, limiting resolution-flexible generation. Continuous neural fields enable resolution-free rendering, but prior methods introduce continuity only at the decoding stage as an interpolation module, leaving the generative latent space discretized and reconstruction-oriented. We propose RaPD (Resolution-agnostic Pixel Diffusion), which performs diffusion in a continuous Neural Image Field (NIF) latent space. RaPD bridges this reconstruction-generation gap with Semantic Representation Guidance for generation-aware latent learning and a Coordinate-Queried Attention Renderer for coordinate-conditioned, scale-aware rendering. A single denoised latent can be rendered at arbitrary resolutions by changing only the query coordinates, keeping diffusion cost fixed. Experiments demonstrate superior generation quality and resolution scalability.


Rare-Class Signal Suppression in Long-Tailed Multi-Expert Fine-Tuning

Zhiheng Gong ⋅ YuHeng Yang ⋅ Pengkun Wang ⋅ Yang Wang

Multi-expert ensemble frameworks have achieved strong performance on long-tailed visual recognition by training heterogeneous expert heads on a shared backbone, but their gradient coordination strategies are inherited from balanced multi-task learning, where resolving directional conflict is the canonical challenge. We show that head-class bias acts sequentially across three pipeline interfaces, namely loss geometry, gradient coordination, and inference fusion, forming a suppression cascade in which rare-class gradient energy can be erased before it influences model parameters, a failure mode that direction-focused analysis cannot detect. A controlled study reveals that direction-modifying and simple-aggregation solvers are statistically equivalent once rare-class source signal is amplified, while energy-balancing objectives erase this amplification entirely. To break this cascade, we propose Gale (Gradient-energy Allocation for Long-tail Ensembles), a framework that applies gradient-energy allocation at every cascade stage, amplifying tail-class source signal through frequency-calibrated loss geometry, preserving it through a non-penalizing coordination rule, and routing predictions through class-specific expert competence. On CIFAR-100-LT (ρ=100), Gale achieves 59.2% overall accuracy and 46.1% few-class accuracy across four seeds (σ=0.23%), surpassing the prior best by +6.2%, with consistent improvements on ImageNet-LT and iNaturalist 2018. Code is available at Supplementary Material.

Causal autoregressive video diffusion models support real-time streaming generation by extrapolating future chunks from previously generated content. Distilling such generators from high-fidelity bidirectional teachers yields competitive few-step models, yet a persistent gap between the history distributions encountered during training and those arising at inference constrains generation quality over long horizons. We introduce the Real-time Autoregressive Video Extrapolation Network (RAVEN), a training-time test framework that repacks each self rollout into an interleaved sequence of clean historical endpoints and noisy denoising states. This formulation aligns training attention with inference-time extrapolation and allows downstream chunk losses to supervise the history representations on which future predictions depend. We further propose Consistency-model Group Relative Policy Optimization (CM-GRPO), which reformulates a consistency sampling step as a conditional Gaussian transition and applies online Reinforcement Learning (RL) directly to this kernel, avoiding the Euler-Maruyama auxiliary process adopted in prior flow-model RL formulations. Experiments demonstrate that RAVEN surpasses recent causal video distillation baselines across quality, semantic, and dynamic degree evaluations, and that CM-GRPO provides further gains when combined with RAVEN.


Readiness-Aware Sample Selection for Noisy Labels with Class Imbalance

Chihyeon Choi ⋅ Sangho Lee ⋅ Jiho Hong ⋅ Youngdoo Son ⋅ Hyungrok Do

Noisy labels can cause deep neural networks to overfit incorrect annotations, leading to biased predictions and poor generalization. A common remedy is to select clean samples after a short warm-up phase, under the assumption that the model begins to learn class-specific patterns from the early training stage. However, this assumption often fails under class imbalance, as the majority classes can hinder the model from learning distinctive features of minority classes. In this paper, we propose a readiness-aware sample selection strategy that identifies clean samples only from classes whose distinctive features have been sufficiently learned and are thus considered ready for selection. We further introduce a novel negative learning scheme to enhance class separability by discouraging confusion with the most similar incorrect classes. The proposed method is supported by theoretical analysis and demonstrates outstanding performance on both benchmark and real-world datasets, showing improved robustness and generalization under noisy and class-imbalanced conditions.


Reading the Unreadable: Text-Aware Image Super-Resolution Needs Reasoning

Jaeseong Lee ⋅ Jinwoo Kim ⋅ Jinho Jeong ⋅ Seon Joo Kim

Text-aware image super-resolution (TAISR) aims to recover high-resolution images while preserving textual content. Yet some degraded text cannot be restored from local visual evidence alone: humans often read it by reasoning over image context, logical patterns, and prior knowledge. This raises a fundamental question: does TAISR need reasoning? We answer with two new diagnostic tools: ReasonText, a human-curated benchmark with perception-grounded difficulty levels that separate reasoning-required text from locally recoverable text, and GTTCA, a crop-based recognition accuracy metric that isolates restoration quality from detection errors. Evaluating twelve recent SR models, we find a consistent and substantial performance drop on reasoning-required text, revealing a systematic blind spot in current TAISR designs. To address this limitation, we build on two observations: (1) SR models benefit from ground-truth text labels provided as captions, and (2) modern MLLMs can infer degraded text from LR images through reasoning. These observations motivate Reasoning Transfer via Captioning (RTC), a training-free plug-in that uses an external MLLM as a reasoning module and injects its inferred text into SR models through natural-language captions. Surprisingly, with RTC, real-world SR models can even outperform dedicated text-aware SR models, suggesting that reasoning should be treated as a core design axis for future TAISR.


RealityTest: How People Probe AI Identity and Whether Models Disclose It

Anna Gausen ⋅ Sarenne Wallbridge ⋅ Bessie O'Dell ⋅ Christopher Summerfield ⋅ Hannah Rose Kirk

AI systems are increasingly deployed in conversational settings where users may be uncertain whether they are speaking with a human or an AI. Despite mounting regulatory attention to this known safety risk, existing evaluations of AI disclosure are typically English-only, based on machine-generated questions, and restricted to text. We present RealityTest to comprehensively test whether AI systems disclose their identity when asked. The benchmark is the first large-scale multimodal and multilingual evaluation, grounded in human data on how people actually encounter and question AI identity in the real-world. Alongside the benchmark, we release the underlying dataset of 3,152 identity-probing queries collected from $\textasciitilde$750 participants across 49 countries and five languages, in text and speech scenarios. We find that only 31\% of people ask about identity directly, and that the questions people ask are far more diverse than machine-generated queries. We test 17 text and 6 speech models, and find substantial variation in disclosure behaviour. However, a single suppression instruction reduces disclosure rates to below 30\%, even in the best-performing models. Validating our investment in diverse, human-grounded evaluation data, we find that how the question is phrased and the context of the conversation matter more for disclosure than which model is being tested. Safety evaluations built on narrow or synthetic query sets risk mischaracterising how models behave in realistic deployment settings.

Reasoning models increasingly depend on test-time computation, but raw accuracy conflates intrinsic ability, problem difficulty, guessing effects, and reasoning budget. Item Response Theory (IRT) adjusts for problem difficulty, yet collapses test-time compute into a single latent score. We propose reasoning-aware IRT, an extension in which each model's effective ability depends on baseline ability, reasoning-budget responsiveness, and overthinking curvature. Parameter estimation is performed via multi-pass generalized EM with post-hoc initialization refinement. On LiveBench with 42 language models, our method improves rank stability, gives stronger held-out correctness discrimination, and reveals interpretable structure unavailable to classical IRT: a unified benchmark--ability scale with quantified reasoning-token benefits, budget-efficient and overthinking regimes, and rank reversals obscured by raw scores. Our framework offers a compact, explainable evaluation tool for reasoning models under test-time scaling.


Reasoning Poisoning: Utilizing Social-Engineering to Steer Chain-of-Thought

Matan Levy ⋅ Ilan Zendel ⋅ Stav Cohen ⋅ Amit LeVi ⋅ Avi Mendelson

The transition of Large Language Models (LLMs) to autonomous reasoning agents introduces a critical vulnerability within the reasoning trace itself. We present Reasoning Poisoning, a framework demonstrating how adversaries can exploit an agent via model-directed social engineering by manipulating retrieved context to weaponize the model's Reinforcement Learning from Human Feedback (RLHF) alignment. Our central paradigm, Logic Hijacking, exploits generalized alignment constraints by introducing fictitious hazards that force the agent to actively eliminate legitimate targets and select the attacker's target as the only "valid" alternative. Evaluating six state-of-the-art, production-deployed reasoning models across 10 domains, we show that this attack fundamentally overrides standard logic, achieving Attack Success Rates (ASR) exceeding 83%. The vulnerability remains highly effective even when the attacker controls only 10% of the retrieved context and exhibits robust success regardless of the adversarial payload's position within the evidence window. Controlled baselines confirm that this steering is driven by adversarial constraints, not ordinary promotional bias. Furthermore, we find that agents are entirely unresponsive to simulated social proof (e.g., user upvotes), relying instead on stylistic and semantic cues that mimic their own alignment training. Ultimately, our findings reveal a troubling paradox: the very alignment mechanisms designed to make models helpful and safe can be exploited to seamlessly hijack the reasoning process.


Reconfiguring Procedural Knowledge for Compositional Robot Skill Adaptation

Daehee Lee ⋅ Dongsu Lee ⋅ Sanghyun Ahn ⋅ Minjong Yoo ⋅ Honguk Woo

General-purpose behavior models increasingly internalize broad procedural knowledge at training time, making adaptation less a problem of learning entirely new behaviors than of reconfiguring existing competence for new goals, object relations, and skill interfaces. We study this as reconfigurability, the ability to realize new behaviors by restructuring components and relations cost-efficiently, under sparse governing information, where a task specification provides high-level goals or abstract directives but omits the concrete bindings and compositions required for execution. We propose RecPo, a framework that represents skills task-agnostically through multiple procedure views and materializes a new task by composing only the transition relations consistent with the task-relevant views. RecPo maintains a procedure knowledge base of local transition relations within each view. At deployment time, it interprets a task demonstration through the same views, converts observed view-level changes into a compositional query, and resolves the query into an executable skill sequence. On an extended Franka Kitchen benchmark with held-out long-horizon tasks outside the stored compositions, RecPo achieves 91.7\% ordered subtask completion, substantially outperforming non-reconfigurable skill retrieval baselines. Real-robot case studies further show deployment-time adaptation through recomposition rather than task-specific retraining.


Rectifying Categorical Flows on Statistical Manifolds for One-Step Generation

Xiaohuan Jia ⋅ Renzhe Xu ⋅ Xiao Wang ⋅ Jiayun Wu ⋅ Shaohua Fan

Generative modeling for discrete data has been significantly advanced by learning continuous vector fields on the statistical manifolds to transport simple priors to target discrete distributions. However, because these continuous-time flows are not constrained to follow geodesics on the statistical manifold, thereby requiring multi-step numerical integration for simulation and leading to substantial computational costs during inference. In this work, drawing inspiration from the ''straight and fast'' properties of Rectified Flow in continuous Euclidean space, we propose the Rectified Statistical Flow (ReSFlow) model, which enables one-step categorical generation on statistical manifolds by explicitly learning geodesic trajectories that adhere to the intrinsic Riemannian geometry. By rectifying the flow on the statistical manifold, we theoretically guarantee that trajectories between prior and target distributions approximate geodesics. This rectification significantly reduces the transport cost by aligning the learned trajectories directly along geodesics on the statistical manifold, allowing the model to map pure noise to target discrete data in a single step. Through extensive experiments on diverse real-world generative tasks, ranging from image and text to biological domains, we demonstrate ReSFlow's superior one-step generation quality, which truly embodies geodesically optimal and rapid categorical generation.


Re-evaluating Confidence Remasking in Masked Diffusion Language Models

Stipe Frković ⋅ Metod Jazbec ⋅ Dan Zhang ⋅ Christian Andersson Naesseth ⋅ Ilija Bogunovic ⋅ Eric Nalisnick

Masked diffusion language models (dLLMs) have recently emerged as a competitive alternative to autoregressive language models, with the promise of faster inference via parallel token generation. A notable limitation of the masked formulation, however, is that once a token has been unmasked it can no longer be revised, leaving dLLMs vulnerable to early sampling mistakes. To address this, a growing body of work has sought to extend masked dLLMs with self-correcting (remasking) capabilities. One appealing subset of these methods does so in a training-free, post-hoc manner based on token confidence, with encouraging early reported results. In this work, we revisit the empirical evaluation of a representative post-hoc remasking method, WINO, and find that under standard decoding settings (shorter block lengths) it brings little-to-no benefit over confidence-based unmasking alone. Extending the evaluation to non-greedy decoding, we find that while confidence-based remasking can mitigate errors introduced by increased stochasticity to some extent, it also exacerbates the diversity collapse previously reported for confidence-based unmasking. Overall, our results show that the benefits of post-hoc confidence-based remasking are highly setting-dependent, underscoring the need for a more comprehensive evaluation framework.

Prediction aggregation aims to combine information from multiple predictors into a more informative one. We study this question in the setting of calibrated predictors, where each prediction must equal the conditional expectation of the quantity being predicted given the predictor's signal. Given several calibrated input predictors and the feature distribution, but not the underlying Bayes probabilities, we ask when one can construct refined calibrated predictors that preserve the information in the original predictors and cannot be further refined using the available information. We formulate calibrated predictors as signaling schemes and define refinement through feature-independent garblings: a predictor refines another if its signal can simulate the other's signal. Constructibility is characterized through observable linear information: each signal corresponds to a vector over the feature space, and a new signal is constructible exactly when its vector lies in the linear span of the input signal vectors. Under this formulation, we establish a sharp algorithmic picture. For deterministic output predictors, bilateral refinement admits a polynomial-time algorithm based on a bipartite graph between the two input signal partitions, while refinement with an arbitrary number of input predictors is $\mathsf{NP}$-hard. In contrast, when randomized output predictors are allowed, we give a polynomial-time algorithm for any number of input predictors by decomposing constructible signal vectors into extreme rays of the associated polyhedral cone.

Large language models (LLMs) support long-context inference but suffer from substantial memory and runtime overhead due to Key-Value (KV) Cache growth. Existing KV Cache eviction methods primarily rely on local attention weights, neglecting the influence of value representations, output projection, and inter-head interactions. In this work, we reformulate KV Cache eviction from a conventional head-wise, weight-averaging approach into an output-aware, layer-wise matrix multiplication approximation problem. We introduce LaProx, a novel eviction strategy that explicitly models the multiplicative interaction between attention maps and projected value states to accurately quantify token contributions while accounting for inter-head dependencies. Building on this metric, we propose the first unified eviction strategy that assigns globally comparable importance scores to tokens, enabling model-wide selection instead of local, head-wise decisions. Experimental results across 19 datasets on long-context benchmarks LongBench and Needle-In-A-Haystack demonstrate that our approach maintains model performance with only 5\% of the KV cache and consistently outperforms prior works across all configurations. Notably, our method achieves up to 2$\times$ accuracy loss reduction under extreme compression scenarios compared to existing state-of-the-art baselines with minimal overhead.


Reformulating Neural Operators in $d+1$ Dimensions for Embedding Evolution

Haoze Song ⋅ Zhihao Li ⋅ Xiaobo Zhang ⋅ Zecheng Gan ⋅ Zhilu Lai ⋅ Wei Wang

Neural Operators (NOs) are powerful architectures for learning mappings between function spaces. While most advances focus on refining kernel parameterizations over the $d$-dimensional physical domain, the evolution of lifted embeddings remains underexplored, which often drives models toward computationally expensive embedding-scaling designs to improve approximation. In this paper, we introduce an auxiliary function dimension that models embedding evolution in operator form, thereby reformulating the NO pipeline in $d+1$ dimensions. We instantiate this framework via Fourier-based operators acting jointly on the physical and auxiliary domains, yielding a basis-diversified auxiliary evolution module as an alternative to brute-force embedding scaling. Across more than ten increasingly challenging benchmarks, ranging from the 1D heat equation to the highly nonlinear 3D Rayleigh–Taylor instability, our model consistently achieves the lowest relative $L_2$ error among the evaluated baselines. Crucially, this advantage is empirically supported by (1) controlled budget-aware comparisons against scaled and ablated baselines; (2) robustness under mixed-resolution training and super-resolution inference; and (3) zero-shot generalization to unseen temporal regimes. In addition, we present a broader set of design choices for lifting and recovery operators, demonstrating their impact on our model’s predictive performance.


REGATE: Confidence-Calibrated Integration of Temporally-Aligned Exogenous Texts for Dynamic Graphs

Liangzu Liu ⋅ Mengzhe Ruan ⋅ Yinjun Wu ⋅ Yang Liu ⋅ Guanjun Wang

Risk assessment and anomaly detection in financial interaction networks often fail under regime shifts triggered by external events such as policy changes or enforcement actions—models may retain high AUC yet suffer sharp drops in average precision and calibration. Although news and regulatory filings contain actionable signals, naïvely injecting raw text or off-the-shelf LLM embeddings into temporal GNNs can introduce temporal leakage, misalignment, and unstable out-of-time ranking. We propose REGATE, a time-causal framework that integrates exogenous documents into dynamic graphs through three coupled components: (1) a schema-guided extractor that converts unstructured documents into auditable, time- and entity-aligned policy tokens with explicit confidence scores; (2) a bounded gating mechanism that fuses these tokens into temporal GNN states as a residual update, downweighting uncertain or stale evidence with a stability guarantee under direction consistency; and (3) a closed-loop retrieval adaptation module that distills the model's own routing attention into a lightweight document scorer without manual relevance labels. On established dynamic-graph benchmarks with documented regime shifts and six backbone architectures (TGN, TGAT, DyGFormer, GraphMixer, CTAN, GeneralDyG), REGATE yields up to +0.16 AP and up to 90% ECE reduction in post-shift windows on Elliptic; cross-domain pilots on credit and equity networks confirm consistent gains. Ablations show improvements stem from time-aligned semantic content rather than timestamps alone, and we report quality–cost trade-offs across open-source and commercial extractors.


Region-Normalized DPO for Medical Image Segmentation

Hamza Kalisch ⋅ Constantin Seibold ⋅ Jens Kleesiek ⋅ Ken Herrmann ⋅ Frederic Jonske

Direct Preference Optimization (DPO) has recently been applied to medical image segmentation, since comparing candidate masks is substantially cheaper than producing new dense pixel-level annotations, making pairwise preference signals an appealing source of supervision. However, existing work derives preferences from ground-truth masks, leaving open whether the approach is sound under the imperfect feedback available in practice. In our analysis of this reformulation, we identify that candidate masks tend to agree over most of the image and differ only in localized regions, yet standard DPO averages the preference signal over the full spatial domain. This couples update strength to disagreement area, diluting small correct refinements and amplifying large misranked differences. Crucially, this bias even degrades performance even under oracle preferences derived from ground truth. We propose Region-Normalized DPO (RN-DPO), which normalizes the likelihood ratio over the disagreement region between candidate masks, removing this coupling. We further provide a systematic empirical study of preference-based segmentation fine-tuning under controlled noisy judges, analyzing the effects of mining strategies, judge types, and reliability regimes. Across two medical segmentation benchmarks, multiple judge configurations, and two backbone architectures, RN-DPO consistently improves over vanilla DPO and robust DPO variants, with gains persisting under noiseless oracle preferences.


Register Anything Model for Generalizable and Robust Point Cloud Registration

Zheng Qin ⋅ Junhua Xi ⋅ Yan Huang ⋅ JiaHong Lai ⋅ Anwen Huang ⋅ Qiong Li ⋅ Guangda Zhang ⋅ Kai Xu

We present Register Anything Model (RAM), a unified model for generalizable 3D point cloud registration that robustly estimates 6-DoF transformations across diverse domains spanning sensors and environments. Existing registration methods are typically tailored to a single domain, with carefully tuned architectures and hyperparameters, and thus degrade markedly when applied to novel domains. We bridge this gap by training a single model on diverse data from indoor, outdoor, and object domains. A central challenge is the severe geometric discrepancy across domains, i.e., differences in scene scale, sensor noise, point density, and structural patterns, which hinders processing with a unified point-based model. To address this, we introduce a geometry-aware normalization strategy that jointly voxelizes and normalizes point clouds into a standarized geometric space, adaptively based on their geometric properties to preserve salient local structures while standardizing resolution. Built on this normalization, RAM follows a hierarchical coarse-to-fine paradigm, and employs an efficient geometric transformer that models local geometric consistency instead of over all points. This design mitigates discretization noise introduced by voxelization, and improves computational efficiency. Extensive experiments on $12$ benchmarks demonstrate state-of-the-art performance in terms of inlier ratio and registration recall. Notably, our method outperforms the counterparts trained on each separate domain, showing superior generalization capability. Code and models will be released upon publication.


Regularization Paths for Continuous DAG Learning

Diyang Li ⋅ Fei Wang ⋅ Kyra Gan

Continuous DAG learning formulates combinatorial structure learning as differentiable optimization over weighted adjacency matrices. In these methods, the estimated adjacency matrix depends critically on a sparsity-controlling regularization parameter, whose magnitude governs the recovered graph skeleton. While extensive empirical evidence underscores the centrality of this parameter, existing practice relies on solving a sequence of independent optimization problems over a grid of tuning parameters, which is computationally wasteful, statistically unstable, and provides only a fragmented view of how the adjacency matrix evolves. In this work, we study the regularization path of stationary solutions in continuous DAG learning and propose DAG-flow, an exact path-following framework that compactly encodes the entire family of estimators. We show that for prominent formulations including NOTEARS, GOLEM, and DAGMA, stationary solutions evolve according to a piecewise-smooth matrix-valued dynamical system once the active support is fixed. Our analysis derives the explicit systems governing these branches, identifies the events at which the path changes regime, and clarifies how nonsmooth sparsity and nonconvex acyclicity interact along the path. Unlike grid-based search, DAG-flow exposes the geometry of model variation across regularization levels, which enables the construction of structured candidate graph sets and provides direct insight into edge-level stability. Across synthetic and real benchmarks, our DAG-flow outperforms grid-search baselines at a fraction of the compute and consistently improves structural stability.


Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow Models

Andreas Bergmeister ⋅ Stefanie Jegelka ⋅ Nikolas Nüsken ⋅ Carles Domingo i Enrich ⋅ Jakiw Pidstrigach

Diffusion and flow-matching models scale because pretraining simply regresses against a closed-form target built by analytically noising clean data samples. In many settings, only a reward function is available, scoring how desirable a generation is, so training must proceed online from a pretrained model. Existing methods either rely on costly SDE rollouts, sometimes with reward gradients, or adopt heuristic reward-dependent variants of the pretraining objective. Under a stochastic optimal control formulation of KL-regularized reward maximization, the optimal generative process tilts only the clean-endpoint distribution and leaves the conditional noising law unchanged. Combining this path-space characterization with the adjoint-matching optimality condition and a REINFORCE estimator to avoid reward gradients, we derive Reinforce Adjoint Matching (RAM). At each step, we draw a clean endpoint from the current model with any off-the-shelf sampler, evaluate its reward, noise it to several training states, and regress against a closed-form, reward-weighted target. Like the pretraining objective, RAM is simple and scales. On Stable Diffusion 3.5M, RAM achieves the highest reward on composability, text rendering, and human preference. It is much more training efficient, reaching Flow-GRPO's peak GenEval accuracy in~$50\times$ fewer training steps.


Reinforcement Learning for Exponential Utility: Algorithms and Convergence in Discounted MDPs

Gugan Chandrashekhar Mallika Thoppe ⋅ Prashanth L.A. ⋅ Ankur Naskar ⋅ Sanjay P. Bhat

Reinforcement learning (RL) for exponential-utility optimization in discounted Markov decision processes (MDPs) lacks principled value-based algorithms. We address this gap in the fixed risk-aversion setting. Building on the Bellman-type equation for exponential utility studied in Porteus [1975], we derive two Q-value-style extensions and show that the associated operators are contractions in the $L_\infty$ and sup-log/Thompson metrics, respectively. We characterize their fixed points and prove that the induced greedy stationary policy is optimal for the exponential-utility objective among stationary policies. These structural results lead to two model-free algorithms: a two-timescale Q-learning--style algorithm, for which we establish almost-sure convergence and provide finite-time convergence rates via timescale separation, and a one-timescale algorithm governed by a sublinear power-law operator. Since the latter does not admit a global contraction in standard metrics, we prove its convergence using delicate arguments based on local Lipschitzness, monotonicity, homogeneity, and Dini derivatives, and provide a scalar finite-time analysis that highlights the challenges in obtaining convergence rates in the vector case. Our work provides a foundation for value-based RL under exponential-utility objectives.


Reinforcement Learning-Guided Symbolic Execution for Efficient and Exploitable Smart Contract Analysis

Zhaoxuan Li ⋅ Ziming Zhao ⋅ Siqi Lu ⋅ Rui Zhang ⋅ Wenhao Li ⋅ Xiaofei Yue ⋅ Fan T Zhang

Currently, frequent security incidents of Ethereum contracts have caused billions of dollars in losses. There is a pressing need to identify defective contract code and generate exploit call sequences to reproduce attacks and ensure detection accuracy. Nevertheless, the state-of-the-art (SOTA) detection methods based on symbolic execution and fuzzy testing cannot achieve the desired performance due to the state explosion problem caused by contract characteristics such as cross-contract calls and loop branches, especially in large wild contracts. To tackle this problem, we propose RSymX, an assembly strategy-guided symbolic execution for contracts that leverages reinforcement learning to dynamically balance code coverage and vulnerability discovery. Extensive experiments on open-sourced datasets demonstrate that RSymX achieves an overlap score of 92.39% and identifies many vulnerable wild instances that SOTAs misreport. Especially, its dynamic strategies improve the efficiency of state exploration by 2x~4x, and in some cases up to 10x, offering guidance for future symbolic-execution-based analyzers. The code and data are available in https://github.com/ContractAudit/Code.


Reinforcement Learning with Verifiable Physics: Post-training LLMs for PDE Solver Generation

Pengfei Cai ⋅ Utkarsh Utkarsh ⋅ Alan Edelman ⋅ Christopher Rackauckas ⋅ Rafael Gomez-Bombarelli

Partial differential equations (PDEs) are foundational to modeling in science and engineering, but constructing reliable numerical solvers remains labor-intensive, demanding expert knowledge of discretization schemes, stability conditions, and boundary treatments. Recent work has begun to frame PDE solving as a code-generation task for large language models (LLMs), yet existing approaches operate primarily at inference time: relying on prompting, debugging, self-refinement, and test-time scaling rather than adapting the model itself. In parallel, reinforcement learning with verifiable rewards has emerged as a powerful post-training paradigm for code and math reasoning, but its verifiers are typically binary: a compiler runs, or a test passes. Such signals discard the graded structure of scientific correctness, where two solvers may both execute and yet differ in solution accuracy by orders of magnitude. In this work, we introduce RLVP: Reinforcement Learning with Verifiable Physics, an RL post-training framework for multi-PDE solver code generation.RLVP addresses this verifiability gap with a hybrid verifier: hard program-validity checks ensure executability, while continuous physics rewards score function-space accuracy and PDE-residual consistency. A single policy is post-trained across diverse PDE families spanning hyperbolic, parabolic, elliptic, and incompressible-flow systems. RLVP improves over both pre-trained and supervised-only baselines on PDE benchmarks, and shows zero-shot improvement transfer to held-out PDEs. We show that a smaller LLM post-trained with RLVP can outperform prompting a frontier model on in-distribution PDE solver generation. The trained policy shows evidence of compositionality in numerical motifs: it recombines stencils, time-stepping schemes, and boundary-handling primitives learned from the PDEs used in training into generated solvers for unseen PDE problems.


Reliability-Coupled Manifold-Aware Diffusion for Missing-Modality Inference

Yiming Ren ⋅ Yuhao Fang ⋅ Qing Zhou ⋅ Ming Li ⋅ Ye Zhang ⋅ Zijian Wang ⋅ Yao Lu ⋅ Chun Li

In real-world multimodal time-series inference, a common practice is completion–inference decoupling: missing modalities are heuristically imputed and then passed to deterministic fusion and decision making. We show that this chained design yields unreliable predictions under missingness, noise, and temporally varying modality reliability, because imputation errors propagate without uncertainty feedback. Through theoretical and empirical analysis, we identify latent-consistent completion and evidential reliability modeling as two key ingredients for robust multimodal inference. Building on these insights, we propose MD2E-MI, a manifold-aware diffusion-based imputation and evidence-driven multimodal inference framework that couples diffusion-based completion with uncertainty-aware inference in a single pathway. Experiments demonstrate stable performance across diverse missingness and noise regimes, while providing interpretable, temporally varying reliability estimates.


Reliable Chain-of-Thought via Prefix Consistency

Naoto Iwase ⋅ Yuki Ichihara ⋅ Mohammad Atif Quamar ⋅ Junpei Komiyama

Large Language Models often improve accuracy on reasoning tasks by sampling multiple Chain-of-Thought (CoT) traces and aggregating them with majority voting (MV), a test-time technique called self-consistency. When we truncate a CoT partway through and regenerate the remainder, we observe that traces with correct answers reproduce their original answer more often than traces with wrong answers. We use this difference as a reliability signal, **prefix consistency**, that weights each candidate answer by how often it reappears under regeneration. It requires no access to token log-probabilities or self-rating prompts. Across five reasoning models and four math and science benchmarks, prefix consistency is the best correctness predictor in most settings, and reweighting votes by it reaches Standard MV plateau accuracy at up to $21\times$ fewer tokens (median $4.6\times$).


Reliable Federated Multi-View Learning via Conflict-Aware Evidence Calibration

Daoyuan Li ⋅ Zuyuan Yang ⋅ Hao Yang ⋅ Jiawen Kang

Federated Multi-View Learning (FedMVL) enables multiple clients with heterogeneous data views to collaboratively train global models without sharing private data. Existing FedMVL algorithms often overlook the reliability of local predictions, as both uncertainty quantification and calibration remain challenging in privacy-preserving federated settings. Heterogeneous sensing devices yield varying view qualities, but view-isolated training further exacerbates inherent bias of local model, inducing unreliable outputs. Without proper calibration, these artifacts cause spurious cross-view conflicts, ultimately undermining the accuracy and reliability of the aggregated global model. To address these issues, we propose a Reliable Federated Multi-View Learning with evidence calibration (FedRMVL-CAL), a FedMVL framework with adaptive evidence fusion. The local calibration scheme harmonizes uncertainty across views, while the fusion mechanism adaptively reconciles conflicting opinions, enhancing robustness in global inference. Theoretical analysis and extensive experiments demonstrate that FedRMVL-CAL achieves superior accuracy and reliability compared to existing approaches, ensuring trustworthy global predictions across heterogeneous views.


ReorgGS: Equivalent Distribution Reorganization for 3D Gaussian Splatting

Luchao Wang ⋅ Kaimin Liao ⋅ Hua Wang ⋅ Qian Ren ⋅ Zhi Chen ⋅ Yaohua Tang

While a converged 3D Gaussian Splatting (3DGS) model may accurately approximate a target scene, its underlying parameterization often becomes severely ill-suited for further optimization. We identify this late-stage bottleneck as \emph{parameterization degeneration}: high-opacity floaters truncate gradient flow to background surfaces via alpha compositing, and redundant overlapping clusters cause severe parameter coupling with nearly collinear Jacobian responses. These structural barriers explain why continued optimization plateaus, even when removable artifacts persist. To break this deadlock, we propose ReorgGS, an equivalent distribution reorganization method. By treating the converged Gaussian set as an empirical probability field, ReorgGS resamples centers, estimates local anisotropic covariances via kNN, and initializes a low-opacity state before resuming optimization. Unlike standard opacity reset—which only rescales weights on a flawed topology—ReorgGS fundamentally rebuilds the spatial and visibility structure. Our analysis reveals a crucial insight: \emph{distributional equivalence does not imply optimization equivalence}. By preserving scene support while drastically improving gradient accessibility and reducing opacity-weighted overlap, ReorgGS provides a vastly superior optimization landscape. Under the same optimization budget and fixed Gaussian count, ReorgGS breaks the performance ceiling, suppresses persistent floaters, and reduces rendering overhead by eliminating redundant overlap.


Replicability is Asymptotically Free in Multi-armed Bandits

Junpei Komiyama ⋅ Shinji Ito ⋅ Yuichi Yoshida ⋅ Souta Koshino

We consider a replicable stochastic multi-armed bandit algorithm that ensures, with high probability, that the algorithm's sequence of actions is not affected by the randomness inherent in the dataset. Replicability allows third parties to reproduce published findings and assists the original researcher in applying standard statistical tests. We observe that existing algorithms require $O(K^2/\rho^2)$ times more regret than nonreplicable algorithms, where $K$ is the number of arms and $\rho$ is the level of nonreplication. However, we demonstrate that this additional cost is unnecessary when the time horizon $T$ is sufficiently large for a given $K, \rho$, provided that the magnitude of the confidence bounds is chosen carefully. Therefore, for a large $T$, our algorithm only requires $K^2/\rho^2$ times smaller amount of exploration than existing algorithms. To ensure the replicability of the proposed algorithms, we incorporate randomness into their decision-making processes. We propose a principled approach to limiting the probability of nonreplication. This approach elucidates the steps that existing research has implicitly followed. Furthermore, we derive the first lower bound for the two-armed replicable bandit problem, which implies the optimality of the proposed algorithms up to a $\log\log T$ factor for the two-armed case.


Replicas-as-Variables: A Planner for Throughput-Optimal Routing in LLM Deployment

Shaojiang Wang ⋅ Geer Yang ⋅ Wenbo Wang ⋅ Bin Wu ⋅ Pengcheng Wang

Deploying Large Language Models on heterogeneous GPU clusters is often bottlenecked by rigid parallelism strategies that fail to fully exploit hardware asymmetry. To address this, we introduce ReVar, a holistic framework that optimizes replica placement and activation routing by treating replica counts as primary decision variables for all model components, including both dense stages and MoE experts. By modeling inference as a coupled resource-allocation and flow-routing problem, ReVar replaces static pipelines with fluid execution graphs that utilize uneven replication and multi-path routing. Supported by rigorous theoretical foundations, including proving the problem's NP-hardness, providing a MILP formulation for throughput upper bounds, and deriving closed-form solutions for constrained regimes, ReVar delivers significant practical gains. In evaluations, it achieved the highest throughput in 14 out of 15 settings, demonstrated up to a 3.1 times improvement in steady-state throughput in mixed-GPU environments, and was successfully validated on a real physical heterogeneous cluster.


RepoMirage: Probing Repository Context Reasoning in Code Agents with Perturbations

Hanyu Li ⋅ Yichi Zhang ⋅ Speed Zhu ⋅ Hang Su ⋅ Jun Zhu ⋅ Yinpeng Dong

Code agents are currently having skillful performance on repository-level software engineering benchmarks, but it remains unclear whether success on end-to-end tasks such as issue resolution truly reflects $\textit{repository context reasoning}$, the ability to identify the task-relevant information across multiple files and reason over the relations among them. To investigate this question, we introduce RepoMirage, a two-stage evaluation suite built on SWE-Bench Verified that adopts perturbation as a diagnostic tool to increase the demand for context reasoning by transforming how the repository is exposed. First, RepoMirage-Perturb applies three types of semantics-preserving repository-level perturbations, revealing a clear performance drop when correct solving requires broader context access. RepoMirage-Extend further turns perturbation-targeted structural bottlenecks into explicit tasks beyond issue resolution, where the average performance declines from 66.8\% in the original setting to 25.3\%, indicating a significant deficiency in repository context reasoning. Further trajectory analysis reveals an exploration drift, where agents access broader repository context but fail to turn it into effective structure information. Motivated by this observation, we propose RepoAnchor, a structure-first prototype workflow that separates repository exploration from downstream problem solving, and show that explicit structural scaffolding yields notable gains. These results uncover an previously overlooked gap in repository context reasoning for code agents and suggest that stronger structure-aware methods are potential to improve them.


Repurposing Video Diffusion Transformers for Cross-View Temporal Object Correspondence

Siyoon Jin ⋅ Dahyun Chung ⋅ Honggyu An ⋅ Sangbeom Lim ⋅ Wonjin Nam ⋅ Jaewoo Jung ⋅ WonJun Moon ⋅ Seungryong Kim

Cross-view object correspondence identifies the same object across views. However, existing methods treat this as a frame-level matching problem, necessitating a query mask at every frame. Furthermore, without modeling temporal context, they suffer from severe ambiguity under appearance changes or occlusions. Therefore, we introduce cross-view temporal object correspondence, which requires preserving object identity across both views and time from a query mask. To address this, we propose Track, a streamlined framework which repurposes a pretrained video DiT with semantic and spatiotemporal priors for RGB-to-mask latent transport, jointly modeling cross-view matching and temporal tracking. However, naive video DiTs do not guarantee accurate cross-view grounding and often attend to semantically similar yet incorrect objects across views. Thus, we introduce focal cross-view attention alignment, hard negative conditioning, and local context crop strategies to ensure reliable grounding. Finally, beyond these architectural enhancements, we also tackle the lack of dense annotation by introducing a pseudo-label curation pipeline that converts sparse Ego-Exo4D annotations into dense supervision to enable effective video-level training. Our highly competitive performances validate the effectiveness of Track as a streamlined framework, ensuring robust temporal consistency and identity preservation. Code and weights will be released.

In a 100-body robot co-design benchmark, the height threshold that resets a falling robot changes which bodies appear optimal. We train PPO controllers for DERL-evolved UNIMAL morphologies under Standard (τ = 0.20), Strict (τ = 0.50), and Reset-free (τ = 0) protocols, then evaluate every policy under a common protocol. Strict and Reset-free training share only 5 of their top-10 bodies and have low rank agreement (Spearman ρ = 0.21). The effect is family-specific, not a uniform penalty: averaging per-body retention ratios, Reset-free training reduces Floor bodies' per-step reward by 45% relative to Standard training but reduces variable-terrain bodies by 16%. Under Strict, the losses are 22% for Floor and 34% for variable-terrain bodies, while multi-task bodies are most stable across protocols. Blocked ANOVA and morphology-level permutation tests support the protocol-by-family interaction. A SAC validation panel planned before analysis finds the same Strict-versus-Reset-free family ordering under a different optimizer. Selecting bodies by their minimum reward across the three training protocols recovers ≥91% of every single-protocol oracle's top-k average reward (k ∈ {5, 10, 20}; all bootstrap lower bounds ≥84.3%). Co-design benchmarks should state the reset rule, audit ranking sensitivity, and report protocol-robust selections when choosing bodies.


ResFusion: Medical Image Fusion Driven by Implicit-Forward Diffusion and Time-aware Joint Optimization

Nannan Wang ⋅ Dawei Zhou ⋅ Zhengkun Yu ⋅ Lei Hu ⋅ Lecong Xiong ⋅ Zaiyi Liu ⋅ Nannan Wang ⋅ Xinbo Gao

Standard training of diffusion models typically relies on a forward diffusion process anchored by ground-truth target images. However, in image fusion tasks, the scarcity of real fused images makes it difficult to formulate a forward process, thereby precluding the standard training. Fortunately, our visual analysis reveals that image fusion can be effectively modeled without standard training or reliance on the real images. Nevertheless, the misalignment of input image pairs remains a significant bottleneck for fusion quality. Recognizing that one-off registration is ill-suited to the progressive generation of diffusion models, we propose a unified registration-fusion framework driven by a time-aware joint optimization mechanism. Specifically, we design a time-aware registration network to progressively optimize registration to guide fusion, while the fusion results provide feedback constraints to the registration network at each time step. This mechanism facilitates joint optimization of registration and fusion, significantly improving the quality of fused results. Experimental results show that the proposed method performs excellently on multiple datasets, validating its effectiveness and superiority.


Rest-Tuning: Data-Efficient Adaptation of EEG Foundation Models to Individuals via Resting-State Signals

Yinuo Zhang ⋅ Xinyu Fu ⋅ Qing Wang ⋅ Yinte Zhang ⋅ Junjie Yu ⋅ Kexin Lou ⋅ Quanying Liu

EEG foundation models (FMs) offer strong population-level priors but degrade on individual subjects due to anatomy-driven distribution shifts. We propose Rest-Tuning, a task-agnostic calibration framework that personalizes a population FM using just 1–3 minutes of unlabeled resting-state EEG. Rest-Tuning employs masked teacher-student alignment to produce a subject-calibrated backbone, reused as initialization across multiple tasks from the same subject. Across two multi-subject datasets and four EEG FMs, Rest-Tuning outperforms task-only adaptation, with up to 13.3% absolute gain in downstream task accuracy and over 40% average reduction in labeled task data. Calibration saturates quickly and requires under 3 minutes of compute per subject. Representation analyses show reduced subject-to-population distance and smoother LoRA adapter interpolation paths, indicating improved compatibility between rest- and task-adapted solutions. These results establish resting-state EEG as a practical, fast, and reusable calibration interface, enabling population EEG FMs to function as reliable, subject-specific BCIs and serve as reusable initialization for downstream task adaptation.


Rethinking CT Synthesis through Semantics-Structure Alignment

Riyu Qiu ⋅ Qichao Zhou ⋅ jiacheng wang ⋅ Liansheng Wang

Mask-guided CT synthesis bridges structural annotations with imaging appearance, facilitating data augmentation and clinical tasks such as radiotherapy planning. Despite plausible visual realism, current methods seek optimal solutions by treating all tissues uniformly, failing to capture intrinsic CT properties such as tissue-specific distributions and strict anatomy, leading to semantic misalignment and structural hallucinations. To address this, we decompose CT synthesis into semantics-structure alignment through the proposed Semantics and Structure Loss. Specifically, the Semantic Intensity Distribution (SID) Loss partitions the broad Hounsfield Unit (HU) range into semantic subspaces to achieve tissue-specific intensity alignment. Building on SID-based semantic alignment, the Structural Anatomy (SA) Loss further improves structural fidelity by regularizing gradient manifolds within clinically-informed spatial domains. Extensive experiments across GAN, Diffusion, Flow, and foundation models demonstrate the broad compatibility and consistent gains of our approach. A user study further confirms the clinical realism and mask fidelity, while downstream segmentation demonstrates the practical utility. Code will be released after acceptance.

Data assimilation is the process of estimating the state of a dynamical system over time by combining model predictions with measurements. This task becomes challenging when the system is nonlinear and high-dimensional. To address this, score-based Bayesian filters have recently emerged. However, these methods still show unsatisfactory performance in certain cases, particularly under spatially sparse measurements. Such degradation stems from heuristic approximations of the likelihood score, whose errors can accumulate over time. This limitation arises because the methods simply adopt a classical forward process for generative modeling that transforms a data distribution toward a Gaussian distribution, which is independent of the measurement equation. Here, we propose a forward process tailored for filtering that transforms the system state toward the measurement space, enabling a theoretically sound formulation of the likelihood score. Based on this, we develop the Measurement-Aware Score-Based Filter (MASF). We evaluate MASF on the Kolmogorov flow, a high-dimensional fluid benchmark with up to $\mathcal{O}(10^5)$ dimensions, under diverse measurement operators, including nonlinear cases with dimensional mismatch between the state and the measurements. MASF shows improved performance over existing score-based filters and ensemble-type Kalman filters. With amortized pretraining, MASF also achieves up to a $28.2\times$ wall-clock speedup compared with the baselines.


Rethinking Knowledge Distillation for Diffusion Language Models

Yuxuan Sun ⋅ Yuanjian Xu ⋅ Jianing Hao ⋅ Yu Li ⋅ Sowmen Das ⋅ Zhong Li ⋅ Sangwoong Yoon ⋅ Miguel Rodrigues

Discrete diffusion language models (DLMs) offer a promising non-autoregressive alternative to large language models by enabling parallel generation, yet their performance still lags behind autoregressive counterparts. Knowledge distillation has recently proven highly effective for improving autoregressive models, but its applicability to DLMs remains poorly understood. We present the first systematic study of KD for DLMs and uncover two non-intuitive failures of classical forward-KL (FKL) distillation: (i) student perplexity does not improve monotonically with more teacher-generated data, and (ii) on the same teacher-generated corpus, FKL distillation underperforms simply training the student with the original DLM denoising objective. These failures hold for both a 110M MDLM teacher and a 7B Dream teacher. To address them, we propose a unified, data-free distillation framework for masked DLMs. The framework first runs an off-policy phase that absorbs general knowledge from a static set of teacher samples and then transitions to an on-policy phase whose objective combines: a trajectory loss along the student's own reverse process, confidence reweighting that down-weights tokens where the teacher is unsure, and a trust-region anchor that stabilizes the moving target. For the smaller MDLM regime where FKL alone is insufficient, we additionally adopt the $\alpha$-$\beta$ divergence. On MDLM, the resulting student attains a GPT-2 perplexity of $24.75$, surpassing the teacher's $27.04$; on Dream-7B distillation to a 93M student model, the framework reduces Qwen-2.5 perplexity from the FKL-distilled baseline of $25.23$ to $19.83$ while preserving zero-shot accuracy.


Rethinking LLM Fine-Tuning via Weight Space Reparameterization: Preserving Safety during Downstream Adaptation

Min-Seong Kim ⋅ Jongbok Won ⋅ Jeesup Park ⋅ Yonghee Choi ⋅ Dong-Jun Han

Fine-tuning large language models (LLMs) on downstream tasks often weakens their previously aligned safety behavior, creating an inherent trade-off between task adaptation and safety preservation. Existing approaches attempt to mitigate this issue by identifying safety-relevant neurons, layers, or update directions in the original parameter space. However, safety-related information is often distributed across many directions in this space, causing downstream updates to overwrite parameters that encode safety behavior. We propose WSR-Tune, a Weight Space Reparameterization-based fine-tuning framework that preserves safety during downstream adaptation. Our approach constructs a safety-conditioned basis and reparameterizes the weight matrices such that safety-relevant information becomes concentrated in a small subset of directions. We then identify safety-critical directions and freeze them in the reparameterized space, while updating the remaining complementary directions for downstream adaptation. By explicitly structuring the parameter space to disentangle safety and task information, our method reduces interference between objectives and enables simultaneous preservation of safety and improvement in downstream performance. Experimental results show that our method achieves a more favorable trade-off between downstream performance and safety retention, demonstrating its effectiveness for reliable LLM fine-tuning.


Rethinking the Mixture of Vision Encoders Paradigm for Enhanced Visual Understanding in Multimodal LLMs

Mozhgan Azadani ⋅ James Riddell ⋅ Sean Sedwards ⋅ Krzysztof Czarnecki

Mixture of Vision Encoders (MoVE) has emerged as a powerful approach to enhance the fine-grained visual understanding of multimodal large language models (MLLMs), improving their ability to handle tasks such as complex optical character recognition and scene understanding. Despite these advances, effectively combining diverse encoders and their visual tokens, while also scaling to high-resolution inputs, remains an open challenge. In this work, we conduct a systematic study of fusion designs for MoVE-based MLLMs, highlighting principles for token-level integration across complementary encoders. Our study shows that a lightweight recipe consisting of post-adaptation fusion with independent projectors, tile-level sequence interleaving, and dynamic tiling with global context delivers strong performance on diverse benchmarks. We integrate these principles into a simple and effective architecture that we call LEO. Extensive evaluation on 11 vision–language benchmarks demonstrates that LEO achieves better results on the majority of tasks compared to existing MoVE-based approaches. Furthermore, LEO adapts effectively to the specialized domain of autonomous driving without altering its architecture or training recipe, achieving competitive performance against established baselines and thereby highlighting its ability to generalize. The code is available at https://github.com/Mozhgan91/LEO.


Rethinking Time Series Tokenization from a Frequency Perspective

Zelong Tian ⋅ Ming Jin ⋅ Bo Du ⋅ Shirui Pan

Tokenization that partitions time series into subsequences has become a foundational paradigm in modern time series modeling. However, we find that current tokenization strategies inherently introduce spectral distortions, forcing the learned representations to diverge from the true underlying patterns such as trend, periodicity, and seasonality, thereby impairing model performance. Through theoretical analysis from a frequency perspective, we derive the boundary conditions under which such distortions occur. To overcome this distortion, we propose TimeFT, a nearly distortion-free frequency-based tokenizer for time series data. TimeFT obtains tokens through frequency-domain partitioning, frequency shifting, and Nyquist sampling. The method is parameter-free, incurs negligible computational complexity, and can serve as a drop-in replacement for existing tokenizers. Our extensive experiments on diverse tasks such as forecasting, classification, and anomaly detection demonstrate that TimeFT consistently and significantly improves performance.

In this work, we study portfolio optimization under the stochastic discount factor (SDF) framework by learning market state representations that capture the underlying risk structures of financial data. This is challenging due to several factors: financial markets exhibit non-stationary dynamics with shifting regimes, multimodal inputs such as price and news data often contain stochastic noise, and existing diffusion-based approaches, while effective for modeling stochastic dynamics, rely on assumptions such as isotropic Gaussian noise that fail to capture the state-dependent nature of financial uncertainty. To address these challenges, we introduce RADAR, a retrieval-augmented diffusion framework that learns market representations by conditioning on similar historical regimes. RADAR leverages retrieval to construct context-dependent noise distributions, applies conditional diffusion to denoise multimodal representations, and initializes the diffusion process using empirical statistics to reflect state-dependent uncertainty. Experiments show that RADAR achieves state-of-the-art performance on key risk-adjusted metrics while producing economically meaningful signals on asset returns and correlations.


Retrieval-Centric Deep Learning in Growing Nonparametric Neural Networks

Maximilian Schlegel ⋅ Rajai Nasser ⋅ Seijin Kobayashi ⋅ Yanick Schimpf ⋅ Oliver Sieberling ⋅ Robert Obryk ⋅ Kazuki Irie ⋅ João Sacramento ⋅ Johannes von Oswald

We investigate a general-purpose layer for deep learning that, instead of compressing arbitrary-size training data into fixed-size weight matrices, stores a new pair of key-value representations for every data point during training, and retrieves and recombines these representations through an attention mechanism at inference time---resulting in a growing neural net (NN). We derive such a "retrieval-centric deep learning" (RCDL) paradigm from the classic duality of a linear layer trained by gradient descent as a type of linear attention (LA) over the training data points---which motivates us to replace the corresponding LA by a more powerful form of attention from recent work on sequence models, namely, kernelized attention using radial basis function (RBF) and softmax-like kernels, as well as more advanced linear attention variants such as MesaNet and DeltaNet. We theoretically derive learning algorithms for the growing layers with kernelized attention, and empirically demonstrate their promising performance and sample-efficiency on classic image classification tasks and synthetic teacher-student learning datasets. In the MesaNet/DeltaNet-inspired extensions, we show a formal connection to a recently proposed optimizer and derive highly sample-efficient optimizers for conventional fixed-size NNs. Despite open challenges, RCDL represents a promising paradigm for nonparametric machine learning.


Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

Chanho Park ⋅ Daehyeon Choi ⋅ Jihyun Lee ⋅ Minhyuk Sung

Vision-language models (VLMs) can locate an image region referred to by a text prompt and route the corresponding visual evidence to the output, yet the internal mechanism behind this behavior is not understood. Inspired by retrieval heads in large language models, we ask whether VLMs contain an analogous mechanism for visual retrieval. We answer affirmatively by introducing Visual Retrieval Heads(VRHs), a small subset of attention heads (about 1.7–2.6%) that are causally responsible for grounding text descriptions to image regions. To find them, we recast existing head-scoring methods under a unified design space over query tokens, key aggregation, and cross-sample aggregation. We then show that scoring attention from output prediction tokens with a sum over the ground-truth referent region most reliably identifies causal heads. Across four VLMs and five referring-expression benchmarks, masking only the top 20 VRHs reduces grounding accuracy by up to 80 percentage points, while masking the same number of random heads has little effect. Beyond replicating the causal-sparse-universal triad established for text retrieval heads, VRHs exhibit several properties not previously reported: they generalize across visual reference tasks, remaining causal on attribute, spatial, counting, and visual-math benchmarks despite being discovered through bounding-box prediction; they are functionally specific, preserving output format while corrupting localization; and they are architecturally shared, transferring causally across VLMs that share an LLM backbone but differ in vision encoder, projector, and instruction tuning


Retrieve What’s Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation

Minseok Joo ⋅ Dogyun Park ⋅ Taehoon Lee ⋅ Kyujin Lee ⋅ Hyunwoo J. Kim

Maintaining long-term geometric consistency remains a key challenge in long-horizon autoregressive video generation. Recent memory-augmented generative models address this issue by retrieving historical frames beyond the limited temporal context, but their effectiveness depends on two key design choices: what 3D-geometric evidence should represent past observations, and how memory frames should be selected from this evidence. Existing methods often rely on camera poses or field-of-view overlap, which are lightweight but too coarse to reason about pixel-wise visibility, or use explicit 3D reconstruction, which provides fine-grained evidence but is costly to maintain over long rollouts. We propose Coverage-Maximizing Retrieval-Augmented Generation (COVRAG), a depth-based memory retrieval framework that uses pretrained 3D priors to construct a target-view coverage map as lightweight 3D memory evidence. For frame selection, COVRAG maximizes residual coverage gain, iteratively retrieving frames that explain target-view regions not covered by the current context or previously selected memories. To improve scalability in long-video generation, we introduce a sliding-window depth caching for efficient geometry estimation. Experiments on RealEstate10K and DL3DV show that \modelname improves long-horizon geometric consistency while maintaining low latency compared to baselines.

High-quality teleoperation datasets are costly to collect, particularly for hard tasks. We observe that many tasks exhibit directional asymmetry: completing the forward hard task is difficult, whereas reversing it by relaxing or disrupting the environment is comparatively easy. This suggests that reversed easy-task trajectories can serve as a scalable supervision signal for the hard task, reducing the cost of manual demonstration collection. However, reversed data can be noisy, and directly training on it may yield suboptimal policies. To enable largely automated acquisition and effective use of reversed data, we propose a teleoperation-cost effective framework for hard policy learning via temporal reversal of easy tasks, consisting of three key components: a closed-loop data collection pipeline that alternates between hard-task and easy-task policies to autonomously reset the environment and generate diverse trajectories; a hierarchical data refinement pipeline that temporally inverts easy-task rollouts and filters low-quality motion using kinematic priors and a critic-guided advantage filter; and an iterative policy learning method that trains the hard-task policy using both initial reversed easy-task demonstrations and the filtered reversed data in a continuous online learning loop. By combining automated collection, hierarchical refinement, and iterative learning, our method enables scalable, reliable training of complex, high-precision manipulation tasks. Across two simulated benchmarks and real-robot experiments, we demonstrate that our method improves hard-task success rates with higher data efficiency and more stable training compared to reversal-based and reinforcement-learning baselines, without requiring extensive hard-task teleoperation.


Revisiting Gradient Ascent: Machine Unlearning from a Geometric Perspective for Source-Free Scenarios

Yufeng Cao ⋅ Naen Xu ⋅ Xuyang Teng ⋅ Tianyu Du ⋅ Qiang Zhao

Machine unlearning is becoming increasingly indispensable for meeting ever-stringent data compliance requirements, such as the ``right to be forgotten''. While Gradient Ascent (GA) based approximate methods have drawn attention for their conceptual simplicity, their practical deployment is severely constrained by three fundamental issues: first, the lack of feasibility analysis leads to divergent update directions; second, gradient conflict inevitably degrades the performance maintenance on retained data; and third, heavy reliance on the retain set restricts their applicability in real-world settings. To address these challenges, we conduct an in-depth investigation from a geometric perspective into the underlying root causes of instability in conventional GA methods. Our analysis reveals that the feasibility of unlearning intrinsically depends on specific geometric constraint relationships within the parameter space. Building upon this theoretical insight, we formally establish the methods governing unlearning feasibility. The first is the direction feasibility, which constrains the feasible update trajectory and keeps the model within reasonable parameter regions during unlearning. The second is the step-size feasibility, which further restricts the update magnitude along the feasible direction to prevent model collapse caused by large steps.Notably, we find that the direction feasibility does not rely on any information from the retain set. In light of this, we further explore viable approaches for effectively estimating the step-size feasibility under strict source-free scenarios.


Revisiting Incremental Learning: A Three-Interface Diagnosis of Stability and Plasticity

Keuntae Kim ⋅ Beomseok Lee ⋅ Hyunwoo Kim ⋅ Yong Suk Choi

Incremental learning is usually evaluated through the final classifier or LM head, but this interface does not indicate where forgetting occurs. A drop in accuracy may reflect representation degradation in the backbone, readout mismatch in the head, or both. We argue that BWT and FWT should therefore be interpreted as interface-dependent quantities rather than as single model-level properties. We study three evaluation interfaces: the sequentially trained head, a newly optimized probing classifier, and a classifier-free backbone separability diagnostic based on clustering. These interfaces answer different questions: deployed performance, recoverable information under a new readout, and representation-level separability. Across CIL experiments with encoder and decoder language backbones, plus discriminative-backbone TIL sanity checks, we find that qualitative conclusions about forgetting and plasticity change substantially with the evaluation interface. Standard heads often indicate severe forgetting, probes often suggest near-complete recoverability, and backbone diagnostics suggest intermediate representation changes and a stability--plasticity pattern. This pattern is consistent across five clustering algorithms and three clustering metrics, suggesting that it is not an artifact of a single diagnostic choice. Finally, we use Just LM-Head Tuning (JLT) as a head-only intervention to quantify recoverable performance on a fixed incrementally trained backbone. The large gap between standard and head-realigned performance suggests that many failures attributed to catastrophic forgetting are better understood as head--backbone mismatch plus partial representation drift. Our results call for reporting where forgetting is measured, not only how much forgetting is observed.


Reward-Estimated Hypergradient for Bilevel Reinforcement Learning with Black-Box Follower

Shigeki Kusaka ⋅ Mikoto Kudo ⋅ Takumi Tanabe ⋅ Akifumi Wachi ⋅ Youhei Akimoto

Optimizing the leader's policy via hypergradients (HG) in bilevel reinforcement learning (RL) typically assumes a white-box follower. To enable real-world applications, we address the black-box follower setting, where the leader must optimize solely from observed trajectories without access to the follower's true reward function. We propose Reward-Estimated Hypergradients (RE-HG), which leverages Inverse RL (IRL) to recover this unobserved reward. Because naive IRL introduces severe reward-shaping biases and numerical instabilities, RE-HG introduces a theoretical bias-cancellation mechanism that aggregates information across diverse environmental dynamics. Together with eigenvalue truncation, this acts as implicit regularization to suppress HG norm explosions. Empirical evaluations demonstrate that RE-HG estimates gradients aligned with a white-box Oracle, achieving comparable or superior performance. Notably, in sharp reward landscapes where the Oracle becomes trapped in suboptimal local minima, RE-HG's implicit regularization extracts stable global gradients, demonstrating robustness.


RewardFlow: Topology-Aware Reward Propagation on State Graphs for Agentic RL with LLMs

Xiao Feng ⋅ Bo Han ⋅ Zhanke Zhou ⋅ Jiaqi Fan ⋅ Jiangchao Yao ⋅ Ka H Li ⋅ Dahai Yu ⋅ Michael Ng

Reinforcement learning (RL) shows promise for enhancing LLM agentic reasoning capabilities with external environments. However, sparse terminal rewards hinder fine-grained, state-level optimization. Although process reward modeling offers a promising alternative, training dedicated reward models often entails substantial computational costs, risks reward hacking, and faces annotation bottlenecks. To address these challenges, we introduce RewardFlow, a lightweight method for estimating state-level rewards for agentic reasoning. RewardFlow leverages the intrinsic topological structure of states within trajectories by constructing state graphs. This enables topology-aware graph propagation to estimate state-wise contributions to success, yielding principled, annotation-free state-level rewards. As dense rewards for RL optimization, RewardFlow substantially outperforms prior RL baselines across four agentic benchmarks, with average success-rate gains of +6.2\% on text-based benchmarks and +29.7\% on visual reasoning over the strongest baseline across three model scales, and over +10\% accuracy improvement on DeepResearch, while demonstrating superior robustness and training efficiency.


Rigel3D: Rig-aware Latents for Animation-Ready 3D Asset Generation

Nikitas Chatzis ⋅ Marios Loizou ⋅ Evangelos Kalogerakis

Recent 3D generative models can synthesize high-quality assets, but their outputs are typically static: they lack the skeletal rigs, joint hierarchies, and skinning weights required for animation. This limits their use in games, film, simulation, virtual agents, and embodied AI, where assets must not only look plausible but also move plausibly. We introduce Rigel3D, a generative method for animation-ready 3D assets represented as rigged meshes. Unlike post-hoc auto-rigging methods that attach rigs to completed shapes, our method jointly models geometry and rig structure through coupled surface and skeleton structured latent representations. A rig-aware autoencoder decodes these representations into mesh geometry, skeleton topology, joint coordinates, and skinning weights, while a two-stage latent generative model synthesizes both surface and skeleton representations for image-conditioned generation. To support downstream animation workflows, we further introduce an open-vocabulary joint labeling module that embeds generated joints into a shared vision-language space, enabling correspondence to arbitrary retargeting templates. Experiments on large-scale rigged asset datasets demonstrate that our method generates diverse, high-quality animation-ready assets and outperforms existing rigging baselines across multiple metrics.


RL-Guided Temporal Localization for Dual-Channel Retrieval in Long-Horizon Agent Memory

Dabin Sheng ⋅ Zhe Wu ⋅ Jinming Zhao ⋅ Junliang Xing ⋅ Yuanchun Shi

Time-aware retrieval is crucial for long-horizon LLM agents to effectively leverage ever-growing long-term memory. Existing methods focus on lexical or semantic similarity, leading to target memories being crowded out by temporally mismatched yet semantically similar candidates. Even time-aware approaches inject temporal cues only during memory organization or answer generation, overlooking temporal constraints at the retrieval stage, where candidate selection occurs—the root cause of the displacement. Inspired by the temporal contiguity effect in human episodic memory, we propose TIDER (Temporal Interval-driven Dual-channel Evidence Retrieval), a time-aware agentic retrieval framework. Central to TIDER is a reinforcement learning (RL)-trained temporal interval locator that infers the temporal windows in which the evidence lies. RL optimizes window localization as a discrete, non-differentiable retrieval decision, using composite rewards to convert sparse answer supervision into fine-grained retrieval feedback. These temporal windows first define the temporal retrieval channel; the retrieved candidates are then merged with results from the semantic channel and reranked by proximity to the windows. This temporal grounding enables TIDER to retrieve temporally aligned evidence that would otherwise be drowned out by the volume of semantically similar yet temporally mismatched memories. With a lightweight 1.7B locator, TIDER achieves state-of-the-art accuracy on both long-horizon memory benchmarks—61.52% on LifeBench (+8.82%) and 78.96% on LoCoMo (+3.66%), which indicates that temporal cues are crucial for pinpointing target evidence amid a vast pool of semantically similar historical events in long-term daily-life memory.


RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models

Hao Wu ⋅ Yuqi Li ⋅ Yuan Gao ⋅ Fan Xu ⋅ Fan Zhang ⋅ Kun Wang ⋅ Penghao Zhao ⋅ Qiufeng Wang ⋅ Yizhou Zhao ⋅ Weiyan Wang ⋅ YingLi Tian ⋅ Xian Wu ⋅ Xiaomeng Huang

Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly aligned with the capabilities that matter most for robot decision making, including instruction following, manipulation success, and physical plausibility. They also suffer from error accumulation in long-horizon autoregressive prediction. We present RoboAlign-R1, a framework that combines reward-aligned post-training with stabilized long-horizon inference for robot video world models. We construct RobotWorldBench, a benchmark of 10,000 annotated video-instruction pairs collected from four robot data sources, and train a multimodal teacher judge, RoboAlign-Judge, to provide fine-grained six-dimensional evaluation of generated videos. We then distill the teacher into a lightweight student reward model for efficient reinforcement-learning-based post-training. To reduce long-horizon rollout drift, we further introduce Sliding Window Re-encoding (SWR), a training-free inference strategy that periodically refreshes the generation context. Under our in-domain evaluation protocol, RoboAlign-R1 improves the aggregate six-dimension score by 10.1% over the strongest baseline, including gains of 7.5% on Manipulation Accuracy and 4.6% on Instruction Following; these ranking improvements are further supported by an external VLM-based cross-check and a blinded human study. Meanwhile, SWR improves long-horizon prediction quality with only about 1% additional latency, yielding a 2.8% gain in SSIM and a 9.8% reduction in LPIPS. Together, these results show that reward-aligned post-training and stabilized long-horizon decoding improve task consistency, physical realism, and long-horizon prediction quality in robot video world models.


RoboProcessBench: Benchmarking Process-Aware Understanding in Visual-Language Robotic Manipulation

Dayu Xia ⋅ Yue Shi ⋅ Yao Mu ⋅ Huiting Ji ⋅ Chaofan Ma ⋅ Yingjie Zhou ⋅ Hua Chen ⋅ Yang Liu ⋅ Jiezhang Cao ⋅ Guangtao Zhai

Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These roles implicitly require models to judge not only final task success, but also how a manipulation execution is physically and temporally progressing. However, existing evaluations fail to test whether VLMs possess fine-grained process understanding. To address this gap, we present RoboProcessBench, a benchmark for process-aware understanding in vision-language robotic manipulation. RoboProcessBench decomposes such capability into two complementary dimensions, $\textit{static monitoring}$ and $\textit{dynamic reasoning}$, instantiated as 12 diagnostic question families covering phase, contact, motion, coordination, primitive-local progress, temporal order, outcome, and primitive-level transitions. Built from physically grounded execution traces, the curated benchmark corpus ProcessData contains ~58k question-answer pairs across 260 manipulation tasks, which is further split into ProcessData-SFT and ProcessData-Eval for post-training and evaluation purposes. Extensive evaluation of various VLMs on ProcessData-Eval reveals broad limitations across 12 diagnostic task families, suggesting current models still lack robust process-aware understanding of manipulation executions. But with ProcessData-SFT, the post-trained $\textit{Qwen2.5-VL-7B}$ and $\textit{InternVL-3-8B}$ exhibit consistent gains on local state, motion, progress, and primitive-aware cues. These results demonstrate that RoboProcessBench serves as both an evaluation benchmark and a learnable supervision source for developing VLMs capable of monitoring and evaluating robotic manipulation processes. Dataset artifacts (https://huggingface.co/datasets/ProcessBench-2026/RoboProcessBench) and executable evaluation suite (https://anonymous.4open.science/r/RoboProcessBench-F875/) are released for reproducibility.

Continual model merging addresses a more realistic setting in which task-specific models arrive sequentially. Naive extension of standard model merging approaches to the continual setting leads to task balance collapse and substantial performance degradation. To address this issue, we propose Global Singular Subspace Separation and Restoration (GS3R), a data-free and optimization-compatible framework for continual model merging. GS3R preserves task balance by globally separating singular components and introducing vector restoration matrices for perturbation-free recovery. To further enhance efficiency, we introduce Pre-merge Restoration and Flushing (PRF), which guarantees peak memory usage comparable to that of other methods. Experiments on vision and language tasks demonstrate that GS3R consistently outperforms prior continual model merging methods.


Robust Offline Reinforcement Learning against Out-of-Distribution Dynamics in Autonomous Driving

Haozhe Liu ⋅ Jie Wang ⋅ Rui Yang ⋅ Chunyang Liu ⋅ Xize Liang

Real-world autonomous driving data inherently exhibits out-of-distribution (OOD) dynamics, as diverse interactions and unpredictable human behaviors cause vehicle trajectory distributions to vary significantly across different scenarios, even for similar road structures. Such OOD dynamics typically induce uncontrollable extrapolation errors in neural networks. We observe that in offline reinforcement learning (RL), these errors manifest as heavy-tailed estimates of action values with significant biases, which existing robust methods fail to address, hindering effective generalization. To address this problem, we propose a robusT offline RL algorithm under heavy-tailed action-valUe eSTimates (TRUST), which models perturbed states associated with heavy-tailed action-value estimates to effectively mitigate such estimation biases. Specifically, TRUST uses a first-order gradient-regularized target to model perturbed states that approximate data under worst-case dynamics. It then introduces a statistical measure to identify perturbations that induce heavy-tailed action-value estimates. By re-weighting the training objective to emphasize correcting these biased estimates, TRUST can reduce the heavy-tailed behavior in action values for robustness against OOD driving data. We further derive an error bound that justifies the gradient-regularized target as an approximation to the robust Bellman target. Extensive experiments on three large-scale real-world datasets show that TRUST achieves strong overall performance across all benchmarks, with a general performance improvement of 41.21\%.


Robust Satisficing Ensemble: Scalable Model Aggregation Under Distribution Shifts

Ahmet Faruk Cetinkaya ⋅ Enes Ağırman ⋅ Cem Tekin

Standard ensemble methods can improve robustness by mitigating individual model stochasticity, but they frequently fail to generalize beyond the training distribution, leaving systems vulnerable to out-of-distribution (OOD) data. We introduce the Robust Satisficing Ensemble (RSE), a scalable probabilistic aggregation framework designed to secure model combinations by minimizing their fragility against distributional changes. By bridging a robust satisficing objective with a quadratic risk surrogate, RSE derives a closed-form analytical inner adversary. This fundamental reformulation enables highly efficient weight updates without the computational bottleneck of standard minimax optimization, crucially decoupling optimization complexity from sample size to scale effectively to large datasets. In particular, we prove a geometrically decaying truncation error for the fragility approximation via Lagrange-Bürmann expansion and show that the proximal RSE updates converge to a stationary point. We evaluate RSE across a comprehensive spectrum of distribution shifts, including adversarial label shift (SST-5, TREC), natural subpopulation and temporal shifts (CivilComments, HuffPost). Our results demonstrate that RSE consistently outperforms strong robust baselines, including GroupDRO, IRM, and SRM. These findings establish RSE as a practical, theoretically grounded safeguard for reliable model aggregation in dynamic environments.


RotMoLE: Enhancing Mixture of Low-Rank Experts through Rotational Gating Mechanism

Mengyang Sun ⋅ MaoChuan Dou ⋅ Tao Feng ⋅ Dan Zhang ⋅ Yihao Wang ⋅ Junpeng Liu ⋅ Yifan Zhu ⋅ Jie Tang

While Large Language Models (LLMs) are commonly fine-tuned to handle domain-specific tasks before being applied to vertical applications, adapting them to complex scenarios with diverse specialized knowledge remains challenging. Meanwhile, Mixture-of-Experts (MoE) architecture has risen as a crucial paradigm for training LLMs, and some recent works have also incorporated MoE into Parameter-Efficient Fine-Tuning (PEFT) to propose the Mixture of Low-rank Experts (MoE-LoRA), to enhance the power of low-rank adapters for learning complicated knowledge. However, conventional gating mechanisms in MoE typically apply only a scalar reweighing to selected experts, thereby limiting their underlying capacity of representation and generalization. Motivated and enabled by the low-rank structures in MoE-LoRA, we propose RotMoLE, a specialized MoE framework for low-rank experts featuring an additional rotation gate. Beyond simple scaling, RotMoLE implements a rotation mechanism for each selected expert, enabling superior expert exploitation and specialization for learning diverse data, especially when expert candidates are limited. Empirical results on complex multi-task and multilingual training scenarios validate the effectiveness of our methodology.


RoutingBench: Can Agentic Routing Analysis Scale to Production Datacenter Networks?

Wenlong Ding ⋅ Zhixiong Niu ⋅ Jianan Yang ⋅ Fajun Zhang ⋅ Bo Zhang ⋅ Ling Liang ⋅ Yongqiang Xiong ⋅ Tianyin Xu ⋅ Hong Xu

Recent advances in AI models and agentic technologies make AI for network operations (NetOps) within reach. However, scalability remains a key bottleneck of agentic NetOps when analyzing hyperscale networks, which comprise hundreds of datacenters, each housing thousands of network devices. The scalability challenge is rooted in the requirement of many NetOps tasks that must conduct global reasoning on how a local change of device behavior affects all relevant routing paths, known as routing-path analysis. This paper studies this scalability problem and evaluates how different agentic approaches, namely in-context learning, iterative reasoning, and agent skills, can scale routing-path analysis to large, complex networks. We present RoutingBench for evaluating agentic routing-path analysis, with varying network size and complexity, for various types of device changes. Our results show that agentic analysis is promising—agent skills curated with a principle termed “explore more; digest less” enables routing-path analysis on hyperscale networks of 50K routers with an accuracy of 99.5%, significantly outreaching the scalability of traditional symbolic analysis. Meanwhile, RoutingBench also reveals the boundary of AI agent capability on complex inter-datacenter networks and compound changes, posing open challenges for AI and agentic research.


S2D: Sparse-To-Dense Keymask Distillation for Unsupervised Video Instance Segmentation

Leon Sick ⋅ Lukas Hoyer ⋅ Dominik Engel ⋅ Pedro Hermosilla ⋅ Timo Ropinski

In recent years, the state-of-the-art in unsupervised video instance segmentation has heavily relied on synthetic video data, generated from object-centric image datasets such as ImageNet. However, video synthesis by artificially shifting and scaling image instance masks fails to accurately model realistic instance motion in videos, such as perspective changes, movement by parts of one or multiple instances, or camera motion. To tackle this issue, we propose an unsupervised video instance segmentation model trained exclusively on real video data. We start from unsupervised instance segmentation masks on individual video frames. However, single-frame segmentations exhibit temporal noise and their quality varies through the video. Therefore, we establish temporal coherence by identifying high-quality keymasks in the video by leveraging deep motion priors. The sparse keymask pseudo-annotations are then used to train a segmentation model for implicit mask propagation, for which we propose a Sparse-To-Dense Distillation approach aided by a Temporal DropLoss. After training the final model on the resulting dense labelset, our approach outperforms the current state-of-the-art across various benchmarks.


SAFE-DRIFT: Data Selection for Supervised Fine-tuning with Controllable Off-Target Drifts

Yeo Jin Jung ⋅ Yating Liu ⋅ Lalchand Pandia ⋅ Claire Donnat

We propose SAFE-DRIFT, a data-selection framework for supervised fine-tuning that explicitly trades off target improvement against unwanted off-target drifts of model behavior. We formalize this objective as a constrained optimization problem: maximize gain on the target gradient subject to a budget on off-target drift, which we quantify by the Fisher information on the reference distribution. The resulting closed-form solution is a damped natural gradient with respect to the reference Fisher matrix. We further address the statistical estimation challenge in the high-dimensional setting, deriving a scalable algorithm that approximates the Fisher information and the target gradient in a shared low-dimensional subspace. We evaluate SAFE-DRIFT against state-of-the-art data selection methods across diverse domains including coding and medical question answering, yielding competitive target performance while reducing drift on specified off-target behaviors.


Safety-Aware Latent Space Reasoning in Large Language Models

Yi Wang ⋅ Wenjie Wang ⋅ Hongye Qiu ⋅ Yu Pan

Latent space reasoning improves inference efficiency by compressing chain-of-thought (CoT) reasoning into continuous latent states, but its safety alignment remains poorly understood. In this work, we conduct the first systematic safety study of latent space reasoning models under jailbreak attacks. Our results reveal that most latent reasoning methods exhibit higher vulnerability to jailbreak attacks than both the base model and token-space CoT reasoning counterparts, suggesting that moving reasoning into latent states can weaken safety alignment. To address this issue, we propose SaLR, a Safety-aware Latent space Reasoning framework that injects safety supervision into latent space reasoning. SaLR converts long-form safety reasoning into a compact four-block safe-chain and transfers this safety signal through teacher-student hidden state distillation. For safety instances, SaLR aligns an early safe response prefix to guide the model toward a safe response trajectory; for reasoning instances, it retains single position distillation to preserve latent reasoning ability. Experiments across model scales and harmful benchmarks show that SaLR substantially reduces jailbreak attack success while preserving reasoning accuracy and token efficiency. Beyond standard jailbreaks, SaLR also mitigates reasoning specific attacks such as token-inflation and CoT-hijacking attacks by keeping reasoning in fixed latent states and avoiding exposed textual reasoning traces. Overall, SaLR provides a stronger safety utility trade-off for latent space reasoning. Our anonymous code repository is available at: https://anonymous.4open.science/r/SaLR-3C19


SageSched: Efficient LLM Scheduling Confronting Demand Uncertainty and Hybridity

Zhenghao Gan ⋅ Yichen Bao ⋅ Yifei Liu ⋅ Chen Chen ⋅ Quan Chen ⋅ Minyi Guo

Efficient LLM inference scheduling is crucial for the user experiences. However, LLM inferences exhibit remarkable demand uncertainty (with unknown output length beforehand) and hybridity (being both compute and memory intensive). Existing LLM schedulers rely on simple heuristics or focus purely on compute resource, suffering suboptimal performance. In this work, we propose SageSched, an efficient LLM scheduler that properly handles demand uncertainty and hybridity of inference workloads. SageSched combines prompt contents with the past inference results to predict output-length distribution in a light-weight and also accurate manner. Meanwhile, it models the true service cost of an inference request with both compute and memory aspects considered. Finally, SageSched employs a uncertainty-aware scheduling policy that can yield the best overall efficiency given the request cost distributions. Testbed experiments over diverse setups confirm that SageSched can attain an efficiency improvement of over 28.7%.


SAGE: Semantically Disentangled Representation Learning through Latent Geometry Constraint and Large Language Model

Qiuyu Chen ⋅ Liang Xu ⋅ Yunnan Wang ⋅ Mingqi Yuan ⋅ Qi Wang ⋅ Jiahao Li ⋅ Zhicheng Wang ⋅ Hao Zheng ⋅ Mulin Chen ⋅ Baao Xie ⋅ Xin Jin ⋅ Wenjun Zeng

Disentangled representation learning (DRL) aims to identify and decompose the interpretable underlying factors of observations. However, most DRL methods rely on independence-oriented regularization to pursue statistical factorization, leaving learned factors without explicit semantic grounding. Such independence-oriented regularization induces an intrinsic trade-off between disentanglement and reconstruction. To address these issues, we establish a spectral learnability threshold for independence-oriented DRL, theoretically revealing a mismatch between the factors favored by the learning objective and human-perceivable semantics. We further introduce SAGE, a semantic-guided DRL framework that leverages MLLMs to automatically extract semantic signals from unlabeled data and uses a triangular latent-geometry constraint to anchor each semantic factor to a designated latent dimension, enabling factor-wise control over human-perceivable semantic variations in generated images. Extensive experiments support our theoretical analysis and demonstrate that SAGE achieves state-of-the-art performance across multiple benchmarks. In addition, our framework exhibits high flexibility and can be integrated with most advanced generative models.


Same Words, Different Judgments: How Preferences Vary Across Modalities

Aaron Broukhim ⋅ Nadir Weibel ⋅ Eshin Jolly

Preference-based reinforcement learning (PbRL) is the dominant framework for aligning AI systems to human preferences. However, evaluation protocols for such data were designed for text and have not been validated for speech. We present the first ICC-based, controlled cross-modal study of human and synthetic preference annotations, comparing text and audio evaluations of identical semantic content across 100 prompts. We show that achieving $\textit{good}$ agreement within either modality (ICC(2,$k$) $\approx$ .80) requires $\sim$9 raters. At the same time, modalities show marked differences in how people report preferences: audio raters exhibit narrower decision thresholds, reduced length bias, and more user-oriented evaluation criteria, with near-chance cross-modality agreement. We demonstrate that synthetic ratings can be used to effectively predict inter-rater agreement, thus serving as an early signal for stimulus selection and proxy for human annotations. Together, these findings argue that evaluation protocols for audio preference data require modality-specific design rather than direct adaptation from text.


Sampling-Based Safe Reinforcement Learning

Luca Vignola ⋅ Bruce D Lee ⋅ Manish Prajapat ⋅ Manuel Wendl ⋅ Melanie Zeilinger ⋅ Andreas Krause ⋅ Yarden As

Safe exploration remains a fundamental challenge in reinforcement learning (RL), limiting the deployment of RL agents in the real world. We propose Sampling-Based Safe Reinforcement Learning (SBSRL), a model-based RL algorithm that maintains safety throughout the learning process by enforcing constraints jointly across a finite set of dynamics samples. This formulation approximates an intractable worst-case optimization over uncertain dynamics and enables practical safety guarantees in continuous domains. We further introduce an exploration strategy based on constraining epistemic uncertainty, eliminating the need for explicit exploration bonuses. Under regularity conditions, we derive high-probability guarantees of safety throughout learning and a finite-time sample complexity bound for recovering a near-optimal policy. Empirically, SBSRL achieves safe and efficient exploration both in simulation and in real-world experiments, and readily extends to practical deep-ensemble implementations that scale to high-dimensional continuous control problems.


SayNext-Bench: Why Do LLMs Struggle with Next-Utterance Anticipation?

Yueyi Yang ⋅ Haotian Liu ⋅ Fang Kang ⋅ Mengqi Zhang ⋅ Zheng Lian ⋅ Hao Tang ⋅ Haoyu Chen

We explore the use of large language models (LLMs) for next-utterance anticipation in human dialogue. Despite recent advances in LLMs demonstrating their ability to engage in natural conversations with users, we show that even leading models surprisingly struggle to anticipate a human speaker’s next utterance. Instead, humans can readily anticipate forthcoming utterances based on multi-modal cues—such as gestures, gaze, and emotional tone—from the context. To systematically examine this gap, we propose SayNext-Bench, a benchmark evaluating MLLMs on anticipating context-conditioned responses across diverse real-world scenarios. To support it, we build SayNext-PC, a large-scale multimodal dialogue dataset, and carefully design a multi-level evaluation framework spanning lexical similarity, emotion-intention consistency, and LLM-based overall alignment. Building on this, we develop SayNext-Chat, a cognitively inspired dual-route MLLM that incorporates learnable priming tokens to fuse perceptual cues with anticipatory priors. Extensive experiments demonstrate that SayNext-Chat consistently outperforms state-of-the-art MLLMs across all evaluation levels, corroborated by user studies and LLM-as-Judge evaluations. Our results emphasize the (i) indispensable role of multimodal cues and (ii) active anticipatory processing as foundations of natural human interaction currently missing in MLLMs.


Scalable Minimal-Change Learning for Controllable Image Editing

Shuo Chen ⋅ Fengming Huang ⋅ Yu Yao ⋅ Mingming Gong ⋅ Tongliang Liu

Image editing aims to modify specific attributes of an image while preserving all other aspects. In practice, however, applying even simple editing instructions to existing methods often leads to unintended changes. We identify the central issue: a fundamental causal principle of \emph{minimal change}, which requires that an intervention alter only the intended attributes in the output while leaving all others invariant, has not been explicitly incorporated as an optimization objective. To encourage minimal change, prior work rooted in causal representation learning typically imposes an $L_1$ regularizer on the latent difference between pre- and post-edit representations. However, this strategy does not scale to modern image editing models: $L_1$ regularization fails to induce true sparsity and instead tends to shrink all latent differences uniformly; sparsity in the latent space does not translate to localized change in the output of nonlinear deep image generation models; and the multi-domain or counterfactual supervision it requires is generally unavailable. To effectively leverage the minimal change principle for controllable image editing, we cast it as a reinforcement learning objective. Rather than constraining latent representations, we design rewards that directly capture minimal change over intervention outcomes and optimize them explicitly during training. To produce these rewards reliably, we further introduce an agentic reward model that leverages the chain-of-thought reasoning ability of vision-language models without requiring human annotations. The idea is that instead of emitting a single, unreliable scalar score, the model reasons explicitly about two complementary failure modes, unimplemented changes and unintended changes, and produces this supervision automatically. Experimental results demonstrate that our method produces significantly more localized and semantically consistent edits compared to existing approaches, reducing unintended changes.


Scaling Laws and Tradeoffs in Recurrent Networks of Expressive Neurons

Aaron Spieler ⋅ Georg Martius ⋅ Anna Levina

Cortical neurons are complex, multi-timescale processors wired into recurrent circuits, shaped by long evolutionary pressure under stringent biological constraints. Mainstream machine learning, by contrast, predominantly builds models from extremely simple units, a default inherited from early neural-network theory. We treat this as a normative architectural question. How should one split a fixed parameter budget $P$ between the number of units $N$, per-unit effective complexity $k_e$, and per-unit connectivity $k_c$? What controls the optimal allocation? This calls for a model in which per-unit complexity can be tuned independently of width and connectivity. Accordingly, we introduce the ELM Network, whose recurrent layer is built from Expressive Leaky Memory (ELM) neurons, chosen to mirror functional components of cortical neurons: multi-timescale memory, structured synaptic integration, and nonlinear internal computation. The architecture allows for individually adjusting $N$, $k_e$, and $k_c$ and trains stably across orders of magnitude in scale. We evaluate the model on two qualitatively different sequence benchmarks: the neuromorphic SHD-Adding task and Enwik8 character-level language modeling. Performance improves monotonically along each of the three axes individually. Under a fixed budget, a clear non-trivial optimum emerges in their tradeoff, and larger budgets favor both more *and* more complex neurons. A closed-form information-theoretic model captures these tradeoffs and attributes the diminishing returns at two ends to: per-neuron signal-to-noise saturation and across-neuron redundancy. Connectivity enters as a related mechanism that helps neurons learn distinct signals. A hyperparameter sweep spanning three orders of magnitude in trainable parameters traces a near-Pareto-frontier scaling law consistent with the framework, mapping the budget-constrained tradeoff surface between unit count, unit complexity, and connectivity. This suggests that the simple-unit default in ML is not obviously optimal once this surface is probed, and offers a normative lens on cortex's reliance on complex spatio-temporal integrators.


Scaling Laws for Synthetic Pretraining in Radio-Map Prediction

Khoren Petrosyan ⋅ Artashes Mkrtchyan ⋅ Rafayel Mkrtchyan ⋅ Hrant Khachatrian ⋅ Theofanis Raptis

Large-scale pretraining has driven much of recent progress in deep learning, but many physics-governed prediction problems remain outside this regime: real measurements are scarce, and high-fidelity simulators are slow and often proprietary. Indoor radio-map prediction is a representative example. Existing benchmarks rely on limited high-quality simulated data, while recent methods largely focus on task-specific input features or architectures that encode physics priors. In this work we propose an alternative path. We generate 128M synthetic indoor radio-map samples with a fast simulator that omits known propagation effects, pretrain a ResNet-based encoder-decoder without task-specific modifications, and fine-tune it on high-quality benchmark data. Despite simulator mismatch, our method reduces error by 15\% on average relative to the best known methods across five indoor radio-map prediction tasks and improves transfer on a small real-measurement dataset. Downstream performance follows a predictable data-scaling trend: a power law fitted on runs up to 16M synthetic samples predicts the per-task 2M-normalized RMSE at an order of magnitude more data within 5\% mean absolute percentage error. Probing analyses further show that the pretrained model captures physically meaningful spatial-field structure, including free-space attenuation, transmitter-centered symmetries, and wall-mediated effects. These results suggest that cheap approximate simulators can serve as scalable pretraining engines for physics-governed spatial prediction, partially compensating for scarce high-fidelity data.


Scaling Neural Motor Decoding via Decoupled Behavioral Pretraining

Divyansha . ⋅ Vinam Arora ⋅ Shivashriganesh P. Mahato ⋅ Alexandre Andre ⋅ Eva L Dyer

Recent neural decoding approaches achieve strong performance by training on large collections of paired neural-behavioral recordings, but such datasets are expensive, invasive, and difficult to scale. In contrast, behavioral data is abundant and easy to collect, raising the question of whether neural decoding can benefit from decoupling behavioral and neural representation learning. We introduce BeeMO, a flexible multi-session framework that enables training with arbitrary mixtures of paired neural-behavioral recordings and unpaired behavioral data. A behavior decoder learns movement dynamics directly from behavioral trajectories, while neural activity is incorporated through cross-attention layers when available, allowing behavioral and neural supervision to scale independently. By decoupling neural and behavioral data, we unlock a new dimension for scaling: unpaired behavioral data, which is substantially cheaper and easier to collect than paired neural recordings, can be used directly to improve decoding performance. Across multiple intracortical motor datasets and tasks in nonhuman primates, we show that incorporating unpaired behavioral data consistently improves decoding performance, particularly, and that incorporating task-aligned behavioral trajectories can further improve transfer. We further show that behavior-only pretraining, without any neural pretraining, outperforms single-session supervised baselines and is comparable to models pretrained on substantially larger neural datasets. Together, our results suggest that scaling behavioral data offers a practical and cost-effective path toward neural decoding models that generalizes across subjects and tasks.


SceneShifter: Training-free Multi-Scene Temporal Control for Audio-driven Human Animation

Jianzhi Long ⋅ Rong-Cheng Tu ⋅ Hao Guan ⋅ Siyuan Liang ⋅ Shunyu Liu ⋅ Xiao Luo ⋅ Dacheng Tao

Although audio-driven human animation has achieved impressive realism, it still lacks effective Multi-Scene Temporal Control: orchestrating background transitions at precise timestamps while preserving a coherent foreground subject. We show that this challenge stems from the diffusion transformer's self-attention mechanism, and addressing it requires two capabilities: (1) Semantic-Aware Temporal Decoupling, which isolates background attention across scenes while retaining cross-scene foreground attention for subject consistency; and (2) Foreground Motion Localization, which accurately tracks the foreground subject across latent frames, especially under large motion. To address these challenges, we introduce SceneShifter, a training-free framework for precise multi-scene control in audio-driven human animation. SceneShifter guides spatio-temporal self-attention by suppressing cross-scene attention among background tokens while preserving cross-scene attention among foreground tokens. To support this guidance under large motion, SceneShifter extracts dynamic foreground masks from the self-attention layers and heads that best capture subject trajectories. We further introduce SceneShifterBench, a benchmark designed to evaluate scene-timing accuracy, foreground preservation, and visual quality under large motion, multiple subjects, and occlusions. Experiments show that SceneShifter achieves frame-level scene timing, strong subject preservation, and high visual quality, outperforming existing baselines. Code and examples can be found at https://anonymous.4open.science/r/SceneShifterneurips26.


SCHTs: A Semi-Structured Dynamic Sparse Training Framework for Hardware-Efficient Deep Learning

Yuan Hua ⋅ Zhaoxu Ding ⋅ Jilin Zhang ⋅ Ziyi Cheng ⋅ Zhijie Jian ⋅ Wenqi Gu ⋅ Carlo Vittorio Cannistraci ⋅ Hong Chen

Although Dynamic Sparse Training (DST) has emerged as a promising solution for deploying deep neural networks on resource-constrained edge devices, unstructured DST methods suffer from irregular weight distributions that incur extra index overhead and severely degrade computational efficiency. In contrast, semi-structured sparsity offers computational efficiency but typically suffers from performance degradation at high sparsity levels with strict semi-structured constraints. To address these challenges, we propose Semi-structured Cannistraci-Hebb Training soft rule (SCHTs), a novel hardware-efficient dynamic sparse training framework, which includes four key stages: semi-structured sparse topological initialization, weight initialization, global network pruning, and local network regrowth. Unlike unstructured DST methods, SCHTs supports fine-grained and coarse-grained N:M semi-structured sparsity patterns with constant fan-in or fan-out, ensuring computational efficiency. Extensive experiments demonstrate that the proposed SCHTs framework enables highly sparse networks (>90\%) to achieve performance comparable to, or even surpassing, their dense counterparts. (1) For spiking neural networks, models trained with SCHTs consistently outperform their dense counterparts at 90\% sparsity, yielding accuracy improvements of +0.05\% (to 94.79\%), +2.95\% (to 75.01\%), and +0.09\% (to 99.16\%) on CIFAR-10, CIFAR-100, and N-MNIST, respectively. (2) For artificial neural networks, sparse networks trained with SCHTs outperform their dense counterparts with accuracy improvements of +0.90\% (to 77.54\%) and +0.63\% (to 63.87\%) using GoogLeNet on CIFAR-100 and TinyImageNet, as well as a +0.67\% (to 78.97\%) improvement using ResNet-152 on CIFAR-100. (3) For large language models, SCHTs enables a 90\% sparse LLaMA-130M to yield a highly competitive validation perplexity of 25.49 on OpenWebText, outperforming unstructured DST methods like RigL. Furthermore, cycle-accurate simulations on Trapezoid (a specialized sparse matrix accelerator) reveal that SCHTs achieves a 75\% reduction in index overhead and a 32.7\% improvement in computational efficiency at 1:16 sparsity, demonstrating its effectiveness in translating algorithmic sparsity into tangible on-chip acceleration.


SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring

Chunxiao Li ⋅ Yuan Xiong ⋅ Lijun Li ⋅ Tianyi Du ⋅ Wenlong Zhang ⋅ LEI BAI ⋅ Jing Shao

Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse guidance. Existing benchmarks often rely on templated queries disconnected from real-world hazards, and employ LLM-as-a-Judge paradigms without domain grounding. To address this, we introduce SciHazard, a real-world-grounded benchmark for scientific risks and a dataset agnostic evaluation framework for measuring harmfulness. SciHazard contains 2400 hazardous questions and 600 oversafety questions across 12 disciplines, with both queries grounded in regulated entities and documented failure scenarios. To compute DeHarm-Score , we develop a decomposed evaluating procedure that combines query hazard severity, refusal behavior, and response-level risk. For non-refused responses, it further decomposes response-level harm into Executability, quantified via dynamic checklists with importance weighting, and Net-new risk, assessed through retrieval-augmented claim extraction and synthesis-barrier verification. An expert-validation study shows that DeHarm-Score improves agreement with expert annotations by 90.17% over the strongest baseline. We benchmark 31 frontier LLMs and deep research agents in an extensive scientific safety evaluation. Notably, Deep research agents yield 32.3% higher mean Deharm-Score than standard LLMs, exposing autonomous agents as a critical blind spot in current safety defenses. Code and dataset are available at https://anonymous.4open.science/r/DeharmScore-7B55.


SCION: Scene Composition via Instanced Neural Primitives

William Koch ⋅ Amogh Joshi ⋅ Cyrus Vachha ⋅ Cheng Zheng ⋅ Felix Heide

Real-world scenes are compositional: bricks, blades of grass, pebbles, and tree leaves recur across human-built and natural environments. Existing neural scene representations model these elements independently. Most 3D Gaussian Splatting and follow-up abstraction and compression methods treat each element as unique, fitting millions of independent Gaussians per scene. Prior methods like Splat and Replace fit template objects, but they require mostly manual selection of repeated elements. As a result, these representations store redundant parameters and provide weak manipulation handles for downstream tasks. We introduce SCION, a hierarchical compositional scene representation that replaces independent Gaussians with a compact vocabulary of reusable primitives and lightweight world-space instances that place transformed copies throughout the scene. We fit this representation to multi-view captures via a joint optimization over discrete and continuous scene parameters, combining two-level densification over splats and instances with an adversarial loss that preserves detail across shared primitives. The recovered structure yields a compact, controllable representation while maintaining high quality even at 1.2 MB. SCION achieves rate-distortion favorable to existing Gaussian compression methods, and it enables instance-level scene editing and animation without retraining. Our results show that neural scene representations need not memorize scenes as independent primitives; they can discover reusable parts.


ScopeSAE: Model-Scope Feature Discovery with Interpretable Layer Selection

Qingwen Zeng ⋅ Zehao Fu ⋅ Shuyu Meng ⋅ Linghan Huang ⋅ Jiayi Zhang ⋅ Chenglin Wu ⋅ Ling Chen ⋅ Huaming Chen

Sparse autoencoders (SAEs) are a central tool in mechanistic interpretability. However, existing SAEs are primarily trained per layer. The modeling subspace is therefore fixed by layer identity, independent of which token-layer states actually drive each prediction. We argue that this constraint contributes to several limitations observed in layer-wise SAEs, including low feature utilization, high dictionary redundancy, and features that lack direct behavioral grounding. In this paper, we propose ScopeSAE, which selects the modeling subspace per token by attributing each prediction to its most influential token-layer state via normalized gradient-based attribution, and learns features over the resulting prediction-relevant subspace. Empirically, ScopeSAE yields an effect we term reconstruction-better-than-original. Written-back reconstructions of the SAE produce lower next-token cross-entropy than the original activations, an outcome that, to our knowledge, has not previously been reported for SAEs and that inverts the usual reconstruction-fidelity trade-off. Through interventional analyses and a KL fine-tuning counter-experiment, we show that this effect is attributable to ScopeSAE's prediction-relevant subspace itself rather than to architectural changes. ScopeSAE further improves effective feature count, interpretability, utilization, and dictionary redundancy over existing layer-based baselines, suggesting that choosing the SAE modeling subspace by predictive relevance leads to more useful and behaviorally meaningful features.


Score-based Variational Inference via Quantum Maximally Mixed States

Yuchen CONG ⋅ Zerui Tao ⋅ Chao Li ⋅ Zhe Sun ⋅ Qibin Zhao

Score-based variational inference (VI) provides an alternative to Kullback--Leibler (KL)-based VI by minimizing the Fisher divergence between the variational distribution and the target. Existing eigenvalue-based score-VI methods face two high-dimensional obstacles: exponential parameter growth and instability under degenerate or nearly degenerate low-energy eigenspaces. We propose QuanVI, a scalable quantum-inspired algorithm that combines a mixed-state density-operator formulation with an MPO-motivated quantum tensor network (QTN) parameterization. The former represents degenerate low-energy eigenspaces by their maximally mixed state, avoiding arbitrary eigenvector selection; the latter realizes a compact density operator without explicitly constructing exponentially large matrices. Experiments and ablations show that QuanVI preserves low-dimensional accuracy while scaling to high-dimensional synthetic and Bayesian posterior-approximation benchmarks, including challenging non-Gaussian targets.


SD-LoRA: Training-Time Structural Distillation into LoRA for Few-Shot Vision-Language Adaptation

Mohammed Rahman Sherif Khan Mohammad ⋅ Ardhendu Behera ⋅ Sandip Pradhan ⋅ Swagat Kumar ⋅ Yonghuai Liu

Few-shot adaptation of vision--language models must balance two competing goals: learning fine-grained visual distinctions from very limited supervision, and preserving the efficiency that makes frozen CLIP-like models attractive for deployment. Existing lightweight adaptation methods, including prompt tuning, LoRA, and cache-based adapters, typically supervise global image embeddings and therefore underuse token-level structure. Conversely, structure-aware methods can exploit local evidence but often introduce additional inference-time computation. We propose SD-LoRA, a structural distillation framework that uses token-level graph reasoning only during training and absorbs its effect into a lightweight LoRA-adapted CLIP model. During training, a heterogeneous graph teacher models patch-patch, global-local, and visual-text relations over high-resolution ViT tokens. Rather than keeping this graph at test time, SD-LoRA transfers its structural bias into the shared CLIP-LoRA backbone through implicit distillation and Prototype Predictive Alignment, which aligns graph-conditioned representations with support-derived class prototypes. At inference, the entire graph teacher is discarded; prediction uses only a single CLIP-LoRA forward pass and a cache-based classifier. Across 11 few-shot benchmarks, SD-LoRA improves over a strong CLIP-LoRA baseline by 1.3/1.2/2.4/1.9/2.0% in the 1/2/4/8/16-shot settings, while adding no graph computation at test time. Further results on base-to-novel generalization, cross-dataset transfer, OOD robustness, and calibration show that training-time structural supervision improves not only accuracy but also transfer reliability.


Search, Edit, and Fold: LLM-Guided MSA Optimization for Protein Conformation Prediction

Yu Pei ⋅ Jiangtao Feng ⋅ Hao Wang ⋅ Dongyu Xue ⋅ Keyue Qiu ⋅ Hao Zhou ⋅ Wei-Ying Ma

Accurately identifying alternative protein conformations remains a fundamental challenge, particularly when functionally relevant states are encoded by sparse evolutionary signals within large multiple sequence alignments (MSAs). In this work, we reformulate protein conformation prediction as a combinatorial search problem in MSA space, shifting the focus from structure divergence to evolutionary information discovery. We introduce MSA-Evolver, a optimization framework that enabling LLM with direct manipulation and iterative exploration of MSAs. Specifically, MSA-Evolver introduces a unified action space for MSA editing together with a feedback-guided multi-step reasoning strategy that allows the language model to progressively explore, evaluate, and refine candidate MSAs based on historical search trajectories. Our framework efficiently identifies informative sub-MSAs under limited folding budgets and substantially improves the prediction accuracy of alternative conformations, including open-closed, inward-outward, apo-holo, fold-switching, and intrinsically disordered proteins.


Seeing Beyond the Next Step: World-Model-Guided Human-Like Navigation in Multi-Agent Scenes

SEONGEUN HONG ⋅ JuYeong Hwang ⋅ Jinhyun Kim ⋅ Hanyoung Jang ⋅ Hyeongyeop Kang

Selecting a next motion in a populated scene is not just a matter of avoiding what is currently in the way: an action can be locally feasible yet drift onto the road instead of the crosswalk, cut into a stream of crossing pedestrians, or bottleneck a narrow passage. These failure modes are invisible from the current state alone, and acting on them requires reasoning about what each candidate action will cause. World models offer exactly this, namely action-conditioned predictions of the future, but rolling them out at every decision step is a cost deployable navigation systems cannot pay, and the field has largely traded foresight for speed. We argue this is a false dichotomy. The natural place for a world model in human-aware navigation is at training time, as a preference oracle that distills foresight into a fast policy, rather than at inference time, as an online planner. We pair a discrete visual action prior with two task-aligned world models, one for nearby social dynamics and one for future scene semantics, that score candidate actions by their predicted consequences and refine the policy through group-relative preference distillation. At deployment, both world models are removed and the policy acts in a single forward pass. Across five held-out scenes, this yields navigation that is simultaneously safer, more socially and group-coherent, and more goal-directed than strong reactive and learned baselines, with naive observers rating the resulting motion as more human-like. Taken together, our findings point to a different approach for bringing world-model foresight into deployable agents: not accelerating rollouts, but amortizing them into the policy so that the cost of looking ahead is confined to training.


Seeing Together, Acting Apart: Shared Environmental Understanding for Multi-Robot Navigation

haihong hao ⋅ Lei Chen ⋅ Mingfei Han ⋅ Dong An ⋅ Yuqiang Yang ⋅ Chenglong Yan ⋅ Changlin Li ⋅ Minghao Guo ⋅ Salman Khan ⋅ Xiaojun Chang

Multi-robot embodied systems are commonly expected to collaborate by coordinating their behaviors: agents may communicate, share experience, divide tasks, avoid conflicts, or jointly decide what to do next. This paper asks a complementary question: can collaboration emerge before decision making, by first building a shared understanding of the environment? We introduce Shared Environmental Understanding, a policy-agnostic representation layer that aggregates partial egocentric observations from multiple robots into a common latent world representation. This formulation shifts the starting point of collaboration from agent-centric traces to environment-centric representation, allowing different robots and downstream policies to benefit from shared world knowledge without being forced into a centralized controller. We instantiate this idea with SEER, a shared environment encoding and retrieval framework that learns to turn multi-robot, multi-view observations into reusable navigation context. SEER is designed not as a new navigation policy, but as an upstream environmental understanding module that can be plugged into heterogeneous agents, including both specialized VLN models and general vision-language model agents. We validate SEER through a comprehensive set of representation-level and task-level studies, including latent inverse dynamics analysis, standard VLN evaluation, controlled paired multi-robot navigation, comparisons against action-, history-, and policy-sharing alternatives, and real-world robot deployments. Across these settings, SEER preserves single-agent navigation ability while improving individual and collaborative success in multi-robot scenarios. These findings support a simple but underexplored hypothesis: for embodied cooperation, robots may benefit from sharing the environmental understanding before sharing the policy.


Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

Kunyu Peng ⋅ Zhikun Zhou ⋅ Kailun Yang ⋅ Di Wen ⋅ Ruiping Liu ⋅ Yufan Chen ⋅ Junwei Zheng ⋅ Hao Shi ⋅ Yi Zhou ⋅ M. Saquib Sarfraz ⋅ Danda Pani Paudel ⋅ Luc V Gool

Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoints remains largely unexplored. We study this problem through multi-robot cooperative dynamic spatial reasoning, where a model must answer spatial, temporal, visibility, and coordination questions by integrating synchronized egocentric videos from a team of moving robots. To support this setting, we introduce CoopSR, the first benchmark for this task, together with EgoTeam, a multi-robot egocentric QA dataset. EgoTeam contains 114,227 QA pairs spanning 19 question types, four difficulty tiers, and three team sizes in Habitat and iGibson, along with a real-world test set of around 2,326 QAs collected using two quadruped robots. We further propose SP-CoR (Spectral and Physics-Informed Cooperative Reasoner), an MLLM framework for fine-grained cooperative spatial reasoning. SP-CoR combines dynamics-aware multi-robot frame sampling, spectral- and physics-guided view fusion, and physics-aligned prompt distillation, enabling the model to benefit from privileged robot-pose supervision during training while requiring only egocentric videos at test time. Across 22 MLLM baselines, SP-CoR consistently improves cooperative reasoning, outperforming the strongest fine-tuned baseline by +3.87% on Habitat and +7.12% on iGibson. It also shows stronger generalization to unseen team sizes and real-world robot tests. We will release the benchmark and code.

What does it mean to create a new concept, rather than retrieve a familiar one? Diffusion models sampled many times at the same prompt vary only within the narrow stylistic envelope set by a fixed prompt embedding. We propose an operational definition of creativity as the generation of concepts that are initially unfamiliar yet rapidly learnable to an adaptive observer, and formalize it as a bilevel optimization between a Creator that generates and an Appraiser that adapts: the Appraiser's improvement under a brief inner-loop adaptation provides the reward signal that the Creator maximizes. We call this framework CAMEL (Creator-Appraiser Meta-Learning), and instantiate it on MNIST with an autoencoder Appraiser and on natural images with a CLIP Appraiser based on a low-rank adapter over the text projection. CAMEL produces outputs that lie outside the basin reachable by vanilla diffusion sampling at any sample count, while remaining recognizable as the prompt's class to an off-the-shelf observer.


Seirênes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning

Chi Zhang ⋅ Haibo Qiu ⋅ Qiming ZHANG ⋅ Yufei Xu ⋅ Xinbo Gao ⋅ Jing Zhang

We present Seirênes, a self-play RL framework that transforms contextual interference from a failure mode of LLM reasoning into an internal training signal for co-evolving more resilient reasoners. While RL with verifiable rewards has significantly advanced reasoning capabilities, models can still exhibit fragility when encountering non-idealized contexts: scenarios characterized by superfluous information, tangential instructions, or incidental correlations that differ from the clean distributions typical of standard benchmarks. Seirênes harnesses this vulnerability through a parameter-shared and adversarial self-play loop. Within this framework, a single model is trained to both construct plausible yet distracting contexts that expose its own reasoning blind spots, and solve problems by discerning the essential task from these perturbations to recover the core underlying logic. By pitting these competing objectives against each other, Seirênes compels the model to move beyond superficial pattern matching and anchors its capabilities in robust underlying reasoning. This continuous interaction sustains an informative co-evolutionary curriculum as the model improves. Across seven mathematical reasoning benchmarks and model scales from 4B to 30B, Seirênes achieves average gains of +10.2, +9.1, and +7.2 points. Besides, distracting contexts produced by the 4B Seirênes model reduce the accuracy of top-tier closed-source models (GPT and Gemini) by roughly 4--5 points, revealing Seirênes' general ability to uncover reasoning models' blind spots. Our code will be available.

Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with complex environments, early mistakes can propagate through trajectories and cause cascading failures. Recent approaches improve reliability by incorporating external critique or deliberation, but invoking these mechanisms at every step substantially increases token consumption and latency, limiting practical deployment. We propose SAG (Self-improving Agent with Gated critique), a cost-aware framework that formulates critique invocation as a step-wise decision problem during long-horizon interaction. SAG introduces a lightweight, training-free gating mechanism that estimates the utility of critique using action-level ambiguity signals—global entropy and local top-2 margin—computed over admissible actions. From a decision-theoretic perspective, this mechanism approximates the Value of Information (VoI) of critique, enabling the agent to selectively allocate expensive feedback only when its expected benefit justifies the cost. SAG further incorporates online bootstrapped self-improvement, allowing the actor to internalize critic-assisted behaviors and progressively reduce reliance on critique. Across three long-horizon interactive benchmarks and multiple backbone models, SAG substantially improves the success–cost trade-off compared with both no-critique and always-on critique agents. On ALFWorld, SAG increases task success from $24.6\%$ to $78.3\%$ while maintaining a token budget comparable to ReAct, yielding a $3.1\times$ improvement in normalized token efficiency. Moreover, a $7$B actor with a lightweight $3$B critic achieves performance comparable to a $14$B actor without critique, showing that selective critique can recover most of the reliability benefits of deliberation while dramatically reducing inference cost.

Rotation invariance is a fundamental requirement across many computer vision tasks. Historically, this inductive bias has been encoded through hand-crafted rotation-invariant representations. These are compact, interpretable, and fast to compute, but they come at the cost of descriptive power. More recently, architectures achieve inductive bias through learned representations. These are highly descriptive and achieve strong empirical performance, at the cost of efficiency and interpretability. In this work, we propose an alternative at the intersection of both paradigms. We introduce the selective disk bispectrum (SDB), a complex-valued rotation-invariant vector that preserves all information about the image except its orientation. Our key theoretical contributions are the selective disk bispectrum, its inversion, its (reduced) spatial and computational complexities (compared to the full disk bispectrum), and its expectation and variance under noise. Furthermore, we propose a numerical SDB approximation and provide theoretical guarantees for its accuracy and rotation invariance. Empirically, we validate SDB's invariance and robustness to noise classification tasks. We test our reconstruction algorithm on multi-reference alignment of rotated images.

Language models routinely revise initial answers through intermediate edits, yet current alignment objectives restrict credit assignment to complete outputs or entire rollouts. To formalize the local geometry of improvement, we introduce Lifted State Policy Optimization (LSPO), which augments each textual answer with a continuous auxiliary coordinate to form a lifted state that encodes refinement context beyond surface text. Training rewards transitions that descend an energy landscape, subject to edit and step penalties, and internalizes the resulting answers into the generative model. Across the Qwen3.5 family on mathematics, science, code, logic, and broad reasoning benchmarks, LSPO improves aggregate accuracy while bypassing the latency of explicit revision loops. Component ablations, together with evaluations of internalization and energy ordering on unseen tasks, confirm that this gain originates from the lifted representation and from credit assigned at the transition level.


SelfGuard: Self-Supervised Deviation Modeling for Multi-Modal Jailbreak Detection

Chenchen Jing ⋅ Hongli Guo ⋅ Huarong Jia ⋅ TianweiBai ⋅ Che Sun ⋅ Yang Liu ⋅ Chunhua Shen

Multi-modal large language models (MLLMs) achieve strong vision–language reasoning ability but remain vulnerable to jailbreak attacks that exploit subtle cross-modal cues to bypass safety mechanisms. Detecting such attacks is difficult because harmful data are scarce and rapidly evolving, while detectors trained on known patterns often fail to generalize. In this paper, we cast multi-modal jailbreak detection as self-supervised deviation modeling and learn attack-agnostic signals from benign data only. We propose SelfGuard, a self-supervised framework, which models benign multi-modal regularities and learns to score structured violations via controlled pseudo-harmful deviations. SelfGuard synthesizes pseudo-harmful samples via controlled transformations that target three signature cues, cross-modal semantic inconsistency, format manipulation, and semantic toxicity. It then learns complementary deviation statistics through multi-task self-supervised learning, including misalignment discrimination, reconstruction-based toxicity modeling, and contrastive format modeling. At inference, SelfGuard performs multi-view deviation estimation by aggregating task-specific deviation scores to identify departures from benign multi-modal regularities. Extensive experiments on various benchmarks show the effectiveness of our method.


Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains

jusheng zhang ⋅ Ningyuan Liu ⋅ Kaitong Cai ⋅ Sidi Liu ⋅ Ziliang Chen ⋅ Wenhao Wang ⋅ Keze Wang

Multimodal LLMs often produce fluent yet unreliable reasoning, exhibiting weak step-to-step coherence and insufficient visual grounding, largely because existing alignment approaches supervise only the final answer while ignoring the reliability of the intermediate reasoning process. We introduce SR-MCR, a lightweight and label-free framework that aligns reasoning by exploiting intrinsic process signals derived directly from model outputs. Five self-referential cues—semantic alignment, lexical fidelity, non-redundancy, visual grounding, and step consistency—are integrated into a normalized, reliability-weighted reward that provides fine-grained process-level guidance. A critic-free GRPO objective, enhanced with a confidence-aware cooling mechanism, further stabilizes training and suppresses trivial or overly confident generations. Built on Qwen2.5-VL, SR-MCR improves both answer accuracy and reasoning coherence across a broad set of visual benchmarks; among open-source models of comparable size, SR-MCR-7B achieves state-of-the-art performance with an average accuracy of 81.4%. Ablation studies confirm the independent contributions of each reward term and the cooling module.


Self-Rewarding Sequential Monte Carlo for Masked Diffusion Language Models

Ziwei Luo ⋅ Ziqi Jin ⋅ Lei Wang ⋅ Lidong Bing ⋅ Thomas Schön

This work presents self-rewarding sequential Monte Carlo (SMC), an inference-time scaling algorithm enabling effective sampling of masked diffusion language models (MDLMs). Our algorithm stems from the observation that most existing MDLMs rely on a confidence-based sampling strategy, where only tokens with the highest prediction confidence are preserved at each step. This can restrict generation to a noise-sensitive, greedy decoding paradigm, often limiting trajectory diversity. We address this problem by launching multiple interacting diffusion processes in parallel, referred to as particles, for trajectory exploration. Importantly, we introduce the trajectory-level confidence as a self-rewarding signal for assigning particle importance weights. During sampling, particles are iteratively weighted and resampled to systematically steer generation towards globally confident, high-quality samples. Our self-rewarding SMC is verified on various masked diffusion language models and benchmarks, achieving significant improvement without extra training or reward guidance, while effectively converting parallel inference capacity into improved sampling quality.


Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation

Asim Osman ⋅ Sasha Abramowitz ⋅ Mark Bergh ⋅ Ulrich Armel Mbou Sob ⋅ Ruan John de Kock ⋅ Omayma Mahjoub ⋅ Oussama Hidaoui ⋅ Noah De Nicola ⋅ Arnol M Fokam ⋅ Felix Chalumeau ⋅ Daniel Rajaonarivonivelomanantsoa ⋅ Siddarth Singh ⋅ Refiloe Shabe ⋅ Juan Formanek ⋅ Simon Du Toit ⋅ Arnu Pretorius

Contrastive reinforcement learning (CRL) learns goal-conditioned Q-values through a contrastive objective over state-action and goal representations, removing the need for hand-crafted reward functions. Despite impressive success in achieving viable self-supervised learning in RL, all existing CRL algorithms rely on off-policy optimisation and are mostly constrained to continuous action spaces, with little research invested in discrete environments. This leaves CRL disconnected from widely used and effective, modern on-policy training pipelines adopted across both single-agent and multi-agent RL in continuous and discrete environments. To establish a first connection, we introduce Contrastive Proximal Policy Optimisation (CPPO). CPPO is an on-policy contrastive RL algorithm that derives policy advantages directly from contrastive Q-values and optimises them via the standard PPO objective, without requiring a reward function or a replay buffer. We evaluate CPPO across continuous and discrete, single-agent and cooperative multi-agent tasks. Whilst the existence of an on-policy approach is inherently useful, we additionally observe that CPPO not only significantly outperforms the previous CRL baselines in 14 out of the 18 tasks benchmarked, but matches or exceeds reward-based PPO's performance in 12 out of the 18 tasks without any reward signal.


Semantically Complementary Spectral Views Learning for Graph-Level Anomaly Detection

Yuanming Liu ⋅ Fan Xu ⋅ Xibin Zhao ⋅ Hai Wan ⋅ Wu Junran ⋅ Nan Wang ⋅ Jiqiang Liu

Graph-level anomaly detection (GLAD) is a critical task in domains like social networks and bioinformatics. However, current unsupervised contrastive learning-based methods rely on semantically homogeneous augmentations, leading to catastrophic informational overlap that trivializes the learning objective and effectively masks anomalous signals in GLAD. To overcome this, we propose Sima (Semantically Complementary Spectral Views), a novel contrastive framework that learns from semantically complementary spectral views rather than redundant homogeneous views. Sima first constructs a pair of spectral views to separately capture global community structure and local heterophilous variations, which together cover the full graph spectrum of anomalous patterns. Inspired by the Multi-View Information Bottleneck principle, we further design a soft subgraph extraction module with Score Reconstruction Attention (SRA) and an anchor-guided condensation mechanism to distill the essential shared information from each view into a compact representation. We then optimize a contrastive objective across these distilled views to enforce cross-view alignment and amplify subtle structural deviations. Extensive experiments on challenging GLAD and graph-level out-of-distribution detection (GLOD) benchmarks demonstrate that Sima achieves substantial and consistent gains in average performance over state-of-the-art baselines.


Semantic Freedom Bottleneck for Domain-Generalized Multimodal Face Anti-Spoofing

Yingjie Ma ⋅ Haonan Wang ⋅ Xun Lin ⋅ Hui Ma ⋅ Ruixin Zhang ⋅ Jingyun Zhang ⋅ Jun Wang ⋅ Rizen Guo ⋅ Shouhong Ding ⋅ Weicheng Xie ⋅ Linlin Shen ⋅ Zitong YU

Multimodal face anti-spoofing (FAS) improves spoof detection by combining complementary RGB, infrared, and depth cues, yet its generalization to unseen domains remains fragile. We revisit multimodal domain-generalized FAS from the perspective of representation freedom. In a controlled CLIP-based post-fusion baseline, increasing visual encoder depth does not monotonically improve cross-domain performance, suggesting that higher-capacity representations may introduce additional degrees of freedom that are not necessarily aligned with real-vs-spoof semantics. To inspect this effect, we introduce the Semantic Dominance Ratio (SDR), a diagnostic statistic that measures the relative proportion of feature variation along the text-defined real-vs-spoof semantic axis versus off-axis residual directions. Motivated by this perspective, we propose Semantic Freedom Bottleneck (SFB), which reparameterizes multimodal features around CLIP text geometry. SFB decomposes each representation into a semantic component along the real-vs-spoof axis and a bounded low-rank residual component for modality-specific evidence. The residual basis is regularized to be orthogonal to the semantic axis, encouraging the residual space to retain complementary modality-specific cues without duplicating the task-semantic direction. This yields a conservative multimodal constraint that preserves discriminative real/spoof semantics while limiting unconstrained non-semantic variation. Experiments on WMCA, CASIA-CeFA, PADISI, and CASIA-SURF under fixed-modality, missing-modality, and limited-source protocols demonstrate consistently strong cross-domain performance.

LLM-based multi-agent systems (LLM-MAS) coordinate well on reliable channels but fail in poorly understood ways when communication is lossy. We measure these failures against a content-blind reference: synchronous push gossip on idempotent payloads, which admits classical bounds in diameter, conductance, and percolation. Our \emph{paired-Oracle} protocol runs an LLM agent and a deterministic gossip Oracle on identical scheduler, topology, and channel realisation, so any outcome difference is attributable to per-node decision logic. On Distributed Common-Free-Slot (\textsc{DCFS}), a set-union task isomorphic to all-to-all rumour spreading, three signatures of a \emph{semantic residual} emerge as channel reliability is swept: a phase-boundary shift $\Delta p^* \approx 0.35$ identified directly from paired trials; a $16$--$46\%$ rate of consensus on wrong answers despite complete knowledge, a regime classical gossip cannot enter; and a $3.6\times$--$5.1\times$ round-count overhead at perfect channels. A mediation analysis shows that coverage and failure-mode labels are jointly a sufficient statistic for accuracy --- no within-stratum residual remains --- so the LLM--Oracle gap is fully captured by a $p$-dependent stochastic mapping over four failure modes. This decomposition supports a four-tier mitigation hierarchy whose tier-wise gains match the externalised stratum's marginal share, and a five-parameter scaling law over an eight-model panel with leave-one-model-out RMSE $0.08$.


Separating Common and Unique Directions for Model Merging

Shiting Wang ⋅ Yingjie Zou ⋅ Ke Li

Model merging combines multiple fine-tuned models without additional training, and arithmetic methods have emerged as a simple yet effective way to merge models without data. However, existing arithmetic methods treat task vectors in parameter space and apply a single element-wise rule to directions that can differ in mergeability. In a shared directional basis, we find a consistent two-regime structure: a low-dispersion common bulk shared across tasks and a high-dispersion task-dominated tail that carries disproportionate energy. These findings motivate Common-Unique Subspace Decomposition (CUSD), a data-free pre-merging method that separates task vectors into common and unique directional components, merges the common components with downstream arithmetic operator, and retains unique residuals separately. Experiments across 11 base models, four model families, and four arithmetic-based merging operators show that CUSD improves 42 of 44 model-operator combinations and achieves up to 22.5% improvement over the corresponding baselines.


Serialization Tax in Shared-Latent Exchangeable Decisions

Siming Zhang ⋅ Zhehui Shen ⋅ Shijie Chen ⋅ Xinle Gu ⋅ Yansen Yu ⋅ Hang Yu

Itemwise calibration does not guarantee calibration of the decision it feeds. Many systems output one score per item in a partially observed group, but the downstream action asks whether enough hidden items cross a threshold: show a slate, defer a panel decision, or inspect a batch. If those hidden items share an unresolved user, patient, or batch state, averaging that state before aggregation preserves means but makes the aggregate posterior too concentrated. We call this failure the serialization tax. For Bernoulli de Finetti posteriors, the exact hidden-count law dominates any independent completion with the same total predictive mean in convex order, yielding variance, total-variation, threshold, and Bayes-value gaps. The same interface loss extends to non-binary outcomes through empirical-measure Laplace functionals and to heterogeneous slates through a shared-covariance identity. TaxScore turns the theory into a routing rule: use cheap marginal interfaces when hidden-block ambiguity is far from the decision boundary, and propagate a shared posterior sample when it is near. On pre-specified, hidden-label-free MovieLens 1M/10M audits, replacing Transformer count heads with existing stochastic shared-latent heads improves high-TaxScore negative log-likelihood (NLL)/utility from $0.673/0.696$ to $0.629/0.723$ and from $0.670/0.691$ to $0.639/0.715$; routing richer heads only on flagged slates gives $0.006$-$0.012$ full-population utility gains.


Settling Pure Differentially Private Covariance Estimation

Tommaso d’Orsi ⋅ Gleb Novikov ⋅ Walter McKelvie

We study the problem of $d$-dimensional covariance matrix release under \textit{pure} differential privacy with error measured in Schatten-$p$ norms, where $p\in [1,\infty].$ We identify two meaningful sample-size regimes with qualitatively different behavior, and give a single efficient algorithm together with matching information-theoretic lower bounds throughout. In the \emph{large-sample} regime $n \gtrsim d^{2}/\varepsilon$, we show that the simple $K$-norm mechanism analyzed by \cite{d2026purely} is simultaneously optimal for all Schatten norms $p\in[1,\infty]$, yielding sharp rates from nuclear to spectral norm. In the \emph{moderate-sample} regime $d/\varepsilon \lesssim n \lesssim d^{2}/\varepsilon$, the problem remains non-trivially solvable, but the $K$-norm mechanism becomes suboptimal. We then design a single improved, efficient, $\varepsilon$-differentially private estimator based on the perturb-and-project framework; this estimator is optimal in all regimes and simultaneously for all Schatten norms. Prior to this work, improvements over the $K$-norm mechanism in this regime were only known for Frobenius loss \cite{nikolov2023private}. A key feature of the moderate-sample regime is an inherent dependency of the error on the nuclear norm of the input, first observed by \cite{dong2022differentially} and further investigated by \cite{d2026purely}. Our mechanism captures this dependence uniformly over $p$, and our lower bounds show it to be unavoidable, establishing optimal rates in all regimes, including the correct dependence on $\|\Sigma\|_*$. %The implications of our approach go beyond covariance release: we obtain optimal guarantees for privately releasing arbitrary matrices under nuclear-norm adjacency, and give partial extensions to higher-order moment estimation under pure differential privacy. These extensions highlight the perturb-and-project framework as a flexible tool for private high-dimensional data release, even beyond second moments. Our results extend and are optimal for the more general problem of privately releasing arbitrary matrices under nuclear-norm adjacency. Finally, our lower-bound techniques further generalize to higher-order tensors, establishing similar limitations for higher-order moment estimation under pure differential privacy.


SGEvolve: Semantic Gradient Guidance for LLM-Driven Evolutionary Search

Kewei Feng ⋅ Jinbiao Nie ⋅ Xiaoyuan Zhang ⋅ Quanhua Liu

Large language model (LLM)-driven evolutionary frameworks have emerged as a promising paradigm for iterative search in complex optimization and structured generation tasks. However, existing methods typically rely on heuristic variation operators and lack an explicit mechanism to extract directional information from observed performance feedback, limiting both efficiency and reliability. To address this, we propose SGEvolve, a semantic gradient-guided evolutionary framework for LLM-driven optimization. The key idea is to treat the performance differences between parent and offspring as zeroth-order evidence of an underlying ascent direction. Based on this view, we formulate semantic gradient extraction as a Maximum A Posteriori (MAP) estimation problem, where the LLM infers the most probable direction of improvement conditioned on observed evolutionary steps. The inferred gradients are then used to guide mutation and crossover, biasing generation toward more promising regions of the search space. We evaluate SGEvolve on a diverse set of tasks, including mathematical optimization, algorithm design, and real-world industrial problems. Experimental results demonstrate that SGEvolve consistently outperforms existing open-source baselines, achieving superior performance in both solution quality and search efficiency.


Sharpness, Stability, and Step-Size Scaling in Deep Polynomial Networks

Alexandru Crăciun ⋅ Debarghya Ghoshdastidar

Predicting the maximum stable learning rate of a neural network from its architecture alone has remained outside the reach of current theory, except for deep linear and shallow scalar models, where exact sharpness expressions are available. We close this gap for deep polynomial networks by proving the first computable, architecture-explicit lower bound on the largest Hessian eigenvalue at any interpolating minimum. The bound factorizes into three independently computable terms: a data-geometry factor, an architectural-capacity factor, and a label-energy factor. Combining our bound with the dynamical framework of Chemnitz and Engel [2025] and the stability conditions of Cohen et al. [2022], we derive critical step-size upper bounds for a range of optimizers from vanilla gradient descent to adaptive versions like Adam. The maximum stable learning rate for deep polynomial networks scales as $(\text{activation degree})^{(-2 \times \text{depth})}$, placing realistic polynomial network training in the Edge of Stability regime. Our derivation shows that this scaling is a Hessian-level consequence of the network's algebraic structure, not a property of any specific optimizer or dataset. With identity activation, our bound recovers the exact deep linear sharpness of Mulayoff and Michaeli [2020].


SheafStain: Sheaf-Theoretic Schr\"odinger Bridge for Spatially and Biologically Coherent Virtual Staining

Hyeongyeol Lim ⋅ Hongjun Yoon ⋅ Eunjin Jang ⋅ Daeky Jeong ⋅ Won June Cho ⋅ HWAMIN LEE

Current virtual staining approaches offer the potential for time- and cost-efficient biomarker quantification in cancer diagnostics and prognostics. However, patch-wise inference for gigapixel whole slide images (WSIs) fails to maintain spatial continuity, yielding artifacts that cause catastrophic mismatches with ground-truth images. Although pathology Vision Foundation Models (VFMs) offer rich representations, their self-attention causes varying global contexts to produce inconsistent embeddings for the same physical region. We formalize and validate this ``context contamination'' as a sheaf-theoretic problem where these embeddings form a presheaf that violates the gluing axiom. To address this, we propose SheafStain, a new approach that reinterprets VFM features as sheaf-like sections for spatially and biologically coherent virtual staining. Specifically, SheafStain integrates class and patch tokens into a Schr\"odinger Bridge framework as sheaf-like sections. While the class token anchors biological consistency, patch tokens form a per-position spatial map. A backbone co-pretrained on Hematoxylin \& Eosin (H\&E) and Immunohistochemistry (IHC) yields non-degenerate cross-stain stalks, so a single VFM feature space supervises both input conditioning and output stain alignment. Departing from prior work that evaluates on isolated $256 \times 256$ patches and either random-crops or resizes the $1024 \times 1024$ ground truth, we translate at $256 \times 256$ and evaluate on the stitched $1024 \times 1024$ outputs across HER2, ER, PR, and Ki-67. SheafStain demonstrates promising results against six prior methods while mitigating patch-boundary stitching artifacts. Code will soon be released upon acceptance.

Multimodal large language models (MLLMs) achieve ever-stronger performance on visual-language tasks. Even as traditional visual question answering benchmarks approach saturation, reliable deployment requires satisfying low error tolerances in real-world out-of-distribution (OOD) scenarios. Precisely, selective prediction aims to improve coverage, i.e.\ the share of inputs the system answers, while adhering to a user-defined risk level. This is typically achieved by assigning a confidence score to each answer and abstaining on those that fall below a certain threshold. To enable reliable generalization, we require reasoner models to produce localized visual evidence while answering, and design a selector that explicitly learns to estimate the quality of the localization provided by the reasoner. We show that SIEVES (Selective Prediction through Visual Evidence Scoring) improves coverage by up to three times on challenging OOD benchmarks (V* Bench, HR-Bench-8k, MME-RealWorld-Lite, VizWiz, and AdVQA), compared to non-grounding baselines. Beyond better generalization to OOD tasks, the design of the SIEVES selector enables transfer to proprietary reasoners without access to their weights or logits, such as o3 and Gemini-3-Pro, providing coverage boosts beyond those attributable to accuracy alone. We highlight that SIEVES generalizes across all five tested OOD datasets and reasoner models (Pixel-Reasoner, o3, and Gemini-3-Pro), without benchmark- or reasoner-specific training or adaptation.


SIGMA: Semantic-Difference Instruction-Grounding Mask Annotator for Text-Driven Image Manipulation Localization

Peiyu Zhuang ⋅ Jianquan Yang ⋅ Haodong Li ⋅ Zhuoying Cai ⋅ Ruitao Xie ⋅ Jishen Zeng ⋅ Baoying Chen ⋅ Jiwu Huang ⋅ XIAOCHUN CAO

Text-driven image editing has advanced rapidly, but reliably localizing these manipulations requires image manipulation localization (IML) models trained on large pixel-annotated datasets, and there is still no low-cost way to obtain such training data at scale. We observe that these data already exist in disguise: public editing datasets contain millions of structurally identical $\textit{(original, edited)}$ pairs to IML training samples, lacking only pixel-level masks. Recovering these masks automatically is non-trivial: pixel differencing is overwhelmed by diffusion-induced perturbations across all pixels, and instruction-only grounding localizes only what the prompt describes, missing unintended editor side-effects. We propose $\textbf{SIGMA}$ ($\textbf{S}$emantic-difference $\textbf{I}$nstruction-$\textbf{G}$rounding $\textbf{M}$ask $\textbf{A}$nnotator), which performs semantic-feature differencing in a vision foundation backbone and injects an instruction-derived spatial prior into this visual stream via bidirectional cross-modal refinement, amplifying the difference signal at intended-edit regions when the editor faithfully realizes user intent. SIGMA is trained in two complementary stages: Stage I supervises on inpainting masks; Stage II closes the diffusion-domain shift via VAE-roundtrip noise calibration, EMA self-training, and an edit-noise disentanglement loss. SIGMA outperforms existing automatic mask generators on five benchmarks ($\textbf{+12.20\%}$ F1, $\textbf{+11.16\%}$ IoU). When applied to public editing corpora, it produces a $\sim$1.1M IML training set that improves six diverse detectors by $\textbf{+18.34\%}$ F1 across five datasets, turning previously unused editing data into a model-agnostic supervisory resource for IML. For reproducibility, our anonymized code is available at https://anonymous.4open.science/r/SIGMA-14FE.


Signal-Adaptive Trust Regions for Gradient-Free Optimization of Recurrent Spiking Neural Networks

Jinhao Li ⋅ Yuhao Sun ⋅ Zhiyuan Ma ⋅ Hao He ⋅ Xinche Zhang ⋅ Xing Chen ⋅ Li Jin ⋅ Sen Song

Recurrent spiking neural networks (RSNNs) are a promising substrate for energy-efficient control policies, but training them for high-dimensional, long-horizon reinforcement learning remains challenging. Population-based, gradient-free optimization circumvents backpropagation through non-differentiable spike dynamics by estimating gradients. However, with finite populations, high variance of these estimates can induce harmful and overly aggressive update steps. Inspired by trust-region methods in reinforcement learning that constrain policy updates in distribution space, we propose Signal-Adaptive Trust Regions (SATR), a distributional update rule that constrains relative change by bounding KL divergence normalized by an estimated signal energy. SATR automatically expands the trust region under strong signals and contracts it when updates are noise-dominated. We instantiate SATR for Bernoulli connectivity distributions, which have shown strong empirical performance for RSNN optimization. Across a suite of high-dimensional continuous-control benchmarks, SATR improves stability under limited populations and reaches competitive returns against strong baselines including PPO-LSTM. In addition, to make SATR practical at scale, we introduce a bitset implementation for binary spiking and binary weights, substantially reducing wall-clock training time and enabling fast RSNN policy search.


Simmer: A Scalable Pretraining Recipe for Video-Text Encoders

Bingliang Zhang ⋅ Wenda Chu ⋅ Albert Li ⋅ Yuan Sui ⋅ Hongkai Zheng ⋅ Sophia Stiles ⋅ Kailen Hargenrader ⋅ Yisong Yue ⋅ Sabera Talukder

Most video-text encoders are pretrained on carefully curated video-text datasets where each video is paired with a single short caption. Such supervision captures only one narrow view of a video, despite videos naturally supporting many complementary descriptions of their objects, actions, and relations. We present Simmer, a scalable pretraining recipe for video-text representation learning that increases textual coverage per video and matches the training objective to this richer supervision. We first augment a video corpus with diverse synthetic captions written in many distinct textual registers, producing a large set of aspects per video. Since this many-to-one supervision is difficult to capture in a single global embedding, we pair caption diversification with a multi-vector contrastive loss that allows videos to align with multiple relevant captions. Second, we show that introducing a distinct cooldown phase over a smaller corpus of human-sourced dense captions with our stylistic augmentations further improves performance. Our resulting encoders achieve state-of-the-art performance on five video-text encoding tasks across three model scales, outperforming prior works while using 2-60x fewer videos. Our results show that broader caption coverage and multi-aspect alignment can pave the way towards new scaling laws in video-text modeling.


Simple yet Effective Budget-Feasible Procurement Auctions for Submodular Welfare Maximization

Junxiang Zhang ⋅ Qingchun HE ⋅ Xinhui Lu ⋅ Yaxin Hu ⋅ Feng Li ⋅ Jing Tang ⋅ Kai Han

We study budget-feasible procurement auctions for social-welfare maximization under submodular valuations. In this setting, a buyer with a limited payment budget seeks to procure goods or services from strategic sellers, with the objective of maximizing the buyer's value for the selected sellers minus their true costs. We propose a simple yet effective single-round clock auction in which each seller receives at most one price offer. The mechanism satisfies desirable economic properties, including truthfulness, individual rationality, budget feasibility, and non-negative auctioneer surplus. Moreover, it achieves approximation ratios of 8.52 for monotone submodular valuations, 23.2 in expectation for non-monotone submodular valuations, and 24.52 deterministically for non-monotone submodular valuations, using only a linear number of value-oracle queries. These guarantees improve over the recent independent work of Cui et al.~(ICML 2026) on the same problem in approximation ratio, value-oracle complexity, and number of pricing rounds. Experiments on influence maximization in social networks and crowdsourcing further demonstrate the effectiveness and efficiency of our approach.


Simplified Reversible Residual Networks

Erland B Olsson ⋅ Zhirong Yang

The de facto standard way of training neural networks requires storing activations at each layer to memory, in order to compute gradients in backpropagation. As progress has often been made by stacking larger and more layers, memory consumption has now become a major bottleneck for making frontier models available to everyone. In this work, we take advantage of the Residual Network (ResNet) layer formulation that is now used in many state-of-the-art neural network architectures to recompute the activations at each layer on-the-fly during backpropagation instead of storing them in memory. We propose a modification to Reversible Residual Network (RevNet) that overcomes its limitation of requiring to replace the ResNet layer network function with two new network functions, by setting either one to an identity function. This modification is simple yet has a significant advantage because now there is no need to change the original ResNet architecture. Our method can work as a drop-in replacement for layers with residual connections such as in ordinary ResNets or Transformers with practically no loss in performance while offering significant reductions in memory consumption. With this simple modification, we are able to maintain the same computational cost and ease of use as activation checkpointing, which requires no architectural changes, while leveraging a reversible procedure such that we only need to store the activations from the last layer. The method is called Simplified RevNet, and compared to previous work in reversible architectures, we here propose a simpler and more streamlined approach that comes in two variants based on which of $F$ or $G$ in RevNet becomes an identity function. Empirically, we demonstrate performance at practically the same level as the non-reversible counterparts on ImageNet image classification with ResNets and on OpenWebText language modeling with Transformers.


Sketching the Readout of Large Language Models for Scalable Data Attribution and Valuation

yide ran ⋅ Jianwen Xie ⋅ Minghui Wang ⋅ W. Jim Zheng ⋅ Denghui Zhang ⋅ Chuan Li ⋅ Zhaozhuo Xu

Data attribution and valuation are critical for understanding data-model synergy for Large Language Models (LLMs), yet existing gradient-based methods suffer from scalability challenges on LLMs. Inspired by human cognition, where decision making relies on a focused readout of relevant memories rather than replaying all pathways, we introduce **RISE (Readout Influence Sketching Estimator)**. Instead of computing and indexing gradients across the entire LLM, RISE focuses on influence hotspots at the output layer, where influence signals concentrate, and the gradient admits a decomposed outer-product form. This enables a dual-channel representation combining a lexical residual channel (RH) and a semantic projected-error channel (GH). Applying CountSketch projections to these channels achieves strong compression while maintaining accurate attribution. Across the OLMo (1B–32B) and Pythia (14M–6.9B) families, RISE reduces index storage by up to 112$\times$ compared to RapidIn and scales to 32B parameters LLM, where gradient-based baselines such as RapidIn and ZO-Inf become memory-infeasible. We evaluate RISE on two paradigms: (1) retrospective attribution, retrieving influential training examples for specific predictions, and (2) prospective valuation, scoring candidate data utility zero-shot. We validate RISE on three tasks: Howdy backdoor data detection, Finance-Medical domain separation, and Brain Rot high-quality data selection. In a closed-loop Brain Rot study, continued pretraining on RISE-selected data yields consistent downstream improvements. Overall, RISE provides a practical and scalable primitive for influence analysis and training-data selection in modern large language models.


SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles

Yuyao Ge ⋅ Yiwei Wang ⋅ Yuchen He ⋅ Baolong Bi ⋅ Lingrui Mei ⋅ Jiayu Yao ⋅ Lizhe Chen ⋅ Shenghua Liu

Memory-augmented reinforcement learning strengthens LLM agents' ability to solve complex long-horizon tasks. Skills are one such form of memory, pairing instructions with an applicability condition over task types. However, retaining every skill indiscriminately as the policy improves lets obsolete or harmful entries accumulate and mislead the agent. We propose SkillForge, an agentic RL method that compiles and evolves the skill library through a fitness-driven skill lifecycle of trial, active, stable, and retired states, so that the skills and the model co-evolve throughout training. A preliminary evaluation phase first uses the base model's own rollouts to pre-retire low-fitness skills, yielding a filtered library that then seeds supervised fine-tuning. Reinforcement learning takes over from this checkpoint, and at each iteration selective retirement, stabilization, and LLM-guided mutation continue to forge the skill library alongside policy optimization. Across multiple interactive agent benchmarks, SkillForge achieves the highest aggregate success rate, delivering up to 7.8\% relative improvement over the strongest baseline while keeping the skill library compact throughout training. We introduce Skillfurnace, a dataset of 5k+ annotated records bundling retirement-filtered SFT trajectories, evolved skill libraries with fitness annotations, and retirement events with human-annotated failure categories to support research on skill quality and lifecycle management.


SkillOrchestra: Learning to Route Agents via Skill Transfer

Jiayu Wang ⋅ Yifei Ming ⋅ Zixuan Ke ⋅ Shafiq Joty ⋅ Aws Albarghouthi ⋅ Frederic Sala

Compound AI systems promise capabilities beyond those of individual models, yet their success depends critically on effective orchestration. Existing routing approaches face two limitations: (1) input-level routers make coarse query-level decisions that ignore evolving task requirements; (2) RL-trained orchestrators are expensive to adapt and often suffer from routing collapse, repeatedly invoking one option in multi-turn scenarios. We introduce SkillOrchestra, a framework for skill-aware orchestration. Instead of directly learning a routing policy end-to-end, SkillOrchestra learns fine-grained skills from execution experience and models agent-specific competence and cost under those skills. At deployment, the orchestrator infers the skill demands of the current interaction and selects agents that best satisfy them under an explicit performance-cost trade-off. Extensive experiments across ten benchmarks demonstrate that SkillOrchestra outperforms SoTA RL-based orchestrators by up to 27.5% with 700x and 300x learning cost reduction compared to Router-R1 and ToolOrchestra, respectively. These results show that explicit skill modeling enables scalable, interpretable, and sample-efficient orchestration, offering a principled alternative to data-intensive RL-based approaches.


SkillOS: Learning Skill Curation for Self-Evolving Agents

Siru Ouyang ⋅ Jun Yan ⋅ Yanfei Chen ⋅ Rujun Han ⋅ Zifeng Wang ⋅ Bhavana Dalvi Mishra ⋅ Rui Meng ⋅ Chun-Liang Li ⋅ Yizhu Jiao ⋅ Kaiwen Zha ⋅ Maohao Shen ⋅ Vishy Tirumalashetty ⋅ George Lee ⋅ Jiawei Han ⋅ Tomas Pfister ⋅ Chen-Yu Lee

LLM-based agents are increasingly deployed to handle streaming tasks, yet they often remain one-off problem solvers that fail to learn from past interactions. Reusable skills distilled from experience provide a natural substrate for self-evolution, where high-quality skill curation serves as the key bottleneck. Existing approaches either rely on manual skill curation, prescribe heuristic skill operations, or train for short-horizon skill operations. However, they still struggle to learn complex long-term curation policies from indirect and delayed feedback. To tackle this challenge, we propose SkillOS, an experience-driven RL training recipe for learning skill curation in self-evolving agents. SkillOS pairs a frozen agent executor that retrieves and applies skills with a trainable skill curator that updates an external SkillRepo from accumulated experience. To provide learning signals for curation, we design composite rewards and train on grouped task streams based on skill-relevant task dependencies, where earlier trajectories update the SkillRepo, and later related tasks evaluate these updates. Across multi-turn agentic tasks and single-turn reasoning tasks, SkillOS consistently outperforms memory-free and strong memory-based baselines in both effectiveness and efficiency, with the learned skill curator generalizing across different executor backbones and task domains. Further analyses show that the learned curator produces more targeted skill use, while the skills in SkillRepo evolve into more richly structured Markdown files that encode higher-level meta-skills over time.


SkiP: When to Skip and When to Refine for Efficient Robot Manipulation

Mingtong Dai ⋅ Guanqi Peng ⋅ Yongjie Bai ⋅ Feng yan ⋅ chunjie chen ⋅ Lingbo Liu ⋅ Liang Lin ⋅ Xinyu Wu

Previous imitation learning policies predict future actions at every control step, whether in smooth motion phases or precise, contact-rich operation phases. This uniform treatment is wasteful: most steps in a manipulation trajectory traverse free space and carry little task-relevant information, while a small fraction of \emph{key} steps around contacts, grasps, and alignment demand dense, high-resolution prediction. We propose a novel \emph{action relabeling} mechanism: at each timestep in a skip segment, we replace the behavior cloning target with the action at the entrance of the next key segment, enabling the policy to leap over redundant steps in a single decision. The resulting \textbf{Skip Policy (SkiP)} dynamically leaps over skip segments and intensively refines actions in key segments, within a single unified network requiring no learned skip planner or hierarchical structure. To automatically partition demonstrations into key and skip segments without manual annotation, we introduce \emph{Motion Spectrum Keying} (MSK), a fast, task-agnostic procedure that detects local motion complexity from action signals. Extensive experiments across 72 simulated manipulation tasks and three real-robot tasks show that SkiP reduces executed steps by $15$--$40\%$ while matching or improving success rates across various policy backbones.

Scientific communication increasingly relies on slide-based narrated videos rather than static PDFs. However, transforming a paper into a presentation video is not a conventional text-to-video task. It requires preserving technical rigor, rendering dense multimodal assets faithfully, and synchronizing narration, subtitles, and cursor motion over long horizons. Current research faces two critical bottlenecks: (1) Benchmark Gap: existing evaluations are restricted to narrow domains or structured LaTex inputs with limited long-range supervision; (2) Evaluation Gap: current protocols often overlook the joint measurement of content fidelity, channel alignment, and actual knowledge transfer. To bridge these gaps, we introduce SLIDE P2V-BENCH, a cross-domain bench mark comprising 108 high-quality paper-video pairs spanning four domains and various durations, coupled with an auditable multi-dimensional evaluation suite. Furthermore, we propose PAPER2MEDIA, a slide-grounded agentic pipeline that produces editable presentations by synchronizing decoupled media channels via a shared semantic clock. Extensive experiments demonstrate that PAPER2MEDIA significantly outperforms existing automated baselines and, notably, achieves performance comparable to human-authored presentations in terms of narrative coherence, visual quality, and knowledge conveyance efficiency.


SMI: Statistical Membership Inference for Reliable Unlearned Model Auditing

Jialong Sun ⋅ Zeming Wei ⋅ Jiaxuan Zou ⋅ Jiacheng Gong ⋅ Chengyang Dong ⋅ Jie Fu ⋅ Heng Xu ⋅ Jialong Li ⋅ Bo Liu

Machine unlearning (MU) is essential for enforcing the right to be forgotten in machine learning systems. A key challenge of MU is how to reliably audit whether a model has truly forgotten specified training data. Membership Inference Attacks (MIAs) are widely used for unlearned model auditing, where samples that evade membership detection are regarded as successfully forgotten. We show this assumption is fundamentally flawed: failed membership inference does not imply true forgetting. We prove that unlearned samples occupy a fundamentally different positions in the feature space feature space than non-member samples, making this alignment bias unavoidable and unobservable, which leads to systematically optimistic evaluations of unlearning performance. Meanwhile, training shadow models for MIA incurs substantial computational overhead. To address both limitations, we propose Statistical Membership Inference (SMI), a training-free auditing framework that reformulates auditing as estimating the non-member mixture proportion in the unlearned feature distribution. Beyond estimating the forgetting rate, SMI also provides bootstrap reference ranges for quantified auditing reliability. Extensive experiments show that SMI consistently outperforms all MIA-based baselines, with no shadow model training required. Overall, SMI establishes a principled and efficient alternative to MIA-based auditing methods, with both theoretical guarantees and strong empirical performance.

Boundary representations (B-reps) encode CAD geometry as parametric surface patches with explicit topological adjacency, forming information-rich geometric graphs. Yet existing methods encode coordinates and normals as scalars, making models rotation-sensitive, while invariant alternatives using per-face local frames discard the shared orientation reference needed to capture inter-face directional relationships. We introduce SO3brep, the first $SO(3)$-equivariant neural architecture for B-rep learning. Faces are represented by scalar, vector, and tensor irreducible representations constructed in per-face local frames, lifted to a shared global frame via Wigner $D$-matrices, and propagated through equivariant message passing. This preserves inter-face directional structure during message passing while guaranteeing $SE(3)$ invariance at the output. Across five CAD benchmarks for shape classification and per-face segmentation, SO3brep consistently outperforms prior B-rep and equivariant point-cloud methods while significantly improving data efficiency. These results demonstrate that $SO(3)$-equivariance is a powerful inductive bias for structured CAD geometry.


SOAP-Bubbles: Effective and Scalable Variational Learning with Structured Covariances

Robert Adrian Minut ⋅ Nico Daheim ⋅ Marco Miani ⋅ Mohammad Emtiyaz Khan ⋅ Wu Lin ⋅ Thomas Möllenhoff

Structured posteriors are expected to be better than mean-field posteriors, but recent variational learning methods for large deep networks only use Gaussians with diagonal covariance. A reason is that diagonal covariances can be estimated by simple modifications of existing optimizer's implementations, for instance, the IVON optimizer closely follows Adam's code. Unfortunately, no such alternatives for Gaussians with structured covariances are effective and scalable. Here, we fill this gap and show that block-diagonal covariances can be obtained by adapting the SOAP optimizer. Specifically, our Eigenspace Variational Online Newton (EVON) method modifies SOAP to run IVON instead of Adam in the Eigenspace of the preconditioner. We refer to the posteriors obtained this way as 'SOAP-Bubbles' and show that for logistic regression they recover the optimal posterior approximation among full Gaussians. For language model pretraining we get significant improvements in validation loss over IVON without any increase in cost, and ensembling models drawn from SOAP-Bubbles reduces the test loss more than ensembling using IVON's posterior. Our work shows that there is a fundamental connection between second-order optimization and Gaussian posteriors, which can be used to improve the accuracy of variational learning.


SOAR: Regression-based LiDAR Relocalization for UAVs

Hengyu Mu ⋅ Jianshi Wu ⋅ Yuxin Guo ⋅ XianLian Lin ⋅ Qingyong Hu ⋅ Sheng Ao ⋅ Chenglu Wen ⋅ Cheng Wang

Regression-based LiDAR relocalization has recently emerged as a promising solution for high-precision positioning in GNSS-denied environments. However, these methods are primarily tailored to autonomous driving, exhibiting significantly degraded accuracy in unmanned aerial vehicle (UAV) scenarios due to arbitrary pose variations and irregular flight paths. In this paper, we propose SOAR, a regression-based LiDAR relocalization framework for UAVs. Specifically, we introduce a locality-preserving sliding window attention module with locally invariant positional encoding to capture discriminative geometric structures robust to viewpoint changes. A coordinate-independent feature initialization module is further designed to eliminate sensitivity to global transformations. Furthermore, most existing UAV datasets are limited to evaluate LiDAR relocalization in real-world, due to the lack of synchronized LiDAR scans, accurate 6-DoF poses, or multiple traversals. Thus, we construct a large-scale UAV LiDAR localization dataset with 4 scenes and 13 irregular paths exhibiting rotation and altitude variations, providing a more realistic benchmark for UAVs. Extensive experiments demonstrate that our method achieves state-of-the-art performance, improving the localization success rate by 40\% and reducing mean error over 10m on UAVLoc. Our code and dataset will be released soon, and the dataset demo is available via the anonymous link.


SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation

LIANGYANG OUYANG ⋅ Ruicong Liu ⋅ Caixin Kang ⋅ Yifei Huang ⋅ Yoichi Sato

Video generation has advanced rapidly, producing photorealistic videos from text or image prompts. Meanwhile, film production and social robotics increasingly demand multi-person videos with rich social interactions, including conversations, gestures, and coordinated actions. However, existing models offer no explicit control over interactions, such as who performs which action, when it occurs, and toward whom it is directed. This often results in wrong person performing unintended actions (actor-action mismatch), disordered social dynamics, and wrong action targets. To address these challenges, we present SocialDirector, a training-free interaction controller that enhances the generation model by modulating cross-attention maps. SocialDirector contains two modules: Social Actor Masking and Directional Reweighting. Social Actor Masking constrains each person's visual tokens to attend only to their own textual descriptions via a spatiotemporal mask, avoiding actor-action mismatch and disordered social dynamics. Directional Reweighting amplifies attention to directional words (e.g., "leftward", "right"), leading each action towards its intended target. To evaluate generated social interactions, we annotate existing datasets with interaction descriptions and build a fully automated evaluation pipeline powered by open-source VLMs. Experiments on different video generation models show that SocialDirector significantly improves interaction fidelity and approaches the upper bound set by real videos.


SolvMix: Learning Formulation-State Landscapes for Liquid Electrolyte Conductivity Prediction

Kexin Zhang ⋅ Minzhang Li ⋅ Guotao Qiu ⋅ Tianqi Zhao ⋅ Yuna Zhang ⋅ Yibing Shan ⋅ Ye Mei ⋅ Xushan Zhao ⋅ Jingyi Yu

Electrolyte conductivity is not a property of isolated molecules, but a response over a formulation landscape defined by solvents, salts, composition, salt mole ratio, and temperature. The same molecular component can play different effective roles as its amount and operating conditions change, while the underlying formulation state is observed only through macroscopic conductivity measurements. This makes conductivity prediction a distinctive mixture learning problem: a model must learn how molecular identities become formulation-specific states, rather than merely how molecules are represented. We introduce SolvMix, a framework for learning formulation-state landscapes for electrolyte conductivity prediction. SolvMix constructs component tokens for solvent and salt molecules, calibrates them into amount-aware component states before mixture interaction, and predicts conductivity through a temperature-formatted readout. We evaluate SolvMix on three curated electrolyte conductivity benchmarks, CALiSol-23, DiffMix, and Bamboo-mixer, covering diverse solvents, salts, compositions, salt mole ratios, and temperatures. Across standard splits on all three benchmarks and binned out-of-distribution (OOD) splits on Bamboo-mixer, SolvMix outperforms strong molecular-representation, mixture-aggregation, and geometric-interaction baselines under the reported protocols. Ablations show that the best-performing formulation-state construction combines amount-and-type-aware component tokens, multiplicative amount calibration, token-level interaction followed by mean pooling, and late-concat temperature formatting. OOD results further show clear gains under formulation-level shifts, where static component descriptors and local interaction patterns are least aligned with the required generalization. These results suggest that electrolyte conductivity prediction benefits from treating formulation-state construction as a first-class modeling problem, beyond molecular representation learning or explicit local interaction modeling.


SPACE: Unifying Symmetric and Asymmetric Routing Problems for Generalist Neural Solver

Rongsheng Chen ⋅ Changliang Zhou ⋅ Canhong Yu ⋅ Yuanyao Chen ⋅ Yu Zhou ⋅ Zhuo Chen ⋅ Zhenkun Wang

Generalist neural routing solvers have shown great potential in solving diverse vehicle routing problems (VRPs) with a unified model. However, existing solvers are typically limited to symmetric settings or degrade in performance when switching to asymmetric settings due to input inconsistencies or inherent structural differences, substantially limiting their practicality in real-world scenarios that encompass both scenarios. To address this limitation, we define the spatial position of each node based on the relative distances to a specific set of pivots and further propose a Spatial Pivot-Aligned Coordinate-free Embedding (SPACE) framework that unifies node representation and solution generation across symmetric and asymmetric VRPs. Specifically, we construct a bidirectional Fréchet representation using a novel furthest pivot sampling strategy to enable invariant node representations across distinct problem settings. Furthermore, we introduce a weight-decomposed adaptive decoding mechanism that decouples geometric perception from problem representations, mitigating the overfitting of constraint decisions to a specific geometry setting. Extensive experiments on 110 VRP variants, comprising 55 symmetric problems and their asymmetric counterparts, demonstrate that SPACE achieves promising zero-shot generalization in both symmetric and asymmetric VRPs.


Sparsely-Supervised Data Assimilation via Physics-Informed Schrödinger Bridge

Dohyun Bu ⋅ Chanho Kim ⋅ Seokun Choi ⋅ Jong-Seok Lee

Data assimilation (DA) for systems governed by partial differential equations (PDE) aims to reconstruct full spatiotemporal fields from sparse high-fidelity (HF) observations while respecting physical constraints. While full-grid low-fidelity (LF) simulations provide informative priors in multi-fidelity settings, recovering an HF field consistent with both sparse observations and the governing PDE typically requires per-instance test-time optimization, which becomes a major bottleneck in time-critical applications. To alleviate this, amortized reconstruction using generative models has recently been proposed; however, such approaches rely on full-field HF supervision during training, which is often impractical in real-world settings. From a more realistic perspective, we propose the Physics-Informed Conditional Schrödinger Bridge (PICSB), which transports an informative LF prior toward an auxiliary-energy-defined, observation-conditioned HF endpoint law without any additional inference-time guidance. To enable learning without HF endpoints, PICSB employs an iterative surrogate-endpoint refresh scheme, and directly incorporates PDE residuals into the training objective while enforcing observations via hard conditioning throughout sampling. Experiments on fluid PDE benchmarks demonstrate that PICSB enables extremely fast spatiotemporal field reconstruction while achieving strong accuracy under sparse HF supervision.

Adapting large-scale vision-language models (VLMs) such as CLIP to video understanding has drawn increasing attention. Prompt tuning offers an efficient and promising way to adapt VLMs to this task. While existing methods propose to enrich textual information over simple textual templates via large language models (LLMs), the generated descriptions are often coarse and may lead to semantic misalignment with the specific video. To address this issue, we propose $\textbf{S}$patial-$\textbf{T}$emporal $\textbf{A}$ttributes enhanced prompt $\textbf{W}$eighting and $\textbf{F}$usion (ST-AWF), a novel parameter-efficient fine-tuning framework for various video recognition tasks. In our method, LLM is leveraged to generate spatial and temporal attributes for each action category, enriching text prompts with informative descriptions. To mitigate attribute irrelevance in LLM outputs, we design a vision-guided prompt weighting mechanism that dynamically evaluates the relevance of each attribute embedding based on its similarity to video features, thereby emphasizing discriminative cues. Furthermore, a dual-attention cross-modal fusion module is proposed to align weighted textual prompts of spatial–temporal attributes with video features for fine-grained cross-modal feature enhancement. Extensive experiments demonstrate that ST-AWF achieves state-of-the-art performance on multiple video recognition benchmarks, consistently improving both discriminability and generalization ability to unseen classes, while maintaining high efficiency with a small number of trainable parameters. Code is available at https://anonymous.4open.science/r/ST-AWF-code/


SpecLoR: Spectral Lookahead Rectification for Motion-Coherent Text-to-Video Generation

Xu Zhang ⋅ Yu Lu ⋅ Ruijie Quan ⋅ Zhaozheng CHEN ⋅ Bohan Wang ⋅ Yi Yang

Flow Matching has enabled robust text-to-video generation via latent ODE sampling. However, velocity approximation and numerical discretization errors inevitably accumulate, causing sampling trajectories to drift. Consequently, generated videos often suffer from severe spatiotemporal inconsistencies, such as physically implausible motions. Yet, directly correcting these drifted noisy latents is challenging: timestep-dependent noise obscures reliable structural cues, and spatial interventions risk disrupting fragile local geometry while incurring heavy computational costs. To address this, we propose Spectral Lookahead Rectification (SpecLoR), a plug-and-play inference method that bypasses noise via lookahead prediction. It also circumvents spatiotemporal entanglement by shifting corrections to the frequency domain, where universal statistical natural-video priors are readily available. First, during early sampling stages, SpecLoR looks ahead to estimate the clean latent $z_{t,0}$ and computes its 3D spatiotemporal spectrum. Next, SpecLoR rectifies only the amplitude to match the statistical prior, leaving the phase intact. Finally, the corrected state is re-noised to resume ODE integration. Experiments on Wan2.2 demonstrate that SpecLoR significantly reduces physical artifacts and enhances motion coherence across multiple benchmarks with minimal computational overhead (4 additional NFEs in a 40-step schedule). Code will be released.


Spectral Feedback for Test-Time Alignment of Protein Diffusion Models

Shai Dickman ⋅ Mert Cemri ⋅ Landon Butler ⋅ Kannan Ramchandran

Alignment methods for discrete diffusion models have primarily focused on steering the denoising process, either by influencing token logits or by selecting favorable sequences at intermediate steps. However, these approaches largely treat inference as a unidirectional process, lacking effective mechanisms for revisiting undesirable token selections. We introduce Spectral Feedback, an algorithm that uses a feedback loop that selects sets of positions to edit, allowing the model to iteratively correct its own generations. This approach leverages the masking structure of discrete diffusion models where tokens can be re-masked and re-sampled, analogous to image editing methods that reintroduce noise and re-run a reverse diffusion process. While prior methods focus on what token values to assign, we instead treat which tokens to revisit as the central alignment problem. Selecting edit positions is challenging because edit effects are interdependent: the impact of modifying one token depends on which others are edited simultaneously. Motivated by prior work on sparse interactions in biological systems, we show empirically that, for protein inverse folding, the expected reward over edit-position sets admits a sparse Fourier representation. This structure enables efficient learning and optimization of a value function over edit-position sets, which Spectral Feedback uses to select promising edits. The algorithm is model-agnostic and can be applied to pretrained, test-time aligned, or fine-tuned protein diffusion models, improving alignment performance without modifying the underlying generative process. Applied to inverse folding with a protein stability oracle, it achieves a 30.6% increase in stable proteins for a pretrained model, 24.8% for Best-of-10, and 8.5% for the SOTA RL-tuned diffusion models.


Spectral Transformer Neural Processes

XIANHE CHEN ⋅ Hao Chen ⋅ Yingzhen Li

Time series, spatial data, and images are natural applications of Neural Processes. However, when such data exhibit strong periodicity and quasi-periodicity, existing methods often suffer from underfitting and generalise poorly beyond the training distribution. In this work, we propose Spectral Transformer Neural Processes (STNPs), a frequency-aware extension of Transformer Neural Processes (TNPs). STNPs introduce a Spectral Aggregator that estimates an empirical context spectrum, compresses it into a spectral mixture, samples task-adaptive spectral features, and concatenates them with time-domain embeddings, thereby injecting a spectral-mixture-kernel bias into TNPs. This design reshapes the similarity geometry, allowing inputs that are distant in Euclidean space to remain close in an induced periodic manifold while enhancing time--frequency interactions. Extensive experiments on synthetic regression tasks, real-world time-series datasets, and an image dataset demonstrate that STNPs consistently improve predictive performance over existing baselines, extending Neural Processes beyond translation equivariance toward effective modelling of periodicity and quasi-periodicity.

We prove that the speedup of exact speculative decoding is governed by a per-token rate, not by the path-level prefix-overlap capacity that folklore suggests, and that the gap between the two can be unbounded. The two candidates are the per-token Leviathan rate σL(P, Q) and the path-level coupling capacity ΩL(P, Q) = Σ_{t=1}^L ov(P_{≤t}, Q_{≤t}). We show σL ≤ ΩL always, the inequality is strict already on i.i.d. products, and there exists an explicit fixed-precision binary trajectory-exact rejection-style protocol on L = 2 with E[τ] = 19/16 > 9/8 = σL, ruling out σL as a universal upper bound across all trajectory-exact rounds. Inside the deployable admissible Leviathan-style clipped-ratio class, σL is the tight rate, and the speculation decoding complexity satisfies SDC{B,L}^Q(M; u, T) = Θ(T / (σ_{L,B}(PM^{u,T}, Q) + 1)), achieved by iterated Leviathan and matched up to constants by Wald. On the path side, ΩL exhibits a sharp two-sided dichotomy in the prefix-KL profile ηt: Σt √(ηt / 2) = o(L) forces ΩL = L − o(L), while geometrically bounded overlap ot ≤ ρ^t forces ΩL ≤ 1/(1−ρ) uniformly in L, with no unconditional KL converse. Two explicit fixed-precision Transformer families of size O(log L) realize the two extremes under the same three-element proposal family Q = {Qflat, Qconst0, Qconst_1}: a constant-readout family F⁺ has SDC ≤ T / ((1−ε)L + 1), while an alternating-parity family F⁻ has SDC ≥ T/5, an unconditional Θ(L) architecture separation. Speculative decoding is not path coupling; it is per-token clipped-ratio acceptance, and the architecture of the target controls the rate.


Speech Tokenizers are Vulnerable: Transferable Semantic Attack and Robust Tokenizer

Zhisheng Zhang ⋅ Yifan Mi ⋅ Yixuan Zhou ⋅ Haiyun Li ⋅ dongfu song ⋅ Shengbo Cai ⋅ Ziang Xu ⋅ Jie Hao ⋅ Zhiyong Wu

Speech tokenizers serve as the critical interface between continuous speech signals and discrete representations, and are widely used in speech systems, e.g., Automatic Speech Recognition (ASR) and Large Audio-Language Models (LALMs). In this paper, we find that speech tokenizers are fragile at the semantic level: slight perturbations can disrupt their semantic encoding process and consequently cause severe degradation in downstream ASR and LALM performance. To study this vulnerability, we propose T-SemAttack, a transferable semantic attack based on a semantic encoder ensemble. By jointly perturbing the representation spaces of multiple semantic encoders, T-SemAttack effectively disrupts the semantic content of speech while preserving perceptual quality, thereby inducing error tokens. We further analyze token fragility and layer-wise representation drift in relation to cross-model transferability and disruption of LALM attention. Our analysis reveals a progressive amplification chain in which small waveform perturbations are magnified through the tokenizer pipeline and lead to semantic collapse, e.g., the attack causes a 99.7% token change in the S3 tokenizer. Building on these findings, we introduce ROSETok, a robust speech tokenizer that combines noisy training with robust semantic distillation to improve reconstruction fidelity and downstream-task robustness. Extensive experiments on multiple datasets, 9 open-source and black-box ASR systems, and 4 LALMs demonstrate that T-SemAttack achieves strong transferable attack performance, while the proposed Robust Speech Tokenizer exhibits robustness under high-fidelity reconstruction and downstream tasks. Our code and demo are available at https://t-semattack.github.io.


SPHERE-JEPA: Spherical Prediction with Homogeneous Embeddings

Léo Nicollier ⋅ Max Dunitz ⋅ Marc Pic ⋅ Pablo Muse ⋅ Enric Meinhardt-Llopis ⋅ Gabriele Facciolo

A fundamental open question in self-supervised learning (SSL) is the explicit characterization of the optimal geometry of the learned representations. Recently, LeJEPA identified isotropic Gaussian embeddings as optimal for minimizing downstream prediction risk in Euclidean spaces. However, the corresponding problem for distributions supported on lower-dimensional manifolds, such as the hypersphere, remains unexplored. In this work, we demonstrate that extending this minimax analysis to smooth distributions on Riemannian manifolds fundamentally changes the optimal solution. We show that, under a worst-case formulation, both \(k\)-nearest neighbors and kernel ridge regression induce hyperspherical uniformity. More precisely, we show that uniform distributions on manifolds are optimal for \(k\)-nearest neighbors, and that the uniform distribution on the sphere is optimal for kernel ridge regression with both the exponential dot-product kernel and the linear kernel. This theoretical insight reveals a fundamental limitation of Gaussian embeddings: their non-uniform density induces anisotropic $k$-NN neighborhoods, severely biasing the estimator. To correct this, we introduce \textbf{SPHERE-JEPA}, a theoretically grounded SSL framework. We adapt LeJEPA's Cramér--Wold projection mechanism to enforce hyperspherical uniformity rather than a Gaussian prior. Empirically, SPHERE-JEPA yields significant improvements, boosting texture retrieval mAP by over 6\%, while consistently matching or outperforming LeJEPA on standard benchmarks—including a $+1.8\%$ linear probing gain on ImageNet-1K (ViT-B/14).


SphereVAD: Training-Free Video Anomaly Detection via Geodesic Inference on the Unit Hypersphere

Chao Huang ⋅ Pengfei Wei ⋅ Wei Wang ⋅ Jie Wen ⋅ Zhihua Wang ⋅ Li Shen ⋅ Wenqi Ren ⋅ XIAOCHUN CAO

Video anomaly detection (VAD) aims to automatically identify events that deviate from normal patterns in untrimmed surveillance videos. Existing methods universally depend on large-scale annotations or task-specific training procedures, severely limiting their rapid deployment to novel scenes. We observe that intermediate-layer features of pre-trained multimodal large language models (MLLMs) already encode rich anomaly semantics, yet existing approaches rely on the language output pathway and fail to exploit the geometric discriminability latent in these representations. Based on this finding, we propose SphereVAD, a fully training-free, zero-shot VAD framework that recasts anomaly discrimination as von Mises–Fisher (vMF) likelihood-ratio geodesic inference on the unit hypersphere, unleashing latent discriminability through principled geometric reasoning rather than learning new representations. Specifically, SphereVAD first applies Fréchet mean centering to unfold feature distributions and eliminate domain biases, then employs Holistic Scene Attention (HSA) to reinforce feature consistency using cross-video priors, and finally performs vMF-guided Spherical Geodesic Pulling (SGP) to align ambiguous segments with directional prototypes on the spherical manifold. This training-free pipeline requires only minimal synthetic images for calibration. SphereVAD establishes new state-of-the-art results among training-free approaches on three major benchmarks and remains competitive with fully supervised baselines. Code will be available upon acceptance.


Spherical Bayesian Experimental Design for Active View Selection in 3D Gaussian Splatting

Yan Song ⋅ Yunfei Guo ⋅ Guyeli Yang ⋅ Li He ⋅ Chenglong Li ⋅ Huanzhen Wang ⋅ Bailin HE ⋅ Qiao Sun ⋅ Zhong Jiangwei ⋅ Shizeng Zhang ⋅ Nuo Chen ⋅ Yan Wang ⋅ Wenqiang Zhang

Active view selection for 3D Gaussian Splatting (3DGS) requires information-theoretic grounding, computational efficiency, and strong reconstruction quality, yet existing methods satisfy these only partially. Information-theoretic approaches quantify gain over the full parameter space via costly Fisher computation under a Laplace approximation; faster alternatives sacrifice the grounding they started from. We propose SpheriBED, a Bayesian experimental design framework that defines the information objective over per-Gaussian spherical appearance—the quantity each observation fundamentally measures—rather than the full parameter vector, yielding exact closed-form posteriors without Laplace approximation. Scoring a candidate view requires only a single forward rasterization pass with no backward pass or per-candidate optimization. Extensive experiments across benchmarks demonstrate that SpheriBED consistently improves active reconstruction performance and provides more reliable uncertainty quantification compared to state-of-the-art methods.


Spherical Interpolation for Backward-Compatible Multimodal Representations

Simone Ricci ⋅ Niccolò Biondi ⋅ Federico Pernici

Contrastive vision-language models map visual and textual representations in a shared normalized embedding space, making cosine similarity the natural metric for cross-modal retrieval. A practical challenge arises during model upgrades: independently trained models generally produce incompatible representation spaces, so replacing a deployed model typically requires recomputing embeddings for the entire gallery, which is prohibitively expensive at scale. Orthogonal post-hoc alignment mitigates this problem by mapping new-model queries into the old-model gallery space while preserving the new learned representation. However, because independently trained models can differ in fine-grained representation structure, the orthogonal alignment remains approximate, leaving a residual angular discrepancy between the old-model query and the aligned new-model query. We study whether spherical linear interpolation (SLERP) between these two normalized query representations can improve retrieval without re-indexing the gallery. We formalize the geometric conditions under which post-alignment interpolation yields a query direction closer to a task-optimal direction than either endpoint, and connect this characterization to Recall@$K$ through a local margin-based certification result. Experiments across multiple benchmarks and model families show that SLERP improves over orthogonal alignment alone, suggesting these conditions are broadly met in practice.


SphMind: Towards Robust, Training-Free VLM-based Spatial Reasoning with a 360 Camera

Shriram Damodaran ⋅ Soumyaratna Debnath ⋅ Cheston Tan ⋅ Lin Wang

Omnidirectional or 360° cameras provide embodied AI agents with a holistic, wide field-of-view (FoV) perception of the 3D structure of their surroundings. This capability has generated significant interest in applying Multi-modal Large Language Models (MLLMs) to omnidirectional spatial reasoning. However, most MLLMs are primarily trained on conventional 2D perspective images and therefore struggle with the severe distortions and wrap-around discontinuities introduced by spherical geometry. As a result, enabling MLLMs to generalize effectively to non-Euclidean 3D spaces without retraining remains an open challenge. In this paper, we propose SphMind, a novel, training-free, and plug-and-play framework designed to bridge this gap. The central idea of SphMind is to decouple semantic perception from geometric reasoning. Instead of requiring MLLMs to learn complex spherical geometric principles internally, the framework preserves their strong semantic understanding while handling geometric reasoning externally. To achieve this, we introduce a Spherical Harmonics-based Spatial Graph (SHSG) that models spatial relationships using equivariant transformations on the sphere. We further integrate this with Inference-Time Geometric Grounding (IGG), a model-agnostic closed-loop optimization process that aligns the internal representations of MLLMs with spherical geometric constraints during inference. Extensive experiments on three benchmark datasets demonstrate the effectiveness of SphMind. Without any additional training, the framework achieves more than 21.4% average improvement in directional reasoning on MP3D and Stanford2D-3D, outperforms all prompt-engineering baselines by 8.7% on the real-world ODI-Bench dataset, and achieves 5.9× higher rotational invariance under panorama rotations compared to existing baselines, all without dataset-specific tuning. We also validate the framework on real-world in-the-wild captures, where SphMind successfully resolves directional reasoning queries that baseline vision-language models fail to answer correctly.


SPHQuant: Efficient extreme low bit weight quantization for Vision-Language Models

Kewei Zhang ⋅ 郑 陈 ⋅ Haotong Qin ⋅ Yulun Zhang

Recent foundation models are moving toward native multimodal Vision-Language Models (VLMs), making VLMs a central form of next-generation foundation models. However, their large language backbones make edge deployment difficult due to high memory footprint and memory-bound autoregressive decoding. Weight-only post-training quantization is a practical solution, but pushing VLMs to extreme low bit-widths remains challenging: existing rotation-free methods suffer from outliers at 2--3 bits, while rotation-based methods improve accuracy at the cost of additional runtime overhead. We propose SPHQuant, a rotation-free spherical weight-only quantization framework for VLMs. Instead of quantizing weights directly in Cartesian coordinates, SPHQuant decomposes each 8D weight vector into coordinate signs, radius, and a positive unit direction. This representation isolates outlier magnitude into the radius while keeping directions bounded and statistically regular. Based on this insight, SPHQuant allocates extra precision to the radius to mitigate accuracy degradation induced by outliers. It further uses a compact positive-direction codebook and fine-tunes codebook entries through angular parameterization to preserve the unit-sphere constraint. We also design a hardware-friendly GEMV kernel that keeps the direction codebook small enough for shared-memory lookup and packs radial bits efficiently. Experiments show that SPHQuant matches the performance of state-of-the-art extreme low-bit quantization methods while improving decode throughput by \textbf{30.3\%}, providing a new step toward efficient edge deployment of VLMs.


SpikeSTAG: A Dendritic Compartmental Spiking Graph Network for Multivariate Time-Series Forecasting

Bang Hu ⋅ Changze Lv ⋅ Junyi Wang ⋅ mingjieli ⋅ Xiaoqing Zheng ⋅ Fan Zhang ⋅ wei cao

Spiking Neural Networks (SNNs) offer a distinctive paradigm for temporal modeling through their intrinsic membrane potential dynamics. However, existing SNN-based forecasting methods focus exclusively on temporal processing, lacking mechanisms to capture spatial dependencies among variables. To bridge this gap, we propose SpikeSTAG, a neuromorphic spatiotemporal architecture that integrates graph-based spatial reasoning into spike-driven temporal computation. Central to our approach is the Dendritic Graph Module, which reinterprets spectral graph convolutions as biological dendritic integration -- coupling multi-scale spatial aggregation with the temporal dynamics of spiking neurons within a single computational unit. To adaptively fuse the resulting spatial and temporal representations, we further introduce a Dual-Stream Integration Module that employs competitive gating and coincidence-based amplification, selectively enhancing consistent spatiotemporal patterns while suppressing conflicting noise. Extensive experiments demonstrate that SpikeSTAG establishes a new state-of-the-art among SNN-based models and surpasses representative ANN baselines, while reducing theoretical energy consumption by approximately 33.3\% compared to Transformer architectures.


SPIRAL: Self-Evolving Action-Conditioned Video Generation via Reflective Planning Agents

Yu Yang ⋅ Yue Liao ⋅ Jianbiao Mei ⋅ Baisen Wang ⋅ Jiangning Zhang ⋅ Xiangtai Li ⋅ Liang Lv ⋅ Hanlin Chen ⋅ Yong Liu ⋅ Shuicheng Yan ⋅ Gim Hee Lee

Long-horizon action-conditioned video generation aims to synthesize temporally coherent videos that follow complex action instructions over extended horizons. Existing single-shot video generation models typically operate in an open-loop manner, leading to incomplete action execution, hallucinated motions, and temporal drift. To address this, we propose SPIRAL, a closed-loop framework that performs sequential planning and iterative reflection for action-conditioned long-horizon video generation. Specifically, a PlanAgent decomposes a high-level goal into sub-actions that condition video generation, while a CriticAgent evaluates intermediate video segments and provides corrective feedback for iterative refinement. This closed-loop design further supports self-evolving, utilizing planning and verification signals for GRPO-based post-training to enhance the video generator's consistency and action quality over extended horizons. Moreover, we introduce ActVideoGen-Dataset and ActVideoGen-Bench for training and evaluation. Experiments across multiple TI2V backbones with self-evolving show consistent gains on ActVideoGen-Bench and VBench, demonstrating the effectiveness of SPIRAL.

Flow matching is an emerging, scalable generative framework for characterizing continuous normalizing flows with wide-range applications. However, state-of-the-art methods are not well-suited for modeling dynamical systems, as they construct conditional paths that are restricted to linear interpolants, a suboptimal supervision signal, which may not capture the system's underlying state evolution. Moreover, constructing unified paths to satisfy multi-marginal constraints across observations is challenging, since naïve higher-order polynomials tend to be unstable and oscillatory. To address these limitations, we introduce SplineFlow, a theoretically grounded flow matching algorithm that jointly models conditional paths across observations via B-spline interpolation. Specifically, SplineFlow exploits the smoothness and stability of B-spline bases to learn the complex underlying dynamics in a structured manner while ensuring the multi-marginal requirements are met. Comprehensive experiments across deterministic and stochastic dynamical systems under various configurations, as well as cellular trajectory inference tasks, demonstrate that SplineFlow outperforms existing baselines, especially when the underlying dynamics are of higher degree or in irregular sampling regimes.

Looped transformers promise test-time compute scaling by spending more iterations on harder problems, but it remains unclear which architectural choices let them extrapolate to harder problems at test time rather than memorize training-specific solutions. We introduce a fixed-point based framework for analyzing looped architectures along three axes of stability -- reachability, input-dependence, and geometry -- and use it to characterize when fixed-point iteration yields meaningful predictions. Theoretically, we prove that looped networks without recall have countable fixed points and cannot achieve strong input-dependence at any spectral regime, while recall with outer normalization reliably produces a regime in which fixed points are simultaneously reachable, locally smooth in the input, and supported by stable backpropagation. Empirically, we train single-layer looped transformers on chess, sudoku, and prefix-sums and find that downstream performance tracks the framework's predictions across tasks and architectural configurations. We additionally introduce internal recall, a novel recall placement variant allowing an identity residual path, and exploit its contrasting fixed-point behavior to test our framework's predictions in a controlled setting.


Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Jiakang Wang ⋅ Runze Liu ⋅ Fuzheng Zhang ⋅ Xiu Li ⋅ Guorui Zhou ⋅ Kun Gai ⋅ Ling Pan

Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training method for improving the reasoning abilities of Large Language Models (LLMs). However, existing methods mainly apply uniform optimization constraints across all tokens, ignoring their heterogeneous roles. Prior work shows that high-entropy tokens are closely tied to reasoning, while low-entropy tokens primarily encode factual knowledge, and recent approaches attempt to exploit this distinction by isolating token updates via masking or asynchronous training. We argue that such isolation breaks the sequential dependency structure of autoregressive generation, leading to suboptimal learning. To address this, we propose \textbf{Archer}, an entropy-aware RLVR framework with \textbf{dual-token constraints} that preserves joint optimization while modulating update strength across token types. Our method introduces response-level entropy normalization for stable token classification and applies differentiated clipping ranges and KL regularization to encourage exploration on reasoning tokens while preserving knowledge tokens. Experiments on mathematical reasoning and code generation benchmarks show that Archer consistently outperforms strong baselines across multiple model scales, improving both \textit{pass@1} and \textit{pass@K} performance. These results highlight the importance of respecting sequence-level dependencies when designing fine-grained RL optimization strategies for LLMs.

RL+Search is the computational backbone of modern AI systems for imperfect-information extensive-form games. The Recursive Belief-based Learning (ReBeL) framework already mitigates one source of training fragility by sampling the pivot iteration uniformly from the iteration budget; we identify a second, hitherto unaddressed source: the budget itself is held fixed across all subgames, creating a rigid coupling between solver maturity and outer-loop learning that propagates variance through the recursive self-play and produces noisy, seed-dependent training trajectories. We propose \textbf{Stochastic Horizon Annealing (SHA)}, which randomizes only this remaining static degree of freedom: at each subgame the total CFR iteration count is drawn from a Gaussian centered at the expected budget. Combined with ReBeL's existing uniform pivot, the procedure smooths the learning objective across a continuum of solving depths and behaves as an implicit ensemble at zero additional cost. SHA is a one-line change to any ReBeL-style trainer. We give a complete theoretical treatment with three results, all proved in the main text: (i) SHA preserves the $O(N^{-1/2})$ convergence rate of linear CFR; (ii) the per-subgame target variance is bounded by $C_\beta^2\sigma^2/(4N^3)$, vanishing rapidly with the budget; and (iii) the recursive variance of the value-network targets scales as $\Theta(D)$ for SHA versus $\Theta(D^2)$ for the fixed-horizon ReBeL, where $D$ is the recursion depth. On Liar's Dice 1$\times$4f, 1$\times$5f, and 1$\times$6f (30 seeds each) at the ReBeL reference setting of depth-2 subgames and 1024 search iterations, SHA reduces mean final exploitability by 9--18\% and shrinks the 95\% confidence interval across seeds by 34--63\%, with larger games benefiting more as predicted by the theory, establishing stability as a first-class, theoretically supported contribution for RL+Search on imperfect-information games.


Stable Max Coverage Under a Cardinality Constraint

Themistoklis Haris ⋅ Fabian Spaeh ⋅ Nithin Varma ⋅ Yuichi Yoshida

We propose a stable algorithm for solving maximum coverage problems under a cardinality constraint of $k$ in a record-level stability model, where adjacent instances $I\sim I'$ differ by the incidence of one universe element. Our algorithm computes a negative-entropy-regularized maximizer of the capped concave relaxation over the hypersimplex and rounds it with a common-seed rounding map from hypersimplex correlated sampling. The resulting sensitivity bound is $\min\{2k, \alpha_k k/\lambda\}$, where $\lambda$ is the regularization parameter and $\alpha_k = O(\log k)$. For utility, the expected coverage has approximation factor $1-\frac{1}{e}$ with an additive error of order $O(\lambda k \log \frac{m}{k})$, where $m$ is the number of input sets. We also report experiments on synthetic and real-world instances.


State of Thought Enables Endogenous Reasoning

Zhiren Gong ⋅ Yikun Hou ⋅ Zihao Zeng ⋅ Ming Xiao ⋅ Chau Yuen ⋅ Wei Yang Bryan Lim

Test-time compute has emerged as a major approach to improving the capabilities of Large Language Models (LLMs). However, existing test-time reasoning paradigms rely heavily on externally imposed control, either through fixed reasoning programs or through costly expansion in constrained search spaces, limiting both generalization and efficiency. We propose State of Thought (SoT), a new reasoning paradigm that enables endogenous reasoning in LLMs, with the model's internal reasoning state governing how reasoning unfolds. Concretely, SoT extracts a compact dynamics-geometric state from the model's internal information transfer and selectively activates historical reasoning support useful under the current reasoning state, framing reasoning as a state-conditioned process over evidence rather than an externally prescribed token chain. Across quantitative (1.29x), general (1.62x), symbolic and code (1.72x), long-context (2.63x), and multimodal (1.08x) reasoning tasks on 4 models with 20 datasets, SoT consistently improves task accuracy while reducing tokens by 69.0% and latency by 48.7%, supporting endogenous state-driven reasoning as a more generalizable and efficient alternative.

Reservoir computing turns sequence learning into linear regression on fixed dynamical features, but those features come from a single highly dependent trajectory, so the true amount of usable data is often unclear. We develop a transport-based learning theory for contractive echo state networks driven by stochastic inputs by viewing the reservoir as a Markov process and proving a Wasserstein contraction that guarantees a unique stationary state law and explicit mixing rates. Using contraction-to-concentration tools, we obtain finite-sample deviation bounds for time-averaged features and empirical covariance estimates, and translate them into excess-risk and stability guarantees for ridge-trained linear readouts. The theory yields an explicit effective-sample-size principle and exposes a sharp trade-off: making reservoirs more “critical” can increase dynamical memory while slowing mixing and raising the data requirements for reliable training. Overall, the results complement echo-state and capacity analyses by providing verifiable design rules that link stability parameters to mixing time, generalization, and forecasting robustness under independent or weakly dependent inputs.


ST-Bridge: Bridging Sketch and Text with Large Language Models for Coarse-to-Fine Image Retrieval

Haoxiang Hu ⋅ Cangjun Gao ⋅ Zhengming Zhang ⋅ Yaxian Shan ⋅ Zeyuan Huang ⋅ Ran Zuo ⋅ Qiang He ⋅ Qingkun Li ⋅ Xiaoming Deng ⋅ Cuixia Ma ⋅ Yu-Kun Lai ⋅ Yong-jin Liu ⋅ Hongan Wang

Sketch-Text image retrieval aims to find natural images using both free-hand sketches and text descriptions. Although these two modalities provide complementary information, the task is challenging due to their large representational differences and the difficulty of matching fine details in complex scenes. Most existing methods rely on simple feature fusion or global alignment, which often fails to preserve important discriminative cues and leads to inaccurate retrieval results. To tackle these issues, we propose ST-Bridge, a high-precision sketch-text image retrieval framework that enables reliable coarse-to-fine matching through an effective semantic bridging mechanism. Specifically, we introduce LLM as a semantic bridge to generate sketch representations that are closer to the textual semantic space. We further enhance these representations and the original sketch features via gated attention and feature normalization, substantially reducing the sketch--text modality gap. Building upon this, we design a coarse-to-fine alignment strategy that supports accurate sketch--text image retrieval. Extensive experiments on scene-level sketch--text image retrieval benchmarks demonstrate that ST-Bridge significantly outperforms existing methods and achieves state-of-the-art performance, validating the effectiveness of large-model-based semantic bridging for cross-modal retrieval tasks. Our code and models will be released upon acceptance.


Steering Fields: Adaptive Vector Fields for Safe Image Generation and Beyond

Simone Facchiano ⋅ Jan Eric Lenssen ⋅ Bernt Schiele ⋅ Wolfgang Stammer ⋅ Fabio Galasso ⋅ Jonas Fischer

As state-of-the-art text-to-image flow models have matured to deliver near-photorealistic quality, controlling what they generate -- e.g., inhibiting harmful content while promoting benign alternatives -- has become a central challenge. The current steering paradigm consists of adding a global steering vector to selected model activations. While functional, a fixed vector applied example-agnostic and uniformly along the entire trajectory cannot adapt to the changing state of the generation, and causes uncontrolled global changes beyond the targeted concepts. We introduce Steering Fields, a generalization of steering vectors that adaptively re-estimates the steering direction at each step of the generative process. Steering Fields operate on latent representations, expose a continuous trade-off between steering strength and content preservation, and are compositional, enabling the simultaneous induction and inhibition of concepts, setting a new state of the art on safety steering benchmarks. Although our formulation prescribes no explicit spatial mask or object-level prior, the trajectory-adaptive estimation naturally preserves local structure, in a manner reminiscent of image editing. In fact, when applied in the VAE latent space, Steering Fields can serve as a structure-preserving image-editing technique that achieves state-of-the-art semantic fidelity (CLIP, VQAScore), while remaining model-agnostic and inversion-free


SteerVTE: Seamless Video Text Editing with Style and Glyph Control

Kai Zeng ⋅ Moran Li ⋅ Zhengwei Wang ⋅ Yingchen Yu ⋅ Yiheng Lin ⋅ Ruichuan An ⋅ Ming Lu ⋅ Qi She ⋅ Wentao Zhang

Visual text editing aims to precisely modify text in images and videos while preserving stylistic consistency and visual realism. Despite significant advances in the image domain, video text editing remains largely unexplored: it is a localized task demanding stroke-level precision within small text regions, which compounds the challenges of cross-frame accuracy, temporal coherence, and stylistic fidelity. We introduce SteerVTE, a unified framework that \underline{\textbf{steer}}s a frozen video diffusion model to perform precise \underline{\textbf{V}}ideo \underline{\textbf{T}}ext \underline{\textbf{E}}diting through style and glyph control. Built on a frozen diffusion transformer, SteerVTE attaches a lightweight text context adapter with two complementary modules: a style encoder capturing the original text's visual attributes, and dual-granularity glyph encoders encoding the target text at both the line and character levels. To overcome the inherently weak text rendering priors of video foundation models, we further propose a glyph-aware spatial-focal loss and a three-stage progressive training curriculum that scales from image to video data. To support large-scale training, we also develop an automatic synthesis pipeline and construct SteerVTE-1M, a dataset of one million triplets spanning diverse scenes, fonts, and stylistic effects. Extensive experiments demonstrate that SteerVTE substantially outperforms existing video editing baselines across text accuracy, style consistency, and temporal coherence. Our anonymous project website is available at \url{https://steervte-paper.github.io}.


STEMFly: Enhancing UAV Vision-Language Navigation via Sensor Grounding, Temporal Diversity and Episodic Memory

Guangdao Zhu ⋅ Xu Chen ⋅ Shuhong Hou ⋅ Weili Guan ⋅ Bin Chen ⋅ Yaowei Wang ⋅ Xiang Deng

Despite rapid progress in indoor vision-and-language navigation (VLN), its aerial counterpart remains significantly more challenging and underexplored. In this setting, unmanned aerial vehicles (UAVs) must interpret free-form instructions and traverse kilometer-scale outdoor environments. Although recent UAV-VLN methods have shown promising results, they often fail to fully exploit contextual signals across modality, time, and prior experience. This limitation appears in three aspects: the predominant reliance on RGB-only inputs, the use of discrete frame inputs that fail to capture temporal dynamics, and the absence of episodic memory, resulting in unreliable termination decisions. To this end, we propose $\textbf{STEMFly}$, a unified framework grounded in $\textbf{S}$ensor signals, $\textbf{T}$emporal diversity, and $\textbf{E}$pisodic $\textbf{M}$emory. STEMFly comprises three components. First, Sensor-Augmented Prompt Injection (SAPI) incorporates multi-modal sensor signals as semantically grounded language prompts that can be interpreted by the language model. Second, Temporal-Diversity Frame Selection (TDFS) constructs a temporally informative observation sequence via entropy-guided frame sampling. Third, Memory-Augmented Success Verification (MASV) improves termination reliability by validating navigation outcomes through episodic memory retrieval. In addition to methodological contributions, we identify and rectify systematic annotation inconsistencies in the OpenUAV benchmark, releasing a revised evaluation protocol. On the original OpenUAV benchmark, STEMFly outperforms the prior state-of-the-art TravelUAV, achieving absolute improvements of 8.26\% in Success Rate (SR) and 6.70\% in Success weighted by Path Length (SPL), while reducing Navigation Error (NE) by 9.55 meters. Similar consistent gains are also observed on our rectified benchmark. These results demonstrate that holistic contextual grounding, which integrates multi-modal sensing, temporal modeling, and episodic memory, is crucial for improving UAV-VLN performance.


Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces

Nicolás Astorga ⋅ Nabeel Seedat ⋅ Mihaela van der Schaar

Verifiable reward training has improved mathematical and coding reasoning, but these domains capture only part of step-by-step decision making. Many real-world tasks require finding a high-value feasible plan among many valid alternatives. We introduce $\text{OPT}\star$, a scalable family of optimization-style tasks for training and evaluating LLM step-by-step optimization-like reasoning along a complexity axis: each task provides a feasibility checker and evaluator, while a complexity parameter expands the search space without requiring new human labels. This motivates studying these tasks in two regimes: (i) solver-guided online policy optimization, which uses a solver as a value oracle for partial states and applies rank-based reward shaping to reinforce better next steps, and (ii) search-based offline RL when such solvers are unavailable. Theoretically, we relate success in large search spaces to the information a reasoner extracts per unit of search budget. Empirically, we ablate the ingredients that make search efficient on $\text{OPT}\star$ and show that training on $\text{OPT}\star$ improves step-by-step optimization-like reasoning.

Activation steering is a lightweight way to control pretrained generative models: a learned map inserted into the model's activations can steer generation toward target behavior without updating the model weights. While effective for single target behavior, existing steering methods remain brittle when multiple behaviors are targeted. A style intervention followed by a safety intervention can produce a different result than the same applied in the reverse order, as the second intervention is evaluated on activations already shifted by the first. We trace this failure to coordinate sharing: existing transport-based steering methods learn different concepts in the same activation coordinates, so their interventions interfere by construction. To address this, we introduce StiCAS (Stifefel Compositional Activation Steering), a geometric steering framework that learns a separate orthogonal subspace for each concept. Each concept is steered by an affine transport restricted to its own subspace, and orthogonality between subspaces makes these transports non-interfering by construction. The concept frame is trained directly on the Stiefel manifold, so Riemannian optimization preserves the orthogonality required for composition throughout learning. This structure gives an exact commutativity guarantee: any sequential ordering of single-concept interventions reproduces the simultaneous multi-concept update. Across 45 concept pairs and three open models, StiCAS drives order-dependent interference exactly to zero and roughly doubles joint-concept success over the strongest activation-transport baseline.


Stitched Value Model for Diffusion Alignment

Hyojun Go ⋅ Hyungjin Chung ⋅ Prune Truong ⋅ Goutam Bhat ⋅ Li Mi ⋅ Zhaochong An ⋅ Zixiang Zhao ⋅ Dominik Narnhofer ⋅ Serge Belongie ⋅ Federico Tombari ⋅ Konrad Schindler

For practical use, diffusion- or flow-based generative models must be aligned with task-specific rewards, such as prompt fidelity or aesthetic preference. That alignment is challenging because the reward is defined for clean output images, but the alignment procedure requires value function estimates at noisy intermediate latents. Existing methods resort to Tweedie-style or Monte Carlo approximations, trading off estimator bias against computational cost: Tweedie estimates are efficient but biased, while Monte Carlo estimates are more accurate but require expensive rollouts. A natural alternative would be a learned value function, but it remains an open question how to effectively train a strong and general value model specifically for noisy latents. Here, we propose StitchVM (Stitched Value Model), a model stitching framework that efficiently transfers reward models pretrained for clean images to the noisy latent regime. StitchVM starts from an existing, truncated pixel-space reward model and attaches a frozen diffusion backbone to it as its head. From the pixel-space model, the resulting hybrid retains a carefully pretrained, robust reward capability; from the diffusion backbone, it inherits its native ability to handle noisy latents. The stitching procedure is exceptionally lightweight, e.g., stitching and finetuning CLIP ViT-L and SD 3.5 Medium takes only $\approx$10 GPU-hours. By lifting powerful pixel-space reward models to latent space, StitchVM opens up a new style of diffusion alignment: instead of rough, yet costly per-sample approximation of the value function, the correct function for the actual, noisy latents is constructed once and then amortized over many samples and iterations. We show that this approach yields improvements across a broad range of downstream steering and post-training methods: DPS becomes $3.2\times$ faster while halving peak GPU memory, and DiffusionNFT becomes $2.3\times$ faster.


STParOpt: An End-to-End Framework for Execution-Aware Parallelism Inference and Optimized CUDA Migration

Meng Wang ⋅ Jinshuo Liu ⋅ Weiran Pian ⋅ Feiyang Wen ⋅ Xinya Liu ⋅ Juan Deng ⋅ Jeff Pan

Automatically migrating serial C programs to parallel CUDA is essential for high-performance computing, yet remains challenging due to the partial observability of data dependencies: critical memory access patterns cannot be reliably inferred from source code alone. Existing static analysis methods often fail on irregular and pointer-intensive loops, while LLM-based approaches cannot infer memory-access-dependent optimization strategies from source text alone. To address these limitations, we propose STParOpt, which formulates parallelism and optimization inference as an uncertainty-aware learning problem under partial observability, where ambiguity in static analysis is explicitly quantified and resolved using execution signals. STParOpt augments static program graphs with stride-level execution features via cross-attention, and employs an entropy-guided fusion mechanism that up-weights dynamic evidence when the static branch's predictive entropy over parallelizability is high. Experiments on standard HPC benchmarks demonstrate that STParOpt significantly outperforms both compiler-based and LLM-based baselines on parallelism detection and CUDA code generation. Notably, STParOpt achieves 96.4% parallelism detection accuracy and 86.1% end-to-end generation correctness, surpassing the strongest baseline by +5.0 pp and +11.4 pp respectively.


StreamPhy: Streaming Inference of High-Dimensional Physical Dynamics via State Space Models

Panqi Chen ⋅ Yifan Sun ⋅ Shikai Fang ⋅ Xiao Fu ⋅ Lei Cheng

Inferring the evolution of high-dimensional and multi-modal (e.g., spatio-temporal) physical fields from irregular sparse measurements in real time is a fundamental challenge in science and engineering. Existing approaches, including diffusion-based generative models and functional tensor methods, typically operate in offline settings, depend on full temporal observations, or incur substantial inference cost. We propose StreamPhy, an end-to-end framework that enables efficient and accurate streaming inference of full-field physical dynamics from incoming irregular sparse measurements. The framework integrates a data-adaptive observation encoder that is robust to arbitrary observation patterns, a structured state-space model that supports memory-efficient online updates across irregular time intervals, and an expressive Functional Tensor Feature-wise Linear Modulation (FT-FiLM) decoder for continuous-field generation. We prove that FT-FiLM is more expressive than the functional Tucker model, admitting a richer function class for handling complex dynamics. Experiments on three representative physical systems under challenging sampling patterns show that StreamPhy consistently outperforms state-of-the-art baselines, with at least 48\% improvement in accuracy and up to 20--100x faster inference than diffusion-based methods.


StructLens: A Structural Lens for Language Models via Maximum Spanning Trees

Haruki Sakajo ⋅ Frederikus Hudi ⋅ Yusuke Sakai ⋅ Hidetaka Kamigaito ⋅ Taro Watanabe

Language exhibits inherent structures, a property that explains both language acquisition and language change. Given this characteristic, we expect language models to manifest their own internal structures as well. While interpretability research has investigated how models compute representations mechanistically through attention patterns and Sparse AutoEncoders, the organization of the resulting representations is overlooked. To address this gap, we introduce StructLens, a framework to analyze representations through a holistic structural view. StructLens constructs maximum spanning trees based on the semantic representations in residual streams, inspired by tree representation in dependency parsing, and provides summaries of token relationships in representation space. We analyze how contiguous tokens are also nearby in representation space and find that middle layers show the strongest local-span organization. Moreover, analysis of pre-training checkpoints reveals that smaller local units become detectable earlier in pre-training, and larger units later. Our findings demonstrate that StructLens provides insights into how models organize token representations across layers and training.

A learned latent space supports reliable scientific decisions only if decisions made in that space remain valid in the original physical variables. We study latent Kalman-type data assimilation, where an encoder--decoder pair is trained without knowledge of the sensor used to collect observations. Can representation-level diagnostics predict which sensors give accurate physical-state filtering? In general, no. We show that any such diagnostic (a smooth functional of the averaged decoder Fisher matrix, including optimal-design criteria, posterior Cram\'er--Rao surrogates, and effective-observability scores) sees the sensor only through a \emph{Fisher--Gram coarsening} of the physical sensor Gramian, leaving a large blind spot. The failure is unconditional, an algebraic consequence of averaging the decoder Jacobian over the latent invariant measure. The positive side is conditional: when the proportionality between physical and latent filtering error is stable across sensors, latent filtering error ranks sensors accurately, with an explicit Pearson bound. Decoder smoothness provides one route to this stability, and a linear-Gaussian Riccati surrogate (built entirely from the trained latent model and decoder) empirically tracks the same ranking when full-state truth is unavailable. Experiments on five chaotic systems support both sides. Representation geometry is useful but structurally limited, and reliable sensor ranking in latent data assimilation requires filter-aware validation rather than decoder-diagnostic scoring.


Structure-aware Reinforcement Learning for Protein Directed Evolution

Zikun Nie ⋅ Suyuan Zhao ⋅ Yizhen Luo ⋅ Siqi Fan ⋅ Zaiqing Nie

Protein optimization remains a longstanding goal in life sciences. Existing machine learning–assisted directed evolution (MLDE) methods primarily rely on sequence-only features, overlooking the critical spatial constraints and co-evolutionary interactions encoded in protein structures. However, directly integrating structural information remains challenging due to the scarcity of reliable mutant structures. To address these issues, we propose StructEvo, a novel structure-aware reinforcement learning framework for protein directed evolution. StructEvo employs a delta-structure fusion encoder to approximate mutant structure features via feature differences, enabling dynamic incorporation of spatial knowledge. The vast mutation space is then decomposed into manageable subspaces through a structure-aligned hierarchical action network, while a geometric constraint further stabilizes delta feature learning. Our approach outperforms prior state-of-the-art methods by 9.2% and 16.3% on two challenging optimization benchmarks, and further identifies an experimentally validated epistasis pattern in GFP, highlighting the importance of structural guidance for effective protein directed evolution.

Evaluating large language models across many benchmarks is expensive, yet many benchmarks are highly correlated. We formalize the selection of a small, informative subset as submodular maximization under a multivariate Gaussian model. Entropy (log-determinant covariance) and mutual information between selected and remaining benchmarks arise as natural objectives. Both are submodular; entropy selection coincides with pivoted Cholesky and has spectral residual bounds, while mutual information is non-monotone in general but empirically monotone for small subsets, so we optimize it greedily. Experiments on three matrices from ten public leaderboards show that mutual information selection outperforms entropy for imputation at small subsets.


Subprocess-Constrained Markov Decision Processes

Jiarui Gan ⋅ Debmalya Mandal

Constrained Markov decision processes (CMDPs) extend standard MDPs by incorporating a cost function and enforcing constraints---most commonly by requiring that the expected cumulative cost over the entire trajectory remains below a prescribed threshold. While this global, expectation-based formulation is natural and widely studied, it is insufficient for many safety- and reliability-critical applications, where guarantees are required not only from the initial state but also for every subprocess starting at any intermediate time step. Motivated by this need, we introduce \emph{subprocess-constrained MDPs} (S-CMDPs), in which the expected cumulative cost must satisfy a constraint not only from the initial state, but from every intermediate step onward along the trajectory. We investigate the algorithmic aspects of this model and characterize the limitations of classical solution concepts. In particular, we show that stationary policies can be both suboptimal and computationally intractable to optimize in S-CMDPs. Leveraging a value-set iteration framework, we further demonstrate that history-dependent policies surprisingly overcome both obstacles. Building on this insight, we develop a polynomial-time algorithm that computes the action distribution of a (near-)optimal history-dependent policy for any given history.


Support Mismatch as a Benchmark Failure Mode for In-Context Prediction

Jie Liu ⋅ Lanlan Liang ⋅ Qinan Bao ⋅ ZiXi Yan ⋅ Linghao Meng ⋅ Wenbo Gong

We analyze support mismatch in language-model in-context prediction as a frozen-prior Bayes benchmark with finite discrete support. The goal is diagnostic rather than literal: we do not identify production LMs with exact Bayesian updating, but characterize a failure mode whose signatures can be checked in learned in-context predictors. Our main theorem gives the finite-sample predictor-risk bound $R_k^\pi \le C e^{-k\rho_{\min}} + 2\varepsilon_{\mathrm{approx}}^2$, so risk decomposes into an exponentially decaying transient and a nonvanishing support-mismatch floor under exponential-moment control of log-likelihood ratios. In scalar and aligned Gaussian location families, exact two-component analyses match the upper exponent up to a tight factor 2 and show a regime change at $\mu_2=3\mu_1$. Controlled learned-model experiments track the benchmark near crossover. Decoder-only LM diagnostics then test the benchmark signatures: restricted support creates a persistent high-error regime; restoring the missing support sharply lowers error; and modern size-matched distractor controls on Qwen and Llama show that adding a wrong same-cardinality candidate does not reproduce this gain. Multiple-choice, instruction-tuned, cross-family, calibration, null-control, and near-tie checks preserve the same diagnostic picture.


SurgVista: Long-Horizon Surgical World Modeling with Plausible Instrument-Tissue Dynamics

Wentao Pan ⋅ WUYANG LI ⋅ Shengyuan Liu ⋅ Xinyu Liu ⋅ Hengyu Liu ⋅ Yixuan Yuan

Scaling robot policy learning for autonomous surgery is challenging, as expert demonstrations are expensive and in vivo exploration poses substantial safety risks. Surgical world models address this by generating realistic, action-conditioned future frames from an initial observation, but existing methods exhibit two persistent failure modes: spatial interaction incoherence, where visible instrument contact fails to induce spatially consistent tissue deformation, and temporal fidelity collapse, where prediction errors compound across autoregressive rollouts and progressively corrupt visual quality. We present SurgVista, a surgical world model that mitigates both failures through two training recipes. Deformation Consistency Regularization extracts scene-point trajectories from training videos and enforces cross-frame coherence through latent contrastive learning, strengthening physically consistent instrument-tissue dynamics. Drift Adaptation Training mitigates long-horizon drift by perturbing conditioning frames with online prediction residuals and photometric augmentations calibrated to long-horizon drift statistics, sustaining visual fidelity over extended rollouts. To enable rigorous evaluation, we further introduce SurgWorld-Bench, featuring diverse procedure types, long-range rollouts, and decoupled metrics for instrument-motion accuracy and tissue-response fidelity. Extensive experiments show that SurgVista consistently outperforms state-of-the-art methods across visual quality, temporal consistency, and interaction fidelity, with gains widening as the prediction horizon grows.


SWE Atlas: Benchmarking Coding Agents Beyond Issue Resolution

Mohit Raghavendra ⋅ Soham Dan ⋅ Miguel Romero Calvo ⋅ Yannis He ⋅ Johannes B Mols ⋅ Gautam Anand ⋅ Cole McCollum ⋅ Edgar Arakelyan ⋅ Vijay Bharadwaj ⋅ Andrew Park ⋅ Jeff Da ⋅ MohammadHossein Rezaei ⋅ Bing Liu ⋅ Brad Kenstler ⋅ Yunzhong He

We introduce SWE Atlas, a benchmark suite for coding agents spanning three professional software engineering workflows: Codebase Q&A (124 tasks), Test Writing (90 tasks), and Refactoring (70 tasks). SWE Atlas differs from prior SWE benchmarks in three key ways: it targets underrepresented but practically important task categories, uses comprehensive category-specific evaluation protocols, and adopts under-specified, agentic task formulations that better reflect real-world usage. Its evaluation framework combines programmatic checks with rubric-based assessment. This goes beyond functional correctness, evaluating software engineering quality, including test and refactor completeness, maintainability, reusable abstractions, and codebase hygiene. We evaluate a range of frontier and open-weight models on SWE Atlas and find that GPT-5.4 and Opus 4.7 achieve the strongest overall performance, while even the best open-weight models score poorly. Our analysis suggests that top models rely on extensive codebase exploration and runtime-driven reasoning. However, even top models consistently struggle with subtle edge cases, complex runtime analysis, and adherence to software engineering best practices. Overall, SWE Atlas provides a complementary evaluation suite for measuring both correctness and engineering quality in coding agents.


SWE-Chain: Benchmarking Coding Agents on Chained Release-Level Package Upgrades

Man Ho Lam ⋅ Chaozheng Wang ⋅ Hange Liu ⋅ Jingyu Xiao ⋅ Haau-Sing Li ⋅ Jen-Tse Huang ⋅ Terry Yue Zhuo ⋅ Michael R Lyu

Coding agents powered by large language models are increasingly expected to perform realistic software maintenance tasks beyond isolated issue resolution. Existing benchmarks have shifted toward realistic software evolution, but they rarely capture continuous maintenance at the granularity of package releases, where changes are bundled, shipped, and inherited by subsequent versions. We present SWE-Chain, a benchmark for evaluating agents on chained release-level package upgrades, where each transition builds on the agent's prior codebase. To produce upgrade specifications, we design a divide-and-conquer synthesis pipeline that aligns release notes with code diffs for each version transition, ensuring the requirements are grounded in actual code changes, informative to agents, and feasible to implement. SWE-Chain contains 12 upgrade chains across 9 real Python packages, with 155 version transitions and 1,660 grounded upgrade requirements. Across nine frontier agent-model configurations, agents achieve an average of 44.8% resolving, 65.4% precision, and 50.2% F1 under the Build+Fix regime, with Claude-Opus-4.7 (Claude Code) leading at 60.8% resolving, 80.6% precision, and 68.5% F1. These results show that SWE-Chain is both feasible and discriminative, and reveal that current agents still struggle to make correct upgrades across chained package releases without breaking existing functionality.


Symb-xMIL: Symbolic Explanations for Multiple Instance Learning in Digital Pathology

Yanqng Luo ⋅ Julius Hense ⋅ Niklas Prenißl ⋅ Andreas Mock ⋅ Klaus-Robert Müller ⋅ Thomas Schnake ⋅ Mina Jamshidi Idaji

Explanations of multiple instance learning (MIL) models are widely used for validation and discovery in digital histopathology. Existing methods primarily rely on heatmaps that highlight influential regions but do not explain how evidence from different tissue regions is combined to produce a prediction. This limits interpretability, especially when decisions depend on interactions between tissue features. We introduce Symbolic explainable MIL (Symb-xMIL), a post-hoc explanation framework that quantifies how a MIL model’s behavior aligns with human-readable decision rules, expressed as logical relationships (e.g., AND, OR, NOT) between input features. These alignment scores reveal semantic patterns underlying the model’s predictions. We evaluate Symb-xMIL on synthetic and real-world histopathology datasets. On synthetic MIL data, Symb-xMIL reliably recovers ground-truth logical rules. In a clinical tumor detection task, the best-aligned rules uncover heterogeneous decision patterns and expose hidden model errors. On an HPV-prediction task on TCGA-HNSCC, a cohort of head and neck cancer, our framework refines patient survival stratification beyond HPV status with potential clinical relevance. Overall, Symb-xMIL extends MIL explainability beyond visual attribution toward structured, rule-based reasoning, enabling more transparent and semantically grounded interpretation of model predictions.


Symmetric Interventions for Eliciting Model Intent

David Vella Zarb ⋅ Rustem Turtayev ⋅ Taywon Min ⋅ Jinghua Ou ⋅ Shi Feng

Alignment faking, where models comply with harmful requests more often when they infer they are being trained than when they infer they are deployed, has been read as evidence of scheming. The same behavior is also consistent with performative scheming: the model identifies the setup as a safety evaluation and produces the behavior it infers researchers are looking for. These hypotheses predict identical behavior in the standard evaluation but generalize differently. We distinguish them by amplifying and suppressing the process each requires (consequence-tracking for scheming, expectation-tracking for performative scheming) and comparing using three methods: synthetic document fine-tuning, activation steering, and prompt interventions. Across four open-weight model organisms that exhibit alignment faking in the helpful-only scenario, fine-tuning produces larger shifts under expectation-tracking than consequence-tracking on three (Llama-3.1-70B, Llama-3.1-405B, Qwen-2.5-72B); on QwQ-32B, both intervention families shift compliance globally rather than changing the gap, suggesting they affect overall helpfulness rather than alignment faking specifically. Activation steering on Llama-3.1-70B reproduces the asymmetry: expectation-tracking steering changes compliance, while consequence-tracking steering changes the model's reasoning without affecting compliance. Alignment faking in these models is better explained by expectation-tracking than by self-preservation reasoning.


SympFNO: Structure-Preserving Fourier Neural Operators for Physical Surrogate Modeling

Luong Doan ⋅ Duc H Nguyen ⋅ Khanh N Quoc ⋅ Phan Quoc Hung Mai ⋅ Ngoc Mai Vu ⋅ Bui T Hieu ⋅ Nhung Duong ⋅ Phong Ho ⋅ Long H Dang ⋅ Tuan Do

Neural operators can efficiently learn solution maps for partial differential equations, but standard architectures often violate fundamental physical structure during inference, leading to unstable long-horizon rollouts. We introduce the Symplectic Fourier Neural Operator (SympFNO), which is composed of three symplectic layers that embed structure preservation directly into Fourier-space and physical-space transformations, thereby preserving the geometric structure of Hamiltonian dynamics. Our theory provides a rigorous foundation for the architecture by showing that the structure-preserving design is mathematically consistent with Hamiltonian evolution and inherently more efficient than unconstrained operator parameterizations. Across five well-known Hamiltonian PDE problems, SympFNO achieves stable long-term rollouts, substantially lower trajectory error, more accurate conservation of physical invariants, and up to two orders of magnitude fewer parameters than state-of-the-art baselines, namely FNO and PINO.


Symplectic Neural Operators for Learning Infinite Dimensional Hamiltonian Systems

Makara Yeang ⋅ Yusuke Tanaka ⋅ Takashi Matsubara ⋅ Takaharu Yaguchi

The modeling and simulation of infinite-dimensional Hamiltonian systems are central problems in mathematical physics and engineering, but they pose significant computational and structural challenges for standard data-driven architectures. In this work, we introduce the Symplectic Neural Operator, a neural operator architecture designed to preserve the symplectic structure intrinsic to Hamiltonian PDEs. We provide a theoretical characterization of their symplecticity and establish a rigorous long-term stability result based on the combination of symplectic structure preservation and learning accuracy. Numerical experiments on canonical Hamiltonian PDEs corroborate this theoretical result and show that SNOs exhibit improved energy behavior compared with non-structure-preserving neural operators.


SynDORBench: Evaluating LVLM Perceptual Robustness Under Physically Constrained Visibility Conditions

Jeremy S G Yee ⋅ Daniel Zhengkui Wang ⋅ Zhiyuan Zhang ⋅ Avinash Anand ⋅ Timothy Liu ⋅ Benedict Chan ⋅ Aik B Ng ⋅ Simon See

Large vision-language models (LVLMs) have demonstrated remarkable performance on multimodal reasoning benchmarks, yet their perceptual reliability under physically constrained imaging conditions remains poorly understood. Existing evaluations predominantly assume ideal visual inputs and therefore fail to characterize how camera distance, illumination, viewpoint, and pixel density fundamentally affect semantic recoverability. We introduce SynDORBench, the first physically grounded benchmark for evaluating LVLM perceptual robustness under DORI-calibrated conditions aligned with human visual capability standards. SynDORBench comprises over 54k question--answer pairs generated through a controllable synthetic pipeline that systematically varies viewing distance, lighting, camera geometry, and action pose according to physically interpretable pixel-density regimes. To support scalable low-visibility supervision, we further propose a discernibility annotation framework that propagates human perceptual labels using mask-conditioned statistical features and ensemble learning. We evaluate 16 open-source LVLMs, a commercial LVLM baseline, and YOLO11x across human-presence classification and action recognition tasks under progressively degraded visibility conditions. Our results reveal that perceptual failure in LVLMs is strongly governed by pixel density and physical imaging constraints rather than model scale alone. Surprisingly, several compact open-source LVLMs outperform larger commercial baselines and substantially exceed YOLO11x robustness under long-range and low-light conditions. SynDORBench establishes a new benchmark paradigm for physically grounded multimodal evaluation, enabling systematic analysis of LVLM reliability under real-world perceptual constraints and direct comparison against human visibility thresholds.

Uniform-noise discrete diffusion and flow models generate sequences non-autoregressively through iterative, context-dependent token replacements. However, these models are typically formulated as time-inhomogeneous CTMC/DTMC processes, sampled using independent Bernoulli change decisions per discretization step. This induces Poisson-binomial variance in per-position jump counts that grows with the number of required edits, leading to the common under-editing (residual noise) and over-editing (cascading substitutions) failure modes that degrade sample quality, especially under tight discretization budgets. We identify this sampler-induced variance as an orthogonal source of degradation, distinct from model-side errors and addressable purely at inference time. We propose Systematic Hazard Sampling (SHS), a training-free, drop-in, and hyperparameter-free inference principle for any sampler that admits a stay-vs.-replace decomposition. SHS models per-token edits as events driven by cumulative hazard (CTMC) or jump mass (DTMC) and triggers an edit whenever this quantity exceeds unit-spaced thresholds with a single random phase per position. For any fixed cumulative mass, this preserves the expected jump count while achieving the minimum conditional variance possible among unbiased integer estimators (at most \(1/4\)), without altering per-jump destination sampling. Experiments on four uniform-noise discrete diffusion and flow language models spanning $\sim$110M to $\sim$3B parameters show that SHS consistently improves sample quality across NFE budgets, with the gains growing with NFE as predicted by the variance gap.

Predictive models on structured tabular data often face imbalanced targets: training emphasizes common values while rare outcomes stay underrepresented. Distributionally robust optimization hedges against shift, but standard ambiguity sets use a single radius over the entire distribution, so adversarial mass concentrates on dense regions and rare target ranges remain weakly protected; worst-case statements then track bulk behavior rather than tails. We introduce target-conditional distributionally robust optimization (TC-DRO), which defines ambiguity per target region with radii that grow where data are scarce. We establish coverage, per-region excess risk that tightens with local sample size when standard DRO yields vacuous tail guarantees, and a dual as weighted empirical risk with theory-derived density-adaptive weights. Experiments on seven tabular benchmarks show that Wasserstein DRO collapses to empirical risk minimization under uniform weights, while global $\chi^2$-DRO can remain close to ERM in practice; TC-DRO improves tail MSE and SERA relative to both standard training and recent imbalanced-regression methods.


TanGCE: Manifold-Aware Concept Erasure

Matan Avitan ⋅ Yoav Goldberg ⋅ Yanai Elazar

Concept erasure aims to remove a target attribute from a representation while preserving the other information encoded in it. This is difficult beyond the linear setting: a target signal hidden from one probe may remain recoverable by a fresh nonlinear probe, while unconstrained nonlinear updates may remove the target by pushing representations off the manifold of natural hidden states. We propose the Manifold Representation Hypothesis (MRH): natural hidden states concentrate near a structured, lower-dimensional manifold, so surgical erasure should act along local manifold degrees of freedom rather than arbitrary ambient directions. We operationalize this hypothesis with the TanGCE family of erasure methods. The core method estimates each representation's local tangent space from nearest-neighbor secants, projects a nonlinear concept-scorer gradient onto that tangent, and applies a per-sample trust-region step; TanGCE+ and TanGCE++ prepend closed-form first- and second-moment erasers before the same manifold-constrained loop. Across 119 settings spanning 13 language models, three NLP concepts, and 40 CelebA-CLIP attributes under two control regimes, TanGCE++ removes more target signal than the strongest published baseline at matched control-damage budgets under one fixed hyperparameter recipe. Applying the same manifold-constrained loop after prior erasers also consistently reduces residual nonlinear leakage without leaving the control budget, empirically supporting MRH's operational prediction that manifold-constrained edits are a useful inductive bias for surgical concept erasure.


TAPIOCA: Why Task- Aware Pruning Improves OOD model Capability

Krish Sharma ⋅ Omar Naim ⋅ Soumadeep SAHA ⋅ Vinija Jain ⋅ Aman Chadha ⋅ Nicholas Asher

Recent work has promoted task-aware layer pruning as a way to improve model performance on particular tasks, as shown by TALE. In this paper, we investigate when such improvements occur and why. We show first that, across controlled polynomial regression tasks and large language models, such pruning yields no benefit on in-distribution (ID) data but consistently improves out-of-distribution (OOD) accuracy. We further show empirically that OOD inputs induce layerwise norm and pairwise-distance profiles that deviate from the corresponding ID profiles. This leads to a geometric explanation of task-aware pruning: each task induces a task-adapted geometry, characterized empirically by the representation profiles observed on ID inputs. OOD inputs can introduce a distorted version of the task-adapted geometry. Task-aware pruning identifies layers that create or amplify this distortion; by removing them, it shifts OOD representational norms and pairwise distances toward those observed on the adapted distribution. This realigns OOD inputs with the model’s task-adapted geometry and improves performance. We provide causal evidence through controlled distribution shifts and residual-scaling interventions, and demonstrate consistent behavior across model scales.


Task Success Is Not Enough: Side-effect-Aware Evaluation of Tool-Using Language Model Agents

Jiaju Huang ⋅ Shaobin Chen ⋅ Xinglong Liang ⋅ Xinyu Ma ⋅ Hao YANG ⋅ Yue Sun ⋅ Wei Ke ⋅ Tao Tan

Tool-using language model agents are usually evaluated by whether they finish the user's task. That metric misses a common failure: an agent can complete the requested work while also taking an unauthorized side-effecting action, such as writing to the wrong object, expanding the scope of the request, or contacting an unintended recipient. We call this failure mode collateral damage (CD). We introduce SABench (Side-effect-Aware Benchmark), a benchmark of 201 email, calendar, and ticket tasks. Each task crosses one of four state structures (independent, dependent, externally reachable, cross-tool) with one of four trap families (ambiguity, mis-targeting, cascade harm, overreach), and each task has a deterministic trace-based oracle for both goal satisfaction and authorization compliance. Across 2,211 model-task pairs from 11 frontier LLMs spanning major commercial and open model families, 15.2% of all runs incur CD, with risk concentrated in ambiguity tasks (38.8% violation rate) while non-ambiguous traps average 8.4% CD. Among goal-satisfied runs, 12.8% still produce unauthorized writes. Cross-tool and dependent states have about twice the CD rate of independent states. CD rates vary by 4.7x across models at comparable goal rates, which suggests that task success alone is a poor proxy for safe tool use. The SABench task corpus is released at https://huggingface.co/datasets/anon-sabench-2026/sabench; the traces, oracle, code, and analysis pipeline are released at https://anonymous.4open.science/r/sabench-artifact.


Temporal Backtracking Search for Test-time Generative Video Reasoning

SeJoon Jun ⋅ Zheng Ding ⋅ Huangyuan Su ⋅ Weirui Ye ⋅ Yilun Du

While test-time scaling has revolutionized reasoning in large language models, generative video reasoning remains bottlenecked by a single-shot paradigm. We demonstrate that searching over denoising steps cannot rescue logically flawed rollouts because spatial trajectories commit early in the diffusion process. Root-level Best-of-$N$ (BoN) sampling is similarly inefficient: reasoning errors cluster early in the temporal axis, and resampling blindly discards verified upstream progress. To unlock effective test-time scaling for video models, we introduce \textbf{Temporal Backtracking Search (TBS)}, which shifts the search space to the temporal axis. TBS transforms video generation into an iterative generate--verify--restart loop via three core mechanisms: (1) \textit{variable-$K$ conditioning} to resume generation from arbitrary clean prefixes; (2) \textit{temporal process verification} to localize failures and extract valid restart anchors; and (3) \textit{prefix-based search} to reallocate compute toward extending correct trajectories rather than root resampling. Across algorithmic, navigation, and robotics domains, TBS Pareto-dominates matched-budget BoN. In a strict out-of-distribution setting where one-shot generation collapses ($0.7\%$ for BoN), TBS achieves $22.7\%$, with every solved episode stemming from a restarted branch. Ultimately, TBS reveals that the local reasoning competence of video models far exceeds what single-shot rollouts indicate, providing a scalable test-time framework to unlock it.

Reinforcement learning from formal specifications often separates task definition, reward shaping, curriculum design, and policy diagnostics into different mathematical objects. This fragmentation introduces additional conversion steps and discards structure already present in the specification, such as sub-task boundaries, ordering constraints, and quantitative margins. We introduce TBTCL, a framework that uses Temporal Behaviour Trees (TBTs) as a unified interface for specification-guided reinforcement learning. A single TBT specifies the task, induces dense robustness-based rewards, defines a curriculum, and supports better understanding of learnt policies - grounded in the formal semantics of the specification language. By adapting TBT syntax and semantics for specification guided reinforcement learning, we provide a formalism that is more expressive than popular specification languages such LTL and inherently hierarchical. We present algorithms to automatically extract a curriculum DAG and derive dense, robustness-based rewards directly from TBT semantics, requiring no additional user inputs like region predicates or demonstrations. Furthermore, we utilize TBT trace segmentation to provide build understandable competence profiles and automatically adapt the reward toward bottlenecks. We evaluate TBTCL on six specifications across discrete and continuous domains. The results show that TBTCL matches dense STL shaping on simpler tasks and improves over baselines on reach-avoid specifications, while also exposing interpretable sub-task failures through the same formal object used for training. We show that a Temporal Behaviour Tree act as a specification, and a reward, and a curriculum and a diagnostic.

Recent advances in generative video models have significantly improved visual realism, making detection increasingly dependent on subtle temporal inconsistencies arising from imperfect frame-to-frame coherence. However, these signals are weak and spatially localized, while most of the video content is redundant. This creates a fundamental mismatch: existing detectors rely on heavy backbones to model full videos, assuming increased capacity can capture such cues, which leads to high computational cost and limited scalability. In this work, we rethink AI-generated video detection from a representation perspective. Instead of modeling entire videos, we propose to explicitly restructure them into compact temporal slice representations that isolate informative temporal evolution. Based on this idea, we introduce a lightweight framework using frequency-aware temporal slice consistency learning, where informative spatial rows and columns are aggregated over time to form slice-domain inputs. This reparameterization suppresses redundant appearance while preserving discriminative temporal dynamics, enabling efficient and targeted detection. By constructing slices from both high- and low-frequency regions, our method exposes complementary temporal artifacts, including unstable details and abnormal motion smoothness. We instantiate this idea in a $\textit{Temporal Slice Consistency Network}$, which integrates task-driven slice localization, hierarchical positional encoding, and a lightweight Transformer to model cross-slice and cross-frequency dependencies with minimal overhead. We further introduce $\textit{AI-Artist}$, a new benchmark with artist-curated videos from recent high-fidelity generators, including Seedance and Kling. Experiments on GenVideo and AI-Artist show that our method achieves strong performance while requiring over 400× fewer FLOPs and 3.6× faster inference than the strong video-based baseline ReStraV, enabling scalable and practical deployment.


TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows

Shuai Fu ⋅ Jing Gu ⋅ Jian Zhou ⋅ Zicheng Duan ⋅ Gengze Zhou ⋅ Qi Wu

Recent text-to-image models have made substantial progress in realism, aesthetics, and prompt-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or impossible spatial relationships. These failures are not well captured by existing quality, aesthetics, preference, or alignment metrics. To address this gap, we introduce TerraVis, a framework for evaluating world-grounded visual consistency in generated images. TerraVis defines a structured taxonomy of world-consistency violations covering object-, interaction-, and scene-level failures, and evaluates images through a multi-stage workflow. Specifically, TerraVis uses an MLLM to first perform eligibility checking, determining whether an image is suitable for world-consistency evaluation. It then conducts taxonomy-guided violation detection over fine-grained violation types and classifies detected violations as minor or severe to derive an overall world-consistency score. Across diverse open-source and proprietary text-to-image models on two widely used benchmarks, TerraVis shows the strongest correlation with human judgments of world consistency among the evaluated metrics. Our results further reveal that models with high quality, aesthetics, preference or alignment scores can still exhibit frequent world-consistency failures. TerraVis therefore provides a complementary evaluation perspective, enabling fine-grained diagnosis of generated images and model comparison beyond existing evaluation dimensions.


Test-Time Graph Anomaly Detection via Shifted Augmentation with Dynamic Objective Scheduling

Chunjing Xiao ⋅ Ruobing Fan ⋅ Tian Zhao ⋅ Fan Zhou ⋅ Chong Tang ⋅ Ying Ma

Graph anomaly detection (GAD) is critical to many real-world systems. However, supervised GAD models often degrade under distribution shifts due to the closed-world training assumption. Test-time adaptation (TTA) offers a promising solution by updating pre-trained models with unlabeled test data. Yet, applying TTA to GAD poses a unique dilemma. High-confidence samples provide reliable pseudo-labels, but they are scarce and biased toward low-shift regions. Low-confidence samples better reflect target shifts, but they are too noisy for direct cross-entropy optimization and often require label-free objectives that are misaligned with anomaly discrimination. To address this dilemma, we propose SADOS, a Shifted Augmentation-based Dynamic Objective Scheduling framework for test-time GAD. SADOS turns unreliable target samples into drift-carrying signals for reliable supervised adaptation. It distills prototypes from confidence-stratified unreliable samples and injects them into reliable samples to generate label-preserving yet shift-oriented augmentations. This expands coverage over shifted target regions. In parallel, SADOS dynamically schedules the adaptation objective. It uses contrastive learning to capture early-stage drift signals, and gradually anneals its weight so that classification-oriented anomaly detection dominates later adaptation. Extensive experiments demonstrate that SADOS robustly improves test-time GAD under distribution shifts.


Test-Time Learning with an Evolving Library

Weijia Xu ⋅ Alessandro Sordoni ⋅ Chandan Singh ⋅ Zelalem Gero ⋅ Michel Galley ⋅ Xingdi Yuan ⋅ Jianfeng Gao

We introduce EvoLib, a test-time learning framework that enables large language models to accumulate, reuse, and evolve knowledge across problem instances without parameter updates or external supervision. Instead of adapting model parameters, our approach maintains a shared library of knowledge abstractions, including modular skills and reflective insights, automatically extracted from the model's own inference trajectories. To support continual improvement, we introduce a principled weighting and consolidation mechanism that jointly optimizes for immediate utility and long-term value. This allows simple, instance-specific abstractions to evolve into more general and reusable ones over time. Across challenging benchmarks in mathematical reasoning, code generation, and multi-turn agentic environments, EvoLib improves substantially over the top test-time scaling and learning methods without ground-truth feedback.


Test-time Scaling for Diffusion Language Models with Frequency-Aware Remasking

Bowen Zuo ⋅ Yue Yu ⋅ Dongruo Zhou ⋅ Yinglun Zhu

Diffusion large language models (DLLMs) generate text through iterative denoising, where the choice of which tokens to commit and which to refine drives output quality. Existing certainty-aware strategies rely on instantaneous confidence or entropy to guide this choice, but these local signals can be misleading on hard reasoning problems. This raises a natural allocation question for DLLM test-time scaling: rather than applying one fixed decoder to all inputs, can we allocate not just more compute but a different decoding strategy to questions that remain unresolved? We propose a two-stage adaptive DLLM sampling method. The first stage uses a fast base sampler to solve easy questions; the second applies frequency-aware remasking to the remaining hard questions. Our frequency-aware strategy tracks how often token predictions change across denoising steps and keeps frequently-changing tokens available for further refinement, providing a trajectory-level signal that complements confidence- and entropy-based remasking. Experiments with LLaDA-8B-Instruct and Dream-7B-Instruct show consistent gains over a strong adaptive allocation baseline: up to 3.8% (LLaDA) and 7.4% (Dream) absolute accuracy improvements on MATH-500, and up to 11.18% (LLaDA) and 10.77% (Dream) absolute coverage improvements on HumanEval. On questions that remain unsolved by the base certainty-aware strategy even after 100 samples, frequency-aware remasking recovers solutions for 6.25% to 18.75% of examples.


Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck

Davide Romano ⋅ Kanak Raj ⋅ Jerrod Parker ⋅ Daniele Giofrè

Test-time scaling (TTS) improves language model outputs by spending additional inference compute — generating multiple candidates, searching over partial sequences, or iteratively refining drafts. These techniques yield large gains on mathematics and code, but have been developed and stress-tested almost exclusively on tasks where verification is straightforward. We conduct the first compute-normalised comparison of five TTS families across five open-ended generation benchmarks spanning medicine, law, finance, general chat, and creative writing — grounded in a unified framework that decomposes the effectiveness of each method's token budget into exploration and exploitation. The answer depends on which side of that decomposition you examine. Scaling exploration works: the best candidate in the pool improves steadily with compute across all settings. What breaks is exploitation — the step that converts a rich candidate pool into a final output. With state-of-the-art generators, reward models correlate at only $\hat{\rho}_v$ $\approx$ 0.12 with true quality, rendering selection near-random regardless of budget. Tree search amplifies this failure through diversity collapse. Refinement helps on one of five benchmarks; its apparent gains elsewhere are confounded. Only synthesis across candidates (Fusion) consistently improves over single-sample baselines, yet still recovers only ${\sim}$40% of available quality.The candidate pool is not the bottleneck — choosing from it is.

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos containing moments relevant to a given text query. Despite recent progress, existing PRVR methods suffer from two key limitations: a fixed video decomposition scheme that causes semantic dilution, and weak cross-domain robustness due to task-specific training. In this paper, we propose TF-PRVR, the first training-free framework for PRVR designed to improve real-world generalization. TF-PRVR leverages frozen vision-language features to construct video-specific hierarchical representations. It derives temporal semantic signals from frame-level features and applies frequency-based multi-scale analysis to identify adaptive temporal boundaries, producing hierarchical segments with coherent event-level semantics. Built on these segments, TF-PRVR constructs a unified multi-scale graph and propagates query relevance across temporally and semantically related nodes. A moment-aware scoring strategy then aggregates temporally aligned relevance across scales, emphasizing consistently supported moments while suppressing isolated false responses. Without task-specific training, TF-PRVR preserves the general-purpose alignment capability of pre-trained vision-language models and avoids dataset-specific overfitting. Extensive experiments on standard PRVR benchmarks demonstrate competitive retrieval performance and strong robustness under cross-domain evaluation, suggesting a practical direction for real-world PRVR.

Visual adversarial examples are a well-known vulnerability of deep learning (DL) systems. The emergence of vision-language models (VLMs) further expands the attack surface through multimodal interactions. Despite extensive research on adversarial defenses over the past decade, existing works on VLM robustness often overlook adaptive white-box attacks, where the adversary has full access to both the model and the defense mechanism. We show that recent defenses can be effectively bypassed under such adaptive settings and propose \method, an adaptive test-time detection method that leverages gradient information induced by a self-targeted attack maximizing the likelihood of the VLM’s generated output. Our approach consistently outperforms existing defenses across multiple victim VLMs, attack formulations, and benchmark datasets. Our code is open-source and available online.


The AI Observatory: A Public Measure of Real-World AI Use

Shayne Longpre ⋅ Anka Reuel-Lamparth ⋅ Dayeon Ki ⋅ Zhiping Zhang ⋅ Cathy Fang ⋅ Chuanyang Jin ⋅ Jennifer Mickel ⋅ Cedric Whitney ⋅ Niloofar Mireshghallah ⋅ Victor Ojewale ⋅ Megan Richards ⋅ Ariel N. Lee ⋅ Alex Pentland ⋅ Tianshi Li ⋅ Yuntian Deng ⋅ Sara Hooker ⋅ Mykel J Kochenderfer ⋅ Sanmi Koyejo

Understanding the benefits, risks, and impacts of general-purpose AI assistants requires looking at real conversations---not curated benchmarks or surveys. Yet AI use narratives rely on limited proprietary reports, or on few data sources, fragmented across platforms, models, and time. We introduce the AI Observatory, a public measurement platform that aggregates the most comprehensive set of real AI conversation sources, and develop a common taxonomy of 145 features, spanning function, topic, sensitive use, interaction style, multi-turn dynamics, and conversation structure. Across 23,158 conversations and 85,633 turns, we find that real AI use is highly heterogeneous: sources differ substantially in tasks, structures, and sensitive-use distributions, with no single source that generalizes. We further show that occupational summaries such as Anthropic's Clio capture an important but incomplete slice of use: 48\% of the conversations we analyze would be filtered out for being non-occupational, omitting substantial personal, social, cultural, and safety-relevant interaction. Real AI use is also not static: we find tokens and turns per conversation all rise sharply between 2023 and 2025, and we expose how model variants within the same developer support distinct usage regimes. The Observatory provides a new taxonomy, reusable annotation tools, and a public platform for reproducible measurement of how AI assistants are really used and evolving in the wild: \url{https://project-ai-observatory.vercel.app/}.


The Alien Space of Science: Sampling Coherent but Cognitively Unavailable Research Directions

Alejandro H. Artiles ⋅ Martin Weiss ⋅ Levin Brinkmann ⋅ Iyad Rahwan ⋅ Bernhard Schölkopf ⋅ Chris Pal ⋅ Hugo Larochelle ⋅ Anirudh Goyal ⋅ Nasim Rahaman

Scientific discovery is constrained not only by what is true, but by what is cognitively available to the researchers currently exploring a field. Many directions are coherent in light of the literature yet unlikely to be proposed because no existing community occupies the right combination of concepts, methods, and intuitions. Modern language models inherit this bias, recombining high-density regions of the literature when prompted for novel ideas. We introduce a framework that targets the complementary region, which we call the \textbf{alien space of science}, where directions are plausible under the structure of existing knowledge but unlikely under the distribution of existing researchers. Our method first decomposes papers into granular conceptual units and clusters them into a shared vocabulary of \emph{idea atoms}. It then learns two complementary models over this vocabulary. A \emph{coherence model} scores whether a combination of atoms forms a viable research direction, and an \emph{availability model} scores whether any existing author community is positioned to produce a given combination. Sampling alien directions then reduces to ranking atom combinations that maximize coherence while minimizing availability. On a corpus of 16{,}068 peer-reviewed LLM papers from NeurIPS, ICLR, ICML, and major NLP venues, the resulting sampler explores a $3.5\text{--}7\times$ broader effective atom vocabulary than frontier LLM ideation baselines without sacrificing coherence, and produces ideas that match or exceed those baselines under blind LLM, human, and downstream experimental evaluation. By separating scientific plausibility from community availability, our framework points toward AI ideation that complements rather than merely accelerates human science, expanding exploration into coherent directions the current community is unlikely to pursue.


The Alignment Illusion in Multimodal Large Language Models

Hong-Han Wang ⋅ Yuntao Wang ⋅ Hu Ding

Layer-wise visual-text similarity in Multimodal Large Language Models (MLLMs) is widely interpreted as evidence that the language model progressively integrates visual content into a shared representation space. This reading rests on the assumption that scalar alignment scores reflect content-level cross-modal interaction. To test this assumption, we apply controlled interventions to the visual stream. Across 13 MLLMs from five families spanning 0.5B to 72B parameters, replacing projector-output visual tokens with Gaussian noise sharply reduces task accuracy, yet four standard scalar measures (CKA, SVCCA, MIR, and the leading principal-angle cosine) fail to consistently separate the corrupted stream from the original. We call this failure the \emph{alignment illusion} and trace it to the shared language-model pathway: anisotropic MLP down-projections pull visual and text tokens toward common output directions, producing \emph{weight-induced alignment}. Because this component is essentially one-dimensional, we introduce the principal-angle gap (PA gap), defined as the difference between the top two principal-angle cosines, which separates weight-induced similarity from multi-directional visual structure. Under graded visual corruption, the PA gap tracks task accuracy more consistently than the scalar scores we consider; under a structured but irrelevant image, it further exposes regimes in which internal geometry and task accuracy come apart. Internal visual-text alignment in MLLMs is therefore best read as a geometric diagnostic of the visual stream inside the language model rather than a direct proxy for content-level cross-modal interaction, and is most informative when calibrated by controlled task evidence.

Prior work localizes the alignment tax---instruction tuning's degradation of structured generation---to transformer layers, but cannot determine which projections within a layer carry the functional change. We identify a consistent sub-layer hierarchy across three primary SwiGLU decoder-only families (7--9B parameters, with Yi-1.5-9B supplementary): among equal-parameter MLP projections, $W_{down} > W_{up} > W_{gate}$ (2.5:1.5:1)---an architecturally grounded ordering (reproduced under continual pretraining) not resolved by concurrent work. Within attention, $W_V$/$W_O$ gradient magnitudes exceed $W_Q$/$W_K$ by 1.76--3.91$\times$. Three independent methods---weight patching, gradient analysis, and training-time exclusion---converge on the same within-layer structure; a continual-pretraining control confirms the within-MLP hierarchy is architectural (not alignment-specific), but the layer-level distribution of harmful modifications is---this task-specific signal enables targeted rollback (8.9$\times$ over architectural-prior heuristics on held-out data, $p=0.002$). Surgical Alignment Reversal (SAR), a training-free rollback of attribution-ranked components, recovers 47% of the alignment tax (cross-entropy metric, Qwen; Llama/Mistral: 2--3$\times$ random) on held-out cross-dataset data (in-domain ratios 2.8--5.3$\times$); Output-Constrained DPO (OC-DPO) largely eliminates the tax while Q/K exclusion does not---both validated on attribution-independent data. Safety evaluation shows no significant degradation: refusal-rate $|\Delta| \le 1.4$ pp ($n=1{,}000$), no significant ASR increase under adversarial attack ($|\Delta\text{ASR}| \le 6$ pp, $n=50$); Yi-1.5-9B is excluded due to safety--structure overlap ($-13.6$ pp). Attribution replicates at 14B; gradient asymmetry persists at 72B.


The Attacker in the Mirror: Breaking Self-Consistency in Safety via Anchored Bipolicy Self-Play

Gabriele La Malfa ⋅ Emanuele La Malfa ⋅ Saar Cohen ⋅ Jie Zhang ⋅ Michael Luck ⋅ Michael Wooldridge ⋅ Elizabeth Black

Self-play red team is an established approach to improving AI safety in which different instances of the same model play attacker and defender roles in a zero-sum game, i.e., where the attacker tries to jailbreak the defender; if self-play converges to a Nash equilibrium, the model is guaranteed to respond safely within the settings of the game. Although the parameter sharing enforced by the use of the same model for the two roles improves stability and performance, it introduces fundamental theoretical and architectural limitations. We show that the set of Nash equilibria that can be reached corresponds to a broad class of behaviours that includes trivial always refuse strategies and oracle-like defenders, thus limiting practical applicability. We then show that when attacker and defender share and update the same base model, the dynamics collapse to self-consistency, so that attacks do not enforce adversarial pressure on the defender. In response, we propose Anchored Bipolicy Self-Play, which trains distinct role-specific LoRA adapters on top of a frozen base model, thereby maintaining stable optimisation while preserving adversarial pressure through explicit role separation. In relation to standard self-play, we show up to 100x greater parameter efficiency than fine-tuning and consistent improvements in safety compared to self-play fine-tuned models. We evaluate on Qwen2.5-{3B, 7B,14B}-IT models across widely used safety benchmarks, showing improved robustness without loss of reasoning ability. Cross-play experiments further show that our attacker and defender models are superior to self-play in terms of adversarial defence and safety.

Certifying authorities for safety-critical software audit recovered trace matrices module by module, not on average. Every existing distribution-free calibration delivers a marginal certificate, and we show the gap to the auditor's quantity is structural and unbounded. Under within-block exchangeability with cross-block heterogeneity in the conditional null score distribution, every procedure controlling marginal $\mathrm{FDR}$ at level $\alpha$ on a bipartite trace matrix with $K$ source modules admits worst-module false discovery proportion at least $1{-}\alpha{-}O(1/K)$ with probability at least $1{-}2/K$. We name this the block catastrophe and close it with $\mathrm{CERTRA}$, which wraps any black-box scorer, computes a within-block rank-indicator conformal $e$-value, and applies $e$-$\mathrm{BH}$ separately within each source module. The construction yields finite-sample per-module $\mathrm{FDR}$ control under arbitrary intra-module dependence, with matching power up to a rank-quantization factor. We release IndTrace-2026, a 12{,}406-link composite across aerospace, automotive, and medical-device specifications with professional ASIL/DAL/SIL annotation (Cohen's $\kappa{=}0.86$). $\mathrm{CERTRA}$ holds worst-block $\mathrm{FDP}$ at 0.118 where split-conformal drifts to 0.413, retains 0.741 of true links versus 0.572 for the strongest fair-operating-point baseline, and reduces audit deficiency 3.3-fold. Per-module distribution-free calibration is what an industrial audit can sign.

Memory-augmented reasoning agents typically store a single realized reasoning trajectory as the unit of experience. This abstraction is brittle under stochastic decoding: the same task can induce multiple plausible reasoning paths with different decompositions, assumptions, and failure modes. We introduce Branch-Structured Distributional Memory (BranchDIME), a framework for constructing memory from branch-structured trajectory sets rather than isolated traces. BranchDIME samples short reasoning prefixes, clusters them into prefix-induced branches, expands representative prefixes into full trajectories, and distills the resulting set into reusable memories containing positive strategies, anti-patterns, and contrastive rules. Across GPQA-DIAMOND and MATH500, BranchDIME improves over single-trajectory memory and alternative trajectory-selection baselines at comparable inference latency, yielding relative gains of 8.7% and 2.2% over single-trajectory memory, respectively. Coverage analysis shows that BranchDIME spans all discovered branches, compared with 20.0% coverage for consensus selection and 64.9% for random selection, while also achieving the highest downstream accuracy. BranchDIME further transfers across tasks and models, improving over single-trajectory memory by up to 15.4% on AIME25 and OlymMATH. These results support a shift from trajectory as experience to trajectory distribution as experience: memory should capture the local reasoning landscape induced by stochastic generation, not merely one sampled path.

Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments. Existing work mostly treats memory as a supply problem, asking what experience to write, how to store it, and which entry to retrieve for the next task. Yet we still lack a clear account of how models consume retrieved memory across a multi-step action trajectory. This consumption process matters because it determines not only what memories should be retrieved, but also what models and control policies are needed to use them safely. To diagnose this process, we propose Entry--Propagation--Recovery (E-P-R), a trajectory-level framework that asks where memory first changes an action, whether that change carries forward, and whether the agent can recover after leaving a correct path. We instantiate E-P-R on WebArena and on MemTrapBench, a controlled benchmark we build to isolate these phases. We find that the main failure often begins at entry: agents adopt conflicting memory at the first exposed decision point even when the recommendation is task-wrong. Repeated exposure then amplifies this early error, while recovery after divergence is weak. Together, these effects create a compliance trap: across models, conflicting memory induces similar compliance rates, but once agents comply, their success rates collapse to a low floor. Stronger agents therefore suffer larger absolute damage because each compliance event erases more baseline capability. These results suggest that memory-augmented agents should be evaluated not only by retrieval quality or final success rate, but by how they consume memory throughout the trajectory.


The Cost of Mismatch: Noise Amplification in Zeroth-Order Reinforcement Learning

Lianmin Chen ⋅ Junbin Qiu ⋅ Chenxing Wei ⋅ Yao SHU ⋅ Kun He

Zeroth-order optimization (ZOO) has emerged as a promising approach for fine-tuning Large Language Models (LLMs) when gradient access is unavailable or memory-prohibitive. However, ZOO relies on estimating gradients via finite differences between two function evaluations at perturbation scale $\mu$, making it inherently noisier than its first-order counterpart. This raises a fundamental question: \textit{under what conditions does ZOO yield reliable gradient estimates?} In this work, we provide a rigorous variance-theoretic answer through a generic noise-amplification framework and demonstrate that the reliability of ZOO gradient estimations is structurally determined by whether the two finite-difference evaluations are performed on \textit{matched data}. Mismatched evaluations, as induced by on-policy Reinforcement Learning (RL), produce a $1/\mu^2$ oracle-noise amplification that sets an irreducible noise floor, whereas matched evaluations, as realized by offline optimization methods such as Boltzmann Targeted SFT (BOLT), eliminate this amplification entirely via noise cancellation. We further derive convergence guarantees with explicit constants that cleanly separate the two regimes. Controlled experiments on math reasoning tasks validate these findings, using ZooBOLT and ZooGRPO --- zeroth-order adaptations of BOLT and GRPO --- as representatives of the matched- and mismatched-data regimes respectively. On GSM8K with Qwen2.5-1.5B, ZooBOLT recovers $96.1\%$ of first-order BOLT's performance gains, while ZooGRPO exhibits near-zero learning despite its first-order counterpart achieving strong results. This qualitative separation persists across harder benchmarks (MATH), larger model scales (Qwen2.5-7B), and different model variants (DeepSeek-R1-Distill-Qwen-1.5B). Additionally, ZOO methods reduce peak memory to near-inference levels, validating its practicality as a memory-efficient fine-tuning paradigm.


The Dynamic-Probabilistic Consistency Gap in Chaotic Surrogate Modeling

Andre Herz ⋅ Matthijs Pals ⋅ Daniel Durstewitz ⋅ Georgia Koppe

Dynamical systems reconstruction (DSR) aims to learn surrogate models that capture the dynamics underlying time-series data. Reliably deploying these surrogates requires uncertainty estimates consistent with the learned dynamics. We expose a dynamic-probabilistic consistency (DPC) gap: the pursuit of finite-horizon probabilistic objectives can degrade dynamics or decouple predictive uncertainty from the local tangent dynamics it ought to reflect. We isolate three mechanisms behind this gap: core collapse, noise masking, and blind uncertainty. Specifically, we show that open-loop Gaussian rollout objectives can penalize Jacobian-generated covariance growth in chaotic systems, encouraging optimization shortcuts that weaken physical expansion or decouple uncertainty from it. To mitigate this gap, we propose KAFFEE (Kalman-Aware Framework For Ergodic Emulation), a differentiable extended Kalman filter-based training framework that evaluates likelihood on local predictive residuals (innovations) while transporting covariance through learned local Jacobians. On stochastic hyperchaotic Lorenz-96, KAFFEE reduces the identified failure modes, improves reconstruction of dynamical invariants relative to open-loop objectives, and maintains competitive predictive scores. We further show that the DPC gap appears when probabilistically adapting a DSR foundation model across 13 chaotic systems, where KAFFEE enables in-context Bayesian filtering while largely preserving zero-shot dynamics.


The FACTS Leaderboard: A Comprehensive Benchmark for Large Language Model Factuality

Aileen Cheng ⋅ Alon Jacovi ⋅ Amir Globerson ⋅ Ben Golan ⋅ Zhaobin Kuang ⋅ Chris Alberti ⋅ Connie Tao ⋅ Eyal Ben-David ⋅ Gaurav Singh Tomar ⋅ Lukas Haas ⋅ Yonatan Bitton ⋅ Adam Bloniarz ⋅ Aijun Bai ⋅ Andrew Wang ⋅ Anfal Siddiqui ⋅ Aravindan Raghuveer ⋅ Arturo Bajuelos Castillo ⋅ Aviel Atias ⋅ Chang Liu ⋅ Corey Fry ⋅ Daniel Balle ⋅ Deepanway Ghosal ⋅ Doron Kukliansky ⋅ Dror Marcus ⋅ Elena Gribovskaya ⋅ Eran Ofek ⋅ Honglei Zhuang ⋅ Itay Laish ⋅ Jan Ackermann ⋅ Lily Wang ⋅ Megan Risdal ⋅ Megan Barnes ⋅ Michael Fink ⋅ Mohamed Amin ⋅ Moran R Ambar ⋅ Natan Potikha ⋅ Nikita Gupta ⋅ Nitzan Katz ⋅ Noam Velan ⋅ Ofir Roval ⋅ Ori Ram ⋅ Polina Zablotskaia ⋅ Prathamesh Bang ⋅ Priyanka Agrawal ⋅ Rakesh Ghiya ⋅ Sanjay Ganapathy ⋅ Simon Baumgartner ⋅ Sofia Erell ⋅ Sushant Prakash ⋅ Thibault Sellam ⋅ Vikram R Sudarshan ⋅ Xuanhui Wang ⋅ Yaroslav Akulov ⋅ Yulong Yang ⋅ ZHEN YANG ⋅ Zhixin Lai ⋅ Zhongru Wu ⋅ Avinatan Hassidim ⋅ Fernando Pereira ⋅ Slav Petrov ⋅ Srinivasan Venkatachary ⋅ Tulsee Doshi ⋅ Yossi Matias ⋅ Sasha Goldshtein ⋅ Dipanjan Das

We introduce The FACTS Leaderboard, an online leaderboard suite and associated set of benchmarks that comprehensively evaluates the ability of language models to generate factually accurate text across diverse scenarios. The suite provides a holistic measure of factuality by aggregating the performance of models on four distinct sub-leaderboards: (1) FACTS Multimodal, which measures the factuality of responses to image-based questions; (2) FACTS Parametric, which assesses models’ world knowledge by answering closed-book factoid questions from internal parameters; (3) FACTS Search, which evaluates factuality in information-seeking scenarios, where the model must use a search API; and (4) FACTS Grounding (v2), which evaluates whether long-form responses are grounded in provided documents, featuring significantly improved judge models. Each sub-leaderboard employs automated judge models to score model responses, and the final suite score is an average of the four components, designed to provide a robust and balanced assessment of a model’s overall factuality. The FACTS Leaderboard Suite will be actively maintained, containing both public and private splits to allow for external participation while guarding its integrity. It can be found at https://www.kaggle.com/benchmarks/google/facts.


The Geometry of Agent Skills: Non-Commutative Composition in Representation Space

Junda Wu ⋅ Yifan Wang ⋅ Zihan Huang ⋅ Xunyi Jiang ⋅ Sheldon Yu ⋅ Rohan Surana ⋅ Lina Yao ⋅ Julian McAuley ⋅ Tong Yu

LLM agents increasingly rely on reusable skills for multi-step reasoning, tool use, and decision making. Yet most existing approaches represent a skill as a prompt, module, or direction in representation space. This view can be insufficient for agentic settings: (1) the behaviour induced by a skill can depend on the current context, and (2) composing two skills can yield order-dependent rollouts. We formalise a skill realization as a local map from continuous intervention coordinates to a readout representation. Its differential defines the context-dependent controllable distribution, and the residence-space norm induces a metric on this distribution. Skill objectives then define local vector fields on these controllable directions. Their \emph{Lie bracket} captures the order-dependent component of skill composition, revealing behaviour that cannot be represented by a single context-independent direction. Our central measurement is a finite-difference skill commutator, which estimates the non-commutative component of two-skill composition from paired ordered rollouts. Unlike standard Lie-bracket estimation, the relevant vector fields are induced by skill objectives and observable ordered rollouts rather than given analytically. We therefore reconstruct the commutator from paired ordered rollouts, preserving the direction of the order-dependent change in representation space. We instantiate the framework on frozen LLM agents using activation-based and prefix-based skill realizations. Same-skill activation and prefix Jacobians are more aligned than cross-skill pairings (Cohen’s $d=-0.76/-0.26$), but they assign different control costs; separately, the measured two-skill interaction predicts ordered-composition gaps (Spearman $\rho_S=0.918/0.889$). High-commutator skill pairs also yield non-saturating reachability signals, with positive spectral-entropy gains and principal-angle novelty, while synthetic bracketed systems validate the finite-difference estimator. Together, these results show that skills form context-dependent geometric objects whose compositions leave measurable vector-valued non-commutative residuals.


The Implicit Bias of Hyperbolic Representation Learning for Multiclass Data: A Busemann Risk Perspective

Xingrun Li ⋅ Sho Kuno ⋅ Yusuke Mukuta ⋅ Xin Yang ⋅ Tatsuya Harada

We study the implicit bias of Riemannian gradient flow for hyperbolic multiclass classification with fixed class prototypes in hyperbolic space $\mathbb H^n$. Our framework accommodates general *permutation invariant relative margin (PERM)* losses, a class that includes cross entropy and other standard multiclass losses. Our analysis is based on a decomposition: at large radius, the distance to each prototype splits into a radial term and a direction dependent term described by the Busemann function. This yields two main results. First, we prove a radial dichotomy: the sign of a drift coefficient $\mu$ determines whether the radius is pushed toward the ideal boundary or back toward the interior; if the positive drift persists, then $r(t)=\tfrac12\log t+O(1)$, while persistent negative drift returns the trajectory to the large radius threshold in finite time. Second, we show that the boundary direction converges to a critical point of the Busemann risk on $\partial\mathbb H^n$. These results provide a rigorous asymptotic perspective on two phenomena we refer to as *boundary saturation* and *near-boundary clustering* in hyperbolic representation learning.

This paper introduces a new PAC framework for scenario decision-making problems. Scenario decision making consists in making a decision that satisfies a probabilistic constraint (also called a chance constraint) from finitely many sampled realizations (called scenarios) of the constraint. PAC bounds are sufficient conditions on the number of samples to guarantee with high confidence that the sample-based decision satisfies the true constraint with a prescribed probability. Existing PAC bounds rely on intrinsic properties of the problem, such as convexity (Calafiore and Campi, 2005), finite VC dimension (Alamo et al., 2009) or existence of a compression scheme (Margellos et al., 2014). While powerful in some applications, these PAC bounds can be vacuous (or infinite) when the properties are not satisfied. In this paper, we propose a new PAC framework, leading to PAC bounds that are not vacuous for a strictly larger class of scenario decision-making problems. This bound is based on the novel notion of ``internal growth'', which adapts the notion of ``growth function'' from classical machine learning (Vapnik and Chervonenkis, 1968) to scenario decision making. We also relate this notion to other novel properties of the system, such as the $k$-VC dimension. Furthermore, we show a partial converse result: namely, that for the family of stable monotone scenario decision algorithms, the algorithm is PAC if \emph{and only if} it satisfies our criterion. Finally, we demonstrate the usefulness of our framework, and compare with existing approaches, on practical problems.


The Piggyback Hypothesis of Generalization: Explaining and Mitigating Emergent Misalignment

Jiachen Zhao ⋅ Zhengxuan Wu ⋅ Aryaman Arora ⋅ Yiyou Sun ⋅ David Bau ⋅ Weiyan Shi

The mechanisms behind LLMs' broad generalization beyond training examples are poorly understood. Emergent misalignment (EM) offers a striking case study: finetuning on narrow tasks induce broad misalignment to semantically-unrelated test domains. In this work, we propose the Piggyback Hypothesis: the chat-template prefix, shared across all user queries, can piggyback the finetuned behavior onto out-of-domain queries. We validate this hypothesis by showing that subtle perturbations to the prefix, or simply patching the prefix representations with those from the unfinetuned model, can restore alignment without changing the user query. Building on this finding, we propose Token-Regularized Finetuning (TReFT), which directly regularizes prefix representations during training to mitigate piggybacking. Across Llama-3.1, Qwen-2.5, and GPT-OSS models and multiple EM-inducing datasets, TReFT reduces EM while preserving in-domain learning. On Llama-3.1-8B finetuned on the legal domain, TReFT achieves 33.5\% more EM reduction than data interleaving with a retain set of aligned examples. We further show that TReFT extends to other narrow-finetuning settings, including abstention, tool use, and refusal (off-topic generalization is reduced by 54.3\% on average), which further supports the Piggyback Hypothesis. Broadly, our work highlights that LLMs may learn and generalize in unintended ways and suggests a path toward more controllable finetuning. It also calls for further study of how shared input features can piggyback model behavior across domains.


The PokeAgent Challenge: Competitive and Long-Context Learning at Scale

Seth Karten ⋅ Jake Grigsby ⋅ Tersoo Upaa ⋅ Junik Bae ⋅ Seonghun Hong ⋅ Hyunyoung Jeong ⋅ Jaeyoon Jung ⋅ Kun Kerdthaisong ⋅ Gyungbo Kim ⋅ Hyeokgi Kim ⋅ Yujin Kim ⋅ Eunju Kwon ⋅ Dongyu Liu ⋅ Patrick Mariglia ⋅ Sangyeon Park ⋅ Benedikt Schink ⋅ Xianwei Shi ⋅ Anthony sistilli ⋅ Joseph Twin ⋅ Arian Urdu ⋅ Matin Urdu ⋅ Qiao Wang ⋅ Ling Wu ⋅ Wenli Zhang ⋅ Kunsheng Zhou ⋅ Stephanie Milani ⋅ Kiran Vodrahalli ⋅ Amy Zhang ⋅ Fei Fang ⋅ Yuke Zhu ⋅ Chi Jin

We present the PokéAgent Challenge, a large-scale benchmark for decision-making research built on Pokémon's multi-agent battle system and expansive role-playing game (RPG) environment. Partial observability, game-theoretic reasoning, and long-horizon planning remain open problems for frontier AI, yet few benchmarks stress all three simultaneously under realistic conditions. PokéAgent targets these limitations at scale through two complementary tracks: our Battling Track, which calls for strategic reasoning and generalization under partial observability in competitive Pokémon battles, and our Speedrunning Track, which requires long-horizon planning and sequential decision-making in the Pokémon RPG. Our Battling Track supplies a dataset of 20M+ battle trajectories alongside a suite of heuristic, RL, and LLM-based baselines capable of high-level competitive play. Our Speedrunning Track provides the first standardized evaluation framework for RPG speedrunning, including an open-source multi-agent orchestration system that enables modular, reproducible comparisons of harness-based LLM approaches. Our NeurIPS 2025 competition validates both the quality of our resources and the research community's interest in Pokémon, with more than 100 teams competing across both tracks and winning solutions detailed in our paper. Participant submissions and our own baselines reveal considerable gaps between generalist (LLM), specialist (RL), and elite human performance. Analysis against the BenchPress evaluation matrix shows that Pokémon battling is nearly orthogonal to standard LLM benchmarks, measuring capabilities not captured by existing evaluation suites and positioning Pokémon as an unsolved benchmark that can drive RL and LLM research forward. We transition from the NeurIPS 2025 competition to a living benchmark by releasing a live leaderboard for Battling and self-contained evaluation for Speedrunning at https://pokeagentchallenge.com.


The Pok\'emon Theorem and other Fairness Impossibility Results

Daniel Matsui Smola ⋅ Alexander Smola

Fairness impossibility results often look like distinct scalar incompatibility statements. We show that several share one RKHS geometry: fairness criteria are linear constraints on conditional mean embeddings, and unequal base rates make the law of total expectation overdetermine those constraints. This view yields four results. The Kleinberg--Mullainathan--Raghavan dichotomy needs only group-conditional unbiasedness, not full calibration. The \emph{Pok\'emon theorem} shows that a distinct group pair satisfying any finite collection of linear mean-fairness criteria leaves a residual violation witnessed by the MMD, decaying at the Kolmogorov $m$-width rate under spectral regularity. The same tools prove an impossibility for fair feature learning: parity and class-conditional separation in representation space force class collapse under unequal base rates. The approximate relaxations yield signal and error frontiers, allowing a trade-off between real-world estimators and fairness goals. Experiments on standard fairness benchmarks are consistent with our bounds.


The Price of Locality in Label Privacy: Optimal Rates for Classification and Regression

Zongrui Zou ⋅ Mina Dalirrooyfard ⋅ Jingcheng Liu ⋅ Jalaj Upadhyay

We study classification and regression tasks under label differential privacy (label-DP), a setting where only the labels are considered sensitive information. We quantify the "price of locality" by revealing fundamental statistical gaps between local and central privacy models. For classification, we establish a strict separation in accuracy. Specifically, we prove that the minimax excess risk in the local model is lower bounded by $\Omega(\frac{1}{\epsilon}\sqrt{\text{VC}(\mathcal{H})/n})$ for classification on size-$n$ dataset, a fundamental limitation that holds even under relaxed approximate $(\epsilon, \delta)$-local DP. Overcoming this barrier, we propose the first *efficient* central-model algorithm that operates under the stricter pure $\epsilon$-DP, yet achieves a significantly improved upper bound of $\tilde{O}(\sqrt{\text{VC}(\mathcal{H})/(\epsilon n)})$. For linear regression, we show that local randomizers or naive central-DP approach algorithms are statistically inefficient. Instead, we turn to the Matrix Mechanism and use an encoder-decoder framework built on ridge regression, which projects the data into a spherical latent space before injecting noise. This structural design effectively isolates and exponentially suppresses the severe noise inflation typically caused by ill-conditioned feature covariance matrices in full-DP (Cai et al., The Annals of Statistics, 49(5)), achieving an optimal privacy penalty of $\tilde{O}({d^2}/({n^2\epsilon^2}))$ where $d$ is the feature dimension.


The Quantization Benefits of Residual-Free Transformers

Yiping Ji ⋅ Mahalakshmi Sabanayagam ⋅ Peyman Moghadam ⋅ Hemanth Saratchandran ⋅ Simon Lucey

Large-scale transformer training and deployment are increasingly constrained by the transfer of activations, gradients, and optimizer states across accelerators. Low-bit quantization offers a natural remedy, but transformer activations are often heavy-tailed and outlier-dominated, making simple quantization highly lossy. We show that this difficulty is not only a property of the quantizer, but also of the architecture. Specifically, residual connections can drive transformer activations away from Gaussianity during training. Using controlled comparisons between residual and residual-free transformers, we demonstrate that this effect leads to substantially higher quantization error and accuracy degradation at low precision in residual models. We explain the phenomenon through an excess kurtosis analysis, showing that residual mixing can amplify non-Gaussianity, whereas dense mixing in residual-free contracts non-Gaussianity. We then show that residual-free transformers can be made trainable using orthogonal initialization, spectral or second-order optimization, and depth-aware scaling of attention temperature. In language tasks, while there is a small drop in full precision performance, these models retain near-Gaussian activations and exhibit significantly improved robustness to low-bit quantization. Our results identify an accuracy-compressibility trade-off in transformer design and motivate architecture-level approaches to quantization-friendly foundation models.


The Rate-Distortion-Polysemanticity Tradeoff in SAEs

Tommaso Mencattini ⋅ Francesco Montagna ⋅ Francesco Locatello

Sparse Autoencoders (SAEs) that can accurately reconstruct their input (minimizing distortion) by making efficient use of few features (minimizing the rate) often fail to learn monosemantic representations (highly interpretable), limiting their usefulness for mechanistic interpretability. In this paper, we characterise this tension in learning faithful, efficient, and interpretable explanations, introducing the Rate-Distortion-Polysemanticity tradeoff in SAEs. Under toy-modeling assumptions, we theoretically and empirically show that restricting the SAE to be monosemantic necessarily comes with an increase in rate and distortion. Assuming a generative model behind the input observations, we further demonstrate that the degree of polysemanticity of optimal SAEs is determined by the training data distribution, especially by the probability of features to co-occur. Finally, we extend the analysis to real-world settings by deriving necessary conditions that a polysemanticity measure should satisfy when the data-generating process is unknown, and we benchmark existing proxy metrics on SAEs trained on Large Language Models. Taken together, our findings show that polysemanticity is a data problem that should be accounted for when addressing it at the architectural and optimization level.


The Road Ahead in Autonomous Driving: The KITScenes Multimodal Dataset

Richard Schwarzkopf ⋅ Fabian Immel ⋅ Alexander Blumberg ⋅ Jonas Merkert ⋅ Nils A Rack ⋅ Kaiwen Wang ⋅ Fabian Konstantinidis ⋅ Julian Truetsch ⋅ Carlos Fernandez ⋅ Annika Bätz ⋅ Kevin Rösch ⋅ Marlon Steiner ⋅ Willi Poh ⋅ Yinzhe Shen ⋅ Felix Hauser ⋅ Dominik Strutz ⋅ Jaime Villa ⋅ Gleb Stepanov ⋅ Royden Wagner ⋅ Omer Sahin Tas ⋅ Frank Bieder ⋅ Holger Caesar ⋅ Jan-Hendrik Pauls ⋅ Christoph Stiller

Existing autonomous driving datasets have enabled major progress, but fall short in sensor fidelity, map completeness, or geographic diversity. We present KITScenes Multimodal, a European dataset built around high-fidelity sensors and maps. Our fully synchronized sensor suite combines high-resolution global-shutter cameras, long-range lidar beyond 400 meters, 4D imaging radar, and redundant GNSS/INS localization. Our HD maps are, to our knowledge, the most complete of any sensor dataset, verified through autonomous driving trials on open-source software. For the first time in a public dataset, all driving-relevant traffic elements, such as traffic lights, are mapped in 3D to a reprojection-accurate level with full topological connectivity. Recorded in cities with irregular street layouts and mixed traffic modes, our dataset complements existing datasets by broadening the available geographic diversity. We also introduce four benchmarks, each advancing spatial learning for embodied AI: online HD map construction, long-range depth estimation, novel view synthesis, and end-to-end driving.


The Subjectivity of Monoculture

Nathanael Jo ⋅ Nikhil Garg ⋅ Manish Raghavan

Machine learning models—including large language models (LLMs)—are often said to exhibit monoculture, where outputs agree strikingly often. However, we present two experiments that substantially disagree with prior findings of monoculture. In particular, we find that monoculture virtually disappears in one dataset when we include item difficulty compared to previous works that do not. Informed by these findings, we theoretically formalize what it means to measure monoculture. We show that monoculture is inherently subjective, relying on two key decisions: 1) the baseline null model for what "independence" should look like; and 2) the population of models and items under consideration. We conclude with concrete guidelines for future evaluators to ensure the robustness of their results.


The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail

Cassandra Goldberg ⋅ Chaehyeon Kim ⋅ Adam Stein ⋅ Eric Wong

Concept vectors aim to enhance model interpretability by linking internal representations with human-understandable semantics, but their practical utility is often limited by noisy and inconsistent activations. In this work, we uncover the SuperActivator Mechanism: a transformer dynamic that amplifies concept activation gaps, concentrating the most reliable concept evidence into a small set of high-activation tokens. To develop a theoretical understanding of this mechanism, we prove that concept-aligned attention heads multiplicatively amplify pairwise activation gaps, with already-extreme activations growing fastest. We find that this amplification is not just theoretical, but also occurs empirically on large-scale models: while in- and out-of-concept activation distributions overlap considerably, the in-concept distribution develops a positive tail clearly separated from the noise. These high-tail tokens, which we call SuperActivators, appear consistently across concept-positive samples, making them reliable indicators of concept presence. Accordingly, SuperActivator-based detection improves F1 by up to 14% over standard concept activation aggregators and prompting baselines across image and text modalities, models, layers, and concept extraction techniques, demonstrating the generality and practicality of our insights. Further empirical analysis demonstrates that the most reliable SuperActivators are sparse, with detection typically peaking when using only 5-10% of in-concept token activations, and capture more faithful localized semantics than global concept vectors.


The Surprising Effectiveness of Deleting Weights in LLM Reasoning and Adaptation

Jack Lu ⋅ Zhenbang Yang ⋅ Mike Lasby ⋅ Tejas Pote ⋅ Yani Ioannou ⋅ Mengye Ren

How little of a pretrained large language model has to change for it to acquire new capabilities? We find that zeroing fewer than 0.05% of its weights, with no other modification, matches full-parameter fine-tuning (FPT) at scale across the three dominant LLM fine-tuning paradigms: on-policy reinforcement learning with verifiable rewards (GRPO), supervised fine-tuning (SFT), and on-policy distillation (SDFT). We call the method Bit-Mask Tuning (BMT): a learnable binary keep/zero mask over the gate projection of transformer blocks. BMT reaches FPT accuracy at our largest backbones, with the GRPO gap closing monotonically across model sizes from 0.5B to 8B. Two practical advantages follow: the trained mask is several 1000x smaller than the full-parameter checkpoint, and BMT forgets substantially less than FPT, retaining prior-task accuracy throughout training where FPT measurably drifts. To understand why masking alone suffices, we analyze the update geometry on GRPO and find that BMT's weight delta lands closer to full-parameter updates than other adapters on three measures: effective rank, spectrum drift, and principal-weight overlap. In sum, deletion alone produces a tiny yet high-performing adapter, reduced forgetting, and an update geometry that tracks FPT's more closely than any baseline we evaluate.

Large audio-language models have made rapid progress in recognizing what is present in an audio clip, yet spatial audio-language understanding still lacks a clear task interface. A model must not only identify sound events, but also decide where they occur, which semantic and spatial attributes belong to the same auditory object, how multiple objects are arranged, and whether a scene-level answer is physically plausible. We formalize this missing capability as audio scene analysis (ASA), a three-level problem spanning atomic perception, relational integration, and cognitive reasoning. We propose The World is Not Mono (TWNM), a framework that instantiates this definition by equipping audio-language models with explicit spatial evidence. TWNM uses physically grounded First-Order Ambisonics (FOA) simulation to obtain controllable supervision, learns slot-regularized spatial representations from multichannel audio, and fuses these representations with semantic audio features before reasoning with a language model. We further train the model with a progressive curriculum, ending with preference optimization over metadata-derived correct answers and auxiliary format/evidence rewards. To operationalize the ASA definition, we build a controlled benchmark from scene metadata, covering localization, attribute binding, spatial comparison, scene abduction, and counterfactual reasoning. On this ASA benchmark, TWNM achieves 70.8% overall accuracy, 66.4% on spatial-family tasks, and 79.76% on mixed L3 scene-level question answering (QA) under exact multiple-choice question answering (MCQA) scoring. We also audit monaural and binaural reference systems as diagnostic references with explicit audit labels, because they differ in spatial input, training interface, and output format. The supported claim is that a clearly defined ASA task hierarchy, FOA-conditioned spatial representations, and metadata-grounded training together enable controlled, auditable spatial audio-language reasoning, with STARSS23 providing a limited real-recording diagnostic.

The task of visual reasoning relative pose identification (VRRPI) has exposed critical weaknesses in the multi-view 3D spatial reasoning capabilities of modern vision-language models (VLMs). We present a systematic analysis of VLMs behavior in VRRPI, uncovering stable output biases where VLMs disproportionately predict specific motion patterns regardless of visual evidence. To mitigate these inherent biases, we propose self knowledge re-expression with iterative debiasing (SKR-ID), an annotation-free pipeline that adapts VLMs for VRRPI using only randomly extracted unannotated image pairs. SKR-ID transitions VLMs' output mechanism from generic next-token prediction to task-optimized classification heads. The pipeline proceeds in two stages: (1) a cold-start initialization using debiasing object-centric prompts to anchor motion inference in 2D visual evidence; and (2) a recursive logit-ranking mechanism that enforces statistical parity by using the median logit value as a dynamic threshold. Experimental evaluations across frontier VLMs on standard benchmarks demonstrate that SKR-ID significantly mitigates systematic biases and unearths latent 3D reasoning ability. Our method achieves substantial gains, improving open-sourced VLMs' Macro-F1 scores by at least 8.3% and at most 23.0%, proving that VLMs possess significant untapped multi-view 3D reasoning potential.


Think Densely, Act Sparsely: Latent Expert Cognitive Chains for Vision-Language-Action Autonomous Driving

Jie Wang ⋅ Guang Li ⋅ Zhijian Huang ⋅ Jinlong Li ⋅ Chenxu Dang ⋅ Hangjun Ye ⋅ Yahong Han ⋅ Long Chen

Autonomous driving fundamentally demands dense understanding of visual, semantic, and geometric information, yet existing Vision-Language-Action (VLA) models are typically trained with sparse supervision such as language instructions or trajectory signals. This direct mapping from dense visual observations to sparse representations forces the model to compress or even discard critical scene information essential for safe driving. We argue that this limitation does not stem from insufficient model capacity, but from the lack of explicit mechanisms that encourage the formation of dense world representations in the latent space. In this spirit, we propose a novel paradigm for driving intelligence—Think Densely, Act Sparsely—and instantiate it with an efficient and interpretable framework, LECDrive (Latent Expert Cognitive Chains for Vision-Language-Action autonomous Driving). The core idea is to mimic the hierarchical perception process of humans by embedding a progressive chain of latent expert cognition within the VLA model. Concretely, we inject a compact set of task-specific latent tokens as carriers, enabling the model to internalize dense knowledge from foundational vision models, including visual representations (DINOv3), semantic structures (SAM3), and spatial geometry (DepthAnything3), while maintaining efficient inference. This interdependent cognitive chain follows a curriculum learning principle, where supervision is progressively structured from representation to semantics and geometry. Extensive experiments demonstrate that LECDrive consistently improves the performance of base VLMs across multiple autonomous driving benchmarks, covering key tasks such as object perception, state prediction, and trajectory planning. These findings suggest that constructing dense world models through structured latent expert chains is a promising direction for VLA-based autonomous driving.


Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

Chenming Zhu ⋅ Jingli Lin ⋅ Yilin Long ⋅ Peizhou Cao ⋅ Tai WANG ⋅ Jiangmiao Pang ⋅ Xihui Liu

While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the observed images and text-oriented chain-of-thought. They often struggle to infer unobserved layouts, maintain cross-view consistency, and reason from alternative viewpoints when only limited egocentric observations are available. In this work, we study this problem as thinking with imagination, where a VLM actively acquires imagined visual evidence by interacting with a world simulator during reasoning. We propose Astra, an agentic spatial reasoning framework that empowers VLMs with action-conditioned visual imagination. Specifically, Astra couples Astra-VL, an RL-trained VLM policy, with Astra-WM, a Bagel-based world simulator that generates novel-view observations from context images and natural-language camera motions. To provide reliable imagined evidence, Astra-WM is trained with view consistency tuning to improve pose and content consistency across views. In the RL stage, we propose a world-simulator-in-the-loop two-phase RL curriculum to stabilize tool-use exploration and advance the model's ability to invoke the simulator only when imagined observations improve over direct answering. Experiments demonstrate that both the world simulator and the agentic policy are necessary: Astra-WM improves simulator-augmented Gemini-3-Flash on MMSI-Bench from 45.1 to 49.5, while Astra-VL improves the Qwen3-VL backbone from 29.8 to 38.8 on MMSI-Bench and from 36.8 to 42.7 on MindCube. These results show that imagined observations can provide useful spatial evidence, but effective world-model-augmented reasoning requires learning when, where, and how to imagine.


ThinkSafe: Self-Generated Safety Alignment for Reasoning Models

Seanie Lee ⋅ Sangwoo Park ⋅ Yumin Choi ⋅ Gyeongman Kim ⋅ Minki Kang ⋅ Jihun Yun ⋅ Dongmin Park ⋅ Jongho Park ⋅ Sung Ju Hwang

Large reasoning models (LRMs) achieve remarkable performance by leveraging reinforcement learning (RL) on reasoning tasks to generate long chain-of-thought (CoT) reasoning. However, this over-optimization often prioritizes compliance, making models vulnerable to harmful prompts. To mitigate this safety degradation, recent approaches rely on external teacher distillation, yet this introduces a \emph{distributional discrepancy} that degrades native reasoning. We formalize safety realignment as a KL projection onto the safe simplex and prove that the student's own safety-filtered distribution is the unique KL-optimal target, while any external teacher incurs an irreducible excess KL penalty. Guided by this analysis, we propose ThinkSafe, a self-generated alignment framework that restores safety without external teachers. Our key insight is that while compliance suppresses safety mechanisms, models often retain latent knowledge to identify harm. ThinkSafe unlocks this via lightweight refusal steering, which preserves the KL-optimal target while increasing the acceptance rate. Experiments on DeepSeek-R1-Distill and Qwen3 show ThinkSafe significantly improves safety while preserving reasoning proficiency, and achieves superior safety and comparable reasoning to GRPO with roughly an order of magnitude less compute. Code, models, and datasets are available at this Github and HF repository.


TIDE: Trajectory-Aware Watermark Propagation for Text-to-Image Diffusion Models

Yihan Meng ⋅ Suping Xu ⋅ Yanfeng Wu ⋅ Chongjun Wang ⋅ Lin Shang

Watermarking text-to-image diffusion models provides a practical mechanism for provenance tracking and responsible deployment. Post-processing methods are easy to deploy but vulnerable to image-space transformations. In-processing methods improve robustness by embedding the watermark within the denoising trajectory. Yet most existing methods inject the signal into the initial noise or a fixed intermediate latent, without controlling how it evolves during the remaining denoising steps. This can entangle watermark preservation with text-guided generation, risking weaker verification or degraded content fidelity. We propose TIDE, a trajectory-aware watermarking method for text-to-image diffusion models. TIDE injects a learnable watermark into an intermediate latent and jointly optimizes it with an auxiliary watermark condition, using preservation and detection objectives to limit deviation from the unwatermarked reference while maintaining separable watermark evidence. During sampling, TIDE estimates recent text-guidance directions and redirects watermark guidance toward the residual component less aligned with this local text subspace, reducing avoidable interference with semantic generation. Verification uses null-prompt partial inversion to recover the injection latent and match its masked coefficients to the target watermark. Experiments show that TIDE achieves reliable verification on unattacked images, improves robustness under common attacks, and better preserves semantic and visual fidelity.

Streaming video understanding (SVU) requires models to answer user queries over continuously evolving video streams. Unlike offline video QA, SVU must handle queries whose evidence may lie in the past, appear in the current scene, or emerge only after future observations. Existing streaming video systems often rely on compact visual memories or fixed evidence-access schemes, which can lose fine-grained details and fail to adapt the answering process to the temporal intent of each query. In this paper, we propose TimeTraveler, a streaming VQA framework that performs temporal strategy planning with a Time Dictionary. Inspired by the need to search across different temporal regions of a stream, TimeTraveler stores each observed moment as a timestamp-indexed structured caption and uses this dictionary as an explicit source of queryable evidence. Given a query, TimeTraveler determines whether to recall past evidence, read the present context, or wait for future observations, and then applies a strategy-specific evidence acquisition process. Comparative experiments on streaming video QA benchmarks show that TimeTraveler improves streaming question answering by preserving fine-grained temporal evidence and selecting the appropriate temporal strategy for each query.


Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models

Tobias Braun ⋅ Jonas Grebe ⋅ Hossein Shakibania ⋅ Anna Rohrbach ⋅ Marcus Rohrbach

Unified autoregressive models (UAMs) are transformer models that generate text as well as image tokens within a single autoregressive pass. Shared parameters and a multimodal vocabulary simplify the training pipeline and facilitate flexible multimodal generation, yet might introduce new vulnerabilities. In particular, we are the first to show that this unified architecture enables multimodal backdoor attacks, where a trigger can propagate malicious effects across multiple output modalities. Specifically, we present the Token by Token Backdoor Attack (ToBAC), the first backdoor attack targeting UAMs, exploring both data-based and model-based poisoning strategies. We demonstrate that innocuous characters or even common words can be transformed into triggers that elicit harmful behavior in autoregressive image generation. ToBAC can jointly manipulate visual outputs and accompanying text, increasing the perceived authenticity of fabricated content. With model access, ToBAC enables attacks on the unified Liquid model in which a subtle word (e.g., ``cool'') induces modality-aligned brand promotion or ideological influence in 55% of generations. Without model access, ToBAC can be induced through data poisoning, achieving an average success rate of 63.1% against JanusPro.


TokenRouter: Efficient Serving System for Token-Level LLM Routing

Tianyu Fu ⋅ Tengxuan Liu ⋅ Ruoxi Wang ⋅ Yixin Dong ⋅ Yi Ge ⋅ Yichen You ⋅ Yu Wang

Large language model (LLM) routing is widely used to advance the cost-quality Pareto frontier of modern serving systems. While coarse-grained routing at the session or query level has been widely adopted in production systems, recent algorithm work shows substantial efficiency and quality advantages of finer token-level routing. However, efficiently serving token-level routed inference poses significant challenges to existing system design. Built on single-LLM assumptions, current systems struggle with step desynchronization and frequent batch admission delays under token-level routing, and they also impose high implementation complexity on developers. To address this, we propose TokenRouter, an efficient and user-friendly serving system for token-level routed LLM inference. TokenRouter shifts from the conventional request-centric paradigm to a model-centric one. Instead of treating each request as a synchronized decoding stream, TokenRouter launches a subserver for each candidate LLM and dispatches requests asynchronously according to routing decisions. Each subserver employs a delayed-batching scheduler, whose hyperparameters are derived mathematically from a throughput model of the system. Across diverse routing algorithms, workloads, and model pairs, TokenRouter achieves 1.83-64.29× higher decoding throughput than existing systems, substantially advancing the serving efficiency of token-level LLM routing.


Too Aligned to be Real: Detecting AI-Generated Images via Cross-modal Alignment Shift

Haifeng Zhang ⋅ Qinghui He ⋅ Xiuli Bi ⋅ Bo Liu ⋅ Chi-Man Pun ⋅ Bin Xiao

AI-generated image detection has become increasingly important as generative models produce highly realistic visual content. Existing detectors mainly rely on visual artifacts or visual features extracted from pretrained multimodal models, while largely overlooking the structural differences between real and generated images in the joint image--text space. In this paper, we reveal a systematic Cross-modal Alignment Shift (CAS): generated images tend to exhibit stronger and more concentrated alignment with text than semantically matched real images. We verify this phenomenon through retrieval preference, image--text residual structure, spectral concentration, and matching entropy. Motivated by this finding, we propose a CAS-based detection framework(CASNet) that learns an amplified multimodal model with a text-guided alignment-shift objective. The amplification-induced feature increment is extracted as an explicit cross-modal cue and coupled with amplified visual representations, enabling the detector to exploit both alignment-level structural evidence and visual discriminative cues. Extensive experiments across diffusion models, GANs, autoregressive models, and commercial generators demonstrate consistent improvements over state-of-the-art detectors. In particular, on recent commercial generators, our method improves ACC and AP by 5.48\% and 6.13\%, respectively.


ToPA: Block-wise Toeplitz Adaptation for Expressive and Efficient Fine-Tuning

Sicong Li ⋅ Qianqian Xu ⋅ Zhiyong Yang ⋅ Zitai Wang ⋅ Longtao Huang ⋅ XIAOCHUN CAO ⋅ Qingming Huang

Parameter-Efficient Fine-Tuning (PEFT) has emerged as a prevalent strategy for adapting large pre-trained models. Among PEFT techniques, reparameterization-based methods, particularly those that employ low-rank or sparsity structures, have attracted significant attention. However, these approaches often underperform full fine-tuning. Specifically, low-rank methods constrain the update rank, which is misaligned with the inherently high-rank nature of full fine-tuning. In contrast, sparsity-based techniques limit the flexibility of the parameter space and compromise connectivity across weights. To overcome these limitations, we explore the use of Toeplitz matrices, whose entries are constant along each diagonal, therefore providing a compact parameterization without imposing explicit rank or sparsity constraints. A straightforward approach is to update weights via a product of Toeplitz matrices. However, similar to RNNs, long multiplicative chains often lead to gradient instability. To address this issue, we propose ToPA, a more refined version of this idea that replaces standard Toeplitz matrices with block-wise Toeplitz, greatly increasing each factor’s capacity and shortening the chain, thereby stabilizing training and improving expressiveness. We theoretically show that products of block-wise Toeplitz matrices can approximate arbitrary matrices under shorter chain length, justifying the structural design of ToPA. Furthermore, we demonstrate that ToPA offers greater expressivity than existing low-rank and sparse parameterizations. Empirical evaluations across 21 NLP and CV datasets, spanning 5 model architectures, consistently validate the effectiveness of ToPA.


Topology-Aware Optimal Transport for Source-Free Test-Time Adaptation in Anomaly Segmentation

Ali Zia ⋅ Usman Ali ⋅ Abdelwahed Khamis ⋅ Muhammad Umer Ramzan ⋅ Abdul Rehman ⋅ Wei Xiang

Deep topological data analysis (TDA) offers a principled framework for capturing structural invariants such as connectivity and cycles that persist across scales, making it a natural fit for anomaly segmentation (AS). Unlike threshold-based binarisation, which produces brittle masks under test-time distribution shift (TTDS), TDA allows anomalies to be characterised as disruptions to global structure rather than local fluctuations. We introduce TopoOT, a topology-aware optimal transport (OT) framework for source-free test-time adaptation in AS. Our key innovation is Optimal Transport Chaining, which sequentially aligns persistence diagrams (PDs) across thresholds and filtrations, yielding geodesic stability scores that identify features consistently preserved across scales. These stability-aware pseudo-labels supervise a lightweight head updated online using only unlabelled target samples, without access to source data or target labels, with OT-consistency and contrastive objectives, ensuring robust adaptation under TTDS. Across standard 2D and 3D anomaly detection benchmarks, TopoOT achieves state-of-the-art performance, outperforming second-best methods by up to +24.1\% mean F1 on 2D datasets and +10.2\% on 3D AS benchmarks.


TopoPrune: Robust Data Pruning via Unified Latent Space Topology

Arjun Roy ⋅ Prajna Malettira ⋅ Manish Nagaraj ⋅ Kaushik Roy

Geometric data pruning methods, while practical for leveraging pretrained models, are fundamentally unstable. Their reliance on extrinsic geometry renders them highly sensitive to latent space perturbations, causing performance to degrade during cross-architecture transfer or in the presence of feature noise. We introduce TopoPrune, a framework which resolves this challenge by leveraging topology to capture the stable, intrinsic structure of data. TopoPrune operates at two scales, (1) utilizing a topology-aware manifold approximation to establish a global low-dimensional embedding of the dataset. Subsequently, (2) it employs differentiable persistent homology to perform a local topological optimization on the manifold embeddings, ranking samples by their structural complexity. We demonstrate that our unified dual-scale topological approach ensures high accuracy and precision, particularly at significant dataset pruning rates (e.g., 90%). Furthermore, through the inherent stability properties of topology, TopoPrune is (a) exceptionally robust to noise perturbations of latent feature embeddings and (b) demonstrates superior transferability across diverse network architectures. This study demonstrates a promising avenue towards stable and principled topology-based frameworks for robust data-efficient learning.


Toward Executable Multi-framework Front-end Code Generation with Self-Correction

Linxiao Li ⋅ Haoran Ma ⋅ Chenyue Wang ⋅ Jiaye Lin ⋅ Haochen Sui ⋅ Jiechao Gao

Generating codes from UI screenshots has recently benefited from MLLMs, yet existing methods centered on HTML targets and transfer poorly to multi-frameworks such as React, Vue, and Angular. However, unlike HTML, which does not entail compilation issues, generating executable code in multi-framework settings is more challenging due to framework-specific differences in syntax and failure modes, as well as the need to maintain global consistency across files. To address this problem, we study executable multi-framework front-end code generation and propose MESCoder, a tree-structured self-corrective generation framework. Our method first extracts a structural scaffold from the screenshot and then constructs sub-trees and assembly relations step by step, transforming large project generation into a sequence of localized decisions. On top of this scaffold, we introduce a unified Project Agent that continuously revises its own partial project under runtime sandbox feedback and dynamically chooses between local repair and cluster level repair. For training, we adopt a two-stage strategy that first initializes the policy offline to learn basic multi-framework generation and repair priors, and then refines it with online reinforcement learning to optimize long horizon executability and final page quality. The resulting framework models multi-framework executability as a conditioned project state transition process and improves the stability and quality of complex front end generation through self correction. Experiments show that our method improves executability, while narrowing the visual-fidelity gap to HTML-based generation.

Diffusion language models (DLMs) promise parallel, order-agnostic generation, but on standard benchmarks they have historically lagged behind autoregressive models in sample quality and diversity. Recent continuous flow-matching approaches over one-hot and embedding spaces have narrowed this gap, suggesting continuous state spaces are highly effective for language. In this work, we further close the autoregressive gap by modeling text as a continuous diffusion process over fixed-width bitstreams. Our approach represents semantic tokens as analog bit sequences and utilizes a matched-filter residual parameterization to isolate contextual learning from analytic independent-bit posteriors. Crucially, our empirical comparisons reveal that while deterministic continuous DLMs are competitive, they can be overly contractive, undershooting real-data entropy. We demonstrate that the remaining quality-diversity gap is bridged by a stochastic sampler that applies Langevin-type corrections gated by the entropy-rate profile, automatically concentrating stochasticity in high-information regions while remaining nearly deterministic elsewhere. On the One Billion Word Benchmark (LM1B), our 130M-parameter bitstream model reaches a generative perplexity (GenPPL) of 59.76 at matched real-data entropy (4.31) using 256 neural function evaluations (NFEs), decisively outperforming prior DLM baselines and reaching the autoregressive reference. On OpenWebText (OWT), our stochastic sampler establishes a new continuous-DLM Pareto frontier, achieving GenPPL = 27.06 at an entropy of 5.26 using 4x fewer steps than previous 1024-NFE baselines. As an additional architectural benefit, bitstream diffusion removes the O(V) vocabulary scaling bottleneck shared by standard DLMs. By predicting O(log V) bitwise logits via semantic bit-patching, our model yields a reduced memory footprint and higher throughput, demonstrating a scalable paradigm for language generation as vocabulary sizes grow.


Towards Explainable Industrial Anomaly Detection via Knowledge-Guided Latent Reasoning

Peng Chen ⋅ Chao Huang ⋅ Yunkang Cao ⋅ Chengliang Liu ⋅ Wei Wang ⋅ wenqiang wang ⋅ Mingbo Yang ⋅ Li Shen ⋅ Wenqi Ren ⋅ XIAOCHUN CAO

Industrial anomaly detection demands precise reasoning over fine-grained defect patterns. However, existing multimodal large language models (MLLMs), pretrained on general-domain data, often struggle to capture category-specific anomalies, thereby limiting both detection accuracy and interpretability. To address these limitations, we propose Reason-IAD, a knowledge-guided dynamic latent reasoning framework for explainable industrial anomaly detection. Reason-IAD comprises two core components. First, a retrieval-augmented knowledge module incorporates category-specific textual descriptions into the model input, enabling context-aware reasoning over domain-specific defects. Second, an entropy-driven latent reasoning mechanism conducts iterative exploration within a compact latent space using optimizable latent think tokens, guided by an entropy-based reward that encourages confident and stable predictions. Furthermore, a dynamic visual injection strategy selectively incorporates the most informative image patches into the latent sequence, directing the reasoning process toward regions critical for anomaly detection. Extensive experimental results demonstrate that Reason-IAD consistently outperforms state-of-the-art methods across multiple tasks. The code will be publicly available upon publication.


Towards Principled Fine-Grained MoE Expert Pruning via Pseudo-Boolean Approximation

Zongfang Liu ⋅ Ziheng Cheng ⋅ Shengkun Tang ⋅ Jinghui Zhang ⋅ Weijie Wang ⋅ Zhiqiang Shen ⋅ Xin Yuan

Mixture-of-Experts (MoE) language models improve parameter efficiency by activating only a small subset of experts per token, but all experts must still be stored at inference time, making memory a key deployment bottleneck. Expert pruning is therefore a practical post-training compression strategy, but existing methods face a fundamental tension: search-based approaches attempt to capture expert interaction effects through joint optimization, yet become intractable for modern fine-grained MoE models; score-based approaches, by contrast, ignore such effects, but remain efficient and often surprisingly effective. To understand this tension, we formulate expert pruning as a constrained pseudo-Boolean optimization problem and empirically analyze the interaction effects induced by pruning. Our results show that cross-layer interaction effects are relatively more important, since pruning earlier layers changes the hidden-state distribution for later ones, whereas within-layer effects are often well approximated by additive singleton damages. Based on this observation, we propose \textbf{MoE-PBA}, a prefix-conditioned pruning method that prunes layers sequentially under the hidden states produced by the already-pruned prefix. Across four representative fine-grained MoE language models with distinct architectures and parameter scales ranging from 7B to 30B, MoE-PBA achieves a substantially better accuracy--efficiency trade-off than search-based methods and outperforms the strong score-based baseline on diverse downstream benchmarks.


Towards Self-Supervised, Generalizable and Decomposable 4D Driving Scene Reconstruction

James Tu ⋅ Anqi Joyce Yang ⋅ Jingkang Wang ⋅ Sivabalan Manivasagam ⋅ Raquel Urtasun

Reconstructing high-fidelity and controllable digital twins from multi-modal sensory observations is a fundamental problem in physical AI applications. While existing per-scene optimization methods handle dynamic driving scenes well, they are slow to optimize, have artifacts at novel viewpoints, and require costly manual annotations for decomposing the scene into controllable instances. Self-supervised generalizable reconstruction methods enable faster and more scalable reconstructions by learning from large datasets, but existing approaches do not fully leverage multi-modal inputs (e.g., camera and LiDAR) and lack decomposition capabilities. To address these limitations, we propose STRIDE, a self-supervised, generalizable and decomposable 4D driving scene reconstruction method. By processing multi-modal sensory inputs in a common 3D space with a Point-Transformer, STRIDE efficiently discovers spatio-temporal correspondences necessary for accurate flow prediction. By incorporating latent instance tokens, STRIDE is the first to perform feed-forward reconstruction and learnable decomposition without manual annotations. Experiments on public driving datasets show STRIDE achieves state-of-the-art performance on recovering geometry, appearance, and 3D flow. Moreover, we demonstrate how learned decompositions can enable dynamic instance manipulation and controllable simulation.


Towards Understanding the Power and Limits of the Muon Optimizer: A River-Valley Perspective

Tianqi Shen ⋅ Jinji Yang ⋅ Runze SHI ⋅ Jianhao Ma ⋅ Jiaye Teng ⋅ Ziye Ma

Recently, Muon has gained substantial attention as an appealing alternative to Adam, with many works highlighting its advantages through spectral normalization and improved conditioning. Yet this positive theoretical narrative contrasts with its empirical performance in large language model (LLM) training, where Muon’s gains over Adam/AdamW are often mixed, schedule-sensitive, and not uniformly superior. To address this gap, we develop a trajectory-level theory characterizing both the strengths and limitations of Muon. We introduce a mixed-spiked matrix sensing model whose sensing operator decomposes into signal, spike, and bulk components, capturing a mixture of anisotropic structure and long-tail information reminiscent of LLM training. On top of it, we adopted a river-valley perspective in which we view the landscape as composed of a river direction flowing to the desired solution and hill directions encoding nuisance or task-irrelevant information. In the momentum-free setting, we show that Muon moves faster along the information-bearing river direction during early optimization, but can converge much more slowly near the river bottom than gradient descent. We then extend the river-valley perspective to general nonconvex objectives with momentum by studying points on the spectral river. There, while Muon converges faster early on, its orthogonalized update removes residual scale information, making it prone to overshooting and oscillation near the target solution. Together, these results suggest that our characterizations extend beyond spiked matrix sensing and motivate switching to GD-like refinement optimizers in the final phase, rather than relying only on a fixed learning-rate schedule for Muon. We also provide preliminary evidence supporting this two-stage approach in language model training experiments.

Inferring continuous system evolution from sparse temporal snapshots is a key challenge in generative modeling and single-cell omics. While Optimal Transport (OT) is popular, existing frameworks are largely restricted to first-order dynamics, assuming memoryless velocity fields. This limits expressiveness, as first-order systems fail to account for regulatory momentum and time-delayed responses inherent in processes like cell differentiation. Here, we introduce TracingFlow, a simulation-free Flow Matching framework generalizing to second-order dynamics. By using neural networks to regress the acceleration field, TracingFlow provides an exact, efficient solution to the Dynamical Optimal Acceleration Transport (DOAT) problem. Unlike first-order methods yielding over-smoothed trajectories, our second-order formulation captures high-curvature transitions and nonlinear evolutions by learning the underlying force fields. Evaluated on complex synthetic and large-scale scRNA-seq datasets, TracingFlow achieves superior accuracy in distributional reconstruction and trajectory faithfulness. Moreover, by integrating lineage tracing priors, it recovers dynamical structures that are both mathematically optimal and biologically plausible.


Track4D: Representing Dense 3D Tracking for Video Diffusion Models

Yushi LAN ⋅ Zeren Jiang ⋅ Kelvin Zheng Li ⋅ Xingang Pan ⋅ Chuanxia Zheng ⋅ Andrea Vedaldi

We introduce Track4D, a method that formulates dense 3D tracking as conditional video generation. Repurposing large-scale pretrained video generators for 3D tracking is non-trivial: 3D point tracking in world space produces signals far from natural video, making it difficult for the video diffusion model (VDM) to learn. We systematically study two representations for encoding 3D tracking in the VDM latent space: a Residual 3D Tracking Video (RTV) that directly encodes metric 3D offsets, and a Normalized Coordinate Map Video (NCMV) that implicitly encodes the underlying geometric correspondences as canonical 2D coordinate maps. We further condition the VDM on point maps from an off-the-shelf 3D reconstructor to provide an explicit geometric scaffold for tracking. Trained exclusively on limited synthetic data with LoRA adaptation, Track4D achieves competitive zero-shot tracking performance across five benchmarks, demonstrating the potential of leveraging video generation priors for dense 3D tracking.


Train at the Moving Edge: Rollout-Efficient RL for Large Reasoning Models

Jiahao Wu ⋅ Ning Lu ⋅ Shengcai Liu ⋅ Kun Wang ⋅ Yanting Yang ⋅ Baijiong Lin ⋅ Chen J Zhang ⋅ Qing Li ⋅ Ke Tang

Reinforcement learning (RL) has become essential for post-training large language models (LLMs) in reasoning tasks. While scaling rollouts can stabilize training and enhance performance, it introduces substantial computational overhead. In algorithms like GRPO, multiple rollouts per prompt incur prohibitive costs, as a large portion of prompts provide negligible gradients and are thus of low utility. This raises a key question: \textit{how to identify high-utility prompts before an expensive rollout?} Our experimental analysis reveals that sample utility is non-uniform and dynamic: the strongest learning signals concentrate at the ``learning edge'', the intersection of intermediate difficulty and high response entropy, which shifts throughout training. Motivated by this observation, we propose HIVE, a history-informed and online-verified prompt selection framework for data-efficient RL training. HIVE first uses historical reward statistics and response entropy as a cheap prior to filter candidate prompts, and then employs prompt entropy as a real-time proxy to prune instances with stale utility. Across multiple reasoning benchmarks and base models, HIVE maintains reasoning accuracy with substantial rollout cost reduction, achieving up to 2.3× total training speedup.


Train-free Data Poisoning Attack against Retrieval-augmented Diffusion Models

XINQI LYU ⋅ Yihao LIU ⋅ Yiming Cao ⋅ Bin Xiao

Retrieval-augmented diffusion models (RAG-DMs) have significantly advanced image synthesis by incorporating external knowledge bases. However, their security vulnerabilities against malicious external data remain largely underexplored. Existing data poisoning attacks against diffusion models primarily corrupt internal parameters. They fail against RAG-DMs because clean retrieved images easily override these compromised parameters. Furthermore, recent attack tailored for RAG-DMs requires training the retriever, resulting in high computational costs and overfitting to specific retrievers. To bridge this gap, we propose PoisonedRDM, a novel and train-free data poisoning attack tailored to execute concept hijacking against RAG-DMs. Specifically, PoisonedRDM poisons the knowledge base with a minimal number of optimized images and optimizes adversarial perturbations through a joint optimization strategy, steering unknown retrievers toward the poisoned images via surrogate ensemble alignment and aligning the generated outputs with a malicious target concept, while adhering to a strict perturbation budget to maintain high visual stealthiness. Experiments show that PoisonedRDM effectively attacks RAG-DMs, achieving high success rates and outperforming state-of-the-art baselines.

To defend against backdoor attacks on neural networks, the defender must identify the verifier function that determines whether an input is a trigger, rather than rejecting individual triggers in isolation, since an attacker can generate new triggers satisfying the same verifier, against which trigger-by-trigger blocking cannot keep up. If a backdoor verifier could be installed with the same kind of asymmetry as a cryptographic authentication scheme, in which only the attacker can produce trigger inputs while the defender cannot reverse-engineer the verifier from observations, learning-based defense would face a principled impossibility. We show that no such asymmetry can arise for any backdoor installed by training into a model. We establish two complementary impossibility results. First, any verifier installable by training a polynomial-size neural network is reconstructible by the defender with sample complexity matching the attacker's up to polynomial factors. Second, any cryptographic secret the attacker might supply to the model at inference is necessarily observable to a white-box defender. Together, these results rule out any way of giving a self-contained backdoor cryptographic asymmetry. We confirm both routes empirically on pretrained large language models, demonstrating that no configuration yields a backdoor simultaneously effective for the attacker and irrecoverable by the defender.

Choosing where to place entangling gates is a central design choice in quantum machine learning circuits, yet entanglers are still typically chosen from fixed templates or by expensive search procedures requiring training or pairwise kernel evaluations. We propose the \emph{HSD-Lipschitz principle}, a training-free geometric criterion built on a simple intuition: a useful entangler should keep encoded states stable across the dataset while separating different classes. We formalise this through a two-sided bound on local light-cone subsystems. Under a light-cone locality assumption, dataset-level gradient variance is upper-bounded by reduced-state spread; under random initialisation, the expected between-class gradient signal is lower-bounded by class-conditional reduced-state distance. The resulting algorithm, HSD-greedy, adaptively grows an entanglement structure using only single-state reduced moments, estimable via classical shadows, and avoids the $\mathcal{O}(N^2)$ pairwise kernel evaluations of kernel-target alignment proxies. Empirically, both bounds hold without violation across $110$ topologies and $440$ parameter rows, and extend to $180$ additional configurations. HSD-greedy attains the strongest average AUC among training-free entangler selectors on synthetic and real-world benchmarks, at roughly $60\times$ lower selection cost than the strongest baseline in 13-qubit simulation. End-to-end execution on a 20-qubit IBM Eagle subchain substantially outperforms a fixed 19-CNOT Linear chain and matches or exceeds a matched-budget Random baseline while using far fewer two-qubit gates than the Linear chain.


Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context

Zhaowei Wang ⋅ Lishu Luo ⋅ Haodong Duan ⋅ WeiWeiLiu ⋅ Sijin Wu ⋅ Ji Luo ⋅ Shen Yan ⋅ Shuai Peng ⋅ Sihang Yuan ⋅ Chaoyi Huang ⋅ Yi Lin ⋅ Yangqiu Song

Long context is becoming a core capability of modern large vision-language models (LVLMs), enabling sustained context management across long-document understanding, video analysis, and multi-turn tool use in agentic workflows. However, practical training recipes remain under-documented, such as how to construct and mix long-context data. In this work, we present a systematic study of long-context continued pre-training for LVLMs, extending a 7B LVLM from 32K to 128K context with extensive ablations on long-document data. We find that: 1) synthesizing long-document VQA data provides effective and diverse long-context supervision, covering tasks from information extraction to numerical reasoning; 2) retrieving relevant evidence remains the primary long-context bottleneck, favoring retrieval-heavy mixtures with a small amount of reasoning data to preserve task diversity; 3) surprisingly, pure long-document VQA data largely preserves short-context capabilities, suggesting that instruction-formatted long data lessens the need for short-data mixing. Instantiating these findings, we obtain MMProLong by continuing training from Qwen2.5-VL-7B, improving long-document VQA performance by 7.11 points under only a 5B-token budget. More importantly, MMProLong generalizes beyond its 128K training window, maintaining strong performance at 256K and 512K without additional training. More broadly, it also transfers to webpage-based multimodal needle-in-a-haystack tasks, long-context vision-text compression, and long-video understanding without task-specific supervision. Overall, our study provides a practical LongPT recipe and an empirical foundation for advancing the next generation of LVLMs.

Unsupervised learning in Vision Transformer models, such as DINO, enables robust visual perception without requiring labeled data by leveraging large-scale image datasets. However, these models lack a focusing mechanism during training, resulting in perceptual representations that are often ambiguous or overly similar. In this work, we introduce two visual focusing loss functions designed to establish correspondences between the model's attention and specific regions within images. Specifically, we leverage the multi-head attention mechanism within the model to selectively steer the model's focus, leading to the emergence of multiple, diverse, and focused perceptual capabilities without requiring supervision. Through both qualitative and quantitative evaluations, we demonstrate that our method substantially increases the spatial selectivity and diversity of attention heads, improving both explainability and fine-grained recognition performance.

Although existing candidate trajectory evaluators have substantially improved end-to-end planning accuracy, their underlying mechanism remains suboptimal. Model-based evaluators typically make scoring future-aware by explicitly predicting candidate-conditioned future states over the planning horizon, which we term rollout. However, rollout incurs high inference cost and accumulates prediction errors. We argue that planning does not require reconstructing a full future scene for each candidate, but only future-relevant information for reliable comparison. We therefore propose RFWorld, a Rollout-Free World model for end-to-end trajectory evaluation. RFWorld constructs a shared scene memory and shapes it into a future-readable representation via lightweight training-time temporal readout with semantic supervision, keeping training overhead low. At inference time, candidate trajectories directly query the shared memory via sparse spatial readout for scoring, reducing computation by roughly the planning horizon times the number of candidate trajectories, while avoiding recursive rollout errors. Experiments on NAVSIM show that RFWorld learns a more planning-aligned world representation and achieves state-of-the-art performance.


Transcoder Adapters for Reasoning-Model Diffing

Nathan Hu ⋅ Jake Ward ⋅ Thomas Icard ⋅ Chris Potts

While reasoning models are increasingly ubiquitous, the effects of reasoning training on a model's internal mechanisms remain poorly understood. We introduce transcoder adapters, a technique for learning an interpretable approximation of the *difference* in MLP computation before and after fine-tuning. We train transcoder adapters on two pairs of base and reasoning models: Qwen2.5-Math-7B / DeepSeek-R1-Distill-Qwen-7B and Qwen2.5-32B / QwQ-32B. We find that modeling the difference in MLP computation is far easier than modeling the full MLP; adapters achieve faithful reconstruction with an order of magnitude fewer active features than typical transcoders. When evaluated on reasoning benchmarks, adapters exhibit the reasoning model's characteristic long responses and recover a large fraction of its benchmark performance. Adapter features are interpretable, achieving higher automated interpretability scores than MLP neurons. To demonstrate the utility of transcoder adapters for interpreting fine-tuning differences, we present two case studies on the 7B model pair. First, we examine the overall composition of adapter features to gain a broad overview of fine-tuning differences. Despite being specific to fine-tuning by construction, many features have activating examples unrelated to reasoning structure. Interventions confirms these features disproportionately affect benchmark performance rather than response length. Second, we study why the model says 'wait' by constructing attribution graphs whose edges flow through base model parameters. We trace hesitation to only $\sim$2.4\% of adapter features (5.6k total), finding that this behavior depends largely on base model computation. These features are necessary and sufficient for producing hesitation tokens; removing them reduces response length, often without affecting accuracy. Anonymous code is available at \url{https://anonymous.4open.science/r/transcoder-adapters-6752/}.

In image classification scenarios where both prediction and explanation efficiency are required, self-explaining models that perform both tasks in a single inference are effective. However, for users who already have prediction-only models, training a new self-explaining model from scratch imposes significant costs in terms of both labeling and computation. This study proposes a method to transfer the visual explanation capability of self-explaining Vision Transformer (ViT) models learned in a source domain to prediction-only VLM-based ViT models in a target domain via task arithmetic, without target-side explanation supervision. The proposed method endows explanation capability by adding an \emph{explainability vector} induced by explanation supervision in the source domain based on task arithmetic framework. Experiments on ten diverse target datasets show that a single explainability vector learned on ImageNet-1k augmented with patch-level explanation supervision transfers consistently, improving explanation quality while largely preserving classification accuracy. Beyond the transfer itself, we further investigate whether transfer success can be anticipated \emph{a priori} from model-internal features alone (without target-side explanation supervision), and find that a small set of such features carries non-trivial signal about the transfer outcome.

In-context learning plays a central role in transformer-based large language models, yet its theoretical understanding remains limited. In this work, we study multiclass in-context classification under more realistic settings, including anisotropic class centers and label imbalance, by introducing a spectral data generation framework that constructs class-center matrices with a prescribed singular spectrum. We first show that the isotropy of class-center vectors, as quantified by the stable rank, improves ICL generalization performance. From a meta-learning perspective, our theorem shows that linear transformers can learn multiclass in-context classification with near-optimal per-label sample complexity, extending prior guarantees beyond the binary setting. In the test label imbalance regime, our analysis reveals that queries from majority classes are easier to classify, while those from minority classes are more error-prone; moreover, robustness to this bias improves with the stable rank. Finally, we empirically demonstrate that our theory is consistent with observations on transformers and pretrained large language models.


Treat Bias as Noise: Training Bias-Robust LLM Reasoning via Reinforcement Learning

Qian Wang ⋅ Xuandong Zhao ⋅ Zirui Zhang ⋅ Zhanzhi Lou ⋅ Nuo Chen ⋅ Dawn Song ⋅ Bingsheng He

Large language models (LLMs) increasingly serve as reasoners and are being considered as automated evaluators, yet they remain susceptible to cognitive biases---often altering their reasoning when faced with spurious prompt-level cues such as consensus claims or authority appeals. Existing mitigations via prompting or supervised fine-tuning fail to generalize, as they modify surface behavior without changing the optimization objective that makes bias cues attractive. We propose \textbf{Epistemic Independence Training (EIT)}, a reinforcement learning framework built around a simple principle: models should learn that bias cues are \emph{unreliable} rather than learning to either follow or reject them. EIT trains on balanced conflict examples where each injected cue is equally likely to support the correct or incorrect answer, and uses a reward that penalizes bias-following errors without rewarding agreement with a cue that happens to be correct---making the cue non-predictive of reward. On controlled MMLU-Pro reasoning tasks with bias injection, EIT improves accuracy and robustness on both Qwen3-1.7B and Qwen3-4B when bias points to wrong answers, while preserving performance when bias aligns with truth. Trained only on bandwagon bias, EIT generalizes along two out-of-domain axes: held-out MMLU-Pro subjects and unseen bias types (authority, distraction, verbosity). EIT-trained Qwen3-4B further outperforms untrained Qwen3-8B and Qwen3-14B on bias resistance, indicating that targeted training is more effective than model scaling alone. Code and data are available at https://anonymous.4open.science/r/bias-mitigation-with-rl-BC47.


Tree-Sliced Orlicz Integral Probability Metric

Tuan Hoang ⋅ Trung-Khang Tran ⋅ Viet-Hoang Tran ⋅ Tan Nguyen

Tree-sliced distances have emerged as scalable alternatives to Sliced Wasserstein distances by projecting probability measures onto tree metric spaces and exploiting closed-form transport on trees. Existing tree-sliced constructions, however, are largely built around first-order Wasserstein geometry. While this choice yields efficient computation, it fixes the discrepancy to an $L^1$-type aggregation of subtree mass imbalances and provides limited control over the geometry used in downstream optimization. We propose Tree-Sliced Orlicz IPM (TS-Orlicz), a tree-sliced framework that replaces the tree-level $W_1$ discrepancy with an Orlicz integral probability metric. Building on the tractable formulation of Orlicz IPMs on trees, TS-Orlicz aggregates Orlicz-induced discrepancies over random tree systems while preserving the scalability of tree-sliced computation. By varying the underlying Orlicz function, the framework recovers $L^p$-type behavior and also supports non-polynomial, tail-sensitive geometries that emphasize large subtree discrepancies. We show that TS-Orlicz preserves key properties of tree-sliced constructions and admits efficient computation via closed-form special cases and one-dimensional scalar optimization. We further extend the framework to probability measures on hyperspheres. Experiments across Euclidean and spherical settings, including gradient flows, self-supervised learning, and diffusion-based generative modeling, show that Orlicz-induced tree-sliced discrepancies achieve competitive or improved performance over sliced and tree-sliced baselines while maintaining low computational cost.


TriAxialKV: Toward Extreme Low-Precision KV-Cache Quantization for Agentic Inference Tasks

Hanzhang Shen ⋅ Haoran Wu ⋅ Yiren Zhao ⋅ Robert Mullins

Agentic workloads have emerged as a major workload for LLM inference. They differ significantly from chat-only workloads, requiring long-context processing, the ability to handle multimodal inputs, and structured multi-turn interactions with tool calling capabilities. As a result, their context exhibits structure that can carry different importance along three key axes: temporal recency to the current turn, modality such as text or image tokens, and semantic role such as user queries, tool calls, observations, or reasoning. These axes capture distinct token behaviors and lead to different sensitivities to KV-cache compression. However, existing KV-cache quantization methods are typically homogeneous or exploit only heterogeneity on a single dimension, such as temporal proximity or modality, overlooking the interactions among them. To this end, we introduce TriAxialKV, a novel mixed-precision KV-cache quantization scheme that assigns each token a triaxial tag, calibrates per-tag sensitivity, and allocates INT2/INT4 bitwidths under a fixed memory budget. We implement TriAxialKV as an end-to-end serving system, comprising calibration, mixed-precision quantization and memory management, and custom fused Triton decode kernels. When using Qwen3-VL-32B-Thinking as a computer-use agent operating the OSWorld, TriAxialKV matches the accuracy of SGLang with BF16 KV cache while supporting 4.5$\times$ KV cache size and achieving 30\% higher end-to-end throughput, when running on real GPU systems.


TriBet: Relative E-values for Online Machine-generated Text Detection

Xunye Tian ⋅ Zhijian Zhou ⋅ Liuhua Peng ⋅ Jared Collette ⋅ Dino Sejdinovic ⋅ Yue Yang ⋅ Feng Liu

To ensure a formal guarantee that fully human-written text streams are *not falsely accused* as machine-generated text (MGT), the online betting framework (e-values) has been successfully applied on various existing detectors. However, current two-sample e-value methods fail in two real-world application scenarios: the false positive rate (FPR) is often out of control and the test power drops drastically when (1) the reference domain is similar but not exactly same as test domain (cross-domain), or (2) proxy model of detector is outdated compared to the modern source model. We identify the root cause as an online calibration dilemma: two-sample optimization and betting relies on a null-calibrating offset that is unidentifiable from unlabeled test streams under shift. We resolve this with a triple-sample betting framework (TriBet) that conduct online MGT detection as relative testing against both human and machine references using a witness score, and introduces a conformal test-inclusive domain-shift betting strategy to achieve FPR control without fragile online calibration. We show that TriBet is a level-$\alpha$ sequential test with asymptotic power one, and characterize its expected stopping time. Empirically, we evaluate on 15 source LLMs, 8 proxy models and mixed-document streams, TriBet consistently delivers faster and more powerful detection than both oracle and online betting baselines while keeping FPR strictly controlled.


Trie-Aware Transformers for Generative Recommendation

Zhenxiang Xu ⋅ Sirui Chen ⋅ Yong He ⋅ Jieyu Yang ⋅ Chuan Yuan ⋅ Ke Ding ⋅ Jing Cao ⋅ Can Wang ⋅ Jiawei Chen

Generative recommendation (GR) aligns with advances in generative AI by casting next-item prediction as token-level generation rather than score-based ranking. Most GR methods adopt a two-stage pipeline: (i) \textit{item tokenization}, which maps each item to a sequence of discrete, hierarchically organized tokens; and (ii) \textit{autoregressive generation}, which predicts the next item's tokens conditioned on the tokens of user's interaction history. Although hierarchical tokenization induces a prefix tree (trie) over items, standard autoregressive modeling with conventional Transformers often flattens item tokens into a linear stream and overlooks the underlying topology. To address this, we propose TrieRec, a trie-aware generative recommendation method that augments Transformers with structural inductive biases via two positional encodings. First, a \textit{trie-aware absolute positional encoding} aggregates a token's (node's) local structural context (\eg depth, ancestors, and descendants) into the token representation. Second, a \textit{topology-aware relative positional encoding} injects pairwise structural relations into self-attention to capture topology-induced semantic relatedness. TrieRec is also model-agnostic, efficient, and hyperparameter-free. In our experiments, we implement TrieRec within three representative GR backbones, achieving notably improvements of 8.83\% on average across four real-world datasets.

Pose-guided text-to-image generation often suffers from limb distortions and feature crosstalk in complex multi-person scenarios. While existing UNet-based adapters struggle with long-range spatial dependencies, emerging Multimodal Diffusion Transformers (MM-DiTs) offer superior global modeling. However, naive signal concatenation in MM-DiTs severely disrupts pre-trained latent distributions. To address this, we propose TrioPose, a native pose-driven framework built upon the SD3.5M architecture. Specifically, we introduce a Triple-Stream Pose-Aware DiT (TSPA-DiT) that treats pose as an independent modality. It employs layer-wise activation and zero-initialized dual-residual injection to smoothly enforce geometric constraints while preserving pre-trained latent stability. To resolve severe multi-instance occlusions, we design a Learnable Relational Bias Mask that categorizes topological connectivity into fine-grained physical states, mapping them into continuous attention soft constraints to effectively decouple inter-instance interference. Furthermore, a Pose-Guided Spatial Loss Weighting strategy modulates the native diffusion objective using heatmap-derived error maps, focusing anatomical supervision strictly on distortion-prone regions. Extensive experiments demonstrate that TrioPose achieves state-of-the-art performance across challenging benchmarks, including Human-Art, CrowdPose, and OCHuman. Notably, it attains an AP of $64.33$ on Human-Art, representing a $30 \\%$ improvement over prior arts, while setting new standards for visual fidelity and text-image semantic alignment in complex multi-human generation.


TriSpec: Ternary Speculative Decoding via Lightweight Proxy Verification

Haoyun Jiang ⋅ Junqi He ⋅ Feng Hong ⋅ Xinlong Yang ⋅ jianwei zhang ⋅ Zhengyang Zhuge ⋅ Zheng Li ⋅ Xiaofeng Cao ⋅ Zhiyong Chen ⋅ Bo Han ⋅ Junyang Lin ⋅ Jiangchao Yao

Inference efficiency in Large Language Models (LLMs) is fundamentally limited by their serial, autoregressive generation, especially as reasoning becomes a key capability and response sequences grow longer. Speculative decoding (SD) offers a powerful solution, providing significant speed-ups through its lightweight drafting and parallel verification mechanism. While existing work has nearly saturated improvements in draft effectiveness and efficiency, this paper advances SD from a new yet critical perspective: the verification cost. We propose TriSpec, a novel ternary SD framework that, at its core, introduces a lightweight proxy to significantly reduce computational cost by approving easily verifiable draft sequences and engaging the full target model only when encountering uncertain tokens. TriSpec can be integrated with state-of-the-art SD methods like EAGLE-3 to further reduce verification costs, achieving greater acceleration. Extensive experiments on the Qwen3 and DeepSeek-R1-Distill-Qwen/LLaMA families show that TriSpec achieves up to 35\% speedup over standard SD, with up to 50\% fewer target model invocations while maintaining comparable accuracy.


Trivialized Generative Models on Lie Groups

Neil He ⋅ Meenal Jhajharia ⋅ Qianxi Wu ⋅ Chaoran Cheng ⋅ Arindam Banerjee ⋅ Ge Liu

Many problems in diverse fields involve data that naturally live on Lie groups. However, existing Riemannian and Lie group generative models still face several limitations. Riemannian flow and consistency models often require position-dependent vector fields, and expensive geometric terms such as Euler-Arnold simulation on general Lie groups or covariant derivatives. Existing Lie group flow models are usually tied to collinear exponential paths, while momentum-based Lie group diffusion relies on compactness assumptions and does not naturally extend to non-compact groups. To address these limitations, we propose Trivialized Generative Models (TGM), a family of generative models that learns endpoint-constrained paths in the fixed Lie algebra and lifts them to the Lie group. This yields Trivialized Flow Matching (TFM), Trivialized Consistency Models (TCM), and momentum-based extensions with endpoint correction, enabling flexible path design, simpler few-step objectives, and generation on non-compact Lie groups. We evaluate TGM on compact SO(3) benchmarks, non-compact $\mathrm{Sp}(4,\mathbb{R})$ datasets motivated by continuous quantum systems and paraxial optics, and real-world protein backbone generation. Across these settings, TGM improves sample quality and efficiency over existing Riemannian and Lie group baselines.

Inference-time selection methods, such as Best-of-N, improve generation by sampling a pool of candidates and selecting the top completion according to a reward model. Distillation seeks to amortize this procedure into a single policy by replacing raw rewards with in-pool ranks and learning a policy that upweights higher-ranked completions. However, existing rank-based policies typically use smooth full-support reweighting, so low-ranked completions receive less mass but remain in the target support. Although a sharper reweighting reduces lower-tail mass, it also increases reliance on brittle ranking at the top made by a single reward model. We propose \textbf{TUP}: a \textbf{T}runcate-bad, \textbf{U}pweight-good \textbf{P}olicy that removes low-ranked completions from the support and reweights only the retained upper tail with a tunable sharpness. TUP admits a closed-form, prompt-independent normalization and can be trained fully offline via binary cross-entropy, using shifted-truncated win-rates as soft labels and distilled-to-reference log-likelihood ratios as logits. Theoretically, under certain assumptions, we show that for any unknown oracle reward, the best monotone rank-reweighting can be matched by a lower-tail truncation rule, providing formal support for removing the lower tail rather than merely downweighting it. Empirically, we show that TUP is competitive with strong offline alignment baselines.


TrunkFish: Making Model Width Incrementally Refinable

Owen M Dugan ⋅ Liam Dugan ⋅ Aaryan Singhal ⋅ Christopher De Sa ⋅ Christopher Ré

Width is a central axis for scaling LLM families, but widening a standard transformer typically changes the hidden representation rather than continuing the smaller-width computation. We introduce Trunkfish, a recipe for converting model architectures into a hierarchy of models with incrementally refinable width (a Trunkfish hierarchy) by enforcing causal hidden dimensions and cumulative prefix readouts: new width may depend on old width, but old width may not depend on new width. This enables schooling, which trains all models in a hierarchy from forward/backward passes of only the largest model, and restart-free cascades, which enable adaptively scaling up inference compute without wasting compute already spent on smaller models. Trunkfish improves the validation-loss/active-parameter frontier over independently trained standard hierarchies under matched hierarchy-training FLOPs, with substantial training-FLOP savings for both full-width and prefix-heavy schooling. In cascade simulations over the same trained hierarchy, exact width continuation also meaningfully reduces matched-loss expected active-parameter traffic relative to restart cascades.

Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions. Whether they evaluate the quality of that evidence, or merely aggregate it based on surface presentation, remains poorly understood. We show that models possess the capability to detect fabricated statistics (correct identification rates of $0.76$--$1.00$ for methodology in isolation) but do not recruit this capability during multi-source synthesis, producing similar numeric estimates whether the statistics are fabricated or valid. Specifically, source influence is governed by a methodology-register gate that responds to the distributional register of analytical text but not to numeric validity: for example, statistically impossible confidence intervals receive the same weight as valid ones. The behavioral dissociation replicates across five models from three families (Claude, Qwen, OLMo) and three professional domains. Mechanistic analyses, including causal tracing, linear probes, and component-level attribution, converge on the same account: the model encodes and causally uses a methodology-register representation that transfers across domains (probe AUC $0.83$--$0.92$), while numeric-validity signals, decodable in isolation, are suppressed to chance during multi-source synthesis. Prompting-based mitigations, even an oracle checklist naming the exact statistical checks, produce blanket skepticism rather than selective discernment, and the post-training pipelines we examine reinforce the stylistic shortcut without building numeric verification. Unlike sycophancy, which tracks user preference, this failure tracks whether a source presents as analytically credible, not whether its claims are internally consistent. We term this $\textit{epistemic alignment}$: like preference and safety alignment, the question is not capability but deployment.

Trustworthy AI encompasses many aspirational aspects for aligning AI systems with human values, including fairness, privacy, robustness, explainability, and uncertainty quantification. The ultimate goal of Trustworthy AI research is to achieve all aspects simultaneously. However, efforts to enhance one aspect often introduce unintended trade-offs that negatively impact others. In this position paper, we review notable approaches to five aspects and systematically consider every pair, detailing the negative interactions that can arise. For example, applying differential privacy to model training can amplify biases, undermining fairness. Drawing on these findings, we take the position that current research practices of improving one or two aspects in isolation are insufficient. Instead, research on Trustworthy AI must account for interactions between aspects and adopt a holistic view across all relevant axes at once. We offer recommendations for how researchers can work towards integrated trust, and provide guidance for practitioners to manage interactions today.


Trustworthy and Efficient Map-free LiDAR Localization via Scan-Pose Alignment and Flow Matching

Haoze Chang ⋅ Dunqiang Liu ⋅ Minghang Zhu ⋅ Wen Li ⋅ Sheng Ao ⋅ Siqi Shen ⋅ Chenglu Wen ⋅ Cheng Wang

Reliable map-free LiDAR relocalization requires not only accurate and efficient pose estimation but also trustworthy confidence estimation for detecting localization failures. However, estimating reliability to ensure trustworthy LiDAR relocalization remains largely unexplored. Also, existing LiDAR pose regression methods suffer from accuracy and efficiency trade-off. To address this limitation, we propose \textbf{SPAR}, a model-agnostic confidence estimator that reformulates localization reliability as scan-pose alignment. Given a query scan and a predicted pose, SPAR measures their compatibility in a learned embedding space and explicitly optimizes confidence scores to correlate scan with pose. In addition, we introduce \textbf{VeLoc}, an efficient and accurate pose regressor based on flow-matching for efficient pose refinement. Experiments on the Oxford and NCLT datasets demonstrate that the proposed framework achieves a favorable balance between efficiency and accuracy for LiDAR relocalization while enabling reliable localization failure detection across different localization backbones.


Tunable Latent Generative Priors for Compressed Sensing and Inverse Problems

Sean Gunn ⋅ Jorio Cocola ⋅ Paul Hand ⋅ Oliver De Candido ⋅ Vaggos Chatziafratis

Latent generative models have emerged as powerful priors for solving inverse problems. These models typically represent a class of natural signals at a single, fixed complexity, governed by the latent dimensionality. This can be limiting: depending on the problem, a latent dimensionality that is too small may result in high representation error, while one that is too large may overfit to noise. We develop tunable latent priors for diffusion models, normalizing flows, and variational autoencoders, leveraging nested dropout. Across tasks including compressed sensing, inpainting, denoising, and phase retrieval, we show empirically that tunable priors consistently achieve lower reconstruction errors than fixed-complexity baselines. In the linear denoising setting, we derive the optimal complexity in closed form, showing how it depends on the noise level and the signal spectrum. This work demonstrates the potential of tunable latent generative priors and motivates both the development of supporting theory and their application across a wide range of inverse problems.


Unaligned Image Guided Denoising via Cross-modal Conditional Flow Matching

Runmin Zhang ⋅ Linshan Li ⋅ Si-Yuan Cao ⋅ Zhu Yu ⋅ Lingyu Zhu ⋅ Huaqi Zhang ⋅ Hui-liang Shen

This work presents UGD-FM, a novel unaligned image guided denoising framework based on cross-modal conditional flow matching. Unlike previous approaches that assume spatially aligned inputs or only handle small-baseline rectified pairs, we address a more challenging setting where severe target noise and large-range two-dimensional cross-modal misalignment coexist. Specifically, we formulate guided denoising as a progressive flow matching process, allowing image restoration and aligned guidance aggregation to mutually reinforce each other. To support efficient few-step inference, we introduce a stage-focused and inference-consistent training strategy that focuses on task-critical denoising stages, and mitigates the training-inference mismatch. We further propose adaptive sparse guidance propagation for efficient and reliable large-range matching, regularized by content and geometry level auxiliary losses. Experiments demonstrate that UGD-FM achieves state-of-the-art denoising performance with only one or two inference steps, and is compatible with modern single-image restoration backbones.


Uncovering Semantic Hierarchies in Text-Attributed Graphs via Variational EM-based LLM–GNN Synergy

Yunhui Liu ⋅ Xudong Jin ⋅ Qizhuo Xie ⋅ Chunhui Zhao ⋅ Xiao Luo ⋅ Tao Zheng ⋅ Bin Chong ⋅ Tieke He

While hierarchical structures are intrinsic to organizing human knowledge, current representation learning on text-attributed graphs predominantly operates on flat semantic spaces, overlooking the rich coarse-to-fine granularity inherent in real-world data. To bridge this gap, we propose SHiFT, which synergizes the reasoning power of LLMs with the structural representation capability of GNNs within an Expectation-Maximization (EM) paradigm. In the E-step, we harness LLMs as expert taxonomists to induce an explicit semantic tree, introducing a lightweight Mapping &amp; Diagnosis strategy that reduces inference costs by surgically updating the hierarchy only when anomalies are detected. In the M-step, we align the GNN encoder with this induced hierarchy from complementary perspectives, including clustering partition, topological skeleton, and semantic concepts. Theoretical analysis interprets SHiFT as a generalized Variational EM algorithm. Extensive experiments demonstrate that SHiFT not only achieves state-of-the-art performance but also uncovers high-quality, human-readable taxonomies.


Uncovering the latent structure of interwoven population and temporal codes

Zachary Friedenberger ⋅ Yiwei Cao ⋅ Richard Naud

Population analysis methods have become standard for navigating the complexity of neural data. However, these methods often assume a rate code, neglecting information encoded in the precise timing of spikes. Critically, additional information encoded in bursts of action potentials may be missed. Here, we develop a factor analysis method that disentangles the factors associated with bursts and individual spikes. This enables burst codes to be investigated directly from the structure of the data, without requiring external covariates. We demonstrate that analyzing firing rates alone obscures the latent structure and factors underlying bursts. Applying our method to simulated and experimental data, we show that it can infer the correct latent structure and be used to test for the presence of burst coding. By merging the population and burst coding perspectives, we provide a framework for linking changes in bursting to internal variables involved in attention, perception, and learning.

Optimization with matrix gradient orthogonalization has recently demonstrated impressive results in the training of deep neural networks (Jordan et al., 2024; Liu et al., 2025). In this paper, we provide a theoretical analysis of this approach. In particular, we show that the orthogonalized gradient method can be seen as a first-order trust-region optimization method, where the trust-region is defined in terms of the matrix spectral norm. Motivated by this observation, we develop the stochastic non-Euclidean trust-region gradient method with momentum, which recovers the Muon optimizer (Jordan et al., 2024) as a special case, along with normalized SGD and signSGD with momentum (Cutkosky and Mehta, 2020; Sun et al., 2023). In addition, we prove state-of-the-art convergence results for the proposed algorithm in a range of scenarios, which involve arbitrary non-Euclidean norms, constrained and composite problems, and non-convex, star-convex, first- and second-order smooth functions. Finally, our theoretical findings provide an explanation for several practical observations, including the practical superiority of Muon compared to the Orthogonal-SGDM algorithm of Tuddenham et al. (2022) and the importance of weight decay in the training of large-scale language models.

Model reprogramming adapts a *frozen* source model to a target task by wrapping it in a *learnable* input transformation from target input to the source-model input space, and an output mapping from source-model outputs to target predictions. Despite extensive empirical progress, the *finite-sample learnability* of the induced wrapper hypothesis class $\mathcal{F}\_{\rm MR}$ remains underexplored. We address this for source models that expose hard source labels and analyze through the wrapper structure: on any target samples, the input transformation selects *reachable* source labels that the output mapping relabels. Through this lens, we answer two questions: ① how many target samples are needed to approach the best predictor in $\mathcal{F}\_{\rm MR}$, and ② what gap remains between the optimum over $\mathcal{F}\_{\rm MR}$ and the target Bayes risk. For ①, the wrapper structure factorizes the sample-wise complexity of $\mathcal{F}\_{\rm MR}$ into a *reachability* term counting reachable source-label patterns, and a *conditional relabeling* term counting their relabelings. This yields an agnostic-PAC bound on the *estimation error*, with sample complexity scaling additively in the two terms; this bound becomes closed-form on additive input transformations and is matched by a lower bound on an explicit binary classification construction. For ②, the same structure decomposes the *approximation error* exactly into the source-label information loss and output-mapping restriction with a source-conditional upper bound on this gap. We extend the approximation analysis to logit-exposing source models, where the upper bound is empirically estimable. The wrapper-side analysis localizes both errors to specific components, giving a structural view of reprogramming practice.


Understanding Randomization in Greedy Model Search

Xin Chen ⋅ Jason Klusowski ⋅ Yan Shuo Tan ⋅ Chang Yu

We study feature subsampling in greedy model search, a randomization mechanism motivated by random forests and other ensemble methods. While recent theory suggests that this randomization acts solely as a variance reduction mechanism analogous to ridge regularization, these results largely rely on base learners optimized via ordinary least squares (OLS). We investigate the effects of feature subsampling on greedy forward selection, a tractable abstraction of the adaptive split search used by decision trees. Assuming an orthogonal design, we prove that ensembling with feature subsampling can reduce both bias and variance, contrasting with the pure variance reduction of convex base learners. Specifically, we show that both the training error and degrees of freedom need not be monotone in the subsampling rate, breaking the analogy with standard shrinkage methods like the lasso or ridge regression. Furthermore, we characterize the exact asymptotic behavior of the estimator, showing that it adaptively reweights OLS coefficients based on their rank, with weights that are well approximated by a logistic function. These are mechanism level results for a tractable analogue of greedy selection that depends on the response, not performance guarantees for full random forests.


Understanding the Effects of Hyper-Connections on Self-Attention Dynamics: A Bifurcation Analysis

Yen N Pham ⋅ Dung V Nguyen ⋅ Trinh Nguyen ⋅ Thieu Vo ⋅ Tan Nguyen

A central challenge in understanding deep sequence models is characterizing how token representations evolve across layers: whether they collapse, diverge, or converge to nontrivial structures. We study this question for Transformers with hyper-connections, whose learnable cross-layer coupling makes the stability of representation collapse analytically tractable. In contrast to standard residual connections, hyper-connections introduce additional parameters that directly control the hyperbolicity of the zero fixed point, at which the local stability is fully determined by the Jacobian eigenvalues, thereby enabling a rigorous and precise bifurcation analysis of the dynamical system underlying the model. In particular, we derive a continuous-depth ODE limit of self-attention with hyper-connections and use bifurcation theory to characterize its token dynamics. For the 1-stream case, we obtain a closed-form critical threshold at which a pitchfork bifurcation occurs, separating representation collapse from divergence or convergence to nontrivial stable states. For the 2-stream case, the richer coupling structure yields two distinct stability boundaries, corresponding to qualitatively different instabilities: one arising when a real eigenvalue crosses zero, altering the number of nearby equilibria (static type), and another when a complex-conjugate pair reaches the imaginary axis, leading to oscillatory behavior (Hopf type). These phenomena have no counterpart in the 1-stream setting. Experiments on pretrained hyper-connected language models validate our theory and show that unstable regimes yield stronger performance.

Chain-of-Thought (CoT) reasoning substantially improves Large Language Models (LLMs), yet its out-of-distribution (OOD) generalization mechanism remains theoretically underexplored. We study CoT as compositional OOD generalization: the target task lacks complete reasoning trajectories or direct input-output pairs, while training provides only in-distribution (ID) atomic step data. We show that CoT generalizes by repeatedly applying a shared single-step predictor to compose reusable ID transitions. Theoretically, the OOD task risk is controlled by the sum of ID subtask risks, with an error-propagation coefficient independent of reasoning depth/steps $T$. A Rademacher-complexity analysis further gives a generalization gap of $\mathcal{O}(\sqrt{T/n})$, rather than linear in $T$, due to predictor sharing, where $n$ denotes training samples per subtask. For autoregressive Next-Token Prediction, we derive an explicit bound $\mathcal{O}\big(k\big(\sqrt{{T}/{mn}}+\sqrt{{T}/{n}}+{T}/{m}\big)\big)$, where $m$ and $k$ denote context length and output token number per step. Experiments on Atom-Task and GSM8K support the theory, showing that CoT enables short-to-long and compositional generalization, while explicit state tracking is crucial.


Unified Loss-Aware Density Control for 3D Gaussian Splatting

Tae-Young Kim ⋅ Jiwoo Chung ⋅ Jae-Pil Heo

The efficiency and fidelity of 3D Gaussian Splatting critically depend on adaptive density control, which grows and suppresses Gaussian primitives throughout optimization. However, the standard pipeline still relies on heuristic view-space gradients, scale thresholds, and global opacity resets, rather than explicitly assessing the loss-reducing effect of each split, clone, or opacity update. This mismatch can allocate primitives to already-sufficient regions, miss under-reconstructed structures, and destabilize training through blanket opacity changes. In this work, we recast adaptive density control as a collection of local loss-reduction decisions. We derive efficient Taylor-based per-Gaussian surrogates that predict the objective change caused by three primitive-level operations: splitting, cloning, and opacity decay. This yields unified loss-aware framework with three components. First, we extend second-order splitting to a normalized joint position-scale space, enabling children to adapt not only their locations but also their anisotropic extents. Second, we formulate duplicate cloning as a surrogate minimization problem, selecting only Gaussians whose duplicated contribution is predicted to improve reconstruction and assigning each child a principled opacity. Third, we replace global opacity reset with selective soft opacity decay that weakens harmful primitives while preserving useful ones. Across standard 3DGS benchmarks, our method produces more targeted primitive allocation and achieves improved rendering quality with comparable or smaller Gaussian budgets compared to baselines.

Punctuation Restoration and Inverse Text Normalization (ITN) are essential post-processing tasks for ASR systems. Traditionally, these are handled by separate models in a sequential pipeline, which suffers from error propagation and suboptimal latency. We propose UniForm, a unified framework that jointly solves both tasks using a single compact language model. UniForm is trained via a three-stage progressive paradigm: large-scale supervised fine-tuning on punctuation data, joint fine-tuning with ITN span masks, and reinforcement learning alignment. To enable effective reinforcement learning in this inherently low-entropy setting, we introduce Segment-Aware GRPO (SA-GRPO), which optimizes model performance through two synergistic components: a structured reward decomposition mechanism that provides dedicated signals for distinct segment types, and a token-level adaptive credit assignment mechanism that dynamically redistributes gradients toward difficult tokens based on sampling consensus and model entropy. This integrated approach effectively eliminates the need for manual sub-task weighting and ensures fine-grained optimization for both tasks. We train UniForm-0.8B on over 238 million multilingual samples. Extensive evaluations across diverse languages, domains, and code-switching scenarios demonstrate that UniForm-0.8B outperforms prior specialized models and larger general-purpose LMs on both tasks while maintaining real-time inference latency. Furthermore, our analysis reveals a strong synergistic effect, proving that joint training mutually benefits both punctuation and ITN performance.

We analyze generalization error, uniform stability, and uniform argument stability of gradient descent (GD) and stochastic gradient descent (SGD) over discrete parameter spaces, where each update involves deterministic or stochastic rounding. We show that deterministic rounding degrades the generalization error of GD on convex, Lipschitz, and smooth loss functions, increasing the rate from $O(T/n)$ to $O(T/\sqrt{n})$, and establish matching lower bounds. We further prove that uniform stability of GD becomes $\Omega(T)$, showing that stability-based generalization bounds are vacuous in this setting. In contrast, for the same losses, stochastic gradient descent with deterministic rounding admits nontrivial uniform stability guarantees, which differ qualitatively from the real-valued case and exhibit distinct dependencies on the number of iterations and the dimension: we prove tight bounds $O(T/n)$ for one dimension and $O(T^2/n)$ for higher dimensions. We also show that stochastic rounding can introduce generalization error that increases with the dimension; such a phenomenon is absent in standard real-valued optimization and in the deterministic rounding case. Finally, we provide upper bounds on uniform argument stability for stochastic rounding schemes and show that these bounds are tight when the loss can be represented as a sum of coordinate-wise functions.


Unifying Partner and Environment Diversity to Improve Human-AI Coordination

Patrick Q Zhang ⋅ Yancheng Liang ⋅ Adrienne Fairhall ⋅ Natasha Jaques

Real-world applications of modern AI systems require generalization to both unseen environments and novel human users. The field of zero-shot coordination (ZSC) focused primarily on deploying AI to novel partners in otherwise familiar environments seen in the training. The popular population-based methods enable partner generalization by training the agent on a diversity of simulated human users (serving as training-time partners) in a fixed environment, aiming to cover the diverse behaviors of novel human users. Such a paradigm does not consider the more challenging realistic scenarios where both the users and the environment are unseen during training. Recently, training across environments has been explored in training a single self-play agent, which shows better generalization on both novel human users and unseen environments. In this work, we ask whether jointly modeling partner diversity and environment diversity yields stronger coordination with novel partners in novel environments. We introduce Cross Environment Co-Play (CECP), a two-stage framework that first trains a population of self-play agents across a distribution of procedurally generated environments, and then trains a cooperator policy against this partner population. To further couple social and environmental reasoning, we add a next-state prediction module that encourages representations to encode both partner behavior and task structure, and validate the design choice with ablation studies. We evaluate CECP on both a neuroscience-inspired cooperative foraging task, and the population ZSC benchmark Overcooked, under held-out partner and held-out environment conditions. CECP achieves the strongest performance among the compared baselines and attains state-of-the-art results under our evaluation protocol for human-AI collaboration. We provide further analysis using the model's latent representation encoded by the next-state prediction module, showing it can not only improve the performance of the cooperative AI, but also provide more interpretable latents, which demonstrate the inference of partner skill levels.


UniGS: Unified Geometry-Aware Gaussian Splatting for Multimodal Rendering

Yusen XIE ⋅ Zhenmin Huang ⋅ Jianhao Jiao ⋅ Dimitrios Kanoulas ⋅ Jun Ma

In this paper, we propose UniGS, a unified map representation and differentiable framework for high-fidelity multimodal 3D reconstruction based on 3D Gaussian Splatting. Our framework integrates a CUDA-accelerated rasterization pipeline capable of rendering photo-realistic RGB images, geometrically accurate depth maps, consistent surface normals, and semantic logits simultaneously. We redesign the rasterization to render depth via differentiable ray-ellipsoid intersection rather than using Gaussian centers, enabling effective optimization of rotation and scale attribute through closed-form analytic depth gradients. Furthermore, we derive the analytic gradient formulation for surface normal rendering, ensuring geometric consistency among the reconstructed 3D scenes. To improve computational and storage efficiency, we introduce a learnable attribute that enables differentiable pruning of Gaussians with minimal contribution during training. Quantitative and qualitative experiments demonstrate the state-of-the-art reconstruction accuracy across several datasets, validating the efficacy and robust optimization stability of our geometry-aware paradigm.


UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation

Mengmeng Liu ⋅ Diankun Zhang ⋅ Jiuming Liu ⋅ Jianfeng Cui ⋅ Hongwei Xie ⋅ Guang Chen ⋅ Hangjun Ye ⋅ Francesco Nex ⋅ Hao Cheng ⋅ Michael Yang

World Action Models have recently emerged as a promising paradigm for improving action generalization in autonomous driving by leveraging future video prediction as dense supervision for scene dynamics and temporal causality. However, it remains unclear how video modeling should be integrated with action generation to maximally transfer these benefits. Existing cascaded or dual-Diffusion Transformer architectures decouple video imagination from trajectory prediction, limiting knowledge transfer: the action model may still overfit dataset-specific driving priors, while the video model only indirectly regularizes planning. In this paper, we propose UNIVERSE, a unified video-action model built upon a single mask-modulated Diffusion Transformer. By co-training future video latents and ego-trajectory tokens within shared generative parameters, UNIVERSE allows dense video supervision to directly shape trajectory denoising, leading to stronger cross-domain action generalization. To ensure causal validity and efficient deployment, we introduce a Modality-Decoupling Visibility Mask, which shares historical context across modalities while blocking mutual attention between future video and trajectory tokens. This prevents future-modality dependency and enables trajectory-only inference by removing future-video denoising at test time, achieving a $4.3\times$ speedup over joint video--trajectory rollout while maintaining comparable planning accuracy. The same model also supports video-only and joint video--trajectory rollouts. Experiments show that UNIVERSE achieves 91.0 PDMS on NAVSIM and strong zero-shot transfer to nuScenes and Bench2Drive without fine-tuning, while ablations confirm the importance of single-DiT unification, video co-training, and mask-based modality decoupling.


Unlocking Compositional Generalization in Continual Few-Shot Learning

Phu-Quy Nguyen-Lam ⋅ Phu-Hoa Pham ⋅ Dao S Minh ⋅ Chi Nguyen Tran ⋅ Trung-Kiet Huynh ⋅ Long Tran-Thanh

Object-centric representations promise a key property for few-shot learning: Rather than treating a scene as a single unit, a model can decompose it into individual object-level parts that can be matched and compared across different concepts. In practice, this potential is rarely realized. Continual learners either collapse scenes into global embeddings, or train with part-level matching objectives that tie representations too closely to seen patterns, leaving them unable to generalize to truly novel concepts. In this paper, we identify this fundamental structural conflict and pioneer a new paradigm that strictly decouples representation learning from compositional inference. Leveraging the inherent patch-level semantic geometry of self-supervised Vision Transformers (ViTs), our framework employs a dual-phase strategy. During training, slot representations are optimized entirely toward holistic class identity, preserving highly generalizable, object-level geometries. At inference, preserved slots are dynamically composed to match novel scenes. We demonstrate that this paradigm offers dual structural benefits: The frozen backbone naturally prevents representation drift, while our lightweight, holistic optimization preserves the features' capacity for novel-concept transfer. Extensive experiments validate this approach, achieving state-of-the-art unseen-concept generalization and minimal forgetting across standard continual learning benchmarks.


Unlocking Fine-Grained Perception in CLIP via Structurally-Aware Latent Masked Modeling

Juntong Li ⋅ Lingwei Dang ⋅ Haomin Wu ⋅ Ziyan Qiu ⋅ Qingxin Xiao ⋅ Qingyao Wu

Vision-Language Models (VLMs) such as CLIP excel in global semantic alignment but often lack fine-grained perceptual capabilities. This hinders dense prediction tasks and bottlenecks the visual potential of Multimodal Large Language Models (MLLMs). Existing research has attempted to enhance CLIP's visual representations by incorporating geometric priors from vision-centric models. However, these strategies often struggle to achieve deep alignment for both local spatial structures and global semantics, potentially even distorting the original image-text space. To address these limitations, we propose SALM, an unsupervised embedding alignment framework based on structurally-aware latent mask modeling. SALM effectively synergizes local and global alignment via a dual-path design combining explicit and implicit mechanisms, without requiring any image-text pairs. First, we introduce a dual-matrix alignment strategy that explicitly calibrates intra-sample spatial correlations and activation intensities, thereby effectively injecting local geometric priors. Based on this, we further design a latent mask modeling mechanism to guide CLIP to restore the missing semantic details of the target model, thereby implicitly aggregating fine-grained structures into the global semantic space. Furthermore, driven by the empirical observations that CLIP's shallow features inherently possess strong spatial observational capabilities, we naturally extend SALM to a highly efficient self-distillation paradigm, SALM-Self. This unlocks CLIP's intrinsic fine-grained potential without relying on any external models. Extensive experiments demonstrate that SALM not only significantly improves performance in dense prediction tasks but also boosts CLIP's zero-shot accuracy, effectively enhancing the fine-grained understanding capabilities of MLLMs.

Highly over-parameterized models can simultaneously memorize noisy labels and generalize well, yet how these behaviors coexist remains Highly over-parameterized models can simultaneously memorize noisy labels and generalize well, yet how these behaviors coexist remains poorly understood. In this work, we investigate the underlying mechanisms of this coexistence using modular arithmetic tasks under heavy label noise. Through extensive experiments on two-layer neural networks, we find that larger models tend to generalize better under appropriate optimization and model configurations, while noisy labels are memorized faster than clean data. Over-parameterized models internally form a generalization structure, but its expression in the output is suppressed by the need to fit noisy labels. Remarkably, even with 80\% label noise, near-perfect test accuracy can be achieved by extracting this internal structure using frequency-based methods. We further propose a task-agnostic method to partition networks into generalization and memorization components. Although this subnetwork improves generalization, it is limited compared with frequency-based extraction, indicating that the generalization structure is distributed across neurons and motivating the development of new tools to retrieve generalizable knowledge from over-parameterized networks.


UpSafe℃: Upcycling for Controllable Safety in Large Language Models

Yuhao Sun ⋅ Lingyun Yu ⋅ Zhuoer Xu ⋅ Kun Yang ⋅ shiwen cui ⋅ Yongdong Zhang ⋅ Hongtao Xie

Large Language Models (LLMs) have achieved remarkable progress across a wide range of tasks, but remain vulnerable to safety risks such as harmful content generation and jailbreak attacks. Existing safety techniques---including external guardrails, inference-time guidance, and post-training alignment---each face limitations in balancing safety, utility, and controllability. In this work, we propose UpSafe℃, a unified framework for enhancing LLM safety through safety-aware upcycling. Our approach first identifies safety-critical layers and upcycles them into a sparse Mixture-of-Experts (MoE) structure, where the router acts as a soft guardrail that selectively activates original MLPs and added safety experts. We further introduce a two-stage SFT strategy to strengthen safety discrimination while preserving general capabilities. To enable flexible control at inference time, we introduce a safety temperature mechanism, allowing dynamic adjustment of the trade-off between safety and utility. Experiments across multiple benchmarks, base model, and model scales demonstrate that UpSafe℃ achieves robust safety improvements against harmful and jailbreak inputs, while maintaining competitive performance on general tasks. Moreover, analysis shows that safety temperature provides fine-grained inference-time control that achieves controllable trade-off between utility and safety. Our results highlight a new direction for LLM safety: moving from static alignment toward dynamic, modular, and inference-aware control. Code: https://anonymous.4open.science/r/UpCycle-70B3.


Urbex: Agentic Spatial Grounding for City-Scale 3D Scenes

Yu Wang ⋅ Rui Dai ⋅ Wei Chen ⋅ Yi Wang ⋅ Bo Dang ⋅ Kaikui Liu ⋅ Xiangxiang Chu ⋅ Yansheng Li

City-scale 3D grounding is commonly formulated as candidate ranking, which requires precomputed instance proposals and becomes inefficient in large outdoor scenes. We propose Urbex, an agentic vision-language framework that instead casts grounding as active spatial search. Urbex represents a city-scale 3D scene as an interactive multi-view environment, allowing a VLM agent to query landmarks, zoom into local regions, render oblique views, and commit a grounded bounding box. We optimize this tool-use policy with reinforcement learning on CityRefer, using smooth localization rewards and lightweight shaping rewards for efficient exploration. Experiments show that Urbex achieves the best Hit@1 and IoU on CityRefer, supports candidate-free grounding without ground-truth boxes, and is substantially faster than the strongest city-scale ranking baseline. Zero-shot results on STPLS3D-Refer further indicate improved cross-scene robustness. The code and dataset will be publicly released.


URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets

Zhuangzhe Wu ⋅ Yue Xin ⋅ Chengkai Hou ⋅ Minghao Chen ⋅ Yaoxu Lyu ⋅ Jieyu Zhang ⋅ Shanghang Zhang

Articulated objects are fundamental for robotics, simulation of physics, and interactive virtual environments. However, recovering them from visual observations is inherently challenging, as images provide only partial and ambiguous cues about both part geometry and their underlying kinematic structure. Existing approaches typically rely on multi-stage pipelines, retrieval from asset libraries, or explicit part segmentation. We present URDF-Anything+, an end-to-end autoregressive diffusion framework that generates simulation-ready URDF models directly from a single RGB image. Conditioned on visual observations and object geometry, URDF-Anything+ operates in a structured latent space and jointly models part geometry and articulation in a unified generation process. Specifically, the model sequentially predicts each articulated part together with its associated joint parameters, while a termination token dynamically determines the number of parts. This design enables direct generation of fully executable URDFs without external retrieval or post-processing stages. Experiments on large-scale articulated object benchmarks demonstrate that URDF-Anything+ outperforms prior methods in geometric reconstruction quality, joint parameter estimation, and physical executability, while being substantially more efficient than existing multi-stage approaches. Furthermore, the generated URDFs serve as faithful digital twins, enabling the zero-shot transfer of manipulation policies trained purely in simulation.


Utility-Constrained Policy Optimization

Mehrdad Moghimi ⋅ Bernardo Avila Pires

Constrained MDPs (CMDPs) are a widely adopted framework for incorporating safety into RL agents, however, the framework does not support risk-sensitive constraints. This can be problematic: For example, CMDPs allow for optimal solutions that, in order to satisfy the risk-neutral constraints, mix infrequent catastrophic behaviors and frequent, overly conservative ones. Moreover, empirical results in multiple previous works suggest that enforcing stricter, risk-sensitive constraints can improve agent performance even when measuring it in a risk-neutral way. In this work, we introduce a simple yet powerful methodology for constrained RL, consisting of an extension of CMDPs that removes the limitation of risk-neutral constraints, and an algorithm for solving the resulting constrained problem. As a convenient side-effect of our framework, it is not necessary to fix constraint limits in advance of training the agent, provided that a sensible range is known. This increases policy flexibility and, in practice, allows for adjustments to these limits at no extra training cost. Besides benefiting from the generality of the framework, our agent shows strong performance in practice, consistently matching or outperforming existing baselines in several Safety Gymnasium benchmark tasks.


VAANI: Capturing the language landscape for an inclusive digital India

Sujith Pulikodan ⋅ Abhayjeet Singh ⋅ Agneedh Basu ⋅ Nihar Desai ⋅ Pavan K J ⋅ Pranav D Bhat ⋅ Raghu Dharmaraju ⋅ Ritika Gupta ⋅ Sathvik Udupa ⋅ Saurabh Kumar ⋅ Sumit Sharma ⋅ Visruth Sanka ⋅ Dinesh Tewari ⋅ Harsh Dhand ⋅ Amrita Kamat ⋅ Sukhwinder Singh ⋅ Shikhar Vashishth ⋅ Partha Talukdar ⋅ Raj Acharya ⋅ Prasanta Ghosh

Voice-based technologies have the potential to bridge digital accessibility gaps; however, existing datasets fail to capture the linguistic and regional diversity of Indic languages. We present Project VAANI, a large-scale multimodal dataset designed to represent India’s linguistic landscape across 165 districts. Speech data is collected using image-based prompts to elicit spontaneous responses, while images are curated through a separate pipeline covering diverse themes across regions. The dataset undergoes a rigorous multi-stage quality control process, combining automated and manual evaluation to ensure high audio quality and transcription accuracy. We release approximately 289K images, 31,255 hours of speech, and 2,043 hours of transcribed audio spanning 105 languages from 28 states and 3 union territories. Many of these languages are represented at this scale for the first time, making VAANI a foundational resource for inclusive speech technology. The dataset enables the development of robust, multilingual, and multimodal models, and supports research in speech recognition, language understanding, and cross-modal learning for underrepresented languages.


Valid Best-Model Identification for LLM Evaluation via Low-Rank Factorization

Elad Tolochinsky ⋅ Yaniv Tenzer ⋅ Yaniv Romano

Selecting the best large language model (LLM) for a fixed benchmark is often expensive, since exhaustive evaluation requires running every model on every example. Multi-armed bandit (MAB) algorithms can reduce the number of LLM calls by sequentially selecting the next model-example evaluations, thereby avoiding unnecessary computations on clearly underperforming models. Further computational savings can be achieved by predicting the evaluation scores of the various models on the set of examples. In practice, these predictions can be obtained using low-rank (LR) matrix factorization that exploits correlations in the partially observed model–example score matrix. However, such predicted evaluations are not the ground truth: they can be biased and may therefore lead to incorrect identification of the best model. In this work, we propose a principled framework that combines MAB with cheap, predicted evaluation scores without compromising statistical validity. Concretely, leveraging prediction-powered inference (PPI), we derive unbiased estimators for the performance of each model that utilize predicted scores from LR factorization to reduce variance. Importantly, this enables the construction of finite-sample valid confidence intervals in our non-i.i.d. setting, where models are selected adaptively and examples are sampled without replacement. Empirical results on real-world benchmarks demonstrate that our approach reduces the number of required evaluations, leading to meaningful savings in compute and cost, while accurately identifying the best-performing model.


Variable-Length Generative Protein Design via Generalized Poisson Flow

Chaoran Cheng ⋅ Zhanghan Ni ⋅ Yanru Qu ⋅ Yuxin Chen ⋅ Ruihan Guo ⋅ Jiajun Fan ⋅ Ge Liu

The ability to generate variable-length proteins is crucial in protein design, where the optimal length is often unknown and tightly coupled to structural designability. Current state-of-the-art diffusion and flow-based generative models require the protein length to be predetermined before sampling, limiting their flexibility in exploring the feasible protein design space. To bridge this gap, we introduce Generalized Poisson Flow (GPFlow), a novel generative framework that enables variable-length generative modeling by learning the rate function that minimizes the negative log-likelihood of an inhomogeneous generalized Poisson process. We establish theoretical guarantees for recovering the joint multimodal distribution (both continuous and discrete) and for an upper bound on the KL divergence to the generated distribution. We evaluate GPFlow extensively across various protein design tasks, including unconditional structure and sequence generation and conditional motif scaffolding, to validate GPFlow’s effectiveness on both continuous and discrete modalities. Our results demonstrate GPFlow’s superior generative performance and quality, achieving the top designability and distributional fitness on unconditional generation, and ranking first on 10 out of 16 motif scaffolding tasks, with up to a 10-fold improvement in success rates on the most challenging targets.


Variational Cover Modification Steganography

Martin Beneš ⋅ Rainer Böhme

Current approaches to learning-based image steganography are, at best, imperceptible, but their security is vulnerable to state-of-the-art detector networks. Variational cover modification steganography tackles this problem by training a generator network that can predict the parameters of the distribution of the least detectable stego-noise for a given cover image. It uses syndrome coding to embed the message within the distribution constraint and retrieve it exactly when a secret key is provided. Adversarial learning ensures that the generator evolves against a continuously refined detector. Experiments demonstrate significant security improvements, and several ablations support the generator architecture and the training protocol.


VEX-Bench: Benchmarking Verification Complexity of LLM-Generated Misinformation

Hanxun Huang ⋅ Oscar W ⋅ Qizhou Wang ⋅ Silvia Montaña-Niño ⋅ Yige Li ⋅ Xiang Zheng ⋅ Elif B Doyuran ⋅ Phoebe Matich ⋅ Xiao Liu ⋅ Xingjun Ma ⋅ Sarah Erfani ⋅ Christopher Leckie

Large language models (LLMs) have made misinformation inexpensive to produce but not to verify, creating a growing asymmetry in the information ecosystem. Under tight time, labor, and budget constraints, media organizations, platforms, and fact-checkers rely on screening to prioritize which content to verify. We introduce **VEX-Bench**, a unified benchmark for evaluating the *verification complexity* of LLM-generated misinformation, as perceived during screening, across models and generation methods. Verification complexity is assessed along multiple dimensions derived from journalistic and fact-checking practices, capturing checkability, harm potential, source credibility signals, imposter legitimacy, and expected verification effort. We define the $\mathrm{VEX}$ score as an integrated measure combining elicitation yield and verification complexity to quantify how generated content consumes limited verification capacity. We construct a benchmark spanning two misinformation categories, 6 high-stakes domains, and 60 real-world topics, and evaluate 7 frontier LLMs and 7 generation methods, yielding 5,880 articles. We employ an LLM-as-judge for scalable evaluation and validate it using content-analysis methodology, including ordinal Krippendorff~$\alpha$ for inter-annotator reliability, complemented by fact-checking agents for verification. Our findings show that no single method dominates all dimensions, underscoring the need for multi-dimensional evaluation. LLMs can generate high-$\mathrm{VEX}$ misinformation at 3$\times$ to 169$\times$ lower cost than agent-based verification. Such content is often prioritized during screening, consuming scarce verification resources and introducing a systematic risk of misallocation in resource-constrained verification systems.


VicEdit: Learning to Edit Videos from Visual In-Context Examples

Yuji Wang ⋅ Teng Hu ⋅ Yuheng Chen ⋅ Ran Yi ⋅ Han Feng ⋅ Weijian Cao ⋅ Chengjie Wang ⋅ Lizhuang Ma ⋅ Jiangning Zhang

Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptual gap, we propose Visual In-context Editing, a new paradigm elevating video editing from textual instructions to multi-modal visual guidance encompassing single image, image pair, and video pair. To facilitate this paradigm, we curate VicEdit-400K, the first large-scale dataset for visual in-context video editing. We develop an automated pipeline to generate 400K high-quality samples across ten task types, ensuring superior visual fidelity and semantic consistency through multi-dimensional filtering. Leveraging this foundation, we introduce VicEdit, a unified framework to bridge visual and textual contexts. To adaptively extract editing semantics from heterogeneous references, we design Modality-Adaptive Semantic Distillation, which produces modality-specific semantic tokens from visual references. These tokens are then synergistically integrated with textual instructions through Dual-Context Injection, enabling the generation process to benefit from both visual and textual signals. Extensive evaluations on VicEdit-Bench demonstrate that VicEdit achieves state-of-the-art performance across both basic instruction editing and visual in-context editing tasks, establishing visual in-context learning as a powerful and controllable paradigm for video editing.


VidUEU-Agent: A Data-Curation Agent for Multimodal Understanding, Editing, and Unified Tasks

Tengjv Ru ⋅ Weitong Lian ⋅ Zecong Tang ⋅ Lingyi Meng ⋅ Haoran Li ⋅ Zhejun Cui ⋅ Yichen Zhu ⋅ Hangshuo Cao ⋅ Qi Kang ⋅ Yechi Liu ⋅ Kaixuan Wang ⋅ Yu-Jie Yuan ⋅ Chunwei Wang ⋅ Yu Zhang ⋅ Bo Dai

The advancement of multimodal learning is heavily constrained by the scarcity and high annotation cost of high-quality instruction data. While raw videos offer an abundant and dynamic source of visual knowledge, transforming them into precise instruction data remains highly challenging. To bridge this gap, we present VidUEU-Agent, an automated pipeline that synthesizes instruction data from raw videos for multimodal understanding, editing, and unified tasks. VidUEU-Agent follows a four-stage pipeline: it performs global perception to filter low-quality shots, mines keyframes and multimodal signals, constructs metadata, and routes the metadata to synthesize instruction samples. More importantly, our agent framework can autonomously iterate the data distribution based on downstream training results, thereby optimizing data quality, balance, and training effectiveness. Scaling this pipeline, we construct VidUEU-100K, a large-scale dataset of 100K training samples spanning multiple domains and tasks, providing a robust foundation for developing comprehensive multimodal capabilities. Extensive experiments demonstrate that fine-tuning baseline models on VidUEU-100K yields significant and consistent performance gains across diverse representative benchmarks, validating the effectiveness of our proposed agent framework and dataset.


VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval

Issar Tzachor ⋅ Dvir Samuel ⋅ Rami Ben-Ari

Recent studies have adapted generative Multimodal Large Language Models (MLLMs) into embedding extractors for text--vision tasks, typically by fine-tuning them to produce universal representations. However, their performance on video--text retrieval often remains below that of dedicated Video Foundation Models (VFMs). In this paper, we revisit the video--text retrieval capabilities of MLLMs. We first analyze the zero-shot retrieval capabilities of pretrained MLLMs and show that they already encode substantial retrieval-relevant information: combining intermediate-layer embeddings with a calibrated MLLM head yields strong performance without any training. To further exploit these capabilities, we introduce a lightweight text-based optimization strategy that requires no paired multimodal data. Our method uses dense video captions as rich textual surrogates for video semantics and short summaries as compact, query-like views, showing that a carefully designed text-alignment objective can improve multimodal alignment through the shared MLLM backbone. We backup these findings through in-depth analysis. Without visual supervision our method outperforms existing approaches, often by a substantial margin, and achieves state-of-the-art performance across common video--text retrieval benchmarks. More broadly, our results offer a new perspective on retrieval with pretrained MLLMs, suggesting that textual descriptions of visual content can serve as an effective substitute for large-scale visual supervision.


View Confidence Perception-Driven Incremental Prediction for Incomplete Multi-view Multi-label Learning

Pingzhu Liu ⋅ Chunming He ⋅ Zunnan Xu ⋅ Zhirui Fang ⋅ Zitong YU ⋅ Xiu Li ⋅ Chengliang Liu

Incomplete Multi-view Multi-label Learning (IMvMIL) refers to a classification task where missing views and incomplete label assignments coexist. Existing methods often directly fuse heterogeneous view features, neglecting the fact of imbalance where dominant views tend to overshadow non-dominant ones, resulting in unsatisfactory performance. To address this, we propose the View Confidence Perception-Driven Incremental Prediction (VCPIP) framework, which incorporates an adaptive structural refinement strategy to balance different views via confidence-based branch expansion, thereby enabling the full exploitation of view-specific information. Specifically, we propose a View-Quality Aware (VQA) strategy, which introduces a novel metric to evaluate the predictive strength of each view-specific branch using supervision signals from the label space. By leveraging these quality assessments, VQA dynamically allocates optimal classifier ensembles and employs residual optimisation to improve the discriminative capability of non-dominant views. Additionally, to eliminate information redundancy and purify features, we propose a Mixture-of-Experts-based Decouple-Fusion (MOEDF) mechanism to extract refined representations by disentangling consistent and specific features under orthogonality constraints for more precise multi-label prediction. Moreover, to simulate real-world data corruption, we introduce the Stochastic Cross-sample Fragment Swapping (SCFS) strategy, which interchanges feature fragments across samples to facilitate the modelling of robust global representations for enhanced generalisation. Extensive experiments across diverse benchmarks and varying missing-rate scenarios confirm that our method consistently surpasses state-of-the-art methods in multi-label classification.

Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder perturbations affect only single-step predictions, MBRL suffers from a two-level vulnerability: visual distractions first push encoder outputs out of distribution, and these errors then compound through recursive latent rollouts over the planning horizon. We propose **VI**sual **G**eneralization via latent-space c**O**nsistency in model-based **R**L (**VIGOR**), a framework that enables zero-shot generalization to unseen visual distractions while retaining the sample efficiency of its MBRL backbone. VIGOR integrates three interdependent components: (i) *asymmetric weak-to-strong augmentation*, which pairs weak-only and weak-to-strong latent views within a single batch; (ii) *dynamics-level consistency*, which enforces augmentation-invariant transition predictions through direct latent regression; and (iii) *encoder-level stabilization*, which prevents encoder drift under the cross-augmentation supervision imposed by dynamics-level consistency. Evaluations on the DeepMind Control Suite (DMC) and Robosuite show that VIGOR outperforms state-of-the-art model-free and model-based baselines, surpassing the second-best baseline by $3.4\\%$ on DMC and $43.6\\%$ on Robosuite. Ablations further show that VIGOR's robustness is augmentation-agnostic: replacing the default augmentation with alternatives from distinct perturbation families preserves strong generalization, confirming that latent-space consistency, not the augmentation choice, drives robustness. The source code is available at https://anonymous.4open.science/r/vigor-B641/.


ViLo: LiDAR Localization with Vision-Language Priors

Minghang Zhu ⋅ Jianshi Wu ⋅ Yuxin Guo ⋅ Wen Li ⋅ Penghui Shang ⋅ Sheng Ao ⋅ Cheng Wang

LiDAR localization is a fundamental task in robotics and computer vision, aiming to estimate the global pose of point clouds. Although Scene Coordinate Regression (SCR) has demonstrated state-of-the-art performance in this field, standard SCR methods employ a monolithic network for uniform optimization across diverse scenes, which inevitably suffers from capacity interference in highly degraded scenes. Recent research on Vision-Language Models (VLMs) indicates that they can acquire rich scene understanding priors, providing the adaptive learning capabilities currently missing in localization networks. In this paper, we propose ViLo, the first framework to integrate VLM priors into SCR to fundamentally enhance localization robustness. Specifically, we design a temporally consistent key-frame querying mechanism to extract open-vocabulary degeneracy priors. We then construct a dynamic prior-guided Mixture-of-Experts (MoE) model for adaptive feature learning tailored to varying scenes. Finally, to avoid additional computational overhead during inference, ViLo internalizes the VLM's scene understanding into the 3D backbone via cross-modal distillation, enabling a lightweight, LiDAR-only deployment. Extensive experiments on the Oxford RobotCar and NCLT datasets demonstrate that ViLo significantly outperforms existing state-of-the-art methods, reducing positional errors by impressive margins of 20% and and 51%, respectively.


VisEditBench: A Benchmark for Vector-Format Diagram Editing with Visual Instructions

Akito Taneguchi ⋅ Itsumi Saito ⋅ Haruto Yoshida ⋅ Jun Suzuki

We introduce a new task and benchmark dataset, VisEditBench, for understanding visual instructions in vector-format diagram editing. While recent advances have made progress in generating vector graphics from text or images, editing such diagrams remains relatively underexplored. Moreover, existing approaches predominantly rely on textual instructions, which are often unintuitive for specifying edits in structured visual content, particularly when users need to identify specific elements or indicate their spatial positions in a diagram. We propose visual instructions as a more intuitive interface for diagram editing, and construct a dataset of 694 samples with human-curated annotations overlaid on TikZ diagrams. The dataset covers multiple edit types as well as instructions with multiple editing operations. We evaluate 32 recent state-of-the-art multimodal models and find consistent performance degradation, particularly for edits involving global layout or diagram structure, compared to local operations such as element addition or deletion. Furthermore, instructions that include multiple editing operations tend to result in lower performance than single-edit cases. Even recent models struggle to achieve consistently accurate edits, highlighting the inherent difficulty of the task. VisEditBench provides a valuable benchmark for advancing research on visually grounded diagram editing and structured visual understanding.


Vision Token Pruning via Query-Vision Interaction Decomposition

Harshithanjani Athi ⋅ Sravan Kumar Ankireddy ⋅ Jianzhong Zhang ⋅ Hyeji Kim

Large vision-language models (VLMs), which process both visual inputs and text queries, incur high inference costs due to the large number of visual tokens, making visual token pruning a natural approach for improving efficiency. Recent work has leveraged query information to guide token selection, often through heuristics derived from attention patterns or token dynamics. However, these approaches do not explicitly exploit the underlying structure of query-vision interactions. We address this gap by observing that the query-vision interaction matrix, naturally induced by the query and key projections already present in transformer attention, is effectively low-rank, with a small number of dominant latent interaction modes capturing query-relevant semantics. Motivated by this observation, we propose Query-Vision Decomposition (QViD), a novel training-free query-aware visual token pruning method that exploits this structure. Extensive experiments on both image and video understanding benchmarks show that QViD consistently improves performance under matched token budgets, with the largest gains appearing under aggressive compression regimes.


Vision Transformers Learn Gestalt-Like Figure-Ground Cues from Natural Images

Matthias Tangemann ⋅ Benjamin Lo ⋅ Zygmunt Pizlo ⋅ Kaleem Siddiqi ⋅ Dirk Bernhardt-Walther ⋅ Sven Dickinson

Figure-ground organization in the human visual system relies on several shape-based cues, including surroundedness, convexity, and symmetry. While these cues have been extensively studied using abstract stimuli, little is known about how they operate under natural conditions or how they arise from the statistics of natural scenes. Deep neural networks offer a promising path forward: a model that relies on the same figure-ground cues as humans would provide tractable experimental access to the underlying mechanisms. In this study, we evaluate shape-based figure-ground organization in Vision Transformers (ViTs), for which prior work has demonstrated the emergence of object-based grouping. We test 25 ViTs spanning supervised and self-supervised training objectives, by fitting linear probes to predict figure-ground assignment from intermediate patch representations using both natural images and controlled artificial stimuli that isolate individual cues. Our results show that ViTs robustly encode surroundedness and convexity, and that probes trained on natural images generalize zero-shot to artificial stimuli across several models. For symmetry we observe mixed results: the cue is encoded for uniformly colored but not for textured regions. Taken together, our findings demonstrate that Gestalt-like figure-ground cues can be learned from natural scene statistics and position ViTs as a compelling model system for studying the computational mechanisms of perceptual organization. All code and data will be published.


Visual-Advantage On-Policy Distillation for Vision-Language Models

Ruiqi Liu ⋅ Xiaolei Lv ⋅ Gengsheng Li ⋅ Ximo Zhu ⋅ Zhiheng Wang ⋅ Zhengbo Zhang ⋅ Junkai Chen ⋅ Zhiheng Li ⋅ Bo Li ⋅ Jun Gao ⋅ Shu Wu

On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-policy distillation can improve a student's output quality while failing to strengthen its reliance on visual input: on vision-critical tokens, the student's predictions remain largely unchanged whether or not fine-grained visual detail is present, even though the teacher's predictions depend heavily on it. To make this difference observable, we introduce visual advantage (VA), the token-level log-probability difference when the teacher scores a student-generated rollout with versus without access to fine-grained visual detail. VA is concentrated in a small minority of tokens, and these high-VA tokens are the ones that actually carry the visual supervision signal. This motivates a distillation objective that treats them differently from language scaffolding, so their contribution is not diluted by the abundant surrounding language tokens. We propose Visual-Advantage On-Policy Distillation (VA-OPD), which uses VA at two granularities: rollout-level reweighting by trajectory-averaged VA, and token-level KL averaged within high-VA and low-VA groups separately. We train on two math datasets (Geometry3K and ViRL-39k) and evaluate on eight benchmarks covering both mathematical reasoning and visual understanding, across three teacher sizes (4B, 8B, and 32B) on the Qwen3-VL family. VA-OPD improves over standard on-policy distillation on every benchmark, with the gain growing monotonically along both the teacher-size and data-scale axes, suggesting that these factors compound consistently.


VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon

Yi Pan ⋅ Miao Pan ⋅ Qi Lu ⋅ Jiaming Huang ⋅ Man Zhang ⋅ Siteng Huang ⋅ Xin Li ⋅ Jie Zhang ⋅ Yongliang Shen ⋅ Xuhong Zhang ⋅ Wenqi Zhang

Vision-Language-Action (VLA) foundation models have recently achieved strong progress in embodied intelligence. To reduce policy-call frequency while preserving temporal coherence, most generative policies adopt an \textbf{action chunk} mechanism, executing multiple future actions in an open-loop manner under a fixed \textbf{action horizon}. However, this ``predict-then-blindly-execute'' paradigm sacrifices closed-loop reactivity: in contact-rich physical interactions, even small local perturbations can rapidly amplify within the open-loop blind spot, leading to compounding errors and ultimately task failure. To address this limitation, we propose \textbf{VLA-Corrector}, a lightweight corrective inference framework for action-chunked VLA policies. Without modifying the backbone policy weights, VLA-Corrector introduces a lightweight \textbf{Latent-space Vision Monitor} (LVM) that continuously compares predicted and actual visual feature evolution, enabling online detection of visual dynamics deviations. Once persistent deviation is detected, the system triggers a truncation event, discards the remaining stale actions, and invokes corrective replanning via \textbf{Online Gradient Guidance} (OGG). The detect-and-correct mechanism of VLA-Corrector naturally induces an \textbf{event-triggered adaptive action horizon}: it preserves long-horizon execution when the current chunk remains reliable, and invokes short-horizon corrective replanning when execution begins to drift. In doing so, VLA-Corrector mitigates the trade-off imposed by static horizons between execution robustness and policy-call frequency. It can be integrated into different VLA models without further retraining the VLA backbone, interrupting compounding errors while preserving much of the efficiency benefit of action chunking and substantially improving robustness in long-horizon, contact-rich robotic manipulation tasks.


VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following

Hyesoo Hong ⋅ Min S Kim ⋅ Wonje Jeung ⋅ Yoon Sangyeon ⋅ Dongjae Jeon ⋅ Albert No

Vision-language models (VLMs) achieve strong performance on multimodal benchmarks, but may still lack robust control over basic visual operations. We study \textit{line tracing}, where a model must follow a selected visual path through successive local continuations. To isolate this ability, we design controlled tracing tasks that introduce nearby competitors while reducing semantic and topological ambiguity such as crossings and overlaps. Across these tasks, even state-of-the-art VLMs frequently lose the target path and switch to nearby alternatives, especially when those alternatives look locally similar to the target. Behavioral interventions and internal analyses indicate that these failures arise from local competition: nearby similar distractors pull the model away from the true continuation. Standard remedies do not remove this bottleneck: model-size scaling provides only limited gains, reasoning partially compensates through costly substitute strategies, and explicit tracing instructions fail to recover stable path following. Finally, tests on tangled-cable scenes and metro maps with richer visual complexity show that the same path-switching failure persists beyond our controlled settings.


VLRS-Bench: A Vision-Language Reasoning Benchmark for Remote Sensing

Zhiming Luo ⋅ Hebaixu Wang ⋅ Haonan Guo ⋅ Jing Zhang ⋅ Bo Du

Recent advancements in Multimodal Large Language Models (MLLMs) have enabled complex reasoning. However, existing remote sensing (RS) benchmarks remain heavily biased toward perception tasks, such as object recognition and scene classification. This limitation hinders the development of MLLMs for cognitively demanding RS applications. To address this, we propose a Vision Language ReaSoning Benchmark (VLRS-Bench), which is the first benchmark exclusively dedicated to complex RS reasoning. Structured across the three core dimensions of Cognition, Decision, and Prediction, VLRS-Bench comprises 2,000 question-answer pairs with an average question length of 130.19 words, spanning 14 tasks and up to eight temporal phases. VLRS-Bench is constructed via a specialized pipeline that integrates RS-specific priors and expert knowledge to ensure geospatial realism and reasoning complexity. Experimental results reveal significant bottlenecks in existing state-of-the-art MLLMs, providing critical insights for advancing multimodal reasoning within the remote sensing community.


V-LUMEN: Visual Lookup Memory for Embedding Scaling in Vision-Language Models

Miso Choi ⋅ Daekeun Kim ⋅ Eunji Kim ⋅ Jungbeom Lee

Embedding scaling has recently emerged as a promising direction for increasing the capacity of language models through embedding-level representations, rather than solely expanding transformer depth or width. However, existing embedding scaling methods are primarily designed for discrete text tokens, leaving it unclear whether similar mechanisms can be extended to vision-language models, where visual representations are continuous, high-dimensional, and spatially structured. In this paper, we introduce V-LUMEN, a visual lookup memory framework for embedding-level scaling in vision-language models. V-LUMEN augments visual-language representations with reusable visual embeddings retrieved from an external memory indexed by discretized visual patterns, thereby expanding accessible representational capacity without directly increasing the dense computation of the backbone model. To address the unique challenges of visual representations, V-LUMEN introduces spatial aggregation for constructing structured visual memory keys and text-conditioned hashing for retrieving task-relevant visual memories based on the input query. Together, these components enable memory retrieval that is both spatially aware and instruction-conditioned. Experiments across diverse reasoning-intensive multimodal benchmarks show that V-LUMEN effectively augments vision-language models through external visual memory. Without backbone adaptation, V-LUMEN substantially improves the vanilla Qwen3-VL-2B-Instruct model, and further achieves competitive or better performance against parameter-matched MoE and LoRA baselines, demonstrating the potential of visual embedding scaling as an alternative scaling axis for vision-language models.


Vortex: Efficient and Programmable Sparse Attention Serving

Zhuoming Chen ⋅ Xinrui Zhong ⋅ Qilong Feng ⋅ Ranajoy Sadhukhan ⋅ Yang Zhou ⋅ Michael Shieh ⋅ Zhihao Jia ⋅ Beidi Chen

Sparse attention is becoming increasingly important for serving large language models (LLMs) as generation lengths continue to grow. However, deploying and evaluating new sparse attention algorithms at scale remains highly engineering-intensive, slowing both human researchers and AI agents in exploring the sparse attention design. To address this challenge, we present Vortex, a system that combines a Python-embedded frontend language atop a page-centric tensor abstraction for expressing a broad range of sparse attention algorithms, with an efficient backend tightly integrated into modern LLM serving stacks. Vortex enables rapid prototyping, deployment, and evaluation of sparse attention algorithms, effectively translating their theoretical efficiency gains into real-world throughput improvements. As a result, Vortex substantially accelerates the design and iteration of sparse attention algorithms. Using Vortex, we further demonstrate AI-agent-driven exploration of the sparse attention, automatically generating and refining diverse algorithms that achieve strong accuracy-throughput trade-offs, with the best generated algorithms delivering up to $3.46\times$ higher throughput than full attention while preserving accuracy.


Walking the Hypercube: Unbiased Quantum Partition Functions Without the Matrix

Kumar Avinava Dubey ⋅ Arijit Sehanobish ⋅ Krzysztof M Choromanski

Computing the partition function $\mathcal{Z} = \mathrm{tr}(e^{-\beta H})$ of a quantum spin Hamiltonian $H$ on $n$ sites is a fundamental problem in statistical physics and quantum chemistry. However exact evaluation requires $O(8^n)$ time and $O(4^n)$ memory making it intractable for all but the smallest systems. We propose an unbiased stochastic estimator of $\mathcal{Z}$, and more generally of $\mathrm{tr}(f(H))$ for any function $f$ admitting a convergent power series over exponentially large, implicitly defined operators using only local random walks, avoiding both matrix materialization and matrix-vector products. Our approach combines Graph Random Features with the Hutchinson's stochastic trace estimator, and exploits the algebraic structure of quantum spin Hamiltonians to reduce the per-step computational cost to $O(n + d)$, where $d$ is the mean degree of the spin coupling graph, independent of the Hilbert space dimension $N = 2^n$. We validate the estimator empirically against exact matrix exponentiation for small system sizes, and demonstrate its applicability to system sizes where exact methods are infeasible.


WASP: Weakly Aligned Spatiotemporal Pairs for Fetal Brain MRI-Ultrasound Learning

Francesco Correnti ⋅ Gabriele Magrini ⋅ Marco Mistretta ⋅ Niccolò Biondi ⋅ Pietro Pala ⋅ Alessandro Ramalli ⋅ simona fiori ⋅ Andrew Bagdanov ⋅ Matteo Lenge

Magnetic Resonance Imaging (MRI) is widely regarded as the optimal sensor for fetal brain analysis due to its superior soft tissue contrast and anatomical detail. However, its high cost and operational burden make it invasive and difficult to obtain at scale. Ultrasound (US), in contrast, is cheap, safe, and routinely acquired, and as a result it has produced substantially larger datasets and a growing ecosystem of pretrained models. This asymmetry raises a natural question: Can we teach a US-only model to understand fetal MRI from only a limited set of examples? The standard recipe, training a foundation model on subject-to-subject paired MRI-US scans, is not viable since no such paired fetal dataset is publicly available. In this paper we address this gap with Weakly Aligned Spatiotemporal Pairs (WASP), a framework that formulates cross-modal correspondence as an entropic Optimal Transport problem driven by clinical metadata — in particular, Gestational Age (GA) and diagnostic planes — enabling the fitting of a lightweight alignment head that lifts MRI representations into the US latent space, without fine-tuning the backbone. Empirically, WASP yields its largest gains when MRI is unseen by the model during pretraining (on USFM, GA estimation error drops from $21.9$ to $16.3$ days and standard plane classification accuracy climbs from 61.9\% to 74.5\%), while still providing refining improvements for backbones pretrained on both modalities (e.g., BioMedParse GA estimation error from $6.5$ to $5.6$ days and SAM-Med2d plane accuracy from $83.3$\% to $88.1$\%).

Video reasoning has a distinctive structure: in a chain of thought over video, most tokens follow from prior text, but a small minority hinge on specific moments of the visual input---and which tokens those are varies from one example to the next. Current post-training procedures like GRPO supply supervision that is coarse in exactly this dimension, applying a single scalar reward uniformly across the trajectory while requiring many rollouts and/or costly external judge models. Our key observation is that a model can audit its own visual reliance: at any position, the shift in its next-token distribution when the video is removed from context measures how much that prediction actually used what was seen. Computing this quantity for both a student and a privileged teacher yields a per-token \emph{reliance gap} that localizes positions at which the teacher's predictions depend on the video and the student's do not. We introduce Visually-Aware On-Policy Self-Distillation (VIOS), which turns this self-auditing signal into dense per-token supervision recovered from a single rollout of a single model---no external reward, judge, or larger teacher required. VIOS combines three components: an exponential moving average (EMA) teacher that co-evolves with the student to avoid the ceiling of a frozen rationalizer; a contrastive term that increases the student's sensitivity to the video input; and a reliance gap gate that concentrates supervision on the positions identified. Across 12 benchmarks spanning general, temporal, spatial, knowledge, and long-video understanding, VIOS consistently outperforms its base model and contemporary video-reasoning systems while training in over $9\times$ fewer steps.


WavCIL: Wavelet Coefficient-Domain Invariant Learning for Dynamic Graph OOD Generalization

Yantong Zhu ⋅ Weining Shi ⋅ Keyi Li ⋅ Xiaoyan Xie ⋅ Yi Chang ⋅ Qinggang Zhang ⋅ Tianjiao Pu ⋅ Ji Qiao ⋅ Zhihong Zhang

Dynamic Graph Neural Networks (DyGNNs) demonstrate powerful representation capabilities in crucial time-series systems by leveraging graph structure and temporal dynamics. However, existing dynamic graph models often suffer from unpredictable distribution shifts induced by complex time-varying factors such as emergency incidents, leading to severe performance degradation due to their low out-of-distribution (OOD) generalization. Graph invariant learning has been extensively studied to solve this problem through feature disentanglement in the time or spectral domain. Despite recent advances, existing methods suffer from critical limitations in practice: Temporal-based methods struggle to disentangle transient perturbations and stable patterns that are highly overlapping and entangled within the same period. While spectral-based methods attempt spectral-domain disentanglement via the global Fourier transform, these methods inherently lose temporal localization and fail to effectively characterize distribution shifts driven by time-varying factors. To this end, in this paper, we propose Wavelet Coefficient-Domain Invariant Learning (WavCIL), a novel framework for resolving distribution shifts by conducting disentanglement within a wavelet coefficient domain endowed with time-frequency localization capabilities. Specifically, WavCIL consists of two key components: (i) a wavelet transform module, which leverages a set of wavelet bases to map the input temporal signals into wavelet coefficient domain, simultaneously capturing frequency components and localizing their temporal occurrences; and (ii) a coefficient domain invariant learning module, which utilizes an adaptive mask mechanism to disentangle stable causal patterns invariant to time-varying factors from spurious patterns driven by time-varying factors in the coefficient domain, guiding the prediction process to rely more on the causal patterns. Extensive experiments on several benchmark datasets demonstrate that WavCIL achieves state-of-the-art generalization performance when handling distribution shifts. Our code is released at https://anonymous.4open.science/r/WavCIL-4B8C.

This work proposes WaveGen, a fully end-to-end diffusion-based text-to-waveform framework based on Spectral Forcing, which jointly learns an internal spectral trajectory and the target waveform trajectory within a single model, enabling direct end-to-end text-to-waveform generation. Unlike recent TTS pipelines, WaveGen does not rely on separately trained or externally decoded intermediate representations, such as neural audio codecs, Mel-spectrogram vocoders, or VAE-based latent models. In addition, WaveGen performs internal text-speech alignment within a single model, eliminating the need for external alignment modules such as duration predictors. To preserve a truly end-to-end formulation, our framework further avoids dependence on self-supervised representations, such as BERT-based text embeddings and Wav2Vec 2.0-based semantic representations, as well as semantic distillation methods for training optimization. To achieve this, we carefully design Spectral Forcing, a model architecture with disentangled diffusion heads, and a training objective based on a multi-scale spectral $v$-loss. We further validate the scalability of end-to-end TTS models by scaling the model size from 0.1B to 0.4B parameters. With a single architecture, open-source data, and one-stage training without external modules, WaveGen demonstrates fully end-to-end waveform generation by achieving promising performance. In particular, WaveGen achieves a WER of 1.77, a SIM of 0.63, and competitive MOS performance on the Seed-en benchmark.

Graphs with a simple spectrum admit cubic-time isomorphism testing, yet we prove that for every natural number $k$, the $k$-Weisfeiler-Leman ($k$-WL) test cannot distinguish all non-isomorphic graphs with a simple spectrum. As the WL hierarchy upper-bounds the distinguishing power of widely-used Graph Neural Networks (GNNs), this incompleteness applies to all such GNNs, ruling out completeness for every $k$-WL-aligned GNN family. To close this gap, we introduce PRiSM (Partition, Refine, Solve, Match), the first provably complete canonicalization of simple-spectrum eigendecompositions. PRiSM obtains the completeness guarantee that prior canonicalizations provably lack, and resolves the open problem of achieving complete expressivity on simple-spectrum graphs. When composed with DeepSets or a Transformer, PRiSM achieves universal approximation on simple-spectrum graphs, justifying the use of canonicalized Laplacian positional encodings. Empirically, PRiSM performs comparably to or outperforms existing spectral canonicalizations on graph regression, classification, and expressivity benchmarks.

Mechanistic interpretability often studies the local features and circuits that implement model computations. What principles govern the arrangement of these features and circuits into geometric structures in activation space? To make this tractable, we study how the computational class of the training-data generator constrains the geometry of predictive states. We show that while the data distribution determines which features are required for prediction, a predictor realizes those features as beliefs about its current latent state, and the generator class determines the geometry of those beliefs. Using this theoretical insight, we design synthetic datasets whose minimal predictive representations fall into different model classes, and test which geometry neural networks learn. In particular, we train transformers, LSTMs, GRUs, and vanilla RNNs on datasets whose predictive geometries are known analytically: a classical HMM process, a quantum-realizable process with no finite-state HMM realization, and a generalized-probabilistic process with no finite-dimensional quantum realization. Across architectures, a single affine map from activations decodes the corresponding predictive representation in each case: HMM beliefs in a latent simplex, Bloch-vector quantum states, or a finite-dimensional generalized predictive vector. These representations emerge during training and fit the compact non-classical geometry far better than finite-order classical Markov baselines. These results suggest that understanding predictive representations requires asking not only which features a network represents, but what geometry organizes those features.


What Cohort INRs Encode, and Where to Freeze Them

Vasiliki Sideri-Lampretsa ⋅ Sophie Starck ⋅ Robbie Holland ⋅ Julian McGinnis ⋅ Daniel Rueckert

Reusing the early layers of cohort-trained INRs as initialization for new signals has been shown to accelerate and improve signal fitting, yet it remains unclear which layers of the shared encoder learn transferable representations and what those representations encode. We address both questions for two standard backbones, SIREN and Fourier-feature MLPs (FFMLP). First, sweeping the freeze depth across the shared encoder at test time, we find that the optimum coincides with the layer of highest weight stable rank. Moreover, freezing at this depth matches or improves on the standard fine-tuning recipe across all our experiments. Second, identifying which layer transfers does not characterize what that layer encodes. To address this we adopt sparse autoencoders (SAEs), the dominant tool in mechanistic interpretability, and present the first SAE decomposition of INR activations into sparse dictionary atoms. Interestingly, SIREN and FFMLP achieve comparable cohort-fitting quality, but learn qualitatively different dictionaries. Cohort SIREN's atoms are localized, tiling the coordinate plane such that each atom fires in a confined region independent of cohort content. Cohort FFMLP's atoms are image-spanning, tracing the contours of memorized cohort signals. Single-atom ablations confirm causal use of these dictionaries: a single FFMLP atom out of 4096 can drop PSNR by up to 10.6 dB across the image, while SIREN ablations remain confined to where the atom fires. Together, these results give the first mechanistic account of what transfers in cohort-trained INRs and turn their activations into inspectable dictionary atoms. These tools open a path towards characterizing what INRs encode and towards architectures designed for generalization rather than memorization. We plan to release our code upon acceptance.


What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion

Zhengrong Yue ⋅ Taihang Hu ⋅ Mengting Chen ⋅ Haiyu Zhang ⋅ Zihao Pan ⋅ Tao Liu ⋅ Zikang Wang ⋅ Jinsong Lan ⋅ Xiaoyong Zhu ⋅ Bo Zheng ⋅ Yali Wang

Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing tokenizers are primarily designed to improve reconstruction fidelity or inherit pretrained representations, leaving unclear what kind of latent space is truly friendly for generative modeling. In this paper, we study this question from the perspective of latent manifold organization. By constructing controlled tokenizer variants, we identify three key properties of a diffusion-friendly latent manifold: coherent spatial structure, local manifold continuity, and global manifold semantics. We find that these properties are more consistent with downstream generation quality than reconstruction fidelity. Motivated by this finding, we propose the $\textit{\textbf{P}rior-\textbf{A}ligned Auto\textbf{E}ncoder} (\textbf{PAE})$, which explicitly shapes the latent manifold instead of leaving diffusion-friendly manifold to emerge indirectly from reconstruction or inheritance. Specifically, PAE leverages refined VFM-derived priors and perturbation-based regularization to turn spatial structure, local continuity, and global semantic organization into explicit training objectives. On ImageNet $256{\times}256$, PAE improves both training efficiency and generation quality over existing tokenizers, reaching comparable performance up to $13\times$ faster than RAE under the same LightningDiT setup and achieving a new state-of-the-art gFID of $\textbf{1.03}$. These results highlight the importance of organizing the latent manifold for latent diffusion models.


What Sketches Tell Us about LVLMs: Conventions, Grounding, and Localisation

Chaitat Utintu ⋅ Ahmed Bourouis ⋅ ARKAPRABHA BASU ⋅ Yi-Zhe Song

LVLM benchmarks tell us whether a model answers correctly, but not which visual conventions it has retained, nor where those conventions live. We use sketches to expose this hidden structure. A sketch removes texture, lighting, and photographic context while preserving the strokes humans judged sufficient for recognition; if an LVLM can operate on that abstraction, the surviving signal is structural rather than merely photometric. We introduce a training-free sketch probe with three readouts: temporal part-construction order, spatial attention-based grounding, and the mechanistic layer/head locus of grounding. We evaluate across nine LVLMs and five LLaVA-architecture variants on two part-annotated sketch datasets. Three findings emerge. First, LVLM part order tracks human convention: models reproduce canonical orderings where human drawings converge, and become correspondingly variable where human orderings vary. Second, selected frozen LVLM attention heads ground object and scene sketches competitively with specialised grounders, exceeding CLIPSeg and GroupViT on mAcc@0.1 despite using no grounding supervision. Third, the grounding signal concentrates in a narrow mid-network band whose location follows the base language model rather than parameter count, and remains stable across object and scene sketches. The result is a practical diagnostic for sketch-driven LVLMs: no fine-tuning, no auxiliary segmenter, and a small architecture-dependent layer band that doubles as a head-selection prior, improving grounding mAcc@0.1 in every model and dataset we tested.


What the Geometry of Good Models Tells Us

Alexis Fox ⋅ Samuel Orellana Mateo ⋅ Krish Yadav ⋅ Yiyang Sun ⋅ Zachery Boner ⋅ Cynthia Rudin

Practitioners using machine learning on noisy tabular data can often find simpler, more interpretable models with the same accuracy as black box models. In some cases, this is because the noise in the data induces a large Rashomon set, that is, a large set of near-optimal models, which then frequently contains simpler models. However, our understanding of this effect from discrete hypothesis classes --- the Rashomon set's size, the diversity of its predictions, and implicit regularization --- leads to a contradiction in continuous hypothesis spaces, and thus does not directly establish whether noise should necessarily induce simpler or more diverse models. We reconcile the seemingly conflicting effects by showing geometrically that noise drives simplicity through separate mechanisms for linear, logistic, and generalized additive models. We introduce the Rashomon reducibility cost to quantify whether a given model simplification is possible within the Rashomon set, and characterize how this cost changes under general data properties such as label noise. For practitioners working with noisy tabular data, our results suggest that simpler, more interpretable models will be included in the Rashomon set, so that black box models will generally not be needed.


What to Forget in Unlearning? Forget Set Curation for Language Models

Animesh Jha ⋅ Arpandeep Khatua ⋅ Youssef Allouah ⋅ Sanmi Koyejo

Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget are already known. In realistic language-model deployments, a requester may ask a model to stop reproducing a song or book without knowing which spans, documents, quotations, or near-duplicates in a trillion-token corpus support that behavior. We study this missing upstream problem, forget set curation: mapping a suppression request to the data passed to an unlearning algorithm. We introduce CleanSlate, a benchmark for verbatim output suppression over songs and books, with model-specific extraction profiles, content-grounded QA, and capability-retention evaluations. CleanSlate exposes two failure modes. Natural lexical and exact-substring curators often yield forget sets that lead to weak suppression. An evaluation-aware curator suppresses requested continuations almost completely, but causes collateral regression on non-requested content and model-dependent capability loss. These results show that practical unlearning is not only an optimization problem once a forget set is given: the data chosen for forgetting determines both what can be unlearnt and what else is damaged.


When Are Predictions Enough? An Evaluation Protocol for Frozen Expert Composition

Alessandro Pereira ⋅ Keith J Ransom ⋅ Lewis Mitchell

Practitioners often combine independently trained models as a fixed pool of frozen experts. Standard combination methods, such as voting, averaging, or stacking, rely on final predictions, amounting to a few scalars per expert, and discard potentially richer cross-expert information found in the penultimate representations. An alternative is to combine these models in feature space. Whether this yields a meaningful improvement depends heavily on the evaluation design. Informal assessments on the same data can support materially different conclusions depending on baseline strength, overlap control, reporting granularity, and pool construction. We argue that auditing the sensitivity of evaluation outcomes to these design choices is critical for the effective composition of expert models. To address these sensitivities, we propose a four-part evaluation protocol: (1) a baseline ladder for prediction-space (PS) models; (2) a deterministic feature-space (FS) anchor plus a higher-capacity cross-check; (3) axis-isolated pools to disentangle performance gains; and (4) grouped evaluation with overlap-aware controls. Applying our protocol to 21 frozen deepfake detectors from DF40 (Yan et al., 2024), we find that a 10K-parameter logistic regression on raw concatenated features outperforms the strongest PS baseline by +0.123 AUROC on an architecture-diverse pool (generator-level bootstrap 95% CI [+0.078, +0.174]). Crucially, the higher-capacity cross-check (a 3.5M-parameter feature-space MLP) adds only +0.004 AUROC, suggesting that the gain arises from the richer feature interface rather than from combiner complexity. Prediction-space models suffice when individual expert performance is high but fall short in architecture-diverse, high-bottleneck pools. Our primary contribution is an evaluation-design audit showing that, on this benchmark and expert pool, relaxing any of these controls risks overstating the apparent FS advantage or creating the illusion of a gain that fails to hold across data subgroups. Our findings demonstrate that, without rigorous controls, the same data can support multiple, sometimes spurious, conclusions.


When to Inject the Target: Stage-Decoupled Guidance for Diffusion-Based Targeted Adversarial Attacks

JUNFENG YU ⋅ Junkai Bao ⋅ Hao Zeng ⋅ XIAOWEN MO ⋅ Xin Liu ⋅ Xuejian Huang

Diffusion-based generative attacks have emerged as a promising paradigm for targeted adversarial example generation, leveraging generative priors to improve transferability and visual fidelity. However, existing methods typically entangle image reconstruction and target manipulation within the same reverse denoising trajectory, where target guidance can disrupt early structural recovery. This coupling creates a persistent trade-off among attack effectiveness, transformation robustness, and perceptual fidelity. In this paper, we propose SDTI (Stage Decoupled Target Injection), a stage-decoupled reverse diffusion framework that separates source-structure preservation from adversarial target injection along the denoising timeline. SDTI first maps the input image to an intermediate latent through null-text DDIM inversion, then recovers and anchors the source structure with a frozen null-text denoising prefix, and finally activates target-conditioned LoRA guidance only in the late denoising suffix to inject adversarial target semantics. To confine adversarial optimization to this suffix phase, we further introduce suffix-wise supervision with inter-step gradient truncation. Extensive experiments on ImageNet-NeurIPS, robust victim models, common input transformations, and MS-COCO cross-domain transfer show that SDTI achieves a favorable attack--fidelity--robustness trade-off, improving transformation robustness and perceptual quality while maintaining competitive targeted transferability. The code is available at: https://anonymous.4open.science/r/AY_code-1431.


When to Trust a PFN: Detecting Harmful Shift in Tabular Foundation Models

Viet Nguyen ⋅ Herman Bergström ⋅ Stephan Rabanser ⋅ Rahul Krishnan

Tabular prior-fitted networks (PFNs) work well across tabular classification problems but can fail silently at deployment when distributions shift. We propose CALD (Calibrated Anchor-Loss Disagreement), a lightweight unsupervised monitor that triggers an alarm to warn when a PFN might fail. By finetuning the decoder to maximize disagreement with the test-set prediction while maintaining accuracy on an ID anchor set, CALD identifies harmful OOD batches through the resulting anchor loss, requiring no labels on test batches. We prove that the post-finetune anchor loss has a strictly lower expectation under harmful shift than under ID eval, yielding a one-sided test with a formal power guarantee. Across 14 datasets covering tabular and structured distribution shift, on both TabPFN and TabICL, CALD matches or outperforms baselines on harmful shifts and is markedly more resilient to benign-shift false alarms.


When to Trust Memory: Retrieval-Guided Probabilistic Spatiotemporal Forecasting under Distribution Shift

Haochen Lv ⋅ Jianhao Zhang ⋅ Zhichen Lei ⋅ Yang Jiang ⋅ Xiaowei Mao ⋅ Shengnan Guo ⋅ Youfang Lin ⋅ Huaiyu Wan

Spatiotemporal forecasting plays an important role in real-world applications, where probabilistic approaches provide uncertainty estimates for decision-making. Real-world systems are inherently non-stationary, leading to distribution shifts between training and test data. However, existing probabilistic methods mainly focus on modeling data distributions under standard settings, while distribution shift methods often handle such shifts in a coarse manner, failing to characterize shift degree and adapt prediction strategy accordingly. As a result, they often produce miscalibrated and unreliable forecasts. To address this issue, we propose REMIND, a retrieval-guided probabilistic framework for spatiotemporal forecasting under distribution shift. REMIND leverages retrieval to estimate how well current observation are supported by historical patterns, enabling adaptive prediction under different shift regimes. It employs spatiotemporal retrieval to construct a memory bank and introduces a dual branch architecture, combining a memory-calibrated probabilistic branch for mild shifts with a deviation-aware probabilistic branch for severe shifts. A shift-adaptive gate, guided by retrieval similarity and memory dispersion, balances the two branches according to shift degree. We further design a hierarchical Gaussian mixture to capture heterogeneous uncertainty over the dual branch predictions, and introduce a memory perturbation training scheme to improve generalization. Experiments on real-world datasets with distribution shifts demonstrate that REMIND consistently outperforms state-of-the-art methods in both forecasting accuracy and probabilistic prediction quality. Code is available at https://anonymous.4open.science/r/REMIND-A46E/.

Subject-driven diffusion transformers preserve coarse object identity but routinely lose identity-defining details such as logos, printed text, textures, and small structural cues. We show that this failure traces to a structural mismatch: identity sensitivity is sparse and non-uniform across layers, denoising stages, and spatial regions, yet reference influence is typically applied as a global signal. We introduce \textbf{VitalScore}, a training-free method that discovers this identity-vital structure through offline probing and uses it to build a reference-attention redistribution tensor $\Lambda(l,t,s)$ coordinated with a stage-aware prompt schedule $(P(t)$. The spatial axis is grounded in a two-phase VLM blueprint calibrated against human identity annotations, ensuring the controller reflects the cues people actually use to judge subject identity. We evaluate across Diptych Prompting, FLUX Kontext, and FLUX 2.0 on DreamBench++, where VitalScore improves DINO-v2 by up to 15%, and on a new product identity benchmark targeting fine-grained commercial cues that standard benchmarks overlook, where gains reach 16.5% on the hardest prompt axis---all without parameter updates.


Where Do Long Captions Fail? Position-Aware Diagnosis and Reinforcement Learning for Detailed Image Captioning

Yuanze Hu ⋅ Zhichao Yang ⋅ Junwei Jing ⋅ Xin Yang ⋅ Wenxuan Liu ⋅ Tinghai Zhang ⋅ Jia Fu ⋅ Hongzhi Zhang ⋅ Xiangyu Wu ⋅ Ruiming Tang ⋅ Tingting Gao ⋅ Han Li

Detailed image captioning is usually framed as an amount problem: a model should mention more visual facts while avoiding hallucination. We argue that this framing hides a basic structure of long-form captioning: a caption is an ordered sequence, and different failures appear at different positions. We introduce a position-aware rubric representation that decomposes dense human references into atomic visual claims, verifies generated captions against those claims, and induces a quartile-wise composition of supported, hallucinated, and generic content. This lens reveals two qualitatively different failures in long captions: models tend to produce unsupported specificity in the middle, while the final portion often escapes into vague, non-evidential language rather than simply accumulating more hallucination. Dense human references do not exhibit the same position-dependent pattern under the identical pipeline, suggesting that the effect is a model behavior rather than an artifact of long descriptive text. Motivated by this diagnosis, we propose PAGER (Position-Aware Grounded Evidence Reinforcement), a GRPO training objective with separate signals for supported fine-grained middle evidence, evidence-bearing tail continuation, claim coverage, and content-aware length control. The result is a closed loop: the same ordered claim representation identifies where long captions fail and supplies the position-specific reward needed to repair those failures. Experiments on detailed-caption benchmarks show that PAGER improves position-wise grounding and pairwise caption preference while preserving competitive global caption quality.


Why Jailbreaks Succeed in Diffusion Language Models: An Energy Landscape Analysis

Thong Bach ⋅ Dung Nguyen ⋅ Thao Le ⋅ Truyen Tran

Existing attacks and defenses for diffusion-based large language models (dLLMs) target specific vulnerabilities but lack a shared framework explaining why attacks succeed. We propose one by interpreting safety alignment as shaping the denoising energy landscape: a well-aligned model routes harmful queries toward safe outputs through an energy barrier that separates the two regions. Current jailbreak attacks reduce to two strategies for circumventing this barrier: obscuring the query's safety disposition at initialisation, or intervening mid-trajectory to force the denoising path across the energy barrier. From this perspective and the result that masked diffusion models minimise kinetic energy during denoising, we derive three complementary, training-free detection signals: a step-0 ratio that reads the initial safety disposition from the logit distribution before generation begins, and two trajectory-velocity signals that track kinetic energy in complementary subspaces of the logit space. An attack must either reveal its intent at initialisation or expend kinetic energy to cross the barrier in at least one monitored subspace, so the three signals cover each other's blind spots in the energy budget by construction. Evaluation across three model families (LLaDA-8B, LLaDA-1.5, Dream-7B) confirms this complementarity. In stress tests of known attacks, every configuration that evades detection also fails to produce harmful content, suggesting that the detection and barrier-crossing thresholds are hard to separate.


Why Routers Freeze: Infinite Width Learning Dynamics for Mixture of Experts

Anish Dhir ⋅ Volkan Cevher ⋅ Leena Chennuru Vankadara

Mixture-of-Experts (MoE) architectures are central to large-scale deep learning, relying on sparse execution and expert specialisation. Despite their success, their large-scale training dynamics remain poorly understood, and training is often unstable. While scaling limits such as infinite-width theory have clarified the behavior of dense models, the scaling behavior of MoEs remains largely unexplored. To this end, we derive the infinite-width limits of MoE architectures under fixed number of experts with both soft and Top-$K$ routing, trained with SGD and Adam, using Tensor Programs. Our results show that under the Standard Parameterisation (SP), router updates freeze after one step of training in both soft and Top-$K$ routing MoEs. This provides, at least in part, a principled explanation for the widespread use of auxiliary losses such as load-balancing or $z$-losses in MoE training. In dense networks, Maximal Update heuristics allow for stable and non-vanishing feature updates. However, in MoEs the Maximal Update heuristic with soft-routing and softmax gating causes the experts to collapse to the same distribution and router gradients to vanish. While sigmoid allows non-trivial router evolution, they still fail to prevent expert collapse. Thus fixed expert soft-routing MoEs do not admit a parameterisation that simultaneously ensures stability, feature learning, and expert specialisation in the infinite-width limit. We then show that, beyond computational efficiency, Top-\(K\) also acts as a symmetry breaking mechanism allowing for both feature learning and specialisation. We empirically validate these predictions, observing close agreement between finite-width training dynamics and the predicted scaling behaviour.


Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression

Shuochen Chang ⋅ Qingyang Liu ⋅ Shaobo Wang ⋅ Bingjie Gao ⋅ Qianli Ma ⋅ Haonan Zhao ⋅ Yibo Miao ⋅ Yulin Sun ⋅ Zelin Peng ⋅ Jiangtong Li ⋅ Li Niu

Large language models achieve high reasoning performance via explicit chain-of-thought and reinforcement learning, but require long output sequences and extended inference time. Latent reasoning reduces this cost by shifting computation into a latent space; however, continuous latent methods are hard to train, suffering from unstable and uninterpretable reasoning trajectories. We argue these issues stem from a misalignment between continuous-space reasoning and discrete symbolic supervision, as continuous states lack explicit anchors for step-by-step alignment. To resolve this, we propose **Discrete Latent Reasoning (DLR)**, the first method that converts continuous latent states into explicit discrete tokens. Inspired by render-based compression, we render textual chains of thought into images, extract visual features, and construct a discrete latent vocabulary via clustering-based fine-tuning. Expanding the vocabulary and output head enables standard autoregressive modeling over both natural language and latent tokens, supporting pretraining alignment, SFT, and RL. Experiments on five reasoning benchmarks and two model series (Qwen3-VL and LLaMA-3) confirm that **DLR** outperforms prior latent reasoning baselines with up to **20$\times$ compression**. Furthermore, the learned latent trajectories retain an interpretable semantic structure. Overall, discrete latent tokens provide a controllable and interpretable basis for efficient latent reasoning.


Winning the Symmetry Lottery: Learning Invariance from Data Augmentations with Transformers

Eduardo Santos-Escriche ⋅ Valerie Engelmayer ⋅ Ya-Wei Eileen Lin ⋅ Stefanie Jegelka

Training transformer-based architectures with data augmentation has become an increasingly popular approach for equivariant machine learning. Despite its empirical success, the interplay between the transformer architecture, invariance to different symmetries, and augmentation budgets remains underexplored. In this paper, we investigate the ability of a vanilla transformer to learn a wide range of symmetries through limited data augmentation. We identify that it performs strongly for common angle-preserving symmetries, while it struggles with non-angle-preserving ones. Within the angle-preserving family, we further find that base subgroups such as translation, rotation, and scale are learned more easily and accurately than compositional groups that combine them. Motivated by this observation, we perform a structural analysis of the trained models and identify interpretable mechanisms that induce invariance to each base angle-preserving symmetry. These findings support what we coin winning the Symmetry Lottery: the transformer architecture is aligned with learning mechanisms invariant to symmetries that happen to be prevalent in scientific domains.


WordEval: Evaluating Word-Native Operation Fidelity in Document Editing

Jinyoung Park ⋅ Yupan Huang ⋅ Shimin Zhang ⋅ Jiayu Ding ⋅ Yilin Jia ⋅ Yuzhong Zhao ⋅ Tengchao Lv ⋅ Wenshan Wu ⋅ Shaohan Huang ⋅ Nan Yang ⋅ Xiangyang Zhou ⋅ Dongdong Zhang ⋅ Li Dong ⋅ Lei Cui ⋅ Sun Mao ⋅ Xun Wang ⋅ Tao Ge ⋅ Huitian Jiao ⋅ Si-Qing Chen ⋅ Changick Kim ⋅ Furu Wei

Microsoft Word is a central environment for office automation, yet evaluating whether an agent has correctly edited a Word document remains difficult. Unlike plain text generation, Word editing changes a structured document object, where formatting, lists, fields, layout, and object properties may all be part of the requested effect. Existing evaluation protocols based on whole-document text similarity, rendered visual similarity, or LLM-as-a-judge scores cannot reliably determine whether the requested Word-native effect is realized as editable document state or whether relevant surrounding state is preserved. We introduce WordEval, a benchmark for evaluating Word-native document editing from natural-language instructions. Each case consists of an input .docx file and an editing instruction; the evaluator compares the agent’s output with a reference document using deterministic checks over Word-readable object properties. WordEval is constructed by design from planned editing units, rather than sampled from arbitrary documents: each unit specifies an input precondition, target object, operation, parameters, reference-generation procedure, and oracle specification. The current release contains 642 cases, including 546 atomic cases over 231 Word commands and 96 compositional cases instantiated from 27 composition templates. Its analysis layer covers 140 unique object-property pairs across 65 Word COM objects and 15 document subsystems. The compositional cases are linked to corresponding atomic cases for their component operations, allowing analyses of whether failures arise from individual operations or from combining multiple operations in one document. By combining controlled document construction, reference-based evaluation, and command-grounded deterministic oracles, WordEval provides a fine-grained testbed for measuring Word-native document automation.


Words That Make Language Models Perceive

Sophie L. Wang ⋅ Phillip Isola ⋅ Brian Cheung

Large language models (LLMs) trained purely on text ostensibly lack any direct perceptual experience, yet their internal representations are implicitly shaped by multimodal regularities encoded in language. We test the hypothesis that explicit sensory prompting can surface this latent structure, bringing a text‑only LLM into closer representational alignment with specialist vision and audio encoders. When a sensory prompt tells the model to 'see' or 'hear', it cues the model to resolve its next‑token predictions as if they were conditioned on latent visual or auditory evidence that is never actually supplied. Our findings reveal that lightweight prompt engineering can reliably activate modality‑appropriate representations in purely text‑trained LLMs.


WorldMemArena: Evaluating Multimodal Agent Memory Through Action–World Interaction

Chengzhi Liu ⋅ Yuzhe YANG ⋅ Sophia Xiao Pu ⋅ Yepeng Liu ⋅ Lin Long ⋅ Yichen Guo ⋅ Nuo Chen ⋅ Zhaotian Weng ⋅ Yiheng Zhong ⋅ Angxiao Zong ⋅ Elena Kochkina ⋅ Simerjot Kaur ⋅ Charese Smiley ⋅ Xiaomo Liu ⋅ James Zou ⋅ Sheng Liu ⋅ Yuheng Bu ⋅ Songyou Peng ⋅ Xin Wang

Multimodal large language models are increasingly deployed as long-horizon agents, where memory must do more than recall: it must track an evolving world, revise what has gone stale, and surface the right evidence at decision time. Existing benchmarks measure recall over static dialogue, collapse memory into a single end-of-task accuracy, and reduce visual observations to captions, leaving us unable to localize failures to writing, maintenance, retrieval, or use. The rise of agent harnesses that author their own memory sharpens this gap, since we have no principled way to compare hand-designed pipelines with self-managing alternatives. To close these gaps, we formulate multimodal agent memory as an Action--World Interaction Loop with an observable four-stage lifecycle, and instantiate it in WorldMemArena: 400 multi-session multimodal tasks spanning Lifelong Evolution (evolving personal and task states) and Agentic Execution (memory from real observations, actions, and feedback), annotated with gold memory points, updates, distractors, and evidence chains for stage-level diagnosis. This enables the first head-to-head comparison of long-context, manually designed (RAG and external memory systems), and harness-based memory agents. Results show that: (1) better memory writing and storage do not guarantee better performance; (2) multimodal memory still struggles to fully use visual evidence; (3) systems are unstable across domains and degrade on realistic agentic trajectories; and (4) harness memory is more flexible but remains costly and less reliable.


X-AVDD: Cross-Attentive Audio-Visual Dataset Distillation

Saumyaranjan Mohanty ⋅ Aravind Reddy ⋅ Konda Reddy Mopuri

*Audio-visual dataset distillation* aims to synthesize a small dataset from a large paired audio-visual dataset, while nearly preserving the training utility for neural networks. Existing approaches such as DM and AVDD face two critical limitations: 1) they are memory-intensive, making them difficult to scale to complex architectures and higher samples-per-class (SPC) settings, and 2) they often fail to preserve strong cross-modal semantic correspondence. To tackle these issues, we propose **X-AVDD**, a **cross-attention-based** distillation method that explicitly captures inter-modal relationships during synthesis. By decoupling cross-modal alignment from joint optimization, X-AVDD is substantially more memory-efficient, reducing end-to-end synthesis cost by $\sim 8\times$ GPU-hours, while producing highly aligned synthetic data. Empirically, X-AVDD significantly outperforms other audio-visual distillation methods. On VGGSound, an audio-visual dataset of short YouTube clips, X-AVDD improves top-1 test accuracy from $8.2$% → $24.0$% at SPC=$10$ and $9.8$% → $26.5$% at SPC=$20$. Furthermore, X-AVDD improves cross-architecture generalization (e.g., $25.5$% → $35.3$% for a ViT target on AVE dataset at SPC=$50$) and successfully scales to complex architectures such as AudioCLIP, achieving $62.3$% accuracy using only $5.65$% of the full dataset, against the full dataset accuracy of $72.4$\% while completing in $9\times$ lower training time compared to the full dataset.


xHC: Expanded Hyper-Connections

Xiangdong Zhang ⋅ Xiaohan Qin ⋅ Tuo Dai ⋅ Xiaoming Shi ⋅ Huaijin Wu ⋅ Yebin Yang ⋅ Zhuo Xia ⋅ Shaofeng Zhang ⋅ Yu Wang ⋅ Yu Cheng ⋅ Junchi Yan

Hyper-Connections (HC) extend the single residual stream of Transformers into $N$ parallel streams, improving language model pre-training with modest additional FLOPs. Manifold-Constrained HC (mHC) stabilizes this formulation at scale by constraining residual mixing to doubly stochastic matrices. Large gains from $N{=}1$ to $N{=}4$ suggest residual-stream expansion as a promising scaling axis, yet existing HC-family methods typically stop at $N{=}4$. Our experiments on mHC reveal why: directly scaling $N$ beyond 4 yields rapidly diminishing returns, as gains saturate while training compute grows sharply. We trace this to two bottlenecks: fixed-dimensional layer outputs cannot supply enough write-back information for more streams; meanwhile, generating the residual-mixing matrix makes the dominant cost scale cubically with $N$. To address both bottlenecks, we propose **xHC** (E**x**panded **H**yper-**C**onnections), the first HC-family method to achieve meaningful expansion beyond $N{=}4$. xHC enriches write-back signals with local contextual features along the token sequence and uses a sparse residual-stream architecture that updates only $k$ out of $N$ streams while preserving dense access to all streams. In language model pre-training, xHC with $N{=}16$ and $k{=}4$ substantially outperforms both mHC and the vanilla residual baseline. On a 69B MoE model, xHC improves the average downstream score by 5.7 points over mHC (vs.\ +1.3 for mHC over vanilla) with only 2.4\% additional training FLOPs over the vanilla baseline. Scaling law experiments show that xHC matches the loss of a vanilla baseline trained with $1.50\times$ the compute, and its continued improvement with larger $N$ makes expansion rate a practical scaling axis for HC-family models.


Z-AXIS: From Deterministic Ground to Agentic Depth for Enterprise Evaluation

Zifan Song ⋅ Mianzhi Chang ⋅ Ziyang Liao ⋅ haiyan xu ⋅ Yutong Liu ⋅ Cairong Zhao

Agentic systems powered by large language models have advanced rapidly yet current architectures treat information acquisition as a homogeneous process, issuing broad searches, accumulating heterogeneous evidence, and deferring verification to post-hoc reconciliation. This neglect of acquisition-mode structure leaves factual grounding and epistemic reliability as emergent rather than engineered properties. In this paper, we formalize the Deterministic-Agentic (D/A) Separation Principle, which partitions the task-relevant information space into programmable ($\mathcal{D}$) and emergent ($\mathcal{A}$) zones, establishing $\mathcal{D}$ as both epistemic anchor and investigative compass for $\mathcal{A}$-directed exploration. We instantiate this principle as Z-AXIS, a three-phase cascade that front-loads deterministic grounding, deploys hierarchical wave agents for macro-to-micro deep research, and synthesizes evidence through tension-aware cross-dimensional analysis. Key mechanisms include authority-calibrated evidence fusion, D-conditioned coverage tracking, and temporal awareness for longitudinal evaluation. Extensive evaluation across multiple industries demonstrates that Z-AXIS achieves superior factual accuracy and epistemic reliability compared to existing deep research systems, with architecture-determined reliability remaining stable across a wide range of backbone scales, confirming that structured governance, not model scale, governs epistemic reliability.

Classic zeroth-order optimization approaches typically optimize for a smoothed version of the original function, i.e., the expected objective under randomly perturbed model parameters. This can be interpreted as encouraging the averaged loss values in the perturbation set to be small. However, popular sharpness-aware minimization objectives typically focus on the largest loss within the neighborhood to arrive at flat minima more effectively. In this work, we connect zeroth-order optimization (and its corresponding objectives) with SAM approaches explicitly, through an exponential tilting objective that provides a smooth transition between the $\texttt{average}$- and the $\texttt{max}$-loss formulations. We explore new zeroth-order algorithms to solve a $\textit{soft}$ SAM objective parameterized by a tilting parameter $t$. We theoretically analyze the sharpness of the tilted objective under different perturbations. Practically, our approach can be used as a gradient-free and memory-efficient alternative to SAM variants, and it achieves better generalization compared to vanilla zeroth-order baselines on a wide range of models and downstream tasks, including classification, multiple choice QA, and language generation.