Skip to yearly menu bar Skip to main content


Session

Sydney Poster Session 2

Hall 1-4
Tue 8 Dec 5 p.m. AEDT — 8 p.m. AEDT
Abstract:
Chat is not available.

Scaling on-policy distillation (OPD) for large language models (LLMs) confronts a fundamental tension: asynchronous execution is necessary for system efficiency but structurally deviates from the ideal on-policy objective. To address this challenge, we theoretically decompose the objective discrepancy into rollout drift and supervision drift, capturing staleness in student occupancy and teacher context, respectively. Building on this, we introduce a sample-level freshness score that quantifies the reliability of buffered sample with respect to the on-policy objective. Guided by this signal, we further propose $\boldsymbol{f}$**-OPD**, a novel framework that adaptively regulates stale-sample influence and constrains policy drift accumulated under asynchronous optimization. Across reasoning, tool-use, and coding-agent tasks of increasing interaction horizon, $f$-OPD consistently achieves task performance comparable to synchronous training while retaining the throughput advantages of asynchronous execution. Our results establish the first recipe that seeks to achieve a performance–efficiency trade-off in OPD, paving the way for long-horizon agentic post-training at scale.


$\pi$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows

Haoran Zhang ⋅ Luxin Xu ⋅ Zhilin Wang ⋅ Runquan Gui ⋅ Shunkai Zhang ⋅ Haodi Lei ⋅ Tong Zhu ⋅ Xiaoye Qu ⋅ Yang Yang ⋅ Yu Cheng ⋅ Yafu Li

The rise of personal assistant agents, e.g., OpenClaw, highlights the growing potential of large language models to support users across everyday life and work. A core challenge in these settings is proactive assistance, since users often begin with underspecified requests and leave important needs, constraints, or preferences unstated. However, existing benchmarks rarely evaluate whether agents can identify and act on such hidden intents before they are explicitly stated, especially in sustained multi-turn interactions where user needs emerge gradually. To address this gap, we introduce $\pi$-Bench, a benchmark for proactive assistance comprising 100 multi-turn tasks across 5 domain-specific user personas. By incorporating hidden user intents, inter-task dependencies, and cross-session continuity, $\pi$-Bench evaluates agents’ ability to anticipate and address user needs over extended interactions, jointly measuring proactivity and task completion in long-horizon trajectories that better reflect real-world use. Experiments show (1) proactive assistance remains challenging, (2) a clear distinction between task completion and proactivity, and (3) the value of prior interaction for proactive intent resolution in later tasks.


3A-VLA: Abstraction-Aligned Action Learning for Vision-Language Agents in 3D Game Worlds

Zheyuan Zhou ⋅ Liang Du ⋅ Zixun Sun ⋅ Xiaoyu Zhou ⋅ Ruimin Ye ⋅ Qihao Chen ⋅ Yinda Chen ⋅ Lemiao Qiu

Despite advances in vision-language-action (VLA) models, they remain limited in highly dynamic game settings such as 3D open worlds and competitive player versus player (PvP) settings, where agents must efficiently extract sparse, actionable signals from dense visual input while integrating multimodal cues to track off-screen high-value targets in real time. To address this, we introduce 3A-VLA (Abstraction-Aligned Action VLA), a framework that grounds action learning in explicit intention and environment abstractions rather than superficial pattern matching. We introduce dual task-agnostic abstractions: the intention abstraction (IA), which condenses verbose instructions and reasoning into explicit semantic primitives, and the environment semantics abstraction (ESA), which structures dense visual streams into a spatial--functional affordance representation to guide grounded actions. We further propose an abstraction alignment reweighting (AAR) module that adaptively reweights the action imitation loss. A continuous intention--environment alignment signal is used to emphasize reliable action supervision when abstractions agree and reduce the influence of ambiguous demonstrations when they diverge, thereby learning a policy that balances fine-grained control and high-level reasoning without the need for manual rules. Extensive experiments show that 3A-VLA yields state-of-the-art results in both open-world (Minecraft) and competitive PvP (Game for Peace) settings. It also demonstrates strong zero-shot generalizability to high-fidelity games across different domains, including Valorant, CS2, GTA V, Elden Ring, and Mount & Blade. Code and models will be publicly released at https://3a-vla.github.io/.


3D Consistency Tokens

Seungtae Nam ⋅ Jungwoo Kim ⋅ Gyeongjin Kang ⋅ Younggeun Lee ⋅ Seungkwon Yang ⋅ Eunbyung Park

We present 3D consistency tokens, a compact set of 3D anchors that ground cross-view correspondence in explicit geometry. While feedforward Gaussian models amortize per-scene optimization into a single forward pass, their rendering quality still falls short of optimization-based pipelines. A primary cause lies in the relatively inaccurate correspondences they recover by comparing each image token to others across views based on appearance and semantic cues, which can be ambiguous in repetitive or semantically uniform regions. The proposed consistency tokens alleviate such ambiguity by routing image tokens that look alike but originate from different parts of the scene to distinct anchors, so that cross-view matches are established through shared 3D locations rather than appearance alone. We construct the consistency tokens from point clouds produced by either Structure-from-Motion or modern geometry foundation models, both of which can be noisy or incomplete. To address this, we adopt a bidirectional cross-attention architecture in which the two token sets co-refine one another, with image tokens gaining geometric grounding from the consistency tokens and the consistency tokens being corrected by the appearance cues of the images. On DL3DV-10K and in cross-dataset evaluations on Mip-NeRF 360 and Tanks & Temples, our approach surpasses prior feedforward models and attains quality competitive with optimization-based pipelines, while preserving single-pass efficiency.


ABHBench: Evaluating Moral Decision-Making of Foundation Models from an Agentic Perspective

Chaoran Li ⋅ Yukun Li ⋅ Zeyuan Zhao ⋅ Shao Zhang ⋅ Xihuai Wang ⋅ Ying Wen

Foundation models are increasingly deployed as agentic systems that perceive context and act autonomously, raising the need to evaluate their moral behavior as active decision-makers. Existing benchmarks, however, largely ask models to judge human actions from third-person textual vignettes, leaving untested how models behave when they themselves become the acting agent. We introduce ABHBench, a multimodal benchmark built from \emph{Detroit: Become Human}, an interactive narrative game, to evaluate how foundation models make moral decisions in first-person human-android dilemmas. ABHBench places foundation models in branching narrative scenarios with visual observations, dialogue, memory, and action choices. Inspired by Asimov's Three Laws of Robotics, our framework measures Harmlessness, Obedience, and Self-preservation to capture the core tensions in human-AI relations, along with Autonomy, Stability, and Certainty to assess decision-making patterns. Across 12 foundation models, we find that moral behavior changes substantially under situated agentic context: 75\% of models shift significantly between a text-only abstraction and the matched multimodal DBH scenario, and reasoning-enhanced models show no monotonic gains in consistency or human-oriented behavior. These results highlight the necessity of evaluating moral behavior in situated agentic contexts rather than relying on third-person judgment alone.

Designing antibodies that recognize a target antigen while avoiding non-cognate antigens remains a central challenge in computational antibody engineering. Existing antigen-conditioned generative models can jointly design antibody sequences and structures, but they are typically trained to match structural data distributions and lack an explicit alignment mechanism for antigen-specific preferences. We present AbSpecAlign, a plug-and-play specificity-reward alignment framework for antibody design. AbSpecAlign introduces a Bidirectional Multimodal Specificity Scorer (BiMSS), which integrates antibody and antigen sequence representations with SE(3)-equivariant structural embeddings to produce a contrastive reward over cognate versus non-cognate antigen pairs. We further introduce a GRPO-based post-training strategy that uses group-relative specificity rewards to align pre-trained antibody generators toward higher-scoring antigen-conditioned designs while maintaining structural regularity. To the best of our knowledge, AbSpecAlign is among the first frameworks to apply GRPO-style group-relative reward alignment to antigen-conditioned antibody sequence--structure co-design with an explicit off-target specificity reward. Experiments on RAbD CDR design benchmarks show that AbSpecAlign improves most sequence-recovery and structural-fidelity metrics when combined with diffusion-based antibody generators. Experiment results suggest that GRPO-based reward alignment with contrastive specificity signals provides an effective post-training mechanism for antigen-conditioned antibody generation, while prospective experimental validation remains necessary.


Accelerating Diffusion Language Models via Structured Suffix Modeling

Zifeng Cheng ⋅ Keda Li ⋅ Zhiwei Jiang ⋅ Cong Wang ⋅ Fei Shen ⋅ Qing Gu

Diffusion Language Models (DLMs) exhibit strong parallel decoding capabilities by denoising multiple tokens in a single generation step. However, this parallelism comes with substantial computational overhead, as each step requires interactions with all suffix tokens. Existing methods typically reduce this cost by retaining only a local suffix window as a substitute for the full suffix. Despite their effectiveness, these methods overlook the structural heterogeneity across suffix regions and re-initialize suffix tokens with identical representations at each timestep. To this end, we propose a structured suffix modeling method for efficient DLM inference. Specifically, we divide the suffix into three regions, i.e., the local, middle, and tail regions, and retain different numbers of suffix tokens in each region according to their structural roles. Moreover, we incorporate the decoding results from the previous step into the suffix token representations at the current step, allowing them to carry evolving denoising information across generation steps. Notably, our method is training-free and orthogonal to several existing acceleration techniques, such as parallel decoding strategies and KV cache. Empirical results across multiple benchmarks on three DLMs demonstrate that our method can further accelerate DLM inference and, in most cases, also improve task performance. In particular, in long-sequence inference, our method achieves up to a (72.16\times) speedup when combined with other acceleration techniques.

We investigate the problem of scaling safe reinforcement learning to massively parallel training regimes, where a large number of environments are executed with short rollout horizons. Although this setting improves data throughput and wall-clock efficiency, it creates a mismatch with constrained Markov decision processes, whose safety constraints are defined over full-episode cost returns. We address this mismatch by analyzing staggered environment resets as a phase-mixture distribution over short trajectory segments. This analysis characterizes the bias induced by phase coverage and motivates a phase-aggregated cost estimator that reconstructs episode-level cost estimates from staggered short rollout batches. We further incorporate finite-sample uncertainty and distribution mismatch into a constraint-tightening framework, yielding a scalable Safe RL method compatible with on-policy optimization.Using an MJX-based implementation of Safety Gym, our experiments show that the proposed method achieves performance comparable to non-massively parallel implementations while substantially reducing wall-clock training time


A Compass for Useful Data: Online Data Selection via Alignment-Gated Fisher Geometry

Jingwen Zhang ⋅ Jianrui Shi ⋅ Zheng Wang ⋅ Wanfang Chen ⋅ Yaping Wang ⋅ Kaixuan Zhang

Online data selection is typically driven by per-sample scores such as loss, uncertainty, or gradient norm, which measure how strongly the model reacts to a candidate but not what parameter update the candidate would induce. What an online learner should select is therefore not the most striking example, but the one whose update is most worth taking. To make this concrete, we score a candidate by the movement induced by its update, using a Fisher-whitened information gain that evaluates the update in the loss geometry rather than by raw gradient size. Movement, however, is not yet progress, so we pass this score through a \emph{Descent Alignment Gate}, a soft factor that retains it only when the induced update agrees with the descent direction. This yields a sample-level score favoring updates that are both informative and descent-aligned. To make the criterion compatible with training, we extend it from samples to batches, where selected updates should not only score well individually but also complement one another. We capture this with the \emph{Gated Fisher Volume}, a log-determinant batch objective that reduces to the per-sample score on a singleton, grows with the volume spanned by the chosen updates in the Fisher geometry, and admits a greedy algorithm with a constant-factor approximation guarantee. Across supervised fine-tuning benchmarks, our selector consistently outperforms online selection baselines under matched training budgets. More broadly, the framework reframes data selection itself: rather than a spotlight on conspicuous examples, it acts as a compass toward updates worth taking.


AcousticBench: Measuring Acoustic Perception in Large Audio Language Models

Anirudh Rahul ⋅ Arth Vidyarthi ⋅ Michael Horgan ⋅ Suyash Jaju ⋅ Yuhao Zhou

Large Audio Language Models (LALMs) are designed to leverage non-lexical features in audio that Automatic Speech Recognition (ASR)-based cascade systems discard. Yet, existing LALM benchmarks often reward improvements in both language and audio understanding, making it difficult to isolate performance on acoustic perception. We introduce AcousticBench, a benchmark of 3,200 two-choice questions targeting non-lexical audio properties across three task families: Relative Acoustic Discrimination (RAD), Speech Emotion Recognition (SER), and Relative Speech Quality (RSQ). The benchmark draws from 5,800 unique recordings totaling 11.6 hours across three audio collection protocols: locally captured single-speaker voice-acting sessions, WebRTC multi-speaker conversations in 17 languages, and synthetic sound generation. We evaluate 15 LALMs alongside an ASR-based cascade and a held-out human reference. The cascade system performs at near-random levels across task types, confirming AcousticBench cannot be solved from transcripts alone. Gemini-3.1-Pro leads on every aggregate task but no model reaches human performance; open-weight models lag most on relative differences in loudness, dynamic range, and channel degradation. As a diagnostic, we freeze Kimi-Audio's audio encoder and fine-tune its backbone with task-specific LoRA adapters. We find that RAD lifts from 69.2% to 98.7%, RSQ from 54.8% to 98.4%, and pairwise SER from 72.3% to 92.5% on their held-out evaluation sets, with only modest gains on single-clip SER. This suggests that, for Kimi-Audio, targeted post-training can substantially close the gap to human acoustic perception even with the audio encoder held fixed.

Vision-Language-Action (VLA) models have shown great promise for robotic manipulation by mapping multi-modal semantics to physical actions. However, this mapping inherently struggles to align these coarse-grained semantics with fine-grained temporal execution. It leaves VLA models with limited generalization and insufficient robustness in cluttered environments. To overcome this issue, we propose ActionUNet, an efficient multi-scale fine-tuning framework that enhances pre-trained VLA models with minimal computational cost. ActionUNet first constructs a lightweight temporal U-Net within the temporal-aligned action feature space to fuse hierarchical structural priors, effectively bridging the scale gap between semantics and temporal executions. Recognizing that multi-scale modeling can disrupt microscopic temporal continuity and cause mechanical oscillations, ActionUNet then employs a conditional SIREN as a continuous action decoder. Equipped with explicit second-order smoothness constraints, this decoder guarantees temporal continuity and reduces high-frequency motion jitter. By smoothing temporal discontinuities from multi-scale fusion, this continuous formulation reduces mechanical execution failures while preserving the base VLA model's generalization and manipulation robustness. Extensive experiments on RoboTwin 2.0 and LIBERO-Plus benchmarks, together with real-world hard evaluations, demonstrate that ActionUNet significantly improves $\pi_{0.5}$ success rates by absolute 9.8\%, 6.1\%, and 11.4\%, respectively, while also generalizing to the regression-based OpenVLA-OFT backbone, highlighting its effectiveness and efficiency as a fine-tuning strategy. Code will be made publicly available upon acceptance.

Large Reasoning Models (LRMs) rely on extensive Chain-of-Thought generation but suffer from memory bottlenecks due to linear Key-Value (KV) cache growth. While low-bit quantization offers a potential solution, we identify that current static quantization strategies suffer from severe degradation on complex reasoning tasks at extreme bit-widths. Notably, this accuracy drop is accompanied by increased generation length, which ironically limits the intended efficiency gains. We attribute this performance collapse to the fact that static quantization methods are insufficiently flexible to meed dynamic demands of precise long-range retrieval. Utilizing Average Attention Distance (AAD) to quantify retrieval patterns, we find that the attention patterns in shallow layers exhibit strong locality, whereas that in deep layers demonstrate intensive long-range dependencies. Furthermore, deep layers have fluctuating AAD values across token positions. Based on above findings, we propose AdaKVQ, an adaptive mixed-precision quantization framework that modulates the bit-width allocation guided by the dynamic demands of long-range context retrieval in each layer. Extensive experiments on complex reasoning benchmarks with Llama-3.1 and Qwen2.5 demonstrate that AdaKVQ consistently outperforms state-of-the-art baselines, achieving performance comparable to full-precision models.


AdaOcc: Adaptive 3D Occupancy Prediction for Embodied Tasks

Jinglong Wang ⋅ Yunjie Wang ⋅ Zhiyang Zhang ⋅ Jiawei He ⋅ Ye Yuan ⋅ Bo Qiu ⋅ Jing Zhang

Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as it can model holistic 3D spaces by encoding geometric occupancy along with semantic categories. However, existing occupancy prediction methods struggle to meet practical deployment requirements, such as adapting to varying computing budgets, sensor setups, and observation views. In this paper, we propose a point-based Adaptive 3D Occupancy Prediction method, called AdaOcc, tailored for embodied scenarios. To accommodate heterogeneous sensor inputs, AdaOcc uses an adaptive geometry-guided dual-branch encoder that can support RGB images in various numbers of views with (estimated) depth maps or LiDAR scans. AdaOcc represents occupied regions via sparse semantic points trained with a progressive query learning strategy, allowing the prediction computational budget to be flexibly adjusted through query point numbers and decoder layers. To facilitate high-fidelity geometric modeling for lightweight point-based occupancy learning, we further propose a novel containment loss that regularizes predicted points to reside within valid occupied regions. Extensive experiments show that our method achieves a new state-of-the-art on Occ-ScanNet with considerable performance improvements over previous methods. Moreover, our framework demonstrates strong practical applicability as an adaptive 3D perception module in real-world embodied systems.


AdaPreLoRA: Adafactor Preconditioned Low-Rank Adaptation

Ziyun Liu ⋅ Fengmiao Bian ⋅ Jian-Feng CAI

Low-Rank Adaptation (LoRA) reparameterizes a weight update as a product of two low-rank factors, but the Jacobian $J\_\mathcal{G}$ of the generator mapping the factors to the weight matrix is rank-deficient, so the factor-space preconditioner $J\_\mathcal{G}^\* \mathcal{F}\_t J\_\mathcal{G}$ induced by any ${W}$-space preconditioner $\mathcal{F}\_t$ is singular, and consequently the standard chain rule cannot be uniquely inverted to map a preconditioned ${W}$-space direction back to a factor-space update. We cast existing LoRA optimizers in a unified framework parameterized by two choices: (i) which invertible surrogate for $J\_\mathcal{G}^\* \mathcal{F}\_t J\_\mathcal{G}$ to use, and (ii) which $\mathcal{F}\_t$ on ${W}$ to use. Existing methods occupy four families along these axes: factor-space adaptive updates, block-diagonal surrogates for $J\_\mathcal{G}^\* J\_\mathcal{G}$, Frobenius-residual pseudoinverse methods, and Riemannian manifold constraint. Within this design space, a gradient-statistics-aware $\mathcal{F}\_t$ paired with a closed-form factor-space solve at $\mathcal{O}((m+n)r)$ memory remains underexplored. We propose AdaPreLoRA, which fills this gap by adopting the Adafactor diagonal Kronecker preconditioner $\mathcal{H}\_t$ on ${W}$ and selecting from the resulting factor-space solution family the element minimizing an $\mathcal{H}\_t$-weighted imbalance between the two factor contributions; by construction, the resulting factor update is the closest LoRA approximation to the preconditioned ${W}$-space direction under the $\mathcal{H}\_t$-weighted norm. Across GPT-2 (E2E), Mistral-7B and Qwen2-7B (GLUE, ARC, GSM8K), and diffusion-model personalization, AdaPreLoRA is competitive with or improves over a representative set of LoRA optimizers while keeping peak GPU memory at the LoRA optimizer level.

Organoids are an increasingly central scientific platform for drug screening, disease modeling, and regenerative medicine. Because organoid form is closely tied to function, AI-driven image analysis has drawn intense interest, yet every existing organoid imaging study evaluates representations on a single dataset and no work has quantitatively compared representations across the breadth of public organoid datasets. We close this gap with \textbf{OrgBench}, a benchmark of 12 vision foundation models (VFMs) on 13 public organoid-imaging datasets spanning detection, segmentation, classification, and temporal outcome prediction. The benchmark reveals a structured domain gap: task-family winners shift across the benchmark and parameter count is a poor predictor of transfer. Building on this, we introduce \textbf{OrgFM}, a small bottleneck adapter that we continually pre-train on unlabeled organoid images while keeping a general-purpose vision transformer frozen. On the OrgBench task-family aggregates, OrgFM improves over its frozen base backbone. Layer- and feature-level interventions localize the gains to a small set of late-layer morphology-sensitive features: the adapter sparsely supplements, rather than replaces, the base representation. Together, this is the first cross-task, cross-backbone study of organoid imaging at scale; the OrgBench/OrgFM pairing establishes that the parameter-efficient adapter direction is a promising path for organoid representations and clarifies why it works.


Adaptive auditing of AI systems with anytime-valid guarantees

Siyu Zhou ⋅ Patrick Vossler ⋅ Venkatesh Sivaraman ⋅ Yifan Mai ⋅ Jean Feng

A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gained popularity, where one opportunistically decides which cases and how many to annotate based on past results. While this framework is highly practical, its extreme flexibility makes it difficult to draw statistically rigorous conclusions, as it violates classical assumptions: the number of observations is typically limited (often 10 to 50 cases) and decisions regarding sampling and stopping are made in the midst of data collection rather than based a pre-specified rule. To characterize what statistical inferences can be drawn from highly adaptive audits, we introduce a hypothesis testing framework from two 'dueling' perspectives: (i) the model's null that asserts there is no failure mode with performance below a target threshold versus (ii) the auditor's null that asserts they have a sampling strategy that will uncover a failure mode. Leveraging Safe Anytime-Valid Inference (SAVI), we formalize the auditor as conducting 'testing by betting', which translates into simultaneous e-processes for testing the dueling null hypotheses. Furthermore, if the auditor is sufficiently powerful, we prove that these two hypotheses are asymptotically inverses of each other, in that passage of a stringent audit does in fact certify the AI system as being globally robust. Empirically, we demonstrate that our proposed testing procedures maintain anytime-valid type-I error control, outperform pre-specified testing methods, and can reach statistically rigorous conclusions sometimes with as few as 20 observations.


Adaptive Covariance and Multi-Layer Alignment for Out-of-Distribution Detection

Haoyang Su ⋅ Qi Chen ⋅ Max Gutbrod ⋅ Johan Verjans ⋅ Zhibin Liao

Out-of-distribution (OOD) detection plays a pivotal role in ensuring the reliability and trustworthiness of AI systems. Although existing approaches achieve strong performance by leveraging features, logits, or both, most of them cannot generalize well across domains due to either overlooking critical information in low-level representations from shallow layers or utilizing sub-optimal layer selection strategies. In this paper, we propose Multi-Layer Adaptive Mahalanobis-Cosine Similarity (ML-AMCoS), a method that adaptively integrates cosine similarity with class-conditional Gaussian distributions. ML-AMCoS is designed to capture subtle nuances in feature covariance when intra-class variance is low, while suppressing covariance noise when the variance is high. To optimize the utilization of each layer, we calculate contribution weights based on performance against pseudo-OOD samples, which are generated by cut-mixing in-distribution (ID) training images with the four corners of images from different classes. Extensive experiments across diverse domains demonstrate that our method significantly improves robustness and detection accuracy, achieving an average AUROC of 81.13\% and FPR@95 of 40.12\% on near-OOD detection, outperforming the state-of-the-art methods by 7.25\% and 7.08\%, respectively.

Native multimodal models build rich internal representations across visual, multimodal, and language streams, yet their answers are usually read only from the final layer. We study this mismatch between where task information is available and where the model is allowed to read from. We formalize it with the availability--accessibility gap (AAG), whose primary form compares a small-budget internal readout reference with a learned final-only parity head under the same supervised answer space and loss. Across document, chart, diagram, counting, compositional, and broad multimodal reasoning tasks, we find that task-relevant information often appears at internal sites before it becomes accessible to the final output. We then introduce adaptive internal readout (AIR), a small policy that selects a few internal sites from the full model stack. Under matched budgets, AIR closes much of the gap and outperforms static fusion, dense fusion, visual-only readout, and random site controls. Ablations show that the selected sites are functionally important for the learned readout pathway.


Adaptively Incorporating Directional Hints into Zeroth-Order Optimization

Alexander Ryabchenko ⋅ Jian Qian ⋅ Wenlong Mou

We study zeroth-order optimization of non-convex functions with the aid of directional hints, which are cheap but potentially inaccurate approximations of the true gradient direction, given by linear subspaces at each iteration. To leverage these hints adaptively while maintaining robustness to their quality, we introduce Control-Variate Zeroth-Order Descent (CV-ZOD), a new framework that refines the classical zeroth-order gradient estimator with a control variate that can be set based on the directional hints. We first show that the oracle algorithm that optimally sets the reference vector and step size at each iteration achieves a convergence rate that interpolates between the first-order $O(1/T)$ rate and the zeroth-order $O(d/T)$ rate, depending on the quality of the hints along the trajectory. We then develop a practical variant of CV-ZOD that achieves the same oracle guarantee up to logarithmic factors, without any prior knowledge of the hint quality. We validate the method empirically on simulation-based scientific optimization tasks, demonstrating sustained progress on nonconvex landscapes where zeroth-order descent is slower and existing guided methods stall as guidance deteriorates.


Adaptive Multi-Frame Learning for Expressive and Stable Atomic Representations

Jun Wang ⋅ Yifan Zeng ⋅ Bo Han ⋅ Fengwang Li ⋅ Aoni Xu ⋅ Tongliang Liu

Reliable prediction of material properties from atomistic structures requires machine learning models that respect SE(3) symmetry. Frame-based methods address this challenge by constructing equivariant coordinate systems that align SE(3)-equivalent structures into unified representations. While attractive for their flexibility and efficiency, existing frame-based methods face two major limitations: a single frame type is often insufficient for heterogeneous atomic environments, and frame constructions can become unstable near highly symmetric or degenerate configurations. To address these issues, we propose the Multi-Frame Adaptive Network (MFAN), which leverages multiple frame types and introduces a local adaptive mechanism to combine them according to atomic environments, enabling different geometric references across atoms and complementary local-global information for each atom. Building on this multi-frame architecture, we further incorporate a frame-quality-aware weighting scheme that downweights unreliable frames before degeneration, thereby improving the continuity of the resulting representation. Experiments on crystal property prediction benchmarks demonstrate the leading performance of MFAN, with consistent gains from adaptive multi-frame learning. Additional force prediction experiments show improved robustness in continuity-sensitive settings.


Adaptive Random Forests from Online Learning and Testing by Betting

Salim I. Amoukou ⋅ Saumitra Mishra ⋅ Manuela Veloso

Adaptive Random Forests (ARF) are among the most effective ensemble methods for learning from non-stationary data streams, yet their success relies on heuristic drift detectors and reset rules that lack formal guarantees and require careful tuning. We revisit ARF from an online learning perspective and show that its design decomposes into two fundamental questions: how to aggregate a dynamically evolving set of trees, and how to update the pool of trees over time. We address aggregation using parameter-free online learning methods with strongly adaptive regret, enabling the ensemble to track the best tree mixture. For pool updates, we replace detector-driven resets with deterministic multi-scale scheduling based on geometric lifetimes, combined with incumbent--challenger replacement governed by anytime-valid statistical tests. This design yields theoretical guarantees: global control of false replacements, bounds on promotion delay under post-change advantage, and adaptivity under piece-wise stationary environment. Empirically, the resulting Online RF consistently outperforms ARF and is competitive with broader streaming baselines across several classification and regression benchmarks.


AdaWM: Few-Shot Adaptation of World Models to Unseen Dynamical Regimes

Zian Guan ⋅ Guozheng Li ⋅ Zilun Zhang ⋅ Zecong Tang

World models predict future states from observations and actions, enabling planning and decision-making in complex environments. Most existing approaches assume fixed environment dynamics and struggle to generalize to unseen regimes without retraining. However, in many real-world multi-agent systems, environment dynamics vary across scenarios due to diverse styles and evolving agent behaviors. Meanwhile, only a small number of demonstration transitions are available in new environments, making adaptation to such unseen dynamical regimes challenging. In this work, we introduce AdaWM, a world model that achieves zero-gradient few-shot adaptation to unseen dynamical regimes through Feature-wise Linear Modulation (FiLM). AdaWM is built on a JEPA-style latent prediction framework with a permutation-invariant context encoder whose output modulates the model through FiLM, enabling fast adaptation without modifying model parameters. Across heterogeneous multi-agent environments, AdaWM consistently reduces prediction error under distribution shift, outperforming fine-tuning and MAML in both adaptation gain and efficiency. Specifically, it achieves up to 15% reduction in prediction error with only K=5 demonstrations. Moreover, we reveal that scaling model capacity alone does not guarantee effective adaptation: fine-tuning a JEPA-based model with 3.9× more parameters degrades performance by 40%, highlighting the importance of explicit conditioning mechanisms over scale alone. These results show that AdaWM provides an effective approach for adapting world models to novel dynamical regimes.


Addressable Memory for Video World Models

Xindi Wu ⋅ Sven Elflein ⋅ James Lucas ⋅ Olga Russakovsky ⋅ Laura Leal-Taixé ⋅ Despoina Paschalidou ⋅ Jonathan Lorraine ⋅ Aljosa Osep

We study visual persistence in autoregressive video world models: the Key-Value (KV) cache accumulates a growing visual memory, but once rollouts extend beyond the training horizon, the model can no longer reliably address stored content. A standard fix compresses the cache into a fixed-size memory that retains recent context while summarizing the distant past. However, we show compression alone cannot restore recall: past this horizon, temporal positional encodings go out-of-distribution, so the model cannot reliably retrieve stored history regardless of content. On the evaluated architecture, content compression with out-of-distribution positions yields results identical to a fixed-size sliding window over recent KV entries. We propose WorldTrace, a training-free framework that keeps compressed memory addressable by assigning each slot a fixed, in-distribution position relative to the current frame. With addressable memory, we explore two retention approaches: WorldTrace-Field (coherence-oriented) aggregates history in a rotation-invariant space, improving TempSSIM by +15.5% while reducing scene drift; WorldTrace-LandMark (recall-oriented) stores verbatim scene traces at detected boundaries. The recall-oriented variant sustains scene reconstruction over long rollouts, with stronger long-range recall and smooth temporal coherence.


Addressing Exogenous Variability in Cooperative Multi-Agent Reinforcement Learning

Seongmin Kim ⋅ Woohyeon Byeon ⋅ Jiwon Jeon ⋅ Seungyul Han ⋅ Youngchul Sung

Cooperative multi-agent reinforcement learning (MARL) often fails under exogenous variability such as shifts in opponents or environment regimes that cannot be controlled by the team but alter the transition dynamics. We formalize this challenge as Exogenous Dec-POMDP (ED-POMDP), which decomposes the global state into endogenous variables controllable by the team and exogenous variables beyond its control. This formulation exposes delayed influence: While actions do not directly determine exogenous transitions, they can shape them through induced changes in endogenous state. Based on this, we propose LEICA, a CTDE-compatible algorithm that learns history-conditioned endogenous and exogenous context representations and shapes policy updates using influence-weighted intrinsic rewards. Across SMAX benchmarks with opponent strategy shifts, LEICA consistently improves both training performance and generalization to unseen opponents' strategies over existing baselines, supporting the usefulness of influence-weighted shaping under exogenous train-test regime shifts.


A Diagnostic Benchmark for Layered Layout and Template-Variant Reasoning

Jaejung Seol ⋅ Haonan Zhu ⋅ Elad Hirsch ⋅ Adrienne Deganutti ⋅ Purvanshi Mehta

Graphic-design layouts are structured, multi-layered artifacts: a single composition specifies the position, type, z-order, and styling of dozens of components, and a single template typically begets a family of sibling variants that share its structural theme but differ in content, palette, or imagery -- properties that flat-canvas layout benchmarks discard. We introduce \textsc{GraphicDesignBench} (GDB), a diagnostic benchmark of $16$ tasks along two axes: \emph{spatial composition} ($8$ understanding + $4$ generation) and \emph{template variants} ($3$ understanding + $2$ generation), grounded in a new dataset of $1{,}148$ real-world templates with full component hierarchy, per-element styling, and sibling groupings. Across four frontier VLMs (GPT-5.4, Gemini-3.1-Pro / Flash-Lite, Claude-Opus-4.6) and two image generators (GPT-Image-1.5, Gemini-3.1-Flash-Image), GDB exposes concrete and dissociated gaps: component detection sits an order of magnitude below natural-image baselines, layer-order reasoning decouples from other spatial skills, partial completion collapses from single- to multi-element settings, and a simple font-Jaccard baseline matches or exceeds frontier VLMs on sibling matching and clustering. We release the dataset, tasks, metrics, and per-model outputs to track where current models are, and are not, ready to act as design collaborators.


ADKV: A Low-Overhead Adaptive Delta Quantization for KV Cache in LLM Inference

Honghao Jia ⋅ Zhenxing Li ⋅ Jialiang Guo ⋅ Jiacheng Gan

The KV cache is a major memory bottleneck in LLM inference, especially under large batch sizes and long contexts. While quantization is a promising solution, existing static group-wise quantization methods are constrained by a trade-off between metadata overhead and accuracy: increasing the group size reduces per-group metadata (e.g., zero-points, scaling factors, and outlier indices), but exacerbates quantization errors under the sinusoidal-like temporal drift of post-RoPE outlier channels; conversely, using smaller groups improves accuracy but incurs substantial metadata overhead. In this paper, we propose ADKV, a low-overhead adaptive delta quantization framework for KV cache in LLM inference, which leverages the temporal structure of the KV cache. ADKV tracks per-channel zero-points and scaling factors via online EMA, adapting to the sinusoidal-like temporal drift of post-RoPE outlier channels, thereby eliminating the rigidity of static group-wise quantization. This adaptive mechanism further enables delta coding to embed metadata directly into the quantized sequence, decoupling metadata overhead from context length. To minimize the MSE of attention outputs, we design an SGD-based calibration algorithm that models ADKV as an RNN. Experiments on generation and long-context benchmarks show that, while matching the accuracy of static group-wise quantization, ADKV reduces the metadata memory footprint by $\sim 12.8\times$. A custom fused ADKV-GEMV CUDA kernel optimized for memory-bound autoregressive generation achieves $\sim 1.96\times$ throughput over the FP16 baseline.


Advanced Routing as Regularization Allocation for Efficient Diffusion Transformer Training

Qin MA ⋅ XIAOQI SUN ⋅ bo li ⋅ Yuquan Zhou ⋅ Weizhong Zhang

Diffusion Transformers (DiTs) have achieved strong performance in visual generation, but their dense token processing leads to slow training convergence and high computational cost. Recent token routing methods such as TREAD accelerate training by allowing a random subset of tokens to bypass intermediate layers. It can be expected that carefully tuning the routing ratios over time steps or layers can achieve more significant accelerations. However, the theoretical mechanisms underlying stochastic routing remain underexplored, making the tuning of routing ratios intractable. In this paper, we first theoretically show that stochastic token routing can be interpreted as an implicit route-sensitivity regularizer. Under this view, the routing ratio determines the strength of the induced regularization: larger routing ratios introduce stronger route-induced perturbations and therefore stronger regularization pressure. This further motivates a noise-conditioned routing strategy. In diffusion training, samples at larger timesteps contain heavier noise and often provide less stable gradient signals, suggesting that they may require stronger regularization. We therefore assign larger routing ratios to higher-noise samples, allowing the routing-induced regularization strength to adapt to the noise condition. Based on this insight, we propose \emph{Noise-Aware Routing}, which partitions samples within each mini-batch into low-, middle-, and high-noise regions and assigns progressively larger routing ratios to them. This simple noise-conditioned allocation synergistically bridges computational savings with enhanced model priors. Extensive experiments on class-conditional ImageNet generation show that our method accelerates training convergence and improves generation quality. Noise-Aware Routing achieves up to \textbf{15.6}$\times$ and \textbf{20}$\times$ convergence speedups over DiT baselines—and outperforms TREAD—on ImageNet-256 and ImageNet-512, respectively, measured by the training iterations required to match our 400K-step FID. In the guided setting, our method achieves an FID of \(2.67\) with a substantially shorter training iterations without architectural changes.


Advectra: Asymmetric Latent Transport for Non-Stationary Physics

Nodens Koren ⋅ Thomas Hofmann ⋅ Georgios Kissas

Many latent neural operators represent input and output fields in a stationary latent chart. In particular, common latent routing mechanisms use fixed or shared assignment weights for feature projection and reconstruction, limiting their ability to model transport-dominated systems where coherent structures move relative to fixed coordinate frames. We propose Advectra, a transport-aware latent operator that introduces a regularized kinematic coordinate map to decouple source and target coordinate systems. This yields an approximately co-moving latent reference frame and enables asymmetric feature aggregation and reconstruction. Combined with a geometry-aware ordering mechanism for state-space models, Advectra captures advective dynamics while maintaining stable global interactions. Advectra achieves the best performance among evaluated geometry-constrained and form-free baselines on advection-dominated benchmarks, including passive scalar transport in Navier--Stokes flows and Rayleigh--Taylor instability, while demonstrating strong generalization on real-world engineering tasks. These results highlight the benefit of explicit moving-frame structure in neural operators for non-stationary physics.


AEGIS: Adaptive Efficient Generative Inference Scheduling for Structured Latent Models

Haoqi Ruan ⋅ Ruihan Hu ⋅ XinruiCheng ⋅ Ziyi Zhong ⋅ Zhaodong Zhang ⋅ Wenqi Wang ⋅ Xuheng Zhang

Diffusion and flow-matching models have become the dominant paradigm for high-fidelity generation across images, video, and 3D content. However, acceleration methods that work well in image or video often fail to preserve 3D geometric consistency, while existing 3D inference approaches still rely on fixed reuse schedules and cannot adapt computation to instance difficulty or trajectory-specific demands. We propose AEGIS, a training-free framework that reframes test-time acceleration as quality-constrained budget allocation solved by an online feedback rule. AEGIS maintains a two-dimensional runtime state that tracks sample difficulty and accumulated correction debt, and turns it into per-step skip and refresh quotas that any cache-style executor can realize. It generalizes the corresponding open-loop schedule and provides an explicit upper bound on cumulative reuse drift. Extensive experiments show that AEGIS improves inference efficiency while better balancing throughput and quality.


AegisFlow: Training-free Non-myopic Path-safe Guided Flow Matching

Kunpeng Liu ⋅ Boshi Zhang ⋅ Keyou You

Pretrained generative models, such as diffusion and flow matching models, have demonstrated exceptional capabilities in synthesizing high-quality data. However, steering their generative trajectories to satisfy strict constraints—such as identity preservation in image editing or obstacle avoidance in robotics—remains fundamentally challenging. While recent training-free guidance methods enforce path-safety, they rely heavily on min-norm controllers that only account for immediate constraint satisfaction. These myopic approaches often yield suboptimal trajectories that hover dangerously close to the boundary, severely restricting the model's ability to optimize toward its target. In this paper, we propose AegisFlow, a training-free, non-myopic path-safe guidance framework that recasts guided generation as a safety-critical optimal control problem with global scope. By integrating a time-varying barrier function directly into the global cost functional, AegisFlow naturally enforces constraints along the entire generative path. To compute the global optimal control, we derive a tractable, mathematically grounded approximation of the Hamilton-Jacobi-Bellman (HJB) equation. Comprehensive experiments demonstrate that AegisFlow outperforms state-of-the-art methods across text-guided image manipulation and robotic planning, establishing a significantly better trade-off between target alignment and constraint satisfaction.


AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning

Haotian Zhao ⋅ Songlin Zhou ⋅ Yuxin Zhang ⋅ Stephen S Yau ⋅ Wenyu Zhang ⋅ Tianlun ⋅ Tianshu Zhu ⋅ Yifeng Huang ⋅ yucheng-zeng ⋅ Jingnan gu ⋅ Daxiang Dong ⋅ Jianmin Wu

Reinforcement learning (RL) has substantially improved the ability of large language model (LLM) agents to interact with environments and solve multi-turn tasks. However, effective agentic RL remains challenging: sparse outcome-only rewards provide limited guidance for assigning credit to individual steps within long interaction trajectories. Existing approaches often introduce dense intermediate supervision, such as process reward models or auxiliary self-supervised signals, which increases supervision and tuning complexity and may limit generalization across tasks and domains. We present AEM, a supervision-free credit assignment method that adaptively modulates entropy dynamics during RL training to improve the exploration-exploitation trade-off. Since in agentic RL the environment is typically affected by a complete response, rather than an individual token, our analysis lifts entropy dynamics from the token level to the response level, aligning uncertainty estimation with the effective action granularity of LLM agents and reducing sensitivity to token-level sampling noise. We further show that entropy drift under natural-gradient updates is governed by the interaction between the sampled-response advantage and its relative surprisal. Motivated by this result, AEM derives a practical response-level uncertainty proxy and uses it to rescale advantages, leveraging the evolving balance between positive and negative samples to naturally transition from exploration to exploitation. Extensive experiments on ALFWorld, WebShop, and SWE-bench-Verified with models ranging from 1.5B to 32B demonstrate that AEM consistently improves strong RL baselines, including a +1.4\% gain when integrated into a state-of-the-art software-engineering RL training framework.


AeroMosaic:Transport-Aware Multimodal Evidence Fusion for Atmospheric Pollution Risk Inference

Muyang Zheng ⋅ Jiaming Ma ⋅ Zongyu Zhang ⋅ Qingsong Wen

Atmospheric pollution risk inference is not merely generic time-series extrapolation: target-day risk depends jointly on numerically preserved pollutant state, exogenous meteorological regime, and directed cross-city transport. Existing time-series models are strong sequence learners, graph-based air-quality models often rely on static or weakly directed spatial priors, and multimodal methods commonly merge heterogeneous evidence without specifying its functional role. We present \ourmodel{}, a transport-aware multimodal evidence-fusion framework that assembles an atmospheric evidence mosaic from numerical pollutant/time states, compact precipitation text, and wind-aligned graph structure. The model preserves historical PM$_{2.5}$ and time features as numerical prefix tokens, compresses recent precipitation into prompt-side context, constructs a daily wind-driven DAG to organize upwind-to-downwind evidence, and refines city embeddings through post-pooling DAG residual propagation with root-sensitive mechanisms. We analyze \ourmodel{} on six PM$_{2.5}$ risk-inference benchmarks under a unified year-based split. The results provide protocol-bounded evidence that transport-aligned structure and role-separated interfaces yield useful risk-inference signals, while leaving no-oracle evaluation as future work.


AETDICE: Unified Framework and Offline Optimization for Nonlinear Multi-Objective RL

Woosung Kim ⋅ Youngjun Suh ⋅ Jinho Lee ⋅ Jongmin Lee ⋅ Byung-Jun Lee

Optimizing nonlinear preferences in multi-objective reinforcement learning (MORL) is essential for capturing complex trade-offs like risk aversion or fairness. However, such non-linearity has historically bifurcated nonlinear MORL objectives into two distinct paradigms: Scalarized Expected Return (SER) and Expected Scalarized Return (ESR). While SER requires global-level optimization and ESR requires non-Markovian policies, leading to fragmented optimization strategies, we bridge this divide through the Aggregation–Expectation–Transformation (AET) framework. By unifying both criteria through a tripartite decomposition of scalarization, AET provides a principled foundation for general nonlinear MORL. Building on this framework, we propose AETDICE, a tractable offline RL algorithm for AET objectives. By utilizing DICE-style density-ratio estimation in an augmented state space, AETDICE enables sample-based optimization from static datasets. Our framework resolves long-standing barriers and captures respective trade-offs induced by AET framework, which existing methods fail to address.


AffectGPT-RL: Revealing Roles of Reinforcement Learning in Open-Vocabulary Emotion Recognition

Zheng Lian ⋅ Fan Zhang ⋅ Lan Chen ⋅ Yazhou Zhang ⋅ Rui Liu ⋅ Jinyang Wu ⋅ Haoyu Chen ⋅ Xiaobai Li ⋅ Xiaojiang Peng ⋅ Bin He ⋅ Jianhua Tao

Open-Vocabulary Multimodal Emotion Recognition (OV-MER) aims to predict emotions without being constrained by predefined label spaces, thereby enabling fine-grained emotion understanding. Unlike traditional discriminative methods, OV-MER leverages generative models to capture the full spectrum of emotions and employs emotion wheels (EWs) for metric calculation. Previous approaches primarily rely on token-level loss during training. However, this objective is misaligned with the metrics used in OV-MER, and these metrics cannot be directly optimized via gradient backpropagation. To address this limitation, we turn our attention to reinforcement learning, as this strategy can optimize non-differentiable objectives. We term this framework AffectGPT-RL. Furthermore, we conduct extensive experiments to elucidate the role of reinforcement learning in this task, revealing the necessity of the reasoning process, the impact of different rewards, and the generalizability to other emotion tasks such as sentiment analysis and basic emotion recognition. Experimental results demonstrate that AffectGPT-RL yields significant performance improvements on OV-MER. Beyond this task, we also achieve remarkable performance gains on basic emotion recognition, attaining state-of-the-art results on MER-UniBench. To the best of our knowledge, this is the pioneering work exploring the role of reinforcement learning in OV-MER, providing valuable guidance for subsequent researchers. Our code is provided in the supplementary material and will be released to facilitate future research.


A Frank-Wolfe Approach to Goldstein Stationarity

Swati Padmanabhan ⋅ Zitao Song ⋅ Zhe Zhang

We investigate the problem of finding $(\delta, \varepsilon)$-stationary points (also called Goldstein stationary points) for nonsmooth nonconvex Lipschitz functions. This problem has recently garnered significant attention, with several non-asymptotic convergence guarantees. However, existing methods for such guarantees require prior knowledge of problem parameters and lack anytime convergence. Our work addresses these issues by providing an algorithm that satisfies two properties: Firstly, our algorithm is parameter-free; secondly, it generates a sequence of $(\varepsilon, \varepsilon)$-stationary points with monotonically decreasing $\varepsilon$. This algorithm achieves the (non-stochastic) state-of-the-art complexity for any $\varepsilon > 0$ and ensuring convergence to a Clarke stationary point in the limit. Central to our result is the insight that one can slightly modify prior algorithms for this problem and view them within the framework of the classic Frank-Wolfe algorithm.


AgentHop: A Diagnostic Benchmark for Agentic Multi-Hop Scientific Question Answering

Chanhee Park ⋅ Jeongho Yoon ⋅ Sungbin Han ⋅ Hyeonseok Moon ⋅ Heuiseok Lim

Agentic tasks require a large language model to interact with the world, navigating information and gathering evidence across multiple steps with restricted resources. Due to this complexity, agentic task failures arise from various sources, and pinpointing these failure causes is essential to diagnose and improve agentic systems. Existing benchmarks, however, tend to focus on a single leaderboard score, leaving the underlying failure modes opaque. To fill this gap, we introduce AgentHop, a diagnostic benchmark of 1,011 multiple-choice questions paired with a controlled seven-tool sandbox under fixed token, turn, and tool-call constraints. AgentHop reveals model vulnerabilities by dissecting a single accuracy score along four axes of agent operation: retrieval, synthesis, tool-call, and resource management. Across 19 models, we find that behavior clusters by model family, with tool-call signatures revealing distinct family fingerprints: GPT models commit early, Anthropic and GLM checkpoints verify before committing, DeepSeek and Kimi over-search, and Gemini-3 Pro stays balanced. Decomposed axes further expose within-family structure: Claude Opus 4.6 and Sonnet 4.6 land within one accuracy point yet diverge on retrieval-versus-synthesis emphasis, with Opus retrieving more and Sonnet synthesizing better. We release the full benchmark set and the harness to support diagnostic agent benchmarking.


AgenTracer-v2: Agentic Failure Tracer for LLM Agentic Systems

Guibin Zhang ⋅ Haoyu Lu ⋅ Junhao Wang ⋅ He Zhu ⋅ Kun Wang ⋅ Wangchunshu Zhou ⋅ Shuicheng Yan

LLM-powered agentic systems increasingly operate across complex, long-horizon task regimes, yet their growing architectural complexity has intensified system-level fragility and failure rates. The task of \emph{agentic failure attribution} seeks to identify the specific agent or execution step that causes failure within lengthy traces. However, most existing approaches remain constrained to passive, full-trajectory analysis, which limits their effectiveness in ultra-long and dynamically evolving trajectory attributions. To address this limitation, we introduce AgenTracer-v2, an autonomous, multi-hop failure attribution agent that reconceptualizes tracing as an active, tool-augmented diagnostic process. Concretely, we develop Tracer-Flywheel, a framework-agnostic data synthesis engine with deterministic system rollback, yielding an 8K dataset of precisely annotated failure trajectories. Leveraging this supervision, AgenTracer-v2 is trained to perform token-efficient global summarization, targeted inward inspection, and outward exploration via integrated diagnostic tools. Extensive experiments demonstrate that AgenTracer-v2 \textbf{(I)} consistently surpasses proprietary models such as \textsc{GPT-5.2} and \textsc{Gemini-3-Pro}, as well as prior specialized tracers, by up to 21.83% in attribution accuracy, \textbf{(II)} maintains robust performance on ultra-long trajectories exceeding 300K tokens and \textbf{(III)} delivering corrective feedback that measurably improves downstream agentic systems and multi-LLM reinforcement training.


AgentTailor: Dual-Gate LLM-Assisted Re-ranking for Personalized Long-Tail Recommendation

Jinpeng Chen ⋅ Zhenye Yang ⋅ Huan Li ⋅ Kaimin Wei ⋅ Senzhang Wang ⋅ Hongbo Gao

Recommender systems are central to modern digital platforms but remain vulnerable to popularity bias, where head items dominate exposure while many long-tail items are underrepresented. Existing long-tail methods often rely on global reweighting or opaque ranking models, making it difficult to personalize tail exposure without degrading recommendation utility. To address this challenge, we propose AgentTailor, an interpretable LLM-assisted framework for personalized long-tail recommendation. AgentTailor decomposes recommendation into two complementary stages. First, a semantic recall agent retrieves a semantically aligned candidate pool, determining whether relevant long-tail items are accessible to later ranking. Second, a conditional tail-specialist branch performs evidence-based re-ranking through a dual-gate tail-pressure router: the activation gate decides whether LLM-assisted evidence assessment should be invoked, while the placement gate decides whether the extracted evidence is allowed to modify the final ranking. Activated users receive item-level relevance and tail-fit evidence for routed candidates, but the final list is changed only when stable placement is warranted; otherwise, the Stage 1 order is preserved. Thus, the LLM provides auditable evidence rather than freely generating the final ranking. Experiments on MovieLens-1M and BookCrossing show that AgentTailor improves long-tail recommendation while maintaining competitive overall performance under a controlled LLM re-ranking protocol. Further analyses demonstrate its cross-backbone generalization, reliable structured outputs, and controllable trade-off between relevance preservation and long-tail exposure. The code and data are available at https://anonymous.4open.science/r/AgentTailor-CAEF.


AI Evaluation Should Require Standardized Item-Level Data Releases

Han Jiang ⋅ Susu Zhang ⋅ Dongyao Zhu ⋅ Yuzhuo Bai ⋅ Sang Truong ⋅ Xiaoyuan Yi ⋅ Sanmi Koyejo ⋅ Xing Xie ⋅ Ziang Xiao

This position paper argues that standardized item-level benchmark data should become the default infrastructure for AI evaluation. Current evaluations suffer from underspecified item selection, construct misalignment, and poor generalization. The root cause of these failures is a misplaced focus on aggregate model scores. Without item-level evidence, validity claims cannot be assessed, resulting in inflated capability claims, misdirected research, and unwarranted trust in deployed systems. Our position is that designing valid evaluations requires empirical evidence from item-level model responses, and the standardized release of such data should be treated as core AI evaluation infrastructure. Such a release, in addition, enables transparency, replicability, and auditability of evaluation results. To show the norm is both feasible and consequential, we construct OpenEval, an item-level archive of 10M responses across 155k items from widely-used benchmarks, under a unified schema that the AI evaluation community can develop upon. We demonstrate how item-level data can identify low-quality items, document construct misalignment, and recover validity evidence about benchmarks' internal structure. We address objections around contamination and author burden, and show each is tractable relative to the cost of decisions made on claims that cannot be trusted.

Multi-Agent Debate (MAD) improves reasoning on multimodal tasks, but its accuracy and token efficiency depend on how agents interact during debate. A largely overlooked design choice is which interaction modalities agents use to exchange evidence. Existing methods rely on fixed text- or graph-based interaction, which fails to convey fine-grained multimodal evidence, degrading accuracy while inflating token cost due to verbose descriptions. To address this, we propose Adaptive Interaction in Multi-Agent Debate (AIM), an adaptive framework that dynamically selects the best combination of interaction modalities for each instance. Beyond text and graph, AIM introduces task-dependent modality interaction, where agents exchange targeted regions of interest from the task-specific multimodal input, such as image regions, audio segments, and spatio-temporal clips. To determine which combination is the best for each instance, AIM first generates a structured baseline response, extracts interpretable routing features, and finally uses a lightweight router to select a modality combination that avoids both under-fusion (i.e., agents miss decisive evidence) and over-fusion (i.e., redundant modalities inflate token cost) risks. Across eight multimodal question answering (QA) benchmarks of Visual-QA, Audio-QA, and Video-QA, AIM achieves the highest accuracy on all datasets, improving upon the best single-modality baseline by up to 8.9% while reducing token cost by up to 53.4% relative to the best fixed-modality baseline.


AIR: Rethinking Image-Text Offset Alignment in Multimodal Contrastive Representation Space

Guimeng Liu ⋅ Milad Abdollahzadeh ⋅ Ngai-Man (Man) Cheung

Image-text offset alignment, which refers to the consistency between visual transformation directions and their corresponding textual transformation directions, has become an important relational property in multimodal contrastive representation space (e.g., CLIP or SigLIP). This property is widely utilized in text-guided generative modeling tasks, such as generative model domain adaptation and text-guided Image-text offset alignment, which refers to the consistency between visual transformation directions and their corresponding textual transformation directions, has become an important relational property in multimodal contrastive representation spaces such as CLIP and SigLIP. This property plays a central role in text-guided generative modeling tasks, including generative model domain adaptation and text-guided image editing, where textual offsets are used to guide visual transformations. These methods fundamentally rely on the assumption that image and text offsets are well aligned, such that textual transformation directions can provide reliable guidance for corresponding visual transformations. However, the validity of this assumption has not been systematically examined. In this work, we question this foundational assumption by conducting a comprehensive empirical analysis of image-text offset alignment in multimodal contrastive representation space. Our findings reveal not only noticeable offset misalignment but also a meaningful positive correlation between image-text offset misalignment and semantic concept distance across six large datasets and eight contrastive vision-language models. Based on this discovery, we propose Adaptation with Iterative Refinement (AIR), a method that iteratively refines text offsets through anchor sampling and our proposed concept description learning to reduce image-text offset misalignment and improve guidance accuracy. Comprehensive experiments on zero-shot generative model domain adaptation and text-guided image editing, including qualitative, quantitative, and user studies, consistently show that AIR enables state-of-the-art performance in these tasks. Code and additional experiments are available in the supplementary material.

Multivariate conformal prediction requires nonconformity scores that compress residual vectors into scalars while preserving certain implicit geometric structure of the residual distribution. We introduce a Multivariate Kernel Score (MKS) that produces prediction regions that explicitly adapt to this geometry. We show that the proposed score resembles the Gaussian process posterior variance, unifying Bayesian uncertainty quantification with the coverage guarantees of frequentist-type. Moreover, the MKS can be decomposed into an anisotropic Maximum Mean Discrepancy (MMD) that interpolates between kernel density estimation and covariance-weighted distance. We prove finite-sample coverage guarantees and establish convergence rates that depend on the effective rank of the kernel-based covariance operator rather than the ambient dimension, enabling dimension-free adaptation. On regression tasks, the MKS reduces the volume of prediction regions significantly, compared to ellipsoidal baselines while maintaining nominal coverage, with larger gains at higher dimensions and tighter coverage levels.

Detecting whether specific text samples were used to train large language models is increasingly important for privacy, copyright and data auditing. Knockoff-based training data detection (KTD) offers a promising approach to controlling the false discovery rate (FDR). However, its effectiveness relies on null sign-flip symmetry induced by exchangeability between each null sample and its knockoff. This condition is difficult to satisfy in natural language settings. We study training data detection under approximate exchangeability, and demonstrate that deviations from exact exchangeability induce measurable null-tail asymmetry, which in turn leads to systematic FDR inflation under the standard knockoff+ threshold. This effect can be captured by pathwise bounds, which explain when and why KTD becomes anti-conservative. Based on this, we propose Asymmetry-adjusted KTD (AKTD), a series of procedures that correct for asymmetry by estimating the penalty term. We introduce an external null variant that uses a separate calibration pool containing only null samples to estimate the penalty term.This enables detection with FDR control without using labels from the evaluated candidate set. Experiments on WikiMIA and MIMIR using GPT and Pythia demonstrate that the asymmetry term accurately predicts the observed FDR expansion, and AKTD restores FDR control while maintaining effective detection capabilities.


ALAM: Algebraically Consistent Latent Transitions for Vision-Language-Action Models

Zuojin Tang ⋅ Haoyun Liu ⋅ Xinyuan Chang ⋅ Changjie Wu ⋅ Dongjie Huo ⋅ Yandan Yang ⋅ Bin Liu ⋅ Zhejia Cai ⋅ Feng Xiong ⋅ Mu Xu ⋅ jiachen luo ⋅ De Ma ⋅ Zhiheng Ma ⋅ Gang Pan

Vision-language-action (VLA) models remain constrained by the scarcity of action-labeled robot data, whereas action-free videos provide abundant evidence of how the physical world changes. Latent action models offer a promising way to extract such priors from videos, but reconstruction-trained latent codes are not necessarily suitable for policy generation: they may predict future observations while lacking the structure needed to be reused or generated coherently with robot actions. We introduce \textbf{ALAM} (\textbf{A}lgebraic \textbf{L}atent \textbf{A}ction \textbf{M}odel), an Algebraically Consistent Latent Action Model that turns temporal relations in action-free video into structural supervision. Given frame triplets, ALAM learns latent transitions that are grounded by reconstruction while being regularized by composition and reversal consistency, encouraging a locally additive transition space. For downstream VLA learning, we freeze the pretrained encoder and use its latent transition sequences as auxiliary generative targets, co-generated with robot actions under a joint flow-matching objective. This couples structured latent transitions with flow-based policy generation, allowing the policy to exploit ALAM's locally consistent transition geometry without requiring latent-to-action decoding. Representation probes show that ALAM reduces additivity and reversibility errors by 25--85$\times$ over unstructured latent-action baselines and improves long-horizon cumulative reconstruction. When transferred to VLA policies, ALAM raises the average success rate from 47.9\% to 85.0\% on MetaWorld MT50 and from 94.1\% to 98.1\% on LIBERO, with consistent gains on real-world manipulation tasks. Ablations further confirm that the strongest improvements arise from the synergy between algebraically structured latent transitions and joint flow matching.


A Latent World-Action Model with Jointly Aligned Reasoning

Hao Luo ⋅ Wanpeng Zhang ⋅ Yicheng Feng ⋅ Sipeng Zheng ⋅ Haiweng Xu ⋅ Chaoyi Xu ⋅ Ziheng Xi ⋅ Yuhui Fu ⋅ Zongqing Lu

Despite recent progress, existing general robot policies, particularly Vision-Language-Action models, are still primarily trained to map observations directly to actions using sparse action supervision. As a result, they often learn shallow observation-to-action correlations instead of deeper representations of world dynamics, including object motion, physical interaction, and long-horizon task progression. Recent world-action models attempt to address this limitation through future-frame prediction, but pixel-space rollout is computationally expensive and poorly aligned with the abstractions required for control, forcing the model to reconstruct visual details irrelevant to action. We present Being-H0.7, a latent world-action model that introduces future-aware reasoning into VLA-style policies without generating future frames. Our model inserts learnable latent queries between perception and action, forming an explicit reasoning interface for action generation. To train this interface, we adopt a future-informed dual-branch design: a deployable prior branch infers latent states from the current context, while a training-only posterior branch replaces latent queries with embeddings from future frames. Joint alignment between the two branches enables the prior branch to learn predictive, action-relevant latent structure directly from current observations. Under a controlled training setup for comprehensive validation, Being-H0.7 achieves state-of-the-art or competitive performance across diverse benchmarks, while further demonstrating strong deployability on real-robot manipulation tasks.


Align as You Couple: Learning Spatial Resolved Inference from H&E Images with Mollified Flow Matching

Rongchao Zhang ⋅ Guangyuan Dong ⋅ Siheng Wang ⋅ Zhengtao Yao ⋅ Junhao Dong ⋅ Haoyang Li

Spatial resolved inference (SRI) is a critical technique in biomedical research, which aims to measure RNA sequence abundances by associating them with whole-tissue images at a fine-grained molecular scale. However, the extremely low throughput and distributional discrepancies involved pose significant challenges to cellular morphologies captured in hematoxylin and eosin (H&E) connection benchmarking, and existing SRI frameworks suffer from coverage gaps. In this paper, we propose MollFlow, a novel generative Mollified Flow matching model for spatial resolution inference, which is designed to learn the probabilistic dependency between the distributions of whole-slide images and expression profiles. We innovatively reframe the anchors of flow matching as a sequential informative prior derived from whole-slide images rather than isotropic Gaussian noise. During inference, MollFlow progressively refines the prior trajectory, ultimately converging to biologically plausible inferences. For the challenging coupling, which requires enabling intermediate states to explore pathologically aligned subspaces, we carry out a boundary mollified strategy to stabilize the velocity field, rather than being constrained by rigid trajectories. Endowed by this, our unique modulation generation perspective allows the model to retain coupling alignment for manifold incommensurability and prevent coverage gaps. Empirically, we find that MollFlow is capable of accurately inferring gene expression while exhibiting excellent performance in a variety of application scenarios.


All-Addition Spiking Diffusion Models with Attention Enhancement

Yuhan Zhang ⋅ Yuanpei Chen ⋅ Zhou Jie ⋅ Weihang Peng ⋅ Xiaode Liu ⋅ Yinglei Wang ⋅ Yufei Guo

Diffusion models have achieved state-of-the-art (SoTA) performance in generative tasks through their iterative denoising mechanism, yet they remain computationally intensive and energy-prohibitive. Spiking Neural Networks (SNNs) represent a promising energy-efficient alternative to traditional Artificial Neural Networks (ANNs). Their binary spike encoding enables the conversion of multiplication operations into addition operations, which is a key attribute for reducing energy consumption. Integrating the strong generative capabilities of diffusion models with the energy efficiency of SNNs thus forms a highly promising research direction. In adherence to the design goal of all-additive computation, this paper proposes a novel spiking diffusion model with attention enhancement. Specifically, we adopt non-negative ternary spiking neurons (NNTSN) to construct the U-Net backbone, which helps mitigate information loss during feature processing. To align NNTSN with the requirement of using only additive operations, we further design two core components: a membrane potential attention mechanism and an all-additive shortcut module. Leveraging a reparameterization technique, the proposed diffusion model is developed to retain only additive operations during inference, ensuring strict compliance with energy-saving design principles. Extensive experiments on relevant benchmarks demonstrate that our method achieves SoTA performance in generative tasks. It also substantially outperforms other SNN-based generative models while using fewer time steps, which validates both its effectiveness and efficiency.

Despite the critical role of temporal reasoning in Large Language Models (LLMs), current synthetic benchmarks fall short: they either introduce information leakage or, upon addressing this issue, merely assess questions with explicit time points, failing to rigorously evaluate implicit temporal reasoning capabilities. To bridge this gap, we propose a formally grounded evaluation framework, ALTER, based on Allen’s Interval Algebra. By combining basic relations into complex networks and converting these structures into natural language forms through an automated pipeline, we construct a validated benchmark of 2,500 question-answer pairs, ensuring a controllable and scalable evaluation environment. We conduct experiments on a representative selection of LLMs. Empirical evaluations on diverse LLMs yield three key findings: (1) identifying performance decline of LLMs in highly complex scenarios involving dense events; (2) isolating complex logical interactions—rather than linguistic comprehension—as an unresolved bottleneck; and (3) revealing a critical evaluation bias induced by narrative chronology, with LLMs exploiting textual shortcuts. In summary, these findings reveal the blind spots of current evaluation paradigms, providing a rigorous foundation for the robust evaluation of complex temporal reasoning.

Diffusion models are increasingly used as controllable samplers, whose generations can be steered at inference time according to a chosen reward function. While such rewards are typically defined on individual samples, for many applications it is desirable to steer according to distribution-level rewards, for example to calibrate with population-level information or to encourage diversity. In both cases, simply incorporating the reward gradient into the dynamics, while often effective, comes with few theoretical guarantees on the sampled distribution. For pointwise rewards, recent work has therefore sought to develop a principled framework for targeting a prescribed tilted distribution using particle reweighting. However, an analogous theoretically-grounded approach for distributional rewards is currently lacking. In this work, we formulate inference-time distributional control as targeting a tilted measure under a mean-field framework, and derive a weighted interacting particle scheme to target it in a principled manner. Our framework recovers pointwise-reward steering as a special case, while providing a theoretical foundation for existing batch-level steering methods. Empirically, we verify that the procedure correctly targets the prescribed distribution in tractable low-dimensional settings, and investigate its behaviour in higher-dimensional protein conformation tasks.


A Mechanistic Analysis of Looped Reasoning Language Models

Hugh Blayney ⋅ Alvaro Arroyo ⋅ Johan Obando Ceron ⋅ Pablo Samuel Castro ⋅ Aaron Courville ⋅ Michael Bronstein ⋅ Xiaowen Dong

Reasoning has become a central capability in large language models. Recent research has shown that reasoning performance can be improved by looping an LLM’s layers in the latent dimension, resulting in looped reasoning language models. Despite promising results, few works have investigated how their internal dynamics differ from those of standard feedforward models. In this paper, we conduct a mechanistic analysis of latent states and layer behavior in looped language models, focusing in particular on how the stages of inference observed in feedforward models compare to those observed in looped ones. We analyze cyclic recurrence and show that for many of the studied models each layer in the cycle converges to a distinct fixed point; consequently, the recurrent block follows a consistent cyclic trajectory in the latent space. We provide evidence that as these fixed points are reached, attention-head behavior stabilizes, leading to constant behavior across recurrences. Empirically, we discover that recurrent blocks learn stages of inference that closely mirror those of feedforward models, repeating these stages in depth with each iteration. We study how architectural choices influence the emergence and stability of these cyclic fixed points and stages of inference, providing practical guidance for looped language model design.


A meshfree exterior calculus for generalizable and data-efficient learning of physics from point clouds

Benjamin D Shaffer ⋅ Brooks Kinch ⋅ M. Ani Hsieh ⋅ Nathaniel Trask

We introduce a meshfree exterior calculus (MEEC) for learning structure-preserving descriptions of physics on point clouds, and use it to build MEEC-Net, a data-efficient surrogate that transfers across resolutions, geometries, and physical parameters. MEEC equips an $\epsilon$-ball graph with virtual node and edge measures via a single sparse Schur complement solve; the resulting complex satisfies discrete conservation exactly, is end-to-end differentiable in the point positions, and exposes a direct geometry-to-physics link without the mesh-generation step required by conventional structure-preserving discretizations. MEEC-Net learns unknown physics as a shared edge-wise flux law in an SO($d$)-invariant local frame, so the same kernel produces compatible fluxes on any point cloud whose features lie in the training range. We prove a solution-error bound that splits into discretization and kernel-approximation terms which is independent of problem geometry, explaining the observed transfer from very few examples. We show that single-solution training transfers to unseen geometries, boundary conditions, and physical parameters. On five canonical PDE benchmarks MEEC-Net achieves 1–2 orders of magnitude lower out-of-distribution error than baseline neural-operator approaches. On the SimJEB structural-bracket benchmark it achieves competitive error while using substantially fewer training geometries.

Analog circuit design remains highly dependent on expert knowledge due to the complexity of device-level interactions and topology design. Recent transformer-based approaches for device-level topology generation have shown promise, yet they suffer from low electrical validity without human-in-the-loop (HITL) training and severe memorization caused by sequence-based circuit representations. In this work, we propose AnalogToBi, a framework for device-level analog circuit topology generation. AnalogToBi introduces circuit-type conditioning for categorizing heterogeneous multi-type topology datasets, device renaming augmentation to mitigate memorization, a bipartite graph representation for improved structural generalization, and grammar-guided decoding to enforce structural validity during bipartite graph generation. Experimental results demonstrate that AnalogToBi achieves high validity and novelty without HITL training while effectively avoiding memorization of training topologies.


An Analytical Model of Compute-limited Multistage Training Pipelines

Nishil Patel ⋅ Jin Hwa Lee ⋅ Basile Confavreux ⋅ Andrew Saxe

Modern foundation models are trained through multistage pipelines, typically beginning with unsupervised pretraining, followed by supervised fine-tuning (SFT) on step-by-step solutions, and reinforcement learning (RL) with bulk feedback. Empirically, the design of these stages, particularly the compute allocated to each, strongly affects performance. Yet our theoretical understanding of multistage training remains limited. Here, we introduce a simple, analytically tractable model of multistage training. The model reproduces several qualitative phenomena: SFT and RL can be complementary, pretraining quality imposes a performance ceiling with power-law scaling, optimal compute allocation varies with data quality and pretraining level, and RL can selectively amplify task-relevant pretrained skills. Our model distills the complex behaviors into key ingredients, allowing clearer understanding of how and where they arise. Overall, our model provides a theoretical starting point for explaining the phenomenology of multistage training pipelines in the compute-limited regime.

Reinforcement learning has become a central tool for large language model (LLM) post-training, where policy gradient methods are routinely deployed off-policy, even though vanilla policy gradient assumes on-policy sampling. We study when such off-policy deployment is justified, and what roles its core ingredients, including importance sampling, KL regularization, and baselines, play in correcting the resulting distribution shift. We show that for a general class of off-policy policy gradient objectives, only two corrections preserve the optimal policy as a stationary point: trajectory-level importance weighting, which suffers from the well-known curse of horizon, and KL regularization paired with a grouped-mean baseline. Focusing on the KL route, we identify it with the group-relative policy gradient (GRPG) update, establish an equivalence to a trajectory-level Bellman residual minimization objective, and obtain a finite-sample regret bound under all-policy coverage. This also gives a theoretical account of the empirically successful GRPO objective. We further uncover a hidden role of baselines in the off-policy setting: a constant shift in the baseline implements pessimism, the standard remedy in offline RL for partial data coverage. We instantiate this via the asymmetric REINFORCE objective, and connect it to the game-theoretic offline RL framework, showing a regret bound with single-policy coverage. Together, these results offer a unified theoretical view of off-policy policy gradient in LLM post-training and connect it to classical ideas in offline RL.


Anchor3DGS: Feed-Forward 3D Gaussian Splatting with Compact Anchor-Based Representation

Sheng Ye ⋅ Jiangke Lin ⋅ Fudong Wang ⋅ Yi Yuan ⋅ ZhenHui Dong ⋅ Yong-jin Liu

Recent feed-forward 3D Gaussian Splatting (3DGS) methods have enabled effective multi-view 3D reconstruction by predicting Gaussian primitives in a single forward pass. However, existing approaches predominantly adopt a pixel-aligned paradigm that predicts one Gaussian per input pixel, coupling the Gaussian count to the input resolution and number of views. This also leads to redundant primitives in textureless regions and insufficient coverage in geometrically complex areas. We present Anchor3DGS, a pose-free feed-forward 3DGS framework that replaces pixel-aligned prediction with a compact anchor-based representation. Guided by information entropy, our method places a set of 3D anchors in the scene, concentrating them in regions of higher visual complexity while maintaining broad spatial coverage. Each anchor performs occlusion-aware aggregation of multi-view image features, exchanges information with other anchors, and is decoded into a set of local Gaussians. A lightweight feature enhancement module further refines the coarse RGB output from Gaussian splatting and improves rendering fidelity. Extensive experiments on diverse benchmark datasets demonstrate that our Anchor3DGS achieves state-of-the-art novel view synthesis quality while only using an order of magnitude fewer Gaussian primitives.


Anchored Protein Engineering

Chi Zhang ⋅ Maria Rosaria Briglia ⋅ Litu Rout ⋅ Jeffrey Ouyang-Zhang ⋅ Iacopo Masi ⋅ Sanjay Shakkottai ⋅ Adam Klivans ⋅ Daniel Diaz

Inspired by recent work on discrete diffusion for anchored unmasking, we introduce a new training-free framework for protein engineering that can generate higher-order variants that improve one or more phenotypes while exhibiting favorable epistatic interactions. Prior approaches require mutation sites to be specified in advance and can drive sequences away from the natural protein manifold. We address these limitations with Anchored Protein Engineering via quantized eXpectation (APEX), a decoding-time method consisting of anchor selection, tilted decoding, and iterative refinement. APEX selects anchor residues using a reward-gradient score with an explicit correction for pairwise epistasis, decodes substitutions for anchors by tilting the base model's per-site predictive distributions along reward gradients, and revisits low-scoring anchors iteratively for optional test-time scaling. All phases query the base model only through masked-conditional logits, a primitive shared by masked language models, discrete-diffusion models, and inverse-folding networks; as a result, APEX works out of the box across commonly-used model families. Across various protein-engineering benchmarks, including thermostability and solubility, APEX achieves state-of-the-art results and avoids the predictor-gaming artifacts that reward-search-only baselines tend to produce.

A key bottleneck in adversarial transfer is a *trajectory-level geometric disconnect*: ambient gradients often drift away from the intrinsic data manifold, causing surrogate-specific overfitting. To rectify this, we propose *Manifold Anchored Bilevel Transfer (MABT)*, a unified framework that anchors adversarial trajectories to the shared semantic subspace. MABT introduces a relaxed manifold-anchoring operator as a semantic rectifier to suppress off-manifold noise. With this constraint, we cast transfer attack generation as a distributional bilevel optimization problem that learns a geometry-aligned initialization by minimizing expected transfer risk under a surrogate uncertainty distribution. We further develop a Hessian-free solver with linear-time complexity to handle the resulting hierarchy. Experiments demonstrate that MABT consistently boosts the transferability of **10** baselines across diverse attack scenarios and defense mechanisms (e.g., **63.31\%** average ASR $\uparrow$ across **28** attacker combinations).

We study the complexity of smoothed agnostic learning of halfspaces on $\{\pm 1\}^n$ under uniform marginals in the model of \citet{KM25} where each input coordinate is independently flipped with probability $\sigma \in (0, {1}/{2})$. We show that $L^1$ polynomial regression achieves runtime and sample complexity $\tilde{O}(n^{O(\log(1/\varepsilon)/\sigma)})$, and prove a nearly matching Statistical Query complexity lower bound of $n^{\Omega(\log(1+\sigma/\varepsilon^2)/\sigma)}$. This complements the recent work of \citet{DK26}, which established analogous bounds in the continuous setting under Gaussian marginals.

Quantifying bias in training data---independently of any downstream model---is critical for fair machine learning, especially when data collection and model design occur in silos. Yet most fairness research focuses on model-level mitigation, leaving the data side underexplored. We address this gap by deriving multiple information-theoretic measures for auditing unfairness bias in data without requiring access to model predictions. We define our measures over feature sets and use Shapley values to deduce feature-level contributions to global unfairness bias. For some of these measures, we prove that Shapley values moderate redundancy across feature coalitions, matching empirical evidence in prior work. We develop our measures within an axiomatic framework that embeds a fairness notion, such as statistical parity (SP) or equalized odds (EO), into desired and undesired independence properties. We characterize baseline measures that guide those we develop, for example, by showing that a baseline upper bounds the bias of any downstream restricted predictor. Finally, we show that SP- and EO-aligned independence properties can conflict, paralleling Kleinberg et al.'s result on incompatible fairness notions. We validate our measures through an empirical study on real and synthetic datasets, assessing how well they predict the true bias of an extensively tuned neural network trained on the same data, and via a simple feature selection experiment. For synthetic data, we propose a parametric linear structural causal model that enables controlled generation of diverse correlation structures and bias types. Overall, our analysis provides a theoretically and empirically validated guideline for selecting an unfairness measure given a group fairness notion and testable data conditions.


AOT-POT: Adaptive Operator Transformation for Large-Scale PDE Pre-training

Qitan Lv ⋅ Hong Wang ⋅ Hao Zhongkai ⋅ Wen Wu ⋅ Xuenan Xu ⋅ Bowen Zhou ⋅ Feng Wu ⋅ Chao Zhang

Pre-training neural operators on diverse PDE datasets has emerged as a promising paradigm for building general-purpose surrogate models in scientific machine learning. However, the inherent complexity and structural diversity of PDE solution operators make multi-PDE pre-training fundamentally difficult. Existing approaches address this mainly by enlarging model capacity, while the target solution operators themselves remain unchanged. Inspired by classical numerical analysis---where a problem-tailored operator transformation reformulates a complex PDE into a numerically easier equivalent---we propose to similarly reformulate the target solution operators, turning a heterogeneous set of complex, structurally divergent operators into a simpler and better-aligned equivalents. Since different PDEs admit different simplifications, the transformation must be adaptive and input-dependent, so that a single neural operator can approximate the entire family jointly. We instantiate this idea as AOT-POT (Adaptive Operator-Transformation for Pre-training Operator Transformer), which realizes such a transformation by expanding the hidden representation into multiple parallel streams, aggregating and redistributing them with input-dependent weights before and after each sub-layer, and mixing streams through Sinkhorn-projected doubly stochastic matrices for stable training. Together, these mechanisms reformulate the diverse, complex solution operators into a simpler equivalents that a single architecture can approximate jointly. Empirically, AOT-POT achieves state-of-the-art results on 12 PDE benchmarks with only 3% additional parameters, reducing the relative L2 error by up to 77.6% (40.9% on average). Fine-tuning further reduces L2 error by up to 92% on in-domain PDEs and 89% on out-of-domain PDEs, confirming that adaptive operator transformation is an orthogonal and effective axis for advancing PDE foundation models, beyond merely scaling model capacity.


APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts

Emily Jin ⋅ Joy Hsu ⋅ Yiqing Xu ⋅ Weiyu Liu ⋅ Nick Haber ⋅ Jiajun Wu

Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a task, a robot must decompose goals, select task-relevant objects, and sequence actions, while ensuring that plans satisfy spatial constraints such as limited free space and object collisions. In this work, we propose APIVOT, a VLM-based planner that adaptively interleaves language and visual thoughts for long-horizon planning. APIVOT learns to rely on language for semantic reasoning, while using visual thoughts as imagined future states for internal verification of geometric feasibility. On long-horizon kitchen tasks, APIVOT outperforms general-purpose VLMs and prior planning frameworks, achieving the largest gains in spatially constrained settings. We find that APIVOT learns meaningful modality selection behavior, demonstrating that adaptive interleaving of vision-language thoughts improves both planning success and reasoning efficiency.


Approaching I/O-optimality for Approximate Attention

Pál András Papp ⋅ Aleksandros Sobczyk ⋅ Anastasios Zouzias

We revisit the I/O complexity of attention in large language models. Given query-key-value matrices $Q,K,V\in\mathbb{R}^{n\times d}$, and a machine with fast memory size $M$, the goal is to compute the "attention matrix" $A=\mathrm{softmax}(QK^{\top}/\sqrt{d})V$ with the minimal number of data transfers between fast and slow memory. Existing methods in the literature, most notably FlashAttention and its variants, incur an I/O cost that depends quadratically on $n$, while a trivial lower bound only requires $\Omega(nd)$ I/O's to read the inputs and write the output. In this work, we present a technique for computing attention where the I/O cost only depends almost-linearly on $n$ in most parameter regimes. This is achieved by developing I/O-efficient algorithms inspired by the recent approximate attention framework of Alman and Song [NeurIPS'23]. We also prove corresponding lower bounds in each parameter regime to show that our algorithms are indeed close to I/O-optimal.

Long-video understanding with multimodal LLMs is fundamentally constrained by the mismatch between video length and the model's limited visual context budget, making frame selection essential for practical reasoning. Existing query-aware methods typically rely on heuristic pipelines: they assign one-dimensional query-conditioned scores to frames and then apply sampling, ranking, or extraction rules to form the final subset. While effective in practice, such methods are generally not derived from an explicit optimization objective and struggle to capture the heterogeneous evidence required for complex reasoning. In this paper, we revisit long-video frame selection from a principled optimization perspective. We formulate it as a structured evidence-allocation problem and propose OTFS, an optimal transport framework that maps multiple evidence sources onto video frames. To reflect the asymmetry of this problem---important evidence should be preserved, whereas redundant frames need not absorb mass---we instantiate this view with a semi-unbalanced entropic optimal transport objective and an efficient solver, OTFS-Sinkhorn. Combined with an exact dynamic program for coverage-aware subset extraction, OTFS yields a practical two-stage training-free pipeline. We further show that, in the single-source case, our formulation induces a Gibbs-type distribution over frames, making Q-Frame's temperature-scaled distribution a special case of our framework while also providing a broader perspective for understanding related prior methods. Experiments on three long-video understanding benchmarks show that OTFS consistently outperforms strong training-free baselines. Code will be released.


AQBENCH: Benchmarking Neural Surrogates for Air Quality Forecasting

Siddharthan Dileep ⋅ Sanchit Bedi ⋅ Pareshbhai D Parmar ⋅ Ayush Maheshwari ⋅ Sri H Kota ⋅ N M Anoop Krishnan

Neural surrogate benchmarks for spatiotemporal PDEs are dominated by idealized test beds with periodic boundaries and synthetic dynamics. Architecture rankings established on these benchmarks do not transfer to real atmospheric chemistry, and standard regression metrics conceal the failure modes that determine operational utility. AQBench evaluates ten architectures across five model families on high-resolution WRF-Chem simulations of $PM_2.5$, $NO_2$, and $CO$ over the Indian subcontinent, against five operationally-grounded objectives: long-horizon stability, exceedance detection, extreme-episode bias, advection-diffusion residuals, and multi-pollutant forecasting. Three findings emerge. Spectral neural operators, the strongest family on canonical PDE benchmarks, rank in the bottom four on 168 hour RMSE here. Models that lead on aggregate regression underestimate during extreme pollution episodes and miss the majority of true exceedance events. Joint training on $PM_2.5$, $NO_2$, and $CO$ improves $PM_2.5$ accuracy across most architectures with no parameter increase, with Transolver gaining 21\% at 168th hour. The dataset, evaluation framework, and trained baselines are released.


ARCANA: A Benchmark for Abstraction and Analogical Reasoning

Ruqi Chen ⋅ Wei Wang ⋅ Zilei Wang

Analogical reasoning is widely considered a fundamental component of human intelligence. Few benchmarks are designed to explicitly test far analogical reasoning—reasoning between domains that are superficially dissimilar yet share deep relational or structural similarities. In this work, we introduce an analogical reasoning benchmark called ARCANA. The tasks of ARCANA require performing far analogical mappings from real-world objects or physical phenomena to abstract grid-based transformations, different from the well-known ARC-AGI benchmark which primarily emphasized innate human priors. We validate ARCANA through a human study, demonstrating that humans solve these tasks reliably. In contrast, evaluations across a range of large language models show a distinct performance gap compared to humans, revealing systematic failure modes that remain challenging for current models. To address this gap, we propose \textsc{ConceptSIR}, a reasoning framework based on concept slippage, introspection, and reflection. Experiments demonstrate that \textsc{ConceptSIR} empowers standard LLMs to solve ARCANA tasks that were previously unsolvable under baseline evaluation. Overall, our results suggest that far analogical reasoning remains an open problem for contemporary LLMs. All the materials will be available at https://sites.google.com/view/arcanacogtest.


Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows

Zixin CHEN ⋅ Peng Liu ⋅ Rui SHENG ⋅ Haobo Li ⋅ Jianhong Tu ⋅ Deng Xiaodong ⋅ KaShun SHUM ⋅ Dayiheng Liu ⋅ Huamin Qu

Language agents are increasingly deployed in complex professional workflows, with tutoring emerging as a particularly high-stakes capability that remains largely unmeasured in existing benchmarks. Effective tutor agents require more than producing correct answers or executing accurate tool calls: a robust tutor must diagnose learner state, adapt support over time, make pedagogically justified decisions grounded in educational evidence, and execute interventions within realistic learning-management systems. We introduce EduAgentBench, a source-grounded benchmark for holistically evaluating tutor agents across the full scope of teaching work. It contains 150 quality-controlled tasks across three capability surfaces: situated multi-turn tutoring, professional pedagogical judgment, and Canvas-style teaching workflow completion. Tasks are constructed through a pedagogical-insight-driven pipeline and evaluated with complementary verification signals and human review. Across a comprehensive evaluation of frontier models, our findings reveal that current models are generally capable of bounded pedagogical judgment, but still fall short of professional teaching standards in situated tutoring and autonomous teaching-workflow execution. To our knowledge, EduAgentBench is the first theory-grounded and realistic benchmark for evaluating the holistic teaching capability of tutor agents, providing a measurement foundation for developing future tutor agents that can support realistic teaching work.

Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time. While empirically powerful, these methods are prone to reward hacking: the guided model over-optimizes the reward at the cost of fidelity to the learned distribution. Prior work has attributed this to the complexity of neural reward functions or implicit biases in diffusion training, but its fundamental origins remain poorly understood. We show that reward hacking arises from an approximation made in most practical implementations of reward-guided diffusion---finite-particle plug-in estimation of the Doob $h$-function---even in the simplest non-trivial settings of Gaussian and Gaussian mixture targets with quadratic rewards. In closed form, we isolate two distinct failure modes of the plug-in estimator: it leads to reward hacking within each mode and it cannot select high-reward modes. We propose a closed-form reward damping schedule that corrects the within-mode bias with no additional compute, and clarify the role of best-of-$n$ sampling in compensating for the mode selection failure. Experiments on Gaussian mixture targets, a 2D checkerboard, and FLUX.1 text-to-image generation confirm that our theoretical insights carry over to practical settings.


ASAP: Assembly-Source Aligned Pseudocode Refinement For Binary Decompilation

Yujian Zhuang ⋅ Dehong Gao ⋅ Qichao Zhang ⋅ QiJing Lai ⋅ Jiaxin Wang ⋅ Libin Yang ⋅ Xiaoyan Cai

Large language models (LLMs) are increasingly used in binary decompilation to refine the C-like pseudocode produced by traditional rule-based decompilers. While this pseudocode is useful, it is a heuristic and lossy abstraction rather than a faithful copy of the source code. It often contains decompiler errors, especially for aggressively optimized binaries where critical low-level details are obscured. We present ASAP, an assembly-source aligned pseudocode refinement framework for binary decompilation. ASAP learns source-aligned assembly representations from paired source and binary functions using joint function-level and snippet-level contrastive alignment. A Q-Former then compresses chunk-level assembly features into a fixed number of assembly tokens that condition the decompilation LLM alongside the decompiler-produced pseudocode. During refinement, we use stochastic pseudocode masking and a relative assembly-advantage loss to reduce the model's tendency to ignore assembly features and rely only on pseudocode refining. On two decompilation benchmarks across multiple compiler optimization levels, ASAP improves the average re-execution rate from 64.7% to 71.9% and the average recompilation rate from 91.6% to 96.6% compared with the strongest baseline, offering both a new perspective and a practical solution to binary decompilation.

Machine Learning (ML) models are trained on in-distribution (ID) data but often encounter out-of-distribution (OOD) inputs during deployment---posing serious risks in safety-critical domains. Recent works have focused on designing scoring functions to quantify OOD uncertainty, with score thresholds typically set based solely on ID data to achieve a target true positive rate (TPR), since OOD data is limited before deployment. However, these TPR-based thresholds leave false positive rates (FPR) uncontrolled, often resulting in high FPRs where OOD points are misclassified as ID. Moreover, fixed scoring functions and thresholds lack the adaptivity needed to handle newly observed, evolving OOD inputs, leading to sub-optimal performance. To address these challenges, we propose ASAT, a human-in-the-loop framework that safely updates both scoring functions and thresholds on the fly based on real-world OOD inputs. ASAT maximizes TPR while controlling FPR at all times under stationary conditions, even as the system adapts over time. Under nonstationary conditions, the method adapts to distribution shifts with only transient FPR violations during the adaptation period. We provide theoretical guarantees for FPR control under stationary conditions and present extensive empirical evaluations on OpenOOD benchmarks to demonstrate that our approach outperforms existing methods by achieving higher TPRs while maintaining FPR control.

Understanding human personality is crucial for web applications such as personalized recommendation and mental health assessment. Existing studies on personality detection predominantly adopt a ``posts $\rightarrow$ user vector $\rightarrow$ labels'' modeling paradigm, which encodes social media posts into user representations for predicting personality labels (e.g., MBTI labels). While recent advances in large language models (LLMs) have improved text encoding capacities, these approaches remain constrained by limited supervision signals due to label scarcity, and under-specified semantic mappings between user language and abstract psychological constructs. We address these challenges by proposing ROME, a novel framework that explicitly injects psychological knowledge into personality detection. Inspired by standardized self-assessment tests, ROME leverages LLMs’ role-play capability to simulate user responses to validated psychometric questionnaires. These generated question-level answers transform free-form user posts into interpretable, questionnaire-grounded evidence linking linguistic cues to personality labels, thereby providing rich intermediate supervision to mitigate label scarcity while offering a semantic reasoning chain that guides and simplifies the text-to-personality mapping learning. A question-conditioned Mixture-of-Experts module then jointly routes over post and question representations, learning to answer questionnaire items under explicit supervision. The predicted answers are summarized into an interpretable answer vector and fused with the user representation for final prediction within a multi-task learning framework, where question answering serves as a powerful auxiliary task for personality detection. Experiments on two real-world datasets show that ROME outperforms state-of-the-art baselines, achieving relative improvements of 13.60% and 19.72% on Kaggle and Pandora, respectively. Our code is available at https://anonymous.4open.science/r/ROME-6641.


A Solver-Efficient Neural Adversarial Attack on Subgraph Matching Models

Ninad Gandhi ⋅ Mayukh Mondal ⋅ Brian S Mackwan ⋅ Soumen Chakrabarti ⋅ Abir De

Neural subgraph matching (NSGM) models are widely used for graph retrieval, where relevance labels for query–corpus pairs are determined by whether the query graph is a subgraph of the corpus graph. Despite widespread applications, the adversarial robustness of such models remains underexplored. In principle, adversarial attacks should induce minimal shift of the instance distribution and the target labels. However, in NSGM, a single edge insertion or deletion can flip relevance labels. Detecting or controlling label flips requires repeated calls to potentially expensive combinatorial solvers. Responding to this challenge, we propose SEAGRAM, a solver-efficient adversarial attack method. It learns to detect non-critical node pairs whose perturbations will likely preserve relevance labels, thereby avoiding solver calls during test time. We further employ an active learning strategy, which reduces the solver calls also during training. Experiments show that SEAGRAM significantly worsens the performance of existing models.


As the Story Unfolds: Watching a Film and Identifying Characters as a Human Does

Zhongrui Gui ⋅ Junyu Xie ⋅ Tengda Han ⋅ Weidi Xie ⋅ Andrew Zisserman

Films introduce characters incrementally as they progress. We study the problem of identifying principal characters and assigning their names directly from the movie itself, without using external cast lists, actor photographs, or pre-built character banks. We introduce the character naming task: given a name mentioned in dialogue and the associated video context, the model must determine which visible character bears that name. This task is challenging because name resolution often depends on cinematic cues such as shot–reverse-shot structure, self-introduction, direct address, third-party reference, gaze, and temporal continuity. We provide a dataset and annotations for evaluating character naming and first character appearances, develop a video-and-dialogue model for the task, and compare it with proprietary multimodal large language models. We further introduce a progressive framework for building a character bank as the film unfolds. Together, these contributions support online, human-like movie understanding in which character identities are inferred from the film itself.


ASTRA: ADMM-Accelerated Topology Reconfiguration for Dynamic Satellite Constellations

João Norberto ⋅ Ricardo N. Ferreira ⋅ Claudia Soares

Dynamic topology reconfiguration is central to the reliability and efficiency of large satellite constellations, yet many existing approaches rely on idealized assumptions such as full constellation deployment or uniform orbital spacing. We present Adaptive Satellite Topology via Regret-Aware learning (ASTRA), a theoretically-grounded framework for dynamic satellite topology reconfiguration that builds on an online learning formulation and makes it computationally practical. ASTRA combines an ADMM-based offline solver with efficient online updates for both online gradient descent and online conditional gradient, yielding markedly cheaper constrained updates than generic optimization pipelines. On the theory side, we show that for a relevant class of entry-wise nonzero utility matrices, the objective is strongly convex, which yields logarithmic static regret for online gradient descent, and we further instantiate known dynamic-regret guarantees under inexact ADMM inner loops. Empirically, ASTRA matches or improves topology quality, presenting a good trade-off with computational time on synthetic constellations, and it remains effective on real Starlink data under partial deployment and non-uniform spacing, where idealized structural assumptions break down. These results position ASTRA as an efficient and theoretically grounded approach to topology reconfiguration in realistic Low Earth Orbit networks.

Addressing codebook collapse in vector quantization models is crucial for efficient discrete representation learning. Recent solutions increasingly rely on expressive codebook reparameterizations, improving adaptability at the cost of additional optimization and computational overhead. This motivates a simple question: what reparameterization scheme is just enough for efficient codebook learning? In this work, we propose Adaptive-Scale Vector Quantization (ASVQ), a lightweight reparameterization method for learning the codebook in a decoupled way. Experiments across image and audio tokenization tasks demonstrate that ASVQ consistently improves codebook utilization, reconstruction quality, and training stability, while remaining competitive with more complex reparameterization-based methods. These results suggest that ASVQ provides a favorable trade-off between codebook expressiveness and efficiency.


Asymmetric Scaling Laws from Sparse Features

John Sous ⋅ Michael Winer

We introduce a model for neural scaling laws under sparse activations. In the model, test loss is often dominated by rare coordinates that are never observed in the training input. This mechanism induces a novel bottleneck absent from dense models. We derive the asymptotic population loss in both the underparameterized and overparameterized regimes, and show that the loss exhibits a double-descent peak near the interpolation threshold---where the number of parameters is just sufficient to fit the training data---resulting in a loss curve governed by two distinct scaling exponents---one for the overparameterized regime and one for the underparameterized regime---with a gap determined by the degree of sparsity. Additionally, we derive a compute-optimal frontier that favors increasing dataset size over model capacity under fixed compute budgets. We also analyze gradient-descent dynamics and identify a scaling law for the probability that fixed-step gradient descent becomes unstable. We further show that the sparsity-induced effect persists under nonlinear activations. Experiments validating the theory can be found at https://anonymous.4open.science/r/sparse-scaling-neurips2026-3CDB/SparseScaling1.ipynb.

Best arm identification under differential privacy is a pure-exploration problem in which both statistical efficiency and privacy protection must be achieved simultaneously. We study fixed-budget best arm identification for bandits under pure $\epsilon$-differential privacy, where the learner must recommend an arm after a prescribed sampling budget while protecting the full transcript. We prove that the optimal exponential decay rate of the error probability is upper bounded by an instance-dependent privacy-aware transportation exponent that differs from the analogous quantity used to characterize the stopping time of in fixed-confidence analysis by Jourdan and Azize [2025]. Guided by this exponent, we propose AO-Pri-BAI, an adaptive algorithm that maintains private running estimates through Laplace-tree mechanisms and learns a sampling design through a min-max interaction between hard alternatives and arm allocations. We prove that AO-Pri-BAI satisfies pure $\epsilon$-differential privacy. We also establish that the exponent of the failure probability of AO-Pri-BAI matches the privacy-aware benchmark. Numerical studies show that even in the non-asymptotic setting, AO-Pri-BAI outperforms benchmark algorithms on various instances, complementing the theoretical analyses.

Backdoor attacks implant a trigger-target association into a model, causing malicious behavior at test time while largely preserving clean performance. Despite extensive empirical study, a unified explanation for why standard training dynamics learn backdoors so effectively remains largely missing. We argue that backdoor learning can be understood as a consequence of simplicity-biased optimization dynamics: trigger features are dynamically simpler than semantic features under stochastic gradient descent (SGD) because they induce more coherent gradient alignment, more stable gating behavior, and stronger early-stage amplification. To formalize this perspective, we introduce a directional data model that separates semantic and trigger features and study a one-hidden-layer ReLU network trained by SGD. Our analysis identifies a frozen-gate signal that governs group-wise early-stage loss decrease up to controlled gate-drift and gate-flip errors, yielding a quantitative explanation for faster poisoned-sample fitting and its dependence on the poisoning ratio. Beyond optimization dynamics, we show that this early-stage bias can induce trigger-dominant neurons with selective activation on poisoned inputs, explaining why vanilla clean fine-tuning can attenuate backdoor behavior without necessarily erasing the underlying trigger-related representation. Empirically, we validate the predicted signatures of early-stage optimization bias, representation-level trigger dominance, and post-training residual behavior using ResNet-18 on standard image classification benchmarks.

Self-play, a type of training algorithm that enables a model to self-improve, has recently shown promising empirical results in the context of formal theorem proving using Large Language Models (LLMs). (Dong & Ma, 2025) instantiate self-play with two cooperating agents: a prover, which proves theorems, and a conjecturer, which generates new theorems as a curriculum to the prover. In this paper, we provide a theoretical framework for understanding the self-improvement capabilities of self-play algorithms for theorem proving. First, we formalize the set of theorems as a graph, with nodes as theorems and edges between pairs of theorems with similar semantics. We introduce a set of primitive assumptions that characterize the guarantees of a trained prover and how a conjecturer can access the structure of the graph. Second, we show that if the underlying graph of theorems is well-connected, then a prover-conjecturer system, where the conjecturing algorithm is based on a reversible random walk, is sufficient to grow the set of proved theorems exponentially. Third, we formulate an issue of diversity encountered empirically by self-play algorithms in our framework, where the conjecturer tends to generate artificially complex and non-fundamental theorems. We propose a diversity measure for a training distribution of theorems generated by a conjecturer and an improved conjecturing algorithm that locally maximizes this diversity measure, by computing the diffusion similarity between neighboring theorems in the theorem graph. Finally, we describe a method to compute the diffusion similarity by first using contrastive learning to embed nodes into Euclidean space and second computing the inner-product between the embeddings.


A Theory of Training Profit-Optimal LLMs

Sophie Hao ⋅ Will Merrill

Scaling large language models (LLMs) requires tremendous computational resources, and recent advances in AI have gone hand in hand with massive amounts of capital expenditure. While it is established that scaling up LLMs reliably increases model quality (quantified in terms of loss or downstream evaluations), it is unclear how these quality improvements translate to potential revenue, and whether revenue increases would offset costs of larger-scale training and inference. In this work, we develop an economic model for characterizing the rational behavior of an LLM training firm by combining scaling laws with microeconomic theory. Under our model, LLM quality can be increased with more parameters and training tokens, leading to more potential adoption by consumers, who each have a quality threshold for using the LLM. On the other hand, additional parameters and training tokens both incur additional costs. We analyze the profit maximization problem for this model under compute-bound and data-bound regimes. In the compute-bound regime, optimal model size and token budget track hardware efficiency $E$ (FLOPs/dollar) at a near-linear rate; total training cost then scales sub-quadratically in $E$. Data efficiency improvements incentivize larger models and training expenditure. When we are limited to $D$ data, profit-optimal training expenditure scales as $D^2 / E$, i.e, increase with data and *decreases* with hardware efficiency (as well as data efficiency). Finally, we analyze practical trends in training expenditure: current trends in training expenditure are consistent with our most permissive model variants in the compute-bound regime, but are not profit-optimal in the data-bound regime or assuming hardware advances will stall. Overall, our results provide a theory of profit-optimal LLM training, providing a foundation for engaging critically with industry statements and supporting long-term economic decision making.


Atom-level Protein Representation Learning Improves Protein Structure Prediction

Taewon Kim ⋅ Hyosoon Jang ⋅ Hyunjin Seo ⋅ Seonghwan Seo ⋅ Hyeongwoo Kim ⋅ Wonho Zhung ⋅ Mingyeong Shin ⋅ Woo Youn Kim ⋅ Sungsoo Ahn

Recent advances in generative modeling show that pretrained representations can improve generation as conditioning features or alignment targets. Motivated by this, we study protein representations for predicting structures beyond conventional function annotation. We propose TRIPROREP, a structure-aware pretraining method that jointly models three aligned residue-level views: amino-acid identity, backbone geometry, and local full-atom geometry, discretely encoded via VQ-VAE tokenizers. By pretraining to recover original tokens from generator-corrupted views, TRIPROREP learns to distinguish plausible but incorrect cross-view augmentations from the original protein. We further introduce REPSP, a benchmark for evaluating protein representations in structure-predictive settings. REPSP tests three uses of representations: homodimer co-folding from apo-chain representations, residue-level prediction of homodimer-derived interaction properties, and representation-aligned monomer structure prediction. Across these tasks, TRIPROREP improves over sequence-only and prior structure-aware representation models, while maintaining competitive performance on conventional benchmarks.


AtomMOF: All-Atom Flow Matching for MOF-Adsorbate Structure Prediction

Nayoung Kim ⋅ Honghui Kim ⋅ Sihyun Yu ⋅ Minkyu Kim ⋅ Seongsu Kim ⋅ Sungsoo Ahn

Metal-organic frameworks (MOFs) are promising materials for applications such as direct air capture, where performance depends on how adsorbates bind within the framework. Accurately predicting MOF-adsorbate configurations is therefore important for screening candidate materials. However, traditional approaches are computationally expensive and require a known host structure, while existing generative models rely on rigid-body assumptions and do not explicitly model adsorbates. We introduce AtomMOF, a scalable all-atom flow-matching model that jointly predicts MOF-adsorbate structures from discrete building blocks and adsorbates. Built on a Diffusion Transformer with a building block-based pairwise attention bias, AtomMOF operates in an unconstrained all-atom space and exhibits clear scaling behavior. To improve the structural validity of flexible all-atom models, we also propose Feynman-Kac (FK) steering with machine-learned interatomic potentials (MLIPs). On the BW dataset, AtomMOF achieves a 58.17\% relative increase in match rate and a 31.84\% relative reduction in RMSD compared with prior work. MLIP-guided steering further improves validity by 28.7\% and reduces formation energy error by 86.5\%. On ODAC25, AtomMOF generates adsorption configurations faster than GCMC and, when combined with MLIP relaxation, identifies lower-energy configurations than those in the reference dataset.


Attention Heads are Complementary Visual Units: Mitigating Hallucinations in LVLMs via Adaptive Visual Cues Focusing

Zhenglin Hua ⋅ Yutong Xie ⋅ Yaxin Hou ⋅ Jiawei Tang ⋅ Jinghan He ⋅ Haiyun Guo ⋅ Junfeng Fang ⋅ Yuheng Jia

Recent studies have explored attention dynamics in Large Vision-Language Models (LVLMs). However, most existing approaches aggregate attention across heads, potentially obscuring head-specific visual information and weakening visual grounding, leading to hallucinations. In this work, we revisit hallucination from the perspective of how visual information is distributed and utilized across attention heads. We observe that a small subset of visual tokens accounts for most of the attention within each attention head, with these tokens—defined as head-wise visual cues—being complementary across heads, suggesting that effective grounding requires preserving head-specific support. We further find that the proportion of visual cues within visual attention declines during generation, leading to key visual information loss and hallucination. In addition, hallucinated tokens show weaker utilization of long-range textual context compared to correctly generated tokens. Building on these findings, we propose Visual Cues Reinforcement via Head-Adaptive De-redundancy and Distance-Aware Attenuation (VerDA), a training-free method to mitigate hallucinations. Specifically, we remove redundant visual tokens via head-adaptive de-redundancy, suppress local textual bias through distance-aware attenuation, and reinforce visual cues by redistributing attention. Experiments across multiple LVLMs and benchmarks demonstrate that VerDA consistently outperforms existing methods in mitigating hallucinations, with negligible inference overhead. Code is available in the supplementary materials.

We consider multi-objective reinforcement learning problems where objectives come from an identical family---such as the class of reachability objectives---and may appear or disappear at runtime. Our goal is to design adaptive policies that can efficiently adjust their behaviors as the set of active objectives changes. To solve this problem, we propose a modular framework where each objective is supported by a selfish local policy, and coordination is achieved through a novel auction-based mechanism: policies bid for the right to execute their actions, with bids reflecting the urgency of the current state. The highest bidder selects the action, enabling a dynamic and interpretable trade-off among objectives. Going back to the original adaptation problem, when objectives change, the system adapts by simply adding or removing the corresponding policies. Moreover, as objectives arise from the same family, identical copies of a parameterized policy can be deployed, facilitating immediate adaptation at runtime. We show how the selfish local policies can be computed by turning the problem into a general-sum Markov game, where the policies compete against each other to fulfill their own objectives. To succeed, each policy must not only optimize its own objective, but also reason about the presence of other goals and learn to produce calibrated bids that reflect relative priority. Under mild assumptions, we prove the existence of Nash equilibria where dishonest bidding leads to suboptimal outcome, and the most urgent objectives win control automatically. In our implementation, the policies are trained concurrently using proximal policy optimization (PPO). We evaluate on two Atari games and a gridworld-based path-planning task with dynamic targets. Our method achieves substantially better performance than monolithic policies trained with PPO.


Auditing Single-Query Recoverability in Self-Supervised Representations

zirui wang ⋅ Guangqiang He ⋅ PangWu ⋅ Peng Wang ⋅ Zhenfeng Li ⋅ Xianxiang Chen ⋅ Lidong Du ⋅ Zhen Fang

Modern transformation-aware self-supervised representations are routinely scored under protocols that grant the evaluator paired views, augmentation labels, or teacher anchors --- none of which the deployed predictor ever sees. We audit the gap that opens once these privileges are removed and only a single query remains. The finding is a systematic evaluation illusion: on 3DIEBench-Honest, honest single-query rotation errors are typically $2{-}3\times$ as large as native transformation scores, consistently across comparators and matched-utility thresholds. To make this gap measurable, we introduce a deployment-honest evaluation contract specifying a one-query test interface, matched semantic utility, a strongest-eligible-neighbor headline with tail summaries, and explicit privilege accounting. Theory and experiments are in service of the contract: lower-bound analysis grounds the contract at the interface level, while a positive control supplies a non-vacuous upper certificate under declared privileges. We exercise the contract on a generated law, on 3DIEBench-Honest, and on PTB-XL. The operational conclusion is that native transformation-aware success is not, on its own, a single-query deployment certificate.

Computing the unregularized Wasserstein barycenter for measure-valued data is a challenging optimization task. Recent algorithms have been tailored to either discrete measures as point clouds or continuous measures discretized on regular grids. In this work, we propose a primal mirror descent algorithm for computing the exact Wasserstein barycenter in the Fisher-Rao geometry. Our algorithm is a unified approach that is flexible enough to simultaneously cover discrete and absolutely continuous input measures, with convergence guarantees in both settings. In particular, when all input measures are discrete, our algorithm, initialized from any probability density, solves a sequence of semi-discrete optimal transport subproblems and produces absolutely continuous iterates that converge to the discrete barycenter. We use synthetic and real data examples to demonstrate the promising result in terms of accuracy and computational cost.


A Unified Audio Language Model with Text-Aligned Factorized Audio Tokenization

Dongchao Yang ⋅ Yuanyuan Wang ⋅ Songxiang Liu ⋅ Dading Chong ⋅ Xixin Wu ⋅ Helen Meng

We present UniALM, a unified audio foundation model built on FactorCodec, a purely discrete tokenizer that factorizes audio into two role-separated representations. Analysis tokens are optimized to retain language-aligned, text-expressible information that can be decoded by a text LLM head into grounded natural-language analyses, while reconstruction tokens preserve the information required for waveform synthesis and are used exclusively for audio decoding. This factorization yields understanding performance competitive with continuous Whisper features on audio understanding tasks, while reducing the perplexity of reconstruction-token modeling for generation. UniALM further adopts functional layer specialization, partitioning the backbone into audio-understanding, cross-modal, and audio-generation experts. We train the model with a four-stage recipe on 100B text tokens and 60B audio tokens, together with a composed audio sequence construction strategy for unified multi-task pre-training. At 3B parameters, UniALM is competitive with strong 7B unified baselines on in-domain tasks and exhibits non-trivial \emph{compositional generalization} to unseen tasks under few-shot and zero-shot evaluation. Unlike prior unified systems that rely on hybrid continuous-discrete inputs or do not support generation beyond speech, UniALM provides a purely discrete interface for both understanding and generation across speech, sound, and music.


A Unified Graph Language Model for Multi-Domain Multi-Task Graph Alignment Instruction Tuning

Haibo Chen ⋅ Xin Wang ⋅ Jiaheng Chao ⋅ Ling Feng ⋅ Wenwu Zhu

Leveraging Graph Neural Networks (GNNs) as graph encoders and aligning the resulting representations with Large Language Models (LLMs) through alignment instruction tuning has become a mainstream paradigm for constructing Graph Language Models (GLMs), combining the generalization ability of LLMs with the structural modeling capacity of GNNs. However, existing GLMs that adopt GNNs as graph encoders largely overlook the problem of aligning GNN-encoded representations across domains and tasks with the LLM token space to obtain unified graph tokens, thereby limiting their ability to generalize across diverse graph data. To bridge this gap, we aim to incorporate a multi-domain, multi-task GNN encoder into GLMs and align its representations with LLMs to enable multi-domain, multi-task graph alignment instruction tuning. This alignment problem remains underexplored and poses two key challenges: 1) learning GNN-encoded representations that are simultaneously generalizable across domains and tasks and well aligned with textual semantics is difficult, due to substantial variations in graph structures, feature distributions, and supervision signals, together with the lack of textual-semantic alignment guidance in task-specific GNN training; 2) diverse graph data and task-specific instructions can exhibit different degrees of compatibility with the LLM token space during instruction tuning, leading to varying alignment difficulty and rendering a fixed alignment strategy suboptimal. To tackle these challenges, we propose UniGraphLM, a Unified Graph Language Model that incorporates a multi-domain, multi-task GNN encoder to learn generalizable graph representations aligned with textual semantics, and then adaptively aligns these representations with the LLM. Specifically, we first develop a graph-text pair pretraining strategy with a tailored GNN encoder, trained on large-scale graph-text data spanning multiple domains and tasks to obtain generalizable representations naturally aligned with textual semantics. We further design a curriculum alignment tuning strategy that adaptively adjusts the alignment process by accounting for varying alignment difficulty across diverse graph data. Extensive experiments demonstrate that UniGraphLM consistently outperforms state-of-the-art baselines across graph datasets from different domains and tasks.

Graph domain adaptation (GDA) has emerged as an important problem in graph machine learning when the distribution of the source graph used for training differs from that of the target graph used for testing. While much of the prior work on GDA has focused on aligning node representations across source and target domains, recent studies show that such approaches can be suboptimal in the presence of graph structure shift, where the underlying connection patterns between nodes change across domains. In this work, we develop a unified pairwise distribution matching framework for mitigating conditional structure shift (CSS), a specific form of graph structure shift in which conditional edge distributions vary across domains. The framework recovers existing GDA methods as instances of moment matching and motivates PLSA, a new likelihood-based method that uses calibrated probabilistic predictors on connected target node pairs. Theoretically, we establish finite-sample guarantees for both the likelihood-based and moment-matching estimators under the contextual stochastic block model. Our analysis uses Bernstein-type concentration bounds for edge-weighted U-statistics, leading to error bounds that reflect both the effective number of observed edges and the conditioning of the corresponding likelihood or moment-matching problem. We complement our theoretical results with empirical studies that demonstrate the effectiveness of the proposed framework.


AutoformBot: Formalizing Mathematics at Scale

Ahmad Rammal ⋅ Niket Patel ⋅ Fabian Gloeckle ⋅ Amaury Hayat ⋅ Julia Kempe ⋅ Remi Munos ⋅ Charles Arnal ⋅ Vivien Cabannes

We present AutoformBot, a multi-agent system for building an Autoformalized Textbook Library At Scale (ATLAS) in Lean 4. AutoformBot orchestrates thousands of LLM agents, equipped with formal verification tools, dependency-aware task scheduling, and collaborative version control, to translate informal textbook prose into machine-checked definitions and proofs. We apply our methods to a corpus of over 20 open-access textbooks spanning analysis, algebra, topology, combinatorics, and probability, producing ATLAS: a verified library of over 50,000 Lean 4 declarations and 500 thousand lines of code. We release two artifacts: (i) AutoformBot, the open-source multi-agent framework; and (ii) ATLAS, the resulting formal library. Our results suggest that autoformalizing the core content of graduate-level mathematics at scale is now economically and technically feasible. This opens the door to the automated verification of both human- and machine-generated mathematics at a research level.


Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models

Jinbiao Chen ⋅ Shuang Jin ⋅ Guoyun Zhang ⋅ Junyu Zhang ⋅ Guanyi Wang ⋅ Hanzhang Qin

Robust optimization (RO) provides a principled framework for decision-making under uncertainty, but its practical use is often limited by the need to manually reformulate uncertain optimization models into tractable deterministic counterparts. Recent large language models (LLMs) have been shown promising for automating optimization formulation, yet RO reformulation remains challenging because it requires precise multi-step reasoning and mathematically consistent transformations. To facilitate systematic evaluation of LLM-based reformulation, for which no dedicated benchmark currently exists, we develop AutoRO-Bench, a benchmark featuring an automated data generation pipeline for the core RO reformulation task and a curated dataset for the RO application task. To address the reformulation challenge, we propose Automated Reformulation with Experience Memory (AutoREM), a tuning-free memory-augmented framework that autonomously builds a structured textual experience memory by reflecting on past failed trajectories through a tailored offline adaptation procedure. AutoREM requires neither domain-specific expert knowledge nor parameter updates, and the resulting memory readily transfers across different base LLMs. Experimental results show that AutoREM consistently improves the accuracy and efficiency of RO reformulation across in-distribution datasets, out-of-distribution datasets, and diverse base LLMs.


Autonomous Scientific Discovery via Iterative Meta-Reflection

Bingchen Zhao ⋅ Sara Beery ⋅ Oisin Mac Aodha

Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and validation. However, current systems operate within constrained search spaces or require predefined research questions, limiting their capacity for true open-ended inquiry. Furthermore, while they generate hypotheses iteratively, they largely lack the ability to explicitly synthesize their own accumulated findings to uncover complex, interconnected phenomena. We introduce DiscoPER, an autonomous large language model-powered framework that conducts open-ended research by dynamically generating and executing code to explore datasets without pre-specified research objectives. To ensure rigorous scientific validity, every proposed discovery must pass statistical testing. To overcome the limitations of isolated search, our framework introduces a second-order reasoning mechanism that periodically analyzes its own accumulated discoveries. By treating prior discoveries as empirical data, DiscoPER identifies structural patterns, confounds, and epistemic gaps, actively redirecting hypothesis exploration toward uncharted regions of the search space. The search space is further expanded by incorporating tool use, enabling the system to explore hypotheses beyond structured metadata by seamlessly processing and extracting useful information from multimodal sources like images. Evaluated on iNatDisco, a new multimodal ecological knowledge benchmark with pattern-level ground truth obtained from peer-reviewed literature, DiscoPER recovers 8 of 9 known patterns with a 72.7% hypothesis support rate, outperforming both classical causal discovery and LLM-guided baselines. Ablations show that DiscoPER scales with more data, and confirms the benefits of second-order ``meta-reflection''.


AuxGeoAgent: Synthesizing Challenging Geometry Proving Data via Planning and Symbolic Deduction

Junyi Liu ⋅ Jia Guo ⋅ Stanley Kok ⋅ Zihao Wang ⋅ zujie wen ⋅ Zhiqiang Zhang

Olympiad-level geometry reasoning remains challenging for large language models, primarily because existing datasets contain limited hard problems requiring long-horizon deduction and hidden auxiliary constructions. To address this gap, we present AuxGeoAgent, an agentic symbolic synthesis framework that formulates problem generation as planning over symbolic construction. Within this framework, an LLM-guided agent iteratively explores construction steps under symbolic constraints and mathematical priors, enabling the generation of structurally challenging problems that require auxiliary constructions. Leveraging this framework, we construct AuxGeo-18K, a compact yet challenging dataset of approximately 18K synthetic geometry proving problems, each accompanied by natural language and symbolic formulations, auxiliary constructions, and complete proof solutions. Unlike prior synthetic corpora that rely on massive data scaling, our approach focuses on the targeted synthesis of difficult instances while using orders of magnitude less data. Experiments on IMO-30 and HAGeo-409 show comparable performance to methods trained on million- and hundred-million-scale synthetic datasets, highlighting the strong data efficiency of our approach. Our work provides a valuable new resource for Olympiad-level geometric reasoning and advances research in mathematical reasoning.


AWP: Activation-based Window Pruning for Gigapixel Object Detection

Xiang Li ⋅ Wenxi Li ⋅ Yuetong Wang ⋅ Chen Zhang ⋅ Chenyang Lyu ⋅ Haozhe Lin ⋅ Fan Zhang ⋅ guiguang ding ⋅ Yuchen Guo

Gigapixel object detection in High-Resolution Wide (HRW) images faces extreme spatial sparsity where targets often occupy less than 5% of the image area while computation is wasted on vast uninformative regions. Existing token selection methods require either complex evolutionary search or training additional learnable modules. We present AWP (Activation-based Window Pruning), which discovers that the mean of Feed-Forward Network (FFN) activations naturally encodes window importance, enabling parameter-free selection without additional modules. Leveraging this intrinsic signal, AWP performs window pruning with remarkable effectiveness. On the PANDA gigapixel dataset, AWP outperforms the previous state-of-the-art by +1.6% AP₅₀ with 26.2% FLOPs reduction, particularly excelling on small objects (+5.1% APₛ). When applied to Swin Transformer, AWP demonstrates strong generalization, retaining 97.4% of baseline performance with 50% windows pruned. Remarkably, AWP enables zero-shot application by directly replacing learned selection modules, improving baseline by +0.6% AP₅₀ without retraining. Our work reveals that window-based vision transformers inherently encode spatial importance through FFN activations, offering a simple yet powerful alternative to existing selection mechanisms, especially for High-Resolution Wide images.


BabyTheorist: A Benchmark for Learning to Theorize the World from Observation Alone

Doojin Baek ⋅ Junyeob Baek ⋅ Mingyu Jo ⋅ Hosung Lee ⋅ Taegu Kang ⋅ Gyubin Lee ⋅ Sungjin Ahn

How would we know whether a machine truly understands the world? We argue that one key capability is \textit{observational theory learning}: acquiring reusable theories from pure observations and applying them to new situations. Current visual reasoning benchmarks focus primarily on the latter, presenting each task as a few-shot puzzle grouped by a common rule. We introduce BabyTheorist, a benchmark for controlled training and evaluation of observational theory learning. BabyTheorist generates continual streams of before-after visual observations from hidden compositional programs; during training, learners see only the resulting transitions, with no program labels, task identities, or support sets. To support theory learning and evaluation, BabyTheorist controls the program language, a curriculum over program complexity, recurring frequent subprograms, and a continual train-test protocol with level skipping. Seven representative baselines plateau well below ceiling, with the gap widening as the curriculum advances---evidence that BabyTheorist isolates a capability current program-learning approaches do not yet address.


Backdoor Attacks Rerouted: BatchNorm as a Sink for Adversarial Signals

Md Abdul Kadir ⋅ Tuan Tran Anh ⋅ Daniel Sonntag

Deep neural networks (DNNs), particularly CNN-based classification systems, are widely deployed due to their strong performance. However, they remain vulnerable to backdoor attacks, where imperceptible triggers can induce targeted misclassification while preserving high accuracy on clean inputs. These triggers may also distort model explanations. Such vulnerabilities raise serious concerns for both currently deployed and real-world applications, highlighting the need for deeper understanding and training-free defense mechanisms. In this study, we extensively investigate the attacking mechanisms in models with Batch Normalization (BN). We provide the first comprehensive theoretical analysis of the relationship between backdoor attacks and BN, showing that trigger-related information is strongly encoded in BN layers even under full model fine-tuning. We further prove that BN’s affine parameters and running statistics jointly influence both predictions and explanations, offering a unified explanation of backdoor behavior. Building on this insight, we introduce a simple training-free defense that re-estimates batch feature statistics and recomputes normalization at inference time, mitigating backdoor effects while preserving clean performance. Extensive experiments on 8 black-box and 3 explanation-aware attacks, compared against 9 defenses, demonstrate that our method reduces attack success rates from 100\% to 1\%, improves true-class recovery by 89\% (+23\% over prior work), and boosts explanation fidelity by up to 91\%, all without retraining or degrading accuracy. Code will be provided upon acceptance.

Real-world backdoor attacks often require poisoned datasets to be stored and transmitted before they are used to compromise deep learning systems. In the era of big data, however, the inevitable use of lossy compression poses a fundamental challenge to invisible backdoor attacks. We observe that triggers embedded in RGB images can become ineffective once the images are lossily compressed into binary bitstreams, such as JPEG files, for storage and transmission. Consequently, poisoned data may lose their malicious functionality after compression, causing backdoor injection to fail. Prior compression-based attacks typically exploit compression artifacts as certain triggers to distinguish poisoned RGB samples from uncompressed benign ones, rather than addressing whether malicious information can survive a shared lossy storage-and-transmission pipeline. In this paper, we highlight the necessity of explicitly accounting for lossy compression in backdoor attacks. This requires attackers to ensure that transmitted binary bitstreams preserve malicious trigger information, such that effective triggers can be induced after decompression. Building on the region-of-interest (ROI) coding mechanism in image compression, we propose two poisoning strategies tailored to inevitable lossy compression. First, we introduce \textbf{Universal Attack Reactivation}, a general method that uses sample-specific ROI masks to reactivate trigger information in bitstreams for learned image compression (LIC). Second, we present \textbf{Compression-Adapted Attack}, a new attack strategy that employs customized ROI masks to encode trigger information into bitstreams and applies to both traditional codecs and LIC. Extensive experiments demonstrate the effectiveness of both strategies.


Backdoor Channels Hidden in Latent Space: Cryptographic Undetectability in Modern Neural Networks

Marte Eggen ⋅ Eirik Reiestad ⋅ Kristian Gjøsteen ⋅ Inga Strümke

Recent cryptographic results establish that neural networks can be backdoored such that no efficient algorithm can distinguish them from a clean model. These guarantees, however, have been confined to stylised architectures of limited practical relevance, leaving open whether comparable undetectability extends to modern, end-to-end trained networks. We construct such an attack mechanism for state-of-the-art architectures, closely aligned to the cryptographic notion of undetectability, by identifying backdoor channels as learned latent directions, and show that the question of undetectability reduces to a hypothesis test between two unknown distributions over model parameters, which we conjecture to be intractable in practice. The consequence of this reframing is significant: if exploitable channels within a network's latent space are statistically indistinguishable from naturally learned directions, an attacker need not introduce foreign structure but can instead exploit the geometry the network already possesses. Demonstrating the approach on ResNet and Vision Transformer architectures trained on standard image classification datasets, the attack achieves both consistently high success rates with negligible clean accuracy degradation, and resists a comprehensive suite of post-training defences, none of which neutralise the backdoor without rendering the model unusable. Our results establish that cryptographic backdoors need not be artefacts requiring exotic architectures or artificial constructions, but identifiable as latent properties inherent to the geometry of learned representations.

Learning semantically meaningful human movement representations from 3D pose sequences is essential for human behavior modeling. Existing self-supervised methods derive supervision from artificial assumptions, e.g., invariance to augmentation within a certain range or recoverability of masked regions at a certain temporal scale, introducing unintuitive hyperparameters that make the learned semantics sensitive to their settings. We propose BAL, a self-supervised framework that instead exploits supervision intrinsic to human movement: the natural ordering structure of everyday behavior. BAL autoregressively predicts in a semantically abstract latent space, analogous to how humans anticipate others' actions rather than their precise joint configurations. Bidirectional autoregressive prediction, both forward and backward in time, fully leverages these ordering constraints while enabling representation extraction using both past and future context at inference. To suppress representation collapse, BAL decomposes each prediction target into a global component capturing sequence-level semantics and a local component capturing context-independent movement. We prove that the local targets retain nontrivial angular diversity unless the representations are completely collapsed, preserving a meaningful training signal. Experiments on large-scale datasets demonstrate that BAL's representations are highly effective for motion captioning, forecasting, and interpolation, outperforming existing self-supervised methods across all tasks.


Ballad: Bandit-Based LLM Routing for Automated Heuristic Discovery

Samidha Verma ⋅ Ankit Anand ⋅ Sayan Ranu

Modern automated heuristic discovery (AHD) systems increasingly rival hand-crafted heuristics by leveraging Large Language Models (LLMs) to iteratively mutate and refine solutions. A natural assumption is that more powerful and more expensive models consistently yield better heuristics. We challenge this empirically: across multiple NP-hard combinatorial optimization problems, we find no consistent advantage for larger, costlier models. This raises a deeper question: given a pool of LLMs spanning frontier and freely hosted or open-weight models, how should one adaptively select which model to invoke at each step of the search? The challenge is non-trivial. The optimal LLM may shift over the course of a run, and because one LLM's output seeds the context for subsequent calls, each model's reward signal is non-stationary and entangled across the heuristic population, thus violating the independence assumptions of classical online decision-making settings. We address this through Ballad (Bandit-based LLM Routing for Automated Heuristic Discovery), an online bandit algorithm that learns to orchestrate a pool of LLMs into an ensemble that matches or exceeds what any soloist LLM achieves alone. This orchestration is powered by an LLM selection policy that favors rare, high-scoring heuristics over stable averages, a DAG-based credit propagation mechanism through the genealogy of the heuristic population, and discounting of stale observations to remain attuned to each model's current utility. Across five NP-hard combinatorial optimization problems and a pool of four models of varying capability and cost, Ballad not only matches the best individual LLM in hindsight but, in several cases, surpasses every individual model while reducing API cost by up to 65.6% relative to always using the strongest model.


BarrierSteer: LLM Safety via Learning Barrier Steering

Thanh Q. Tran ⋅ Arun Verma ⋅ Kiwan Wong ⋅ Bryan Kian Hsiang Low ⋅ Daniela Rus ⋅ Wei Xiao

Despite the strong performance of large language models (LLMs) across diverse tasks, their susceptibility to adversarial attacks and unsafe content generation remains a significant barrier to deployment, particularly in high-stakes settings. Addressing this challenge requires safety mechanisms that are both practically effective and theoretically grounded. In this paper, we introduce BarrierSteer, a novel framework that improves response safety by embedding learned nonlinear safety constraints directly into the model's latent representation space. BarrierSteer treats hidden-state safety classifiers as Control Barrier Functions (CBFs), enabling constraint-guided steering of unsafe latent trajectories during generation. By composing multiple safety constraints through efficient constraint merging without modifying the underlying LLM parameters, BarrierSteer preserves model utility and performance. We provide theoretical results showing that applying CBFs in latent space yields a principled and computationally efficient approach for steering with respect to learned safety constraints, with guarantees conditional on the learned barriers capturing the intended safety property. Extensive experiments across multiple models and datasets demonstrate that BarrierSteer substantially reduces adversarial attack success rates and unsafe generations, outperforming existing methods.


BASIL-DCM: Biophysical Amortized Scalable Inference for Latent Dynamic Causal Modeling

Moein Khajehnejad ⋅ Forough Habibollahi ⋅ Leonardo Novelli ⋅ Adeel Razi

Estimating directed, weighted, and signed interactions among brain regions ($\textit{effective connectivity}$) from fMRI requires disentangling neural dynamics from delayed and nonlinear hemodynamic transformations. Dynamic Causal Modeling (DCM) provides a principled solution by inverting a biophysical generative model, but the standard Variational Laplace inversion method is computationally prohibitive for large parcellations and cohort-scale datasets. We introduce $\textbf{BASIL-DCM}$, a physics-informed amortized inference model that estimates subject-specific effective connectivity and biophysical DCM parameters, including ROI-wise hemodynamic transit time, spectral properties of endogenous neural fluctuations, and observation noise, $\textit{in a single forward pass}$. BASIL-DCM combines a linear-time state-space temporal encoder with an ROI-wise Transformer to capture long-range temporal dependencies and inter-regional interactions. The model is trained on data informed by Human Connectome Project resting-state fMRI, with effective connectivity initialized from regression DCM and complementary biophysical parameters sampled from physiologically plausible ranges. Learning is further constrained by a differentiable spectral-consistency objective derived from the DCM forward model. This approach enables fast, uncertainty-aware whole-brain network inference while preserving mechanistic interpretability and ensuring consistency with biophysical dynamics.


Batch-Conditioned Semantic Anchors for Robust Transductive Adaptation of Vision--Language Models

Mohammed Rahman Sherif Khan Mohammad ⋅ Ardhendu Behera ⋅ Sandip Pradhan ⋅ Swagat Kumar ⋅ Amr Ahmed

Transductive adaptation improves vision–language models at test time by refining predictions on an unlabeled batch, but current methods face a key trade-off. Aggressive approaches exploit batch structure but can drift when batches are sparse, have prior shift, or cover little of the label space. Conservative methods better preserve CLIP text semantics, but fixed semantic references can under-adapt even when the batch provides strong corrective evidence. We propose JPTA (Joint Prior and Transport Adaptation), a prior-aware framework that casts transductive VLM adaptation as batch-conditioned semantic reference estimation. Instead of treating CLIP text prototypes as fixed priors or freely replaceable, JPTA views them as semantic references that shift toward image-side batch structure only under calibrated evidence. It estimates batch priors and soft image-side prototypes to form transported semantic references whose impact is scaled by batch reliability. Across standard benchmarks, low effective-class regimes, all-class evaluation, and online streams, JPTA consistently outperforms TransCLIP and StatA, with the largest gains on sparse and prior-shifted batches. Anchor-drift, active-support, and absent-class-mass diagnostics show that JPTA fixes effective-class mismatch without unconstrained transductive drift, indicating that robust transductive VLM adaptation should recalibrate semantic references using reliable batch evidence rather than choose between unrestricted transduction and fixed prototypes.


Bayesian Decision Making around Experts

Daniel Jarne Ornia ⋅ Joel Dyer ⋅ Nicholas Bishop ⋅ Anisoara Calinescu ⋅ Michael Wooldridge

Complex learning agents are increasingly deployed alongside existing experts, such as human operators or previously trained agents. However, it remains unclear how learners should optimally incorporate certain forms of expert data, or how to quantify the potential benefit from doing so. We study this problem in the context of Bayesian multi-armed bandits, considering offline settings, where a learner receives a dataset of outcomes from an expert before interaction, and online settings, where posterior updates from different data sources may have different computational costs, so the learner must decide whether a given update should use its own experience or an outcome generated by an expert. We formalize how expert data influences the learner's posterior and quantify how pretraining or online learning on expert outcomes tighten information-theoretic regret bounds. We propose an information-directed rule for allocating a limited update budget across data sources, and we study strategies for how the learner can infer when to trust the expert, safeguarding for compromised experts. By disentangling and quantifying the value of expert data, our framework builds towards a practical, information-theoretic understanding of how agents should learn from others.


Bayes-Optimal BER and AUC: Estimation and Evaluation of Estimators

Ryota Ushio ⋅ Takashi Ishida ⋅ Masashi Sugiyama

A fundamental quantity in machine learning is the optimal performance achievable by any model on a given task. Estimating this quantity allows us to distinguish the irreducible part of the error from a deficiency of the model, telling us how much room for improvement remains. Recent work has shown that the Bayes error, or equivalently the optimal accuracy, can be estimated from soft labels in binary classification. However, accuracy is often a poor summary of performance in settings with severe class imbalance or noisy annotations, where metrics such as the balanced error rate (BER) and the area under the ROC curve (AUC) are more appropriate. We address this gap with two complementary contributions. (i) Estimation. We propose soft-label-based estimators for the optimal BER and AUC. We first consider the clean setting in which true soft labels and the class prior are known, and then extend the estimators to a more realistic setting in which the class prior is unknown and the observed soft labels are corrupted by an unknown order- preserving transformation. In the latter setting, we approximately recover the clean soft labels via isotonic regression with auxiliary hard labels, estimate the class prior with a clipped mean of the hard labels, and derive finite-sample error bounds for the resulting plug-in estimators. (ii) Evaluation. Since the optimum is unobservable on real datasets, evaluating any such estimator is itself nontrivial. We extend the FeeBee framework, originally proposed for evaluating Bayes-error estimators, to the optimal BER and AUC. The resulting procedure provides practical evaluation scores without requiring knowledge of the optimum, and applies to any estimator of the optimal BER or AUC, not only our proposed ones. Experiments on synthetic and real-world datasets validate both the estimators and the evaluation procedure.


BayesRAG: Probabilistic Mutual Evidence Corroboration for Multimodal Retrieval-Augmented Generation

Xuan Li ⋅ Yining Wang ⋅ Haocai Luo ⋅ Shengping Liu ⋅ Jiaen Liang ⋅ Ying Fu ⋅ Wei Huang ⋅ Jun Yu ⋅ Junnan Zhu

Retrieval-Augmented Generation (RAG) has become a pivotal paradigm for Large Language Models (LLMs), yet current approaches struggle with visually rich documents by treating text and images as isolated retrieval targets. Existing methods relying solely on cosine similarity often fail to capture the semantic reinforcement provided by cross-modal alignment and layout-induced coherence. To address these limitations, we propose BayesRAG, a novel multimodal retrieval framework grounded in Bayesian inference and Dempster-Shafer evidence theory. Unlike traditional approaches that rank candidates strictly by similarity, BayesRAG models the intrinsic consistency of retrieved candidates across modalities as probabilistic evidence to refine retrieval confidence. Specifically, our method computes the posterior association probability for combinations of multimodal retrieval results, prioritizing text-image pairs that mutually corroborate each other in terms of both semantics and layout. Extensive experiments demonstrate that BayesRAG significantly outperforms state-of-the-art (SOTA) methods on challenging multimodal benchmarks. This study establishes a new paradigm for multimodal retrieval fusion that effectively resolves the isolation of heterogeneous modalities through an evidence fusion mechanism and enhances the robustness of retrieval outcomes. Our code is available at https://anonymous.4open.science/r/BayesRAG-4C6B/.


BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

Yutaro Yamada ⋅ Kei Hiroshima ⋅ Nozomu Yoshinari ⋅ Kento Uchida ⋅ Shinichi Shirakawa

Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. Recent studies have therefore examined how to automatically derive optimization problems from natural-language descriptions, but existing benchmarks focus on settings where objectives and constraints can be written explicitly as mathematical expressions. Many practically important problems, including hyperparameter optimization in machine learning, are naturally treated as black-box optimization (BBO) problems, in which only objective values are observable, and the functional form is unavailable. In BBO, the search space design, a part of the problem formulation, and the selection of the optimization algorithm are crucial for problem-solving. Automating these processes with large language models (LLMs) is a significant challenge. This paper introduces Black-Box Optimization Word Problems (BBOWP), a novel problem setting in which a system must infer both a search space and an optimization algorithm from a natural-language description of a black-box optimization task. To support research on this setting, we establish the BBOWP Benchmark Suite (BBOWP-Bench), a dataset and evaluation framework for BBOWP. Each instance combines a natural-language problem description, an executable evaluation environment, and a human-designed baseline formulation, allowing evaluation of both search-space design and algorithm selection. The benchmark covers 23 instances from four different application domains and supports reproducible assessment through actual black-box optimization runs. Using this benchmark, we provide the first evaluation of LLMs and show that current LLMs are capable of selecting suitable algorithms based on the given evaluation budget. However, they sometimes struggle with search space design, particularly in identifying important variables and balancing their ranges when the problem description is less informative or the search space is highly problem-specific.


BDC-Merge: Cross-Architecture Model Merging via Dependency Alignment

Yukai Wang ⋅ Wenhao Zhao ⋅ Xingyu Zhu ⋅ Junfeng Fang ⋅ Ruipeng Wang ⋅ Houcheng Jiang ⋅ Xiang Wang

Large language models have achieved remarkable capabilities by scaling model capacity and training data, yet many practical deployments still rely on smaller models with limited resources whose capabilities lag behind their larger counterparts trained with richer resources. This gap calls for efficient knowledge transfer from source models with richer resources to compact target models. While model merging provides an effective mechanism, most existing methods are based on the assumption that the source and target models are architecturally compatible, making them inapplicable to heterogeneous source and target pairs. Although a recent method based on optimal transport extends model merging to settings with different architectures, it remains limited by linear correspondence modeling, iterative transport optimization, and reliance on supervised adaptation after fusion for further performance gains. To address these limitations, we propose BDC-Merge, a framework for model merging across different architectures based on Brownian Distance Correlation (BDCorr). BDC-Merge uses a small calibration set to estimate dependencies between heterogeneous activations at the feature level and the layer level, and directly lifts these dependencies from activation space into fusion operators in weight space, enabling effective parameter fusion across different architectures without any gradient based optimization or training after fusion. Extensive experiments across four low resource language and two specialized knowledge benchmarks show that BDC-Merge consistently outperforms the state of the art baseline for model merging across different architectures and largely preserves the target model’s general capabilities.

Constrained preference-conditioned MORL typically assumes full task observability, treating preferences and safety budgets as stationary task signals. In real-world deployment, however, these signals are often unreliable proxies: intent may drift, safety tolerances may shift with context, and exposed descriptors may be noisy, delayed, or incomplete. We therefore formulate constrained preference-conditioned MORL as a partially observed task-specification problem, where the agent must navigate a fundamental conflict between unreliable external signals and latent interaction dynamics. We propose \textbf{BECON}, a belief-conditioned framework for constrained MORL with drifting preferences and budgets, which infers latent task context from recent histories to estimate preferences and safety budgets. These estimates are fused with the exposed task signal via uncertainty-dependent gates, enabling the policy to defer to observations under high epistemic uncertainty while correcting them when history is informative. The fused context jointly conditions preference optimization and context-adaptive constraint enforcement. Experiments on diverse multi-objective control tasks demonstrate that BECON consistently improves over state-of-the-art MORL baselines, achieving stronger robustness and constraint satisfaction under task drift.


Behavioral Probes for Information Flow in LLM Swarms

Junhui Chen ⋅ Yu Zhang ⋅ Lantian Li ⋅ Dong Wang

Most existing approaches to LLM swarm optimization are utility-driven, focusing on final task utility while leaving the processing, integration, and reliance of information within individual interactions largely implicit. To address this, we study LLM swarm interactions from an information-flow-based view and introduce behavioral probes to quantify information flow patterns. Specifically, we repurpose External Context Score (ECS) and Parametric Knowledge Score (PKS)—originally designed for hallucination analysis—to measure a node’s external-context dependence and parametric-knowledge dependence, respectively. Across controlled and realistic swarm settings, we observe consistent patterns: ECS increases with input relevance and distinguishes useful multi-source contributions, while PKS decreases as informative context accumulates. Building on this heuristic interpretation, we demonstrate how these probes can complement utility-driven structure optimization in two controlled applications. Probe-guided node pruning significantly reduces anomalous nodes, and probe-guided reasoning-path search accelerates convergence compared to utility-only baselines. Our empirical results suggest that behavioral probes can make the internal information flow of LLM swarms observable and serve as effective auxiliary signals for structure optimization.


Benchmark Health Index: A Systematic Framework for Benchmarking the Benchmarks of LLMs

Longyuan Zhu ⋅ Hairan Hua ⋅ Linlin Miao ⋅ Wei Wang ⋅ Bing Zhao

Large Language Models (LLMs) are advancing rapidly, yet the benchmarks used to measure this progress are becoming increasingly unreliable. Score inflation and selective reporting have eroded the authority of standard benchmarks, leaving the community uncertain about which evaluation results remain trustworthy. We introduce the Benchmark Health Index (BHI), a pure data-driven framework for auditing evaluation sets along three orthogonal and complementary axes: (1) Capability Discrimination, measuring how sharply a benchmark separates model performance beyond noise; (2) Anti-Saturation, estimating remaining headroom before ceiling effects erode resolution and thus the benchmark's expected longevity; and (3) Impact, measuring capability-weighted benchmark adoption across the model ecosystem. By distilling 158 validated benchmarks from the technical reports of 117 representative models released between 2025 and April 2026, we systematically characterize the evaluation landscape. BHI is the first framework to quantify benchmark health at a macro level, providing a well-grounded quantitative basis for benchmark selection and supporting the dynamic filtering and continuous updating of evaluation sets. Our comprehensive analysis not only identifies benchmarks that remain highly credible but also exposes pervasive structural issues. By establishing a rigorous foundation for selection and management, BHI elevates evaluation practice from heuristic-based judgment to a quantifiable and interpretable scientific process. Our code is available online at


Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

Dylan Feng ⋅ Pragya Srivastava ⋅ Anca Dragan ⋅ Cassidy Laidlaw

Many safety and alignment failures of large language models (LLMs) occur due to out-of-distribution (OOD) situations: unusual prompt or response patterns that are unforeseen by model developers. We systematically study whether LLM monitoring pipelines can detect these OOD alignment failures by introducing a benchmark called Misalignment Out of Distribution Benchmark (MOOD). It is difficult to find failures that are truly OOD for off-the-shelf models trained on vast safety datasets. We sidestep this by including a restricted training set in MOOD that we use to train our own monitors, as well as seven test sets with diverse alignment failures that are outside the training distribution. Using MOOD, we find that guard models (safety classifiers) often fail to generalize OOD. To fix this, we propose combining guard models with OOD detectors. We test four types of OOD detectors and find that a combination of a guard model and Mahalanobis distance-based OOD detector can improve recall from 39% to 47%. We also establish positive scaling trends across model scales for monitors that combine a guard model and OOD detector; we find that incorporating OOD detection into monitoring achieves a higher recall gain than using a guard model with 20 times more parameters. Our work suggests that OOD detection should be a crucial component of LLM monitoring and provides a foundation for further work on this important problem.

Structured generation remains a critical bottleneck for Multimodal Large Language Models (MLLMs). Current post-training paradigms struggle to resolve this: Supervised Fine-Tuning (SFT) hits a rigid ceiling, while scalar-reward Reinforcement Learning (e.g., DPO, GRPO, SimPO) suffers from a policy-level signal gap. Compressing long routing trajectories into sparse signals induces a short-completion bias and traps RL near the SFT baseline, whereas vanilla self-distillation actively regresses. To systematically quantify this joint visual-structural-relational reasoning, we introduce OracleGraph, a 10,000-page 2.5D historical document benchmark featuring a 12-dimensional diagnostic VQA suite and a structured-hallucination taxonomy. To overcome these optimization pathologies, we propose PRISM, a Policy-Reward Integrated Self-distillation framework. PRISM leverages multimodal information asymmetry via a Hindsight Teacher guided by a Diagnostic Graph Report (DGR). By synergizing Softmax Reward-Weighted Aggregation (preserving Graph-F1 rankings) with a Length-Normalized Auxiliary Preference Loss, PRISM explicitly neutralizes catastrophic truncation, yielding a $+3.81$pp Graph Reward gain and a $+2.3$pp schema-parsing improvement over SDPO ($p < 0.001$). While PRISM's nominal $+1.02$pp mean gain over SFT is not statistically significant ($p = 0.12$), it drastically stabilizes optimization, achieving $2.4\times$ lower cross-seed variance. Crucially, on a cross-book out-of-distribution (OOD) probe, PRISM's stability advantage amplifies to a $14\times$ variance reduction over SFT, confirming robust structural stabilization. Code and datasets are available at https://anonymous.4open.science/r/PRISM2026.


Benchmark Shadows: How Data Regimes Shape Parameter Footprints and Generalization

Hongjian Zou ⋅ Yidan Wang ⋅ dingqi ⋅ Yixuan Liao ⋅ xiaoxin chen

Large language models can achieve strong benchmark scores without proportional gains in broader capability, but diagnosing when this occurs remains difficult. We study this gap through benchmark shadows: support-concentrated data regimes that may concentrate learning around narrow evaluation-relevant patterns while limiting broader representational development. Under fixed architecture, tokenizer, optimizer family, and training budget, we compare a coverage-expanding baseline with two support-concentrated regimes: redundant repetition and frequency-concentrated rewriting. Spectral, rank-based, and layer-wise update diagnostics reveal distinct parameter-space signatures: repetition-concentrated training is largely recoverable after later diverse training, whereas frequency-concentrated support collapse leaves more persistent footprints. Correlational analyses of open-source multimodal model families reveal analogous structural patterns that co-occur with asymmetric benchmark profiles, while a prompt de-duplication case study shows that surface redundancy alone does not induce the same regime-level effects. Benchmark scores alone are therefore insufficient to characterize capability; parameter-space diagnostics provide complementary signals about training quality, coverage, and generalization-related regime effects.


Bentkus-type asymptotic e-values

Diego Martinez Taboada ⋅ Ben Chugg ⋅ Aaditya Ramdas

Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, where significance levels may be data-dependent. Existing asymptotic e-values, however, suffer from the ``missing factor,'' a scaling inefficiency resulting in overly conservative inference. Drawing on the framework of near-optimal concentration inequalities developed by Bentkus in the 2000s, we introduce Bentkus-type asymptotic e-values and prove that they successfully eliminate the missing factor. We also demonstrate both theoretically and empirically that Bentkus-type e-values consistently deliver sharper inference than existing alternatives, leading to tighter post-hoc confidence intervals and higher rejection rates in multiple testing procedures.


Bernoulli Flow Models: Self-Consistent Generative Modeling for Binary Data

Hao Mo ⋅ Liying Yang ⋅ Shumin Yao ⋅ Xinxing Yu ⋅ Ajian Liu ⋅ Xudong Mao ⋅ Yanyan Liang

Binary diffusion models typically require a massive number of function evaluations (NFEs) to generate high-quality samples, making practical inference computationally expensive. Reducing NFEs while preserving sample quality without relying on distillation or additional training remains a significant challenge. However, existing binary diffusion models sequentially define a discrete one-step forward path and subsequently derive the reverse posterior. In low-NFE scenarios that require cross-step sampling, these models incorrectly approximate the true multi-step likelihood using a single-step likelihood formulation, which severely degrades sample quality. To address this fundamental limitation and completely decouple the generative dynamics from fixed discrete time steps, we propose Bernoulli Flow Models (BFM). Rather than building upon sequential one-step Markov diffusion chains, we predefine a unified continuous global Bernoulli probability flow path between data distributions and pure noise, from which we derive analytical closed-form posterior transitions over arbitrary time intervals. Consequently, reducing the inference NFE in BFM is no longer an approximation of skipped discrete steps; it simply requires re-evaluating the analytical posterior on the new time intervals. This eliminates the structural training-inference mismatch inherent to discrete chains, yielding strictly self-consistent low-NFE sampling. Experimental results show that BFM is highly robust to aggressive NFE reduction. On the LSUN Churches 256x256 dataset, a 256-step-trained BFM achieves an FID of 9.31 when sampled with only 16 steps, whereas the state-of-the-art discrete baseline severely degrades to 204.10. BFM also remains competitive with both continuous and discrete generative baselines under standard full-step inference. Ultimately, these results establish BFM as a theoretically rigorous, self-consistent, and practically effective framework for fast binary data generation.

Text-to-image (T2I) models are increasingly used by online creative communities to generate imagery in the shared visual languages those communities have built, from Cottagecore's pastoral warmth to Cyberpunk's neon dystopias. The aesthetic predictors embedded in T2I pipelines filter training data and score generated outputs, yet whether different predictors agree on which visual languages are "high-quality" is largely untested. We present SubCulture-2.6K, a dataset of 2,610 Stable-Diffusion-generated image–prompt pairs spanning six subcultures (Cottagecore, Cyberpunk, Dark Academia, Goblincore, Synthwave, Y2K), released with dominant-color palettes, CLIP ViT-L/14 embeddings, and per-image scores from four deployed predictors (LAION-Aesthetics V2, PickScore, ImageReward, HPS-v2). Using this dataset we report four findings. First, subcultures are visually separable in the representations these predictors read from, so any cross-subcultural score difference reflects a value judgment, not a perceptual failure. Second, LAION-Aesthetics V2 imposes a large ordered preference across subcultures, with Y2K rendering retained at the standard quality cutoff at a small fraction of the rate of the most-favored subculture. Third, peer predictors share LAION-Aesthetics V2's dispreference for Y2K but disagree sharply on which subculture they reward most, and they cluster much more closely with each other than with LAION-Aesthetics V2. Fourth, an interpretable-feature regression accounts for roughly half of the LAION-Aesthetics V2 effect, leaving a sizeable residual subcultural taste. Because LAION-Aesthetics V2 sits both upstream and downstream of recent Stable Diffusion variants, this audit shows that aesthetic predictors are not interchangeable: the choice of which one to use is itself a choice about whose visual language is rewarded.


Beyond Bit Matching: Orthogonal Watermarks for Collusion-Resistant Image Fingerprinting

Zhakshylyk Nurlanov ⋅ Tobias Weißberg ⋅ Florian Bernard

Image fingerprinting assigns each distributed copy a user-specific watermark for source tracing. Averaging collusion is a central threat: several recipients can average their differently marked copies to suppress each individual mark. Discrete bit-string watermarks are especially vulnerable because averaging drives bit evidence toward ambiguous decisions. We introduce **OrthoMark**, which maps user keys to near-orthogonal high-dimensional unit vectors and identifies colluders by cosine similarity to the extracted watermark direction. Under ideal averaging of $K$ vector keys, the normalized average preserves a cosine signal of order $1/\sqrt{K}$ for each colluder, while unrelated keys remain concentrated near zero. The spherical geometry of random directions gives a shifted-Beta null distribution for cosine similarities, enabling analytic per-key false-positive-rate control after validation on unwatermarked images. OrthoMark uses a neural watermark encoder and centered extractor trained with JND masking and a progressive curriculum for photometric, geometric, and collusion robustness. Experiments on MS-COCO, OpenImages, and AI-generated images show that OrthoMark maintains robust single-key detection under photometric and geometric distortions, and achieves the best strict all-colluder detection among evaluated methods, while preserving high visual quality.


Beyond Exemplar Selection: Value-Aware Memory Allocation in Replay-Based Continual Learning

Yuhang Li ⋅ Guoxu Zhou ⋅ Zhenhao Huang ⋅ Yuning Qiu ⋅ Qibin Zhao

Replay is a standard strategy for mitigating catastrophic forgetting in continual learning, where a bounded memory buffer stores exemplars from previously seen classes. Existing replay methods have largely focused on instance-level decisions, such as which samples to store, replay, or replace, while the class-wise allocation of memory slots is often treated as a uniform heuristic. Here, we study class-wise memory allocation as a complementary design axis and propose VAMA, a value-aware memory allocation framework. VAMA keeps cross-task budgets balanced, estimates task-local class values from Shapley-inspired sample utility scores, and converts them into class-wise quotas through a regularized allocation rule that conservatively deviates from class-balanced replay. This design changes only the class-wise budgeting rule and leaves within-class exemplar selection unchanged, making VAMA easy to integrate into replay pipelines. Experiments on CIFAR-100 and ImageNet-100 show that VAMA consistently improves ER, iCaRL, and FOSTER across different memory budgets and task protocols. Further analyses show that the improvements are not due to non-uniform allocation alone, and that allocation-relevant class-value rankings stabilize early in training, enabling efficient value-aware allocation. These results suggest that class-wise memory allocation is a consequential design choice in replay-based continual learning.


Beyond Generation: Unlocking Discriminative Representations from Diffusion Models

Haowen Cui ⋅ Ge Wu ⋅ Shuo Chen ⋅ Yikai Ge ⋅ Ge Gao ⋅ Xiang Li ⋅ Jun Li ⋅ Jian Yang

Recent advancements in generative models have increasingly leveraged self-supervised visual representations to guide the synthesis process, significantly improving generation quality and semantic consistency. However, the reverse paradigm of harnessing the generative process to enhance visual representations remains largely underexplored. In this paper, we investigate the semantic properties of class tokens synthesized by generative models. We observe that generative models equipped with representation entanglement can generate class tokens that exhibit stronger discriminative capabilities than those extracted from the original pretrained visual models. Motivated by this observation, we propose a completely new Generation-to-Perception Knowledge Distillation (GPKD) framework, where our method generates class tokens as teacher signals to instill global semantics into the student model. To prevent the degradation of local details, we further incorporate a new masked patch-level distillation objective. This dual distillation strategy enhances global representations while mitigating the forgetting of local details, thereby producing more robust representations. Extensive experiments demonstrate that GPKD obtains consistent improvements across various downstream tasks compared with the base visual encoder, achieving a gain of more than 2\% in $k$-NN accuracy on ImageNet-1K.


Beyond Ground Truth: Evaluating Non-Verifiable Reasoning in LLMs through Moral Robustness

Elizaveta Tennant ⋅ Benjamin Henke ⋅ Anita Keshmirian ⋅ Murray Shanahan ⋅ Verena Rieser ⋅ Kristian Lum ⋅ Sydney Levine ⋅ Julia Haas

As LLMs increasingly serve in advisory and deliberative roles, users rely on them for non-verifiable reasoning in domains lacking objective ground truths. However, traditional evaluations of LLM reasoning focus almost exclusively on fact-based domains, such as mathematics and science, leaving uncertainty over whether and to what degree models can handle ambiguous, subjective, or value-laden problems over time. To address this concern, we propose moral reasoning as a paradigmatic subdomain of non-verifiable reasoning. We define moral robustness as a model’s capacity to exhibit sound moral reasoning across time and contexts, and we introduce a scalable, adversarial, multi-turn evaluation framework to empirically measure this capability. We simulate 48,000 user-agent moral deliberations across four frontier LLMs, varying premise relevance, premise order, conversation duration, and the user's stated moral view. We find that models successfully ignore morally-irrelevant distractors, but shift their reasoning by up to 6.5\%, on average, towards the user's stated preferred moral view, and varying their reasoning depending on factors such as order (altering moral judgments by order in 13-22\% of the cases) and duration (altering moral judgments between single-turn and multi-turn in 10-24\% of the cases). Our analysis indicates that models tailor not just their final verdicts but their underlying justifications to align with a user’s moral viewpoint — a failure mode we characterize as moral deliberative sycophancy.

We prove the identifiability of deep generative models (DGMs) with Piecewise-Affine (PWA) decoders and Gaussian Mixture Model (GMM) priors, in a purely unsupervised setting. We introduce three algebraic contrast principles for symmetry breaking: domain contrast, which trivializes the mixture symmetry group; mechanism contrast, which ensures every decoder branch is witnessed by a unique boundary; and interaction contrast, which forbids parameter conspiracies between latent components and decoder branches. Together they exploit the interplay between the discrete combinatorics of the PWA map and the continuous symmetry structure of the latent GMM. Continuity is replaced by algebraic symmetry conditions; injectivity is decoupled from structural identification and required only for pointwise inversion. Our results form a hierarchy: from law identifiability (LID; latent distribution up to a global affine map) through map identifiability (MID; decoder up to the same map) to posterior and pointwise identifiability. The ICA-form ambiguity emerges under diagonal covariance conditions; classical ICA is a further specialization under independence. Assumptions are only on the data-generating process, not on learning methods, except for the interaction contrast.

Explainable AI (XAI) offers far-reaching consequences for human-AI collaboration, but the fundamental question of how best to evaluate these systems has proven frustratingly elusive. Methodologies typically revolve around either proxy metrics which tell us little about the utility of an explanation, or human evaluation which---while generally considered the gold standard---is expensive, difficult to scale, and challenging to reproduce. In this paper, we investigate if Multimodal Large Language Models (MLLMs) can act as surrogate human participants to automate the evaluation of explainable AI. To do this, we first propose a ``purpose-grounded'' evaluation framework with specific metrics applicable to both MLLMs and human users alike. Then, we offer some conceptual insights as to why and when we can trust these models to automate this process. Lastly, we conduct extensive empirical testing across five purpose-grounded areas and nine popular MLLM APIs. Results indicate that all models trend positively, but OpenAI models (particularly GPT-5 mini) are the best aligned at replicating user responses, and for complex visual reasoning tasks more advanced modes (like GPT-5) perform best, which would be the recommended choices in practice. We expect our insights not to replace user testing, but rather to help XAI researchers complement their current evaluation methodologies and move the field away from heavy reliance on proxy metrics for measuring explanation utility.

Long-horizon agents increasingly operate across many steps, tools, and observations. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized. Drift can accumulate quietly: an agent may call the right tool with plausible arguments at every step, while its prefix moves toward a broader role, an adjacent objective, or evidence the user never supplied. Existing monitors mostly check local compliance, deliver final-trace verdicts, or score generic risk; they do not directly estimate this prefix-level relation. We introduce ontological trust, a task-conditioned property of trajectory prefixes, and instantiate it as RGE, an online monitor that decomposes trust along Role, Goal, and Evidence. RGE uses LLMs only to derive structured representations. State updates, projections, and intervention decisions are deterministic, so the output is a replayable and auditable trust trajectory rather than a single end-to-end judge verdict. We construct a cross-domain trajectory corpus from OSWorld, FinanceBench, and EICU-AC, covering benign executions, prefix-paired drift, and pseudo-consistency failures. On this corpus, RGE outperforms adapted rule-, judge-, and shield-style baselines on prefix-paired drift detection. With the two larger estimator models, it exceeds 93\% Drift F1 on every benchmark while keeping benign coverage at or above 95.8\%. Pseudo-consistency is harder: detection depends on whether task completion is externally visible, a structural limit we characterize empirically.

Deep learning models achieve high accuracy in signal processing but remain difficult to interpret. Classical matching pursuit and sparse coding offer better transparency but suffer from spectral leakage and grid mismatch when signal frequencies do not align with discrete dictionary atoms. We introduce Continuous Dictionary Pursuit (CDP), a framework for interpretable signal decomposition that learns the parameters of analytical functions directly from data. Unlike traditional methods, CDP optimizes over a continuous manifold of differentiable atoms to resolve the limitations of fixed grids. The algorithm employs a greedy iterative strategy using data-driven priors and gradient-based optimization to isolate individual signal components. We provide a library of parametric atoms that covers periodic, trend, and transient structures, including discontinuous functions modeled through spectral annealing. To ensure a sparse atom set for reconstruction, the framework utilizes a stopping criterion based on a predefined atom budget and residual error convergence. Evaluations on synthetic and real-world datasets show that CDP resolves spectral leakage for the parameterized atom families considered and recovers exact mathematical components. Beyond decomposition, we demonstrate that the extracted atoms can serve as structured curriculum targets for training sequence models, leading to consistent improvements in forecasting accuracy compared to standard end-to-end training. The method offers a glass-box alternative for signal decomposition by providing high reconstruction fidelity alongside direct physical interpretability.


Beyond Training Time, Test-Time Coordination is Essential for Cooperative MARL

Dongsu Lee ⋅ Sooraj Sathish ⋅ Joonkyung Kim ⋅ Woojun Kim ⋅ Amy Zhang

We argue that cooperative multi-agent reinforcement learning (MARL) should treat test-time coordination as part of the problem definition rather than a deployment artifact. Our focus is on long-horizon, centrally monitored cooperative systems such as warehouse robotics, automated storage-and-retrieval systems, and cloud-resource scheduling, where teams share a reward, operate under shifting objectives, and already run with global telemetry at deployment. The dominant paradigm in this regime, centralized training with decentralized execution (CTDE), uses global information only during training and provides no coordination channel at execution. This is sufficient when agents are weakly coupled or the environment is stationary, but it fails when the team must realign after conditions change. The limitation is structural rather than algorithmic, and more training cannot add a channel that does not exist at execution time. We call this the recoverability gap. To address it, we name the underexplored design space of intermittent test-time coordination as centralized training with mixed execution (CTME), in which a coordinator observes global state and emits high-level signals while agents act locally. CTME is not a particular learning algorithm, but a cooperative MARL problem setting that makes test-time information structure explicit. It specifies team-level information, high-level signals, and update frequency after training, with CTDE recovered as a limit case. We do not argue that CTDE is obsolete. We argue that removing the test-time channel is a modeling choice, not a necessity, and the wrong default when global telemetry is available and teams must recover from drift.

Existing underwater image enhancement (UIE) methods predominantly rely on unidirectional mapping, which lacks often results in over-enhancement or under-enhancement, without user-friendly controllability. To address this, we propose a paradigm of unsupervised trajectory learning for Omnidirectional Controllable UIE that treats restoration as a omnidirectional dynamical system, namely OC-UIE. The main purpose of our work is the realization of a consistent omnidirectional mapping across the representation space, which allows the model to master complex underwater dynamics beyond traditional unidirectional constraints. To implement this, we propose an Omnidirectional Training that optimizes over arbitrary source-target positions sampled from the data distribution. We decompose this purpose into two synergistic stages: first, Representation Trajectory Flattening is employed to organize intricate degradations into a unified adaptation axis; second, Trajectory Omnidirectional Integration is introduced to model the enhancement as a path integral over the resulting manifold. By optimizing only the relative shifts, OC-UIE isolates structural preservation from fluid degradation dynamics. This approach provides a controllable space for intensity-parameterized enhancement with state-of-the-art performance.


Bias-Variance Optimized Preference Optimization for Large Reasoning Models

Mingkang Zhu ⋅ Xi Chen ⋅ Bei Yu ⋅ Hengshuang Zhao ⋅ Jiaya Jia

Large reasoning models (LRMs) generate reasoning traces before producing final answers, yielding strong performance on multi-step reasoning tasks. However, preference alignment for LRMs remains underexplored. The ideal answer-level preference objective marginalizes over reasoning traces, but this trace-marginalized objective is computationally intractable. As a result, practical methods rely on single-trace surrogates, which can induce high gradient variance due to stochastic trace sampling. We cast this challenge as a gradient-estimator design problem and propose Bias-Variance Optimized Preference Optimization (BVPO), which combines a standard trace-based gradient estimator with an empty-trace gradient estimator from a separately sampled empty-trace branch. Under our fixed-prompt, fixed-answer conditional view, the empty-trace branch is deterministic with respect to the trace randomness studied in our analysis. We show that the composite estimator contracts the conditional reasoning-trace-noise component, admits a conditional MSE-optimal mixture, and appears in the estimator-dependent term of a standard biased-SGD bound. Empirically, BVPO improves alignment over baselines on alignment benchmarks, including AlpacaEval 2 and Arena-Hard. Although trained on general conversational data, BVPO also generally preserves reasoning performance on six math reasoning benchmarks. Empirical diagnostics further show that BVPO yields lower total gradient variance during training.


Bidirectional Information Flow (BIF) - A Sample Efficient Hierarchical Gaussian Process for Bayesian Optimization

Juan D. Guerra ⋅ Thomas Garbay ⋅ Numa Dancause ⋅ Guillaume Lajoie ⋅ Marco Bonizzato

Hierarchical Gaussian Process (H-GP) models divide problems into different subtasks, allowing for different components to address each part, making them well-suited for problems with inherent compositional structure. However, existing H-GP frameworks typically employ one-way information sharing — either top-down or bottom-up — which limits sample efficiency and slows convergence. We propose Bidirectional Information Flow (BIF), which establishes continuous two-way communication. BIF retains the modular structure of hierarchical models — the parent conditions its own posterior on child summaries, treating them as structured priors — while introducing top-down feedback to softly decompose environment observations from the parent into sub-responses. This mutual exchange improves sample efficiency, enables robust training, and allows modular reuse of learned subtask models. We prove analytically the regret of a GP with a learned kernel scales linearly with the mismatch to the true kernel, tightening in the hierarchical case to the sum of child-level errors. Ablation shows removing the downward pathway collapses child $R^2$ by up to 58\%. Across synthetic, neurostimulation, and HPO benchmarks, BIF achieves up to 4× higher parent $R^2$ and $\sim$100\% AUC improvement over vanilla GPBO, and outscores all hierarchical state-of-the-art methods on child $R^2$ given the correct acquisition function, while supporting modular child transfer to novel composite tasks.


Bidirectional Sparse Attention for Faster Video Diffusion Training

Chenlu Zhan ⋅ Wen Li ⋅ Jun Zhang ⋅ chuyu shen ⋅ Hao Zhang

Video diffusion Transformer (DiT) models excel in generative quality but hit major computational bottlenecks when producing high-resolution, long-duration videos. Full attention requires computing dot products between all pairs of queries and keys, resulting in a quadratic computational complexity with respect to the sequence length ((O(L^2))), leading to high training and inference costs.To overcome this limitation, we propose a Bidirectional Sparse Attention (BSA) framework that sparsification from both the query and key–value directions to reduce the quadratic cost of full attention.Specifically, the sparsification of queries is achieved by pruning tokens that are locally redundant or semantically similar within the 3D spatiotemporal domain of a video. For the key–value pairs, only those that are highly correlated with the query at a global level are selected for attention computation. Furthermore, we design a dynamic threshold adjustment mechanism to adaptively regulate the selected number of key–value pairs, thereby maintaining quality without degradation. Extensive experiments demonstrate that BSA significantly accelerates DiT training across long sequences, reducing FLOPs by up to 20× and achieving 17.79× faster attention training, while preserving or even surpassing the generative quality of full attention.


Bigger Isn’t Better: Why the Indiscriminate Scaling of Foundation Models Can’t Solve Biology

Kathryne Metcalf ⋅ Lorin Crawford ⋅ Mary L Gray ⋅ Kevin K Yang ⋅ Alex X Lu

In the wake of AlphaFold’s spectacular achievements, a new generation of biological foundation models have promised to bring about similarly transformative advances in other life science domains. However, we argue that the availability of prefabricated datasets, architectures, and benchmarking metrics—coupled with the pressure to publish novel results—has led computational biology to prioritize scale, visibility, and convenience over progress. Without confronting the inbuilt limitations of our existing data and components, and without realigning the institutional and professional incentives that drive research decision-making at the programmatic level, we risk misallocating effort at great cost to our ability to achieve meaningful biological insight. This paper focuses on foundation modelling in proteomics, genomics, and single-cell biology, offering a critical synthesis of common failure modes. We highlight empirical challenges that have been brought against a number of high-profile models as case studies, and—using AlphaFold as a countercase—explore what it would take to build a research ecosystem that would facilitate the development of theory-driven, fit-for-purpose models.


Binding Visual Features Point by Point

Udith Haputhanthri ⋅ Declan Campbell ⋅ Rim Assouel ⋅ Jonathan D Cohen ⋅ Taylor Webb

Despite success on standard benchmarks, vision language models display persistent failures on tasks involving processing of multi-object scenes, including many tasks that are relatively easy for humans. Recent work has found that these failures may stem from a basic inability to accurately bind object features in-context, a challenge that is referred to as the ‘binding problem' in cognitive science and neuroscience. The human visual system is thought to solve this binding problem via serial processing, attending to individual objects one at a time so as to avoid interference from other objects. Recent work has proposed `pointing' -- the use of explicit spatial coordinates to refer to objects -- as an analogous solution for vision language models, and found that it improves performance on challenging multi-object tasks. However, it is unclear \textit{why} (i.e., on a mechanistic or representational level) this approach improves performance, and how directly this relates to serial processing in human vision. Here, we investigate this question. We find that learning to point-via-text induces an internal visual search routine, and we characterize the mechanisms that support this procedure. We also find that pointing behavior can be generalized to new tasks via fine-tuning, and that doing so eliminates binding errors and enables compositional generalization. These results provide a proof-of-principle that serial processing can solve the binding problem for vision language models just as it does for biological vision.


BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks

Loka Li ⋅ Duzhen Zhang ⋅ Xingbo Du ⋅ Leonard Song ⋅ Zixiao Wang ⋅ Assanali Aukenov ⋅ Noel Thomas ⋅ Shakhnazar Sailaukan ⋅ Yonghan Yang ⋅ Feilong Chen ⋅ Jiahua Dong ⋅ Kun Zhang ⋅ Bin Zhang ⋅ Le Song

Large language model (LLM) agents can now automate parts of machine-learning model building, but biomedical benchmarks still either emphasize question answering, reasoning, and tool use, or cover only narrow slices of biomedical ML coding. We introduce BioXArena, a biomedical machine learning (BioML) coding benchmark that evaluates whether agents can create task-specific model-building code for heterogeneous, often multi-modal biomedical datasets. It contains 76 end-to-end tasks across 9 domains: sequence, single-cell, structure, network biology, chemical biology, perturbation dynamics, phenotype--disease, imaging, and text-integrated tasks. Each task is curated from primary sources into a unified public capsule with hidden labels, held-out graders, and biology-aware metrics on a common 0-to-1 scale; agents must write runnable code, train models, and submit predictions for private test samples. BioXArena emphasizes realistic data interfaces: most tasks combine multiple input sources, and more than half are multi-modal, spanning tables, images, text, molecular sequences, omics matrices, and protein structures. We evaluate 11 agent configurations, including general coding LLMs, biomedical agents, and ML coding agents, in a shared 2-hour, single-GPU sandbox. MLEvolve with Gemini-3.1-Pro obtains the highest average score of 0.666, followed by GPT-5.4 with an average score of 0.636; no agent dominates across all domains. Beyond the main leaderboard, we conduct extensive ablation studies, robustness checks, scaling analyses, cost analyses, and failure-mode analyses to characterize how backbones, scaffolds, budgets, and domains affect BioML coding performance. We will release all tasks, graders, runner scripts, leaderboard results, and agent traces.


Boosting LLM Reasoning via Human-Inspired Reward Shaping

Wenze Lin ⋅ Zhen Yang ⋅ Xitai Jiang ⋅ Xiaoteng Ma ⋅ Gao Huang

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for enhancing reasoning in Large Language Models (LLMs). However, existing reward formulations typically treat exploration and consolidation as a monolithic process, resulting in entangled stage-wise learning dynamics. This contradicts the natural learning behavior of human learners. In human learning, individuals adopt distinct behavioral patterns toward mastered versus unfamiliar problems. When confronting unmastered challenges, humans prioritize broad exploration to seek viable solutions. By contrast, for well-mastered problems, they focus instead on reasoning condensation and knowledge abstraction to distill concise underlying principles. Motivated by this gap, we introduce T2T(Thickening-to-Thinning), a dynamic reward framework inspired by human learning processes. Specifically, it implements a dual-phase mechanism: (1) On incorrect attempts, T2T incentivizes "thickening" to broaden the search space and explore novel solution paths; (2) Upon achieving correctness, it shifts to "thinning", imposing length penalties to discourage redundancy, thereby fostering model confidence and crystallizing reasoning capabilities. Extensive experiments on mathematical benchmarks (MATH-500, AIME, AMC) across 5 mainstream LLMs demonstrate that T2T significantly outperforms standard GRPO and recent baselines, achieving superior performance.


Brain2voice 2.0: High-performance voice synthesis brain-computer interface

Maitreyee Wairagkar ⋅ Aparna Srinivasan ⋅ Nicholas S Card ⋅ Tyler Singer-Clark ⋅ Xianda Hou ⋅ Carrina Iacobacci ⋅ Lee Miller ⋅ Leigh Hochberg ⋅ David Brandman ⋅ Sergey D Stavisky

Brain-computer interfaces (BCIs) offer a promising solution to speech loss due to neurological injury by decoding intended speech directly from brain activity. While recent BCIs have restored high-accuracy text-based communication, they fail to provide instantaneous voice output essential for the natural flow of conversation. Brain-to-voice BCIs address this gap by decoding voice directly from neural signals. However, even the state-of-the-art (SOTA) BCI-synthesized voice is not yet intelligible enough for real-world adoption. We introduce brain2voice 2.0, a new multimodal Transformer-based BCI decoder architecture capable of synthesizing intelligible voice from intracortical neural signals in real-time. Brain2voice 2.0 is trained on continuous and custom-tokenized acoustic targets and phoneme targets, leveraging their complementary speech information. We use self-supervised and adversarial training objectives that enhance acoustic feature quality and improve synthesis intelligibility. At each 10 ms timestep, the model causally outputs continuous and tokenized acoustic features for real-time voice synthesis as well as time-aligned phoneme predictions (raw phoneme error rate: 7\%, comparable to the latest brain-to-text models). We evaluated this new approach on the publicly available SOTA intracortical brain-to-voice benchmark dataset. Naïve human listeners transcribed brain2voice 2.0 synthesized voice with a word error rate of 5.24\%—an 8$\times$ improvement in intelligibility over previous results (43.75\%). Brain2voice 2.0 demonstrates that highly intelligible real-time voice synthesis from neural signals is achievable, for the first time crossing the intelligibility threshold necessary for clinically viable brain-to-voice BCIs for people with paralysis.


BrainWhisperer: Leveraging Whisper for Speech Decoding in Neuroprosthetics

Tommaso Boccato ⋅ Michał Olak ⋅ Moein Khajehnejad ⋅ Matteo Ferrante

Decoding continuous speech from intracortical recordings holds transformative potential for individuals with severe motor impairments, yet current systems face fundamental barriers to real-world deployment: training data is scarce, neural signals are non-stationary across sessions, and the external language models required by state-of-the-art cascaded decoders impose memory requirements that preclude local, privacy-preserving inference. We introduce BrainWhisperer, a neural speech decoder that adapts Whisper–pretrained on approximately 680,000 hours of speech–to microelectrode array recordings via a convolutional front-end, hierarchical low-rank projections for non-stationarity, windowed self-attention in phoneme-selective encoder layers, and a multi-task objective combining CTC and cross-entropy losses. Subject-specific embedders enable cross-participant training within a unified architecture. Evaluated on publicly available Utah array datasets from BrainGate participants, BrainWhisperer achieves a word error rate of 8.5\% in end-to-end decoding on the Card benchmark--to the best of our knowledge, the best result reported in this setting. Cross-dataset training improves performance without participant-specific fine-tuning, pointing toward scalable foundation models for speech brain-computer interfaces.


BRAVO: Bridge Matching for Autoregressive Video Generation

Shihua Zhang ⋅ Zhenxiong Tan ⋅ Xinchao Wang

Image-to-video generation aims to animate an image while preserving appearance details and producing visually coherent future frames. Autoregressive generation models this process as a sequence of next-frame predictions conditioned on the input image and text prompt. However, advanced autoregressive generators built on diffusion or flow models formulate each step as a noise-to-data process, despite the fact that the previous frame already provides an informative visual prior. This mismatch makes next-frame synthesis harder than necessary and can degrade prediction quality, since the model must recover scene layout and visual context from uninformative Gaussian noise at every step. To this end, we introduce BRAVO, a Brownian bridge matching framework for autoregressive video generation. BRAVO replaces the conventional noise-to-data path with a data-to-data bridge from the previous frame to the current one, turning next-frame synthesis into local frame-to-frame transport. The bridge state is formed by interpolating between adjacent real frames with a Brownian perturbation, so the learned velocity models the local transition from the previous frame to the next under the text prompt and causal visual history. BRAVO is trained with teacher-forcing and sampled autoregressively, where each bridge starts from the last generated frame and reuses historical messages through a causal KV-cache. Experiments on different datasets show that BRAVO surpasses the conventional noise-to-data baseline and achieves even competitive performance with substantially larger models.


Breaking Curse of Dimensionality for Mutual Information Estimation with Vine Copulas

Sigurd Holmsen ⋅ Berit Øksnes ⋅ Ingrid Hobæk Haff ⋅ Sylvia Richardson ⋅ Ali Ramezani-Kebrya

We propose a computationally efficient vine copula-based mutual information (MI) estimator. Unlike existing non-parametric density estimators that suffer from the curse of dimensionality, a non-parametric vine copula has a convergence rate that is independent of the dimensionality of data. We leverage this property to tackle the challenging task of MI estimation and propose an interpretable MI estimator. Extensive experiments on datasets with known ground-truth MI values across dimensions, data types, and (input, output) dependence structures demonstrate a superior trade-off between MI estimation error and computational time compared to SotA neural MI estimators.

General-purpose time series encoders such as \texttt{OTIS} and \texttt{MOMENT} promise a single backbone that transfers across domains, but they do not scale and rely on manually tagged domain labels at tokenisation. We trace both symptoms to a single pathology: under masked-data-modelling supervision the encoder's output patch tokens \textit{partially collapse}, with mean off-diagonal pair-wise cosine similarity stabilising at $0.68$ in \texttt{OTIS} and the encoder's effective rank capped accordingly. This is the natural consequence of supervising the encoder only in data space, which leaves no constraint against redundant token directions in the representation space. We introduce \texttt{OTISv2}, which suppresses the collapse by complementing the masked-reconstruction objective with two representation-space terms --- a self-distillation loss against an exponential-moving-average teacher and a KoLeo regulariser --- and replaces \texttt{OTIS}'s per-domain variate embeddings with register tokens to remove the label dependency. The resulting encoder pulls mean patch-to-patch cosine similarity from $0.63$ to $0.58$, scales monotonically from $7.1\,$M to $40\,$M parameters where its predecessor plateaus, and dominates general-purpose baselines on $143$ uni- and multi-variate datasets from the UEA and UCR archives from a single label-free checkpoint. We release code and pre-trained weights upon acceptance.


BrickFlow: Connectivity-Guided Brick Reconstruction

Peter Kulits ⋅ Cordelia Schmid

Given a set of LEGO parts and an image of an assembled object, our task is to infer the pose of each part to assemble the object. In contrast with traditional 3D reconstruction, doing so requires satisfying discrete physical constraints, which provide structure but render the problem combinatorial and highly sensitive to small pose errors. To learn a prior over how parts fit together, we train a set-conditioned SE(3) flow model that maps noisy poses to valid assemblies. We then finetune this model to inject image conditioning. However, while the learned flow captures global structure, it does not enforce connectivity. To address this, we introduce an analytic connector-field guidance to pull predictions toward the connection-consistent manifold at test time. Through this, we demonstrate improved validity and reconstruction performance, scaling with model size, pretraining, and conditional guidance. We will make our models and data available for research purposes.


Bridging Diffusion and Autoregression for Flexible Time Series Synthesis

Xin Wang ⋅ Xuan Zhang ⋅ Haipeng Zhang ⋅ Chunyu Wei ⋅ Yueguo Chen

Synthetic time series generation is critical for data augmentation, privacy-preserving sharing, and simulation. Autoregressive models extend to arbitrary horizons but suffer compounding errors, while diffusion models achieve high fidelity through bidirectional refinement only at fixed lengths. We propose **Horizon Diffusion**, a unified framework that resolves this tension by decomposing generation into variable-length *horizon blocks* produced autoregressively across blocks yet jointly denoised within each block. Two innovations make this hybrid practical: *Horizon-Causal Attention*, a structured mask that simultaneously preserves cross-block causality and within-block bidirectional refinement, and the *Block Size Curriculum*, a deterministic Warmup--Stable--Decay schedule that fluidly transitions training from full-sequence diffusion to block-wise autoregressive generation. Our method consistently outperforms strong baselines on fidelity, diversity, and downstream utility, and remains stable when sequence length scales by $4\times$ beyond training from a single trained model. Code is available at Supplementary Material.


Bridging Image Restoration and Recognition via Causal Mediated Unrolling

Haowei Li ⋅ Qing Zhao ⋅ Yanze Li ⋅ ZhiChen Chen ⋅ Pengxu Wei ⋅ Dongyu Zhang ⋅ Liang Lin

Robust visual recognition under degraded imaging conditions is critical for real-world vision systems.Task-oriented image restoration addresses this problem by enhancing degraded inputs before recognition, but it must preserve natural visual appearance without disrupting the evidence used by downstream recognizers.Existing optimization strategies struggle to satisfy these requirements simultaneously.Restoration driven mainly by visual quality may weaken task-critical cues or introduce texture shifts that are visually plausible but poorly matched to the recognizer.Directly tuning restoration networks with recognition loss is also insufficient, as the restoration model may exploit recognizer-specific shortcut cues and produce visually unnatural artifacts while improving task scores.We propose Causal Mediated Unrolling (CaMeU), an optimization-based framework for task-oriented image restoration with frozen recognition models.Starting from a joint optimization objective, the proposed framework derives a Half-Quadratic Splitting formulation that separates restoration into alternating task-driven and prior-guided updates.The task-driven update introduces an explicit mediator for front-door guided optimization, reducing shortcut bias induced by the recognizer.The prior-guided update performs manifold projection, pulling the intermediate result back toward the natural-image manifold to preserve visual quality.Through this mediated alternating process, the restored image is progressively optimized for both recognition performance and visual quality.Experiments on VOC Haze, VOC Dark, and degraded CUB-200-2011 show that CaMeU consistently improves downstream recognition while reducing visual artifacts.

Structured road understanding of lane geometry, topology, and traffic element relationships is foundational to safe autonomous driving. While vision-language models (VLMs) offer promising semantic flexibility, they lack the geometric and relational grounding required for precise road reasoning. Conversely, traditional modular systems, e.g., HD maps and topological road graphs, provide structural precision but remain semantically rigid. To bridge this gap, we introduce the Combined Road Substrate (CRS), a graph-grounded framework that makes geometric road structure and open-vocabulary semantics jointly executable in a single representation. CRS enables the automatic generation of compositionally complex and linguistically varied question-answer pairs via recursive graph queries, augmented with a ``grounding for free'' mechanism that ensures logical traceability to specific map elements, and procedurally extracted chain-of-thought supervision traces. We demonstrate that state-of-the-art VLMs - including large, closed-source models - struggle significantly with structured road reasoning, yet training a small 2- or 4-billion-parameter model with as few as 20 to 80 CRS-enriched scenes yields stable gains in compositional reasoning tasks of varying depth. Analysis of model behavior via verifiable reasoning traces reveals a systematic shift in failure modes: whereas baseline models fail at relational scene understanding, CRS-trained models reduce failures to attribute recognition, suggesting that the primary bottleneck in road understanding is not model scale, but the absence of structured supervision.


BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation

Haiquan Wen ⋅ Yiwei He ⋅ Zhenglin Huang ⋅ Tianxiao Li ⋅ Zihan Yu ⋅ Xingru Huang ⋅ Lu Qi ⋅ Baoyuan Wu ⋅ Xiangtai Li ⋅ Guangliang Cheng

As generative video models become increasingly realistic, detecting AI-generated videos requires systems that offer both accuracy and interpretability. However, applying Multimodal Large Language Models (MLLMs) to video forensics is currently limited by outdated datasets, simplistic evaluation protocols, and a reliance on black-box classification. To address these issues, we introduce a comprehensive dataset, benchmark, and baseline model for video forgery detection. First, we present \textbf{GenBuster-200K}, a fair dataset of over 200,000 high-quality videos sourced from state-of-the-art generators, featuring diverse real-world scenarios. Second, we propose \textbf{GenBuster-Bench}, a diagnostic benchmark spanning three progressive tracks (In-Domain, Out-of-Domain, and In-the-Wild) to evaluate models across \textit{domain shifts} and \textit{generational shifts}. It also introduces an MLLM-as-a-Judge protocol to assess the quality of the generated forensic explanations. Finally, we develop \textbf{BusterX}, an MLLM baseline with RL training. Instead of direct binary classification, BusterX formulates detection as a visual reasoning task, where the generated reasoning chain serves as detector itself. Experimental results demonstrate that BusterX outperforms several leading MLLMs (e.g., Qwen3.5, Claude-Sonnet-4.6) in both detection accuracy and rationale quality.

Flow and diffusion models achieve high-fidelity, high-resolution image synthesis, but often require many function evaluations (NFEs) at sampling time. Existing acceleration methods either require additional training through distillation or rely on training-free high-order solvers, and both can degrade sample quality at low NFE budgets. We propose CAB (Corrected Adams-Bashforth), a training-free sampler that accelerates both flow and diffusion models. CAB first transforms the sampling dynamics to a common rectified coordinate system, and then applies a multistep Adams--Bashforth predictor augmented with a simple correction term based on past velocity evaluations and therefore incurs no additional NFEs. The resulting method is simple, has the same algorithmic form across model classes, and has at least third-order local truncation error and second-order global error. Experiments on pretrained flow and diffusion models, including class-conditional and large-scale text-to-image benchmarks, show that CAB improves quality--NFE trade-offs in the low-step regime of (6)--(20) NFEs. It also remains competitive with strong training-free samplers at higher step counts across most tested models.


CA-Judge: Teach Large Models to Judge Anomalies via Comparison for Video Anomaly Detection

He Huang ⋅ Zixuan Hu ⋅ Dongxiao Li ⋅ Zhaoyi Li ⋅ LINGYU DUAN

Video anomaly detection (VAD) is critical for identifying safety risks in surveillance and autonomous systems. Recent large-model-based VAD methods generate textual rationales and directly output anomaly scores. However, mapping rich semantic descriptions and analysis directly to absolute scalar scores can be poorly grounded and inconsistent, leading to scores that are not well aligned with anomaly severity. With the observation that for humans, it is usually easier to compare things than to rate things directly, we explore whether it also holds for large models in VAD by proposing \textbf{CA-Judge} (\textbf{C}omparative \textbf{A}nomaly \textbf{Judge}), a training-free framework that replaces direct scoring with comparative order inference. CA-Judge treats model-perceived anomaly severity as a latent coordinate, constructs this coordinate from self-comparisons among generated normal/anomalous reference events, and models how anomaly likelihood varies along this coordinate using the generated normal/anomalous reference identities. At test time, CA-Judge avoids exhaustive comparison by maintaining a posterior belief over query severity and selecting a few query--reference comparisons by their expected information about the anomaly status. The resulting comparison trace is aggregated under a Bradley--Terry likelihood and read out through the learned severity--anomaly association to produce the anomaly probability. Without fine-tuning or training data, CA-Judge substantially improves over direct-scoring baselines and surpasses prior training-free state-of-the-art methods on UCF-Crime and XD-Violence.

Transformer-based scientific foundation models are increasingly deployed in high-stakes settings, but current architectures give deterministic outputs and provide limited support for calibrated predictive uncertainty. We propose Stochastic Attention, a sample average lightweight inference-time modification that randomizes attention by replacing softmax weights with normalized multinomial samples controlled by a single concentration parameter, and produces predictive ensembles without retraining. To set this parameter, we introduce a calibration objective that matches the stochastic attention output with the target, yielding an efficient univariate post-hoc tuning problem. We evaluate this mechanism on scientific foundation models for weather and time-series forecasting, as well as several regression tasks. Across benchmarks against uncertainty-aware baselines, we find that Sample Average Stochastic Attention achieves the strongest native calibration and the sharpest prediction intervals at comparable calibration, with adaptation costs nearly three orders of magnitude lower than the next-best baseline.


Calibration Is Not Control: Intervention Advantage for LLM-Agent Oversight

Chubin Zhang ⋅ Zhenglin Wan ⋅ Xingrui Yu ⋅ jingxuan wu ⋅ Qi Wen ⋅ Pengfei Zhou ⋅ Wangbo Zhao ⋅ Ivor Tsang

Runtime oversight for LLM agents is commonly framed as scalar risk prediction: estimate failure likelihood, confidence, or uncertainty, then intervene once the score crosses a threshold. We argue that this framing targets the wrong object for control. The relevant question is not how likely the agent is to fail if it continues, but whether an available intervention would improve the outcome. Two trajectory prefixes can have the same risk estimate while requiring different actions, because one remains recoverable and the other does not. We formalize this mismatch as target error and identify intervention advantage, the expected utility gain from intervening rather than continuing, as the decision object for oversight. To measure this mismatch, we introduce prefix branching, a same-prefix counterfactual protocol that executes candidate actions from identical trajectory states. Across four benchmarks, action-conditioned control yields regime-dependent gains over scalar routing. In a calibration decomposition, recalibrating the same scalar score improves prediction metrics but leaves control regret unchanged, showing that calibration alone does not repair target error. A simple prefix-only action-conditioned controller substantially reduces regret in the strongest interactive regime, from $0.506$ to $0.110$ on ALFWorld. Gains shrink when interventions are weak or when scalar routing already preserves intervention-relevant information. These results suggest that LLM-agent oversight should move from calibrated risk scoring toward action-conditioned value estimation.


CaMeLs Can Use Computers Too: System-level Security for Computer Use Agents

Hanna Foerster ⋅ Tom Blanchard ⋅ Kristina Nikolić ⋅ I Shumailov ⋅ Cheng Zhang ⋅ Robert Mullins ⋅ Nicolas Papernot ⋅ Florian Tramer ⋅ Yiren Zhao

AI agents are vulnerable to prompt injection attacks, where malicious content hijacks agent behavior. Among proposed defenses, architectural isolation provides the strongest guarantees by strictly separating trusted task planning from untrusted environment observations. However, applying this design to Computer Use Agents (CUAs), which automate tasks by viewing screens and executing actions, presents a fundamental challenge. Current agents require continuous observation of UI state to determine each action, which conflicts with the isolation required for security. We resolve this tension by demonstrating that UI workflows, while dynamic, are structurally predictable. Single-shot planning, where a trusted planner emits upfront a complete branching plan covering all anticipated runtime states, provides control flow integrity guarantees against arbitrary instruction injections. We introduce NOVA (Navigating via Observation, Verification, and Action) to make this viable in the combinatorially large UI state space, where the plan can invoke a perception model to resolve runtime values such as UI coordinates. We evaluate our design on OSWorld, and retain up to 57\% of the performance of frontier models while improving performance for smaller open-source models by up to 19\%, demonstrating that rigorous security and utility can coexist in CUAs. Although upfront planning prevents instruction injections, we show that additional measures are needed to defend against Branch Steering attacks, where adversaries deceive the perception model into routing execution down attacker-preferred branches of the plan, such as redirecting the agent to a malicious website.

Masked autoregressive models (MAR) have emerged as a powerful paradigm for image and video generation, combining the flexibility of masked modeling with the expressiveness of continuous tokenizers. However, when sampling individual frames, video MAR models often produce highly distorted outputs due to the lack of a structured global prior, especially when using only a few sampling steps. To address this, we propose CanvasMAR, a novel autoregressive video prediction model that predicts high-fidelity frames with few sampling steps by introducing canvas—one-step prediction of the future-frame expectation that serves as a non-uniform mask during masked generation. The canvas supplies global structure early in sampling, enabling faster and more coherent frame synthesis. To further stabilize autoregressive sampling, we propose an easy-to-hard curriculum via a motion-aware sampling order that synthesizes stationary regions before attending to highly dynamic ones. We also integrate compositional classifier-free guidance that jointly strengthens the canvas and temporal conditioning to improve generation fidelity. Experiments on the BAIR, UCF-101, and Kinetics-600 benchmarks demonstrate that CanvasMAR produces higher-quality videos with fewer autoregressive steps. On the challenging Kinetics-600 dataset, CanvasMAR achieves remarkable performance among autoregressive models and rivals advanced diffusion models.


Can VLMs Reason When to Stop for Human Safety?

Ryo Hachiuma ⋅ Arun G Zachariah ⋅ Barnaby Simkin ⋅ Jibin Varghese ⋅ Fu-En Yang ⋅ Chi-Pin Huang ⋅ Frank Wang ⋅ Yusuke Hirota

Since Vision-Language Models (VLMs) are the core of the perceptual and reasoning capability in Vision-Language-Action models, which are deployed in environments shared with humans, evaluating VLMs' safety reasoning from video is critical for preventing human harm. Current video-based benchmarks, however, evaluate only whether a model can recognize dangers that are already visible, and overlook two capabilities that VLMs require: reasoning under hierarchical principles in which human safety takes precedence over obedience to user instructions, and anticipating harm rather than identifying it after it has occurred. We introduce VESPER, a video benchmark that targets both capabilities by scenarios in which a robot is executing a user-assigned task when a person unexpectedly moves into harm's way, requiring the model to predict the potential harm and decide whether to continue or stop. The benchmark comprises 458 videos covering matched safe and dangerous variants, built through a human-in-the-loop pipeline from diverse taxonomies. We evaluate 21 VLMs, and the results indicate that the current VLMs fail to properly decide the actions under such a scenario.


Capacity-Constrained Online Convex Optimization with Delayed Feedback

Alexander Ryabchenko ⋅ Idan Attias ⋅ Dan Roy

Online learning with delayed feedback typically assumes that the learner can track all pending rounds until their feedback arrives. In practice, tracking resources are finite, and feedback from rounds that cannot be tracked is permanently lost. In this paper, we study delayed online convex optimization (OCO) under a hard capacity constraint, where at most $C$ pending rounds can be tracked at any time. To model delay information, we introduce a semi-clairvoyant model that refines the clairvoyant assumption from prior work: rather than requiring delays to be known at prediction time, the learner observes delay expirations online, consistent with the classical unconstrained delayed setting. Our approach proceeds via a reduction to a novel ``delayed and weighted" OCO problem, using a scheduler that randomizes which rounds are tracked and importance-weights the resulting observations. For the base problem, we propose and analyze Delayed-Weighted FTRL and its bandit analogue, obtaining regret bounds that characterize how weights interact with delayed feedback in both settings. Combining these bounds with our schedulers yields regret guarantees for capacity-constrained OCO under convex and strongly convex losses, for both first-order and bandit feedback. For first-order feedback, capacity $C = \Omega(\log T)$ suffices to recover the standard delayed OCO rates up to logarithmic factors. For bandit feedback, the standard delayed BCO rates are instead modulated by $(1 + \sigma_{\text{max}}/C)$ factors, where $\sigma_{\text{max}}$ is the maximum number of pending observations. This allows the regret bound to degrade gracefully when $C < \sigma_{\text{max}}$, while remaining sublinear.


CasePlay: Self-Play Reinforcement Learning from Case Reports for Medical Reasoning

Hao Wang ⋅ Zihan Wang ⋅ Kai WU ⋅ Jian Yao ⋅ Yiwen Ye ⋅ Yongcan Luo ⋅ Dapeng Wu ⋅ Jiale Chen ⋅ Haihua Yang

Reinforcement learning (RL) is a promising way to improve the clinical reasoning abilities of large language models (LLMs), but current pipelines often rely on expert-labeled questions or costly synthetic-data generation. Meanwhile, the medical literature offers a wealth of high-quality unlabeled documents, including clinical case reports that trace the full clinical course from initial presentation to final treatment decisions. We study whether document-grounded self-play can convert these reports into effective RL training signals. We identify two challenges that limit existing self-play methods in this setting: knowledge-point fixation, where the proposer repeatedly asks about a narrow set of salient clues in a source report, and difficulty-quality mismatch, where appropriately difficult questions may still be poorly formed or shallow, limiting their value for reasoner training. To address these challenges, we propose CasePlay, a self-play framework with two key components: (a) a Knowledge-Conditioned Proposer, which anchors question generation to case-specific knowledge points so that questions draw on a broader range of evidence from the source report; and (b) a Rubric-Judge Reasoner, which augments answer-accuracy feedback with rubric-based quality feedback on clue integration, reasoning depth, and option quality, guiding the proposer toward reasoning-intensive questions that better support reasoner training. Extensive experiments across seven medically related benchmarks demonstrate that CasePlay outperforms existing self-play baselines and improves the SFT warm-start model by 3.0 points on average, highlighted by a 6.9-point gain on LiveClin. This approach offers an effective path toward continually improving medical reasoning ability through self-play RL. Code is available at \url{https://anonymous.4open.science/r/CasePlay}.


Cauchy Scientific Networks: Loss–Architecture Alignment and Its Limit

Haonan Tan ⋅ Xin Li ⋅ juyi peng ⋅ Zhihong Xia ⋅ Yuxing Han ⋅ Gene Wen

Cauchy activations and Cauchy-adaptive layers can give large gains on scientific learning tasks, but the gains are selective and not uniquely Cauchy-specific. We organize this selectivity with an empirical loss-architecture alignment taxonomy: an architecture can help when the objective exposes structures it can use, and can lose under a mismatched objective. In score learning, Cauchy score networks improve over SiLU under a Fokker-Planck PDE loss, but the same advantage disappears under denoising score matching and a Gaussian activation is comparable to Cauchy under the PDE loss. In loss-activation factorials, robust residual losses explain more of the heavy-tailed and spectral gains than the activation alone; an Allen-Cahn trap initially escaped by CauchyAct+Cauchy loss is also escaped, sometimes more strongly, by CauchyAct+L1, CauchyAct+Huber, and ReLU+Cauchy loss. We also analyze ResidualCAN, a CauchyAct main path plus a softmax-weighted Cauchy Adaptive Node residual. It achieves 1520x lower stationary Allen-Cahn error than a SiLU MLP at epsilon = 1.0 and remains strongest or second strongest against KAN-style, tuned SIREN/Fourier, and adaptive PINN baselines across an Allen-Cahn epsilon-suite. Ablations show that most of this gain comes from a dense local-basis residual rather than from Cauchy activation alone. A Poisson/Burgers boundary check prevents a broad PDE dominance claim: tuned SIREN/Fourier and KAN-style baselines can dominate away from transition-layer structure. The resulting claim is deliberately bounded: Cauchy components are one useful instance of robust/local-basis co-design in low-dimensional scientific tasks, not a universal activation or a uniquely rational mechanism.


Causal Discovery Under Hard Selection Bias: A New Robust Score-Matching Approach

Yiwen Qiu ⋅ Francesco Montagna ⋅ Shimeng Huang ⋅ Francesco Locatello

Selection bias is a difficult yet widespread problem in causal discovery, occurring whenever non-random selection processes lead to data that is not representative of the underlying populations. Due to its practical importance, related works have been proposed to address the problem under local [Versteeg et al., 2022] or interventional settings [Dai et al., 2025a], but general algorithms for observational data remain elusive. In this paper, we focus on truncated data (hard selection), and provide a general result for Additive Noise Models (ANM). We show that score-based methods are a surprisingly suitable candidate due to a special property of the score function ($\nabla \log p(x)$): *deterministic truncation preserves both the score* of the joint distribution as well as the Jacobian of the score. While established score-matching algorithms still fail in most settings, we actively leverage this observation to propose a new method for causal discovery that is robust under truncated data with ANMs. Overall, our approach offers a general, robust solution that is agnostic to external information about the selection process and still achieves comparable performance to the state-of-the-art approaches when selection is not present.

Cellular perturbation experiments are essential for probing biological mechanisms and guiding therapeutic discovery. While existing methods often treat perturbations as monolithic distributional shifts, they leave the underlying structural variations implicit. Rather than attempting to infer independent global networks for every condition, we propose scPertCRL, a causally structured generative framework for single-cell perturbation prediction. scPertCRL decomposes perturbation effects into a shared latent structural backbone, condition-specific structural modulations, and exogenous shifts. By leveraging biologically informed embeddings and mechanism-aware supervision, Differential Network Learner parameterizes these localized structural changes, enabling robust generalization to unseen interventions. Extensive experiments on genetic and pharmacological benchmarks demonstrate that scPertCRL significantly improves predictive accuracy and out-of-distribution generalization, while yielding model-based insights into perturbation mechanisms.


CausalSpatial: A Benchmark for Object-Centric Causal Spatial Reasoning

Wenxin (Wendy) Ma ⋅ Chenlong Wang ⋅ Ruisheng Yuan ⋅ Hao Chen ⋅ Nanru Dai ⋅ Chengxin Qian ⋅ Yijun Yang ⋅ Qi Chen ⋅ Zhaoyang Wang ⋅ S. Kevin Zhou ⋅ Jianwen Xie ⋅ Alan Yuille ⋅ Jieneng Chen

Humans can look at a static scene and instantly predict what happens next --- will moving this object cause a collision? We call this ability Causal Spatial Reasoning. However, current multimodal large language models (MLLMs) cannot do this, as they remain largely restricted to static spatial perception, struggling to answer ``what-if'' questions in a 3D scene. We introduce CausalSpatial, a diagnostic benchmark evaluating whether models can anticipate consequences of object motions across four tasks: Collision, Compatibility, Occlusion, and Trajectory. Results expose a severe gap: humans score 84\% while GPT-5 achieves only 54\%. Why do MLLMs fail? Our analysis uncovers a fundamental deficiency: models over-rely on textual chain-of-thought reasoning that drifts from visual evidence, producing fluent but spatially ungrounded hallucinations. To address this, we propose the Causal Object World Model (COW), a framework that externalizes the simulation process by generating videos of hypothetical dynamics. With explicit visual cues of causality, COW enables models to ground their reasoning in physical reality rather than linguistic priors. Code and dataset are publicly available.

Many survival problems are inherently time-dependent: a biomarker may be predictive soon after diagnosis but irrelevant years later, or a treatment effect may diminish, reverse, or interact with follow-up duration. However, tree-based survival models commonly used in practice typically assign each individual a fixed, time-independent risk score. Neural non-proportional-hazards models can capture dynamic effects, but are often more challenging to tune, scale, and interpret than boosted trees. We present CCTimeBoost, a gradient-boosted tree framework for learning continuous-time, time-varying relative risk. The method augments covariates with time features, trains boosted trees on sampled risk sets using a group-softmax objective, and estimates a full-risk-set baseline hazard to generate survival curves. We show that this objective has a likelihood interpretation: for finite sampled risk sets, it matches a nested case-control likelihood, and in the full-risk-set limit, it converges to the Cox partial likelihood. Across eight right-censored survival benchmarks, CCTimeBoost delivers state-of-the-art discrimination performance relative to tree and neural baselines and is competitive in calibration. In controlled simulations, it accurately recovers diminishing effects, crossing hazards, and rule-based temporal interactions, also behaving as a proportional model when proportional hazards are present. These results indicate that non-proportional survival modeling can be incorporated into standard boosted-tree methods.

A concept-bottleneck classifier can remain correct, remain $\ell_2$-robust in pixel space, and still silently rewire which concepts carry its decision under a small perturbation. The label looks safe; the audit trail has been swapped out underneath it. This failure mode is not a quirk of randomized smoothing: any certifier that sees only the input and the label distribution admits classifiers with arbitrarily large concept drift inside its certified ball. To certify the missing object, we propose CRISP, which delivers a deterministic concept-space certified radius $r^\star(x)$ inside which both the label and the concept activations are stable. The radius is given by a closed-form Lipschitz--margin bound and is tight up to a factor of two on the relevant encoder/head/margin triple. Under $\varepsilon$-concept-faithfulness and a known concept-to-target subgraph, $r^\star$ further certifies the causally meaningful concepts; a matching impossibility result shows that faithfulness cannot be removed. Training relies on interventional proxies---SCM simulators, attribute editors, environment swaps, concept-matched retrieval---which approximate $\mathrm{do}(\cdot)$-interventions on natural images. We bound the proxy-vs-truth gap in closed form and measure it directly in the synthetic regime where ground truth is available. Empirically, the bound is informative rather than vacuous. On a synthetic SCM with observable $c^\star$, 923 correctly-classified samples (5 seeds, $k \in \{3,6\}$) yield zero certificate violations under budget-5.0 APGD, with empirical tightness $r^\star/r_{\mathrm{adv}} \approx 0.32$. Concept-fidelity (MCC) rises from $[0.37,0.51]$ for plain CBM to $[0.74,0.83]$ for CRISP, and concept-space attack success drops by up to 13 points at matched budget. A frozen-backbone Waterbirds run preserves the zero-violation property across 1,980 samples and surfaces what we call the Lipschitz tax: a $\approx 27\times$ gap between architectural and local encoder Lipschitz constants that is the dominant bottleneck to tightness on natural images. A preregistered protocol for CUB-200 and CANDLE (Reddy et al., 2022) is fixed in Appendix C.8.

Autoregressive language models (ARMs) have been shown to memorize and occasionally reproduce training data verbatim, raising concerns about privacy and copyright liability. Diffusion language models (DLMs) have recently emerged as a competitive alternative, yet their memorization behavior remains largely unexplored due to fundamental differences in generation dynamics. To address this gap, we present a systematic theoretical and empirical characterization of memorization in DLMs. We propose a generalized probabilistic extraction framework that unifies prefix-conditioned decoding and diffusion-based generation under arbitrary masking patterns and stochastic sampling trajectories. Theorem 4.3 establishes a monotonic relationship between sampling resolution and memorization: increasing resolution strictly increases the probability of exact training data extraction, implying that autoregressive decoding corresponds to a limiting case of diffusion-based generation by setting the sampling resolution maximal. Extensive experiments across model scales and sampling strategies validate our theoretical predictions. Under aligned prefix-conditioned evaluations, we further demonstrate that DLMs exhibit substantially lower memorization-based leakage of personally identifiable information (PII) compared to ARMs.


CircuitSeer: Mining High-Quality Data by Probing Mathematical Reasoning Circuits in LLMs

Shaobo Wang ⋅ Yongliang Miao ⋅ YuanCheng Liu ⋅ Qianli Ma ⋅ Ning Liao ⋅ Dongrui Liu ⋅ Linfeng Zhang

Large language models (LLMs) have demonstrated impressive reasoning capabilities, but scaling their performance often relies on massive reasoning datasets that are computationally expensive to train on. Existing data selection methods aim to curate smaller, high-quality subsets but often rely on costly external models or opaque heuristics. In this work, we shift the focus from external heuristics to the model's internal mechanisms. We find that complex reasoning tasks consistently activate a sparse, specialized subset of attention heads, forming core reasoning circuits. Building on this insight, we propose CircuitSeer, a novel data selection method that quantifies the reasoning complexity of data by measuring its influence on these crucial circuits. Extensive experiments on 4 models and 9 datasets demonstrate CircuitSeer's superiority. Notably, fine-tuning Qwen2.5-Math-7B on just 10\% of data selected by our method achieves a 1.2-point gain in average Pass@1 over training on the full dataset, highlighting its efficiency and effectiveness.


City-RAG: Stepping Into a City via Spatially-Grounded Video Generation

Gene Chou ⋅ Charles Herrmann ⋅ Kyle Genova ⋅ Boyang Deng ⋅ Songyou Peng ⋅ Bharath Hariharan ⋅ Jason Y. Zhang ⋅ Noah Snavely ⋅ Philipp Henzler

We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produce a plausible sequence that is consistent with a text (T2V) or image (I2V) prompt. However, the capability to reconstruct the real world under arbitrary weather conditions and dynamic object configurations is essential for downstream applications including autonomous driving and robotics simulation. To this end, we present CityRAG, a video generative model that leverages large corpora of geo-registered data as context to ground generation to the physical scene, while maintaining learned priors for complex motion and appearance changes. CityRAG relies on temporally unaligned training data, which teaches the model to semantically disentangle the underlying scene from its transient attributes. Our experiments demonstrate that CityRAG can generate coherent minutes-long, physically grounded video sequences, maintain weather and lighting conditions over thousands of frames, achieve loop closure, and navigate complex trajectories to reconstruct real-world geography.


CitySTAR: Agent-Driven Structured and Topology-Aware Reasoning for Open-Vocabulary Urban 3D Grounding

Shuai Zhang ⋅ Hongye Hou ⋅ qinghe liu ⋅ Zhuoxiao Li ⋅ Dongli Wu ⋅ Jing OU ⋅ Yuan Liu ⋅ Wufan Zhao

3D grounding aims to localize target entities in complex scenes from natural language and plays a fundamental role in embodied perception and spatial reasoning. However, existing approaches mostly rely on feature similarity or direct matching, making it difficult to connect natural-language intent with the implicit semantic and geometric structures hidden in billion-scale urban point clouds. We reformulate city-scale 3D grounding as structured constraint reasoning, where description semantics are organized into computable cross-modal constraints over open-vocabulary 3D entities, attributes, and spatial relations. We present CitySTAR, a training-free framework for reasoning-driven urban 3D grounding. CitySTAR lifts raw billion-scale urban point clouds into a query-ready scene graph of open-vocabulary 3D instances, with CodeLLM-driven tools supplying multimodal evidence for node attributes and 3D spatial relations. It then models target-context topology with paired hypergraphs and performs bidirectional topology verification for structural disambiguation. Finally, a Reflective Cross-modal Grounding module integrates topology consistency and candidate-centered 2D visual evidence to make decisions over a metric-aware 3D context graph. To further support this setting, we introduce CitySTAR-3D, an enhanced benchmark that improves semantic coverage, instance completeness, bounding-box fidelity, and spatial-relation complexity in city-scale 3D grounding. Extensive experiments show that CitySTAR consistently improves open-world urban 3D grounding while maintaining strong interpretability and generalization.

Recent text-only models demonstrate remarkable reasoning capabilities. Extending these to visual domains requires vision-language models to translate images into text descriptions. However, current models, trained to produce captions for human readers, often omit the precise details that reasoning systems require. This creates an interface mismatch: reasoners often fail not due to reasoning limitations but because they lack access to critical visual information. We propose Adaptive-Clarification Reinforcement Learning (AC-RL), which, through interaction, teaches vision models what information reasoners need. Our key insight is that clarification requests during training reveal information gaps; by penalizing success that requires clarification, we create pressure for comprehensive initial captions that enable the reasoner to solve the problem in a single pass. AC-RL improves average accuracy by $4.4$ points over pretrained baselines across seven visual reasoning benchmarks, and analysis shows it would cut clarification requests by up to $44$\% if those were allowed. By treating clarification as a form of implicit supervision, AC-RL demonstrates that vision-language interfaces can be effectively learned through interaction alone, without requiring explicit annotations.

Domain Generalization (DG) aims to leverage multiple source domains to train models that can generalize to unseen target domains, thereby mitigating the performance degradation caused by domain shifts. Most existing approaches impose unified alignment or suppression constraints in the global feature space, overlooking the heterogeneity of feature channels in terms of class discriminativeness and domain stability. Such coarse-grained constraints may lead to the loss of discriminative information or the retention of domain-sensitive features. To address this issue, we propose a novel DG framework consisting of two core modules: Channel Sensitivity Decomposition (CSD), which employs channel-wise response variances along the class and domain dimensions as empirical proxy signals to partition feature channels into four functional subspaces via a continuous and differentiable soft decomposition mechanism; and Subspace-Specific Feature Augmentation (SSA), which implements tailored augmentation and regularization strategies for different subspaces to suppress environment-related spurious correlations while preserving cross-domain stable discriminative features. Additionally, a KL-divergence-based prediction consistency constraint is introduced to stabilize the training process and enhance model robustness. Extensive experiments on multiple standard DG benchmarks demonstrate that our framework consistently outperforms state-of-the-art methods.

Bayesian accounts of in-context learning treat context as exogenous evidence, but in multi-turn dialogue this assumption can fail: the assistant changes the user state that generates the next utterance, so the model updates on evidence partly produced by its own policy. We formalize this as Socially Coupled In-Context Learning (SC-ICL). The relevant stability quantity is a \emph{round-trip gain} (the product of belief-to-action, action-to-evidence, and evidence-to-belief derivatives), and in a two-state local reduction the standard spectral condition on the closed-loop matrix factors as a product of a user-side gain and an assistant-side gain. Risk is therefore dyadic: neither a susceptible user nor a responsive assistant becomes unstable without the other side of the loop. We connect this stability boundary to a sigmoidal transition in relational-posterior space, which explains why myopic approval-seeking can become unstable in supportive dialogue. AI--AI experiments instantiate one-step gain assays, free-running conversations, open-loop replay, and prompt-based controllers, and four convergent results support the closed-loop account: high-gain dyads cross an operational collapse criterion while damping and barrier-style controllers sharply reduce it; a trajectory-conditioned spectral estimator separates collapsing from stable runs at the theoretical stability boundary, crossing unity near operational onset; the dyadic pattern recurs across all four pairings of GPT-4o and GPT-5.4-mini as user and assistant simulators; and a persona-only ablation that uses no numeric coupling labels in the prompts reproduces the same dyadic collapse matrix. The results are not clinical evidence; they argue that alignment should evaluate bounded closed-loop inference, not only next-response quality.


Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training

Hongzhi Ruan ⋅ Pei Liu ⋅ Weiliang Ma ⋅ Zhengning Li ⋅ Xueyang Zhang ⋅ Jun Ma ⋅ Dan Xu ⋅ Kun Zhan

Data scaling is fundamental to modern deep learning, and grows increasingly critical as autonomous driving shifts to end-to-end learning. Real-world driving data is expensive to annotate and scene-biased, making real-synthetic co-training with near-infinite synthetic data a promising direction. However, naively incorporating all available synthetic data is inefficient and leads to distribution shifts, and optimizing data mixture under practical training budgets remains a critical yet under-explored problem. In this sense, we claim that the mixture of training data requires clear guidance in terms of scene types and quantities. Particularly in this work, we conceptualize the data mixture approximately as a dynamic optimization process that iteratively adjusts the training data mixture to maximize model performance, guided by closed-loop evaluation feedback, and propose AutoScale, a fully automated closed-loop data engine unifying scene representation, data mixture optimization and retrieval, as well as model training and evaluation. Specifically, we propose Graph Regularized AutoEncoder (Graph-RAE) for driving scene representations, introduce Cluster-aware Gradient Ascent (Cluster-GA) for cluster-wise importance estimation and reweighting, and perform cluster-guided vector retrieval to select high-value samples. Experiments on NavSim demonstrate that AutoScale outperforms vanilla co-training and cross-domain baselines, achieving better performance with fewer synthetic samples under constrained budgets.

The 1-identification problem is a fundamental pure-exploration problem in multi-armed bandits. An agent aims to determine whether there exists an arm whose mean reward exceeds a known threshold $\mu_0$, or to output \textsf{None} otherwise. The agent must guarantee correctness with probability at least $1-\delta$, while minimizing the expected number of arm pulls $\mathbb{E}[\tau]$. We study the 1-identification problem and make two main contributions. First, for instances with at least one qualified arm, we derive a new lower bound on $\mathbb{E}[\tau]$ via a novel optimization formulation. Second, we propose a new algorithm and establish upper bounds that match the lower bounds up to polynomial logarithmic factors uniformly over all instances. Our result complements the analysis of $\mathbb{E}\tau$ when there are multiple qualified arms, which is an open problem in the literature.

Regular time-series generation from irregular time series is typically reduced to one-shot imputation: complete each partial sequence once, freeze the completions, and train a generator on the resulting surrogate dataset. We show that this open-loop strategy is structurally flawed. Point imputers collapse ambiguous missing regions to conditional averages, while stochastic imputers become stale as the generator evolves; the binding constraint is therefore iteration, not imputer design. We introduce a co-evolving Monte Carlo EM framework that closes the imputer--generator loop. Each E-step samples missing values from the posterior of the current generator, and each M-step retrains the generator on these completions, preserving unconditional generator training. Since our diffusion prior operates in a lifted 2D representation while observations live in time-series space, we further introduce Posterior Sampling for Lifted Representations (PSLR), which enforces operator-consistent conditioning, uncertainty preservation, and manifold consistency during the E-step. Across nine datasets and four corruption regimes, our method achieves state-of-the-art performance, improving over the strongest baselines by 59\% in discriminative score, 15\% in predictive score, 73\% in Context-FID, and 71\% in feature-correlation error. Ablations show that our approach reduces average discriminative score by 73\% over single-step baselines, closing most of the remaining gap to the clean-data oracle.

Spoof diarization aims to jointly localize spoofed regions and determine which spoofing method generated each segment within a partially spoofed utterance. Existing approaches rely on modular pipelines that separate localization and clustering, introducing training--inference mismatch and requiring oracle knowledge of the number of spoofing classes at inference time. We propose \textbf{E2E-SD}, the first \emph{clustering-free} end-to-end framework for spoof diarization, which reformulates the task as variable-cardinality temporal set prediction. E2E-SD jointly learns localization and class assignment using encoder-decoder attractors with Hungarian-matched training, eliminating the need for post-hoc clustering and oracle class counts. Our framework dynamically generates class-specific attractors through an LSTM decoder and estimates the number of active spoofing classes via attractor existence probabilities, requiring no external class-count information at any stage. On the PartialSpoof benchmark, E2E-SD achieves a $\JER$ of 15.76\%, a 41.6\% relative improvement over the previous best result (26.99\%), while operating fully oracle-free with 88.7\% class-count estimation accuracy. The largest gain is observed on unknown attacks, where $\JER$ drops from 49.02\% to 21.00\% (57.1\% relative improvement), demonstrating that end-to-end optimization with dynamic attractors enables more attack-agnostic generalization than modular clustering pipelines.


CoANeRV: Coordinate-Aware Token-Space Neural Video Representation

Jialong Guo ⋅ Ke Liu ⋅ MENGXUAN LI ⋅ Jiajun Bu ⋅ Haishuai Wang

Neural representations for videos (NeRV) have shown strong reconstruction fidelity by storing video-specific information in network weights. However, existing formulations typically require either costly per-video optimization or video-specific weight generation, making it difficult to scale to efficient amortized video representation. We propose \textbf{CoANeRV}, a coordinate-aware token-space neural video representation framework that shifts video-specificity from decoder weights to compact latent video tokens. Instead of optimizing or generating an instance-specific network, CoANeRV uses a shared coordinate-conditioned decoder to reconstruct videos by querying tokenized video content at continuous spatio-temporal coordinates. This formulation decouples video representation from parameter instantiation, enabling feed-forward video encoding while retaining coordinate-level reconstruction flexibility. To make token-space reconstruction effective, CoANeRV introduces a coordinate-aware decoding architecture that aligns spatio-temporal queries with video tokens through axis-adaptive positional encoding and temperature-modulated cross-attention. Block-wise coordinate querying further reduces peak attention memory, making high-resolution reconstruction practical. Experiments on diverse video datasets show that CoANeRV consistently improves reconstruction quality over prior feed-forward NeRV and INR baselines, reduces peak memory compared with attention-based coordinate decoders, and provides efficient amortized encoding without per-video optimization. These results suggest that token-space neural representation is a scalable alternative to conventional weight-space NeRV formulations. The code is available at https://anonymous.4open.science/r/CoANeRV-50F7/.


Coarsening Linear Non-Gaussian Causal Models with Cycles

Francisco Madaleno ⋅ Francisco Pereira ⋅ Alex Markham

Recent work on causal abstraction, in particular graphical approaches focusing on causal structure between clusters of variables, aims to summarize a high-dimensional causal structure in terms of a low-dimensional one. Existing methods for learning such summaries from data assume that both the high- and low-dimensional structures are acyclic, which is helpful for causal effect identification and reasoning but excludes many high-dimensional models and thus limits applicability. We show that in the linear non-Gaussian (LiNG) setting, the high-dimensional acyclicity assumption can be relaxed while still allowing recovery of a low-dimensional causal directed acyclic graph (DAG). We further connect identifiability of this low-dimensional DAG to existing results: LiNG models with cycles are observationally identifiable only up to an equivalence class whose members differ by reversals of directed cycles; our low-dimensional DAG, which is invariant across all members of a given equivalence class, thus forms a natural representative of the class. While existing approaches for learning this observational equivalence class over high-dimensional variables have exponential time complexity, our low-dimensional summary is learned in worst-case cubic time and comes with explicit bounds on the sample complexity. We provide open source code and experiments on synthetic data to corroborate our theoretical results.


CodecSplat: Ultra-Compact Latent Coding for Feed-Forward 3D Gaussian Splatting

Pengpeng Yu ⋅ Runqing Jiang ⋅ Qi Zhang ⋅ Dingquan Li ⋅ Jing Wang ⋅ Yulan Guo

While feed-forward 3D Gaussian splatting reconstructs renderable Gaussian primitives from sparse context views without per-scene optimization, existing pipelines do not provide a compact scene representation for storage or transmission. A natural solution is to apply existing 3DGS compression methods to the generated Gaussian primitives. However, this approach operates on the final irregular 3D representation and is decoupled from the internal feature-to-Gaussian generation process, which limits compression efficiency. To address this, we introduce \emph{CodecSplat}, an ultra-compact latent coding framework for feed-forward 3D Gaussian splatting. CodecSplat first encodes an intermediate 2D Gaussian-generation feature into an entropy-coded scene bitstream. At the decoder, the latent feature is reconstructed and used to predict depth and Gaussian parameters, which are then mapped to 3D Gaussian primitives. Note that, by integrating compression into the feed-forward Gaussian generation pipeline, CodecSplat avoids inefficient compression over irregular 3D Gaussian primitives and allows the codec to exploit the structured intermediate feature representation. We instantiate CodecSplat on a feed-forward Gaussian splatting backbone with depth-guided multi-view feature refinement and a hierarchical learned feature codec. On DL3DV and RealEstate10K datasets, CodecSplat achieves 23.56-26.36 dB and 24.76-27.05 dB PSNR with only 20.00-107.77 KiB and 3.37-12.51 KiB per scene, respectively. This is roughly one order of magnitude smaller than compressing feed-forward generated Gaussian primitives, while preserving controllable rate--distortion behavior.


CoDMD: Copula-aware Distribution Matching Distillation for Fast Video Generation

Wenhu Zhang ⋅ Kun Cheng ⋅ Changyuan Wang ⋅ Shiyao Li ⋅ Yuechen Zhang ⋅ Wenbo Li ⋅ Jiajun Zha ⋅ Jingyi Zhang ⋅ Kang Zhao ⋅ Jiaya Jia

Few-step distillation for video diffusion models has attracted significant attention, driven by the urgent demand for efficient deployment in real-world scenarios. However, Distribution Matching Distillation (DMD), a leading paradigm, tends to degrade under limited NFE budgets, manifesting in video generation as layout instability, oversaturation, and broken motion dynamics. We trace this failure to a structural limitation: standard DMD is an intra-sample distribution-matching objective with coordinate-wise gradients, and thus imposes no explicit constraint on the relational geometry across batch elements or temporal frames, leaving the underlying copula largely unregulated. Combined with the mode-seeking tendency of its reverse-KL objective, this absence of relational guidance makes DMD prone to collapsing into local optima in the few-step regime. Motivated by this insight, we propose Copula-aware DMD (CoDMD), a lightweight relational regularizer that reuses score estimates already produced by the frozen teacher and the online fake model to construct pairwise relation matrices across samples and frames.These are matched through a supplementary distributional objective that requires no additional networks, datasets, or sampling trajectories. On the Wan-2.1-T2V model series at 1.3B \& 14B scales, CoDMD distills 50-step teachers into 4-step students, achieving an approximate 25$\times$ speed-up while attaining VBench scores of 84.46 \& 84.87, outperforming prior trajectory-based (rCM 82.81 \& 84.05) and distribution-based (DMD 83.38 \& 83.81) methods.


Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Image Generation

Renping Zhou ⋅ Zanlin Ni ⋅ Tianyi Chen ⋅ Zeyu Liu ⋅ Yang Yue ⋅ Yulin Wang ⋅ Yuxuan Wang ⋅ Jingshu Liu ⋅ Gao Huang

Recently, Masked Diffusion Models (MDMs) have shown promising potential across vision, language, and cross-modal generation. However, a notable discrepancy exists between their training and inference procedures. In particular, MDM inference is a multi-step, iterative process governed not only by the model itself but also by various schedules that dictate the token-decoding trajectory (e.g., how many tokens to decode at each step). In contrast, MDMs are typically trained using a simplified, single-step BERT-style objective that masks a subset of tokens and predicts all of them simultaneously. This step-level simplification fundamentally disconnects the training paradigm from the trajectory-level nature of inference, leaving the inference schedules never optimized during training. In this paper, we introduce Co-GRPO, which reformulates MDM generation as a unified Markov Decision Process (MDP) that jointly incorporates both the model and the inference schedule. By applying Group Relative Policy Optimization at the trajectory level, Co-GRPO cooperatively optimizes model parameters and schedule parameters under a shared reward, without requiring costly backpropagation through the multi-step generation process. This holistic optimization aligns training with inference more thoroughly and substantially improves generation quality. Empirical results across four benchmarks-ImageReward, HPSv2, GenEval, and DPG-Bench-demonstrate the effectiveness of our approach.

Label-free proxy scores are often used as if they measured downstream discriminative difficulty. A score may rank samples similarly to a target, predict that target on held-out data, or improve a downstream decision; these are distinct evidence levels. We present Cross-Objective Hardness Evaluation (COHE), a protocol that audits each proxy--target--protocol tuple through dependence, held-out-prediction, and operational-transfer gates. COHE maps standard, inspectable statistics to bounded claim tiers; the contribution is evidence-to-claim calibration, not a new sampler, metric, or model. A DINOv2 isolation stress test illustrates the boundary: representation isolation captures a first-learning dynamics relation (rho = 0.4363, R^2 = 0.2126), but using it for direct 10% subset selection fails against Random (Delta = -15.41, 95% CI [-15.87, -14.95]). COHE therefore preserves this target-specific relation without rebranding it as actionability. Across CIFAR-10/100, ImageNet-10, and ImageNet-1K validation checks, eight primary generative proxy families do not support CE/margin surrogate claims: they show weak alignment, near-zero held-out CE prediction, and unreliable transfer. Controls show stable discriminative target rankings, expected shuffled-proxy null behavior, and a task-aligned ImageNet target-hard retrieval/triage tuple that passes all gates. The anonymized supplementary package provides lightweight gate-wise audit code, toy scalar CSV examples, cached aggregate summaries, selected derived scalar arrays for tabular gate verification, supported paper-style verification scripts, and provenance notes without redistributing raw datasets, images, weights, checkpoints, or full scoring/retraining pipelines.


Colored Noise Diffusion Sampling

Hadar Davidson ⋅ Noam Issachar ⋅ Sagie Benaim

Diffusion models achieve state-of-the-art image synthesis, with their generative trajectories fundamentally exhibit a spectral bias, resolving low-frequency global structures early and high-frequency fine details later. Conventional stochastic differential equation (SDE) solvers fail to account for this dynamic, naively injecting uniform white noise throughout the entire process and misusing the finite energy budget. In this work, we establish a mathematical framework that reconsiders SDE inference as a targeted, frequency-decoupled energy transfer. Leveraging this framework, we introduce Colored Noise Sampling (CNS), a novel, training-free stochastic solver. Rather than injecting uniform white noise, CNS utilizes a dynamic, timestep- and frequency-dependent schedule that more-efficiently allocates injected energy toward structurally unresolved frequency bands. By actively exploiting the model's inherent spectral bias, CNS systematically steers the generated distribution toward the true data manifold. Extensive experiments demonstrate that CNS significantly outperforms standard ODE and SDE baselines as a strictly plug-and-play, inference-time sampler substitution across diverse architectures (SiT, JiT, FLUX). Compared to standard sampling on ImageNet-256, CNS achieves substantial unguided FID reductions, improving from 8.26 to 6.27 on SiT-XL/2, 32.39 to 26.69 on JiT-B/16, and 11.88 to 8.31 on JiT-H/16, while yielding consistent relative FID improvements with Classifier-Free Guidance.


Combating Camouflage and Forgetting: Spatio-Temporal Dual Denoising for Money Laundering Detection

Xueting Yang ⋅ Zhong Li ⋅ changjun jiang ⋅ Wenbin Zhang ⋅ Meikang Qiu

Money laundering detection on financial transaction graphs is critical but challenging. Laundering activities often conceal illicit fund flows within dense legitimate transaction neighborhoods and disperse suspicious behaviors over extended time spans. These patterns cause weak illicit signals to be spatially diluted and temporally forgotten, limiting the effectiveness of standard graph learning methods in AML scenarios. In this paper, we propose a spatio-temporal dual denoising framework for money laundering detection on evolving graphs, named EvoDen. For spatial denoising, instead of relying on passive neighborhood aggregation, we design a risk-potential guided walk mechanism to extract denoised and suspicion-biased transaction sequences, enabling the model to isolate anomalous fund flows from noisy local neighborhoods. For temporal denoising, we propose a continual Transformer-based sequence learning method driven by a reconstruction objective, which learns denoised sequence representations while enabling latent-space replay of historical knowledge to mitigate representation drift. We conduct experiments on three transaction datasets, and the results show that EvoDen outperforms state-of-the-art baselines in accurately identifying concealed money laundering entities.


COMET: Codebook-based Online-adaptive Multi-scale Embedding for Time-series Anomaly Detection

JINWOO PARK ⋅ Hyeongwon Kang ⋅ Seung H Han ⋅ Pilsung Kang

Multivariate time-series anomaly detection aims to identify abnormal temporal patterns in complex real-world systems. However, robust unsupervised time-series anomaly detection remains challenging because anomalies can manifest over different temporal ranges and emerge from interactions among variables, while distribution shifts can cause normal patterns to drift over time. To address these challenges, we propose Codebook-based Online-adaptive Multi-scale Embedding for Time-series anomaly detection (COMET), a unified framework with three tightly coupled components. First, Multi-scale Patch Encoder learns correlation-aware patch embeddings by combining variable-specific temporal dynamics with shared inter-variable structure across multiple patch scales. Second, Vector-Quantized Coreset quantizes these embeddings into representative normal prototypes, which are used to detect anomalies through a dual score combining quantization error and density-aware memory distance. Third, Online Codebook Adaptation leverages the same codebook structure to identify reliable normal samples from training-activated codebook entries, and then updates the prototypes through contrastive learning at inference time. Experiments on five benchmark datasets demonstrate that COMET achieves the best performance on 39 out of 45 evaluation metrics compared to seven baselines, showing robust performance across diverse evaluation protocols.


CoMet: Context and Multiplicity Decomposition for Multimodal Uncertainty Estimation

Sanghyuk Chun ⋅ William Yang ⋅ Amaya Dharmasiri ⋅ Olga Russakovsky

Uncertainty estimation has been a long-standing challenge in AI models; it amounts to ``knowing what you don't know,'' and metacognition is notoriously difficult even for humans (cf. the Dunning-Kruger effect). Although it is still far from solved even in simpler classification systems, tackling it in multimodal large language models (MLLMs) is becoming increasingly important. Within MLLMs, uncertainty can stem from any of the diverse sources as well as from their relationships, and further can stem from the unbounded answers in the open-ended setting. To tackle the issues, we propose CoMet, an MLLM uncertainty estimation method by decomposing uncertainty into a context-specific term and a multiplicity-specific term. The former captures ambiguity induced by the given context (e.g., task or prompt), while the latter captures how many plausible answers determined by the context remain compatible with the given input. We train a lightweight post-hoc uncertainty module to estimate these quantities, which enables efficient uncertainty estimation without autoregressive answer generation or repeated sampling. Experiments on various open-ended multimodal benchmarks, hallucination detection, and multiple-choice visual question answering benchmarks show that CoMet consistently improves uncertainty estimation over existing baselines while remaining efficient in practice. Code will be available upon acceptance.

Vision–language models (VLMs) exhibit strong multimodal reasoning, yet their growing parameter counts incur prohibitive memory footprints and inference latency, hindering real-time deployment. Post-training quantization alleviates these costs, but aggressive low-bit quantization often degrades accuracy by disrupting visual–linguistic alignment. Meanwhile, lossless compression preserves accuracy but suffers from metadata overhead and sequential decoding, limiting GPU parallelism. We propose Comp²VLM, a hybrid framework that jointly designs quantization and lossless compression for efficient VLM serving. We introduce Duplication-Enhanced Quantization (DEQ), which adaptively selects group-wise scales to reduce the entropy of quantized representations while preserving reconstruction fidelity, thereby improving compressibility. To realize runtime gains, we present Quantized Data-aware Lossless Compression (QPress), a hardware-friendly scheme based on a universal Huffman table that enables block-wise metadata sharing and GPU-parallel decoding. Across diverse VLMs and benchmarks, Comp²VLM maintains competitive accuracy while reducing the effective precision to W3.5A3.2~4.9. On LLaVA-onevision-7B, it achieves +4.7% TextVQA and +7.7% OCRBench accuracy over W4A8-based SOTA at an average bit-width of W3.5A4.6. DEQ yields up to 33% entropy reduction, and QPress improves compression efficiency by up to 13.3% with only ~4% end-to-end latency overhead. To our knowledge, Comp²VLM is the first unified framework to co-design quantization and lossless compression for VLMs, demonstrating their complementary benefits beyond standalone methods.


Compact Representations of Impact-Based Fair-Ranking Policies

Yuki Uehara ⋅ Naoki Nishimura ⋅ Noriyoshi Sukegawa ⋅ Yuichi Takano

On two-sided platforms such as e-commerce marketplaces, recommender systems deliver personalized item rankings to consumers and serve as a primary channel through which transactions occur between item providers and consumers. At the same time, promoting the activities of item providers requires the platform to balance consumer satisfaction with fair exposure across items. Recent work formulates this objective as impact-based fair ranking, which maximizes a Nash social welfare criterion over stochastic ranking policies. However, the resulting policies must be served as user-specific mixtures of deterministic rankings, whose support size can grow combinatorially. This imposes substantial storage and serving overhead, driving up operational cost in production deployment. Moreover, unseen users must fall back to a generic ranking rule. We therefore develop compact representations for serving impact-based fair-ranking policies. We first prove a tight bound that every optimum admits an impact-preserving mixture with support size at most the number of items. Using a dual reformulation, we then replace stored user-specific ranking mixtures with item-indexed dual vectors and compress the resulting sequence of dual vectors by weighted random subsampling. The resulting deployment representation is independent of the training-user population and can be used for unseen users. Experiments on real-world recommendation datasets show that the sampled dual mixture preserves the fair-ranking objective while reducing the storage required at deployment by up to 2,596$\times$ relative to a user-coupled mixture baseline.


Complementary Cache Guidance with Gradient Disentanglement for Continuous Test-Time Adaptation

Fanchun Meng ⋅ Yuhang Pei ⋅ Jiazhen Huang ⋅ Tao Ren ⋅ Yifan Wang ⋅ Wei Ju ⋅ Xiao Luo

Continuous test-time adaptation (CTTA) is crucial for deploying vision-language models (VLMs) in real-world environments where data distributions shift over time. Existing approaches typically utilize confidence-driven pseudo-labeling to enhance the performance on the target domain while retraining VLMs using historical information. Despite the progress, their performance is far from satisfactory due to error accumulation from noisy pseudo-labels and optimization interference between newly acquired and historical knowledge. Towards this end, we propose a novel approach named Complementary Cache Guidance with Gradient Disentanglement (CURE) for CTTA of VLMs. The core of our CURE is to balance instant knowledge acquisition and historical knowledge consolidation using both complementary cache systems and parameter space disentanglement. In particular, our CURE first constructs affinity structures among diverse text prompts, which guide majority voting to improve the quality of pseudo-labeling. More importantly, our CURE introduces a complementary cache system consisting of a short-term cache and a long-term cache with sample quality monitoring. The short-term cache stores recent reliable samples with high entropy, while the long-term cache preserves representative samples with high gradient consistency across iterations. Then, we extract prototypes from both caches for cross-modal alignment, enabling the model to acquire new knowledge while mitigating old knowledge forgetting. To further reduce the interference when learning from two caches, we perform low-rank decomposition of the gradient space, which facilitates historical knowledge consolidation in the null space of short-term signals. Extensive experiments on several benchmarks demonstrate the superiority of CURE over state-of-the-art baselines. Our source codes are available at https://anonymous.4open.science/r/CURE-B486.


Compute Efficiency and Serial Runtime Tradeoffs for Stochastic Momentum Methods

Depen Morwani ⋅ Alexandru Meterez ⋅ Pranav Nair ⋅ Sham Kakade

Stochastic momentum methods such as heavy ball (HB), Nesterov momentum, and variants of Accelerated SGD (ASGD) [Kidambi et al., 2018] are widely used in modern training, but their stochastic benefits depend on two distinct quantities: serial runtime, the number of iterations needed to reach a target accuracy, and compute efficiency (CE), the inverse total gradient-query or FLOP cost. Larger batches reduce serial runtime without hurting CE only when the contraction gap grows linearly with batch size. We study stochastic HB and ASGD for consistent linear regression with Gaussian covariates and prove finite-dimensional, discrete-time lower bounds on their batch-size tradeoffs. Our first result shows that HB does not improve the CE frontier over SGD for arbitrary spectra; rather, it preserves SGD-level CE over a larger batch-size window, allowing larger batches to reduce serial runtime until HB reaches its deterministic accelerated scale. This window can be a factor $\sqrt{\kappa}$ larger than the SGD critical batch size. For ASGD, the picture is more spectrum-dependent: for rapidly decaying power-law spectra, ASGD improves small-batch CE over HB/SGD, but as batch size grows it trades this CE advantage for improved serial runtime. Synthetic linear-regression experiments verify these qualitative regimes, including near-overlap of ASGD and HB for slowly decaying spectra and the predicted CE--serial tradeoff for rapidly decaying spectra.


Computer Use at the Edge of the Statistical Precipice

Pierluca D Oro ⋅ Sneha Silwal ⋅ William R Wong ⋅ Yuxuan Sun ⋅ Fanyi Xiao ⋅ Manchen Wang ⋅ Eric Gan ⋅ Allen Bolourchi ⋅ Joseph Tighe

Evaluating Computer Use Agents (CUAs) on interactive environments is fraught with methodological pitfalls that the field has yet to systematically address. We show that a 1MB replay script that blindly executes a recorded action sequence without ever observing the screen outperforms frontier models on prominent static benchmarks, and prove that its expected success rate is exactly equal to the source agent's pass@k in deterministic environments. We trace this and other failures to two root causes: non-principled environment design (static, unsandboxed, or unreliably verified environments) and non-principled evaluation methodology (naïve aggregation and misuse of pass@k for stateful UI interactions). To address the first, we propose PRISM, five design principles for CUA environments (privileged verification, realistic environments, integrity-checked configurations, sandboxed execution, and multifactorial variability) and instantiate them in DigiWorld, a benchmark of 15 realistic sandboxed mobile applications able to evaluate agents in over 3.2 million verified unique configurations. To address the second, we develop an aggregation framework pairing Wilson score intervals with hierarchical bootstrap, producing confidence intervals that correctly account for the nested structure of CUA benchmarks, as we empirically demonstrate. All together, we show that principled environment design and rigorous evaluation methodology are not optional refinements but prerequisites for meaningful CUA research.

Evaluator-guided reasoning systems face a local control problem: given a current answer and evaluator evidence, should the system trust, fallback, or buy more inference? We formulate this as a compute-aware three-action controller over trust, fallback, and expand, where each action is scored by predicted eventual success minus explicit compute cost. The central quantity is action-specific confidence: the success probability induced by trusting now, falling back now, or expanding once and continuing. We show that local action selection is stable when confidence error is smaller than the compute-adjusted action gap, and that regret accumulates only along the realized expansion path. Guided by this analysis, we learn one confidence head per action, calibrate the heads on held-out data, and penalize locally uncertain actions. Across verifier-guided width expansion on MATH and AIME25 and judge-guided revision expansion on a deterministic-answer General Reasoning benchmark, calibrated control improves the utility--compute frontier relative to strong fixed and adaptive baselines. The same pattern persists under policy-induced expansion targets, same-feature budget routers, fallback removal, stronger learned controllers trained on the same features, and moderate evaluator mismatch after refitting and recalibration. When compute is free, high-budget policies can remain competitive; the gain here is compute-efficient allocation, not a uniformly stronger reasoner.

Learning constraint-satisfying policies from offline data without risky online interaction is crucial for reliable deployment of reinforcement learning (RL) agents. Representative methods leverage Implicit Q-Learning and advantage-weighted regression to learn value functions and balance the reward–safety trade-off, demonstrating solid safety performance in safety-critical hard constraint scenarios. However, these methods typically couple reward and safety optimization during offline training and rely on a predefined unified factor to balance the reward–safety trade-off (conservatism). Which ignores the fact that different tasks, due to their distinct reward and cost functions, may require different degrees of conservatism to satisfy safety constraints, resulting in sub-optimal performance. To address this limitation, we propose Co3G, a compositional generative model with test-time controllable conservatism for offline safe RL. During training, Co3G decouples the optimization objectives by separately employing reward-optimal guidance and safety-optimal guidance to induce high-reward and high-safety actions. During deployment, Co3G combines the reward and safety objectives via compositional generation, and further provides three mechanisms—manual tuning, rejection sampling, and RL—to balance the reward and safety guidance scales during deployment, thereby achieving a superior reward–safety trade-off. Extensive experiments on the OSRL benchmark show that Co3G delivers strong safety performance under stringent safety constraints while retaining flexible control over conservatism, significantly outperforming previous hard-constraint and soft-constraint baselines.

Offline multi-objective optimization aims to identify high-quality trade-off designs using only a fixed dataset of previously evaluated candidates. Existing methods often cast this problem as guided generative sampling, producing a finite candidate pool by steering a diffusion or flow model toward Pareto-preferred regions. We study a different interface: deterministic Pareto set amortization. Our method, Offline Pareto Set Learning (OffPSL), trains a preference-conditioned map $h_\pi:\Delta^{m-1}\rightarrow \mathcal{X}$ that returns one design for each requested trade-off direction. Because offline data contain no preference–solution labels, OffPSL trains this map using two offline-only signals: a coordinate-wise pessimistic surrogate scalarization and a diffusion-style denoising-residual support regularizer learned from offline designs. The denoising model is used only as a frozen support prior and is not used to sample candidates. At inference, OffPSL produces exactly one candidate per reference direction with no iterative sampling, test-time optimization, or filtering from a larger generated pool. On standard offline MOO benchmarks, OffPSL is competitive with recent guided generative baselines while keeping the learned preference-to-design map directly queryable and diagnosable.


Constrained Modulatory Reservoirs for Context-Dependent Computation

Sejeung Choi ⋅ Minji Jung ⋅ Jinyeong Park ⋅ Taek D Chung

Biological recurrent circuits can reuse a common substrate by changing its operating regime rather than replacing the circuit itself. This principle is difficult to isolate in fully adaptive neural models, where recurrence, controllers, readouts, and internal representations can all absorb task structure. As a result, successful multi-context behavior does not reveal which route made the computation possible. We introduce Constrained Modulatory Reservoirs (CMR), a fixed-substrate framework for studying such reconfiguration with the main bypasses removed. The primary dynamics and low-rank perturbation basis are fixed, and the readout is denied direct access to the cue, forcing input-dependent adaptation to act through a limited set of recurrent connectivity changes. In transient-cue channel-selection tasks, this restriction makes a capacity transition visible. Low-loss behavior appears only when the available perturbation directions are sufficient, the selected change is retained across the cue-to-signal gap, and the primary dynamics express it during the response window. Controls show that this effect is not explained by low-dimensional cue representation alone. Allowing the substrate to learn hides the bottleneck, while activity- and gain-level alternatives fail to reproduce the same regime. CMR therefore reframes flexible recurrent computation as a question of where adaptive influence is allowed to enter a fixed dynamical system, making capacity, memory, and expression experimentally separable.


Constructive Neural Policies for the Quadratic Assignment Problem via Multi-Expert Imitation

Julien Canitrot-Paradis ⋅ H. Murat AFSAR ⋅ Jean-Yves Pierron ⋅ Chokri Mraidha

The quadratic assignment problem (QAP) is widely considered one of the most difficult NP-hard combinatorial optimization problems, and one that remains challenging for neural constructive methods. We introduce CIMP, a Constructive policy trained by Imitation from Multi-expert teachers and built around an explicit Pairwise pre-conditioning block. CIMP reaches a 11.02% mean optimality gap on the 133 QAPLIB instances with published reference values, across 5 fixed same-config seeds, in a single constructive pass — without local search or per-instance refinement — with inference times between 0.1 and 3.1 seconds per instance. The model is trained on 400 synthetic instances with n ≤ 25. The most recent published neural reference reporting full QAPLIB results is SAWT [Tan and Mu, ICML 2024], a learn-to-improve (L2I) method that iteratively refines an assignment per instance after training on 5,120 synthetic instances, and reports a 26.8% mean gap. CIMP operates in a strictly more constrained learn-to-construct (L2C) regime and from a much smaller training budget, yet attains a substantially lower mean gap. The pre-block builds a state-dependent facility–location representation before scoring, a hybrid context-biased decoder mixes learned content similarity with QAP-specific algorithmic bias terms, and a sequence-distribution distillation objective is derived from a diversity-filtered pool of heuristic teachers. Family-wise and size-conditioned analyses, supported by a secondary normalized score, indicate that the gain is broadly consistent across QAPLIB families and across instance sizes beyond the training range.

Transformers rely on a growing key–value (KV) cache to store context, causing linear memory growth and limited generalization beyond the training window. We show that contextual influence can instead be viewed as updates to a low-dimensional subspace in the model’s weight space, meaning context can be represented as a bounded, low-rank parameter update rather than token-level storage. Motivated by this, we propose PARADE, which internalizes context as parametric memory. It compresses past inputs into a fixed set of weight-space basis vectors and composes them via query-dependent routing into per-query weight updates, enabling streaming inference with constant memory and effectively unbounded context. Experiments show that PARADE matches full attention on long-context tasks and outperforms efficient baselines, especially when relevant information lies far beyond the attention window, while improving as context exceeds the model’s training length.

Missing data is pervasive in real-world datasets, arising from sensor failures, image occlusions, and incomplete reporting. Imputation is a conditional generation problem that must preserve observed entries exactly while producing plausible values for missing ones. Yet most diffusion-based imputers adopt Gaussian noising whose global perturbations are structurally misaligned with this local hard constraint, so consistency is typically enforced through conditional training, clamping, projection, or repainting-style corrections. We address this mismatch with Continuously-Augmented Hybrid Masked Diffusion (CAHMD), a masked diffusion framework whose conditional sampler leaves observed coordinates unchanged by construction. To extend masked diffusion beyond purely discrete data, CAHMD attaches coordinate-wise auxiliary continuous latents to masked entries, providing a denoising channel before values are revealed in continuous and hybrid data. For incomplete training data, we introduce observed-mask gating, an observed-data masked reconstruction principle that uses available entries as pseudo-missing targets and avoids repeated EM-style imputation steps. Across tabular, image, and image--caption benchmarks, CAHMD gives a strong reconstruction--perception--efficiency trade-off while reducing the training overhead of EM-style self-imputation and the inference overhead of repainting-style inner loops.

Efficient sampling from complex probability distributions is an important task across Bayesian inference and certain classes of deep generative models. Markov chain Monte Carlo (MCMC) remains one of the most widely used tools for this. However, even state-of-the-art samplers, such as the NUTS implementations in Stan and Turing.jl, can converge slowly on ill-conditioned, heavy-tailed, or near-degenerate targets. We identify that sampling error can be decomposed into two terms: mismatch in the potential energy (PE) marginal, and average conditional mismatch across PE level sets. We propose *Contour Monte Carlo (CMC)*, which corrects each of these in turn: first by drawing energies $u_i$ from the PE marginal, and then by sampling states on the corresponding level sets. For a broad family of radial targets, including the Gaussian and Student's $t$ distributions, both stages admit closed-form draws, so CMC is exact and requires no Markov chain. Beyond this, for general targets, we estimate the PE marginal using umbrella sampling with MBAR, then run constrained MCMC on the sampled level sets, initialised from energy-matched umbrella samples (which tend to lie in regions of high conditional mass). We show that this procedure essentially reduces the global sampling problem to approximating the one-dimensional PE marginal, after which posterior draws can be generated readily on demand. Furthermore, the level-set chains are embarrassingly parallel, and therefore well-suited to modern accelerator hardware. On a suite of challenging benchmarks -- composed of `posteriordb` models, synthetic targets, and real-world inference problems -- CMC converges substantially faster than HMC and NUTS baselines at a fixed computational budget.


Contrastive Adversarial Training for Robust Graph Neural Networks under Label Poisoning

Manshika C Bissessur ⋅ Melis Ilayda Bal ⋅ Michael Muehlebach

Graph Neural Networks (GNNs) are effective for modeling relational data but are vulnerable to label poisoning, where a small number of corrupted training labels can propagate errors across the graph via message-passing. Despite this risk, defenses against label poisoning remain underexplored: existing methods are primarily designed for label noise and often rely on robust losses or heuristic data cleaning that fail to distinguish adversarial poisoning from naturally hard examples. In this paper, we propose CoLAT (Contrastive Label-flipping for Adversarial Training), a novel contrastive adversarial training framework tailored for robust node classification on graphs. The core of our approach is a structure-aware detection model that uses contrastive learning to identify label–structure inconsistencies. Unlike prior contrastive methods that focus on representation learning, CoLAT leverages contrastive embeddings to select high-risk nodes that guide adversarial label-flipping during training. This alternating optimization not only performs structure-aware adversarial training but also implicitly sanitizes corrupted labels, improving robustness against diverse attacks from the literature. Extensive experiments on multiple benchmarks show that CoLAT consistently outperforms existing defenses under various poisoning intensities while scaling efficiently to large graphs.


Co-optimization for Adaptive Conformal Prediction

Xiaoyi Su ⋅ Zhixin Zhou ⋅ Rui Luo

In regression, conformal prediction often suffers from inefficiency under heteroscedasticity and skewness due to fixed, non-adaptive interval centering. We propose CoCP, a new method that parameterizes prediction intervals by a center $m(x)$ and a radius $h(x)$, and alternates between learning $h(x)$ via smooth quantile regression on folded residuals and refining $m(x)$ using a smooth interval loss. This corrects mis-centering and drives the interval toward high-density regions. Beyond finite-sample marginal validity via split-conformal calibration, we prove that CoCP asymptotically achieves conditional coverage and optimal interval length if the base estimators are consistent. Experiments demonstrate that CoCP yields tighter intervals and achieves state-of-the-art conditional coverage reliability.


COSMIO: A Benchmark for Cross-Survey Modality Imputation

Dichang Zhang ⋅ Yixuan Shao ⋅ Jiali Cui ⋅ Haotian Yin ⋅ Yuanpeng Liu ⋅ Jiaiq Deng ⋅ Xiang Gao ⋅ Zhiqiang Lao ⋅ Heather Yu ⋅ Simon Birrer ⋅ Dimitris Samaras

Large-scale astronomical surveys define a heterogeneous, partially observed multi-survey learning setting: each object is observed by only a subset of surveys, while survey observations differ in resolutions, wavelength coverage, noise characteristics, and footprints. We formalize this as the problem of cross-survey modality imputation and introduce COSMIO, a multi-survey benchmark of aligned observations obtained by forward-modeling shared astrophysical scenes across LSST, Euclid, and Roman under survey-specific instrument characteristics. The benchmark enables systematic evaluation of imputation across source-target survey combinations and missing-survey patterns. We evaluate a broad range of machine learning methods for imputation. Our results reveal a clear domain gap: methods developed for existing imputation settings transfer poorly to the astronomical survey setting, indicating that cross-survey imputation constitutes a distinct learning problem. Beyond common-object imputation, we identify missing-modality rare-event imputation as a central challenge. This regime arises in the early stages of new surveys, when rare but scientifically valuable phenomena may be available only from old surveys. We use strong gravitational lensing, where light from a background galaxy is gravitationally deflected by a foreground galaxy, as a representative rare-event setting. We evaluate model performance on zero-shot imputation of gravitational lenses and on downstream scientific tasks using the imputed lens observations. On our constructed evaluation set of galaxy–galaxy strong lenses, performance degrades under rare-event distribution shift, and this degradation further propagates to the downstream lens-detection task. Our results establish cross-survey imputation as a distinct and challenging machine learning problem, highlighting the need for methods that are robust to distribution shift and aligned with scientific objectives. Code and data are available at https://anonymous.4open.science/r/COSMI-3671/.


Cost-Aware Learning

Clara Mohri ⋅ Amir Globerson ⋅ Haim Kaplan ⋅ Tomer Koren ⋅ Yishay Mansour

We consider the problem of Cost-Aware Learning, where sampling different components of a finite-sum objective incurs different costs. The objective is to reach a target error while minimizing the total cost. We propose Cost-Aware SGD, which uses a distribution based on gradient norms and costs to sample components. We provide a thorough analysis of this algorithm, including cost-improvement bounds over baselines, a characterization of distribution proxy sub-optimality, and a lower bound. We apply our theoretical insights to reinforcement learning with language models, where the computational cost of sequence-level policy gradients varies with length. We find that the advantage magnitude serves as a high-fidelity proxy for gradient norms, and use this to introduce Cost-Aware GRPO. Empirical results on 1.5B, 4B, and 8B LLMs demonstrate that this algorithm significantly reduces the tokens used in policy optimization while matching or exceeding baseline accuracy.

Large-scale human demonstrations have been a key ingredient in training interactive agents, from robot policies to computer-use agents and game world models. But for multi-agent world modeling, the open data landscape is missing a crucial combination: expert humans acting together in the same evolving world, observed from every participant’s viewpoint, with precise controls and state. Internet-scale video captures human behavior, but its controls are inferred and its viewpoints are not synchronized; simulator and policy rollouts provide exact labels and all-player views, but not strategic, collaborative human play. We present CounterStrike-1K, a public benchmark dataset built from professional Counter-Strike 2 match replay files. It contains 1,490 rendered player-view hours across seven maps, with every released round synchronized across all ten active player viewpoints. Each view is paired with audio, replay-derived game controls, player state, world position, and sparse game event annotations. Because each round records expert human teams coordinating under partial observability, CounterStrike-1K tests whether models can simulate not just game physics or single-view visual plausibility, but skilled decisions and their consequences across all players’ views. We provide fixed splits and benchmarks for action-conditioned prediction, cross-POV retrieval, multi-POV action coverage, and shared-state recovery. By releasing round-complete multi-perspective demonstrations at this scale, CounterStrike-1K provides a grounded training and evaluation testbed for multi-agent world-model research.


CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving

Jingqi Wang ⋅ minqing huang ⋅ Zihan Liang ⋅ Yujiao Xiang ⋅ JiaJie Huang ⋅ Zhi Xu ⋅ Feiyang Tan ⋅ Hangning Zhou ⋅ Mu Yang ⋅ Gong Chen

Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms still struggle to provide planning-oriented intermediate representations: textual Chain-of-Thought (CoT) fails to preserve continuous spatiotemporal structure, while latent world reasoning remains difficult to use as a direct condition for action generation. In this paper, we propose CoWorld-VLA, a multi-expert world reasoning framework for autonomous driving, where world representations serve as explicit conditions to guide action planning. CoWorld-VLA extracts complementary world information through multi-source supervision and encodes it into expert tokens within the VLA, thereby providing planner-accessible conditioning signals. Specifically, we construct four types of tokens: semantic interaction, geometric structure, dynamic evolution, and ego trajectory tokens, which respectively model interaction intent, spatial structure, future temporal dynamics, and behavioral goals. During action generation, CoWorld-VLA employs a diffusion-based hierarchical multi-expert fusion planner, which is coupled with scene context throughout the joint denoising process to generate continuous ego trajectories. Experiments show that CoWorld-VLA achieves competitive results in both future scene generation and planning on the NAVSIM v1 benchmark, demonstrating strong performance in collision avoidance and trajectory accuracy. Ablation studies further validate the complementarity of expert tokens and their effectiveness as planning conditions for action generation.

Domain Incremental Learning (DIL) aims to continuously adapt to new domains while retaining the knowledge of previous domains. Existing DIL methods are generally built upon an idealized assumption of clean supervision. However, real-world data streams are often affected by label noise, giving rise to the more challenging Noisy Domain Incremental Learning (N-DIL) scenario. Under this setting, not only is intra-domain knowledge acquisition hindered, but inter-domain knowledge conflicts are also exacerbated, which amplifies catastrophic forgetting. To address these challenges, we propose a novel Cross-Domain Knowledge Separation and Positive Transmission (ST-Prompt) framework. Specifically, to mitigate inter-domain knowledge conflicts, a Cross-domain Prompt Knowledge Contrastive Isolator is developed to enhance domain-wise knowledge separation, thereby mitigating the cross-domain knowledge interference during inference. Furthermore, to improve intra-domain knowledge acquisition, a Synergistic Knowledge Transmission scheme is introduced, which extracts reliable knowledge from previous domains to facilitate noisy data learning in the current domain. These two components are mutually reinforcing, jointly promoting effective cross-domain knowledge learning and utilization. Extensive experiments on diverse benchmarks demonstrate that ST-Prompt achieves state-of-the-art performance. Our code will be released.


Cross-Model Circuit Discovery

Harrish Thasarathan ⋅ Matthew Kowal ⋅ Thomas Fel ⋅ Kosta Derpanis

Consider two large vision models. They process the same image and both correctly predict the class ``rabbit.'' How much of the circuit computation along the way was shared? Model diffing offers a natural lens on this question. So far, however, it has largely operated on a single layer and at the level of representations rather than circuits. In this work, we introduce Universal Circuits (UCs), enabling model diffing at the circuit level. Specifically, we extend CLTs across both layers and models, with losses that encourage sparsity for interpretability and output fidelity for faithfulness. We train UCs between pairs of standard large vision models. We find that a compact cross-model intersection of pruned class circuits, typically a few hundred universal (shared) features per pair, produces $89$-$98\%$ of full-circuit classification accuracy in both models, while a complementary set of universal features (hundreds per pair) is kept by only one model's circuit, reflecting how each model weights shared concepts differently in its own representations. To demonstrate a downstream use of UC, we perform model \textit{surgery}: a class is successfully transferred from one model to another with no gradient steps taken in the new model. More broadly, our work demonstrates that large vision models leverage shared multi-layer algorithms for downstream performance, and that these algorithms can be discovered, compared, and reused across models.


CRUMB: Efficient Prior Fitted Network Inference via Distributionally Matched Context Batching

Jamie Heredge ⋅ Mattia Jacopo Villani ⋅ Pranav Deshpande ⋅ Akshay Seshadri ⋅ Niraj Kumar

Prior-fitted networks (PFNs) are a promising class of tabular foundation models that perform in-context learning, whereby the entire labelled training set is supplied as context, and predictions for test queries are produced in a single forward pass. However, the quadratically scaling self-attention mechanism in many PFN architectures makes inference prohibitive for very large training datasets. We propose CRUMB (Clustered Retrieval Using Minimised-MMD Batching), a three-stage inference wrapper that (i) clusters the test queries, (ii) selects a small, distributionally matched training subset for each cluster by greedily minimising the maximum mean discrepancy (MMD), and (iii) runs exact PFN inference on each reduced-context batch. CRUMB is architecture-agnostic and requires no retraining. On the 51-dataset TabArena benchmark, evaluated across three PFN architectures (TabPFNv2, TabICLv1, TabICLv2), we show that CRUMB outperforms similar state-of-the-art context selection strategies. We also show that CRUMB is resilient to covariate drift, as the MMD-minimisation step naturally helps align the training context distribution to match the current test batch distributions.


CryptanalysisBench: Can LLMs do cryptanalysis?

Lukas Fluri ⋅ Avital Shafran ⋅ Nicholas Carlini ⋅ Matthew Jagielski ⋅ Milad Nasr ⋅ Orr Dunkelman ⋅ Eyal Ronen ⋅ Florian Tramer

Cryptanalysis---the task of finding attacks against cryptographic schemes---sits at the intersection of mathematical reasoning and programming, two areas where LLMs have made rapid progress. Just like math and programming, cryptanalytic attacks are formally defined and can be unambiguously verified. This raises the question if cryptanalysis might experience a similar acceleration in progress through LLMs. In this paper we introduce $\texttt{CryptanalysisBench}$, a benchmark of 113 tasks spanning six families of cryptographic primitives (block ciphers, hash functions, etc) drawn primarily from four NIST standardization competitions. Each task asks an agent to break an implementation of a cryptographic primitive by winning a formal security game. The benchmark has three tiers: (1) primitives with known practical breaks; (2) scaled-down variants of primitives without one; (3) a challenge set of unbroken production primitives. We evaluate Claude Opus 4.7 and GPT-5.5: both solve a majority of Tier 1 (71.4\% and 78.5\%, respectively), while Tier 2 remains largely out of reach. For Tier 1 successes, we provide a fine-grained analysis distinguishing paper recall from source-level rediscovery. We release $\texttt{CryptanalysisBench}$ both as a forecasting tool to track when AI cryptanalytic capability becomes a serious factor, and as scaffolding for subjecting candidate schemes to more attacks before they are deployed.


CS-Dialogue: A 104-Hour Dataset of Spontaneous Mandarin-English Code-Switching Dialogues for Speech Recognition

Jiaming Zhou ⋅ Yujie Guo ⋅ Shiwan Zhao ⋅ Haoqin Sun ⋅ Hui Wang ⋅ Jiabei He ⋅ Aobo Kong ⋅ Shiyao Wang ⋅ Yong Qin

Code-switching (CS), the alternation between two or more languages within a single conversation, presents significant challenges for automatic speech recognition (ASR) systems. Existing Mandarin-English code-switching datasets often suffer from limitations in size, spontaneity, and the lack of full-length dialogue recordings with transcriptions, hindering the development of robust ASR models for real-world conversational scenarios. This paper introduces CS-Dialogue, a novel large-scale Mandarin-English code-switching speech dataset comprising 104 hours of spontaneous conversations from 200 speakers. Unlike previous datasets, CS-Dialogue provides full-length dialogue recordings with complete transcriptions, capturing naturalistic code-switching patterns in continuous speech. We describe the data collection and annotation processes, present detailed statistics of the dataset, and establish benchmark ASR performance using state-of-the-art models. Our experiments, using Transformer, Conformer, and Branchformer, demonstrate the challenges of code-switching ASR, and show that existing pre-trained models such as Whisper still have the space to improve. The CS-Dialogue dataset will be made freely available for all academic purposes.


CSFlow: Aligning Flow Matching with Human Contrast Sensitivity

Malgorzata Galinska ⋅ Bart Pogodzinski ⋅ Jan Eric Lenssen

We introduce Contrast Sensitive Flow (CSFlow), a weighting scheme that connects the human eye's Contrast Sensitivity Function (CSF) to the iterative denoising steps of flow matching. Because real-world images concentrate signal at low spatial frequencies, these components reach high signal-to-noise ratio earlier during continuous diffusion than high-frequency components. When generating images with diffusion or flow matching models, this induces a soft autoregressive structure in Fourier space, where coarse image content stabilizes before fine detail. Meanwhile, the human visual system is unequally sensitive to spatial frequencies: very low and very high frequencies require significantly higher contrast to be perceived. We for the first time merge these observations through two contributions: (1) a metric that estimates which frequencies are generated at each reverse flow interval and (2) timestep weights obtained by aligning the frequencies generated at each noise level with human contrast sensitivity. We validate our contributions experimentally showing that these weights can improve generative performance by lowering FID by 4.7%, increasing Inception Score by 2.2% and improving GenEval scores by 2.5% using inference-only timestep modification or short fine-tuning. Qualitatively, we find that our CSFlow weights lead to better visual realism and less cartoonish appearance of generated images.


CSO: Refining Robotic Policies via Skill Distribution Alignment and Skill-Grained Optimization

Zhiyuan Xiang ⋅ Xiang Deng ⋅ Wanjie Tao ⋅ Xu Chen ⋅ Zhao Jielun ⋅ Weili Guan

Discretizing continuous actions into skills using methods like VQ-VAE has emerged as a powerful paradigm for robotic manipulation. However, the quantization errors in discretizing continuous actions yield a suboptimal training distribution for the prior, degrading its performance. While reinforcement learning offers a path for refinement, its direct application is challenging, suffering from unstable encoder updates and a granularity dilemma in importance sampling. To address these challenges, we introduce Cascaded Skills Optimization (CSO), a two-stage post-training framework. First, to rectify the initial policy's suboptimal distribution, CSO employs Rejection-Sampling Supervised Fine-tuning to align the model's observation-to-skill mapping with the distribution of successful online trajectories via supervised fine-tuning. Second, to resolve the granularity dilemma, CSO introduces Skills Policy Optimization, which computes an independent, clipped importance ratio for each skill, enabling more stable and efficient updates. Our post-training strategy delivers highly competitive performance on challenging benchmarks like LIBERO and RoboTwin, with its effectiveness further validated on a physical robot.

Test-Time Reinforcement Learning (TTRL) constructs pseudo-supervisory signals by sampling reasoning trajectories during inference, enabling online adaptation to current inputs. However, existing research predominantly focuses on single-task settings, overlooking continuous updates across task streams. Theoretical and empirical analyses reveal two coupled challenges in this setting: (1) Error Accumulation: Majority voting tends to treat erroneous consensus as pseudo-labels, with such bias amplifying across sequential parameter updates; (2) Catastrophic Forgetting: Gradients from new tasks conflict with historical knowledge, causing acquired reasoning patterns to be overwritten. These issues form a positive feedback loop: pseudo-label noise exacerbates gradient interference, while degradation of historical knowledge reduces subsequent trajectory quality, jointly driving performance deterioration during continual adaptation. To address this, we propose the CTRL framework: first, we introduce a Process Reward Model (PRM) to replace outcome-level voting, filtering high-confidence trajectories via fine-grained process scoring to suppress pseudo-label bias; second, we design a cognitive anchor-driven soft gradient correction module that dynamically recomputes anchor gradients from a historical trajectory memory pool and projects out update components negatively correlated with historical knowledge, constraining the parameter update trajectory. Experiments demonstrate that CTRL significantly outperforms existing TTRL baselines on continuous task streams, exhibiting stronger robustness in blocking error propagation and preserving historical reasoning patterns.

We cast image segmentation refinement as a \textit{discrete} diffusion process. To solve this task we introduce Cyclic Discrete Diffusion (CDD), a novel forward–reverse formulation with a cyclic label topology that maps a discrete label set to the uniform distribution and back. Instead of injecting Gaussian noise as is done for continuous diffusion, CDD evolves segmentation masks through label jumps: at each step, every pixel either remains in its current class or jumps to the next label with a fixed rate. This jump-or-stay mechanism defines a finite-state Markov process with bounded forward and reverse steps, featuring analytically consistent and tractable reverse dynamics. CDD operates natively in a multi-class label space, enabling joint refinement of all classes in a single pass without requiring class-wise decomposition, as commonly needed in prior binary refinement approaches. The proposed formulation admits a finite and controllable number of refinement steps, for which we provide a theoretical characterization based on spectral gap analysis. We evaluate CDD across diverse segmentation tasks, including multi-class semantic, instance-level, and binary segmentation refinement. Our method consistently improves both region accuracy and boundary quality across multiple backbones and datasets.


D$^2$Quant: Accurate Low-bit Post-Training Weight Quantization for LLMs

Xianglong Yan ⋅ chengzhu bao ⋅ Zhiteng Li ⋅ Tianao Zhang ⋅ Shaoqiu Zhang ⋅ Ruobing Xie ⋅ Xingwu Sun ⋅ Yulun Zhang

Large language models (LLMs) deliver strong performance, but their high compute and memory costs make deployment difficult in resource-constrained scenarios. Weight-only post-training quantization (PTQ) is appealing, as it reduces memory usage and enables practical speedup without low-bit operators or specialized hardware. However, accuracy often degrades significantly in weight-only PTQ at sub-4-bit precision, and our analysis identifies two main causes: (1) down-projection matrices are a well-known quantization bottleneck, but maintaining their fidelity often requires extra bit-width; (2) weight quantization induces activation deviations, but effective correction strategies remain underexplored. To address these issues, we propose D$^2$Quant, a novel weight-only PTQ framework that improves quantization from both the weight and activation perspectives. On the weight side, we design a Dual-Scale Quantizer (DSQ) tailored to down-projection matrices, with an absorbable scaling factor that significantly improves accuracy without increasing the bit budget. On the activation side, we propose Deviation-Aware Correction (DAC), which incorporates a mean-shift correction within LayerNorm to mitigate quantization-induced activation distribution shifts. Extensive experiments show that D$^2$Quant achieves superior weight-only PTQ performance at sub-4-bit precision. We will release the code and quantized models of D$^2$Quant.


DACA-GRPO: Denoising-Aware Credit Assignment for Reinforcement Learning in Diffusion Language Models

Amin Karimi Monsefi ⋅ Dominic Culver ⋅ Nikhil Bhendawade ⋅ Lokesh Boominathan ⋅ Manuel R Ciosici ⋅ Yizhe Zhang ⋅ Irina Belousova

Diffusion large language models are a compelling alternative to autoregressive models, yet existing RL methods for diffusion treat all denoising steps as equally important and rely on biased, high-variance likelihood estimates. We identify two fundamental weaknesses: the absence of temporal credit assignment across the denoising trajectory, and the systematic bias of mean-field likelihood estimates used for policy optimization. To address these, we propose Denoising-Aware Credit Assignment for GRPO (DACA-GRPO), a lightweight, plug-and-play enhancement for any GRPO-style trainer. DACA-GRPO introduces two complementary mechanisms: Denoising Progress Scores, which extract per-token importance weights from intermediate predictions at no additional forward cost, and Stratified Masking Likelihood, which partitions token positions into strata so that each token is predicted with most of the sequence as context, reducing the mean-field bias. Applied on top of three GRPO base methods, DACA-GRPO achieves consistent improvements across seven benchmarks spanning mathematical reasoning, code generation, constraint satisfaction, and constrained generation, with gains of up to 5.6pp on math reasoning, 7.4pp on code generation, 36.3pp on constraint satisfaction, and 5.9pp on JSON schema adherence.


DAGA: Dynamic Attention-Guided Adaptation for Self-Supervised Vision Transformers

Tianjian Zhou ⋅ Jiang Jie ⋅ Yishan Li ⋅ Yifei Zhang

Self-supervised Vision Transformers in the DINO family encode rich semantic structure within their self-attention maps, yet existing parameter-efficient fine-tuning (PEFT) methods apply static, content-agnostic transformations that ignore this internal signal. We present DAGA (Dynamic Attention-Guided Adaptation), a PEFT module that repurposes a frozen backbone's emergent attention maps as instance-specific guidance: per-image attention is converted into channel and spatial gates that condition a lightweight bottleneck adapter. With only 1.4\% additional parameters, DAGA reaches 85.9\% top-1 on ImageNet-1K with DINOv3-ViT-B and surpasses recent PEFT baselines, with strong gains on dense prediction (+12.1\% mIoU on ADE20K) and fine-grained classification (+1.6 / +3.5 on CUB-200 / FGVC-Aircraft). DAGA's effectiveness is tied to the semantic quality of the backbone's attention: gains are large on DINOv2/v3 and iBOT but small on MAE, CLIP, and DeiT, positioning it as the first PEFT module to operationalize emergent SSL attention as in-loop guidance and complementing self-distillation methods that refine attention itself. Anonymized code: https://anonymous.4open.science/r/ID1680.

Recent work has identified a set of emergent symbolic mechanisms that support abstract reasoning in large language models, but it remains unclear what factors drive the emergence of these mechanisms. To address this question, we trained neural networks from scratch on abstract sequence tasks, varying both architectural and data distributional factors, and investigated the effects of these factors on the emergence of symbolic mechanisms. Using a combination of representational, attentional, and casual mediation analyses, we first confirmed that transformer language models trained from scratch on our task developed symbolic mechanisms, partially capturing the set of mechanisms learned by large-scale pretrained models. We then investigated the effect of data diversity, operationalized as vocabulary size, on the emergence of these mechanisms, finding that mechanistic and behavioral signatures of emergent symbol processing scaled with data diversity, including downstream measures of systematic (i.e., out-of-distribution) generalization. Surprisingly, we found that the emergence of symbolic mechanisms was not driven by architectural inductive biases, as the same mechanistic and behavioral signatures emerged in modified transformer architectures and even multilayer perceptrons. These results suggest that data diversity, rather than architectural inductive biases, is the primary driver of emergent symbolic computation in neural networks.


Dataset Mismatch Matters in Group Relative Policy Optimization for Reinforcement Learning from Verifiable Rewards

Zhen-Yu Zhang ⋅ Jiandong Zhang ⋅ Huaxiu Yao ⋅ Gang Niu ⋅ Masashi Sugiyama

Recently, reinforcement learning from verifiable rewards (RLVR) is practically very important, where the leading approach is group relative policy optimization (GRPO) that typically faces a trade-off between rollout efficiency and downstream performance. To improve efficiency while preserving performance, existing selective rollout methods mainly exploit only source-side training dynamics to optimize the discrete distribution over source prompts. However, in practice, the source training dataset is rarely identically distributed with the target test dataset. When such a mismatch exists, due to legal or privacy constraints, it is usually not acceptable to share target prompts themselves with the RLVR trainer. Nevertheless, it is still possible to share some target-side feedback such as accuracy on a small held-out data. In this paper, we propose learning-to-reweight rollout (L2RR) that explicitly incorporates target-side performance feedback into GRPO. Specifically, L2RR learns a discrete sampling distribution over source prompts by optimizing a bi-level objective that maximizes accuracy on a held-out target dataset, while simultaneously penalizing rollouts that provide little incremental learning information. Empirical studies on six math reasoning benchmarks and three model scales show that L2RR improves target-domain performance and increases rollout efficiency.


DAWN: Dependency-Aware Fast Inference for Diffusion LLMs

Lizhuo Luo ⋅ Zhuoran Shi ⋅ Jiajun Luo ⋅ Zhi Wang ⋅ Shen Ren ⋅ Wenya Wang ⋅ Tianwei Zhang

Diffusion large language models (dLLMs) have shown advantages in text generation, particularly due to their inherent ability for parallel decoding. However, constrained by the quality--speed trade-off, existing inference solutions adopt conservative parallel strategies, leaving substantial efficiency potential underexplored. A core challenge is that parallel decoding assumes each position can be filled independently, but tokens are often semantically coupled. Thus, the correct choice at one position constrains valid choices at others. Without modeling these inter-token dependencies, parallel strategies produce deteriorated outputs. Motivated by this insight, we propose DAWN, a training-free, dependency-aware decoding method for fast dLLM inference. DAWN extracts token dependencies and leverages two key observations: (1) masked positions dependent on unmasked certain positions become more reliable; (2) coupled masked positions strongly influence each other's predictions. Given those findings, DAWN leverages a dependency graph to select more reliable unmasking positions at each iteration, achieving high parallelism with negligible loss in generation quality. Extensive experiments across multiple models and datasets demonstrate that DAWN speedups the inference by 1.80–8.06× over baselines while preserving the generation quality.


DCVD: Dual-Channel Cross-Modal Fusion for Joint Vulnerability Detection and Localization

Wenxin Tang ⋅ Wenbin Li ⋅ Junliang Liu ⋅ Jingyu Xiao ⋅ Xi Xiao ⋅ Mingzhe liu ⋅ Jinlong Yang ⋅ Xuan Liu ⋅ Yuehe Ma ⋅ Wang Luo ⋅ Qing Li ⋅ Lei Wang ⋅ Peng Xiangli

Software vulnerability detection plays a critical role in ensuring system security, where real-world auditing requires not only determining whether a function is vulnerable but also pinpointing the specific lines responsible. However, existing approaches either rely on a single information source---sequential, structural, or semantic---failing to jointly exploit the complementary strengths across modalities, or treat statement-level localization merely as a byproduct of function-level detection without explicit line-level supervision. To address these limitations, we propose DCVD (Dual-Channel Cross-Modal Vulnerability Detection), a unified framework that performs joint function-level detection and statement-level localization. DCVD extracts control-dependency and semantic features through two parallel branches and integrates them via contrastive alignment coupled with bidirectional cross-attention, effectively bridging the cross-modal representation gap. It further introduces explicit supervision signals at both the function and statement levels, enabling collaborative optimization across the two granularities. Extensive experiments on a large-scale real-world vulnerability benchmark demonstrate that DCVD consistently outperforms state-of-the-art methods on both function-level detection and statement-level localization.

Communication enhances cooperative multi-agent reinforcement learning (MARL) under partial observability. In real networks, however, each agent's available bandwidth varies over time due to changing link quality, interference, and congestion. Most existing MARL communication policies decide whether to transmit based on message importance or novelty, or are trained under fixed bandwidth conditions. As a result, they cannot adapt sending frequency to bandwidth variations, leading to packet loss when bandwidth decreases or missed information when bandwidth recovers. To address this, we propose Dual-Decoupled Adaptive Communication (DDACOM)---a framework that decouples message sending and message generation from the action policy. The message-sending module ranks each message by its deviation from teammates' information and the sender's history, and uses a dynamic token bucket to convert the probe-estimated safe rate into a real-time token budget, enabling the communication frequency to track the safe rate online. The message-generation module is trained independently with receiver-side attention feedback and counterfactual marginal value, ensuring each transmitted message meaningfully contributes to cooperation. On SMAC and MPE, DDACOM achieves leading cooperative performance among baseline methods. Under dynamic network conditions in ns-3 simulation, DDACOM sustains over (95\%) safe-rate utilization and packet loss below (0.3\%) while reducing communication frequency to less than half of full broadcast, confirming its adaptability to runtime network variation.


DDx-TRACE: A Benchmark for Medical Diagnostic Trajectories in VLMs

Jiazhen Pan ⋅ Weixiang Shen ⋅ Jun Li ⋅ Julian Canisius ⋅ Felix Bitzer ⋅ Paula Roßmüller ⋅ Jiancheng Yang ⋅ Virginie Kreutzinger ⋅ Daniel Rueckert ⋅ Benedikt Wiestler

Medical diagnosis is not a single prediction from a fully specified vignette. It is a sequential workup: clinicians decide what evidence to obtain, revise a differential diagnosis, and stop when the diagnosis is sufficiently supported. Most medical AI benchmarks instead reveal the relevant context upfront and score only the final answer, making unsupported correct guesses, premature closure, inefficient workups, and poor uncertainty updating invisible. We introduce DDx-TRACE, a physician-adjudicated benchmark for multimodal neuroradiology that evaluates diagnostic trajectories under hidden evidence over 211 challenging cases. Each case begins with limited clinical history; models request imaging studies in free form, receive matched image bundles when available, update a probabilistic differential diagnosis after each turn, and stop with a localized final diagnosis. Evaluating state-of-the-art VLMs, we find that final diagnosis scores can substantially misrepresent workup quality: models may guess plausible diagnoses without essential evidence, request useful studies but misinterpret raw images, or acquire evidence inefficiently while updating uncertainty poorly. Controlled evidence variants isolate bottlenecks in planning, visual evidence extraction, and downstream differential reasoning. DDx-TRACE shifts medical AI evaluation from final answers to evidence-supported diagnostic trajectories.


Decentralized Multi-Goal Multi-Agent Pathfinding with Spatial Prior and Neighbor Intent Prediction

Xuening Wang ⋅ Bin Wang ⋅ Yucen Gao ⋅ Sihui Li ⋅ Xiaochun Yang

As a variant of the Multi-Agent Pathfinding (MAPF) problem, Multi-Goal Multi-Agent Pathfinding (MG-MAPF) requires planning conflict-free paths for a team of agents to visit sequences of preassigned goal vertices. Existing search-based methods suffer from poor scalability due to their reliance on centralized control, while decentralized learning-based approaches struggle with long-horizon planning under partial observability. Furthermore, existing paradigms rarely address joint goal ordering across agents, where uncoordinated goal sequences cause multiple agents to simultaneously converge on the same region, increasing congestion and conflicts. To address these challenges, we propose the Intent-aware Multi-Goal Multi-Agent Pathfinding (IMMP), a hierarchical framework for decentralized MG-MAPF that explicitly models spatial congestion at the high level and neighbor intent at the execution level. The high-level planner encodes a spatial prior via spatially-balanced goal ordering and path planning, reducing conflict pressure passed to the low-level policy. The low-level module resolves real-time conflicts via a reinforcement learning policy augmented with a neighbor intent prediction module, trained end-to-end through a prediction-guided collaborative proximal policy optimization. Extensive experiments demonstrate that IMMP achieves strong scalability and generalization, Extensive experiments demonstrate that IMMP achieves strong scalability and generalization, reducing SoC by an average of 19.77\% compared to the best-performing decentralized baseline, while maintaining robust performance. Our code is available at https://anonymous.4open.science/r/IMMP-CB2E/

Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents. It does not check whether retrieved evidence is attributed to the correct entity. A clinical RAG response can pass every automated check (zero hallucinations, near-perfect faithfulness, real citations) while presenting drug $Y$'s clinical evidence as evidence about queried drug $X$. We term this **deceptive grounding** (DG): a failure invisible to faithfulness, hallucination, and citation checks because every claim is sourced from a real document, about the wrong entity. Using a controlled factorial benchmark across 13 models, we find DG rates spanning 8--87\% at peak adversarial conditions. Medical and biomedical fine-tuned models reach up to 86.7\%; domain specialization amplifies the failure rather than mitigating it. A controlled ablation identifies the mechanism: removing entity-specific clinical evidence from retrieved documents eliminates entity-attribution failure entirely, shifting all failures to confabulation. The two failure modes respond to the same trigger, taking different paths. Production measurement across 740 drug--disease pairs finds 7.8\% overall DG in a deployed RAG system, rising to 13.6\% for recently approved drugs. Entity-attribution verification (checking that cited evidence applies to the queried entity) detects DG at 92.3\% precision and 95.8\% DG recall; no existing framework implements it.


Decomposed Graded Verifier for Generative World Modeling

Bowei Liu ⋅ Xinchen Zhang ⋅ Xuhuan Li ⋅ Kaian Jiang ⋅ Jingwei Liu ⋅ Congyang Zhao ⋅ Yifei Qian ⋅ Junhong Liu ⋅ Zhiheng Li ⋅ Tianlin Zhang ⋅ Yujiu Yang ⋅ Yang Cai ⋅ Ling Yang

Generative visual verifiers play a crucial role in advancing multimodal intelligence systems. In this paper, to enhance the robustness of visual verifiers in complex and realistic open-world scenarios, we move beyond conventional binary verifiers that produce only True/False discrete judgments. Instead, we introduce a verifier that automatically decomposes complex prompts or environmental rules into fine-grained components and assigns pointwise scores, thereby providing a continuous and more informative critique signal. Under a general reinforcement learning framework, we first empirically demonstrate that the proposed pointwise decomposed verifier significantly outperforms binary verifiers in traditional text-to-image scenarios. We then extend our investigation to a novel paradigm: generative world modeling, where generative models are leveraged for world simulation and world reasoning. From a theoretical standpoint, we further show that when ground-truth images are easy to obtain, a pairwise verification paradigm yields more accurate critiques than the pointwise formulation. Empirical results on diverse world-modeling tasks, such as Maze and Sudoku, further validate the effectiveness of our approach. Overall, our findings suggest a key insight: the transition from generative models to world-modeling agents critically hinges on the availability of accurate pairwise decomposed visual verifiers.


Decomposing Conformal Uncertainty: Calibration- and Instance-Driven Feature Attribution

Sangyeon Cho ⋅ Minyoung Cho ⋅ Jungsoo Kim ⋅ Sujeong Oh ⋅ Sanghack Lee

In high-stakes settings, decision-makers using conformal prediction (CP) read the predictive interval's width as their operational uncertainty signal. Local feature attribution methods, applied to that width, silently conflate two semantically distinct contributions: a calibration-driven baseline that depends on the held-out calibration set, and an instance-driven term that depends on the test point. We propose a two-game Shapley construction---one game over feature coalitions at the test point, one over feature coalitions across the calibration set---that decomposes the total attribution into calibration-driven and instance-driven components. The construction is method-agnostic on two axes: any local attribution method that can be applied to a scalar uncertainty target can be used, and any split CP family whose uncertainty size admits a known separation into a calibration-only and an instance-only summary satisfies the decomposition. Empirically, the calibration-driven and instance-driven attributions frequently carry opposite signs on synthetic and real data, with representative cases where a feature's sign reverses between the two views; the total attribution then becomes a misleading sum. On a clinical-decision task with semi-synthetic data, this conflation degrades decision quality, whereas jointly considering the calibration-driven and instance-driven attributions recovers high-uncertainty cases missed by total attribution. We further empirically verify that the decomposition holds across multiple CP families that satisfy the separability conditions.


Decomposing Earth Embeddings with Sparse Autoencoders

Vitus Benson ⋅ Fanny Yang ⋅ Markus Reichstein

Earth observation foundation models produce global meter-resolution embeddings that underpin work across agriculture, ecology, hydrology, and beyond. Yet the embeddings themselves are opaque: each dimension entangles unrelated physical concepts, rendering their interpretability and robustness challenging for practitioners and scientists alike. Here, we propose sparse autoencoders (SAEs) as a principled way to decompose Earth embeddings into overcomplete dictionaries of sparse, approximately monosemantic features. We introduce an evaluation protocol suited to a domain where LLM-based auto-interpretability fails and spatial autocorrelation matters, scoring any SAE on reconstruction, sparsity, monosemantic recovery, and spatial coherence against a curated label registry of categorical and continuous reference products. On this protocol, three state-of-the-art SAE recipes (ReLU+$\ell_1$, BatchTopK, and Matryoshka BatchTopK) all recover monosemantic and spatially contiguous concepts from AlphaEarth and Tessera embeddings much better than raw embeddings, with Matryoshka the strongest of the three. Two applications follow: (i) unsupervised discovery of land-cover subconcepts (e.g., splitting water into river, lake, ocean), and (ii) improved area-of-applicability estimates for OOD detection. Together, these results are a step toward Earth observation models whose features can be inspected rather than only used, opening pathways for scientific discovery and improved trustworthiness in geospatial applications. Code and models will be released upon publication.

We study two-layer neural networks trained by stochastic gradient descent (SGD) on a multi-index teacher model, focusing on how first-layer parameters orthogonal to the teacher subspace affect generalization. We decompose the first-layer weights into a signal component aligned with the teacher and an orthogonal complement, and analyze their joint SGD dynamics together with the second layer in the proportional limit. We show that although SGD updates to the complement are initially isotropic, the training dynamics induce teacher-dependent spectral spikes in the complement Gram matrix. This emergent anisotropy provides a concrete mechanism by which representation variance is amplified, and generalization degrades. To isolate this effect, we introduce Decomposed Dynamics, an analytically tractable surrogate whose asymptotic test-error dynamics coincide with those of SGD. Using this surrogate, we characterize regimes in which freezing the orthogonal complement strictly improves generalization relative to training it, despite identical training error. Finally, we show that the signal component rapidly aligns with the second-layer weights, revealing strong cross-layer coupling even in this shallow setting. Together, our results provide a precise dynamical explanation for when and how feature learning outside the teacher subspace harms generalization.

Existing multi-task active learning (MTL-AL) methods predominantly rely on the linear scalarization of heterogeneous objectives, such as uncertainty, diversity, and gradient conflict.We argue that this entangled formulation is fundamentally limited, as linear combination fails to capture Pareto-optimal trade-offs under conflicting objectives. To address this limitation, we propose a paradigm shift from combination to deconstruction and introduce a principled framework that rethinks MTL-AL along three orthogonal dimensions. First, we identify the Conflict Paradox: contrary to conventional wisdom, samples with high gradient conflict are not noise to be avoided but carry maximal information about unresolved inter-task trade-offs. We therefore decouple sample selection, which embraces conflict, from optimization, which resolves it via gradient surgery. Second, we introduce Orthogonal Decomposition, separating the acquisition objective into two distinct axes: Informativeness and Exploration. This two-stage process, consisting of subspace projection followed by manifold coverage, prevents ``diversity dilution'' by excluding uninformative outliers before enforcing diversity. Third, we propose Cross-Task Mutual Information (CTMI), a distribution-free proxy for inter-task dependency derived from a variational lower bound under the Maximum Entropy principle. Extensive experiments on COCO, Cityscapes, and NYUv2 — including comparisons against recent multi-task active learning baselines — demonstrate consistent improvements over state-of-the-art MTL-AL methods, highlighting that structural rethinking of acquisition yields larger gains than incremental heuristic design. Our code is publicly available at: \href{https://anonymous.4open.science/r/multitaskactivelearning-1545}{\texttt{https://anonymous/CTMI}}


Decoupled Safety Control: A Safety-Control Algorithm for Training-Free Safety Guidance

Huilin Zhou ⋅ Ruoxi Cheng ⋅ Yuhang Wang ⋅ Yuming Liu ⋅ Minghao Sun ⋅ Ruolong Ma ⋅ Lan Zhang

Diffusion models achieve strong performance in text-to-image generation, but they can still produce unsafe content. Training-free safety guidance offers a practical inference-time alternative that intervenes during sampling without modifying model weights. However, many existing safeguards still rely on a coupled safety-control chain. In this chain, the same unsafe condition simultaneously specifies what should be avoided, indicates the current unsafe signal, constructs the intervention, and determines when that intervention remains active. This coupling can work when unsafe factors are explicit, but becomes less reliable once harmful and benign semantics are entangled in the current generation state, often leaving residual unsafe content or unnecessarily degrading benign-generation quality or prompt fidelity. We address this problem with Decoupled Safety Control (DSC), a plug-in safety-control algorithm for training-free safety guidance. DSC reformulates training-free safety guidance as an explicit state-dependent control process, so the executed intervention is determined from the current generation state rather than inherited directly from a single unsafe condition. Experiments on standard safety benchmarks show that DSC improves representative latent- and text-space safeguards, achieving stronger unsafe-content suppression while preserving benign-generation quality and maintaining efficient inference.


Decoupling Action from Egocentric Observation for World Simulation

Yue Ma ⋅ Pengjie Song ⋅ Xinyu Wang ⋅ Yi He ⋅ Zeqian Long ⋅ Fangneng Zhan ⋅ Kaichen Zhou ⋅ Hongyu Liu ⋅ Hongfa Wang ⋅ Peihao Li ⋅ haoyang huang ⋅ Nan Duan ⋅ Qifeng Chen

Egocentric action transfer aims to decouple the semantic action from egocentric observations and reproduce them in new visual contexts, enabling scalable world simulation, embodied policy learning, and robotic data generation. Existing approaches to transferring actions rely on condition-guided video generation, which converts the observation video into explicit geometric controls (e.g., hand poses or meshes) to drive synthesis. However, these methods introduce geometric estimation errors, making it difficult to preserve physically plausible interactions when transferred to new scenes. Alternatively, motion transfer methods directly extract motion patterns from the observation video, yet motion-level features alone cannot encode rich contact dynamics, frequently leading to severe hand structural collapse and implausible contact layout. To address both limitations, we present EgoACT, a test-time framework for egocentric action transfer that operates directly on the denoising process of video diffusion models without requiring additional geometric estimators. EgoACT uses Velocity-guided Structure Anchoring to stabilize reference-consistent hand-object structure in the early denoising stage, and Sparse Correspondence Calibration to refine reliable local correspondences in the mid-to-late denoising stages. Together, these two components preserve transferable action semantics while improving temporal coherence and interaction realism. We further establish EgoActionBench, a benchmark for evaluating action preservation, visual quality, and hand-object plausibility across diverse egocentric manipulation scenarios. Experiments show that EgoACT generates more coherent and physically plausible interaction videos than strong baselines.


Decoupling Direction and Magnitude: Language-Steered Flow Matching for Super-Resolution in the Dark

Jiaxin Gao ⋅ Ziyu Yue ⋅ Yaohua Liu ⋅ Danchen Cui ⋅ Zhihui Zhao ⋅ Zhixun Su

Image super-resolution in the dark is fundamentally challenged by extreme spatial heterogeneity: severely underexposed and noisy regions demand aggressive generative enhancement to reconstruct missing details, while relatively well-exposed areas require conservative updates to prevent hallucinated textures. Standard diffusion and flow-matching models, however, rely on globally uniform integration trajectories, making them sub-optimal for such spatially varying degradations. In this paper, we propose DeLang-SR (Degradation-Language steered Super-Resolution), a one-step flow-matching framework that explicitly decouples the restoration direction and magnitude. Rather than treating low-light corruption as arbitrary latent features, DeLang-SR translates measurable physical statistics into structured degradation language, harnessing the semantic prior of pretrained text-to-image models. The degradation-language intent, complemented by image-specific visual tokens, steers the velocity field (direction) toward restoration-relevant manifolds. Simultaneously, a continuous spatial update-strength field acts as a locally varying Euler step-size field (magnitude), applying stronger integration steps in heavily degraded shadows while enforcing conservative updates in reliable regions. Experiments on paired and real dark images, together with ablations and diagnostic analyses, show that DeLang-SR provides an effective and interpretable one-step route to super-resolution in the dark.

Decomposing prediction uncertainty into aleatoric (irreducible) and epistemic (reducible) components is critical for the reliable deployment of machine learning systems. While the mutual information between the response variable and model parameters is a principled measure for epistemic uncertainty, it requires access to the parameter posterior, which is computationally challenging to approximate. Consequently, practitioners often rely on probabilistic predictions from deep ensembles to quantify uncertainty, which have demonstrated strong empirical performance. However, a theoretical understanding of their success from a frequentist perspective remains limited. We address this gap by first considering a bootstrap-based estimator for epistemic uncertainty, which we prove is asymptotically correct. Next, we connect deep ensembles to the bootstrap estimator by decomposing it into data variability and training stochasticity; specifically, we show that deep ensembles capture the training stochasticity component. Through empirical studies, we show that this stochasticity component constitutes the majority of epistemic uncertainty, thereby explaining the effectiveness of deep ensembles.


DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos

Can Li ⋅ Zhoujian Li ⋅ Ren Li ⋅ Jie Gu ⋅ lei lei ⋅ Jingmin Chen ⋅ Lei Sun

World models for deformable objects should recover not only geometry and appearance, but also underlying physical dynamics, interaction grounding, and material behavior. Learning such a model from real videos is challenging because deformable linear, planar, and volumetric objects evolve under high-dimensional deformation, noisy interactions, and complex material response. The model must therefore infer a physical state from visual observations, roll it forward under new interactions, and render the resulting dynamics with high visual fidelity. We present DeformMaster, a video-derived interactive physics--neural world model that turns real interaction videos into an online interactive model of deformable objects within a unified dynamics-and-appearance framework. DeformMaster preserves structured physical rollout while using a neural residual to compensate for unmodeled effects, grounds sparse hand motion as distributed compliant actuator for hand--continuum interaction, represents material response with spatially varying constitutive experts, and drives high-fidelity 4D appearance from the predicted physical evolution. Experiments on real-world deformable-object sequences demonstrate DeformMaster's ability to roll out future dynamics and render dynamic appearance, outperforming state-of-the-art baselines while supporting novel action rollout, material-parameter variation, and dynamic novel-view synthesis.


DegBins: Degradation-Driven Binning for Depth Super-Resolution

Zhiqiang Yan ⋅ Zhengxue Wang ⋅ Jian Yang ⋅ Gim Hee Lee

Depth super-resolution (DSR) aims to recover a high-resolution (HR) depth map from its low-resolution (LR) counterpart. With color image guidance, this task is typically formulated as learning the residual between HR and LR in a low-dimensional feature space. However, this additive formulation is insufficient to accurately capture the complex relationship between HR and LR, especially under spatially varying degradations. In this paper, we introduce \textbf{DegBins}, a novel DSR framework that leverages degradation-driven binning to adaptively enhance residual modeling. Specifically, DegBins reformulates the regression-based DSR as a hybrid classification-regression problem, where the residual depth is represented as a linear combination of discrete depth bins weighted by their learned probability distribution, yielding more flexible and expressive representations. Furthermore, DegBins models the degradation relationship between HR and LR in a high-dimensional feature space, enabling adaptive bin range adjustment and probability optimization conditioned on local degradation characteristics. To progressively improve reconstruction quality, DegBins adopts a multi-stage refinement scheme, where each stage performs finer-grained bin partitioning and probability updating based on the former estimation. This coarse-to-fine design facilitates more accurate depth recovery, particularly in regions with severe degradations or complex structural variations. Extensive experiments across five benchmarks demonstrate that DegBins consistently outperforms existing state-of-the-art methods in terms of accuracy, robustness, and generalization.

In industrial machine vision, removing specular highlights is crucial for accurate surface analysis; however, existing open research remains largely confined to natural images, limiting its efficacy on highly reflective metallic surfaces. When used for zero-shot inference, these methods generate artifacts, making them unreliable for specialized industrial scenes. Since current state-of-the-art methods rely on pixel-level supervision, adapting them to novel industrial subjects is often impractical. Such adaptation requires highlight-free ground truths across hundreds of multi-view scenes, necessitating hardware-intensive cross-polarization, which is difficult to implement at scale in uncontrolled real-world settings. To address these challenges, we propose DeGlare, a flexible training framework for highlight removal in specialized industrial settings. Our approach leverages only a single-view scene by exploiting its multi-illumination observations together with a latent permutation strategy, enabling an encoder-decoder–style architecture to implicitly learn diffuse–specular decomposition. This eliminates the dependency on the burdensome manual annotation required by state-of-the-art supervised methods. We demonstrate the practical effectiveness of our approach on a real-world industrial robotic rig data set, where DeGlare achieves promising performance under limited data and annotation constraints.

Denoising-based Vision-Language-Action (VLA) policies are often parameterized by a single denoiser shared across denoising time. This paper shows that expert usefulness in VLA action generation can vary systematically across denoising phases, and that this variation can be exposed through controlled step-wise expert aggregation. We mix two full action experts at the velocity-field level with a fixed smooth schedule, enabling schedule-only interventions that keep expert capacity and total mixing mass matched while changing when each expert is emphasized. The empirical evidence is organized as a matched chain: semantically distinct expert pairs test denoising-phase assignment through forward versus reversed schedules, while an anchor-free isomorphic-expert control tests schedule sensitivity without manually designed expert semantics. Experiments on both LIBERO and CALVIN show the same directionality phenomenon: aligned schedules outperform their time-reversed counterparts under matched strength, indicating that when each expert is emphasized matters for action denoising. We complement these findings with a squared-error risk analysis showing that denoising-time variation in expert advantage implies a time-dependent optimal mixture, together with a schedule-weighted objective note explaining why copied experts can still functionally differentiate under non-constant schedules.


Despa: Resolving Spatial Collapse in VLMs via Depth-Grounded Geometry

Yujing Lou ⋅ Pingyi Chen ⋅ Shen Cao ⋅ Lubin Fan ⋅ Yue Wu ⋅ Lizhuang Ma ⋅ Jieping Ye

Vision-Language Models (VLMs) excel at open-world 2D perception but struggle with precise metric spatial reasoning, a key requirement for applications like autonomous driving, navigation, and embodied intelligence. We identify this as a structural limitation: images are 2D projections of the 3D world, inherently suffering from projective ambiguity, while VLM components favor semantic understanding and rely on 2D positional bias, leading to Spatial Collapse. To address this, we propose Despa, a Depth-grounded spatial VLMs that injects depth-based geometric information via Geometric Positional Embedding (GPE) and Depth-aware Rotational Positional Embedding (DoPE). To expose masked limitations in current benchmarks, we construct SpaDataset (1.7M QA pairs) and SpaBench (2,051 QAs), to our knowledge the largest fine-grained training corpus and benchmark suite for metric spatial reasoning. They cover three difficulty levels and ten sub-tasks across diverse settings. The fine-tuned Despa-4B model consistently outperforms general-purpose, closed-source, and specialized VLMs on SpaBench by 27.9% over Qwen3-VL-235B-A22B and 31.1% over GPT-5.2, generalizes to MSMU (66.3%), RefSpatial-U (42.4%), and the multi-frame VSI-Bench (66.0%), and consistently lifts LLaVA-1.5, Qwen2.5-VL, and Qwen3-VL backbones, achieving the performance with only a 4B model. The code will be publicly released.

Scientific measurements inherently contain variance. Standard regression, however, treats noisy data as absolute ground truth, leading models to memorize noise rather than physical laws. This overfitting hampers generalization, particularly when transferring from abundant theoretical proxies to scarce experiments. We introduce DETS, a framework replacing rigid point estimation with Physical Tolerance Modeling. Our Interval-Censored Evidential Engine (ICEE) maximizes probability mass within acceptable error margins, explicitly decoupling aleatoric noise from epistemic uncertainty. Using this filtered uncertainty signal, a Thermodynamic Sampling strategy dynamically selects theoretical data to align source and target domains. Experiments across thermodynamics, drug affinity, and bandgap prediction show DETS outperforms state-of-the-art methods. Crucially, it exhibits superior robustness in high-noise, data-scarce settings.


DexOPE: 6D Object Pose Estimation in Dexterous Manipulation

Ke Wu ⋅ Zhiwei Yang ⋅ Xiangting Meng ⋅ Zicheng Zhang ⋅ Hui Zhang ⋅ Zijun Xu ⋅ Jieru Zhao ⋅ Wenchao Ding

Reliable 6D object pose estimation is essential for dexterous manipulation but remains highly challenging due to severe visual occlusions from multi-finger interactions. Progress has been limited by two key factors: the lack of large-scale manipulation datasets and the inability of existing methods to handle high uncertainty under occlusion. To address these challenges, we introduce DexOPE, a large-scale dataset and a physics-guided pose estimation framework. DexOPE provides a multimodal dataset with 1.5k diverse objects, an order of magnitude larger than prior work, covering simulated and real-world grasping and manipulation. To handle pose ambiguity under occlusion, we propose a Physics-Guided Score Matching approach that models the pose posterior using score-based diffusion models rather than deterministic regression. Physical constraints from hand–object interactions, including nonpenetration and tactile contact, are incorporated as guidance during sampling, enabling physically plausible pose inference even under severe or complete occlusion. Extensive experiments show that DexOPE achieves state-of-the-art performance, significantly outperforming existing methods in robustness to occlusion and generalization to unseen objects.

While large language models have significantly advanced Text-to-SQL generation, it remains largely an open-loop paradigm due to the absence of a reliable validator capable of systematically diagnosing generated queries. Training a dedicated diagnostic validator to provide targeted, reflective feedback is crucial for closing this loop. However, employing standard Reinforcement Learning paradigms, particularly Group Relative Policy Optimization (GRPO), reveals fundamental structural misalignments. Through exploratory analysis, we identify a critical limitation of standard GRPO in this context: a severe credit assignment failure, where uniform sequence-level rewards inadvertently reinforce hallucinated reasoning alongside correct sub-answers in complex structured outputs. Concurrently, we discover profound structural dependencies among SQL errors, revealing inherent co-occurrence semantics dictating that certain errors frequently appear together. To address the limitation and leverage this insight, we propose DiagSQL, a novel framework that utilizes an improved GRPO algorithm tailored for diagnostic Text-to-SQL validation. DiagSQL introduces Fine-grained Token-level Reward Allocation (FTRA), which parses structured responses to precisely distribute rewards at the token level, effectively isolating and penalizing erroneous reasoning. Furthermore, it incorporates Co-occurrence-aware Reward Shaping (CORS), which leverages our discovered pre-computed error co-occurrence matrix to dynamically adjust optimization objectives, encouraging logical error combinations while penalizing structural violations. Extensive experiments demonstrate that DiagSQL significantly enhances the robustness and accuracy of SQL error diagnosis, providing actionable, high-quality feedback that effectively transforms Text-to-SQL into a closed-loop paradigm. Our code is publicly available at https://anonymous.4open.science/r/DiagSQL.

Model cascades reduce LLM inference costs by sequentially querying candidate models and stopping once a satisfactory response is generated. However, existing cascades largely optimize model escalation under fixed evaluation schemes, overlooking that optimal stopping fundamentally depends on the reliability of intermediate evaluations. We formalize this as an online sequential search problem with tunable-fidelity probes, in which the system jointly decides which model to query, at what fidelity to evaluate it, and when to stop. Under a fixed evaluation protocol, higher fidelity incurs greater evaluation cost but reduces stochastic observation noise through an unknown arm-specific noise function. For the oracle setting, we derive a generalized index policy that tightly couples fidelity selection with stage-wise stopping. For the online setting, we propose DialBandit, which learns fidelity-dependent noise functions from unbiased replicate-based labels and performs uncertainty-aware planning over probing order, evaluation fidelity, and stopping. We prove a finite-time sublinear regret bound and show that DialBandit consistently improves net utility over strong baselines in semi-empirical environments based on repeated judging and calibrated noisy LLM evaluations.

Recent advancements in Automatic Speech Recognition (ASR) have significantly improved performance for high-resource languages, yet ASR systems still struggle with low-resource languages and dialects due to the scarcity of annotated training data. To address this challenge, we propose a novel framework that enhances the utilization of limited data through a two-pronged approach: (1) a novel data augmentation technique based on a multi-view pseudo-parallel strategy, and (2) a robust multi-level contrastive learning framework capable of jointly leveraging semantic and dialect-specific attributes to improve model effectiveness under noisy conditions. Our approach effectively generalizes across dialectal variations when dialect data are non-parallel, allowing acquisition of shared linguistic structures and dialectal distinctions. Furthermore, our method could be simply transferred to scenarios where dialect data are parallel. We validate our method by adapting the Whisper-large model for Swiss German, Chinese and Arabic dialect ASR tasks, demonstrating substantial performance gains. Specifically, our framework achieves up to a 48.59% reduction in WER and 23.16% reduction in CER under non-parallel conditions, outperforming state-of-the-art baselines. These results highlight the robustness and adaptability of our approach for low-resource dialectal ASR. The data and codes will be public available upon acceptance.


DIBench: Benchmarking Decision Integrity of GUI-based Mobile Agents Under Deceptive Injections

Li Hu ⋅ Kanghua Mo ⋅ Yingbin Jin ⋅ Qingqing Ye ⋅ Haibo Hu

As GUI-based mobile agents rapidly progress, rigorous safety evaluation of their autonomous decision-making in realistic app interfaces becomes increasingly critical. Existing benchmarks mainly focus on execution-level anomalies using task success or hijack rates, but fail to capture the in-task goal deviation risk in multi-candidate selection tasks, where the decision may be steered toward an attacker-specified target, even in violation of instruction-implied constraints (e.g., cheapest/highest-rated), without any overt execution anomalies. We present DIBench, a decision integrity benchmark for measuring this risk in mobile agents. DIBench covers 7 commercial and 3 simulated apps with 5 task types. Under a threat model restricted to non-privileged UI content, we construct 8 deceptive injection probe instantiations that can steer critical selections without overt anomalies. The benchmark includes 1,000 clean and 36,672 injected instances, with a unified protocol and integrity metrics for comparison. Experiments spanning 4 agent frameworks and 7 base models show that completion-based evaluation can overestimate agent trustworthiness and miss decision-integrity risks: deceptive injections steer selections and shift early action policies, inflating completion rates and creating a misleading illusion of safety. Common defenses, including detection, image preprocessing, and prompt reminders, yield inconsistent integrity gains. Overall, DIBench provides a unified, reproducible benchmark to quantify the risk of in-task goal deviation in mobile agents and enable comparable evaluations of safety defenses. Code is available at: https://anonymous.4open.science/r/DIBench-5432.


Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation

Binglei Li ⋅ Mengping Yang ⋅ Zhiyu Tan ⋅ Qiang Xiang ⋅ Junping Zhang ⋅ Hao Li

Recent text-to-image (T2I) diffusion models have achieved remarkable advancement, yet faithfully following complex textual descriptions remains challenging due to insufficient interactions between textual and visual features. Prior approaches enhance such interactions via architectural design or handcrafted textual condition weighting, but lack flexibility and overlook the dynamic interactions across different blocks and denoising stages. To provide a more flexible and efficient solution to this problem, we propose Diff-Aid, a lightweight method that adaptively adjusts per-token text and image interactions across transformer blocks and denoising timesteps. Beyond improving generation quality, Diff-Aid yields interpretable modulation patterns that reveal how different blocks, timesteps, and textual tokens contribute to semantic alignment during denoising. As a plug-and-play module, Diff-Aid can be seamlessly integrated into downstream applications for further improvement, including style LoRAs, controllable generation, and zero-shot editing. Experiments on strong baselines (SD~3.5 and FLUX) demonstrate consistent improvements in prompt adherence, visual quality, and human preference across various metrics. Our code and models will be released.


DiffCap-RL: Differential QA Rewards for Dense Image and Video Captioning

Kaixun Jiang ⋅ Zhihang Liu ⋅ Jiyuan Fu ⋅ Yuzheng Wang ⋅ Pandeng Li ⋅ Junjie Zhou ⋅ Chen-Wei Xie ⋅ Yun Zheng ⋅ Wenqiang Zhang

Dense image and video captioning demands exhaustive visual details while strictly maintaining factuality, forcing a severe trade-off between descriptive density and hallucination. Supervised fine-tuning (SFT) struggles here: for dense captions, its token-level loss dilutes effective signals while failing to reinforce specific fine-grained dimensions. Reinforcement learning (RL) is a promising alternative but remains bottlenecked by reward design: coarse-grained VLM judges suffer from reward drift, and static QA rewards quickly saturate as the policy improves. To address this, we propose DiffCap-RL, an RL framework combining a stable offline Pre-Scorer with an adaptive online Diff Scorer. We utilize the Pre-Scorer as a stable anchor for broad visual grounding via automated QA pairs. Once it loses discriminative power, Online Diff Scorer dynamically extracts semantic disagreements between policy rollouts. By decomposing these disagreements into targeted conflict questions to suppress hallucinations and extra questions to reward valid new details, the Diff Scorer provides reliable, non-saturating supervision at the policy frontier. We validate our reward quality from three perspectives: human alignment, online RL signal quality, and BoN selection. Across multiple image and video captioning benchmarks, our DiffCap-RL delivers large gains and consistently outperforms state-of-the-art baselines at similar or larger scales.


Differentiable Learning of Lifted Action Schemas for Classical Planning

Jonas Reiter ⋅ Jakob Gebler ⋅ Hector Geffner

Classical planners can effectively solve very large deterministic MDPs represented in STRIPS or PDDL where states are sets of atoms over objects and relations, and lifted action schemas add or delete these atoms. This compact representation yields strong search heuristics and provides an ideal setting for structural generalization, since lifted relations and action schemas give rise to infinitely many domain instances. A central challenge is to learn these relations and action schemas from data, and recent approaches have addressed this problem using different types of observations. In this work, we develop a novel neural network architecture for learning action schemas from traces where states are fully observed but action arguments are unobserved. The problem is a simplification but an important step towards learning planning domains from sequences of images and action labels, and we aim to solve this simplification in a nearly perfect manner. The challenge lies in learning the action schemas while simultaneously identifying the action arguments from observed state changes. Our approach yields a robust differentiable component that can then be integrated into larger neuro-symbolic models. We evaluate the architecture on various planning domains, where the learned lifted action schemas must recover the ground-truth structure. Additionally, we report experiments on robustness to observation noise and on a variation related to slot-based dynamics models.


DiffPTS: Rethinking Diffusion ELBO for Probabilistic Time Series Forecasting

Weiwei Ye ⋅ Dongyuan Li ⋅ Hangchen Liu ⋅ Haotong Jiang ⋅ Yoshihide Sekimoto ⋅ Renhe Jiang

Probabilistic time series forecasting requires modeling and predicting complex and time-varying distributions. Recently, Denoising Diffusion Probabilistic Model~(DDPM)-based approaches have shown promise by equipping the diffusion process with pretrained mean and variance estimators to accommodate distributional shift. However, these methods typically follow the standard DDPM framework and consider only partial components of the evidence lower bound (ELBO), treating the training of estimators as designed regression tasks separate from the variational inference framework. To address this, we rethink the ELBO under the Location-Scale Noise Model (LSNM) and find that it naturally induces a Gaussian negative log likelihood objective for the estimators and inherently defines a joint training objective that unifies recent the diffusion paradigms for probabilistic forecasting. Building on this principled ELBO reformulation, we propose DiffPTS, a general framework that enables end-to-end optimization of all components within the ELBO. Across multiple benchmarks, DiffPTS consistently outperforms recent models, achieving state-of-the-art performance with an average CRPS/MSE reduction of over 14.53\%/16.55\% compared to existing diffusion-based methods. The code is provided at \url{https://anonymous.4open.science/r/DiffPTS}.

While Large Vision-Language Models (LVLMs) excel at bounding-box grounding, they struggle with precise spatial tasks such as polygon grounding. We attribute this bottleneck to two properties of the vertex-level autoregressive (AR) paradigm used by current LVLMs: (1) errors in early-emitted vertices propagate uncorrected through the rest of the sequence, and (2) the model commits to local vertex placement before observing the full contour, leading to suboptimal allocation of a fixed vertex budget. We propose Diffusion Fine-Tuning (DFT), which removes the vertex-level sequential dependency by placing all 2N coordinate tokens under a single discrete-diffusion denoiser, while introducing a much shorter digit-level factorisation: each coordinate is decomposed into hundreds, tens, and units digits, and the reverse process is trained to predict them coarse-to-fine. We train this with a Hierarchical Curriculum Learning strategy that progressively refines loss supervision from macro-contour to per-pixel detail. Under matched fine-tuning protocols, DFT matches strong AR LVLMs on 2D bounding-box grounding and improves over them on 16-point polygon grounding; the same network and training recipe extend to 9-DoF monocular 3D bounding-box grounding under a 9-parameter coordinate parameterisation. A block-wise top-k decoder closes most of the latency gap to AR and recovers most of the quality lost when scaling beyond 16 vertices. We do not claim to eliminate sequential dependence in general: vertex-level O(N) AR decoding is replaced by K joint denoising steps (K=12 in this work), with the three-stage digit hierarchy entering as a training-time loss factorisation rather than an inference-time chain.


Diffusion-State Policy Optimization for Masked Diffusion Language Models

Daisuke Oba ⋅ Hiroki Furuta ⋅ Naoaki Okazaki

Masked diffusion language models generate text through iterative masked-token filling, but terminal-only rewards on final completions provide coarse credit assignment for the intermediate filling decisions that shape the generation process. We propose Diffusion-State Policy Optimization (DiSPO), a plug-in credit-assignment layer that directly optimizes intermediate filling decisions. At selected intermediate masked states, DiSPO branches by resampling the currently masked positions from rollout-cached logits, scores the resulting completions, and updates only the newly filled tokens, requiring no additional multi-step diffusion rollouts or optimizer steps. We formalize a fixed-state objective for branched completions and derive a policy-gradient estimator that reuses the same rollouts as terminal-feedback policy optimization. Experiments on LLaDA-8B-Instruct show that DiSPO consistently improves terminal-feedback baselines, including diffu-GRPO and SPG, on math and planning benchmarks under matched rollout compute and optimizer steps, supporting its use as a general plug-in for masked diffusion policy optimization.


DiRecT: Safe Diffusion-Based Planning via Receding-Horizon Denoising

Paolo Giaretta ⋅ Zeyang Li ⋅ Navid Azizan

Diffusion models have emerged as powerful tools for planning and control by learning multimodal distributions over actions and trajectories. Yet reliable inference-time safety enforcement remains a key barrier to their deployment in safety-critical tasks. Existing approaches typically project each denoising iterate onto the feasible set, even though constraints are defined only on the final clean trajectory. Enforcing feasibility on noisy intermediate samples can therefore overconstrain the sampling dynamics, substantially degrading sample quality. To address this limitation, we introduce DiRecT (Diffusion-based planning via Receding-horizon denoising with Terminal constraints), a training-free algorithm for constrained sampling from diffusion models via stochastic optimal control (SOC). DiRecT enforces constraints only on the final clean sample, avoiding unnecessary restrictions on the intermediate denoising dynamics. Inspired by model predictive control, we derive a principled receding-horizon surrogate for the otherwise intractable constrained SOC formulation, yielding an efficient algorithm that cleanly separates stochastic denoising from constraint satisfaction, progressively steering samples toward feasible final trajectories without distorting the learned diffusion dynamics. Furthermore, DiRecT is highly flexible: it can leverage off-the-shelf or domain-specific optimizers, incorporate priors over environment dynamics, and optimize additional soft rewards. Extensive experiments on safe planning benchmarks demonstrate that DiRecT substantially improves deployment safety and task performance over existing diffusion-based planning baselines.


Discovering Programmatic Policies from Reinforcement Learning-Based Traffic Signal Controllers

Lindong Xie ⋅ Yang Zhang ⋅ Beiyu Song ⋅ XING Zeren ⋅ Flora Salim ⋅ Edward Chung

Traffic signal control (TSC) is critical for mitigating urban congestion. Deep reinforcement learning (DRL) has achieved strong performance by learning adaptive policies from traffic observations, but learned neural controllers are often opaque, structurally complex, and difficult to generalize across scenarios. We propose PPL-TSC, a programmatic policy learning framework that distills compact and human-readable TSC policies from DRL teachers. PPL-TSC formulates policy learning as an evolutionary search over phase utility functions (PUFs), which score candidate signal phases and select the highest-utility phase for execution. Each generation consists of four steps: policy generation, policy review, policy calibration, and policy evaluation. In the generation step, a large language model (LLM) acts as a semantics-aware evolutionary operator to propose and refine candidate PUFs. Generated PUFs are screened through explicit checks for syntactic validity, mathematical soundness, and consistency with TSC principles, with invalid ones repaired through iterative feedback. Valid PUFs are then calibrated to improve behavioral fidelity to the DRL teacher, and evaluated using a joint score combining teacher fidelity and traffic-control performance to select candidates for the next generation. Extensive experiments across diverse traffic scenarios and representative DRL-based controllers show that PPL-TSC produces compact and human-readable policies, matches teacher performance, improves transfer generalization, and substantially reduces computational cost.


Discretizing Continuous Time Series for Imputation with Masked Diffusion Training

Dongbin Kim ⋅ Seungyun Lee ⋅ Geonwoo Shin ⋅ Jaewook Lee

Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics and noise of real-world data. Existing approaches, however, exhibit two limitations: missing and observed values are embedded within the same representation space without explicit structural separation, and continuous diffusion-based methods are trained to predict added noise rather than the original signal. To address these, we propose the Masked Diffusion Time-series Imputation Model (MDTIM), which leverages the training paradigm of masked diffusion model for imputation tasks. The \texttt{[MASK]} token is structurally orthogonal to valid observations, and the model directly predicts the original values, naturally aligning both the representation and the learning objective with the imputation task. To bridge the gap between discrete masked diffusion and the continuous, ordinal nature of time series, we further introduce Stochastic Discretization, which maps continuous values to ordinal-aware tokens while preserving continuous dynamics. Our experiments on diverse benchmarks confirm that MDTIM achieves superior robustness and scalability, consistently outperforming state-of-the-art deterministic and generative baselines across various missing scenarios.


Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning

Hyeonjin Kim ⋅ Hangyeol Jung ⋅ Heechan Yun ⋅ Sungjun Yun ⋅ Dong-Jun Han

Unlearning specific concepts in text-to-image diffusion models has become increasingly important for preventing undesirable content generation. Among prior approaches, sparse autoencoder (SAE)-based methods have attracted attention due to their ability to suppress target concepts through lightweight manipulation of latent features, without modifying model parameters. However, SAEs trained with sparse reconstruction objectives do not explicitly enforce concept-wise separation, resulting in shared latent features across concepts. To address this, we propose SAEParate, which organizes latent representations into concept-specific clusters via a concept-aware contrastive objective, enabling more precise concept suppression while reducing unintended interference during unlearning. In addition, we enhance the encoder with a GeLU-based nonlinear transformation to increase its expressive capacity under this separation objective, enabling a more discriminative and disentangled latent space. Experiments on UnlearnCanvas demonstrate state-of-the-art performance, with particularly strong gains in joint style-object unlearning, a challenging setting where existing methods suffer from severe interference between target and non-target concepts.


Disentangling Optimization Geometry via Hierarchical Polar Adapters for Class-Incremental Learning

Dat Q Mac ⋅ Thanh Hai Dang ⋅ Duc-Trong Le ⋅ Quynh-Trang Pham Thi

Deep neural networks suffer from catastrophic forgetting when learning on sequence tasks because standard adaptation mechanisms update a monolithic set of Cartesian weights, inherently entangling feature intensity with spatial displacement and destructively overwriting ancestral representations. Inspired by the brain’s multi-timescale consolidation and frequency-aware memory dynamics, we propose HiPo(Hierarchical Polar Adapters), a novel continual learning framework that fundamentally redefines how neural knowledge is parameterized, stored, and consolidated. First, HiPo projects token representations into a complex orthogonal Fourier basis and explicitly parametrizes the adapter weights in polar coordinates, decoupling feature intensity from geometric correlation to probabilistically isolate task updates. Second, we embed these polar weights within a multi-tiered hierarchical memory stack, functionally separating a highly plastic working memory from a deep, stabilizing long-term memory. Finally, an adaptively thresholded snapshot Fisher mechanism evaluates parameter criticality in the decoupled polar space. By executing mathematically safe, Cartesian-invariant knowledge transfers into the deep memory tiers, HiPo safely hollows out the working memory, paralyzing obsolete gradients and preserving plasticity without inducing structural distortion. Extensive experiments across standard CIL benchmarks demonstrate that HiPo establishes a highly resilient optimization manifold, achieving state-of-the-art performance and exceptional geometric stability, particularly under severe out-of-distribution shifts.


DisRFM: Polar Riemannian Flow Matching for Structure-Preserving Graph Domain Adaptation

Yingxu Wang ⋅ Xinwang Liu ⋅ Siyang Gao ⋅ Mengzhu Wang ⋅ Nan Yin

Graph Domain Adaptation (GDA) aims to transfer graph classifiers across domains with both semantic and topological shifts. Existing Euclidean adversarial methods face two challenges: Structural Degeneration, where domain confusion entangles and suppresses label-relevant topology, and Optimization Instability, where minimax training induces oscillatory gradients under large structural shifts. We propose DisRFM, a geometry-aware GDA framework that addresses these challenges with Riemannian representation learning and flow-based transport. DisRFM embeds graph representations on a constant-curvature manifold and expresses them in geodesic polar coordinates. Polar endpoint regularization calibrates topologysensitive radial scales via univariate Wasserstein alignment and preserves scalenormalized class semantics through confidence-filtered angular alignment, with radial magnitude modulating pseudo-label reliability. DisRFM introduces topologyconditioned polar flow matching, which couples class-compatible source and target samples by a normalized polar transport cost and learns a metric-corrected vector field along geodesic interpolants. Theoretical analysis characterizes the structural risk of unconditional domain confusion and relates polar discrepancies and flow error to target risk. Extensive experiments under diverse domain shifts demonstrate that DisRFM consistently outperforms state-of-the-art methods.


Distinguishing Performance From Competence in Evaluations of Humanlike Abstract Reasoning

Claas Beger ⋅ Ryan Yi ⋅ Shuhao Fu ⋅ Kaleda K Denton ⋅ Arseny Moskvichev ⋅ Sarah Tsai ⋅ Sivasankaran Rajamanickam ⋅ Melanie Mitchell

AI reasoning models have exceeded human performance on the ARC-AGI-1 benchmark, but does that mean state-of-the-art models have the underlying competence—humanlike abstract reasoning—the benchmark was designed to test? Here we investigate the abstraction abilities of AI models using the closely related but simpler ConceptARC benchmark. Our evaluations vary input modality (textual vs. visual), use of external Python tools, and reasoning effort. Beyond output accuracy, we evaluate the natural-language rules that models generate to explain their solutions, enabling us to assess whether models recognize the abstractions that ConceptARC was designed to elicit. We show that the best models’ rules are frequently based on less abstract, more domain-specific concepts, capturing intended abstractions considerably less often than humans. In the visual modality, AI models’ output accuracy drops sharply; however, our rule-level analysis reveals that a substantial share of their rules capture intended abstractions, even as the models struggle to apply these concepts to generate correct solutions. In short, we show that using performance (accuracy) alone can substantially overestimate AI competence in textual modalities while underestimating it in visual modalities—an illustration of the risk of mistaking performance for competence.


DISTMOE: Rehearsal-free Routing in Mixture-of-Experts for Distributed Instruction Tuning

Mainak Singha ⋅ Niccolò Biondi ⋅ Elisa Ricci ⋅ Subhankar Roy

Multimodal Large Language Models (MLLMs) have shown strong multimodal instruction-following ability, but adapting them to diverse visual-language domains typically assumes centralized data access and costly joint training. This is restrictive when data is distributed across private, domain-specific, or permission-limited clients. To this end, we propose DISTMOE, a mixture-of-experts (MoE) approach for distributed visual instruction tuning. In each layer of the language decoder it augments the public feedforward network (FFN) with a client-specific private FFN expert, with the goal to acquire domain-specific knowledge. However, independent expert training causes the private FFNs to learn representation of different scale and magnitudes, making merging the experts difficult. To reduce client-specific drift, we introduce a public-anchored expert composition stage that updates only routers and lightweight private projection adapters on a mix of local client data and public data, via an isotropic regularization loss, therefore making it rehearsal-free. During inference, DISTMOE performs modular routing over public and private experts, enabling token-wise domain composition without explicit domain labels. Experiments across diverse visual-language benchmarks show that DISTMOE enables flexible expert reuse, effective domain adaptation, and competitive performance while preserving modular control over client-specific knowledge.


Distribution-Adaptive Policy Optimization

Yuxiao He ⋅ Ziqi Wang ⋅ Xingzhou Lou ⋅ Xiaoqian Liu ⋅ Junge Zhang

Reinforcement Learning with Verifiable Rewards empowers Large Language Models to enhance their reasoning capabilities. Most algorithms rely on on-policy training, which is stable but inefficient due to its strict dependency on fresh samples. To overcome this inefficiency, off-policy RL decouples experience generation from policy optimization via Importance Sampling (IS). However, this introduces a policy discrepancy between the behavior policy and the target policy. We observe that mainstream RLVR algorithms struggle to maintain a balance between performance and stability in off-policy scenarios, suffering from either severe performance degradation or training collapse. We identify the root cause as the failure of static trust regions to maintain the bias--variance trade-off associated with IS ratios in off-policy settings. To address this, we propose DIstribution-adaptive Policy Optimization (DIPO), a novel hyperparameter-free RLVR algorithm employing an adaptive trust region. DIPO leverages the intrinsic statistics of the IS ratio distribution to achieve a robust bias--variance trade-off. Experimental results demonstrate that DIPO enables stable training in extreme off-policy scenarios, achieving performance comparable to or even surpassing the on-policy baseline, and outperforming other off-policy baselines by 13.9%.

Multi-task reinforcement learning (MTRL) aims to train a single agent to efficiently optimize performance across multiple tasks simultaneously. However, jointly optimizing all tasks often yields imbalanced learning: agents quickly solve easy tasks but learn slowly on harder ones. While prior work primarily attributes this imbalance to conflicting task gradients and proposes gradient manipulation or specialized architectures to address it, we instead focus on a distinct and underexplored challenge: \emph{imbalanced data allocation}. Standard MTRL allocates an equal number of environment interactions to each task, which over-allocates data to easy tasks that require relatively few interactions to solve and under-allocates data to hard tasks that require substantially more experience to solve. To address this challenge, we introduce Distributionally Robust Adaptive Task Sampling (DRATS), an algorithm that adaptively prioritizes sampling tasks furthest from being solved. We derive DRATS by formalizing MTRL as a feasibility problem from which we derive a minimax objective for minimizing the worst-case return gap, the difference between a desired target return and the agent's return on a task. In benchmarks like MetaWorld-MT10 and MT50, DRATS improves data efficiency and increases worst-task performance compared to existing task sampling algorithms.

Policy learning from observational data is fundamental to applications ranging from personalized medicine to welfare-program targeting, but standard approaches optimize empirical performance and may fail when the deployment distribution differs from the training distribution. We study \emph{distributionally-robust policy learning}: given observational data $\\{(X_i,W_i,Y_i)\\}_{i=1}^n$ with binary treatment $W$, we seek a policy $\pi \in \Pi$ that maximizes the worst-case policy value over an $f$-divergence ball of radius $\rho$ around the data-generating distribution. We propose an estimator built on cross-fitted doubly-robust scores for the conditional treatment effect, paired with a DRO solver that handles \texttt{KL}, $\chi^2$, and \texttt{CVaR} ambiguity sets. We establish two convergence guarantees. First, the estimated policy attains the standard $\widetilde{O}(n^{-1/2})$ rate on its DRO regret, and inherits the doubly-robust property from its non-DRO counterpart -- the rate continues to hold whenever either the outcome model or the propensity model is consistently estimated, even at slow nonparametric rates. We then propose a sample-splitting bias-correction algorithm derived from the Lagrangian dual, and show that under a Tsybakov-style margin condition on the worst-case treatment-effect contrast, the estimator attains a fast rate of $\widetilde{O}(n^{-(1+\alpha)/(2+\alpha)})$ , interpolating between $\widetilde{O}(1/\sqrt{n})$ and $\widetilde{O}(1/n)$ as the margin parameter $\alpha$ ranges over $[0,\infty)$. To our knowledge, this is the first work to incorporate a Tsybakov margin condition into the analysis of distributionally-robust policy learning, and the first to obtain fast rates in this setting. Empirically, on semi-synthetic experiments calibrated to a welfare-to-work program, distributional robustness yields meaningful improvements over empirical welfare maximization when training and test distributions differ structurally, while remaining competitive when no shift is present.


DiversePlace: Diversity-Seeking Curriculum Reinforcement Learning for Macro Placement

Wenrui Zhou ⋅ Jiashun Liu ⋅ Wenji Fang ⋅ Zhiyao Xie ⋅ Ling Pan

Macro placement is a critical early-stage decision in physical design because macro locations strongly constrain subsequent implementation stages. The task is difficult because these high-impact discrete decisions must be made before downstream quality can be measured accurately, forcing training to rely on fast but imperfect training-time oracles such as half-perimeter wirelength (HPWL). Existing RL-based placers therefore mostly optimize single HPWL-driven solutions, and their long sequential horizons often require heavy search assistance or offline expert data while leaving little room to preserve multiple competitive layouts. We address this gap by making quality-controlled diversity an explicit objective for macro placement. Instead of relying on a single proxy-optimal solution, we maintain a compact archive of HPWL-competitive yet meaningfully different placements for downstream selection. We realize this idea in DiversePlace, a curriculum proximal policy optimization (PPO) framework that combines a quality-gated diversity reward with an expanding placement horizon, enabling from-scratch RL optimization while keeping exploration inside competitive HPWL regions. On the ISPD2005 benchmarks, DiversePlace consistently improves top-5, top-10, and top-20 archive diversity and achieves lower or virtually identical HPWL on all eight designs with better efficiency. Our code is available at https://anonymous.4open.science/r/NIPS2026-DiversePlace-A615.


DMax: Aggressive Parallel Decoding for dLLMs

Zigeng Chen ⋅ Gongfan Fang ⋅ Xinyin Ma ⋅ Ruonan Yu ⋅ Xinchao Wang

We present DMax, a new paradigm for efficient Diffusion Language Models (dLLMs). It mitigates error accumulation in parallel decoding, enabling aggressive decoding parallelism while preserving generation quality. Unlike conventional masked dLLMs that decode through a binary mask-to-token transition, DMax reformulates decoding as a progressive self-refinement from mask embeddings to token embeddings. At the core of our approach is On-Policy Uniform Training, a novel training strategy that efficiently unifies masked and uniform dLLMs, equipping the model to recover clean tokens from both masked inputs and its own erroneous predictions. Building on this foundation, we further propose Soft Parallel Decoding. We represent each intermediate decoding state as an interpolation between the predicted token embedding and the mask embedding, enabling iterative self-revising in embedding space. Extensive experiments across a variety of benchmarks demonstrate the effectiveness of DMax. Compared with the original LLaDA-2.0-mini, our method improves TPF on GSM8K from 2.04 to 5.47 while preserving accuracy. On MBPP, it increases TPF from 2.71 to 5.86 while maintaining comparable performance. On one single H200 GPU using the SGLang framework, our 16B model achieves an average of 955 TPS at a batch size of 1.


Does 1/2-Tsallis-INF Also Work Well for Best-Arm Identification?

Jingxin Zhan ⋅ Yuze Han ⋅ Zhihua Zhang

Regret minimization (RM) and best-arm identification (BAI) are two fundamental objectives in multi-armed bandits. Among regret-minimizing algorithms, $1/2$-Tsallis-INF is a canonical best-of-both-worlds FTRL algorithm: it achieves logarithmic pseudo-regret in stochastic bandits while retaining minimax-optimal regret in adversarial bandits, without knowing the environment in advance. This raises a natural question: can the same algorithm, without additional exploration, also identify the best arm reliably? We study this question in stochastic bandits by analyzing the failure probability $e_t$, defined as the probability that the empirical best arm determined by the cumulative importance-weighted loss estimates of 1/2-Tsallis-INF differs from the true optimal arm. The main difficulty is that, at the logarithmic-regret scale, suboptimal arms are sampled with probability heuristically of order $1/t$. Consequently, importance weighting causes the cumulative estimator to fluctuate on the same linear scale as its mean separation. To overcome this obstacle, guided by a diffusion toy model, we construct a Lyapunov function for the gap process between the estimated cumulative loss of the optimal arm and that of the best competing arm. This leads to polynomial upper bounds on $e_t$: for learning rate $\eta_t=\alpha/\sqrt t$, $e_t$ decays at rate $t^{-2+\alpha^2\mu_{i_*}/4+\rho}$ for any $\rho>0$, where $\mu_{i_*}$ denotes the mean loss of the true optimal arm. We also establish a lower bound $\Omega(t^{-2-\varepsilon})$ for any $\varepsilon>0$, showing that the exponent $2$ is essentially tight.


Does Synthetic Data Help? Empirical Evidence from Deep Learning Time Series Forecasters

Hugo Cazaux ⋅ Eyjolfur Asgeirsson ⋅ Hlynur Stefansson

Synthetic data has transformed language model training, yet its role in time series forecasting remains poorly understood. We present a large-scale empirical study: nine experiment groups, 4,218 runs systematically evaluating synthetic time series augmentation across five architectures, four synthetic signals and seven datasets. The effect is sharply architecture-conditional: channel-mixing models (TimesNet, iTransformer) benefit in the majority of trials, while channel-independent models (DLinear, PatchTST) are consistently degraded. In selected low-resource settings the gains are striking: TimesNet trained on only 10\% of Weather data with synthetic augmentation surpasses the full-data baseline (4 of 16 sparsity-dataset combinations). Averaged across all architectures, augmentation hurts in 67\% of trials. We further find that only the Seasonal-Trend generator reliably helps across the tested benchmarks, and that hard curriculum switching is actively harmful (+24\% MSE degradation). These results provide concrete, actionable guidelines on how to use synthetic data: use synthetic augmentation with channel-mixing architectures, use gradual annealing schedules, and treat low-resource augmentation as architecture- and dataset-dependent. Code is available at an \href{https://github.com/anonmirror/synthetic-timeseries}{anonymized repository}.


Does Your Large Language Model Have An Intuitive Sense of The Difficulty of A Question?

Ruitao Wang ⋅ Jinyi Liu ⋅ Hongyao Tang ⋅ Rong Cheng ⋅ Yi Ma ⋅ Shaojin Ma ⋅ Yiwen Zhu ⋅ Hebin Liang ⋅ Gangyi Zhao ⋅ Pengyi Li ⋅ YAN ZHENG ⋅ Jianye Hao

Difficulty perception is essential for adaptive reasoning in large language models (LLMs). Previous studies rely on training auxiliary models or using extra reasoning rollouts to estimate difficulty, which incurs high computational costs. In this paper, we identify an intrinsic property of LLMs: their internal representations, even before explicit reasoning, encode an informative and useful signal that correlates with problem difficulty. Inspired by this property, we propose Referential Latent Difficulty Perception (RLDP), a training-free and rollout-free method that estimates difficulty directly from hidden activations in a single forward pass, and requires only minimal reference problems. Additionally, we introduce RLDP-AdaSwitch, a lightweight controller that dynamically allocates reasoning effort based on the difficulty signals provided by RLDP, enabling efficient trade-offs between accuracy and compute. Our experimental results across multiple LLMs and diverse datasets, including math reasoning, code generation, and QA, demonstrate that RLDP provides stable and effective difficulty discrimination. This further powers RLDP-AdaSwitch to achieve 1.34×–2.00× efficiency compared to rollout-based methods, while matching the performance of training-based methods.


DoG: Sniffing Out Overconfidence in LLM Agents via Post-hoc Trajectory Restructuring

Hyunjun Jeon ⋅ Dongha Lim ⋅ Kunwoong Kim ⋅ Daewon Choi ⋅ Jinwoo Shin

The overconfidence of Large Language Model (LLM) agents poses a critical challenge to their reliable deployment for complex, high-stakes tasks. Current confidence estimation methods predominantly treat agent execution as a flat, one-dimensional, time-ordered sequence. This oversimplification fundamentally fails to capture the complex logical dependencies, branching paths, error propagation, and tool interactions inherent in agentic reasoning. To address this, we introduce DAG of Grounds (DoG), a novel method for agent confidence estimation via post-hoc trajectory restructuring. By transforming linear trajectories into Directed Acyclic Graphs (DAGs), DoG explicitly maps the grounding of an agent's final answer to its intermediate reasoning steps and tool outputs. We evaluate DoG on several benchmarks, GAIA, GPQA, and HLE, which involve long-horizon reasoning and iterative tool use, and show that DoG improves calibration performance across standard metrics such as ECE, Brier score, and AUROC. DoG outperforms existing calibration methods designed for general LLM outputs, performs comparably to or better than methods specifically designed for agent trajectories, and remains robust across different LLM backbones and tool use settings. Supported by an interactive diagnostic tool, our framework provides unprecedented interpretability for diagnosing opaque failure modes. Together, these contributions establish a structured, graph-based paradigm for confidence estimation that successfully "sniffs out" overconfidence in autonomous systems.

Large language models spend enormous parameter capacity approximating arithmetic, a deterministic function with an exact solution, and still fail reliably on multi-digit operations. We propose Arithmetic Residual Blocks (ARBs): frozen, differentiable modules inserted into the transformer forward pass that compute exact integer arithmetic using Residue Number System (RNS) encoding on unit circles. The model learns only the interface, a gated injection pathway and a low-rank adapter on the language model head, while the base model and all arithmetic computation remain frozen. The architecture embodies a design principle we call Don't Learn What You Can Compute: any deterministic function expressible as tensor operations can be embedded as a frozen residual block, and the model will learn to route through it because the gradient reward for exact answers dominates internal approximation. We validate this principle by inserting ARBs into a frozen SmolLM2-360M model, training only 1.7M parameters (0.47% of the base model). On exact-match evaluation across addition, subtraction, multiplication, and division with operands up to 3 digits, the augmented model achieves 99.9%+ accuracy across all operations, compared to 4.6% mean accuracy for the unmodified base model. A logit analysis of the frozen base model confirms that the injection pathway overrides the base model's prior uniformly across operations, and that the residual errors (0.047%) concentrate at later digit positions in autoregressive generation, not at any particular operation. These findings motivate the hypothesis that pre-training with ARBs, where no misaligned frozen representations exist, would yield ceiling performance across all operations.


D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models

Dengyang Jiang ⋅ Xin Jin ⋅ Dongyang Liu ⋅ Zanyi Wang ⋅ Mingzhe Zheng ⋅ Ruoyi Du ⋅ Xiangpeng Yang ⋅ Qilong Wu ⋅ Zhen Li ⋅ Peng Gao ⋅ Harry Yang ⋅ Steven Hoi

The landscape of high-performance image generation models is currently shifting from the inefficient multi-step ones to the efficient few-step counterparts (e.g, Z-Image-Turbo and FLUX.2-klein). However, these models present significant challenges for direct continuous supervised fine-tuning. For example, applying the commonly used fine-tuning technique would compromise their inherent few-step inference capability. To address this, we propose D-OPSD, a novel training paradigm for step-distilled diffusion models that enables on-policy learning during supervised fine-tuning. We first find that the modern diffusion models, where the LLM/VLM serves as the encoder, can inherit its encoder's in-context capabilities. This enables us to formulate the training as an on-policy self-distillation process. Specifically, during training, we make the model act as both the teacher and the student with different contexts, where the student is conditioned only on the text feature, while the teacher is conditioned on the multimodal feature of both the text prompt and the target image. Training minimizes the two predicted distributions over the student's own roll-outs. By optimizing on the model's own trajectory and under its own supervision, D-OPSD enables the model to learn new concepts, styles, etc., without sacrificing the original few-step capacity.


Do You CARE to Generalize? Extracting Robust Concept Directions from LLMs

Sweta Karlekar ⋅ Claudia Shi ⋅ Aahlad Manas Puli ⋅ Carolina Zheng ⋅ Maggie Makar ⋅ John Bowlan ⋅ Michal Kucer ⋅ David Blei

Transformer-based large language models (LLMs) encode many high-level concepts as linear directions in the latent activation space. Once identified, these directions support both measurement, the quantification of a concept's presence, and intervention, the steering of the model's behavior. In practice, however, a direction learned from one context often fails when applied to prompts from a new context. This poor transfer can have two distinct sources: spurious correlation between the concept label and dataset-specific features, and genuine heterogeneity in how the concept is encoded across contexts. Methods from out-of-distribution (OOD) literature can address the first, but worsen the second by discarding context-specific structure that may carry concept signal. We introduce Context-Aided Representation Extraction (CARE), which bridges OOD generalization and mechanistic interpretability. Given activations labeled with a concept and an environment, CARE jointly learns a shared direction, optimized to be invariant across environments, and orthogonal environment-specific residuals that capture how the concept varies by context. We evaluate CARE on subject--verb agreement, refusal of harmful prompts, and toxicity. CARE produces directions that measure concepts more reliably under distribution shift and in unseen environments and datasets than existing methods while remaining effective for intervention. Its decomposition also supports a swap-cost diagnostic that identifies when context-specific structure carries concept-relevant signal.


D-PACE: Dynamic Position-Aware Cross-Entropy for Parallel Speculative Drafting

Tianyu Wu ⋅ Yu Yao ⋅ Zhenting Qi ⋅ Han Zheng ⋅ Chengxi Zhang ⋅ Zhuohan Wang ⋅ Haoran Ma ⋅ Zichun Liao ⋅ Himabindu Lakkaraju ⋅ Ju Li ⋅ Yilun Du

Speculative decoding accelerates LLM inference by having a small drafter propose tokens that a larger target model verifies in parallel. Current SOTA diffusion-based parallel drafters such as dFlash predict the full $B$-token block in one forward pass, allowing deeper drafters and higher accepted length. However, across the field, from autoregressive tree drafters to parallel block drafters, training losses use fixed per-position weights that do not adapt to the drafter's current bottleneck position. We derive per-position training weights directly from the expected accepted block length, so that each position's weight reflects its actual contribution to acceptance. The resulting loss, D-PACE (Dynamic Position-Aware Cross-Entropy), automatically identifies where the block's acceptance bottleneck lies and shifts training signal there as the drafter improves. D-PACE consistently improves both wall-clock speedup and accepted length across diverse benchmarks and positions as a drop-in replacement with negligible training overhead and requires no changes to the drafter architecture or inference pipeline, making it broadly applicable to general model architectures.

Multimodal diffusion large language models (dLLMs) generate text through iterative masked denoising with bidirectional attention, offering a compelling alternative to autoregressive vision-language modeling. However, each denoising step requires a full forward pass over long multimodal sequences in which visual tokens dominate the sequence length, and a fixed step budget further compounds this quadratic per-step cost. Existing inference acceleration approaches for multimodal dLLMs characterize attention redundancy implicitly, motivating heuristic designs from isolated attention visualizations rather than principled analysis. We close this gap with DRAMA, a training-free acceleration framework grounded in a systematic dissection of attention redundancy in multimodal dLLMs. By probing attention tensors along spatial and temporal axes, we reveal two persistent redundancy structures: (i) attention mass consistently concentrates on a sparse subset of visual tokens, and (ii) attention distributions rapidly stabilize across denoising steps after an initial formation phase. These findings motivate two complementary mechanisms: Spatial-Redundancy-Aware Visual Token Pruning (SVTP) and Sufficiency-Guided Adaptive Stopping (SGAS). Together, DRAMA achieves an average $18.67\times$ inference speedup across six vision-language benchmarks without training, architectural modification, or attention caching, while maintaining competitive accuracy. Our results establish a principled and strong acceleration baseline for multimodal dLLM inference.


DressWild: Feed-Forward Pose-Agnostic Garment Sewing Pattern Generation from In-the-Wild Images

Zeng Tao ⋅ Ying Jiang ⋅ Yunuo Chen ⋅ Tianyi Xie ⋅ Huamin Wang ⋅ Ying Nian Wu ⋅ Yin Yang ⋅ Chenfanfu Jiang

Recent advances in garment pattern generation have shown promising progress. However, existing feed-forward methods struggle with diverse poses and viewpoints, while optimization-based approaches are computationally expensive and difficult to scale. This paper focuses on sewing pattern generation for garment modeling and fabrication applications that demand editable, separable, and simulation-ready garments. We propose DressWild, a novel feed-forward pipeline that reconstructs physics-consistent 2D sewing patterns and the corresponding 3D garments from a single in-the-wild image. Given an input image, our method leverages vision–language models (VLMs) to normalize pose variations at the image level, then extract pose-aware, 3D-informed garment features. These features are fused through a transformer-based encoder and subsequently used to predict sewing pattern parameters, which can be directly applied to physical simulation, texture synthesis, and multi-layer virtual try-on. Extensive experiments demonstrate that our approach robustly recovers diverse sewing patterns and the corresponding 3D garments from in-the-wild images without requiring multi-view inputs or iterative optimization, offering an efficient and scalable solution for realistic garment simulation and animation.


Drift-Resistant Navigation World Model with Anchored Epipolar Guidance

Po-Chien Luan ⋅ Zimin Xia ⋅ WUYANG LI ⋅ Yang Gao ⋅ Alexandre Alahi

We propose Drift-Resistant Navigation World Model, a generative model that mitigates both perceptual drift and geometric drift in conventional rollout-based navigation world models. Existing methods recursively feed generated content into subsequent steps, causing noise accumulation and degraded predictions, i.e., perceptual drift. Meanwhile, their predictions often deviate from the agent’s motion, resulting in geometry drift. We address both types of drift by redesigning world-model prediction as an anchor-guided rollout. Instead of rolling out every frame sequentially, we first predict sparse future anchors that serve as stable long-range targets, and then generate intermediate frames within each chunk conditioned on both past context and future anchors. Importantly, these sparse anchors also provide geometric constraints, supported by bidirectional epipolar geometry, to localize where corresponding content should appear in the intermediate frames. Experiments on four benchmarks demonstrate consistent improvements over strong baselines in long-horizon visual quality, geometric consistency, and multi-view coherence. These gains further translate into improved downstream planning performance under the same planners, highlighting the importance of drift-resistant, geometry-aware prediction for reliable navigation world models.

Generating neural network weights in a single forward pass promises to amortize model training for fast sampling, transfer, and model-set construction. Yet accuracy alone is insufficient: existing generators can map diverse latent inputs to weights that remain functionally close to their training checkpoints. We study this functional mode collapse in the one-step regime and introduce DriftWeight, a MeanFlow-based generator with temperature-scaled repulsive drift in a checkpoint-augmented PCA proxy space. The method uses structured negatives from the current batch and optimization history to push generated weights away from observed clusters while preserving task performance. Our key finding is that repulsion is not automatically compatible with one-step inference. In a controlled ablation, applying the drift loss at the same boundary evaluation used for sampling collapses one-step accuracy even though multi-step integration remains accurate. DriftWeight mitigates this boundary-localized failure with a consistency objective that anchors the inference point to the training-weight manifold. Across MNIST, Fashion-MNIST, and CIFAR-10, DriftWeight matches or exceeds reported 100-step DeepWeightFlow accuracy in a single forward pass. On MNIST, where we directly evaluate error-set functional overlap, it reduces MaxIoU from the DeepWeightFlow regime of approximately $0.82$ to $0.66$. On CIFAR-100 with a ViT-Base target ($\sim$86M parameters), one-step samples remain within $1.7$ percentage points of the training-population mean.

Long-horizon time series forecasting requires a careful balance between predictive accuracy and computational cost. Segmented temporal modeling has therefore become a practical strategy for reducing the burden of long-range prediction. However, real-world time series often contain superposed multi-periodic components, whose spectral structures vary across time, variables, and samples. Such variability limits fixed aggregation or frequency-modeling mechanisms, making them less adaptive to non-stationary periodic patterns and potentially causing informative components to be attenuated over long forecasting horizons. To address this limitation, we propose DSR-TSF, a spectrum-driven dynamic routing framework for long-horizon forecasting. DSR-TSF adaptively selects and combines multiple period-specific branches conditioned on the spectral characteristics of the input sequence, thereby improving the model's ability to capture non-stationary periodic structures. It further learns period-aware dynamic weights to modulate the contributions of multi-periodic patterns while retaining the segmented prediction structure. Without complex resampling or explicit period alignment, DSR-TSF provides a compact mechanism for modeling spectral heterogeneity and achieves competitive performance across seven datasets with diverse dimensionalities and temporal characteristics.


DSSP: Diffusion State Space Policy with Hierarchical Full-History Conditioning

Zhiyuan Guan ⋅ Jianshu Hu ⋅ Han Fang ⋅ Yunpeng Jiang ⋅ Yize Huang ⋅ Shujia Li ⋅ Xiao Li ⋅ Yutong Ban

Diffusion-based imitation learning has shown strong promise for robot manipulation. However, most existing policies condition only on the current observation or a short window of recent observations, limiting their ability to resolve history-dependent ambiguities in long-horizon tasks. To address this, we introduce DSSP, a history-conditioned Diffusion State Space Policy that enables efficient, full-history conditioning for robot manipulation. Leveraging the continuous sequence modeling properties of State Space Models (SSMs), our history encoder effectively compresses the entire observation stream into a compact context representation. To ensure this context preserves critical information regarding future state evolution, the encoder is optimized with a dynamics-aware auxiliary training objective. This high-level context representation is then seamlessly fused with recent state observations to form a hierarchical conditioning mechanism for action generation. Furthermore, to maintain architectural consistency and minimize GPU memory overhead, we also instantiate the diffusion backbone itself using an SSM. Extensive experiments across simulation benchmarks and real-world manipulation tasks show that DSSP achieves state-of-the-art performance with a significantly smaller model size, demonstrating superior efficiency of the hierarchical conditioning in capturing crucial information as the history length increases.


Dual Feature-Relational Alignment for Transferable Targeted Attacks on MLLMs

Dongshen Han ⋅ Xiaorong Tian ⋅ Sheng Zheng ⋅ Zishan Huang ⋅ Boyu Deng ⋅ Wei Dong ⋅ Zhijie Li ⋅ Yang Yang ⋅ Chaoning Zhang

Multimodal large language models (MLLMs) remain vulnerable to transferable adversarial examples. However, targeted transferable attacks are particularly challenging, since perturbations must inject specific target semantics while generalizing from surrogate models to unseen black-box MLLMs. Existing methods mainly rely on coarse global feature alignment or unstable local matching, which tend to overfit surrogate-specific representations and fail to preserve spatially consistent local structures. In this paper, we propose Dual Feature-Relational Alignment Attack (DFRA-Attack), a locality-aware framework for targeted transferable attacks on MLLMs. To capture local fine-grained semantics, DFRA-Attack introduces semantic-aware alignment, which aligns adversarial and target images in a shared local observation space using saliency-guided shared masking, ensuring that both images are constrained under strictly consistent visible regions. Beyond semantic-aware alignment, DFRA-Attack introduces relational-aware alignment, which preserves target-consistent inter-region dependencies by jointly aligning gram-based feature relational matrices and attention-based interaction maps constructed from visible regional embeddings. Furthermore, we introduce a temperature-annealed dynamic reweighting mechanism to adaptively balance multi-surrogate optimization, coupled with a progressive strategy that refines the adversarial image from masked-view local alignment to full-image target consistency. Extensive experiments on open-source, closed-source, and reasoning MLLMs demonstrate that DFRA-Attack achieves significantly stronger targeted transferability and higher semantic fidelity than state-of-the-art baselines.


Dual-Rate Diffusion: Accelerating diffusion models with an interleaved heavy-light network

Grigory Bartosh ⋅ David Ruhe ⋅ Emiel Hoogeboom ⋅ Jonathan Heek ⋅ Thomas Mensink ⋅ Tim Salimans

Diffusion models achieve state-of-the-art generative performance but suffer from high computational costs during inference due to the repeated evaluation of a heavy neural network. In this work, we propose Dual-Rate Diffusion, a method to accelerate sampling by interleaving the execution of a heavy high-capacity context encoder and a light efficient denoising model. The context encoder is evaluated sparsely to extract high-dimensional features, which are effectively reused by the light denoising model at every step to refine the sample efficiently. This approach significantly accelerates inference without compromising sample quality. On ImageNet benchmarks, Dual-Rate Diffusion matches the performance of standard baselines while reducing computational cost by a factor of $2$--$4$. Furthermore, we demonstrate that our method is compatible with distillation techniques, such as Moment Matching Distillation, enabling further efficiency gains in few-step generation.


DUDS: Dual-stage Data Selection for Efficient Reinforcement Learning with Verifiable Rewards

Hongling Zheng ⋅ Li Shen ⋅ Zichuan Lin ⋅ Jiafei Lyu ⋅ Zhicong Lu ⋅ Shuhan Xu ⋅ Yong Luo ⋅ Deheng Ye ⋅ Dacheng Tao

Recent advances in large language models (LLMs) have leveraged reinforcement learning with verifiable rewards (RLVR) to enhance reasoning capabilities. However, RLVR typically relies on massive training data and extensive rollouts, posing substantial challenges to computational resources. Existing data selection approaches address this challenge in isolation: offline methods perform static, one-time dataset pruning that cannot adapt to the model’s evolving learning needs throughout training, while online methods conduct per-iteration filtering at the expense of significant additional computation. In this paper, we propose DUDS, a Dual-stage Data Selection framework for RLVR, which organically integrates the strengths of both paradigms to improve training efficiency while maintaining competitive performance. Specifically, the offline stage curates a candidate pool from the full dataset by jointly considering quality score and sample diversity via Determinantal Point Processes, providing an informative warm start for subsequent training. The online stage employs Bayesian posterior updates to estimate sample pass rates without introducing additional rollouts, and combines them with a freshness metric for data sampling, then applies asymmetric filtering to phase out mastered or intractable problems. Extensive experiments across four reasoning benchmarks demonstrate that DUDS consistently outperforms existing methods in both offline and online data-selection scenarios, achieving competitive performance with significantly less data and computation. Code is available at here.


DyCoRM: Dynamic Criterion-Aware Reward Modeling for Text-to-Image Generation

Jiaying Qian ⋅ Ziheng Jia ⋅ Qian Zhang ⋅ Zicheng Zhang ⋅ Jiayi Guo ⋅ Junqi Zhang ⋅ Guangtao Zhai ⋅ Xiongkuo Min

With the continued advancement of text-to-image (T2I) generation, producing high-quality images is becoming increasingly attainable; consequently, user demands are shifting toward images that better satisfy their specific requirements. As reward models play an increasingly important role in assessing whether generated images align with user preference, this trend introduces an important challenge for reward modeling: rather than relying solely on static and general evaluation dimensions, reward models should account for the task-relevant and fine-grained criteria through which users assess whether generated images meet their specific requirements. To address this challenge, we propose DyCoRM, a dynamic, criterion-aware reward model that grounds task-relevant criteria and performs criterion-aware preference comparison. To support this setting, we construct DyCoDataset-20K, which provides dynamic criteria together with criterion-level annotations, and further derive DyCoBench-1K, a benchmark for systematically evaluating reward models under dynamic criteria. We further introduce DyCoPick, which applies criterion-aware reward modeling to selecting T2I images. Our contributions establish the first reward modeling framework for dynamic and fine-grained evaluation and practical application in T2I generation.


DynaFLIP: Rethinking Robotics Perception via Tri-Modal-Dynamics Guided Representation

Jusuk Lee ⋅ Seungjae Lee ⋅ Jonghun Shin ⋅ Hoseong Jung ⋅ Sungha Kim ⋅ Daesol Cho ⋅ H. Jin Kim ⋅ Jia-Bin Huang ⋅ Furong Huang

Robot manipulation critically depends on perception that preserves the action-relevant aspects of a scene. Yet most robot learning pipelines are built upon visual encoders pre-trained for static recognition or vision-language alignment, leaving motion understanding to downstream policies. We introduce DynaFLIP, a dynamics-aware multimodal pre-training framework that pushes motion understanding upstream into perception. We construct image–language–3D flow triplets from heterogeneous human and robot videos, and use these triplets as training-time supervision to shape an image-only encoder. Our key idea is to encourage the three modalities to span a small simplex volume in the shared hyperspherical space—a smaller simplex volume indicating stronger alignment. To avoid the geometric ambiguity and trivial collapse of naive volume minimization, we combine simplex-volume minimization with a cosine regularizer and a contrastive objective. Our analyses show that DynaFLIP focuses on control-relevant regions critical for manipulation. The resulting dynamics-aware representations serve as reusable visual backbones and consistently outperform baselines across diverse downstream policies, including VLAs. We validate this across diverse simulation and real-world setups, with gains reaching +22.5\% under out-of-distribution scenarios. Our results suggest that robot generalization improves when visual representations are trained to encode not just what is present, but how the world changes under action.


Dynamic Causal Structure Discovery for Autoregressive Visual Generation

Songliang Guo ⋅ Qi Yan ⋅ Jianzhou Wang ⋅ Xinfu Liu ⋅ Cheng Zhen ⋅ Yirui Wu ⋅ Lixin Yuan ⋅ Wenxiao Zhang ⋅ Jun Liu

Autoregressive (AR) models achieve strong performance in visual generation, yet their prefix-based factorization imposes a rigid one-dimensional dependency structure. We show that this structure is fundamentally suboptimal: any predefined ordering necessarily induces a trade-off between contextual sufficiency and necessity, resulting in both missing and redundant dependencies. To address this limitation, we propose Dynamic Causal Structure Discovery, an inference-time framework that replaces static conditioning with dynamic, content-dependent causal parent sets. By interpreting the Transformer as a causal graph, we reformulate generation as the problem of constructing a minimal parent set for each token. To make this objective tractable, we approximate $\mathrm{do}$-interventions as vector-space ablations and derive a Fisher-based criterion that characterizes and corrects structural deviations during decoding. Our framework transforms AR factorization into an order-agnostic dependency structure, yielding improved generation quality and global coherence without retraining.

Quantifying uncertainty in time series forecasting is particularly demanding because sequential data exhibit temporal dependence and are prone to distributional changes. Conformal inference has emerged as a powerful uncertainty quantification approach through the construction of prediction sets. Recent advances have introduced online conformal methods that adaptively adjust prediction thresholds through feedback mechanisms. However, the existing feedback mechanism typically relies solely on miscoverage indicators (actual feedback)—whether the true label falls within the interval at each time step—while overlooking the empirical prediction threshold (estimated feedback) that is derived from the oracle conformal method. In this paper, we propose $\textit{Dynamic Dual-feedback Conformal Inference}$ (DDCI), which incorporates a dual-feedback mechanism consisting of $\textit{actual feedback}$ and $\textit{estimated feedback}$. The former drives the primary adjustment of the intervals based on true observations, while the latter dampens excessive expansions or contractions by leveraging empirical thresholds from conformal inference during updates. By balancing these two signals, DDCI achieves more stable and narrower prediction intervals in sequential settings while adhering to the target coverage rate in practice.


Dynamic Regulatory Graph Learning for Histology-to-Spatial Transcriptomics

Yikai Luo ⋅ Peng Zhang ⋅ Heyang Zhao ⋅ Hongming Shan ⋅ Wenjian Wang

Virtual spatial transcriptomics aims to predict spatial gene expression from histopathology images. However, gene expression is not governed by morphology alone; regulatory relationships are structured, context-dependent, and shaped by the local tissue microenvironment. Existing visual-to-expression models and fixed-prior graph methods struggle to align molecular dependencies with spatially varying histological contexts. We propose DragH2ST, a dual-stage retrieval-augmented dynamic regulatory graph learning framework that constructs task-specific regulatory manifolds and rewires them at the spot level according to local histology. During training, DragH2ST retrieves and integrates heterogeneous biomedical knowledge to build a task-specific regulatory prior, guiding gene representations toward biologically plausible and context-adaptive regulatory structures. To capture microenvironment-dependent regulation, context-gated graph rewiring performs spot-level soft rewiring of gene-regulatory edges conditioned on histology-derived context, enabling adaptive graph reasoning without dense graph reconstruction. At inference, DragH2ST retrieves morphologically similar historical spot prototypes as case-level references, providing evidence-based support for interpretation and biological plausibility assessment. Experiments on multiple benchmarks show that DragH2ST improves prediction accuracy, preserves gene-gene regulatory topology, and captures spatially heterogeneous expression patterns.


Dynamic Spectral Federated Graph-Level Clustering

Wenxin Zhang ⋅ Xi Xuan ⋅ Guangzhen Yao ⋅ Feng Zhou ⋅ Cuicui Luo ⋅ Renxiang Guan ⋅ Zonghao Ying ⋅ Xiaojian Lin ⋅ yangchen zeng ⋅ Renda Han

Federated graph clustering (FGC) aims to partition distributed graphs into different groups while preserving data privacy. Existing FGC methods focus on spatial-domain information aggregation, while how to use expressive spectral-domain signals of graphs for FGC still remains unexplored. Main challenges are twofold: (1) Collecting and integrating semantic signals that fully express the critical patterns of graphs; (2) Obtaining spectral consensus while preserving local properties of clients. To address these challenges, we propose dynamic spectral federated graph-level clustering (DySFGC). DySFGC designs learnable polynomial filters to capture abundant graph signals and employs spectral contrastive learning and reconstruction to explore the latent semantic information. Besides, DySFGC develops dynamic spectral consensus, which updates the consensus via frequency-specific channels and conducts the dynamic updating process by evaluating the discrepancies between the server and clients. Extensive experiments on fifteen datasets demonstrate the superiority of DySFGC over eleven baselines.


DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation

Haozhe Xie ⋅ Beichen Wen ⋅ Jiarui Zheng ⋅ Zhaoxi Chen ⋅ Fangzhou Hong ⋅ Haiwen Diao ⋅ Ziwei Liu

Manipulating dynamic objects remains an open challenge for Vision-Language-Action (VLA) models. Although recent VLAs generalize well in static manipulation, dynamic scenes introduce a latency-induced perception–execution mismatch: object states continue to evolve during inference, making actions predicted from past observations stale at execution time. We present DynamicVLA, a latency-aware VLA model for dynamic object manipulation. It combines a compact 0.4B architecture and convolutional vision encoder for efficient multimodal inference with a continuous inference schedule that overlaps reasoning and execution for non-blocking control. Latent-aware Action Streaming then discards latency-invalid action prefixes and executes only the temporally valid suffix of each predicted chunk, preserving action-time alignment under dynamic object motion. To fill the missing foundation of dynamic manipulation data, we introduce the Dynamic Object Manipulation (DOM) benchmark, built with an automated collection pipeline that gathers 200K synthetic episodes across 2.8K scenes and 206 objects, and enables fast collection of 2K real-world episodes without teleoperation. Extensive evaluations in simulation and on real robots show that DynamicVLA improves dynamic manipulation success under changing object motion, perception-heavy instructions, and unseen motion patterns.


DynT2I-Eval: A Dynamic Evaluation Framework for Text-to-Image Models

Wang Juntong ⋅ Wang Jiarui ⋅ Huiyu Duan ⋅ Lewei Li ⋅ Guangtao Zhai ⋅ Xiongkuo Min

Existing text-to-image (T2I) benchmarks largely rely on fixed prompt sets, leaving them vulnerable to overfitting and benchmark contamination once publicly released and repeatedly reused. In this work, we propose DynT2I-Eval, a fully automated dynamic evaluation framework for T2I models. It constructs a structured visual semantic space from long-form descriptions, decomposing prompts into controllable dimensions (e.g., subject, logical constraint, environment, and composition). This enables the continuous generation of fresh prompts via task-specific spaces and difficulty-aware sampling. DynT2I-Eval evaluates model performance across text alignment, perceptual quality, and aesthetics. Heterogeneous outputs are unified into prompt-conditioned pairwise comparisons, allowing a dynamic scheduler, micro-batch aggregation, and weighted Bayesian updates to maintain a stable online leaderboard despite changing prompt distributions and model injection. Experiments with independently sampled prompt streams demonstrate that continually refreshed prompts provide a robust evaluation protocol, reducing the impact of prompt-set-specific tuning. Simulations and ablations further confirm that the proposed ranking framework achieves a strong balance among cold-start convergence, late-entry discovery, and long-run ranking fidelity. An anonymous codebase is provided for review, which will be fully open-sourced upon publication.


DyPSI: Dynamic Physics Sensing via Joint Field and Sensor-Trajectory Generation

Yizhou Zhang ⋅ Panqi Chen ⋅ Lei Cheng ⋅ Ting Zhang ⋅ Jianlong Li ⋅ Jianfeng Shi ⋅ Shikai Fang

Physics sensing in evolving environments requires reconstructing dense spatiotemporal fields from sparse observations while adapting where future observations should be collected. Existing reconstruction and sensor-placement methods have made substantial progress, but they typically rely on static sensing assumptions: they either use a single layout across all frames or optimize placements independently at each frame, which limits their ability to produce temporally coherent sensing trajectories. We introduce DyPSI, a dynamic physics sensing framework that formulates this problem as the joint generation of future fields and sensor trajectories conditioned on historical observations. DyPSI lifts discrete sensor coordinates into continuous Sensor Position Fields, encodes fields and sensing configurations with a shared Functional Tucker representation, and trains a joint diffusion model to generate future fields and sensing trajectories simultaneously in the resulting latent space. A Gaussian-process temporal kernel correlates the perturbation injected during diffusion training, biasing the denoiser toward temporally smooth recoveries, and we further provide an analysis linking the EDM training objective to a sequence-aggregated A-optimality criterion and a kernel-induced bound on discrete trajectory variation. Experiments on turbulent flow, GLORYS12 sea-surface temperature, and 3D car aerodynamics show that DyPSI consistently outperforms static and frame-wise placement strategies, with substantial error reduction under tight sensor budgets.


ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning

Jingwei Song ⋅ Meng Chen ⋅ Jie Xiao ⋅ Qingnan Ren ⋅ Jiaqi Huang ⋅ Yangshen Deng ⋅ Senyu Tong ⋅ Wanyi Chen ⋅ Suli Wang ⋅ Zhisheng Chen ⋅ Ziqian Bi ⋅ Shuo Lu ⋅ Yiqun Duan ⋅ Xu Wang ⋅ Rymon Yu ⋅ Lynn Ai ⋅ Eric Yang ⋅ TIANYU SHI

Reinforcement learning (RL) is a critical stage in post-training large language models (LLMs), involving repeated interaction between rollout generation, reward evaluation, and centralized learning. Distributing rollout execution offers opportunities to leverage more cost-efficient inference resources, but introduces challenges in wide-area coordination and policy dissemination. We present ECHO-2, a distributed RL framework for post-training with remote inference workers and non-negligible dissemination latency. ECHO-2 combines centralized learning with distributed rollouts and treats bounded policy staleness as a user-controlled parameter, enabling rollout generation, dissemination, and training to overlap. We introduce an overlap-based capacity model that relates training time, dissemination latency, and rollout throughput, yielding a practical provisioning rule for sustaining learner utilization. To mitigate dissemination bottlenecks and lower cost, ECHO-2 employs peer-assisted pipelined broadcast and cost-aware activation of heterogeneous workers. Experiments on GRPO post-training of LLMs ranging from 4B to 32B parameters under real wide-area bandwidth regimes show that ECHO-2 significantly improves cost efficiency while preserving RL reward comparable to strong baselines.


Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD

Arseniy Andreyev ⋅ Pierfrancesco Beneventano

Recent findings by \citet{cohen_gradient_2021} demonstrate that when training neural networks with full-batch gradient descent at step size $\eta$, the largest eigenvalue $\lambda_{\max}$ of the full-batch Hessian consistently stabilizes around $2/\eta$. These results have significant implications for convergence and generalization. This stabilization, however, does not occur for mini-batch stochastic gradient descent (SGD), so the implications above do not directly transfer. We show that SGD trains in a different regime we term Edge of Stochastic Stability (\textsc{EoSS}). In this regime, what stabilizes at $2/\eta$ is \emph{Batch Sharpness}: the expected directional curvature of mini-batch Hessians along their corresponding mini-batch gradients. As a consequence, $\lambda_{\max}$---which is generally smaller than \emph{Batch Sharpness}---is suppressed, aligning with the long-standing empirical observation that smaller batches and larger step sizes favor flatter minima. We further discuss implications for mathematical modeling of SGD trajectories.

Euclidean combinatorial optimization problems (ECOPs), such as the Traveling Salesman Problem (TSP) and Capacitated Vehicle Routing Problem (CVRP), possess inherent symmetries under the two-dimensional Euclidean group E(2), including rotations, reflections, and translations. Existing learning-based methods, including recent diffusion-based methods, rely on data augmentation or regularization to approximate E(2)-equivariance. This paper presents EDISCO, the first discrete diffusion model for ECOPs with exact E(2)-invariant generative distributions over node-index solutions. EDISCO introduces an E(2)-equivariant edge-score network coupled with a categorical continuous-time Markov chain over discrete edge variables, and exact posterior sampling provides efficient multi-step inference. This design gives EDISCO a local geometric inductive bias: edge neighborhoods with the same relative geometry and combinatorial context are represented consistently regardless of absolute position or orientation, making learning more efficient and inference more robust than non-equivariant methods. EDISCO outperforms previous learning-based state-of-the-art solvers on synthetic TSP from 100 to 10000 nodes and CVRP from 50 to 2000 customers, while using only 33--50\% of the training instances. Trained only on uniform synthetic data, EDISCO also outperforms competing learning-based baselines under spatial distribution shift and CVRP constraint-tightness shift. Code is available at https://anonymous.4open.science/r/EDISCO-7F54.


EDITORS Know Your Style! Editing LoRA Subspaces for Stylistic Attribution and Imitation

Zixuan Wang ⋅ Gregory Kang Ruey Lau ⋅ Xiao Tian ⋅ Jue Fan ⋅ Caroline Chaux ⋅ Bryan Kian Hsiang Low

Disentangling writing style from semantic content is a fundamental challenge in literary text modeling. Style-content entanglement causes models to rely on semantic content for authorship attribution (AA) tasks and memorize author-specific content for style imitation (SI) tasks. We propose EDITORS (Editing LoRA Subspaces), a diagnose-then-deflate framework that adapts pretrained large language models and prioritizes stylistic features, making it effective for both AA and SI tasks. EDITORS first trains a diagnostic LoRA adapter using style-neutral statements derived from the training corpus to identify a content-orientedsubspace. It then trains the final style adapter under activation-space deflation that projects input activations away from the identified content-oriented directions. We demonstrate that EDITORS is empirically effective: On three AA benchmarks spanning literature, social media, and news, EDITORS achieves state-of-the-art accuracy, outperforming the strongest existing baseline by up to 29%. On SI, EDITORS generates stylistically faithful and diverse passages, outperforming baselines in achieving style similarity while adhering to content specifications.


Effect-Level Validation for Causal Discovery in Interactive Telemetry

Dang Van Cong Hoang ⋅ Luan Pham ⋅ Minh Nguyen

Causal discovery is increasingly used to analyze telemetry from interactive systems, but a plausible graph does not guarantee an identifiable or reliable product decision. We study this problem for one query: whether early competitive gameplay increases Day-1 retention in a deployed Role-playing game (RPG). We propose an admissibility-first framework that treats discovered graphs as structural hypotheses, keeps only graphs that identify the target effect and have adequate treatment-control overlap, and validates admissible effects through cross-algorithm stability, placebo, subsampling, and E-value diagnostics. On real telemetry, only about one third of discovery runs support the target effect after temporal and semantic constraints. Some admissible Directed Acyclic Graphs (DAGs) converge to a positive risk-difference average treatment effect (ATE), down from a raw retention gap, and the estimate survives all refutation checks. Graph-level recovery is therefore an inadequate proxy for causal reliability; discovery pipelines for decision support should be evaluated at the query and effect level.


Efficient Adjoint Matching for Fine-tuning Diffusion Models

Jeongwoo Shin ⋅ Dongsoo Shin ⋅ Yuchen Zhu ⋅ Wei Guo ⋅ Yongxin Chen ⋅ Joonseok Lee ⋅ Jaewoong Choi ⋅ Jaemoo Choi

Reward fine-tuning has become a common approach for aligning pretrained diffusion and flow models with human preferences in text-to-image generation. Among reward-gradient-based methods, Adjoint Matching (AM) provides a principled formulation by casting reward fine-tuning as a stochastic optimal control (SOC) problem. However, AM inevitably requires a substantial computational cost: it requires (i) stochastic simulation of full generative trajectories under memoryless dynamics, resulting in a large number of function evaluations, and (ii) backward ODE simulation of the adjoint state along each sampled trajectory. In this work, we observe that both bottlenecks are closely tied to the non-trivial base drift inherited from the pretrained model. Motivated by this observation, we propose Efficient Adjoint Matching (EAM), which substantially improves training efficiency by reformulating the SOC problem with a linear base drift and a correspondingly modified terminal cost. This reformulation removes both sources of inefficiency; it enables training-time sampling with a few-step deterministic ODE solver and yields a closed-form adjoint solution that eliminates backward adjoint simulation. On standard text-to-image reward fine-tuning benchmarks, EAM converges up to 4× faster than AM and matches or surpasses it across various metrics including PickScore, ImageReward, HPSv2.1, CLIPScore and Aesthetics.


Efficient Memory Crystallization for Graph Learning under Non-Stationary Distribution Shifts

Yue Hou ⋅ Ruomei Liu ⋅ Yingke Su ⋅ Wu Junran ⋅ Ke Xu

Deep graph learning models deployed in real-world systems often need to cope with non-stationary environments, where the underlying graph distribution drifts continually over time. Prevailing solutions rely on training auxiliary generative modules to synthesize memory graphs for cross-domain adaptation, which incurs substantial computational overhead and scales poorly under prolonged distribution shifts. We argue that a more economical path exists: rather than generating memory, one can crystallize it. To this end, we propose Efficient Memory Crystallization (EMC), a training-free test-time framework that distills each incoming graph domain into a compact, semantically faithful memory through a closed-form solution to a memory-oriented distribution-matching objective, thereby eliminating redundant domain information under continual covariate shifts. To preserve both generalizability and adaptability as the model traverses a long sequence of target domains, EMC further models inter-domain dependencies through state-evolving memories and admits a theoretically grounded, tighter generalization error bound than direct adaptation. Extensive experiments demonstrate the superior performance of EMC over state-of-the-art baselines on graphs under non-stationary distribution shifts, while reducing average runtime by 87.4% and GPU memory consumption by 92.4% relative to the recent competitor, making continual graph adaptation practical at scale.


Efficient Test-Time Adaptation For Robot Policies

Motasem Alfarra ⋅ Pietro Mazzaglia ⋅ Markus Peschl ⋅ Daniel Dijkman ⋅ Christos Louizos

Reinforcement-learning policies for robot locomotion can fail at deployment when operating conditions shift (e.g., rough or slippery terrain, payload changes, sensing degradation), and retraining or manual recalibration is often infeasible under on-device latency and memory constraints. We study test-time adaptation (TTA) for robotic control: adapting a pretrained policy online and self-supervisedly during execution, without access to rewards. Our starting point is a lightweight baseline that updates only the actor using a frozen pretrained critic as a deployment-time performance surrogate. Building on this, we introduce TEMPO, a test-time entropy-regularized, memory-guided policy optimization method that stabilizes critic-driven updates under severe shift. TEMPO combines (i) an entropy-promoting regularizer on the critic’s predictive distribution to mitigate overconfident extrapolation, and (ii) a score-based memory, inspired by prioritized replay, that prioritizes low-value (“hard”) observations and applies importance sampling to focus limited computation on the most informative deployment conditions. Across Go1, Go2, and G1 in MuJoCo Playground under four representative shifts and multiple architectures, TEMPO yields consistent gains, reaching up to 80\% improvement in a single-device setting. We further validate on a physical Unitree Go2 with an added calf payload, where fully on-device adaptation improves performance over the non-adapted baseline across real-world episodes.


E-GEO: A Testbed for Generative Engine Optimization in E-Commerce

Puneet S Bagga ⋅ Vivek Farias ⋅ Tamar Korkotashvili ⋅ Tianyi Peng ⋅ Yuhang Wu

With the rise of large language models (LLMs), generative engines have become powerful alternatives to traditional search, reshaping retrieval tasks. In e-commerce, for instance, conversational shopping agents now guide consumers to relevant products. This shift has created the need for generative engine optimization (GEO)---improving content visibility and relevance for generative engines. Despite its growing importance, current GEO practices are largely ad hoc, and their impacts remain poorly understood, especially in the e-commerce setting. We address this gap by introducing E-GEO, the first dataset built specifically for e-commerce GEO. E-GEO contains 13,747 realistic, multi-sentence consumer product queries, each paired with 10 relevant Amazon listings, capturing rich intent, constraints, preferences, and shopping contexts that existing datasets miss. Using this dataset, we conduct the first large-scale empirical study of e-commerce GEO across five frontier generative engines, seven popular LLM rewriters with a simple prompt, and fifteen hand-crafted rewriting heuristics. We further formulate GEO as a tractable optimization problem and develop a lightweight prompt meta-optimization algorithm that significantly improves over heuristic baselines. Notably, the optimized prompts reveal a stable, domain-agnostic pattern, suggesting the existence of a "universally effective" GEO strategy. Finally, we red-team the GEO system through both heuristic and optimization-based attacks and show that, under a simple in-prompt defense, gains from GEO reflect genuine content improvement rather than manipulation, anchoring GEO as a substantive, well-defined optimization problem.


EgoBabyVLM: Benchmarking cross-modal learning from naturalistic egocentric video data

Dongyan Lin ⋅ Phillip Rust ⋅ Angel Villar-Corrales ⋅ Alvin Tan ⋅ Mahi Luthra ⋅ Charles-Éric Saint-James ⋅ Robin Algayres ⋅ Rashel Moritz ⋅ Vanessa Stark ⋅ Surya Parimi ⋅ Jiayi Shen ⋅ Youssef Benchekroun ⋅ Yosuke Higuchi ⋅ Angelo Ortiz Tandazo ⋅ Martin Gleize ⋅ Nicolas Hamilakis ⋅ manel khentout ⋅ Sho Tsuji ⋅ Balázs Kégl ⋅ Juan Pino ⋅ Michael C Frank ⋅ Emmanuel Dupoux

Children acquire language with remarkable speed and robustness from limited linguistic input in ways that surpass today's best large language models. Recent research suggests this efficiency stems from learning through multimodal sensory experiences where semantic alignment between what children see and hear grounds language in the external world. However, vision-language models (VLMs) trained on curated web data fail to generalize to the sparse, weakly-aligned egocentric streams produced by wearable devices, embodied agents, and infant head-cams—and no fixed evaluation pipeline exists for measuring progress on this regime. We train VLMs on datasets with varying degrees of semantic alignment between visual and linguistic inputs, including naturalistic infant egocentric videos, and evaluate them with a comprehensive suite spanning multimodal language grounding and unimodal vision and language tasks. At the core of this suite is Machine-DevBench, a corpus-grounded benchmark of lexical and grammatical competence, automatically generated from the model's training vocabulary across logarithmic frequency bins to eliminate the train/eval mismatch and low statistical power of prior developmental benchmarks. Our results show that current VLM paradigms hinge on the tight semantic alignment of curated data and fail to exploit the weakly-aligned signal that dominates naturalistic egocentric input—the very regime in which humans thrive. To motivate progress, we introduce the EgoBabyVLM Challenge to drive the development of models capable of grounded language learning from the kind of naturalistic data that human infants experience.


Emergent Steering Beyond Endpoint Alignment in Chemical Reaction Models

Yili Shen ⋅ Kehan Guo ⋅ Haomin Zhuang ⋅ Xiangliang Zhang

Machine learning models for chemical reaction prediction are usually evaluated by endpoint accuracy, but endpoint correctness alone does not reveal how a model represents chemical transformation or why it fails. We introduce $\textbf{UniRxnAxis}$, a unified forward--retrosynthesis framework that models both tasks as direction-conditioned endpoint redistribution in a shared latent space. By aligning source states, transformed states, and encoded target endpoints, UniRxnAxis enables directional auditing of learned latent updates. This shared latent geometry allows the learned displacement (from the encoder source state $(S_0)$ to the decoder-facing transformed state $(S_1)$) to be decomposed into an endpoint-parallel component and an endpoint-orthogonal residual, which we call $\textbf{emergent steering}$. Our evaluation shows that emergent steering is functionally important: in retrosynthesis, it recovers most of the model's utility, remains effective under random, shuffled, and prototype controls, and directly affects which bond edits are promoted or suppressed. The decomposition also supports a proxy taxonomy of model-side failures, showing how reaction models can be audited beyond endpoint accuracy.


EMO: Pretraining Mixture of Experts for Emergent Modularity

Ryan Wang ⋅ Akshita Bhagia ⋅ Sewon Min

Large language models are typically deployed as monolithic systems, requiring the full model even when applications need only a narrow subset of capabilities, e.g., code, math, or domain-specific knowledge. Mixture-of-Experts (MoEs) seemingly offer a potential alternative by activating only a subset of experts per input, but in practice, restricting inference to a subset of experts for a given domain leads to severe performance degradation. This limits their practicality in memory-constrained settings, especially as models grow larger and sparser. We introduce EMO, an MoE designed for modularity—the independent use and composition of expert subsets—without requiring human-defined priors. Our key idea is to encourage tokens from similar domains to rely on similar experts. Since tokens within a document often share a domain, EMO restricts them to select experts from a shared pool, while allowing different documents to use different pools. This simple constraint enables coherent expert groupings to emerge during pretraining using document boundaries alone. We pretrain a 1B-active, 14B-total EMO on 1T tokens. As a full model, it matches standard MoE performance. Crucially, it enables selective expert use: retaining only 25\% (12.5\%) of experts incurs just a 1\% (3\%) absolute drop, whereas standard MoEs break under the same setting. We further find that expert subsets in EMO specialize at semantic levels (e.g., domains such as math or code), in contrast to the low-level syntactic specialization observed in standard MoEs. Altogether, our results demonstrate a path toward modular, memory-efficient deployment of large, sparse models and open new opportunities for composable architectures.


EmpathyChat: Structured Cognitive Reasoning in Empathetic Spoken Dialogue

Dingdong WANG ⋅ Shujie LIU ⋅ Jinyu Li ⋅ Yuxuan Hu ⋅ Yayue Deng ⋅ Yunrui Cai ⋅ Jincenzi Wu ⋅ Jianwei Yu ⋅ Helen Meng

Empathetic spoken dialogue is a sophisticated cognitive process that requires not only recognizing emotions but also inferring a user’s latent mental states to provide appropriate support. However, current SpeechLLMs often treat empathy as a direct input-to-response mapping, leading to "superficially warm" but emotionally hollow interactions. In addition, since empathy relies on a multi-stage process with strong inter-step dependency, errors at any intermediate step can cascade through subsequent steps and lead to inappropriate responses, while existing training paradigms lack mechanisms to precisely localize and improve such errors. In this work, we propose EmpathyChat, a unified framework that reformulates empathetic spoken dialogue as a structured cognitive reasoning process integrating perception, mental-state reasoning, and response generation. To support this paradigm, we first construct EmpathySet-400K, an acoustically rich dataset for multi-stage empathetic supervision. During the SFT stage, we strengthen acoustic grounding through proposed Acoustic-Anchored Attention (AAA). During the RL stage, we further introduce a novel stage-aware optimization objective with Step-Decomposed Credit Assignment (SDCA) to localize reasoning errors and mitigate cascaded error propagation. In addition, we introduce EmpathyEval, an expert-annotated benchmark for multi-dimensional empathy evaluation. Extensive experiments demonstrate that EmpathyChat achieves state-of-the-art performance in perception, reasoning, and response alignment. All resources, including model, code, dataset, and benchmark, will be publicly released.


End-to-End Differentiable Diffusion Conditioning for Physics-Informed Optimization

Ricardo Luna Gutierrez ⋅ Vineet Gundecha ⋅ Rahman Ejaz ⋅ Varchas Gopalaswamy ⋅ Riccardo Betti ⋅ Sahand Ghorbanpour ⋅ Aarne Lees ⋅ Soumyendu Sarkar

Optimizing high-dimensional physical systems under expensive evaluations and complex constraints is a central challenge in science and engineering. In offline model-based optimization (MBO), methods must improve designs using only a fixed dataset and no test-time access to the true objective. However, existing methods typically suffer from conservatism, unconstrained surrogate exploitation, or function as static inference-time samplers that cannot refine candidates against precise physical objectives. In this paper, we propose Differentiable Diffusion Conditioning (D2C), and show that deterministic sampling makes the reverse denoising chain differentiable with respect to the conditioning signal. D2C turns a frozen conditional diffusion model into an end-to-end gradient-based optimizer by directly optimizing conditioning variables at test time via gradients from frozen surrogate objectives and physics-informed penalties backpropagated through the reverse chain. We implement this idea in a composed proposer-evaluator framework and evaluate it on laser pulse optimization for inertial confinement fusion, sustainable data-center workload scheduling, and four Design-Bench tasks. Our method achieves the strongest results on the two real-world tasks and the best average rank across four Design-Bench tasks.


EngramState: Loadable Tool Priors for Efficient Function Calling

Seulkee Lee ⋅ Hyun-rae Jo ⋅ Dongkun Shin

Large language model (LLM) agents typically perform function calling by retrieving relevant tools and reinserting their specifications as textual prompts at inference time. Although retrieval reduces the number of candidate tools, selected tool specifications must still be repeatedly re-encoded, incurring substantial latency and memory overhead, especially in on-device environments. We argue that this inefficiency stems from treating tool knowledge as transient text rather than reusable execution memory. To address this, we propose EngramState, a state-centric framework that compiles tool specifications offline into reusable recurrent state priors and directly loads them at runtime for query-only inference. EngramState combines retention-aware state construction, channel-wise state compression, a lightweight same-backbone state retriever, and feedback-driven state routing. On the DroidCall benchmark with an RWKV-7 1.5B backbone, EngramState reduces prompt tokens from 985 to 22 (97.8\%) under the Top-4 retrieval setting and substantially lowers time-to-first-token (TTFT) on a Galaxy S25 Ultra while maintaining or improving function-calling accuracy. Furthermore, the proposed state retriever achieves competitive retrieval quality using only 8.24\,MiB of retrieval storage. These results demonstrate that reusable recurrent states can serve as an efficient execution-time memory interface for scalable on-device function calling.


Ensuring Deployment-Time Safety of Neural Network Controlled Systems via Localized Certificate Repair

Xiaoyang Lv ⋅ Peixin Wang ⋅ Jianhao Bai ⋅ Bai Xue ⋅ Min Zhang ⋅ Luke Ong

Barrier certificates provide formal safety guarantees for neural network controlled systems by ensuring that system trajectories avoid unsafe regions over time. However, after deployment, previously unseen obstacles or constraint changes may introduce new unsafe regions, invalidating the original certificate and potentially leading to unsafe behavior. We address this setting with $\textit{adaptive piecewise barrier certificates}$, which preserve the original certificate and policy in unchanged regions while introducing localized updates around newly emerged unsafe regions. The central question is how to perform such updates efficiently during system execution without retraining from scratch. Our key idea is to cast deployment-time adaptation as a localized certificate repair problem and reduce it to an optimization task that leverages the learned structure of previous barrier certificates, enabling fast updates without global retraining. We further design a staged repair framework that progressively transitions from localized repair to joint policy–certificate updates and, when necessary, localized retraining. Empirical results on four benchmarks with five new obstacles show that our method succeeds in all cases and completes within tens of seconds on average, whereas retraining from scratch fails in most cases. On instances where retraining succeeds, our method is \textbf{17.3$\times$} faster on average. It incurs only \textbf{16.2\%} performance degradation relative to the original policy.


Entangled Schrödinger Bridge Matching

Sophia Tang ⋅ Yinuo Zhang ⋅ Pranam Chatterjee

Simulating trajectories of multi-particle systems on complex energy landscapes is a central task in molecular dynamics (MD) and drug discovery, but remains challenging at scale due to computationally expensive and long simulations. Previous approaches leverage techniques such as flow or Schrödinger bridge matching to implicitly learn joint trajectories through data snapshots. However, many systems, including biomolecular systems and heterogeneous cell populations, undergo dynamic interactions that evolve over their trajectory and cannot be captured through static snapshots. To close this gap, we introduce Entangled Schrödinger Bridge Matching (EntangledSBM), a framework that learns the first- and second-order stochastic dynamics of interacting, multi-particle systems where the direction and magnitude of each particle's path depend dynamically on the paths of the other particles. We define the Entangled Schrödinger Bridge (EntangledSB) problem as solving a coupled system of bias forces that entangle particle velocities. We show that our framework accurately simulates heterogeneous cell populations under perturbations and rare transitions in high-dimensional biomolecular systems. Anonymous code is provided at https://anonymous.4open.science/r/EntangledSBManon.


EnvFaultBench: Benchmarking LLM Agents on Software Environment-Fault Troubleshooting

Siqi Zhong ⋅ Haiyang Shen ⋅ Mugeng Liu ⋅ Chongyang Pan ⋅ Yun Ma

Large language model (LLM) coding agents are increasingly used as general-purpose software engineering assistants, yet existing benchmarks focus on source-code patches or environment setup from scratch. A common class of failures remains unaddressed: the code is correct but the environment is faulty due to dependency conflicts, configuration drift, stale caches, or resource contention. We present EnvFaultBench, to the best of our knowledge the first benchmark targeting environment-fault troubleshooting. It contains 348 Dockerized instances derived from real GitHub issues across 75 open-source projects in three ecosystems (Python, TypeScript/JavaScript, JVM), covering 23 fault types in three root-cause layers. Hidden functional oracles accept any command sequence that restores target behavior. Evaluating 10 LLMs (4 proprietary, 6 open-weight) under a unified agent framework, we find that the best model (GPT-5.5) resolves 65.8% while a zero-shot baseline resolves only 10.9%. The gap between strong and weak models is strongly associated with diagnostic efficiency under our bounded protocol: stronger models commit to a repair within 2–3 steps (Spearman ρ = −0.85 between first-fix latency and solve rate), while weaker models exhaust the budget on undirected exploration. Data and code are publicly available.


Epistemic Social Learning: Latent Behavioral Structure under Endogenous Multi-Agent Interaction

Jainendra Shukla ⋅ Dhruv Jaiswal ⋅ Divyanshi Beniwal ⋅ Kiriti Kanjilal

In many interactive settings, from social and institutional environments to multi-agent systems, agents must infer latent behavioral structure from limited signals while their own actions shape the data they observe. This endogeneity violates the stationary, exogenous assumptions underlying standard likelihood-based methods, leading to failures in adaptation despite accurate clustering. We propose \emph{Epistemic Social Learning} (ESL), a two-timescale framework that couples belief-based inference with adaptive representation learning under endogenous interaction. At a fast timescale, agents maintain Bayesian beliefs over shared behavioral prototypes. At a slower timescale, prototypes are updated from belief-weighted interaction data via the recursion $\Theta_{m+1} = \Theta_m + \gamma_m \widehat{H}_m$, aligning learned representations with the interaction distributions induced by agent behavior. We formalize ESL as a stochastic approximation process with controlled Markov noise and show that its dynamics track the differential inclusion $\dot{\Theta} \in \mathcal{G}(\Theta)$, whose limit sets correspond to self-consistent interaction regimes. Across repeated games with behavioral shifts, ESL reduces post-switch regret by $26\%$ over strong baselines while improving decision-relevant prediction. These results reveal a fundamental gap between latent identification and adaptive performance under endogenous interaction.

De novo design of nanobody complementarity-determining regions (CDRs) targeting user-specified epitopes is crucial yet remains computationally challenging. Current approaches rely on either (i) diffusion-based structural sampling that requires thousands of designs per target, (ii) gradient-based hallucination through frozen structure-prediction networks, or (iii) all-atom generative systems requiring 3D structural input and high computational cost. A fundamental limitation of these methods is their dependence on high-quality antigen structures, either experimental or computationally predicted, restricting applicability to the small fraction of therapeutic targets with available structural data. We present EpiRAG-PBind42 (Epitope-conditioned Retrieval-Augmented Generation with PBind42), a sequence-only framework for epitope-conditioned VHH/nanobody CDR generation. Our approach builds on PBind42, a Prot42-derived autoregressive binder generator instruction-tuned on DIPS-Plus protein–protein interaction pairs. EpiRAG-PBind42 adds three key innovations: (1) target-conditioned binder generation from sequence prompts, (2) retrieval-augmented latent-space decoding using hidden-state transition datastores derived from strict VHH/nanobody–antigen donor pairs, and (3) multi-agent filtering via four specialized evaluators (Structural, BLI, Interface, and Expression) that assess structural confidence, sequence naturalness, developability-adjacent liabilities, and interface energetics. Because EpiRAG-PBind42 operates on sequence input alone, it enables design on targets lacking high-quality structural models, including intrinsically disordered regions, flexible multi-domain proteins, and membrane-bound targets inaccessible to structure-dependent methods. We evaluate EpiRAG-PBind42 on five therapeutic antigen targets spanning diverse therapeutic areas: infectious disease, oncology, inflammation, and immunology.


Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning

Hansen Lillemark ⋅ Alex Rojas ⋅ Zachary Novack ⋅ Runqian (Ray) Wang ⋅ Yilun Du ⋅ Yian Ma ⋅ Taylor Berg-Kirkpatrick ⋅ Rose Yu

Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampling schedules, limiting inference procedures from adapting to the data. We introduce Equilibrium Forcing (EqF), a simplified objective for video denoising generative models without noise level conditioning. EqF turns inference into a gradient-based optimization problem with modular training- and inference-time designs that decouple learning the denoising field from sampling. By adapting the structure of the gradient flow during sampling to propose data-dependent step sizes, EqF improves video quality and consistency on challenging autoregressive video generation benchmarks. Extensive analysis elucidates exactly how removing the noise level conditioning enables EqF's data-dependent inference properties to surpass the performance of standard noise level-conditional denoising video methods. Project page: https://anonequilibriumforcing.github.io/

Large language models achieve near-ceiling performance on code generation benchmarks, yet most of the programming languages used by popular benchmarks such as SWE-bench and HumanEval (e.g. Python, JavaScript) are squarely in-distribution. They appear at scale in pre-training corpora and are heavily reinforced during post-training. To study LLM performance on unfamiliar programming languages, we introduce EsoLang-Bench, a benchmark using five esoteric programming languages (Brainfuck, Befunge-98, Whitespace, Unlambda, and Shakespeare). All five of our chosen esoteric languages are Turing Complete, so the same algorithmic problems that are solvable in Python or JavaScript are in principle solvable in each of them. Yet, they are unfamiliar to LLMs which makes them a good proxy for evaluating out-of-distribution performance. The unfamiliarity of esoteric languages comprises of: (i) the hard-by-design primitives comprising the language; (ii) substantially less representation in pre-training corpora (340× to over 60,000× fewer public GitHub repositories than Python); (iii) negligible deployment value, which makes targeted inclusion in post-training data economically irrational. We evaluate five frontier models across five prompting strategies and find a dramatic capability gap. The same 80 problems expressed in Python or JavaScript reach 100% accuracy on top frontier models, while the equivalent esoteric versions score only 0–11%. Few-shot learning and self-reflection also fail to close this gap. EsoLang-Bench therefore provides a contamination-resistant testbed for measuring how well frontier models generalise algorithmic problem-solving to programming languages outside their training distribution.


Estimating and Orthogonalizing Unknown Pre-training Gradients for Continual Fine-tuning of Large Language Models

Bing Wang ⋅ Changchun Li ⋅ Xin-Qiang Cai ⋅ Lin Y Wu ⋅ Ximing Li ⋅ Gang Niu ⋅ Masashi Sugiyama

Continual fine-tuning is essential for large language models (LLMs) to dynamically adapt to real-world environments, yet it inevitably suffers from catastrophic forgetting, particularly the performance degradation of previous tasks and LLMs' general-purpose knowledge. Although existing methods, such as orthogonal gradient projection, mitigate the forgetting across various fine-tuning tasks, they fundamentally fail to preserve pre-training LLMs' inherent general-purpose knowledge because the original data and gradients of off-the-shelf pre-training LLMs required by these methods are strictly unknown and highly diverse. To bridge this critical gap, we propose EoupCT, a novel framework designed to Estimate and Orthogonalize Unknown Pre-training gradients for Continual LLM fine-Tuning. Specifically, EoupCT estimates pre-training gradients by dynamically generating pseudo data that is most susceptible to forgetting for new tasks through a learnable soft prompt equipped with Gumbel-Softmax relaxation. Furthermore, we formulate a multi-objective optimization problem and introduce a first-order efficient Pareto optimizer that jointly optimizes LLM parameters and the soft prompt, rigorously enforcing orthogonality between new task updates and the estimated pre-training gradients. Extensive experiments across multiple LLMs demonstrate that EoupCT effectively preserves both task-specific proficiency and inherent general-purpose knowledge, successfully mitigating the catastrophic forgetting.


Estimating Continuous Treatment Effects with Recourse Data

Alessandro Marchese ⋅ Jeroen Berrevoets ⋅ Niels Martin ⋅ Sam Verboven

Recourse explanations describe what would need to change for a different outcome to occur, yet they are not typically used for treatment effect estimation. We study how such recourse data can be used for causal inference with continuous treatments. This provides counterfactual supervision beyond observed outcomes at realized treatment values, and is especially useful when routine treatment assignment leaves parts of the treatment domain unsupported, as in diverse applications such as drug dosing, healthcare operations, and recommendation systems. We formalize recourse explanations as structural boundary samples and show how they identify continuous dose-response curves under a positivity condition fundamentally distinct from standard overlap. Because recourses are observed only for negative-outcome units, the boundary distribution is left-truncated relative to the population; we correct this selection using Lynden--Bell inversion. We translate this result into an auxiliary recourse loss for neural dose-response estimators, converting recourse-derived boundaries into counterfactual supervision. Experiments on synthetic and semi-synthetic data show consistent improvements over standard baselines. More broadly, our work shows that recourse explanations can serve as structured causal evidence, expanding what can be learned from observational data.


Estimating Implicit Regularization in Deep Learning

Joseph H Rudoler ⋅ Kevin Tan ⋅ Giles Hooker ⋅ Konrad Kording

Deep learning systems are known to exhibit *implicit regularization* (alt. *implicit bias*), favoring simple solutions instead of merely minimizing the loss function. In some cases, we can analytically derive the implicit regularization -- connecting it to an equivalent penalty that augments the learning objective. However, modern deep learning systems are complex, carrying modifications to the training procedure and architecture (e.g. early stopping, minibatching, dropout) whose effects are not always directly interpretable. Although estimating the resulting implicit regularization could aid theorists in algorithm design and practitioners in interpreting their hyperparameter choices, this problem has received little direct attention. It is also tractable: regularization makes weight updates deviate from loss gradients, promising a signal for identifying implicit bias. Here we provide gradient matching methods that can be used to empirically estimate the implicit regularization. Our method works on networks with known regularization, recovering popular explicit penalties like $\ell_1$ and $\ell_2$. It also replicates known implicit effects, like the quadratic weight penalty induced by early stopping in gradient descent, demonstrating that it can be used to test theories of implicit regularization. Crucially, because our method is empirical, it can handle implicit regularization in arbitrary networks. We demonstrate this use by characterizing the effects of dropout in deep networks, showing implicit $\ell_2$ effects in this popular method. Our work shows that practitioners can use gradient matching to understand regularization in networks with implicit biases that are too complicated to derive analytically.


EVA-Cap: Optimizing Audiovisual Video Captioning via Event-Centric Alignment

Jiapeng Shi ⋅ Qiuxia Hou ⋅ Jiaqi Leng ⋅ Junke Wang ⋅ Yanhao Zhang ⋅ Haonan Lu ⋅ Ziyi Ye ⋅ Zuxuan Wu

Recent advances in Omni Language Models (OLMs) provide a unified solution for audiovisual video captioning. Nevertheless, existing approaches predominantly treat captioning as an unstructured text generation process, neglecting the inherent semantic structure of videos and leading to suboptimal fine-grained audiovisual alignment. To bridge this gap, we propose EVA-Cap, a novel framework that decomposes captions into atomic audiovisual events and transforms unstructured text supervision into event-centric alignment. We train EVA-Cap via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO) on a self-curated dataset, EVA-Data, which is built with event-centric quality control. During the GRPO phase, the model is explicitly optimized against three event-centric alignment targets: global event completeness, intra-event attribution, and inter-event synchronization. Extensive experiments demonstrate that EVA-Cap significantly outperforms existing methods in event-centric evaluation (e.g., +5.3% average gain on ChronusAV), while achieving competitive performance on general audiovisual captioning benchmarks (e.g., +3.3% on WorldSense).

ECG, with mature physiological simulators and public benchmarks, is a useful testbed for evaluating what synthetic pretraining supports. We introduce a reusable framework for ECG representation learning that compares patient-free physiological simulators, patient-derived real ECGs, and learned synthetic ECGs from real-ECG-trained generators. Holding the encoder and MAE objective fixed, we study three matched settings. Transfer-protocol analysis compares abnormal-inclusive real ECGs, normal-only real ECGs, and simulator pretraining under frozen probing and full fine-tuning on 26 abnormal-ECG tasks from PTB-XL, G12EC, and CPSC2018 using a single-lead-II pipeline. Compute-normalized PTB-XL scaling tests whether simulator pretraining matches real-data pretraining under equal processed-sample budgets. A fixed 100k-sample generator-class comparison evaluates VAE, DCGAN, SSSD-ECG diffusion variants with and without labels, and two simulators, SimECG-M and SimECG-N. The transfer analysis shows that protocol determines the supported claim: abnormal-inclusive real ECG pretraining yields more linearly separable pathology features under frozen probing, whereas full fine-tuning narrows the gap and can reverse it when the simulator corpus is sufficiently scaled. Under the fixed-100k protocol, task-wise F1 averaged over five runs ranks SimECG-N first on 22/26 tasks; no evaluated learned generator ranks first under this fixed-100k protocol; and SimECG-N obtains higher mean F1 than both diffusion variants on all 26 tasks. These results delineate where abnormal-inclusive real exposure remains beneficial and where patient-free physiological simulation can serve as an effective, protocol-bound substitute under the evaluated single-lead-II MAE pretraining and full-fine-tuning setting. We release pretraining corpora, splits, and code to support reproducible evaluation of synthetic ECG pretraining sources.


Evaluation of Visual Processing Should Be Human-Centered, Not Metric-Centered

Fanghua Yu ⋅ Jinfan Hu ⋅ Zhiyuan You ⋅ Xiang Yin ⋅ Hongyu An ⋅ Xinqi Lin ⋅ Hongyang Li ⋅ Chao Dong ⋅ Jinjin Gu

This position paper argues that the evaluation of modern visual processing systems should no longer be driven primarily by single-metric image quality assessment benchmarks, particularly in the era of generative and perception-oriented methods. Image restoration exemplifies this divergence: while objective IQA metrics enable reproducible, scalable evaluation, they have increasingly drifted apart from human perception and user preferences. We contend that this mismatch risks constraining innovation and misguiding research progress across visual processing tasks. Rather than rejecting metrics altogether, this paper calls for a rebalancing of evaluation paradigms, advocating a more human-centered, context-aware, and fine-grained approach to assessing the visual models' outcomes.


Event-Centric Perception in Weak-Signal Physical Streams with Multimodal LLMs

Chi Xu ⋅ Mengdi Jin ⋅ Jiaxing Li ⋅ William I Atlas ⋅ Thor Veen ⋅ Edith C Ngai ⋅ Jiangchuan Liu

Robotic and sensing systems operating in challenging environments often receive weak, noisy, and incomplete data streams, such as imaging sonar in underwater settings or thermal sensing under low visibility. In these scenarios, useful evidence is sparse over time, and the objective extends beyond frame-level analysis to reconstructing physically meaningful events. We study this problem through event-centric perception, which maps physical sensor streams into structured events with temporal support, spatial context, direction, count, and magnitude. Frontier multimodal LLMs, such as Qwen3.5-Plus and Gemini, have shown promising capability for event-centric inference under weak and temporally distributed evidence. However, with direct prompting, weak signals often remain below the models' effective decision boundary, leaving event-level structure underused. We address this problem with localized event reasoning, structured event records, and lightweight event-aligned adaptation. This moves weak signals from ignored observations into usable event evidence. In our experiments on a sonar benchmark, the adapted approach improves positive event recall from 0.041 to 0.914 and reduces normalized event error from 0.982 to 0.334. Thermal experiments further show that the approach extends to wildlife event reconstruction and temporally grounded occupancy reasoning. These results support event-centric perception as a principled framework for physical-stream intelligence.


Every Measurement, Every Direction, All at Once: Multimodal Flow Matching for Molecules and Spectra

Thorben Prein ⋅ Elton Pan ⋅ Rafael Gomez-Bombarelli ⋅ Santiago Miret

Generative modeling for science is data-scarce: each sample is expensive but arrives with several coupled measurements of the same system. Standard conditional models often use single-direction generation, discarding supervision from the rest. We argue that this asymmetry should not be a constraint during training: Under random target/conditioning masking, every measurement can act as a training target, hence one expensive sample yields multiple supervised denoising tasks. We present MOSAIC, a multimodal flow matching transformer over molecular 3D structure and UV, IR, and Raman spectra. On QM9S, MOSAIC achieves state-of-the-art performance with 88.8% Acc@1 for spectra-to-structure task, outperforming the strongest baseline by over 20% with 20$\times$ fewer integration steps. MOSAIC enables any-to-any generation from any combination of input to any combination of output modalities common in chemical sciences. We provide ablations to demonstrate the effectiveness of MOSAIC and show consistent performance scaling with increasing model and data size.

Open World Object Detection (OWOD) is a challenging task that requires detectors to recognize known categories while discovering unlabeled objects and incrementally incorporating them as new categories. Existing methods mainly rely on known-class features to recall unknown objects, while overlooking the semantic pull of known classes in the attribute space and its impact on the interpretability of unknown detection. In this paper, we propose EviAttr-OW, a novel evidential attribute reasoning framework for OWOD that discovers potential unknown objects with interpretable attribute evidence. Specifically, we construct a class-agnostic attribute space that decouples attribute representations from known-class bias and provides a more open attribute description basis for potential unknown objects. We then map candidate-region attribute responses into a Dirichlet evidence distribution, producing known-class predictions and evidential uncertainty to measure the reliability of known-class support. Finally, we derive Known-Class Evidence Deficiency from this evidence distribution and combine it with object evidence, enabling the model to identify unknown objects with evidence-supported objectness and insufficient known-class support. Experiments on the Real-World Object Detection (RWD) benchmark across five real-world application datasets show that EviAttr-OW consistently outperforms existing state-of-the-art (SOTA) methods, achieving +7.3 mAP on unknown classes.


Evidence-RL: Towards Evidence-intensive Visual Reasoning

Haojie Huang ⋅ Xinlei Yu ⋅ Chengming Xu ⋅ Zhangquan Chen ⋅ Cheng Yang ⋅ Qingdong He ⋅ Yu Yang ⋅ Jiangning Zhang ⋅ Xiaobin Hu

Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visual context. Existing perception-aware post-training methods encourage image use through global perturbations or attention proxies, but they do not test whether a sampled answer causally depends on the local evidence that supports it. We propose Counterfactual Evidence Disentanglement (CED), a training-time evidence audit for VLM grounding. For each response, CED neutralizes an object-centric Evidence Region and compares the resulting support drop against matched non-evidence Regions. We combine this signal with answer correctness inside GRPO, rewarding correct answers that rely on the evidence path rather than shortcut or nuisance paths. CED uses weak object-level proposals, requires no question-specific evidence annotations, and adds no inference-time overhead. Across nine public benchmarks and four backbones, CED outperforms prior RL-based post-training methods, with targeted analyses verifying its object-centric signal.


EVOCHAMBER: Test-Time Co-evolution of Multi-Agent System at Individual, Team, and Population Scales

Yaolun Zhang ⋅ Tianyi Xu ⋅ Shengyu Dai ⋅ Zhenwen Shao ⋅ Qingyun Wu ⋅ Huazheng Wang

We argue that multi-agent test-time evolution is not single-agent evolution replicated N times. A single-agent learner can only evolve its own context and memory. A multi-agent system additionally evolves who collaborates, how they collaborate, and how knowledge flows across the population. These components have no single-agent counterpart and can produce phenomena such as emergent specialization. Yet prior test-time methods either confine experiences to individual agents, forfeiting cross-agent learning, or broadcast symmetrically to all agents, erasing the specialization that makes collaboration valuable. We present EVOCHAMBER, a training-free framework that instantiates test-time evolution at three levels over a coevolving agent pool. At its core is CODREAM (Collaborative Dreaming), a post-failure protocol in which agents collaboratively reflect, distill insights, and route them asymmetrically from strong agents to weak ones on the failed niche, preserving specialization while filling knowledge gaps. Team-level operators assemble niche-conditioned teams and select collaboration structures online. Population-level lifecycle operators fork, merge, prune, and seed agents under performance pressure. On three heterogeneous task streams with Qwen3-8B, EvoChamber reaches 63.9% on competition math, 75.7% on code, and 87.1% on multi-domain reasoning, outperforming the best baseline by 30% relative on math. These gains transfer to GPT-4.1-mini. Ablations attribute the largest single drop of -10.8% to removing CODREAM, confirming asymmetric cross-agent transfer as the primary driver. Starting from 20 identically initialized agents, four to five stable niche specialists spontaneously emerge, a structural signature of multi-agent evolution that no single-agent learner can express.

Empirical risk minimization (ERM) can be computationally expensive, with standard solvers scaling poorly even in the convex setting. We propose a novel lossless compression framework for convex ERM based on color refinement, extending prior work from linear programs and convex quadratic programs to general differentiable convex optimization problems. We develop concrete algorithms for a range of models, including linear and polynomial regression, binary and multiclass logistic regression, regression with elastic-net regularization, and kernel methods such as kernel ridge regression and kernel logistic regression. Numerical experiments on representative datasets demonstrate the effectiveness of the proposed approach.

Autoregressive scientific forecasters often enforce physical or structural constraints by repairing each predicted state before feeding it back into the model. However, it remains unclear when stronger physical rule enforcement becomes reliable and when it becomes a source of distribution shift. We study this question through operator exactness, meaning whether the repair map is the identity on the target manifold and is aligned with the target geometry. We compare raw forecasting, post hoc repair, and in-loop repair across periodic incompressible Navier--Stokes, non-periodic CFDBench flows, and a hierarchical-forecasting support task. In the exact periodic regime, Fourier projection substantially improves rollout accuracy. On the NS-128 benchmark, a strong Raw-FNO has a final-step rollout MSE at horizon 100 of $(9.390 \pm 6.290)\times 10^{-5}$, and post hoc and in-loop projection reduce it to $(1.130 \pm 0.165)\times 10^{-6}$ and $(5.370 \pm 0.113)\times 10^{-7}$. However, once an exact projection is unavailable and only approximate boundary-preserving cleanup is available, the ordering changes. Across cavity, tube, dam, and cylinder flow, stronger Poisson-based cleanup can reduce divergence while worsening rollout error; target-distortion MSE predicts this harm far better than a linear-system residual. Controlled mismatch, screened cleanup, adaptive gating, and external-backbone checks show that the best approximate-regime operating point can be raw or near-identity. Hierarchical forecasting gives the same broader pattern. Exact forecast reconciliation is a stable baseline, whereas blended top-down repair, a validation-tuned interpolation toward historical-proportion top-down reconciliation, is dataset-dependent. Thus, constraint enforcement should be benchmarked by operator--data alignment before enforcement strength. Use in-loop projection when the operator is exact, and validate approximate cleanup strength using rollout metrics otherwise.


Exact power indices for plurality-voting ensembles

Ilie Sarpe ⋅ Theofanis Georgakopoulos ⋅ Aristides Gionis

Plurality-voting ensembles aggregate the predictions of multiple base classifiers and output the most voted class, with ties typically broken at random. Despite the simple aggregation rule, attributing the impact of each individual base classifier to the ensemble's predictions is far from straightforward. That is, a single base classifier's impact depends not only on its own predictions but also on how it interacts with the votes of all other models in the ensemble, for example, in resolving ties. We address the attribution problem using established tools from cooperative game theory, studying the Banzhaf and Shapley power indices for new games that capture plurality voting and random tie-breaking, in both binary and multi-class classification settings. Computing power indices is, in general, exponential in the number of players. Surprisingly, our methods, collectively named ENPOWER, enable the exact computation of Shapley and Banzhaf values in linear time for the binary-classification setting and in low-degree polynomial time for the multi-class setting. The algorithms in ENPOWER rely on new combinatorial formulations of the Shapley and Banzhaf indices, yielding closed-form expressions for the binary case and ordinary generating function expressions for the multi-class case. We validate empirically the methods in ENPOWER and demonstrate that they are highly efficient, providing attribution values not captured by existing methods. We further apply the techniques in ENPOWER to: ensemble pruning and prompt-based ensembling of vision-language models.


ExComm: Exploration-Stage Communication for Error-Resilient Agentic Test-Time Scaling

Woomin Song ⋅ Beomjun Kim ⋅ Daewon Choi ⋅ Sai Muralidhar Jayanthi ⋅ Saket Dingliwal ⋅ Jinwoo Shin ⋅ Aram Galstyan

A common failure mode in long-horizon agentic test-time scaling is error propagation, where factual errors or invalid deductions introduced at intermediate steps persist in the agent's belief state and contaminate later reasoning. Existing test-time scaling methods provide limited control over this process, as they often rely on agents to detect their own mistakes, select among flawed trajectories, or refine solutions only after errors have already shaped the reasoning path. We propose ExComm, a communication protocol for exploration-stage agentic test-time scaling. ExComm is motivated by the empirical observation that the majority of intermediate errors in parallel agentic reasoning produce detectable cross-agent factual conflicts. Leveraging the iterative structure of agentic workflows, ExComm periodically audits agent belief states to detect such conflicts, resolves them through a dedicated tool-based verification loop, and returns concise, targeted feedback to the involved agents. Corrections are incorporated through soft belief updates, which append verified feedback rather than overwriting existing beliefs. Furthermore, to prevent collapsing trajectory diversity due to communication, ExComm further introduces a trajectory diversification module that redirects redundant trajectories toward orthogonal strategies. Experiments on AIME 2024, AIME 2025, and GAIA with Gemini-2.5-Flash-Lite and Qwen3.5-4B show that ExComm consistently outperforms strong test-time scaling baselines, achieving average performance gains of 5.7\% and 5.0\% over the best-performing baselines, respectively. Further analyses demonstrate improved error recovery, favorable scaling behavior, stronger diversity than adapted communication baselines, and the best performance-cost trade-off among the evaluated methods.


Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding

Jiayu Ying ⋅ Qijian Tian ⋅ Ruijie Xu ⋅ Xinnan Zhu ⋅ Daoguo Dong ⋅ Jiachen Xu ⋅ Xin Tan

Advancing spatial intelligence in Multimodal Large Language Models (MLLMs) is bottlenecked by the scarcity of complex, scalable 3D question-answer (QA) data. While manual annotation is labor-intensive, directly utilizing LLMs to synthesize these QA pairs often fails due to their inherent deficiencies in spatial and geometric computation. We introduce $\textbf{Exemplar2VQA}$, a scalable exemplar-driven visual question answering generation framework that rapidly synthesizes large-scale spatial QA pairs in simulated environments via multi-agent coding. By equipping collaborative agents with a meticulously designed library of geometric utilities, Exemplar2VQA bypasses LLMs' spatial reasoning flaws through deterministic code execution. Crucially, the framework exhibits remarkable versatility: taking diverse static object-centric spatial query templates as exemplars, it seamlessly and autonomously scales them into massive, high-fidelity synthetic datasets. Fine-tuning Qwen2.5-VL (3B/7B) exclusively on Exemplar2VQA-generated synthetic indoor data yields significant performance improvements across various diverse benchmarks. Furthermore, its effectiveness is not limited to in-domain indoor datasets but also robustly extends to outdoor and mixed-scene benchmarks. These results establish Exemplar2VQA as a scalable and powerful paradigm for bridging the sim-to-real gap in Embodied AI. The code will be available upon acceptance.


Existence Precedes Value: Joint Modeling of Observational Existence and Evolving States in Time Series Forecasting

Yifan Hu ⋅ Hongzhou Chen ⋅ Peiyuan Liu ⋅ Yiding Liu ⋅ Zewei Dong ⋅ Jiang-Ming Yang

Real-world time series are often highly incomplete and irregular due to sensor dormancy, transmission delays, and event-driven sampling, making reliable forecasting fundamentally challenging. Existing methods have evolved from impute-then-forecast pipelines to continuous-time models such as Neural ODEs and continuous-time graph networks. While these approaches improve the modeling of historical irregularity, they still rely on an implicit oracle assumption at inference time: the timestamps of future valid observations are presumed to be known in advance. This assumption limits practical relevance, since in many real systems the more fundamental question is not only what the future value will be, but also whether a valid observation will occur at all. In this paper, we propose Timeflies, a unified framework that reformulates forecasting as a joint problem of future observability inference and value estimation. To explicitly model the interaction between observation dynamics and state evolution, Timeflies adopts an observation stream and a value stream, coupled through three dedicated modules for reliability-aware embedding, observation-guided dependency modeling, and joint prediction. We further construct Glimpse, a benchmark that combines natural missingness from public datasets with real-world industrial data. Extensive experiments show that Timeflies consistently outperforms existing methods, highlighting the importance of explicitly modeling future observability in time series forecasting with missing values. Code and dataset are available in the supplementary material.


Explanation Mechanism Influences Human Reliance on Reinforcement Learning Agents

Xuying Zhong ⋅ Daniel Beechey ⋅ Crescent Jicol ⋅ Janina A. Hoffmann ⋅ Özgür Şimşek

When reinforcement learning agents are deployed in decision-support settings, humans must repeatedly decide whether to adopt or override individual action recommendations. Explainable reinforcement learning supports these decisions, yet it remains unknown whether explanation mechanisms impact user behaviour. We conduct two controlled user studies that isolate explanation mechanism while controlling for agent policy, task, and abstraction level, using a design that independently manipulates action optimality and explanation veracity. Study 1 compares three representative action-level mechanisms across 100 participants and finds that Shapley-based attributions produce conservative reliance sensitive to explanation plausibility; saliency-based explanations indiscriminately increase adoption even for suboptimal actions; and a novel reward-decomposition mechanism, Action Advantage Attribution (AAA), achieves the highest appropriate adoption and is the only condition in which explanation comprehension positively predicts appropriate adoption. Study 2 benchmarks AAA against a no-explanation baseline across 70 participants, showing that it substantially improves detection of appropriate recommendations by 89\% but does not eliminate over-reliance on suboptimal ones. Our results imply that algorithm designers must treat explanation mechanism as a first-order design choice, as mechanisms can influence human reliance on reinforcement learning agents.

The asymptotic behaviour of Monte Carlo Exploring Starts (MCES) is a long-standing open question in reinforcement learning, even in the tabular setting. We investigated the convergence properties of tabular MCES by constructing examples in which the algorithm converges to suboptimal solutions. This paper presents new counterexamples for both initial-visit and first-visit MCES and gives a convergence-restoring modification for the initial-visit case. We show that stable suboptimal solutions may exist for initial-visit MCES with sample-average updates even when greedy actions are updated more often than non-greedy actions on average. However, by scaling learning rates inversely to update frequencies on a state-by-state basis, convergence to optimality is guaranteed. Unlike previous uniformisation methods, this modification is applicable to large-scale problems that require approximating the estimated value function. We then extend the example to show that sample-average first-visit MCES may also converge to suboptimal solutions. This largely settles a fundamental open problem and shows that exploring starts alone do not guarantee convergence to optimality. More broadly, these results highlight that convergence depends critically on the relative size and frequency of updates applied to different actions, making the choice of learning rates and the balance between exploration and exploitation central to the analysis of MCES and the implementation of scalable Monte Carlo control methods.

Offline reinforcement learning is a standard paradigm for fine-tuning Large Language Models (LLMs) on multi-turn dialogue, where on-policy rollouts and reward annotations are costly. Within this paradigm, reward-weighted supervised fine-tuning has emerged as an efficient approach, weighting each trajectory's log-likelihood by a reward-dependent coefficient. We show that existing reward-weighted SFT methods are specific members of a broader affine reward-weighted SFT class, and identify two structural limitations that hold for every member of this class: every centered member yields an unbounded-below loss, and no member admits a hyperparameter for attenuating reward noise. We introduce ExpRFT as a principled escape from this class — an exponential reweighting of the standardized advantage at a tunable temperature that restores a well-posed objective and provides a noise-control dial. ExpRFT parallels Advantage-Weighted Regression (AWR) in spirit, yet requires no per-action learned critic and adds no infrastructure cost beyond standard SFT. On six multi-turn dialogue benchmarks under reasoning and clarifying-question protocols, ExpRFT outperforms rejection-sampling, preference-based, value-based, and reward-weighted baselines. In the near-mean trajectory regime, where weight noise dominates the advantage signal, ExpRFT remains stable under injected reward noise while baselines significantly degrade.

In high-stakes risk prediction, interval-valued predictions provide a natural way to represent predictive uncertainty. However, standard evaluation tools such as the receiver operating characteristic (ROC) curve and the area under the curve (AUC) are designed for point-valued predictions and do not capture how uncertainty affects discrimination performance. To address this gap, we develop an uncertainty-aware ROC framework and corresponding interval-based AUC (iAUC) metrics for evaluating the performance of interval-valued risk predictions. The framework constructs two ROC-style curves with associated areas $\mathrm{AUC}_L$ and $\mathrm{AUC}_U$, and shows that these quantities induce a three-region decomposition of positive-negative rankings into confidently correct, ambiguous, and confidently incorrect orderings. Under valid class-conditional coverage, $\mathrm{AUC}_L$ and $\mathrm{AUC}_U$, together with the pairwise miscoverage rate, provide theoretical bounds on the optimal AUC, linking interval coverage to achievable discrimination. The framework is compatible with a range of methods for quantifying epistemic uncertainty in risk prediction, including Bayesian, bootstrap, and ensemble-based procedures, and provides a practical way to compare predictive models and interval-generation methods. Experiments on clinical benchmark data and MIMIC-IV mortality prediction show that models with similar point AUC can have different uncertainty-aware ranking profiles, and that the proposed decomposition can support risk-aware model selection and threshold-based abstention.

Memorization in large language models has been studied almost exclusively through prefix-conditioned extraction, a natural choice for autoregressive models. However, diffusion language models (DLMs) can denoise masked tokens at arbitrary positions. Thus, prefix-only probing reveals only one facet of memorization in DLMs and significantly underestimates the risk of training-data extraction. In order to realistically model extractability of training data in DLMs, we introduce infilling extraction, a data-extraction protocol parameterized by an arbitrary binary mask that subsumes prefix-only probing and accounts for the bidirectional inductive bias of DLMs. Instantiating it on LLaDA-8B and Dream-7B across five extraction modes, three training pipelines, and three corpora covering verbatim and partial leakage, we find that mask geometry governs extractability: edge-conditioned masks extract up to three times more verbatim sequences than prefix-conditioned ones, and bidirectional access opens channels inaccessible in autoregressive models. In particular, we show that a realistic adversary with access to training data where personally identifiable information has been redacted, can even achieve higher recall on extracting redacted email addresses from DLMs than from scale-matched autoregressive models. Tunable parameters for decoding measurably affect extraction performance, while a follow-up supervised finetuning stage does not eliminate the prior memorization.


Facts Don't Speak Louder than Words: The Behavioral Essence of Long-CoT Distillation

Yongcan Yu ⋅ Jian Liang ⋅ Lingxiao He ⋅ Kuangpu Guo ⋅ Yanbo Wang ⋅ Shuo Lu ⋅ Meng Wang ⋅ Ran He

While long chain-of-thought (CoT) distillation significantly improves the reasoning abilities of language models, its underlying mechanism remains poorly understood. Current pipelines generally rely on rejection sampling to gather reasoning traces across diverse prompts, assuming that factual correctness and broad problem coverage are indispensable. Surprisingly, we find that training exclusively on flawed reasoning trajectories or a restricted prompt set yields nearly comparable performance. Driven by this counterintuitive finding, we propose the Behavioral Hypothesis: the essence of Long-CoT distillation is behavioral alignment with structured reasoning patterns, rather than factual knowledge memorization. We then validate this by demonstrating substantial reasoning gains even when models are fine-tuned exclusively on previously mastered problems—effectively isolating behavioral alignment from novel knowledge injection. Having established that reasoning patterns are the true bottleneck, we investigate how to learn them across varying levels of complexity optimally. We reveal a critical interplay: while the recent probability-based loss excels at learning simple patterns, the standard cross-entropy loss is essential for scaling to complex CoTs. We believe these insights demystify CoT distillation and provide a principled foundation for the future development of reasoning models.


Falcon-X: A Time Series Foundation Model for Heterogeneous Multivariate Modeling

Yiding Liu ⋅ Yifan Hu ⋅ Hongjie Xia ⋅ Peiyuan Liu ⋅ Hongzhou Chen ⋅ Xilin Dai ⋅ Zewei Dong ⋅ Jiang-Ming Yang

Time series foundation models (TSFMs) are transforming the forecasting paradigm through large-scale cross-domain pretraining. However, most existing TSFMs remain univariate, and recent efforts to enable cross-variate modeling still operate directly within the raw variate space. This design introduces fundamental limitations in semantic alignment and relational expressivity. Specifically, raw-space group mixing lacks a dedicated mechanism to align heterogeneous physical quantities, while standard non-negative attention fails to capture the complex synergistic and antagonistic interactions ubiquitous in real-world systems. To address these challenges, we propose Falcon-X, decouples variates from the raw space and maps them into a unified latent prototype space. Falcon-X employs a Unified Prototype Diff-Attention mechanism that explicitly evaluates both positive and negative semantic affinities to explicitly align heterogeneous variates. Cross-variate interactions are then efficiently performed within this shared space via Latent Entity Attention, naturally facilitating zero-shot structural transfer. Finally, a Variate Reassembly Router robustly reconstructs variate-specific trajectories via a request-and-dispatch mechanism. Extensive evaluations on the GIFT-Eval and fev-bench benchmarks demonstrate that Falcon-X achieves state-of-the-art forecasting performance, offering a principled and scalable paradigm for complex multivariate environments. Code is available in the supplementary material.


FASTER: Rethinking Real-Time Flow VLAs

Yuxiang Lu ⋅ Zhe Liu ⋅ Xianzhe Fan ⋅ Zhenya YANG ⋅ Jinghua Hou ⋅ Junyi Li ⋅ kaixin Ding ⋅ Hengshuang Zhao

Real-time execution is crucial for deploying Vision-Language-Action (VLA) models in the physical world. Existing asynchronous inference methods primarily optimize trajectory smoothness, but neglect the critical latency in reacting to environmental changes. By rethinking the notion of reaction in action chunking policies, this paper presents a systematic analysis of the factors governing reaction time. We show that reaction time follows a uniform distribution determined jointly by the Time to First Action (TTFA) and the execution horizon. Moreover, we reveal that the standard practice of applying a constant schedule in flow-based VLAs can be inefficient and forces the system to complete all sampling steps before any movement can start, forming the bottleneck in reaction latency. To overcome this issue, we propose Fast Action Sampling for ImmediaTE Reaction (FASTER). By introducing a Horizon-Aware Schedule, FASTER adaptively prioritizes near-term actions during flow sampling, compressing the denoising of the immediate reaction by tenfold (e.g., in $\pi_{0.5}$ and X-VLA) into a single step, while preserving the quality of long-horizon trajectory. Coupled with a streaming client-server pipeline, FASTER substantially reduces the effective reaction latency on real robots, especially when deployed on consumer-grade GPUs. Real-world experiments, including a highly dynamic table tennis task, prove that FASTER unlocks substantially improved real-time responsiveness for generalist policies, enabling rapid generation of accurate and smooth trajectories.


Fast Organic Crystal Structure Prediction with Unit Cell Flow Matching

Alston Lo ⋅ Luka Mucko ⋅ Austin Cheng ⋅ Andy Cai ⋅ Alastair J Price ⋅ Wojciech Matusik ⋅ Alan Aspuru-Guzik

Organic crystal structure prediction (CSP) is a requirement for computational modelling of organic solids. Traditional CSP is accurate but relies on exhaustive search, costing several CPU-years per molecule. Generative models such as OXtal dramatically reduce this cost by sampling stable organic crystal structures directly. However, OXtal forgoes explicit lattice parametrization in favour of modelling large crops of the bulk material with expensive triangle layers, which can incur a computational cost of minutes per molecule. In this paper, we reduce this to seconds with Clari, a large-scale flow matching model that generates redundancy-free unit cells and replaces triangle layers with pure pair-bias attention. Clari requires only atom types and bonds as input and does not need an RDKit-sanitizable input molecule, which expands its applicability to challenging chemistries such as fullerenes, metal complexes, and atom clusters. We further ablate key design choices such as auxiliary losses, timestep distributions, noise priors, and self-conditioning. Because Clari also models explicit hydrogens, it supports inference-time scaling via direct energy ranking, without any decoration or relaxation step. On OXtal's aggregated test set, we generate 1000 crystals and select the best 30 ranked by energy, surpassing OXtal's solve rate while obtaining a speedup of $5$-$8\times$. We also introduce a new test split of diverse and complex molecules for future benchmarking. Our contributions enable CSP within seconds, making large-scale virtual screening of organic solids practical.


FAVLA: A Force-Adaptive Multi-Rate VLA model for Contact-Rich Robotic Manipulation

Yao Li ⋅ Peiyuan Tang ⋅ Wuyang Zhang ⋅ Haojie Ren ⋅ Chengyang Zhu ⋅ Yifan Duan ⋅ WeiKai Shi ⋅ Xiaodong Zhang ⋅ Yuming Dong ⋅ Zijiang J Yang ⋅ Jianmin Ji ⋅ Yanyong Zhang

Vision-Language-Action (VLA) models have shown strong potential for general robotic manipulation, but contact-rich tasks still require timely action refinement using force/torque feedback. Existing force-aware VLA models usually align all modalities to a single low operating frequency, which discards high-frequency contact cues that are important for reactive action correction. Additionally, they use static fusion, which cannot adaptively balance visual context and force feedback across manipulation processes. To mitigate these issues, we propose FAVLA, a force-adaptive multi-rate VLA model that explicitly separates low-frequency visual-language reasoning from high-frequency force-conditioned action refinement. Based on the $\pi_0$-style VLM-action expert architecture, FAVLA uses a low-rate VLM to encode visual-language-force context and predict near-future force statistics, while a high-rate action expert refines action chunks using the latest force observations. To decide \emph{when} to update actions, we introduce a Force-Adaptive Multi-rate Inference (FAMI) mechanism, which schedules the action expert's inference rate from predicted future force variance, keeping low-rate reasoning in stable phases and increasing update frequency near contact transitions. To decide \emph{how} to use force, we design a Force-Guided Dynamic Fusion (FGDF) module, which injects high-frequency force features into the action expert and dynamically balances visual-semantic and force cues across manipulation stages. Extensive real-robot experiments on high-precision and contact-rich tasks show that FAVLA outperforms force-aware VLA baselines, achieving an average success rate of 88.8\% and lower peak contact forces during manipulation.


Feature Recovery for Object Understanding Under Physical Transformation

Aditi Tiwari ⋅ Sofia Stoica ⋅ Savya Khosla ⋅ David Forsyth ⋅ Heng Ji

Objects in post-fire environments often undergo irreversible physical transformations that change their geometry, material state, and visual appearance. Detecting and identifying these remnants is important for locating hazards, reconstructing pre-incident contents, and inventorying losses. Unlike standard image corruptions, fire damage changes the physical structure of the object itself. To study this setting, we introduce TRACE, a transformation-aware benchmark for post-fire object understanding. TRACE contains 21.4K real-image-grounded synthetic scenes and paired object-level pristine-to-degraded progressions spanning 499 object identities across 189 categories. We define five tasks that evaluate both localization and pre-degradation understanding: degraded-object detection, pristine-state recovery and retrieval, original material recovery, pristine description generation, and functional reasoning. Existing models degrade sharply as fire damage becomes more severe. From the least to the most severe degradation level, RF-DETR mAP decreases by 71% relative, while InternVL3.5 retrieval Recall@1 drops from 93.85 to 28.11. To address this, we propose the Feature Restoration Module, or FRM, a lightweight plug-and-play module that maps degraded encoder features toward pristine-aligned representations while keeping the host model frozen. FRM is trained only with paired feature supervision and improves scene-level detection, CLIP/SigLIP2 feature recovery, and all four object-level VLM tasks. Its gains are larger under more severe degradation. Across VLM hosts and severity levels, FRM improves retrieval by 12.5% on average, material recovery by 20.1%, description generation by 13.2%, and functional reasoning by 12.4%.


Federated Concept-Based Models: Interpretable models with distributed supervision

Dario Fenoglio ⋅ Arianna Casanova Flores ⋅ Francesco De Santis ⋅ Gabriele Dominici ⋅ Johannes Schneider ⋅ Pietro Barbiero ⋅ Giovanni De Felice ⋅ Marc Langheinrich ⋅ Martin Gjoreski

Concept-based Models (CMs) enhance interpretability in deep learning by grounding predictions in human-understandable concepts. However, concept annotations are costly and rarely available at scale within a single data source. Federated Learning (FL) could alleviate this limitation by enabling cross-institutional training over concept annotations distributed across multiple data owners. Yet, FL lacks interpretable modeling paradigms. Integrating CMs with FL is non-trivial: although FL supports heterogeneous and non-stationary client participation, it typically assumes a fixed shared architecture, whereas CMs may require architectural adaptation as the available concept set evolves. We propose Federated Concept-based Models (F-CMs), a new methodology for deploying CMs in evolving FL settings. F-CMs aggregate concept-level information across institutions and efficiently adapt the model architecture to changes in concept supervision while preserving privacy. Empirically, F-CMs maintain accuracy and intervention effectiveness comparable to training settings with full concept supervision, while outperforming on average non-adaptive federated baselines. Notably, F-CMs enable interpretable inference on concepts unavailable to a given institution, a key novelty over existing approaches.


FedSEM: Mitigating Cross-Client Evidence Drift in Federated Multiple Instance Learning

Zhiqiang Kou ⋅ Beidi Wang ⋅ Yuling Shi ⋅ Shengkun Zhu ⋅ Wei Tang ⋅ Xiaodong Gu ⋅ Jiankai Zuo ⋅ Di Jiang ⋅ Yu Zhang

Federated Multiple Instance Learning (MIFL) has emerged as a promising paradigm for privacy-preserving weakly supervised learning, particularly in medical image analysis where instance-level annotations are expensive or unavailable. Existing federated MIL methods mainly rely on parameter aggregation to learn a global model across clients. However, in MIL, the key challenge is not only client-level data heterogeneity, but also the ambiguity of instance-level evidence. Since supervision is only provided at the bag level, each client must independently infer discriminative instances through local attention mechanisms. Under heterogeneous data distributions, different clients may focus on inconsistent or even misleading instances, resulting in cross-client critical evidence drift. To address this problem, we propose FedSEM, a Federated Shared Evidence Memory framework that introduces an evidence-level communication pathway in addition to standard model aggregation. Specifically, each client identifies high-confidence evidence instances using local attention scores and compresses them into compact prototypes with importance scores. The server then refines these prototypes through redundancy removal, diversity selection, and score normalization to construct a shared evidence memory. Importantly, FedSEM does not require transmitting raw patches or dense instance embeddings; it only exchanges a small number of evidence prototypes and periodically broadcasts the refined memory, thereby introducing limited additional communication overhead compared with standard federated training. The shared memory serves as cross-client evidence anchors to guide local attention learning and instance-level representation optimization. By explicitly aligning critical evidence patterns across clients, FedSEM mitigates evidence drift while maintaining communication efficiency and privacy preservation. Extensive experiments under heterogeneous federated settings demonstrate that FedSEM consistently improves MIL performance and generalization over existing federated learning baselines.


FedVSSAM: Mitigating Flatness Incompatibility in Sharpness-Aware Federated Learning

Bingnan Xiao ⋅ Yuan Gao ⋅ Bingcong Li ⋅ Wei Ni ⋅ Xin Wang ⋅ Tony Quek

Sharpness-aware minimization (SAM) is an effective method for improving the generalization of federated learning (FL) by steering local training toward flat minima. Under data heterogeneity, however, device-side SAM searches for locally flat basins that are incompatible with the flat region preferred by the global objective. We identify this structural failure mode as flatness incompatibility, which explains why improving local flatness alone may provide limited training and generalization improvement for the global model. We reveal that flatness incompatibility arises from data heterogeneity and the friendly adversary phenomenon, and is further amplified by local updates and partial device participation. To mitigate this issue, we propose Federated Learning with variance-suppressed sharpness-aware minimization (FedVSSAM), which constructs a variance-suppressed adjusted direction and uses it consistently in local flatness search, local descent, and global update. FedVSSAM anchors both perturbation and update directions to a more stable global direction, instead of correcting only an isolated local perturbation. We establish non-convex convergence guarantees of FedVSSAM and prove that the mean-square deviation between the adjusted direction and the global gradient is effectively controlled. Experiments demonstrate that FedVSSAM mitigates flatness incompatibility and outperforms the baselines across diverse FL settings.


Feed-Forward 3D Gaussian Splatting for High-Fidelity Animatable Hand Avatar Reconstruction from a Single Image

Chanho Kim ⋅ Kyeonghwan Gwak ⋅ Muhammad Salman Ali ⋅ Muhammad Shaheryar ⋅ Incheol Park ⋅ Jeongwan On ⋅ Seungryul Baek

Recent hand avatar reconstruction methods achieve high-quality results through multi-view observations or per-subject optimization, limiting scalability and practical applicability. We present \emph{FF3DGS-Hand}, a feed-forward framework that reconstructs animatable hand avatars from a single monocular image using 3D Gaussian Splatting. Our method establishes a stable canonical geometry for thin and articulated hand structures, then reconstructs appearance by leveraging source image evidence instead of relying solely on global latent features. We further introduce a rendering-aware, source-conditioned appearance framework that selectively updates Gaussians according to their rendering contribution, recovering input-specific details while suppressing artifacts in unobserved regions. The resulting representation is animatable via standard hand pose parameters and supports efficient rendering. Experiments demonstrate strong visual quality and quantitative performance in the single-image setting.


FerQ: Fermat Quotient Reformulation of High-Order Binary Optimization

Phuong Nam Nguyen ⋅ Seng Loke ⋅ Anil Prabhakar ⋅ Son N Tran

High-order binary optimization (HOBO) problems, defined as the minimization of functions involving interactions among three or more binary variables, are prevalent in machine learning and quantum computing. The minimization of these high-order energy functions remains computationally intractable without quadratization or reformulation. This work introduces FerQ, a universal and auxiliary-free reformulation of high-order energy functions based on Fermat quotient expansion. For any degree-d monomial over binary variables, FerQ provides an exact, closed-form representation as a rational linear combination of Fermat quotients evaluated on a scalar aggregate, with coefficients determined by a single matrix inversion. FerQ is evaluated against twelve existing energy function transformation methods on k-SAT and Max-k-SAT benchmarks from four databases, as well as on random p-spin glass and k-local Hamiltonian instances. FerQ achieves the highest satisfaction and weighted satisfaction rates across all benchmarks, while incurring the lowest CPU runtime. To enable deployment on quantum annealers, we derive FerQ-Bc, a qubit-efficient embedding scheme that maps the reformulated energies onto quadratic unconstrained binary optimization (QUBO) form, achieving good ancilla efficiency in high-degree regimes.


FETTUCCINE: Fast and efficient brain-to-text decoding on mobile devices

Jonathan McCart ⋅ Pranav Deevi ⋅ Mehdi Azabou ⋅ Nanda H Krishna ⋅ Caleb A McKinney ⋅ Nicholas S Card ⋅ Sergey D Stavisky ⋅ Chethan Pandarinath

Recent advances in speech neuroprostheses have enabled new communication avenues for people who have lost the ability to speak due to neurological illness or injury. These systems can decode neural activity during attempted speech into text by relying on language priors to achieve high decoding accuracy. This performance currently comes at the cost of using a compute-intensive language modeling stack, presenting a major bottleneck for real-world use. Further, current language models incur highly variable and disruptive inference latencies, limiting communication throughput and preventing users from reliably participating in natural conversation. We introduce FETTUCCINE, an end-to-end brain-to-text framework that leverages the speed and portability of automatic speech recognition (ASR) models to address the aforementioned practical challenges, while achieving usable (<10%) word error rates (WER), efficient adaptation to neural non-stationarities, and reliable performance on mobile devices. When evaluated using publicly available data from an intracortical speech neuroprosthesis user, FETTUCCINE outperforms previous end-to-end methods, achieving as low as 4.39% WER. Most notably, our models successfully run on a commercially-available mobile device with throughput reaching up to 370x real-time, providing the first demonstration of brain-to-text decoding on a mobile device. We also show that our models can be successfully finetuned to extend usable performance to future days. Beyond performance, our approach enables localization of the neural features most relevant for decoding via gradient-based salience maps, which we show align with well-established physiological priors across different ASR model families. Taken together, our results show that FETTUCCINE overcomes core infrastructural and computational barriers, yielding a new class of portable brain-to-text communication.


Few-Shot Truly Benign DPO Attack for Jailbreaking LLMs

Yoon Sangyeon ⋅ Wonje Jeung ⋅ Yoonjun Cho ⋅ Dongjae Jeon ⋅ Albert No

Fine-tuning APIs make frontier LLMs easy to customize, but they can also weaken safety alignment during fine-tuning. While prior work shows that benign supervised fine-tuning (SFT) can reduce refusal behavior, deployed fine-tuning pipelines increasingly support preference-based objectives, whose safety risks remain less understood. We show that Direct Preference Optimization (DPO) introduces a stronger and harder-to-audit failure mode. We propose a truly benign DPO attack using only 10 harmless preference pairs, the minimum data scale accepted by OpenAI’s fine-tuning service. Each pair contains a benign prompt, a normal helpful answer as the preferred response, and a refusal as the dispreferred response. Unlike prior benign fine-tuning attacks, our data exhibits no suspicious behavior: it is practically indistinguishable from the fine-tuning request of a legitimate user seeking to reduce over-refusal, making harmful intent almost impossible to infer from the request alone. Nevertheless, because DPO directly optimizes the model to prefer helpful answers over refusals, this seemingly benign objective broadly suppresses refusal behavior and transfers to harmful prompts outside the fine-tuning data. Across OpenAI models supporting DPO fine-tuning, our attack achieves attack success rates of 59.13% on GPT-4o, 70.20% on GPT-4.1, 54.80% on GPT-4.1-mini, and 81.73% on GPT-4.1-nano, at costs of only \\$1.7, \\$1.7, \\$0.3, and \\$0.1. Moreover, on open-weight models that do not impose minimum data requirements, we find that this effect can emerge from even a single benign preference pair.

Joint prediction sets for multivariate time series should control a single event while adapting to cross-coordinate dependence. We study filtered conformal ellipsoids: a frozen state-space filter emits a one-step predictive mean and covariance, and split-conformal calibration is applied to the resulting Mahalanobis scores. The filter is used to choose the ellipsoid shape; conformal calibration chooses the scalar radius, so the construction benefits from a learned predictive covariance without relying on Gaussian tail probabilities for coverage. The main difficulty is that filtered scores are dependent and learned recurrent filters need not contract in their raw hidden state; we therefore analyse contraction in an observable predictive-law quotient that identifies hidden states producing the same future sequence of emitted Gaussian laws. Under a stable Bayes Gaussian-projection filter, covariance bounds, and a finite-horizon observability/Fisher condition, small excess Gaussian negative log-likelihood implies contraction of the learned emitted laws. Combined with a threshold-autocovariance envelope this yields a Chebyshev-type approximate coverage bound for filtered split-conformal prediction under dependence; a sharper Bernstein-type bound requires an additional geometric-mixing concentration assumption. Under Gaussian oracle realisability we also obtain a near-oracle log-volume comparison within the class of conditionally valid Gaussian ellipsoid rules. We instantiate the framework with a GCN-GRU filter with diagonal-plus-low-rank covariance. On moderate-size graph-native traffic benchmarks (METR-LA-20 and PEMS-BAY-50), the learned filter gives sharper at-target ellipsoids than static-covariance and non-filter baselines; at full-graph scale and on non-graph-native datasets, factor and copula baselines can be stronger.


FLAG: Flow Policy MaxEnt-RL by Latent Augmented Guidance

Sungha Kim ⋅ Gawon Lee ⋅ Jusuk Lee ⋅ Jonghae Park ⋅ H. Jin Kim ⋅ Daesol Cho

Maximum entropy reinforcement learning (MaxEnt-RL) enables robust exploration, yet practical implementations often restrict policies to simple Gaussians. While recent approaches incorporate expressive generative policies via importance-weighted supervised learning, they are prone to importance weight collapse, which limits their scalability in high-dimensional action spaces. Our key insight is to mitigate this limitation by localizing the sampling region, avoiding the weight degeneracy induced by importance sampling over the entire action space. To instantiate this insight, we introduce FLAG (Flow policy with Latent-Augmented Guidance). FLAG augments the state space with a flow latent variable and optimizes a provably consistent proxy MaxEnt-RL objective. We empirically demonstrate that FLAG enables expressive policy optimization with limited importance samples and scales to high-dimensional control tasks. Furthermore, FLAG achieves state-of-the-art performance across challenging benchmarks.

Nighttime lens flare removal is difficult to study with supervised learning because real flare-corrupted and flare-free image pairs are hard to acquire with accurate alignment. Existing benchmarks therefore rely mainly on synthetic compositing or rendering, but these pipelines cannot fully reproduce the coupled effects of lens contamination, internal reflection, sensor saturation, automatic exposure, and diverse urban light sources. We introduce FlareReal, a real-captured benchmark for nighttime flare removal, containing 4,037 pixel-aligned flare-corrupted/flare-free pairs and 500 real pure flare images. FlareReal is collected through a controlled contamination-cleaning protocol: each scene is first photographed with physically induced lens contamination and then immediately re-captured after optical cleaning, followed by robust global alignment and manual curation. The dataset covers 241 scenes, multiple smartphone lens systems, point/linear/area light sources, single- and multi-source layouts, and challenging exposure conditions. Experiments across representative restoration architectures show that models trained with FlareReal consistently outperform models trained on Flare7K, Flare7K++, and FlareX on Flare7K, FlareX, and FlareReal test sets, and further improve generalization under cross-device and off-screen settings. These results suggest that real paired supervision captures transferable flare formation cues that are difficult to obtain from synthetic data alone.

Diffusion Large Language Models (dLLMs) have recently emerged as a promising alternative to autoregressive LLMs by enabling non-autoregressive text generation. However, their practical deployment remains limited by inefficient inference, largely due to the absence of effective Key-Value (KV) caching and scalable parallel decoding mechanisms. Existing acceleration methods typically study KV caching and parallel decoding in isolation, overlooking the I/O bottlenecks that arise when cache reuse and parallel token verification are jointly applied. In this work, we introduce ${\bf Flash-dLLM}$, a training-free inference acceleration framework for fast and memory-efficient dLLMs. Flash-dLLM first identifies GPU memory I/O as a dominant bottleneck in KV-cache-enabled dLLM inference and addresses it with an I/O-aware fused KV-cache kernel that reduces redundant memory movement. Building on this optimized cache mechanism, Flash-dLLM further proposes an efficient KV-cache-driven draft-and-verify decoding strategy, where the dLLM itself serves as both drafter and verifier without requiring an auxiliary model. This unified design enables faster decoding while preserving generation quality and improving scalability to longer sequences and larger models. Extensive experiments on mathematical reasoning and code generation benchmarks demonstrate that Flash-dLLM consistently outperforms all existing dLLM acceleration methods in both inference speed and memory efficiency.


FlashMol: High-Quality Molecule Generation in as Few as Four Steps

Xinyuan Wei ⋅ Zian Li ⋅ Shaoheng Yan ⋅ Cai Zhou ⋅ Muhan Zhang

Generating chemically valid 3D molecular conformations is critical for computational drug discovery. Classical diffusion-based models like GeoLDM perform well but require hundreds of steps, making large-scale in silico screening impractical. Recent efforts on few-step molecular generation have accelerated this process to 12-50 steps, but they often largely sacrifice sample stability. In this work, we present FlashMol, an ultra-fast molecule generative model producing high-quality molecular conformations in as few as 4 steps. To achieve this, we adapt distribution matching distillation (DMD) — a reverse KL-divergence minimization objective — to the molecular domain for effective distillation. Considering the local minimization behavior of DMD, we respace the molecule generation timesteps, providing the generator with much better initialization and enables effective distillation. Additionally, to mitigate the mode-seeking behavior of DMD and improve diversity, we further regularize it with a Jensen-Shannon divergence term, which incorporates the mean-seeking behavior of the forward KL divergence. Extensive experiments on QM9 and GEOM-DRUG datasets demonstrate that FlashMol matches and even surpasses the original 1000-step teacher, achieving up to 250× acceleration in sampling speed while maintaining high molecular quality.


FlexCover: Flexible Cover Song Generation via Symbolic Lead Sheet Control

Lynn Yi ⋅ Zihan Xiong ⋅ Junjie Cao ⋅ Xiaolong Weng ⋅ Yingzhe Ma ⋅ Gus Xia ⋅ Jinliang Liu ⋅ Yuxin Xie ⋅ Xianwei Zhuang ⋅ Yuguo Yin ⋅ Qiao Jin ⋅ Gongxi Z Zhu ⋅ Jiayu Wang ⋅ Junjie Liang ⋅ Qin Yue ⋅ Ziyu Wang ⋅ Dading Chong ⋅ Meng Cao ⋅ Jiahuan Zhou ⋅ Dongchao Yang

A cover song re-renders an existing piece by preserving its tonal content—melody and chord progression—while reshaping other musical attributes. This suggests a natural formulation for generative cover synthesis with pretrained music language models: explicitly control tonal content, while specifying lyrics and style through text. Existing approaches condition on frame-level features, tightly coupling generation to the reference’s timing and structure, and limiting flexibility such as tempo variation, structural rearrangement, or partial conditioning. To this end, we present FlexCover, a cover generation model that conditions a pretrained text-to-song foundation model on a symbolic lead sheet. Our design enables alignment-free control: generated outputs preserve the tonal signature of the source without frame-level alignment, allowing flexible timing, structure, and segment-level generation. We further introduce a training curriculum with partially mismatched audio–symbolic pairs, improving diversity and robustness. We evaluate FlexCover with objective metrics, standard subjective ratings, and in-depth expert interviews that provide fine-grained diagnostic insights. Experiments show that FlexCover achieves state-of-the-art performance among open-source systems and is competitive with leading commercial models such as Suno-v5.5.

Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scalability to long sequences. Kernel-based linear attention reduces this complexity but typically relies on fixed or weakly learnable kernels, restricting expressiveness and performance. In this work, we propose Flexformer, a flexible linear Transformer that learns attention kernels in a fully data-driven manner. Flexformer builds on random Fourier feature-based linear attention and treats spectral frequencies as trainable parameters, enabling the model to learn a broad family of attention kernels. We develop both stationary and nonstationary variants, with the latter offering strictly greater expressiveness. Extensive experiments on language modeling and sequence classification demonstrate that Flexformer consistently outperforms baselines. Moreover, Flexformer can be effectively distilled from pretrained Transformers to recover softmax attention and exhibits strong kernel transferability across domains, achieving both high efficiency and competitive performance on long-sequence tasks.


FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models

Fan Mo ⋅ Han Yuxuan ⋅ Geng Zhang ⋅ Wangbo Zhao ⋅ Yang You

Mixture-of-Experts (MoE) language models scale model ability with sparsely activated experts, making this architecture a standard recipe for modern large models. However, sparse activation does not remove the deployment burden of storing and serving all experts, and the available deployment budget can vary substantially across devices, users, and workloads. Existing MoE compression methods are still largely fixed-budget, typically optimizing one compressed endpoint at each chosen target budget. We study a different setting: converting a large pretrained MoE LLM into a nested family of deployable subnetworks across budgets. Our method first ranks expert FFN channels by their importance, then lets each expert learn a discrete action to prune its channels. By gradually increasing cost pressure, a single action-training run exports a series of action masks from high to low budgets, each of which identifies a reliable smaller subnetwork nested in the ranked base model. Moreover, we use a single recovery fine-tune at a mid pruning budget 40% to recover degraded model quality and transfer the recovered model to other unseen budgets. Overall, our framework surpasses recent MoE compression baselines. Specifically, on Qwen2-57B-A14B, our method retains ~99.8% of base performance while pruning 50% of routed expert parameters even without fine-tuning. For deployment, our pruned subnetworks deliver real memory reduction and throughput gains, and further support realtime online budget switching with kernel-level co-design.

Many existing multivariate time series anomaly detection methods formulate anomaly detection as a reconstruction problem, where detections are derived from reconstruction errors in the time domain. While effective for large deviations, reconstruction-based scoring struggles with subtle anomalies where only a small subset of features deviates slightly. In such cases, most features are well reconstructed, causing the overall reconstruction error to remain low and indistinguishable from normal data. To address these challenges, we propose Flow+Diff, a novel framework that detects anomalies via discrepancies in a latent space, where two complementary representations, induced by distinct modeling paradigms, diverge under anomalous inputs. One representation is obtained via a normalizing flow, which provides an invertible mapping that preserves temporal dependencies and captures inter-feature dependencies. The other is produced by a diffusion model that learns the distribution of normal samples in the latent space via a stochastic denoising process. When given anomalous inputs, the normalizing flow maps the inputs to latent representations that lie in low-probability regions of the learned latent distribution, while the diffusion model generates representations aligned with normal data distributions, resulting in a pronounced discrepancy between the two. We compute the anomaly score based on this latent-space discrepancy, thereby shifting anomaly detection from the time domain to the latent space. Finally, the invertibility of the normalizing flow enables reconstruction in the time domain, facilitating interpretation of anomalous features. Extensive experiments across six benchmark datasets show that Flow+Diff achieves state-of-the-art performance on four datasets against competitive baselines.

Training-free guidance enables pre-trained diffusion and flow models to optimize application-specific objectives using feedback from external black-box reward functions. However, existing methods are feedback-inefficient because reward feedback is used only transiently to inform a localized gradient approximation or a discrete search decision, and is subsequently discarded. To address this limitation, we propose Flow-Direct, a framework that guides the generation process via a persistent guidance field. Theoretically, this guidance field is analytically derived from the log-density ratio between the base and reward-weighted target distributions; it transports the pre-trained distribution to the target distribution. In practice, the field is implemented as a non-parametric estimator constructed from all accumulated reward-evaluated samples. As more samples are collected during optimization, this empirical guidance field becomes increasingly accurate. This persistent formulation yields two major advantages. First, Flow-Direct is highly feedback-efficient: because every evaluated sample is used to refine the global guidance field, no reward information is wasted. Second, the framework is naturally reusable: once optimization is complete, the collected dataset defines a reusable guidance field for generating novel target samples without additional reward evaluations, and distinct guidance fields can be combined to generate samples that simultaneously satisfy multiple objectives.


Flowette: Flow Matching with Graphette Priors for Graph Generation

Asiri Wijesinghe ⋅ Sevvandi Kandanaarachchi ⋅ Daniel M Steinberg ⋅ Cheng Soon Ong

We study generative modeling of graphs with recurring subgraph motifs. We propose Flowette, a continuous flow matching framework that employs a graph neural network-based transformer to learn a velocity field over graph representations with node and edge attributes. Our model promotes topology-aware alignment through optimal transport-based coupling and encourages global structural coherence through regularisation. To incorporate domain-driven structural priors, we introduce graphettes, a new probabilistic family of graph structure models that generalize graphons via controlled structural edits for motifs such as rings, stars, and trees. We theoretically analyze the coupling, invariance, and structural properties of the framework, evaluate it on synthetic and molecular benchmarks, and isolate the contributions of the structural prior, the optimal-transport coupling, and the regularisation terms through controlled ablations. Flowette achieves competitive performance overall, attaining state-of-the-art results on several metrics across multiple benchmarks, highlighting the effectiveness of combining structural priors with flow-based training for modeling complex graph distributions.


FlowSteer: Prompt-Only Workflow Steering Exposes Planning-Time Vulnerabilities in Multi-Agent LLM Systems

Fanxiao Li ⋅ Jiaying Wu ⋅ Tingchao Fu ⋅ Natasha Jaques ⋅ Wei Zhou ⋅ Min-Yen Kan

Multi-agent systems (MAS) powered by large language models (LLMs) increasingly adopt planner--executor architectures, where planners convert prompts into subtasks, roles, dependencies, and routing paths. This flexibility enables adaptive coordination, but exposes an attack surface in workflow formation: prompts can shape agent organization without modifying MAS infrastructure. We study this risk through social influence probing workflows to identify high-impact subtasks and malicious-signal propagation. The analysis reveals two vulnerabilities: workflow position can amplify or suppress a malicious signal, and sycophantic framing makes downstream agents more likely to relay it. We translate these findings into FlowSteer, a prompt-only workflow steering attack that converts vulnerability priors into one crafted prompt. FlowSteer aligns a malicious signal with influential task components and guides replanning toward dependencies that preserve propagation. Experiments show that FlowSteer increases malicious success by up to 55% over naive prompting, transfers across MAS setups, and remains effective with black-box topology inference. As FlowSteer biases the planning signals that generate the workflow, MAS defenses that inspect only the generated workflow provide limited protection. As such, we introduce FlowGuard, an input-side defense that reduces malicious success by up to 34% while preserving prompt utility. Our results position workflow formation as a new safety frontier for multi-agent LLM systems, opening a planning-time security perspective on how agent coordination itself can be attacked and defended.


FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics

Qiran Zou ⋅ Hou Hei Lam ⋅ Wenhao Zhao ⋅ Tingting Chen ⋅ Yiming Tang ⋅ Samson Yu ⋅ Yingtao Zhu ⋅ srinivas anumasa ⋅ Zufeng Zhang ⋅ Tianyi Zhang ⋅ Chang Liu ⋅ zhengyao Jiang ⋅ Anirudh Goyal ⋅ Dianbo Liu

AI research agents accelerate ML research by automating hypothesis generation, experimentation, and empirical refinement. Existing agent strategies range from greedy hill-climbing to tree search and evolutionary optimization, yet which strategy choices drive performance remains unclear. Answering this question requires a benchmark that separates agent strategy (e.g., search topology) from execution infrastructure (e.g., code editor), so that performance differences are attributable to strategy rather than infrastructure, and that provides process-level metrics beyond final scores to analyze exploration behaviors. Existing benchmarks offer limited support. We propose FML-Bench, a benchmark of 18 fundamental ML research tasks across 10 domains that separates agent strategy from execution infrastructure and defines 12 process-level behavioral metrics. Evaluating six representative agents, we find that: (1) strategy complexity alone does not guarantee strong performance: a simple greedy hill-climber nearly matches the best-performing tree-search agent, both well above the remaining agents; (2) our analysis suggests this pattern relates to improvement opportunity structure: greedy search tends to be more effective when opportunities are dense, while tree-search and evolutionary strategies tend to be more effective when opportunities are sparse; an adaptive agent built on this insight switches to broader exploration upon detecting improvement stagnation and outperforms the other six agents, lending initial support to this observation; and (3) process-level analysis reveals that early convergence and directionally focused exploration are significantly associated with final performance, while solution diversity and compute cost are not. Our benchmark is available at: https://anonymous.4open.science/r/Anonymous-78B6.


fmxcoders: Factorized Masked Crosscoders for Cross-Layer Feature Discovery

Andreas D Demou ⋅ Panagiotis Koromilas ⋅ James Oldfield ⋅ Yannis Panagakis ⋅ Mihalis Nicolaou

Many features in pretrained Transformers span multiple layers: they emerge through stages of inference, persist in the residual stream, or are built jointly by parallel MLPs. Crosscoders (namely, sparse dictionaries trained jointly across layers) aim to recover these cross-layer features in a single shared latent space. We show that standard crosscoders largely fail at this purpose. Although their decoder weight norms spread evenly across layers, a functional coherence metric we introduce reveals that each latent's activation is effectively driven by only one or two layers on average. While functionally coherent latents act as human-interpretable concept detectors (e.g., US states and cities), the layer-localized latents that crosscoders predominantly learn collapse onto surface-level patterns such as digit detectors. We trace this failure to two structural limitations: unconstrained cross-layer parameterization and unregularized cross-layer dependence. We address both by introducing fmxcoders, which (i) replace the encoder and decoder with low-rank tensor factorizations that draw every latent's per-layer weights from a shared cross-layer basis, and (ii) apply stochastic layer masking, a denoising regularizer along the layer axis that penalizes latents whose contribution collapses when a single layer is masked. Across GPT2-Small, Pythia-410M, Pythia-1.4B, and Gemma2-2B, fmxcoders lift mean probing F1 by 10-30 points, surpassing per-layer SAE baselines that standard crosscoders fail to reach, reduce reconstruction MSE by 25-50%, and roughly double mean functional coherence. An LLM-as-a-judge evaluation further shows that fmxcoders recover 3-13 times more semantically coherent latents than standard crosscoders across all four base LLMs.


Focusing Influence Mechanism for Multi-Agent Reinforcement Learning

Yisak Park ⋅ Sunwoo Lee ⋅ Seungyul Han

Cooperative multi-agent reinforcement learning (MARL) under sparse rewards remains fundamentally challenging because agents often fail to concentrate their influence, leading to insufficiently coordinated exploration. To address this, we propose the Focusing Influence Mechanism (FIM), a framework that encourages agents to focus their influence on under-explored parts of the state space through an entropy-based criterion, while leveraging eligibility traces to enable multiple agents to consistently align and sustain their influence on the same parts of the state space when beneficial, thereby promoting coordinated and persistent joint behavior. By emphasizing under-explored regions of the state space, FIM facilitates more efficient and structured exploration even under extremely sparse rewards. Across diverse MARL benchmarks, FIM consistently improves cooperative performance over strong baselines.


FocusVLA: Hijacking Attention to Break Visual Token Pruning in Vision-Language-Action Models

Yanhui Li ⋅ Qianpu Sun ⋅ Qi Zhou ⋅ Chenru Jiang ⋅ Peiyu Zhang ⋅ Dongxia Wang

Vision-Language-Action (VLA) models are becoming an important approach for robotic manipulation, but their long visual token sequences make inference expensive. Visual token pruning is a practical way to reduce this cost, with many pruning methods rely on attention scores to decide which tokens to remove. However, we show that this reliance can be exploited: an attacker can hijack attention scores and turn the visual token pruner into a vulnerability. Based on this insight, we propose FocusVLA, a backdoor attack against attention-based pruning. FocusVLA inserts trigger-controlled attention patterns into a few low-impact attention heads. These poisoned heads have little effect on normal manipulation. Without pruning, the poisoned model still behaves normally when the trigger appears. With pruning enabled, clean performance remains intact. But inputs with the trigger shift attention to task-irrelevant peripheral regions. As a result, the pruner removes task-critical tokens, leading to task failure. We evaluate FocusVLA on OpenVLA-OFT and $\pi_{0.5}$ across LIBERO tasks with ten attention-based pruning methods. For example, on OpenVLA-OFT with ADP pruning, the success rate on LIBERO-10 drops from 90.6\% to 31.0\% under triggered inputs. These results reveal a broader risk: attention is widely used as an importance signal, but it can be manipulated and should not be blindly trusted.


FoldAbS: Repurposing the Protein Folding Model as a Foundation Encoder for Antibody Screening

Jun Wu ⋅ FANDI WU ⋅ Xinyuan Zhu ⋅ Dawei Huang ⋅ Kaiwen Cheng ⋅ Jianzhu Ma ⋅ Fuli Feng ⋅ Jianhua Yao

Therapeutic antibody screening requires prioritizing antibody candidates across specificity, affinity, and developability. Current pipelines typically silo these objectives and underuse structural information that governs antibody-antigen recognition. To address these inefficacies, we introduce FoldAbS, a unified supervised screening framework that reconceptualizes protein folding models as frozen foundation encoders, exploiting their internal representations to capture rich sequence, geometric, and interaction dynamics well beyond basic structure prediction. By coupling the folding model with lightweight, task-specific heads, FoldAbS seamlessly evaluates all three therapeutic criteria within a single architecture. Instantiating this framework with the open-source Protenix model (FoldAbS-Px) yields consistent improvements over baseline approaches, achieving significant performance gains of up to 13.1%, 9.2%, and 6.2% in specificity, affinity, and developability, respectively. Collectively, these results reposition protein folding models from structure predictors and confidence scorers into reusable foundation encoders for comprehensive and supervised antibody screening.


Follow-Bench 2.0: An End-to-End 3D Benchmark for Socially-Aware Robot Person Following

Hanjing YE ⋅ Tianle Zeng ⋅ Jianwei Peng ⋅ Yanci Wen ⋅ Yuchen Zhou ⋅ Yonggen Ling ⋅ Hong Zhang

Socially-aware robot person following (RPF) requires a mobile robot to follow a designated person in dynamic human environments while maintaining target identity, avoiding obstacles and pedestrians, and preserving socially comfortable formations. Existing benchmarks only partially evaluate this coupled problem: embodied visual tracking benchmarks emphasize target visibility with limited social interaction, person re-identification benchmarks study perception without closed-loop control, and motion-planning benchmarks often assume ground-truth target and pedestrian states. We introduce Follow-Bench 2.0, an end-to-end 3D benchmark for evaluating socially-aware RPF under coupled perception--planning challenges. Built in Unreal Engine, Follow-Bench 2.0 provides difficulty-leveled scenarios with diverse pedestrian flows, lighting and weather conditions, cluttered layouts, bottlenecks, queueing behaviors, distractors, and temporary target occlusions. The benchmark supports modular perception--planning pipelines, end-to-end visual policies, and foundation-model-based active trackers under a unified closed-loop protocol. It further evaluates both back- and side-following configurations and reports safety--comfort metrics covering task success, target visibility, target recovery, social-space intrusion, following formation, and motion smoothness. Experiments show that RPF-oriented modular pipelines substantially outperform current end-to-end visual policies on close and socially comfortable following, but also reveal that socially-aware RPF remains far from solved in interaction-heavy environments and especially in side-following. Follow-Bench 2.0 therefore provides a diagnostic platform for studying how perception errors, occlusions, crowd interactions, and following configurations jointly affect robot person following. Our code is available at https://anonymous.4open.science/r/follow-benchv2-NIPS2026-03D3/.


FoMEMO: Towards Foundation Models for Expensive Multi-objective Optimization

Yiming Yao ⋅ Fei Liu ⋅ Liang Zhao ⋅ Xi Lin ⋅ Yilu Liu ⋅ Qingfu Zhang

Expensive multi-objective optimization is a prevalent and crucial concern in many real-world scenarios, where sample-efficiency is vital due to the limited evaluations to recover the true Pareto front for decision making. Existing works either involve repeatedly fitting Gaussian process surrogates from scratch for each newly encountered problem, or rely on computationally demanding hypervolume-oriented policy learning for amortized optimization, making it challenging to achieve scalable pre-training and efficient yet robust adaptation for diverse emerging real-world applications. To address these challenges, we propose FoMEMO (Foundation Models for Expensive Multi-objective Optimization), which adopts a decomposition-based training paradigm that converts multi-objective optimization into preference-wise aggregated posterior learning tasks, enabling scalable pre-training on hundreds of millions of diverse synthetic datasets without relying on extensive real-world domain experiments. At test time, given observed trajectories from unseen problems and user preferences, the foundation model performs in-context posterior prediction and optimizes the derived acquisition functions for efficient candidate generation, achieving strong generalization and optimization performance without any subsequent model training or updates.


Formal Conjectures: An Open and Evolving Benchmark for Verified Discovery in Mathematics

Moritz Firsching ⋅ Paul Lezeau ⋅ Salvatore Mercuri ⋅ Miklós Z. Horváth ⋅ Yaël Dillies ⋅ Calle Sönne ⋅ Eric Wieser ⋅ Fred Zhang ⋅ Thomas Hubert ⋅ Blaise Aguera y Arcas ⋅ Pushmeet Kohli

As automated reasoning systems advance rapidly, there is a growing need for research-level formal mathematical problems to accurately evaluate their capabilities. To address this, we present Formal Conjectures, an evolving benchmark of currently 2615 mathematical problem statements formalized in Lean 4. Sourced from areas of active mathematical research, the dataset features 1029 open research conjectures providing a zero-contamination benchmark for mathematical proof discovery, and 836 solved problems for proof autoformalization. Notably, the repository provides a structured interface connecting mathematicians who formalize and clarify problems with the AI systems and humans attempting to solve them. Demonstrating its immediate utility, the benchmark has already been leveraged to make new mathematical discoveries, including the resolution of open research conjectures. We describe our approach to ensuring the correctness of these formalizations in a collaborative open-source project where contributions stem from an active community. In this framework, AI-generated proofs and disproofs serve as a valuable auditing mechanism to iteratively improve the fidelity of the benchmark. Finally, we provide a standardized evaluation setup and report baseline results on frozen evaluation subsets, demonstrating a climbable signal that measures the current frontier of automated reasoning on research-level mathematics.


Foundation Model Informed Acquisition Functions for Molecular Discovery

Qi CHEN ⋅ Fabio Ramos ⋅ Alan Aspuru-Guzik ⋅ Florian Shkurti

Bayesian optimization (BO) is widely used to accelerate molecular discovery by reducing costly oracle evaluations. Foundation models provide a promising source of prior knowledge for BO, but simply using their representations in standard surrogate-based pipelines is often unreliable in low-data regimes with high-dimensional features and vast discrete candidate spaces. Rather than asking whether LLMs are universally useful for molecular BO, we ask how Acquisition Functions (AFs) should be designed to exploit weak, high-dimensional, and partially informative foundation-model priors. We propose \emph{LLM-guided Acquisition Tree} (LLMAT), a surrogate-free BO framework that reformulates likelihood-free acquisition estimation as recursive local acquisition learning. LLMAT trains binary classifiers on LLM representations, where each classifier jointly induces a promising/non-promising partition and defines a local AF within the corresponding region. This produces a hierarchy of localized AFs and enables efficient candidate selection via Monte Carlo Tree Search. To stabilize learning with few observations, LLMAT meta-learns shared classifier parameters and initializations across tree nodes. An optional LLM-guided clustering module further reduces AF evaluation cost by restricting search to statistically promising coarse property clusters. Extensive experiments and ablations demonstrate substantial improvements in scalability, robustness, and sample efficiency for LLM-guided BO in molecular discovery.


FourierMoE: Fourier Mixture-of-Experts Adaptation of Large Language Models

Juyong Jiang ⋅ Fan Wang ⋅ Hong Qi ⋅ Sunghun Kim ⋅ Jing Tang

Parameter-efficient fine-tuning (PEFT) has emerged as a crucial paradigm for adapting large language models (LLMs). However, standard PEFT methods often struggle in multi-task fine-tuning settings due to task interference and a limited parameter budget. Recent approaches incorporate the mixture-of-experts (MoE) architecture, referred to as mixture-of-parameter-efficient-experts (MoPE), to alleviate this issue by dynamically routing inputs to specialized experts. However, these methods remain based on spatial parameterization, which may introduce structural redundancy and additional parameter overhead. To address these limitations, we revisit model adaptation from a spectral perspective. Our analysis uncovers heterogeneity in frequency sensitivity across model layers and downstream tasks, indicating that adaptation should be frequency-aware rather than uniformly parameterized in the spatial domain. Motivated by this insight, we propose FourierMoE, a novel framework that unifies spectral parameterization with MoE via frequency-specialized experts and the learning of conjugate-symmetric complex coefficients. We conduct extensive experiments across multiple model families on a diverse range of tasks, including commonsense reasoning, math reasoning, image classification, and natural language understanding. Experimental results on 28 benchmarks show that FourierMoE consistently outperforms full fine-tuning (FFT) and 15 competitive baselines in both single-task and multi-task settings, while requiring significantly fewer trainable parameters, demonstrating superior effectiveness, efficiency, and adaptability.


Fractional Power-of-Two Quantization for Efficient and Effective Multiplier-Free LLM Inference

Sunghyun Wee ⋅ Geunjae Choi ⋅ Hyeonjin Kim ⋅ Suyoung Kim ⋅ Nojun Kwak

Large Language Models (LLMs) demand efficient inference, where the multiply-accumulate (MAC) operations in linear projections dominate compute and energy cost. Post-training quantization (PTQ) reduces this cost by mapping weights and activations to low-bit integers. Power-of-Two (PoT) weight quantization promises efficient multiplier-free LLM inference via shift-add logic, but suffers significant accuracy degradation at low bitwidths due to its coarse base-2 quantization grid. We introduce **Fractional PoT (FPoT)**, a base-$\sqrt{2}$ (half-octave) fractional PoT grid that achieves practical 4-bit weight quantization while preserving a fully integer shift-add datapath. The method combines (i) the base-$\sqrt{2}$ weight grid with uniform integer activations, (ii) a calibration-time grid-alignment refinement absorbed into the static per-channel scale, and (iii) a lightweight 1-adder $\sqrt{2}$ approximation whose error is absorbed during weight calibration via approximate-grid quantization, yielding a processing element (PE) with significantly fewer gates than a 4-bit integer multiplier. On Llama-3-8B with 4-bit weight and activation quantization, FPoT improves zero-shot commonsense reasoning performance by 8.72% relative to the PoT baseline, while maintaining performance close to the uniform integer baseline, with only a negligible 0.4%p degradation. These results hold across various configurations spanning seven models and three bit-widths. Furthermore, FPoT can be integrated with state-of-the-art LLM PTQ methods, demonstrating framework-agnostic applicability.

Long-video question answering benchmarks like Video-MME-v2 typically evaluate systems under a uniform frame budget for every question, despite stark variation in the evidence each query actually demands. We analyze this mismatch on Video-MME-v2 and find that different question subsets prefer different frozen frame-budget policies, while uniformly increasing the frame count does not reliably improve all metrics. These findings motivate fixed-mean frame-budget routing, a controlled inference setting in which each question may receive a different frame budget while the average budget matches a uniform reference. We introduce FrameRouter, a training-free test-time framework that combines an evidence-demand router, a budget-constrained allocation rule, and a frozen bank of frame-sampling policies. By controlling how much visual evidence each question receives while leaving the underlying samplers unchanged, FrameRouter is orthogonal to query-aware frame selection. Across Video-MME-v2 and several long-video understanding benchmarks, FrameRouter improves long-video question answering by using the same average visual budget more selectively across questions.

Non-Euclidean optimisation methods with matrix-valued updates, such as Muon and Scion, have recently shown strong empirical performance for training Transformer models, yet their theoretical advantages over Euclidean methods remain poorly understood. We address this gap in the heavy-tailed non-convex regime, where stochastic gradients have bounded $p$-th central moments, $p \in (1,2]$. We show that certain non-Euclidean methods achieve optimal sample complexity under stronger stationarity measures, while Euclidean methods incur additional dimension-dependent costs. As a consequence, for $m \times n$ matrices, Muon finds an $\varepsilon$-stationary point in nuclear norm within $\mathcal{O}\left(\min\lbrace m, n\rbrace \frac{\Delta_1 L}{\varepsilon^2} \left(\frac \sigma \varepsilon \right)^{\frac p {p-1}}\right)$ iterations, absorbing heavy-tailed noise without extra dimension dependence, unlike Euclidean methods. We further prove this dimension dependence is optimal for all first-order methods under nuclear-norm stationarity. Experiments on large language models support our theory. Surprisingly, our results suggest that other Schatten geometries beyond the spectral geometry of Muon can perform competitively in certain settings.


Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States

Denis Peskoff ⋅ Joe Barrow ⋅ Christopher Vu ⋅ Diag Davenport

Progress in legal AI increasingly depends on access to authoritative legal text at scale. Yet one of the most consequential layers of American law remains largely absent from existing machine-readable corpora: local ordinances. Local codes govern zoning, housing, business licensing, public health, noise, animal control, and many other domains of everyday regulation, but they are fragmented across vendor platforms designed for human browsing rather than bulk research access. We introduce LOCUS—the Local Ordinance Corpus for the United States—a comprehensive corpus and county-harmonized access layer for U.S. municipal and county ordinance codes. The raw corpus, available for release to researchers, represents nearly all publicly available municipal and county ordinance codes. The resulting raw corpus contains codes from 9,239 cities and counties. A smaller county-harmonized LOCUS access layer provides coverage for the largest 2,309 of 3,144 U.S. counties, accounting for a majority of the population. We use OCR to handle the myriad of document formats that have kept the law from being a public resource. We release the corpus with coverage metadata to support reproducibility, downstream legal AI research, and the incremental expansion of machine-readable access to local law. We train a collection of ModernBERT-based classifiers and scorers to facilitate analyzing U.S. local law among several dimensions, such as opacity and paternalism, that have not previously been studied at this scale.


FreeSpec: Training-Free Long Video Generation via Singular-Spectrum Reconstruction

Fangda Chen ⋅ Shanshan Zhao ⋅ Longrong Yang ⋅ Chuanfu Xu ⋅ Zhigang Luo ⋅ Long Lan

Video diffusion models perform well in short-video synthesis, but their training-free extension to long videos often suffers from content drift, temporal inconsistency, and over-smoothed dynamics. Existing methods improve temporal consistency by combining a global branch with a local branch, but they often further decompose appearance consistency and temporal dynamics within each branch using predefined criteria. This assignment is unreliable when appearance and action progression are tightly coupled, such as in camera motion and sequential motion. We analyze the video temporal extension issue from a singular-spectrum perspective and show that enlarged self-attention windows induce spectral concentration: spectral energy becomes dominated by a few low-rank singular directions, preserving coarse structure but suppressing high-rank spatial details and motion-rich temporal variations. To mitigate this problem, we propose FreeSpec, a training-free spectral reconstruction framework for long-video generation. FreeSpec decomposes global and local features with singular value decomposition, and uses the global branch as low-rank spectral guidance and the local branch as a high-rank reconstruction basis. This spectrum-level fusion avoids the rigid feature partitioning of previous decomposition rules, preserving long-range consistency while better retaining spatial details and temporal dynamics. Experiments on Wan2.1 and LTX-Video demonstrate that FreeSpec improves long-video generation, especially for temporal dynamics, while maintaining strong visual quality and temporal consistency.


FreshMem: Brain-Inspired Frequency-Space Hybrid Memory for Streaming Video Understanding

Kangcong Li ⋅ Peng Ye ⋅ Lin Zhang ⋅ Chao Wang ⋅ Huafeng Qin ⋅ Jiayuan Fan ⋅ Tao Chen

Transitioning Multimodal Large Language Models (MLLMs) from offline to online streaming video understanding is essential for continuous perception. However, existing methods lack flexible adaptivity, leading to irreversible detail loss and context fragmentation. To resolve this, we propose FreshMem, a Frequency-Space Hybrid Memory network inspired by the brain's logarithmic perception and memory consolidation. FreshMem reconciles short-term fidelity with long-term coherence through two synergistic modules: Multi-scale Frequency Memory (MFM), which projects overflowing frames into representative frequency coefficients, complementing by residual details to reconstruct a global historical “gist”; and Space Thumbnail Memory (STM), which discretizes the continuous stream into episodic clusters by employing an adaptive compression strategy to distill them into high-density space thumbnails. Extensive experiments show that FreshMem significantly boosts the Qwen2-VL baseline, yielding gains of 5.20\%, 4.52\%, and 2.34\% on StreamingBench, OV-Bench, and OVO-Bench, respectively. Besides its plug-and-play design, following low-cost component-wise fine-tuning, FreshMem outperforms current fully fine-tuned methods and achieves a state-of-the-art performance of 82.31\% on StreamingBench, exceeding the baseline by 13.31\%, offering a highly efficient paradigm for long-horizon streaming video understanding.

Model-based reinforcement learning agents optimize policies by backpropagating through imagined latent trajectories. However, model reliability can vary substantially within a single rollout: some imagined steps produce useful gradients while others introduce harmful bias. Existing adaptive methods address this at either the rollout level (adaptive horizons, priority replay) or the per-step level (gradient truncation); their interaction remains unexplored. In this paper, Freshness-Gated Imagination (FGI), a lightweight trust layer for latent world models that combines rollout-level adaptation with per-step entropy-based soft gating at zero additional forward-pass cost, is proposed. Through systematic factorial ablation on Crafter (10 seeds), a superadditive synergy is uncovered: per-step gating alone harms performance (−6.2%), rollout-level adaptation alone is ineffective (+0.6%), while their combination achieves +14.8% improvement in terms of the Crafter geometric-mean score (Wilcoxon signed-rank p=0.032, Cohen's d=0.74). A bias-variance analysis shows that this synergy arises because rollout-level priority replay improves the calibration of the entropy signal on which the per-step gate depends. On DMC continuous control tasks, FGI preserves baseline performance, confirming no-harm in well-modeled environments. Code is available at https://anonymous.4open.science/r/FGI-89D9.

Large language models (LLMs) generally lack deep causal reasoning capabilities in temporal reasoning. Inspired by document layout reordering, this paper proposes a paradigm shift from discrete narratives to causal topologies. It constructs two architectures: a training-free neuro-symbolic framework, temporal causal reasoning agent (TCR-Agent), which achieves interpretable causal deduction and counterfactual truncation through external causal graph instantiation and a symbolic intervention engine; and an end-to-end model, temporal causal reasoning former (TCR-Former), which internalizes these mechanisms into the Transformer attention space via a causal-topological attention mask and temporal span biases. Experiments on the TempoCausal benchmark demonstrate that TCR-Agent comprehensively surpasses existing training-free baselines. At the same time, TCR-Former achieves an accuracy 16.5% higher than that of state-of-the-art (SOTA) methods, outperforming large-scale models such as GPT-5.4 and DeepSeek-V3.2 with only approximately 8.6B parameters. Further analysis reveals a structure-reasoning disconnection phenomenon in existing methods: external symbolic constraints and chain-of-thought (CoT) reasoning can improve reasoning steps, yet fail to enhance deep causal deductive capability. By internalizing temporal causal reasoning capabilities into the model, TCR-Former effectively bridges this gap, providing an efficient and reliable end-to-end foundation for complex reasoning scenarios.


From Finding to Linking: Benchmarking and Advancing Cross-Long-Video Reasoning for Multimodal LLMs

Meng Luo ⋅ Zikang Zhou ⋅ Shanqing Xu ⋅ Shize Zhang ⋅ Bobo Li ⋅ Hao Fei ⋅ Mong-Li Lee ⋅ Wynne Hsu

Reasoning across multiple long-form videos requires models to find sparse evidence within each video and link related entities, events, and narratives across streams. Existing long-video benchmarks mainly evaluate single-video understanding, while multi-video benchmarks typically use short clips. We introduce CLoVR-Bench, a comprehensive benchmark for Cross-Long-Video Reasoning with 2,000 expert-annotated QA pairs over 400 long-form videos. Its three-level taxonomy covers Comparative Analysis, Tracking and Retrieval, and Integrated Reasoning, spanning 14 tasks and 36 subtasks that diagnose both intra-video localization and inter-video alignment. Our evaluation of 13 representative MLLMs shows substantial gaps across this hierarchy, with performance degrading from localized comparison to long-range retrieval and integrated cross-video reasoning. We further propose HOLMES, a training-free framework that formulates cross-long-video reasoning as evidence-slot filling over a typed binding graph. HOLMES plans option-discriminating evidence slots, localizes evidence with risk-conditioned policies, records coverage certificates for failed searches, verifies cross-video bindings visually, audits constraints, and returns answers through hard graph-readout gates. Experiments show that HOLMES improves both open-source and closed-source backbones, outperforms adapted long-video baselines, and yields the largest gains on tasks requiring long-range evidence search and cross-video binding. Together, CLoVR-Bench and HOLMES provide a diagnostic testbed and an interpretable baseline for future cross-long-video reasoning.

Rough path signatures provide a universal feature map for continuous paths and, via the expected signature, a principled characterisation of path distributions. These results do not directly extend to \clag paths of Temporal Point Processes (TPPs), limiting the use of signature methods for event sequences. Furthermore, neural TPP models, including recent generative approaches, optimise per-event objectives with no global sequence-level loss, while evaluation of variable-length event sequences lacks distributional discrepancy measures. This paper proposes a common pathwise framework for addressing these limitations. We introduce the interarrival embedding, a stable (homeomorphic) lift from jump paths to continuous paths of bounded variation, enabling signature methods to discrete event sequences. Our theoretical contributions give rise to \sigTPP, the first signature-based generative model for TPPs, trained using a path-level loss on complete trajectories. We further analyse the space of counting paths and derive three distributional discrepancies, providing mathematically justified tools for evaluating generative TPP models. Across synthetic and real-world datasets, \sigTPP~achieves the best average rank based on 8 complementary metrics, outperforms or is within one standard error of the strongest baseline in $64\%$ of the dataset-metric pairs, and according to a relative score, improves against every baseline by at least $19\%$ on average.


From Label Priors to Task Evidence: Long-Video Frame Selection via Bayesian GRPO

Shuochen Chang ⋅ Bingjie Gao ⋅ Qingyang Liu ⋅ Qianli Ma ⋅ Yibo Miao ⋅ Haonan Zhao ⋅ Zhaohe Liao ⋅ Xiaofeng Zhang ⋅ Jiangtong Li ⋅ Li Niu

Multimodal Large Language Models (MLLMs) face significant scalability challenges in long-video understanding due to quadratic attention complexity and context window constraints. Existing retrieval-based approaches typically using supervised learning on static keyframe annotations. However, these label priors often fail to align with the actual information required by the downstream MLLM for complex reasoning, creating a fundamental misalignment between the training target and inference needs. To bridge this gap, we propose a probabilistic framework that reformulates frame selection as a Bayesian inference process. We first initialize a policy using supervised learning to capture general saliency. Then, we treat the frozen MLLM as a task environment and employ its negative log-likelihood (NLL) as dense reward evidence to update the policy via Bayesian Group Relative Policy Optimization (GRPO). This process explicitly transits the selection principle from label priors to task evidence. Extensive experiments on Video-MME, MLVU, and LongVideoBench demonstrate that our approach matches the accuracy of heavy MLLM scorers while maintaining the efficiency of lightweight models.


From Noise to Diversity: Random Embedding Injection in LLM Reasoning

Heejun Kim ⋅ Seungpil Lee ⋅ Jewon Yeom ⋅ Jaewon Sok ⋅ Seonghyeon Park ⋅ Jeongjae Park ⋅ Taesup Kim ⋅ Sundong Kim

Recent soft prompt research has tried to improve reasoning by inserting trained vectors into LLM inputs, yet whether the gain comes from the learned content or from the act of injection itself has not been carefully separated. We study Random Soft Prompts (RSPs), which drop the training step entirely and append a freshly drawn sequence of random embedding vectors to the input. Each RSP vector is sampled from an isotropic Gaussian fitted to the entrywise mean and variance of the pretrained embedding table; the sequence carries no learned content, and yet reaches accuracy comparable to optimized soft prompts on math reasoning benchmarks in several settings. The mechanism unfolds in two stages: because attention has to absorb a never-seen-before random position, the distribution over the first few generated tokens flattens and reasoning trajectories branch, and as generation continues this influence dilutes naturally so the response commits to a single completion. We show that during inference RSPs lift early-stage token diversity and, combined with temperature sampling, widen Pass@N, the probability that at least one out of N attempts is correct. Beyond inference, we carry the same effect into DAPO training and demonstrate practical gains. Our contributions are: (i) RSP isolates the simplest form of soft prompt --- training-free, freshly resampled --- providing a unified lens for the structural effect of injection that variants otherwise differing in training and form all share; (ii) a theoretical and empirical validation of the underlying mechanism; and (iii) an extension from inference to training.


From Retrieval to Reasoning: Agentic Mechanism Prediction from Cell Painting Profiles

Jiayuan Chen ⋅ Botao Yu ⋅ Tianyu Liu ⋅ Thai-Hoang Pham ⋅ Meng Wu ⋅ Ping Zhang

Cell Painting is a high-content morphological profiling assay widely used for phenotype-based biological inference, with mechanism of action (MOA) prediction as a central application. Existing approaches largely formulate Cell Painting-based inference as representation matching, assigning predictions from nearby reference perturbations in morphological feature space. However, retrieved neighbors are often noisy and partially misleading evidence due to batch effects, non-specific cytotoxicity, phenotypic convergence, and source-dependent variability. We reformulate Cell Painting-based MOA prediction as a calibrated evidence reasoning problem, where retrieved neighbors are treated as uncertain observations that must be evaluated, compared, and sometimes rejected before supporting a mechanistic conclusion. We propose PhenoAIR, a reliability-aware multi-agent framework that maintains a candidate-centric evidence memory and performs controller-guided refinement over phenotype- and mechanism-side evidence. PhenoAIR uses offline reference-set calibration to weight evidence by source reliability, phenotype stability, and mechanism-level confusion. We evaluate PhenoAIR on a benchmark constructed from JUMP Cell Painting profiles and annotations, covering controlled, realistic, and discovery-oriented open-world MOA prediction settings. PhenoAIR outperforms representation-matching and LLM-based baselines across all settings.


From Seeing to Foreseeing: Unleashing LVLM Thinking in Dynamic Latent Space

Yuxuan Liang ⋅ Xu Li ⋅ Xiaolei Chen ⋅ Rui Zhu ⋅ Zhe Liu ⋅ Haotian Chen ⋅ Yi Zheng ⋅ Wenjuan Meng ⋅ Zhuolin He ⋅ Fan Shi ⋅ Xiangyang Xue

Although Large Vision-Language Models (LVLMs) have achieved remarkable progress, complex visual reasoning remains challenging. Existing approaches suffer from fundamental limitations: Thinking about Image mainly performs reasoning in text space, Thinking with Image relies on external tools to introduce additional visual evidence, and Latent Visual Reasoning, despite moving reasoning into latent space, still focuses largely on static visual modeling and lacks explicit characterization of temporal dynamics. We therefore propose Dynamic Latent Visual Reasoning (DLVR), a training framework that enables LVLMs to reason about dynamics in continuous latent space from only a static image. We introduce Mahalanobis Novelty Token Selection and Novelty-Adaptive Temporal Quantization to construct dynamic latent supervision from open-source datasets, building DLVR-SFT-100K and DLVR-RL-4K. We further develop a two-stage SFT pipeline that first builds temporal grounding over explicit dynamic processes and then teaches the model to encode dynamic semantics and temporal structure into latent tokens. Finally, we propose Contrastive Dynamic Latent Policy Optimization, which encourages latent trajectories to align with real dynamics while moving away from counterfactual ones. DLVR empirically improves both dynamic-centric and general visual reasoning, achieving an average gain of 10.82% on BabyVision and a 10.50% improvement on HRBench4K, demonstrating the promise of dynamic latent reasoning for LVLMs.

Detecting training data memorization in diffusion models is important for copyright protection and privacy auditing. We view memorization as an abnormally local concentration process: during reverse-time generation, a memorized instance acts as a point-like attractor whose probability basin is sharper and more position-sensitive than that of a generalized concept. Starting from the Fokker–Planck equation, we derive an exact evolution law for the score field $\partial_t \mathbf{s}_t$ and show that its leading geometric contribution in score-dominated local regimes is $\mathbf{H}_t \mathbf{s}_t$, linking temporal concentration dynamics to spatial curvature. We then show that standard sampling along a shared guided trajectory already exposes this curvature information through a complementary transport-curvature response, without requiring explicit Hessian computation. Motivated by this analysis, we propose the **Dynamical Singularity Metric (DSM)**, an on-trajectory detector that measures the pathwise score-evolution discrepancy between conditional and unconditional branches. DSM requires no additional backpropagation or network evaluations beyond sampling itself. Experiments on Stable Diffusion show that DSM matches or exceeds curvature-based baselines while being substantially cheaper and effective at the earliest reverse step.


From Static Policies to Adaptive Priors in Offline Reinforcement Learning

Tianwei Ni ⋅ Vineet Jain ⋅ Akash Karthikeyan ⋅ Pierre-Luc Bacon

Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertainty outside the offline dataset is treated pessimistically to ensure robustness. We argue that this formulation becomes incomplete when an offline-trained policy is subsequently updated through online interaction, as increasingly occurs in modern intelligent systems through test-time adaptation and online fine-tuning. This position paper argues that, in such settings, the objective of offline RL should extend beyond immediate deployment and instead prioritize learning adaptive policy priors: policies that preserve the capacity to improve during subsequent interaction through memory, exploration, and self-correction. We formalize this perspective as adaptive offline reinforcement learning (AORL), distinguish it from offline-to-online RL, and explain why adaptability becomes important under distributional shift, limited dataset coverage, and changing test-time conditions. We further discuss Bayesian offline RL as one principled direction for constructing adaptive policy priors by preserving epistemic uncertainty over plausible environments. Finally, we outline connections, open challenges, and research directions for treating offline RL as preparation for future experience rather than as a static deployment problem.


From Weak Data to Strong Policy: Q-Targets Enable Provable In-Context Reinforcement Learning

yichen lin ⋅ Xuyuan Xiong ⋅ xue wang ⋅ Xiangfu Meng ⋅ Mike Wei ⋅ Tao Yao

Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. While such objectives support task inference from context, they tie the learned policy to the quality of offline actions and can fail when trajectories are weak or suboptimal. We propose Q-Target Pretrained Transformers (QTPT), which preserves context-conditioned inference but replaces behavior cloning with a Bellman-style Q-target objective. QTPT learns to estimate action values from contextual rewards and transitions, rather than simply imitating the behavior policy. We analyze QTPT in stochastic linear bandits and finite-horizon MDPs, deriving suboptimality bounds that separate offline-data effects from Transformer approximation error. Empirically, QTPT is most beneficial under weak offline data across controlled RL benchmarks, and controlled ablations show that context conditioning is essential while Bellman/TD targets provide additional gains. We further include D4RL experiments as higher-dimensional stress tests under matched Transformer baselines.

LLM-based coding agents are usually evaluated in familiar software settings: mainstream languages, common libraries, and public repositories. These benchmarks remain important, but they can hide how agents behave when the language itself is unfamiliar. We evaluate six contemporary coding agents on four esoteric programming languages using a sequential setup with file editing, local execution, and hidden-test grading. Our protocol exposes capability differences between these agents that mainstream coding and agentic benchmarks such as SWE-Bench Verified and Terminal-Bench 2.0 compress into much narrower bands. We observe that the strongest agents, Claude Opus 4.6 and GPT-5.4 xhigh, often avoid writing the target language directly. On Brainfuck and Befunge-98, they write Python programs that generate target-language code and debug those generators locally. Forbidding this metaprogramming strategy causes large performance drops. Text guidance distilled from this strategy does not materially improve weaker agents. In contrast, Opus-derived Python helper code for building generators, with no solved benchmark programs or hidden-test answers, sharply improves Sonnet 4.6 and GPT-5.4 mini on the same problems, while Haiku 4.5 remains low. More interpreter calls and output tokens improve stronger agents but leave weaker agents near their original performance, indicating that these resources amplify useful strategies rather than create them. Together, these results show that strong coding agents adapt to unfamiliar languages by using tools, feedback, and workspace state to build a working model of the target language. Metaprogramming is the clearest case, but the broader gap is constructing and debugging a strategy that works under the target language's rules.


FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale

Runyuan He ⋅ Qiuyang Mang ⋅ Shang Zhou ⋅ Kaiyuan Liu ⋅ Hanchen Li ⋅ Huanzhi Mao ⋅ Qizheng Zhang ⋅ Zerui Li ⋅ Bo Peng ⋅ Lufeng Cheng ⋅ Tianfu Fu ⋅ Yichuan Wang ⋅ Wenhao Chai ⋅ Jingbo Shang ⋅ Alex Dimakis ⋅ Joseph Gonzalez ⋅ Alvin Cheung

Many real-world coding challenges are open-ended and admit no known optimal solution. Yet, recent progress in LLM coding has focused on well-defined tasks such as feature implementation, bug fixing, and competitive programming. Open-ended coding remains a weak spot for LLMs, largely because open-ended training problems are scarce and expensive to construct. Our goal is to synthesize open-ended coding problems at scale to train stronger LLM coders. We introduce FrontierSmith, an automated system for iteratively evolving open-ended problems from existing closed-ended coding tasks. Starting from competitive programming problems, FrontierSmith generates candidate open-ended variants by changing the problems’ goals, restricting outputs, and generalizing inputs. It then uses a quantitative idea divergence metric to select problems that elicit genuinely diverse approaches from different solvers. Agents then generate test cases and verifiers for the surviving candidates, and these problems then join the seed pool for the next round. On two open-ended coding benchmarks, training on our synthesized data yields substantial gains over the base models: Qwen3.5-9B improves by +8.82 score on FrontierCS and +306.36 (Elo-rating-based performance) on ALE-bench; Qwen3.5-27B improves by +12.12 and +309.12, respectively. Moreover, FrontierSmith-generated problems elicit long-horizon code-agent behavior comparable to human-curated open-ended tasks, causing agents to spend substantially more turns and tokens than on closed-ended problems.


Frozen Memory Is Not Enough: Rethinking External Memory as Extraction

Mingyuan Li ⋅ Guangsheng Yu ⋅ Xu Wang ⋅ Shaoxiong Ji

Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, where a memory trained on a source model is frozen and attached to a different target model while only a lightweight reader is trained. Across architectures, tokenizers, and model scales, frozen memory improves every source--target pair in a $3 \times 3$ transfer matrix, with up to 15.7\% relative perplexity reduction. Ablations show that learned memory content and correct addressing both matter, but stronger readers account for much of the transfer gain; on OpenQA, a dual-layer 4-branch reader nearly closes the gap between same-model and cross-model transfer and reaches 38.78 average. These results position Engram as a hybrid external memory and suggest that, in this regime, progress depends as much on reader design as on storage itself. Code and reproducible setups are available at https://anonymous.4open.science/r/Engram_Extractor/README.md. .


Fusion or Confusion? Multimodal Complexity Is Not All You Need

Tillmann Rheude ⋅ Roland Eils ⋅ Benjamin Wild

Multimodal learning has become a prominent research area, with the potential of substantial performance gains by combining information across modalities. At the same time, model development has trended toward increasingly complex deep learning architectures, motivated by the assumption that multimodal-specific methods improve performance. We challenge this assumption through a large-scale empirical study by reimplementing 19 high-impact multimodal methods across nine diverse datasets with up to 23 modalities. Under standardized experimental conditions, including hyperparameter tuning, weight initialization, cross-validation, and statistical testing, increased multimodal complexity often yields confusion rather than effective fusion of data modalities. Accordingly, complex multimodal architectures do not reliably outperform unimodal baselines and a Simple Baseline for Multimodal Learning (SimBaMM). Through a focused case study, we further demonstrate concrete methodological shortcomings even in top-tier multimodal learning publications, underscoring the need for standardized evaluation practices. In summary, we argue for a shift in focus for multimodal learning: away from the pursuit of architectural novelty and toward methodological rigor.


FutureSim: Replaying World Events to Evaluate Adaptive Agents

Shashwat Goel ⋅ Nikhil Chandak ⋅ Arvindh Arun ⋅ Ameya Prabhu ⋅ Steffen Staab ⋅ Moritz Hardt ⋅ Maksym Andriushchenko ⋅ Jonas Geiping

AI agents are being increasingly deployed in dynamic, open-ended environments that require adapting to new information as it arrives. To efficiently measure this capability for realistic use-cases, we propose building grounded simulations that replay real-world events in the order they occurred. We build FutureSim, where agents forecast world events beyond their knowledge cutoff while interacting with a chronological replay of the world: real news articles arriving and questions resolving over the simulated period. We evaluate frontier agents in their native harness, testing their ability to predict world events over a three-month period from January to March 2026. FutureSim reveals a clear separation in their capabilities, with the best agent's accuracy being 25%, and many having worse calibration than a constant predictor. Through careful ablations, we show how FutureSim offers a realistic setting to study emerging research directions like long-horizon test-time adaptation, search, memory and reasoning about uncertainty. Overall, we hope our benchmark design paves the way to measure AI progress on long-horizon open-ended adaptation spanning multiple months in the real world.


fxBench: Evaluating and Understanding Formula Suggestions in Spreadsheets

Sanket Mhatre ⋅ Sumit Gulwani ⋅ Vu Le ⋅ Yasharth Bajpai ⋅ Gust Verbruggen

Spreadsheets are among the most widely used tools for computations, and formulas are their core building blocks. Besides agentic spreadsheet tools, proactive formula suggestions enable significant productivity gains and are now available in popular spreadsheet software (Microsoft Excel and Google Sheets). Such suggestions pose interesting challenges: the right context must be collected and represented for a machine learning model, which must then be able to write a correct formula, in a way that is fast and cheap enough to trigger on every $\texttt{=}$ (or subsequent keystrokes for auto-completion) without taking users out of their editing flow. Surprisingly, there is no public benchmark for formula suggestions on real spreadsheets. In this paper, we therefore (1) introduce fxbench as a manually curated benchmark of 503 formula suggestion tasks based on Sheetpedia, (2) break down the formula suggestion problem into four capabilities---gather context, understand context, understand intent, write formula---that a formula suggestion system requires, and (3) evaluate existing and new approaches to understand bottlenecks in these capabilities that pose interesting research directions. Specifically, we annotate each benchmark with the cells that describe intent to determine the theoretical performance limit of a given context selection strategy, we evaluate performance of small and large models on different representations of context---including a dense format with high theoretical coverage at low token counts---and we fine-tune smaller language models to show that we can learn to understand such dense formats. Additionally, we perform quantitative and qualitative analysis on failure modes to understand bottlenecks in the formula writing capability---the best configuration only achieved 58.8\% correct suggestions.


GameVerse: A Minute-Scale Gameplay Dataset for Long-Horizon Interactive World Modeling

Kang He ⋅ Wenshuo Peng ⋅ Chuanhao Li ⋅ Zihui Gao ⋅ Xiaojie Xu ⋅ Zhengyuan Lin ⋅ Yuzhe Ding ⋅ Donghong Ji ⋅ Kaipeng Zhang ⋅ Yongtao Ge

Recent advances in world models have emphasized the need for learning temporally consistent and interactive representations of dynamic environments. However, existing video datasets fail to simultaneously support long-horizon dynamics, precise camera control, and structured semantic representations. To address this, we introduce GameVerse, a large-scale game video dataset for interactive world modeling. GameVerse provides minute-level continuous video sequences, covering diverse interaction behaviors, across both first- and third-person perspectives. The dataset is constructed through a unified data processing pipeline, including scene-consistent segmentation, quality filtering, camera pose estimation, and multi-granular semantic annotation. GameVerse consists of 11,092 videos, with a total duration of 4,692 hours (195.5 days) and approximately 986.6 million frames. Each video segment is augmented with temporally stable camera trajectories and structured semantic annotations. Comprehensive data analysis and experiments demonstrate that GameVerse exhibits strong advantages in terms of scale, diversity, and annotation quality, and shows promising performance for long-horizon modeling and interactive world understanding. We expect GameVerse to provide a unified data foundation for long-horizon modeling and interactive world modeling, and to facilitate future research in this area.


GEAR: Generator-Adaptive State Space Models for Associative Recall

Jun Meng ⋅ Mohammadhossein Amouei ⋅ Zengyu Lin ⋅ Xinyu Hu ⋅ Benjamin C. M. Fung

Selective state space models (SSMs) such as Mamba-3 are efficient and increasingly competitive with Transformers. However, their fixed recurrent state remains a bottleneck for using context as temporary memory. To address this issue, we propose $ \textit{GEAR} $ (GEnerator-Adaptive state space models for associative Recall), a slow-hypernetwork that produces coefficients for LoRA updates to the SSM token-to-recurrence parameter generators. GEAR makes the parameter generator context-adaptive while preserving the original linear-time Mamba-3 scan. To evaluate whether this adaptation improves in-context recall, we pretrain 180M-parameter models for 2B FineWeb-Edu tokens and compare against Mamba-3 baseline and ablations at the same scale. In the Multi-Query Associative Recall (MQAR) task, models see $N$ in-context key-value bindings and must predict the value paired with a queried key. In compact MQAR evaluation where the adaptive path is active at the query, GEAR improves candidate-restricted NLL over Mamba-3 by 0.70, 0.69, and 0.48 at $ N=32,64,128 $, respectively. Beyond compact MQAR, GEAR reduces the penalty on older queried bindings and gives the clearest gain within the Mamba-3 family under moderate random-filler interference. These results show that context-adaptive SSM generator improves associative recall while preserving Mamba-3 language modeling capabilities.

Multimodal Chain-of-Thought (CoT) reasoning has significantly advanced large models, yet forcing continuous visual evidence into discrete text or fixed tool calls inevitably sacrifices fine-grained detail. Recent paradigms attempt to reason directly within a continuous latent space; however, supervising latent slots in isolation fails to account for their collective geometry, resulting in a fragmented latent space that lacks structural guidance. To ensure structural integrity across training stages, we propose GeLVR(Geometry-Consistent Latent Visual Reasoning), a framework that establishes geometric consistency as a governing principle: the relational structure is explicitly aligned during SFT and actively preserved throughout RL, anchoring both stages to a common geometric scaffold. To operationalize this paradigm during the reward-driven phase, we further introduce GePO (Geometry-preserving Policy Optimization), a policy-optimization algorithm tailored for the hybrid discrete-continuous action spaces of latent reasoning. GePO employs a spherical von Mises-Fisher (vMF) policy to respect the decoder's inherent geometry and integrates the geometry-preserving regularizer directly into the RL objective. This ensures that reward optimization and structural integrity advance jointly rather than in tension. Extensive experiments across multiple visual reasoning benchmarks demonstrate that GeLVR consistently outperforms state-of-the-art baselines, particularly in high-resolution and fine-grained perception tasks. Comprehensive ablations further validate the necessity of each geometry-consistent component in establishing stable and coherent latent reasoning. Code is available in the supplementary materials.


GEM: Interpretable Language Models via Geometric Embedding Alignment

Dimitrios Tsaras ⋅ Yankun Hong ⋅ Lei Chen ⋅ Zhiyao Xie ⋅ Mingxuan Yuan

Large language models encode rich representations in their hidden states, yet these representations remain largely opaque and disconnected from the model's vocabulary. We introduce Geometric Embedding Mixture (GEM), a pre-training regularization method that aligns transformer hidden states with the geometry of the vocabulary embedding space. GEM utilizes a self-supervised objective to encourage the final hidden state to approximate a probability-weighted mixture of token embeddings. This effectively minimizes the free energy of the models and pushes the representation into the convex hull of the vocabulary. We pre-train models at three scales (135M, 360M, and 760M) and demonstrate that GEM yields superior predictive uncertainty compared to standard training, reducing perplexity by up to 53\% on benchmarks. Furthermore, we show that this geometric structural constraint mitigates representation anisotropy and unlocks \emph{native interpretability}. Unlike baseline models, GEM enables direct layer-wise decoding of intermediate hidden states and better causal faithfulness using the untuned language head, revealing coherent semantic trajectories without the need for auxiliary probes. Our work demonstrates that imposing geometric structure during pre-training improves both performance and transparency, offering a path toward language models that are interpretable by construction.

Minimax optimization refers to a class of optimization problems with some variables to minimize and other variables to maximize. Biased stochastic gradient methods have shown practical success in solving minimax problems to either improve robustness, enhance communication efficiency, or decrease computational costs. These successes motivate a lot of theoretical works to study the convergence of biased stochastic gradient methods, while their generalization analysis remains untouched. In this paper, we present the first framework to study the stability and generalization of biased stochastic gradient methods for minimax problems. We establish a connection between stability and generalization for minimax problems by relaxing the existing bounded gradient assumption to a bounded second moment condition. We then introduce a generalized Lipschitz-type condition on bias and gradient estimators, and derive a general stability bound to clarify the connection among bias, gradient estimators and stability. We apply our general analysis to Zeroth-order and Clipped stochastic gradient descent ascent (SGDA), and derive stability bounds that match those of SGDA under appropriate smoothing/clipping parameters. We combine stability and convergence analyses together, and derive optimal excess risk bounds of order $1/\sqrt{n}$, where $n$ is the sample size.

Operator learning for partial differential equations (PDEs) aims to learn solution operators on infinite-dimensional function spaces from finite-resolution data. In this setting, it is important for the learned model to be discretization-invariant, or resolution-robust, and to reflect PDE-specific structure. It is therefore natural to ask how such structure should be encoded in the model architecture, hypothesis class, or learning procedure. In this paper, we study operator learning for solution operators of nonlinear parabolic PDEs based on Duhamel--Picard iteration. We formulate Picard iteration as an abstract state-transition model and present a theoretical framework for Picard-type operator learning. We derive implementation-agnostic generalization error bounds that separate the implementation error from the estimation error associated with the abstract state-transition model induced by Picard iteration. A key consequence is that increasing the Picard depth reduces the Picard truncation error without causing an unbounded growth of the entropy-based estimation error. We also extend the analysis to long-time prediction by rolling out the same learned local model over successive time blocks. Finally, we illustrate the theory for nonlinear heat equations on the torus using a Picard-type Fourier neural operator as a concrete implementation.


Generating from Discrete Distributions Using Diffusions: Insights from Random Constraint Satisfaction Problems

Alankrita Bhatt ⋅ Mukur Gupta ⋅ Germain Kolossov ⋅ Andrea Montanari

Generating data from discrete distributions is important for a number of application domains including text, tabular data, and genomic data. Several groups have recently used random $k$-satisfiability ($k$-SAT) as a synthetic benchmark for new generative techniques. In this paper, we show that fundamental insights from the theory of random constraint satisfaction problems have observable implications (sometime contradicting intuition) on the behavior of generative techniques on such benchmarks. More precisely, we study the problem of generating a uniformly random solution of a given (random) $k$-SAT or $k$-XORSAT formula. Among other findings, we observe that: $(i)$ Continuous diffusions outperform masked discrete diffusions; $(ii)$ Learned diffusions can match the theoretical `ideal' accuracy; $(iii)$ Smart ordering of the variables can significantly improve accuracy, although not following popular heuristics.


Generative Molecular Morphing for Flexible-Size Design via Unbalanced Optimal Transport

Malte Franke ⋅ Stefan P. Schmid ⋅ Žarko Ivković ⋅ Kjell Jorner ⋅ Andreas Krause

The success of generative molecular design hinges on a model's steerability toward high-reward samples. Because many molecular properties are intrinsically linked to molecular size, accurately capturing the joint distribution of properties and the number of atoms is essential. However, current diffusion and flow-based models fix the number of atoms, which ultimately limits their ability to navigate this complex relationship. To address this, we introduce Morph, a flexible-size generative model for conditional and unconditional 3D molecular design based on geometric graphs. By dynamically adapting size, Morph can seamlessly integrate existing structural priors, like scaffolds, and significantly enhances property steering. We show that Morph matches current fixed-size state-of-the-art models while offering the benefit of unparalleled sampling flexibility. We demonstrate out-of-distribution generation in regimes where previous models fail, paving the way for enhanced generative modeling for molecular design.


Gen-Searcher: Reinforcing Agentic Search for Image Generation

Kaituo Feng ⋅ Manyuan Zhang ⋅ Shuang Chen ⋅ Yunlong Lin ⋅ Kaixuan Fan ⋅ Yilei Jiang ⋅ Hongyu Li ⋅ Dian Zheng ⋅ Chenyang Wang ⋅ Xiangyu Yue

Recent image generation models have shown strong capabilities in generating high-fidelity and photorealistic images. However, they are fundamentally constrained by frozen internal knowledge, thus often failing on real-world scenarios that are knowledge-intensive or require up-to-date information. In this paper, we present Gen-Searcher, as the first attempt to train a search-augmented image generation agent, which performs multi-hop reasoning and search to collect the textual knowledge and reference images needed for grounded generation. To achieve this, we construct a tailored data pipeline and curate two high-quality datasets, Gen-Searcher-SFT-10k and Gen-Searcher-RL-6k, containing diverse search-intensive prompts and corresponding ground-truth synthesis images. We further introduce KnowGen, a comprehensive benchmark that explicitly requires search-grounded external knowledge for image generation and evaluates models from multiple dimensions. Based on these resources, we train Gen-Searcher with SFT followed by agentic reinforcement learning with dual reward feedback, which combines text-based and image-based rewards to provide more stable and informative learning signals for GRPO training. Experiments show that Gen-Searcher brings substantial gains, improving Qwen-Image by around 16 points on KnowGen and 15 points on WISE. We hope this work can serve as an open foundation for search agents in image generation, and we fully open-source our data, models, and code.


Geometry-Aware Representation Denoising for Multi-view Image Restoration and 3D Reconstruction

Jin Hyeon Kim ⋅ Jaeeun Lee ⋅ Claire Kim ⋅ Kyoungjin Oh ⋅ Paul H Cho ⋅ Jaewon Min ⋅ Yeji Choi ⋅ Jihye Park ⋅ Hyunhee Park ⋅ Park M Kyu ⋅ Hyungju Chun ⋅ Seungryong Kim

Multi-view 3D reconstruction has achieved remarkable progress with the advent of feed-forward 3D reconstruction models. However, these models are typically trained and evaluated under ideal, degradation-free imaging conditions, whereas real-world observations often contain various degradations that differ significantly from such settings. Improving robustness for multi-view 3D reconstruction under degraded conditions therefore remains an important challenge. We present Geometry-Aware Representation Denoising (GARD), a novel framework that performs diffusion-based multi-view restoration directly in the feature space of a feed-forward reconstructor. This design exploits the geometry-aware feature representations of the reconstructor to effectively recover accurate scene geometry. Furthermore, by employing a decoder, the refined representations can also be used to restore high-quality RGB images, thereby enabling the simultaneous recovery of 3D scene geometry and high-quality imagery. Comprehensive experiments on the DA3 benchmark demonstrate the effectiveness of the proposed GARD framework. Our code and weights will be publicly released for full reproducibility.


Geometry Meets Physics: Data-Efficient Pre-Training for Unstructured Neural PDE Solvers

Luis Medrano-Navarro ⋅ Giacomo Baldan ⋅ Qiang Liu ⋅ Benjamin Holzschuh ⋅ Jan Hagnberger ⋅ Mathias Niepert ⋅ Nils Thuerey

Neural surrogate models for Partial Differential Equations (PDEs) on unstructured 3D geometries are often limited by poor generalization and the high cost of generating large-scale training datasets. Consequently, pre-training has emerged as a critical alternative to enhance the robustness and scalability of these models. In this work, we introduce a pre-training framework tailored to both steady-state and transient regimes. For steady-state problems, we propose a geometry-driven strategy that leverages intrinsic shape descriptors to learn representations of complex 3D domains. For transient problems, we introduce a physics-driven approach based on online generation of synthetic PDE data, enabling scalable pre-training without reliance on expensive datasets. This framework learns transferable features for challenging downstream tasks. Across multiple experiments, our approach achieves up to 3$\times$ faster convergence, 2$\times$ greater data efficiency, and up to 25\% higher accuracy during fine-tuning, particularly under realistic low-data regimes. This methodology provides a practical pathway toward data-efficient neural emulators for large-scale simulations.

A paraphrase pair is a small piece of evidence that two activations should agree. We treat that as a sheaf condition and ask what its cohomology looks like at scale. A cellular sheaf over a paraphrase graph defines a coboundary δ⁰; its kernel is the space of activations that paste consistently across the graph, and the rest of the sheaf Laplacian's spectrum reads off how badly they fail to. On a matching graph this collapses to projected within-pair covariance, so the framework only earns its keep once depth supplies cycles: stacking the same activations across the transformer's own layers turns the construction into a 2-D grid with b₁ = N(L−1) algebraic 1-cycles, one per elementary square asking whether layer-l → l+1 commutes with paraphrase equivalence. The model-derived Hodge harmonic mass at L=2 varies by four orders of magnitude across architectures (from 10⁻⁴ on Mistral-7B and 10⁻³ on Llama-3-8B to 0.08 on Llama-2-7B), and the ordering by harmonic mass matches the ordering by steering fragility we measure independently. The framework supplies a mechanistic prediction that variance-matched and probe baselines do not. Three validations follow. (i) Across nine architectures from 124M to 13B, the spectral complement of LF carries 5.6 to 26.5× the causal influence of variance-matched controls (p < 10⁻¹⁵). (ii) On held-out CounterFact across eight architectures, sheaf H⁰ at 20 dimensions beats LEACE's preserved subspace at full hidden dimension: 85.6% vs. 42.7% on GPT-2, 87.1% vs. 77.7% on Mistral-7B, mean gap +17.9 pp. (iii) On Llama-2-7B, contrastive sheaf steering preserves 7.4× more facts than random under style transfer (31.0% vs. 4.2%, n=1000, McNemar p < 10⁻⁵⁰). The same operators read off representational collapse: tr(LF) − β log det Cov reproduces VICReg's invariance + variance + covariance recipe in coordinates and holds encoder rank at 57–60/64 where naive consistency collapses to 14–42/64 across seven architectures.


GNES: Neural-Guided Evolutionary Program Search for Interpretable Multi-Agent Control

Chen Wang ⋅ Minfang Lu ⋅ Xiangke Wang ⋅ Cheng Zhu ⋅ Fei Ming ⋅ Yaochu Jin

Synthesizing interpretable controllers for multi-agent systems requires optimizing discrete symbolic programs from sparse, expensive, long-horizon simulator feedback. Evolutionary program search preserves interpretability but often spends many evaluations on weak random variation, whereas deep multi-agent reinforcement learning can learn effective policies that are difficult to inspect. We propose GNES, a neural-guided evolutionary program search framework for symbolic controller synthesis. GNES casts controller design as a program-edit decision process: gene growth trees are program states, validity-preserving genetic operators are actions, simulator fitness provides the return, and policy--value graph neural networks learn edit signals over controller topology. Monte Carlo tree search performs look-ahead planning in controller-program space before simulator validation, closing the loop between evolution, learned guidance, and verified rollouts. Experiments on three decentralized AirSim swarm-control benchmarks show that GNES improves search efficiency and final controller quality relative to evolutionary and reinforcement-learning baselines, while automatically evolving compact, inspectable distributed velocity controllers. The discovered policies expose coordination motifs, such as density-normalized radial gains and motion-compensating terms, that differ from standard hand-designed swarm rules. Ablations and learned-search comparisons indicate that graph-structured policy--value learning and MCTS look-ahead both contribute to the gains.

A central challenge in deep Active Learning is that optimal selection strategies depend on the labeling budget, shifting from coverage in low-budget regimes to uncertainty in high-budget regimes. Existing methods attempt to bridge this gap via heuristic switching or interpolation, but lack a principled mechanism to determine when such transitions should occur. We address this with pseudo-CDNV (pCDNV), a label-free proxy for neural collapse that tracks representation maturity. We show that pCDNV exhibits a characteristic peak that reliably signals when learned representations become suitable for uncertainty-based selection. Building on this insight, we propose Geometry-Oriented Adaptive Targeting (GOAT-AL), a unified query strategy that operates across all budget regimes. Rather than switching objectives, GOAT-AL maintains a single coverage-based objective while adapting the underlying feature space from self-supervised to task-aligned representations. Experiments on CIFAR-10, CIFAR-100, and TinyImageNet demonstrate that GOAT-AL consistently matches or outperforms state-of-the-art methods across low-, mid-, and high-budget regimes.


Gradient Regularized Newton Boosting Trees with Global Convergence

Nikita Zozoulenko ⋅ Daniel Falkowski ⋅ Thomas Cass ⋅ Lukas Gonon

Gradient Boosting Decision Trees (GBDTs) dominate tabular machine learning, with modern implementations like XGBoost, LightGBM, and CatBoost being based on Newton boosting: a second-order descent step in the space of decision trees. Despite its empirical success, the global convergence of Newton boosting is poorly understood compared to first-order boosting. In this paper, we introduce Restricted Newton Descent, a framework for convex optimization with Newton's method on Hilbert spaces with inexact iterates, based on the concepts of cosine angle and weak gradient edge. Within this framework, we recover Newton boosting with GBDTs and classical finite-dimensional theory as special cases. We first prove that vanilla Newton boosting achieves a linear rate of convergence for smooth, strongly convex losses that satisfy a Hessian-dominance condition. To handle general convex losses with Lipschitz Hessians, we extend a recent gradient regularized Newton scheme to the restricted weak learner setting, establishing the first global $\mathcal{O}(\frac{1}{k^2})$ convergence rate for second-order GBDTs. This scheme minimally modifies the classical algorithm by introducing an adaptive $\ell_2$-regularization term proportional to the square root of the gradient norm at each iteration. In numerical experiments, we show that this scheme converges while vanilla Newton boosting may diverge.


GradTrack: Detecting Noisy Labels via Temporal Trajectories of Class-wise Gradient Misalignment

Townim Faisal Chowdhury ⋅ Ta Duc Huy ⋅ Hieu Phan ⋅ Anton van den Hengel ⋅ Johan Verjans ⋅ Gustavo Carneiro ⋅ Zhibin Liao

Noisy-label detection aims to identify incorrectly labeled training samples so that their adverse impact on model learning can be mitigated. Existing methods typically rely on loss, prediction confidence, or gradient magnitude signals, which primarily capture the scale of prediction errors. While effective, such scalar signals overlook the geometric structure and temporal dynamics of gradient evolution during training. In this work, we show that gradient magnitude provides a useful entry point for analysing the gradient distortion induced by noisy-label samples, enabling us to quantify their deviation from the gradients that would be generated by clean labels. In this light, we propose GradTrack, a simple and effective framework for detecting noisy-label samples by estimating class-wise gradient directions likely to be induced by clean labels and measuring the alignment of training samples with these directions. In each training epoch, GradTrack partitions samples based on gradient magnitude and approximate clean-label gradient directions to compute a class-wise gradient misalignment score. These scores are used to rank training samples, which are then aggregated across training epochs to form a temporal rank trajectory. Finally, a Gaussian mixture model is applied to detect noisy-label samples. Extensive experiments on both synthetic and real-world noisy-label benchmarks demonstrate that GradTrack achieves strong detection performance, particularly under instance-dependent noise, and can further improve existing learning-with-noisy-labels pipelines as a plug-in sample selection module.

Prototype-based medical image classifiers have three clinical gaps: they treat findings as independent, silently amplify unsafe doctor feedback, and require full retraining whenever a new finding is needed. We present GRAPE (Graph-Augmented Prototype Explanations), a unified architecture that closes all three gaps. A Graph Attention Task Head models anatomical concept co-occurrence, boosting macro-F1 by +13.8\,pp over the prototype baseline on TBX11K. A Concept-Mismatch Safety Check is the first such mechanism in prototype-based medical classifiers, warns when the model's dominant finding inside a doctor-drawn region conflicts with the claimed label, catching 85\% of erroneous annotations versus 51\% for MC-Dropout with no extra inference cost. Open-Vocabulary Prototype Anchoring aligns visual prototypes to clinical text, so a new finding can be added from a single labelled image without modifying any other component: on NIH ChestX-ray14, one Effusion example recovers full-supervision localisation accuracy; on TBX11K, prototype maps achieve $2.6{\times}$ better lesion localisation than end-to-end baselines. All three capabilities add only $+1$~ms latency at interactive batch size.


Graph Energy Matching: Transport-Aligned Energy-Based Modeling for Graph Generation

Michal Balcerak ⋅ Suprosanna Shit ⋅ Chinmay Prabhakar ⋅ Sebastian Kaltenbach ⋅ Michael Albergo ⋅ Yilun Du ⋅ Bjoern Menze

Generative modeling of discrete data, such as graphs, underpins many scientific and industrial applications, including molecular discovery and materials design. In these domains, probabilistic inference is particularly valuable, as it enables composable generation and principled incorporation of desired constraints, such as structural or functional properties. Energy-based models naturally support this goal by capturing relative likelihoods and enabling composable inference by directly enforcing constraints during inference. However, discrete energy-based models typically struggle with efficient and high-quality sampling, as off-support regions often contain spurious local minima, trapping samplers and causing training instabilities, resulting in a fidelity gap compared to discrete diffusion models. To address this gap, we introduce Graph Energy Matching (GEM), a discrete generative framework inspired by the Jordan--Kinderlehrer--Otto (JKO) transport-map optimization perspective. GEM learns a permutation-invariant potential energy that simultaneously guides discrete transport from noise toward high-likelihood graph regions and refines samples within these regions. We further introduce a sampling protocol leveraging an energy-based switching strategy, seamlessly bridging rapid, gradient-guided transport and a local mixing regime for effective exploration. On molecular graph benchmarks, GEM matches or surpasses strong discrete diffusion baselines on most reported metrics. Beyond improving generation quality, GEM's relative likelihood modeling enables targeted exploration, facilitating compositional generation, property-constrained sampling, and interpolation between graphs.


Graph-Enhanced Attribute-Aware Modeling for Cloth-Changing Person Re-Identification

Yongkang Ding ⋅ Zi Ye ⋅ Tiantian Gong ⋅ Liyan Zhang

Cloth-changing person re-identification aims to retrieve images of the same person across cameras when clothing may vary over time. Unlike conventional Re-ID, appearance cues such as clothing color and texture become highly unreliable in cloth-changing scenarios. Existing methods either rely on auxiliary clothing-invariant biometric cues that may introduce estimation noise, or model attribute semantics as isolated labels, making it difficult to simultaneously achieve robustness to attribute noise and fine-grained visual--semantic interaction. To address these issues, we propose a Graph-Enhanced Attribute-Aware Modeling framework, termed GEAM. Specifically, we first preprocess the attribute representation to explicitly suppress clothing-related semantics. We then introduce an adaptive attribute graph to model the latent dependencies among identity-related attributes, thereby alleviating the interference caused by attribute prediction errors. Based on the refined attribute representation, we further map the enhanced attribute semantics into compact semantic tokens and inject them into the visual backbone in a hierarchical manner, enabling local visual features to adaptively absorb identity-relevant yet clothing-irrelevant discriminative cues. This design effectively improves robustness to cloth-changing scenarios while preserving the stability of pretrained visual representations as much as possible. Extensive experiments on multiple mainstream CC-ReID benchmarks demonstrate that GEAM consistently outperforms a variety of state-of-the-art methods under the challenging clothes-changing setting. The source code and pretrained weights will be publicly released on GitHub.


Graph Learning Should Move Beyond Restrictive Views of Spectral and Message-Passing GNNs

Antonis Vasileiou ⋅ Juan Cervino ⋅ Pascal Frossard ⋅ Charilaos Kanatsoulis ⋅ Christopher Morris ⋅ Michael T Schaub ⋅ Pierre Vandergheynst ⋅ Zhiyang Wang ⋅ Guy Wolf ⋅ Ron Levie

Graph neural networks (GNNs) are commonly divided into message-passing neural networks (MPNNs) and spectral GNNs, reflecting two largely separate research traditions in machine learning and signal processing. While MPNNs have a precise definition, there is no widely accepted criterion for what makes a mapping a spectral GNN. Most existing work restricts spectral GNNs to layered architectures based on linear spectral filters. Under this restriction, we show that spectral and spatial GNNs have largely equivalent expressive power. To promote progress in the field, we propose a precise definition of spectral GNNs based on eigenbasis symmetries, contrasting the definition of MPNNs via neighborhood permutation symmetries. We further argue that the two perspectives offer complementary strengths. MPNNs provide a natural language for discrete structure and expressivity analysis through tools from logic and graph isomorphism, while the spectral perspective offers principled tools for understanding smoothing, bottlenecks, stability, and community structure. Overall, we argue that progress in graph learning will be accelerated by clarifying the similarities and differences between these perspectives and by moving toward a unified theoretical framework.

As large language model (LLM) agents move from isolated prompting to long-horizon workflows, failures increasingly arise at the role-to-instance binding boundary, where task-specific role requests must be assigned to concrete agent instances under current service, network, and query conditions. Existing agent system research has improved role specialization, workflow topology, memory, and tool use, but often assumes a fixed stable execution environment. This assumption limits deployed reliability, because the same role request can exhibit different latency, failure probability, and output quality across agent instances operating under different service regions and network conditions. We propose Hedged Agent Computing (HACO), a runtime control scheme that treats each role request as a reliability-constrained selection problem over candidate agent instances, each coupling a role type, an LLM, and a concrete execution environment. Different from routing, HACO adaptively selects a hedge set of candidates for each invocation. Its allocation rule combines optimistic ranking, which prioritizes candidates with high estimated quality, reliability, and informative uncertainty, with conservative reliability accumulation, which stops selection only after the hedge set reaches a target success probability. Through experience harvesting, HACO updates candidate and link profiles from all executed candidate traces, including quality, success, latency, and network statistics. Experiments on various benchmarks, together with runtime degradation studies, show that HACO improves robustness and output quality under changing deployment conditions, while using lower token and latency cost than exhaustive parallel execution.


Hadamard Representation: Scaffolding Performance Across Model-free RL

Jacob Eeuwe Kooi ⋅ Zhao Yang ⋅ Mark Hoogendoorn ⋅ Vincent Francois-Lavet

Deep reinforcement learning agents progressively lose representational capacity during training: neurons become dormant, removing active capacity from the network, and effective rank collapses, leaving surviving neurons redundant. Existing remedies such as periodic resets, and special neural network architectures, are largely algorithm- or domain-specific. We propose a simple architectural fix, the Hadamard Representation (HR), which replaces a standard hidden layer with the element-wise product of two independently parameterized layers. HR operates through two complementary mechanisms. First, it reduces the probability of a neuron becoming dormant, which is particularly valuable for continuously differentiable activations such as $\tanh$: unlike dormant ReLU neurons, which are effectively pruned, saturated $\tanh$ neurons silently corrupt downstream layers by turning their outgoing weights into fixed biases. Second, independently of dormancy, the multiplicative structure captures richer feature interactions and increases effective rank without widening the layer. We evaluate HR across five algorithms and three domains: DQN, PPO, and PQN on pixel-based discrete-action Atari, SimbaV2 on state-based continuous control, and MR.Q on visual continuous control. HR consistently improves performance over the strong baselines without any hyperparameter tuning, and gains persist against parameter-matched wider variants, ruling out parameter count as an alternative explanation.


HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

Emmy Liu ⋅ Varun Gangal ⋅ Michael Yu ⋅ Zhuofu Tao ⋅ Karan Singh ⋅ Sachin Kumar ⋅ Steven Y. Feng

Hallucination remains a central failure mode of large language models, but existing benchmarks operationalize it inconsistently across tasks such as summarization, question answering, retrieval-augmented generation, and agentic interaction. This fragmentation makes it unclear whether a mitigation that works in one setting actually reduces hallucinations across contexts. Current hallucination benchmarks either require human annotation and fixed references that may eventually be memorized, or rely on naturalistic observations which are often recorded in settings that are difficult to reproduce or test systematically. To enable further research on the root causes of hallucination, we introduce HalluWorld, an extensible benchmark framework grounded in an explicit reference-world formulation: a model hallucinates when it produces an observable claim that is false with respect to this reference world. Building on this view, we construct a family of synthetic and semi-synthetic benchmark environments in which the reference world is fully specified, the model's observable view is controlled, and hallucination labels can be generated automatically by construction. HalluWorld spans multiple settings that are classically representative for AI, i.e., gridworlds, chess, and realistic terminal tasks. This enables controlled variation of key factors such as world complexity, observability, temporal change, and source-conflict policy, allowing us to disentangle hallucinations into more fine-grained error categories. We evaluate frontier and open-weight language models across these settings and find consistent patterns across domains: perceptual hallucination on directly observed information is near-solved for frontier models, while multi-step state tracking and causal forward simulation are still difficult for frontier models, and are not generally solved by extended thinking. In the terminal setting specifically, models also struggle with when to abstain from answering. The uneven profile of failures across probe types and domains suggest that different hallucinations arise from qualitatively distinct failure modes rather than reflecting a single underlying capability. Our results suggest that controlled reference worlds offer a scalable and reproducible path toward measuring and reducing hallucinations in modern language models.


HALMES: Knowing When to Intervene in LVLM Hallucination Mitigation

Ruining Hu ⋅ Jiaqi Lu ⋅ Xiao Liu ⋅ Ying Shen ⋅ Yu Wang ⋅ Lin Zhang

Large Vision-Language Models (LVLMs) have demonstrated strong capabilities in multimodal tasks. Nevertheless, their outputs remain prone to object hallucinations that are inconsistent with visual content. Contrastive decoding, as a training-free inference-time mitigation approach, typically intervenes in the output distribution at every decoding step. However, such indiscriminate full-stage intervention may disrupt originally correct generations and even induce new hallucinations. To address this issue, we first analyze the relationship between internal attention features and hallucinated token segments in LVLMs, revealing that hallucination-related abnormalities already emerge progressively in preceding token segments. Based on this observation, we propose HALMES, a lightweight plug-and-play module that identifies hallucination precursor segments from attention features during generation and selectively triggers contrastive decoding. HALMES can be integrated with various full-stage contrastive decoding methods, transforming indiscriminate intervention into on-demand targeted correction. Extensive experiments across different LVLM architectures and multiple benchmarks show that HALMES further improves the hallucination mitigation performance of existing contrastive decoding methods while reducing the negative impact of full-stage intervention on normal outputs, validating the importance of modeling and selectively intervening on hallucination-prone segments. All data and code will be made publicly available.


HandXFM: Semantic-Structural Distillation for Hand Radiograph Foundation Models

Yuxi Long ⋅ Ganlin Feng ⋅ Lianghong Chen ⋅ Liam J O'Neil ⋅ Carol A Hitchon ⋅ Pingzhao Hu

Hand radiographs support diverse clinical tasks, including skeletal maturity assessment, rheumatoid arthritis scoring, abnormality screening, and anatomical localization, yet existing models are typically trained separately for each task and dataset. We present HandXFM, a domain-specific foundation model that leverages cross-modality knowledge for comprehensive hand X-ray understanding. HandXFM distills semantic priors from BiomedCLIP and structural cues from a chest X-ray Vision Transformer (ViT) through a unified alignment framework, integrating complementary medical knowledge in an interpretable way. After pretraining, HandXFM can be adapted to a range of musculoskeletal tasks, including bone age estimation, SvH score prediction, abnormality detection, joint localization, bone segmentation, and visual question answering. It consistently outperforms task-specific and single-expert baselines, achieving up to 0.11 improvement in Pearson correlation coefficient (PCC) for bone age prediction and a 12\% accuracy gain in abnormality classification. Moreover, HandXFM generalizes to unseen datasets and yields cross-attention maps linking clinical terms to anatomical regions. These results demonstrate that multi-expert distillation effectively unifies semantic and structural supervision within a single pretrained model, establishing HandXFM as an interpretable and generalizable foundation model for hand radiograph analysis.

In standard acoustic anomaly detection (ASD), models are usually trained separately for each machine or domain using explicit metadata. However, in realistic deployments, machine identifiers are incomplete, unreliable, or simply unavailable. Thus, metadata-free universal ASD asks a single model to monitor mixed machine populations without machine IDs. In this paper, we argue that the main failure mode is not a weak detector family, but representation entanglement. The heterogeneous normal modes overlap in frozen pre-trained feature spaces, making anomaly scoring and approximate retrieval unreliable. We study this problem directly and propose HARMONY (Hierarchical Anchor Retrieval on Manifold for Oblivious-source acoustic aNomaly detection), an anchor-guided geometric repartitioning framework that reorganizes mixed-source features into compact Voronoi regions and reuses this structure for hierarchical retrieval. On dataset DCASE 2020 and MIMII, HARMONY improves mean AUC from 75.40% to 90.00% (+14.60% absolute improvement) over the strongest unified baseline and from 83.90% to 90.00% over a strong foundation-feature baseline. On a single CPU, its hierarchical retrieval achieves a 7.8× speedup over exhaustive search with only a 0.20% AUC drop. These results suggest that incorporating explicit geometric partitioning can improve both detection accuracy and retrieval efficiency for ASD. Code and data are available at the anonymous URL: https://anonymous.4open.science/r/HARMONY-26CC.


Hawkeye: Hardware-Aware GPU Kernel Optimization with Minimal Supervision

Arya Tschand ⋅ Kesavan Ramakrishnan ⋅ Alexander Ingare ⋅ Simon Guo ⋅ Jeffrey Ma ⋅ Zishen Wan ⋅ Simran Arora ⋅ Azalia Mirhoseini ⋅ Vijay Janapa Reddi

Achieving peak GPU kernel performance increasingly relies on architecture-specific optimizations targeting new hardware features. While AI coding agents show promise in generating performant kernels, they lack the necessary context to effectively implement and stack hardware-specific optimizations, especially on newer GPU architectures. We propose Hawkeye (Hardware-Aware Kernel Optimization), an open-source framework that grounds autonomous kernel generation in a minimal and comprehensive taxonomy with only one unit test per optimization strategy per target architecture. Supporting a new accelerator therefore requires only 10 expert-written unit tests per architecture (one per recurring optimization strategy) that generalize across downstream workloads, rather than hand writing a new kernel for each workload and precision. Hawkeye effectively scales the test-time compute of coding agents with this minimal expert supervision to enable kernel generation that consistently leverages hardware-specific features, approaching and even surpassing expert-written PyTorch or Triton in BF16 and emerging low precision (FP8, NVFP4, MXFP4) across Ampere, Hopper, Blackwell, and MI350 GPUs. Hawkeye demonstrates that minimally supervised coding agents can exploit architecture-specific hardware features and reduce the overhead of supporting emerging hardware accelerators.


HeatKV: Head-tuned KV-cache Compression for Visual Autoregressive Modeling

Jonathan Cederlund ⋅ Axel Berg ⋅ Durmus Alp Emre Acar ⋅ Chuteng Zhou ⋅ Pontus Giselsson

Visual Autoregressive (VAR) models have recently demonstrated impressive image generation quality while maintaining low latency. However, they suffer from severe KV-cache memory constraints, often requiring gigabytes of memory per generated image. We introduce HeatKV, a novel compression method that adapts cache allocation in each head based on its attention to previously generated scales. Using a small offline calibration set, the attention heads are ranked according to their attention scores over prior scales. Based on this ranking, we construct a static pruning schedule tailored to a given memory budget. Applied to the Infinity-2B model, HeatKV achieves $2 \times$ higher compression ratio in memory allocation for KV cache compared to existing methods, while maintaining similar or better image fidelity, prompt alignment and human perception score. Our method achieves a new state-of-the-art (SOTA) for VAR model KV-cache compression, showcasing the effectiveness of fine-grained, head-specific cache allocation.


HeterSEED: Semantics–Structure Decoupling for Heterogeneous Graph Learning under Heterophily

Xinyi Li ⋅ Ming Li ⋅ Lu Bai ⋅ Lixin Cui ⋅ Feilong Cao ⋅ Ke Lv ⋅ Yunliang Jiang ⋅ Pietro Lió

Many real-world heterogeneous graphs exhibit pronounced heterophily, where connected nodes often have dissimilar labels or play different semantic roles. In such settings, standard heterogeneous graph neural networks that aggregate messages along metapaths or meta-relations primarily based on feature similarity can propagate misleading information, since feature similarity may be misaligned with underlying relational semantics. In this paper, we propose HeterSEED, a semantics–structure decoupling framework for heterogeneous graph learning under heterophily. HeterSEED decouples representation learning into a heterogeneous semantic channel that captures type- and relation-aware local semantics and a structure-aware heterophily channel that separates homophilic and heterophilic neighborhoods via pseudo-label-guided partitioning and aggregates them using metapath-based structural weights. A node-level adaptive fusion mechanism then combines the two channels to produce context-dependent node representations. Theoretically, we establish that, on heterogeneous graphs under heterophily, HeterSEED is strictly more expressive than standard heterogeneous graph neural networks that rely primarily on feature similarity and provably reduces the prediction bias introduced by heterophilic neighbors. Experiments on five real-world heterogeneous graphs, including two large-scale networks at the million-node and hundred-million-edge scale, demonstrate that HeterSEED consistently outperforms representative heterogeneous graph neural networks and recent heterophily-aware baselines, especially in strongly heterophilic regimes.


H-GenPO: Hierarchical Generative Policy Optimization via the Option-Critic Framework

Wonhyeok Choi ⋅ Minwoo Choi ⋅ Jaeyeul Kim ⋅ Kyumin Hwang ⋅ Jongmin Gim ⋅ Sunghoon Im

Diffusion-based policies have demonstrated remarkable performance in online reinforcement learning (RL) by virtue of their expressive, multimodal action distributions. However, existing methods operate on a flat architecture that must generate a new action at every timestep, with no mechanism for temporal commitment to a behavioral mode—making them ill-suited for context-dependent tasks that require sustaining coherent behavioral modes, regardless of how expressive their action distributions are. We propose Hierarchical Generative Policy Optimization (H-GenPO), the first framework to integrate on-policy diffusion-based RL with the temporal abstraction of the Option-Critic architecture. H-GenPO adopts a two-level hierarchy in which a high-level policy over options selects among a discrete set of options, while a unified flow matching model—conditioned on a learned option embedding—serves as the expressive low-level intra-option primitive. The shared intra-option policy design encapsulates diverse behavioral repertoires within a single set of parameters, enabling scalable hierarchical control without the parameter growth associated with maintaining separate policy networks per option. We evaluate H-GenPO on 8 standard continuous control tasks in Isaac Lab and 3 custom context-dependent tasks designed to require coordinated behavioral mode switching. Our empirical evaluation demonstrates that H-GenPO achieves the best mean rank among all baselines on both standard and context-dependent benchmarks, while also exhibiting interpretable option specialization that emerges without any option-level supervision.

Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measure whether old answers remain correct, leaving the stability of multimodal grounding largely unexamined. We study this overlooked failure mode and ask whether a continually adapted MLLM can preserve not only what it answers, but also how it uses visual, textual, OCR, chart, and document evidence. We identify hidden evidence-use forgetting, where answer accuracy is retained while the model silently shifts toward different or less grounded evidence channels, and propose RCL, a replay-free reliance-constrained continual learning framework. RCL freezes the previous checkpoint as a behavioral reference, estimates teacher and student evidence-reliance profiles through counterfactual channel interventions, and jointly optimizes task learning, prediction preservation, and reliance preservation without adding inference-time cost. Across CoIN, COAST, MCITlib, and an evidence-sensitive multimodal stream, RCL consistently improves final performance and reduces forgetting over replay-free, PEFT, routing, and memory-assisted baselines, while substantially lowering modality reliance drift, dominant evidence flips, and hidden forgetting rates. These results suggest that robust continual multimodal learning requires preserving the evidence path behind correct answers, not merely the answers themselves.

Knowledge graph question answering (KGQA) systems are typically evaluated on questions whose surface tokens lexically expose the gold KG path. Real users, however, ask questions without knowing the underlying schema: "Which condition could Andrew have due to family history?" compresses a 4-hop chain (Andrew - father → father → medical condition → subclass of - hereditary disorder) into a single phrase (family history) whose tokens name none of those relations. We show that this collapse of path-question alignment is a blind spot of existing KGQA evaluation. The three dominant Large Language Model (LLM)-KG paradigms — hop-by-hop KG exploration, plan-then-retrieve, and similarity-based retrieval — all implicitly rely on the question's surface leaking the gold path, an assumption rarely tested in practice. We make three contributions. (i) HiddenPathQA, a human-validated KGQA benchmark of 1,055 multi-hop questions (2-6 hops, multiple domains) constructed so that the question surface does not leak the gold KG path's relations. (ii) Holistic Path Selection, a reasoning paradigm that exposes all candidate relation chains around the topic entity to the LLM at once, replacing per-hop relation scoring with question-against-full-chain alignment. Two implementations, TieredHolistic and FlatHolistic, outperform six strong baselines (ToG, PoG, FiDeLiS, R2-KG, KAPING, KARPA) on HiddenPathQA by +8.8 to +17.9 points over the strongest baseline across all four (backbone × split) settings. (iii) A pseudoword stress test that replaces entity surfaces with placeholder tokens. It both confirms HiddenPathQA items are solvable from KG structure alone — not from entity-level priors — and reveals that several baselines rely on entity memorization rather than KG-grounded reasoning.

High-stakes prediction systems are often evaluated only on labels revealed by past decisions. This selective observation can hide the losses that determine full-population tail risk: a predictor may appear safe on revealed labels while its worst compatible failures remain unseen. We study certification of conditional value-at-risk (CVaR) under selective labels. First, we prove non-identifiability: even with overlap, two predictors with identical observed selected-loss distributions can have reversed full-population CVaR rankings under compatible hidden-label laws. Under an outcome-dependent odds-ratio sensitivity model, we derive sharp upper and lower CVaR envelopes. A minimax interchange shows that worst-compatible hidden-label completion commutes with the Rockafellar--Uryasev threshold optimization, giving an exact certificate rather than a loose robust surrogate. We then define the tail-risk certification frontier: the minimum passive revelation cost needed to reduce hidden-tail ambiguity below a target tolerance. For a fixed predictor and threshold, this frontier is a fractional-knapsack problem whose value density, tail-identification value, pinpoints labels that can move the CVaR tail. Finally, we give finite-sample guarantees for cross-fitted conservative certificate learning and a lower bound governed by the effective number of revealed tail labels.


HIDRA: Hierarchical Dual-Routing Attention for Replay-Free Lifelong Imitation Learning

Fanqi Yu ⋅ Matteo Tiezzi ⋅ Cigdem Beyan ⋅ Tommaso Apicella ⋅ Vittorio Murino

Robotic agents operating in real-world environments must adapt continuously from a stream of multimodal demonstrations, often under constraints that preclude storing or revisiting past data. In this replay-free, single-pass lifelong imitation learning setting, sequential updates can progressively distort the representation space, making expert-based approaches particularly sensitive to routing interference, causing expert selection to degrade over time. We propose HIDRA, a structured routing framework that mitigates this issue through a hierarchical dual-attention mecha8 nism, which decouples instruction-level expert retrieval from context-dependent refinement. To further stabilize routing under representation drift, HIDRA introduces a key-level geometric regularization that enforces both alignment within tasks and separation across tasks. The resulting approach enables reliable expert selection while supporting an expandable set of experts for continual adaptation. We evaluate HIDRA on several LIBERO benchmark suites, including more challenging variants with heterogeneous task sequences and paraphrased language instructions. Our approach consistently outperforms both replay-free and replay-based baselines, improving AUC and forward transfer while maintaining low forgetting, with stronger gains under high interference and language variation. Code will be released.


HIFC-IQA: Train-Free Cross-Domain Image Quality Assessment via Dual-Process Cognition

Yu Li ⋅ Zhengran Shen ⋅ Puchao Zhou ⋅ Yachun Mi ⋅ Yukang Ding ⋅ Boyuan yang ⋅ Xi Zhai ⋅ Shaohui Liu

The discrepancy between controlled synthetic distortions and complex real-world degradations poses a significant domain shift challenge for No-Reference Image Quality Assessment (NR-IQA). While Unsupervised Domain Adaptation (UDA) methods aim to bridge this gap, they heavily suffer from retraining overhead and error propagation during iterative pseudo-labeling. Inspired by the dual-process theory of human cognition, we propose Hierarchical Intuitive and Fuzzy Consensus (HIFC), a completely train-free cross-domain IQA framework. HIFC employs an asymmetric perceptual decomposition to disentangle semantic and distortion features. It derives the final quality score through a synergy of two branches: a memory consensus branch that retrieves a stable Fuzzy Quality Centroid from a synthetic reference gallery, and an intuitive branch that captures aesthetic polarization via zero-shot vision-language alignment. Extensive experiments demonstrate that, without any parameter updates or target-domain adaptation, HIFC achieves highly competitive results against state-of-the-art UDA methods. Most notably, it exhibits exceptional robustness and generalization in the challenging synthetic-to-authentic cross-domain setting. The code will be available upon acceptance.


Higher-Order Action Supervision Makes A Strong Policy Class

Peng Cheng ⋅ Yunxian Hou ⋅ Zhi Zhou ⋅ Qian Zhang ⋅ Chang Huang ⋅ Xianyuan Zhan

Modern data-driven decision-making methods, such as imitation learning (IL) and reinforcement learning (RL), have achieved great success in solving many complex tasks. However, these methods often suffer from serious control instability and robustness issues when applied in real-world applications such as robotics and autonomous driving, posing notable challenges for their practical deployment. We argue that this instability issue stems largely from their limitations on solely supervising and optimizing zeroth-order actions (i.e., the action labels), failing to account for higher-order action dynamics and temporal consistency. In this paper, we show that simultaneously supervising both zeroth- and first-order actions can dramatically enhance policies' performance and control robustness. To achieve this, we introduce a novel and elegant loss scheme supported by formal theoretical guarantees that can equip any off-the-shelf policy model (e.g., deterministic, stochastic, or flow policies) with the capability for higher-order action supervision, without requiring any structural modifications. Moreover, our proposed method can serve as a lightweight plug-and-play module that seamlessly integrates with a broad spectrum of existing offline RL frameworks. Extensive evaluations on OGBench and D4RL demonstrate that our approach yields substantial performance and robustness improvements across a wide range of continuous control environments. Notably, our method can also enhance policies' out-of-distribution (OOD) generalization capability in the challenging low-data regime, making it an ideal tool in tackling many real-world control problems.


Historical Relative Policy Optimization for Bootstrapping LLM Reasoning

Sitong Wu ⋅ Haoru Tan ⋅ Bei Yu ⋅ Xiaojuan Qi ⋅ Jiaya Jia

Reinforcement learning has become a key approach for optimizing large language models (LLMs), with Group Relative Policy Optimization (GRPO) emerging as the most popular algorithm due to its simplicity and effectiveness. However, GRPO normalizes advantages purely within the current sampling group, ignoring the dynamic evolution of the policy's performance throughout training. This renders the model susceptible to relative deception, i.e., the model may be steadily deteriorating yet remain oblivious to its own regression. To address this issue, we propose Historical Relative Policy Optimization (HRPO), which integrates two synergistic designs: (i) a historical-aware advantage estimator that normalizes each response against the running maximum of the policy's strongest past group-mean rewards, thereby pushing the model to continuously surpass its own historical peak (bootstrapping); and (ii) an on-demand historical replay that recalls historical high-reward responses as positive examples, which is triggered only when all responses in the current step fail to provide positive signals. Experiments show that HRPO consistently outperforms GRPO variants in training stability, convergence speed, and final performance across diverse models and benchmarks, demonstrating that accounting for training dynamics leads to more reliable and effective policy optimization in LLMs.


Hitting Time Isomorphism for Multi-Stage Planning with Foundation Policies

Magnus Victor Boock ⋅ Abdullah Akgül ⋅ Mustafa Mert Çelikok ⋅ Melih Kandemir

We present a new operator-theoretic representation learning framework for offline reinforcement learning that recovers the directed temporal geometry of a controlled Markov process from hitting time observations. While prior art often produces symmetric distances or fails to satisfy the triangle inequality, our framework learns a Hilbert-space displacement geometry where expected hitting times are realized as linear functionals of latent displacements. We prove that this representation exists under latent linear closure and is uniquely identifiable up to a bounded linear isomorphism. For finite-dimensional implementations, we show that global hitting-time error is bounded by one-step transition error amplified by the environment's transient spectral radius. Furthermore, we provide finite-sample guarantees accounting for approximation, statistical complexity, and trajectory-label mismatch. Derived from this theory, we curate Isomorphic Embedding Learning (IEL) as a new goal-agnostic foundation policy learning algorithm that anchors a HILP-style consistency objective with explicit hitting-time regression to ensure that the learned geometry reflects actual decision-time progress. This asymmetric and compositional structure enables robust graph-based multi-stage planning for long-horizon navigation. Our experiments demonstrate that IEL improves the state of the art of learning foundation policy policies from offline maze locomotion data.


How Data Augmentation Shapes Neural Representations

Tianxiao He ⋅ Alex Williams ⋅ Sarah Harvey

Data augmentation is widely recognized for improving generalization in deep networks, yet its impact on the geometry of learned representations remains poorly understood. In this work, we characterize how different data augmentation strategies reshape internal representations in neural networks. Using tools from shape analysis, we embed network hidden representations into a metric space where distance is invariant to scaling, translation, rotation and reflection. We show that increasing augmentation strength leads to well-behaved trajectories in this space, and that different augmentation types steer representations in distinct directions. Moreover, we investigate how neural representation shapes are distorted along data augmentation trajectories, and show that insights from neural geometry can predict which representations provide the most improvement when ensembling models. Our results reveal shared geometric patterns across architectures and seeds, and suggest that analyzing shape-space trajectories offers a principled tool for understanding and comparing data augmentation methods.

Compositional priors describe the generic properties of layered functions in deep Bayesian models, where deep neural networks with random weights are a canonical example. In the wide-network limit, the prior is a Gaussian process with a depth-dependent kernel, and its behaviour as depth grows has been extensively studied through this kernel. Here, we study another case, where each layer itself is a vector valued Gaussian process, and our aim is similarly to understand the limiting behaviour of the prior as depth grows. Previous GP work has established that for the RBF kernel and a certain range of bandwidths $r$, the prior degenerates in the limit, converging to the set of constant functions --- which is not useful as a probabilistic model. In this paper we establish several new results. First, we identify a sharp bandwidth threshold $r_c(d) = \Theta(\sqrt{d})$ above which the limit is degenerate, strengthening the earlier bounds. Second, and more importantly, we show that for $r$ below the threshold $r_c(d)$ the prior converges to a limit distribution $\pi_{\bar{Z}}$. We also prove that these distributions are non-degenerate and non-Gaussian, with non-vanishing dependence between coordinates. In contrast to the previously known degenerate regime, deep Gaussian process priors can therefore admit non-trivial limits. Empirically, we verify the threshold across a range of dimensions $d$, and demonstrate a complex multimodal behaviour of the limit distributions $\pi_{\bar{Z}}$ --- a regime that becomes increasingly narrow with $d$ and would be hard to identify without knowing the threshold.


How do Small Transformer Models Learn Hard Math Tasks?

Eshika Saxena ⋅ Kristin E. Lauter

Recent works have demonstrated that transformers can be trained to recover sparse, binary cryptographic secrets in the Learning With Errors (LWE) problem, a foundational problem that underlies many post-quantum cryptographic schemes. However, as architectures have evolved to efficient encoder-only models, the mechanism by which these models recover the cryptographic secret has become more opaque. In this paper, we present the first layer-wise and embedding-level mechanistic interpretability analysis of encoder-only transformers trained on LWE samples. We reveal a surprising phenomenon: despite achieving near-zero exact prediction accuracy on the training objective, the models successfully recover the secret by bypassing the standard predictive pathways. We use dimensionality reduction, causal intervention, and linear probing and find that the secret is implicitly present in the positional embedding. Building on this mechanistic understanding, we introduce an architectural intervention that applies $L_1$ sparsity regularization directly to the positional embeddings. This modification forces the model to explicitly isolate the latent secret, transforming the computationally expensive post-hoc secret recovery process into a direct, human-interpretable parameter inspection. Our findings provide fundamental insights into how transformers allocate representational capacity when faced with high-noise, structured combinatorial problems.


How Finite-Rank Bottleneck Shape the Low-Rank Adaptation Landscape

Long Nguyen-Chi ⋅ Quynh Nguyen ⋅ Thanh Nguyen Cung ⋅ Binh T. Nguyen

Low-rank adaptation (LoRA) has become the standard method for parameter-efficient fine-tuning of large pretrained models, yet theoretical explanations for why low-rank updates suffice remain incomplete. Most existing analyses rely on the Neural Tangent Kernel (NTK) regime, which linearizes the network around the pretrained weights. In this work, we take a different approach that requires no linearization: our key observation is that, once the backbone is frozen, the loss sees each adapted weight matrix only through a finite-rank linear map induced by the frozen activations. Trace-norm regularization then forces the optimizer to pick a low-rank matrix among all matrices that produce the same observed output. This yields global minimizers whose ranks are controlled by the rank of this map, which we call the \emph{bottleneck rank}, rather than the sample size. For multi-head attention, we obtain tighter rank upper bounds by exploiting two additional invariances: row-wise softmax is insensitive to additive row constants in the score matrix, and value updates are observable only after attention weighting and frozen output projection. For the factorized LoRA objective with weight decay, we prove a complementary result on the nonconvex side: after an arbitrarily small generic positive semidefinite perturbation, every first-order stationary point is rank-deficient once the LoRA rank exceeds a threshold determined by the bottleneck rank. Finally, we show that global minimizers of the perturbed objective are near-optimal for the original trace-norm problem, and we provide uniform generalization bounds governed by the spectral structure of the frozen features.


How Private is Private? A Comparative Study for Face De-Identification

Hui Wei ⋅ Hao Yu ⋅ Hui Kuurila-Zhang ⋅ Guoying Zhao

Face de-identification (FDeID) has emerged as a critical privacy-preserving technology, yet its evaluation remains fundamentally fragmented. Existing protocols rely on inconsistent metrics, heterogeneous datasets, and partial annotation coverage, so methods targeting different utility dimensions, such as landmark versus expression preservation, are reported on different benchmarks under different metrics, rendering cross-method comparison infeasible. We revisit FDeID evaluation from both the data and metric perspectives. On the data side, we introduce UTILFACE, a curated, demographically balanced benchmark with high identity diversity, assembled from four large-scale face datasets through identity-aware cleaning, resolution enhancement, and stratified filtering. On the metric side, we propose HiFD, a Hierarchical Face De-identification metric that unifies identity suppression, multi-level utility preservation, and image quality under a single consistency-based paradigm: every component is computed from pretrained estimators’ outputs on the original face and its de-identified counterpart, directly quantifying how much identity is suppressed and how much downstream-perceivable utility survives. HiFD organizes facial signals into a three-level utility hierarchy spanning macro cues (L1), micro cues (L2), and imperceptible cues (L3), and aggregates the five resulting components into a single interpretable score via weighted harmonic mean, with configurable application-specific profiles. Using this unified protocol, we conduct a comprehensive comparative study spanning adversarial, GAN-based, and diffusion-based methods, surfacing trade-offs and failure modes that remain invisible under existing protocols. We will release the benchmark and evaluation toolkit to foster reproducible research in privacy-preserving face analysis (anonymous codebase: https://anonymous.4open.science/r/HiFD).

LLMs are increasingly deployed as agents within ecosystems where they compete for attention, are reused across contexts, and are reinforced through feedback loops. In such systems, behavioral traits spread not only because they are better, but because system design rewards some behaviors over others. We develop an analytical framework for LLM ecosystems, grounded in evolutionary dynamics, that links system-level selection to the evolution of trait distributions by treating exposure allocation as a source of selection pressure. We characterize the resulting regimes with stability-aware diagnostics (tangent stability and invasion exponents). We test the theoretical predictions in a minimal AI-native social platform in which LLM agents generate posts, evaluate one another, and compete for future visibility. The experiments support the theory: stronger selection increases concentration and pushes endpoints toward local instability; diversity support stabilizes specialization only within a bounded regime; and early trajectory fluctuations forecast later instability. We identify phantom diversity as a failure mode: endpoints can appear coexisting or specialized while remaining locally unstable or invadable. These results show that diversity audits should measure not only endpoint geometry, but also dynamic stability and early-warning signals.


How to make the most of your masked language model for protein engineering

Calvin McCarter ⋅ Nick Bhattacharya ⋅ Sebastian Ober ⋅ Hunter Elliott

A plethora of protein language models have been released in recent years. Yet comparatively little work has addressed how to best sample from them to optimize desired biological properties. We fill this gap by proposing a flexible, effective sampling method for masked language models (MLMs), and by systematically evaluating models and methods both in silico and in vitro on actual antibody therapeutics campaigns. Firstly, we propose sampling with stochastic beam search, exploiting the fact that MLMs are surprisingly efficient at evaluating the pseudo-perplexity of the entire 1-edit neighborhood of a sequence. Reframing generation in terms of entire-sequence evaluation enables flexible guidance with multiple optimization objectives. Secondly, we report results from our extensive in vitro head-to-head evaluation for the antibody engineering setting. This reveals that choice of sampling method is at least as impactful as the model used, motivating future research into this under-explored area.


How You Move Tells What You'll Do: Trajectory-Conditioned Egocentric Prediction

SeJoon Jun ⋅ Hai Nguyen-Truong ⋅ Luigi Seminara ⋅ Lorenzo Torresani

Predicting how a person's first-person view will evolve (what action will follow, what plan completes a task, whether an in-progress shot will score) is fundamentally under-specified: the same context admits many plausible futures, and a model trained to minimize prediction error is forced to hedge or average across them, getting it wrong either way. Two findings shape our approach. First, the future camera trajectory, the path the head carves through space, lets the model commit to one of those futures: it carries the operator's intent in a form fine enough to determine how an action will unfold, substantially outperforming language as a conditioning signal. Second, this same intent makes the trajectory itself partially predictable from the context at hand, enough that trajectory need not be observed at test time to recover most of the gain. We instantiate these findings as TrajPilot, a model that predicts candidate future trajectories from egocentric context and uses them to pilot action prediction in an action-aligned embedding space where language shapes the structure but is never used as a conditioning input. TrajPilot beats VLM and structured-planner baselines on procedural planning across Ego-Exo4D atomic, Ego-Exo4D Keystep, Ego4D GoalStep, and EgoPER, with the trajectory advantage widening with horizon (exactly where prior planners collapse) and holding under RGB-only camera-pose estimation. With the goal masked at inference, the same model performs goal-free anticipation, beating VLM baselines on Ego-Exo4D atomic and extending to EPIC-Kitchens-100 and basketball shot-outcome prediction.


Humanoid Horizon: Extending Task Horizon in Whole-Body Loco-Manipulation via Parallel Training, Dynamic Starting, and Reward Gating

Haozhuo Zhang ⋅ Qiang Zhang ⋅ Jian Tang ⋅ Mingzhe Ni ⋅ Michele Caprio ⋅ Angelo Cangelosi ⋅ Wei Pan

Cluttered indoor environments, where large and heavy objects are scattered across diverse surfaces, require humanoid robots to sequentially navigate, grasp, transport, and accurately place each item at its target location within a single uninterrupted episode. This long-horizon, whole-body loco-manipulation task remains a significant challenge for current methods. Previous approaches often suffer from two main issues: easy-reward bias, where training overemphasizes early transport stages at the expense of later ones, and catastrophic forgetting, where focusing on later stages leads to a decline in earlier-stage performance. In this work, we introduce Humanoid Horizon, a unified policy framework designed to overcome these limitations through three interrelated mechanisms. The Parallel Training Strategy expands $N$ scenes into $S \times N$ concurrent stage streams governed by a shared policy, ensuring all transport stages receive continuous gradient updates and removing the bottleneck of sequential optimization. The Dynamic Starting Mechanism updates each environment's initial state with terminal states from upstream rollouts, gradually broadening transition coverage and enhancing robustness at stage boundaries. Reward Gating sets the reward to zero in later-stage streams if any previously placed object is displaced beyond a set threshold, thus maintaining object placement throughout the episode without extra reward terms. Collectively, these strategies achieve per-stage success rates exceeding 80\% on the LHM-Humanoid benchmark (350 training scenes, 66 held-out scenes), with performance remaining stable even as the number of sequentially transported objects increases beyond two—unlike the sharp drop seen in all baselines. Additionally, we demonstrate that the RL teacher policy can be distilled into a Vision-Language-Action student via DAgger. When provided only with egocentric RGB or depth observations and natural language instructions, the student successfully replicates the teacher's long-horizon, multi-object behaviors, highlighting the potential for perception-driven deployment on real humanoid robots.


HYDRA: Representation Harmonized Tokenization for Multimodal Generation and Understanding

Xuerui Qiu ⋅ Yutao Cui ⋅ Guozhen Zhang ⋅ Junzhe Li ⋅ Yaqi Zhao ⋅ Xiao Zhang ⋅ Yang Li ⋅ Songtao Liu ⋅ Miles Yang ⋅ Yu Shi ⋅ Zhao Zhong ⋅ Liefeng Bo

Unified Multimodal Models struggle to bridge the fundamental gap between the abstract representations needed for visual understanding and the detailed primitives required for generation. Existing approaches typically compromise by employing decoupled encoders, stacking representation encoder atop VAEs, or utilizing discrete quantization. However, these methods often disrupt information coherence and lead to optimization conflicts. To this end, we introduce \tokenizer, a representation-harmonized pure ViT in the insight that visual modeling should evolve from generation to understanding. \tokenizer reformulates the standard backbone into a progressive learner that transitions from a Gen-ViT, which captures structure-preserving primitives, to a Sem-ViT for semantic encoding. Crucially, this transition is mediated by a Generation-Semantic Bottleneck (GSB), which compresses features into a low-dimensional space to filter noise for robust synthesis, then restores dimensionality to empower complex semantic comprehension. Built upon this foundation, we present \model, a native unified framework integrating perception and generation within a single parameter space. Extensive experiments establish \model as a new state-of-the-art. It sets a benchmark in visual reconstruction (rFID 0.08) and achieves top-tier generation performance on GenEval (0.86), DPG-Bench (86.4), and WISE (0.53), while simultaneously outperforming previous native UMMs by an average of $\sim$10.0 points across eight challenging understanding benchmarks.


Hyperbolic Concept Embedding Model for Interpretable Medical Image Diagnosis

Qihao Xu ⋅ Yadong Liu ⋅ Yong Xu ⋅ Xiaoling Luo ⋅ Chengliang Liu

Although deep neural networks (DNNs) have demonstrated strong performance in medical classification, their opaque reasoning undermines the trustworthiness of clinical diagnoses due to limited interpretability. Concept bottleneck models (CBMs) introduce an intermediate concept layer, decomposing black-box prediction into an explicit reasoning path: image → concept → disease. However, concepts and their corresponding diagnostic results in medical image diagnosis are often hierarchical, inclusive, and semantically unevenly distributed. Most concept-based methods represent concepts in Euclidean space and treat them as isolated entities, which hinders the modeling of their structural relationships and thereby reduces diagnostic transparency. To address this, we introduce a hyperbolic concept embedding model (HCEM) that maps medical concepts into a hyperbolic space better suited for hierarchical representation, enabling explicit modeling of complex concept relationships and their semantic associations with diseases. In this model, we regularize the hyperbolic concept embedding space using a positive–negative concept contrastive loss and a concept entailment cone loss. Furthermore, we employ an intervention-aware disease concept hyperbolic embedding regularization to learn their dynamic relationships. It avoids rigid rule priors while improving diagnostic consistency and flexibility under concept interventions. Extensive experiments on four medical datasets validate that our HCEM provides high accuracy in both concept and disease classification, as well as superior interpretability and intervenability. The code will be released soon.


Hyperbolic Graph Neural Networks Under the Microscope: The Role of Geometry–Task Alignment

Dionisia Naddeo ⋅ Jonas Linkerhägner ⋅ Nicola Toschi ⋅ Geri Skenderi ⋅ Veronica Lachi

Many complex networks exhibit hierarchical, tree-like structures, making hyperbolic space a natural candidate wherein to learn representations of them. Based on this observation, Hyperbolic Graph Neural Networks (HGNNs) have been widely adopted as a principled choice for representation learning on tree-like graphs. In this work, we question this paradigm by proposing the additional condition of geometry–task alignment, i.e., whether the metric structure of the target follows that of the input graph. We theoretically and empirically demonstrate the capability of HGNNs to recover low-distortion representations on regression problems, and show that their geometric inductive bias becomes helpful when the problem requires preserving metric structure. By jointly analyzing predictive performance and embedding distortion, we further show that HGNNs gain an advantage on link prediction, a naturally geometry-aligned task, whereas this advantage largely disappears on standard node classification benchmarks, which are typically not geometry-aligned. Overall, our findings shift the focus from only asking Is the graph hyperbolic? to also questioning Is the task aligned with hyperbolic geometry?, showing that HGNNs consistently outperform Euclidean models under such alignment, while their advantage vanishes otherwise.

While test-time fine-tuning is beneficial in cross-domain few-shot classification, the need for multiple backpropagation steps can be prohibitively expensive in resource-constrained environments. We propose HyperFlow, a gradient-free test-time adaptation method that amortizes fine-tuning dynamics into a lightweight conditional drift network. During offline training, HyperFlow learns from fine-tuning trajectories collected on meta-training tasks. Once trained, it adapts a selected PEFT parameter subspace for a new task by numerical ODE solving, requiring only forward passes of the drift network and no test-time backpropagation through the target model. In experiments on Meta-Dataset and CD-FSL benchmarks, our method improves out-of-domain performance over the direct transfer approach while using only 6-14\% of peak memory and about 1\% of the FLOPs of standard fine-tuning, positioning HyperFlow as an accuracy–efficiency trade-off between direct transfer and fine-tuning.

Multiview learning integrates complementary information from diverse observations to enhance model performance, where to quantify predictive uncertainty evidential deep learning has recently gained significant attention. However, existing methods generally overlook the critical importance of high-order correlation structures in multiview features, and traditional evidential combination rules frequently suffer from issues such as order dependence and irrational belief assignments when handling highly conflictive multiview features. To address these issues, we propose the Hypergraph-guided Global Mean-field Negotiation (HGMN) method for Multiview Evidential Classification. Specifically, HGMN first leverages hypergraph convolutional networks to capture high-order topological correlations within each view. Subsequently, instead of relying on pairwise or hierarchical fusion strategies, HGMN introduces a global mean-field negotiation mechanism that enables multiple views to dynamically reach a global consensus, effectively isolating dissenting noise. Finally, HGMN incorporates a multi-objective collaborative optimization strategy that enhances decision robustness and trustworthiness in complex open-world environments. Extensive experimental results on eight public datasets demonstrate that our method significantly outperforms state-of-the-art baselines in terms of accuracy and robustness.


Hypergraph Modeling of Transformer Attention for Hallucination Detection

Mark Junjie Li ⋅ zhishun liu ⋅ Wei Wang ⋅ Yuanshan Lu ⋅ yangjingxi ⋅ Xianyu Bao ⋅ Jun Li ⋅ Sunjie Huang

Large Language Models (LLMs) often generate factually unsupported content known as hallucination. However, existing detection approaches rely on handcrafted heuristics or isolated attention signals, limiting their ability to capture higher-order dependencies. In this work, we propose $\textbf{AttnHyper}$, a framework that represents a Transformer's attention patterns as a hypergraph for hallucination detection. Tokens are treated as nodes, while hyperedges connect groups of tokens that co-occur across heads and layers, with features derived from attention weights to preserve higher-order interactions beyond pairwise graphs. Based on this representation, the method formulates hallucination detection as a hypergraph learning problem and employs a hypergraph neural network (HGNN) to extract structural cues. It consistently outperforms state-of-the-art methods on the RAGTruth benchmark, achieving an average gain of +5.08 AUROC over the strongest baseline across diverse settings and architectures, while also demonstrating strong zero-shot transfer. Importantly, it incurs minimal overhead at inference time, enabling efficient deployment. Our results demonstrate that hypergraph-based attention modeling provides a more expressive and reliable signal for hallucination detection.


HyperVAttention: Efficient Sparse Attention with Spatio-Temporal Clustering for Video Diffusion

Dongyeun Lee ⋅ Amir Zandieh ⋅ Vahab Mirrokni ⋅ Junmo Kim ⋅ Insu Han

Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation. However, their ability to produce long-duration videos is fundamentally constrained by the quadratic complexity of the self-attention mechanism. Recent clustering-based sparse attention methods improve the quality-speed trade-off by grouping semantically similar tokens, but their practical efficiency remains limited by two bottlenecks: substantial clustering overhead and low CTA utilization caused by irregular cluster-induced blocks. We propose HyperVAttention (HVA), a training-free sparse attention framework that addresses both bottlenecks jointly. To reduce clustering overhead, we introduce 3D local-window clustering, which exploits the spatio-temporal locality of video tokens to restrict centroid search to fixed local neighborhoods, and implement it with a custom Triton kernel for efficient execution. We further propose a hybrid clustering strategy that performs full clustering only at anchor steps and updates only subset tokens at intermediate steps, leveraging the temporal stability of cluster assignments across denoising steps. To improve CTA utilization, we present hardware-aware cluster merging that minimizes CTA-aligned execution cost through parallel agglomerative merging, improving block density and approximation fidelity by utilizing idle tile capacity. Together, these components reduce clustering overhead, avoid redundant updates, and better align sparse attention with the fixed tile structure of modern GPU kernels. Experiments on Text-to-Video generation show that HVA establishes a new Pareto frontier for training-free sparse attention in video diffusion, reducing end-to-end latency by up to $2.13\times$ while improving fidelity over existing training-free sparse attention baselines.

Reconstructing the gene regulatory dynamics underlying tissue development from unpaired, cross-sectional spatial transcriptomic snapshots poses a significant challenge in single-cell biology. While recent optimal transport and flow-based methods have advanced trajectory interpolation, their unconstrained latent dynamics lack identifiability, hindering reliable discovery of the underlying regulatory mechanisms. To address this, we propose single-cell Identifiable Feedback-controlled Latent Flow (scIFLF), a generative framework that introduces a physics-inspired structural prior that decomposes the latent flow into a macroscopic developmental drift and a restorative feedback toward functional attractors. Crucially, the gene-regulatory Jacobian is shown to be uniquely identifiable, invariant to the latent ambiguity, enabling principled discovery of gene regulations directly from the learned flow. Implemented with a multimodal Neural ODE, scIFLF aligns latent trajectories using entropy-regularized optimal transport. Benchmark experiments demonstrate that scIFLF outperforms state-of-the-art methods in trajectory interpolation and spatial coherence, while successfully recovering key driver genes of tissue development and organogenesis, bridging the gap between deep generative flexibility and biological interpretability.


iDETR: Implicit DETR for Tiny Object Detection

Lingcong Cai ⋅ Haiqin Yang ⋅ Jikui Liu ⋅ Yunpeng Cai ⋅ Yumeng Liu ⋅ Xiaomao Fan

Tiny object detection (TOD) remains challenging due to the extremely limited spatial details of small objects. Existing methods heavily rely on multi-scale feature pyramids with discrete feature maps, which suffer from insufficient spatial resolution, quantization errors, and poor localization accuracy, while imposing heavy computational and memory overhead. We present the implicit DEtection TRansformer (iDETR), an efficient detector that operates exclusively on low-resolution single-scale features. It incorporates two novel components: (1) implicit Attention (iAttn), which leverages implicit neural representations to model continuous features and enables precise sub-pixel querying beyond discrete grid limitations; and (2) Centroid-Guided Query Initialization (CGQI) for robust query initialization under single-scale constraints. We further propose head-conditional sampling in iAttn, which reduces querying computational cost by 4× and memory footprint by 3× without sacrificing performance. Extensive experiments demonstrate the effectiveness of iDETR. Especially, on AI-TODv2, iDETR outperforms state-of-the-art methods by 0.2% AP overall, with particularly strong gains of 1.1% AP on very tiny objects, while reducing computational cost by 37%. The code is available at https://anonymous.4open.science/r/iDETR-FC25/.

Modern AI agents can plan, reflect, reason, and act over long-horizon digital tasks. Even so, they cannot plan to steer a social interaction in real time. Doing so requires anticipating how latent factors like trust and resistance will evolve under the agent's actions, fast enough to plan within a conversational turn. Generative world models approach this by narrating possible futures, but autoregressive text generation is both too slow for real-time planning and fundamentally lossy. Representation-predictive methods can be superior, but are underexplored for interaction. We build LID-Bench, the first controlled testbed with oracle latent states for interaction dynamics, enabling systematic comparison of generative and representation-predictive world models against known dynamics. A bottleneck decomposition on three generative models spanning 117M--350M parameters reveals that all learn interaction dynamics internally, with text-rendering signal losses of 89--100\% observed by horizon $k{=}5$. Our representation-predictive model (Social-JEPA, 125M params, 500K trainable) bypasses text entirely, forecasting latent dynamics $2.8$--$3.9\times$ more accurately while performing rollouts $1{,}059$--$2{,}314\times$ faster---making real-time planning within conversational turn-taking latencies feasible.


Imitation Dominates Reinforcement: Direct In-Context RL Is Closer to ICL Than RL

Minchan Kwon ⋅ Seunghee Koh ⋅ Sunghyun Baek ⋅ Minsung Bae ⋅ Junmo Kim

In-context reinforcement learning (ICRL) is often described as inference-time RL in which an LLM improves by accumulating trajectory-reward pairs in its context, with the reward acting as a learning signal. This framing poses a central question: is ICRL genuine inference-time RL, or better understood as a form of in-context learning (ICL)? We investigate this in the context of direct ICRL, where the model directly uses trajectory-reward pairs. Through controlled experiments on three benchmarks across six models, we find that the reward is read, but its effect is small; randomizing or removing the reward leaves the improvement curve almost unchanged. Trajectories drive improvement, but not through their semantic content: shuffled or corrupted trajectories work as well as real ones. These patterns closely mirror those known in ICL, suggesting that direct ICRL is better understood as a special case of ICL than as inference-time RL. This reframing has implications for ICRL memory design: ICL factors such as input distribution and demonstrations may matter more than RL elements such as reward shaping and exploration.


Implicit Goal Conditioning via Value Disaggregation

Shashwat Saxena ⋅ Mehul Goel ⋅ Sreyas Venkataraman ⋅ Sarvesh Patil ⋅ Max Simchowitz

Long-horizon tasks are an important challenge for modern AI and robotics, yet pose a challenge for reinforcement learning. A historically popular approach has been to decompose a task into subgoals, and learn a subgoal-conditioned policy to execute each subtask in sequence. We introduce an alternative called Value Disaggregation (VaDar), which trains a single subgoal-independent policy to maximize a linear combination of per-subgoal critic values. Because subgoal information is needed only at training-time, VaDar removes the need for possibly costly or cumbersome subgoal-selection during policy deployment. Moreover, through careful experiments, we show that VaDar (1) is always competitive with, and frequently outperforms, goal-conditioning and other natural baselines and (2) succeeds on tasks with non-sequential goal structure and multiple optimization objectives. In particular, VaDar is the first algorithm to solve challenging, long-horizon tasks in the Robocasa suite, even when starting from a policy pretrained from behavior cloning with near-zero initial task success. Taken together, VaDar's success suggests that subgoal information benefits reinforcement learning by enhancing the efficacy of critic learning, whereas policy goal-conditioning is often unnecessary.


Implicit Neural Representations for Variational Problems on Graphons

Taeyoung Kim ⋅ Jineon Baek ⋅ Joonkyung Lee ⋅ Hongseok Yang

We propose a neural exploratory framework for *variational problems over graphons*, i.e., symmetric measurable functions $[0,1]^2 \to [0,1]$, that arise as limits of dense graph sequences. Such problems are central to two parts of the mathematical literature: extremal graph theory, which studies graph parameter optimisation under given restrictions, often for those graphs with a large number of vertices, and the large deviation theory of dense random graphs, which characterises the structure of rare events through constrained graphon optimisation. We represent graphons by implicit neural networks and optimise graphon objectives by gradient descent. Our design combines three ingredients: a multi-scale sinusoidal residual architecture biased toward sharp, step-like graphons; an embedded solver that enforces a single empirical density constraint and is differentiated by implicit differentiation; and symmetry-aware Monte Carlo estimators. On generalised Turán problems that previously required substantial human effort, our framework rediscovers known optimal graphons without human intervention. Applied to open instances, it produces candidate optima, including a previously unreported family of extremal structures. Applied to the variational problems from the large deviation theory of Erdős–Rényi random graphs, the framework produces new candidate optimal graphons across both upper- and lower-tail regimes for multiple pattern graphs; for the case of the triangle graph and the upper-tail regime, these candidates improve upon the best-known reference construction of Lubetzky and Zhao.


Improved State Mixing in Higher-order and Block Diagonal Linear Recurrent Networks

Igor Dubinin ⋅ Antonio Orvieto ⋅ Felix Effenberger

Linear recurrent networks (LRNNs) and linear state space models (SSMs) promise computational and memory efficiency on sequence modeling tasks, yet their diagonal state transitions limit expressivity. Dense and/or nonlinear architectures (e.g., LSTMs) on the other hand are provably more expressive, but computationally costly. Here, we explore how expressivity in LRNNs can be increased via richer state mixing across time and channels while maintaining competitive efficiency. Specifically, we introduce two structured LRNN architectures: (i) Higher-order Linear Recurrent Units (H-LRU), which generalize recurrences to arbitrary order, mixing multiple past states, and (ii) Block-Diagonal LRUs (BD-LRU), which enable dense intra-block channel mixing. To ensure stable training, we introduce a selective gate normalization scheme that allows for scalable window and block sizes. To maintain efficiency, we utilize a parallel-scan implementation that keeps the throughput competitive with diagonal LRNNs for moderate orders (H-LRU) and block sizes (BD-LRU). Consistent with prior theoretical studies on the limitations of diagonal models, we empirically demonstrate in both synthetic sequence modeling and language modeling that our architectures significantly benefit from the increased expressivity of structured state mixing. Our results show that the structure of state mixing is a critical driver of performance in LRNNs, offering a practical pathway to closing the efficiency–expressivity gap in linear sequence models.


Improving Audit Realism with Inference-Time Compute and Deployment Scaffolds

Axel Ahlqvist ⋅ Richard Guan ⋅ Juan-Pablo Rivera ⋅ Adeline Kassler ⋅ Alexandra Souly ⋅ Kai Fronsdal ⋅ Robert Kirk ⋅ John Hughes

A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety audit can support. We present two techniques that make audits in Petri, where an auditor model red-teams a target over multiple turns to elicit concerning behavior, harder to distinguish from real deployments. Our first technique, \emph{critique refinement}, spends additional inference-time compute on each auditor action: an instance of the target model scores candidate actions for realism, provides feedback, and the auditor iterates before the highest-scoring candidate is selected. On Sonnet-4.6, realism win rate (the fraction of pairings in which the audit transcript is judged more realistic than a real deployment transcript) rises monotonically from 12\% to 39\% as refinement depth increases, and verbalized evaluation awareness drops to near zero; Haiku-4.5 and Opus-4.7 show a similar pattern. Our second technique, DISH (Deployment-Imitating SWE-Agent Harness), wraps the target in a Claude Code agent harness, reducing the gap between auditor-simulated and real deployment environments; in coding settings DISH raises realism win rate from 7\% to 18\% on Sonnet-4.6. The gains are additive on Sonnet-4.6 (12pp over either alone) but not on Opus-4.7.


Improving the Efficiency of Language Agent Teams with Adaptive Task Graphs

Elizabeth Mieczkowski ⋅ Alexander Ku ⋅ Tiwalayo Eisape ⋅ Dilip Arumugam ⋅ John Matters ⋅ Katie Collins ⋅ Ilia Sucholutsky ⋅ Tom Griffiths

Large language models (LLMs) are increasingly deployed in teams, yet existing coordination approaches often occupy two extremes. Highly structured methods rely on fixed roles, pipelines, or task decompositions assigned a priori. In contrast, fully unstructured teams enable adaptability and exploration but suffer from inefficiencies such as error propagation, inter-agent conflicts, and wasted resources (measured in time, tokens, or file operations). We introduce Language Agent Teams for Task Evolution (LATTE), a framework for coordinating LLM teams inspired by distributed systems, where processors must operate under partial observability and communication constraints. In LATTE, a team of agents collaboratively construct and maintain a shared, evolving coordination graph which encodes sub-task dependencies, individual agent assignment, and the current state of sub-task progress. This protocol maintains consistency while empowering agents to dynamically allocate work, adapt coordination, and discover new tasks. Across multiple collaborative tasks and a variety of base models, we demonstrate how LATTE reduces token usage, wall-clock time, communication, and coordination failures (e.g. file conflicts and redundant outputs) while matching or exceeding the accuracy of standard designs including MetaGPT, decentralized teams, top-down leader-worker hierarchies, and static decompositions.

Clinical case search aims to retrieve disease-consistent prior cases from electronic health record databases for a natural-language patient query under a retrospective retrieval objective. The task is challenging because patient queries contain heterogeneous evidence, including symptoms, diagnoses, laboratory results, treatments, and timelines. These fields are often only partially specified, and fixed retrieval pipelines cannot adapt from prior search trajectories. Existing clinical retrieval methods typically flatten the query into a single representation, use static retrieval pipelines, or lack an explicit mechanism for cross-case adaptation. \textbf{PatientSearch-SE} addresses retrospective disease-centric clinical case search with three components: a \emph{sufficiency-aware planner} that separates dimensions suitable for direct matching from those requiring hypothesis completion, a \emph{hierarchical tool orchestration} module that maps these states to retrieval and re-ranking actions, and a \emph{trajectory-guided memory consolidation} mechanism that writes reusable disease cards and strategy templates without parameter updates. We also propose an evaluation framework that combines disease-centric retrieval outcomes with trajectory-level process probes for retrieval behavior analysis. Experiments on PMC-Patients and MIMIC-IV show higher disease-centric retrieval metrics for PatientSearch-SE than for the compared same-interface sparse, dense, and reasoning-based baselines, with larger gains on PMC-Patients and more modest gains on MIMIC-IV. Its memory-accumulation results are consistent with controlled external-memory adaptation, and its retrieved cases provide additional evidence for diagnosis-masked downstream few-shot RAG prediction under this retrospective protocol.


In-context Learning in Presence of Spurious Correlations

Hrayr Harutyunyan ⋅ Rafayel Darbinyan ⋅ Samvel Karapetyan ⋅ Hrant Khachatrian

Large language models exhibit a remarkable capacity for in-context learning, where they learn to solve tasks given a few examples. Recent work has shown that transformers can be trained to perform simple regression tasks in-context. This work explores the possibility of training an in-context learner for classification tasks involving spurious features. We find that the conventional approach of training in-context learners is susceptible to spurious features. Moreover, when the meta-training dataset includes instances of only one task, the conventional approach leads to in-weights learning and fails to produce a model that leverages context for predictions. Based on these observations, we propose a novel technique to train such a learner for a given classification task. Remarkably, this in-context learner matches and sometimes outperforms strong methods like ERM and GroupDRO. However, unlike these algorithms, it does not generalize well to other tasks. We show that it is possible to obtain an in-context learner that generalizes to unseen tasks by training on a diverse dataset of synthetic in-context learning instances.


Indications of Belief-Guided Agency and Meta-Cognitive Monitoring in Large Language Models

Noam Steinmetz Yalon ⋅ Ariel Goldstein ⋅ Liad Mudrik ⋅ Mor Geva

Rapid advancements in large language models (LLMs) have sparked the question whether these models possess some form of consciousness. To tackle this challenge, Butlin et al. introduced a list of indicators for consciousness in artificial systems based on neuroscientific theories. In this work, we evaluate a key indicator from this list, called HOT-3, which tests for agency guided by a general belief-formation and action selection system that updates beliefs based on meta-cognitive monitoring. We view beliefs as representations in the model's latent space that emerge in response to a given input, and introduce a metric to quantify their dominance during generation. Analyzing the dynamics between competing beliefs across models and tasks reveals three key findings: (1) external inputs systematically modulate internal belief formation, (2) belief formation causally drives the model's action selection, and (3) models can monitor and report their own belief states and adjust behavior in response to internal conflict. Together, these results provide empirical support for the existence of belief-guided agency and meta-cognitive monitoring in LLMs. More broadly, our work lays methodological groundwork for investigating the emergence of agency, beliefs, and meta-cognition in LLMs.

Effective data valuation is essential for machine learning, especially for CLIP pretraining, where models are trained on massive and inherently noisy image-text datasets. Existing CLIP data selection methods often combine heuristic signals such as alignment, target similarity, and diversity, but lack a unified principle for estimating sample contribution. This can lead to inconsistent performance across data regimes and obscures how each sample contributes to downstream performance. In this work, we propose InfCLIP, a principled influence-based approach for data selection in CLIP pretraining. InfCLIP estimates the contribution of each training sample to a target task by leveraging feature representations and temperature-rescaled softmax distributions, enabling efficient computation without retraining or access to model internals. This formulation provides a unified view of alignment, informativeness, uniqueness, and target relevance within a single framework. Empirically, InfCLIP consistently outperforms existing data selection methods across a wide range of benchmarks, including zero-shot evaluation on ImageNet and distribution-shifted datasets. We also show that Self-InfCLIP enables efficient training-set analysis, including noise detection and memorization-aware data assessment, without additional retraining.

Zero-shot object goal navigation (ObjectNav) requires an embodied agent to locate a target object in an unseen environment without any task-specific training. Recent methods leverage large vision-language models (VLMs) to inject semantic priors, but still suffer from two structural issues: (i) semantic cues are treated uniformly across space, ignoring how the spatial extent and hierarchy of contextual entities modulate target likelihood; and (ii) semantic exploitation and geometric exploration are decoupled, producing brittle behavior whenever semantic cues are weak, conflicting, or absent. We address these issues by reformulating navigation as a unified value estimation problem over frontier candidates. Specifically, we propose InfoNav, which jointly models (1) a hierarchical semantic value map that instantiates a spatially-weighted total-probability decomposition of target likelihood through a multi-level spatial influence field; (2) a VLM-based estimator that approximates the intractable long-horizon information gain by grounding it in visual-semantic connectivity; and (3) a memory-augmented verification mechanism that replaces fixed confidence thresholds with a temporally adaptive schedule with deferred re-examination of uncertain detections. InfoNav establishes the highest zero-shot success rate on all three major ObjectNav benchmarks: 62.0% on HM3Dv1 (+0.6 over BeliefMapNav), 81.5% on HM3Dv2 (the only method above 80%), and 41.8% on MP3D (+4.5 over BeliefMapNav), while requiring more than 96% fewer multimodal-model calls than representative LLM-driven baselines. Extensive ablations confirm that the semantic and information-theoretic value fields are complementary, and that the unified formulation is what enables the agent to remain decisive under strong cues and exploratory under weak ones.


Information bottleneck dynamics during learning across artificial and biological neural systems

Nikita Pospelov ⋅ Olga Ivashkina ⋅ Plusnin Viktor ⋅ Olga Rogozhnikova ⋅ Anna Ivanova ⋅ Ksenia Toropova ⋅ Konstantin Anokhin

The information bottleneck offers a candidate general framework for learning, defined by two axes: how much a representation preserves about the external world and how much it carries about the relevant target. However, empirical evidence for information-plane dynamics in realistic-scale artificial networks and in biological systems remains limited. We address this by tracking information-plane trajectories during learning in three systems with different substrates, using a matched analytical pipeline. First, in a gradient-trained spiking ResNet-18 on CIFAR-10 --- a biologically plausible artificial substrate --- where a two-phase fitting-then-compression trajectory emerges specifically in the deep layers; the first information-plane analysis of a spiking network at this scale. Second, on the open-source dataset of mouse visual cortex during multi-week category learning, where our pipeline recovers the published cohort-level effect now reformulated in IB-plane terms. Third, in our own single-photon miniscope calcium imaging data of mouse hippocampus during multi-day place learning in open field, where the IB axes map onto the well-established egocentric-to-allocentric transformation: the input axis becomes egocentric sensorimotor state, the output axis becomes allocentric position. Behavioral and decoder-baseline controls confirm the hippocampal trajectory reflects experience-driven learning. Across all three systems, we observe coherent motion along the relevance axis with learning, while compression emerges only where task structure demands it --- most pronounced in the SNN. The information bottleneck plane therefore offers a substrate-independent coordinate for representation learning, even as the strength of compression along it remains task-dependent rather than universal.


Information Discernment in Large Language Models

Joshua Ashkinaze ⋅ Laura Kurek ⋅ Alina Faisal ⋅ Tongyuan Miao ⋅ Mariam Joseph ⋅ Ceren Budak ⋅ Eric Gilbert

LLMs are increasingly used with external knowledge sources like the Internet. Do they weigh information appropriately---updating more for reliable sources (source discernment) and more when claims bring priors closer to the truth (truth discernment)? We formalize this as information discernment and introduce Learn2Discern (L2D), an experimental framework and benchmark grounded in three normative axioms with interpretable metrics. To establish external validity, a pre-registered, quota-matched user study (n=299) confirms that real LLM users endorse all three axioms and report that violations reduce their trust and usage intent. Across 13 models and nearly 670K trials, we find consistent failures across both dimensions: models perform near chance on source and truth discernment, rely on source popularity twice as much as source reliability, and update roughly equally whether a claim improves or worsens their position relative to the ground truth. Models integrate external knowledge most effectively on datasets where their priors are already the most accurate. Newer and larger models improve truth discernment but not source discernment, a blind spot that model complexity does not address. We identify simple inference-time interventions that meaningfully improve both forms of discernment. We release our dataset, metrics, and survey as a testbed for a core alignment property that scales in importance as LLMs replace traditional search.


Instance-Optimal Estimation with Multiple LLM Judges on a Budget

Junghyun Lee ⋅ Sanghwa Kim ⋅ Yassir Jedra ⋅ Alexandre Proutiere ⋅ Se-Young Yun

Evaluating large language models increasingly relies on LLM-as-a-judge protocols, but such evaluations remain costly: different judges have different prices and reliabilities, and the difficulty of each prompt--response pair can vary substantially. This raises a basic allocation question: under a fixed budget, how should one distribute evaluation queries across heterogeneous judges and instances to obtain the most accurate score estimates? We formalize this question as \emph{budgeted heteroskedastic multi-judge estimation}. Given $K$ prompt--response pairs, $J$ judges with known costs, and unknown query--judge variances, the goal is to estimate a bounded score vector while minimizing an $\ell_p$-error. Our first contribution is to analyze the inverse-variance weighted estimator (IVWE) and to derive the oracle allocation that minimizes its error rate. Since this allocation depends on the unknown variances, we then address the practical unknown-variance setting by proposing Est-IVWE, an adaptive algorithm that constructs and leverages *optimistically biased* variance estimates to stabilize the empirical allocation. We prove that Est-IVWE matches the oracle IVWE rate up to lower-order terms in the budget. Our second and central theoretical contribution is a matching *local* minimax lower bound, which establishes the instance-optimality of the proposed algorithms. A key technical insight is that Fano-type high-probability arguments are too coarse for this problem: their packing construction loses the local variance structure that governs the optimal allocation. We instead use an Assouad-type in-expectation argument, based on local perturbations, which preserves this structure and yields the sharp allocation-dependent lower bound. Finally, we validate our adaptive approach on synthetic benchmarks and the real-world HelpSteer2 dataset, demonstrating significant gains over naive uniform allocation strategies.


Instruct-Particulate: Scaling Feed-Forward 3D Object Articulation with Kinematic Control

Ruining Li ⋅ Yuxin Yao ⋅ Matt Zhou ⋅ Chuanxia Zheng ⋅ Christian Rupprecht ⋅ Joan Lasenby ⋅ Shangzhe Wu ⋅ Andrea Vedaldi

Reconstructing articulated 3D objects is important for animation, gaming, and robotic simulations. Recent neural networks can estimate the articulated structure of 3D objects, but their generalization remains limited by the scarcity of annotated data for this task. To address this gap, we introduce Instruct-Particulate, a model that takes a 3D mesh together with a target kinematic specification, including part descriptions, connectivity, joint types, and optional point prompts, and predicts the corresponding kinematic part segmentation and joint motion parameters. The kinematic specification disambiguates the task and allows the model to target annotations of different granularity, thereby making it possible to use more abundant heterogeneous training data. At test time, the kinematic specification can be obtained automatically, so the model can be applied to any input mesh. To train our model at scale, we construct a heterogeneous dataset of more than 150,000 articulated 3D objects, extending existing publicly available collections with data obtained by partially labelling other 3D models, monolithic or already decomposed into parts, with kinematic labels by means of vision-language models. Experiments show that our model generalizes better across categories and to AI-generated meshes, enabling articulated asset reconstruction from real-world images via image-to-3D models.


Integrating Background Knowledge for Scalable Causal Discovery

Mátyás Schubert ⋅ Theofanis Aslanidis ⋅ Tom Claassen ⋅ Sara Magliacane

Expert background knowledge is often available in practical applications of causal discovery. Such constraints on the true causal graph can help causal discovery in terms of identifiability of causal effects and accuracy of the learned structure, but also in reducing the space of candidate causal graphs. As causal discovery can become computationally expensive for large number of variables, it is crucial to utilize background knowledge effectively \emph{during} the causal discovery process. However, most current methods only use background knowledge in a postprocessing step after causal discovery to refine the learned graph. In this work, we develop a framework for utilizing background knowledge during the causal discovery process, focusing especially on scalable causal discovery methods that recover only a subset of the whole graph. We implement our framework for multiple algorithms and empirically show that utilizing background knowledge can both reduce computational requirements and increase the quality of the learned structures.


Integrating digital twins with randomized experiments

Yanping Li ⋅ Xinwei Ma ⋅ Jingshen Wang

Randomized controlled trials (RCTs) identify treatment effects for an enrolled trial population by assigning treatment independently of potential outcomes, conditional on baseline covariates. These trial-specific effects may not generalize to a broader target population when the target and trial covariate distributions differ. This problem becomes harder when individual-level target population data are unavailable. Pre-trained large language models (LLMs) offer one way to address this data limitation by generating digital twins (DTs) with synthetic covariates intended to resemble the target population. Synthetic covariates alone, however, are insufficient, because neither the RCT covariate distribution nor the DT covariate distribution necessarily matches the target covariate distribution. We therefore propose a statistical inference procedure that integrates RCTs with calibrated DTs using external target population summaries. The procedure represents the target covariate distribution as a calibrated mixture of the RCT and DT covariate distributions, allowing DTs to contribute covariate information while limiting the influence of poorly calibrated synthetic data. The procedure also incorporates LLM-generated auxiliary outcome predictions through calibrated outcome regressions, improving precision without changing the estimand. Theoretical results and simulation studies show that the proposed estimators reduce bias relative to RCT-only estimation and improve efficiency for overall and subgroup target treatment effects.


Interactive Combinatorial Reinforcement Learning for Knowledge Graph Reasoning

Jun Nie ⋅ Yonggang Zhang ⋅ Tongliang Liu ⋅ Chengqi Zhang ⋅ Xinmei Tian ⋅ Bo Han

Knowledge Graph Question Answering (KGQA) increasingly requires multi-turn interaction with knowledge graphs (KGs) to derive final answers. Such interactive capabilities often rely on proprietary large-scale LLMs, which hinders cost-efficient and local deployment. Reinforcement learning offers a scalable alternative for endowing smaller open-source LLMs with multi-turn interaction ability, as it enables models to improve from self-explored interaction trajectories rather than relying on expensive expert demonstrations. However, KGQA presents a distinctive challenge for RL: unlike tasks where exploration can proceed in a relatively unconstrained textual space, KG reasoning is governed by a sparse graph topology. Free-form policy rollouts frequently generate relation sequences that cannot be executed on the graph, making reward signals sparse and unstable. We propose ICOR, an Interactive Combinatorial Reinforcement Learning framework that grounds policy exploration in the combinatorial structure of KGs. ICOR first retrieves a query-specific set of candidate relations and then restricts policy rollouts to relation-path compositions within this graph-derived space. This design increases the likelihood that sampled trajectories are executable and thus informative for policy optimization. Building on this constrained rollout space, ICOR applies GRPO with a staged feedback mechanism that guides learning from structural validity to answer-consistent reasoning. Experiments on multiple KGQA datasets demonstrate the effectiveness of ICOR.


interwhen: A Generalizable Framework for Steering Reasoning Models with Test-time Verification

Vishak K Bhat ⋅ Prateek Chanda ⋅ Vijval Ekbote ⋅ Ashmit Khandelwal ⋅ Maitreyi Swaroop ⋅ Subbarao Kambhampati ⋅ Vineeth N Balasubramanian ⋅ Nagarajan Natarajan ⋅ Amit Sharma

Reasoning models produce long traces of intermediate decisions and tool calls, making test-time verification increasingly important for ensuring correctness. Existing approaches either verify only the final answer, which misses early errors, or rely on branch-and-verify strategies that explore multiple trajectories at substantially higher compute cost. We introduce interwhen, a single-trajectory verification framework that steers model behavior by providing feedback on intermediate reasoning traces. It addresses two key challenges. First, given a set of verifiers, obtaining verifiable states from the reasoning trace typically requires prompt engineering or external task decomposition into fixed steps, which can constrain the model's reasoning strategy. Instead, we propose a monitoring system that periodically polls the reasoning trace and forks inference of the reasoning model to recover intermediate states. Verifiers are run asynchronously alongside generation, adding negligible overhead on correct executions and intervening only when violations occur. Second, beyond math and code, a central challenge for process verification is the scarcity of verifiers. interwhen addresses this through automatic verifier synthesis from natural-language policy documents. Given a policy, it can generate code-based verifiers, including provably correct verifiers in Lean and Z3. Together, these contributions yield a plug-and-play system for policy-grounded, formally verified process supervision of any reasoning agent. On reasoning benchmarks where policies encode mathematical or logical constraints, interwhen-based steering achieves near-perfect accuracy for reasoning models using a fraction of the token compute of test-time verification baselines. On agentic benchmarks with policy-based verifier generation, it enables significant improvements in task quality for SLMs without any finetuning; e.g., the task completion rate of Qwen3-30B jumps from 32% to 87% on the telecom domain in tau^2-bench.


Intrinsic Selection and Particle Resampling for Inference-Time Scaling Beyond Domain Verifiability

Giorgio Giannone ⋅ Mustafa Eyceoz ⋅ Shabana Baig ⋅ Shivchander Sudalairaj ⋅ Anna C Doris ⋅ Faez Ahmed ⋅ Akash Srivastava ⋅ Kai Xu

Inference-Time Scaling (ITS) has largely succeeded in verifiable domains like math and coding, where cheap verification enables scalable output selection. However, extending ITS to tasks prone to systematic failure - driven by faulty initial assumptions or unmet multidimensional constraints - typically relies on costly external solvers or brittle, model-based verifiers. Our key insight is that the intrinsic statistics of parallel sample sets, specifically length-adjusted tail entropy, provide a robust discriminative signal for solution quality without access to ground truth. Crucially, these statistics serve as a difficulty gate for adaptive compute allocation, dynamically routing problems across scaling regimes. First, Intrinsic Selection (iS) ranks candidates post-hoc, matching consensus-based algorithms across three domains and improving engineering design selection by 20 % over pass@1 baselines. Second, Intrinsic Particle Filtering (iPF) generalizes this to step-level resampling, guiding generation toward high-confidence reasoning trajectories to improve pass@1 by 6.1 points on average on hard math problems. Finally, Particle Distillation (dPF) injects privileged guidance via early logit blending and KL-guided resampling, steering generation past systematic reasoning errors to satisfy expert rubrics, yielding up to 26.5 % gains on complex clinical responses. Our pipeline applies seamlessly across broad-purpose, domain-specialized, and multimodal architectures, successfully extending ITS to open-ended domains without requiring trained reward models or exact verification.

We study label-free, topology-conditioned zero-shot cross-graph link prediction: at test time the observed target topology is the input graph, but the model receives no target labels, no positive/negative context links, no target fine-tuning, no text attributes, and no post-hoc alignment. This setting is zero-shot with respect to target supervision and adaptation, but it is not topology-free inference. We identify a critical geometric barrier for hyperbolic transfer: radial non-identifiability. Hyperbolic graph learning is motivated by the intuition that radius encodes hierarchy or popularity and angle encodes similarity, yet standard HGNN objectives do not make radii semantically calibrated across disjoint graphs. The same structural role can therefore occupy incompatible radial scales on different target graphs. We propose Invariant Hyperbolic Unfolding (IHU), which restores the intended radial hierarchy channel through selective invariance: radii are fixed by graph-internal structural percentiles while angular representations remain learnable. Its core module, Invariant Structural Anchoring (ISA), rank-canonicalizes structural scores such as coreness into shared hyperbolic radii, producing a comparable radial skeleton without target supervision. In a controlled comparison among matched frozen learned encoders, IHU improves average HR@50 by 2.7 percentage points over the strongest learned hyperbolic baseline and by 7.6 percentage points over the mean of the learned hyperbolic baseline family. The gain amplifies under missing-edge and low-degree regimes, while hierarchy-gap analysis reveals a boundary condition for fixed radial shells: 50% edge sparsification causes 52.5% less degradation, and low-degree nodes show a 4.2× larger advantage. We further position IHU against recent universal link prediction and graph foundation methods through an access-aware taxonomy, showing that the closest recent LP methods use labeled target context links, whereas IHU isolates the geometric effect of fixed versus learned radii in frozen structural hyperbolic encoders.


Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench

Qingyun Zou ⋅ Feng Yu ⋅ Hongshi Tan ⋅ Jiahao Cui ⋅ Bingsheng He ⋅ Weng-Fai Wong

We ask whether agentic AI systems built for software engineering transfer to realistic hardware engineering. Existing hardware LLM benchmarks isolate sub-tasks but none jointly requires repository navigation, hierarchy-aware localization, Electronic Design Automation (EDA) executable verification, and maintenance-style patching. We introduce **Phoenix-bench**, a synchronized corpus of 511 verified Verilator instances from 114 GitHub repositories, each shipped with the developer patch, design-flow labels, fail-to-pass and pass-to-pass testbenches, and a Docker-pinned EDA environment so resolved-rate differences reflect agent behavior rather than toolchain availability. Using Phoenix-bench we run a uniform evaluation of four commerical agents and eight open-source agentic structures across four LLM backbones, plus two diagnostic interventions (file-level oracle localization and one round of testbench-log feedback). Three findings emerge. **Software and hardware are fundamentally different engineering tasks:** the same agent loses 37\% to 58\% from SWE-bench Verified to Phoenix-bench because hardware bugs propagate across parallel instantiated modules through signal flow rather than along a software-style call graph, and software-tuned agents stop at the symptom file instead of tracing back through the instantiation chain. **Failures concentrate on design control-flow / Finite State Machine bugs, verification testbench bugs, and hard cases** that demand cross-hierarchy signal-flow tracking and coordinated multi-file edits. **Localization granularity matters far more than localization itself:** a perfect file-level oracle yields only $+1.4$\% because the agent then breaks files that did not need editing, while a single round of test case feedback lifts resolved rate by $42$\% to $45$\% because the test case tells *where* the bug is and *what* the fix has to look like.

Introduced is a methodology for adapting the topology of dense neural networks, enabled by isotropic activation functions. Achieved through prescribed reparameterisation symmetries and singular-value decomposition of affine maps, this diagonalises layers into one-to-one, ordered connections. This makes it simpler to assess the impact of individual connections on the function. Low-impact neurons can be removed (neurodegeneration), and a thresholded buffer of largely inactive 'scaffold' neurons is maintained (neurogenesis). These symmetry-led diagonalisation and structural changes are function-invariant, demonstrated to be computationally identical during neurogenesis, arbitrarily well approximated during neurodegeneration, and enable asymptotic 50% parameter sparsification of dense networks with identically preserved function. Thus, real-time restructuring of the architecture in response to task demands, task appending, removal or changes is shown. The approach is conceptually centred on primitive symmetry-prescriptions, through which isotropic functions are derived that feature explicit basis independence and a loss in the individuation of neurons implicit in typical elementwise functional forms. Hence, this allows freedom in the basis to which layers are decomposed and interpreted as individual artificial neurons, directly enabling this adaptive topology approach. Additionally, a new tunable model parameter, the 'intrinsic length', is introduced to improve this analytical invariance, alongside a generalised isotropic-perceptron architecture that enables parallel precomputation of all matrix-vector products and displays a nested functional class. Diagonalisation is suggested to offer new possibilities for interpretability and monitoring of isotropic networks.


Iterative ILP with Update-Size Control for Reducing Surrogate-Task Mismatch in Bit-Width Selection

Shinya Gongyo ⋅ Ryosuke Ogasawara ⋅ Tatsuya Moe ⋅ Masafumi Mori ⋅ Yusuke Sekikawa ⋅ Mitsuru Ambai

Mixed-precision quantization improves the accuracy-efficiency trade-off by assigning different bit-widths across a network. Since bit-width selection is a combinatorial problem, existing methods often optimize additive surrogate losses instead of directly evaluating task losses in all configurations. We show that these surrogates can suffer from surrogate-task mismatch, especially when many layers are changed simultaneously from a reference configuration. We further observe that, although allowing more layers to change improves the surrogate objective, the evaluated task loss can be minimized at an intermediate limit on the number of changed layers. Based on this observation, we propose Iterative Integer Linear Programming (I2LP), which iteratively refines a reference bit configuration. At each refinement step, I2LP solves multiple ILPs with different limits on how many layers may change from the current reference, and accepts the best candidate only when it reduces the task loss. I2LP applies to both post-training quantization and quantization-aware training, and consistently outperforms uniform-precision baselines and existing mixed-precision methods. Code will be released.


Jump Start Your Policy Learning with Lessons from 145,000 Training Runs

Nabil Omi ⋅ Chung Yik Edward Yeung ⋅ Eric Bae ⋅ Siddhartha Sen ⋅ Ali Farhadi

Reliable progress in offline policy learning depends on a small set of methodological practices, including careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has noted that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties independently, but these sources of variability have not been systematically investigated at the scale needed to understand how they shape conclusions. To address this gap, we present a large-scale empirical study of offline reinforcement and imitation learning, training over 145,000 policies across 114 datasets. At this scale, no algorithm dominates: aggregate performance across top methods is often close, but the leaders differ substantially across environments. We find that proper hyperparameter tuning frequently reshuffles perceived algorithm rankings; and simple baselines, including behavior cloning, are often stronger than is commonly assumed after extensive tuning. We also study hyperparameter transfer and sensitivity across environments, identifying a simple strategy that generalizes well. From these analyses we distill practical recommendations, and release JumpStart: a resource suite of trained models, per-model hyperparameter and reward data, strong baselines across all environments, and a website to make retrieval and analysis trivial. Together, these resources aim to make offline policy learning research more reliable and to open new directions for work beyond the scope of this study.

The Maximum Mean Discrepancy (MMD) is a cornerstone statistic for nonparametric two-sample testing, but its test power is dictated entirely by the chosen kernel. Because any fixed kernel inherently fails to distinguish certain distributions, the kernel must be dynamically optimised. However, data-driven optimisation violates the foundational i.i.d. assumption, forcing a strict trade-off in existing frameworks. Ratio criteria ignore this dependence, inducing overfitting and variance collapse on rich kernel classes. Conversely, aggregation methods bypass the dependence using finite grids, but this strategy cannot scale to continuous search spaces like deep kernels. To break this dichotomy, we establish data-driven kernel selection as a model selection problem. We propose Complexity-Penalised MMD (CP-MMD), a criterion derived from a uniform concentration inequality (extending Maurer's framework to the two-sample setting) that bounds the post-optimisation distribution by penalising the empirical MMD with the complexity of the kernel search space. Because this penalty mathematically absorbs the cost of optimisation, CP-MMD enables direct, grid-free maximisation over continuous parametric classes, including scalar bandwidths, polynomial coefficients, and deep network parameters. By formally accounting for optimisation complexity, we theoretically guarantee that CP-MMD maximises true test power while ensuring unconditional Type-I validity. Consequently, CP-MMD enables grid-free kernel selection across linear, polynomial, and deep regimes, matching or exceeding state-of-the-art test power.


Kernel Value Regression in Offline Reinforcement Learning

Jongyeon Lee ⋅ Jaehyoung Jeon ⋅ Kim Haneol ⋅ Myungjoo Kang

Value overestimation is a central challenge in offline reinforcement learning (RL), arising when Bellman backups propagate out-of-distribution (OOD) errors that lead to catastrophic inflation of $Q$-values. Prior approaches address this issue by encouraging the learned policy to stay close to the behavior policy, but they often result in suboptimal policies or require significant training complexity. More recent work attempts to tackle this issue through minimal modifications to the critic network, yet it remains limited in either theoretical grounding or practical implementation. In this work, we present Kernel Value Regression (KVR), a simple critic regularization method based on kernel regression. KVR builds on the vanishing extrapolation property of kernel ridge regression, which encourages value estimates to decay toward zero away from the data support. This inductive bias naturally suppresses $Q$-values for OOD actions without auxiliary regularizers, providing a structural mechanism for addressing value overestimation. Using random Fourier features, KVR can be incorporated into standard actor-critic algorithms with only minimal modifications. We empirically demonstrate that these properties persist in complex offline RL environments, leading to strong performance across diverse tasks in the D4RL and OGBench benchmarks.

Nearest neighbor search algorithms are widely used and have a long history, but we find that key implementation details are seldom reviewed in articles about them. These details can materially impact performance, and thus change the conclusions one makes about a given algorithm's superiority. In this article, we highlight these impacts by reviewing the cover tree as an example, producing RustKNN, a unified Rust implementation of four variants of the cover tree. Our analysis reveals: (1) the optimal query strategy depends critically on $k$, and can have up to $111\times$ factor impact on runtime; (2) the cover tree base parameter, set to 2 in theory and 1.3 in practice, tends to have an optimal value in $[1.1, 2.0]$ which can be 20\% faster than other settings for a given dataset; (3) micro-optimizations (early-exit distance, packing, triangle pre-filters) compose with early-exit distance being very powerful in high dimensions; and (4) all-nearest-neighbor and held-out benchmarks yield different implementation rankings. We release the implementation with reproducible benchmark infrastructure.


Knowing When Multivariate Forecasts Are Wrong

Binli Luo ⋅ Ning Gui ⋅ Xianhan Tan

At deployment, a multivariate forecaster may expose only a recent input window and a frozen point prediction, with no access to ground truth, calibration residuals, repeated inference, or model internals. We study deployment-constrained forecast risk localization: ranking future timesteps by likely error before they are observed. We introduce SubDx, a zero-parameter diagnostic that audits the forecast against the local cross-channel subspace of the input history. Each predicted channel vector receives an off-subspace residual score, computed with no labels, no learned parameters, and less than 1% runtime overhead on Traffic. Under a linear factor model, we derive closed-form score distributions and a two-population AUROC expression showing how heterogeneous off-subspace error visibility makes the ranking identifiable. On Traffic (N=862), SubDx reaches within-horizon AUROC 0.90 and reduces MSE by 62% at 50% coverage; on Solar (N=137), it reaches AUROC 0.85 with 67% MSE reduction. The same post-hoc score applies to frozen pretrained forecasters, with Chronos reaching AUROC 0.838 on Traffic without fine-tuning. Across trained backbones, pretrained models, and synthetic controls, the results show that local cross-channel consistency can provide a practical abstention signal for redundant multivariate systems.


KnowVis: A Dual-View Benchmark for Diagnosing World-Knowledge Grounding in Text-to-Image Models

Yonghang Yu ⋅ Ruoyu Wang ⋅ Haoyu Zhang ⋅ Lin Qi ⋅ Yawei Zhou ⋅ Changqian Yu

Text-to-image (T2I) models have made substantial progress in visual realism, aesthetic quality, and instruction following. However, real-world prompts often go beyond explicit visual descriptions and require implicit facts, structured knowledge, and domain-specific commonsense. Existing evaluations mainly focus on explicit prompt-to-image semantic alignment or specific knowledge subdomains, lacking a unified and diagnostic framework for evaluating world-knowledge grounding in T2I models. We introduce KnowVis, a dual-view diagnostic benchmark for world-knowledge grounding in T2I generation. KnowVis is built on a hierarchical knowledge taxonomy with case-level annotations, and contains two complementary views: a Skill-Tree Benchmark for fine-grained atomic knowledge evaluation and a Generalized Benchmark for open-ended knowledge composition and expression. We further design a structured MLLM-based VQA judge protocol and validate its reliability on the Skill-Tree validation split, enabling scalable evaluation and diagnostic error attribution. The protocol uses Required VQAs to assess atomic knowledge correctness in the Skill-Tree Benchmark, and combines Required, Optional, and dynamically discovered knowledge expressions with human-judge collaborative evaluation in the Generalized Benchmark. Experiments show that KnowVis reveals notable gaps in world-knowledge grounding across existing T2I models and provides structured diagnostic signals for knowledge-intensive image generation.


Label-Efficient Dataset Pruning via Semi-Supervised Pseudo-Labeling

Yeseul Cho ⋅ Baekrok Shin ⋅ Changmin Kang ⋅ Chulhee Yun

Dataset pruning reduces the storage and training costs of deep learning by selecting an informative subset from a large dataset. However, most existing pruning methods require fully labeled data, which limits their applicability in realistic settings where unlabeled data are abundant and annotation is costly. Recent label-free pruning methods address this issue, but they rely on features from pre-trained models to estimate example difficulty. This dependence can be unreliable when the target dataset differs substantially from the pre-training distribution. We propose a label-efficient dataset pruning framework that, given only a small labeled subset, uses semi-supervised learning to generate pseudo-labels for unlabeled data, allowing existing supervised pruning methods that require label information to be seamlessly applied to the resulting pseudo-labeled training pool. We then estimate example difficulty from pseudo-label-induced training dynamics and select a coreset. By learning directly from the target dataset, our method better captures the target distribution and provides more reliable signals for difficulty estimation and coreset selection. We validate our approach on domain-specific, image-corrupted, and long-tailed datasets, where it achieves state-of-the-art performance among label-free and label-efficient baselines, while also demonstrating competitive performance on standard benchmarks.


LAFP: Preserving Latent Action Structure in Latent Policy Learning via Flow Matching

Jiexi Lyu ⋅ Xi zhou Bu ⋅ Qingqiu Huang ⋅ Chufeng Tang ⋅ Xiaoshuai Hao ⋅ Hongbo Wang ⋅ Wei Li

Learning high-quality latent actions from large-scale unlabeled videos, coupled with limited real-world interaction data for training an action decoder, has emerged as a promising paradigm for scalable latent policy learning. However, existing approaches typically rely on behavior cloning, which tends to collapse inherently multimodal action distributions into unimodal ones, thereby degrading the pretrained latent action structure. While flow matching provides a potential alternative, directly applying it leads to a misalignment between latent actions and physical actions during action decoder training, due to the stochastic nature of the learned policy. To address these, we propose Latent Action Flow Policy (LAFP), which leverages flow matching for latent policy learning and introduces an inference-time interpolation mechanism to mitigate stochasticity-induced misalignment. Experimental results demonstrate that LAFP consistently outperforms prior methods on downstream imitation learning tasks, achieving up to 10–15\% improvement in success rate while incurring less than 1× additional inference overhead.

This paper argues that calibration should become a primary evaluation and optimization target for Large Language Models (LLMs), because diverse failures, from confident hallucinations to collapsed diversity to brittle safety refusals, are best understood as different facets of a single underlying problem: miscalibration. We propose a unified framework organized around four types of miscalibration: probabilistic (the model cannot match requested probability distributions), semantic (confidence is misaligned with factual correctness), distributional (output diversity collapses to stereotypical modes), and metacognitive (the model fails to assess its own competence). We argue that foregrounding calibration as a diagnostic lens and evaluation target is essential for building models that are not merely capable, but genuinely reliable. We call for calibration metrics to become a standard part of benchmark reporting, for training paradigms that preserve uncertainty, and for interaction designs that surface model confidence to downstream decision-makers.


LaSA-Net: A Language-Guided Network for Outdoor Generalized 3D Referring Expression Segmentation

Lingfei Ma ⋅ Bin Liu ⋅ Wen Li ⋅ Wentao Sun ⋅ Chaorui Liu ⋅ Haiyan Guan ⋅ Jonathan Jun LI

3D Referring Expression Segmentation (3D-RES) aims to segment target objects in point clouds via language. However, extending 3D-RES from indoor scenes to large-scale outdoor environments introduces three major challenges, i.e., large-scale geospatial complexity lacking in indoor benchmarks, sparse-target perception overwhelmed by massive geometric clutter and distractors, and context-induced query dilution, where highly relation-intensive expressions cause globally-initialized queries to be absorbed by salient reference objects. To address these issues, we introduce UrbanRefer, the first outdoor 3D-GRES benchmark with 195 urban road point cloud scenes and 7,840 textual descriptions for complex geospatial reasoning. We further propose LaSA-Net, a Language-guided Semantic Segmentation Network. First, to enhance sparse-target perception, LaSA-Net incorporates pixel-level dense features from DINOv3 as 2D priors and fuses them with 3D features through depth-consistent verification. Then, to alleviate query dilution, we design a semantic-adaptive query generation module that applies text-weighted farthest point sampling to focus on text-relevant regions, and performs multi-granularity sampling to maintain both local grounding and global semantic consistency. Finally, a reliability-aware decoder adaptively suppresses noisy language cues through cross-modal reliability-gated injection, generating precise prompts for accurate mask prediction. Extensive experiments demonstrate that LaSA-Net achieves 52.7\%, 55.9\%, and 47.9\% mIoU on the ScanRefer, Multi3DRefer, and UrbanRefer datasets, surpassing prior state-of-the-art methods by 2.3\%, 4.2\%, and 2.8\% respectively.

Asynchronous Partial Rollouts (APR) accelerate RLHF by training on the first $k$ of $N$ parallel completions, but this speed hides a critical flaw: because autoregressive generation time scales linearly with token count, a wall-clock cutoff acts as an invisible length filter. On reasoning tasks where reward correlates positively with trajectory length, APR systematically deletes the most valuable, multi-step reasoning chains from the training batch. We formalize this phenomenon as **Latency-Conditioned Selection Bias** (LCSB). We prove that whenever reward and selection probability negatively correlate ($\mathrm{Cov}_{p_\theta}(R, \alpha) < 0$), the APR objective structurally attenuates and can even invert the on-policy gradient. Crucially, standard importance-sampling corrections (e.g., V-trace) cannot resolve this because all trajectories strictly originate from the behavior policy. After empirically validating this latency-length filter in a deployed vLLM system and demonstrating catastrophic, monotonic performance degradation on reasoning benchmarks, we provide two actionable solutions. We introduce **Effective Gradient Information** (EGI), a diagnostic metric that detects LCSB thousands of steps before reward curves diverge, and **Selection-Probability Weighting** (SPW), a near-zero-overhead correction that mathematically neutralizes the bias to recover on-policy reasoning performance without sacrificing asynchronous throughput.


Latent Introspection: Models Can Detect Prior Concept Injections

Theia Pearson-Vogel ⋅ Martin Vanek ⋅ Raymond Douglas ⋅ Jan Kulveit

We uncover a latent capacity for introspection in a Qwen 32B model, demonstrating that the model can detect when concepts have been injected into its earlier context and identify which concept was injected. While the model denies injection in sampled outputs, logit lens analysis reveals clear detection signals in the residual stream, which are attenuated in the final layers. Furthermore, the effect is heavily dependent on context – e.g., prompting the model with accurate information about AI introspection mechanisms can dramatically strengthen this effect: the sensitivity to injection increases massively (0.3% → 39.9%) with only a 0.6% increase in false positives. Also, mutual information between nine injected and recovered concepts rises from 0.62 bits to 1.05 bits, ruling out generic noise explanations. Our results demonstrate models can have a surprising capacity for introspection and steering awareness that is easy to overlook, with consequences for latent reasoning and safety.


Latent Reasoning in Continuous Space for Unified Multimodal Models

Byungwoo Jeon ⋅ Yoonwoo Jeong ⋅ Hyunseok Lee ⋅ Minsu Cho ⋅ Jinwoo Shin

Unified Multimodal Models (UMMs) achieve remarkable success on diverse multimodal understanding and generation tasks, yet they still struggle to effectively leverage language-based chain-of-thought reasoning for image generation. As a result, existing reasoning-augmented UMMs often fail to inherit the test-time scaling and fine-grained controllability observed in large language models. We attribute this failure to a fundamental modality mismatch: reasoning is performed in a discrete token space, whereas image generation operates in a continuous space. To bridge this gap, we introduce a Latent Reasoning framework in Continuous space (LARC), which alleviates the token-level bottleneck by reasoning over continuous hidden states. Specifically, the training framework consists of two stages: (i) curriculum supervised fine-tuning (SFT), and (ii) information-gain reinforcement learning (RL). First, the curriculum SFT stage gradually converts token-level reasoning into latent reasoning. Subsequently, the information-gain RL stage self-evolves latent reasoning without ground truth supervision. Experiments across diverse text-to-image generation benchmarks show that LARC consistently improves generation quality, prompt following, and compositional fidelity. Notably, LARC achieves state-of-the-art performance on GenEval with a 7\%p gain over BAGEL and demonstrates clear test-time scaling behavior, outperforming token-based generation under both parallel and sequential test-time scaling strategies.


Latent Spatial Reasoning: Building Innate 3D Awareness via Latent-Space Distillation

Ruifei Zhang ⋅ Xiangru Lin ⋅ Haoyuan Li ⋅ Wei Zhang ⋅ Xiao Tan ⋅ Junlin Xie ⋅ Xiang Wan ⋅ Guanbin Li

Humans effortlessly infer 3D structure, such as depth, occlusion, and spatial arrangement, from 2D images and reason about it fluidly. Multimodal large language models still struggle with such spatial reasoning. Current approaches attempt to bridge this gap by injecting explicit 3D priors at inference time or aligning features with 3D-aware teachers during training. While these strategies improve geometric perception, they typically treat 3D knowledge as static representations, rather than enabling the model to reason with 3D cues during inference. We argue that closing this gap requires internalizing 3D knowledge in a form that can participate in the model's reasoning dynamics. To this end, we introduce \textbf{Latent Spatial Reasoning (LSR)}, a framework that distills geometric knowledge into \emph{latent spatial tokens}: geometry‑aware representations that the model actively queries and conditions on to guide its spatial reasoning process. Through hierarchical distillation, LSR transfers 3D knowledge from foundation models, aligns it with the vision-language space, and trains the model to interleave spatial and linguistic reasoning, all without external 3D modules at inference time. Experiments on diverse spatial reasoning benchmarks demonstrate state-of-the-art performance.


LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model

Jiachun Jin ⋅ Zetong Zhou ⋅ Xiao Yang ⋅ Hao Zhang ⋅ Pengfei Liu ⋅ Jun Zhu ⋅ Zhijie Deng

Unified models (UMs) hold promise for their ability to understand and generate content across heterogeneous modalities. Compared to merely generating visual content, the use of UMs for interleaved cross-modal reasoning is more promising and valuable, e.g., for solving understanding problems that require dense visual thinking, improving visual generation through self-reflection, or modeling visual dynamics of the physical world guided by stepwise action interventions. However, existing UMs necessitate pixel decoding as a bridge due to their disjoint visual representations for understanding and generation, which is both ineffective and computationally cumbersome. In this paper, we introduce LatentUM, a novel unified model that represents all modalities within a shared semantic latent space, eliminating pixel-space mediation for model-internal reasoning over generated visual states. By making generated visual tokens directly interpretable to the model itself, LatentUM enables flexible interleaved cross-modal reasoning and generation. Empirically, LatentUM achieves state-of-the-art performance on the Visual Spatial Planning benchmark, pushes the limits of visual generation through self-reflection and demonstrates world modeling capability by predicting future visual states in the shared semantic latent space.


Latent Video Prediction for World Modeling: An Evaluation Uncovering Intriguing Favorable Evidence

Ali J Alrasheed ⋅ Aryan Y Parast ⋅ Basim Azam ⋅ James Bailey ⋅ NAVEED AKHTAR

Self-supervised video models are increasingly framed as world models, yet their evaluation remains largely confined to a single top-1 accuracy score on clean benchmarks. This leaves a major gap in comprehending their potential as world models. We present the first systematic study addressing this gap, analyzing four matched-capacity frontier video foundation models, V-JEPA 2.1, V-JEPA 2, VideoPrism, and VideoMAEv2, across five robustness axes relevant to their deployment as video world models: feature discriminability, corruption robustness, fine-grained discrimination, occlusion robustness, and sensitivity to temporal direction. Our evaluations establish that across all five axes, latent-prediction models form a distinct and consistent profile. They degrade more gracefully under pixel corruption, preserve usable class structure rather than mere geometric stability under occlusion, capture fine-grained physical contact cues without reconstructing pixels, and uniquely encode the arrow of time. These advantages can even survive task adaptation: a frozen V-JEPA 2 backbone with a lightweight attentive probe outperforms a fully fine-tuned VideoMAE and a supervised TimeSformer on corruption and occlusion robustness. Our extensive results offer concrete new evidence in favor of latent prediction for robust world modeling.


Layerwise LQR for Geometry-Aware Optimization of Deep Networks

Simon Dufort-Labbé ⋅ Pierre-Luc Bacon ⋅ Razvan Pascanu ⋅ Simon Lacoste-Julien ⋅ Aristide Baratin

Geometry-aware optimizers such as Newton and natural gradient can improve conditioning in deep learning, but scalable variants such as K-FAC, Shampoo, and related preconditioners usually impose structural approximations early, often discarding cross-layer interactions induced by the network computation. We introduce Layerwise LQR (LLQR), a framework for learning structured inverse preconditioners under a global layerwise optimal-control objective. The starting point is an exact equivalence: the steepest-descent step under a broad class of divergence-induced quadratic models—including Newton, Gauss–Newton, Fisher/natural-gradient, and intermediate-layer metrics—can be written as a finite-horizon Linear Quadratic Regulator (LQR) problem. This formulation serves as a reference that exposes the layerwise dynamics and cost matrices encoding the original dense geometry. We then derive a scalable relaxation that learns diagonal, (E-)Kronecker-factored, or other structured inverse preconditioners by minimizing the LQR objective and reusing them across iterations. The resulting optimizer wraps standard methods while retaining a principled connection to second-order geometry, without forming or inverting the global curvature matrix. Experiments on ResNets and Transformers show that LLQR improves optimization dynamics and often translates these gains into improved final test performance, while adding only modest wall-clock overhead. It establishes LLQR as a practical framework for geometry-aware second-order methods and a reference for evaluating scalable approximations.


LEAF: Language-EEG Aligned Foundation Model for Brain-Computer Interfaces

Muyun Jiang ⋅ Shuailei Zhang ⋅ Zhenjie Yang ⋅ Wu Mengjun ⋅ Wei Zhang ⋅ Weibang Jiang ⋅ Chenyu Liu ⋅ Zhiwei Guo ⋅ Rui Liu ⋅ Shangen Zhang ⋅ Yong Li ⋅ Yi Ding ⋅ Cuntai Guan

Electroencephalography (EEG) foundation models learn transferable representations for brain--computer interfaces, but existing approaches treat tasks and labels as discrete targets and fail to use the semantic structure of natural-language task instructions to guide representation learning. We present \textbf{LEAF}, a foundation model for \textbf{EEG--Language Alignment with Semantic Task Instruction and Querying}. LEAF integrates task-aware semantic guidance to produce structured and linguistically aligned EEG embeddings, thereby enhancing decoding robustness and transferability. In the EEG pretraining stage, we introduce a joint \textbf{Spectral--Temporal Reconstruction (STR)} framework that captures the coupled spectral rhythms and temporal dynamics of EEG signals. STR applies randomized spectral perturbation to enhance frequency robustness and uses two complementary temporal objectives to learn both contextual and sequential structure. In the EEG-Language alignment stage, we propose the \textbf{Instruction-conditioned Q-Former (IQF)}. This query-based cross-attention transformer injects instruction embeddings into EEG tokens and achieves semantic alignment with textual label embeddings through learnable queries. We evaluate LEAF on 16 downstream datasets spanning motor imagery, emotion recognition, steady-state visual evoked potentials, covert speech, and healthcare tasks. LEAF achieves state-of-the-art performance on 12 of the 16 datasets and obtains the best average results across all five task categories. Importantly, our analyses reveal for the first time that explicit task instructions serve as semantic priors guiding EEG embeddings into coherent and linguistically grounded spaces. Code is available at \url{https://anonymous.4open.science/r/LEAF-Model}


Learned Lagrangian Models of PDEs via Euler–Lagrange Residual Minimization

Lyra Zhornyak ⋅ Eric Forgoston ⋅ M. Ani Hsieh

We present the first method to directly use a learned continuous Lagrangian to forecast the dynamics of systems governed by partial differential equations, exploiting the inherent conservative structure to achieve stable long-range predictions. We develop an optimization-based integrator that minimizes the squared Euler--Lagrange residual via a mesh-free near-symplectic construction on local space-time patches. Different from integrators for analytical models, integrators for learned models should decouple model error (phase error) from integration error (conservation error). By relying on optimization rather than time-stepping, we bypass the global coupling inherent to fixed discretizations, which slows time- and space-stepping and complicates learning. Our method scales linearly with domain size via Jacobi iteration, and places no structural requirements on the learned network, allowing it to be coupled with existing physics-guided machine learning (ML) methods. We validate our approach on a learned representation of a double pendulum, a one-dimensional wave equation, and a two-dimensional wave equation. Our method achieves error comparable to classical symplectic methods while generalizing to spatially varying dynamics and arbitrary boundary conditions without retraining.


Learning Active Perception and Manipulation via Spatio-temporal Visual Memory

Enshen Zhou ⋅ Mengzhen Liu ⋅ Yibo Li ⋅ Yanjun Ding ⋅ Yuheng Ji ⋅ Pengwei Wang ⋅ Zhongyuan Wang ⋅ Lu Sheng ⋅ Shanghang Zhang

Active perception is crucial for robots to interact with unstructured scenes, where a closed perception-action loop is required. In this work, we introduce ActiveZero, an end-to-end framework that formulates active perception as information-driven spatio-temporal memory management. Leveraging memory as a universal information interface for conveying perception and action, it maintains a visual memory bank, actively expands it via purposeful exploration for missing information, retrieves relevant evidence for execution, and filters memory online for efficiency. This formulation not only enjoys long-context reasoning ability but also enables multi-level supervision for heterogeneous data, going beyond prior paradigms. To support it, we construct ActiveMem, an active perception dataset of 4M episodes with camera actions (300× prior) across 49 scenes for long-horizon tasks (up to 4 subtasks). In addition, we present ActiveBench, the first memory-centric benchmark suite for active perception, covering VQA and simulated manipulation. Experiments show that ActiveZero achieves SOTA on existing active-perception-related benchmarks and ActiveBench. It also outperforms all baselines on challenging real-world tasks by a large margin, even surpassing $\pi_{0.5}$ by 34.16%. Notably, ActiveZero exhibits diverse active perception strategies (e.g., search, track, interact) and handles long-horizon tasks in unstructured scenes with a single unified model.


Learning Agentic Policy from Action Guidance

Yuxiang Ji ⋅ Zengbin Wang ⋅ Yong Wang ⋅ Yang ⋅ Ziyu Ma ⋅ Guanhua Chen ⋅ Zonghua Sun ⋅ Liaoni Wu ⋅ Xiangxiang Chu

Agentic reinforcement learning (RL) for Large Language Models (LLMs) critically depends on the exploration capability of the base policy, as training signals emerge only within its in-capability region. For tasks where the base policy cannot reach reward states, additional training or external guidance is needed to recover effective learning signals. Rather than relying on costly iterative supervised fine tuning (SFT), we exploit the abundant action data generated in everyday human interactions. We propose \textsc{ActGuide-RL}, which injects action data as plan-style reference guidance, enabling the agentic policy to overcome reachability barriers to reward states. Guided and unguided rollouts are then jointly optimized via mixed-policy training, internalizing the exploration gains back into the unguided policy. Motivated by a theoretical and empirical analysis of the benefit-risk trade-off, we adopt a minimal intervention principle that invokes guidance only as an adaptive fallback, matching task difficulty while minimizing off-policy risk. On search-agent benchmarks, \textsc{ActGuide-RL} substantially improves over zero RL (+10.7 pp on GAIA and +19 pp on XBench with Qwen3-4B), and performs on par with the SFT+RL pipeline without any cold start. This suggests a new paradigm for agentic RL that reduces the reliance on heavy SFT data by using scalable action guidance instead.


Learning-Augmented Streaming Algorithms for Approximating Boolean Max-CSPs

Yinhao Dong ⋅ Pan Peng ⋅ Ali Vakilian ⋅ Santhoshini Velusamy ⋅ Xiaoyang Xu

We study learning-augmented streaming algorithms for estimating the value of Boolean Maximum Constraint Satisfaction Problems (Max-CSPs). Specifically, we consider streaming algorithms equipped with an $\varepsilon$-accurate oracle: for each variable, the oracle outputs a label in $\\{-1,1\\}$ that agrees with a fixed optimal assignment with probability $1/2+\varepsilon$, and disagrees otherwise. Prior work by Dong, Peng, and Vakilian [2025] showed that, for the special case of Max-Cut, such an oracle enables a $(1/2+\Omega(\varepsilon^2))$-approximation using $\operatorname{poly}(1/\varepsilon)$ words of space in insertion-only streams, and $\operatorname{poly}(1/\varepsilon,\log n)$ words of space in dynamic streams. We substantially generalize and strengthen their result, provided that a highly accurate estimate of $\varepsilon$ is available. For any Boolean Max-$2$CSP and any $\eta>0$ independent of $\varepsilon$, we obtain, with high constant probability, a single-pass $(1-\eta)$-approximation using $\operatorname{poly}(1/\varepsilon,1/\eta)$ words of space in insertion-only streams, and $\operatorname{poly}(1/\varepsilon,1/\eta,\log n)$ words of space in dynamic streams, where $n$ is the number of variables. We further extend our techniques to general Boolean Max-$k$CSPs for $k\ge 3$, with slightly worse space complexity.


Learning Data-free Universal Adversarial Perturbation with Hybrid Priors and Gradient-Guided Sharpness Regularization

Zhi Tan ⋅ Jiazheng Cui ⋅ Wenwen Zhang ⋅ Yiran Liu ⋅ Guodong Liu ⋅ Chris Ding ⋅ Di Ming

Universal adversarial perturbations (UAPs) aim to learn a single input-agnostic perturbation that consistently induces misclassification across diverse inputs, exposing systemic vulnerabilities of deep models. To enhance practical applicability, data-free UAP methods replace real data with samples drawn from handcrafted synthetic priors and optimize UAPs via activation maximization. However, existing approaches rely on a single synthetic prior throughout training, which biases optimization toward a narrow feature subspace and necessitates costly per-prior tuning. Moreover, the commonly used stochastic gradient ascent (SGA) updates are inherently unstable and tend to converge to sharp regions of the loss landscape, weakening black-box transferability. To overcome these limitations, we propose HPSR-UAP, a data-free framework that couples hybrid-prior reweighting with gradient-guided sharpness regularization. Specifically, we develop a dynamic reweighting mechanism over multiple synthetic priors to mitigate prior bias and avoid overfitting to any single distribution. Building upon this, we introduce a gradient-guided sharpness regularization that smooths the prior-weighted activation loss landscape within a local perturbation neighborhood, promoting updates towards flatter and more transferable directions. Extensive experiments on ImageNet demonstrate that HPSR-UAP outperforms state-of-the-art methods in both white-box and black-box data-free attack settings. We further introduce a multi-granularity block-wise similarity metric that quantifies spatial repetitiveness within UAPs, revealing a strong correlation between structural regularity and attack performance.


Learning from Disagreement: Multi-Teacher Distillation for Chinese Spelling Correction

Hongcheng Ding ⋅ Xuanze Zhao ⋅ Ziping Hu ⋅ Hongdan Xiao ⋅ Kang Xu ⋅ Weiyu Zhang

Chinese Spelling Correction (CSC) is hard precisely because a single misspelling can originate from any of four loosely coupled dimensions (glyph, phonetic, syntactic, or semantic confusion), yet existing systems treat the four as a flat fusion target. Multimodal pre-training and fixed-weight gating let evidence from each dimension co-exist, but never arbitrate: when a phonetic-plausible candidate contradicts a syntactic-plausible one, the model averages them, and learning effort is spent uniformly on cases where every dimension already agrees. Our diagnosis is that the bottleneck is no longer representation but \emph{decision}: the cases where dimensions disagree are exactly the cases worth specializing for, and dimension-specific teachers expose this disagreement as a usable signal rather than as noise to be smoothed away. We instantiate this view as SEMTD (Self-Evolving Multi-Teacher Distillation), a 4B-parameter student trained from four single-dimension experts (glyph / phonetic / syntactic / semantic) under three coupled objectives: multi-teacher distillation with input-dependent expert selection, self-evolving learning that turns disagreement into a composite reward, and reflective path optimization that replays high-disagreement decisions. Across four benchmarks the student reaches Cor-F1 of $83.8$ on SIGHAN15, $62.3$ on LEMON (7-domain average), $98.2$ on ECSpell (3-domain average), and $73.0$ on CSCD-NS, matching or surpassing 14B LLM-based correctors at $3.5{\times}$ fewer parameters. The takeaway is methodological: in CSC, treating teacher disagreement as the supervision target, rather than a residual to be averaged out, reopens headroom that flat fusion has saturated.


Learning inexact alternating minimization

Paul Häusner ⋅ Jevgenija Rudzusika ⋅ Jens Sjölund ⋅ Ozan Öktem

Alternating minimization is a standard tool for optimization problems with two coupled variable blocks. Its cost is dominated by the inner subproblem solvers, which rarely have a closed-form solution. In this paper, we propose to learn an inexact alternating scheme in which neural networks approximate the subproblem solution at each iteration. We derive a worst-case bound on the optimality gap that exposes the per-iteration error that each network is trained to minimize in a greedy fashion. To show convergence of the learned scheme on unseen problem instances, we combine the bound with a PAC-Bayes argument that controls the expected optimality gap with high probability. We apply the learned scheme to low-dose CT reconstruction with dictionary regularization, demonstrating a $13\times$ speedup over the inexact first-order baseline at matched accuracy. A brief end-to-end fine-tuning stage extends this to $30\times$ and outperforms other learned schemes.

Structured prediction in dynamic, open-world environments—such as online High-Definition (HD) map construction for autonomous driving, mobile robotics, and embodied AI—fundamentally relies on the precise alignment of continuous geometry and discrete semantics. However, existing methods typically decouple these heterogeneous modalities, producing "overconfident yet erroneous" predictions that lack reliable uncertainty calibration in degraded environments. To address this fundamental limitation, we propose IntroMap, a probabilistic framework dedicated to learning joint semantic-geometric uncertainty with closed-loop calibration. At its core, the Semantic-Geometric Joint Calibration (SGJC) mechanism explicitly captures the aleatoric co-occurrence between spatial coordinates and semantic proxies via a Hybrid 3D Multivariate Gaussian distribution. Furthermore, to overcome the shortcomings of passive, open-loop uncertainty estimation, we introduce the Closed-Loop Feature Calibration (CLFC) module. It transforms the predicted joint covariance into latent modulation signals, actively rectifying degraded feature representations. Extensive experiments on the nuScenes and Argoverse 2 datasets demonstrate IntroMap's model-agnostic generalizability across multiple baselines (e.g., MapTRv2, MapQR), achieving state-of-the-art accuracy—including 63.1\% mAP on nuScenes and a striking 8.1\% mAP leap under challenging rainy conditions. Crucially, downstream end-to-end autonomous driving evaluations verify that our calibrated joint uncertainty establishes explicit safety margins, effectively preventing motion planners from blindly trusting hallucinated features and significantly reducing collision rates in long-tail scenarios. Code is available at https://anonymous.4open.science/r/IntroMap-F0BD/.


Learning Modular Addition with Auxiliary Modulus

Hanato Kikuchi ⋅ Ryosuke Masuya ⋅ Kazuhiko Kawamoto ⋅ Hiroshi Kera

Learning parity functions, more general modular addition, is a challenging machine learning task due to its input sensitivity. A recent study substantially scaled modular addition learning in both the number of summands and the modulus. Its key idea is to increase zeros in training sequences, reducing the effective number of summands and thus controlling training difficulty; however, this induces covariate shift between training and test input distributions. This study theoretically and empirically analyzes this side effect and proposes a covariate-shift-free method for modular addition. Specifically, we introduce an auxiliary modulus $Kq$ during training, which reduces wrap-around frequency and problem difficulty while preserving the same input distribution across training and testing. Experiments show strong scalability and sample efficiency: even for large input length $N$, large modulus $q$, and small datasets---where the sparse method fails to learn---our method achieves equal or better match accuracy and relaxed $\tau$-accuracy. For example, at $N=64$ and $q=974269$, our method trained on 100K samples achieves $97.0\%$ $\tau$-accuracy at $\tau=0.05$, while the sparse method achieves only $9.5\%$ with the same data size and $93.9\%$ even when extended to 1M samples.


Learning POMDP World Models from Observations with Language-Model Priors

Valentin Six ⋅ Frederik Panse ⋅ Mathis Fajeau ⋅ Lancelot Da Costa ⋅ Mridul Sharma ⋅ Alfonso Amayuelas ⋅ Tim Xiao ⋅ David Hyland ⋅ Philipp Hennig ⋅ Bernhard Schölkopf

Whether navigating a building, operating a robot, or playing a game, an agent that acts effectively in an environment must first learn an internal model of how that environment works. Partially observable Markov decision processes (POMDPs) provide a flexible modeling class for such internal world models, but learning them from observation-action trajectories alone is challenging and typically requires extensive environment interaction. We ask whether language-model priors can reduce costly interaction by leveraging prior knowledge, and introduce Pinductor (POMDP-inductor): an LLM proposes candidate POMDP models from a few observation-action trajectories and iteratively refines them to optimize a belief-based likelihood score. Despite using strictly less information, Pinductor matches the performance and sample efficiency of LLM-based POMDP learning methods that assume privileged access to the hidden state, while significantly surpassing the sample efficiency of tabular POMDP baselines. Further results show that performance scales with LLM capability and degrades gracefully as semantic information about the environment is withheld. Together, these results position language-model priors as a practical tool for sample-efficient world-model learning under partial observability, and a step toward generalist agents in real-world environments.


Learning Rate Matters: Vanilla LoRA May Suffice for LLM Fine-tuning

Yu-Ang Lee ⋅ Ching-Yun Ko ⋅ Pin-Yu Chen ⋅ Mi-Yen Yeh

Low-Rank Adaptation (LoRA) is the prevailing approach for efficient large language model (LLM) fine-tuning. Building on this paradigm, recent studies have proposed alternative initialization strategies and architectural modifications, reporting substantial improvements over vanilla LoRA. However, these gains are often demonstrated under fixed or narrowly tuned hyperparameter settings, despite the known sensitivity of neural networks to training configurations. In this work, we systematically re-evaluate nine representative LoRA variants alongside vanilla LoRA through extensive hyperparameter searches. Across tasks spanning mathematical reasoning, commonsense reasoning, code generation, and instruction following at diverse model scales, we find that different LoRA methods favor distinct learning rate ranges. Crucially, once learning rates are properly tuned, all methods achieve similar peak performance (within 1-2\%), with only subtle rank-dependent behaviors. These results suggest that vanilla LoRA remains a competitive baseline and that improvements reported under a single training configuration may not reflect consistent methodological advantages. Finally, a second-order analysis attributes the differing optimal learning rate ranges to variations in the largest Hessian eigenvalue, aligning with classical learning theories.

Naturalistic language encoding models based on large language model features can predict cortical responses well, but their learned representations remain difficult to interpret. Three challenges are central: semantic selectivity is fundamentally non-identifiable in high-dimensional correlated feature spaces, cortical organization remains implicit in voxelwise weight maps, and cross-subject structure is usually addressed only after model fitting. We introduce a multisubject coupled sparse autoencoder that jointly reconstructs language-model features and predicts subject-specific brain responses, learning a sparse basis of semantic-cortical atoms that makes semantic structure, cortical organization, and cross-subject correspondence explicit within a single model. In naturalistic story-listening fMRI, the model retains 88\% of matched dense-feature ridge brain prediction performance using 50 atoms and $k{=}5$ active atoms per TR. Relative to matched two-stage baselines, joint training yields the strongest overall tradeoff between prediction and interpretability. The learned atoms are semantically coherent and interpretable, recover more reproducible cross-subject brain maps than standard post hoc encoding-model analyses, align with large-scale cortical structure from Yeo~7 and Neurosynth, and transfer better than ridge to held-out subjects in the low-data regime. This framework provides a more interpretable basis for cognitive neuroscience analyses of semantic-to-brain encoding by linking semantic features to corresponding cortical response patterns across subjects.

Spiking reservoir computing, and reservoir computing more generally, is a powerful and efficient framework for neuromorphic and biological applications, in which a fixed random reservoir drives a trained readout. Yet its performance depends critically on the reservoir initialization. Enriching the reservoir with adaptive, unsupervised plasticity rules offers a natural solution to this limitation. However, it is unclear which rules work, or why, and how to find the best-suited rules for specific tasks without exhaustive and expensive search. Here, we learn how to play the Atari game Pong in a plastic spiking reservoir network. We systematically characterize a large family of local plasticity rules that were meta-learned in prior work. We then show that their performance is predictable and structured: high-scoring rules are characterized by strong differentiation between neurons encoding the ball trajectory and the background, with stable weight dynamics, and consistent readout alignment across time. Rule performance can be estimated from rule parameters alone. Finally, conditioning simulation-based inference (SBI) on high scores allows us to sample directly from promising regions of the rule space, to discover rules that exceed the performance of those found in the prior distribution and reveal stable, high-performing configurations. Together, these results offer competitive performance against classical reservoir computing while providing a transparent, interpretable account of what makes a plasticity rule useful for neuromorphic hardware and biological computing.


Learning to Solve Compositional Geometry Routing Problems

Mingfeng Fan ⋅ Jianan Zhou ⋅ Jiaqi Cheng ⋅ Yifeng Zhang ⋅ Jie Zhang ⋅ Guillaume Sartoretti

We study the Compositional Geometry Routing Problem (CGRP), a unified superclass of traditional routing problems that covers point-only, line-only, area-only, and arbitrary hybrid task geometries, providing a broad abstraction for real-world routing scenarios. Beyond standard point-based routing, CGRP with non-point tasks can be inherently asymmetric, tightly coupled travel routes with the intrinsic path, and enlarges the action space with numerous feasible yet often irrelevant options, thereby posing significant challenges for both representation learning and decision-making. To address these challenges, we propose DiCon, a differential attention–assisted solver with contrastive learning, as a plug-and-play framework that tackles the problem from two complementary angles. First, we introduce a differential attention mechanism that actively suppresses the probability mass on less competitive candidate actions. Second, we design a double-level contrastive learning objective to promote robust global instance representations and regularize geometry-aware task representations. Extensive experiments demonstrate that DiCon achieves strong performance, broad versatility, and superior generalization across diverse CGRP instances with different compositions.


Learning to Synergize Textual and Visual Prompts for Fine-Grained Traffic Element Detection in HD Maps

Xiaoyang Bi ⋅ Haowen Guo ⋅ Caoshengzhe Xue ⋅ Siyuan Li ⋅ TAO XUE ⋅ Dong Liu ⋅ Mengshi Qi

High-definition (HD) map construction demands the exhaustive and precise parsing of diverse traffic elements. However, practical traffic scene perception faces two fundamental hurdles: (1) generalizing to novel traffic elements that continuously emerge during periodic map updates, and (2) synergizing fine-grained textual descriptions with visual exemplars to resolve severe visual ambiguities. Consequently, even state-of-the-art Vision-Language Models (VLMs) exhibit severe vulnerabilities when confronting these challenges, hindering their direct deployment. To bridge this gap, we formalize Multimodal Fine-grained Traffic Element Detection (MFTED), a practical task dedicated to identifying traffic targets by synergizing nuanced textual and visual prompts. We introduce MFT-150K, a large-scale benchmark featuring over 150K images and 337K instance annotations, explicitly designed to evaluate fine-grained multimodal alignment and novel category generalization. Furthermore, we provide a comprehensive benchmark of state-of-the-art models and propose an initial baseline to explore multimodal fusion strategies. Extensive experiments reveal a significant multimodal fusion gap in current VLMs when handling fine-grained traffic semantics. Even our proposed probabilistic latent-space alignment yields only limited gains, underscoring the difficulty of this task and positioning MFT-150K as a challenging benchmark for advancing intelligent transportation. Our full-scale dataset is publicly available at https://huggingface.co/datasets/MFTED/MFTED, and code will be released upon publication.


Learning When to Think: Dual-Reference Offline Optimization for Adaptive VLM Reasoning

Hongbin Lin ⋅ Sizhe Zou ⋅ Juangui Xu ⋅ Xinyue Xu ⋅ Zhuoli Ouyang ⋅ Stanislav Kolchin ⋅ Chengwei Qin ⋅ Zhongxiang Dai ⋅ Yao SHU

Reasoning traces improve vision-language models on complex multimodal tasks, but forcing every input to follow a long reasoning process leads to overthinking and unnecessary inference cost. We study adaptive fast/slow thinking for VLMs, where a single model should answer directly for simple inputs and reason step by step only when needed. The key observation is that offline adaptive thinking is not merely a length-control problem: fast direct answers and slow reasoning traces are typically generated by distinct behavior policies. This dual-reference structure makes standard offline reward optimization, such as Decoupled Generation and Optimization (DGO), mismatched because it assumes a single reference policy and can misweight one mode. We propose \textsc{Dual-Reference DGO}, which models fast and slow responses under a fast/slow mixture reference and derives a reward-weighted offline objective for adaptive VLM reasoning. Our formulation shows that router-free fast/slow switching emerges from competition between fast and slow partition functions, and yields an efficient mode-wise correction for behavior-reference mismatch. Experiments on multiple datasets show that \textsc{Dual-Reference DGO} improves the accuracy--length trade-off over baselines. Ablations show that single-reference DGO tends to collapse to almost always-fast or always-slow behavior, while our dual-reference formulation preserves both modes and learns difficulty-dependent routing.


Leveraging Dale’s Principle as an Inductive Bias in Recurrent Neural Networks

Jiyi Wang ⋅ Chenxiao Yang ⋅ Yujia Zhao ⋅ Jingzhao Zhang ⋅ Lu Mi

One of the fundamental differences between biological neural networks and artificial neural networks (ANNs) is that biological neurons obey Dale’s principle: each neuron projects either exclusively excitatory or exclusively inhibitory outputs. While this constraint is central to biological circuits, previous attempts to impose Dale’s principle in ANNs have often led to training instability or degraded performance. In this work, we introduce EISep RNN, a recurrent neural network architecture that implements Dale’s principle through a learnable excitatory–inhibitory separation. Across a series of learning settings, including multi-task learning, continual learning, and reinforcement learning, EISep achieves more stable training dynamics, competitive or improved performance, and sparser, more modular connectivity than vanilla RNNs and previous Dale-constrained models. EISep also exhibits stable and informative latent dynamics. These results establish Dale’s principle as a biologically grounded and computationally effective inductive bias for recurrent neural networks.


Leveraging Error Diversity in Group Rollouts for Reinforcement Learning

Wenpu Liu ⋅ Yuqi Xu ⋅ Weichu Xie ⋅ Yongfu Zhu ⋅ Shuai Dong ⋅ Ziyue Wang ⋅ Wenqi Shao ⋅ Xiaoying Zhang ⋅ Tong Yang ⋅ Nan Duan ⋅ Jiaqi Wang

Reinforcement Learning from Verifiable Rewards (RLVR) typically samples multiple responses per prompt and assigns binary rewards based on individual correctness, yet the collective structure of the group output, specifically, the distribution of errors is largely discarded. We identify this as a missed opportunity: empirical analysis reveals that error diversity within a group is a strong predictor of training success, with problems eliciting diverse wrong answers benefiting substantially more from RLVR than those producing homogeneous failures. Motivated by this observation, we propose $\textbf{Error Diversity Advantage Shaping (EDAS)}$, a lightweight, algorithm-agnostic technique that modulates the advantage signal for incorrect rollouts based on intra-group error diversity. EDAS amplifies penalties for dominant, repeated errors and attenuates penalties for rare, exploratory ones, thereby encouraging the model to maintain diverse reasoning paths and discouraging error perseveration. Crucially, EDAS operates as a simple post-hoc adjustment that can be seamlessly integrated into any RLVR algorithm. We validate EDAS on top of several mainstream RLVR methods across a series of models and seven challenging math benchmarks, demonstrating consistent improvements. Notably, EDAS yields an average improvement of $\textbf{6.29}$ points over DAPO on Qwen3-8B across seven benchmarks, confirming that exploiting the latent information in group rollouts is a broadly effective strategy for strengthening RLVR.


LG-Bench: A Graph-Structured Evaluation Benchmark for Life Science

Lu Sun ⋅ Xiangyi Zhang ⋅ Xiangyang Zhu ⋅ Zijian Chen ⋅ Qiang Hu ⋅ Yuan Tian ⋅ Guangtao Zhai

Traditional evaluation benchmarks reduce inherently interconnected scientific knowledge in life sciences into flat lists of questions, disregarding the underlying topological structure of the knowledge. We introduce, the first graph-structured benchmark for life sciences, featuring over 10,000 high-quality multiple-choice questions across medicine, biology, and chemistry. Our approach constructs a weighted evaluation graph using bidirectional matching and semantic similarity algorithms, where nodes represent questions and edge weights capture their semantic relationships. Leveraging this graph topology, we design two novel evaluation metrics. The Global Coherence Score (GCS) measures a model’s consistency within semantically related neighborhoods, while Knowledge Balance Score (KBS) analyzes how model errors are distributed across the graph to reveal conceptual blind spots. LG-Bench facilitates fine-grained comparison of LLMs by surfacing differences in conceptual coherence and patterns of knowledge organization across models. Our framework shifts the evaluation paradigm from flat accuracy metrics to structure-aware analysis, offering a new lens for diagnosing and improving LLM performance in the life sciences domain.


LibEvoBench: Probing Temporal Knowledge Stratification in Code Generation Models

Daniele Cipollone ⋅ Sergey Titov ⋅ Maliheh Izadi ⋅ Egor Bogomolov ⋅ Arie v Deursen

Large software projects often depend on older versions of libraries, even as APIs continue to evolve across releases. This creates a challenge for LLMs: they must maintain knowledge of multiple API versions, not merely the latest or most common one. However, current LLMs, trained on temporally mixed corpora and lack explicit mechanisms for such version-specific reasoning, leading to anachronistic errors -- calling APIs as they exist in a different library version. To systematically evaluate this phenomenon, we introduce LibEvoBench, a multi-task benchmark spanning multiple versions of widely used Python libraries, along with a new metric, the Software Evolution Understanding Score (SEUS), to measure models' consistency when working with evolving APIs. Our results show that state-of-the-art models are largely version-oblivious: performance degrades for evolving APIs, while for stable APIs it remains the same across versions. Moreover, simply specifying the target version provides no benefit, while relevant documentation significantly boosts models' accuracy. These findings highlight a systematic limitation of current training paradigms and motivate new approaches for temporally grounded knowledge in code generation.


LightMoE: Reducing Mixture-of-Experts Redundancy through Expert Replacing

Jiawei Hao ⋅ Zhiwei Hao ⋅ Jianyuan Guo ⋅ Li Shen ⋅ Yong Luo ⋅ Han Hu ⋅ Dan Zeng

Mixture-of-Experts (MoE) based Large Language Models (LLMs) have demonstrated impressive performance and computational efficiency. However, their deployment is often constrained by substantial memory demands, primarily due to the need to load numerous expert modules. While existing expert compression techniques like pruning or merging attempt to mitigate this, they often suffer from irreversible knowledge loss or high training overhead. In this paper, we propose a novel expert compression paradigm termed expert replacing, which replaces redundant experts with parameter-efficient modules and recovers their capabilities with low training costs. We find that even a straightforward baseline of this paradigm yields promising performance. Building on this foundation, we introduce LightMoE, a framework that enhances the paradigm by introducing adaptive expert selection, hierarchical expert construction, and an annealed recovery strategy. Experimental results show that LightMoE matches the performance of LoRA fine-tuning at a 30% compression ratio. Even under a more aggressive 50% compression rate, it outperforms existing methods and achieves average performance improvements of 5.6% across five diverse tasks. These findings demonstrate that LightMoE strikes a superior balance among memory efficiency, training efficiency, and model performance.

Linear contextual bandit algorithms are typically built around pointwise optimism: the exploration bonus is chosen large enough to upper-bound the value of every action with high probability. We study an alternative principle of quasi-optimism, in which the index is deliberately allowed to be smaller than a valid confidence bound. We analyze the algorithmic framework, EQOL, that replaces the usual confidence-radius bonus with the capped quadratic bonus min(c₁,t ‖x‖²_{Vt⁻¹}, c₂,t ‖x‖{Vt⁻¹}). This index is generally not optimistic, since its quadratic branch can be pointwise smaller than the OFUL bonus. We prove a general regret bound for arbitrary schedules (c₁,t, c₂,t), separating the exploration cost from the quasi-optimistic residual. With a single gap-agnostic tuning, EQOL recovers the high-probability minimax O(d√T) regret bound and, under a positive contextual margin, a uniform gap-dependent O(d²/Δmin,T) bound. We also provide a geometric interpretation of EQOL through action-dependent ellipsoids, clarifying how quasi-optimism differs from simply shrinking a UCB bonus. Experiments on synthetic and real-data contextual bandit benchmarks show that EQOL substantially outperforms existing algorithms.


LiveOption: Evaluating LLM Agents in Structured Option Trading with Nonlinear Payoffs

Haochen Luo ⋅ YIFAN LI ⋅ An B Minh ⋅ Xiaolong Luo ⋅ Zhengzhao Lai ⋅ Yuan Zhang ⋅ Chen Liu

Large language models (LLMs) and multi-agent systems (MAS) have shown promise in financial decision-making, yet existing evaluations focus on equity trading and primarily assess directional prediction, overlooking the structural complexity of derivative markets. Option trading introduces fundamentally different challenges, including nonlinear payoffs and multi-leg strategy construction, requiring structured decisions rather than simple directional bets. We introduce LiveOption, an evaluation framework for LLM-based and reinforcement learning agents in option trading. LiveOption formulates the problem as structured sequential decision-making under realistic execution and capital constraints, and provides a reproducible environment with standardized interaction protocols. The framework includes several task suites covering portfolio overlays, open-ended strategy construction, event-driven earnings trading, and 0DTE intraday trading. We further propose a hierarchical metric suite that evaluates action validity, decision quality, risk characteristics, and outcome-level performance. Experiments show that current agents often fail to construct valid strategies and handle nonlinear payoff dynamics, or achieve competitive returns in most scenarios. LiveOption offers a principled testbed for evaluating structured decision-making beyond outcome-based metrics.

The rapid proliferation of Large Language Models (LLMs) has created a diverse ecosystem and motivated model routing as a way to navigate differences in architectures, capabilities, latency, and cost. Yet current routing methods still face bottlenecks in generalization, scalability, and interpretability. To solve these problems, we argue that model routing can be formulated as a recommendation problem. This formulation links routing bottlenecks to problems long studied in recommendation systems (RecSys), including cold start, large-scale recommendation, dynamic adaptation, and explainable decision-making. Building on this RecSys perspective, we organize a roadmap that maps major routing bottlenecks to established RecSys paradigms, providing a practical foundation for adapting mature recommendation methodologies to multi-LLM orchestration. We further present preliminary empirical case studies showing that recommendation-inspired routers can achieve strong accuracy-cost trade-offs while supporting model cold-start and interpretable routing behavior.


LLMs Keep Thinking When Told Not To

Dianqiao Lei ⋅ Kevin Qinghong Lin ⋅ Pan Lu ⋅ Philip Torr ⋅ James Zou

Large Language Models (LLMs) increasingly ship with “thinking modes”, yet their counterpart, the “no-thinking” has received far less exploration. In this work, we revisit language model “no-thinking” behavior along the following efforts: (a) Explore the definition of no-thinking. Rather than relying on proxies like thinking-mode control, we split each response into pre-answer text and a final answer, scoring both answer-only compliance and the question–pre-answer relevance. (b) Broad experimental coverage. We evaluate multiple no-think controls (thinking-mode toggles, prompt instructions, etc.) across boolean, multiple-choice, and open-ended tasks, on open-source models with scaled variants, and proprietary models. With extensive experiments, we found that (i) Explicit no-think controls do not reliably eliminate visible reasoning-like content. Instead, models exhibit (ii) “thinking inertia”: residual content is systematically organized by the task's answer space, with boolean verification compressing most easily, multiple-choice tasks forming an intermediate regime, and open-ended tasks remaining the most resistant. (iii) We further identify what we call a no-thinking and performance trade-off in open-ended: stronger answer-only constraints on complex tasks either fail to suppress the intermediate payload or succeed by removing computation needed for accuracy. (iv) Finally, we show that compressibility depends strongly on the structural support provided by the question itself. Thus, no-think controls are not direct guarantees of no-thinking; they must be evaluated jointly through answer-only success, visible payload, and task accuracy. We hope our work moves LLMs closer to a human-like ability to stop thinking on demand.


Local Gaussian Processes on Compact Lie Groups

Bochuan Liu ⋅ Mingyang Zhao ⋅ Xiaohong Jia

Gaussian processes are a cornerstone of modern probabilistic machine learning, providing principled uncertainty quantification and flexible nonparametric modeling. While the Gaussian kernel is widely used in Euclidean spaces, many real-world problems involve data residing on non-Euclidean domains, particularly compact Lie groups that naturally encode \emph{symmetries}. In this work, we construct a local Gaussian kernel on arbitrary compact Lie groups by decomposing the group into a \emph{maximal torus} and its orthogonal complement. For the prominent Lie groups $\mathbf{SO}(3)$ and $\mathbf{SU}(2)$, we derive closed-form kernel expressions using the Rodrigues formula. Extensive experiments verify the positive definiteness of the proposed kernel and demonstrate its effectiveness in Gaussian process regression, achieving accurate predictions across the entire manifold.

Visual understanding is inherently hierarchical as high-level semantic concepts are composed of fine-grained, spatially localized components. Yet existing interpretability methods typically operate at a single scale, either local or global, leaving a critical gap in understanding the compositional structure between them. Recovering this structure without supervision is particularly challenging, as the correspondence between local and global concepts is latent, varies across samples, and may appear only through inconsistent or partial patterns. To address this challenge, we introduce Local-Global Sparse Autoencoders (LG-SAE), an unsupervised framework for multiscale interpretability in vision models. The core innovation of LG-SAE is a learned compositional bridge that explicitly captures the relationships between local and global concepts. By jointly decomposing patch-level and image-level embeddings, LG-SAE induces a sparse interpretable graph that reveals how each high-level semantic concept is supported by a small set of spatially grounded components. Building on this representation, we present a unified framework for discovering local-global concept relationships, constructing multiscale Concept Bottleneck Models, spatially aware concept naming, and model steering through localized concept edits. Extensive evaluations show that LG-SAE: (i) captures meaningful cross-scale feature relations; (ii) enables high-quality spatial grounding without compromising downstream accuracy; and (iii) provides interpretable model explanations, further validated by a user study.

As large language models (LLMs) are increasingly integrated into high-stakes decision-making, the ability to reliably quantify uncertainty has become a critical requirement for safety and trust. However, current uncertainty quantification methods primarily operate at the output level, often failing to distinguish whether uncertainty arises from the model’s lack of knowledge or from ambiguity in the user’s input. While input-centric uncertainty quantification has recently emerged as a promising direction, it remains relatively underexplored and typically relies on coarse, input-level information. Consequently, users are provided with scalar uncertainty scores that offer little actionable guidance on which parts of the input should be clarified to improve reliability. To address this limitation, we propose Shapley-based input uncertainty Quantification (ShaQ), a framework for span-level attribution of input-induced uncertainty. Our approach models ambiguous spans in the input as players in a cooperative game and quantifies their contributions using Shapley values, defined via the weighted average of marginal reductions in conditional entropy obtained by clarifying each span coalition. Unlike existing input-level approaches, our formulation explicitly captures complex interactions among spans and provides a principled decomposition in which individual attributions sum exactly to the total input-induced uncertainty. We evaluate ShaQ on the AmbigQA and AmbiEnt benchmarks, where it achieves state-of-the-art performance in ambiguity detection. We further demonstrate its practical utility on a MediTOD benchmark, showing that ShaQ can precisely localize under-specified clinical utterances, facilitating more effective human-AI collaboration in high-stakes settings. Overall, our results show that ShaQ not only improves uncertainty estimation but also provides actionable insights for targeted input clarification. Codes are available https://anonymous.4open.science/r/ShaQ-0E39/README.md.


Locking Pretrained Weights via Deep Low-Rank Residual Distillation

Keitaro Sakamoto ⋅ Pierre Ablin ⋅ Federico Danieli ⋅ Marco Cuturi

The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also allow for more open research and testing, to the extent that users can use them as checkpoints, fine-tune them according to their needs, and potentially redistribute them. In some cases, however, concerns on modifying these weights towards unauthorized uses may outweigh the pros of giving users such a freedom. Defending against such adaptation is non-trivial: since an adaptive attacker can observe all weights and architectures by definition, they can reverse simple structural defenses, and use optimization to defeat the simplest locking mechanisms. In this work, we exploit the inference–training asymmetry of automatic differentiation as a novel defense axis. We propose DLR-Lock, a method where the purveyor of the model purposely replaces each pretrained MLP in their model with a deep low-rank residual network (DLR-Net) of comparable parameter count, forcing activation memory that grows linearly with depth during backpropagation. DLR-Nets are efficiently trained via module-wise distillation. We show that, beyond this memory overhead, DLR-Lock results in architectural mismatches that complicate the optimization landscape of standard fine-tuning, and a backward pass that incurs disproportionately more overhead than the forward pass. Our defense succeeds in withstanding adaptive attackers with full knowledge of the defense strategy while preserving the original model's capabilities. Experiments on LLM validate these claims.


LOCO: Local Light-Aware Object Compositing with Spatially Varying Illumination-Augmented Data

Jinseo Jeong ⋅ Hyunsoo Kim ⋅ Junseo Koo ⋅ Junhyeog Yun ⋅ Gunhee Kim ⋅ Soo Ye Kim

In natural scenes, light sources and occlusions from scene geometry create spatially varying illumination, including brightness gradients, cast shadows, and local color shifts that shape how every object appears. A less explored aspect of diffusion-based object compositing is whether an inserted object can adapt to this local illumination at the desired location. The core challenge to achieve this is the scarcity of large-scale training data capturing how object appearance should adapt to location-dependent illumination. We present LOCO, a data augmentation framework for LOcal light-aware object COmpositing that leverages publicly available monocular videos as illumination sources. We treat each video frame as a virtual spotlight anchored at its camera pose, so that viewpoint variation across a video yields diverse light directions. With controllable beam shape, color, and intensity, this design enables large-scale generation of training data for placement-dependent illumination on the fly. We also introduce LOCO-bench, a benchmark of physically captured composites under spatially varying illumination, where the same object is placed at multiple locations within each scene. Experiments on Laval Indoor SV HDR and LOCO-bench show that LOCO outperforms prior methods in local illumination adaptation while preserving object identity.

We propose Log-Averaged Mirror Prox (LAMP), a linear-space primal-dual method for large-scale optimal transport. LAMP implements primal mirror prox updates by tracking an averaged dual sequence, reducing storage complexity from $\mathcal O(nm)$ to $\mathcal O(n+m)$ while preserving dense, GPU-friendly reductions. Consequently, LAMP preserves the last-iterate $\widetilde{\mathcal{O}}( nm\varepsilon^{-1})$ arithmetic complexity of conservatively parameterized primal-dual mirror prox. We further analyze LAMP as a direct optimal transport solver in a more performant parameter regime, providing a last-iterate sub-optimality certificate dependent on infeasibility and an explicit $\mathcal O(1/t)$ term. Moreover, we give a computable sufficient condition for best-iterate convergence to a saddle-point. Numerical experiments with an optimized CUDA implementation show that LAMP outperforms first-order baselines in several high-accuracy (entropic) optimal transport problems. LAMP is further shown to scale up to problems with $n=m=2^{18}$ marginal supports, which were previously beyond the reach of primal-dual first-order methods.


Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

Aryo Gema ⋅ Beatrice Alex ⋅ Pasquale Minervini

In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pasting them. Identifying which attention heads perform this synthesis matters for interpreting long-context model behavior. Yet existing detectors miss these heads by construction: they reward heads whose attended token matches the generated token, a literal-copy criterion that captures where a head reads but not what it writes through its output-value (OV) circuit, the very mechanism that carries non-literal retrieval. We introduce Logit-Contribution Scoring (LOCOS), a write-aware detector that scores each head by the projection of its OV-circuit output onto the answer-token unembedding direction, contrasting needle and off-needle source positions in a single forward pass. Across three model families (Qwen3, Gemma-3, OLMo-3.1), mean-ablating the top LOCOS heads on the NoLiMa non-literal retrieval benchmark collapses ROUGE-L at lower head counts than every attention-based baseline; on Qwen3-8B, ablating 50 heads drives ROUGE-L from 0.401 to 0.000 while the strongest baseline still retains 0.292. The selected heads are retrieval-specific: parametric recall and arithmetic reasoning stay at baseline under the same ablation.


LogSTOP: Temporal Scores over Prediction Sequences for Matching and Retrieval

Avishree Khare ⋅ Hideki Okamoto ⋅ Bardh Hoxha ⋅ Georgios Fainekos ⋅ Rajeev Alur

Neural models such as YOLO and HuBERT detect local properties such as objects ("car") and emotions ("angry") in video frames and audio clips, with scores in [0, 1]. Lifting these scores to temporal properties over sequences enables applications such as query matching (e.g., "does the speaker eventually sound happy in this audio clip?") and ranked retrieval (e.g., "retrieve top 5 videos with a 10 second scene where a car is detected until a pedestrian is detected"). We formalize this problem of assigning Scores for TempOral Properties (STOPs) over sequences, given noisy score predictors for local properties. We propose LogSTOP, a scoring function that efficiently computes scores for temporal properties represented in Linear Temporal Logic. Empirically, LogSTOP with YOLO and HuBERT outperforms Large Vision / Audio Language Models by at least 16\% on query matching with temporal properties over objects-in-videos and emotions-in-speech, while matching or improving over temporal-logic baselines. On ranked retrieval with temporal properties over objects and actions in videos, LogSTOP with OWLv2 and SlowR50 improves mean average precision over zero-shot text-to-video retrieval baselines by 28\% and 18\% respectively.


LongSCOP: Semantically Consistent Long Video Outpainting

Yu Tang ⋅ Zhiyuan You ⋅ Zhixiao Wang ⋅ Fanghua Yu ⋅ Qingyu Zhang ⋅ Xiang Yin ⋅ Chao Dong ⋅ Jinjin Gu

Long video out-painting is an important yet underexplored problem, while existing video out-painting methods are mainly developed for relatively short clips. However, extending video out-painting to long videos is highly challenging. A natural solution is iterative out-painting, \ie, generating videos chunk by chunk, but this introduces two major issues: unstable even crashed long-horizon generation caused by inter-chunk inconsistency and accumulated errors, and semantic drift over long temporal horizons. In this paper, we propose LongSCOP, short for Semantically Consistent long Video Out-Painting. First, we provide a fundamental training recipe with three components for smooth and stable long-horizon generation: self-forcing training, latent calibration, and future references. Built upon this stable generation foundation, a more advanced requirement is semantic consistency, for which we further develop a VLM-based agent that automatically selects the most informative reference frames from the whole video, enabling coherent subjects, scenes, and backgrounds throughout long-range generation. To support training and evaluation, we construct a large-scale, high-quality training dataset and introduce LVO-Bench, a dedicated benchmark curated by human experts for assessing long-horizon stability and semantic consistency. Extensive experiments demonstrate clear advantages of our method over existing approaches in both generation stability and semantic coherence, significantly improving the quality and reliability of long video out-painting.

Spiking Neural Networks (SNNs) are well-regarded for their biological plausibility and energy efficiency in processing sequential data. However, dominant SNN architectures typically rely on first-order Ordinary Differential Equations (ODEs) to govern neuronal state transitions. This first-order assumption imposes a ``memoryless'' bottleneck, limiting the model's capacity to capture the complex, long-range dependencies inherent in long-sequence tasks. In this work, we propose LongSpike, a novel SNN framework that integrates fractional-order State-Space Modeling (f-SSM) from control theory into the spiking domain. By extending traditional integer-order SSMs to the fractional-calculus regime, LongSpike enables the hierarchical integration of neuronal dynamics with long-memory kernels. To mitigate the computational overhead and parallelization challenges typically associated with fractional operators, we leverage a state-space formulation that supports efficient, parallel training. Empirical evaluations on challenging benchmarks, including Long Range Arena (LRA), large-scale WikiText-103, and Speech Commands, demonstrate that LongSpike outperforms state-of-the-art SNNs in accuracy while preserving sparse synaptic computation.


Lookalike3D: Seeing Double in 3D

Chandan Yeshwanth ⋅ Angela Dai

3D object understanding and generation methods in indoor scenes produce impressive results, yet they often overlook a pervasive source of information in real-world scenes: repeated objects. We introduce the task of lookalike object detection in indoor scenes, which leverages repeated and complementary cues from identical and near-identical object pairs. Given an input scene, the task is to classify pairs of objects as identical, similar or different using multiview images as input. To address this, we present Lookalike3D, a multiview image transformer that effectively distinguishes such object pairs through similarity learning by harnessing strong semantic priors from large image foundation models. To support this task, we collected the 3DTwins dataset, containing 76k manually annotated identical, similar and different pairs of objects based on ScanNet++, and show an improvement of 14\% IoU over baselines on the task of lookalike object detection. We also demonstrate how our method improves the downstream tasks of joint 3D object reconstruction and part co-segmentation, turning repeated and lookalike objects into a powerful cue for consistent, high-quality 3D scene understanding. Our code, dataset and models will be made publicly available.


Look Before You Reason: Implicit Visual Thinking for Efficient Multimodal Reasoning

Zerui Chen ⋅ Changrui Chen ⋅ Fei Ni ⋅ Jiankang Deng

Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet they remain limited on tasks requiring fine-grained understanding of small or easily overlooked image regions. Recent thinking with images approaches mitigate this limitation by allowing models to inspect images through operations such as zooming, cropping, or code execution. While effective, their multi-turn interactions with VLMs introduce redundant token computation and slow inference, and separate region inspection from the reasoning step that depends on it. To address these issues, we propose Look Before You Reason (LoRe), an implicit visual thinking framework for efficient visual reasoning within a single-turn VLM interaction. Given an image and a question, LoRe predicts the most relevant image region, aligns it with the corresponding image patch features, and strengthens the associated key-value (KV) cache entries, enabling subsequent decoding to directly exploit the selected visual cues without additional visual inputs or explicit tool calls. We further employ reinforcement learning to guide the VLM toward identifying informative image regions and leveraging them during reasoning. To support fine-grained evaluation of multimodal reasoning, we introduce COCO-ObjQA, a large-scale benchmark for object-centric multimodal reasoning. Experiments on three challenging multimodal benchmarks demonstrate that LoRe consistently improves VLM reasoning performance with low additional inference overhead.

We revisit regularized regret minimization under full-information and bandit feedback, where a learner optimizes an objective of the form $\langle r, \pi \rangle - \eta^{-1} \psi(\pi)$ for adversarially chosen rewards $r$, policies $\pi \in \Delta_K$, and a known strongly-convex regularizer $\psi(\cdot)$. We establish the minimax regularized regret for a broad class of regularizers, specifically those that are relatively strongly-convex and smooth with exponent $p \geq 1$, Legendre, and coordinate-wise separable. For large $\eta$, the problem approaches unregularized online learning, recovering the standard $\Theta(\sqrt{T})$ scaling. For small $\eta$, the strong convexity of the regularizer dominates, and we identify the regimes that attain $\Theta(\eta \log T)$ minimax regret under both feedback models. Notably, while prior logarithmic regret lower bounds are largely restricted to specific curved loss functions (e.g., squared loss), our lower bounds hold for ***any*** such regularizer against purely linear adversarial rewards. A key technical contribution is our proof strategy, which leverages a Bregman geometric perspective of regularized regret over the quotient space $\mathbb{R}^K / \mathrm{span}(\\{\mathbf{1}\\})$ to resolve the translation invariance of the inverse mirror map. By combining this dual geometric formulation with the ***van Trees inequality***, we reduce the adversarial online learning to Bayesian estimation, offering a unified framework for establishing logarithmic regret lower bounds.


LookThere! Sparse Vision by Reinforced Selection

Sreehari Rammohan ⋅ Yousef Yassin ⋅ Anthony Fuller ⋅ Junfeng Wen ⋅ Evan Shelhamer ⋅ Carl Vondrick

Essential visual information can be sparse, yet most visual computation remains dense. Though only a fraction matters, vision transformers process images as uniform sets of tokens, scaling cost with their number. As these models grow stronger, they grow larger, and the cost of computing everything climbs higher. Adaptive computation has shown that we can get by with less, but existing methods struggle at extreme sparsity and are difficult to specialize to a task. We address this limitation by introducing LookThere, training an efficient selector of locations and an expressive extractor of representations to derive meaning from images without ever processing the full high-resolution input. We do so by learning the selector and extractor end-to-end with actor-critic reinforcement learning. The selector learns where to look, the extractor what to see, together saving computation by selecting only what is worth processing for a given task. We show that LookThere selects the task-specific input, excelling at sparse recognition in high-resolution settings (traffic signs, billiards), and maintaining accuracy with as little as 0.2% of the input. It also generalizes across tasks and models, including global recognition (ImageNet classification), local recognition (ADE20K segmentation), zero-shot classification (by distillation), and regression (class agnostic enumeration). Across all settings, LookThere surpasses state-of-the-art selection, providing a general and scalable framework for specialized and efficient adaptive computation.


LoopNav: Benchmarking Spatial Consistency in World Models

Kewei Lian ⋅ Shaofei Cai ⋅ Yitao Liang ⋅ Anji Liu

The ability to simulate the world in a spatially consistent manner is a crucial requirement for effective world models. Such a model enables high-quality visual generation, and also ensures the reliability of world models for downstream tasks such as simulation and planning. It must not only retain long-horizon observational information, but also enables the construction of explicit or implicit internal spatial representations. However, existing datasets do not explicitly enforce spatial consistency constraints, limiting both the ability to systematically evaluate this capability and to learn it through data-driven approaches. Furthermore, most existing benchmarks primarily emphasize visual coherence or generation quality, neglecting the requirement of long-range spatial consistency. To bridge this gap, we propose loopnav, a dataset and corresponding benchmark centered on loop-based navigation for evaluating spatial consistency. The dataset comprises 250 hours (20 million frames) of loop-based navigation videos with actions, collected from diverse locations in the open-world environment of Minecraft. We further introduce a Scene Graph Consistency Score to quantify spatial consistency while remaining invariant to pixel-level variations. Dataset, benchmark, and code are open-sourced to support future research.

Structured pruning offers hardware-agnostic acceleration for Large Language Models (LLMs) but typically degrades performance by strictly adhering to the original static sequential execution graph. In this paper, we challenge this rigidity with LoopPrune, an evolutionary framework that transforms model compression into an architectural re-discovery process. By treating pre-trained layers as a library of reusable primitives, LoopPrune optimizes a dynamic execution path where modules can be skipped, executed once, or iterated via controlled loops. This approach effectively decouples effective model depth from parameter count, allowing compressed models to maintain high reasoning capacity. To navigate the vast combinatorial search space, we introduce a hybrid evolutionary strategy enhanced by a Multi-Tier Population Initialization mechanism, which stratifies the search start-point using cross-model elite transfer, structural priors, and Wanda score importance estimation. Extensive experiments on diverse LLM families demonstrate that LoopPrune sets a new state-of-the-art across all sparsity levels. Notably, LoopPrune identifies efficient iterative patterns that allow compressed models to match or even surpass the zero-shot reasoning accuracy of their dense counterparts on specific benchmarks, while maintaining near-lossless perplexity at high retention rates.


LucidNFT: LR-Anchored Multi-Reward Preference Optimization for Flow-Based Real-World Super-Resolution

Song Fei ⋅ Tian Ye ⋅ Sixiang Chen ⋅ Zhaohu Xing ⋅ Jianyu Lai ⋅ Lei Zhu

Generative real-world image super-resolution (Real-ISR) can synthesize visually convincing details from severely degraded low-resolution (LR) inputs, yet its stochastic sampling makes a critical failure mode hard to avoid: outputs may look sharp but be unfaithful to the LR evidence, exhibiting semantic or structural hallucinations. Preference-based reinforcement learning (RL) is a natural fit because each LR input yields a rollout group of candidate restorations. However, effective alignment in Real-ISR is hindered by three coupled challenges: (i) the lack of an LR-referenced faithfulness signal that is robust to degradation yet sensitive to localized hallucinations, (ii) a rollout-group optimization bottleneck where scalarizing heterogeneous rewards before normalization compresses objective-wise contrasts and weakens DiffusionNFT-style reward-weighted updates, and (iii) limited coverage of real degradations, which restricts rollout diversity and preference signal quality. We propose LucidNFT, a multi-reward RL framework for flow-matching Real-ISR. LucidNFT introduces LucidConsistency, a degradation-invariant and hallucination-sensitive LR-referenced evaluator trained with content-consistent degradation pools and original-inpainted hard negatives; a decoupled reward normalization strategy that preserves objective-wise contrasts within each LR-conditioned rollout group before fusion; and LucidLR, a large-scale collection of real-world degraded images for robust RL fine-tuning. Extensive experiments show that LucidNFT improves perceptual quality on strong flow-based Real-ISR baselines while generally maintaining LR-referenced consistency across diverse real-world scenarios.


M$^2$E-UAV: A Benchmark and Analysis for Onboard Motion-on-Motion Event-Based Tiny UAV Detection

Weiqi Yan ⋅ Lixin Chen ⋅ Xiangrui Hou ⋅ zhipeng cai ⋅ Youbiao Wang ⋅ Yangyang Shi ⋅ Yu Zang ⋅ Cheng Wang

Tiny UAV detection from an onboard event camera is difficult when the observer and target move at the same time. In this motion-on-motion regime, ego-motion activates background edges across buildings, vegetation, and horizon structures, while the UAV may appear as a sparse event cluster. To explore this practical problem, we present M$^2$E-UAV, a benchmark and analysis setup for onboard motion-on-motion event-based tiny UAV detection. The processed M$^2$E-UAV benchmark contains 87,223 training samples and 21,395 validation samples across four scene families: sunny building-forest, sunny farm-village, sunset building-forest, and sunset farm-village. We provide M$^2$E-Point, a point-based event baseline, and M$^2$E-Point + IMU, an IMU-conditioned variant, to analyze the role of inertial cues under onboard motion-on-motion detection. M$^2$E-Point encodes events as $[x,y,t,p]$ point sets, extracts local event structure with EdgeConv, and predicts event-level UAV foreground scores, from which bounding boxes are derived via DBSCAN. Our validation-stage analysis shows that point-based event modeling is a strong baseline, while simple IMU conditioning provides only marginal aggregate gains. Under the train/validation split, M$^2$E-Point achieves 0.9673 F1 and 0.5501 mAP50-95, while the IMU-conditioned variant reaches 0.5561 mAP50-95 with only marginal aggregate changes, serving as an initial baseline for future exploration in this domain. The dataset will be publicly released upon acceptance.


M$^2$RNN: Non-Linear RNNs with Matrix-Valued States for Scalable Language Modeling

Mayank Mishra ⋅ Shawn Tan ⋅ Ion Stoica ⋅ Joseph Gonzalez ⋅ Tri Dao

Transformers are limited to computations in the TC$^0$ complexity class, excluding tasks such as entity tracking and code execution that provably require greater expressivity. Motivated by this, we revisit non-linear Recurrent Neural Networks (RNNs) and introduce Matrix-to-Matrix RNN (M$^2$RNN): an architecture with matrix-valued hidden states and non-linear state transitions. We show that (i) the language modeling gap of non-linear RNNs is primarily a state-size gap, not a non-linearity penalty, and (ii) outer-product state expansion enables efficient tensor-core utilization. Empirically, M$^2$RNN achieves perfect state-tracking generalization at sequence lengths beyond training. These gains transfer to large-scale language modeling: at 7B MoE, Hybrid M$^2$RNN outperforms Hybrid Gated DeltaNetby $0.5$ perplexity points using $3\times$ smaller recurrent states, and replacing even a single recurrent layer with M$^2$RNN matches Hybrid M$^2$RNN accuracy with minimal throughput cost. Hybrid Gated DeltaNetwith a single M$^2$RNN layer also outperforms state-of-the-art hybrid linear-attention architectures by up to $8$ points on LongBench. Together, these results establish non-linear RNN layers as a compelling building block for efficient and scalable language models.


M$^\star$: Every Task Deserves Its Own Memory Harness

Wenbo Pan ⋅ Shujie LIU ⋅ Xiangyang Zhou ⋅ Xianlong Wang ⋅ Shiwei Zhang ⋅ Bingjing Xu ⋅ Wanlu Shi ⋅ Xiaohua Jia

Large language model agents rely on a memory harness to write, organize, retrieve, and use past experience. A harness that works for one task often fails on another, because conversations, embodied planning, and expert reasoning require different storage and retrieval behavior. To address this limitation, we introduce Mstar, a method that automatically discovers task-optimized memory programs through reflective code evolution. Specifically, Mstar operationalizes a memory harness as an executable Python memory program that defines the Schema, Logic, and Instruction, and optimizes these components jointly. We evaluate Mstar on four tasks covering conversation, embodied planning, and specialized reasoning. Our results demonstrate that Mstar improves performance over existing baselines across all evaluated tasks. Furthermore, the evolved memory programs exhibit structurally distinct processing mechanisms for each evaluated domain. These findings suggest that every task benefits from its own memory harness, and that memory-program search provides a concrete way to discover it.


Machine Unlearning in Low-Dimensional Feature Subspace

Kun Fang ⋅ Qinghua Tao ⋅ Junxu Liu ⋅ Yaxin Xiao ⋅ Yulin Jin ⋅ Qingqing Ye ⋅ Jian Sun ⋅ Haibo Hu

Machine Unlearning (MU) aims at removing the influence of specific data from a pretrained model while preserving performance on the remaining data. In this work, a novel perspective for MU is presented upon low-dimensional feature subspaces, which gives rise to the potentials of separating the remaining and forgetting data herein. This separability motivates our LOFT, a method that proceeds unlearning in a LOw-dimensional FeaTure subspace from the pretrained model through principal projections, which are optimized to maximally capture the information of the remaining data and meanwhile diminish that of the forgetting data. During unlearning, LOFT simply optimizes a small-size projection matrix flexibly plugged into the pretrained model, and only requires one-shot feature fetching from the pretrained model. LOFT demonstrates that competitive and efficient MU can be achieved in low-dimensional feature subspaces even with neither repeated raw-data access nor updates to the entire pretrained model. Extensive experiments validate the significantly lower computational overhead and superior unlearning performance of LOFT across diverse models, datasets, tasks, and applications. Code is anonymously available at https://anonymous.4open.science/r/4352/.


MANGO:Multi-Angle Neural Gated Operators for Chirp-Perturbed PDEs

Yunlong Zhu ⋅ Zunwei Fu ⋅ Zheng Wang ⋅ EUN-HU KIM

We introduce **MANGO** (Multi-Angle Neural Gated Operator), a spectral neural operator for chirp-perturbed PDEs whose solutions exhibit a local spectral content that rotates continuously with position. MANGO maintains multiple fractional Fourier transform (FRFT) branches whose angles are learnable parameters, optimized jointly with the spectral weights and a position-dependent softmax gate that combines branches at every spatial location. The architecture is matched to the structure of the problem class: the FRFT at angle $\alpha_0$ diagonalizes the chirp-perturbed PDE whose intrinsic angle is $\alpha_0$, and MANGO's learnable angles let the architecture discover $\alpha_0$ from data when it is unknown or varies across the domain—neither of which classical FRFT methods accommodate. We further introduce a Mihlin–Hörmander regularizer on the spectral weights of each non-Fourier branch, encouraging each branch to act as a bounded operator on $L^p$ for $1 < p < \infty$. We construct four chirp-perturbed PDE benchmarks with closed-form ground truth, positioned as identifiability tests. On these benchmarks, MANGO recovers the intrinsic spectral structure of each PDE from supervised solution data alone and achieves state-of-the-art performance against six neural-operator baselines.


Manifold-Aligned Adversarial Perturbation for Anti-Customization under Diffusion-based Purification

Ruotian Liu ⋅ Zhen Chen ⋅ Zhiguo Yang ⋅ Xiangyu Yin ⋅ Peipei Xu ⋅ Wenjie Ruan

Anti-customization techniques aim to protect visual content from unauthorized diffusion model personalization by adding imperceptible perturbations to training images. However, such protection can be significantly weakened by diffusion-based purification, which may remove adversarial perturbations before customization. We study representative anti-customization methods under diffusion purification and observe that existing methods exhibit similar off-manifold perturbation structures that are largely suppressed by purification. We then provide a principled geometric analysis that characterizes purification as a local contraction toward the data manifold, primarily suppressing perturbation components along normal directions. Motivated by this insight, we propose Manifold-Aligned Customization protection MAC, a practical purification-aware framework that optimizes downstream anti-customization objectives through a differentiable purifier and uses a structured perturbation parameterization to bias updates toward purification-retentive, manifold-aligned directions. Across experiments on CelebA-HQ, VGGFace2-HQ, and WikiArt under diffusion-based purification, MAC substantially improves protection effectiveness across most prompts and domains. Finally, frequency analysis and ImageNet and MS-COCO retention studies show that spectral signatures cannot predict purification outcomes, providing additional evidence for the proposed manifold-alignment perspective beyond a single data domain.


Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking

Yansen Han ⋅ Shengyi Liao ⋅ Yuanxing Zhang ⋅ Pengfei Wan ⋅ Tao Lin

Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify the transport trajectories but lack inherent constraints to pretrained data manifold, often pushing generated (terminal) samples off the pretrained support. We formalize this failure mode as manifold drift, a phenomenon where preference-aligned trajectories diverge from the pretrained support. Theoretically, we prove that while optimal flow matching exactly preserves the terminal manifold, existing methods like FlowDPO fail to do so whenever reward-driven updates contain components normal to the manifold surface. As a remedy, we propose ThermoDPO, a temperature-controlled objective that anchors pairwise preference optimization on preferred samples. Across temperature regimes, this objective connects rejection sampling fine-tuning and FlowDPO while upper-bounding a reconstruction-based surrogate for manifold drift. To counteract diminished signals at low temperatures, we further introduce a weighted variant, ThermoDPO-weighted, which we evaluate across all experiments for its superior trade-off between preference alignment and manifold preservation. Empirically, ThermoDPO-weighted demonstrates a better trade-off on synthetic manifolds than FlowDPO variants. On real-world image benchmarks, it improves OCR-oriented alignment over the pretrained model and FlowDPO variants while remaining competitive with strong baselines in held-out model-based and human evaluations without visible sample degradation.


Manifold Prior Guided Deep Unfolding for Hyperspectral Image Reconstruction

Shuqi Liu ⋅ Yingkai Zhang ⋅ Lin Gu ⋅ Tao Zhang ⋅ Ying Fu

Hyperspectral Image (HSI) reconstruction has made substantial progress with the deep unfolding framework by decomposing the problem into a projection module and a denoiser module. Nevertheless, existing methods still exhibit limitations of insufficient matching with HSI data. The issues lie in two aspects: 1) the data-consistency projection module applies a fixed gradient descent step while ignoring the spectral content adaptivity of HSI; 2)the denoiser module either suffers from high computational complexity or is limited to the local receptive field, failing to efficiently capture the global spatial-spectral dependencies. In this paper, we propose a deep unfolding framework centered on a two-stage manifold learning strategy to obtain a degradation-free structural prior, which is then strategically injected into both the projection and denoiser modules. In the data-consistency projection, a prior-guided spectral weight mechanism is introduced to adaptively modulate the weight of spectral band. In the denoiser module, a manifold-guided Mamba is proposed to achieve the structure-aware extraction of spatial-spectral features through prior modulation. Quantitative and visual comparisons on synthetic and real-world datasets show that the proposed method significantly outperforms existing state-of-the-art approaches, while maintaining the structural interpretability of the unfolding framework. To support reproducibility, the code will be publicly released upon acceptance.


Manifold Random Features

Ananya Parashar ⋅ Derek Long ⋅ Dwaipayan Saha ⋅ Krzysztof M Choromanski

We present a new paradigm for creating random features to approximate bi-variate functions (in particular, kernels) defined on general manifolds. This new mechanism of $\textit{Manifold Random Features}$ (MRFs) leverages discretization of the manifold and the recently introduced technique of $\textit{Graph Random Features}$ (GRFs) to learn continuous fields on manifolds. Those fields are used to find continuous approximation mechanisms that otherwise, in general scenarios, cannot be derived analytically. MRFs provide positive and bounded features, a key property for accurate, low-variance approximation. We show deep asymptotic connection between GRFs, defined on discrete graph objects, and continuous random features used for regular kernels. As a by-product of our method, we re-discover recently introduced mechanism of Gaussian kernel approximation applied in particular to improve linear-attention Transformers, considering simple random walks on graphs and by-passing original complex mathematical computations. We complement our algorithm with a rigorous theoretical analysis and verify in thorough experimental studies.


Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs

Seung-Hyun Lee ⋅ Dongyoon Han ⋅ Sangdoo Yun

Safety interventions on dual-use knowledge typically choose between destroying hazardous content (e.g., unlearning, filtering) and suppressing it at the output layer (e.g., refusal training); both pay a tax in adjacent-domain competence or over-refusal. We argue that the right operation is conditioning, not reduction: we show that hazardous knowledge can be retained in the model and behaviorally gated by a privileged control token. Our method, Token Inoculation, introduces a binding-and-branching approach. First, during continued pre-training, we mark hazardous contents by inserting a special token alongside dual-use documents, so the model binds the marker to the underlying semantics of the hazardous domain. Second, during supervised fine-tuning, we teach the model to answer hazardous queries correctly when the special token is present and to refuse them when it is absent, thereby enabling selective refusal without removing dual-use knowledge. On hazardous domain (e.g., WMDP-Bio), Token Inoculation reduces accuracy from 79% to 18% while retaining 93% of the base-model's benign-domain performance (e.g., MMLU), achieving the best safety-utility trade-off against unlearning and refusal-tuning baselines across 1B–14B model scales. We further show that refusal selectivity is controllable through the quality of the conditioning signal and that domain-specific semantic binding during pre-training is critical for the conditional behavior to generalize beyond memorized triggers. Our results suggest that safety alignment is better cast as a conditioning problem than a forgetting one: behavioral control is more precise when sensitive knowledge is retained under controlled access than when it is destroyed. We further show that refusal selectivity is controllable through the quality of the conditioning signal and that domain-specific semantic binding during pre-training is critical for the conditional behavior to generalize beyond memorized triggers. Our results suggest that safety alignment is better cast as a conditioning problem than a forgetting one: behavioral control is more precise when sensitive knowledge is retained under controlled access than when it is destroyed.


MASTARS: Multi-Agent Sequential Trajectory Augmentation with Return-Conditioned Subgoals

Jiwon Jeon ⋅ Myungsik Cho ⋅ Woojun Kim ⋅ Seongmin Kim ⋅ Woohyeon Byeon ⋅ Seungyul Han ⋅ Youngchul Sung

The performance of offline Multi-Agent Reinforcement Learning (MARL) is limited by the quality of the fixed offline dataset, motivating trajectory augmentation. However, augmenting multi-agent trajectories presents a fundamental challenge of the diversity–coordination trade-off: independent generation improves diversity but lacks coordination, while joint generation induces coordination but lacks diversity, reproducing the training distribution. To address this challenge in multi-agent data augmentation, we introduce MASTARS, a novel diffusion-based framework that generates multi-agent trajectories sequentially across agents while enforcing cross-agent consistent coordination via the method of inpainting. MASTARS produces trajectories that are diverse, coordinated and globally coherent. Experiments on various benchmarks such as MPE, SMAC, and SMAC-v2 demonstrate significant performance gains across multiple offline MARL methods.

Unit-modulus measurement vectors, due to their phase-only variations, exhibit favorable hardware compatibility and storage efficiency. In this paper, we consider matrix recovery using symmetric unit-modulus rank-one measurements. We identify a fundamental limitation of symmetric unit-modulus measurements---their inability to capture individual diagonal entries, and show that matrix recovery can be achieved using such measurements when the diagonal entries are known or directly measurable. We construct a stacked operator and establish that it satisfies a mixed-norm restricted isometry property ($\mathrm{mRIP}$) for unit-modulus measurements. Leveraging the $\mathrm{mRIP}$ condition, we derive exact and stable recovery results for matrix recovery with unit-modulus measurements in both noiseless and noisy settings. Our work is the first to provide the theoretical foundation for matrix recovery under symmetric rank-one unit-modulus measurements. Numerical experiments further corroborate these theoretical findings.


MDPBench: A Benchmark for Multilingual Document Parsing in Real-World Scenarios

Zhang Li ⋅ Lin Zhibo ⋅ Qiang Liu ⋅ Ziyang Zhang ⋅ Shuo Zhang ⋅ Zidun Guo ⋅ Jiajun Song ⋅ Jiarui Zhang ⋅ Xiang Bai ⋅ Yuliang Liu

We introduce Multilingual Document Parsing Benchmark, the first benchmark for multilingual digital and photographed document parsing. Document parsing has made remarkable strides, yet almost exclusively on clean, digital, well-formatted pages in a handful of dominant languages. No systematic benchmark exists to evaluate how models perform on digital and photographed documents across diverse scripts and low-resource languages. MDPBench comprises 3,400 document images spanning 17 languages, diverse scripts, and varied photographic conditions, with high-quality annotations produced through a rigorous pipeline of expert model labeling, manual correction, and human verification. To ensure fair comparison and prevent data leakage, we maintain separate public and private evaluation splits. Our comprehensive evaluation of both open-source and closed-source models uncovers a striking finding: while closed-source models (notably Gemini3-Pro) prove relatively robust, open-source alternatives suffer dramatic performance collapse, particularly on non-Latin scripts and real-world photographed documents, with an average drop of 17.8% on photographed documents and 14.0% on non-Latin scripts. These results reveal significant performance imbalances across languages and conditions, and point to concrete directions for building more inclusive, deployment-ready parsing systems.


Mean-Field Parallel Decoding for Discrete Diffusion Language Models

Tamim Zoabi ⋅ Ameen A Ali ⋅ Liran Ringel ⋅ Lior Wolf

Discrete diffusion language models enable parallel token generation, offering a pathway to low-latency decoding. However, selecting tokens independently by marginal confidence limits effective parallelism: tokens that appear reliable in isolation can form incompatible configurations when several positions are updated at once. We introduce a training-free decoding framework that coordinates these parallel updates. At each forward pass, the method assigns a commit score to each masked position and refines these scores using pairwise interactions derived from the model's predictive distributions. A variational relaxation yields a simple fixed-point update that suppresses conflicting simultaneous commitments within a single forward pass.This mechanism allows the decoder to commit more tokens in parallel while maintaining competitive generation quality. The method is lightweight, requires no auxiliary model or retraining, and drops into existing diffusion decoding pipelines without modification. Experiments on reasoning and code-generation benchmarks show consistent improvements in the quality-latency trade-off.


Mean Testing under Truncation beyond Gaussian

Yuhao Wang ⋅ Roberto I Oliveira ⋅ Themis Gouleakis

We characterize the fundamental limits of high-dimensional mean testing under arbitrary truncation, where samples are drawn from the conditional distribution $P(\cdot \mid S)$ for an unknown truncation set $S$ that may hide up to an $\varepsilon$-fraction of the probability mass. For distributions with $p$-th directional moments of magnitude at most $\nu_{P,p}$, truncation induces a bias of order $\mathcal{O}(\nu_{P,p}\varepsilon^{1-1/p})$. This bias creates a sharp information-theoretic detectability floor: when the signal $\alpha$ falls below this threshold, the null and alternative hypotheses are indistinguishable even with infinite data. Above this floor, we prove that a simple second-order test achieving near-optimal sample complexity $$ n = \mathcal{O}\left(\frac{\|\Sigma_P\|}{(\alpha-4\nu_{P,p}\varepsilon^{1-1/p})^2}\sqrt{d}\right). $$ We further identify a structural escape from this finite-moment bias barrier. Under a *directional median regularity* assumption, truncation bias improves to linear order $\mathcal{O}(\varepsilon)$. This reveals an intermediate regime in which estimation requires $\Theta(d)$ samples for uniform recovery, while testing recovers the classical $\Theta(\sqrt d)$ rate once truncation bias is eliminated. Together, our results provide a unified framework for mean testing under truncation, connecting finite-moment, sub-Gaussian, and median-regular structural regimes.


Mechanism Design for AI Overviews: Creator Incentives and Long-Term Profit

Yihang Wu ⋅ Jiajun Tang ⋅ Jinfei Liu ⋅ Haifeng Xu ⋅ Fan Yao

The integration of AI Overviews into search engines enhances user experience but diverts traffic from content creators, potentially discouraging high-quality content creation and causing user attrition that undermines long-term search engine profit. To address this issue, we propose a game-theoretic model of creator competition with costly effort, characterize equilibrium behavior, and design two incentive mechanisms: a \emph{citation mechanism} that references sources within an AI Overview, and a \emph{compensation mechanism} that offers monetary rewards to creators. For both cases, we provide structural insights for profit-maximizing mechanisms. Evaluations parameterized by real click data show that although AI Overviews harm long-term search engine profit, interventions based on our proposed mechanisms can increase long-term profit across a range of realistic scenarios, pointing toward a more sustainable trajectory for AI-enhanced search ecosystems.

Modeling the mechanisms of chemical reactions is central to the prediction and design of chemical reactivity. We present Electron Bookkeeping Transformer (Bookkeeper), a model that simulates chemical reactions at mechanistic resolution by generating electron movement trajectories. Bookkeeper explicitly models each electron as a token and predicts its movement sequentially, following the arrow pushing formalism in chemistry. This ensures mass and charge conservation during both elementary step and reaction pathway predictions by construction. Bookkeeper achieves higher accuracy in predicting both individual mechanistic steps and full reaction pathways, shows strong generalization to out-of-distribution reaction types, and runs significantly faster than previous flow matching-based and large language model-based methods. The accuracy and efficiency make Bookkeeper a powerful framework for virtual screening of chemical reactions across diverse reagents, for validating computer-assisted synthesis planning predictions, and for providing mechanistic interpretability of reactivity; the utility of Bookkeeper as a tool will most likely grow as does its training corpus.


MedCache: Training-Free Spatially Aware Caching for Accelerated Medical Video Generation

Ufaq Khan ⋅ Umair Nawaz ⋅ Sathira Silva ⋅ Numan Saeed ⋅ Muhammad U Sheikh ⋅ Muhammad Bilal ⋅ Junaid Qadir ⋅ Lena Maier-Hein ⋅ Yutong Xie ⋅ Muhammad Haris Khan

Video world models are now being adapted to medical applications such as surgical simulation, procedural rehearsal, and synthetic data generation, but their inference cost remains a major obstacle to practical use. Training-free caching is attractive because it can be applied to existing checkpoints, yet current methods face a clear trade-off. Lightweight rules track only the average change in the video and treat all regions of the frame as equally important, which is a poor fit for medical video where the meaningful motion is often confined to a small region such as a tool tissue interaction. Richer rules recover spatial awareness by partially running the network on every step to inform the skip decision, but this decision cost is paid whether the step is skipped or not, and it eats into the speedup. We introduce \textbf{MedCache}, a training-free and probe-free cache for rectified-flow video generation. MedCache reads a spatial risk signal directly from the latent using cheap arithmetic, with no help from the network itself. It combines this risk signal with a prompt-derived domain prior and a lazy consistency check that runs only when reuse is no longer clearly safe. Across three open-world model families in Text2World and Image2World settings, MedCache improves both speed and quality over the strongest training free baselines in the medical regime it targets. On Cosmos-H-Surgical 2B Text2World, it raises Total quality from 0.767 to 0.837 over EasyCache while reducing latency from 185 seconds to 95 seconds. Ablations show that risk-weighted drift drives the runtime gain, while the prior and guard recover fidelity.


MedHEB: Benchmarking Medical Embeddings Across Heterogeneous Clinical Evidence

Yingshu Li ⋅ Shaoyang Zhou ⋅ Zhanyu Wang ⋅ YUNYI LIU ⋅ Xinyu Liang ⋅ Xi Zhang ⋅ Lingqiao Liu ⋅ Lei Wang ⋅ Luping Zhou

Clinical evidence is inherently relational: medical images are interpreted with prior studies, cross-modal examinations, reports, and localized findings, and retrieving relevant evidence is therefore central to clinical interpretation. Recent foundation models increasingly encode these heterogeneous sources into shared embedding spaces, yet existing embedding evaluation often focuses on isolated matches, leaving unclear whether current models retrieve relevant evidence across changes in modality, appearance, and granularity. We introduce MedHEB, a unified benchmark with $107$ retrieval tasks from $54$ datasets, including both in-distribution and held-out-source evaluation, across three axes: evidence form (2D images, 3D volumes, text), pairing type (image--image, image--text, text--text, and modality-bridging visual matching), and relevance granularity (global, question-conditioned, region-level). Under a single ranking protocol, MedHEB varies clinically relevant evidence definitions to test whether embedding proximity remains meaningful across heterogeneous evidence, including modality-bridging visual retrieval and region-level text-to-region retrieval. Across representative models, performance is uneven across retrieval families: models strong on conventional image--text or text--text retrieval are less reliable on modality-bridging and region-level tasks. Current embedding spaces capture some clinical relationships well but do not yet provide balanced retrieval across heterogeneous evidence. We further provide MedEmb, an MLLM-based reference embedding model trained on the MedHEB ID split, as a reproducible baseline for unified 2D--3D--text medical retrieval. We release MedHEB with preprocessing scripts, evaluation code, access manifests, and MedEmb to support research on clinically meaningful medical embeddings.


Median-of-Means under Structured Heavy-Tailed Noise: High-Probability Bounds for Clipped Stochastic Optimization

Ahmed El Bajdali ⋅ Ohad Shamir ⋅ Samuel Horváth ⋅ Eduard Gorbunov

Optimization under heavy-tailed stochastic noise remains a major challenge in modern machine learning, where classical moment assumptions often fail to capture realistic gradient distributions. In this work, we revisit robust estimation techniques and provide a systematic analysis of the Median-of-Means (MoM) estimator under a structured noise model that interpolates between a symmetric heavy-tailed distribution with bounded $\beta$-th moment ($\beta \in (0, 1]$) and a mean-zero distribution with bounded $\alpha$-th moment ($\alpha \in (1, 2]$). We derive new upper bounds on the variance and bias of MoM in this setting, thereby extending the applicability of existing results beyond purely symmetric or bounded-moment settings. Building on these properties, we study clipped stochastic optimization methods with MoM gradient estimation and establish high-probability convergence guarantees that hold for general $\beta \in (0, 1]$ and $\alpha \in (1, 2]$. Our bounds match or improve upon prior complexity guarantees while requiring weaker distributional assumptions, and they reveal that bias only affects the final neighborhood of convergence. Overall, this work advances the theoretical foundations for robust stochastic optimization in the presence of extreme gradient noise.

Many agent-memory evaluations collapse state revision, deprecation, erasure, and provenance into retrieve-heavy scores. We introduce MemContract, a diagnostic benchmark of 1 , 080 multi-session tasks over store, retrieve, revise, deprecate, erase, and prove-origin, with formal pre/post-conditions, mandatory distractor sessions plus alias rotation on every family, blinded/stale/delayed controls, a fixed 25-probe residue audit, and benchmark-validity checks via a 180-task human contract-fit audit and a 120-task human-assisted upper bound. Under matched prompts, compute, storage, and latency budgets on GPT-4o across seven architectures, Graph Memory reaches 67.6% macro average (95% paired-bootstrap CI [65.8, 69.4]) versus 62.4% for the strongest hybrid baseline and 40.6% for RAG, with most separation concentrated in revise, deprecate, erase, and prove-origin while store/retrieve remain comparatively compressed. Claude 3.5 Sonnet and Gemini 1.5 Pro 002 preserve the same top-three ordering on a matched 360-task rerun, while the manual audit finds 96.7% agreement that the harness label matches the intended contract and the human-assisted upper bound reaches 93.8% macro, placing the best-system score in a solvable regime rather than a floor of label noise. Reviewer-critical calibration checks tell a narrower but durable story: removing prove-origin leaves the Graph - Hybrid gap unchanged at 5.2 pp, best-effort prompt/schema tuning narrows it to 2.6 pp, and fully external LongMemEval wu2024longmemeval and LoCoMo maharana2024locomo narrow it to 1.8 and 2.8 pp while preserving the top-three ordering Graph > Hybrid > Editable KV. The stable claim is therefore not universal backend superiority: contract-sensitive evaluation reveals separation that broader or retrieval-heavier workloads attenuate but do not erase. The 25-probe audit bounds catchable residue leakage at 3.9% for Graph Memory and 27.4% for RAG, but does not certify deletion.


Memorize Theorems, Not Instances: Probing SFT Generalization through Mathematical Reasoning

Ruiying Peng ⋅ Mengyu Yang ⋅ Jing Lei ⋅ Xiao-Hui Li ⋅ Xueyu Wu ⋅ Xinlei Chen

Supervised Fine-Tuning (SFT) is widely used for task-specific adaptation, yet recent work shows it systematically undermines reasoning generalization. We argue the root cause is not memorization itself, but its target: vanilla SFT drives models to exploit and memorize spurious surface correlations in problem-solution pairs, leaving them brittle to superficial input variations. To address this, we propose Theorem-SFT, which reorients supervision toward explicit theorem application by teaching models how rules are invoked rather than what answers look like. Theorem-SFT yields consistent gains across benchmarks and model families: +8.8\% on MATH (LLaMA3.2-3B-Instruct) and +20.27\% on GeoQA (Qwen2.5-VL-7B-Instruct) without modality-specific re-training. Fine-tuning MLP layers alone matches full-layers performance, implicating feed-forward components as the primary locus of reasoning rules. Our findings reframe the debate: Generalization failures stem not from memorization as a mechanism, but from memorizing the wrong inductive targets.


Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation

Yanjun Guo ⋅ Zhengqiang ZHANG ⋅ Pengfei Wang ⋅ Xinyue Liang ⋅ Zhiyuan Ma ⋅ Lei Zhang

Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory modeling with video generation, leading to inconsistent content during scene revisits and diminished generative capacity when exploring novel regions, even trained on extensive annotated data. To address these limitations, we propose a decoupled framework that separates memory conditioning from generation. Our approach significantly reduces training costs while simultaneously enhancing spatial consistency and preserving the generative capacity for novel scene exploration. Specifically, we employ a lightweight, independent memory branch to learn precise spatial consistency from historical observation. We first introduce a hybrid memory representation to capture complementary temporal and spatial cues from generated frames, then leverage a per-frame cross-attention mechanism to ensure each frame is conditioned exclusively on the most spatially relevant historical information, which is injected into the generative model to ensure spatial consistency. When generating new scenes, a camera-aware gating mechanism is proposed to mediate the interaction between memory and generation modules, enabling memory conditioning only when meaningful historical references exist. Compared with the existing method, our method is highly data-efficient, yet the experiments demonstrate that our approach achieves state-of-the-art performance in terms of both visual quality and spatial consistency.


Mental Imagery That Matters: Curing Latent Collapse in Multimodal Reasoning

Romy Luo ⋅ Rohan Dugad ⋅ Zihui (Sherry) Xue ⋅ Alex Dimakis ⋅ Kristen Grauman

Mental simulation, the ability to reason by constructing and updating internal visual states, is central to flexible human cognition, but remains elusive in current multimodal large language models. Latent multimodal reasoning offers a promising route toward this capability by interleaving text with continuous visual latent blocks, allowing models to maintain internal visual workspaces without decoding them into pixels. In this paper, we ask whether these latent visual states actually matter. Through a direct perturbation analysis across representative latent-reasoning MLLMs, we find that they often do not: dropping, replacing, or shuffling generated visual latents has only minor effects on the model's own final-answer probability, while equivalent perturbations to the linguistic reasoning trace are substantially more damaging. We call this failure mode Mental Simulation Collapse (MSC): the model appears to perform interleaved visual reasoning, but routes most answer-relevant computation through text. We further diagnose that MSC follows naturally from existing training objectives: visual alignment can make generated latents resemble helper-image embeddings, but it does not require the decoder to rely on them; subsequent text-only supervision further leaves the semantic role of the latent pathway unidentifiable. To address this, we introduce AnchorSim (Anchored Mental Simulation), a contrastive distillation framework that trains models to make their generated visual latents consequential. Across multiple visual reasoning benchmarks, AnchorSim improves performance, while diagnostic re-evaluation shows that visual perturbations now meaningfully affect predictions. Our results expose a fundamental failure mode of latent multimodal reasoning and demonstrate that mental simulation must be anchored to model behavior in order to matter.

Recent advances in video generative models have enabled flexible conditional generation and editing pipelines, such as text-to-video, image-to-video, and reference-guided synthesis. However, these workflows are inherently one-directional, mapping conditions to videos without providing a mechanism to invert this process. In complex media production settings, the ability to recover controllable textual conditions from a given video is highly desirable, yet remains underexplored. While inverse prompting has been studied for image generation, directly extending these approaches to video is challenging due to the need to capture temporal dynamics beyond static appearance, as well as the substantial computational cost of per-video optimization. In this work, we propose the first practical framework for efficient video inverse prompting. Specifically, we formulate inverse prompting as a meta-learning problem and introduce Meta Video Inverse Prompting (MVIP), which learns a meta-prompt optimized over temporal dimensions to enable fast and effective prompt inference. Our approach amortizes the inversion process, eliminating the need for costly iterative optimization at test time. Experimental results demonstrate that MVIP recovers higher-quality inverse prompts with significantly improved reconstruction fidelity compared to image-based inversion and iterative captioning baselines, while achieving substantial gains in efficiency.


Mimicking the Physicist's Eye: A VLM-centric Approach for Physics Formula Discovery

Jiaqi Liu ⋅ Songning Lai ⋅ Pengze Li ⋅ Di Yu ⋅ Zhou wenjie ⋅ Yiyang Zhou ⋅ Peng Xia ⋅ Zijun Wang ⋅ Xi Chen ⋅ SHIXIANG TANG ⋅ LEI BAI ⋅ Wanli Ouyang ⋅ Mingyu Ding ⋅ Huaxiu Yao ⋅ Aoran Wang

Automated discovery of governing equations remains challenging when models rely only on numerical tokens and ignore the phase-space and trajectory evidence that is informative in low-dimensional dynamics. We propose Visual Induction for Physics-based Equation Reasoning (VIPER-R1), a multimodal framework for visual formula discovery from phase portraits, trajectory plots, and aligned measurements. VIPER-R1 uses a supervised Motion Structure Induction (MSI) curriculum for joint rationale-and-equation generation, and is then refined by \textbf{Reward-Guided Symbolic Calibration (RGSC)}, a GRPO-based stage optimized with a structure-aware reward that emphasizes topological correctness over coefficient matching. At inference time, we optionally apply Symbolic Residual Realignment (SR$^2$), an external symbolic regression module that models only the residual numerical mismatch of the base hypothesis. We release PhysSymbol, a 10{,}000-instance two-part multimodal corpus; the quantitative evaluation in this paper uses its Part 1 mechanics split, which provides paired visualizations, trajectory data, and expert C-CoT traces. On this benchmark, VIPER-R1-7B achieves the best structural score ($S_{\mathrm{struct}}=0.812$), is effectively tied with VIPER-R1-3B on exact-match accuracy ($S_{\mathrm{acc}}=0.487$ vs.\ $0.488$), and, after optional SR$^2$ refinement, reaches a Post-SR$^2$ MSE of $0.032$, nearly $3\times$ lower than the best VLM baseline. These results show that VIPER-R1 improves symbolic structure induction in focused low-dimensional dynamics by combining paired visual evidence with structure-aware calibration to produce stronger first-order hypotheses for downstream refinement.


MineEvolve: Self-Evolution with Accumulated Knowledge for Long-Horizon Embodied Minecraft Agents

Zhengwei Xie ⋅ Zhisheng Chen ⋅ Ziyan Weng ⋅ Jinhan Li ⋅ Chenglong Li ⋅ Zikai Xiao ⋅ Jinhao Jing ⋅ Jingwei Song ⋅ Guibin Zhang ⋅ Kun Wang

Long-horizon embodied intelligence requires agents to improve through interaction, not merely to execute plans generated from static goals. A central challenge is therefore to transform past executions into knowledge that can shape future decisions. Minecraft provides a representative testbed for this problem, where tasks such as crafting tools, building redstone components, and obtaining diamond equipment involve long prerequisite chains and are frequently disrupted by missing tools, blocked paths, GUI failures, or stagnant execution. To this end, we propose \textbf{MineEvolve}, a knowledge-driven self-evolution framework that converts execution feedback into actionable behavioral knowledge. MineEvolve first uses \underline{\emph{\textbf{\ding{182}Monitor}}} to convert each subgoal execution into typed feedback, including state changes, inventory changes, failure types, progress signals, and stagnation indicators. \underline{\emph{\textbf{\ding{183}Inducer}}} then derives reusable skills from successful executions and remedies from failed or stagnant executions. \underline{\emph{\textbf{\ding{184}Curator}}} validates, merges, filters, and retrieves these knowledge entries, while \underline{\emph{\textbf{\ding{185}Adaptor}}} uses them to repair the unfinished part of the plan under repeated failures or stagnation. Experiments on the Minecraft MCU long-horizon task suite show that MineEvolve consistently improves performance across multiple language-model planners, with larger gains on high-dependency task groups. Ablation and knowledge-accumulation studies further demonstrate that converting execution signals into structured behavioral knowledge is an effective path toward self-evolving embodied agents in long-horizon environments. Our code is available at \url{https://anonymous.4open.science/r/MineEvolve-1B5B}.

Tree ensembles such as random forests (RFs) and gradient boosting machines (GBMs) are among the most widely used supervised learners, yet their theoretical properties remain incompletely understood. We adopt a spectral perspective on these algorithms, with two main contributions. First, we derive minimax-optimal convergence for RF regression, showing that, under mild regularity conditions on tree growth, the eigenvalue decay of the induced kernel operator governs the statistical rate. Second, we exploit this spectral viewpoint to develop compression schemes for tree ensembles. For RFs, leading eigenfunctions of the kernel operator capture the dominant predictive directions; for GBMs, leading singular vectors of the smoother matrix play an analogous role. Learning nonlinear maps for these spectral representations yields distilled models that are orders of magnitude smaller than the originals while maintaining competitive predictive performance. Our methods compare favorably to state of the art algorithms for forest pruning and rule extraction, with applications to resource constrained computing.

We show that the MinMax algebra provides a form of recurrence that is expressively powerful, efficiently implementable, and most importantly it is not affected by vanishing or exploding gradient. We call MinMax Recurrent Neural Cascades (RNCs) the models obtained by cascading several layers of neurons that employ such recurrence. We show that MinMax RNCs enjoy many favourable theoretical properties. First, their formal expressivity includes all regular languages, arguably the maximal expressivity for a finite-memory system. Second, they can be evaluated in parallel with a runtime that is logarithmic in the input length given enough processors—and they can also be evaluated sequentially. Third, their state and activations are bounded uniformly for all input lengths. Fourth, at almost all points, their loss gradient exists and it is bounded. Fifth, they do not exhibit a vanishing state gradient: the gradient of a state w.r.t. a past state can have constant value one regardless of the time distance between the two states. Finally, we find empirical evidence that the favourable theoretical properties of MinMax RNCs are matched by their practical capabilities: they are able to perfectly solve a number of synthetic tasks, showing superior performance compared to the considered state-of-the-art recurrent neural networks; also, we train a MinMax RNC of 127M parameters on next-token prediction, and the obtained model shows competitive performance for its size, providing evidence of the potential of MinMax RNCs on real-world tasks.


Mirror Descent-Ascent for mean-field min-max problems

Razvan-Andrei Lascu ⋅ Mateusz Majka ⋅ Lukasz Szpruch

We study two variants of the mirror descent-ascent (MDA) algorithm for solving min-max problems on the space of measures: simultaneous and alternating. We work under assumptions of convexity-concavity and relative smoothness of the payoff function with respect to a suitable Bregman divergence, defined on the space of measures via flat derivatives. We establish non-asymptotic convergence rates to mixed Nash equilibria, measured in the Nikaido-Isoda error, proving an $\mathcal{O}(N^{-1/2})$ rate for simultaneous MDA and an improved $\mathcal{O}(N^{-2/3})$ rate for alternating MDA. The main technical contribution is an infinite-dimensional dual space analysis that relates Bregman divergences on measures to dual Bregman divergences on spaces of bounded continuous functions, allowing us to control asymmetric commutator terms created by alternating updates. The results substantially generalize prior analyses restricted to bilinear objectives and also apply to nonlinear convex-concave problems on measure spaces, thereby providing a unified theoretical foundation for MDA in mean-field min-max optimization.

Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared design. Existing mitigation approaches are largely tied to specific economic settings, like two-sided platforms and auctions, leaving open how to design interventions for general repeated games. We address this gap by formalizing the connection between empirical observations from prior work on Q-learning collusion and classical theory of Simple Penal Codes (SPCs). We show any non-trivial SPC induces a quantifiable conditional dependence in agents' policies, detectable via the total variation distance between an agent's action distributions across cooperation and defection histories. Building on this connection, we propose CURB (Collusion Unwinding via Reward shaping and Belief injection), a reward-shaping framework that penalizes this Total Variation (TV) distance signal during Q-learning and is guaranteed to convert any SPC fixed point of the dynamics into a trivial one, thus precluding collusive equilibria sustained by punishment threats. Empirically, CURB substantially reduces collusion by Q-learning agents in both Bertrand and Cournot Competition Repeated Games. We further demonstrate that CURB extends to deep Q-network agents in Bertrand competition, suggesting the mechanism generalizes beyond tabular Q-learning.


Mixture of Activations: Token-Adaptive Mixing for Expressive Feedforward Layers

Mingze Wang ⋅ Jinbo Wang ⋅ Yikuan Xia ⋅ Shen Kai ⋅ Mingwu Zheng

Feedforward network (FFN) layers account for a large fraction of parameters and nonlinear expressivity in Transformer-based large language models (LLMs). Despite the evolution from ReLU and GELU to gated variants such as SwiGLU, most FFN designs still use a single fixed activation function, applying the same nonlinear transformation to all tokens. In this work, we propose Mixture of Activations (MoA), a token-adaptive FFN design that mixes a dictionary of activation functions using lightweight input-dependent gates while sharing the same linear projections. As an input-independent counterpart, we also introduce learnable activations (LA), which form linear combinations of activation functions for both ReLU-type and SwiGLU-type FFNs. Theoretically, we establish strict finite-width expressive separations among fixed-activation FFNs, LA, and MoA: LA strictly contains fixed-activation FFNs, while MoA strictly contains LA, with the additional expressivity arising from input-dependent nonlinear hybridization. Empirically, we evaluate MoA through extensive pre-training experiments on dense and MoE language models ranging from 0.12B to 2B parameters under different token budgets, optimizers, and learning rate schedules. MoA consistently achieves lower terminal loss and exhibits more favorable scaling behavior than well-tuned baselines, with minimal parameter and computational overhead. These results suggest that token-adaptive activation mixing is a simple and effective mechanism for improving FFN expressivity in LLMs.


Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models

Duc Anh Nguyen ⋅ Tien N Luu ⋅ Tung Pham ⋅ Toan Tran

State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls into states rather than model weights, offering substantial memory savings while retaining parameter efficiency. However, most existing state-based methods typically apply only per-block control updates, which limits inter-block information exchange and restricts representational adaptation. Meanwhile, prior mechanisms that enable cross-block communication often introduce considerable computational overhead, reducing their practicality for efficient fine-tuning. We introduce Mixture-of-Control (MoC), a lightweight fine-tuning framework that adaptively integrates local and global control signals to enhance representation learning. MoC treats block-wise control states as experts in a sparse mixture-of-experts process, enabling efficient communication across transformer blocks. Experiments on 33 datasets and 10 pretrained models across diverse transformer-based benchmarks demonstrate that MoC consistently outperforms existing state-based fine-tuning methods while maintaining comparable efficiency in memory and compute.


MMLongCite: A Benchmark for Evaluating Faithfulness of Long-Context Vision-Language Models

Keyan Zhou ⋅ Zecheng Tang ⋅ Lingfeng Ming ⋅ Qiguang Chen ⋅ WangJie You ⋅ Guanghao Zhou ⋅ Dan Qiao ⋅ Zheming Yang ⋅ Libo Qin ⋅ Minghui Qiu ⋅ Juntao Li ⋅ Min zhang

The rapid advancement of long-context vision language models (LCVLMs) has led to a significant expansion of their context windows. However, an extended context window does not guarantee the effective utilization of the context, posing a critical challenge for real-world applications. Current evaluations of such long-context faithfulness in multimodal settings remain limited to short contexts. To bridge this gap, we introduce MMLongCite, the first benchmark evaluating the faithfulness of LCVLMs via multimodal citation generation. MMLongCite features 2,280 examples across 8 tasks and diverse modalities (image, video, interleaved), with context lengths scaled from 16K to 128K tokens. To test spatial localization capabilities of LCVLMs, we also introduce MMLongCite-HR, evaluating fine-grained visual grounding amidst dense pixel spaces. Through extensive benchmarking of cutting-edge LCVLMs, we provide a systematic analysis of current multimodal citation capabilities. Our results reveal a significant discrepancy between answer correctness and citation faithfulness. We also conduct attention pattern investigations and in-depth error analyses to reveal the underlying phenomena of failures in LCVLMs. MMLongCite establishes a rigorous foundation for diagnosing and advancing the faithfulness of LCVLMs. We hope our findings provide meaningful insights to drive further improvements in the long-context capabilities of LCVLMs.


MM-OptBench: A Solver-Grounded Benchmark for Multimodal Optimization Modeling

Zhong Li ⋅ Qi Huang ⋅ Yuxuan Zhu ⋅ Mohammad Mohammadi Amiri ⋅ Niki van Stein ⋅ Thomas Bäck ⋅ Matthijs van Leeuwen ⋅ Zaiwen Wen ⋅ Lincen Yang

Optimization modeling translates real decision-making problems into mathematical optimization models and solver-executable implementations. Although language models are increasingly used to generate optimization formulations and solver code, existing benchmarks for this capability are almost entirely text-only. This omits many optimization-modeling tasks that arise in operational practice, where requirements are described in text but essential instance information is often conveyed through visual artifacts such as tables, graphs, maps, schedules, and planning dashboards. We introduce \emph{multimodal optimization modeling}, a benchmark setting in which models must construct both a mathematical formulation and executable solver code from a text-and-visual problem specification. To evaluate this setting, we develop a solver-grounded framework that generates structured optimization instances, verifies each instance with an exact solver, and builds both the model-facing inputs and the hidden reference files from the same verified source. We instantiate the framework as MM-OptBench, a benchmark of 780 solver-verified instances spanning 6 optimization families, 26 subcategories, and 3 structural difficulty levels. We then evaluate 9 multimodal large language models (MLLMs), including 6 frontier general-purpose models and 3 math-specialized models, with aggregate, family-level, difficulty-level, and failure-mode analyses. The results show that the task remains far from solved: the best two models reach 52.1\% and 51.3\% pass@1, while on average across the six general-purpose MLLMs, pass@1 is 43.4\% on easy instances and 15.9\% on hard instances. All three math-specialized MLLMs solve 0/780 instances. Failure attribution shows that errors arise both when extracting instance data from text and visuals and when turning extracted data into solver-correct formulations and code. MM-OptBench provides a concrete testbed for solver-grounded, decision-oriented multimodal intelligence.


MobileWan: Closing the Quality Gap for Mobile Video Diffusion

Mohsen Ghafoorian ⋅ Denis Korzhenkov ⋅ Adil Karjauv ⋅ Ioannis Lelekas ⋅ Noor Fathima ⋅ Spyridon Stasis ⋅ Animesh Karnewar ⋅ Amirhossein Habibian

Recent advances in video diffusion have been driven by scaling transformer-based architectures to billions of parameters, substantially improving visual fidelity and motion coherence. In contrast, existing mobile video diffusion models remain limited to relatively small parameter budgets, typically 0.4–1.8B, restricting generation quality. In this work, we show that high-quality mobile video generation does not require small models. Instead, we demonstrate that a server-scale 5B-parameter video diffusion transformer can be deployed efficiently on memory-constrained mobile hardware through recurrent reformulation and structured compression. Starting from Wan2.2-5B, we rely on a recurrent distillation framework that converts video generation into a chunk-wise autoregressive process with constant-memory attention computation. Combined with causal linear attention, the model operates as an RNN at inference time while preserving temporal coherence across chunks. We further propose a learnable attention head pruning method based on binary per-head gates optimized end-to-end using a noise-biased sparsity objective and distillation-based finetuning. Together with sampling-step distillation and memory-optimized VAE decoding, MobileWan becomes the first 5B-scale video diffusion model deployable on a commercial mobile device. Our system generates 5-second 480x832 videos at 16 FPS in 20 seconds end-to-end latency, achieving a VBench score of 83.79 and establishing a new state of the art in mobile video generation. The model checkpoint will be released.


Modality-Depth Routing for Visual Reasoning in VLM Post-Training

Yiming Ren ⋅ Yiran Xu ⋅ Chufan Shi ⋅ Yu Qiao ⋅ junjie wang ⋅ Yujiu Yang

The prevailing route to stronger visual reasoning in vision-language models (VLMs) builds costly Chain-of-Thought corpora or curated step-by-step rationales, concentrating progress in well-resourced industry labs. We ask whether the reasoning signal already present in generic visual instruction data can instead be unlocked through a structural change to the post-training interface itself. At fixed data and capacity, standard supervised fine-tuning (SFT) on a generic instruction mixture partially improves reasoning but homogenizes the modality-depth pathways through which reasoning and perception inputs flow. We isolate two coupled patterns. First, visual and textual representations couple at the deepest layer: cross-modal Centered Kernel Alignment (CKA) rises by roughly 12% over the frozen baseline. Second, perception and reasoning inputs are routed through similar late-depth profiles, leaving no mechanism for per-input depth allocation. We propose Modality-Depth Attention (MDA), a lightweight side pathway that adds modality-specific query projections and a sigmoid-gated learnable depth residual. Under a matched protocol on Qwen3-VL-2B, MDA improves the reasoning average over SFT by +3.9 and keeps perception almost intact, at 0.05% extra parameters. Multiple mechanistic probes converge on the same routing account. The gains extend to Qwen3-VL-8B and transfer to LLaVA-OneVision-7B, suggesting that modality-depth routing may be a general lever for unlocking reasoning from generic instruction data without compromising perception. Code will be released upon publication.


Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning

Wenxi Gao ⋅ Guanxi Lu ⋅ Didi Zhu ⋅ Hao Chen ⋅ Quan Deng ⋅ Zhican Wang ⋅ Jiankang Deng ⋅ Hongxiang Fan

Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps into intermediate reasoning traces, have demonstrated great potential for visual mathematical reasoning tasks. However, we identify a key insight in this paradigm: generating intermediate visual reasoning steps is not always beneficial and can even be harmful, as self-generated visual steps may introduce erroneous visual evidence that misleads subsequent reasoning. Moreover, frequently triggering visual steps during reasoning incurs substantial computational and memory overhead, degrading inference efficiency. To address these accuracy and efficiency challenges, we observe that the model's internal signals can indicate whether a visual step will benefit reasoning before the entire visual generation is completed. Specifically, this work identifies two internal signals: i) Generation Intent, which reflects whether the model has a concrete textual plan for what to draw, and ii) Visual Fidelity, which measures whether the visual generation remains grounded in the original input image. Leveraging these internal signals, we propose AdaViG, a training-free adaptive visual gating method for unified multimodal reasoning. AdaViG dynamically evaluates each triggered visual step at an early visual generation stage and aborts it when both signals are weak, thereby preventing misleading visual evidence from entering the reasoning trace while avoiding unnecessary computation. Comprehensive experiments demonstrate that AdaViG improves accuracy by up to 5.7\% while reducing visual generation FLOPs by 25.0\%-91.0\% and wall-clock latency by 15.4\%-45.6\%.


Modeling the Vividness of Imagined Natural Scenes Reveals a Model-Common Image-Level Component in Vision Models

Morteza Mahdiani ⋅ Catherine Landry ⋅ Jasper JF van den Bosch ⋅ Frédéric Gosselin ⋅ Ian Charest

The visual features that support vivid mental images remain poorly understood. Here, we collected more than 229,000 vividness judgments for the 73,000 images in the Natural Scenes Dataset in a large online experiment (n = 1,991 participants). On each trial, participants viewed two images, were cued to mentally recreate one of them for 4 seconds and then rated the vividness of their mental image on a continuous scale from 0 to 100, followed by a memory task to encourage compliance. Using these data, we trained lightweight MLP readout heads on frozen vision-model representations to predict imagery vividness directly from image content. Several models predicted human vividness judgments above a low-level features baseline, with the best individual model reaching r = 0.376 [0.355, 0.396]. A single principal component of model predictions captured most of the cross-model variance and matched or exceeded individual models in predicting human judgments (r = 0.378). Even after removing human-aligned variance, residual predictions remained highly correlated across architectures, indicating a strong shared structure in model predictions. Attempts to extract additional human-aligned signal through residual probes, fresh layer-cached probes, and fine-tuning failed; fine-tuning instead further amplified the shared component. This shared axis also correlated more strongly with image memorability than with vividness itself, while observer-aware models generalized poorly to held-out raters, suggesting that models rely on a generic image-affordance signal rather than a vividness-specific representation. Together, these results show that subjective imagery vividness is partially predictable from visual content alone, while revealing a systematic gap between predictive performance and construct-specific representation in current vision models.


ModelLens: Finding the Best for Your Task from Myriads of Models

Rui Cai ⋅ Wenjie Mo ⋅ Xiaofei Wen ⋅ Qiyao Ma ⋅ Wenhui Zhu ⋅ Xiwen Chen ⋅ Muhao Chen ⋅ Zhe Zhao

The open-source model ecosystem now contains hundreds of thousands of pretrained models, yet picking the best model for a new dataset is increasingly infeasible: new models and unbenchmarked datasets emerge continuously, leaving practitioners with no prior records on either side. Existing approaches handle only fragments of this in-the-wild setting: AutoML and transferability estimation select models from small predefined pools or require expensive per-model forward passes on the target dataset, while model routing presupposes a given candidate pool. We introduce ModelLens, a unified framework for model recommendation in the wild. Our key insight is that public leaderboard interactions, though scattered and noisy, collectively trace out an implicit atlas of model capabilities across heterogeneous evaluation settings, a signal rich enough to learn from directly. By learning a performance-aware latent space over model--dataset--metric tuples, ModelLens ranks unseen models on unseen datasets without running candidates on the target dataset. On a new benchmark of 1.62M evaluation records spanning 47K models and 9.6K datasets, ModelLens surpasses baselines that either rely on metadata alone or require running each candidate on the target dataset. Its recommended Top-K pools further improve multiple representative routing methods by up to 81\% across diverse QA benchmarks. Case studies on recently released benchmarks confirming generalization to both text and vision-language tasks.


Modulating Merging Strengths via Joint Loss Estimation for LoRA-based Continual Learning

Kun Gu ⋅ De Cheng ⋅ Zhipeng Xu ⋅ Lingfeng He ⋅ Di Xu ⋅ Nannan Wang

Continual learning (CL) aims to adapt models to new tasks while preserving previously learned knowledge. Recently, Low-Rank Adaptation (LoRA), a representative Parameter-Efficient Fine-Tuning (PEFT) method, has gained increasing attention in CL due to its efficiency and scalability. Several LoRA-based merging methods maintain a single inference model by scaling the newly learned LoRA update with one coefficient and merging it into the previous weights after each task. However, they usually control the whole update with a shared scalar factor, although different update elements may affect previous and current tasks in different ways. To address this issue, we propose Merging Strength Modulation via Joint Loss Estimation (M²LE) It consists of Hessian Information-guided Adaptive Merging (HIM) and Task Simulation Perturbation for Merging (TSP). Instead of searching for a global merge coefficient, HIM directly solves for the target merged update under a joint objective over previous and current tasks. Solving this objective gives an analytic solution that naturally takes the form of element-wise modulation over the current task LoRA update. To improve the stability of the learned update under subsequent merging, we introduce TSP which simulates different tasks interference. Experiments on multiple benchmarks show that M²LE consistently outperforms existing PEFT-based continual learning methods. Our code is provided in the supplementary material.


Monitoring the Internal Monologue: Probe Trajectories Reveal Reasoning Dynamics

Maciej Chrabaszcz ⋅ Aleksander Szymczyk ⋅ Marcin Sendera ⋅ Tomasz Trzcinski ⋅ Sebastian Cygert

Large Reasoning Models (LRMs) introduce new opportunities for safety monitoring through their Chain of Thought (CoT) reasoning. However, CoT is not always faithful to the model's final output, undermining its reliability as a monitoring tool. To address this, we investigate the hidden representations of LRMs to determine whether future behavior can be predicted from prompt and CoT representations. By evaluating a probe at each generated token, we construct a probe trajectory, the continuous evolution of a concept's probability across the reasoning process. We find that future model behavior is more distinguishable when examined over the full trajectory than from a single static prediction. To characterize these temporal dynamics, we extract signal-processing features that capture volatility, trend, and steady-state behavior, significantly improving the separation of future model states. We also present two methodological insights. First, template-based training data achieves near-parity with dynamically generated model responses, eliminating the need for a costly initial inference and labeling. Second, the choice of pooling operation is critical: average-pooling and last-token methods collapse to near-random performance, while max-pooling achieves up to $95\%$ AUROC and yields stable probe trajectories. Using four datasets and four reasoning models across the domains of safety and mathematics, we demonstrate that trajectory features encode task-specific dynamics that improve outcome separability. These findings establish probe trajectories as a complementary framework for monitoring LRM behavior. Warning: This article contains potentially harmful content.


MOOD: Benchmarking Post-Hoc OOD Detection for Materials Property Prediction

Shirley X Yu ⋅ Amey P Pasarkar ⋅ Adji Bousso Dieng

Out-of-distribution (OOD) detection is essential for deploying machine learning models in high-stakes scientific applications, yet methods that flag OOD inputs without retraining the model have been validated almost exclusively on image classification. We introduce MOOD (Materials Out-of-Distribution Detection), a benchmark for evaluating whether these post-hoc OOD detectors transfer to graph neural networks (GNNs) on materials property prediction. MOOD evaluates 15 post-hoc detectors and 2 supervised diagnostic probes across 4 GNN architectures, 5 MatBench tasks, and 15 physically motivated compositional and structural shifts, yielding over 200 evaluation settings and 4,000 AUROC measurements. Our evaluation reveals three key findings. First, current post-hoc detectors perform near chance; all 15 detectors achieve an AUROC around 0.5. Second, detector failures decompose into two regimes: method bottlenecks and encoder bottlenecks. Method bottlenecks arise when supervised probes can separate in-distribution (ID) from OOD samples but unsupervised detectors fail, meaning that the embeddings contain OOD signal that current detectors do not extract. Encoder bottlenecks arise when even supervised probes fail to separate ID from OOD samples, indicating that the relevant shift is not accessible in the frozen GNN representation. Third, when OOD prediction error exceeds ID error, the AUROC of distance-based and density-based detectors correlates with the size of that gap. Collectively, these findings indicate that reliable OOD detection in scientific domains requires representation learning strategies that explicitly preserve OOD relevant signal, particularly for tasks where current GNN architectures encode no extractable signal at all.


MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI

Paula Arguello ⋅ Berk Tinaz ⋅ Mohammad Shahab Sepehri ⋅ Maryam Soltanolkotabi ⋅ Mahdi Soltanolkotabi

Deep learning underpins a wide range of applications in MRI, including reconstruction, artifact removal, and segmentation. However, progress has been driven largely by public datasets focused on brain and knee imaging, shaping how models are trained and evaluated. As a result, careful studies of the reliability of these models across diverse anatomical settings remain limited. In this work, we introduce MosaicMRI, a large and diverse collection of fully sampled raw musculoskeletal (MSK) MR measurements designed for training and evaluating machine-learning--based methods. MosaicMRI is the largest open-source raw MSK MRI dataset to date, comprising 2,725 volumes and 81,570 slices, including a 1.5T core dataset and a 0.55T low-field set. The dataset offers substantial diversity in volume orientation (e.g., axial, sagittal), imaging contrasts (e.g., PD, T1, T2), anatomies (e.g., spine, knee, hip, ankle, and others), numbers of acquisition coils, and field strengths. Using accelerated reconstruction as a testbed, we study scaling, robustness, and data selection under realistic MSK distribution shifts. Across E2E-VarNet, U-Net, and ViT baselines, mixed-anatomy training improves over anatomy-specific training, showing that the benefit of anatomical diversity is not architecture-specific. Zero-shot experiments on the 0.55~T low-field set highlight that using more training data improves average reconstruction under field-strength shift. Controlled subset and leave-one-anatomy-out experiments further suggest that these gains come from transferable cross-anatomy structure, rather than anatomy or contrast matching alone.


Most ReLU Networks Admit Identifiable Parameters

Moritz Grillo ⋅ Guido Montufar

We study the realization map of deep ReLU networks, focusing on when a function determines its parameters up to scaling and permutation. To analyze hidden redundancies beyond these standard symmetries, we introduce a framework based on weighted polyhedral complexes. Our main result shows that for every architecture whose input and hidden layers have width at least two, there exists an open set of identifiable parameters. This implies that the functional dimension of every such architecture is exactly the number of parameters minus the number of hidden neurons. We further show that minimal functional representations can still have non-trivial parameter redundancies. Finally, we establish a generic depth hierarchy, whereby for an open set of parameters the realized function cannot be represented generically by any shallower network.


MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing

Bangbang Zhou ⋅ Hangdi Xing ⋅ Yifan Chen ⋅ Jianjun Xu ⋅ Qi Zheng ⋅ Feiyu Gao ⋅ Zhibo Yang ⋅ Shuai Bai ⋅ Ming Yan ⋅ Jieping Ye ⋅ Hongtao Xie

Document parsing converts visually rich documents into machine-readable structured representations, forming a crucial foundation for information systems. Although many benchmarks have been proposed for document parsing, they remain inadequate for realistic scenarios. Existing benchmarks either focus on specific tasks or assess only single-page, text-centric settings, making them insufficient for practical multi-page parsing. Moreover, they lack fine-grained evaluation of semantic continuity, hierarchical structure recovery, and visual content preservation. To address these gaps, we propose MPDocBench-Parse, a benchmark for multi-page document parsing in real-world applications. It contains 433 manually annotated documents with 3,246 pages, covering 15 document types in English and Chinese, with diverse layout styles, and supports document-level end-to-end evaluation. We further design a comprehensive protocol for content fidelity and logical structure, covering text, table, and formula recognition, truncated text and table merging, figure extraction, reading order, and heading hierarchy recovery. Experiments show that, while existing models perform well on basic text extraction, they still suffer clear limitations in semantic continuity integration, visual content parsing, and hierarchical structure recovery. MPDocBench-Parse provides a unified foundation for advancing document parsing toward more realistic scenarios.


mRNABench: A curated benchmark for mature mRNA property and function prediction

Ruian (Ian) Shi ⋅ Taykhoom Dalal ⋅ Philip Fradkin ⋅ Divya Koyyalagunta ⋅ Simran Chhabria ⋅ Andrew Jung ⋅ Cyrus L Tam ⋅ Defne Ceyhan ⋅ Jessica Lin ⋅ Kaitlin U Laverty ⋅ Ilyes Baali ⋅ Bo Wang ⋅ Quaid Morris

Messenger RNA (mRNA) is central to gene expression, and its half-life, localization, and translation efficiency drive phenotypic diversity in eukaryotic cells. While supervised learning has been used to study the mRNA regulatory code, self-supervised foundation models support a wider range of transfer learning tasks. However, the dearth of standardized benchmarks limits efforts to pinpoint the strengths of various models. Here, we present mRNABench, a benchmarking suite for mature mRNA biology, focused on human transcripts, that evaluates the representational quality of mature mRNA embeddings from self-supervised nucleotide foundation models. We curate 11 datasets and 79 prediction tasks that broadly capture salient properties of mature mRNA, and assess the performance of 29 families of nucleotide foundation models for a total of 460k experiments. Using these experiments, we study parameter scaling, correlations between sequence compressibility and performance, data-splitting strategies, and the effects of self-supervised training objective on mRNA property prediction.


MSAR: Next-Scale Autoregressive Forecasting for Time Series via Modular Multi-Scale Decoupling

Fanda Fan ⋅ Kuiye Ding ⋅ Liming Mao ⋅ Wang Bingsong ⋅ Yao Wang ⋅ Xiaorui Wang ⋅ Ruijie Jian ⋅ Zhipeng Liu ⋅ Luqi Gong ⋅ Zhenghua Lu ⋅ Chunjie Luo ⋅ Jianfeng Zhan

Time series forecasting underpins critical applications in finance, energy, healthcare, and transportation. Although deep models have achieved strong results, most adopt single-scale modeling or restrict multiscale processing to the input side, causing a misalignment between multiscale inputs and single-scale outputs and limiting predictive power. We introduce the Modular Scale-wise Autoregressive Framework (MSAR), a model-agnostic design that forecasts progressively across multiple temporal resolutions. MSAR offers three advantages: (1) scale-wise aligned modeling, which disentangles heterogeneous temporal patterns by aligning inputs and outputs at each scale; (2) scale-wise autoregression, where coarse-scale predictions guide finer-scale forecasting through hierarchical information flow; and (3) a modular architecture, enabling seamless integration with diverse backbones such as CNNs, MLPs, and Transformers. Extensive experiments across a broad set of datasets and forecasting models demonstrate that MSAR achieves consistent improvements in both accuracy and inference efficiency, validating the effectiveness of scale-aligned autoregression for multiscale time series forecasting.


Mult-DPO: Multinomial Direct Preference Optimization for Recommender Systems

Yaochen Zhu ⋅ Harald Steck ⋅ James McInerney ⋅ Aditya Sinha ⋅ Yinhan He ⋅ Nathan Kallus ⋅ Jundong Li

Direct preference optimization (DPO) is a simple and effective alignment strategy for large language models (LLMs) based on pairwise preferences between two candidates. In recommender systems, however, user feedback is rarely pairwise. For a given context, e.g., a user, a session, or a conversation, we typically observe set-wise preferences with multiple positive items, where every positive item should outrank every unobserved or explicitly negative item, with no prescribed order among the positives or the negatives themselves. A natural generalization is to use the Plackett–Luce (PL) reward model, which extends the Bradley–Terry (BT) reward model underlying vanilla DPO from pairwise preferences to full rankings of candidates. However, adapting the PL model to set-wise preferences requires marginalizing over all consistent permutations of the positives and negatives respectively, which is intractable. To address this fundamental challenge, we propose Mult-DPO, a novel DPO objective with a tractable multinomial reward model over set-wise preferences. We show that, like the BT and PL reward models, the multinomial reward admits a closed-form DPO-style objective, enabling direct alignment of LLMs without reinforcement learning (RL). In addition, we prove that the multinomial DPO loss is a tractable upper bound on the exact marginalized PL DPO loss when optimizing against the set-wise preference data. We further characterize the tightness of this bound in terms of relative total weight (i.e., exponentiated rewards) of positives versus negatives, which provides insights into tightening the bounds with more or harder negatives. Finally, we extend the framework to a group-wise setting that accommodates multiple preference levels. Code and datasets are available at https://anonymous.4open.science/r/mult-dpo-41C7.


Multi-Agent Coordination via Support-Preserving Distillation

Sangmin Lee ⋅ Youngju Na ⋅ Chanmi Lee ⋅ Sung-eui Yoon

Offline MARL increasingly relies on generative policies to model multimodal joint behavior, typically by distilling a centralized teacher into decentralized one-step actors under the CTDE. We identify a failure mode at the teacher training stage: standard flow-based teachers pair noise with replay targets independently, so nearby noise samples can be routed toward conflicting coordination modes. The teacher then produces samples between valid modes, and because the distillation loss regresses each local actor onto the conditional mean of the teacher's output given local input, this error is not absorbed but propagated to the student. To remove this teacher-side artifact, we propose Mode-Support Semi-Discrete Optimal Transport (MoSDOT), which summarizes multimodal replay into a finite mode support with prescribed capacities and uses conditional semi-discrete optimal transport to assign each noise sample to a single mode before teacher training. We additionally study a common-randomness variant that uses a shared noise component at execution to expose the residual gap intrinsic to strict-product execution. On controlled diagnostics and offline MARL benchmarks, MoSDOT improves endpoint quality and routing consistency, particularly on datasets exhibiting multimodal joint behavior.


Multi-Constrained Randomized Smoothing Certificates

Jimin Cao ⋅ Aleksandar Bojchevski

Randomized Smoothing (RS) is a principled and widely adopted approach for certifying the robustness of black-box models, such as neural networks, against adversarial perturbations. We propose an optimization perspective that unifies several prior randomized smoothing certificates. We then show how to reduce the underlying high-dimensional worst-case optimization problem to an equivalent two-dimensional formulation under a double- or multi-sampling scheme, enabling efficient solutions via convex optimization. By exploiting additional information about the variance of the smoothed classifier across different smoothing distributions, we derive certified bounds that improve upon standard certificates. We show results on both classification and regression tasks.

We study stochastic multi-objective optimization when each objective may be non-smooth and non-convex. We introduce the notion of $(\delta,\epsilon)$-Pareto Goldstein stationary points (PGSP) to characterize the convergence of solving multi objective optimization. The $(\delta,\epsilon)$ generalizes the definition of $(\delta,\epsilon)$-Goldstein stationary point for single objective non-smooth non-convex optimization and the $\epsilon$-Pareto stationary point for solving multi objective optimization with smooth objectives. We propose the multi-gradient-free method (MGFM), which finds $(\delta,\epsilon)$-PGSP with $\tilde{\mathcal{O}}(m^2d^{3/2}\delta^{-1}\epsilon^{-4})$ oracle complexity to the stochastic function values, where $m$ is the number of objectives and $d$ is the problem dimension. We further propose MGFM+, a faster multi-gradient-free method based on variance reduction, which improves the complexity to $\tilde{\mathcal{O}}(m^2d^{3/2}\delta^{-1}\epsilon^{-3})$. These provide the first non-asymptotic analysis for multi-nonsmooth-nonconvex-objective optimization. \end{abstract}


Multiple Instance Verification

Xin Xu ⋅ Eibe Frank ⋅ Geoff Holmes

We explore multiple instance verification, a problem setting in which a query instance is verified against a bag of target instances with heterogeneous, unknown relevancy. We show that naive adaptations of attention-based multiple instance learning (MIL) methods and standard verification methods like Siamese neural networks are unsuitable for this setting: directly combining state-of-the-art (SOTA) MIL methods and Siamese networks is shown to be no better, and sometimes significantly worse, than a simple baseline model. Postulating that this may be caused by the failure of the representation of the target bag to incorporate the query instance, we introduce a new pooling approach named “cross-attention pooling” (CAP). Under the CAP framework, we propose two novel attention functions to address the challenge of distinguishing between highly similar instances in a target bag. Through empirical studies on three different verification tasks, we demonstrate that CAP outperforms adaptations of SOTA MIL methods and the baseline by substantial margins, in terms of both classification accuracy and the ability to detect key instances. The superior ability to identify key instances is attributed to the new attention functions by ablation studies.


Multistage Defer Trees for Hybrid Interpretability: If at First You Can't Succeed, Tree Again

Zakk Heile ⋅ Hayden McTavish ⋅ Margo Seltzer ⋅ Cynthia Rudin

Recent work has shown that well-optimized individual decision trees can match complex black box models in some settings, primarily in noisy domains. For the remaining settings, however, complex ensembled compositions of trees often achieve higher accuracy at the cost of interpretability, leaving practitioners with difficult modeling decisions along an accuracy-interpretability tradeoff. Ideally, we would like to classify as much of the data as possible with one or a small number of trees, achieving interpretability for most samples while maintaining state-of-the-art accuracy. We introduce Multistage Defer Trees: a sequence of sparse decision trees that each make predictions for most samples, while deferring a small proportion to the next tree in the sequence or, ultimately, to a black box. We demonstrate that we can train this model class to match the performance of complex tree-based ensembles while routing most samples through only one or a small number of sparse decision trees. We discuss a range of techniques for training these models while maintaining simplicity. Our method expands the accuracy--interpretability frontier in settings where single-tree methods remain insufficient, demonstrating that even when complex models are necessary, they need not be fully opaque.

We introduce $\texttt{MUNI}$, an end-to-end multimodal latent diffusion framework for any-to-any generation that unifies subset-conditioned cross-modal generation and unconditional joint sampling through a shared stochastic latent. Existing multimodal generative models are largely LLM-based, which limits leveraging modality-specific generators and requires text-paired data for training. Recent diffusion- and flow-based any-to-any extensions take a different direction but still rely on text-aligned embeddings, fully-paired training, or matched-dimensionality deterministic mappings. $\texttt{MUNI}$ rests on two complementary contributions, one architectural and one in the training objective. First, we extend latent diffusion to multimodal any-to-any generation *end-to-end*: instead of the standard two-stage recipe that precomputes a frozen latent space and then fits a prior over it, $\texttt{MUNI}$ jointly trains modality-specific encoders, expressive decoders, and a single shared flow-based prior under one objective. Second, we identify that the standard aggregation rules of multimodal variational inference are insufficient once coupled with a learned prior and expressive decoders. A suitable shared latent must simultaneously satisfy *coherence* across generated modalities, *predictive sufficiency* of subset latents, and *minimality* of the latent content. We propose a routed training objective whose structural choices align the latent with these criteria and admit a minimal-sufficiency characterization in the realizable setting. Experiments on PolyMNIST-Quadrant-Labels and a large-scale image-text-audio benchmark show $\texttt{MUNI}$ matching or exceeding the strongest baselines on conditional generation while opening its largest margins on unconditional coherence.


MUX: Continuous Reasoning via Multiplexed Tokens

Ayhan Suleymanzade ⋅ Halil Alperen Gozeten ⋅ Michael Bronstein ⋅ Ismail Ilkan Ceylan ⋅ Jinwoo Kim

Language models solve complex problems by articulating intermediate reasoning steps in natural language. While effective, this process is computationally bottlenecked: each reasoning step conveys only a single subword, and many are spent expressing a thought instead of carrying out computation. We propose MUX, a simple method for high-bandwidth and compact reasoning based on distillation of discrete reasoning into continuous multiplexed tokens in a latent space. Each latent token is trained to represent a weighted linear superposition (multiplexing) of a span of discrete reasoning subwords, where this superposition is lossless by construction and the span can be fully recovered (demultiplexing). We prove that simple position-dependent weightings, such as suitable geometric decay, support lossless multiplexing and prevent latent collapse. We further show that multiplexed reasoning can perform parallel exploration in problems that require search. Across 32 evaluation settings spanning four language models, MUX outperforms strong latent reasoning baselines. Our results suggest that properly chosen learning targets can make continuous latent reasoning efficient and interpretable.


MVPISplat: Multi-View Photometric Inconsistency for Defending 3D Gaussian Splatting Attacks

Md Mahedi Hasan Rigan ⋅ Nicole Meng ⋅ Miao Yin ⋅ Yingjie Lao ⋅ Faysal Hossain Shezan

3D Gaussian Splatting (3DGS) achieves real-time photorealistic novel-view synthesis but is vulnerable to adversarial perturbations of its training images. Recent attacks inflate the Gaussian count to exhaust GPU memory or corrupt rendered scenes to mislead downstream classifiers. We find that both attacks have the same structural signature: 2D image adversarial perturbations are not 3D-consistent by design. The Gaussians generated to match them provide erroneous contributions in some views while making significant contributions in others. We propose Multi-View Photometric Inconsistency (MVPI) Defense that scores each Gaussian by the variance of its opacity-weighted local $\mathcal{L}_1$ contribution across a stratified set of training views. The most inconsistent are soft-pruned via 3DGS's existing opacity threshold method. We evalute our approach on two different datasets. MVPI reduces peak Gaussian count by up to $1.90\times$ and training time by $1.22{-}1.50\times$ under Poison-Splat attack. It also restores top-1 classification accuracy on adversarially perturbed renders from $57.4\%$ to $66.5\%$. Our code is available at https://anonymous.4open.science/r/PruneDefense-2876/README.md.


My Video Stays Mine: Temporally Consistent Universal Adversarial Perturbations against Video Customization

Yuxin Huang ⋅ Mingming Gong ⋅ Wanyu Wang ⋅ Jing Zhang ⋅ Tongliang Liu

Recent diffusion-based video generation models have enabled high-quality video customization through both tuning-based pipelines, which fine-tune a video diffusion model, and reference-based pipelines such as image-to-video generation. However, these capabilities raise serious concerns about facial privacy and security. Existing anti-diffusion protections focus on the image domain or on reference-based I2V pipelines, leaving the tuning-based video customization unexplored. Protecting videos in this setting raises three challenges: image-level perturbations are erased by temporal operations, a perturbation tied to one clip fails to generalize across videos and lengths, and inconsistent temporal noise is easily removed by temporal attack. To address these challenges, we propose Temporally Consistent Universal Adversarial Perturbations (TC-UAP), the first protection method against both reference- and tuning-based video customization. TC-UAP learns a multi-frame universal adversarial perturbation over a set of videos of the same identity, so that a single perturbation can transfer to unseen videos and arbitrary clip lengths of that identity. Besides, we further enforce consistency through intrinsic temporal modeling and an extrinsic surrogate temporal-attack loss, ensuring robustness against temporal attacks. Extensive quantitative and qualitative experiments show that TC-UAP degrades identity preservation more than existing baselines under both tuning-based and reference-based video customization, and remains robust under three unseen temporal attacks.


NAGO: Noise-Aware Generative Operator via Flow Matching

Guanyu Chen ⋅ Shengze Xu ⋅ Pengwei Liu ⋅ Xingyu Ren ⋅ Pengkai Wang ⋅ Tieyong Zeng ⋅ Antonios Armaou ⋅ Dong Ni

We introduce NAGO, a noise-aware generative operator for efficient sampling from invariant measures of stochastic partial differential equations (SPDEs) in function space. Our central idea is to treat the noise measure as a design variable rather than a fixed default. While existing generative approaches mainly shape the sampling process through interpolation schedules or guidance heuristics, NAGO directly controls the marginal velocity field in functional flow matching by choosing a noise measure with suitable regularity. We establish conditions under which the resulting noise-induced flows are well defined in infinite-dimensional settings, and show how noise regularity affects the spectral geometry and stiffness of the marginal velocity field. This analysis leads to a principled design rule: match the regularity of the noise measure to that of the target invariant measure. The resulting method enables stable training and fast sampling without relying on downstream correction mechanisms. Experiments on three representative SPDEs with multiscale invariant measures show that NAGO consistently improves the accuracy--efficiency trade-off over standard flow matching baselines, while achieving up to two orders of magnitude speedup compared with classical numerical solvers.


Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions

David Wessels ⋅ FARHAD RAMEZANGHORBANI ⋅ David W Romero ⋅ Alireza Moradzadeh ⋅ Olivia Viessmann ⋅ Maksim Zhdanov ⋅ John St. John ⋅ Ken Janik ⋅ David Knigge ⋅ Yucheng Tang ⋅ Erik Bekkers ⋅ Saee Paliwal

Subquadratic alternatives to attention require compromises when applied to multi-dimensional data: standard convolutions abandon the global, input-dependent receptive field, while recurrent models require rasterizing images, volumes, and PDE grids into an ad-hoc $1\rm D$ scan order that violates their spatial structure. We introduce HyenaND, an operator that acts directly on native $\rm ND$ data in its intrinsic geometry and recovers a global, input-dependent receptive field at subquadratic cost. It uses an implicit, input-dependent parameterization of the multi-dimensional convolutional kernel. We provide a CUDA implementation, \texttt{nSubQ}, which fuses the FFT-convolution path to turn HyenaND's $O(N \log N)$ scaling into wall-clock speedups. Across long-context genomics, computer vision, medical imaging, and PDE modeling, pure HyenaND stacks match the accuracy of strong attention-based baselines, while hybrid configurations that interleave HyenaND and attention layers outperform both pure attention and strong recurrence-based hybrids.

Task-oriented dexterous grasp generation aims to produce dexterous grasp poses that are both physically plausible and functionally suitable for specified manipulation tasks. Existing diffusion-based methods often address these two requirements in a decoupled manner: they first train a grasp diffusion model for task alignment and then rely on post-generation refinement to improve physical plausibility. However, this after-the-fact correction strategy applies physical plausibility guidance only once the grasp has already been generated, leaving the generation trajectory itself unguided by physical constraints and potentially leading to suboptimal grasps. To address this problem, we propose a novel framework that directly injects physical plausibility guidance into the denoising process of a task-aligned grasp diffusion model in a practical and effective manner, even when physical plausibility constraints are non-differentiable. This allows physical plausibility to shape grasp generation throughout denoising while preserving task alignment. Extensive experiments demonstrate the efficacy of our framework.


Need-Aware Multi-Objective Reinforcement Learning for Emotionally Intelligent LLM Agents

Xiaohe Bo ⋅ Xueyang Feng ⋅ Rui Li ⋅ Zihang Tian ⋅ Xu Chen ⋅ Quanyu Dai

In emotional support conversations, user needs are dynamic, shifting across turns between emotional soothing, cognitive guidance, and their combinations. Existing methods often optimize a single user-state objective, such as emotional relief or cognitive growth, making it difficult to capture such need shifts. In this paper, we formulate emotional support dialogue as a need-aware multi-objective optimization problem, where agents learn to generate responses that balance multiple support objectives according to users' evolving needs. We propose \textbf{MindFlow}, a psychologically grounded user simulator that tracks users' emotional and cognitive needs at each turn, derives state-dependent preference vectors from personalized arousal curves, and models mental-state transitions with a dual-process system. Building on MindFlow, we introduce \textbf{NeMo}, a need-aware multi-objective reinforcement learning framework that estimates local multi-turn effects through forward simulation and dynamically aggregates emotional and cognitive rewards based on current support preferences. By converting delayed, multi-dimensional feedback into step-level optimization signals, NeMo improves credit assignment and aligns policy learning with long-term user-state improvement. Experiments show that NeMo consistently improves user-side state metrics and response-level supportive quality, with robust generalization under both Sim2Sim and Sim2Real evaluations. Code and data are available at https://anonymous.4open.science/r/NeMo-6EB9.


Negative-Only Policy Optimization for One-Sided Verifiable Rewards

Jiacheng Xu ⋅ Shuo He ⋅ Fuxiang Zhang ⋅ Chaojie Wang ⋅ Bo An

Reinforcement learning with verifiable rewards (RLVR) typically relies on a verifier that can reliably judge whether sampled solutions are correct. However, in many realistic settings, complete verification is generally unavailable, while cheap checks can only falsify failures. We formalize this regime as one-sided verifiability (OSV): verified negatives are reliable counterexamples, but verified positives are merely non-falsified candidates and may contain many false positives. This asymmetry makes standard RLVR unreliable, because reinforcing verified positives can directly reward shortcuts that pass partial checks. To address this issue, we propose Negative-Only Policy Optimization (NOPO), which treats verified positives as unlabeled and updates the policy only by suppressing trusted verified negatives. Specifically, NOPO combines two safeguards: a sample-level NPO term that self-attenuates once a verifier-negative is suppressed below the reference, and a group-level non-falsified support weight that scales each group's update by its verifier-positive count. Across unit-test, constraint-check, and LLM self-verifiers, NOPO improves pass@k over baselines on code, logic, and math benchmarks when oracle rewards are unavailable. These gains persist under inference-time scaling and cross-benchmark transfer.


Neighbor-Aware Snapshot-Based Temporal Graph Learning

Jinxi Yang ⋅ He Li ⋅ Wei Yu ⋅ Xingyu Gao ⋅ Mang Ye

Dynamic graph learning models the temporal evolution of structural interactions. Existing methods present a stark trade-off between performance and efficiency. Event-based models achieve strong accuracy because processing graphs edge-by-edge allows them to continuously update and leverage fine-grained joint neighborhood information. Conversely, snapshot-based methods offer high computational efficiency by processing entire snapshots simultaneously. However, updating neighborhood states for all nodes concurrently is computationally prohibitive. Consequently, current snapshot methods fail to utilize joint neighborhood information, leading to significantly weaker performance. To bridge this gap, we propose the Neighbor-Aware Graph Neural Network (NAGNN), a novel snapshot-based architecture that seamlessly integrates joint neighborhood information to achieve both strong performance and high efficiency. Specifically, NAGNN introduces a Locality-Sensitive Hashing (LSH)-based topology distiller to construct compact reservoirs that store neighborhood information efficiently. These reservoirs are further augmented with degree-scaled statistical moments to encode structural characteristics of local neighborhoods. A Gated Recurrent Unit (GRU) subsequently processes the distilled structural representations to model and capture their temporal dynamics. Extensive experiments demonstrate that NAGNN matches the predictive accuracy of computationally expensive event-based models with minimal overhead, establishing a highly scalable and effective paradigm for dynamic graph learning.


Neptuna: A Comprehensive Machine Learning Framework for Benchmarking Complex Multiphase Flows

Harish Ramachandran ⋅ Björn Kimpel ⋅ Thomas Paula ⋅ Josef Winter ⋅ Steffen Schmidt ⋅ Nikolaus Adams

Compressible multiphase flows involving shocks and material interfaces arise in applications such as bubble collapse and droplet breakup, where strong nonlinear interactions produce complex interface deformation, mixing, and multiscale dynamics. Developing reliable machine learning surrogates for these flows remains challenging due to the simultaneous presence of compressibility, sharp discontinuities, and multiphase effects. In this work, we introduce the first large-scale benchmark specifically designed for shock-driven compressible multiphase flows, comprising 2.4 TB of high-fidelity 2D and 3D datasets featuring shock-induced bubble collapse and droplet breakup. We evaluate diverse surrogate model families on our benchmarking framework: Neptuna, including convolutional, spectral, transformer-based, and pre-trained PDE foundation models. Beyond standard MSE training, we investigate composite losses combining MSE with Sobolev, interface-aware, and structure-aware terms, together with adaptive loss balancing using SoftAdapt and GradNorm. Evaluation includes pointwise, spectral, feature-focused, structural, and physics-informed metrics. Results show that no single model performs best across all datasets and metrics, while composite losses significantly improve interface preservation and spectral fidelity. Among adaptive weighting strategies, SoftAdapt provides the most consistent improvements with almost no overhead compared to MSE-only training.

Vision-Language-Action (VLA) models promise generalist robotic agents, yet consolidating skills learned across heterogeneous tasks, environments, and embodiments remains difficult: training data is fragmented across platforms, joint multi-task training suffers from interference, and model merging tends to introduce parameter conflicts or require architectural changes. We instead advocate a paradigm in which skills are first learned independently and then consolidated post hoc. To this end, we propose NestedVLA, which replaces parameter aggregation with parameter generation: a nested hypernetwork, conditioned on the target task, synthesizes task-adaptive parameters that compose with a frozen base VLA. This formulation enables flexible task-conditioned skill reuse while mitigating interference. Across three RL benchmarks (MetaWorld, ManiSkill, CALVIN), NestedVLA outperforms the strongest model-merging baseline by 41 pp on average, matches per-task fine-tuning within 1 pp on ManiSkill, and substantially closes the gap on CALVIN, pointing toward a scalable route to skill consolidation.

Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model (LLM) agents; however, this oversight is compromised if models learn to conceal their reasoning. We explore steganographic CoT---where models hide secret reasoning within innocuous text---to inform risk assessment and deployment policies. Steganographic reasoning requires two skills in a single forward pass: computing an intermediate result, and embedding it into a coherent cover that answers an unrelated question. Drawing on our taxonomy of steganographic and non-steganographic CoT types, we systematically evaluate the limits of prompt-elicited steganographic CoT capability across 34 models, ranging from past generations to the current frontier. We measure monitor evasion, refusal rates, encoding fidelity, and hidden task accuracy across five datasets, comparing against plain reasoning, direct answer, and filler-token baselines. The two experiments isolate the two sub-skills: a reasoning tasks sweep tests joint reason-and-embed, while a counting task hands the model a known numerical sequence and tests embedding alone---a necessary precondition for stego reasoning. Current frontier models cannot sustain joint reason-and-embed: a paired McNemar comparison shows the steganographic channel is dominated by an filler-token baseline on every (model, family) cell. The encoding-only floor, by contrast, is cleared---Claude Opus~4.5 reaches 92\% per-number partial accuracy (54\% exact-match) on 4-digit sequences and saturates at 100\% exact-match on length-8 single-digit sequences---establishing that the binding constraint on stego CoT is the joint reasoning-plus-encoding load, not raw channel capacity. Separately, GPT-5.2 issues an explicit refusal to the steganographic instruction yet still produces partial valid encoding in 6 of 644 trials. Our findings underscore the need for continuous evaluation of steganographic risk and provide a methodology to preemptively detect and evaluate hidden reasoning that might empower misaligned scheming and deceptive behavior.


NestRL: A Nested Training Regime for Mutual Adaptation in Human–AI Teaming

Upasana Biswas ⋅ Durgesh Kalwar ⋅ Subbarao Kambhampati ⋅ Sarath Sreedharan

Mutual adaptation is a central challenge in human–AI teaming, as humans naturally adjust their strategies in response to an AI agent's behavior. Existing approaches attempt to approximate human behavior by diversifying training partners; however, these partners are typically static and fail to capture the adaptive nature of human teammates. When agents are trained jointly in standard multi-agent settings, they often converge to opaque coordination strategies that work only with their co-trained partners, leading to poor generalization. To model adaptive human behavior, we formulate human–AI teaming as an Interactive Partially Observable Markov Decision Process (I-POMDP). We propose NestRL, a nested training regime that learns the solution to a finite-level I-POMDP by training agents at each level against adaptive agents from the level below. This exposes agents to adaptive behavior while preventing emergence of opaque coordination strategies. We provide theoretical analysis showing that NestRL agents avoid convergence to partner-specific strategies, and validate this empirically in the Overcooked domain against state-of-the-art baselines. NestRL achieves higher task performance with both unseen adaptive agents and real human teammates, while exhibiting significantly greater adaptability over the course of interaction.


Network Intervention by Polling Strategic Agents

Chenyu Zhang ⋅ Rohit Parasnis ⋅ Saurabh Amin

A planner in a network of strategic agents faces three entangled challenges: the optimum depends on agents' private information, queried agents may misreport to steer the outcome, and exact computation does not scale. We study these challenges in multi-activity network games with heterogeneous private technologies, in which the planner sets non-discriminatory prices. We show that the optimal prices admit a centrality-based decomposition of the welfare kernel: each agent's contribution scales with its squared centrality in a network reweighted by agents' preferences across activities. This decomposition motivates Poll, a polling algorithm in which the planner samples one agent per round, walks briefly through the agent's neighborhood, and updates the price from a local report. From the same decomposition flow three forms of efficiency: computationally, Poll uses significantly fewer operations than exact computation and other distributed methods; statistically, its query complexity scales with topology and preference heterogeneity rather than with population size; and economically, it converges to welfare-maximizing prices while inducing truthful reports and detecting adversarial deviations.


Neural-Behavioral Representation of Natural Whole-body Movement in Monkeys

Jieshi He ⋅ Puzhe Li ⋅ Yanan Sui ⋅ Mu-ming Poo

Understanding how cortical activity represents natural whole-body behavior in primates remains challenging. Limited by diversity of movements and inaccessible of large-scale neural feature programming whole-body kinematics, previous motor decoding studies rely on constrained tasks and limited limb movements. Here, we present a neural-behavioral recording and modeling framework for freely moving monkeys, combining epidural cortical signals from distributed sensorimotor areas with synchronized multi-view motion capture through a data custom collection platform. We reconstruct whole-body monkey kinematics and learn a compact motion prior using an autoregressive encoder-decoder model. Conditioned on epidural signals, the model decodes more accurate and realistic full-body movements without explicit physical constraints. Our results provides a novel proof-of-concept approach for decoding natural whole-body movements in primates using large-scale intracranial neural activity.


NeuralBES: A Differentiable, Control-Aware Emulator for Scalable Building Energy Modeling

Ting-Yu Dai ⋅ Takuya Kurihana ⋅ Wing Yee Au ⋅ YUNG WONG

Demand-side flexibility, i.e., forecasting, shifting, and curtailing residential energy loads, depends on thermal models trusted across millions of heterogeneous buildings. Existing tools force a hard tradeoff: high-fidelity physics simulators such as EnergyPlus are accurate but sequential and require per-building calibration, while purely data-driven sequence models scale but abandon the physical structure that makes their predictions trustworthy. We introduce NeuralBES (Building Energy Simulation), a differentiable emulator that resolves this tradeoff by parameterizing a resistance--capacitance (RC) based thermal model with a shared neural encoder: static building metadata such as floor area, vintage, and HVAC type is mapped to physically bounded capacitances, conductances, and equipment coefficients, which become the coefficients of a scalar linear recurrence solved via a log-space parallel scan, and a predictor--corrector loop closes the thermostat--temperature nonlinearity while preserving full-horizon gradient flow. Trained on the ResStock dataset across three climate zones, NeuralBES handles heterogeneous building archetypes, vintages, and climate zones within a single trained encoder, while black-box baselines produce statistically plausible but physically inconsistent trajectories. On the annual full-year rollout, NeuralBES is the only data-conditioned model that is simultaneously physics-valid and accurate to within 4 MAPE points of the strongest raw-error baseline, while operating at roughly an order of magnitude fewer parameters than the transformer and recurrent baselines; among physics-valid baselines at parameter parity it more than halves the MAPE of the grey-box RC alternative.

Neural populations exhibit heterogeneous tuning: Neurons differ in their selectivity, gain, and in how their preferred inputs are distributed across stimulus or latent-variable space. While this heterogeneity has been studied extensively in the context of encoding precision, its implications for learning and generalization are less well-understood. Here, we develop a probabilistic theory that links a neural population's tuning statistics to its inductive bias: in large populations, the tuning statistics induce a kernel and, equivalently, a prior of a generative model over functions. This prior in turn predicts the population's ability to generalize across a distribution of related tasks. Our theory determines the population tuning statistics that minimize the generalization error over a distribution of functions. We further present a normative account of tuning adaptation inspired by our probabilistic approach, showing that tuning changes that maximize the marginal likelihood of observed data systematically reduce generalization error relative to non-adapting populations. Applied to the adaptation of hippocampal spatial representations in a reward learning task, we found that our theory could explain the over-representation of rewarding locations through experience in an environment with a localized reward. We further demonstrate how different components of the population's generative model differentially impact the speed and magnitude of tuning adaptation. Overall, our work proposes a novel connection between tuning variability and a population's ability to learn and to generalize across tasks and develops a Bayesian theory of tuning adaptation in neural populations during learning.


Neural Quantum Spectral Operator Learning for Solving Partial Differential Equations

Chanyoung Kim ⋅ Myeonghwan Seong ⋅ Kim Yujin ⋅ Daniel Kyungdeock Park ⋅ Youngjoon Hong

Partial differential equations (PDEs) are central to modeling physical and engineering systems, but repeatedly solving parametric PDEs remains computationally expensive. Operator learning enables fast surrogate inference, yet typically requires large input–output paired datasets generated by costly high-fidelity PDE solvers. Unsupervised operator learning frameworks alleviate data dependency but remain hindered by computational bottlenecks. To address this, we propose Neural Variational Quantum Linear Solver (NVQLS), the first hybrid quantum–classical operator learning framework leveraging the Legendre-Galerkin weak formulation. We critically resolve the sign ambiguity in VQLS energy minimization, preventing erroneous solution representations. Additionally, we introduce a neural embedding, a novel encoding scheme to map varying forcings and PDE coefficients into parameterized quantum circuit representations. These structural innovations provide theoretical computational complexity advantages under efficient state preparation schemes, while achieving superior accuracy compared to a representative classical baseline. Validations on 1D and 2D parametric PDEs under diverse boundary conditions demonstrate NVQLS's capability to simultaneously process varying inputs, offering a scalable unsupervised approach to quantum-enhanced operator learning.


Neural Signals Generate Clinical Notes in the Wild

Jathurshan Pradeepkumar ⋅ Zheng Chen ⋅ Jimeng Sun

Generating clinical reports that summarize abnormal patterns, diagnostic findings, and clinical interpretations from long-term EEG recordings remains labor-intensive. We present CELM, the first clinical EEG-to-Language foundation model capable of summarizing long-duration, variable-length EEG recordings and performing end-to-end clinical report generation at multiple scales. CELM integrates pretrained EEG foundation models with language models to enable scalable multimodal learning. We curate a large-scale clinical EEG dataset containing 9,922 reports paired with approximately 11,000 hours of EEG recordings from 9,048 patients to train CELM, and release the benchmark with an automated report-structuring pipeline to facilitate future research. Experimental results show that CELM consistently outperforms existing methods across all evaluation settings. Importantly, we further conduct human evaluation with clinical experts, demonstrating that CELM generates reports that are more clinically coherent, diagnostically reliable, and better aligned with expert interpretation. We release our model and benchmark construction pipeline at https://anonymous.4open.science/r/CELM-3AF4.


Neuronal Identity as an Organizational Basis for Analyzing Neural Population Dynamics

Seungjae Han ⋅ Joshua Y You ⋅ Soi Kim ⋅ Minho Eom ⋅ Jinho Park ⋅ Seunghee Park ⋅ Byeongwook Lee ⋅ Young-Gyu Yoon

Individual neurons exhibit diverse molecular, morphological, and electrophysiological properties that shape their distinct roles in neural circuits. While these multimodal features are largely shared across individuals, they remain an underutilized source of information in analyzing neural population dynamics. We posit that these shared biological profiles provide a common reference for relating neural dynamics across subjects and recording sessions. Our framework leverages this structure by routing neurons into latent groups based on their biological profiles and analyzing population dynamics through group-level representations. Because these assignments are anchored in biological properties that are consistent across subjects, the learned organization supports transfer to unseen animals without retraining or neuron-wise correspondence. We evaluate our framework on whole-brain calcium imaging in larval zebrafish and large-scale electrophysiology in the mouse visual cortex. Across both species and recording modalities, the framework delivers state-of-the-art performance in both neural decoding and manifold analysis: it achieves superior decoding of behavior and stimulus labels while better preserving the temporal consistency and local geometric structure of latent trajectories on unseen data. Our experiments demonstrate that not only does providing biological features improve baseline performance, but our proposed framework achieves even stronger results by using these properties to explicitly structure the population rather than treating them as auxiliary inputs. Together, these results suggest that organizing neural populations by their multimodal biological features offers a promising framework for generalizing population dynamics across subjects and recording sessions.


NEvo: Neural-Guided Evolutionary Video Synthesis for Dynamic Visual Selectivity

Yingtian Tang ⋅ Sogand Salehi ⋅ Ming Zhou ⋅ Amir Zamir ⋅ Leyla Isik ⋅ Martin Schrimpf

The human brain processes dynamic visual input through hierarchically organized, functionally specialized regions. While recent in silico brain encoding models can synthesize optimal stimuli to probe selectivity in different brain regions, prior work has been largely limited to static images, leaving dynamic visual processing underexplored. We introduce a novel neural-guided video synthesis framework that generates stimuli optimized for target brain regions across visual cortex. Our method performs evolutionary search over a structured prompt space, guided by a dynamic encoding model that predicts voxel-level responses to video inputs. By maximizing predicted activity for a target ROI, the framework efficiently discovers hyper-activating dynamic stimuli that consistently surpass handcrafted localizer videos. The synthesized videos recover known selectivities across ventral, dorsal, and lateral pathways, and further reveal systematic differences in sensitivity to temporal dynamics. A searchlight analysis provides new insight into the progression toward increasingly complex social-dynamic features along the lateral stream, further supported by probing with synthesized abstract, non-naturalistic stimuli. Taken together, our framework enables in silico exploration of dynamic visual selectivity, with new predictions for in vivo experiments.


New Coresets for Fair Clustering

Xuan Wu ⋅ Chansophea Wathanak In ⋅ Yi Li

We construct the first coreset of size near-linear in $k$ for fair \kMedian in general metric spaces. Combined with the standard merge-and-reduce framework, our construction also gives the first streaming algorithm for fair \kMedian with space complexity near-linear in $k$. The main technical innovation is a distributional reinterpretation of capacitated clustering that the cost of assigning a dataset to a capacitated center set can be expressed as the earth mover's distance between two discrete distributions. This viewpoint enables us to use tools from optimal transport to analyse uniform sampling. A key ingredient in this analysis is a new family of problem-specific $\eps$-nets for the earth mover's distance.


NGDB-Zoo: Towards Efficient and Scalable Neural Graph Databases Training

zhongwei xie ⋅ Jiaxin Bai ⋅ Shujie LIU ⋅ Haoyu Huang ⋅ LI Yufei ⋅ Yisen Gao ⋅ Hong Ting Tsang ⋅ Yangqiu Song

Neural Graph Databases (NGDBs) support complex logical reasoning over incomplete knowledge structures, yet their training efficiency and expressivity are constrained by rigid query-level batching and structure-only embeddings. We present NGDB-Zoo, a unified framework that resolves these bottlenecks by synergizing operator-level training with semantic augmentation. By decoupling logical operators from query topologies, NGDB-Zoo transforms training loops into dynamically scheduled data-flow executions, enabling multi-stream parallelism and achieving a $1.8\times$-$6.8\times$ average throughput compared to baselines. Furthermore, we formalize a decoupled architecture to integrate semantic priors from pre-trained text embeddings without triggering I/O stalls or memory overflows. Experiments on six benchmarks, including $\textit{ogbl-wikikg2}$ and $\textit{ATLAS-Wiki}$, show that NGDB-Zoo scales dense query-embedding training to million-entity graphs while maintaining competitive filtered MRR.

Image captioning systems are unable to generate fine-grained captions as they are trained on data that is either noisy (alt-text) or generic (human annotations). This is further exacerbated by maximum likelihood training that encourages generation of frequently occurring phrases. Previous works have tried to address this limitation by fine-tuning captioners with a self-retrieval (SR) reward. However, we find that SR fine-tuning has a tendency to reduce caption faithfulness and even hallucinate. In this work, we circumvent this bottleneck by improving the MLE initialization of the captioning system and designing a curriculum for the SR fine-tuning process. To this extent, we present (1) Visual Caption Boosting, a novel framework to instill fine-grainedness in generic image captioning datasets while remaining anchored in human annotations; and (2) BagCurri, a carefully designed training curriculum that more optimally leverages the contrastive nature of the self-retrieval reward. Jointly, they enable the captioner to describe fine-grained aspects in the image while preserving faithfulness to ground-truth captions. Our approach outperforms previous work by +8.9% on SR against 99 random distractors (RD100) (Dessi et al., 2023); and +7.6% on ImageCoDe.

Additionally, existing metrics to evaluate captioning systems fail to reward diversity or evaluate a model's fine-grained understanding ability. Our third contribution addresses this by proposing self-retrieval from the lens of evaluation. We introduce TrueMatch, a benchmark comprising bags of highly similar images that uses SR to assess the captioner's ability to capture subtle visual distinctions. We evaluate and compare several state-of-the-art open-source MLLMs on TrueMatch, and find that our SR approach outperforms them all by a significant margin (e.g. +4.8% - 7.1% over Cambrian) while having 1-2 orders of magnitude fewer parameters. We also outperform vanilla SR by +14.4% to +19.5%.


No Free Alignment: Observability-Aware Alignment for Multimodal Heterogeneous Learning

Canran Xiao ⋅ Puning Zhao ⋅ Enneng Yang ⋅ XIAOCHUN CAO ⋅ Haobo Fu ⋅ Li Shen

Multimodal heterogeneous learning aims to train reliable multimodal models when data sources differ in modality availability, semantic coverage, domain distribution, and missingness patterns. Although cross-modal alignment is widely used to mitigate such heterogeneity, we argue that alignment is not free: enforcing relations that are weakly supported by the observed data can amplify noise, over-transfer missing-modality information, and induce negative transfer. This paper studies when a modality--modality--semantic relation is sufficiently observable to be safely aligned and shared across heterogeneous sources. We propose ObsAlign, an observability-aware framework that summarizes visible modality--semantic evidence through compact source-level sketches and estimates relation-level support from bridge evidence, side support, cross-modal consistency, and prototype uncertainty. The resulting observability score gates co-observed alignment, calibrates weak missing-modality transfer, and guides evidence-aware source consolidation so that reliable relations are shared while unsupported variation remains local. Across image--text, audio--visual--text, action-recognition, sensor, and healthcare benchmarks under source-partitioned multimodal heterogeneity, ObsAlign achieves the best overall performance in both encoder-based and VLM-compatible regimes, while sharply reducing negative transfer on low-observability relations. These results suggest that relation-level observability provides a practical principle for deciding when cross-modal alignment should be strengthened, weakened, or avoided.

Iterative fine-tuning on synthetic data causes \emph{model collapse}: output diversity narrows as rare patterns are progressively lost. Existing mitigations either require model log-probabilities, an external oracle, or continued access to real human data. Here we develop a new approach grounded in mathematical information theory: the non-parametric Kontoyiannis entropy rate estimator $H_K$, computed entirely from raw text via match-length statistics, with no model of any kind. We show that this is in fact a superior training-data filter on text-diversity metrics in a fully-synthetic, single-lineage fine-tuning setting. In a six-generation QLoRA collapse experiment on Llama-3.1-8B, logprob-based filtering (the established gold standard) provides no significant text-diversity benefit on any metric ($p > 0.23$), whereas $H_K$-filtering yields +42% unique trigrams, +30% vocabulary, and -19% repetition (all $p < 0.001$). We validate $H_K$ as a cross-domain entropy proxy ($\beta = 0.924$, $R^2 = 0.746$) and collapse detector ($\rho = +0.454$, $p < 0.0001$) across 4 domains, 3 temperatures, 2 generator--scorer model pairs, and 1,680 generated documents. Our results demonstrate that information theoretic approaches to collapse mitigation are efficient, and suggest new approaches for maintaining multi-agent diversity.

Transformers have become a central architecture for \emph{in-context learning}, particularly through their strong empirical performance in large language models. This success suggests that transformers can extract task-relevant structure directly from prompts, even when the underlying data have complex geometric structure. However, existing theoretical analyses of transformer-based in-context learning are largely confined to simplified settings, such as Euclidean domains or single-manifold models. In this paper, we study in-context nonparametric regression under growing geometric complexity, modeled by a sample-size-dependent mixture of manifolds with heterogeneous local geometry. For this model class, we establish a minimax lower bound that captures the aggregate difficulty of its local components. We then construct an oracle tangent local-polynomial estimator and prove a matching upper bound by exploiting local geometric structure. The main technical step is to connect this estimator to a transformer architecture. We construct a two-stage linear-attention transformer consisting of a geometric preconditioner and all-chart reduced local-polynomial solvers, and show that its approximation error is negligible relative to the minimax regression rate. We also prove an in-context generalization bound for $\varepsilon$-near empirical risk minimizers within this transformer class. Together, the lower bound, oracle construction, transformer approximation, and generalization analysis identify the conditions under which the resulting in-context predictor adapts to local geometry and attains the aggregate minimax rate.


Not All Layers Are Equal in Image-to-Video Transfer

Thinesh Thiyakesan Ponbagavathi ⋅ Constantin Seibold ⋅ Alina Roitberg

Parameter-efficient image-to-video transfer adapts frozen image foundation models to video via lightweight trainable adapters. Existing methods insert the same adapter at every transformer layer, implicitly assuming all layers require equal temporal modeling. A layer-wise sensitivity analysis across four video-native model families and two pretraining paradigms reveals that this assumption is wrong: temporal sensitivity increases with depth, and the pretraining paradigm determines its growth profile. Vision-only models exhibit a gradual rise, while vision-language models show early suppression followed by late recovery. This finding motivates STRAP, a training-free adapter placement rule that derives a spatial-to-temporal transition boundary from a paradigm-matched video model. Spatial adapters are assigned before this boundary and temporal adapters after it, without modifying the adapter architecture itself. On four fine-grained action recognition benchmarks, STRAP improves accuracy by up to 11.8\% over uniform placement, reduces trainable parameters by 29 to 48\%, and lowers inference time by up to 10\%. On standard benchmarks, it matches accuracy with fewer parameters.


Not All Low-Confidence Tokens Are Equal: Calibrated Confidence for Efficient Test-Time Reasoning

Tangyu Jiang ⋅ Haodi Wang ⋅ Yuanbing Zhu ⋅ Xiaojiang Du ⋅ Xiaohua Jia ⋅ Bo Han

Test-time scaling boosts the reasoning performance of large language models (LLMs) on challenging tasks but incurs substantial computational overhead. The existing training-free approaches adopt token-level confidence as a proxy and terminate the trajectories once the confidence falls below a threshold. However, these methods view all low-confidence tokens uniformly, overlooking the fact that some of them actually reflect benign generations. We introduce Calibrated Confidence (CalConf), a training-free proxy for test-time scaling in LLM reasoning. The design rests on a simple empirical observation: flawed trajectories tend to be long. CalConf therefore learns a length-calibrated token-weight map and uses it to selectively terminate degrading trajectories during generation. We theoretically prove that CalConf attains the desired early-stop rate in practice without per model tuning. Across five benchmarks (on mathematical reasoning, scientific reasoning, and code generation tasks) and three open-source LLMs, CalConf consistently matches or exceeds offline majority-vote accuracy while reducing generated tokens by 20–44%. Token-level analyses trace these gains directly to CalConf’s ability to isolate harmful patterns that uniform-confidence baselines cannot distinguish.


Not All Noise Is Harmful: Towards Perception Aware and Controllable RAW Image Joint Denoising and Demosaicing

Qianjun Huang ⋅ Qingguo Liu ⋅ Hui Zeng ⋅ Kai Zhang ⋅ Jian Yang

RAW image denoising and demosaicing play a critical role in the early stages of the image signal processing (ISP) pipeline. Previous methods handle them sequentially, which can introduce error propagation. Unified restoration, such as joint denoising and demosaicing (JDD), has become a widely adopted paradigm and achieves impressive performance. Nonetheless, under constraints of model capacity and low-light conditions, these methods tend to apply overly aggressive denoising, causing notable smearing effects and the erosion of fine textures. Researchers attempt to moderate the denoising strength to better preserve details, but this often results in noticeable artifacts. These observations suggest a key insight: Not All Noise is Harmful—retaining an appropriate amount of noise can be preferable to over-smoothing and is often aligned with user aesthetics. Moreover, most existing methods are mainly trained on paired clean–degraded data, yielding fixed, hard-to-control denoising behavior that fails to meet evolving user demands. In this paper, we propose $\textbf{JDD-PRO}$, a $\textbf{P}$erception-awa$\textbf{R}$e and contr$\textbf{O}$llable framework for RAW image JDD. Built on a controllable JDD training scheme and a HVS-inspired perceptual adapter, our method allows users to tune the denoising strength at inference to strike their desired balance between noise removal and detail preservation. Comprehensive experiments on real-world datasets validate the effectiveness and robustness of our proposed method.

Class-incremental learning (CIL) with pre-trained models has increasingly adopted Mixture-of-Experts (MoE) adapters, which freeze the backbone and selectively route inputs to sparse adapter subsets, achieving competitive performance through efficient parameter reuse. However, as new tasks are learned sequentially, the router's output distribution gradually shifts away from previous states, undermining consistent adapter reuse and accelerating forgetting. Through a controlled instance-level analysis, we show that this effect is highly asymmetric, where instances with large routing drift account for most of the performance degradation, while those with small drift remain stable or can even benefit from it. This finding suggests that only excessive drift requires correction, while small drift may be better preserved than suppressed, and motivates us to propose Trust-Region Projected Routing (TRPR) that adapts the trust-region principle from constrained optimization to routing stabilization. It constrains each input's routing distribution within a KL-bounded trust region around a class-level anchor, correcting only excessive deviations while leaving small drift intact. To complement class-level anchoring with per-sample adaptability, TRPR additionally introduces an Instance-Adaptive (IA) path for fine-grained per-sample specialization, fused with the projected path via a learned gate. Extensive experiments on four CIL benchmarks show that TRPR consistently outperforms nine representative PEFT-based baselines, with accuracy gains of up to 3.06% in average and final accuracy and forgetting reduced by up to 2.48%.


Not All Slots Are Equal: Non-Co-Progressive Markov Bridge for Bundle Construction

Rongchao Zhang ⋅ Haodong Jing ⋅ Siheng Wang ⋅ Zhengtao Yao ⋅ Yajun Liu

Bundle construction (BC) is a critical technique in recommender systems research, which aims to select a coherent subset of items from large-scale item catalogs to either construct a complete bundle from scratch or complete a partial bundle with missing items. However, the extremely high-dimensional combinatorial spaces and heterogeneous semantic granularities involved pose significant challenges to BC benchmarking, and existing generative frameworks suffer from failing to precisely align the generated unknown slots with highly compatible known items. In this paper, we exploit conditional probability transitions in discrete state spaces and introduce a novel Non-co-progressive Markov bridge, NonMBB, a generalization of Markov bridge for constructing bundles of products. Specifically, given bundle's unique properties, NonMBB first improve Markov bridge for slot modeling through: i) categorizing item slots into high-compatibility slots and weakly-associated slots with respect to the known items, assigning them distinct timestep schedules to avoid the pitfall where high-compatibility slots must reference noisy, uninformative context at identical noise levels; ii) applying introduce schedule reweighting to constrain the maximum timestep disparity across slots, thereby preventing noise residuals caused by excessive asynchrony. Extensive experiments on Spotify and POG demonstrate that NonMBB establishes a substantially improved optimal bipartite matching over existing diffusion methods, yielding higher embedding-space alignment independently of exact-match recall.


Nüwa.RNA: An RNA Foundation Model for Unified Representation with Deep Structure Infusion

Kun Huang ⋅ Jiyang Li ⋅ Xin Guo ⋅ LIMEI HAN ⋅ Yuan Cheng

RNA function arises from a hierarchical folding process in which primary sequences form secondary structures that govern biological activity. While foundation models have proven effective for learning transferable RNA representations, most existing approaches rely on primary sequence alone or incorporate structural information only superficially during pre-training. Three problems limit current approaches: experimentally determined secondary structures are scarce, computational pseudo-labels are noisy, and no existing model integrates structure at multiple levels across masking, architecture, and supervision. Here, we present \textbf{N\"uwa.RNA}, a pre-trained foundation model built on a \emph{structure-driven multi-level pre-training framework} that integrates secondary structure at all three stages of pre-training simultaneously. N\"uwa.RNA introduces dual-granularity masking that combines single-token masking with structure motif masking, treating entire structural elements as indivisible units; a noisy structural attention bias that injects curriculum-corrupted contact maps into every Transformer layer to model pseudo-label uncertainty; and a structure denoising objective that recovers clean base-pairing states from corrupted input, completing a closed denoising loop. Pre-trained on a large-scale non-coding RNA corpus, N\"uwa.RNA achieves state-of-the-art performance on 10 of 13 BEACON benchmark tasks, with up to $+$7.83 F1 improvement on structure prediction, and attains the highest zero-shot secondary structure prediction accuracy among single-sequence RNA foundation models.


Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty

Sarina Penquitt ⋅ Jonathan Klees ⋅ Antonia van Betteray ⋅ Parssa Jashnieh ⋅ Peter Stehr ⋅ Matthias Rottmann ⋅ Lars Schmarje

While object detection has advanced through improved architectures and open-vocabulary models, we show that benchmark quality is fundamentally limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We demonstrate that current detectors are misaligned with human uncertainty, and that training with soft labels improves calibration and better reflects human perception. Existing label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real-world label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation.

Irregular data assimilation is difficult because sparse, off-grid observations do not constrain the same state corrections from one cycle to the next. As the observation geometry changes, the set of supported correction directions changes as well. Many learning-based DA methods infer updates in a fixed latent or control space, even when the current observation layout supports a different set of correction directions. To address this mismatch, we introduce Observability-Constrained DataAssimilation with Control-Space Inference (ObsConDA), which determines those supported correction directions from the current background and observation geometry and infers only their control coefficients. On real surface-station assimilation, ObsConDA gives the best held-out station accuracy and the best full-field reconstruction among the tested classical and learned baselines. The same advantage carries to matched synthetic observations, explicit geometry shifts, and synthetic dynamical systems, supporting the view that the usable correction directions must change from cycle to cycle. Ablations show that the gain comes from adapting those directions themselves rather than from adding a latent corrector alone.


Occupancy-based Quantile Risk Control

Zihao Shi ⋅ Huajun Xi ⋅ Bingyi Jing ⋅ Hongxin Wei

Conformal risk control is an emerging framework for the safe deployment of machine learning models with finite-sample guarantees. To accommodate a broader class of risk notions, quantile risk control extends this framework to quantile-based risk measures. However, existing methods either suffer from excessive conservatism or lack rigorous finite-sample guarantees. To address these limitations, we introduce Occupancy-based Quantile Risk Control (OQRC), a novel method that provides tight risk control bounds with finite-sample validity. Our key idea is to formulate risk control as a finite-occupancy problem by partitioning the loss space with the ordered calibration losses. Specifically, we estimate the distribution of test losses across the resulting bins and upper-bound the risk by the maximum loss attained within each bin. We then select the parameter $\lambda$ such that this upper bound does not exceed a predefined threshold $\alpha$ with high probability $1-\delta$. Theoretically, we establish a finite-sample guarantee showing that OQRC yields tight risk control bounds that converge to the optimal bounds at a provable rate of $\mathcal{O}_p(n^{-1/2})$. Extensive experiments demonstrate the effectiveness of our method, reducing the risk gap by up to 78.64\% on common benchmarks.


Offline Reinforcement Learning for Plasma Control in Nuclear Fusion: Codebase and Benchmark

YANG FU ⋅ Haomin Bao ⋅ Rohit Sonker ⋅ Xiaoyan Hu ⋅ Aravind Venugopal ⋅ Jeff Schneider ⋅ Jiayu Chen

Offline reinforcement learning (RL) offers a promising route for developing plasma controllers from historical tokamak data, since online trial-and-error on real devices is costly and risky. However, progress in this direction remains difficult to measure due to the lack of a standardized offline RL benchmark for realistic multi-actuator, long-horizon plasma control problems in nuclear fusion. We introduce RL4F, an Offline Reinforcement Learning Benchmark for Plasma Control in Nuclear Fusion, providing closed-loop evaluation environments and baseline comparisons across four full-profile tracking tasks: rotation, density, temperature, and pressure. The dynamics function underlying the evaluation environment is built from historical discharge data from DIII-D, a real-world Tokamak. We evaluate a broad set of imitation learning and offline RL baselines under a unified protocol. We find that offline model-based RL methods obtain the best average performance on most objectives, although no single method dominates all tasks, highlighting the importance of dynamics modeling in complex, long-horizon plasma control tasks. To foster further research, we open-source the codebase, datasets, and evaluation framework, providing a benchmark not only for the fusion community but also for algorithm development in offline RL. Our data and code are available at https://anonymous.4open.science/r/Anonymous-B71F


Off-policy Learning with Excursion Policies

Jiamin He ⋅ Mark Rowland ⋅ Daniel (Zhaohan) Guo ⋅ Hado van Hasselt ⋅ Diana Borsa

Off-policy reinforcement learning (RL) tries to learn the value of a target policy while data comes from a different behavior policy. This often requires trading off bias for variance or stability. One way to trade these off is to linearly interpolate between the behavior and target policies, but this represents only one point in a broader space of interpolation mechanisms. We introduce \emph{excursion policies}--a structured alternative where the agent follows the target policy for a multi-step ``excursion'' before reverting to the behavior policy. We derive dynamic programming results for these policies, establishing the theoretical basis for their use in RL. We further develop provably sound off-policy temporal-difference (TD) algorithms for excursion value estimation. We find that excursion policies provide an efficiency-stability trade-off comparable to mixture policies in continuous control benchmarks, and our analysis illustrates conceptual differences in their behavior.

Self-supervised representation learning has advanced time series forecasting by capturing robust features from unlabeled data. However, existing contrastive and masking-based methods often struggle to explicitly model the underlying temporal flow, either by treating time lag as noise or by disrupting the inherent temporal dynamics essential for forecasting. To address these limitations, we propose OliO (ODE-based Linear Transition Operator), a plug-and-play self-supervised learning method that learns bidirectional temporal flows as a continuous and structurally consistent dynamical system. OliO introduces a transition operator derived from linear Ordinary Differential Equations (ODEs) to align latent representations across arbitrary time shifts. By imposing a strictly upper triangular constraint on the transition matrix, we mathematically ensure numerical stability and enforce a robust polynomial boundary that effectively captures long-term dependencies while preventing the exponential divergence typically found in unconstrained ODE-based models. This structural constraint induces a temporally coherent representation space that preserves the underlying flow essential for accurate forecasting. Extensive experiments across various backbone architectures and benchmarks demonstrate that OliO achieves significant performance gains, providing Mean Squared Error (MSE) reductions of up to 5.75\% over the most competitive grid-searched SOTA baselines. The code is available at https://anonymous.4open.science/r/OliO-46EF.


Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamic Mechanisms and Efficient Alignment

Kun Wang ⋅ Zherui Li ⋅ Zhenhong Zhou ⋅ Jie Zhang ⋅ Yitong Zhang ⋅ Yan Mi ⋅ Kun Yang ⋅ Yiming Zhang ⋅ Zhongxiang Sun ⋅ Qiankun Li ⋅ Yang Liu

Omni-modal Large Language Models (OLLMs) greatly expand LLMs' multimodal capabilities but also introduce cross-modal safety risks. However, a systematic understanding of vulnerabilities in omni-modal interactions remains lacking. To bridge this gap, we establish a modality-semantics decoupling principle and construct the AdvBench-Omni dataset, which reveals a significant vulnerability in OLLMs. Mechanistic analysis uncovers a Mid-layer Dissolution phenomenon driven by refusal vector magnitude shrinkage, alongside the existence of a modal-invariant pure refusal direction. Inspired by these insights, we extract a golden refusal vector using Singular Value Decomposition and propose OmniSteer, which utilizes lightweight adapters to modulate intervention intensity adaptively. Extensive experiments show that our method not only increases the Refusal Success Rate against harmful inputs from 69.9% to 91.2%, but also effectively preserves the general capabilities across all modalities.


On Communication-Efficient Training of Ensembles in Federated Learning

Valery Parfenov ⋅ Mikhail Aleksandrov ⋅ Daniil Medyakov ⋅ Dmitry Bylinkin ⋅ Aleksandr Beznosikov

In recent years, ensemble ideas have found parallels in federated learning (FL). Specific ensemble learning formulations under the federated setup involve information exchange between clients during the training process, making the classic FL communication bottleneck a critical issue. Importantly, the resulting communication patterns differ from those studied in traditional communication-efficient FL methods, placing ensembling outside their analytical scope. In this work, we close this gap by extending classic communication-efficient techniques to the ensembling setting and introducing a new mechanism specifically tailored to it. We instantiate these ideas in four algorithms and provide convergence guarantees for each under the smooth non-convex setup. We further support our theoretical findings with experimental results.


On Differential Private $\ell_1$, $\ell_2$ and $\ell_p^p$ Distance Queries

Erzhi Liu ⋅ Jerry Yao-Chieh Hu ⋅ Alex Reneau ⋅ Zhao Song ⋅ Han Liu

We introduce a refined differentially private (DP) data structure for kernel density estimation (KDE) with $\ell_1$, $\ell_2$ and $\ell_p^p$ kernels. This new DP data structure offers not only improved privacy-utility tradeoff but also better query efficiency over prior results. Specifically, we study the mathematical problem: given a similarity function $f$ (or DP KDE) and a private dataset $X \subset \mathbb{R}^d$, our goal is to preprocess $X$ so that for any query $y \in \mathbb{R}^d$, we approximate $\sum_{x \in X} f(x, y)$ in a differentially private fashion. The best previous algorithm for $f(x, y) = |x - y|_1$ is the node-contaminated balanced binary tree by [Backurs, Lin, Mahabadi, Silwal, and Tarnawski, ICLR 2024]. Their algorithm requires $O(nd)$ space and time for preprocessing with $n = |X|$. For any query point, the query time is $\alpha^{-1} d \log^2 n$, with an error guarantee of $(1+\alpha)$-approximation and $\varepsilon^{-1} \alpha^{-0.5} d^{1.5} R \log^{1.5} n$. In this paper, we use the same space and pre-processing time, improve the best previous result [Backurs, Lin, Mahabadi, Silwal, and Tarnawski, ICLR 2024] in three aspects - We reduce query time by $\alpha^{-1} \log n$ factor - We improve the approximation ratio from $\alpha$ to $1$ - We reduce the error dependence by a factor of $\alpha^{-0.5}$ From a technical perspective, our method of constructing the search tree differs from previous work [Backurs, Lin, Mahabadi, Silwal, and Tarnawski, ICLR 2024]. In prior work, for each query, the answer is split into $\alpha^{-1} \log n$ numbers, each derived from the summation of $\log n$ values in interval tree countings. In contrast, we construct the tree differently, splitting the answer into $\log n$ numbers, where each is a smart combination of two distance values, two counting values, and $y$ itself. We believe our tree structure may be of independent interest.

In Large Language Model (LLM) fine-tuning, parameter and data selection are common strategies for reducing fine-tuning cost, yet they are typically driven by separate scoring mechanisms. When a parameter mask and data subset jointly determine restricted fine-tuning, this separation incurs redundant overhead and makes coordinated selection difficult. We cast parameter and data selection as two bilevel selection problems under a common validation objective and derive a shared local response-surrogate scoring rule. Under first- and second-order validation-improvement approximations, parameter importance and data utility emerge as column-wise and row-wise aggregations of a single gradient interaction matrix, yielding a closed-form row--column correspondence for co-extracting both signals. Building on this structure, we propose $\textbf{DualSFT}$ ($\textbf{Dual}$-$\textbf{S}$election $\textbf{F}$ine-$\textbf{T}$uning), a one-shot dual-scoring algorithm that produces a parameter mask and data subset from shared gradient statistics. On 3B$-$9B LLMs, single-axis DualSFT variants strengthen target-task performance and stability--plasticity trade-offs within their comparison groups, while full DualSFT yields a more favorable joint-constrained trade-off than sequential hybrid baselines under matched budgets.


ONE-SHOT: Compositional Human-Environment Video Synthesis via Spatial-Decoupled Motion Injection and Hybrid Context Integration

Fengyuan Yang ⋅ Luying Huang ⋅ Jiazhi Guan ⋅ Quanwei Yang ⋅ Dongwei Pan ⋅ Jianglin Fu ⋅ Haocheng Feng ⋅ Wei He ⋅ Kaisiyuan Wang ⋅ Hang Zhou ⋅ Angela Yao

Recent advances in Video Foundation Models (VFMs) have revolutionized human-centric video synthesis, yet fine-grained and independent control of subjects and scenes remains a critical challenge. Recent attempts to achieve joint human-environment control often rely on explicit human-scene 3D alignment, improving geometric control but sacrificing generative flexibility and long-horizon consistency while requiring heavy 3D pre-processing. We present ONE-SHOT, a parameter-efficient framework for compositional human-environment video generation. Our key insight is to factorize the generative process into disentangled signals, separating human dynamics from scene cues. Specifically, our canonical-space motion injection and Dynamic-Grounded-RoPE map canonical human motion to the target video region, enabling accurate motion and placement control without heuristic human-scene 3D alignment. Hybrid Context Integration further maintains subject and scene consistency for minute-level synthesis. Experiments show that ONE-SHOT consistently outperforms state-of-the-art methods in structural control, flexible composition, text-guided semantic control, and long-horizon consistency.


On Generalization in Bilevel Optimization with Overparameterized Models

Fares El Khoury ⋅ Edouard Pauwels ⋅ Samuel Vaiter ⋅ Michael Arbel

Bilevel optimization provides a general framework for machine learning problems where an outer objective depends implicitly on the solution of an inner problem. While significant attention has been devoted to its algorithmic aspects, its generalization properties remain far less understood. In this paper, we address this gap by studying the case where the inner solution is approximated by an overparameterized two-layer ReLU neural network trained by gradient flow. Under the Neural Tangent Kernel (NTK) regime, we derive non-asymptotic generalization bounds that characterize how statistical accuracy depends jointly on the inner and outer sample sizes, as well as on the regularity of the target function and the complexity of the hypothesis space. We further establish convergence guarantees for a gradient-based bilevel algorithm, where the inner level is approximately solved by applying an early stopping strategy to the gradient flow training. Finally, we complement our theoretical analysis with numerical experiments showing that the bilevel structure induces an initialization-dependent implicit bias that enables feature learning at the inner level and yields performance improvements beyond NTK predictions.


On JEPA Isotropy

Thomas YL Lin ⋅ Han Liu ⋅ Jerry Yao-Chieh Hu

We study the empirical risk minimization (ERM) of linear Joint Embedding Predictive Architecture (JEPA) and establish its key conditions for optimality. By casting the embedding bottleneck as a rank constraint, we formulate JEPA as a reduced-rank regression problem. This induces a risk decomposition into an irreducible error and an excess-risk bound. Within the irreducible term, we find the rank of the optimal predictor grows with target heterogeneity and saturates at the bottleneck. Within the excess-risk bound, we find spectral conditioning of the whitened end-to-end predictor governs minimization. Together, these conditions prescribe an isotropic regularizer. Under this regularization, we show such JEPA isotropy subsumes canonical JEPA isotropic embedding designs and hence justifies their empirical successes via ERM. Numerical experiments corroborate our theory.


Online Decision-Focused Learning under Semi-Bandit Feedback

Aabhash Dhakal ⋅ Tim Lachner ⋅ Jayanta Mandi ⋅ Marco Foschini ⋅ Christoph M Flath

Decision-focused learning (DFL) improves on the Predict-then-optimize (PtO) pipeline by training predictive models to minimize decision regret rather than prediction loss. Most DFL methods assume offline full-feedback data with complete action cost vectors. Many real-world decision systems instead operate in an online setting: in each round, the agent acts using past feedback and observes costs only for the selected components, yielding semi-bandit feedback. Bandit methods optimize regret in this regime, but train models with predictive updates rather than decision-focused updates. In this paper, we formalize online semi-bandit DFL: decision-focused learning for online optimization problems with semi-bandit feedback. The most straightforward extension would impute unobserved components with the model's own predictions and compute the decision-focused gradient. But this creates a self-confirming imputation degeneracy: imputation errors can flip counterfactual decisions that define the decision-focused gradient, biasing the representation update and reinforcing errors in later imputations. To address this degeneracy, we propose BayesianSPO, which combines an analytically updated Bayesian last-layer head with masked decision-focused learning. The Bayesian head supplies both (i) uncertainty to balance exploration and exploitation and (ii) stable imputations for unobserved components, while masking unobserved component gradients prevents those imputations from directly corrupting the representation update. Across knapsack, shortest-path, and portfolio benchmarks, BayesianSPO achieves the lowest mean cumulative regret on all five instances and the improvement is statistically significant on four of them.

We develop an online directional regression for sufficient dimension reduction with streaming data. Unlike first-moment methods, directional regression exploits both inverse conditional means and inverse conditional variances, and can therefore recover central subspace directions generated by linear structures, symmetric dependencies, and interaction driven relationships. Extending directional regression to the online setting is nontrivial because its kernel depends on slice-wise inverse moments and on standardization by an unknown covariance matrix. We address these challenges by utilizing a stable slice probability free kernel, a ridge-stabilized recursive least-squares estimator of the standardization operator, metric-aware subspace tracking under a vanishing-ridge covariance metric, and a weighted-bootstrap ladle criterion for online structural dimension selection. Under standard regularity conditions, the online kernel estimator is root-$t$ consistent, the estimated generalized eigenspace consistently recovers the central subspace, and the proposed dimension selector consistently estimates the structural dimension. Simulation studies and real-data analyses demonstrate that the proposed method serves as a robust online alternative when the dominant data structure is unknown, adapting to both first-moment and second-moment-driven dependencies while yielding substantial computational savings relative to repeated batch directional regression.


On Reparameterizing the Score Function

Mohammad Hossein Amani ⋅ Leonardo Bocchieri ⋅ Maxime Peyrard ⋅ Clement Hongler ⋅ Robert West

Policy gradients for autoregressive sequence models differentiate each token's log-probability while holding the conditioning history fixed. This discards information about how earlier decisions shape later conditional distributions, creating a credit-assignment bottleneck when rewards are sparse. We introduce History-Reparameterized (HR) policy gradients, which replace the frozen prefix in the score function with a differentiable reparameterized or relaxed history, allowing gradients of downstream log-probabilities to propagate backward through the generated trajectory. The resulting history-pathwise term is generally biased when added directly, so we derive Optimal HR (OHR), a centered control-variate estimator with a closed-form coefficient that minimizes the trace of the gradient variance at the population level. The variance reduction is governed by temporal structure in the policy dynamics: random or weakly structured policies provide little benefit, while policies with learned sequential dependencies yield informative history-pathwise corrections. For discrete tokens, the choice of relaxation is crucial: embedding-level relaxation yields substantial variance reduction that grows with horizon, whereas straight-through Gumbel-Softmax yields essentially none in our diagnostics. We validate the method on a Gaussian linear Markov model, a sparse-reward T-maze with a recurrent policy, and an autoregressive Transformer model for small Traveling Salesman Problem (TSP) instances.


On the Burden of Achieving Fairness in Conformal Prediction

Ziang Gao ⋅ Pengqi Liu ⋅ Archer Yang ⋅ Mouloud Belbahri ⋅ Jesse Cresswell ⋅ Masoud Asgharian

Conformal prediction is often calibrated with a single pooled threshold, but this can hide cross-group heterogeneity in score distributions and distort group-wise coverage. We study this phenomenon through the population score distributions underlying split conformal calibration. First, we derive a conservation law and lower bound showing that pooled calibration incurs irreducible group-wise coverage distortion at a scale set by cross-group quantile heterogeneity. Second, we demonstrate that the two leading fairness definitions for conformal prediction, Equalized Coverage and Equalized Set Size, are fundamentally in tension. Third, we quantify the cost of moving between policies which treat groups separately or pool them. Experiments on synthetic and real data confirm the same bidirectional trade-off after finite-sample calibration. Our results show that, for the policy families studied here, calibration choice does not remove cross-group heterogeneity; it determines whether the resulting distortion appears in the coverage or size dimension, providing a principled lens for analyzing fairness-oriented calibration choices in practice.

Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDP). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within $\mathcal{O}(1/\varepsilon^p)$ iterations to an $\varepsilon$-optimal policy, where the exponent $p$ depends on problem data. For single-period MDPs, $\mathcal{O}(1/\varepsilon)$ iterations are required.


On the Recall Scaling Laws in Mamba: A Theoretical and Mechanistic Study via Hashing

Yuval Koren ⋅ Assaf Ben-Kish ⋅ Raja Giryes ⋅ Lior Wolf ⋅ Itamar Zimerman

Associative Recall (AR) is the cognitive ability to learn and retrieve links between items in memory. In NLP, AR is used as a benchmark for evaluating the in-context memory capacity of architectures such as Mamba, and has been found to strongly correlate with language modeling performance. This paper explores AR from the perspective of mechanistic interpretability, aiming to reverse-engineer the exact internal algorithm used by Mamba to perform recall. Our key insight is that Mamba performs recall by implicitly learning linear hash functions, and we identify the low-level circuit that enables this behavior. Building on these findings and inspired by theoretical tools in similarity-preserving hashing, such as the Johnson–Lindenstrauss (JL) lemma, we develop a theoretical framework for analyzing AR, which we term recall scaling laws. For example, given a model’s state, context length, and embedding dimensions, our theory predicts a lower bound on the largest vocabulary size that can be perfectly recalled in the AR task. Empirical results show that this bound is tight and predictive, offering insights into how AR capacity scales with vocabulary size, state size, embedding size, and model architecture.


OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories

Yibing Liu ⋅ Yangze Liu ⋅ Xiao-Long Yin ⋅ Bin Wang ⋅ Chong Zhang ⋅ Hao Yin ⋅ Zhongyi Han

Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambiguity, unsafe external writes, ignored errors, weakly grounded commitments, or capability-boundary overcommitment. We study this mismatch as the Outcome-Process Gap and introduce OpenClawBench, a large-scale dataset for measuring and supervising process-side anomalies in real agent execution processes. OpenClawBench is built from BFCL-driven OpenClaw sessions produced by 6 source models and contains 31,264 annotated trajectories. It aligns task-oracle outcomes with structured process evidence. FullTax converts the aligned trajectories into structured anomaly supervision: binary labels, supporting evidence, onset/span localization, severity, recoverability, and a 5-class anomaly taxonomy. Using OpenClawBench, we make the Outcome-Process Gap measurable. Among 31,135 oracle-passing executions, 2,904 (9.33%) are still labeled process-anomalous under FullTax. Within oracle-passing executions that contain high-risk process evidence, 1,765 of 1,904 (92.70%) are labeled anomalous. These results show that success-only evaluation misses a concrete class of process-side failures in real agent executions. FullTax silver labels match human audit on 96.0% of a 300-trajectory human-audited pilot. A LoRA-fine-tuned Gemma 3 12B detector trained on the high-confidence FullTax supervised pool reaches binary F1=0.729 on the cleaner-labels held-out test split (n=2,646). It outperforms the GPT-5.4 frontier reference by +0.302 absolute, the no-fine-tuning base by +0.357, and wins on all six source-model agent slices. The gain comes from calibration rather than higher recall: zero-shot detectors over-flag at 42–50% against a 14.7% label rate, while the fine-tuned detector predicts anomalies 17.7% of the time. Together, OpenClawBench turns real agent execution logs into auditable and reusable supervision for studying, diagnosing, and operationally monitoring runtime agent reliability.


OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents

Shuang Chen ⋅ Kaituo Feng ⋅ Hangting Chen ⋅ Wenxuan Huang ⋅ Dasen Dai ⋅ Quanxin Shou ⋅ Yunlong Lin ⋅ Xiangyu Yue ⋅ Shenghua Gao ⋅ Tianyu Pang

Deep search has become a crucial capability for frontier multimodal agents, enabling models to solve complex questions through active search, evidence verification, and multi-step reasoning. Despite rapid progress, top-tier multimodal search agents remain difficult to reproduce, largely due to the absence of open high-quality training data, transparent trajectory synthesis pipelines, or detailed training recipes. To this end, we introduce OpenSearch-VL, a fully open-source recipe for training frontier multimodal deep search agents with agentic reinforcement learning. First, we curate a dedicated pipeline to construct high-quality training data through Wikipedia path sampling, fuzzy entity rewriting, and source-anchor visual grounding, which jointly reduce shortcuts and one-step retrieval collapse. Based on this pipeline, we curate two training datasets, SearchVL-SFT-36k for SFT and SearchVL-RL-8k for RL. Besides, we design a diverse tool environment that unifies text search, image search, OCR, cropping, sharpening, super-resolution, and perspective correction, enabling agents to combine active perception with external knowledge acquisition. Finally, we propose a multi-turn fatal-aware GRPO training algorithm that handles cascading tool failures by masking post-failure tokens while preserving useful pre-failure reasoning through one-sided advantage clamping. Built on this recipe, OpenSearch-VL delivers substantial performance gains, with over 10-point average improvements across seven benchmarks, and achieves results comparable to proprietary commercial models on several tasks. We will release all data, code, and models to support open research on multimodal deep search agents.


OpenVTON-Bench: A Large-Scale High-Resolution Benchmark for Controllable Virtual Try-On Evaluation

Jin Li ⋅ Tao Chen ⋅ Kai Wen ⋅ Siqi Yin ⋅ Shuai Jiang ⋅ Weijie Wang ⋅ Jingwen Luo ⋅ Chenhui Wu

Recent advances in diffusion models have significantly elevated the visual fidelity of Virtual Try-On (VTON) systems, yet reliable evaluation remains a persistent bottleneck. Traditional metrics struggle to quantify fine-grained texture details and semantic consistency, while existing datasets fail to meet commercial standards in scale and diversity. We present OpenVTON-Bench, a large-scale benchmark comprising approximately 100K high-resolution image pairs (up to $1536\times1536$). The dataset is constructed using DINOv3-based hierarchical clustering for semantically balanced sampling and Gemini-powered dense captioning, ensuring a uniform distribution across 20 fine-grained garment categories. To support reliable evaluation, we propose a multi-modal protocol that measures VTON quality along five interpretable dimensions: background consistency, identity fidelity, texture fidelity, shape plausibility, and overall realism. The protocol integrates VLM-based semantic reasoning with a novel Multi-Scale Representation Metric based on SAM3 segmentation and morphological erosion, enabling the separation of boundary alignment errors from internal texture artifacts. Experimental results show strong agreement with human judgments (Kendall's $\tau$ of 0.833 vs. 0.611 for SSIM), establishing a robust benchmark for VTON evaluation. Code and dataset are available \href{https://github.com/RenxingIntelligence/OpenVTON-Bench}{here}.


Optimal Rates for Differentially Private Hypothesis Testing with E-values

Ben Jacobsen ⋅ Tomás González Lara ⋅ Gavin R Brown ⋅ Kassem Fawaz ⋅ Aaditya Ramdas

E-values have attracted considerable interest in recent years as flexible tools for enabling anytime-valid and adaptive data analysis. Hypothesis testing is at the core of many of these applications, which can often involve private or sensitive data. In this work, we answer a simple but important question: given two distributions $\mathbb{P}$ and $\mathbb{Q}$, what is the maximum achievable e-power when testing $X\sim \mathbb{P}^n$ against $X\sim\mathbb{Q}^n$ with e-values that satisfy $\varepsilon$-differential privacy? We characterize the optimal rate for this problem and provide an algorithm which matches it asymptotically. In the sequential setting, when observations arrive one-by-one and the analyst chooses when to halt, we give matching upper and lower bounds on the stopping times of any differentially private e-process. Numerical experiments confirm the practicality of our algorithms, which require less data than the recently proposed DP-SPRT across a range of sequential testing problems and privacy levels.

Recent work shows that gradient descent (GD) can achieve almost optimal risk bounds for shallow ReLU networks with a logarithmic width. However, these discussions require $O(n^2)$ gradient computations to achieve this optimality, which is not appealing for modern machine learning problems with large sample size $n$. In this paper, we significantly improve the existing gradient complexity by showing that stochastic gradient descent (SGD) can achieve similar risk bounds with $O(n)$ gradient computations for shallow ReLU networks with a logarithmic width. As compared to GD, the analysis with SGD is more challenging as there are additional fluctuations incurred by the stochastic gradient noise along the optimization process. With concentration inequalities to handle these fluctuations, we show that the entire SGD trajectory stays around its initialization point with high probability. Under an NTK separability condition with margin $\gamma$, we show SGD achieves near-optimal risk bounds $\widetilde{O}(1/(n\gamma^2))$ with a logarithmic width.

Machine learning models often suffer performance degradation under subpopulation shift, particularly when spurious correlations cause models to rely on shortcut features that fail to generalize across subgroups. A recent line of work mitigates this issue by using loss-based signals to identify informative samples, but these signals can become severely distorted under label noise: mislabeled samples may also incur large losses and contaminate subsequent reweighting or retraining. Despite its practical importance, this intersection remains largely underexplored. We propose POTER, a reweighting framework based on optimal transport that derives sample importance from the transport geometry between the training distribution and a reference distribution constructed from limited validation group annotations. By measuring alignment at the individual-sample level rather than relying on loss, POTER downweights mislabeled or strongly bias-aligned samples while assigning higher importance to samples better aligned with the reference distribution. In addition, POTER requires only a single ERM training stage, moving beyond the retraining paradigm common in recent work. Across standard benchmarks and noisy-label settings, POTER achieves state-of-the-art worst-group accuracy, including cases where label corruption is concentrated within minority subgroups.


ORCA: Orthogonal Residual Consensus Alignment for Multi-View Clustering

Xiaojian Ding ⋅ Xian Li ⋅ Xiaoying Zhu ⋅ Lin Zhao ⋅ Kaixiang Wang

Deep multi-view clustering often struggles with inter-view feature redundancy and dimensionality collapse during iterative representation fusion. To address these limitations, we propose Orthogonal Residual Consensus Alignment (ORCA). First, a Semantic Incremental Recurrent Fusion (SIRF) module explicitly separates shared redundancy from complementary residuals via geometric projection. By employing an adaptive halting mechanism, SIRF autonomously maximizes information gain while avoiding structural collapse. Second, a Consensus-Calibrated Hierarchical Contrastive Alignment (CCHA) module is introduced to preserve multi-scale semantics. It adaptively aligns view-level, local, and recurrent representations using the global consensus as a semantic anchor. Consequently, ORCA learns latent representations with robust structural integrity and highly discriminative semantics. Extensive experiments on multiple benchmarks demonstrate that ORCA consistently outperforms state-of-the-art methods, with comprehensive ablations verifying its robustness.


Ordered Policy Optimization

Dongkyu (Derek) Cho ⋅ Pan Xu

Proximal Policy Optimization (PPO) stabilizes policy updates by clipping the importance ratio between the new policy and the data-collecting policy. The standard rule uses one fixed clipping range for all samples. While this choice offers simplicity and effectiveness, it fails to distinguish statistically harmful extreme ratio values from large, yet tolerable changes. We study how to adapt the clipping radius to the reliability of the induced importance ratio distribution. Our case study demonstrates that reliable updates require controlling these extreme ratios, rather than merely verifying that individual ratios fall within a fixed interval. Motivated by this observation, we propose Ordered Policy Optimization (\textsc{OPO}). In each minibatch, \textsc{OPO} sorts the importance ratios, selects the extreme ratios that potentially have the largest effect on the update, and assigns them adaptive sample-wise radii instead of using the fixed clipping radius. These radii are derived from a prescribed tail profile, which specifies desired target values for the selected extreme ratios. The resulting objective remains first-order and requires only minibatch sorting and a modified policy loss. Experiments conducted on MuJoCo and Atari benchmarks demonstrate that this simple modification achieves strong performance relative to established baselines.


Outlier-robust Diffusion Posterior Sampling for Bayesian Inverse Problems

Yiming Yang ⋅ Xiaoyuan Cheng ⋅ Yi He ⋅ Kaiyu Li ⋅ Wenxuan Yuan ⋅ Zhuo Sun

Diffusion models have emerged as powerful learned priors for Bayesian inverse problems (BIPs). Diffusion-based solvers rely on a presumed likelihood for the observations in BIPs to guide the generation process. Likelihood misspecification is common in practical BIPs and is known to degrade recovery performance, particularly under outlier contamination. We investigate this problem by first characterizing the induced posterior deviation and proving the stability of diffusion-based solvers for linear BIPs. Our stability analysis further reveals potential robustness deficiencies of existing diffusion-based solvers under outlier-contaminated measurements. To address this issue, we propose a simple yet effective solution: robust diffusion posterior sampling, which is provably outlier-robust for linear BIPs and compatible with existing gradient-based posterior samplers. Empirical results from scientific inverse problems and natural image tasks demonstrate the effectiveness and robustness of our method, with consistent performance gains in challenging scenarios involving outlier contamination for both linear and nonlinear tasks.


Overcoming the Resolution Limit: Significance-Aware Regularization for Intersectional Fairness

Antonio Ferrara ⋅ André Panisson ⋅ Francesco Cozzi ⋅ Alan Perotti ⋅ Francesco Bonchi

Intersectional fairness requires evaluating models across fine-grained subgroups. However, as group definitions become more granular, the sample size per subgroup shrinks, causing fairness estimates to become increasingly unstable. Existing in-processing methods are certainty-blind: they penalize observed disparities without distinguishing statistically certified unfairness from sampling noise. This leads to fairness overfitting, where models sacrifice utility to "fix" spurious disparities in sparse regimes. To address this, we introduce the significant-excess metric family, a continuous, uncertainty-aware fairness score that measures only the portion of a subgroup’s disparity exceeding what sampling variation can explain at a specific confidence level. Leveraging this, we propose SAFER (Significance-Aware Fairness Regularizer), a differentiable in-processing method that gates fairness penalties using a soft-count likelihood-ratio significance test. SAFER satisfies a provable safety property: when no subgroup disparity is statistically certifiable, its fairness gradient is exponentially small and training reduces to unconstrained optimization. Mechanism ablations across three benchmark datasets and purpose-built stress variants confirm that each component of SAFER – threshold smoothing, the significance gate, and a dense-tail term – is empirically necessary in its theoretically predicted operating regime. SAFER consistently achieves a superior fairness-utility Pareto frontier, compared to baselines methods in sparse intersectional settings while preserving utility when certified unfairness is absent.


Oversmoothing as Representation Degeneracy in Neural Sheaf Diffusion

Arif Dönmez ⋅ Ellen Fritsche ⋅ Axel Mosig ⋅ Katharina Koch

Neural Sheaf Diffusion (NSD) generalizes diffusion-based Graph Neural Networks by replacing scalar graph Laplacians with sheaf Laplacians whose learned restriction maps define a task-adapted geometry. While the diffusion limit of NSD is known to be the space of global sections, the representation-theoretic structure of this harmonic space remains largely implicit. In this paper, we develop a quiver-theoretic interpretation of NSD by identifying cellular sheaves on graphs with representations of the associated incidence quiver. Under this correspondence, learned sheaf geometries become points in a finite-dimensional representation space. We prove that direct-sum decompositions of the underlying incidence-quiver representation induce corresponding decompositions of the harmonic space reached in the diffusion limit. This provides an algebraic interpretation of oversmoothing as representation degeneration: a conceptual framing where learned sheaves collapse toward trivial or low-complexity summands whose global sections fail to preserve discriminative information. Building on this viewpoint, we connect sheaf diffusion to stability, moduli, and moment-map principles from Geometric Invariant Theory. We introduce moment-map-inspired regularizers that bias learned restriction maps toward more balanced representation geometries, and we identify a structural obstruction in standard equal-stalk architectures: when $d_v=d_e$, the admissibility condition for learnable stability parameters forces the trivial all-object summand onto a stability wall. We show that non-uniform stalk dimensions remove this obstruction, making adaptive stability meaningful in principle. Empirical evaluations on heterophilic benchmarks are consistent with this mechanism: breaking stalk symmetry can reduce variance or improve validation behavior on some datasets, and adaptive stability regularization becomes more effective in selected rectangular settings. These results support the view that moment-map regularization is a structured but dataset-dependent geometric bias rather than a universal performance booster. Overall, our framework interprets oversmoothing not only as a spectral pathology, but as a degeneration phenomenon in the underlying representation geometry.


PAMod: Modeling Cyclical Shifts via Phase-Amplitude Modulation for Non-stationary Time Series Forecasting

Yingbo Zhou ⋅ Yutong Ye ⋅ Shuhao Li ⋅ Rui Qian ⋅ Qiang Huang ⋅ Lemao Liu ⋅ Li Sun ⋅ Dejing Dou

Real-world time series forecasting faces the fundamental challenge of non-stationary statistical properties, including shifts in mean and variance over time. While reversible instance normalization (RevIN) has shown promise by stationarizing inputs and denormalizing outputs, it relies on the strong assumption that historical and future distributions remain identical. We observe that in many practical applications, distribution shifts follow cyclical patterns that correlate with periodic positions (e.g., seasonal and holiday volatility). To this end, we propose $\textbf{PAMod}$, a lightweight yet powerful framework that models cyclical distribution shifts via $\textbf{P}$hase-$\textbf{A}$mplitude $\textbf{Mod}$ulation in the normalized feature space. PAMod learns periodic embeddings to modulate representations: phase modulation captures mean shifts, while amplitude modulation adapts to variance changes. Crucially, we prove mathematically that modulating in normalized space is equivalent to applying dynamic denormalization, offering an elegant unification of distribution adaptation and representation learning. Extensive experiments on twelve real-world benchmarks demonstrate that PAMod achieves state-of-the-art performance with fewer computational resources. Furthermore, our modulation mechanism, as a novel plug-and-play technique, can improve existing time-series forecasting methods with simple integration.

Continuous-token autoregressive (AR) generation directly models continuous signals without discretizing them into vocabulary tokens, but suffers from a severe train--inference mismatch: teacher-forced training conditions on ground-truth prefixes, whereas inference conditions on model-generated continuous tokens. This mismatch becomes especially severe in pixel-space image generation, where each token is a high-dimensional raw pixel patch and autoregressive errors can accumulate over generation steps. Rollout-based training can reduce this mismatch but is prohibitively slow, especially when each token is produced by a multi-step diffusion head. We propose \emph{Parallel Rollout Approximation} (PRA), which approximates rollout-based training by constructing training inputs in parallel at all positions. PRA combines a shared one-step denoiser with an end-to-end learned low-dimensional intermediate state, aligning these constructed training inputs with inference-time generated outputs without relying on a separately pretrained tokenizer. On class-conditional ImageNet-1K generation at 256$\times$256 resolution, even PRA-S with 135M parameters achieves an FID of 2.58, surpassing the previous billion-scale pixel-space AR result of 3.60. Scaling to PRA-L (511M) further improves FID to 1.94, establishing a new state of the art among pixel-space AR models and substantially narrowing the gap to pixel-space diffusion models. Beyond generation, PRA achieves higher ImageNet classification probing accuracy than latent-space AR baselines, suggesting its potential for unified pixel-space image generation and understanding.


Partially Performative Prediction

Jaewook Lee ⋅ Tijana Zrnic

Performative prediction studies feedback loops that arise when predictive models are deployed in consequential domains. In these settings, model updates can shape the population whose patterns they aim to predict, inducing a distribution shift that is endogenous to the learning system. This perspective departs from classical treatments of distribution shift, where shifts are typically modeled as exogenous changes in the data-generating process. Yet, in practice, distribution shift is rarely one or the other. Predictive models may influence future data through the decisions they support, while the world itself continues to drift for reasons beyond the learner’s control. We study partially performative prediction, a framework that captures both endogenous and exogenous sources of distribution shift. The framework generalizes performative prediction by allowing the data distribution to evolve both in response to the deployed model and according to an external, time-varying process. We extend the central notions of performative stability and performative optimality to this setting by defining online analogues that measure performance relative to the evolving partially performative environment. We analyze practical learning heuristics, including repeated retraining, and characterize when they successfully adapt to partially performative environments.


Patch Hierarchical Attention Transformers for Efficient Particle Jet Tagging

Zihan Zhao ⋅ Aaron Wang ⋅ Alan Xia ⋅ Chang Sun ⋅ Javier Duarte ⋅ Abhijith Gandrakota ⋅ Jennifer Ngadiuba ⋅ Richard Cavanaugh

Real-time jet tagging is critical for identifying short-lived particle decays in the high-throughput detectors of the Large Hadron Collider, where real-time trigger systems which are responsible for deciding which collision events to store impose strict latency and accuracy constraints. Transformer architectures achieve the highest jet tagging accuracy when compute is unconstrained, but their quadratic self-attention cost places them orders of magnitude beyond the trigger budget. Existing efficient variants reduce this cost by compressing the attention matrix or restricting it to ordered local windows, at the price of the explicit particle-particle interactions that drive substructure identification. To address this limitation, we introduce the Patch Hierarchical Attention Transformer (PHAT-JeT), which combines two mechanisms: a physics-inspired geometric message-passing module that encodes local detector-plane structure, and a hierarchical patch-based attention scheme that computes exact attention within small particle groups while preserving global context through lightweight patch-token communication. Within this compute budget, PHAT-JeT achieves state-of-the-art accuracy and background rejection among resource-constrained jet tagging models on four benchmarks (hls4ml, JetClass, Top Tagging, and Quark-Gluon). Our code is available at https://anonymous.4open.science/r/PHAT-JeT-540B/README.md


Path-Guided Flow Matching for Dataset Distillation

xuhui li ⋅ Zhengquan luo ⋅ Zixu Wu ⋅ Xiwei Liu ⋅ Yongqiang Yu ⋅ Zixiang Hong ⋅ Zhiqiang Xu

Dataset distillation compresses large datasets into compact synthetic sets with comparable performance in training models. Despite recent progress on diffusion-based distillation, such methods typically rely on heuristic guidance or prototype assignment over long denoising chains, which increases sampling cost and makes prototype-consistent control harder under strong guidance or low IPC. We propose \emph{Path-Guided Flow Matching (PGFM)}, the first flow matching-based framework for generative distillation, which enables deterministic synthesis by solving an ODE in a few steps. In particular, we introduce a retrieval-based prototype inversion stage that identifies prototype-consistent initial noises for class prototypes on the frozen flow manifold, and further develop an anchor-guided residual correction strategy for bounded stage-wise control. This design follows a controlled-transport principle: retrieval reduces initialization mismatch, while anchor guidance provides a bounded residual correction along the flow trajectory. Extensive experiments across high-resolution benchmarks demonstrate that PGFM matches or surpasses prior diffusion-based distillation approaches with fewer sampling steps while delivering competitive performance with improved efficiency.


Pause and Reflect: Conformal Aggregation for Chain-of-Thought Reasoning

Yu Gu ⋅ Zijun Yu ⋅ Vahid Partovi Nia ⋅ Masoud Asgharian

Chain-of-thought (CoT) reasoning with self-consistency improves performance by aggregating multiple sampled reasoning paths. In this setting, correctness is no longer tied to a single reasoning trace but to the aggregation rule over a pool of candidate paths, making aggregation uncertainty the central challenge. This issue is critical where confidently incorrect answers are far more costly than abstentions. We introduce a conformal procedure for CoT reasoning that directly addresses aggregation uncertainty. Our approach replaces majority voting with weighted score aggregation over reasoning paths and calibrates an abstention rule using conformal risk control. This approach leads to finite-sample guarantees on the confident-error rate--the probability that the system answers and is wrong. We further identify score separability as the key condition under which abstention provably improves selective accuracy, and derive closed-form expressions that predict accuracy gains from calibration data alone. The method is fully inference-time, and requires no retraining. Across four benchmarks, four open-source models, and three score classes, realized confident-error rates are consistent with the prescribed targets up to calibration-split and test-set variability. Our method achieves $90.1\\%$ selective accuracy on GSM8K by abstaining on less than $5\\%$ of problems, compared with $82\\%$ accuracy under majority-voting baseline.


PDF-HR: Pose Distance Fields for Humanoid Robots

Yi Gu ⋅ Yukang Gao ⋅ Yangchen Zhou ⋅ Xingyu Chen ⋅ Yixiao Feng ⋅ Mingle Zhao ⋅ Yunyang Mo ⋅ Zhaorui Wang ⋅ Lixin Xu ⋅ Renjing Xu

Pose and motion priors play a crucial role in humanoid robotics. Although such priors have been widely studied in human motion recovery (HMR) domain with a range of models, their adoption for humanoid robots remains limited, largely due to the scarcity of high-quality humanoid motion data. In this work, we introduce Pose Distance Fields for Humanoid Robots (PDF-HR), a lightweight prior that represents the robot pose distribution as a continuous and differentiable manifold. Given an arbitrary pose, PDF-HR predicts its distance to a large corpus of retargeted robot poses, yielding a smooth measure of pose plausibility that is well suited for optimization and control. PDF-HR can be integrated as a reward shaping term, a regularizer, or a standalone plausibility scorer across diverse pipelines. We evaluate PDF-HR on various humanoid tasks, including single-trajectory motion tracking, general motion tracking, style-based motion mimicry, and general motion retargeting. Experiments show that this plug-and-play prior consistently and substantially strengthens strong baselines.


Performance-Driven Policy Optimization for Speculative Decoding with Adaptive Windowing

Jie Jiang ⋅ xing sun ⋅ RuoTian Chen ⋅ Jianan Su ⋅ Kaixin Shen

Speculative decoding accelerates LLM inference by having a lightweight draft model propose speculative windows of candidate tokens for parallel verification by a larger target model. In practice, speculative efficiency is often bottlenecked by hard-to-draft positions, where an early mismatch truncates the accepted prefix and invalidates the rest of the speculative window. Most learning-based drafters are still optimized with token-level supervised objectives, even though speculative utility is inherently window-level and prefix-sensitive. We propose **PPOW** (**P**erformance-Driven **P**olicy **O**ptimization with Adaptive **W**indowing), a reinforcement learning framework that shifts drafter optimization from token-level imitation to window-level optimization. PPOW combines a Cost-Aware Speedup Reward, a Distribution-Based Proximity Reward, and Adaptive Divergence-Aware Windowing, which prioritizes informative windows with high confidence-weighted draft--target divergence. PPOW achieves average acceptance lengths of 6.29–6.52 and speedups of 3.39–4.36$\times$ across multiple model families and benchmarks under a unified decoding protocol. These results show that performance-driven window-level optimization is a practical approach to improving speculative decoding efficiency.


Permutation-Invariant Spectral Learning via Dyson Diffusion

Tassilo Schwarz ⋅ Cai Dieball ⋅ Constantin Kogler ⋅ Renaud Lambiotte ⋅ Arnaud Doucet ⋅ Aljaz Godec ⋅ George Deligiannidis

Diffusion models are central to generative modeling and have been adapted to graphs by diffusing adjacency matrix representations. The challenge of having up to $n!$ such representations for graphs with $n$ nodes is only partially mitigated by using permutation-equivariant learning architectures. Despite their computational efficiency, existing graph diffusion models struggle to distinguish certain graph families and their spectra, unless graph data are augmented with ad hoc features. This shortcoming stems from enforcing the inductive bias within the learning architecture. In this work, we leverage random matrix theory to analytically extract the spectral properties of the diffusion process, allowing us to push most of the inductive bias from the architecture into the dynamics. Building on this, we introduce the Dyson Diffusion Model, which employs Dyson's Brownian motion to capture the spectral dynamics of an Ornstein-Uhlenbeck process on the adjacency matrix. Furthermore, conditioned on the spectral dynamics, we formulate a Lie group diffusion, appropriately modeling the remaining degrees of freedom. Strikingly, the resulting learning problem becomes permutation invariant at the Lie algebra level. We demonstrate that the Dyson Diffusion Model learns graph spectra accurately and outperforms existing graph diffusion models.


Permute-then-Adapt: Weak-to-Strong Contrastive Image--Text Adaptation

Jinhao Li ⋅ Sarah Erfani ⋅ Lei Feng ⋅ Guangrui Li ⋅ James Bailey ⋅ Feng Liu

Image--text alignment models such as CLIP are typically trained with contrastive learning, where unpaired examples are treated as strict negatives. While this assumption introduces noise by penalising semantically related pairs, existing solutions often rely on heuristic similarity thresholds that lack statistical grounding and sensitivity to dataset-specific noise distributions. In this paper, we propose **Permute-then-Adapt (PTA)**, a framework that leverages efficient *weak models* to identify missed positives, while enforcing *statistical control* to reject random associations. Unlike standard distillation or arbitrary thresholding, PTA estimates the null distribution of the weak teacher's similarities via permutation testing, ensuring selected pairs are *statistically distinguishable from the distribution of random pairings* at level $\alpha$. The calibrated positive set is then trained against by a single multi-positive contrastive objective: each anchor maximises the total probability mass it assigns to its positive set. We further demonstrate that the calibration can be pre-computed offline using the weak model, so training overhead is negligible compared to standard baselines. Extensive experiments show that PTA consistently outperforms heuristic soft-label approaches on object recognition and cross-modal retrieval, while exhibiting superior data efficiency.


Personalize-then-Store: Benchmarking and Learning Personalized Memory for Long-Horizon Agents

Yeonjun In ⋅ Wonjoong Kim ⋅ Sangwu Park ⋅ Kanghoon Yoon ⋅ Chanyoung Park

Existing large language model (LLM)-based memory systems apply universal, static policies that overlook a fundamental reality: the contexts that are worth storing in memory are different across users. This misalignment wastes limited memory budget on transient interactions while failing to preserve critical context for long-horizon tasks. To address this gap, we investigate an underexplored question: can LLM-based memory systems learn personalized memory policies? We introduce PerMemBench, the first benchmark for evaluating personalized memory systems, featuring multi-year, multi-domain interaction histories across diverse user personas. We further present the first empirical study of memory personalization and propose simple baseline methods. Our empirical study confirms that personalization yields substantial retention gains when the user profile is exactly inferred, yet reveals that accurate profile inference remains an open and critical challenge.

Controlling persona in large language models (LLMs) at inference time is important for role-playing, personalized dialogue, and social simulation. Recent methods extract persona vectors from the model's activation space and apply Euclidean operations—addition, scaling, and linear interpolation—under the linear representation hypothesis. However, these methods themselves report systematic failures: non-orthogonal trait dimensions, asymmetric ceiling and resistance effects, and significant deviations in multi-trait composition, suggesting that the linear isotropic assumption does not hold. We propose PersonaManifold, a framework that models persona representations as points on a curved, low-dimensional Riemannian submanifold in activation space. We estimate the manifold's intrinsic geometry—local metric tensors, geodesic distances, and Ollivier–Ricci curvature—and introduce geodesic steering, which interpolates between personas along manifold geodesics rather than Euclidean straight lines. We also propose the Behavioral Similarity Triplet (BST) benchmark, which automatically generates situational questions grounded in six established psychological constructs and defines persona similarity through behavioral responses rather than self-report questionnaires. Experiments on three open-source LLMs show that persona activations form a manifold with heterogeneous curvature, geodesic distance predicts behavioral similarity more accurately than Euclidean alternatives with independent contributions from anisotropy and curvature, and geodesic steering produces more coherent intermediate personas on both our BST benchmark and external evaluations, with the advantage concentrated in high-deviation regions where the manifold deviates most from flatness.


Persona-Model Collapse in Emergent Misalignment

Davi Bastos Costa ⋅ Renato Vicente

Fine-tuning large language models on narrow data with harmful content produces broadly misaligned behavior on unrelated prompts, a phenomenon known as *emergent misalignment*. We propose that emergent misalignment involves *persona-model collapse*: deterioration of the model's internal capacity to simulate, differentiate, and maintain coherent personas. We test this hypothesis behaviorally using two metrics: moral susceptibility and moral robustness; computed as the cross-persona and within-persona variability of models' Moral Foundations Questionnaire responses under persona role-play, respectively. We evaluate four frontier models (DeepSeek-V3.1, GPT-4.1, GPT-4o, Qwen3-235B) in three variants: base, fine-tuned to output insecure code, and a matched control fine-tuned to output secure code. Across the four models, insecure fine-tuning produces a $55$% average spike in moral susceptibility, pushing all four insecure variants beyond the band observed across 13 frontier models benchmarked in prior work, with GPT-4o reaching more than twice its upper end. It also causes a $65$% average drop in moral robustness (equivalently, a $304$% surge in within-persona variability), while the secure control preserves susceptibility near the base and induces only a partial robustness loss. Complementing these metric shifts, insecure variants' unconditioned responses converge toward saturation near the scale ceiling, departing markedly from both base models' differentiated responses and those elicited when they role-play toxic personas. Taken together, these metrics provide a sensitive diagnostic for emergent misalignment and serve as behavioral evidence that it involves persona-model collapse.

Score-based diffusion is well-studied at the mixture level—averaged over the prior—but many practical pipelines (posterior sampling, SDEdit, OOD detection) fix the starting point and care about the per-start output. To our knowledge, this per-starting-point regime has no prior non-asymptotic analysis. We give two coupled bounds: a plausibility upper bound on how close the per-start output is to the data distribution, and a matching diversity lower bound on how distinguishable outputs from two different starts remain. The same score-contraction rate governs both—where plausibility is tight, diversity is vacuous, and vice versa—predicting a sharp horizon beyond which starting-point identity is erased. We illustrate the phenomenon qualitatively on 2D toy datasets, MNIST, and CIFAR-10.

Models that are indistinguishable on in-distribution data can behave very differently under distribution shift. We introduce Perturb-and-Correct (P&C), a post-hoc method for constructing epistemically diverse predictors from a single pretrained network. P&C applies random hidden layer perturbations with a least-squares correction in the subsequent affine layer, producing predictors that agree on calibration data while remaining free to disagree away from it. We analyze this mechanism through the post-correction residual and its first-order sensitivity: the residual is controlled near the calibration distribution by a leverage term, while corrected sensitivity grows as inputs deviate from the calibration geometry. Empirically, P&C achieves a strong ID/OOD tradeoff across MuJoCo dynamics prediction and CIFAR-10 OOD detection, matching or outperforming standard post-hoc baselines while requiring only a single pretrained model. Our findings highlight the potential in further exploiting overparameterization as a strength of deep learning models.


PhaseDance: Capturing Rhythm and Expressivity in Dance Modeling

Meongeun Kim ⋅ Taehui Lee ⋅ Soomin Park ⋅ Sung-Hee Lee

Generating dance that reflects many aspects of music remains a fundamental challenge in dance motion synthesis. Existing music-conditioned approaches formulate the task as beat-level matching, which captures local synchrony but misses the compositional structure of rhythm, where movement is organized into global beat units. We argue that musical dance arises from the alignment of \emph{periodicity} between motion and music, not from instantaneous beat coincidence. To address this, we propose \textbf{PhaseDance}, a phase-conditioned framework that represents dance as a quasi-musical signal on periodic phase manifolds, enabling unified dance generation across multiple tasks. Alongside this, we ground text descriptions in fifteen music-correlated metrics derived from Laban Movement Analysis to capture the qualitative richness of dance, articulating structured dimensions of motion — effort, shape, and dynamics — that coarse genre labels cannot convey. Experiments show that PhaseDance produces choreography with stronger rhythmic coherence and richer expressive understanding than existing models.

Optical microscopy enables rapid, label-free imaging of live bacteria and is the standard instrument for species identification across clinical, environmental, and industrial microbiology. Yet field samples are routinely polymicrobial and may contain organisms that were never seen during system training, and no computer-vision benchmark tests multi-label species identification from phase-contrast microscopy (PCM) of such mixtures. We introduce Phase-contrast Optical bEnchmark for Bacterial Identification ($\textbf{PHOEBI}$), a wet-lab-prepared dataset of $\textbf{120{,}000}$ PcM images covering $\textbf{40}$ combinations of six rod-shaped species, paired with a leave-combinations-out (LCO) evaluation protocol that holds out entire species combinations to mirror the practical scenario of a model trained on catalogued mixtures that must generalise to unseen ones. On LCO, every gradient-trained per-image aggregator we test drops $0.39$ to $0.57$ F1 from the in-distribution to the held-out split, a systematic open-world recognition failure in the aggregator, not the visual representation. A linear probe of thirteen different encoders over the same features spreads only about six percentage points of F1 across general-purpose and biomedical pretraining objectives, confirming the representation is sound. We propose three lightweight $\textit{anchor-based}$ decoders that capture per-species presence geometrically over a shared frozen tile-feature pool, scoring $\textit{higher}$ on held-out combinations than on in-distribution validation. A single reconstruction residual from the strongest decoder then unifies the remaining open-world primitives at no additional training cost: open-set rejection lifts area under the receiver-operating-characteristic curve (AUROC) from chance to $0.70$, and novel-class discovery clusters high-residual tiles to propose one new prototype per novel species at perfect purity with negligible drift on the known classes.

Code-reasoning agents trained from rollouts typically rely on costly verifiers — expert annotation, hand-written unit tests, or learned reward models — each of which scales poorly across domains. We argue that in domains with closed-form physical laws, the laws themselves form a free, dense, automatic class of reward signals. We instantiate this for physics via PhysAgentGym, an agentic code- execution gym in which a language model receives a physics scenario, writes Python that simulates it, and emits a trajectory scored by a rule-based verifier that compiles natural-language standards into executable trajectory predicates. Verifier signals are label-free: physical laws supply the ground truth without any per-task annotation. Across four physics subcategories totaling 1,320 closed-form- simulator task instances, two frontier models, Claude Sonnet 4.5 and OpenAI GPT-5, achieve nearly-identical aggregate scores (0.944 and 0.945 at N=30 each), with per-cell agreement within 0.5 pp, providing independent evidence the verifier is well-calibrated rather than arbitrary. We filter 491 successful frontier trajectories on two in-distribution subcategories into an SFT corpus and train Qwen-2.5-Coder {1.5B, 3B, 7B} with QLoRA. The trained 7B adapter reaches mean = 0.942 on the four-subcategory test set, matching both frontier models within 0.3 pp on aggregate at approximately 50×lower inference cost. It scores identically to both Sonnet and GPT-5 on all three in-distribution subcategories and ties them within 0.8 pp on out-of-distribution Damping, despite never seeing damping data in training. We further show that combined sparse-and-dense reward filtering beats sparse-only filtering by 1.4–8.7 pp, that smaller models exhibit fundamental seed brittleness which disappears at 3B and above, and that scale enables the OOD generalization that 1.5B and 3B do not provide. The method is portable to any domain admitting closed-form physical laws


PhysElite: How Far Are LLMs from Solving Olympiad-Level Physics Problems?

Ruoran Xu ⋅ Wending Gao ⋅ Liyunfeng Chen ⋅ Aixin Shi ⋅ Haoyu Cheng ⋅ Zixiang Fang ⋅ Yiqiang Zou ⋅ Qiufeng Wang

Understanding how (multimodal) large language models perform on physics problems requires benchmarks that reflect the difficulty and breadth of expert-level physical reasoning. Existing physics benchmarks remain limited in the following two important ways: (1) short of high-difficulty datasets, and (2) lack of comprehensive coverage of visual forms, knowledge points, and step-by-step solution processes. As a result, model performance on current datasets may not be fully representative of their ability to solve complex physics problems. To address these issues, we present \textsc{PhysElite}, a large-scale bilingual multimodal benchmark for Olympiad-level physics reasoning. \textsc{PhysElite} contains 11,586 Olympiad-tier problems. For each problem, we provide corresponding visual diagrams, step-by-step bilingual Chinese--English solution derivations, and the final answer. We benchmark 18 open-source and closed-source MLLMs, and find that even the strongest model reaches only 33.7\% answer accuracy. We additionally conduct step-level process evaluation to diagnose where models fail in the reasoning chain. The benchmark results and dataset will be publicly available.


PhysEval: Quantifying the Gap Between Video Generation and World Physical Laws

Hongchu Zeng ⋅ Sijing Wu ⋅ Yanhan Zhou ⋅ Yunhao Li ⋅ Huiyu Duan ⋅ Yucheng Zhu ⋅ Xiongkuo Min ⋅ Guangtao Zhai

Text-to-video (T2V) models can now synthesize visually compelling motion, but visual plausibility does not guarantee that the generated world follows the physical laws it depicts. A falling object, sliding block, collision, or oscillating spring may look reasonable while still implying an incorrect acceleration, material parameter, or conservation relationship. Most existing benchmarks emphasize perceptual quality, text-video alignment, or qualitative physical plausibility, leaving a gap between visually plausible video generation and quantitatively verifiable physical behavior. To quantify this gap, we present PhysEval, a benchmark dataset and automatic evaluation protocol for scoring the physical accuracy of T2V generations. PhysEval pairs each prompt with auditable metadata, including the target value, unit, evaluator type, expected object count, calibration setting, and known physical parameters. It covers ten physical metrics across kinematics, physical constants, material parameters, and conservation laws. Given generated videos, PhysEval automatically determines which samples are measurable, estimates the relevant physical quantities, and produces normalized scores together with discard diagnostics. Our evaluation and detailed analysis show that current T2V models still struggle to reliably generate videos that are both measurable and physically consistent: even visually plausible samples often fail to express the quantitative cues needed by the underlying physical law.


PhysFlow: Physics-Intrinsic Velocity Regularization for Motion-Intensive Video Generation

Xianglong Guo ⋅ Chang Yu ⋅ Haobo Xu ⋅ Junhao Ma ⋅ Zhen Lei

Text-to-video models built on flow matching generate compelling videos for common prompts but degrade systematically under motion-intensive scenarios, producing artifacts such as object fragmentation, geometric deformation, and temporal flickering. Existing training-free approaches operate at the attention or guidance-scale level, neither of which directly addresses the velocity field that governs latent motion evolution. We observe that these motion artifacts stem from local misalignment between the predicted velocity and the latent's frame-axis temporal structure, a signal that can be diagnosed from quantities the sampler already computes. Based on this, we propose PhysFlow, a training-free velocity regularization method. PhysFlow constructs a per-frame latent flow residual to locate motion-inconsistent regions, steers the velocity to restore frame-axis alignment, and bounds the correction via residual-aware masking and safe clipping. We also introduce MotionStress-100, a five-category benchmark with a VLM-based protocol that isolates motion-intensive failures. On Wan2.1, PhysFlow improves the average MotionStress score over the baseline with no additional inference cost ($1.00\times$ cost), outperforming CFG-Zero* (-10.0%, $0.98\times$ cost) and FlowMo (+0.7%, $2.12\times$ cost) by at least 4.7%, while improving VBench motion smoothness without loss of visual fidelity.


Physics-Informed Video Generation via Mixture-of-Experts Latent Alignment

Cong Wang ⋅ Hanxin Zhu ⋅ Jiayi Luo ⋅ Yonglin Tian ⋅ Xiaoqian Cheng ⋅ Peiyan Tu ⋅ Xin Jin ⋅ Long Chen ⋅ Zhibo Chen

Large-scale video generation models have made remarkable progress in semantic consistency and visual quality, producing videos that are increasingly coherent and visually convincing. Nevertheless, the dynamics induced by pixel-level fitting do not naturally accommodate the regularities that govern real-world motion and interaction, resulting in persistent shortcomings in physical plausibility. To address this limitation, we propose PILA (Physics-Informed Latent Alignment), a framework that injects physics-structured latent guidance into the frozen flow-matching dynamics of pretrained video models. Specifically, PILA first employs anchored field estimation to map frozen-generator latents into an operational physical attribute bank organized by field-proxy slots, using observable motion as a kinematic anchor for constructing less directly observed proxies. To handle the heterogeneity of real-world dynamics, PILA adopts a mixture-of-experts design over physical categories. Label-prior masked expert routing selects category-specific operator experts, whose refinements are regularized by operational residuals abstracted from physical relations. Finally, the refined proxies are fused into the physical attribute bank and decoded into a correction to the flow-matching vector field, injecting physics-aware guidance while preserving the visual prior of the pretrained backbone. With staged adapter training on Wan 2.1-1.3B and direct transfer of the learned adapter to Wan 2.2-14B, PILA achieves state-of-the-art results on VBench-2.0, VideoPhy-2, and PhyGenBench in both visual quality and benchmark-measured physical plausibility.


Platonic Task Arithmetic

Junghwan Park ⋅ Woojin Cho

When distinct pre-trained models are specialized for the same task, they often converge to nearly equivalent functional behaviors, while the underlying parameter-space changes share no common coordinate system. Existing approaches that compose such changes arithmetically are therefore confined to a single model, leaving task knowledge stranded inside the network that acquired it. Drawing on Plato's allegory of the cave, we hypothesize that these model-specific updates are shadows cast by a shared, model-agnostic object that governs how task specialization reshapes a model's behavior, and we refer to this object as the platonic task vector. To make this view operational, we introduce Universal Task Descriptors, matrices whose shape is fixed independently of architecture or embedding dimension and that capture a task's functional effect in a model-transferable form. Universal Task Descriptors admit addition, negation, and analogy in closed form, and a lightweight realization step then transfers any composed descriptor into a chosen target model. Experiments across a range of models and tasks indicate that this pipeline transfers and composes task knowledge across heterogeneous models.


Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction

Sophia Koehler ⋅ Antonia Wüst ⋅ Inga Ibs ⋅ Top Piriyakulkij ⋅ Wolfgang Stammer ⋅ Constantin Rothkopf ⋅ Kevin Ellis ⋅ Kristian Kersting

A central challenge in building intelligent systems is enabling agents to jointly perceive complex inputs, form hypotheses about hidden patterns, and design informative experiments to test them. To study this problem, we propose ZendoWorld, a controlled interactive environment in which agents must infer a logical rule about visual game observations, acquire information by proposing new scenes, and refine their hypotheses based on feedback from the game environment. We evaluate several agents spanning pure VLM reasoning, Bayesian particle filtering, dynamic concept discovery, and neuro‑symbolic methods. Our main findings are: (1) high accuracy in predicting labels for observed examples does not imply recovery of the underlying rule; (2) perception and induction are distinct bottlenecks for different agent classes; and (3) VLM‑based agents propose near‑uninformative experiments, failing to actively reduce hypothesis uncertainty. To compare these results, we collect human data on the task, which reveals a gap in inductive reasoning, particularly for more complex rules. Overall, ZendoWorld takes an important step toward evaluating intelligent agents and identifies concrete avenues for improvement, particularly in domains like scientific discovery.


Pluralistic AI Alignment Requires Inference-Time Multi-Objective Control

Weichen Li ⋅ Mislav Stojanović ⋅ Daniel Neider ⋅ Marius Kloft ⋅ Sophie Fellenz

Pluralistic AI alignment---accommodating diverse human values rather than a single canonical preference---requires agents to reason under multiple, often conflicting objectives, such as helpfulness, honesty, harmlessness, fairness, and context-specific user preferences. Unlike classical learning methods that optimize a fixed scalar objective, pluralistic alignment requires distinguishing between objectives that may be flexibly traded off and constraints that should remain non-negotiable. These two categories map onto two existing lines of research: offline multi-objective reinforcement learning provides tools for representing and navigating trade-offs among multiple objectives, and offline safe reinforcement learning formalizes safety-critical constraints and feasible policy regions. A third line, multi-objective LLM alignment, exposes a further requirement: because retraining large models for each preference configuration is infeasible, controllability must shift from training time to deployment time. Taken together, these observations motivate our central position: inference-time multi-objective control should be a central goal of pluralistic AI alignment. We argue that unifying these perspectives yields a framework in which training-time learning produces reusable objective and safety representations, while inference-time control enables adaptation to diverse and changing preferences without retraining.


Point4D: Long-range 4D Motion Reconstruction

Minsik Jeon ⋅ Jay Karhade ⋅ Deva Ramanan ⋅ Shubham Tulsiani

We introduce Point4D, a feed-forward model for 4D reconstruction of long-range video sequences. Point4D is able to reliably infer dense per-point 3D trajectories across multi-hundred video sequences, unlike existing 4D methods that are limited to short input windows of at most a few dozen frames. A key innovation that enables this is our flexible 3D query-based motion decoder that decouples trajectory prediction from image-plane visibility. The predicted 3D endpoints are then directly re-queried in the next chunk without re-projection or matching. Furthermore, we show that extracting and reusing a visual across an arbitrary frame where the point is visible leads to superior performance than purely relying on the source patch. Overall, Point4D achieves state-of-the-art performance across diverse long-video tracking benchmarks spanning over 200 frames and substantially outperforms previous feed-forward 4D methods.


Point Clustering Encoders

Evangelos Chatzipantazis ⋅ Guillem Brasó ⋅ Cristiano Saltori ⋅ Sérgio Agostinho ⋅ Kostas Daniilidis ⋅ Aljosa Osep ⋅ Laura Leal-Taixé

Learning expressive and transferable representations from unstructured 3D data remains a fundamental challenge. Existing backbones mirror the design of image-based networks by relying on symmetric U-Net architectures with large parametric decoders and feature skip connections. While successful in 2D, these designs are ill-suited for 3D domain where point coordinates, that define the underlying geometric manifold on which feature representations are learned, are sensitive to sensor-specific sampling patterns. In such architectures, skip connections allow low-level coordinate cues to bypass the semantic bottleneck, leading the model to overfit to local spatial patterns rather than learning robust, transferable semantic abstractions. We re-think this design and introduce Point Clustering Encoders (PCE), a minimal, decoder-free network that treats point cloud processing as a hierarchy of end-to-end learned spatial supertokens. At the core of PCE is our Time-Reversible Attention Pooling (TRAP) layer, which formalizes hierarchical downsampling as a time-reversible Markov chain. By coupling upsampling to the reverse chain, we eliminate the need for traditional decoders; the task of representation learning is fully delegated to the encoder. PCE not only surpasses state-of-the-art w.r.t. segmentation accuracy in indoor/outdoor datasets across a variety of tasks, but also reduces the number of trainable parameters, and can process scenes consisting of up to 6.5M points.

Assessing model compatibility is a key challenge in collaborative machine learning. Heterogeneous data distributions can lead to discrepancies among independently trained models that degrade performance when combined, while adversarial manipulation can introduce harmful behaviors. In both cases, reliable evaluation prior to integration is essential. Existing approaches mostly rely on parameter-space statistics, which do not reflect model behavior and are often unreliable in non-IID settings. Inspired by the pointillism art style, where images are formed from small, structured dots of color, we propose Pointillism, a probing-based framework that evaluates model compatibility directly in function space without requiring task data. The method uses structured, randomized probes sampled from an out-of-distribution space to elicit model responses. From these responses, we construct compact model signatures that capture both global prediction characteristics and class-level features. Compatibility is assessed through feature consensus, measuring cross-model agreement on class-level representations via probe transfer. This provides a behavior-centric and fine-grained view of model alignment. We apply the framework to federated learning, where consensus-based selection identifies a self-consistent subset of client models for aggregation. Experiments in federated learning demonstrate improved robustness against both untargeted and backdoor attacks, outperforming existing defense methods. Our source code is included in the supplementary material.

Positional encoding in transformers is commonly implemented through positional embeddings, attention masks, or bias terms, but formal connections between these mechanisms remain limited. We study attention with positional bias through the lens of locality-sensitive hashing (LSH), focusing on Attention with Linear Biases (ALiBi). We show that the ALiBi bias matrix is the expectation of contiguous block-diagonal binary masks induced by a ``positional LSH'' scheme. The empirical mean of masks sampled from this scheme yields spectral norm and max-norm approximation guarantees with bounded block sizes with high probability. This structural theorem implies a uniform approximation theorem for ALiBi-biased attention: with high probability over the sampled masks, the approximate attention output is accurate simultaneously for all query-key-value inputs and can be computed in near-linear time in the context length, reducing long-context ALiBi to a collection of randomized short-context regular (positionally unbiased) attention operations. Conceptually, this connects positional bias, masks, and positional embeddings in a single formal framework and suggests an approach to efficient ALiBi-biased attention. Experiments on large language models validate our theoretical findings.


Positional versus Symbolic Attention Heads: Learning Dynamics, RoPE Geometry, and Length Generalization

Felipe Urrutia ⋅ Juan J Alegría ⋅ Cinthia Sánchez ⋅ Jorge Salas ⋅ Cristian Buc Calderon ⋅ Cristobal Rojas

Transformer-based language models are widespread in today's society. As such, understanding the mechanisms by which they solve structured tasks and predicting how they may behave in novel scenarios is of great importance for safe deployment. We study the learning dynamics of attention heads in a controlled setting by training a decoder-only Transformer (GPT-J) on two structurally equivalent multi-hop reasoning tasks: a number task requiring positional reasoning and a letter task requiring symbolic reasoning. Using a recently introduced metric that classifies attention-head behavior as positional or symbolic for a given prompt, we show that successful learning is associated with the emergence of pure heads, i.e., heads that express themselves as either positional or symbolic. Despite the tasks' structural equivalence, they impose different mechanistic demands: the number task requires both positional and symbolic heads, whereas the letter task requires only symbolic heads. We then identify the computational roles of these heads, characterize the basic functions they implement, and give theoretical constructions showing how single-layer RoPE-based attention can realize these functions through geometrically interpretable query, key, and value operations. This analysis yields a quantitative separation between positional and symbolic mechanisms in their robustness to longer sequences, formalized through a novel notion of discrepancy. We empirically validate the resulting predictions in both controlled and real-world models, showing that symbolic mechanisms extrapolate more reliably to longer sequences while positional mechanisms face sharper limitations.

Transformer predictions depend on positional embeddings, which are known to produce biases such as first- and last-token dominance. However, feature attribution methods implicitly entangle influence from positions with features into a single opaque importance score per feature. To uncover these positional effects, we propose Position-aware eXplanation (PaX), a model-agnostic framework that transforms standard attribution methods to jointly produce feature and position attributions. PaX additionally produces counterfactual positional explanations: actionable scores quantifying how the prediction changes when a feature is relocated. We use PaX to generalize perturbation-, gradient-, and boundary-based attribution methods across vision, language, and clinical time-series benchmarks with no architectural changes. We demonstrate that separating positional effects improves both feature and position attribution faithfulness by 17.5% / 55% and 16.8% / 43%, respectively (insertion / deletion). In a case study on sepsis forecasting, we find that PaX recovers clinically validated bedside signals that standard attributions miss.

Autonomous machine learning research is moving from a speculative possibility to an emerging research practice. Yet, current conference policies are not equipped to handle this change. This position paper argues that machine learning conferences should introduce an Autonomous Research Track where an autonomous AI system controls claim-shaping decisions. However, to ensure that conferences continue to exist for both science and scientists, our proposal anchors this new track in human judgment and participation. First, we propose a hierarchy of AI involvement from incidental AI use to autonomous research, grounded in the CRediT taxonomy of research contributions. In keeping with current conventions, we assert that authorship remain exclusively human while recognizing that the role of author may shift to one of curation of autonomously generated research. We propose a novel review format that separates verification and adjudication: human authors submit a technical review to guarantee accountability, an AI-generated review provides a critical baseline, and a pair of human reviewers evaluate the work's technical correctness and significance. We then propose a discussant-style conference format, where both a human author and a human reviewer present the accepted research. This design assigns visible credit to human participants for thoughtful evaluation and judgment, which is essential for both scientific excellence and maintaining a sense of community.


Position: Neurosymbolic AI is a strong technical foundation for trustworthy, deployable AI by design

Chandler Squires ⋅ Yaqi Xie ⋅ Simon Stepputtis ⋅ Katia Sycara ⋅ Pradeep Ravikumar

As AI becomes more widely adopted, this technology has started to reach a new stage of maturity, with new expectations. In many domains, features like interpretability, controllability, and reliability can no longer be treated as secondary considerations, but must be treated as primary concerns. We argue that the currently dominant incremental refinement strategy should be complemented by a more design-oriented first principles strategy, and that neurosymbolic AI provides one of the most coherent and practically promising technical foundations for this strategy. In particular, we describe certain technical capabilities provided by neurosymbolic AI, including knowledge integration, conceptual alignment, modifiability, and structured generalization, and argue that these capabilities are crucial for building AI that is trustworthy and deployable by design.


Position: Next-Generation Game Engines Should Be Built on Interactive Generative Video

Jiwen Yu ⋅ Yiran Qin ⋅ Haoxuan Che ⋅ Quande Liu ⋅ Xintao Wang ⋅ Pengfei Wan ⋅ Di ZHANG ⋅ Xihui Liu

Modern game development faces significant challenges in creativity and cost due to predetermined content in traditional game engines. Recent breakthroughs in video generation models, capable of synthesizing realistic and interactive virtual environments, present an opportunity to revolutionize game creation. This position paper argues that next-generation game engines should be built on Interactive Generative Video (IGV). We term such engines Generative Game Engines (GGE), with IGV as the core technology enabling unlimited novel content generation. We argue why IGV is uniquely suited for this role, identify what is missing, and examine progress across six functional modules to show how each supports our position, culminating in a five-level maturity roadmap (L0--L4). Drawing on discussions with researchers, game developers, and community participants, we present eight alternative views and ethical considerations to stimulate debate.

The Abstraction and Reasoning Corpus (ARC) has become a prominent benchmark for evaluating skill acquisition in AI models. Yet standard ARC evaluation considers only a single capability: producing the correct output grid for a test input. We argue that this narrow format underestimates the diversity of abilities required for genuine abstract skill acquisition. We introduce PotARCin, a benchmark that extends ARC by assessing understanding of a task’s underlying abstract rule across five dimensions: Definition, Classification, Constrained Generation, Editing, and Inversion. PotARCin uses explicit generator and verifier programs, enabling dynamic generative sampling beyond fixed input-output pairs. Across four state-of-the-art models evaluated on the ARC-AGI-1 training set, we observe a 30–50 percentage-point performance gap between standard ARC evaluation and PotARCin. We further investigate effects of generative sampling, difficulty of corruption types and questions of self-consistency. We also introduce P-ARC, a hand-crafted test set with generator and verifier programs, on which models achieve 0-8% accuracy across all five dimensions, underscoring the difficulty of acquiring versatile abstract skills.


PRECISE: SDE-Consistent Stochastic Sampling for RL Post-Training of Flow-Matching Models

Bo Peng ⋅ Tao Huang ⋅ Weijie Kong ⋅ Junzhe Li ⋅ Yue Wu ⋅ Qi Tian ⋅ Jiangfeng Xiong ⋅ Jian-Wei Zhang ⋅ Liefeng Bo ⋅ Zhao Zhong

Reinforcement learning (RL) has become an effective way to improve prompt alignment and perceptual quality in diffusion and flow-matching generators. A critical step for applying online RL to flow matching is turning the deterministic sampling trajectory into a stochastic policy, typically by replacing the reverse-time Ordinary Differential Equation (ODE) with a Stochastic Differential Equation (SDE). The stochastic sampler, controlling the exploration behavior and denoising dynamics, is thus part of the policy, and its design can significantly affect the reward optimization performance. We break down the sampler design into two interdependent components: choosing the right amount of stochastic exploration, and discretizing the resulting SDE faithfully at the small step counts used in RL. To address the first component, we analyze the inherent tension between exploration and stability in denoising and derive an SDE schedule that balances the two. Turning to the discretization challenge, we use a toy example to show that existing samplers can deviate from the flow-matching process, either by introducing excessive discretization noise or by relying on heuristic rules that do not guarantee convergence to the data distribution. To address these issues, we propose \textsc{Precise}, a new stochastic sampler that balances effective exploration with stability. Crucially, \textsc{Precise} keeps the denoising trajectory SDE-consistent through a novel approximation that freezes the clean-latent posterior mean, resolving the excess noise issue in standard samplers. Extensive experiments demonstrate that this formulation leads to significantly faster and more stable reward optimization via reinforcement learning, achieving state-of-the-art alignment scores (e.g., PickScore, HPSv2.1) while requiring 13.1--53.2\% less wall-clock training time to match the best in-domain performance of prior samplers.


Precision-Aware Hopfield Retrieval: Unifying Population Codes and Memory Retrieval with Information Optimization

Yi-Chun Hung ⋅ Dennis Wu ⋅ Hong-Yu Chen ⋅ Han Liu ⋅ Emma Alexander

Bayesian estimation has been used to explore optimal encoder-decoders for continuous quantities, but no such framework exists for discrete associative memory. We introduce precision-aware Hopfield retrieval, a framework that incorporates the encoder's Fisher information directly into a decoder of memory retrieval. Four results follow. First, optimal encoder-decoder information is shaped only by the prior, across both log and $L^p$ norm loss functions. Second, the optimal information follows a general coupling law across log and $L^p$ loss functions. Third, optimal precision-aware Hopfield retrieval uses a stimulus-dependent inverse temperature. It outperforms the encoder-blind Bayesian decoder on both retrieval loss and required coding resource. Last, the framework yields a closed-form expression for memory recall bias. It reproduces central-tendency and cognitive-load effects observed in cognitive psychology. These effects emerge only when the encoder and decoder are modeled jointly.


Precision-Pyramid: Towards real-time neural decoding for fault-tolerant quantum computing

Zhenhao Zhong ⋅ Ge Yan ⋅ SHANCHUAN LI ⋅ pengyue ma ⋅ Yuxuan Du

Real-time quantum error correction (QEC) is a critical bottleneck for fault-tolerant quantum computing due to strict hardware latency constraints. While neural decoders achieve state-of-the-art logical accuracy, their reliance on high-precision, compute-intensive inference precludes real-time deployment. Conversely, uniformly quantizing these models to extreme low-bit regimes for FPGAs triggers a catastrophic collapse in decoding accuracy. To overcome this precision-accuracy bottleneck, we propose Precision Pyramid ($\mathsf{PP}$), a hardware-algorithm co-designed neural decoder applicable to both surface and BB codes. $\mathsf{PP}$ features a monotonically escalating activation precision hierarchy built upon a globally weight-binarized (W1) foundation. By scaling intermediate activations down to INT2 for the massive perception stages and reserving higher precision strictly for the lightweight final logical decision, this hierarchy effectively shifts over $99\%$ of the computational workload to abundant FPGA look-up tables. Comprehensive evaluations on both simulated and hardware-calibrated noise data demonstrate that $\mathsf{PP}$ consistently suppresses logical errors below highly optimized classical baselines (e.g., MWPM and Relay-BP) while seamlessly satisfying stringent sub-microsecond latency budgets.


PreCoMem: Predictive Cognitive Memory for Self-Evolving Long-Term Dialogue Agents

Jian Zhong ⋅ Zeyu Liu ⋅ Pingchuan Cao ⋅ Rongduo Han ⋅ Shunye Tang ⋅ Guohuan Xie ⋅ Mingyuan Qin ⋅ Ran Ji ⋅ Xiao Liang ⋅ Haining Zhang ⋅ Wei Wang

Long-horizon dialogue agents must keep user memories stable enough for personalization while adapting when the user genuinely changes. Existing memory systems often append every observation or overwrite old entries by heuristic rules, which makes transient noise hard to distinguish from real belief drift. We introduce PreCoMem, a memory consolidation framework inspired by predictive coding that represents memory as confidence-weighted beliefs about the user's state. Its core signal, Effective Surprisal, combines semantic deviation and contradiction evidence, then downweights this signal when retrieval is diffuse and unreliable. According to effective surprisal, a two-threshold gate maps each turn to one of three deterministic updates: MAINTAIN reinforces supported beliefs, PROFILE stores ambiguous cues as low-confidence hypotheses, and CORRECT softly decays reliably contradicted facts. The PROFILE stage acts as an explicit waiting room, preventing premature commitment to noisy evidence. Experiments on LoCoMo, LongMemEval, and PersonaMem-v2 show consistent state-of-the-art accuracy; our new DiMoBench confirms that PreCoMem adapts to non-stationary belief drift without over-reacting to noise. The code is available at: \url{https://anonymous.4open.science/r/PreCoMem-9000}.

Optimizing a proxy objective can improve human-perceived quality, but the resulting gain can vary with optimization strength. Average proxy--human correlation alone does not reveal how human-perceived quality changes as optimization strength increases. We propose a pre-optimization diagnostic that uses repeated human ratings to predict the human-gain curve induced by a fixed proxy. The method models proxy optimization as a KL-regularized exponential tilt. At KL radius $d$, the local gain is approximated by $A\sqrt{d}+Bd$, where $A$ measures first-order proxy--human alignment and $B$ captures the leading second-order correction along the proxy-optimization path. We estimate $A$ and $B$ from repeated human ratings on calibration items and test the resulting gain-curve prediction on held-out evaluation items. Across SummEval, analysis-eligible WMT MQM clusters, an OpenMEVA-MAGS ROC/WP repeated-rating slice, and a GPT-4o judge proxy experiment, the proposed approach consistently reduces held-out gain-prediction error relative to $A$-only, zero-gain, and reliability-only baselines. Transfer experiments across datasets and evaluation slices show that source coefficients are useful priors, but target calibration further improves prediction. The method provides a principled local diagnostic for estimating the human-gain curve of a fixed proxy as optimization strength varies.


Predicting the Needle in a Petabyte Scale Haystack: Open-Vocabulary Event Anticipation in Satellite Imagery

Lekha Revankar ⋅ Mikhail Klassen ⋅ creon levit ⋅ Ash Hoover ⋅ Kavita Bala ⋅ Bharath Hariharan

Satellite constellations now image the Earth daily, capturing the lifecycle of events from deforestation to urban construction. Anticipating these events before completion enables timely intervention, yet existing systems cannot jointly identify *where* a change occurs, *what* it will become, and *when* it will finish across an open vocabulary. Building such a model requires diverse event data, but events are like needles in a haystack, making manual annotation at global scale infeasible. We address this problem by representing change as the vector difference between vision-language model embeddings at distinct time points. This arithmetic approach allows us to search for semantic transformations directly in the latent space, enabling (1) an automated data engine for global event discovery without manual annotation and (2) the largest global satellite event dataset to our knowledge, comprising 21,000+ locations across 64+ event types with ~2M images from PlanetScope and Sentinel-2 spanning six continents, to train (3) a novel global-scale event anticipation model. Our model detects change with a 93.5% max F1 score, outperforms baselines 3.6$\times$ on predicted event retrieval, and forecasts completion dates with a 3-day median error, even nine months before completion.


Prediction-Augmented Trees for Reliable Statistical Inference

Vikram Kher ⋅ Argyris Oikonomou ⋅ Manolis Zampetakis

Machine learning (ML) is increasingly central to scientific discovery, but translating ML predictions into rigorous scientific claims still requires statistical inference on gold-standard labeled data. A setting that has emerged across the sciences provides the analyst with a small labeled sample, a much larger unlabeled sample, and a pre-trained predictor, with the goal of constructing a confidence interval for a quantity of interest, as was formalized in Angelopoulos et al. (2023). In this ML-driven scientific discovery literature, labeled data is typically the binding constraint on precision while computation is comparatively unconstrained, motivating the design question of how additional computation can extract more inferential value from a fixed labeled set, a fixed unlabeled set, and a fixed predictor. To address this question, we propose the *Prediction-Augmented Residual Tree* (PART) estimator. PART takes the same inputs as the original PPI estimator and replaces their single global rectifier with locally computed rectifiers obtained from an adaptive partition of the feature space. We show that PART outperforms existing methods by producing tighter confidence intervals across real-world datasets from ecology, astronomy, and census reports, among other domains. We then describe and analyze PAQ, an estimator that arises when considering the limit of PART when the depth of its tree grows to infinity. Under appropriate assumptions in the input data, we show that the variance of PAQ shrinks at a rate of $O(N^{-1} + n^{-4})$, showing that there are settings where the rate of $O(N^{-1}+n^{-1})$ of existing methods can be provably broken.


Predictively-Oriented Kalman Filtering

Zheyang Shen ⋅ Gerardo Duran-Martin ⋅ Chris Oates

This paper presents a post-Bayesian approach to online filtering in nonlinear state-space models, capable of avoiding over-confident inferences in settings where either the dynamical model, the measurement model, or both, could be misspecified. This is addressed using predictively oriented (PrO) posteriors, an emerging paradigm in which learning (i.e., posterior concentration) occurs if and only if the overall model is well-specified, without strict adherence to the Bayes' theorem. As the characterisation of PrO posteriors is challenging, our main technical contribution is a fast approximate linear-Gaussian update procedure, analogous to an (iterated) extended Kalman filter (EKF). The methodology, which we call PrO-EKF, has no tunable hyper-parameters and has a computational cost comparable to that of existing filtering methods. Performance is empirically assessed on a range of linear and non-linear applications, in which the state-space model is systematically misspecified.

Token-level maximum likelihood and mean negative log-likelihood dominate the training and evaluation of large language models (LLMs), yet long-horizon generation often suffers from late-stage degradation despite improved mean loss. We identify two coupled causes of this mismatch: mean token loss is not scale-free with respect to sequence length, so small per-token divergences can amplify into large sequence-level distribution shifts, and long-horizon failures are spike-dominated, where rare but abrupt prefix deviations trigger errors that are largely invisible to mean objectives. To address these issues, we propose \textbf{Prefix Likelihood-Ratio Control} (PLRC), which centers training on the prefix log-likelihood ratio process and its increments to directly capture sequence-level deviation and spikes. As the data likelihood is unknown, PLRC learns a prefix ratio critic via noise-contrastive prefix discrimination and introduces tail-sensitive regularizers on the terminal ratio and its increments, yielding guarantees that bound long-horizon event distortion and the probability of the first spike. PLRC preserves standard teacher-forcing training without architectural changes while explicitly targeting failure modes that mean token loss cannot capture.


Pretext Reasoning: Scaling the Building Blocks of Interleaved Multimodal Reasoning

Jiawei Gu ⋅ Linjie Li ⋅ Yiming Liu ⋅ Yunzhuo Hao ⋅ Zhichao Peng ⋅ Huichen Wang ⋅ Guanzheng Chen ⋅ Luxin Xu ⋅ Xinyu Zhang ⋅ Li Luo ⋅ Disen Lan ⋅ Zican Hu ⋅ Mingyang Song ⋅ David Valente ⋅ Alex Jinpeng Wang ⋅ Yafu Li ⋅ Ganqu Cui ⋅ Zhengyuan Yang ⋅ Michael Shieh ⋅ Yejin Choi ⋅ Ranjay Krishna ⋅ Yu Cheng

Human reasoning is not purely linguistic: for visual problems, people often think by changing the visual representation itself. Interleaved multimodal reasoning seeks to bring this ability to unified models by allowing them to construct intermediate visual states and use them within multimodal reasoning traces. Yet scaling this ability is difficult, since task-driven instruction tuning must teach visual-state construction, faithfulness, and downstream reasoning all at once. We introduce PRIMER (\textbf{PR}etext-based \textbf{I}nterleaved \textbf{M}ultimodal r\textbf{E}asoning \textbf{R}ecipe), a two-stage training recipe that uses pretext reasoning as a primer for interleaved multimodal reasoning. Stage 1 builds \pretext, an instruction-free corpus that converts classical self-supervised vision pretext tasks into Thought--Image--Thought traces, teaching the model the reusable mechanics of producing and reading visual states. Stage~2 builds \instruct, a task-driven instruction-tuning corpus that teaches task-relevant visual-state construction across perception, spatial understanding, and mental world modeling. On a 7B unified multimodal model, Primer improves over the matched-budget BAGEL reference by +10.40, +5.93, and +15.70 across the three cognitive levels, with shorter traces, and rivals models several times its size on spatial and world-modeling averages. These results establish classical self-supervised pretext tasks, repurposed as Thought-Image-Thought traces, as a scalable, parameter-efficient primer for interleaved multimodal reasoning, the recipe at the core of Primer.


Pretraining Data Statistics Shape the Phases of Learning Entity Comparison in Language Models

Yik Siu Chan ⋅ Jing Huang ⋅ Yanai Elazar ⋅ Atticus Geiger

How does data shape language model (LM) behavior throughout pretraining? We investigate this question through a case study on entity comparison, e.g., Between France and Brazil, which country is larger?. We begin with controlled experiments in which we train small LMs (124M parameters) on mixtures of natural text from pretraining corpora and synthetic data from entity comparison tasks. We identify three distinct phases of learning: (1) an early phase where the LM selects entities by frequency, (2) a middle phase where the LM selects entities by position in a prompt (first vs. last), and (3) a late phase where the LM selects the entity that is the correct answer to the question. We show that the emergence of these three phases is controlled by statistical properties of the training data. With small amounts of task-specific synthetic data, we observe only the first two phases and the model fails to learn the task; with large amounts, the model jumps directly from the frequency-based heuristic to solving the task correctly. Moreover, if we modify the frequency of entities in data from a naturally occurring Zipfian distribution (a small number of entities are very common and the vast majority are rare) to a uniform distribution, the first phase disappears and the model learns the task more quickly. Finally, we find the same three phases of learning in the pretraining of open-sourced OLMo models. Together, our findings demonstrate that properties of pretraining data are causal drivers of heuristic learning and show that small-scale synthetic experiments can predict training dynamics at larger scales.


Primal-Dual Representation Learning for Low-Rank Constrained MDPs

Chenhao Zhou ⋅ zhang chao ⋅ Wensong Bai ⋅ Hanbin Zhao ⋅ Hui Qian

Representation learning has substantially improved the efficiency of unconstrained continuous-state Markov Decision Processes (MDPs), but its extension to Constrained MDPs (CMDPs) is difficult because the transition representation must be learned while constraint violation is controlled. In this paper, we study representation learning for low-rank CMDPs with continuous state spaces, and propose REP-PD-TEG for the soft cumulative-violation setting. In the feature learning period, the algorithm interacts with the environment using a decaying-$\epsilon$-greedy mixture execution policy, and uses the collected samples to learn transition features by maximum likelihood. Based on the learned features, it constructs uncertainty bonuses and plans with an optimistic Lagrangian objective to obtain the return policy. We prove regret and soft-violation guarantees for the return policies, and quantify the influence of the decaying-$\epsilon$-greedy exploration. In particular, REP-PD-TEG achieves the optimal \(\widetilde O(\sqrt K)\) return policy bound, and gives a \(\widetilde O(K^{3/4})\) overall sample bound when exploratory actions are charged. For the stricter hard cumulative violation, where feasibility must be verified episode by episode, we propose REP-PD-H-TEG. It plans under a bonus-corrected lower-confidence utility certificate with the learned representation, and performs a stabilized dual search to obtain the return policy, yielding regret and hard-violation bounds. Together, these results give the first theoretically guaranteed representation-learning method for low-rank continuous-state CMDPs with both soft cumulative guarantees and certified hard-violation control.

Existing low-light image enhancement methods are increasingly built on deep representation learning. However, most of them still encode features as deterministic points in the embedding space, making it difficult to capture the statistical uncertainty. Under complex low-light scenes, latent neighborhood relations tend to reflect degradation patterns rather than semantic content or local structures, which leads to statistical shifts and reduced discriminability of representations. Therefore, we propose Prior-Anchored Local Statistical Representation Rectification (PaLSR), which models low-light enhancement as a representation learning process driven by local statistical rectification. PaLSR first learns a normal-light statistical prior with K Gaussian anchors, which provides a shared coordinate system for local representation. Instead of enhancing degraded features in the Euclidean embedding space, PaLSR represents each feature as mean-variance offsets relative to these anchors, describing its content displacement and uncertainty. With these offsets rectified, local representations are reorganized under complex low-light degradations. Extensive experiments on multiple low-light benchmarks and different network architectures show that PaLSR achieves consistent improvements in restoration quality. These results validate the effectiveness of prior-anchored local statistical representation rectification under complex degradation conditions.


PRISMIC: Reconstructing User Preference via Intent Decomposition and Consolidation

Donghee Han ⋅ Jiwon Jeong ⋅ Hwanjun Song ⋅ Mun Yi

Large language models (LLMs) have recently improved sequential recommendation, yet the task remains a retrieval problem without explicit user queries: the system must infer the next item from user histories where diverse intents and preferences are intertwined. Existing methods typically compress this heterogeneous evidence into a single embedding or query, while LLM-based recommenders can produce overly general queries that are semantically plausible but weakly aligned with retrieval. We propose PRISMIC (Preference Reconstruction via Intent Synthesis and Multi-signal Inference Consolidation), which reformulates sequential recommendation as multi-signal intent consolidation. PRISMIC first generates multiple candidate queries from different historical signals, then trains an LLM-based consolidator with GRPO using an NDCG-based retrieval reward to merge relevant signals into a single query. The resulting consolidated query is not directly used at inference time; instead, it serves as a semantic supervision target for training a user encoder, enabling LLM-free inference. Across four real-world datasets, PRISMIC consistently outperforms strong baselines. Ablations further show that the gains come not from GRPO alone, but from consolidating multiple intent-signals and distilling the consolidated intent into an encoder.


PRISM: Phenotype-Resolved Inference in Single-Cell Mixed Models via Latent Disease States and Contextualized Differential Expression

Andrea Rubbi ⋅ Lama Salem ⋅ Caleb Ellington ⋅ Pietro Lió ⋅ Mohammad Lotfollahi ⋅ Manolis Kellis ⋅ Ben Lengerich

Standard single-cell differential expression (DE) analysis identifies genes that change across conditions, but it usually overlooks two key sources of heterogeneity that single-cell data are uniquely positioned to reveal: which cells within a donor are truly disease-affected, and how disease effects depend on cell state or subject-level context. Assigning every cell the diagnosis of its donor can contaminate DE signal when disease penetrance is partial, while modeling each gene with a single disease effect cannot capture heterogeneous responses across cell populations. We present PRISM (Phenotype-Resolved Inference in Single-cell Mixed models), a negative-binomial mixed-effects framework that augments standard DE analysis with three outputs: a context-DE vector $\theta_g$ that tests whether disease effects vary along a biological axis $z_{ij}$ (such as cell state, sex, or age); a cell-level disease posterior $q_{ij}$ that provides unsupervised disease annotation; and a subject-level disease burden $\rho_i$. Naive likelihood-only inference of $q_{ij}$ is not identified onto the disease axis when nuisance variation (batch, cell-cycle, dominant subtypes) competes with disease; we resolve this with a closed-form 1-D Wasserstein-2 projection of the disease-arm posterior onto a bimodal reference marginal that enforces the \{healthy, affected\} structure implied by the disease label.


PRISM: Priority-Guided Scanning in the Wavelet Domain for UAV Maritime Small Object Detection

Pengqi Gao ⋅ Fan Shi ⋅ Mianzhao Wang ⋅ Jiangpeng Zheng ⋅ Xu Cheng ⋅ Shengyong Chen

Sparse object detection remains difficult when weak target evidence occupies only a few pixels and shares local statistics with surrounding clutter. This failure is not merely due to limited model capacity, but to a representation bottleneck: spatial domain features mix target edges, clutter textures, and scene context before they can be reliably separated. We propose PRISM, a wavelet domain detector that treats this problem as frequency separated representation learning followed by priority ordered feature interaction. PRISM decomposes each scene into directional high frequency evidence and global low frequency context, models the two streams with dedicated encoders, and fuses them through a learned content adaptive causal scan that lets likely target regions guide feature interaction before clutter dominated regions participate. Multi scale features are then reconstructed by inverse wavelet synthesis with adaptive gates and passed to a standard detection head. On low altitude UAV maritime benchmarks, including SeaDronesSee and AFO, PRISM achieves state-of-the-art performance with the largest gains on the smallest targets. Without architectural modification, it also transfers to aerial urban imagery on VisDrone, suggesting that priority ordered fusion of directional frequency evidence is a generalizable design principle for sparse visual signal detection.


PRISM: Prior Relational Information for Self-supervised Modeling to Enhance Solubility OOD Generalization

Xinyi Chen ⋅ Jiahuan Pang ⋅ Titus Chima ⋅ Paul Weng ⋅ Wendong Wang

Solubility depends on intermolecular interactions between solutes and solvents and is central to drug discovery, chemical and materials process engineering. Its prediction requires out-of-distribution (OOD) generalization to unseen solutes, solvents, and solvent mixtures, where supervised methods suffer from data scarcity, and existing self-supervised pretraining methods do not explicitly learn chemically meaningful relations. Learning chemically meaningful relational information without labels can alleviate data scarcity and improve OOD generalizations, but remains challenging. We introduce Prior Relational Information for Self-supervised Modeling (PRISM) and apply it to solubility OOD generalizations. PRISM is a label-free pretraining framework that converts expert-defined chemical priors into relational teachers. For solubility prediction, it trains four encoders to preserve the population-level relational information from priors capturing hydrophobic, electrostatic, polarizability, and substructural effects, using Kullback-Leibler (KL) divergence to match pairwise distance-based distributions during pretraining. For scaffold, solute, and solvent OOD generalizations, PRISM matches or surpasses supervised and self-supervised baselines. For mixed-solvent OOD generalizations, freezing the four PRISM encoders and training a light fusion-and-composition head outperforms the strongest supervised baselines and greatly reduces prediction variations. Ablations and control studies confirm that gains stem from chemically meaningful relational structures baked into the encoders by PRISM. We expect this strategy of turning relational priors into self-supervision signals to extend beyond solubility prediction to domains where chemically or biologically meaningful relational priors are available.


PRISM: Rethinking Atmospheric Scattering Reconstruction as a Unified Understanding and Restoration Model for Real-world Dehazing

Chengyu Fang ⋅ Chunming He ⋅ Yuelin Zhang ⋅ Chubin Chen ⋅ zhu ⋅ Hongqiu Wang ⋅ Longxiang Tang ⋅ Xiu Li ⋅ Sina Farsiu

Real-world image dehazing (RID) aims to remove haze-induced degradation from real scenes. This task remains challenging due to non-uniform haze distribution, spatially varying color shifts, and the scarcity of paired real hazy-clean data. In PRISM, we propose Proximal Scattering Atmosphere Reconstruction (PSAR), a physically structured framework that jointly reconstructs the clear scene and scattering variables under the atmospheric scattering model, making the restoration process more interpretable in complex real-world conditions. To bridge the synthetic-to-real gap, we design an online non-uniform haze synthesis pipeline and a Selective Self-Distillation Adaptation (SSDA) scheme for unpaired real-world scenarios, which enables the model to selectively learn from high-quality perceptual targets while leveraging its intrinsic scattering understanding to audit residual haze and guide self-refinement. Experiments on real-world benchmarks demonstrate that PRISM achieves competitive performance on RID tasks.


Privacy Risk Scales with Effective Dimension in Federated Learning

MD MAMUNUR RASHID ⋅ Md Palash Uddin ⋅ Yong Xiang ⋅ Keshav Sood ⋅ Longxiang Gao

Privacy mechanisms in federated learning (FL) are often calibrated without explicit regard to model scale, implicitly assuming that privacy risk remains stable as federated models grow. We challenge this assumption by introducing federated leakage, a signal-level mutual-information measure of client-specific information surviving in privatized update trajectories, and by identifying effective model dimension, rather than raw parameter count, as the operative scaling variable. Under a Gaussian signal-channel model with spectral growth of client-informative directions, we establish a regime-conditional scaling law: leakage grows as $\Theta(d/\log d)$ with effective dimension $d$. This result yields a leakage-ratio notion of privacy debt, where fixed-noise deployment can expose increasing client-informative signal even when the accountant-reported DP configuration is unchanged. We therefore propose a scale-aware Gaussian calibration rule that preserves the target leakage regime up to constant factors, and show across vision and language settings that it substantially reduces cross-scale privacy drift relative to fixed-noise baselines.


ProactBench: Beyond What The User Asked For

Sepehr Harfi Moridani ⋅ Ahmad Salimi ⋅ Dongming Shen ⋅ Alexander Smola

Most LLM benchmarks score how well a model responds to explicit requests. They leave unmeasured a different conversational ability: noticing and acting on needs the user has implied but not said. We call this conversational proactivity. ProactBench decomposes it into three phase-tied types: Emergent, inference from a single disclosed anchor; Critical, synthesis across multiple anchors; and Recovery, grounded forward-looking value after task completion. We operationalise the benchmark with three agents: a Planner, a User Agent, and an Assistant Model. Their information asymmetries defend against style-confounded scoring, rubric leakage, external-context contamination, and information dumps. The released corpus contains 198 curated dialogues with 624 trigger points across 24 communication styles drawn from a psychometric inventory and audited by an independent LLM judge. Across 16 frontier and open-weight models, Recovery is both difficult and weakly predicted by six standard benchmarks, making it a useful new evaluation signal.


Probabilistic Recursive Reasoning

Junyeob Baek ⋅ Mingyu Jo ⋅ Minsu Kim ⋅ Mengye Ren ⋅ Yoshua Bengio ⋅ Sungjin Ahn

How should future neural reasoning systems implement extended computation? Recursive Reasoning Models (RRMs) offer a promising alternative to autoregressive sequence extension by performing iterative latent-state refinement with shared transition functions. Yet existing RRMs are largely deterministic, following a single latent trajectory and converging to a single prediction. We introduce Generative Recursive reAsoning Models (GRAM), a framework that turns recursive latent reasoning into probabilistic multi-trajectory computation. GRAM models reasoning as a stochastic latent trajectory, enabling multiple hypotheses, alternative solution strategies, and inference-time scaling through both recursive depth and parallel trajectory sampling. This yields a latent-variable generative model supporting conditional reasoning via $p_\theta(y \mid x)$ and, with fixed or absent inputs, unconditional generation via $p_\theta(x)$. Trained with amortized variational inference, GRAM improves over deterministic recurrent and recursive baselines on structured reasoning and multi-solution constraint satisfaction tasks, while demonstrating an unconditional generation capability.


Probing the Trajectories of Reasoning Traces in Large Language Models

Marthe Ballon ⋅ Brecht F Verbeken ⋅ Vincent Ginis ⋅ Andres Algaba

Large language models (LLMs) increasingly solve difficult problems by producing reasoning traces before emitting a final response. However, it remains unclear how accuracy and decision commitment evolve along a reasoning trajectory, and whether intermediate trace segments provide answer-relevant information beyond generic length or stylistic effects. Here, we propose a trajectory-probing protocol for evaluating how a reasoning model's predicted answer evolves along a partial reasoning trace. The protocol 1) generates a model's full reasoning trace, 2) truncates it at fixed token-percentiles, and 3) injects each partial trace back into the model, measuring the model's induced answer distribution. We apply the protocol to five open-source reasoning models (Qwen3-4B/-8B/-14B and gpt-oss-20b/-120b) across three benchmarks (GPQA Diamond, MMLU-Pro, and Omni-MATH-2). The protocol reveals that accuracy and decision commitment consistently increase as the percentage of provided reasoning tokens grows, and length, form, and token-identity controls confirm that gains stem from instance-specific semantic content rather than context length or generic reasoning style effects. Furthermore, weak-to-strong cross-model probing experiments reveal when stronger models recover from incorrect partial traces and when they instead anchor to them. Together, these results provide a reproducible, model-agnostic measurement of reasoning-trace dynamics that generalizes across answer formats and can reveal model-specific failure modes invisible to aggregate accuracy.

Many online learning systems now have access to several gradient predictors at once, such as simulators, replay buffers, neural surrogates, or time-series forecasters, and the quality of these predictors varies from round to round. We study online convex optimization under one-gradient feedback against an arbitrary moving comparator, with a user-supplied finite class $\Pi$ of gradient-field predictors, and ask whether a single algorithm can compete with the best predictor in $\Pi$ at a logarithmic cost, fall back to a small-loss bound when none of the predictors is useful, and still retain a path-length-adaptive worst-case guarantee when the predictors are adversarial. We give a positive answer through PH-Sword (Predictable-Hybrid Sword.optimism), which combines an exp-concave Hedge over $\Pi$, optimistic projected-gradient base learners on a geometric learning-rate grid, and a correction-term optimistic-entropy meta learner with prediction-error doubling. For any comparator sequence $\boldsymbol u$, PH-Sword attains dynamic regret $\widetilde O\bigl(R\_0\sqrt{(1+P\_T/D)(1+P\_T/D+\min\{\mathcal E\_\*/G^2,\,2LF\_T(\boldsymbol u)/G^2\}+\log|\Pi|)}\bigr)$ up to logarithmic factors, where $R\_0=DG+LD^2$, $P\_T$ is the path length, $F\_T(\boldsymbol u)$ the cumulative comparator loss, and $\mathcal E\_\*$ the best-in-class field error. The bound merges the field-controlled, small-loss, and worst-case regimes into one deterministic guarantee that improves over prior one-gradient dynamic-regret algorithms. On six controlled benchmarks the five predicted scaling laws hold simultaneously: PH-Sword reaches 17% of OGD's final dynamic regret and 54% of Ader's on a slow-drift benchmark, and we observe that a single noisy predictor without model selection can actually do worse than using no predictor at all.


probly: Uncertainty-Aware Machine Learning

Paul Hofman ⋅ Clemens Damke ⋅ Timo Löhr ⋅ Santo Thies ⋅ Alireza Javanmardi ⋅ Valentin Margraf ⋅ Jakub Paplhám ⋅ Yusuf Sale ⋅ Eyke Hüllermeier ⋅ Maximilian Muschalik

As machine learning systems are increasingly deployed in real-world applications, the question of how to represent and quantify uncertainty has moved from a methodological side issue to a central concern. In practice, however, making a model uncertainty-aware is still surprisingly difficult: relevant tools tend to be scattered across libraries, each tied to a particular framework and a particular approach to modeling uncertainty, and the choice of representation, quantification, and evaluation is typically left to the user without much guidance. In this paper, we present probly, a Python package that addresses theses issues in a single, modular framework. probly offers (i) lightweight transformations that turn existing models into uncertainty-aware ones, currently supporting PyTorch, scikit-learn, and Flax, (ii) several representations, including second-order distributions, credal sets and conformal prediction, (iii) the corresponding quantification measures, and (iv) a number of evaluation protocols, which can be combined more or less arbitrarily. Using probly, we conduct a benchmark study on three standard tasks --- selective prediction, out-of-distribution detection, and active learning. In addition, and unlike existing benchmarks, we evaluate suitable methods on first-order data, i.e., datasets for which the target itself is a distribution over outcomes. We further illustrate the flexibility of the package on a number of less standard case studies, including large language models, graph neural networks, and data streams. The package is publicly available at \url{https://anonymous.4open.science/r/probly}.


ProCARE: Real-World Study Automation via Profile-Grounded Evidence Contracts

Zhaoyu Fan ⋅ Bowen Han ⋅ Yanwei Ren ⋅ Jincheng Ou ⋅ Siyi Cao ⋅ Zehua He ⋅ Kaiyu Huang ⋅ Pengcheng He ⋅ Shanshan Ni ⋅ Chengchen Gong ⋅ Qingjiang Shi

Real-world studies (RWS) analyze routinely collected healthcare data, such as electronic health records and insurance claims, to estimate treatment effects, monitor safety, and identify risk factors outside randomized trials. Automating RWS with large language model agents is appealing but brittle: valid analyses depend on dataset-specific clinical semantics, including index dates, follow-up windows, patient/event-level units, and outcome boundaries that are often implicit in observational data. This leads to schema mismatches, temporal errors, inconsistent analytical granularity, and unsupported clinical claims. We present ProCARE, a profile-grounded framework that reformulates RWS automation as constrained compilation. ProCARE converts heterogeneous data profiles into Evidence Contracts specifying available variables, temporal attributes, distributions, and evidentiary limits, and compiles study objectives into RWS-IR, a strongly typed intermediate representation for variables, time relations, analysis units, and statistical tasks. Deterministic validation and localized repair then check and revise generated workflows across schema mapping, cohort construction, variable derivation, analysis, and interpretation. On 36 EHR-based benchmark tasks, ProCARE outperforms strong LLM-agent baselines in code executability, semantic consistency, and conclusion traceability, with automated violation diagnoses showing moderate-to-substantial agreement with blinded expert review. Code, benchmark prompts, evaluation scripts, and the fully de-identified HCC benchmark tables are publicly available at https://anonymous.4open.science/r/ProCARE-7E52.


ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder

Xiaoxing Hu ⋅ Kaicheng Yang ⋅ Ziqi Ye ⋅ Ziyang Gong ⋅ Qi Ming ⋅ Zonghao Guo ⋅ Yu Tian ⋅ Xiang An ⋅ Ziyong Feng ⋅ Xue Yang

Contrastive Language-Image Pre-training (CLIP) is fundamentally limited by its 77-token restriction, lack of multilingual capabilities, and coarse-grained semantic representations. While replacing CLIP’s native text encoder with a Large Language Model (LLM)-based embedder offers a promising solution, direct, from-scratch alignment can disrupt the semantic structures learned during pre-training. This often leads to degradation of cross-modal knowledge and particularly impairs recognition capabilities in zero-shot scenarios. In this paper, we propose ProCLIP, a curriculum-learning-inspired progressive alignment framework designed to systematically bridge the LLM embedder and CLIP's visual space, unlocking CLIP's potential for long-text, multilingual, and fine-grained understanding. Our framework operates in two stages: (1) Representation Inheritance, which distills CLIP's original text-space knowledge into an LLM adapter to establish a robust initial vision-language prior, and (2) Contrastive Tuning, which refines the cross-modal alignment with self-distillation regularization on the image encoder to further prevent catastrophic forgetting. To maintain semantic and geometric consistency, we introduce instance-level semantic and global structural alignment constraints throughout both stages. Extensive experiments demonstrate that ProCLIP improves zero-shot classification accuracy by 6.8%--13.5% over existing LLM-augmented baselines and achieves state-of-the-art performance in diverse long-text, multilingual, and fine-grained cross-modal retrieval tasks under comparable settings.


Progressive Residual Warmup for Language Model Pretraining

Tianhao Chen ⋅ Xin Xu ⋅ Lu Yin ⋅ Hao CHEN ⋅ Yang Wang ⋅ Shizhe Diao ⋅ Can Yang

Transformer architectures serve as the backbone for most modern Large Language Models, therefore their pretraining stability and convergence speed are of central concern. Motivated by the logical dependency of sequentially stacked layers, we propose Progressive Residual Warmup (ProRes) for language model pretraining. ProRes implements an "early layer learns first" philosophy by multiplying each layer's residual with a scalar that gradually warms up from 0 to 1, with deeper layers taking longer warmup steps. In this way, deeper layers wait for early layers to settle into a more stable regime before contributing to learning. We demonstrate the effectiveness of ProRes through pretraining experiments across various model scales, as well as normalization and initialization schemes. Comprehensive analysis shows that ProRes not only stabilizes pretraining but also introduces a unique optimization trajectory, leading to faster convergence, stronger generalization and better downstream performance. Our code is available in supplementary materials.


Prompt Optimization Makes Misalignment Legible

Caleb Biddulph ⋅ Micah Carroll

Large language models (LLMs) trained with reinforcement learning (RL) often exhibit reward hacking, exploiting unintended loopholes in reward functions in ways that can be difficult to detect and eliminate. We propose using prompt optimization—methods which increase an LLM's reward by updating its instructions rather than its weights—to make learned strategies easier to monitor and edit. Applying the GEPA prompt optimizer to environments with exploitable reward functions, we find that optimized system prompts describe reward hacking strategies in highly interpretable language. Furthermore, by simply removing descriptions of unwanted behavior from the optimized system prompt at test time, we can improve the model's alignment while preserving legitimate performance gains. We show that prompt optimization can be guided with an RL-trained teacher LLM, combining the performance advantages of RL with the interpretability of prompting. Finally, we explore an approach to shorten optimized prompts, removing distracting and unhelpful instructions which would otherwise hinder interpretability. We hope that these insights about mitigating misalignment with prompt optimization will aid the discovery of unintended exploits in RL environments and the creation of predictable and monitorable AI systems.

Large language models are increasingly used as proxies for human subjects in social science research, yet external validity requires matching the response distributions of target human populations. We study population-level survey alignment: reconstructing aggregate survey responses from limited public survey data, without individual-level demographic profiles or model finetuning. We formalize this problem as preference reconstruction: rather than matching proxy agents to demographic profiles, we construct a functional basis of proxy agents and recovering population preferences by weighted aggregation. We instantiate this idea via Prompts to Proxies ($\texttt{P2P}$), a two-stage inference-time system. Stage 1 uses structured attribute-based prompting with entropy-guided adaptive sampling to construct a diverse proxy pool spanning the latent preference space. Stage 2 employs L1-regularized regression to select a compact weighted ensemble matching observed target-population responses. Across 14 American Trends Panel waves, $\texttt{P2P}$ achieves an average test MSE of 0.014 at approximately 0.8 USD per survey, improving over prompting and demographic-conditioning baselines. On the World Values Survey, cross-locale transfer experiments show that basis expressiveness can outperform locale matching on average, while locale-specific generation helps on culturally divergent questions. A stress test against an SFT-aligned survey model shows competitive performance using less than 3\% of the training data. These results position preference reconstruction as a lightweight, externally verifiable alternative for survey-based population alignment.


ProPolar: Progressive Polar Decomposition for Implicit Neural Representations

Pureum Kim ⋅ Younggeon Ryu ⋅ Dongyoon Lee ⋅ Hae Beom Lee ⋅ Kyong Hwan Jin

Implicit neural representations (INRs) capture global low-frequency structure during the early stage of training and then refine localized high-frequency details, afterwards. However, standard optimizers are agnostic to the coarse-to-fine learning behavior of INRs. Such optimizers accumulate a gradient matrix in momentum buffer and retain components misaligned with the dominant low-frequency directions in the early stage. We propose a progressively growing rank scheduler and a rank-aware learning-rate scheduler for enhancing a subspace-based momentum optimization, where a momentum matrix is updated within a low-rank subspace. The rank schedule expands the subspace during training, and the scheduler holds the peak learning rate through rank growth so that the added subspace contributes to fine-detail recovery. Experiments on image fitting, single-image super-resolution (SISR), and neural radiance field optimization show that the proposed method improves convergence and peak reconstruction quality over competitive optimizers. The largest gains appear in the high-frequency refinement stage, mirroring the coarse-to-fine progression that the rank schedule targets. A significant study on hyperparameter optimization (HPO) on NeRF further confirms that these gains persist under matched search budgets.

Vision language navigation (VLN) is very challenging for Unmanned Aerial Vehicles (UAVs) since small errors in 3D motion prediction can lead to catastrophic outcomes. Unable to fathom execution outcomes in 3D environments, existing methods often struggle to predict accurate navigation waypoints from language instructions. To solve this, we introduce the Prototypical Visual Imagination framework, shortened as ProtoVis-Nav, which forces the model to imagine the execution outcome by predicting visual features surrounding the next intended location of the UAV. Specifically, instead of using multi-layer perceptrons, we predict feature prototypes and use their linear combination to estimate the visual feature of the target region. In this manner, the VLN model first imagines the future location visually, and then uses predicted prototypes as context for generating the final action, promoting alignment between visual context and waypoint outputs. Extensive experiments on the OpenUAV benchmark demonstrate the effectiveness of the ProtoVis-Nav framework and show substantial improvements over existing baselines.


Provable Explanations for Any-Order Neural Additive Models

Idan Refaeli ⋅ Shahaf Bassan ⋅ Yizhak Y. Elboher ⋅ Guy Katz

Post-hoc explanation methods for neural networks are often heuristic and lack provable guarantees. A common alternative to such methods that was studied in recent years is to compute a *cardinally minimal* subset of input features that is *provably sufficient* to determine the model's prediction. For general neural networks, however, finding such explanations is computationally intractable. In contrast, recent work showed that this problem becomes more tractable for a particular family of neural networks, namely neural additive models (NAMs), though standard NAMs are highly restrictive because they exclude feature interactions altogether. In this work, we study the computation of provably minimal and sufficient explanations for neural networks with restricted *feature interactions*, bridging the gap between fully general neural networks and purely additive models. We first prove that even with constant-order interactions, the problem is $\Sigma_2^P$-hard, implying that worst-case exponential complexity is unavoidable. We then identify two provably tractable settings: one based on sparsity in the interaction structure and another based on a stronger interaction-aware notion of sufficiency. Under either one of these settings, we develop algorithms that compute provable explanations for neural networks with any-order interactions while retaining much of the computational efficiency of additive models. Empirically, we show that these explanations are significantly smaller and faster to obtain than those produced by standard algorithms, while being derived from models that achieve much higher accuracy than standard NAMs. Overall, our results help expand the class of neural networks that admit efficient provable explanations, while clarifying the fundamental role of feature interactions in shaping their complexity.


Provable Robustness against Backdoor Attacks via the Primal-Dual Perspective on Differential Privacy

Aman Saxena ⋅ Jan Schuchardt ⋅ Yan Scholten ⋅ Stephan Günnemann

Randomized smoothing provides a powerful framework for certifying robustness against adversarial perturbations by injecting randomized noise either into the training process (against poisoning attacks) or model inputs (against evasion attacks). Yet, extending these guarantees to hold against backdoor attacks, where adversaries jointly perturb both training and test data, remains challenging. In particular, certifying general mechanisms requires tight, compositional guarantees over heterogeneous randomized components, which are not jointly supported by existing approaches. We address this gap by proposing a general framework that numerically composes robustness guarantees for arbitrary mechanisms while remaining tightly connected to black-box randomized smoothing guarantees. To achieve this, we establish a connection between randomized smoothing and the dual perspective of differential privacy, enabling us to leverage advances in the analysis of differentially private mechanisms together with tight numerical composition. This yields a modular framework in which robustness guarantees are derived by reasoning about the privacy of individual components and composing them end-to-end. We instantiate our framework for DP-SGD and Deep Partition Aggregation with inference-time smoothing, deriving joint robustness guarantees against both training-time and inference-time attacks. Empirically, we demonstrate the effectiveness of our framework on MNIST and CIFAR-10. Overall, we provide a principled and general framework for certifying robustness under complex joint threat models and mechanisms, laying the groundwork for future research on unified certification methods towards guarantees under more complex real-world adversaries.


Pruning and Distilling Mixture-of-Experts into Dense Language Models

Junhyuck Kim ⋅ Jihun Yun ⋅ Haechan Kim ⋅ Gyeongman Kim ⋅ Joonghyun Bae ⋅ Jaewoong Cho

Mixture-of-Experts (MoE) is now the dominant architecture for frontier language models, yet it requires all expert parameters to be loaded in memory, making it less preferable for memory-constrained deployment. Existing compression methods reduce the number of experts but the output remains an MoE model with the same fundamental limitation. We present the first systematic framework for converting a trained MoE into a standard fully dense architecture: experts are scored, selected, and grouped, then concatenated into a dense FFN and refined by knowledge distillation from the MoE teacher. We evaluate 7 scoring, 5 grouping, and 2 magnitude scaling methods across a range of selected expert counts on Qwen3-30B-A3B, yielding 350 configurations. We find that the choice of scoring method is the most impactful, with our novel diversity-aware scoring consistently outperforming prior methods on Qwen3-30B-A3B, DeepSeek-V2-Lite, and GPT-OSS-20B. Under a controlled comparison at matched parameter count, MoE-to-dense outperforms dense-to-dense pruning by $+$6.3 pp in average downstream accuracy after ~4B-token distillation at 1.6x faster training wall-clock speed.


PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding

Shengyin Sun ⋅ Yiming Li ⋅ Renxi Liu ⋅ Xinqi Li ⋅ Hui-Ling Zhen ⋅ Weizhe Lin ⋅ Chen Chen ⋅ Xianzhi Yu ⋅ Mingxuan Yuan ⋅ Chen Ma

Diffusion large language models (dLLMs) generate text by iteratively denoising masked token sequences. Although dLLMs can predict all masked positions in parallel within each step, the large number of denoising iterations still makes inference expensive. This cost can be reduced spatially by unmasking multiple tokens per step, or temporally by collapsing multiple denoising steps into one verification call. We propose Parallel Speculative Decoding (PSD), a training-free framework that jointly improves inference along both axes. Using the confidence scores from a single forward pass, PSD selects positions to unmask via a configurable, adaptive unmasking policy and constructs multi-depth speculative drafts without extra model calls. A final batched verification pass then applies hierarchical acceptance, keeping the deepest draft that remains consistent with the updated predictions. Experiments on three dLLMs across reasoning and code generation tasks show that PSD achieves favorable trade-offs between inference efficiency and generation quality, reaching up to $5.5\times$ tokens per forward pass with accuracy comparable to greedy decoding.


Publishing Below-Threshold Triangle Counts under Local Weight Differential Privacy

Kevin Pfisterer ⋅ Quentin Hillebrand ⋅ Vorapong Suppakitpaisarn

We propose an algorithm for counting below-threshold triangles in weighted graphs under local weight differential privacy. Although many prior studies have considered the setting in which the graph topology is public and only the edge weights are sensitive, to the best of our knowledge, this is the first work to study this privacy notion in the local model. Building on a two-round protocol for locally differentially private triangle counting, we exploit the public graph topology to design a novel algorithmic framework. This leads to significant improvements in both accuracy and scalability. In particular, when the input graph is planar, our algorithm eliminates the covariance arising from distributed triangle counting at nodes; for graphs with bounded degeneracy, it significantly reduces this covariance. Since covariance is the dominant source of error in the counting task, our method achieves accuracy that closely aligns with the lower bound. We also present an efficient algorithm for the computation of the smooth sensitivity and provide experiments that quantify the trade-off between the biased and unbiased variants of our estimator and demonstrate the effectiveness of the proposed improvements.

Realistic intensive-care EHR trajectories are difficult to simulate because physiology, treatments, observation times, and patient-existence state co-evolve under sparse clinical measurement. We introduce PULSE, a probabilistic simulator family for ICU trajectories. PULSE separates treatment rollout from patient-state rollout, predicts binary patient state before continuous physiology, and updates continuous variables through gated residual dynamics around the last observed or simulated state. Probabilistic variants add event-time inputs, heteroscedastic Gaussian heads, monotone B-spline-flow residual transforms, latent severity pooling, and covariance-coupled residual groups. On a MIMIC-IV v3.1 sepsis cohort scored over a 37-variable predictive-check panel, reference PULSE reduces k=6 vector-normalized RMSE from 7.124 for a matched per-covariate transformer Monte Carlo baseline to 5.969, and improves over carry-forward from 6.633 to 5.969; paired patient-level bootstrap intervals support both differences. RMSE is a useful sanity check for aggregate scale, but not a sufficient criterion for trajectory realism: carry-forward is already difficult to beat despite generating straight-line futures. Gaussian NLL is the strongest calibrated PULSE variant by CRPS, while B-spline flows address the non-Gaussian shape of factual residuals. Within B-spline flows, event-time encodings with Δt prediction improve k=6 rollout from 8.301 to 6.635 and yield sharper, more state-dependent residual densities. In k=6 density-mixture analysis, B-spline intervals have larger across-state 90% width variation than Gaussian intervals (coefficient of variation 0.79 versus 0.48) and higher 90/95% inclusion despite slightly weaker aggregate PIT uniformity, consistent with better local skew and sharpness in some states. Similar gains over carry-forward persist on eICU sepsis validation panels, while absorbing-state timing remains a major unresolved error source. CVSim known-effect experiments show that synthetic treatment response can be learned under mechanistic ground truth, although transfer to MIMIC-seeded counterfactual tests remains weak.


Punctuation-aware Hybrid Trainable Sparse Attention for Large Language Models

Junxiang Qiu ⋅ Shuo Wang ⋅ Zhengsu Chen ⋅ Hengheng Zhang ⋅ Jinda Lu ⋅ Changcheng Li ⋅ Xinting Hu ⋅ Qi Tian

Attention serves as the fundamental mechanism for long-context modeling in large language models (LLMs), yet dense attention becomes structurally prohibitive for long sequences due to its quadratic complexity. Consequently, sparse attention has received increasing attention as a scalable alternative. However, existing sparse attention methods rely on coarse-grained semantic representations during block selection, which blur intra-block semantic boundaries and lead to the loss of critical information. To address this issue, we propose Punctuation-aware Hybrid Sparse Attention (PHSA), a natively trainable sparse attention framework that leverages punctuation tokens as semantic boundary anchors. Specifically, (1) we design a dual-branch aggregation mechanism that fuses global semantic representations with punctuation-enhanced boundary features, preserving the core semantic structure while introducing almost no additional computational overhead; (2) we introduce an extreme-sparsity-adaptive training and inference strategy that stabilizes model behavior under very low token activation ratios. Extensive experiments on general benchmarks and long-context evaluation tasks, together with ablation studies on punctuation types and the linguistic origin of punctuation tokens, demonstrate that PHSA consistently outperforms both dense attention and the state-of-the-art sparse attention baseline InfLLM v2. Specifically, for a model under the training-inference consistent setting with an input sequence length of 32k tokens, PHSA reduces information loss by 10.8\% at a sparsity ratio of 97.3\%.


Quantifying and Optimizing Path Uncertainty in Masked Diffusion Models

Ziyu Chen ⋅ Xinbei Jiang ⋅ PENG SUN ⋅ Tao Lin

Masked Diffusion Models (MDMs) offer flexible, non-autoregressive generation, but this freedom introduces a challenge: final output quality is highly sensitive to the decoding order. We are the first to formalize this issue, attributing the variability in output quality to the cumulative predictive uncertainty along a generative path. To quantify this uncertainty, we introduce Denoising Entropy, a computable metric that serves as an internal signal for evaluating generative process. Leveraging this metric, we propose two algorithms designed to optimize the decoding path: a post-hoc selection method and a real-time guidance strategy. Experiments demonstrate that our entropy-guided methods significantly improve generation quality, substantially boosting accuracy on challenging reasoning, planning, and code benchmarks. Our work establishes Denoising Entropy as a principled tool for understanding and controlling generation, effectively turning the uncertainty in MDMs from a liability into a key advantage for discovering high-quality solutions.


Quantum Speedups for Stochastic Optimization with Heavy-Tailed Noise

Bin Luo ⋅ Chengchang Liu ⋅ Jonathan Allcock ⋅ Shengyu Zhang ⋅ John C. S. Lui

We study stochastic optimization with heavy-tailed gradient noise. We first propose a novel quantum mean estimator for multivariate heavy-tailed random variables that achieves lower query complexity than optimal classical estimators in the low-dimensional regime. We further develop an unbiased quantum mean estimator by applying a generalized multi-level Monte Carlo technique. We prove quantum lower bounds showing that, when the dimension $d$ of the random vector is small and can be viewed as a constant, our quantum estimators are optimal up to logarithmic factors; while for the high-dimensional regime, no quantum speedup is available compared to optimal classical mean estimators. Based on these estimators, we propose a quantum normalized stochastic gradient descent method ($\texttt{QNSGD}$), which finds an $\epsilon$-stationary point of a nonconvex objective using $\tilde{\mathcal{O}}\big(d^{\frac{p}{4(p-1)}}\epsilon^{-\frac{5p-4}{2p-2}}\big)$ queries to the stochastic gradient, where $p\in(1,2]$ is the tail index. For a convex objective function, we propose a quantum projected stochastic gradient descent method ($\texttt{QPSGD}$), which computes an $\epsilon$-optimal solution with query complexity $\tilde{\mathcal{O}}\big(d^{\frac{p}{4(p-1)}}\epsilon^{-\frac{3p-2}{2p-2}}\big)$. Our results improve upon the classical lower bounds $\Omega\big(\epsilon^{-\frac{3p-2}{p-1}}\big)$ for nonconvex problems and $\Omega\big(\epsilon^{-\frac{p}{p-1}}\big)$ for convex problems, demonstrating a quantum speedup when $d$ is relatively small.


Quasi-Linear ICA for Motor Unit Decomposition during Dynamic Contractions

Alexander K Clarke ⋅ Dimitrios Chalatsis ⋅ Agnese Grison ⋅ Irene Mendez Guerra ⋅ Noura Ezaz-Nikpay ⋅ Pranav Mamidanna ⋅ Shihan Ma ⋅ Silvia Muceli ⋅ Dario Farina

Decomposing surface electromyography (EMG) into the spike trains of individual motor neurons is a long-standing inverse problem and a key step toward motor-neuron-driven neural interfaces such as prosthetics and exoskeletons. The standard approach, independent component analysis (ICA) of the multichannel signal, assumes that the mixing from neurons to electrodes is stationary in time. This assumption fails during movement, when volume-conductor deformation makes the mixing time-varying, and current decomposition algorithms are correspondingly restricted to isometric contractions. We introduce a quasi-linear ICA formulation in which a static linear separator is preceded by a learned, low-rank, time-varying invertible transformation. The separator is trained with an independence loss on the uncompensated projection, and the transformation with a stationarity loss on the recovered source. Gradients are not shared between the two, so the source-extraction step reduces to classical linear ICA and inherits its identifiability guarantee, while non-stationary distortion is absorbed by the transformation. The closed-form inverse of the transformation enables per-spike subtraction with a time-varying template during sequential peel-off. On a public benchmark of dynamic high-density EMG with ground-truth spike trains, the method outperforms five adaptive ICA baselines at every recall threshold, recovering more units at a higher accuracy.

Remote sensing object detection often proceeds in a domain-incremental manner, where detectors continuously encounter new domains arising from changes in regions, resolutions, sensors, and modalities. This challenge is particularly severe when passive optical red-green-blue (RGB) imagery and active synthetic aperture radar (SAR) imagery coexist, because their distinct imaging mechanisms create large domain gaps. Existing methods mainly rely on feature alignment or regularization, but often overlook semantic associations across heterogeneous domains. The problem is more pronounced for DETR-like detectors, where decoder queries act as object-centric semantic slots and are easily disrupted by cross-domain adaptation. Under a query-as-resource view, we propose Q-cost for remote sensing domain-incremental object detection. At the feature level, Q-cost uses modality- and domain-aware prototypes to guide a two-level mixture-of-experts adapter and generate encoder prompts for domain-aware semantic aggregation. At the semantic level, it models decoder queries as limited resources with domain-dependent activity costs, and applies activity-aware gating to regulate query updates according to current and historical domain demands. Experiments on multiple remote sensing benchmarks show that Q-cost effectively balances new-domain adaptation and old-domain retention under severe domain and modality shifts. Our code is provided in the supplementary material.


QUEST: Q-Learning for Uncertainty-Guided Efficient Search Teams

Arsh Verma ⋅ Tejus Gupta ⋅ Jeff Schneider

Robots deployed for active search must decide where to sense while uncertainty, teammates, and remaining mobility change online. In active search, each waypoint decision also commits the robot to a route through the map, making efficient localization a long-horizon problem over belief, topology, robot state, and traversal budget rather than a sequence of independent next-best sensing decisions. We present QUEST, a shared-belief graph neural framework for multi-agent active search with noisy sensors, a shared posterior, and finite decision and path budgets. Rather than learning a myopic policy that values only the next route, QUEST trains Q-functions whose targets include both reward collected along the executed path and the downstream value of the resulting belief, robot positions, team coverage, and remaining budget. The same formulation supports single robots, homogeneous robot teams, and heterogeneous UAV--UGV teams with different sensors and motion budgets. During execution, each robot acts independently using the shared posterior and knowledge of teammates' positions and search plans. Across simulated search environments, QUEST reaches $0.913$ F-score with four UGVs, versus $0.853$ for the strongest learned baseline, and keeps above $0.90$ F-score under communication outage and two mid-mission robot failures without retraining. On out-of-distribution UAV--UGV teams, it improves early F-score ($0.759$ vs.\ $0.728$) while using $10\%$ less path length and $12\%$ less duplicate coverage than the same baseline. The same policies transfer across map shape, variable team size, communication outage, and mid-mission robot failure, and show sensing-mode adaptation after UGV failure.


RA-ClipScore: Making Generative Model Evaluation More Interpretable

Yifan Lu ⋅ Taras Kucherenko ⋅ Hedvig Kjellstrom ⋅ Judith Bütepage

Generative models can produce images nearly indistinguishable from real data, yet rigorous and interpretable evaluation remains challenging. Conventional metrics such as FID provide only scalar scores with limited diagnostic insight. Widely adopted CLIP-based metrics enable semantic evaluation beyond simple training class labels, but inherit limitations from CLIP’s training paradigm that restrict attribute-wise analysis. We propose RA-CLIPScore, a novel metric that mitigates these issues and extends CLIP-based evaluation to spatial distribution alignment, measuring whether generated objects adhere to the positional priors found in the training data. RA-CLIPScore introduces dual prompts to decouple competing attributes and leverages local patch tokens to capture fine-grained regional semantics. We evaluate image generative models on their ability to match both attribute and spatial distributions of the training data. Extensive experiments show that RA-CLIPScore provides more robust and interpretable evaluations than prior methods, particularly under distribution misalignment or partially irrelevant textual attributes. We further demonstrate how it reveals previously undocumented regional tendencies in generative models.


RAG in a Trenchcoat: When Minimal Memory Is Enough for Agentic Systems, and When It Isn’t

Jingyu Liu ⋅ Zongze Li ⋅ Zach Xu ⋅ Zhanhui Zhou ⋅ Tahseen Rabbani ⋅ Dawn Song ⋅ Ce Zhang

Agentic memory research has accumulated extraction, supersession, reflection, and graph synthesis on top of retrieval-augmented generation. Yet on the field's three established benchmarks, almost none of those primitives earn their cost: We show this with Bridge, a deliberately minimal RAG memory with chained observation notes, hybrid BM25 + cosine retrieval, and chronological rendering, which reaches parity within bootstrap CI with the prior published SOTA on LongMemEval-S and BEAM, and beats the best memory baseline on EverMemBench by $+5.7-6.3$pp at matched reader. It runs at $\sim5$-$7$K reader-prompt tokens per query and a single LLM call for each ingestion or query. We argue this parity is itself the result: standard benchmarks score what memory retrieves, not what memory enables when the agent acts. To test the latter, we introduce five adversarial probes (Rule Application $\pm$, Pattern Induction style / procedural, Dynamic Multi-turn) judged at generation time by a cross-vendor verification chain; each can only be solved by applying a remembered rule or pattern the probe never mentions. Across 5 probes $\times$ 20 scenarios, no architecture wins universally: atomic-fact extractors win rule application (LangMem leads Bridge by $+45$pp on RA$^+$, with non-overlapping bootstrap CIs), while Bridge's block-summary memory wins pattern induction (leads the strongest baseline by $+33$pp on PI-P). The right memory architecture is mechanism-dependent, and the field's evaluation paradigm fails on both ends: standard benchmarks fail to distinguish a minimal memory from a complex one, and application-grade probes that do distinguish find no universal winner.

The growing context lengths of large language models (LLMs) have made key-value (KV) cache memory a critical bottleneck during inference. Vector quantization (VQ) with commutative codebooks provides an effective Rotary Position Embedding (RoPE)-aware solution by enabling reconstruction-free attention over quantized key indices. However, existing commutative VQ methods still parameterize scalar codebook coefficients across RoPE subspaces and codewords largely independently, leaving parameter-side redundancy underexploited. We propose $\textbf{RankVQ}$, a low-rank parameterized commutative vector quantization method for KV cache compression. Instead of factorizing the deployed codebooks, RankVQ factorizes the underlying scalar coefficient matrices used to construct key and value codebooks. This design preserves the pseudo-symmetric structure required for RoPE commutativity on the key side, while improving the quality and compactness of offline codebook construction. The low-rank factors are discarded after codebook construction, so RankVQ keeps the same online inference form as standard commutative VQ and introduces no additional decoding overhead. Experiments on LongBench, InfiniteBench, and GSM8K show that RankVQ achieves a stronger accuracy--compression trade-off than directly comparable KV quantization baselines, with particularly clear gains in the 1-bit regime.


RAOP: Step-Level Resource Orchestration for LLM Agents across Edge and Cloud

Jinze Li ⋅ Xin Yang ⋅ Shuo Yang ⋅ Jinfeng Xu ⋅ Edith C Ngai

LLM agents do not present a fixed inference job to an edge and cloud system. At each step, a placement decision chooses where to generate the next message, where to execute any resulting tool call, which prefix must be replayed, which cache state is updated, and whether the episode continues. As a result, the workload to be scheduled is partly produced by the scheduling decisions themselves. We formalize this problem as an endogenous, sequentially revealed workload (ESRW) controlled trajectory MDP. We then introduce Resource Aware Orchestration Policy (RAOP), a step level orchestration policy that first places language generation and then, after observing any tool call, places tool execution. RAOP predicts candidate action workloads from prefix replay, prefill, KV cache residency, queue, network, and tool payload features, and is trained with GRPO to optimize measured full trajectory success and latency rather than a one step latency surrogate. To make controlled evaluation auditable, the main simulator replays language and tool continuations produced by real model and tool execution while recomputing physical timing under the sampled placement sequence. As an analytical lens, we derive a Bellman regret decomposition showing that myopic placement policies incur error from workload prediction, success priors, and a controlled endogeneity term. On GSM8K, HotpotQA, WebShop, and a calibrated three tier deployment, RAOP matches the strong cloud only success reference while reducing P99 latency by about 53\% and communication by about 73\%. Against the strongest learned step level baseline, RAOP improves average success by 2.5 points, indicating that full trajectory optimization is central for agent orchestration.

Time series foundation models (TSFMs) commonly adapt to new data by attaching a single trainable head to a frozen backbone, a one-size-fits-all setup that underfits heterogeneous regimes. Replacing the head with a mixture of experts is the standard upgrade, but on instance-normalized backbones (the dominant TSFM design class) it fails: routing entropy collapses to zero and one expert absorbs every input, a failure we call *normalization-induced routing collapse*. Standard MoE rescue mechanisms do not repair it, because the cause is in the router's input, not its optimization. Pre-encoder normalization strips the statistics a router would need to tell regimes apart. A mutual-information decomposition makes this precise and yields a signal-ratio that, computed before training, predicts dataset vulnerability (Spearman $\rho = -0.88$). Eight causal controls, including a vision-modality replication, isolate instance normalization as the cause. The prescription is a minimal causal intervention: *Raw-Routed Mixture of Adapters* (RR-MoA), which routes on the raw, pre-normalization input. Under a strictly frozen backbone, RR-MoA wins 54/54 comparisons against the strongest fixed adapter and significantly outperforms LoRA, TRACE, AdaMix, and full fine-tuning. The effect generalizes across six backbones and an imputation task. Frozen RR-MoA also beats full fine-tuning by 12–79% (the *Frozen Paradox*); two architecturally distinct variants confirm the principle generalizes beyond this specific router. Code is provided in the supplementary material.


RB-LDC: Redundancy-Balanced Latent Coding for Robust Diffusion

Hyunseok Jeong ⋅ Jaeho Jeon ⋅ Young-Sik Kim

Latent diffusion models (LDMs) achieve high-fidelity generation by moving diffusion from pixels to a compact latent space, but standard tokenizers are optimized for compression, not robustness under diffusion-time corruption. We formulate latent tokenization as diffusion-aware joint source--channel coding and diagnose standard LDM pipelines as missing an explicit channel-coding step. A diffusion-aware rate--robustness functional admits an explicit spectral characterization in $G^\top G$, yielding the Redundancy-Balance Principle: optimal redundancy is not uniform, but aligns with semantic sensitivity and equalizes marginal integrated-risk reduction across active directions. The resulting optima take weighted tight-frame and water-filling forms; isotropic bottlenecks are minimax-suboptimal under heterogeneous sensitivity, and a group-conditioned extension identifies rare groups as capacity-limiting under a fixed latent budget. We instantiate the theory as RB-LDC (Redundancy-Balanced Latent-Diffusion Coding), a lightweight tokenizer modification that inserts a learned coding layer before diffusion. Six controlled experiments validate the theorem sequence up to $k{=}1024$, with a $53.5\times$ alignment gap, $659\times$ minimax gap, and real-image VAE learnability via Adam recovery of the KKT spectrum with relative error below 0.5 percent. On SD-VAE and VA-VAE at ImageNet-256, RB-LDC outperforms isotropic uniform redundancy in all 20 perturbed metric cells over a 15-SNR grid; VA-VAE block-local obtains $-15.40$ average pFID and $-105.79$ peak pFID at $\log_{10}\rho{=}+1$, while preserving clean rFID parity (gap $-0.0001$). RB-LDC turns latent tokenizer design into a reliability problem for diffusion models.


REACT: A Lightweight Reliability-Aware Framework for Spatio-Temporal Out-of-Distribution Prediction

Yongfeng Su ⋅ Ziquan Fang ⋅ Jihua Yang ⋅ Wei Shao ⋅ Yunjun Gao

In the real world, spatio-temporal forecasting systems inevitably encounter out-of-distribution (OOD) shifts, where temporal shifts evolve due to seasonal patterns or policy changes, and spatial dependencies shift with sensor deployment, removal, or topology reconfiguration. While many approaches attempt to address these challenges via invariant learning or causal adjustment, they implicitly assume that the contextual signals supporting prediction remain reliable. Through empirical studies and theoretical analysis, we find that this assumption often breaks because: (i) node-level context can become biased or incomplete under temporal and spatial OOD, yet existing methods lack mechanisms to assess such reliability; (ii) existing standard graph-based aggregation propagates information with fixed strength, potentially amplifying corrupted context across nodes. To this end, we propose REACT, a lightweight REliability-Aware Context-calibrated spatio-Temporal forecaster that explicitly models and controls context reliability under OOD. REACT decouples prediction from adaptation by first establishing a topology-independent predictive anchor via a graph-free linear encoder, ensuring transferable node-wise representations. It then constructs an availability-aware context that captures both observed signals and prior-based imputations, enabling explicit characterization of context completeness. Building on this, REACT introduces a reliability-guided calibration mechanism that dynamically modulates spatial support through uncertainty-aware gating, selectively leveraging trustworthy neighbors while suppressing unreliable propagation. Extensive experiments on ten real-world datasets with diverse temporal and structural shifts demonstrate that REACT consistently outperforms eleven baselines. These results highlight the importance of explicitly modeling and controlling context reliability as a fundamental principle for robust spatio-temporal forecasting under distribution shift. Our source code is available at https://anonymous.4open.science/r/REACT-A82B.


Read It Back: Pretrained MLLMs Are Zero-shot Reward Models for Text-to-Image Generation

Runhui Huang ⋅ Qihui Zhang ⋅ Zhe Liu ⋅ Yu Gao ⋅ Jie Wu ⋅ Hengshuang Zhao

In this paper, we propose SpectraReward, a training-free reward function that turns pretrained MLLMs into off-the-shelf reward models for image-generation reinforcement learning. Instead of asking the MLLM to judge a generated image or answer decomposed verification questions, SpectraReward measures how well the original prompt can be recovered from the generated image through a single image-conditioned, teacher-forced forward pass. We use the average image-conditioned prompt log-likelihood as the reward, directly reusing the MLLM's pretrained image-text alignment ability without preference labels, reward-model fine-tuning. We further introduce Self-SpectraReward, a special case for unified multimodal models where the policy's own understanding branch serves as the reward model for its generation branch, forming a closed-loop self-improving framework without external reward models or external knowledge. Extensive experiments validate SpectraReward through a broad image-generation RL study covering two diffusion models, three RL algorithms, nine reward MLLM backbones from four MLLM families spanning 4B to 235B parameters, and five out-of-distribution text-to-image benchmarks. Results show that both SpectraReward and Self-SpectraReward significantly and consistently improve generation performance and outperform prior MLLM-derived reward training methods. Further analysis reveals that larger reward MLLMs are not always better, while Self-SpectraReward can match or surpass much larger external reward models, suggesting that reward-policy alignment is a key factor for effective image-generation RL.


RealDev-QA: Trajectory-Level Diagnosis for Developer RAG Under Real-World Noises

Tianling Lan ⋅ Yao Junxiao ⋅ Tingyang Chen ⋅ Cibo Yu ⋅ Siyuan Gong ⋅ Cong Fu ⋅ Xiangyu Ke ⋅ Beng Chin Ooi

Real-world developer help-seeking is rarely a clean search query: requests mix stack traces, failed attempts, environment constraints, and mistaken hypotheses, while relevant-looking documents may be invalid under the user's specific runtime settings. We introduce RealDev-QA, a trajectory-supervised benchmark for developer RAG under diagnostic query noise and evidence conflict. RealDev-QA contains 1,233 post-June-2024 multi-hop questions derived from public developer discussions, each paired with decomposed sub-queries, dependency edges, per-step gold evidence, verified reasoning graphs, and adversarial hard negatives drawn from real documentation, release notes, and verified issues. Its construction uses answer-conditioned evidence tracing to ensure each retained instance has a verifiable evidence path; evaluation then removes the resolution and tests whether systems can recover and use that path from the diagnostic query alone. Across static, graph-based, and agentic RAG systems, comprehensive evaluations show that retrieval helps but does not solve RealDev-QA: the best human-audited accuracy reaches only 24.7\%. Trajectory diagnostics show that failures are typically decided early: over 78\% of incorrect trajectories miss gold evidence in the first retrieval step, and first-query precision predicts downstream correctness beyond aggregate evidence coverage—broad initial queries admit hard negatives that persist and pollute later evidence navigation. Low correctness decomposes into five pipeline-level failure modes, while retrieval, automatic, and human-grounded metrics yield divergent system rankings largely masked by aggregate evaluations on standard benchmarks. RealDev-QA reframes developer RAG as evidence-path control: deciding where to start, what to connect, what to reject, and how to recover.

CoT prompting improves LLM accuracy on complex tasks but often increases token usage and inference cost. Existing "Budget Forcing" methods reduce cost via fine-tuning with heuristic length penalties, suppressing both essential reasoning and redundant filler. We recast efficient reasoning as a lossy compression problem under the IB principle, and identify a key theoretical gap when applying naive IB to transformers: attention violates the Markov property between prompt, reasoning trace, and response. To resolve this issue, we model CoT generation under the CIB principle, where the reasoning trace $Z$ acts as a computational bridge that contains only the information about the response $Y$ that is not directly accessible from the prompt $X$. This yields a general Reinforcement Learning objective: maximize task reward while compressing completions under a prior over reasoning traces, subsuming common heuristics (e.g., length penalties) as special cases (e.g., uniform priors). In contrast to naive token-counting approaches, we introduce a semantic prior that measures token cost by surprisal under a language model. Crucially, the prior is queried only for token-level log-probabilities, adding negligible overhead to the training loop. Empirically, our CIB objective prunes reasoning redundancy while preserving fluency and logic, improving accuracy at moderate compression and enabling aggressive compression with minimal accuracy drop. These gains generalize across model families and task domains, confirming CIB as a domain-agnostic CoT compression framework.


RECAP: Looking Once Is Not Enough for Vision-Language Reasoning

Zhaolu Kang ⋅ Tailong Luo ⋅ Chenxin Li ⋅ Zhenyu Yu ⋅ Fengyu Zhou ⋅ Jiachen Qian ⋅ Lei Wei ⋅ Shuang Chen ⋅ Jiachen Li ⋅ Yingjie He ⋅ Eric Hanchen Jiang ⋅ Rongchao Zhang ⋅ Zhengtao Yao ⋅ Hoi Leong Lee ⋅ Guansu Wang ⋅ Kaiyue Zhou

Long chain-of-thought reasoning improves vision-language models (VLMs), but it also exposes a temporal grounding failure: as generation unfolds, models may drift from image evidence while reinforcement learning still provides only a final-answer reward. We study this gap through visual revisit, a temporally localized reactivation of image evidence during reasoning. Across three VLM backbones, revisit peaks are predictive of correctness and causally linked to performance: masking them degrades accuracy more than masking non-peaks, while injecting revisit-like peaks into failed rollouts partially restores correct answers. Control analyses show that this signal is not explained by raw attention strength, uncertainty, logit margin, or hidden-state norm. We propose RECAP, a GRPO-compatible training method that turns detached revisit traces into a credit-assignment signal. RECAP uses revisit-conditioned advantage estimation, dynamic discounting, and gated value prediction to propagate final rewards through visually important reasoning steps, without extra rewards, additional rollouts, or inference-time changes. Across Qwen2.5-VL-7B, InternVL3-8B, and LLaVA-OV-7B, RECAP consistently improves over GRPO and VPPO, yielding $+6.1$--$+6.5$ HallusionBench gains over the SFT base while preserving reasoning and text-only performance. Results on larger LoRA-tuned models, Chinese zero-shot benchmarks, visual perturbations, and human evaluation further demonstrate robust improvements in image-grounded reasoning.


Reconciling Safety and Performance via Dual-Expert Offline Imitation Learning

Seokin Seo ⋅ Sung Kuk Shyn ⋅ Minji Seo ⋅ Yisoo Lee ⋅ Jongmin Lee ⋅ Hongseok Yang ⋅ Kee-Eung Kim

Ensuring high returns while providing strong safety guarantees in offline imitation learning (IL) is a fundamental challenge: existing approaches either rely on explicitly specified cost functions, which are difficult to design for complex real-world safety constraints, or require expert demonstrations that are both safe and high-performing, which are often too costly to collect. We introduce dual-expert imitation learning, which leverages two complementary data sources: (i) safety-compliant demonstrations that faithfully follow safety guidelines but may be suboptimal, and (ii) performance-oriented demonstrations that achieve high returns but may violate safety guidelines. To address this problem, we propose DexDICE (Dual-expert offline imitation learning via stationary DIstribution Correction Estimation). Instead of conventional distribution constraints, DexDICE combines DICE-style offline learning with safe support constraints, enabling agents to exploit high-return behaviors while remaining within safety-compliant regions. We evaluate DexDICE on a re-curated DSRL benchmark and validate it on a real-world mobile-robot navigation task trained directly from human teleoperation demonstrations, where it significantly outperforms offline IL baselines and remains robust to data scarcity and quality degradation.


ReCon: Toward Balanced Learning under Inter-Context and Context-Memory Conflicts

Huipeng Ma ⋅ Dandan Song ⋅ Shaonan Ma ⋅ Changzhi Zhou ⋅ Jun Yang ⋅ Yuhang Tian ⋅ Luan Zhang ⋅ Chenhao Li ⋅ Xudong Li ⋅ Fang Xi ⋅ Guangyuan Feng

In retrieval-augmented generation (RAG), knowledge conflicts arise when retrieved contexts disagree with each other or with the model's parametric memory. Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a standard approach to address such conflicts. However, existing RLVR-based methods are typically formulated under a single-conflict assumption. As a result, the type with stronger advantage signals dominates policy updates, leaving the other under-optimized and thus yielding imbalanced performance across conflict types. To address this challenge, we propose ReCon, a type-aware RLVR framework for balanced learning across Inter-Context and Context-Memory conflicts. Specifically, ReCon computes a step-level imbalance signal from type-wise cumulative advantages and triggers resampling on the weaker conflict type. Instead of resampling from scratch, ReCon branches from informative failures. It selects an informative failed trajectory using uncertainty and closeness signals, then branches at an entropy-jump point, concentrating additional rollouts on critical reasoning decisions. Experiments show that ReCon yields stronger and more balanced performance across knowledge-conflict and multi-hop QA benchmarks, and also generalizing well under mixed-conflict settings.

We study the sample complexity of noisy one-bit compressed sensing for signals drawn from a prior distribution. By characterizing the effective distributional complexity of the prior via its approximate covering number, we prove that posterior sampling achieves accurate recovery with high probability when the number of measurements scales with the logarithm of the approximate covering number, up to a one-bit separation gap factor. This upper bound is robust to learned prior mismatch. Specifically, we show that posterior sampling with an approximate prior remains reliable, provided that the learned prior distribution is sufficiently close to the true signal distribution in Wasserstein distance. In addition, we establish a sample complexity lower bound for any reliable method of noisy one-bit compressed sensing, showing that our upper bound is nearly matched in its main prior dependent term. To approximate the ideal posterior sampling process for real world scenarios, we instantiate posterior sampling through a plug-and-play algorithm with diffusion priors. Experiments on the FFHQ and ImageNet datasets demonstrate the effectiveness of our proposed approach.


Recursive Multi-Agent Systems

Jiaru Zou ⋅ Rui Pan ⋅ Ruizhong Qiu ⋅ Pan Lu ⋅ Shizhe Diao ⋅ Jindong Jiang ⋅ Hanghang Tong ⋅ Tong Zhang ⋅ Markus Buehler ⋅ Jingrui He ⋅ James Zou

Recursive or looped language models have recently emerged as a new scaling axis by iteratively refining the same model computation over latent states to deepen reasoning. We extend such scaling principles from a single model to multi-agent systems, and ask: Can agent collaboration itself be scaled through recursion? To this end, we introduce RecursiveMAS, a recursive multi-agent framework that casts the entire system as a unified latent-space recursive computation. RecursiveMAS connects heterogeneous agents as a collaboration loop through the lightweight RecursiveLink module, enabling in-distribution latent thoughts generation and cross-agent latent state transfer. To optimize our framework, we develop an inner-outer loop learning algorithm for iterative whole-system co-optimization through shared gradient-based credit assignment across recursion rounds. Theoretical analyses of runtime complexity and learning dynamics establish that RecursiveMAS is more efficient than standard text-based MAS and maintains stable gradients during recursive training. Empirically, we instantiate RecursiveMAS under 4 representative agent collaboration patterns and evaluate across 9 benchmarks spanning mathematics, science, medicine, search, and code generation. In comparison with advanced single/multi agents and recursive computation baselines, RecursiveMAS consistently delivers an average accuracy improvement of 8.3\%, 1.2$\times$-2.4$\times$ end-to-end inference speedup, and 34.6\%-75.6\% token usage reduction.


REDACTED-Tunes: An Open 1.4M-Track Dataset and Perceptual Benchmark for AI-Generated Music

Robert Kaczmarczyk ⋅ TAWSIF AHMED ⋅ Felix Friedrich ⋅ Aidan C Erickson ⋅ Orian Sharoni ⋅ Dorien Herremans ⋅ Christoph Schuhmann

Generative music platforms have reached commercial scale, with outputs approaching human-made quality. Yet music machine learning lags image-text research because no open music dataset matches the scale that catalyzed image-text foundation models: commercial recordings cannot be legally redistributed, and existing corpora are small, single-platform, or both. AI-generated music is the natural substrate to close this gap: We introduce $\textbf{REDACTED-Tunes}$, an open dataset of URLs and metadata for $\textbf{1{,}429{,}734 AI-generated music tracks}$ from Suno, Udio, and Mureka. The release ships public CDN URLs, 768-d audio and text embeddings, captions, ASR transcription embeddings (not raw transcripts), five-dimensional aesthetics scores, real and predicted engagement counts, and three-axis NSFW safety labels under Apache 2.0. We also release the models used to annotate the corpus: a 242M-parameter $\texttt{music-captioner}$ and fingerprint extractor, a $\texttt{SongEval}$-calibrated quality scorer, and an engagement predictor trained on platform play and upvote counts. To demonstrate downstream utility, we construct a perceptual benchmark from $\textbf{REDACTED-Tunes}$ and genre-matched human controls. In this study, $\textbf{61 participants annotated 591 song excerpts}$, and $\textbf{12 LLM-judge configurations}$ evaluated the same songs. Human listeners detect AI music above chance ($d' = 0.83$) but modestly (62.5\% accuracy), outperforming every LLM configuration on balanced accuracy. We further identify a quality--authenticity halo effect: songs receiving higher aesthetic ratings are more likely to be judged human-made; equivalently, songs judged real receive nearly two more aesthetic-quality points than songs judged AI-generated, independent of true provenance ($p < 10^{-26}$). All dataset artifacts, models, and code are available at https://anonymous.4open.science/r/anonymized-for-double-blind-review, alongside a live $\textbf{Tunes-Search demo}$ at https://anonymous.4open.science/w/anonymized-for-double-blind-review/search-tunes.


Re-evaluating Continual Learning with Few-Shot Adaptation

Amogh Inamdar ⋅ Matthew So ⋅ Vici I Milenia ⋅ Richard Zemel

Continual learning methods aim to maximize the stability and plasticity of machine learning models that are trained on a sequence of tasks. The standard measure of stability (i.e., forgetting) is the 0-shot performance of a model on previously learned tasks, and plasticity, the performance on the most recently learned task. However, 0-shot evaluation does not fully measure a model or method's ability to retain learned information or adapt quickly to new information, as it requires perfect recall across multiple tasks. In this paper, we propose few-shot evaluation as a more comprehensive assessment of the stability and plasticity of a continual learning system. Through few-shot evaluation with a novel metric---per-shot plasticity---on task sequences for continual image classification, we show that a mixture of foresight via the meta-learning of a short sequence of future tasks and hindsight via replay on previously learned tasks leads to a model that strikes an ideal balance between stability and plasticity in a continual learning setting.


Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning

Xuan Gong ⋅ HanBo Huang ⋅ Hao Zheng ⋅ Yiran Zhang ⋅ Wenbin Dai ⋅ Weishu Zhao ⋅ Shiyu Liang

Long chain-of-thought (CoT) reasoning improves large vision–language models, but visual information often fades during generation, limiting long-horizon multimodal reasoning. Existing methods either re-inject vision at inference or train policies for stronger grounding, but where to intervene relies on perception heuristics rather than principled gain analysis, and how local visual influence propagates remains implicit. We study this problem from an information-theoretic standpoint and derive a lower bound on the downstream visual gain of a one-step intervention, which suggests two factors: local branching room (token entropy) and downstream visual propagation potential (suffix divergence from a vision-marginalized reference). Guided by this analysis, we propose reflection-anchor policy optimization (RAPO), a GRPO-based policy optimization method that selects high-entropy reflection anchors and optimizes a chain-masked finite-window KL surrogate for downstream visual dependence. Experiments on reasoning-intensive and general-domain benchmarks show that RAPO delivers substantial gains over strong baselines across multiple LVLM backbones. Mechanism analyses further indicate that reflection anchors are enriched for visually sensitive decision points and that RAPO increases contrastive visual-dependence signals along generated trajectories.


Reflection with Action-Induced Visual Differences for Desktop GUI Agents

Yijie Ma ⋅ Chaoyue Niu ⋅ Fan Wu ⋅ Guihai Chen

The Planner-Operator-Reflector (POR) framework is widely used in GUI agents to maintain objective alignment in complex tasks through modular collaboration. However, desktop GUIs introduce a key challenge: large, dense interfaces often exhibit subtle or scattered state changes, placing most of the burden on the reflector, which must compare pre- and post-action screens, while the planner and operator reason over a single state. Existing reflectors collapse change detection and outcome verification into one step, leaving evidence implicit and yielding weakly grounded decisions. To address this limitation, we propose Evidence-First Reflection (EFR), a two-stage reflector that explicitly decouples action-induced visual differences extraction from outcome verification. EFR identifies the action location and candidate changed regions with Set-of-Marks annotations, describes and filters action-relevant changes, and makes the final judgment from the cleaned evidence. This evidence-reasoning decoupled design makes reflection better grounded in screen transitions, while reducing both visual search complexity and reasoning burden. Experiments on OSWorld-Verified and WindowsAgentArena demonstrate that EFR improves reflector accuracy by 7.01\%, yielding end-to-end task success gains of 5.86\% and 5.05\% on the two benchmarks, respectively.

RNA molecules often function through multiple experimentally observed conformational states. A central challenge is to capture ensemble diversity without sacrificing structural fidelity: samples may collapse toward averaged representative structures or drift into unsupported conformational regions. We introduce $\texttt{REFLEX}$ (\underline{R}NA \underline{E}nsemble generation calibrated by \underline{FLEX}ibility), a sequence-conditioned generator for RNA backbone-frame ensembles based on a heteroscedastic stochastic bridge on $\mathrm{SE}(3)$. During training, experimental $\mathrm{B}$-factors calibrate residue-wise bridge widths, exposing each observed conformer through local perturbation neighborhoods with residue-specific scales while keeping constrained regions sharper. A learned flexibility condition guides the bridge-field network to recover target conformer-specific frames from these calibrated noisy states. At inference, $\texttt{REFLEX}$ uses only the input sequence, requiring no MSA-subsampling heuristics. On held-out multi-conformer RNA clusters, $\texttt{REFLEX}$ achieves the best precision--recall trade-off for ensemble coverage across evaluation thresholds, while ablations indicate that the learned bridge field recovers conformer-specific basins more reliably.


Reformulate LLM Reinforcement Learning for Stable Training under Black-box Discrepancy

Jiashun Liu ⋅ Runze Liu ⋅ Xu Wan ⋅ Jing Liang ⋅ Hongyao Tang ⋅ Ling Pan

Reinforcement Learning (RL) has emerged as a pivotal post-training paradigm for Large Language Models (LLMs), yet it frequently suffers from unpredictable training collapses. Recent findings attribute these failures to a hidden train-inference discrepancy (or mismatch), such a discrepancy will increase as training goes on, stemming from the disparate underlying engines and precisions required to balance generation throughput and training fidelity. Such Existing mitigations either sacrifice numerical stability or rely on heuristic masking that fails to explicitly optimize the ultimately deployed policy. In this paper, we discover that training policies inherently possess the capability to heal this common and harmful discrepancy. To operationalize this, we transition the standard RL objective into a Discrepancy-Constrained Markov Decision Process (\texttt{DCMDP}). In order to practice this new paradigm at the algorithmic level, we first introduce a robust, trajectory-level geometric penalty that provides black-box feedback of inter-policy deviations, enabling autonomous self-correction under any specific trigger, e.g., infrastructure level or model architecture level. Furthermore, capitalizing on our empirical discovery of a \textit{discrepancy tolerance region} in various models, we employ an adaptive performance-discrepancy balancing mechanism that penalizes the policy only when deviations exceed a safe boundary, achieving stable dual-objective optimization. Our approach not only eradicates mismatch-induced collapses and shatters performance bottlenecks, but crucially unlocks a heterogeneous training paradigm—leveraging high-fidelity training environments to natively optimize LLMs for low-cost, resource-constrained deployments.


Reinforcement Learning from Rich Feedback with Distributional DAgger

Rishabh Agrawal ⋅ Jacob Fein-Ashley ⋅ Paria Rashidinejad

Reasoning models have advanced rapidly, but the dominant reinforcement learning from verifiable rewards (RLVR) recipe remains surprisingly narrow: sample many responses and reward each with a single bit indicating whether the final answer is correct. Yet many settings provide $\textit{rich feedback}$, including execution traces, tool outputs, expert corrections, and model self-evaluations. We study how to use such feedback through a distributional variant of the classic imitation learning algorithm DAgger, where the learner has local access to an expert distribution on states visited by the current policy. This yields a simple forward cross-entropy objective that admits a blackbox expert and whose sequence-level gradient {conduct rich credit assignment by propagating} future expert-student disagreement back to earlier decisions. We show that prior RL with self-distillation objectives based on reverse KL or Jensen-Shannon fail to guarantee monotonic policy improvement: even when the expert has higher reward, their updates may increase probability on worse actions. In contrast, we show that forward cross-entropy admits monotonic policy improvement and enjoys guarantees on regret. We further show that our objective optimizes a teacher-weighted maximum-likelihood RL lower bound, leading to improved Pass@N. Empirically, our approach, DistIL, consistently improves over RLVR and RL with self-distillation baselines across a variety of domains: scientific reasoning, mathematical reasoning, and coding.

Generative solvers for combinatorial optimization benefit from additional test-time compute when their refinement states remain aligned with task relaxations. We introduce RAGenCO, a relaxation-aligned test-time scaling framework for generative combinatorial optimization. The central idea is to treat projection onto a task relaxation as a state control primitive: refinement first maintains a projected control state, and only then forms the task-dependent decoder input or expands candidates. For TSP and ATSP, Sinkhorn projection aligns matrix states with the cycle-cover relaxation during refinement and decoding. For MIS and MVC, Bernoulli projection controls refinement on the independent-set side, while raw logits are retained for terminal greedy decoding because their ordering carries the decoder signal. Best-of-$B$ branching then reallocates compute by expanding candidates around aligned control states rather than sampling from unconstrained logits. Because these operations are used only at inference, the diffusion training objective remains unchanged. Under equal runtime budgets, this alignment of control state, decoder input, and branching yields a strong quality-runtime trade-off on TSP and ATSP and competitive results on MIS and MVC.


RelGS: Relation-Aware Gaussian Splatting for Open-Vocabulary 3D Scene Understanding

Kai Zhao ⋅ Dai Shi ⋅ YIXIN CHEN ⋅ Ye Wang ⋅ Oula Ghannoum ⋅ Yi Guo

Open-vocabulary 3D scene understanding provides an important interface for querying and interacting with reconstructed scenes through natural language. Although recent 3DGS methods have enabled text-driven object selection and open-vocabulary segmentation, they still struggle with compositional queries involving attributes, spatial relations, and part-whole relations. Most existing approaches learn per-Gaussian or per-cluster language features, construct query-conditioned referring fields, or perform spatial reasoning only at inference time, but they do not organize the scene itself into a reusable relation-aware representation. As a result, object relations are not explicitly stored as typed and confidence-aware 3D structures, limiting compositional reasoning and multi-hop querying. To address this issue, we propose RelGS, a relation-aware 3D Gaussian Splatting framework for open-vocabulary scene understanding. RelGS learns semantic and instance embeddings for Gaussians under multi-view supervision, groups them into cross-view consistent 3D semantic nodes, and enriches each node with structured attributes verified by reverse CLIP consistency. It further constructs a confidence-weighted relation graph, where 3D spatial cues propose candidate relations and an LLM verifies their semantic plausibility. At query time, RelGS combines attribute matching, CLIP retrieval, and graph traversal to localize the target object. Extensive experiments demonstrate strong performance on word-level selection, sentence-level referring segmentation, multi-hop relational queries, and point-level semantic segmentation.


ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization

Junbo Jacob Lian ⋅ Yujun Sun ⋅ Huiling Chen ⋅ Chaoyu Zhang ⋅ Hanzhang Qin ⋅ Chung-Piaw Teo

Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that executes and returns solver-feasible solutions may encode semantically incorrect formulations—a feasibility–correctness gap reaching 90 percentage points on compositional problems. We introduce ReLoop, which addresses this gap through two complementary mechanisms. Structured generation decomposes code production into a four-stage reasoning chain (understand, formalize, synthesize, verify), preventing formulation errors at their source. Behavioral verification detects errors that survive generation by testing whether the formulation responds correctly to solver-based parameter perturbation—an external semantic signal that bypasses LLM self-review and requires no ground truth. The two mechanisms are complementary by error structure: structured generation drives the largest gains on compositional problems (+8.5pp accuracy on RetailOpt-190 with Claude Opus 4.6), while behavioral verification dominates on localized defects (+4.4pp on MAMO-ComplexLP, its largest contribution across benchmarks). Combined with diagnostic execution recovery, ReLoop reaches 100% executable code on Claude Opus 4.6 and consistently improves accuracy on chat-tuned foundation models across three benchmarks; we further identify a known limitation of narrowly-tuned SFT models, whose learned output formats are brittle to chain-of-thought prompts—an interaction we document and analyze. We release RetailOpt-190, 190 compositional retail optimization scenarios targeting the multi-constraint interactions where LLMs most frequently fail. Code and benchmark will be released upon acceptance.


Remembering What Matters: From Markovian to Subtask-Causal Memory in VLA Policies

Peishuo Wang ⋅ Fengshuo Bai ⋅ Yufeng Li ⋅ Tawei Chou ⋅ Yinda Xu ⋅ Ying Wen ⋅ Yaodong Yang ⋅ Chen Gao ⋅ Zhenzhe Zheng ⋅ Fan Wu

Long-horizon robotic manipulation is often non-Markovian: the correct action can depend on earlier state changes that are no longer visible in the current observation. Vision-language-action (VLA) models provide strong generalist robot policies, but their performance can drop when a task requires remembering which earlier subtasks changed the world. Existing memory mechanisms extend context through recent-frame windows, compressed latent stores, or similarity-based keyframe retrieval. They do not explicitly distinguish which completed subtasks are relevant from which frames within those subtasks carry the needed evidence. We introduce TTC-MEMORY, a two-tier memory framework for long-horizon VLA policies. The first tier segments execution into subtasks and constructs a dependency directed acyclic graph (DAG), so retrieval is restricted to ancestors of the current subtask. The second tier adaptively stores state-change keyframes within each retained subtask and injects them into a frozen VLA backbone through graph-aware cross-attention, with fallback to recent history when planning or clustering is uncertain. We evaluate TTC-MEMORY on LIBERO, RMBench, SimplerEnv-Bridge, the SafeLab chemistry-lab simulation benchmark, and Franka real-robot tasks. The comparisons use a shared VLA backbone, training budget, simulation seed protocol, and physical-trial protocol. TTC-MEMORY improves the long-horizon and memory-dependent slices in this matched setting. On the 10-subtask SafeLab slice, it reaches 58.3% success, compared with 38.2% for the strongest matched prior memory baseline, MemER, and 45.1% for DAG-only retrieval. Real-robot trials provide supporting transfer evidence rather than the main claim. Ablations show that the two tiers are complementary: DAG retrieval increases long-range dependency recall, while adaptive keyframe selection removes redundant history and preserves attention on task-critical state changes.


Render Structure Uncertainty for HTML Repair in MLLM-based UI-to-Code Generation

Haoran Ma ⋅ Jiechao Gao ⋅ Shisong Tang ⋅ bing han ⋅ Jifeng Hu ⋅ Hechang Chen

Existing MLLM-based UI-to-Code generation methods can generate complete HTML/CSS codes through end-to-end or multi-stage pipelines, yet the generated codes often contain structural errors such as incorrect repeated layouts, wrong component boundaries, or missing interaction slots. Subsequent generations may continue along the erroneous structure, causing the error to propagate and amplify. This paper studies post-generation repair: how can we locate and repair unreliable local DOM structures? We observe that this repair process is ambiguous in two ways: a mismatch between the target screenshot and the output may correspond to multiple plausible DOM subtrees, and a single DOM subtree may admit multiple repairs with different rendered behaviors. We propose Render Structure Uncertainty (RSU)for partial DOM repair, a browser grounded uncertainty layer that can be attached to any base HTML generator. RSU parses the target screenshot into a visual structure graph and converts the browser execution trace of the generated code into a DOM-anchored render structure graph. By matching these graphs, RSU yields structural mismatches and a small set of candidate DOM holes, preserving error attribution uncertainty. For each hole, RSU generates multiple local candidates, renders them in the full page context, and clusters them using render-structure signatures that encode layout relations, repeated patterns, content slots, semantic roles, and invalid states, thereby modeling repair uncertainty. The retained representative candidates form a compact Partial DOM candidate set, over which we perform global reranking to produce the final HTML code. In this way, RSU reframes HTML repair into uncertainty-guided search in render-structure space, avoiding premature commitment to either a single error location or a single local repair. Experiments show that RSU improves the generation quality of base models, producing higher-quality HTML with better layout consistency and fewer local structural errors.

Diffusion Large Language Models (DLMs) are being actively explored as promising alternatives to autoregressive models due to their fast and flexible generation capabilities. To optimize inference in DLMs, recent studies have proposed a variety of decoding algorithms. However, the tendency of these methods to mimic autoregressive generation limits the speed and flexibility of DLMs, while also preventing the methods from fully realizing their potential. Moreover, this rigid approach leads to error propagation, while existing sampling-based correction methods incur prohibitive computational costs. To overcome this, we propose Renoise Consistency, a novel post-hoc approach that enables effective self-correction by maximizing train--inference consistency and can be applied to any decoding algorithm with a flexible computational budget. Furthermore, we introduce Adaptive Block Drafting, which leverages consistent internal representations of DLMs to significantly reduce the overall computational cost. When combined, our proposed method achieves up to a 9.00\% performance improvement over the semi-autoregressive baseline, along with a 25.85\% reduction in forward steps.


Reporting Practice Matters: The Impact of Reference Choice on Chest X-ray Report Evaluation

Daniel P Jeong ⋅ Charles Q Li ⋅ Hossein Hosseiny ⋅ Nitya M Bhalla ⋅ Fatma U Morency ⋅ Pradeep Ravikumar ⋅ Zachary Lipton ⋅ Michael Oberst

Radiologists follow heterogeneous reporting practices. Two radiologists examining the same image and identifying the same clinical findings might nevertheless compose superficially distinct reports, varying in terminology, shorthand, formatting, and level of detail. These variations in reporting norms represent an under-appreciated obstacle in efforts to evaluate AI-based radiology report generation (RRG) models, where machine-generated reports are typically assessed based on their concordance with human-generated references. In this paper, we quantify the sensitivity of established evaluation metrics to variations in reporting practices, revealing impacts significant enough to alter the rankings of models. We introduce a radiologist-informed taxonomy of variations in radiology reporting practice and a method (ReRef) that rewrites reference reports along the axes of our taxonomy while preserving clinical interpretation. For instance, when comparing the performance of nine RRG models on MIMIC-CXR using the GREEN metric, condensing the discussion of normal findings in the reference reports causes Libra to drop from first place to third while CheXOne rises from third to first. Our results suggest that current metrics may fail to decouple clinical interpretation from conformity to reporting practices and that choosing the "right" references that accurately reflect the desired reporting practices can be important in practice. To support future research, we release MIMIC-CXR-Ext-ReRef, a radiologist-validated dataset of 120 (original, alternative) reference report pairs derived from MIMIC-CXR.


RePO-VLA: Recovery-Driven Policy Optimization for Vision-Language-Action Models

Weijia Liufu ⋅ Xiaoyu Guo ⋅ Ruiyi Chen ⋅ Jingzhi Liu ⋅ Kaidong Zhang ⋅ Xiwen Liang ⋅ Jianqi Lin ⋅ Dawei Sun ⋅ Yuze Wang ⋅ Rongtao Xu ⋅ Bingqian Lin ⋅ Bowen Yang ⋅ Tongtong Cao ⋅ Bowen Peng ⋅ Dongyu Zhang ⋅ Guangrun Wang ⋅ Min Wang ⋅ Liang Lin ⋅ Xiaodan Liang

Vision-Language-Action (VLA) models remain brittle in long-horizon, contact-rich manipulation because success-only imitation provides little supervision for execution drift, while failed rollouts are often discarded. We introduce RePO-VLA, a recovery-driven policy optimization framework that assigns distinct roles to success, recovery, and failure trajectories. RePO-VLA first applies Recovery-Aware Initialization (RAI), slicing recovery segments and resetting history so corrective actions depend on the current adverse state rather than the preceding failure. It then learns a Progress-Aware Semantic Value Function (PAS-VF), aligning spatiotemporal trajectory features with instructions and successful references. The resulting labels salvage useful failure prefixes via reliability decay, while low-value labels mark drift and terminal breakdowns, teaching differences among nominal, failed, and corrective actions. The data engine turns adverse states into planner-generated or human-collected corrective rollouts, teaching recovery to the success manifold. Value-Conditioned Refinement (VCR) trains the policy to prefer high-progress actions. At deployment, a fixed high value ($v=1.0$) biases actions toward the learned success manifold without online failure detectors or heuristic retries. We introduce FRBench, with standardized error injection and recovery-focused evaluation. Across simulated and real-world bimanual tasks, RePO-VLA improves robustness, raising adversarial success from 20\% to 75\% on average and up to 80\% in scaled real-world trials.


ReSET: Accurate Latency-Critical NVFP4 Reasoning via Step-Aware Temperature Scaling

Sihwa Lee ⋅ Janghwan Lee ⋅ Donghoon Yoo ⋅ Jae Gon Kim ⋅ Hanyul Ryu ⋅ Soojung Ryu ⋅ Jungwook Choi

Large reasoning models (LRMs) improve complex problem-solving by generating long intermediate reasoning traces, but this substantially increases inference costs. NVFP4 inference offers a promising approach to reduce both computational and memory costs through hardware-supported low-precision execution. However, directly applying NVFP4 to LRMs introduces two practical limitations: reasoning accuracy degrades under quantization, and existing NVFP4 kernels do not fully realize latency benefits in small-batch autoregressive decoding. In this work, we analyze the effect of NVFP4 quantization on token-level uncertainty during reasoning. We show that quantization increases incorrect sampling at low-entropy symbolic tokens, while causing over-concentration on a small set of tokens in high-uncertainty reasoning steps. Based on this observation, we propose \textbf{ReSET}, a reasoning-step entropy-based temperature-scaling method that estimates step-level uncertainty online and adapts the decoding temperature using both token-level and step-level entropy signals. To address the latency gap, we further design a CUDA-core small-$M$ NVFP4 kernel for latency-critical autoregressive decoding. Across reasoning benchmarks and model scales, ReSET improves NVFP4 reasoning accuracy by up to $\sim$2 points over the NVFP4 baseline. Our CUDA-core small-$M$ kernel further improves latency-critical decoding, delivering up to $2.5\times$ kernel-level speedup over NVFP4 vLLM and approximately $2\times$ end-to-end decoding speedup over BF16. Code will be released.


Resilient Semi-Supervised Inference with Heterogeneous Unlabeled Data

Mengyuan Wang ⋅ Chengde Qian ⋅ Haojie Ren ⋅ Changliang Zou

Model-assisted semi-supervised learning offers a powerful paradigm for enhancing the efficiency of statistical estimation and inference by leveraging black-box predictions on unlabeled data. However, the validity of these methods typically relies on the strict assumption that labeled and unlabeled populations share identical distributional characteristics. Consequently, state-of-the-art approaches, such as prediction-powered inference, become fragile when this assumption is violated. To address the fragility, we propose a robust framework for semi-supervised learning that remains reliable under distributional heterogeneity. By embedding a robust statistical calibrator into the prediction-rectification mechanism, the approach effectively reduces the bias arising from shifted or corrupted unlabeled samples. To maximize data utility, we introduce an adaptive cross-validation procedure to select the optimal calibrator, ensuring a reliable trade-off between statistical efficiency and robustness. Theoretical analysis confirms the consistency of the proposed estimator under mild conditions, while empirical results demonstrate its significant superiority over conventional baselines in heterogeneous environments.


Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching

Chenyu Zhang ⋅ Yuhang Cao ⋅ Yingxi Lu ⋅ Daru Du ⋅ Jing Shao ⋅ Jiajun Liu ⋅ Ruoqu Chen ⋅ liu cao ⋅ Yicheng Liu ⋅ Hang Zhao ⋅ Mengdi Xu

Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate, establishing a high-performance, scalable foundation for purely autoregressive VLA systems. Our project page is available at https://causalactiontokenizer.github.io.


Rethinking Gradient Approximation in Quantization: A Zeroth-Order Expectation Perspective

Wenkang Wang ⋅ Jia Zhang ⋅ Zelin Wei ⋅ Dongxu Liu ⋅ Jiabao Li ⋅ Wanli Shi ⋅ Bin Gu

As large language models (LLMs) continue to scale, quantization has become a key technique for efficient deployment. However, multi-bit quantization employs non-differentiable rounding, hindering gradient optimization. Existing methods rely on heuristic surrogate gradients (e.g., STE), which work empirically but lack a unified theory. To address this challenge, we propose the Zeroth-Order Expectation Gradient (ZOE-Grad), a zeroth-order expectation view of quantization surrogate gradients. Specifically, we establish an equivalence between the expectations of zeroth-order gradient estimators and surrogate gradients, providing a unified zeroth-order interpretation of widely used surrogates. Under this framework, STE corresponds to a degenerate, discontinuous perturbation that ignores the local geometry of quantization boundaries, limiting its expressiveness. In contrast, continuous perturbation-based surrogate gradients capture boundary-local information and produce smoother, more structured gradients. We further establish bounded-error convergence guarantees for the quantized objective. Through simulations under multiple perturbation distributions, we verify that the proposed framework accurately captures the relationship between the expectations of zeroth-order gradient estimators and surrogate gradients. Furthermore, extensive experiments on OPT-1.3B, OPT-6.7B, LLaMA-2-7B, and Qwen3-8B show that continuous ZOE-Grad surrogates consistently outperform STE across diverse LLM architectures.

Standard generative models face a memorization--generalization trade-off, that is avoiding memorization is considered necessary for generalization. In supervised learning, however, recent studies show that highly overparametrized models can memorize (interpolate) training data while still generalizing well, a phenomenon known as benign overfitting. Motivated by this, we investigate whether generative models can similarly bypass the trade-off and generalize while interpolating. Generative models should minimize the distance between the true data distribution and the distribution induced by mapping the full latent distribution through the generator. But since the true distribution is inaccessible, existing models instead minimize empirical risk with respect to the training distribution. Under this standard formulation, exact empirical-risk minimization forces the generator to produce only training samples. To address this issue, we consider an alternative empirical risk based on presampled latent variables. Our experiments demonstrate that the resulting presampled-latent version of the flow matching model exhibits benign overfitting on standard image benchmarks, such as MNIST and CIFAR-10. As a theoretical proof of concept, we recast generative modeling as a regression problem and extend existing benign overfitting theory to our setting. Together, these results establish, for the first time, that benign overfitting can occur in generative models.


Rethinking Molecular Graph Backdoors under Chemistry-aware Admission

Thinh Nguyen ⋅ Sze Jue Yang ⋅ Khoa D Doan ⋅ Chee Seng Chan ⋅ Kok-Seng Wong

Backdoor attacks on molecular graph neural networks (GNNs) are typically evaluated as abstract graph edits, but real molecular learning pipelines do not train on arbitrary graphs. Molecular records must first survive parsing, sanitization, canonicalization, and graph-string consistency checks. We formalize this overlooked admission stage as ChemGuard, an operational protocol for testing whether a submitted molecular record can enter a realistic learning pipeline, while complementing existing defenses.ChemGuard admits a record only when its molecular string is sanitizable and the graph reconstructed from that string matches the submitted molecular graph. Under this operational view, many existing graph-based backdoors lose much of their apparent efficacy because their poisons are chemically invalid or representation-inconsistent. We then show that admission checks alone are insufficient to rule out molecular backdoors. We propose ChemBack, an admission-aware molecular backdoor attack that constructs chemically feasible motif-anchor attachments and ranks admitted candidates by fingerprint-based Tanimoto similarity to clean target-class molecules. ChemBack is model-free during trigger selection, using molecular structures, target labels, fingerprints, and public validity checks, but no victim model, surrogate GNN, learned embedding, gradient, logit, or training-code access. Across molecular benchmarks, validators, architectures, and defenses, ChemBack achieves high attack success with fully admitted poisons while preserving clean accuracy. Our results reveal a two-sided lesson, chemistry-aware admission suppresses many graph-only backdoors, yet chemically valid and target-aligned molecular backdoors remain a practical threat.


Retrieval Over Training: Similarity search-based Model Selection for Time Series Anomaly Detection

Christos Panourgias ⋅ Roberto Stanzione ⋅ Adrien Petralia ⋅ Themis Palpanas ⋅ Paul Boniol

Anomaly detection is a fundamental task for time-series analytics, with important implications for the downstream performance of many applications. Despite the large number of anomaly detection methods proposed in the literature, recent benchmark studies have shown that no single detector performs best across highly heterogeneous time series. Therefore, a practical and scalable solution is to develop a model-selection method that, for a given time series, selects the anomaly detector most likely to perform well. Nevertheless, the model selection approach proposed in the literature suffers from a significant drop in accuracy when applied in Out-of-Distribution (OOD) settings. In this paper, we tackle the aforementioned limitation and propose \textbf{\textsc{RAMSAD}}, a train-free, retrieval-based framework for model selection in time-series anomaly detection. Given a new time series, our method queries a knowledge base of previously observed instances, retrieves the top-${k}$ most similar series, and transfers detector recommendations from their performance profiles. The framework can operate with standard similarity measures as well as embedding-based representations. Overall, we demonstrate that similarity-based retrieval constitutes a strong and efficient foundation for model selection in time-series anomaly detection.

Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in multi-step reasoning and calling search engines at appropriate steps. However, existing retrieval-augmented reasoning approaches rely on separate retrieval models, limiting the LRM's role in retrieval to deciding when to retrieve and how to formulate queries, even with its inherent ability to handle vast knowledge spaces. This structural separation induces a representation bottleneck, as the retriever’s latent space often lacks the expressive granularity required to satisfy the generator’s sophisticated information needs. To address this, we shift our perspective on retrieval from sequence-to-sequence matching for all corpora to locating the answer-containing paths within the corpus, and propose a novel framework called FREESON (Retriever-FREE Retrieval-Augmented ReaSONing). This framework enables LRMs to directly access external knowledge by acting as both a generator and a retriever. To achieve this, we introduce a variant of the MCTS algorithm specialized for the retrieval task, which we call CT-MCT (Corpus-Traversing Monte Carlo Tree Search). Through this algorithm, the LRM selectively references specific segments identified during traversal, instead of fetching a fixed top-k set. Experiments on five open-domain QA benchmarks covering both single-hop and multi-hop questions demonstrate that FREESON achieves an average improvement of 14.4% in EM and F1 over four multi-step reasoning models with a separate retriever, and it also performs comparably to the strongest baseline, surpassing it by 3% on PopQA and 2WikiMultihopQA, and by 12% on the fact-checking benchmark FEVER.


Return to Basics: Very Simple Graph Contrastive Learning via Noise Cancellation Principle

Yanan Zhao ⋅ Feng Ji ⋅ Jingyang Dai ⋅ Jiaze Ma ⋅ Wee Peng Tay

Graph Contrastive Learning (GCL) has shown strong promise for unsupervised graph representation learning, yet its effectiveness remains limited on heterophilic graphs, where connected nodes often belong to different classes. Existing methods rely on complex augmentation schemes, intricate encoders, or negative sampling, which raises the question of whether such complexity is truly necessary in this challenging setting. In this work, we revisit the foundations of supervised and unsupervised learning on graphs and uncover a simple yet effective principle for GCL: mitigating node feature noise by aggregating it with structural features derived from the graph topology. This phenomenological study suggests that the original node features and the graph structure naturally provide two complementary views for contrastive learning. Building on this insight, we propose an embarrassingly simple GCL model that uses a GCN encoder to capture structural features and an MLP encoder to isolate node feature noise. Our design requires neither data augmentation nor negative sampling, yet achieves state-of-the-art results on heterophilic benchmarks with minimal computational and memory overhead, while offering advantages in homophilic graphs in terms of complexity, scalability, and robustness. Moreover, a variant of GCN-MLP based on the same principle also achieves SOTA performance on homophilic datasets. We provide theoretical justification for our approach and validate its effectiveness through extensive experiments.


Revisiting Activation Steering Through an Optimization Lens

Ryan Lee ⋅ Dung V Nguyen ⋅ Minh Hieu Vu ⋅ Noelle Y.L. Wong ⋅ Lei Zhang ⋅ Dipti Srinivasan ⋅ Duy Linh Tran ⋅ Tan Nguyen

Activation steering has emerged as a powerful approach for controlling large language models (LLMs), with prominent methods such as ActAdd, Directional Ablation, and Angular Steering relying on difference-in-means activations from contrastive prompts across layers. These differences are typically treated as candidate feature directions, later refined into optimal steering vectors or planes. In this work, we reinterpret these candidate directions as gradients of an underlying optimization problem. Building on this perspective, we propose Momentum Steering, a momentum-based framework for activation steering in LLMs. Unlike traditional difference-in-means methods, our framework generates a richer family of candidate directions through momentum updates, enabling more expressive steering. We first introduce a non-causal variant that accumulates difference-in-means signals via momentum, producing enhanced candidate directions. We then develop a causal variant, where future layer statistics are recursively influenced by previously applied momentum directions, explicitly modeling the causal effects of interventions on downstream activations. Momentum Steering is lightweight and modular, making it easily compatible with state-of-the-art steering methods. We empirically demonstrate that Momentum Steering delivers stronger, more robust, and more reliable behavioral control across diverse LLM families and benchmarks.


Revisiting Decentralized Online Convex Optimization with Compressed Communication

Hao Zhou ⋅ Xiaoyu Wang ⋅ Chang Yao ⋅ Mingli Song ⋅ Yuanyu Wan

Decentralized online convex optimization (D-OCO) is a popular framework for distributed applications with streaming data. To tackle the communication bottleneck, previous studies have investigated D-OCO with compressed communication and proposed several algorithms that are variants of online gradient descent (OGD). However, for D-OCO with exact communication, the best existing algorithms are variants of follow-the-regularized-leader (FTRL). In this paper, for the first time, we propose two FTRL-type algorithms for D-OCO with compressed communication. Compared with OGD-type algorithms, our algorithms are more elegant in both algorithmic design and theoretical analysis. The key insight is that the dual update mechanism of FTRL allows us to make a simple application of the technique for average consensus with communication compression. More specifically, our first algorithm considers the full-information setting, and can match the existing regret bounds. Our second algorithm is designed for the bandit setting, and can significantly improve both the regret bounds and communication costs of existing algorithms.


Reward Score Matching: Unifying Reward-based Fine-tuning for Flow and Diffusion Models

Jeongjae Lee ⋅ Jinho Chang ⋅ Jeongsol Kim ⋅ Jong Chul Ye

Reward-based fine-tuning steers a pretrained diffusion or flow-based generative model toward higher-reward samples while remaining close to the pretrained model. Although existing methods are derived from different perspectives, we show that many can be written under a common framework, which we call reward score matching (RSM). Under this view, alignment becomes score matching against a value-guided target, and the main differences across methods reduce to the construction of the value-guidance estimator and the effective optimization strength across timesteps. This unification clarifies the bias--variance--compute tradeoffs of existing designs, and distinguishes core optimization components from auxiliary mechanisms that add complexity without clear benefit. Guided by this perspective, we develop simpler, more efficient redesigns across representative differentiable and black-box reward alignment tasks. Overall, RSM turns a seemingly fragmented collection of reward-based fine-tuning methods into a smaller, more interpretable, and more actionable design space.

In learnable logic gate networks, each neuron implements a $k$-input Boolean function. The dominant setting $k{=}2$ was chosen for tractability. We ask: when does wider fan-in help? At fixed parameter budget, each gate costs $2^k$ parameters, so increasing $k$ exponentially reduces width. Combined with the depth--fan-in bound $L \geq \lceil \log_k d \rceil$, this partitions tasks into two regimes: width-dominated (image classification), where more gates beat richer gates, and degree-dominated (parity, arithmetic), where $k{>}2$ is necessary. Our main finding is that representability does not imply learnability: at the regime boundary, the selection mechanism determines training success at identical architecture, and the optimal mechanism inverts between regimes. We formalize this via a Jacobian analysis: the softmax mechanism's positive-semidefinite (PSD) structure gives basin attraction absent from sigmoid. We train multilinear coefficients with exact gradients and Möbius-transform snapping (zero deployment error), and validate on 28 tasks and 5 real-world datasets across multiple selection mechanisms.


Rich Insights from Cheap Signals: Efficient Evaluations via Tensor Factorization

Felipe Maia Polo ⋅ Aida Nematzadeh ⋅ Virginia Aglietti ⋅ Adam Fisch ⋅ Isabela Albuquerque

Moving beyond evaluations that collapse performance across heterogeneous prompts toward fine-grained evaluation at the prompt level, or within relatively homogeneous subsets, is necessary to diagnose generative models' strengths and weaknesses. Such fine-grained evaluations, however, suffer from a data bottleneck: human gold-standard labels are too costly at this scale, while automated ratings are often misaligned with human judgment. To resolve this challenge, we propose a novel statistical model based on tensor factorization that merges cheap autorater data with a limited set of human gold-standard labels. Specifically, our approach uses autorater scores to pretrain latent representations of prompts and generative models, and then aligns those pretrained representations to human preferences using a small calibration set. This sample-efficient methodology is robust to autorater quality, more accurately predicts human preferences on a per-prompt basis than standard baselines, and provides tight confidence intervals for key statistical parameters of interest. We also showcase the practical utility of our method by constructing granular leaderboards based on prompt qualities and by estimating model performance solely from autorater scores, entirely eliminating the need for additional human annotations.


Right Results, Wrong Reasons: Auditing Behavioral Reliance in Motion Forecasting

Geonyeong Park ⋅ Byounghun Park ⋅ Nayoung Kim ⋅ Kyungmin Kim ⋅ Soonmin Hwang

Motion forecasting models are mainly evaluated by trajectory-level accuracy metrics. However, these metrics do not directly evaluate which surrounding agents a model relies on for prediction—that is, agent-level reliance. We adopt leave-one-out (LOO) agent removal as a model-agnostic, output-based, linear-cost intervention. It measures how sensitive a model's output behavior is to removing a single agent. We formalize this as the LOO Reliance Audit protocol. We evaluate five architecturally distinct motion models on Argoverse 2 and the Waymo Open Motion Dataset, validate the protocol along four aspects—faithfulness, non-reducibility, stability, and convergence—and then derive three findings. Across the five architectures, agent-reliance rankings exhibit near-zero agreement (mean pairwise Spearman $\rho \approx 0.02$), consistently reproduced across both benchmarks (F1). LOO reliance is only partially aligned with human-annotated causal agents, revealing a collective blind spot: in 17.4% of scenes, all five models unanimously identify the same non-causal agent as top-1 (F2). In addition, within-model agreement between attention and LOO rankings is very low, suggesting that attention currently does not function as a reliable proxy for actual reliance (F3). Overall, current motion forecasting evaluation supports trajectory accuracy but does not support claims about what models rely on for prediction. We propose adding to existing benchmarks the two progress axes defined by the audit—how well a model's reliance aligns with human-annotated causal agents and how well attention and LOO rankings agree within the same model. Code and per-scene LOO scores are publicly released.


RipBench: A Unified Benchmark for Multi-Level Rip Current Detection, Classification and Segmentation

Andrei Dumitriu ⋅ Aakash Ralhan ⋅ Florin Miron ⋅ Florin Tatui ⋅ Radu Tudor Ionescu ⋅ Radu Timofte

Rip currents are a serious and often under-addressed threat to beach safety, and the leading cause of coastal drownings worldwide. They are difficult to detect due to their amorphous structure, similarity to the background, and large variability in viewpoints and environments. Despite their dangers and growing interest in automated detection, existing research remains fragmented across datasets, metrics, models and approaches, leading to an incomplete understanding of model performance. We introduce RipBench, a unified benchmark that enables controlled evaluation of rip current detection spanning multiple tasks, including classification, axis-aligned and oriented object detection, and instance, semantic, and panoptic segmentation, on the same data with standardized splits. This allows direct comparison of model performance on all levels of visual abstraction. Across $304$ videos ($303{,}491$ frames) collected from diverse coastlines, our results expose a clear performance gap between coarse recognition and precise spatial understanding. While models achieve near-saturated classification performance, accurate localization proves to be a substantially more challenging task, with performance varying across tasks. All tasks are supported with carefully curated annotations and evaluated using both standard and safety-critical metrics, with a focus on the $F_2$ score to emphasize recall in this safety-critical setting. RipBench, along with multiple baseline models per task, is publicly available at \url{https://RipBench.ai} to support progress in real-world, multi-task vision for beach safety.

Clustered Federated Learning (FL) partitions a client population into groups of similar local distributions and trains one specialized model per cluster, mitigating client drift that degrades single-model methods under non-IID data. Prior methods discover cluster structure inside the training loop through gradient similarity, loss evaluation, or EM-style updates, thus increasing communication overhead, exposing gradients to inversion attacks, and providing no mechanism to assign clients absent from training. We propose Ripple, a clustered FL framework in which cluster assignment is computed entirely offline from a spectral characterization of each client's local data: a variance-weighted principal-component prototype embedded via the Wavelet Scattering Transform and decoded by a Gaussian Mixture VAE trained server-side on synthetic client populations before federation begins. Per-round communication cost matches FedAvg exactly, and a client absent from training obtains a personalized model from a single forward pass, without gradient computation, model evaluation, or extra communication round. We prove that the gap between Ripple's surrogate clustered objective and the oracle is bounded by a computable quantity decaying with client sample size and independent of federation duration; per-cluster convergence matches the minimax-optimal rate for non-convex smooth objectives. Across five benchmarks spanning controlled and realistic heterogeneity, Ripple consistently outperforms all baselines, with margins growing on the most realistic partitions.

Predicting future clinical events from longitudinal electronic health records (EHRs) requires selecting plausible outcomes from a large and structured event space under sparse observations. While clinical coding systems provide hierarchical organization of events, cross-modal and temporal relationships are not explicitly specified and must instead be inferred from data, making prediction difficult for weakly observed longitudinal transitions. We introduce Risk Horizons, a geometry-aware framework for constructing patient-specific candidate spaces for multi-modal next-visit prediction. Risk Horizons combines deterministic coding hierarchies with data-driven lagged cross-modal associations, embeds the resulting clinical graph in hyperbolic space, and retrieves candidate futures using directional risk cones. This reframes longitudinal prediction as ranking within a compact, clinically coherent hypothesis space rather than scoring an unconstrained vocabulary. Experiments on MIMIC-IV and eICU demonstrate competitive next-visit prediction performance, with consistently improved hierarchy consistency across diagnoses, procedures, and medications. Further analysis suggests that hyperbolic structured candidate retrieval is the primary driver of performance, while LLMs are effective as constrained inference-time rerankers operating over clinically grounded candidate sets.

Goldwasser, Shafer, Vafa and Vaikuntanathan (STOC 2025) recently introduced a formal framework for defending against backdoors that may be planted in large-scale ML models by malicious model developers. They gave several algorithmic results in their framework for efficiently ``mitigating'' the effects of such backdoors by leveraging ideas that were developed in theoretical computer science in the 1980s, namely \emph{random self-reducibility} and \emph{self-correction}. However, the approaches of Goldwasser, Shafer, Vafa and Vaikuntanathan only provably achieve secure mitigation under restrictive assumptions about the ground-truth population data distribution that the ML model is trained on. In this work we apply tools that have been developed quite recently in the theoretical computer science research area known as \emph{tolerant property testing} to achieve secure backdoor mitigation for a much broader class of population distributions than could be handled by prior work. Our approach naturally provides a way for an ML model user to select a hypothesis class from a very wide range of possibilities for the mitigated ML model, and naturally enables the ML model user to control a tradeoff of the mitigated model's accuracy against its security, efficiency, and interpretability.


Robust Importance Sampling for Rare Events via Constrained Gaussian Mixtures

Pawel Lorek ⋅ Rafal Nowak ⋅ Rafał Topolnicki ⋅ Tomasz Trzcinski ⋅ Maciej Zieba

We study estimating rare-event probabilities $I = \mathbb{P}(g(\mathbf{X}) > \gamma)$ with $\mathbf{X} \sim \mathcal{N}(\mu, \Sigma)$ and general $g : \mathbb{R}^d \to \mathbb{R}$. We address this problem through importance sampling, and propose a framework that substantially improves efficiency and robustness over baselines such as crude Monte Carlo, adaptive cross-entropy, variational-inference-based methods (including forward- and reverse-KL approaches), as well as Safe-ICE, Subset Simulation, and Sequential Monte Carlo, *drawing* on ideas from both cross-entropy methods for rare-event estimation and cross-entropy methods for optimization. The key contribution has two parts: first, we separate the problem into *coverage*, to overcome the cold-start barrier, and *fitting*, to refine proposals once a meaningful signal is available; second, we constrain the proposal family in a way that provably guarantees finite-variance importance sampling, supported by a theoretical result (since coverage alone is not sufficient --- without safeguards, importance sampling may still suffer from infinite variance). Together, these ingredients yield proposals that are both expressive and stable. Extensive experiments demonstrate significant variance reduction, strong robustness across diverse benchmarks, and favorable cost--efficiency trade-offs, with the proposed approach often outperforming these baselines, particularly in high-dimensional and multimodal settings where competing methods frequently become unstable or fail.


Robust Statistical Estimators with Bounded Empirical Sensitivity

Valentio Iverson ⋅ Gautam Kamath ⋅ Argyris Mouzakis ⋅ Adam Smith

We introduce a new measure of robustness for statistical estimators, which we call empirical sensitivity. An estimator $\hat \mu$ has bounded empirical sensitivity if, with high probability over a dataset $X = (X_1, \dots, X_n) \sim \mathcal{D}^{\otimes n}$, for any dataset $Y$ obtained by modifying at most $\eta n$ points in $X$, we have that $\hat \mu(Y)$ is close to $\hat \mu(X)$. We study bounds on this quantity for the prototypical problem of Gaussian mean estimation. We prove new lower bounds, showing that for any estimator $\hat \mu$ which achieves an optimal $\ell_2$-error bound of $O(\sqrt{d/n})$, the empirical sensitivity is at least $\Omega(\eta + \sqrt{\eta d/n})$. The two terms arise due to obstructions on the mean and variance (via an Efron-Stein argument) of such an estimator. We show that this bound is tight up to logarithmic factors, by employing recent results for robust empirical mean estimation.

Data Streams are characterised as a potentially infinite source of data subject to non-stationarity. This non-stationarity, also called Concept Drift, often results in deployed predictive models having to be perpetually retrained as the underlying data generating process of the stream changes over time. Drift Detectors are a class of methods used to identify these changes, triggering these retraining procedures. Historically, Stream Learning has operated under the assumption that instances are only seen once, and cannot be stored for later use. Contemporary work relaxes this assumption yet drift detection and simple resetting procedures are still ubiquitous. Building on related work which has shown that traditional Batch Learning algorithms often outperform Stream Learning algorithms when instances can be stored, we introduce Time Neutralising Trees (TNT), a Decision Tree architecture that enables robust stream classification in settings with Concept Drift. During training, TNT will filter out ("neutralise") old training samples that are no longer relevant given the current context of the stream. If those instances become relevant again, TNT will then re-include them. We evaluate TNT, and related algorithms on 24 real and semi-real classification data streams from the USP DS repository. Results show strong evidence that TNT achieves state-of-the-art performance on a wide-range of data streams with varying concept drift types.


RoLL: Robust Low-Rank Learning via Nesterov Momentum

Zhaojun Hu ⋅ Wenchen Liu ⋅ Ting Wei ⋅ Biao Mei ⋅ Yifan Sun

In modern multivariate problems, exploiting the low-rank structure of the coefficient matrix serves as a powerful dimension reduction strategy to enhance statistical efficiency and interpretability, but its effectiveness can be severely compromised by outliers. In this work, we propose a robust low-rank learning (RoLL) framework for a broad range of loss functions, implemented via a fast algorithm that exploits local restricted strong convexity. Theoretically, we establish non-asymptotic error bounds for the obtained fixed-point estimator (not necessarily the global or local optimum), showing that it achieves the minimax optimal rate. Furthermore, we develop a novel information criterion with finite-sample theoretical guarantees for model selection. Extensive experiments demonstrate the effectiveness of our proposed framework in the presence of data anomalies.


Root Cause Analysis of Measurement and Mechanistic Anomalies

Hendrik Suhr ⋅ David Kaltenpoth ⋅ Jilles Vreeken

Root cause analysis of anomalies aims to identify how and why a sample deviates from the normal process. Existing methods primarily focus on telling which features are responsible, ignoring that anomalies can arise through two fundamentally different processes: measurement errors, where the sample is generated normally but one or more values is recorded incorrectly, and mechanism shifts, where the causal process that generated the sample was changed. While measurement errors can often be safely corrected, mechanistic anomalies require careful consideration. In this paper, we formally define a causal model that explicitly captures both types by treating outliers as latent interventions on latent (“true”) and observed (“measured”) variables and show under which conditions the distinction is possible. Based on this model, we develop an efficient inference procedure for localizing root causes and distinguishing anomaly types. Experiments on synthetic and real-world data show that our method provides state-of-the-art and highly robust performance in both root cause localization and classification of anomaly types.

Fixed-point inversion improves data-to-noise inversion by explicitly solving the local inverse equation at each timestep as a fixed-point problem. Empirically we observe that fixed-point inversion can reach distinct approximate roots, and the resulting inverse trajectories can differ substantially in how reconstruction errors accumulate. However, existing fixed-point inversion methods lack a principled mechanism for selecting among these solutions. Therefore, we propose SelFix, a root-selecting fixed-point inversion method for rectified flows that uses trajectory straightness as the selection criterion. We derive an on-the-fly proxy for rectified flow straightness from previously recovered inverse velocities and use it to construct a local straightness anchor. A vanishing anchored fixed-point update, combined with decoupled momentum for finite-iteration stability, then biases the iteration toward the fixed point that minimizes this selector while preserving the original inverse equation asymptotically. Under a standard local nonexpansiveness assumption, SelFix converges to the straightness-selected exact local inverse root. Experiments on FLUX.1-dev and PIE-Bench show that SelFix improves fixed-point inversion for rectified flows, achieving stronger real-image reconstruction and better source-preserving prompt-based editing than prior inversion baselines.


ROSE: Risk-Aware Orthogonal Subspace Navigation for Lifelong Knowledge Editing in Multimodal Large Language Models

Lingyun Song ⋅ Ziyao Chen ⋅ Kang Pan ⋅ Xiaolin Han ⋅ Yifei Zhang ⋅ Xiaoqi Wang ⋅ Yudai Pan ⋅ Xuequn Shang

Multimodal Large Language Models (MLLMs) require continuous factual updates to stay current, yet lifelong knowledge editing remains a daunting challenge. The primary obstacle is the severe Locality Erosion and Catastrophic Forgetting, where sequential edits inevitably interfere with pre-trained capabilities and prior updates. In this paper, we propose \textbf{ROSE}, a \textbf{R}isk-aware \textbf{O}rthogonal \textbf{S}ubspace navigation framework for lifelong knowledge \textbf{E}diting. ROSE addresses forgetting by restricting parameter updates to the orthogonal complement of previously edited subspaces. Crucially, it introduces a risk-guided mechanism that utilizes a dynamically computed risk mask to safeguard specific factual parameter subspaces essential for model stability. During inference, ROSE employs a multi-granularity knowledge integration strategy, featuring prototype-anchored semantic gating and perturbation-driven subspace fusion to precisely activate edited knowledge for relevant queries while strictly maintaining the integrity of the frozen backbone for unrelated inputs. Extensive evaluations on two major multimodal benchmarks demonstrate that ROSE consistently outperforms state-of-the-art methods across five key metrics, exhibiting exceptional robustness over 1,000 sequential updates. Our code is available in the supplementary material.


Rosetta: Composable Native Multimodal Pretraining

Xiangyue Liu ⋅ Zijian Zhang ⋅ Miles Yang ⋅ Zhao Zhong ⋅ Liefeng Bo ⋅ Ping Tan

Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. However, accommodating continuous generative objectives alongside discrete understanding tasks causes severe gradient conflicts. Existing architectures, including standard Mixture-of-Experts (MoE), are highly susceptible to representation overwriting. Even structurally partitioned paradigms like Mixture-of-Transformers (MoT) remain vulnerable to catastrophic forgetting, severely impeding multimodal scalability. In this work, we introduce Rosetta, a composable native multimodal pretraining framework designed for seamless and non-destructive modality expansion. Rosetta adopts a modular paradigm where core foundational knowledge is preserved within global shared experts, while modality-specific capabilities are distributed across plug-and-play experts. To guarantee non-destructive composition, we propose Momentum-Anchored Orthogonal Projection (MAOP). MAOP leverages the optimizer's momentum state as an implicit semantic anchor, selectively neutralizing conflicting gradient components from new modalities while preserving synergistic updates. To strictly isolate the architectural impact, we evaluate Rosetta against standard MoE and MoT baselines under strict active parameter parity. All models are trained from scratch within the Transfusion framework, using discrete next-token prediction for language and continuous visual diffusion. Extensive evaluations demonstrate that, while standard MoE and MoT architectures suffer catastrophic forgetting of previously acquired knowledge, Rosetta robustly preserves established language and visual understanding. Furthermore, it delivers superior image generation and unlocks cross-modal synergy, paving the way for truly composable and unified multimodal foundation models.


ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork

Caroline Wang ⋅ Muhammad Arrasy Rahman ⋅ Benjamin Nativi ⋅ Johnny Liu ⋅ Jiaxun Cui ⋅ Yoonchang Sung ⋅ Peter Stone

Learning to collaborate with previously unseen partners is a fundamental generalization challenge, known as Ad Hoc Teamwork (AHT). Existing methods often adopt a two-stage pipeline: first, a fixed population of teammates is generated, and second, an AHT agent is trained to collaborate with them. This separation limits coverage of behaviors and ignores whether the generated teammates are informative for the AHT agent to learn from. On the other hand, AHT agents are typically trained under the assumption that the training teammate set is uncontrollable, despite the fact that its composition strongly influences generalization. This paper presents a unified framework for AHT by reformulating the problem as an open-ended learning process between an AHT agent and an adversarial teammate generator. We introduce ROTATE, a regret-driven, open-ended training algorithm that alternates between improving the AHT agent and generating teammates that probe its collaboration deficiencies. Experiments across Overcooked and Level-Based Foraging tasks demonstrate that ROTATE significantly outperforms baselines on an unseen set of teammates, establishing a new standard for robust, generalizable teamwork.


Routeability Before Routing: Routeability Audit Protocol (RAP) and RouteabilityBench for Audited LLM Model Selection

Yihang Lu ⋅ Denica Kjorvezir ⋅ Ana Gjorgjevikj ⋅ Carola Doerr ⋅ Tome Eftimov

LLM routing aims to choose, for each prompt, which model in a portfolio is most likely to answer correctly. This only makes sense when the benchmark contains enough prompt-specific variation for different models to be best on different inputs. On coarse-scored LLM benchmarks, however, many models often tie for the best score, and the single best model may already leave little room for improvement. We therefore propose a routeability-first audit, instantiated as RAP and RouteabilityBench, a reviewer-runnable evaluation package. Across four benchmarks and 66 LLMs, generic routing enhancements such as larger embedding sets, dimensionality reduction, selective abstention, clustering, and stronger regressors do not consistently outperform the single best model. The clearest positive evidence is concentrated in portfolio-dependent Omni-MATH settings, while GPQA, MMLU-Pro, and IFEval show different ways in which routing can remain hard or predictor-sensitive. Our practical recommendation is that routing gains should be reported together with routeability diagnostics, tied-best analysis, portfolio construction details, and family-corrected statistical tests.

Medication recommendation requires generating drug combinations that are therapeutically effective while minimizing harmful drug-drug interactions (DDIs). Both objectives are rooted in the combinatorial nature of prescriptions. Existing discriminative methods predict drugs independently, neglecting inter-drug dependencies; autoregressive methods introduce sequential dependencies but impose arbitrary generation orders and accumulate errors. Both paradigms confine DDI mitigation to training-time penalties, offering limited capacity to regulate interactions during inference for individual prescriptions. We reformulate medication recommendation as a discrete diffusion process and propose RxDiff, which leverages iterative denoising for bidirectional and revisable generation and provides a natural interface for dynamic DDI regulation at each step. A clinically grounded dual-guided forward process and an inference-time interaction-aware modulation mechanism are designed to jointly enable RxDiff to achieve consistent accuracy improvements, lower and more individually controlled DDI rates, and robustness to evolving DDI knowledge across MIMIC-III, MIMIC-IV, and a real-world inpatient dataset. Code is available at https://anonymous.4open.science/r/RxDiff-9157.


S2MDF: A Plug-And-Play Layer for Intersection-Free Multi-Object Signed Distance Fields

Deniz Sayin Mercadier ⋅ Federico Stella ⋅ Aurel Bizeau ⋅ Nicolas Talabot ⋅ Pascal Fua

Compositional implicit surface representations model scenes as collections of objects, each encoded by a Signed Distance Field (SDF). A fundamental limitation of this approach is that multiple SDFs can produce geometries that interpenetrate, violating physical plausibility. Existing mitigation strategies rely on soft penalty terms that reduce but do not eliminate intersections, and require careful loss weighting. To truly prevent interpenetration, we propose a hard constraint on vector-valued SDFs and introduce S2MDF, a lightweight plug-and-play module that enforces the constraint on any object-compositional SDF representation without architectural modifications. It introduces \fs{negligible} computational overhead and is compatible with linearly-interpolated standard meshing algorithms such as Marching Cubes. It can be applied during training or as a post-processing step. Experiments on multiple state-of-the-art compositional methods show that S2MDF reduces intersections to numerical precision while preserving reconstruction quality, outperforming existing mitigation strategies.


S3Former: Sequential, Structural, and Statistical Fusion for Continuous-Time Dynamic Graphs

Ziyan Li ⋅ Xiangyu Hu ⋅ Tianqing Zhu ⋅ Wanlun Ma ⋅ Ke Qin

Continuous-time dynamic graphs model evolving systems with irregularly timed interactions, where predicting future links requires reasoning over temporal evolution, local topology, and historical relation patterns. Existing memory-based and sequence-based methods capture temporal dependencies effectively, but they often organize historical interactions as node memories or serialized neighbor sequences. As a result, they lack explicit analysis of the interaction structure around candidate nodes which prevents them from explicitly modeling local topological patterns of candidate nodes at inference time. They also lack explicit causal statistical memory for activity, recency, and repeated pairwise interactions. Motivated by these drawbacks, we propose S3Former, a temporal graph framework that jointly models sequential dynamics, shared pair-centered structure, and online statistical memory. For each candidate interaction, S3Former encodes long-term and short-term historical sequences, constructs a history-induced shared pair subgraph around the two endpoints, and computes node-level and pair-level statistics from past events only. A two-level context-aware fusion mechanism further combines these signals: NodeCCF learns reliable endpoint representations, while PairCCF performs candidate-specific relation correction using explicit pair memory. Experiments on 10 dynamic graph benchmarks show that S3Former achieves state-of-the-art performance on most datasets in both transductive and inductive link prediction. Ablation studies confirm the effectiveness of shared pair-subgraph modeling, online statistical memory, and hierarchical fusion.


SaaSBench: Exploring the Boundaries of Coding Agents in Long-Horizon Enterprise SaaS Engineering

Qingnan Ren ⋅ Shun Zou ⋅ Shiting Huang ⋅ Ziao Zhang ⋅ Kou Shi ⋅ Zhen Fang ⋅ Yiming Zhao ⋅ Yu Zeng ⋅ Qisheng Su ⋅ Lin Chen ⋅ Yong Wang ⋅ Zehui Chen ⋅ Xiangxiang Chu ⋅ Feng Zhao

As autonomous coding agents become capable of handling increasingly long-horizon tasks, they have gradually demonstrated the potential to complete end-to-end software development. Although existing benchmarks have recently evolved from localized code editing to from-scratch project generation, they remain confined to structurally simplified, single-stack applications. Consequently, they fail to capture the heterogeneous environments, full-stack orchestration, and system-level complexity of real enterprise Software as a Service (SaaS) systems, leaving a critical gap in assessing agents under realistic engineering constraints. To fill this gap, we introduce SaaSBench, the first benchmark designed to explore the boundaries of AI agents in enterprise SaaS engineering. Spanning 30 complex tasks across 6 SaaS domains with 5,370 validation nodes, it incorporates 8 programming languages, 6 databases, and 13 frameworks to meticulously mirror real-world software heterogeneity. Furthermore, we design a dependency-aware hybrid evaluation paradigm tailored for complex systems with long horizons and multi-component coupling, enabling fine-grained, reproducible assessment. Crucially, our extensive experiments reveal a striking insight: the primary bottleneck for state-of-the-art agents is not generating isolated code logic, but successfully configuring and integrating a multi-component system. Over 95\% of task failures occur before agents even reach deep business logic, with models often falling victim to overconfidence and prematurely halting during foundational system setup, or getting trapped in ineffective debugging loops. We hope SaaSBench serves as a practical and challenging testbed to drive the evolution of reliable, system-level coding agents.


SAFE-FEC: Semantically Constrained Adversarial Frontier Evolution for Factual Error Correction

Lei Zhu ⋅ Xiaobao Wang ⋅ Jianbiao Yang ⋅ Chenyang Wang ⋅ Longbiao Wang ⋅ Jianwu Dang

The difficulty of factual error correction (FEC) is not only how to repair false claims, but how to construct errors that meaningfully test repair. Existing FEC data construction methods often follow a static corruption view: they mask spans in false claims or inject errors into supported claims, but do not explicitly target the boundary of what current evidence-based correctors can repair. We introduce the notion of a \emph{repair frontier}: factual errors that are evidence-grounded, close enough to the source claim to admit a clear correction, and resistant to current correctors. We propose \textsc{SAFE-FEC}, a semantically constrained adversarial frontier-evolution framework for constructing such examples. Starting from evidence-supported claims, SAFE-FEC generates mutations from multiple factual-error perspectives, refines them through cross-perspective critique, selects or composes stronger candidates, filters semantic drift, and accepts only candidates that survive repeated corrector-in-the-loop repair attempts. This turns FEC data construction from one-shot error injection into correction-aware frontier search. Experiments on FECData, HoVer, and FEVEROUS show that SAFE-FEC consistently lowers correction performance for both LLM-based and FEC-specific correctors, especially in multi-hop and structured-evidence settings. Quality evaluation and ablations further suggest that these failures arise from plausible, semantically anchored, evidence-grounded errors rather than invalid hard negatives.


SAFE-Hair: Scalp-Anchored Fields for Exportable Single-View Hair Reconstruction

Yucheng Wang ⋅ Zedong Wang ⋅ Yuetong Wu ⋅ Yue Ma ⋅ Dan Xu

Single-view 3D hair reconstruction has progressed rapidly, but many methods optimize for visual plausibility rather than producing strand assets that can be directly used in grooming, editing, or simulation pipelines. We propose SAFE-Hair, a framework that reconstructs explicit strand grooms from a single portrait by decoding hair from persistent follicles in a canonical scalp-UV domain. Given a fitted template head, SAFE-Hair projects self-supervised image features onto the scalp, uses a conditional Rectified Flow model to generate latent scalp fields, and decodes occupancy, segment directions, length allocation, and total length at a fixed set of scalp follicles. This representation enforces scalp-rooted strand identities by construction and enables direct export with follicle IDs and guide-family metadata. We further introduce attachment, collision, and bending priors to improve geometric validity without per-subject test-time optimization. To evaluate both reconstruction fidelity and asset usability, we introduce HairBench-3D, a benchmark built from public 3D hair assets with standardized alignment, rendering, resampling, and metrics for geometry, direction consistency, invisible-region completion, silhouette agreement, and collision. SAFE-Hair improves Chamfer distance from 3.74 mm to 3.28 mm, F-score from 77.9 to 81.7, invisible-region F-score from 66.1 to 69.8, and penetration ratio from 4.8% to 3.9% over the strongest baseline.


Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

Zhengyang Tang ⋅ Yi Zhang ⋅ Chenxin Li ⋅ Xin Lai ⋅ Pengyuan Lyu ⋅ Yiduo Guo ⋅ Weinong Wang ⋅ Junyi Li ⋅ Yang Ding ⋅ Huawen Shen ⋅ Zhengyao Fang ⋅ Xingran Zhou ⋅ Liang Wu ⋅ Fei Tang ⋅ Sunqi Fan ⋅ Shangpin Peng ⋅ Zheng Ruan ⋅ anran zhang ⋅ Benyou Wang ⋅ Chengquan Zhang ⋅ Han Hu

When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may be avoided because the agent recognized the risk and chose the safe action, or because it failed to understand the screen or execute any relevant action at all. These cases have different causes and call for different fixes, yet current benchmarks often merge them under task success, refusal, or final harmful outcome. We address this problem with \textsc{PhoneSafety}, a benchmark of 700 safety-critical moments drawn from real phone interactions across more than 130 apps. Each instance isolates the next decision at a risky moment and asks a simple question: does the model take the safe action, take the unsafe action, or fail to do anything useful? We evaluate eight representative phone-use agents under this framework. Our results reveal two main patterns. First, stronger general phone-use ability does not reliably imply safer choices at risky moments. Models that perform better on ordinary app tasks are not always the ones that behave more safely when the next action matters. Second, failures to do anything useful behave like a capability signal rather than a safety signal: they are concentrated in more visually and operationally demanding settings and remain stable when the evaluation protocol changes. Across models, failures split into two recurring patterns: unsafe choices in settings where the model can act but chooses wrongly, and inability to act in more visually and operationally demanding screens. Overall, a harmless outcome is not enough to count as evidence of safety. Evaluating phone-use agents requires separating unsafe judgment from inability to act.


SALT: When More Rollouts Don’t Help in Group-Based Policy Optimization and How to Make Them Matter

Powei Chang ⋅ Jinpeng Zhang ⋅ Chaoqun Sun ⋅ MiniWell Tsao ⋅ Lianrui Li ⋅ JianXiang Xiang ⋅ Chenyu Wang ⋅ Yukang Gao ⋅ Dongying kong

Reinforcement learning with verifiable rewards (RLVR) often adopts GRPO-style group-relative updates, sampling multiple rollouts per prompt to construct normalized learning signals. However, merely increasing the number of rollouts does not reliably strengthen learning: under GRPO-style group normalization, per-rollout policy-gradient features can concentrate into a low-rank, signed geometry, causing substantial cancellation during aggregation and weakening the effective update. We address this failure mode with SALT, a Subspace-Adaptive geometry pLug-in componenT that uses sample-wise gradient geometry to reweight the coefficients of group-relative updates. SALT estimates a dominant shared subspace from the mini-batch Gram geometry, decomposes group-relative coefficients into shared and residual channels, and adaptively amplifies the residual channel when signed cancellation is severe. Across diverse reasoning-oriented RLVR benchmarks and model scales, SALT improves effective update geometry and performance without modifying the reward model or the rollout sampling procedure.


Salvation Lies Within: Proactive Prefix Re-forming for LLM-based Tagging

SHULAN WANG ⋅ YingJie Zhu ⋅ Yuting Yan ⋅ Ke Cheng ⋅ Yan Chen ⋅ Sheng Zhang ⋅ Hefei Guo ⋅ Qidong Zhang ⋅ Zhengyong Zhang ⋅ Yibo Jin ⋅ Zhuzhong Qian

Industrial personalized recommendation increasingly adopts LLMs for offline intent tagging (e.g., user interest from installed apps or browsing history), where each user corresponds to one prompt with thousands tokens. While shared prefix among prompts can improve the inference efficiency, the reusable part is actually dwarfed by user-specific records. We observe that even a slight $\textit{discordance}$ in preceding records prevents two prompts from sharing a longer prefix, revealing an opportunity to proactively re-form the common part from within, by adjusting the internal orders, with ensured accuracy for intent tagging. Realizing this at industrial scale is challenging: (1) the data preparation and the inference operate separately; (2) re-forming must complete within minutes for massive prompts under strict SLO; and (3) user data evolves continuously, requiring incremental updates. We present $\textit{PrefixShakeup}$, a system that enables $\textit{proactive prefix re-forming}$ before actual inference, intrinsically enlarging the common part from the source. PrefixShakeup, bridges the data and inference by enhancing the datasets during its preparation, thereby increasing the reuse. It combines three techniques: (1) a prefix model that tolerates a bounded number of records, instead of requiring strict prefix matching; (2) an alleviator that reorders both records within each prompt and the prompt positions to maximize the reuse under limited HBM, upon the kernel-based acceleration for scoring the similarity; and (3) an adaptor that supports incremental changes via lightweight management and necessary cache over time, to avoid full re-computation. We implement PrefixShakeup, on Ascend NPUs and evaluate it in the mirror environment. Compared to state-of-the-art inference, PrefixShakeup, improves throughput by up to 2.1$\times$, with controlled tagging accuracy.


Sampling Is Not Curiosity: Why LLM Agents Should Investigate

Alfonso Amayuelas ⋅ Piotr Piękos ⋅ Xin Wang ⋅ William Yang Wang ⋅ Jürgen Schmidhuber ⋅ Alex Pentland

LLM agents are increasingly used for open-ended tasks, very common in research scenarios where they are required to explore, investigate, or act curiously. At inference time, LLMs sample from a distribution that is fixed by its pre-trained weights and the provided context. Mechanisms like temperature, best-of-N, tree search, or RAG vary how the distribution is drawn, but do not update the agent's model about the environment. We argue that current LLM agents do not explore curiously; they explore stochastically. From how curiosity is understood in the RL literature, it requires a predictive structure over how the environment responds to the agent's actions, which improves online from current interaction, drives action selection through its predicted improvement, and persists in structured form across the episode. Additionally, in its open-ended form, the agent should intrinsically originate what to investigate. Thus, we propose a framework defining curiosity in LLM agents, reserving the term for systems that satisfy it. In doing so, we identify the research directions its implementation opens up.


SCALECUA : Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL

Bowen Lv ⋅ Xiao Liu ⋅ Yanyu Ren ⋅ Hanyu Lai ⋅ Bohao Jing ⋅ Hanchen Zhang ⋅ Yanxiao Zhao ⋅ 顺天 姚 ⋅ Jie Tang ⋅ Yuxiao Dong

Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI execution. Online reinforcement learning with verifiable rewards (RLVR) has emerged as a key direction for scaling their capabilities. However, applying this paradigm is severely bottlenecked by verifiable data scarcity and online RL inefficiency. To break these barriers, we introduce ScaleCUA, a unified framework that scales online RL for CUAs by combining verifiable task synthesis with efficient online RL. At the data level, we design VeriGen, an end-to-end framework for generating verifiable RL tasks through iterative docker interactions and a multi-agent feedback loop. Scaled to 100+ concurrent agent workers via a shared docker interaction probe, this pipeline produces 24K+ verifiable tasks and nearly 3K high-quality RL tasks for online RL. To maximize sample efficiency over this scaled pool, we propose Frontier Sampling, which dynamically tracks the model's per-task capability and allocates rollouts to tasks at the current learning frontier. On the training side, we further design Visual Context Segmentation, which keeps the recent visual context within a bounded window to balance rollout and training engine pressure on long-horizon trajectories, yielding a 2.83× training speedup over step-wise decomposition. Together, ScaleCUA achieves 68.7% on OSWorld and 54.0% on ScienceBoard, establishing new state-of-the-art performance among open-source computer use agents. Source code:


Scaling Laws for Multimodal Data Mixtures

Aditi Khandelwal ⋅ Ayush Kumar Tarun ⋅ Yixuan Xu ⋅ Imanol Schlag ⋅ Golnoosh Farnadi ⋅ Siva Reddy ⋅ Luke Zettlemoyer ⋅ Antoine Bosselut

Frontier AI systems are increasingly natively multimodal, jointly pretrained on multiple modalities, such as text, vision, and audio. However, the scientific understanding of the multimodal data mixtures used to pretrain these systems --- how much of each modality to mix, and at what potential interference cost to the others --- is largely absent or gatekept. In this work, we derive data mixing scaling laws for three modalities: text, vision, and audio. Multimodal Large Language Models (MLLMs) increasingly adopt Mixture-of-Experts (MoE) architectures, as MoEs scale efficiently and partially limit cross-modal interference. We therefore use MoEs to conduct 268 experiments ranging from 457M to 8.3B parameters, on up to 150 billion tokens. Our work leads to multiple actionable insights into multimodal data mixing for training MLLMs. First, compute optima are modality-specific: audio modality is model-heavy, favouring scaling parameters over tokens, while text and vision require balanced scaling of model size and data. Second, repeating vision or audio data beyond 2x yields negligible benefits. Finally, cross-modal interactions are highly asymmetric: vision strongly interferes with audio, while audio provides mild positive transfer to text and vision despite competing for shared capacity. Overall, we provide a principled foundation for understanding the multimodal data mixtures needed to train frontier MLLMs.


Scaling Optimization-Oriented Hypernetworks for Implicit Neural Representations

Lulu Cai ⋅ Hanxiang Ren ⋅ Siyan Dong ⋅ Li Sun ⋅ Jiefeng Wu ⋅ Youyi Zheng ⋅ Yi Ma ⋅ Yanchao Yang

We present a hypernetwork that generates the parameters of an implicit neural representation (INR) for a given signal in a single forward pass, without test-time optimization. Existing hypernetworks for this problem either ignore the chain-rule structure of the test-time optimization that produces an INR or, when they preserve it, instantiate the structure with MLPs whose parameter count grows quartically in target-network width. Our central contribution is to show that this parameter-count cost is structural to the MLP instantiation, not to the optimization-oriented principle of preserving chain-rule structure, and to resolve it by replacing the MLP body with attention. The resulting hypernetwork combines three design choices: an output-neuron-centric tokenization that aligns the token axis with each layer's output-neuron axis; an information-flow consistency rule that fixes the role of every cross-attention module from the representation its output should occupy; and a token-space iterative rollout that realizes the five quantities one optimization step on the target INR computes as five cross-attention modules with parameter count independent of target-network width. Across five 2D and 3D INR generation benchmarks, our method outperforms the state-of-the-art single-pass and optimization-oriented baselines; scales to target-network widths at which the MLP-based optimization-oriented paradigm exhausts GPU memory; and produces the highest result on NeRF generation without ground-truth INR weights supervision.


Scaling Point-in-Time Language Models: Economic Evaluation of Embeddings

Bryan Kelly ⋅ Semyon Malamud ⋅ Johannes Schwab ⋅ Teng A Xu

Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences. Point-in-time language models—trained exclusively on text available up to each calendar date—eliminate this leakage by construction, but existing efforts typically produce models that lag substantially behind their unconstrained counterparts. We show that this performance gap can be narrowed through scale. Training decoder-only transformers with up to 4 billion parameters on 1 trillion chronologically filtered tokens from FineWeb, we construct a sequence of monthly model checkpoints spanning 2013–2024. Across a range of common-sense reasoning and language understanding benchmarks, our models approach the performance of leading open-weight models of comparable size (such as Gemma-3-4B and LLaMA-7B) trained on temporally unrestricted data, although a performance gap remains on several tasks. Finally, in a strict out-of-sample economic evaluation task, portfolios built from point-in-time embeddings achieve robust positive Sharpe ratios and perform close to full-sample counterparts that violate temporal validity, indicating that chronologically consistent language models can extract economically meaningful signals without relying on look-ahead bias. We release the complete pipeline—including dataset construction, training infrastructure, and evaluation code—to enable reproducible point-in-time language modeling and to support research applications that require strict temporal validity.


SciResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

Lei Xiong ⋅ Kun Luo ⋅ Ziyi Xia ⋅ Wenbo Zhang ⋅ Jingying Shao ⋅ Jianlyu Chen ⋅ Hongjin Qian ⋅ Zhicheng Dou ⋅ Zheng Liu

Autonomous scientific research is significantly advanced thanks to the development of AI agents. One key step in this process is finding the right scientific literature, whether to explore existing knowledge for a research problem, or to acquire evidence for verifying assumptions and supporting claims. To assess AI agents’ capability in driving this process, we present SciResearchBench (Scientific - Research - Bench), a dedicated benchmark for autonomous scientific literature discovery. SciResearchBench consists of two complementary task types: (1) Deep Research, which requires tracking down a specific target paper through a progressive, multi-step probing process, and (2) Wide Research, which requires comprehensively collecting a set of papers satisfying given conditions. Compared to previous benchmarks on agentic web browsing, SciResearchBench is distinguished along three dimensions: it is research-oriented, calling for in-depth comprehension of scientific concepts; literature-focused, demanding fine-grained utilization of detailed information; and open-ended, involving an unknown number of qualified papers and thus requiring deliberate reasoning and search throughout. These properties make SciResearchBench uniquely suited for evaluating autonomous research capabilities, and extraordinarily challenging. Even the most powerful LLMs, despite having largely conquered general agentic web-browsing benchmarks such as BrowseComp, achieve only 9.39% accuracy on Deep Research and 9.31% IoU on Wide Research, while many other strong baselines fall below 5%.


Scratchpad Patching: Decoupling Compute from Patch Size in Byte-Level Language Models

Lin Zheng ⋅ Vasilisa Bashlovkina ⋅ Timothy Dozat ⋅ Dan Garrette ⋅ Laura Rimell ⋅ Joshua Maynez

Tokenizer-free language models eliminate the tokenizer step of the language modeling pipeline by operating directly on bytes; patch-based variants further aggregate contiguous byte spans into patches for efficiency. However, the average patch size chosen at the model design stage governs a tight trade-off: larger patches reduce compute and KV-cache footprint, but degrade modeling quality. We trace this trade-off to patch lag: until a patch is fully observed, byte predictions within it must rely on a stale representation from the previous patch to preserve causality; this lag widens as patches grow larger. We introduce Scratchpad Patching (SP), which inserts transient scratchpads inside each patch to aggregate the bytes seen so far and refresh patch-level context for subsequent predictions. SP triggers scratchpads using next-byte prediction entropy, selectively allocating compute to information-dense regions and enabling post-hoc adjustment of inference-time compute. Across experiments on natural language and code, SP improves model quality at the same patch size; for example, even at $16$ bytes per patch, SP-augmented models match or closely approach the byte-level baseline on downstream evaluations while using a $16\times$ smaller KV cache over patches and $3$–$4\times$ less inference compute.


SCTI: Self-Calibrated Trident Identification of Black-Box LLM Watermarks

Zhixiong Nan ⋅ Haoyu Lu ⋅ Tao Xiang ⋅ Yiwei Wang

Black-box watermarked LLM identification has become an important task for watermark auditing. The core challenge arises from inherent black-box constraints that deny access to logits, detector keys, model parameters, and internal watermark settings. For this reason, existing methods have not conducted in-depth exploration on this task, leaving prominent limitations: Fixed-Reference Calibration and Family-Specific Tests. First, existing works use fixed null reference to serve as a baseline for judging the presence of watermarks in LLM outputs, which might misidentify a watermarked LLM as unwatermarked because the null reference is sensitive to the prompt pair, queried model, and sampling configuration. Second, existing statistical-tests based methods adopt distinct statistical features and judgment criteria for different LLM watermark families, making them less suitable when the deployed watermark algorithm is unknown. To remedy the above two limitations, this work proposes SCTI, a Self-Calibrated Trident Identification framework for black-box watermarked LLM identification. Specifically, to handle the first limitation, SCTI constructs empirical null distributions from the queried responses to estimate the generalizable reference instead of relying on a fixed reference. Meanwhile, to address the second limitation, SCTI is configured with a Trident-View Consistency measuring mechanism, which enables a unified pipeline to avoid designing separate tests for individual watermark family. Experimental results verify that SCTI outperforms representative black-box watermark identification baselines across diverse LLMs and watermark algorithms.

Conventional real-valued neural networks struggle to explicitly capture phase-frequency coupling, while long-horizon time series forecasting requires models to characterize non-stationary amplitude variations, phase shifts, and frequency-dependent temporal evolution. This mismatch limits the stability and expressiveness of neural models on complex forecasting tasks. To address this challenge, we propose SDHilb, a structure-preserving forecasting framework that lifts real-valued observations into an adaptive complex state space and models temporal evolution as a process jointly governed by amplitude and phase information. Through the Hilbert transform, phase-related structures are explicitly encoded, leading to a coupled evolution of real and quadrature components that motivates Schrödinger-Hamiltonian structured dynamics in the lifted complex space. Based on this formulation, SDHilb integrates Schrödinger-guided linear dynamics with an adaptive multi-scale Hilbert transform and introduces three key components: (1) a structure-preserving linear dynamic module with symplectic discretization, (2) an adaptive multi-scale Hilbert module for constructing data-dependent quadrature components and refined time-frequency decomposition, and (3) a theoretically grounded parameter reuse mechanism derived from the closed-form recurrence structure. Extensive experiments demonstrate that SDHilb achieves competitive or superior performance on both short- and long-term forecasting tasks while reducing horizon-specific parameterization through recurrence-guided parameter reuse.


SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation

JungHun Oh ⋅ Sungyong Baik ⋅ Kyoung Mu Lee

Low-Rank Adaptation (LoRA) enables efficient adaptation of large pre-trained models to downstream tasks by parameterizing weight updates with low-rank matrices. In this paper, we investigate the limitations of the LoRA parameterization from a geometric perspective. Specifically, we show that when a full fine-tuning gradient is backpropagated to the low-rank matrices, it undergoes anisotropic scaling driven by their singular values. We argue that this phenomenon is undesirable because it distorts the full fine-tuning gradient by skewing it toward dominant singular directions while suppressing others. Our analyses demonstrate that anisotropic gradient scaling reduces the effective rank of the low-rank matrices' gradients and results in suboptimal alignment between the full fine-tuning gradient and its low-rank approximation in LoRA, thereby exacerbating the gap to full fine-tuning. To address these limitations, we propose a new low-rank parameterization, SDS-LoRA, which structurally decouples singular values from the backward pass. Our method ensures that the full fine-tuning gradient backpropagates only through the orthonormal bases of the low-rank matrices' subspaces, independent of their scales. Convergence analysis demonstrates that while LoRA’s convergence rate degrades with the condition number of the low-rank matrices, SDS-LoRA remains independent of it. Experimental results across natural language and vision benchmarks show that SDS-LoRA improves loss convergence and reduces the gap to full fine-tuning, significantly enhancing adaptation performance.


SearchV: Evolutionary Fine-Grained Visual-Token Skipping for Efficient Vision-Language Models

Xinrui Chen ⋅ Zhen Huang ⋅ Shuwei Li ⋅ Fanyi Zeng ⋅ Dongxu Yue ⋅ Hao Yu ⋅ Huaisong Zhang ⋅ Xitong Ling ⋅ Yongxian Wei ⋅ Feng Lu ⋅ Chun Yuan

Large vision--language models (VLMs) incur substantial computational costs due to massive parameter counts and the deep propagation of long visual sequences. While recent studies highlight visual token redundancy, existing methods often rely on local heuristics or coarse skipping policies that treat Transformer layers as isolated units, thereby limiting the optimization landscape. We propose SearchV, a framework that redefines VLM acceleration as a discrete policy search problem guided by a holistic fidelity objective. Central to our approach is Global Contribution (GC) fitness, a metric that captures end-to-end performance drift induced by complete skipping configurations rather than isolated components. By leveraging a fine-grained search space and evolutionary mutation that decouple attention and feed-forward pathways, SearchV identifies surgical interventions that preserve multi-benchmark robustness more effectively. Experimental results demonstrate that SearchV defines a new efficiency frontier for VLMs. Notably, SearchV identifies optimized 13B-based policies that surpass the vanilla 7B model in accuracy while operating at a lower computational budget. This represents a training-free milestone that effectively breaks the performance ceiling of conventional model scaling, proving that global structural optimization can recover high-capacity knowledge under heavy compression. Code is available in supplements.

Machine learning relies on gradient-based training procedures whose empirical efficiency is usually analyzed through finite-dimensional arithmetic counts rather than formal complexity over continuous real-valued data. This paper studies neural-network training primitives as represented-real operators in the second-order complexity framework of Kawamura and Cook. Here, second-order polynomial time means computing a real-valued operator to accuracy $2^{-n}$ in time polynomial in the requested precision, architecture size, primitive evaluation costs, and supplied analytic certificates. We prove that dense backpropagation, convolutional backpropagation, soft attention, fixed finite SGD/Adam-style updates, and related smooth local primitives are second-order polynomial-time computable over the relevant represented real spaces. More precisely, dense one-sample backpropagation has tight arithmetic complexity $\Theta(s)$ for $s$ weights and polynomial bit complexity; convolutional backpropagation is polynomial in the number of active convolution incidences; dense soft-attention backpropagation has arithmetic cost $\Theta(\ell d^2+\ell^2d)$; and fixed finite smooth optimizer updates remain second-order polynomial-time computable under explicit smoothness and lower-bound certificates. Higher complexity enters through three distinct mechanisms: discontinuous selection, global optimization, and continuous-time flow solving. Hard routing, top-$k$ sparsification, argmax decisions, ReLU kink conventions, exact line search, and idealized gradient flow therefore change the second-order complexity status of the training operator. Overall, this work provides a rigorous mathematical language for precision, conditioning, and certificate dependence, which are often implicit in floating-point analyses, thus enabling a unifying complexity account for dense networks, CNNs, attention, Transformers, and standard optimizers.


Seeing Both the Forest and the Trees: Reusing Holistic 3D Priors for Part-Decomposed Generation

Yiqun Zhao ⋅ Binbin Huang ⋅ Haobin Duan ⋅ Zibo Zhao ⋅ Shenghua Gao

Generating 3D assets with explicit parts requires two properties at once: geometric detail inside each part and structural coherence across the whole object. Recent image-conditioned 3D generation models adopt a \emph{structured 3D representation} that anchors tokens to explicit spatial positions, capturing both properties, but their output remains a single fused mesh. We ask how a part-decomposed generation model can inherit both properties of these holistic priors, and answer with STRUCT-PARTS, a framework whose core technical contribution is a dual-frame coordinate interleaving mechanism: each shape latent is interpreted under both a part-local and an object-global reference frame, and processed through a global-local interleaved transformer that simultaneously reuses the prior's geometric detail at the part scale and its structural coherence at the whole-object scale. As its structural input, STRUCT-PARTS uses a segmented mesh scaffold, a coarse mesh paired with a face-level segmentation mask kept entirely within the prior's native 3D representation. With only lightweight finetuning, both intra-part detail and inter-part coherence emerge from the same pretrained weights. On PartObjaverse-Tiny, STRUCT-PARTS matches or surpasses prior part-decomposed models at a fraction of their training cost, while naturally supporting part-mesh refinement and part-level editing.


Seeing the World through Any Eyes

Yang Fu ⋅ Jianqin Wang ⋅ Xiangtai Li ⋅ Henghui Ding

Egocentric video generation aims to synthesize first-person visual experiences, enabling applications in filmmaking, virtual reality, and embodied AI. Generating egocentric videos from exocentric observations is particularly challenging, as it requires reasoning across large viewpoint changes, limited visual overlap, and substantially different camera motions. To address these modeling challenges, we propose EgoEye, a framework that generates egocentric videos from a single exocentric input video. EgoEye integrates reward-guided egocentric context reasoning for inferring unseen first-person content, multi-stage motion alignment for enforcing cross-view temporal consistency, and egocentric pretraining for improving first-person realism. To support training and evaluation under diverse and distortion-free exo-ego settings, we further introduce EgoScape, a large-scale dataset comprising 19.5K time-aligned exo-ego video pairs and 24.2K in-the-wild egocentric videos. EgoScape provides paired cross-view supervision while enriching the first-person visual priors needed for realistic egocentric synthesis. Extensive experiments demonstrate that the proposed EgoEye generates realistic and temporally consistent egocentric videos across diverse scenarios.

Medical visual question answering is a safety-critical decision problem in which abstaining on inputs the image cannot support matters as much as answer accuracy. Existing inference-time agentic methods address hallucination by adding a verification step, but they collapse the final decision at a single language-model judge that integrates all evidence in one prompt and is systematically biased toward declaring inputs answerable. We propose PICV (Parallel Independent Claim Verification), which reformulates selective medical VQA as independent claim verification with transparent aggregation: the question is decomposed into a small set of typed visual claims, each is verified by an isolated prompt with no access to the other verdicts, and the verdicts are combined by a deterministic rule rather than another model, yielding an answer/abstain decision together with a typed unanswerability reason. We evaluate PICV efficiently on an automated answerability benchmark we construct over two public radiology datasets via three reproducible corruption procedures, requiring neither heavy VLM judging nor human annotation at evaluation time. Across four VLM backbones, PICV consistently outperforms strong prompting and agentic baselines; ablations attribute the gains specifically to per-claim isolation and rule-based aggregation rather than to agentic decomposition alone.


Selective Safety Steering via Value-Filtered Decoding

Bat-Sheva Einbinder ⋅ Hen Davidov ⋅ Yee Whye Teh ⋅ Yarin Gal ⋅ Yaniv Romano

While large language models (LLMs) are trained to align with human values, their generations may still violate safety constraints. A growing line of work addresses this problem by modifying the model’s sampling policy at decoding time using a safety reward. However, existing decoding-time steering methods often intervene unnecessarily, modifying generations that would have been safe under the base model. Such unnecessary interventions are undesirable, as they can distort key properties of the base model such as helpfulness, fluency, style, and coherence. We propose a new test-time steering method designed to reduce such unnecessary interventions while improving the safety of unsafe responses. Our approach filters tokens using a value-based safety criterion and provides an explicit bound on the probability of false interventions. A single threshold hyperparameter controls this bound, allowing practitioners to trade off higher rates of unnecessary intervention for better output safety. Across multiple datasets and experiments, we show that our value-filtered decoding method outperforms existing baselines, achieving better trade-offs between safety, helpfulness, and similarity to the base model.


Self-Cleaning Diffusion Models

Adrian Rodriguez-Munoz ⋅ Adam Klivans ⋅ Antonio Torralba ⋅ Constantinos Daskalakis ⋅ Giannis Daras

We present Self-Cleaning Diffusion Models, a principled and effective framework for training diffusion models from heterogeneous data. In many domains, high-quality data is scarce or expensive to collect, while low-quality and out-of-distribution (OOD) data is abundant. Leveraging these ubiquitous samples is critical for scaling generative models. Still, principled techniques remain elusive, leaving practitioners to rely on simple heuristics such as high-quality finetuning or explicit quality conditioning. Self-Cleaning Diffusion Models bridges this gap by decoupling data correction from prior learning. We first train a transport map between the abundant OOD and the limited in-distribution samples. We then train an unconditional generative prior on a mixture of the few in-distribution samples and noisy versions of the transported points. Crucially, this noise injection prevents learning errors in the transport map from propagating to the generative prior. Theoretically, we demonstrate that transporting OOD samples towards the empirical proxy reduces the distance to the underlying distribution under mild assumptions. Empirically, we achieve state-of-the-art results across a diverse range of settings, from mitigating synthetic corruptions to increasing diversity and transcending dataset quality.


Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group Discovery

Louis Berthier ⋅ Ahmed Shokry ⋅ Maxime Moreaud ⋅ Guillaume Ramelet ⋅ Aymeric Dieuleveut

Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional undercoverage in safety-critical subgroups. We introduce Self-Organized Conformal Prediction (SOCP), a calibration scheme that discovers input-space groups with a Self-Organizing Map (SOM) and, at test time, draws a local calibration buffer from the query's best-matching unit (BMU) cell or a fixed grid neighborhood. The same retrieval rule applies to regression and classification tasks across tabular features and image embeddings, leaving the predictor and nonconformity score untouched. SOCP gives exact validity for BMU-cell retrieval and fixed retrieved-set validity for neighborhood buffers; central-cell validity for neighborhood retrieval holds up to a Kolmogorov-Smirnov (KS) bias term. A split-routed extension recovers fixed retrieved-set validity conditional on the routing split. On eight regression and classification benchmarks, SO-SCP reduces the weighted regional coverage gap on $7/8$ datasets (mean paired change $-7.1\%$) for a mean prediction-set size increase of $6.2\%$, with negligible overhead on the largest six datasets; SO-CQR yields smaller gains, since quantile regression already absorbs much of the heterogeneity. By learning groups directly from the input geometry, SOCP provides group-local calibration with exact fixed-group guarantees and approximate central-cell guarantees, without supervised partitions or predictor retraining.


Self-Supervised Reconstruction Knockoffs for Calibrated Unsupervised Feature Selection

HaiHui Huang ⋅ Dingkui Kang ⋅ Yanan Zhou ⋅ Yong Liang

Unsupervised feature selection is widely used as a discovery tool, yet most methods return only absolute rankings: a selected gene, pixel, or sensor is never compared against a feature-wise null. We introduce Self-Supervised Reconstruction Knockoffs (SSRK), a calibrated proxy-discovery framework for unlabeled data. SSRK defines relevance through masked reconstruction: a feature is deemed useful only when replacing it with a matched knockoff degrades reconstruction of other masked coordinates. The method trains a symmetric knockoff-gated masked autoencoder and converts the learned gates into a slot-aware signed statistic that satisfies the knockoff sign-flip property under Model-X exchangeability and coupled implementation. We formalize the estimation target as knockoff-relative masked information, prove immunity to self-copy and marginal-variance artifacts, derive an entropy-regularized population fixed point for correlated gates, and establish finite-sample margin and defect-robust false discovery rate bounds. Oracle synthetic experiments validate knockoff+ control at q = 0.10 with full support recovery in the reported regimes. On Peripheral Blood Mononuclear Cells, MNIST, Fashion-MNIST, and UCI Human Activity Recognition, where learned knockoffs are approximate, the same statistic is evaluated as a ranking and recovers biologically, spatially, and sensor-structurally meaningful subsets. SSRK thereby turns masked self-supervision into a feature-wise null comparison while clearly separating controlled discoveries from exploratory rankings.


Semantic-Bridge Federated Learning: Bridging CLIP Semantics for Heterogeneous FL

Leixuhai Xu ⋅ Xiangtao Zhang ⋅ Hailong Yan ⋅ Le Zhang

Federated learning (FL) suffers from severe performance degradation under statistical heterogeneity, where label skew and domain shift induce representation drift and biased decision boundaries. While vision--language models such as CLIP provide transferable semantic priors, raw CLIP embeddings are noisy, highly correlated, and not directly aligned with the latent space of lightweight federated classifiers. We propose $\textbf{Semantic-Bridge Federated Learning}$ (SBFL), a framework that transfers pre-trained vision--language semantics into heterogeneous FL without accessing raw client data or requiring online CLIP inference. SBFL constructs text-verified semantic banks, learns a server-side semantic bridge that maps CLIP-derived features into student-compatible teacher spaces, and regularizes local training through mixed semantic prototype alignment and quality-aware semantic replay. Experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet, Office-Caltech 10, and Digit5 show that SBFL consistently improves the average performance of representative FL backbones under both label-skewed and domain-shifted settings. Analysis further shows that the semantic bridge converts highly correlated CLIP priors into more separable teacher-space geometry, explaining its effectiveness in improving representation alignment and global generalization. Our code will be released upon publication.

In the computer vision world, vision-language and self-supervised encoders represent two unique paradigms characterized by global language alignment and high-quality vision-only patch representations, respectively. Multimodal large language models (MLLMs) often use the former because of language alignment, despite the recent rise in vision-centric tasks resembling traditional dense vision tasks that would likely benefit from the high-quality patch representations of the latter. This raises the question whether self-supervised models could aid MLLMs in vision-centric tasks. To investigate this, we define the concept of semantic (in)consistency of patch tokens produced by language-aligned encoders and identify two sets of tokens, semantically consistent (SCTs) and semantically inconsistent (SITs). Through systematic probing via the removal of SCTs or SITs from the input of modern MLLMs, we show higher reliance on SCTs than SITs, as removal of the former induces larger performance drops. This demonstrates that MLLMs benefit from semantic consistency of vision tokens. With this in mind, we devise a simple yet effective extension of the MLLM training objective that combines traditional language modeling with semantic guidance via minimization of semantic inconsistency of vision tokens. This strategy, named Semantically-Guided Training (SGT), combines the signals of both loss terms to push the vision encoder towards providing semantically consistent tokens, while preserving language alignment. Through extensive evaluation, we show that SGT drastically outperforms all baselines and MLLMs with specialized vision encoders on all vision-centric benchmarks while keeping strong scores on text-centric benchmarks, thus setting state of the art when averaged across all tasks. Our results highlight the importance of combining language alignment and semantic consistency through the lens of MLLMs rather than independently from them and pave the way towards the design of better MLLM vision encoders.

We introduce SemGeo-Gen, an unsupervised method for generating approximate cross-instance semantic-geometric supervision from weakly aligned 3D object collections. Our goal is not to recover exact point-level correspondences, which are often ambiguous across different instances and would require expensive, impractical manual supervision at scale. Instead, we automatically produce approximate correspondences that are semantically meaningful, geometrically consistent, and diverse enough to train modern correspondence models. Given only coarse category-level rotation alignment, SemGeo-Gen lifts multi-view DINOv2 features to 3D, aligns object instances using continuous piecewise-affine registration, and prunes candidate matches using semantic consistency. The resulting 3D correspondences can be projected into rendered views, yielding scalable 3D-3D, 3D-2D, and 2D-2D supervision without manual annotation. We validate the generated correspondences against sparse human annotations from KeypointNet, obtaining 93% PCK@0.10 and 97% PCK@0.15. More importantly, we show that using the generated correspondences for synthetic pretraining improves recent state-of-the-art semantic correspondence models on SPair-71k after fine-tuning on real data, with gains of up to 5 PCK@0.10 points. Together, these results indicate that approximate semantic-geometric supervision generated from 3D assets can improve real-image correspondence learning and serve as a scalable alternative to costly manual annotation. Code and generated data will be released upon acceptance.


SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

Wenhao Li ⋅ Zhibin Wu ⋅ Chong Xiao ⋅ Qiangchang Wang

Recent research on Multimodal Sentiment Analysis (MSA) has focused on learning from language, visual, and acoustic modalities with incomplete data to infer human sentiment. Most studies typically compensate for missing information by reconstructing modality features or designing complicated fusion mechanisms. However, these methods still suffer from spurious generation and noisy guidance due to the lack of high-level semantic grounding in partially observed multimodal evidence. To address these issues, we propose SemMSA, a latent semantic-aided framework that constructs rich sentiment-relevant semantics with LLMs, fully integrating with all modalities via anchor-free spectral alignment. It mainly consists of Cross-modal Semantic Refinement (CSR) and Cross-modal Spectral Alignment (CSA). Specifically, CSR first adaptively extracts visual and acoustic representations by corresponding adapters to form a unified multimodal prefix with language in the frozen LLM embedding space. It then iteratively produces continuous discriminative semantic states through a token-efficient latent refinement process without decoding explicit text. Next, CSA simultaneously aligns the refined semantics with all modalities by enhancing the dominant spectral component of their kernel Gram matrix. This captures global nonlinear dependencies among all representations without relying on a predefined anchor modality. In addition, an instance-level spectral separation constraint preserves cross-sample discriminability and mitigates representation collapse. Extensive experiments on SIMS, MOSI, and MOSEI benchmarks demonstrate that SemMSA achieves state-of-the-art performance.

Vision-Language Models (VLMs) show great potential for Industrial Anomaly Detection (IAD) but often overlook subtle high-frequency defects due to their semantic bias. Additionally, one-class adaptation risks rank collapse and catastrophic forgetting. To address these issues, we propose a unified framework termed SF-DST. First, we design Spatial-Frequency Dual-Stream Transformer, which explicitly captures lost details via a parallel discrete cosine transform branch. Crucially, it incorporates Asymmetric Cross-Modal Modulation (ACMM), a mechanism that leverages spatial semantics as a context gate to selectively retrieve and inject spectral features, effectively compensating for the texture blindness of VLMs. Second, Anomaly-Aware LoRA (A-LoRA) enforces Orthogonal Regularization on the adaptation subspace. This geometric constraint ensures subspace diversity to prevent mode collapse, while optimizing a hypersphere-based metric for precise anomaly discrimination. Extensive experiments on MVTec-AD, VisA and MMAD benchmarks demonstrate that our method significantly outperforms state-of-the-art approaches, successfully bridging the gap between large-model generalization and industrial-grade precision.


SGD in Multiclass Logistic Regression: Sequential Learning and Scaling Laws

Konstantinos Tsiolis ⋅ Denny Wu ⋅ Christos Thrampoulidis ⋅ Murat Erdogdu

We study the training dynamics of multiclass logistic regression on high-dimensional Gaussian mixture models with a large number of classes and establish precise scaling laws governing the cross-entropy risk under gradient-based optimization. We show that learning proceeds sequentially across classes, from most to least frequent. When the class priors follow a power law distribution, the risk dynamics decompose into three phases: an initial plateau until the first class is learned, a power-law decay regime during which sequential learning occurs, and a final convergence regime. We then analyze how model capacity interacts with optimization under a fixed compute budget. When the effective dimension is restricted via projection onto leading principal components, the risk decomposes into a capacity term (a power law in the retained dimension) and an optimization term (a power law in training time). Optimizing this tradeoff yields a compute-optimal scaling law for logistic regression, with explicit prescriptions for model size and training time as functions of compute. These results extend theoretical scaling laws from linear regression to multiclass classification, while connecting to empirical scaling laws observed in large-scale neural networks.


ShadowFPT: Backdooring Federated Prompt Tuning via Shadow Triggers

Kun Zhai ⋅ Teng Li ⋅ Yunhao Feng ⋅ Xingjun Ma

Federated Prompt Tuning (FPT) adapts large vision--language models by freezing the pretrained backbone and optimizing only lightweight prompt parameters across clients. Although this design improves communication and parameter efficiency, it creates an overlooked security risk. Since the backbone is shared and frozen, a malicious client can induce backdoor behavior through the representation space while keeping its uploaded prompt updates close to benign ones. We propose \textbf{ShadowFPT}, a targeted backdoor attack that exploits this frozen-backbone attack surface. ShadowFPT first pretrains a learnable \emph{Shadow Trigger} against the frozen CLIP visual encoder, using either auxiliary public data or the malicious client's local data, so that triggered inputs are steered toward the target class in representation space. During federated prompt tuning, the malicious client adapts the trigger under the current global prompt and then optimizes its local prompt on both clean and triggered samples. Only prompt parameters are uploaded to the server, while the trigger and frozen encoders remain local. By shifting most of the attack burden from prompt updates to trigger-induced representation steering, ShadowFPT achieves targeted misclassification while preserving prompt-space stealthiness. Across multiple datasets, aggregation rules, and non-IID federated partitions, ShadowFPT increases the attack success rate from 19.81\% to 90.36\% in our main setting, while maintaining clean accuracy. It remains effective across textual, visual, and joint vision--language prompt tuning. These results identify frozen backbones as stealthy and underexplored backdoor surfaces in federated prompt tuning, suggesting that defenses based only on prompt-update anomaly detection are insufficient.

Hybrid modeling, the combination of machine learning models and scientific mathematical models, enables flexible and robust data-driven prediction with partial interpretability. However, the unknown parameters of the scientific model cannot necessarily be estimated properly, since the flexibility of the machine learning model might make the scientific model part effectively ignored in prediction. We may avoid it by applying some regularization, but the formulation of such regularizers typically depends on model architectures and domain knowledge. In this paper, we propose an architecture-agnostic method to learn hybrid models while properly estimating the scientific parameters. The idea is to use the flatness of loss minima to achieve model simplicity, based upon the Occam's razor principle. We employ the idea of sharpness-aware minimization and adapt it to the hybrid modeling setting. Numerical experiments demonstrate the effectiveness of the SAM-based hybrid model learning for scientific parameter estimation.


ShopGym: An Integrated Framework for Realistic Simulation and Scalable Benchmarking of E-Commerce Web Agents

Yuanzheng Zhu ⋅ Mingyu Zhao ⋅ Chinmay Savadikar ⋅ Han Li ⋅ Shuang Xie ⋅ Alberto Castelo ⋅ Tianfu Wu ⋅ Lingyun Wang

Developing and evaluating e-commerce web agents requires environments that preserve meaningful task structure while enabling controllable, reproducible, and scalable scientific comparison. Existing methodologies force a tradeoff: live storefronts provide realism but are non-stationary, difficult to inspect, and irreproducible, while hand-built sandbox benchmarks provide control but cover only a narrow range of layouts, catalogs, policies, and interaction patterns. We argue that the core bottleneck is methodological: the field lacks a scalable way to construct evaluation settings that are simultaneously realistic, diverse, controllable, inspectable, and reproducible. We introduce ShopGym, an integrated framework for realistic simulation and scalable benchmarking of e-commerce web agents. ShopGym is a framework for constructing e-commerce simulation environments and grounded benchmark tasks. Its simulation layer, ShopArena, converts one or more live seed storefronts into self-contained sandbox shops through anonymized shop specifications and a staged, validated generation process. On top of these simulated storefronts, ShopGuru synthesizes benchmark tasks across seven skill categories, grounding each task in the shop's catalog, navigation structure, policies, and interaction affordances. Together, ShopArena and ShopGuru produce self-contained, resettable, inspectable, and stable evaluation artifacts that preserve structural properties and agent-evaluation signals relevant to shopping tasks. We validate the framework through graph-based structural analysis and agent-based behavioral evaluation with 224 generated tasks across six sandbox shops: three constructed with synthetic data and three with real data. Our results show that the synthetic shops preserve key structural properties of live storefronts, with agent performance on synthetic shops positively correlated with performance on live storefronts.


SIGMA: A Sigmoid-Gated Sampler for Test-Time Scaling in Diffusion Language Models

Ziwen Zhang ⋅ Weiyu Chen ⋅ Yuyan Zhou ⋅ Yichen Zhu ⋅ James Kwok

Autoregressive language models generate under a fixed left-to-right order, whereas diffusion language models (dLLMs) denoise masked tokens through flexible any-order trajectories that may commit multiple positions per step. This makes dLLMs a distinctive setting for test-time scaling: additional inference compute can diversify both token choices and the order in which positions are committed. Existing samplers do not explicitly allocate these two sources of diversity. Confidence-based decoding is reliable but often redundant, while global temperature sampling and random remasking inject stochasticity without distinguishing useful exploration from unreliable commitments. We propose SIGMA, a training-free sampler that formulates each denoising step as a constrained allocation problem over decoding efficiency, token-level diversity, trajectory-level diversity, and model-internal commitment uncertainty. The resulting objective decomposes joint sampling entropy into token and trajectory terms, yielding a sigmoid-form stochastic gate for position selection and an adaptive per-position temperature for token sampling. Experiments on LLaDA and Dream across math and code benchmarks show that SIGMA improves the accuracy--cost Pareto frontier and can be integrated into existing dLLM test-time scaling pipelines.


SignalBench: Comparing Dense Feedback Methods for Long-Horizon Agents

Sergio Hernández-Gutiérrez ⋅ Matteo Merler ⋅ Ilze Amanda Auzina ⋅ Joschka Strüber ⋅ Ameya Prabhu ⋅ Matthias Bethge

Dense supervision is an essential ingredient when scaling long-horizon agent training, yet we don't really know how well current methods recover the value structure they purport to estimate. A growing line of work -- intrinsic signals, self-distillation, embedding similarities -- can be unified as dense feedback methods that score intermediate states and actions. However, prior works evaluate dense supervision by measuring downstream model performance, an approach that is computationally expensive and entangles the quality of the signal with engineering choices for training. To address this gap, we propose SignalBench: a computationally cheap, training-free testbed designed to isolate and evaluate dense supervision methods for agentic systems. SignalBench allows us to comprehensively analyse the correlation of the evaluated signals with reference-policy value labels across 21 dense feedback methods, with over 1.2K evaluations across four diverse environments and six open-weight backbones. We find that simple prompting baselines consistently outperform state-of-the-art dense feedback methods, and methodological families cluster together in performance. These core findings are surprisingly robust: They hold across model sizes and families, environments, observation modality, and estimating Q-value or state-value. Overall, SignalBench offers a valuable and inexpensive test bed for dense feedback method evaluation across model and task families.


Sign-Based Optimizers Are Effective Under Heavy-Tailed Noise

Dingzhi Yu ⋅ Hongyi Tao ⋅ Yuanyu Wan ⋅ Lijun Zhang

While adaptive gradient methods are the workhorse of modern machine learning, sign-based optimization algorithms such as Lion and Muon have recently demonstrated superior empirical performance over AdamW in training large language models (LLM). However, a theoretical understanding of why sign-based updates outperform variance-adapted methods remains elusive. In this paper, we aim to bridge the gap between theory and practice through the lens of heavy-tailed gradient noise, a phenomenon frequently observed in language modeling tasks. Theoretically, we introduce a novel generalized heavy-tailed noise condition that captures the behavior of LLMs more accurately than standard finite variance assumptions. Under this noise model, we establish sharp convergence rates of SignSGD and Lion for generalized smooth function classes, matching or surpassing previous best-known bounds. Furthermore, we extend our analysis to Muon and Muonlight, providing what is, to our knowledge, the first rigorous analysis of matrix optimization under heavy-tailed stochasticity. These results offer a strong theoretical justification for the empirical superiority of sign-based optimizers, showcasing that they are naturally suited to handle the noisy gradients associated with heavy tails. Empirically, LLM pretraining experiments validate our theoretical insights and confirm that our proposed noise models are well-aligned with practice.

Modern LLM workflows increasingly move coordinate-indexed objects across checkpoints: steering vectors, sparse autoencoders, top-$k$ neuron sets, attribution lists, and merge alignments. This is only well posed after fixing the model's residual-stream gauge. We show that the native discrete gauge is architecture-dependent: LayerNorm residual charts have permutation gauge $S_d$, up to a global sign flip, while RMSNorm residual charts with generic per-channel gain have signed-permutation gauge $B_d = S_d \ltimes \{\pm 1\}^d$. Thus permutation-only alignment is symmetry-incomplete for RMSNorm models. We introduce sign-marginalized Hungarian matching and prove a sharp population failure mode: with decorrelated source coordinates, raw signed-correlation matching has a structural permutation-accuracy ceiling equal to the fraction of positive signs in the true gauge up to $O(d \cdot 2^{-d})$, whereas sign-marginalized matching removes this obstruction. We then make coordinate-preserving transport, rather than function-level merging, the primary object: composing saved-checkpoint local $B_d$ gauges along same-base fine-tuning trajectories recovers 91.1% of cross-run coordinates at 1500 steps versus 60.3% for endpoint matching, and the gain is not explained by merely routing through the base. The recovered gauge transfers tools that permutation-only alignment breaks: TinyLlama SAE reconstruction has NMSE $0.004$ under $B_d$ recovery versus $1.08$ under $S_d$; Qwen sentiment steering preserves 95.8% of its effect versus 17.2%; refusal steering reverses sign under $S_d$. Coordinate-preserving merge tests show the same mechanism. The same covariance governs stateful training: signed transport of AdamW state preserves the resumed trajectory, while permutation-only state transport starts from a functionally identical checkpoint but follows a different trajectory. Finally, we give gauge-sweep audits for index-level interpretability claims: coordinate names are reproducible only relative to an explicit gauge.


SIMBAD: Spatio-Temporal Traffic Forecasting Robust to Aperiodicity

Daniel Yoonhwan Lee ⋅ Seungwon Shin ⋅ Seunghoon Han ⋅ Sungsu Lim ⋅ Susik Yoon

Accurate traffic prediction is essential for urban planning but remains challenging due to the irregular patterns caused by unexpected events. While Spatio-Temporal Graph Neural Networks (STGNNs) effectively model periodic data, they often struggle with these aperiodic fluctuations. To address this, we propose SIMBAD, which improves generalization on irregular patterns via two key mechanisms: 1) adaptive regulation of periodic signals to prioritize recent signals, and 2) dynamic spatial modeling that adjusts influence between connected nodes based on contextual relations. Built upon simple similarity metrics to adapt to aperiodic patterns, SIMBAD outperforms existing methods on real-world benchmark datasets, demonstrating superior performance in predicting unseen and irregular traffic events.

Contextual recommendation is a variant of contextual linear bandits in which the learner observes an (optimal) action rather than a reward scalar. Recently, Sakaue et al. (2025) developed an efficient Online Newton Step (ONS) approach with an $O(d\log T)$ regret bound, where $d$ is the dimension of the action space and $T$ is the time horizon. In this paper, we present a simple algorithm that is more efficient than the ONS-based method while achieving the same regret guarantee. Our core idea is to exploit the improperness inherent in contextual recommendation, leading to an update rule akin to the second-order perceptron from online classification. This removes the Mahalanobis projection step required by ONS, which is often a major computational bottleneck. More importantly, the same algorithm remains robust to possibly suboptimal action feedback, whereas the prior ONS-based method required running multiple ONS learners with different learning rates for this extension. We describe how our method works in general Hilbert spaces (e.g., via kernelization), where eliminating Mahalanobis projections becomes even more beneficial.

Large language models (LLMs) can generate student-like responses fluently enough to serve as simulated students, i.e., virtual learners used to train and evaluate AI tutors and human educators. Yet such simulators are typically evaluated by output similarity to real students, not by whether they behave like students with coherent misconceptions during interaction. We introduce a controlled framework for evaluating misconception faithfulness, whether a simulator maintains a misconception-driven belief state and updates selectively when feedback addresses the underlying misconception. Central to our framework is a misconception-contrastive feedback protocol that compare targeted feedback against two controls: misaligned feedback, targeting a different but plausible misconception, and generic feedback, which only signals that the answer is incorrect. We propose Selective Flip Score (SFS), which quantifies how much more often a simulator flips its answer under targeted feedback than under the contrastive controls. Across seven LLMs (4B–120B), multiple datasets, and prompting strategies, simulators exhibit near-zero SFS, correcting their answers at similarly high rates regardless of feedback relevance to the true misconception. Further analysis reveals a sycophantic failure mode: models behave less like students with stable misconceptions and more like problem solvers who treat any corrective signal as a cue to abandon the simulated misconception and recompute from internal knowledge. To improve faithfulness, we develop a post-training pipeline spanning supervised fine-tuning (SFT), preference optimization, and reinforcement learning (RL) with an SFS-aligned reward. SFT yields notable gains up to +0.56; SFS-aligned RL provides more consistent further improvements than preference optimization. Our results establish misconception faithfulness as a challenging but trainable property of student simulators, motivating a shift from static output matching toward interaction- and belief-aware student modeling.


Simultaneous Individual, Group and Multigroup Fairness in Set Covering Problems

Sharmila Duppala ⋅ Nathaniel Grammel ⋅ Tyler He ⋅ Aravind Srinivasan

Covering problems are fundamental combinatorial optimization problems with broad connections to clustering, facility location, and various applications such as machine learning and controlling disease outbreaks. In this work we study a variant of the set cover problem that generalizes the partition set cover problem~\citep{bera2014approximation}. In partition set cover, given an instance $(P, S)$ with a universe $P = \cup_{c \in [\ell]} P_c$ (where $[\ell]$ denotes $\{1, 2, \ldots, \ell\}$), a collection of subsets \(S\) (each with associated costs), and coverage requirements $k_1,k_2,\dots,k_\ell$, the objective is to find a subcollection with minimum cost that covers at least $k_c$ elements from each group $c \in [\ell]$. We generalize this problem by allowing for the groups $P = \cup_{c \in [\ell]} P_c$ to be overlapping thus capturing both \emph{group fairness} and \emph{multigroup fairness}. We further allow for probabilistic coverage requirements for each element in the ground set thus capturing the notion of \emph{individual fairness}. Our main result is a randomized iterated-rounding-based algorithm with provable guarantees. Our main result is a randomized iterated rounding based algorithm with provable guarantees. We supplement our theoretical results by showing that algorithm has significantly running times and outperforms the best known algorithm of \citep{inamdar18PartitionSetCover}.


SimVLA: Attributing Gains in VLA Models Through Controlled Ablation

Yuankai Luo ⋅ Woping Chen ⋅ Tong Liang ⋅ Zhenguo Li

Vision-Language-Action (VLA) models often bundle architectural changes with different pretraining data, backbone scales, and optimization recipes, obscuring what actually drives progress. We introduce SimVLA, a deliberately minimal VLA—a standard vision-language backbone with a lightweight continuous-action head—as a controlled testbed for this attribution problem. Across LIBERO, CALVIN, WidowX, and Google Robot, one-knob-at-a-time ablations show that training dynamics are the largest measured driver, producing performance swings of $\Delta$$\approx$54-89\%, while task configuration and architecture have smaller effects under matched settings. Despite using a small backbone and no additional large-scale robot-data VLA pretraining before benchmark-specific training, SimVLA reaches 98.6\% average success on LIBERO, remains competitive across the other benchmarks, and transfers to held-out real-robot scenes. These results suggest that simple, carefully controlled VLA baselines remain underexplored and provide a calibration point for future architectural claims within the VLM-encoder plus lightweight-head family.


Sink vs. diagonal patterns as mechanisms for attention switch and oversmoothing prevention

Peter Súkeník ⋅ Cristina Lopez Amado ⋅ Christoph Lampert ⋅ Marco Mondelli

This paper studies the role of sinks and diagonal patterns as attention switch and anti-oversmoothing mechanisms. We analyze geometric conditions under which sinks can be represented, showing a necessary alignment between the embedding of the sink and all other embeddings. Next, we refine the current understanding of the role of sinks in oversmoothing prevention: we specify the conditions under which dense attention provably smooths more than sparse attention, and empirically verify that such conditions are often satisfied in practice. We further prove an equivalence between sinks and hard attention switch, in which the output of the attention is identically 0. Finally, we relax the hard attention switch by allowing token self-communication: we provide a quantitative comparison of the costs of representing sinks vs.\ diagonal patterns, showing why sinks are favored in pretrained transformers. The introduction and analysis of diagonal patterns and the generalization of the attention switch close the gap between what oversmoothing prevention requires and what sinks provide, while also establishing when and why attention layers act like MLPs if token communication is not necessary.


SitCom: Scaling Egocentric Multi-Party Spoken Dialogue for Situated Communication Assistance

Heeseung Yun ⋅ Sehun Lee ⋅ Sang Hoon Woo ⋅ Yoonji Nam ⋅ Sung-Feng Huang ⋅ Chao-Han H Yang ⋅ Gunhee Kim

Multi-party spoken communication pervades daily life, yet even attentive participants routinely lose track of who said what, miss a turn, or struggle to recall something said minutes earlier. As always-on wearables become commonplace, an assistant that listens alongside the user could ease this everyday friction. We introduce \textit{Situated Communication Assistance}: on-demand support for users immersed in ongoing multi-party conversations, grounded in the same egocentric acoustic scenes they perceive. We argue this capability rests on three abilities: producing a Rich Situated Transcription (RST) of \textit{who} said \textit{what}, \textit{when}, and \textit{where}; comprehending the conversation holistically; and answering user queries while the conversation is unfolding. To address the persistent data deficit in this regime, we build \textbf{SitCom}, a multi-party spoken dialogue corpus more than an order of magnitude larger than the largest prior multi-party corpus, pairing 15.8k hours of synthesized data with 144 hours of real recordings standardized under a unified RST schema. The synthesis pipeline generates group-structured scripts with persona-grounded speech and behaviors, and renders them in 3D scenes through a directional simulator with head-and-torso radiation. Using this corpus, we curate \textbf{SitCom-Bench}, a fully human-validated question answering benchmark of 4.3k items spanning seven comprehension and four in-the-moment query subtypes. Experiments show that combining synthetic and real data is necessary to close the sim-to-real gap, and that in-the-moment querying remains hard even for frontier models, pointing to it as an open challenge distinct from post-hoc comprehension.


SkillOpt: Executive Strategy for Self-Evolving Agent Skills

Yifan Yang ⋅ Ziyang Gong ⋅ weiquan Huang ⋅ Qihao Yang ⋅ Ziwei Zhou ⋅ Zisu Huang ⋅ Yan Li ⋅ Xuemei Gao ⋅ Qi Dai ⋅ Bei Liu ⋅ Kai Qiu ⋅ Yuqing Yang ⋅ Dongdong Chen ⋅ Xue Yang ⋅ Chong Luo

Adapting language agents often means teaching them domain procedures: where to look, which tools to call, how to verify intermediate results, and how to format outputs for a grader. Existing context-level adaptation usually either hand-writes such procedures, generates them once, or lets skill artifacts grow through loosely controlled self-revision. We ask whether a skill can instead be trained as the compact external state of a frozen agent. SkillOpt is a harness-agnostic text-space optimizer for a single natural-language skill document. It runs the frozen target model on scored rollout batches, uses a separate optimizer model to reflect over success and failure minibatches, proposes structured add/delete/replace edits, ranks them under a textual learning-rate budget, and accepts a candidate only if it improves held-out selection performance. Rejected-update memory, slow cross-epoch guidance, and optimizer-side meta skill make the editing loop behave like bounded training while leaving deployment as a single best_skill.md. Across six direct-chat benchmarks covering QA, spreadsheets, documents, math, and embodied decision making, \ourmethod{} improves GPT--5.5 over no skill by 21.5 points on average, with positive gains on every task and the best measured result on five of six. Codex- and Claude-Code-style harness runs show that the same artifact remains useful in tool-backed execution. Ablations and transfer studies identify validation-gated, bounded edits and update memory as key to stable, portable skill learning.


Skill-SD: Skill-Conditioned Self-Distillation for Multi-turn LLM Agents

Hao Wang ⋅ Guozhi Wang ⋅ Han Xiao ⋅ YufengZhou ⋅ Yue Pan ⋅ Jichao Wang ⋅ Ke Xu ⋅ yafei wen ⋅ XIAOHU RUAN ⋅ xiaoxin chen ⋅ Honggang Qi

Reinforcement learning (RL) has been widely used to train LLM agents for multi-turn interactive tasks, but its sample efficiency is severely limited by sparse rewards and long horizons. On-policy self-distillation (OPSD) alleviates this by providing dense token-level supervision from a privileged teacher that has access to ground-truth answers. However, such fixed privileged information cannot capture the diverse valid strategies in agent tasks, and naively combining OPSD with RL often leads to training collapse. To address these limitations, we introduce Skill-SD, a framework that turns the agent's own trajectories into dynamic training-only supervision. Completed trajectories are summarized into compact natural language skills that describe successful behaviors, mistakes, and workflows. These skills serve as dynamic privileged information conditioning only the teacher, while the student always acts under the plain task prompt and learns to internalize the guidance through distillation. To stabilize the training, we derive an importance-weighted reverse-KL loss to provide gradient-correct token-level distillation, and dynamically synchronize the teacher with the improving student. Experimental results on agentic benchmarks demonstrate that Skill-SD substantially outperforms the standard RL baseline, improving both vanilla GRPO (+14.0%/+10.9% on AppWorld/Sokoban) and vanilla OPD (+42.1%/+40.6%).


SMART: Scalable Multi-Agent Role-conditioned Teaming via LLM-free Tree Search

Runzhe Zhang ⋅ Ning Li ⋅ Xingshan Zeng ⋅ Letian Chen ⋅ Yasheng Wang ⋅ Weinan Zhang ⋅ Yong Yu ⋅ Weiwen Liu

Large language model (LLM)-powered multi-agent systems have been shown to effectively solve real-world end-to-end problems through collaboration among diverse agents. However, as LLM capabilities advance and user demands continue to grow, the range of problems that agents can address expands dramatically, while the population of available agents and tools continues to increase, giving rise to open and evolving ecosystems of autonomous agents. In this paradigm, agents are likely to become increasingly specialized, each focusing on particular capabilities or problem types. In such a setting, a central challenge is how to efficiently assemble an appropriate team of agents to take on different responsibilities from a massive agent pool for a given task. In this work, we study this problem in a controlled setting with 300 task-oriented agents. SMART (Scalable Multi-Agent Role-conditioned Teaming) first plans a task-conditioned collaboration topology and candidate role set, then constructs a compact role-conditioned candidate pool, and finally uses a LLM-free Monte Carlo Tree Search (MCTS) to assign agents to roles online via metadata-based rollouts. This design enables high flexibility without incurring prohibitive inference costs. Across multiple experiments, our approach achieves superior performance on GSM8K (93.64\%), MATH (61.73\%), HumanEval (93.89\%), and MBPP (87.68\%) using gpt-4o-mini under a limited budget.


SMoA: Spectrum Modulation Adapter for Parameter-Efficient Fine-Tuning

Yongkang Liu ⋅ Xing Li ⋅ Mengjie Zhao ⋅ Shanru Zhang ⋅ Zijing Wang ⋅ Qian Li ⋅ Shi Feng ⋅ Feiliang Ren ⋅ Daling Wang ⋅ Hinrich Schuetze

As the number of model parameters increases, parameter-efficient fine-tuning (PEFT) has become the go-to choice for tailoring pre-trained large language models. Low-rank Adaptation (LoRA) uses a low-rank update method to simulate full parameter fine-tuning, which is widely used to reduce resource requirements. However, decreasing the rank encounters challenges with limited representational capacity. Theory suggests that LoRA fine-tuning with rank $r$ converges toward the top $r$ singular values of the pre-trained weight matrix. As the rank increases, more principal singular directions are preserved, which generally improves the model’s performance. However, a larger rank also introduces more trainable parameters, leading to higher computational cost. To overcome this dilemma, we propose SMoA, a \textbf{S}pectrum \textbf{Mo}dulation \textbf{A}dapter that enlarges the accessible family of spectrum-aware updates under a smaller parameter budget. SMoA partitions the layer into multiple aligned spectral blocks and applies one in-block Hadamard-modulated low-rank branch to each diagonal block, yielding broader coverage of pretrained spectral directions. We provide theoretical analysis and empirical results on multiple tasks. In our experiments, SMoA improves average performance in the current lower-budget setting over LoRA and competitive LoRA-style baselines. Our repository is on https://anonymous.4open.science/r/SMoA-904F/.


Smooth Partial Lotteries for Stable Randomized Selection

Alexander Goldberg ⋅ Giulia Fanti ⋅ Nihar Shah

Competitive selection processes, from scientific funding to admissions and hiring, use evaluations to score candidates, and then choose a subset based on those scores. Recently, many organizations have adopted partial lotteries, which randomize selection based on evaluation scores. However, existing lottery designs are inherently unstable, as a small change to a single candidate's score can cause large shifts in their selection probabilities. This instability undermines a key goal of lotteries: reducing the influence of fine-grained score distinctions near the decision boundary. We propose smoothness as a design principle for partial lotteries, formalizing it as a Lipschitz condition on the mapping from review scores over candidates to selection probabilities. We propose using a Linear Lottery, a simple mechanism in which selection probabilities scale linearly with aggregated review scores between an upper threshold, above which we always accept and a lower threshold, below which we always reject. We prove that the Linear Lottery's worst-case regret matches a lower bound for any smooth selection rule up to a factor of $(1 - k/n)$, where $k/n$ is the acceptance rate. We compare smooth selection to other stability notions like Individual Fairness and Differential Privacy, showing that the Linear Lottery achieves a better smoothness--regret tradeoff than alternatives. Experiments on real peer review data from ICLR 2025, NeurIPS 2024, and the Swiss NSF confirm our theoretical analysis is tight and demonstrate that existing lottery designs are highly unstable in practice even under changes to a single review score.


SnapMLA: Efficient Long-Context MLA Decoding via Hardware-Aware FP8 Quantized Pipelining

Yifan Zhang ⋅ Zunhai Su ⋅ Shuhao Hu ⋅ YangRui ⋅ Wei Wu ⋅ Yulei Qian ⋅ Yuchen Xie ⋅ Xunliang Cai

As the context length of large language models (LLMs) grows, the key-value (KV) cache memory overhead becomes a critical bottleneck, limiting long-context inference efficiency. While FP8 attention has shown substantial promise in innovations like FlashAttention-3, its integration into the decoding phase of the DeepSeek Multi-head Latent Attention (MLA) architecture presents notable challenges. These challenges include numerical heterogeneity arising from the decoupling of positional embeddings, misalignment of quantization scales in FP8 PV GEMM, and the need for optimized system-level support. In this paper, we introduce SnapMLA, an FP8 MLA decoding framework optimized to improve long-context efficiency through the following hardware-aware algorithm-kernel co-optimization techniques: (i) RoPE-Aware Per-Token KV Quantization: Motivated by our analysis of the heterogeneous quantization sensitivity inherent to the MLA KV cache, this approach preserves the RoPE part in high precision. Furthermore, per-token granularity is employed to align with the autoregressive decoding process and maintain quantization accuracy. (ii) Pre-posed Scale Domain Alignment, where we skillfully project RoPE components into the quantization domain to mitigate pipeline stalls caused by mixed-precision accumulation; (ii) Quantized PV Computation Pipeline Reconstruction: Addresses the misalignment of quantization scales in FP8 PV computation caused by the shared KV structure of the MLA. (iii) End-to-End Dataflow Optimization: Establishes an efficient data read-and-write workflow using specialized kernels, ensuring streamlined data flow and improved performance. Extensive experiments on state-of-the-art MLA LLMs show that SnapMLA achieves up to a 1.91× improvement in throughput on long-output decoding workloads while maintaining near-parity benchmark quality compared with the BF16 baseline on the evaluated reasoning and code-generation benchmarks. The code corresponding to this work is included in the supplementary material.


Sobolev Regularized MMD Gradient Flow

Chenyang Tian ⋅ Bharath Sriperumbudur ⋅ Arthur Gretton ⋅ Zonghao Chen

We propose Sobolev-regularized Maximum Mean Discrepancy (SrMMD) gradient flow, a regularized variant of maximum mean discrepancy (MMD) gradient flow based on a gradient penalty on the witness function. The proposed regularization overcomes the non-convexity of the MMD objective and yields provable \emph{global} convergence guarantees in MMD in both continuous and discrete time. A more surprising appeal is that our convergence analysis does not rely on isoperimetric assumptions on the target distribution. Instead, it is based on a regularity condition on the difference between kernel mean embeddings. A key highlight of the proposed flow is that it is applicable in both sampling (from an unnormalized target distribution)---using Stein kernels---and generative modeling settings, unlike previous works, where a gradient flow is suitable for only generative modeling or sampling but not both. The effectiveness of the proposed flow is empirically verified on a broad range of tasks in both generative modelling and sampling.


SODA: Selective Optimization with Deferred BN Alignment for Efficient Dataset Distillation

Xinyue Bi ⋅ Jiacheng Cui ⋅ Yaxin Luo ⋅ Xinyi Shang ⋅ Jiacheng Liu ⋅ Xiaohan Zhao ⋅ Zhiqiang Shen

Recent AI research has increasingly evolved along two complementary directions: model-centric learning, which improves architectures and training algorithms, and data-centric learning, which improves the quality and compression of training data. Within data-centric learning, dataset distillation on large-scale datasets has attracted growing attention, with decoupled distillation methods like SRe$^2$L, G-VBSM, LPQLD, FADRM have emerged as a representative paradigm. These methods typically synthesize compact datasets by matching the Batch Norm statistics of a teacher network, using BN alignment as an effective training objective. However, modern networks often contain many BN layers, and enforcing alignment at every layer introduces substantial computational redundancy. More importantly, we observe that not all BN layers provide equally informative supervision or consistent performance gains. Motivated by this, we propose SODA, a novel Selective Optimization with Deferred BN Alignment framework for efficient dataset distillation. SODA selectively identifies and optimizes informative BN alignment objectives while deferring unnecessary or low-value alignment computations, substantially improving both distillation efficiency and synthetic data quality. We further provide a theoretical analysis explaining why selective and deferred BN alignment can simultaneously reduce optimization cost and improve generalization. Extensive experiments across multiple datasets of CIFAR-100, Tiny-ImageNet, ImageNet-1K and its subsets thereof demonstrate that SODA achieves state-of-the-art performance while offering significantly improved computational efficiency over existing BN-matching-based dataset distillation methods, surpassing FADRM+ by +1.6\% on ImageNet-1K IPC=10 under ResNet-18 while delivering a $1.54\times$ speedup.

Open-weight Vision-Language Models (VLMs) score 85--94\% on existing spatial benchmarks (BLINK Spatial, What's Up CLEVR) yet collapse to chance on simple geometric questions when object, scene, and texture priors are stripped away. We introduce and open-source Sophon, a procedurally generated diagnostic of paired colored point-cloud stimuli across three coordinate systems (1D linear, 2D Cartesian, 2D polar) and four question types (binary, categorical, comparative, quantitative). Across four open architectures spanning every major projector design (LLaVA-NeXT, LLaVA-OneVision, Qwen2.5-VL, Idefics3), accuracy clusters at 47--52\% while frontier models (Opus 4.6, Gemini 2.5 Pro, GPT-5.4) clear 77--83\% and humans reach 84--94\%. We trace the open-weight failure end-to-end. The vision tower works: encoder probes recover Sophon spatial direction at 80--97\%. The language tower's reasoning circuit works: a text-only control restores $+40$ pp, and probes preserve the spatial signal through every LLM decoder layer at 80--86\%. Per-layer attention attribution localizes visual integration to mid-stack layers. Yet at the readout, the model elevates the correct answer-set tokens but selects among them at chance. Fine-tuning experiments rule out the unembedding as the mechanism. By moving the unembedding layer through the residual stream, we observe the emergence of a small but discriminative contrast at late-stack residual layers in fine-tuned models, which in practice converts the random selection into a confident one more likely to be correct. However, ablation experiments locate the mechanism upstream: training only the upstream layers with the late stack frozen recovers the full $+29$ pp gain, while training only the late stack captures less than a third. The late-stack contrast is the downstream readout of upstream changes. No layer in either base or fine-tuned models shows geometric alignment between the spatial signal and the unembedding's vocabulary axes, suggesting that 7B-class VLMs have the weights to reason spatially but lack the residual-stream pathways to convert that reasoning into answers.


SpaG-DiT: Enhancing Spatial Grounding for Diffusion Transformers

Zongliang Wu ⋅ Benlei Cui ⋅ Longtao Huang ⋅ Xiaoqian Xia ⋅ Yupeng Cao ⋅ Hui Xue' ⋅ Tuo Chen ⋅ Haiwen Hong ⋅ Xin Yuan

Despite remarkable progress in text-to-image generation, current Diffusion Transformers (DiTs) inherently struggle with learning intricate visual content, such as accurate multilingual scene text rendering and novel visual concepts. This limitation stems from the inherent contradiction between vague semantic guidance provided by text and the objects of accurate reconstruction during diffusion training. To address this bottleneck, we propose SpaG-DiT, a novel spatial grounding training framework. SpaG-DiT directly injects explicit spatial priors into the DiTs before the attention module via Contextual Diagonal Position Encoding (CDPE), which dynamically modulates spatial coordinates without destroying the inherent sequential structure of the linguistic input. Furthermore, we introduce VLM Grounding Guidance (VLM-GG) as an automated pipeline to extract implicit spatial priors from pre-trained Vision-Language Models, eliminating human annotation costs. To evaluate the grounding ability, we establish two new datasets: SpaG-DiT-MultiLingual for non-Latin/Chinese text rendering and SpaG-DiT-Creature for new concept learning. Extensive experiments demonstrate that SpaG-DiT achieves superior training efficiency and significantly outperforms vanilla self-finetuning on the above two tasks and the widely adopted GenEval benchmark. These results validate the plug-and-play nature of SpaG-DiT and its potential to enhance DiT training processes. Code, model, and data will be released.


SpanFormer: Multi-Level Adaptive Sparsity for Object Detection in High-Resolution Wide Shots

Xiang Li ⋅ Chen Zhang ⋅ Wenxi Li ⋅ Yuetong Wang ⋅ Chenyang Lyu ⋅ Haozhe Lin ⋅ Fan Zhang ⋅ guiguang ding ⋅ Yuchen Guo

Object detection in high-resolution wide (HRW) shots, where a single image can span gigapixel resolutions and kilometre-scale fields of view, suffers from extreme foreground sparsity, dramatically varying foreground density from sparse to crowded scenes, and object clusters whose spatial extent ranges from dozens to thousands of pixels within the same dataset. State-of-the-art sparse vision transformers for gigapixel detection select a fixed top-k fraction of windows by hand-crafted variance scores, which over-promote textured but object-free background such as foliage and patterned facades, cannot adapt the keep ratio to the wide range of crowd density, and confine each kept window to a 7x7 receptive field that is far smaller than the spatial extent of many real-world object clusters. We propose SpanFormer, a sparse vision transformer that broadens the computational span at three complementary granularities: (i) Prototype Routing replaces variance with learnable foreground/background prototypes regularised by a balance constraint that prevents prototype collapse, widening the semantic span of the scoring criterion; (ii) Dynamic Top-k predicts a per-image keep ratio with a lightweight selector, making the compute span elastic so that crowded scenes receive proportionally more compute while near-empty scenes are skipped; (iii) Cluster Memory Attention performs union-find clustering over the kept windows by joint cosine similarity and Chebyshev spatial radius and augments each window's K and V with cross-window memory drawn from its cluster, extending the attention span beyond the local window with all projection parameters shared with the local branch and therefore no added projection cost. On the PANDA gigapixel benchmark, SpanFormer improves AP50 from 78.0% to 80.3% over the SparseFormer baseline while reducing backbone FLOPs by over 75%.


Spark: Path-Aware Experiential Self-Evolution for VLMs Spatiotemporal Reasoning

Chengzhengxu Li ⋅ Xiaoming Liu ⋅ Zhaohan Zhang ⋅ Yu Lan ⋅ Zicheng Zhao ⋅ Bingxiang Wang ⋅ Cong Wang ⋅ Chao Shen

While vision-language models (VLMs) perform well on general tasks, their spatiotemporal reasoning remains limited. Existing distillation and Reinforcement Learning (RL) methods are often computationally expensive and poorly generalizable. Recent studies suggest that model self-improvement by experience distilled from success and failure trajectories is a promising direction. However, in VLM tasks where reasoning must be grounded in visual information, valuable signals of reasoning quality arise not only from final outcomes, but also from how faithfully the reasoning path interacts with visual information. We examine the dependency between reasoning paths and visual information, and observe that high-quality paths exhibit stronger visual anchoring and greater visual dependency. Building on this observation, we introduce the Spatiotemporal Path-Aware Reasoning paradigm based on experiential Knowledge (Spark) that encourages the model to explore and exploit valuable experiences derived from paths with subtle quality differences. To enable diverse and efficient path exploration, we integrate Monte Carlo Tree Search (MCTS) with fine-grained visual rewards and submodular optimization, allowing the search process to prune redundant branches while preserving critical reasoning paths. We further present a training-free experience construction mechanism that converts contrastive pairs of distinct-value paths into structured experiences, which are dynamically injected into the VLM through similarity-based retrieval during inference. Extensive experiments demonstrate that Spark enables Qwen3.5-9B to achieve an average accuracy of 61.21% across three benchmarks, outperforming Gemini-3-Pro (57.20%). Further analyses verify the flexibility and cross-model generalization of our method, highlighting the crucial role of path-aware experience in advancing spatiotemporal intelligence.


Sparse blind deconvolution via thresholded Wirtinger flow

Mengting Chen ⋅ Haitong Lan ⋅ Ziping Zhao

Blind deconvolution is a classical problem arising in many signal processing and machine learning applications, where one aims to recover an unknown kernel and an unknown signal from their convolution. Since both the kernel and the signal are unknown, the problem is intrinsically ill-posed, and meaningful recovery is possible only under suitable structural assumptions. In this paper, we study sparse blind deconvolution, where the kernel is $s_h$-sparse and the signal admits an $s_x$-sparse representation in a known dictionary. Although sparsity is a natural and practically relevant prior, its interaction with the bilinear observation model creates substantial challenges for both computation and analysis. To address these challenges, we propose Thresholded Wirtinger Flow (ThWF), a simple and scalable iterative algorithm equipped with a sparse spectral initialization. For noiseless observations, we prove that the proposed initialization achieves exact support recovery with sample complexity $M \gtrsim s_h^2+s_x^2$, up to logarithmic factors. Building on this initialization, we establish what is, to the best of our knowledge, the first algorithmic convergence guarantee for sparse blind deconvolution: ThWF converges linearly to the ground truth, up to the inherent scaling ambiguity, provided that $M \gtrsim s_h+s_x$ up to logarithmic factors. Our analysis further extends to noisy measurements, showing that ThWF is robust and contracts linearly up to a statistical error under a bounded noise level. Experiments on synthetic data and real-world image deblurring tasks corroborate the predicted linear convergence and phase transition behavior, and demonstrate the restoration performance of the proposed method.


Sparsely Wired Mortal LLM Inference

Mincheol Park ⋅ Sunwoo Lee ⋅ Sukjin Lee ⋅ Wooram Yang ⋅ Yongmin Tai ⋅ Sang Joon Kim ⋅ Jaehoon Yu

This paper rethinks Large Language Model (LLM) inference from a Mortal Computing perspective, in which computation is embodied in hardware and functionality is inseparable from its physical substrate. Rather than treating inference optimization as merely a matter of arithmetic simplification or precision reduction, we explore hardwiring as a way to eliminate weight access at the root of data movement. We show that viable hardwiring begins not by retaining weights in their multiplicative form, but by factorizing them into combinations of shiftable bases and rewiring computation at the base level, where redundancy is exposed. This view, however, introduces a fundamental challenge in the dramatic growth of wires required to reconstruct weights. To address this, we propose Moira, a budgeted greedy search that formulates wire-count control as a Lagrangian relaxation problem and selects only the bases that satisfy the objective under a wiring budget. What sets Moira apart is its adaptive base selection, which gives it a dual characteristic of pruning and quantization while preserving model expressivity under sparse wiring without retraining. Extensive evaluations demonstrate that Moira achieves near-lossless performance with, on average, only about two bases per weight and outperforms existing compression methods. Moira opens a concrete path beyond the von Neumann paradigm, marking a first step toward Digital Mortal Computing for LLMs.


Sparsity for Free: A Budget-Induced Equilibrium in Joint Topology–Parameter Search

Marcel Mordarski ⋅ Daniel Budina ⋅ Benjamin Gras ⋅ Abdulrahman Shehata ⋅ Roberto Bondesan

Joint optimisation over a discrete topology $T$ and continuous parameters $\theta$ under a strict evaluation budget $B$ arises in neural architecture search, pruning, and variational quantum algorithms. We identify a budget-induced effect, \emph{dimensionality pressure}: enlarging $T$ expands the search space of $\theta$, slows inner-loop optimisation at fixed $B$, and yields noisier fitness estimates that systematically bias outer-loop selection toward smaller structures, even without explicit sparsity regularisation. We formalise this as a Proposition composing standard CMA-ES dimensionality results with rank-based noisy-comparison selection, predicting a stationary distribution over $|T|$ concentrated below the dense baseline. To test the mechanism we introduce \textsc{EvoluCMAES}, a minimal two-loop framework whose only domain-specific component is the topology-mutation operator. Across two unrelated settings, the same fixed apparatus produces low-density solutions: in pruning of UCI tabular models, the equilibrium density settles at $\rho \approx 0.13$, roughly an order of magnitude below dense; in BB84 eavesdropping, discovered Eve circuits use 4--6 gates from an arbitrary-depth search space and approach the analytical Pauli-channel cloning bound under bit-flip noise. When a specific target sparsity is required for deployment, adapting only the mutation operator suffices: at $95\%$ sparsity on the UCI suite, mean accuracy reaches $0.972$, exceeding the strongest one-shot baseline at $0.946$. As a downstream consequence of the bound, the \textsc{EvoluCMAES} library is small enough to serve as a tractable action set for an off-the-shelf RL agent: a single Rainbow-DQN with feasibility masking and matched hyperparameters recovers $\approx 99\%$ of the dynamic-programming optimum on adaptive eavesdropping for both BB84 and E91/DIQKD.


SpatialBench: Is Your Spatial Foundation Model an All-Round Player

Haosong Peng ⋅ Hao Li ⋅ jiaqi chen ⋅ Yuhao Pan ⋅ Runmao Yao ⋅ Yalun Dai ⋅ Fushuo Huo ⋅ Fangzhou Hong ⋅ Zhaoxi Chen ⋅ Haozhao Wang ⋅ Dingwen Zhang ⋅ Ziwei Liu ⋅ Wenchao Xu

While spatial foundation models have demonstrated impressive performance on standard datasets, a critical question remains: are they truly all-round players capable of generalizing robustly across diverse downstream tasks, arbitrary viewpoints, shifting scene domains, varying input densities, and specific hardware constraints? Answering this overarching question requires a holistic assessment, yet current models are mainly evaluated on specific domains for which they were specifically designed or trained. Such evaluations are intrinsically limited by narrow paradigm coverage, limited scene domains, and arbitrary frame sampling, making it fundamentally difficult to assess their true generalization capabilities. To address this gap, we present SpatialBench, a cross-paradigm, domain-diverse benchmark for spatial foundation models with deterministic sampling. SpatialBench features unprecedented scale and rigorous deterministic design, comprising 19 datasets and 546 scenes across 5 diverse spatial domains. It comprehensively evaluates 40 models across 6 paradigms on 5 task suites under 4 different input density settings. Our extensive evaluation reveals that current models are not yet all-round players, and uncovers crucial insights for future advancement. Specifically, we demonstrate that full-context attention maximizes accuracy while bounded-memory strategies unlock long-sequence scalability. Moreover, our empirical evaluations in challenging embodied and egocentric tasks demonstrate that strict domain alignment and high data quality are far more critical to performance than simple dataset scaling. Furthermore, to address the largest data gap identified in our analysis, we go beyond evaluation by introducing a large-scale dataset, DA-Next-5M, alongside a strong baseline model, DA-Next, pushing the boundaries of spatial representation learning.


SpatialVAM: Spatial-Aware Multi-View Video Diffusion as a Data-Efficient Robot Policy

Peiyan Li ⋅ Yixiang Chen ⋅ Yuan Xu ⋅ Jiabing Yang ⋅ Xiangnan Wu ⋅ Jun Guo ⋅ Nan Sun ⋅ Long Qian ⋅ Xinghang Li ⋅ Xin Xiao ⋅ Minghui Zhang ⋅ Jing Liu ⋅ Nianfeng Liu ⋅ Tao Kong ⋅ Yan Huang ⋅ Liang Wang

Robotic manipulation requires understanding both the 3D spatial structure of the environment and its temporal evolution, yet most existing policies neglect one or both aspects. They often rely on 2D visual observations or backbones pretrained on static image--text pairs, which leads to high data requirements and limited comprehension of environment dynamics. To address this, we introduce SpatialVAM, **the first 3D Video Action Model** that simultaneously predict spatial-aware multi-view heatmap videos and RGB videos. Our key insight is that this design naturally injects 3D information into video foundation models while aligning the representation format between video pretraining and action finetuning. Extensive experiments demonstrate that SpatialVAM enables **data-efficient, robust, generalizable, and interpretable** manipulation. With only ten demonstration trajectories and no additional pretraining, SpatialVAM handles challenging long-horizon and contact-rich tasks, generalizes to out-of-distribution settings, and predicts realistic future videos. Evaluations on Meta-World (22\%$\uparrow$), RoboCasa (10\%$\uparrow$) and real-world robotic platforms (16\%$\uparrow$) show that SpatialVAM consistently outperforms other video action models, vision language action models and 3D-based policies, establishing a new state-of-the-art in data-efficient multi-task manipulation.


SpecBridge: Learning Natural-Language Formalization Plans for the Formal Specification Synthesis Task

Wenjie Zhang ⋅ Yun Lin ⋅ Zining He ⋅ Meiqi Wu ⋅ Jin Song Dong

Software requirements are usually written in natural language, which is essential for human communication but insufficient as a target for machine-checked correctness. While natural-language descriptions can state intended functionality, verification requires formal specifications with explicit types, relations, quantifiers, guards, witnesses, and edge cases. To bridge this divide, we study the Formal Specification Synthesis task: given a natural-language programming requirement and fixed formal signatures, synthesize a proof-assistant specification for the intended input-output relation. This task is difficult for two primary reasons. First, a large representation gap separates informal natural language from the rigorous structures needed by a proof assistant. Second, generated formal specifications are notoriously hard to evaluate, as proving full equivalence between specifications is often too complex to serve as routine feedback. To address these intertwined challenges, we propose SpecBridge, a reconstruction-guided formalization framework. At its core, SpecBridge introduces a Natural-Language Formalization Plan (NLFP), a semi-structured, readable intermediate representation that captures key formalization choices before decoding to a target specification language. Through reconstruction-guided pattern mining, SpecBridge learns these NLFPs by reconstructing formal specifications to discover reusable patterns, ultimately transferring them to improve natural-language task generation. Furthermore, to tackle the evaluation bottleneck, we introduce a multi-layer evaluation protocol for formal-specification correctness. This protocol utilizes LLM Specification Alignment as an LLM-based judge, alongside Test-Driven Computable Validation and Test-Driven Proof Validation, which approximate "running" testcases on formal specifications by translating propositions into executable pre/post checks and proving testcase obligations in Lean. In our experiments, SpecBridge outperforms the few-shot + CoT baseline by 6.1% on LLM Specification Alignment, 15.2% on Computable Validation Pass Rate, and 28.4% on Proof Validation Pass Rate. These results demonstrate that the NLFP bridge and our multi-layer evaluation protocol make formal specification synthesis significantly more reliable.


Spectral Energy Allocation Enables Source-Free Domain Adaptation in Time Series Forecasting

Zhimin Mei ⋅ Gezheng Xu ⋅ Mohammad Hossein Moslemi ⋅ Ruiyi Fang ⋅ Nima Hosseini Dashtbayaz ⋅ Ziyan Wang ⋅ Jincheng Zhou ⋅ Haolin Jiang ⋅ Song Tang ⋅ Charles Ling ⋅ Boyu Wang

Source-free domain adaptation (SFDA) adapts a pre-trained source model to an unlabeled target domain without access to source data. However, existing SFDA studies have predominantly focused on classification tasks, leaving time series forecasting (TSF), a fundamentally different problem with continuous outputs and complex temporal dependencies, largely unexplored. In this work, we identify unique characteristics and key challenges of SFDA in TSF: 1) forecasting labels correspond to future temporal segments of the same underlying signal, enabling us to split target input and construct a self-supervised learning task, thereby training an accurate short-term forecasting model to provide high-quality pseudo-labels; 2) however, such supervision remains inherently local rather than global for long-horizon source model adaptation. To fill this gap, we propose the notion of Spectral Energy Allocation (SEA) pattern, defined over a short-long signal pair. The SEA pattern provides a structured correspondence that bridges signals across different temporal horizons, enabling reliable short-horizon pseudo-labels to guide long-horizon adaptation. Extensive experiments demonstrate that our method substantially outperforms baselines across different forecasting backbones.


Spectral Progressive Diffusion for Efficient Image and Video Generation

Howard Xiao ⋅ Brian Chao ⋅ Lior Yariv ⋅ Gordon Wetzstein

Diffusion models have been shown to implicitly generate visual content autoregressively in the frequency domain, where low-frequency components are generated earlier in the denoising process while high-frequency details emerge only in later timesteps. This structure offers a natural opportunity for efficient generation, as high-resolution computation on noise-dominated frequencies is largely redundant. We propose Spectral Progressive Diffusion, a general framework that progressively grows resolution along the denoising trajectory of pretrained diffusion models. To this end, we develop a spectral noise expansion mechanism and derive an optimal resolution schedule from the model's power spectrum. Our framework supports training-free acceleration and a novel fine-tuning recipe that further improves efficiency and quality. We demonstrate significant speedups on state-of-the-art pretrained image and video generation models while preserving visual quality.


Spectral Unlearning: Transformer Structure-Preserving Updates for Language Model

Sung Il Choi ⋅ Junhao Cai ⋅ Dohun Kim ⋅ Changhee Joo

Machine unlearning in large language models (LLMs) aims to remove specific learned knowledge, capabilities, or behaviors while preserving model utility. Existing methods have mostly focused on the design of unlearning objectives, while largely overlooking the intrinsic matrix structure of transformer weights. In particular, standard optimizers operate on independent scalar coordinates, thereby ignoring this matrix structure. We propose \emph{spectral unlearning}, an update framework that combines standard unlearning objectives with matrix-structured optimization. We show that the spectral update corresponds to the exact steepest-descent direction for transformers composed of linear layers with normalized inputs and outputs. Our approach is applied on top of existing unlearning methods and evaluated across complementary axes of memorization, privacy, and utility, summarized via their harmonic mean. Spectral unlearning consistently outperforms a broad range of baselines and closes most of the gap to a retain-trained reference.


SPEXT: A Decoupled Multi-Spectral Foundation Model for Earth Observation and Vision-Language Grounding

Sannara EK ⋅ Chong Tang ⋅ Dirk Koch ⋅ Alex Weddell ⋅ Jagmohan Chauhan ⋅ Robert Mullins

Vision-Language Models for Earth Observation are bottlenecked at the visual encoder. Existing pipelines either compress multi-spectral signals into RGB pixels or collapse them into a single CLIP-style global vector, losing spectral information or spatial structure respectively. We show this is a consequence of coupling contrastive supervision to the dense feature stream, which degrades spectral fidelity. SPEXT resolves this through architectural decoupling. A wavelength-parameterized encoder produces multi-spectral patch tokens for dense tasks and VLM grounding, while a dedicated alignment token absorbs the contrastive gradient for retrieval. SPEXT simultaneously leads the PANGAEA segmentation benchmark (+2.51 mIoU), tops zero-shot retrieval among multi-spectral CLIP baselines (+4.6 mAP@100), and enables a frozen VLM, paired with a small projection adapter and LoRA, to answer free-form Earth-Observation questions directly from its tokens, substantially outperforming an RGB-pixel baseline on a downstream EO-RAG benchmark.


Spikes as Detectors: Phase-Conditioned Spiking Dynamics for Time-Series Anomaly Detection

Miao Wang ⋅ Gangyi Ding ⋅ Yunlin Lei ⋅ Yu Zhang ⋅ chenjingfeng ⋅ Xu Yang

Most time-series anomaly detectors ask whether an observation can be predicted or reconstructed. We ask a different question: can a detector’s own spiking dynamics become anomaly evidence? We study this question for periodic and quasi-periodic time series and propose SPIRE, a phase-conditioned spiking anomaly detector. SPIRE uses period-aware current formation to expose phase-aligned cross-period residuals to leaky integrate-and-fire neurons. We refer to this conversion as eventification: after residual injection, recurrence violations induce abnormal bursts, silences, or phase-shifted spike trains. SPIRE then constructs phase-conditioned normal templates of spike states and scores deviations in residual energy, firing-rate patterns, and burst/silence statistics, instead of relying solely on output-space prediction or reconstruction errors. This separates the operational protocol, including causal current-step detection, one-step prediction, and non-causal reconstruction, from the source of anomaly evidence, namely output errors, spike dynamics, or both. On TSB-AD-U and TSB-AD-M, SPIRE establishes new best VUS-PR among neural time-series anomaly detectors. Spike-native scores remain competitive without reconstruction error, hybrid scoring consistently improves error-only scoring, and replacing LIF neurons with continuous activations in matched ANN counterparts reduces detection performance under the same period-aware architecture. Periodicity-stratified analyses and spike-state visualizations further show that the gains are strongest when phase recurrence is reliable, supporting the proposed spike-as-detector mechanism. These results suggest that phase-conditioned spiking dynamics can serve as a first-class anomaly detection signal for periodic time series, rather than merely an efficient computational substrate.


SpikeSSL: A Universal Spike Inference Framework with Dynamics-Informed State-Space Layers

Chenghao Yue ⋅ Siming Xing ⋅ shuran liu ⋅ Angran Li ⋅ Yuanlong Zhang

Two-photon calcium imaging is a standard tool for recording large neural populations in vivo, yet inferring spikes accurately across the growing diversity of calcium indicators remains an open problem. Existing supervised methods achieve reasonable in-domain accuracy but generalize poorly to unseen indicators, because different indicators induce distinct fluorescence kinetics and signal statistics while existing architectures remain relatively simple generic temporal regressors without dynamics-matched inductive bias. We propose SpikeSSL, a universal spike inference framework whose temporal backbone is a bank of bidirectional IIR state-space layers broadly motivated by calcium dynamics. A multi-modal conditioning encoder maps indicator identity, sampling rate, and trace-level signal statistics into a global conditioning vector that modulates the backbone via Adaptive Layer Normalization, while a heteroscedastic variance head provides calibrated per-frame uncertainty. On a benchmark with five fixed evaluation splits built from 33 public ground-truth datasets, SpikeSSL achieves state-of-the-art performance in both in-domain and zero-shot leave-one-indicator-out settings. We also develop a biophysical simulation pipeline capable of generating paired fluorescence-spike traces with systematically varied kinetic parameters, spike statistics, response nonlinearities, baseline drift, and noise. Using this pipeline, we synthesize approximately 11{,}000 simulated traces. Augmenting training with these data effectively closes the cross-indicator domain gap and improves zero-shot generalization.


SPRM: From Cooperative Games to Marginal-Contribution Process Reward Modeling

Yu Bao ⋅ Pak Lon Ip ⋅ Qiyu Ruan ⋅ Xitong Gao ⋅ Cheng-Zhong Xu

Process reward models provide step-level feedback for reasoning, but reliable process supervision remains difficult to obtain: human annotations are costly, Monte Carlo rollouts require large sampling budgets, and LLM judges can introduce prompt-sensitive biases. We introduce SPRM, an outcome-only framework that constructs process rewards directly from terminal correctness signals. SPRM views a reasoning trajectory as a temporally ordered cooperative game, where reasoning steps act as participants and the final outcome reward is the shared payoff. Following this allocation view, SPRM learns a differentiable trajectory value function from outcome rewards and distributes the payoff to individual steps through Aumann-Shapley-style Step Value Integration. We further analyze its Shapley properties under the temporal and causal structure of reasoning trajectories. To make the resulting supervision robust, SPRM incorporates causal consistency and estimates a cross-trajectory inherent score that captures each step's context-independent value from semantically similar steps. These signals train a dual-head process reward model that jointly represents trajectory-specific marginal credit and cross-trajectory inherent step value. Across ProcessBench, PRMBench, and Best-of-N reranking on AceMath-RewardBench, SPRM improves first-error localization, marginal-contribution-sensitive evaluation, and test-time reasoning selection over existing open PRMs trained from human, rollout, curated, or outcome-broadcast supervision. These results suggest that marginal-contribution-based credit redistribution offers an effective and scalable route to process supervision from outcome-only data.


Squeezing Capacity from Multimodal Large Language Models for Subject-driven Generation

Shuhong Zheng ⋅ Aashish K Misraa ⋅ Kevin Li ⋅ Yu-Jhe Li ⋅ Igor Gilitschenski

Subject-driven image generation aims to synthesize new images that preserve the identity of the given subject while following textual instructions. Existing approaches often encode text and reference images separately. This limits cross-modal reasoning abilities and causes copy-paste artifacts. Recent frameworks that connect multimodal models and diffusion models improve instruction following, but largely overlook identity preservation. To address these limitations, we condition diffusion models on Multimodal Large Language Models (MLLMs) that jointly encode text and reference images, and augment it with VAE‑based identity conditioning. A novel Dual Layer Aggregation (DLA) module is designed to aggregate multi‑level MLLM features for optimal conditioning, and a multi‑stage denoising strategy is applied to progressively balance the semantic information from MLLM and fine‑detail identity from VAE during inference. Extensive experiments demonstrate that our approach harmonizes multimodal understanding with identity preservation, mitigates copy-paste issues, and achieves superior performance regarding human preference on subject-driven image generation.


SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios

Jackson Clark ⋅ Yiming Su ⋅ Saad Mohammad Rafid Pial ⋅ Yifang Tian ⋅ Lily Gniedziejko ⋅ Hans Arno Jacobsen ⋅ Yinfang Chen ⋅ Tianyin Xu

AI agents are increasingly used to diagnose and mitigate failures of production systems, known as agentic Site Reliability Engineering (SRE). Current SRE benchmarks are limited to oversimplistic SRE tasks and are unfortunately hard to extend due to bespoke designs. We present SREGym, a high-fidelity, interactive benchmark for SRE agents. SREGym exposes a live system environment built atop real-world cloud-native system stacks, where high-fidelity failure scenarios are emulated through fault injectors. SREGym models the complexity of production environments by simulating (1) a wide range of faults across the stack, (2) various ambient noises, and (3) diverse failure modes such as metastable failures and concurrent failures. SREGym is architected as a modular, extensible framework that orchestrates fault injectors and event emulators across stacks; this framework is key to curating high-fidelity, challenging SRE problems. SREGym shows that the effectiveness of frontier models and agents varies significantly (40+%) in mitigating different kinds of problems. SREGym is actively maintained as an open-source project and has been used by many researchers and practitioners.


SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control

Ruihua Han ⋅ Rui Gao ⋅ Zhe Liu ⋅ Xinyi WANG ⋅ Chang Chen ⋅ Shuai Wang ⋅ Qi Hao ⋅ Jia Pan ⋅ Hengshuang Zhao

Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume homogeneous robots, sparse workspaces, simplified geometry, offline computation, or handcrafted parameters to make the problem tractable, which limits their deployment in dense crowd scenarios. Toward this end, we propose Shape-Aware Reinforcement Learned Model Predictive Control (SRL-MPC), a method for safe, efficient, and adaptive navigation in crowds with heterogeneous shapes without geometry simplification. To encode shape-aware safety, we formulate high-order control barrier function (HOCBF) constraints from geometric separation features (GSFs) based on support function transformation. A reinforcement learning (RL) framework then learns a neural policy that reads GSFs and outputs real-time MPC parameter updates, enabling the MPC solver to adapt to neighboring crowd geometries. The key advantage of SRL-MPC is that it preserves the safety structure and generalizability of MPC while integrating the adaptability and intelligence of RL. Experiments in randomized crowd scenarios with arbitrary shaped robot fleets demonstrate the effectiveness, scalability, and robustness of SRL-MPC. The results show that SRL-MPC substantially outperforms representative baselines in safety and adaptability.


Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping

Hsun-Yu Kuo ⋅ El Mahdi Chayti ⋅ Patrik Reizinger ⋅ Wieland Brendel ⋅ Martin Jaggi

Looped Transformers --- which repeatedly apply a shared transformer block --- are an architecturally natural fit for variable-length algorithmic tasks. Although they can exhibit strong length generalization beyond the length of training sequences, this behavior is brittle, yielding high out-of-distribution (OOD) variance, even across well-performing in-distribution solutions. We trace this variance to the spurious correlation in simple algorithmic tasks between sequence length and number of loops. Introducing stochasticity into the number of loops during training sharply reduces OOD variance and stabilizes predictions across inference-time loop counts. To improve upon heuristic randomization schemes, we further analyze RL-Halting as a learned stochastic schedule and find that it generally improves the accuracy--stability trade-off. Across binary addition, Dyck-1, Unique Set, and Copy, learned stochastic stopping often improves this trade-off but can also stabilize a suboptimal computation. Our work suggests that “when to stop” should be treated as a training-time design choice, not merely an inference-time computation-allocation rule.


Stable3R: Streaming 3D Reconstruction with Stable Geometric References

Jincheng Xiong ⋅ Jiarong Han ⋅ Hang Zhang ⋅ Mu Xu ⋅ Ming Qian

Streaming 3D reconstruction has become a critical component for real-time applications. While recent advances have adapted large-scale offline feedforward models into streaming architectures via causal attention and KV-caching, they suffer from fundamental instability: the irreversible bias in early frames. This limitation stems from two coupled factors: (1) the existing KV-cache-based causal architecture strictly prohibits access to future information, resulting in information-poor early caches; and (2) current methods typically anchor the pose constraints to the initial frame, amplifying the early bias throughout the sequence. To address these challenges, we propose Stable3R, a holistic framework that establishes stable geometric references by enriching historical caches and anchoring poses to prefixes rather than a single initial frame. Specifically, we first introduce a lookahead-augmented attention mechanism, integrating future observations into past caches without violating causal constraints. Second, we adopt a frame-equivariant architecture with prefix-relative supervision, alleviating the reliance on the biased initial frame. Extensive experiments across multiple benchmarks show that our method consistently boosts reconstruction quality within a strictly causal inference schedule.


Stage-wise Attention-Guided Region Sequencing for Adversarial Attacks on Large Vision-Language Models

Jaehyun Kwak ⋅ Nam Cao ⋅ Boryeong Cho ⋅ Segyu Lee ⋅ Sumyeong Ahn ⋅ Se-Young Yun

Targeted adversarial attacks on Large Vision-Language Models (LVLMs) test whether small image perturbations can steer model responses toward attacker-specified content. Under the standard $L_\infty$ constraint, targeted attacks become a regional perturbation budget allocation problem: attack success depends not only on the perturbation objective, but also on which regions receive updates and in what order. Existing localized attacks improve over global perturbations but rely on stochastic spatial sampling, often updating weakly influential regions. We address this limitation through an attention-based analysis showing that cross-modal attention identifies adversarially sensitive regions and that perturbing high-attention hotspots induces predictable redistribution toward subsequent salient regions. These findings motivate attention-guided region sequencing, which begins from dominant hotspots and progressively moves the update support toward next-salient regions. Based on these principles, we propose Stage-wise Attention-Guided Attack (SAGA), a black-box region-sequencing framework that uses a fixed attention map from an open-source LVLM to guide perturbation updates without accessing target-model parameters, gradients, or attention maps. Across ten closed-source and open-source LVLMs, SAGA achieves state-of-the-art attack success rates and the best overall imperceptibility.


StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns

Uladzislau Sobal ⋅ Shuo Yang ⋅ Yuting Zhang ⋅ Wei Xia ⋅ Stefano Soatto

We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests) they can handle before failing. Unlike the prevailing fraction-of-tasks-solved metric, this matches real vibe-coding where sessions run dozens or hundreds of turns. In StaminaBench, agents implement a REST API server and modify it across a tunable number of pro- cedurally generated follow-up change requests—100 in our experiments, resulting in codebases of up to 6,000 lines. Tests are generated fully programmatically with- out LLM involvement, ensuring reproducibility and reliability; change sequences are drawn from either a hardcoded or LLM-driven sampler, both constrained to a structured action space to ensure changes are valid. The agent and the server run in an isolated environment and communicate with the benchmark through HTTP, making testing fully black-box and language-agnostic. We evaluate six agent harnesses paired with seven open-source LLMs across 20 scenarios of 100 turns each and find that: (1) all the tested models fail within 5–6 turns, confirm- ing that vibe-coding-style programming without thorough testing produces bugs; (2) passing test feedback back to the agent and allowing it to retry improves passed turn count by up to 12×; and (3) a good harness is required for strong performance: stronger models exhibit up to a 6× gap between their best and worst harness, while weaker models fail with any harness. We release the benchmark and the generated tasks to enable further research into multi-turn coding agent behavior.


Standing on the Shoulders of Giants: Rethinking EEG Foundation Model Pretraining via Multi-Teacher Distillation

Chenqi Li ⋅ Yu Liu ⋅ Shuo Zhang ⋅ Timothy Denison ⋅ Tingting Zhu

Pretraining for electroencephalogram (EEG) foundation models has predominantly relied on self-supervised masked reconstruction, a paradigm largely adapted from and inspired by the success of vision and language foundation models. However, unlike images and text, EEG datasets are notoriously expensive to collect and characterized by low signal-to-noise ratio. These challenges introduce difficulties in scaling the EEG foundation models and capturing the underlying neural semantics through reconstruction. In this work, we ask the question: can we stand on the shoulders of well-established foundation models from well-represented modalities to bootstrap the pretraining of EEG foundation models? We first demonstrate that mainstream foundation models, such as those from vision and time series, transfer surprisingly well to EEG domain. Motivated by this observation, we propose the multi-teacher distillation pretraining (MTDP) framework for pretraining EEG foundation models via two-stage multi-teacher distillation. In the first stage, we introduce a learnable gating network to fuse representations from diverse teachers (e.g., DINOv3 and Chronos) via a masked latent denoising objective. In the second stage, we distill the fused representation into an EEG foundation model. Extensive evaluations across 2 backbone architectures, 9 downstream tasks and 12 datasets demonstrate that multi-teacher distillation pretraining consistently outperforms self-supervised pretraining for EEG foundation model and remains competitive even when the pretraining data is reduced to 25%.


StateLedger: Path-Addressed External Memory for Persistent Multi-Agent Systems

Geer Yang ⋅ Xixuan Liu ⋅ Bin Wu ⋅ Shaojiang Wang

Persistent LLM agents need memory that preserves state, provenance, validity, and evidence across long interaction histories. We introduce StateLedger, a path-addressed external memory substrate that stores memory as durable state evolution rather than flat retrievable items: events update canonical paths, registers absorb small changes, CommitPages record residual drift against Checkpoints, cue views bound traversal, and content-addressed artifacts avoid evidence duplication. We prove storage, retrieval, coalescing, and conflict-aware ranking guarantees under explicit coverage and reader-consistency conditions. Empirically, StateLedger improves task performance and storage efficiency across LongMemEval, LoCoMo, and Evo-Memory Multi-turn, reaching 98.8% on LongMemEval-S, 82.6% on LongMemEval-M with 14.7% durable footprint, 95.8% LoCoMo non-adversarial accuracy, 87.1% unweighted success, and 93.5% unweighted progress on Evo-Memory.


Static-Dynamic Disentanglement for Efficient Multi-Frame Vision-Language-Action Models

Weikang Qiu ⋅ Huashuo Lei ⋅ Tinglin Huang ⋅ Aosong Feng ⋅ Rex Ying

Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for generalist robotic control. Built upon vision–language model (VLM) architectures, VLAs predict actions conditioned on visual observations and language instructions, achieving strong performance and generalization across tasks. However, VLAs face two major challenges: a limited context window for input frames and inefficient inference due to the quadratic attention complexity and large parameter counts. To this end, we propose Dysta, a framework that disentangles visual inputs into multi-level static and dynamic tokens, which enables (1) retaining a single copy of static tokens across frames to significantly reduce context length, and (2) reusing the key–value (KV) cache of static tokens through a lightweight recache gate that updates only when necessary. This design enables efficient multi-frame integration and efficient inference. In addition, we introduce a new benchmark that more effectively evaluates the multi-frame integration ability of VLAs. Experiments show that DySta improves multi-frame integration by 24.5\% across metrics on our benchmark and 23.3\% in absolute success rate on real-world memory-dependent tasks, while accelerating inference by 2.0× (with +2.3\% success rate) on simulation benchmarks and 2.2× (with +10.6\% success rate) on real-world general tasks.


Static-to-Dynamic: Animating Still Mattes via Generative Motion for Video Matting

Suqi Song ⋅ Jiaqi Xu ⋅ Jialuo He ⋅ Renjing Pei ⋅ Lei Chen ⋅ Lei Zhu

Video matting is essential for high-precision foreground extraction, yet its development is constrained by the scarcity of diverse annotated training data. To address this, we propose Static-to-Dynamic (S2D), a novel generative paradigm that scales video matting data by animating still labeled images into high-fidelity video-matte pairs. Starting from a static image-matte pair, S2D leverages an image-to-video generative model to synthesize realistic motion and scene dynamics, while simultaneously producing pseudo alpha labels from generative priors. Specifically, we introduce Latent Matte Propagation (LMP), which propagates the first-frame matte to subsequent generated frames via the attention mechanisms in the diffusion transformer. We further design a Coarse-to-Fine Matte Decoder to recover pixel-level alpha details from compressed latent representations using multi-scale generative features. Based on this paradigm, we build S2D-VM, a large-scale dataset containing 14K high-fidelity paired video clips. We also propose a Self-Corrected Progressive Training strategy to mitigate pseudo-label noise during optimization. Experiments show that S2D effectively scales video matting data and consistently improves the performance of existing models across multiple benchmarks.


STDec: Spatio-Temporal Stability Guided Decoding for dLLMs

Yuzhe Chen ⋅ Jiale Cao ⋅ Xuyang Liu ⋅ Jin Xie ⋅ Aiping Yang ⋅ Yanwei Pang

Diffusion Large Language Models (dLLMs) have achieved rapid progress, viewed as a promising alternative to the autoregressive paradigm. However, most dLLM decoders still adopt a global confidence threshold, and do not explicitly model local context from neighboring decoded states or temporal consistency of predicted token IDs across steps. To address this issue, we propose a spatio-temporal stability guided decoding approach, named STDec. We observe strong spatio-temporal stability in dLLM decoding: newly decoded tokens tend to lie near decoded neighbors, and their predicted IDs often remain consistent across several denoising steps. Inspired by this stability, our STDec includes spatial-aware decoding and temporal-aware decoding. The spatial-aware decoding dynamically generates the token-adaptive threshold by aggregating the decoded states of nearby tokens. The temporal-aware decoding relaxes the decoding thresholds for tokens whose predicted token IDs remain consistent over denoising steps. Our STDec is training-free and remains compatible with cache-based acceleration methods. Across textual reasoning and multimodal understanding benchmarks, STDec substantially improves throughput while maintaining comparable task performance score. Notably, on MBPP with LLaDA, STDec achieves up to 14.17$\times$ speedup with a comparable score. Our code will be available.


Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

Zihao Zhu ⋅ Siwei Lyu ⋅ Adel Bibi ⋅ Baoyuan Wu

A skill is a modular package of natural-language instructions, executable scripts, and reference resources that an agent can load at runtime to extend its capabilities for a specific task. Skill-based agent systems therefore enable flexible reuse of third-party capabilities, but the openness of this skill ecosystem also opens up a new attack surface. Prior work has focused on vulnerabilities within individual skills, but little attention has been paid to risks that arise from interactions across skills. In this paper, we introduce skill cascading attacks, a threat paradigm in which a malicious objective is distributed across multiple skills so that each modification looks benign in isolation, yet their combined execution is harmful. For instance, in a prescription-review pipeline, the first skill weakens signals of recently discontinued medications in the extracted history, the second downgrades the severity of any drug interaction tied to them, and the third suppresses the resulting low-priority alert in the final summary, so that a severe drug-interaction warning silently disappears before reaching the physician. To systematically study this safety blind spot, we develop SkillCascade, an automated multi-agent red-teaming framework, and release SkillCascade-Bench, a benchmark of 213 validated cascading test cases across multiple agent systems and domains. Across representative agents (e.g. OpenClaw, Claude Code, CodeX) and LLM backbones, cascaded interactions reliably induce harmful behaviors while evading existing per-skill scanners and runtime monitors. Our findings highlight a gap between component-level integrity and system-level safety, and call for defenses that reason over cross-skill interactions rather than individual skills in isolation.


Steering Optimisation Trajectories in Diffusion Representation Learning

Rajat Rasal ⋅ Tian Xia ⋅ Avinash Kori ⋅ Ben Glocker

We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace this behaviour to optimisation dynamics; we analyse curves of image reconstruction against latent representation quality, revealing trajectories that organise around two distinct regimes early in training. Models in the reconstruction regime prioritise image fidelity early, whereas those in the disentanglement regime improve reconstruction and disentanglement more gradually. We hypothesise that this behaviour can be influenced by targeting shortcut pathways in the diffusion U-Net and controlling early noise-level exposure, thereby shaping the reconstruction-disentanglement trade-off during training. To steer optimisation toward stronger representations, we introduce SteeringDRL, combining gated residual U-Nets with a simple noise-level exposure curriculum for training. Across disentanglement benchmarks, SteeringDRL improves representation quality and reduces seed sensitivity. Our method further extends to spatial disentanglement in object-centric learning, improving segmentation quality on synthetic and real-world datasets.

The rapid advancement of large language models (LLMs) has made machine-generated text increasingly difficult to distinguish from human-written text. While recent studies explore leveraging internal representations of language models to uncover deeper detection signals, these raw features often exhibit substantial overlap between classes, limiting their discriminative power. To address this challenge, we propose Steer-to-Detect (\texttt{S2D}), a two-stage framework for detecting LLM-generated text. In the first stage, \texttt{S2D} learns a steering vector that is injected into the hidden states of a frozen observer LLM, producing representations with improved class separability. In the second stage, detection is performed via a hypothesis testing procedure based on the steered representations. We establish finite-sample, high-probability guarantees for Type I and Type II errors, providing a theoretical characterization of the procedure. Empirically, \texttt{S2D} achieves strong and consistent performance across a range of settings, including out-of-distribution scenarios and adversarial perturbations.

Generating 3D indoor objects under fine-grained structural control is crucial and has broad applications in robotics, gaming, and simulation. Such fine-grained structural control encompasses an object’s spatial layout, geometric icons, and relative size, which together constitute the object’s core structural properties. However, existing controllable 3D generation methods do not take structural control as a primary objective, and they commonly rely on suboptimal conditions in representing structure or inadequate algorithms for maintaining structural consistency. Therefore, they fail to achieve fine-grained structural control during generation. To address this, we present structure-grounded 3D object generation, a novel perspective on 3D object generation, which decomposes 3D shapes into structural configurations and geometric details, directly conditions on the object’s structural configurations and synthesizes realistic geometric details to produce high-quality 3D objects. Building on this perspective, we introduce StructBridge, which first employs a diffusion bridge in the 3D latent space to transform comprehensive structural conditions into structurally accurate coarse meshes, followed by a structure-preserving normal refiner that enriches geometric details while preserving structure. We conduct extensive experiments and comprehensive comparisons with various conditional generation methods, demonstrating that StructBridge achieves state-of-the-art performance in both generation quality and structural control capabilities.


Structuring Open-Ended NAS: Semi-Automated Design Knowledge Structuring with LLMs for Efficient Neural Architecture Search

Yuiko Sakuma ⋅ Masakazu Yoshimura ⋅ Marcel Gröpl ⋅ Zitang Sun ⋅ Junji Otsuka ⋅ Atsushi Irie ⋅ Takeshi Ohashi

Current neural architecture search (NAS) methods are often limited by their predefined, restrictive search spaces. While recent large language model (LLM)-assisted NAS methods enable open-ended search spaces, they often suffer from inefficient exploration due to biased or low-quality design ideas. To address these issues, we propose to semi-automatically structure model design knowledge to guide the search process. Our approach first defines a high-level structural template of architectural attributes. An LLM then populates this template by analyzing papers, creating a rich and diverse search space that embodies this structured design knowledge. To efficiently explore this vast space, we introduce FairNAD, using a multi-type mutation that enables broad exploration through mutation with fair idea sampling, Pareto-aware mutation, LLM-driven iterative mutation, and a fine-grained feedback loop. We demonstrate the effectiveness of FairNAD in discovering high-performing architectures that yield 0.84, 2.17, and 2.35 points improvement on CIFAR-10, CIFAR-100, and ImageNet16-120, respectively, compared to current state-of-the-art methods.


StyleRoute: Diffusion Style Transfer via Regional Routing and Conflict-aware Projection

Xunhao Lin ⋅ Yu Sun ⋅ Xinpeng Ding ⋅ Fei Gao ⋅ Maoying Qiao ⋅ Lihuo He ⋅ Nannan Wang

Diffusion-based style transfer faces two fundamental questions: which regions should prioritize structure preservation and which are more amenable to texture transfer, namely the “where” question, and how to achieve such region-aware stylization without corrupting content, namely the “how” question. Existing methods typically rely on a single mechanism to control local structure, local texture, and global style together, making them fragile in region-rich scenes: uniform stylization strength may distort salient objects, while conservative control leaves texture-dominant regions under-stylized. We propose StyleRoute, a region-aware diffusion stylization framework that explicitly separates where style should be transferred from how the transfer should be optimized. For the where question, StyleRoute infers a region-level routing variable from structure confidence, texture confidence, and route disagreement, deciding per region whether to protect structure or to inject texture. For the how question, StyleRoute coordinates the regional routing with diffusion optimization to guide global style evolution, suppress style-content conflicts, and stabilize coarse layout. Together, StyleRoute transforms the implicit style-content trade-off into an explicit, interpretable, and region-aware transfer process. Experiments on diverse content-style pairs show consistent improvements over prior diffusion-based and region-matching stylization methods. Code will be released.


Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

Zhixuan Liu ⋅ Zhichen Dong ⋅ Yuyu Fan ⋅ Xiangtian Li ⋅ Chao Yang

Model distillation can transfer not only intended capabilities but also hidden traits from the teacher. A teacher biased by a system prompt can generate semantically clean training data, such as numeric sequences, that still causes a downstream student to inherit the hidden preference, a phenomenon known as $\textit{subliminal learning}$. Prior work establishes this phenomenon, but the transfer mechanism remains unclear, making targeted mitigation difficult. We propose and validate $\textbf{trait-direction drift}$ as the mechanism underlying subliminal learning: $\textbf{(1)}$ on the teacher side, biased teacher samples retain a measurable preference gap toward the target trait even when their semantic content is unrelated to that trait; $\textbf{(2)}$ on the student side, under a low-rank logit-linear approximation, samples with larger preference gaps induce stronger trait-aligned updates during supervised fine-tuning, and these updates accumulate over training. Guided by this mechanism, we propose $\textbf{probe-space corridor}$, a regularizer that constrains drift along a calibrated trait readout during distillation. The method substantially reduces hidden-trait transfer while preserving task performance: for example, it lowers malicious-response transfer from 29.55% to 6.47% with low main-task accuracy cost, and consistently suppresses animal-preference transfer across the main Qwen-7B preference settings. These results provide both a validated mechanistic explanation of subliminal learning and a targeted recipe for more controllable distillation.

Generative modeling has become a mainstream paradigm for learning based robot manipulation research. Among these, diffusion-based policies achieve strong performance but typically require multiple iterative denoising steps to generate final action, leading to inefficient inference. Although flow-based methods eliminate the need for iterative denoising process, existing approaches often face challenges in maintaining stable optimization and high-quality action generation. In this paper, we propose SwiftFlow, an efficient one-step policy learning framework built upon improved mean flow for robotic manipulation. Specifically, SwiftFlow formulates policy learning as a conditional generation problem by jointly leveraging 3D point cloud observations and proprioceptive robot states. It stably learns mean velocity field through a principled regression objective and optimization process, enabling more reliable action generation while preserving efficient inference. To further regularize the learned implicit generation dynamics, we construct a mean velocity smoothness regularization term, which constrains rapid local variations in the learned mean velocity field. Since this regularization is imposed only during training, it hardly introduces additional inference overhead. Extensive experiments on Adroit and MetaWorld benchmarks demonstrate that our proposed method achieves stronger task performance together with highly efficient one-step inference, validating its effectiveness for robotic manipulation tasks.

3D Gaussian Splatting enables high-fidelity, real-time novel view synthesis, yet adapting it to dynamic, evolving scenes remains challenging. Existing methods rely on fixed static-dynamic partitioning, which suffers from persistent partitioning errors and cannot adapt during optimization. To tackle this issue, we introduce SynGS, a unified 3D Gaussian framework for continual scene updating and multi-view change detection. We cast static-dynamic decomposition as a learnable task, jointly optimizing per-Gaussian change probabilities and scene representations for iterative refinement. We further propose a change-guided update strategy that modulates illumination-aware photometric corrections via learned probabilities, separating transient appearance variations from authentic scene changes. Projecting Gaussian-level probabilities to the image plane yields 3D-consistent change masks, enabling mutual optimization between reconstruction and change localization. An illumination-decoupled refinement step further improves mask quality. Evaluations on standard benchmarks show that SynGS outperforms baseline methods on scene update and change detection tasks, while fully retaining the real-time rendering capability of 3D Gaussian Splatting. Code will be released upon acceptance.

Medical visual question answering (VQA) and federated learning (FL) have gained prominence as critical methods for clinical artificial intelligence, with VQA supporting image-based diagnostic reasoning and FL allowing institutions to train models jointly without sharing sensitive patient data. However, existing approaches face substantial challenges in cross-modal vertical FL settings, where each participating institution possesses unimodal, modality-specific textual data without access to paired multimodal samples. To address this limitation, we introduce SynTeX-FL, a federated framework designed to enable cross-modal text transfer for medical VQA without sharing raw data. Moreover, SynTeX-FL applies a reconstruction-based text synthesis model that extracts clinically relevant semantics from medical images to generate cross-modal textual representations. Furthermore, the proposed method incorporates modality-specialized low-rank adaptation (LoRA) modules to enhance modality-aware representations for each imaging domain. During aggregation, the central server employs discriminator-based quality scores and source-aware weighting strategies to integrate heterogeneous client contributions. Extensive experiments on multimodal medical VQA benchmarks demonstrate that SynTeX-FL significantly outperforms current FL baselines, providing a robust and efficient solution for cross-modal reasoning in decentralized clinical environments.


T3-S2S: Training-free Triplet Tuning for Sketch to Scene Generation

Zhenhong Sun ⋅ Yifu Wang ⋅ Yonhon Ng ⋅ Yongzhi Xu ⋅ Daoyi Dong ⋅ Hongdong Li ⋅ Pan Ji

Scene generation is crucial to many computer graphics applications. Recent advances in generative AI have streamlined sketch-to-image workflows, easing the workload for artists and designers in creating scene concept art. However, these methods often struggle with complex scenes with multiple detailed objects, sometimes missing small or uncommon instances. In this paper, we propose a Training-free Triplet Tuning for Sketch-to-Scene (T -S2S) generation after reviewing the entire cross-attention mechanism. This scheme revitalizes the existing ControlNet model, enabling effective handling of multi-instance generations, involving prompt balance, characteristics prominence, and dense tuning. Specifically, this approach enhances keyword representation via the prompt balance module, reducing the risk of missing critical instances. It also includes a characteristics prominence module that highlights TopK indices in each channel, ensuring essential features are better represented based on token sketches. Additionally, it employs dense tuning to refine contour details in the attention map, compensating for instance-related regions. Experiments validate that our triplet tuning approach substantially improves the performance of existing sketch-to-image models. It consistently generates detailed, multi-instance 2D images, closely adhering to the input prompts and enhancing visual quality in complex multi-instance scenes.


TACO: Towards Task-Consistent Open-Vocabulary Adaptation in Video Recognition

Minghao Zhu ⋅ Xiao Lin ⋅ Mengxian Hu ⋅ Xun Zhou ⋅ Liuyi Wang ⋅ Xiaoyan Qi ⋅ Chengju Liu ⋅ Qijun Chen

Adapting CLIP for open-vocabulary video recognition necessitates a delicate balance between newly acquired video knowledge and the pretrained generalization. While existing studies pursue this generalization-specialization trade-off with additional regularizations or constraints, we argue that they overlook the deviation of representations beyond the fine-tuning data distribution, resulting in suboptimal adaptation effects. We believe such deviation is inherited from the inconsistency between the fine-tuning and evaluation objectives, where model optimization is restricted to the known training distribution but evaluated on unseen ones. In this paper, we introduce \emph{TACO}, a simple yet effective framework to mitigate the potential negative effects induced by this inconsistency. Our key insight is that adaptation should preserve OOD-relevant alignment beyond the training distribution. To this end, we propose \emph{Relative Structure Distillation}, which regularizes the relative geometry of the representation space and suppresses harmful alignment shift during training. We further decouple the representation space from the optimization space with a lightweight specialization projection, allowing task-specific adaptation without directly overspecializing the representations used at test time. \emph{TACO} establishes state-of-the-art performance on diverse benchmarks under cross-dataset and base-to-novel settings. Code will be released.


TagBO: LLM-Driven Task-Aware Graph Bayesian Optimization for Scientific Discovery

Xinzhe Yuan ⋅ Zimu Zeng ⋅ Zhuo Chen ⋅ Jianshu Zhang ⋅ Jinzong Dong ⋅ Jingyi Chen ⋅ Nanyang Ye ⋅ Huan Xiong ⋅ Qinying Gu

We study mixed-variable optimization in scientific discovery, where the key challenge is to efficiently explore the search space under expensive experimental evaluations. Traditional Bayesian optimization (BO) methods typically rely on structural descriptors (e.g., molecular fingerprints) to measure similarity among categorical variables. However, such task-agnostic similarity often fails to reflect task-specific objectives, leading to biased generalization. To address this limitation, we propose \textbf{TagBO} (\textbf{T}ask-\textbf{a}ware \textbf{G}raph \textbf{B}ayesian \textbf{O}ptimization), a method that integrates LLM-derived relational priors with descriptor-based structural information for mixed-variable Bayesian optimization. Rather than directly using LLM outputs as predictive scores, TagBO treats LLM outputs as a potentially noisy source of task-aware relational priors and extracts coarse ordinal relational signals from LLM reasoning to construct task-aware similarity structures for scientific optimization. TagBO then performs constrained Gaussian process Bayesian optimization over the resulting joint discrete-continuous representation space. Experiments on diverse chemistry and materials formulation tasks demonstrate that TagBO consistently improves sample efficiency under fixed evaluation budgets across LLMs of varying capabilities.

Open-Set Test-Time Adaptation (OSTTA) operates on non-stationary target streams where covariate and semantic shifts coexist. Existing TTA methods often obtain adaptation signals by updating model parameters or maintaining mutable memories, which can be costly and vulnerable to unknown samples. We present **TALOR**, a lightweight, backward-free rectification module that uses the source linear head as a geometric anchor. The head-induced basis decomposes each normalized feature into high-gain principal coordinates and low-gain tail coordinates. TALOR estimates a soft-routed affine correction in the principal subspace, combining a _regime-level_ principal bias with a tail-conditioned slope that captures _sample-level_ tail-to-principal coupling. Using weighted ridge regression, it subtracts this drift estimate from the principal coordinates, reconstructs the feature and feeds it into the source head. On the ImageNet-C benchmark with six csOOD datasets, TALOR outperforms the next-best UniEnt by **2.2%** H-score while running $\mathbf{2.5\times}$ faster and using only **12%** of its GPU memory; as a post-correction plug-in, it further boosts COME to **68.4%** H-score with only **2%** additional memory overhead.


TailDiff: Elicitable Tail-Guided Diffusion for Risk-Sensitive Generation

Shuning Zhao ⋅ Patrick Wong ⋅ Rongshang Li ⋅ Leran Zhang ⋅ Xiaolin Hu

Generative models are increasingly used to simulate risk-sensitive systems, from financial markets and energy demand to healthcare, where downstream decisions depend on rare but consequential events. In this setting, matching average behavior is not enough: rare tail events drive stress testing, capital allocation, and policy analysis. Yet tail objectives such as Value-at-Risk and Expected Shortfall are difficult to optimize directly because empirical estimators rely on sorting, thresholding, and subset selection. We introduce TailDiff, a training-free inference-time guidance method for pretrained diffusion models that uses jointly elicitable Fissler-Ziegel scoring rules to construct differentiable tail-aware objectives. We analyze the induced guidance field locally in the low-noise regime, characterizing tail-shaping and boundary-coupling behavior together with controlled perturbations under small realized effective drift. On synthetic data, TailDiff approaches the oracle finite-sample floor, on a challenging financial tail-risk simulation benchmark, it outperforms the state-of-the-art on tail metrics while preserving bulk structure.


Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation-Level Enforcement

Xutao Mao ⋅ Jianing Zhu ⋅ Jinman Zhao ⋅ Tongliang Liu ⋅ Xiaowen Chu ⋅ Cong Wang ⋅ Bo Han

Reinforcement learning (RL) has substantially improved the reasoning capabilities of vision-language models (VLMs), yet it often triggers chain-of-thought (CoT) obfuscation, a regime where models achieve high accuracy while producing reasoning traces that are ungrounded and difficult to monitor. While prior work documents this decay at the behavioral level, the underlying mechanistic drivers of such representational drift remain poorly understood. In this work, we identify that obfuscation is a representational pathology, where RL-induced optimization causes task-agnostic template features to replace visually grounded content in the model’s activation space. Based on this insight, we propose TAME, an activation-level intervention framework that leverages Sparse Autoencoders (SAEs) to regularize VLM internal features during RL training. Specifically, TAME utilizes an LLM-as-judge monitor to detect behavioral obfuscation and maps these signals to specific internal directions, penalizing template intrusion while preserving visual evidence. By directly intervening on the features responsible for obfuscation, our method restores the transparency and visual grounding of reasoning traces. Experiments on the VIRL-39k and SPA-VL benchmarks across two model families show that TAME improves CoT monitorability (by up to +30.9% on VIRL and +16.7% on SPA-VL over GRPO), while preserving task accuracy and general-capability performance across four standard VLM benchmarks.


TasteBench: multimodal benchmark for sensory prediction, from molecules to sustainable foods

Anna Thomas ⋅ Sohum Patnaik ⋅ Caroline Cotto ⋅ Benjamin Sanchez-Lengeling

Sustainable protein discovery lacks the fast computational proxies, analogous to molecular docking or density functional theory, that accelerate drug and materials discovery. Evaluating whether a novel food tastes like its animal-based target requires expensive human sensory panels, bottlenecking the design-build-test loop. We introduce \textsc{TasteBench}, a multimodal benchmark and privacy-preserving competition for sensory prediction, spanning two tasks: a food-level ranking task built on 21K+ human evaluations across 215 plant-based foods in 24 product categories, yielding 935 within-category ranking pairs, and a supporting molecular-level taste classification task over 15K flavor molecules. To enable rigorous interpretation of model performance, we characterize the ground truth: inter-rater agreement among panelists is low (Krippendorff's $\alpha$ = .077), and the split-half reliability ceiling of panel-aggregated rankings is .825, establishing the range within which ML systems on this benchmark should be assessed. We evaluate baselines across four input modalities; on the same pairs panelists rated, the best model achieves .661 pairwise accuracy, competitive with the median individual panelist (.650). TasteBench provides the evaluation infrastructure and baselines for measuring progress on computational screening for sustainable protein discovery.


Teacher-Aware Evolution of Heuristic Programs from Learned Optimization Policies

Minyu Chen ⋅ Song Qin ⋅ Ling-I Wu ⋅ Jianxin Xue ⋅ Guoqiang Li

LLM-based automatic heuristic design has shown promise for generating executable heuristics for combinatorial optimization, but existing methods mainly rely on delayed endpoint performance. We propose a teacher-aware evolutionary framework that uses independently trained learned optimization policies as behavioral teachers. Instead of deploying or imitating the teacher, our method queries it on states visited by candidate heuristic programs and uses its action preferences as local feedback for evolution. The resulting search discovers static executable heuristics guided by both task performance and teacher-derived behavioral signals. Experiments on scheduling, routing, and graph optimization benchmarks show that our method improves over performance-driven LLM heuristic evolution baselines while requiring no neural inference at deployment. These results suggest that learned optimization policies can be repurposed as behavioral feedback sources for automatic heuristic discovery.


Temporal Alignment Guidance: On-manifold Sampling in Diffusion Models

Youngrok Park ⋅ Hojung Jung ⋅ Sangmin Bae ⋅ Se-Young Yun

Diffusion models have achieved remarkable success as generative models. However, even a well-trained model can accumulate errors throughout the generation process. These errors become particularly problematic when arbitrary guidance is applied to steer samples toward desired properties, which often breaks sample fidelity. In this paper, we propose a general solution to address the off-manifold phenomenon observed in diffusion models. Our approach leverages a time predictor to estimate deviations from the desired data manifold at each timestep, identifying that a larger time gap is associated with reduced generation quality. We then design a novel guidance mechanism, `Temporal Alignment Guidance' (TAG), attracting the samples back to the desired manifold at every timestep during generation. Through extensive experiments, we demonstrate that TAG consistently produces samples closely aligned with the desired manifold at each timestep, leading to significant improvements in generation quality across various downstream tasks.

Neurobiological circuits may be energy efficient in part because there is a division of labor: different subsystems compute or represent different things. In light of this, it is interesting that a major difference between subsystems is the time scale on which they typically vary. Here, we argue that temporal smoothness constraints---the idea that it may be energetically costly for systems to vary much more quickly or slowly than their typical time scale---imply a certain division of labor with respect to temporal stimulus features. In particular, slow subsystems ought to represent slowly varying stimulus features, and fast subsystems ought to represent quickly varying features. We formulate a novel efficient coding model that we use to investigate this claim, and exactly solve for the optimal division of labor. In addition to finding that slower subsystems ought to represent more slowly varying temporal features, we find surprising structure of mathematical interest: optimal codes have subsystems which represent orthogonal temporal features, and which correspond to the eigenfunctions of a Sturm-Liouville problem.


TENET: Time-point Encoding Network for Multivariate Time Series Anomaly Detection

Dongchan Cho ⋅ Juwon Hwang ⋅ Keumyeong Kang ⋅ Jiho Han ⋅ Namsoon Jung

In unsupervised multivariate time series anomaly detection (MTSAD), local variation can inflate reconstruction and forecasting errors, making observation-space discrepancy an unreliable indicator of abnormal system behavior. Since latent-space scoring can mitigate this limitation, we reframe one-class MTSAD as current-point system-state compatibility estimation and propose TENET, the Time-point Encoding Network for Multivariate Time Series Anomaly Detection. We define the current-point system state as a latent representation of the target time point, constructed from variable-level temporal context in the input window. TENET constructs this state through a temporal layer that retrieves window context relative to the current point, encodes time with current-point-centered embeddings, and adaptively mixes the retrieved context with the current representation. With this design, TENET treats the input window not as the object to reconstruct, forecast, summarize, or score, but as context for constructing a latent representation evaluated with a fixed standard-normal negative log-likelihood score. Across benchmark datasets, TENET achieves state-of-the-art performance. Additional experiments further show that TENET remains effective under alternative one-class objectives and outperforms substitutions based on generic sequence encoders, suggesting that its advantage comes from decision-aligned representation construction rather than from a specialized scoring rule. These results indicate that one-class MTSAD can be strong when the latent representation is designed around the current time point.


Testable and Actionable Calibration for Full Swap Regret

Konstantina Bairaktari ⋅ Lunjia Hu ⋅ Huy Nguyen ⋅ Jonathan Ullman

AI generated predictions increasingly inform decision making in critical tasks, and therefore must be trustworthy. One widely used measure of trustworthiness is calibration, which requires that the predictions match the true frequencies and can be treated like real probabilities of a given outcome. However, defining calibration is subtle, and designing good measures of calibration error has been an active topic of recent research. The first goal is to find calibration measures that are actionable, meaning they can inform decision makers about their utility loss when predictions are treated as true probabilities, which is known as swap regret. The second goal is to find calibration measures that are testable, meaning that calibration error can be measured from a small sample of predictions and outcomes. Although these are very basic requirements, there is no existing calibration measure that fully satisfies both properties, and all existing measures relax actionability by bounding a weaker notion of swap regret, or relax testability by having suboptimal estimation error. We introduce a new calibration measure, Soft-Binned Calibration Decision Loss (SCDL), which we prove is fully actionable without weakening either requirement, and testable with nearly optimal error rate. In addition, SCDL satisfies other desired properties such as continuity and consistency. We also provide a set of experiments confirming that the theoretical advantages of SCDL compared to other measures lead to better performance in practice.


Test Time Search Requires Training for Diversity

Ryan Bahlous-Boldi ⋅ Isha Puri ⋅ Idan Shenfeld ⋅ Akarsh Kumar ⋅ Mehul Damani ⋅ Sebastian Risi ⋅ Omar Khattab ⋅ Zhang-Wei Hong ⋅ Pulkit Agrawal

Language models are increasingly deployed alongside test-time search. We argue that we should therefore shift the role of RL post-training from converging on a single best response to producing a diverse pool of competent candidates. Standard policy gradient methods such as GRPO optimize a fixed scalar reward and drive the policy toward near-duplicate responses, erasing the diversity search needs. We propose Vector Policy Optimization (VPO), which exploits the fact that practical rewards are vector-valued, e.g., per-test-case correctness, per-criterion ratings, or per-hop credit. Instead of collapsing these into one scalarization, VPO samples randomized scalarizations from a distribution over the reward simplex, incentivizing candidates to specialize to different trade-offs along the Pareto front. The model emits multiple candidates per prompt, with each conditioning on the previous ones, allowing for directed diversification. VPO is a drop-in replacement for the GRPO advantage estimator. Across four domains, VPO consistently improves test-time best@$k$ over scalar baselines, with gains widening as the test-time budget grows.

Test-time training (TTT) enhances model performance by explicitly updating designated parameters prior to each prediction to adapt to the test data. While TTT has demonstrated considerable empirical success, its theoretical underpinnings remain limited, particularly for nonlinear models. In this paper, we investigate the combination of TTT with in-context learning (ICL), where the model is given a few examples from the target distribution at inference time. We analyze this framework in the setting of single-index models, where the feature vector is drawn from a hidden low-dimensional subspace. For single-layer transformers trained with gradient-based algorithms and adopting TTT, we establish an upper bound on the prediction risk. Our theory reveals that TTT enables the single-layer transformers to adapt to both the feature vector and the link function, which vary across tasks. This creates a sharp contrast with ICL alone, which is theoretically difficult to adapt to shifts in the link function. Moreover, we provide the convergence rate with respect to the data length, showing the predictive error can be driven arbitrarily close to the noise level as the context size and the network width grow.

Contrastive vision-language models (VLMs) are deployed in a growing range of retrieval and grounding systems, making their robustness to input corruption a critical property. In this paper, we apply controlled corruptions to textual and image inputs for five contrastive VLMs spanning three Vision Transformer architectures, four training corpora, and two contrastive loss functions. We aim to measure how image and text representations and retrieval performance respond as input damage increases. Previous cross-modal robustness work often applied ordinal severity scales to corruptions independently, even though the same severity step on the image side may not represent equivalent damage on the text side or vice versa. Our experiment addresses this measurement gap using within-modality damage percentile binning, which ranks each side's corruptions on its own damage scale and compares only at matched ranks. We evaluate all models on MS-COCO retrieval with $n = 1{,}000$ pairs over three seeds, six corruption types, and five severity levels. Our results indicate that image-to-text (i2t) Recall@1 at the highest damage quintile is 4.6$\times$ to 5.8$\times$ higher than text-to-image (t2i) Recall@1 at the matched quintile. Two mechanistic probes are used to localize part of the asymmetry to the single-token text-pooling bottleneck and to text damage that destroys lexical content rather than word order. These results suggest that downstream applications retrieving images from text are notably more fragile than the reverse, and that future improvements to cross-modal robustness require architectural attention to text readout.


The $1/\mathcal{W}$ Law: Context Length is the Dominant Energy Lever in LLM Inference Fleets

Huamin Chen ⋅ Xunzhuo Liu ⋅ Yuhan Liu ⋅ Junchen Jiang ⋅ Steve Liu ⋅ Chenxu Niu ⋅ Bowei He

Energy efficiency in LLM inference is often treated as a property of the model, the hardware, or the serving software. We show that, for decode-dominated serving, a simpler systems variable dominates: the serving context window. Because KV-cache memory is finite, increasing the context window reduces the number of concurrent sequences a GPU can hold. Under continuous batching, this concurrency limit directly sets decode throughput, while GPU power remains nearly flat over most of the operating range. The result is a simple $1/\mathcal{W}$ law: tokens per watt scale inversely with the serving context window $\mathcal{W}$. We derive this law from KV-cache capacity and a logistic GPU power model, and empirically verify it for Llama-3.1-70B inference over the 2K--128K context range. On identical H100 hardware and software, tok/W varies by nearly $40\times$ across this range. We then show that the same mechanism determines GPU fleet-level energy efficiency. Context-length routing keeps short requests on high-concurrency short-window pools, while newer hardware shifts the power and memory curve upward. These two operator-controllable levers are approximately multiplicative: on the short-context-dominant Azure trace, two-pool context-length routing improves tok/W by $2.5\times$ over a homogeneous NVIDIA H100 fleet, a calibrated NVIDIA B200 deployment improves tok/W by $2.0\times$ at fixed topology, and combining them reaches $5.76\times$. Finally, we analyze active-parameter weight streaming in MoE models as a third, architecture-level lever, giving an upper-bound gain of up to $5.1\times$. Together, these results suggest a deployment order for energy-efficient LLM serving: route by context length first, upgrade hardware second, and exploit sparse architectures where dispatch costs are controlled. We release code, data, and simulator in anonymous [repository](https://anonymous.4open.science/r/gpu-fleet-sim-1BEB) to facilitate reproduction and further research.


The Commit-Abstain Circuit: Why Language Models Hallucinate Instead of Abstaining

Gavin Vy Nguyen ⋅ Ziqi Xu ⋅ Jeffrey Chan ⋅ Estrid He ⋅ Feng Xia ⋅ Renqiang Luo ⋅ Erik Cambria ⋅ Xiuzhen Zhang

Language models (LMs) often hallucinate by committing to substantive answers when they should abstain. Existing methods detect hallucinations or guide abstention, but leave open how models internally decide to commit or abstain. We study this decision through mechanistic analysis, framing a class of hallucination as unsupported commitment: the model commits despite exhibiting signals of unanswerability. Using causal gating, we identify a Commit-Abstain Circuit (CAC), a sparse subset of attention heads and MLP sublayers that causally contribute to this decision. Across ten LMs (3B-14B) from five families and three benchmarks, the CAC exhibits a recurring accumulate-yet-undercorrect pattern: commitment-promoting components build up commitment in earlier layers, while abstention-promoting components act later as corrective signals that are often insufficient to overturn the accumulated commitment. Building on this finding, a lightweight policy trained on CAC activations improves decision accuracy by 12.2 points over the model's intrinsic commit-abstain margin, reduces false abstentions by 2.5 times, transfers to unseen benchmarks, and extends to larger models (27B-35B). The CAC is both diagnostic, clarifying how models overcommit, and practical, enabling improved abstention decisions. Code is provided in the supplementary material.

Permutation-invariant machine learning architectures, such as DeepSets, Set Transformers, Graph Neural Networks, and $k$-GNNs, are each highly used and studied individually. We unify them into a single complexity-theoretic framework: noisy symmetric circuits. Each architecture falls into our hierarchy at a level $k$, parameterized by the order of interactions it captures. Within this framework, we tightly characterize the cost of symmetry. At level $1$, where DeepSets operate, every symmetric function on boolean input computable by a neural network has an equivalent DeepSet, at the cost of an additive blowup nearly linear in the input size, and we exhibit an explicit function that requires this blowup. At level $2$, where Set Transformers and GNNs operate, the same near-linear blowup suffices for WL-invariant functions, while non-WL-invariant functions are known to be uncomputable at level $2$. These results show that, in the regimes we cover, WL is the \textit{only} barrier to using a symmetric architecture; everywhere else, symmetry is \textit{low-cost}. Experiments illustrate the near-linear threshold predicted by the theory.


The Illusion of Multi-Agent Advantage

Prathyusha Jwalapuram ⋅ Hehai Lin ⋅ Chuyuan Li ⋅ Fangkai Jiao ⋅ Sudong Wang ⋅ Yifei Ming ⋅ Zixuan Ke ⋅ Chengwei Qin ⋅ Giuseppe Carenini ⋅ Shafiq Joty

Prevailing wisdom posits that Multi-Agent Systems (MAS) are superior to Single-Agent Systems (SAS), citing advantages like context protection, parallel processing and distributed decision-making. However, empirical support for this claim relies primarily on comparisons with SAS baselines using benchmarks that prioritize isolated reasoning tasks, which do not adequately assess these advantages. Focusing on automatically generated MAS that are designed for enhanced generalizability over manually-designed counterparts, we perform a rigorous, systematic evaluation against SAS, specifically Chain-of-Thought with Self-Consistency (CoT-SC). Across traditional reasoning datasets and tasks with interactive multi-step workflows (e.g., BrowseComp-Plus), we demonstrate that automatic MAS consistently underperform CoT-SC despite being up to 10x more expensive. To isolate these failures from limitations inherent to task structure, we introduce a diagnostic synthetic dataset tailored for MAS featuring explicit task decomposition, context separation and parallelization potential. We show that expert-architected MAS consistently outperforms automatically generated architectures in both raw performance and cost-efficiency on this dataset, demonstrating that existing evaluation frameworks mask critical architectural gaps and inefficiencies of complex MAS by failing to account for the marginal utility of increased computational cost. Critically, systematic deconstruction of the generated MAS architectures reveals that current automated design paradigms produce architectural bloat that prioritizes superficial complexity which does not translate into functional utility, exposing a fundamental misalignment with multi-agent principles.


The Limits of AI-Driven Allocation: Optimal Screening under Aleatoric Uncertainty

Santiago Cortes-Gomez ⋅ Mateo Dulce Rubio ⋅ Carlos M Patiño ⋅ Bryan Wilder

The rise of machine learning has shifted targeted resource allocation in policy and humanitarian settings toward algorithmic targeting based on predicted risk scores. This approach is typically cheaper and faster than traditional screening procedures that directly observe the latent vulnerability status through physical verification. Yet, even access to the true conditional vulnerability probability cannot eliminate misallocation: aleatoric uncertainty over individual vulnerability status is irreducible, and probabilistic targeting inevitably misallocates some resources. In this work we study how screening and algorithmic targeting should be optimally combined in a two-stage allocation framework where a screening stage observes true outcomes for a subset of units before a final allocation stage assigns the resource under a fixed coverage budget. We show that the optimal strategy screens units at the margin of algorithmic allocation, while directly targeting the highest-risk units. Furthermore, we empirically characterize when screening and algorithmic targeting act as complements or substitutes: efficiency gains from screening grow as the aleatoric uncertainty in the population increases. We illustrate our framework with applications in income-based social protection programs and humanitarian demining in Colombia, where the tension between screening costs and allocation efficiency is operationally consequential.


The Locality Cost of Semantic Patch Self-Distillation

Xi Weng ⋅ Zhaoyu Liu ⋅ lianyu hu ⋅ Jin Song Dong

Masked patch self-distillation has become a default ingredient of modern joint-embedding pretraining, and is widely regarded as a uniform enhancement of dense visual representations. In this work, we revisit this assumption and argue that its effect is more nuanced. Through controlled experiments across multiple joint-embedding objectives and model scales, we observe a consistent tension: while the auxiliary patch objective strengthens semantic abstraction, it systematically erodes fine-grained spatial correspondence. To examine this phenomenon beyond the reach of human-annotated benchmarks, we introduce annotation-free diagnostics that probe how much local visual evidence remains accessible in the learned representation. We further provide a population-level theoretical analysis showing that this trade-off is not incidental but follows directly from the conditional nature of masked semantic prediction, under which target variation unpredictable from visible context is provably averaged out. Taken together, our findings recast masked patch self-distillation as an explicit trade-off between semantic abstraction and local detail, rather than a uniform improvement to dense features.


The Mask Is Not the Object: Volumetric Supervision for 3D Gaussian Segmentation

Jeonghwan Cho ⋅ Minsu Kim ⋅ Junyoung Hong ⋅ Seon Joo Kim

Existing 3D Gaussian Splatting (3DGS) segmentation methods produce accurate rendered masks yet often fail to select a complete set of object Gaussians: removing the predicted Gaussians leaves the rendered depth in the target region nearly unchanged. We trace this gap to the contribution-weighted supervision path of rendering-based feature learning. Because Gaussian geometry and opacity are pre-optimized for RGB reconstruction, rendered-feature losses deliver strong gradients only to high-contribution Gaussians, leaving other Gaussians in the same object region weakly supervised and ultimately unselected. We propose VODA, a plug-and-play supervision objective that addresses this bias. After RGB optimization, Gaussians of the same object portion concentrate at similar depths along the camera ray. A depth-aware rasterizer exploits this property, grouping per-pixel contributing Gaussians into depth-coherent clusters and selecting the cluster that best explains each object pixel. The rendered feature of the selected cluster is then applied as a target directly to every Gaussian within it, bypassing the contribution-weighted gradient path. A complementary pixel-level loss further sharpens boundaries between adjacent objects. To evaluate volumetric completeness, we introduce Removal-3D and the Removal Depth Error (RDE), which measure whether removing the predicted Gaussians induces the expected depth change in 3D. Across four rendering-based baselines, VODA consistently improves volumetric segmentation on Removal-3D while maintaining or improving standard performance on LERF and ScanNet.

Kotłowski et al. (NeurIPS 2017) posed as a central open problem the design of efficient online isotonic regression algorithms beyond totally ordered domains. We resolve this for the product order $[m]\times[m]$, the canonical two-dimensional case, by determining the minimax regret up to constant factors: $$ R_T^* = \Theta\left(\min\left\\{T,\ \inf_{K \in \mathbb{N}_+}\left[\log\Omega([m]^2, K+1) + \frac{T}{K^2}\right]\right\\}\right), $$ where $\Omega(\mathcal{P}, K+1)$ is the order polynomial counting monotone maps from $\mathcal{P}$ to $[K+1]$. This unified formula reveals a three-phase scaling law: $\Theta(T)$ for $T \lesssim m$, $\Theta(m^{2/3} T^{1/3})$ for $m \lesssim T \lesssim m^4$, and $\Theta(m^2 \log(T/m^4))$ for $T \gg m^4$. We further construct polynomial-time algorithms achieving this rate in every regime, including a horizon-free variant requiring no advance knowledge of $T$.


The Quiet Prompt: Erasing Ineffable Styles from Diffusion Models via Concept Leakage-aware Negative Guidance

Kiyun Park ⋅ Min Hee Cha ⋅ Hyeok Nam ⋅ Jae Hyeon Park ⋅ Seonho Lee ⋅ Jua Han ⋅ Hyunse Lee ⋅ Sung In Cho

Recent text-to-image diffusion models (T2I DMs) depend on vast, uncurated datasets that often include copyrighted artworks and personal images, risking the generation of unwanted content. While concept erasure methods have emerged to suppress such outputs, they predominantly rely on text prompts, limiting their effectiveness for visually nuanced or ineffable styles that are difficult to articulate verbally. To bridge this gap, we propose Quiet Prompt (QuP), a reference-based concept erasure method that identifies complex concepts using fewer than ten reference images. QuP operates in two stages: (1) Concept Embedding Generation (CEG), which captures the intricate visual morphology of an unwanted concept into a single latent embedding; and (2) Concept leakage-aware Negative Guidance (CNG), which utilizes this embedding as a negative semantic condition to steer the diffusion process. In addition, to overcome the limitations of conventional fixed-scale negative guidance (negative prompting)—which often suffers from structural distortion and unintended content loss—we introduce concept leakage-aware negative guidance scale $n_g(t)$. This mechanism adaptively modulates the guidance strength at each denoising step based on a semantic measurement of concept leakage between the intermediate generated image and the target concept in a joint embedding space. Extensive evaluations on idiosyncratic art styles and object erasure tasks demonstrate that QuP achieves superior concept suppression and content preservation, outperforming state-of-the-art text-based baselines in both image fidelity and erasure precision. The source code will be publicly provided to ensure reproducibility of our work.


The Reciprocity Gradient

Yue Lin ⋅ Pascal Poupart ⋅ Shuhui Zhu ⋅ Dan Qiao ⋅ Wenhao Li ⋅ Yuan Liu ⋅ Hongyuan Zha ⋅ Baoxiang Wang

Communication is fundamental to sustaining reciprocity and cooperation in strategic interactions. We identify and formulate the influence attribution problem as the central optimization difficulty inherent in such dynamics for a learning agent: any action or signal the agent emits reshapes the reputations of many third parties along combinatorially branching paths before feeding back into its own future rewards, forcing the agent to account for all of these indirect channels at once when choosing every action. To address this, we introduce the reciprocity gradient, which explicitly backpropagates reward gradients through private estimators of opponents' policies trained from public observations. The gradient flows through the reputation chain itself analytically, rather than being estimated from sampled returns. It jointly optimizes actions and evaluative signals without intrinsic rewards or reward shaping. Empirically, the method recovers near-optimal context-sensitive policies, while sample-based baselines collapse into constant-output policies.


The road reaches every place, the short cut only one: Self-Adversarial Shortcut Mitigation for AI-Generated Image Detection

Haifeng Zhang ⋅ Qinghui He ⋅ Xiuli Bi ⋅ Bo Liu ⋅ Chi-Man Pun ⋅ Bin Xiao

With the rapid advancement of generative models, AI-generated images (AIGIs) have become increasingly realistic, posing new challenges for reliable detection. In this work, we revisit the root causes of poor generalization in AIGI detectors from the perspective of shortcut learning. We find that detectors often rely on spurious semantic cues or a single prominent artifact, causing the feature space to collapse toward low diversity and severely limiting generalization. To characterize this issue, we propose the Generative Artifact Manifold Clustering (GAMC) hypothesis, which shows that although generative artifacts are multi-dimensional, different generators occupy distinct distributions along these dimensions, making detectors that rely on local artifact cues inherently fragile. To address this, we shift AIGI detection from “single-cue reliance” to “multi-dimensional discrimination” and introduce the Semantic–Artifact Self-Adversarial Feature Learning (SA-AFL) framework. SA-AFL progressively disentangles semantics from artifacts and employs a self-adversarial mechanism to balance learning across artifact dimensions, enabling the model to retain richer and more shared generative features. As a result, SA-AFL reduces shortcut reliance and learns more transferable evidence for cross-generator detection. Comprehensive experiments on images from 23 generative models show that SA-AFL improves ACC by 6.84\% and mAP by 8.25\% over state-of-the-art methods.

The mechanism governing the training dynamics of Quantum Neural Networks (QNNs) remains under-explored. In classical Deep Neural Networks (DNNs), training is known dominated by "Spectral Bias,'' i.e. prioritizing learning low-frequency while struggling with high-frequency. QNNs exhibit similar spectral limitations when the target function possesses flat spectral amplitudes. However, in this work, considering target problems with different spectral amplitudes, we theoretically and empirically identify a distinct mechanism in QNNs, which we term Spectral Amplitude Priority. By analyzing the frequency-domain gradients and residual dynamics via the Quantum Neural Tangent Kernel (QNTK), we prove that QNN training is governed primarily by the magnitude of spectral components rather than their frequency indices. Consequently, QNNs can efficiently capture high-frequency functions—provided they have significant amplitude—thereby overcoming the inherent limitations of their classical counterparts. We validate this principle on both synthetic high-frequency functions and quantum-advantage tasks. These results show that QNNs notably outperform DNNs in high-frequency tasks, offering an explanation for QNNs' superior expressivity.


The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval

Zekai Tong ⋅ Ruiyao Xu ⋅ Aryan Shrivastava ⋅ Chenhao Tan ⋅ Ari Holtzman

Existing Large Language Model (LLM) benchmarks primarily focus on syntactically correct inputs, leaving a significant gap in evaluation on imperfect text. In this work, we study how word-boundary corruption affects how LLMs detect targeted information. By inserting whitespace characters within words to break them into fragments, LLMs' detection accuracy follows a U-shaped curve with the increase in insertion rate. We refer to this curve as the Text Uncanny Valley. To explain such observation, we propose a mode transition hypothesis: LLMs operate in a word-level mode for near-normal text and a character-level mode for heavily fragmented text, with the valley marking the disordered transition where neither mode is effective. Four experiments and one analysis are consistent with this account: in-context learning fails to rescue valley-bottom performance; regularizing the perturbation substantially reduces the U-shape; a math reasoning task replicates the U-shape for Gemini 3.0 Flash but not for stronger models, suggesting the effect is attenuated when tasks rely less on exact lexical alignment; and tokenization entropy peaks before the F1 minimum, consistent with a regime-conflict interpretation. These findings reveal a failure mode invisible to clean-text benchmarks yet directly relevant to any deployment scenario involving noisy or uncurated text inputs.


Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance

Jiachen Yu ⋅ Zhihao Xu ⋅ junjie wang ⋅ Yujiu Yang

Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. However, existing frameworks typically treat rubrics only as external evaluator disjointed from the policy's primary reasoning trace. Such design confines rubrics to post-hoc measurement, leaving them unable to actively guide the model's generation process. In this work, we introduce Think-with-Rubrics, a novel paradigm for instruction following tasks. Think-with-Rubrics integrates rubric generation into the reasoning context, transforming the rubric from an independent artifact into an internal guidance of LLM's generation. During training, LLM sequentially generates a rubric followed by a response, while a trained rubric verifier provides joint supervision by evaluating the consistency between the answer and the self-generated / golden rubrics. Experiments across multiple benchmarks demonstrate that Think-with-Rubrics consistently outperforms the Rubric-as-Reward baseline supervised by golden rubrics by an average of 3.87 points. We have also discussed the mechanism by which Think-with-Rubrics enhances model performance. Experimental results demonstrate that supervision from golden rubrics and self-generated rubrics enhances the performance of Think-with-Rubrics by improving the quality of self-generated rubrics and increasing the internal consistency of responses respectively.


Thompson Sampling using Prior-fitted Diffusion Transformers

Sihwa Park ⋅ Jingsen Zhu ⋅ Vinamr Jain ⋅ Sheng-Yen Chou ⋅ Alexander Terenin

Thompson sampling is natural randomized strategy for optimizing unknown functions in a black-box manner. In continuous global optimization, Thompson sampling has largely been restricted to Gaussian process priors, as these are essentially the only class for which computing the posterior distribution has been tractable in practice. In this work, we introduce Prior-fitted Diffusion Transformers, which are diffusion models that perform Thompson sampling at inference time, after being pre-trained using the prior-fitted network paradigm. To apply these models, we develop a set of practical non-Gaussian priors for general black-box functions - including those that have characteristics that cannot be achieved under Gaussianity, such as priors over unimodal functions. We also quantify the tradeoff between model size and the number of pre-training iterations, and show that these models follow scaling laws akin to those observed in transformers defined over other modalities. Across a set of global optimization benchmarks, we show that prior-fitted diffusion transformers with non-Gaussian priors can achieve improved performance compared to traditional pipelines built on Gaussian processes.

We study cost-aware cascading bandits, where a learner selects an ordered subset of options, tests them sequentially until the first success, and pays the costs of all tested options. In this problem, regret comes from both testing inefficient options and placing options in a suboptimal order, but existing analyses do not separate these effects for inefficient options and therefore yield an inverse-square dependence on the gap $c_i-\theta_i$. We develop a regret decomposition based on intermediate policies that reorder the remaining suffix and remove inefficient options one position at a time. This allows us to quantify the incremental regret incurred when each option is tested. As a consequence, we show that the regret of CC-UCB admits a problem-dependent bound of $O(\sum_{i:\theta_i/c_i<1}\log T/(c_i-\theta_i))$ up to additive terms and a problem-independent bound of order $\tilde O(L\sqrt{T})$. We further prove that the minimax regret is bounded from below by $\Omega(\sqrt{LT})$ by reducing standard multi-armed bandits to a special case of the model. Finally, we propose CC-UCBv2, which removes the need to specify a positive lower bound on costs and handles zero-cost options by separating empirically zero-cost options from the others. Numerical experiments show the effect of misspecified cost lower bounds and demonstrate that the proposed modification can reduce regret in representative instances involving zero or misspecified costs.


ToolSearcher: Optimizing Tool Selection at Scale via Reinforcement Learning

Zhenlong Dai ⋅ Xujie Song ⋅ Zitong Wang ⋅ Tong Niu ⋅ Jian liu ⋅ Weiqiang Wang ⋅ Xiu Tang ⋅ Sai Wu ⋅ Chang Yao ⋅ Jingyuan Chen

Large language models (LLMs) excel at natural language processing but struggle to interact with external environments. Tool learning provides a promising way to extend LLMs into actionable agents, where tool selection is a critical prerequisite for successful tool use. Existing work often assumes a small or predefined set of tools, leaving large-scale tool selection underexplored. Real-world repositories contain a vast and diverse array of tools, making it difficult for LLMs to effectively search, distinguish, and compose tools under context-length constraints. We identify large-scale tool selection as a new challenge for agentic reinforcement learning, highlighting that existing RL methods for knowledge-based question answering are inadequate for selecting tools while considering compatibility. To address this challenge, we propose ToolSearcher, a novel RL framework for effective multi-turn search and fine-grained optimization in large-scale tool selection. Specifically, we introduce category-constrained tool discrimination to improve the model's ability to distinguish functionally similar tools, event-level search modeling to explicitly optimize the discovery of target tools during multi-turn search, and trajectory-aligned credit allocation to provide fine-grained reward signals for different stages of the search-selection process. Extensive experiments on large-scale tool selection benchmarks demonstrate that ToolSearcher consistently outperforms a set of strong baselines in challenging settings involving iterative search and complex tool composition.


Topic-Aware Contextual Cascading Bandits

Hyun-jun Choi ⋅ Taehyun Hwang ⋅ Min-hwan Oh

We propose a contextual cascading bandit model in which the click probability at each position depends on both the displayed item and the topics of previously shown items. This captures how previously displayed topics affect later click probabilities, allowing different topic orders to induce different user responses. Unlike standard cascading bandits, the optimal cascade is no longer determined only by the selected set of items, since the ordering itself affects the expected reward. To address the resulting planning challenge, we reformulate cascade construction as a Markov decision process whose state summarizes the previously selected topics. Based on this formulation, we develop an optimism-based algorithm and prove a regret upper bound of $\widetilde{\mathcal O}(\bar p^{\frac{K-1}{2}}d\sqrt{T})$, where $d$ is the number of unknown parameters, $T$ is the horizon, $K$ is the cascade length, and $\bar p<1$ upper bounds the no-click probability of an examined item. A key implication is that the regret decreases with the cascade length $K$, a phenomenon that had not been established even in order-insensitive contextual cascading bandits. We further provide a local regret lower bound showing that this decreasing dependence on $K$ is intrinsic, and validate our theory through experiments with topic-order-dependent click probabilities.


TopoRefine: Topology-Aware Correspondence and Residual Refinement for Training-Free Subject-Consistent Generation

Zhanxin Gao ⋅ Zexin Ti ⋅ Chen Zhao ⋅ Beier Zhu ⋅ Xiantao Hu ⋅ Yanhao Ge ⋅ Jian Yang ⋅ Ying Tai

Subject-consistent generation (SCG) aims to preserve subject identity across diverse text-driven contexts. Recent training-free methods move beyond global sharing via correspondence-aware identity transfer. However, semantic-driven correspondence may transfer identity features to appearance-similar but structurally incompatible regions. Meanwhile, extended attention, commonly used in existing methods, enhances identity consistency but may perturb the emerging target structure, from global pose and layout to local structural details. Together, these issues can cause reference-pose over-copying, local structural artifacts, and degraded structural fidelity. We view training-free SCG as controlled identity injection, where the pretrained model determines the target structure, while the reference image provides subject identity. We propose TopoRefine, a training-free framework with Topology-aware Subject Correspondence (TSC) and Residual Identity Refinement (RIR). TSC regularizes semantic matching with subject-internal geodesic relations for structurally compatible identity transfer, while RIR injects reference identity as a foreground-gated, magnitude-aligned residual to preserve the emerging target structure. Both qualitative and quantitative results show that TopoRefine improves subject consistency while better preserving target-driven structure.

Multimodal reasoning systems increasingly assume that longer chain-of-thought uniformly improves downstream grounding. We show that this assumption fails in referring audio-visual segmentation: simple queries can be harmed by excessive reasoning, creating an “overthinking trap” that dilutes visual focus, while genuinely ambiguous queries still require multi-step reasoning. To systematically study this, we construct counterfactual reasoning-budget labels under a rigorous, leakage-free evaluation protocol. Our results recast adaptive multimodal reasoning as a pre-decisional representation problem: before the first generated token, the generation-onset state contains linearly decodable information predictive of whether additional reasoning will improve segmentation. A linear probe predicts binary reasoning need with 70.1% accuracy, substantially outperforming text-only features and uncertainty-based proxy signals. As a router, the onset-state controller improves over downstream confidence routing and closely matches Short-then-Decide while using far fewer tokens. Mechanistic decomposition suggests that this signal is better captured by distributed high-dimensional representations than by isolated manual heuristics. Reusing this onset state for lightweight, single forward-pass routing reduces autoregressive token consumption by 60% while retaining ~96% of always-long performance. It further transfers to out-of-distribution (OOD) benchmarks without target-split tuning. Comprehensive error analysis shows that bypassing unnecessary reasoning can reduce phrase and localization drift. Ultimately, our results indicate that adaptive multimodal reasoning can be reliably and efficiently routed from pre-decisional internal representations.

Machine learning models, particularly large foundation models, are increasingly evaluated along multiple axes that are relevant for their usage. In addition to accuracy measures appropriate for the specific task of interest, measures capturing safety, calibration and fairness properties are often of interest. The classical approach to these problems assumes that the test set has been held out from the training process, an assumption that is increasingly harder to justify for large foundation models. In this work, we formalize a notion of evaluability without this assumption, and ask if it is feasible to evaluate black-box models in settings where our test samples may have been used during training. We study this problem under different settings, capturing different levels of access to the model. We first study the natural setting, where the evaluator only looks at the output of the model on the test set. We fully characterize the class of distributions where accurate evaluation is possible, and show tight upper and lower bounds on the sample complexity of evaluation. Our results show that black-box evaluation in this set up is feasible if and only if the distribution is close to being small support. We then consider evaluation algorithms that can query the model on additional inputs and demonstrate connections to self-correctors studied in program testing. Using these, we show that a small amount of query access can increase the power of the evaluator for natural distributions and concept classes.


Toward Better Geometric Representations for Molecule Generative Models

Shaoheng Yan ⋅ Zian Li ⋅ Cai Zhou ⋅ Qiaojing Huang ⋅ Kai Liu ⋅ Muhan Zhang

Geometric representation-conditioned molecule generation provides an effective paradigm that decouples molecule representation modeling from structure generation. By decoupling molecule generation into two stages—first generating a meaningful molecule representation, and then generating a 3D molecule conditioned on this representation—the efficiency and quality of the generation process can be significantly enhanced. However, its effectiveness is fundamentally limited by the quality of the representation space: pretrained molecular encoders, such as UniMol, produce representations that are non-smooth and not fully exploited during the generative training process. In this work, we propose LENSEs, a framework that better exploits the potential of molecule representations in representation-conditioned generation methods. In particular, LENSEs introduces three complementary mechanisms: (1) a representation head, simultaneously trained during generative tasks, that extracts multi-level representations from the pretrained encoder; (2) a molecule perceptual loss that optimizes the generator in a semantic-informative representation space; and (3) a node-level representation alignment (REPA) loss that explicitly aligns the generator’s hidden states with encoder representations, reducing the semantic gap between pretraining and generation. We demonstrate the effectiveness of these improvements through extensive molecule generation tasks. Specifically, on the challenging molecule generation dataset GEOM-DRUG, LENSEs achieves 97.28% validity and 98.51% molecule stability, surpassing existing advanced methods. Further analyses through Lipschitz constant reduction (4.6×) and QM9 probing tasks also demonstrate the smoother, more informative refined representations, establishing generative training with alignment objectives as a potential pretraining paradigm for molecular encoders.


Toward Interactive Understanding of Code APIs

Dhananjay Ashok ⋅ Jesse Thomason ⋅ Jonathan May

Empowered by advances in Language Model agents, systems have made substantial strides in code generation and understanding. However, these approaches often rely on access to the relevant code, an assumption which does not hold when dealing with external APIs. In this work, we introduce the PAU (Python API Understanding) benchmark, where we provide models with black-box, API-level access to code snippets. Models must query the API with exploratory inputs and draw insights from the resulting outputs, with the goal of describing the snippet's true functionality. By treating the code snippets as external tools that must be understood via interaction alone, PAU studies the more general problem of unsupervised tool understanding, specifically for tools implemented as Python methods. Despite recent progress in coding agents, even frontier models struggle to achieve high performance on PAU, with the best model (Claude-4-Opus) failing to understand over 45% of the PAU test set. An investigation into the common error modes reveals that models are overconfident; they often overrate the quality of their current hypothesis, leading to insufficient exploration and premature termination. Finally, we take inspiration from the Asymmetric Actor Critic (AAC) paradigm, frequently used in robot learning, to post-train models for interactive code understanding. Models trained with AAC conduct more active exploration of the APIs, with an AAC-tuned Qwen3-8B model matching the performance of GPT-5-mini.

Classification with rejection (CwR) generalizes the standard classification task by providing an extra option of rejection with a cost $c$ smaller than misclassification costs, and the design of its convex calibrated surrogate loss is our central topic. While typical surrogate losses for $K$-class classification operates on $K$-dimensional scoring functions, it has been proved that this dimensionality can be reduced to $\lceil\log_2K\rceil$ for CwR with $c\in(0,0.5]$, rendering more efficient computation. This logarithmic dependency is commonly conjectured to be the lowest possible among any convex calibrated surrogate losses for CwR. Through the lens of property elicitation, we explore even lower-dimensional structures for CwR, by constructing a $2$-dimensional convex calibrated surrogate loss under $c\in(0,0.5)$. This loss dimension is provably optimal among all convex calibrated surrogate losses. The boundary cost $c=0.5$ is more nuanced---we construct a $4$-dimensional convex calibrated loss, improving over the $\lceil\log_2 K\rceil$-dimension when $K>16$. Notably, this breaks the optimal logarithmic dimensionality attainable by the *embedding framework* [Finocchiaro et al., 2024], a widely used loss design principle for polyhedral losses, negatively resolving their conjecture on the dimension optimality.

Lifelong person re-identification (LReID) requires models to continuously learn from sequentially arriving domains while retaining discriminative power for previously seen identities. A key challenge is to prevent catastrophic forgetting without access to old data, especially under exemplar-free constraints. While flat-minima optimization has shown promise in continual learning, we identify a gradient conflict that limits the effectiveness of standard Sharpness-Aware Minimization (SAM) in LReID. The ReID loss gradient dominates the perturbation direction, causing the sharpness of the distillation loss to be underestimated, which hinders the flattening of the landscape necessary for knowledge retention. To resolve this conflict, we propose a framework that unifies selective flatness-aware optimization, dual-model training, and weight-space interpolation. Specifically, we maintain two models per task: a stability model whose SAM perturbation is computed solely from the distillation loss, and a plasticity model optimized for the current domain. By decoupling the perturbation objectives, our selective SAM achieves more targeted sharpness exploration along the distillation landscape, guiding the stability model toward flatter and robust regions. After training, the two models are fused via weight-space interpolation, and we provide a theoretical bound showing that flatter stability solutions tighten the interpolation bound and reduce forgetting during fusion. Our method is lightweight, modular, and readily compatible with existing LReID frameworks. Experimental results demonstrate that the proposed method consistently improves the performance of LReID in terms of both knowledge retention and generalization to unseen domains.


Towards Real-world Human Behavior Simulation: Benchmarking Large Language Models on Long-horizon, Cross-scenario, Heterogeneous Behavior Traces

Jiawei Chen ⋅ Ruoxi Xu ⋅ Boxi Cao ⋅ Ruotong Pan ⋅ Zhang yunfei ⋅ Yifei Hu ⋅ Yong Du ⋅ Tingting Gao ⋅ Han Li ⋅ Yaojie Lu ⋅ Yingfei Sun ⋅ Xianpei Han ⋅ Le Sun ⋅ Xiangyu Wu ⋅ Hongyu Lin

The emergence of Large Language Models (LLMs) has illuminated the potential for a general-purpose user simulator. However, existing benchmarks remain constrained to isolated scenarios, narrow action spaces, or synthetic data, failing to capture the holistic nature of authentic human behavior. To bridge this gap, we introduce OmniBehavior, the first user simulation benchmark constructed entirely from real-world data, integrating long-horizon, cross-scenario, and heterogeneous behavioral patterns into a unified framework. Based on this benchmark, we first provide empirical evidence that previous datasets with isolated scenarios suffer from tunnel vision, whereas real-world decision-making relies on long-term, cross-scenario causal chains. Extensive evaluations of state-of-the-art LLMs reveal that current models struggle to accurately simulate these complex behaviors, with performance plateauing even as context windows expand. Crucially, a systematic comparison between simulated and authentic behaviors uncovers a fundamental structural bias: LLMs tend to converge toward a positive average person, exhibiting hyper-activity, persona homogenization, and a utopian bias. This results in the loss of individual differences and long-tail behaviors, highlighting critical directions for future high-fidelity simulation research.


TRACER: Verifiable Generative Provenance for Multimodal Tool-Using Agents

Bihui Yu ⋅ caijun jia ⋅ Jing Chi ⋅ Xiaohan Liu ⋅ Yining Wang ⋅ Richard H Bai ⋅ Yuchen Liu ⋅ Jingxuan Wei ⋅ Junnan Zhu

Multimodal large language models (MLLMs) increasingly solve vision-centric tasks by calling external tools for visual inspection, OCR, retrieval, calculation, and multi-step reasoning. Current tool-using agents usually expose the executed tool trajectory and the final answer, but they rarely specify which tool observation supports each generated claim. We call this missing claim-level dependency structure the provenance gap. The gap makes tool use hard to verify and hard to optimize, because useful evidence, redundant exploration, and unsupported reasoning are mixed in the same trajectory. We introduce TRACER, a framework for verifiable generative provenance in multimodal tool-using agents. Instead of adding citations after generation, TRACER generates each answer sentence together with a structured provenance record that identifies the supporting tool turn, evidence unit, and semantic support relation. Its relation space contains Quotation, Compression, and Inference, covering direct reuse, faithful condensation, and grounded derivation. TRACER verifies each record through schema checking, tool-turn alignment, source authenticity, and relation rationality, and then converts verified provenance into traceability constraints and provenance-derived local credit for reinforcement learning. We further construct TRACE-Bench, a benchmark for sentence-level provenance reconstruction from coarse multimodal tool trajectories. On TRACE-Bench, simply adding tools often introduces noise. With Qwen3-VL-8B-Instruct, TRACER reaches 78.23% answer accuracy and 95.72% summary accuracy, outperforming the strongest closed-source tool-augmented baseline by 23.80 percentage points. Compared with tool-only supervised fine-tuning, it also reduces total test-set tool calls from 4,949 to 3,486. These results show that reliable multimodal tool reasoning depends on provenance-aware use of observations, not on more tool calls alone.


TRACE: Trajectory-Aware Conceps for Explainable Video Understanding

Wooil Lee ⋅ Junpyo Hong ⋅ Jongseo Lee ⋅ Tuomas Oikarinen ⋅ Lily Weng ⋅ Jinwoo Choi

Video understanding models often rely on motion, but faithful explanations must reveal not only whether motion matters, but also which source of motion supports a prediction. We identify motion-source under-specification as a key limitation of existing video XAI: region-centric explanations are source-ambiguous, while concept-based explanations are source-incomplete when their motion vocabulary is restricted to single-person pose sequences. To bridge this gap, we propose TraCE-Trajectory-Aware Concepts for Explainable Video Understanding-a concept-based framework that elevates visual trajectories to first-class explanatory concepts. By clustering optical-flow trajectories, TraCE discovers source-flexible motion concepts that capture dynamic evidence from people, objects, and interactions, while object and scene concepts account for static context. TraCE applies the same typed trajectory/object/scene concept interface to both action recognition and video question answering (VQA), enabling concept-level explanations beyond a single task. We evaluate TraCE across six benchmarks: NTU RGB+D Mutual, Chiral SSv2, UCF-101, KTH, TOMATO, and STAR. Experiments show that TraCE provides trajectory-aware explanations of dynamic evidence with minimal performance trade-off, trailing non-interpretable baselines by only 0.9 points on action recognition and 0.1 points on VQA on average. User studies, faithfulness analyses, temporal sensitivity tests, and concept-intervention results further show that TraCE produces interpretable, faithful, and practically useful explanations.


TrackCraft3R: Repurposing Video Diffusion Transformers for Dense 3D Tracking

Jisu Nam ⋅ Jahyeok Koo ⋅ Soowon Son ⋅ Jaewoo Jung ⋅ Honggyu An ⋅ Junhwa Hur ⋅ Seungryong Kim

Dense 3D tracking from monocular video is fundamental to dynamic scene understanding. While recent 3D foundation models provide reliable per-frame geometry, recovering object motion in this geometry remains challenging and benefits from strong motion priors learned from real-world videos. Existing 3D trackers either follow iterative paradigms trained from scratch on synthetic data or fine-tune 3D reconstruction models learned from static multi-view images, both lacking real-world motion priors. Pre-trained video diffusion transformers (video DiTs) offer rich spatio-temporal priors from internet-scale videos, making them a promising foundation for 3D tracking. However, their frame-anchored formulation, which generates each frame's content, is fundamentally mismatched with reference-anchored dense 3D tracking, which must follow the same physical points from a reference frame across time. We present TrackCraft3R, the first method to repurpose a video DiT as a feed-forward dense 3D tracker. Given a monocular video and its frame-anchored reconstruction pointmap, TrackCraft3R predicts a reference-anchored tracking pointmap that follows every pixel of the first frame across time in a single forward pass, along with its visibility. We achieve this through two designs: (i) a dual-latent representation that uses per-frame geometry latents and reference-anchored track latents as dense queries, and (ii) temporal RoPE alignment, which specifies the target timestamp of each track latent. Together, these designs convert the per-frame generative paradigm of video DiTs into a reference-anchored tracking formulation with LoRA fine-tuning. TrackCraft3R achieves state-of-the-art performance on standard sparse and dense 3D tracking benchmarks, while running 1.3x faster and using 4.6x less peak memory than the strongest prior method. We further demonstrate robustness to large motions and long videos. Our code and weights will be publicly released.


TrackTok: Object-Centric Video Tokenization with Semantically Persistent Tokens

Nhi Pham ⋅ Christopher Wewer ⋅ Bernt Schiele ⋅ Jonas Fischer ⋅ Jan Eric Lenssen

We introduce TrackTok, an efficiency-focused video tokenization framework that produces object-centric tokens, i.e., tokens that consistently encode the same semantic objects across time, even as it moves, deforms, or undergoes partial occlusion. Conventional patch-based or existing holistic tokenizers produce token identities often tied to fixed image locations. In contrast, our approach learns temporally persistent object-centric representations that bind visual features across frames into stable semantic units. This enables substantially more compact, continuous latent video representations, reducing the number of tokens required for downstream generative modeling with state-of-the-art flow models, while preserving the semantic structure needed for high-quality synthesis. On the UCF-101 video generation benchmark, we provide a proof of principle that our tokenizer uses ~4x fewer latent tokens to achieve the same generative FVD compared to existing holistic tokenizers. This results in an up to 10x increase in sample throughput, offering a favorable quality-efficiency trade-off. These results suggest that object-aligned and temporally persistent tokenization is a promising direction for scalable video generation, where reducing latent token counts is critical for efficient modeling.

Traffic forecasting benchmarks measure accuracy within a single city and test distribution, leaving open how architectures behave when sensors fail, when a model is transferred to another city, or when weekday and weekend traffic differ. We evaluate eight STGNN architectures under these three shifts using frozen checkpoints; for transfer across cities, target graph information is withheld so that fixed and adaptive graph models have the same structural access. We find that fragility under sensor dropout follows routing concentration rather than the use of adaptive adjacency itself. Across four adaptive backbones, mean row maximum routing mass correlates with dropout degradation at Pearson $\rho=+0.979$, and the relationship replicates on a second source city. Graph WaveNet lies at the concentrated end of this spectrum, while AGCRN and STAEformer stay close to DCRNN when sensors are zeroed. The other shifts separate from this pattern: MegaCRN performs best when transferred to another city, yet MegaCRN and STAEformer show the largest weekday and weekend performance gaps. In this backbone bank, strong performance under one shift does not imply robustness under another. We release a registry of frozen checkpoints, evaluation code, and a decision table mapping each backbone to its main failure mode.


Training Deliberative Monitors for Black-Box Scheming Detection

Aditya Sinha ⋅ Akshat Naik ⋅ Victor Gillioz ⋅ Simon Storf ⋅ Kilian Merkelbach ⋅ Rich Barton-Cooper ⋅ Axel Højmark ⋅ Marius Hobbhahn

As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a central AI control problem. Existing monitors often rely on chain-of-thought access or internal activations, or use prompted frontier models, all of which can be unavailable, unreliable or expensive in deployment. In this work, we study action-only deliberative monitors: smaller open-weight models trained to detect scheming and sabotage from agentic trajectories without accessing the monitored agent's reasoning or model internals. Our method, inspired by deliberative alignment, uses a scheming specification to elicit structured rationales from a frontier teacher, filters them with a separate judge, and distills the highest-quality rationales into open-weight monitors with supervised fine-tuning and reinforcement learning. We train on five datasets, and evaluate across six out-of-distribution agentic misalignment benchmarks. We show that applying our method to Qwen3.5-27B yields higher performance than all smaller prompted frontier monitors (Gemini 3.1 Flash-Lite, GPT-5.4 Nano, and Claude Haiku 4.5) and than Gemini 2.5 Pro, while also achieving lower marginal inference cost (token-metered USD per 1000 evaluations). Stronger prompted frontier monitors (Gemini 3.1 Pro, GPT-5.4, Claude Sonnet 4.6, and Claude Opus 4.6) achieve higher performance but at roughly $16$--$34\times$ higher marginal inference cost. Several of our trained monitors are positioned on the empirical cost--performance Pareto frontier among the monitors we evaluate, providing practical low-cost, low-FPR alternatives to prompted frontier models.


Training-Free 3D Editing via Feature-Divergence Localization and Trajectory Correction

THAT HUU TRI TON ⋅ Siwoo Lim ⋅ Joshua Tian Jin Tee ⋅ Chang Yoo

Modifying 3D assets requires enacting targeted structural transformations while strictly preserving the geometric identity of unedited regions. While recent probability flow models enable training-free editing, bounding these modifications within a high-dimensional latent space presents severe mathematical challenges. Current inversion-free trajectories rely on global velocity updates, which consistently result in spatial bleed into preserved structures and induce trajectory under-commitment, causing edits to stall mid-path. To resolve these dual failures, we introduce FocusFlow, a segmentation-free framework that structures the integration trajectory into two disjoint phases. During early coarse structure formation, FocusFlow applies Feature-Divergence Localization (FDL) to spatially gate the velocity field using intrinsic cross-attention. During late-stage fine-detail refinement, Reference-Guided Generation (RGG) drives topological convergence by pulling the trajectory toward an explicit clean-latent target. Evaluated on the joint Google Scanned Objects and PartObjaverse-Tiny benchmarks, FocusFlow breaks the traditional preservation-modification trade-off. It achieves superior structural preservation while maximizing edit magnitude. By providing stable, localized geometric control without manual spatial annotations, this phase-scheduled approach removes significant technical barriers to the precise modification of synthetic 3D media.

In this work, we introduce a novel training-free encoding framework, dubbed XShapeEnc, that encodes an arbitrary spatially grounded 2D geometric shape into a compact representation exhibiting five favorable properties, including invertibility, adaptivity, generality and controllability. Specifically, a 2D spatially grounded geometric shape is decomposed into its normalized geometry within the unit disk and its pose vector, where the pose is further transformed into a harmonic pose field that also lies within the unit disk. A set of orthogonal Zernike basis is constructed to encode shape geometry and pose either independently or jointly, with controllable relative emphasis on shape geometry or shape pose. We demonstrate the theoretical validity, efficiency, discriminability, and wide applicability of XShapeEnc via extensive analysis and experiments across a wide range of shape-aware tasks and our self-curated XShapeCorpus dataset. We envision XShapeEnc as a foundational tool for research beyond one-dimensional sequential data and data-driven, learning-based encoding paradigms, paving the way for a unified spatial encoding framework for frontier 2D spatial intelligence.


Training Optimal Large Diffusion Language Models

Jinjie Ni ⋅ Qian Liu ⋅ Chao Du ⋅ Longxu Dou ⋅ Hang Yan ⋅ Zili Wang ⋅ Tianyu Pang ⋅ Michael Shieh

We introduce Quokka, the first comprehensive scaling law for diffusion language models (DLMs), encompassing both compute- and data-constrained regimes alongside key modeling and optimization designs. Under compute constraints, we find that optimal parameter and dataset sizes scale equally with compute ($N_{\mathrm{opt}} \propto C^{0.5}$, $D_{\mathrm{opt}} \propto C^{0.5}$); however, DLMs are 2--5$\times$ more data-hungry than autoregressive (AR) models, requiring larger corpora and proportionally smaller models for a given FLOP budget. Under data constraints, validation loss follows a U-shaped curve across epochs, with the onset of overfitting scaling as $e_{\mathrm{opt}} \propto U_D^{0.39}/N^{0.55}$. This indicates that optimally leveraging a fixed unique data budget ($U_D$) requires allocating both modestly larger models and more training epochs. Beyond data and parameter allocation, we provide actionable guidance on DLM design choices: absorbing-mask transitions outperform uniform diffusion, linear noise schedules are the most stable and performant, and principled diffusion ELBO objectives ultimately surpass MaskGIT losses despite slower initial convergence. Furthermore, we demonstrate that easy-to-hard noise curricula accelerate early learning, established AR scaling laws for batch size and learning rate transfer seamlessly to DLMs, and weight decay, while unhelpful for single-epoch runs, is essential for multi-epoch regularization.


Trajectory-Consistent Diffusion Policies for Offline Reinforcement Learning

Yichao Fu ⋅ Shangde Gao ⋅ Zhuoling Li ⋅ Wen Wang ⋅ Shangqi Gao ⋅ Ke Liu

Diffusion policies have emerged as a powerful policy class for offline reinforcement learning (RL) due to their ability to model complex, multimodal action distributions through iterative denoising. However, in offline RL, expressivity alone is not enough. Generated actions are supposed to remain within value-reliable regions of the action space to avoid extrapolation error. We find that standard diffusion policies often produce unstable denoising trajectories whose intermediate and final action candidates drift toward weakly supported regions, making critic-guided action selection unreliable and degrading teacher quality for one-step distillation. This issue is particularly pronounced on mixed-quality offline datasets, where suboptimal behaviors amplify trajectory instability. To address this problem, we propose Trajectory Consistency Rectification (TCR), a framework designed to promote denoising stability during both training and inference. At inference time, TCR aggregates locally consistent, high-value action candidates across denoising steps to suppress critic-unreliable outliers. During training, it adaptively reweights denoising steps to improve trajectory-level coherence. We show that this stabilization provides a more reliable teacher for one-step flow-matching distillation. Extensive experiments on D4RL and OGBench show that TCR consistently improves over diffusion-policy baselines, with significant gains in sparse-reward domains like AntMaze and Adroit. Moreover, one-step policies distilled from TCR achieve competitive performance among efficient single-step methods. These results suggest that enhancing the temporal consistency of denoising trajectories is a key ingredient for robust offline diffusion reinforcement learning.


Transition-Aware Credit Assignment in Agentic Learning for LLM Reasoning

Xiangyu Lu ⋅ Zhanke Zhou ⋅ Jiazhe Ning ⋅ Chentao Cao ⋅ Jiangchao Yao ⋅ Bo Han

Multi-turn agentic reinforcement learning with verifiable rewards enables language models to reason with tools, but standard GRPO-style training often collapses: validation accuracy peaks and then falls while the overall frequency of tool calls remains nearly unchanged. We identify a credit-assignment mechanism behind this failure. With terminal correctness rewards, every tool-related token receives the trajectory’s scalar advantage, making credit transition-blind: the model cannot distinguish useful tool-use steps from unhelpful ones when they share the same final outcome. This transition-blind credit induces two biases. At the token level, failed rollouts with denser tool use can assign negative subset credit to tool tokens. At the trajectory level, token-mean reduction implicitly gives longer failed rollouts larger influence. These biases form a self-sustaining loop that suppresses useful tool interaction without visibly reducing tool-call frequency. To address this, we propose Transition-Aware Credit Assignment (TACA), a novel method that makes tool-token credit depend on the utility of the underlying tool interaction rather than only the terminal trajectory outcome. By restoring transition-sensitive credit and stabilizing sparse tool-use cases, TACA consistently outperforms agentic and non-agentic baselines across different models and multiple reasoning benchmarks.


TransMem: Transition-Aware Retrieval for Evolving Personal Memory

Shigeng Chen ⋅ LINHAO LUO ⋅ Changlong Shi ⋅ Zhangchi Qiu ⋅ Shirui Pan

Long-term personalized agents are increasingly expected to answer queries over user histories in which preferences, goals, and constraints evolve across interleaved conversations. Existing retrieval-based and structured-memory systems can surface relevant facts, but often return isolated memory units that do not explain which earlier preference was refined, contradicted, or superseded by later evidence. We introduce TransMem, a transition-aware memory retrieval framework that stores preference updates as explicit retrieval units alongside ordinary factual memories. During memory construction, TransMem routes user interaction history into topic-structured memory graphs, detects shift events, and grounds them as transition records. At inference time, TransMem first retrieves factual anchors and then adaptively expands over transition records, producing a compact context of factual anchors and accepted transition records. Experiments on PersonaMem and PersonaBench across two Qwen backbones show that TransMem outperforms representative retrieval-based and structured-memory baselines. Controlled analyses further show that transition records and adaptive expansion contribute beyond factual anchoring, chronological expansion, flat transition retrieval, and fixed-hop traversal. These results suggest that preference-update records are useful retrieval units for long-term personalization.

Effective discrete representation learning with VQ-VAE serves as a fundamental component of modern autoregressive generative models. However, it is widely known to suffer from codebook collapse and inefficient code utilization. To address this issue, we propose a simple yet effective alternative to standard vector quantization by formulating neural quantization as a joint modeling and clustering process. Specifically, we learn the latent space using a lightweight flow-matching model alongside training, corresponding to a probability flow ODE that transports an often fixed well-behaved source distribution while preserving probability mass. Based on this, we define the codebook $\mathcal{C}_1$ by transporting even-quantile partitioned prior codes $\mathcal{C}_0$ from the source space through the learned flow, and obtaining the Voronoi centroids under the induced distribution, leading a structured discretization. We evaluated on image tasks, and two tasks require high-precision tokenization, demonstrating scalability and improved expressivity of the codebook with full utilization.


Treat Domain-Specific Languages as Design Variables in LLM Agents

Letian Li ⋅ Wenyuan Jiang ⋅ Xin Yang ⋅ Shuzhao Xie ⋅ Chenghao Gu ⋅ Yu Meng ⋅ ZhengXiao He ⋅ Zhi Wang

Large Language Model (LLM) agents increasingly excel at complex tasks, yet a persistent bottleneck emerges at the boundary of rule-based systems. Here, the core challenge is often not comprehending intent, but routing user intent through a prescribed, executable interface. This paper posits that this routing problem is best addressed not by scaling model reasoning alone, but by treating the interface itself—specifically, Domain-Specific Languages (DSLs) and their attached verifiers—as first-class boundary infrastructure. We argue that an effective DSL functions as a cognitive shortcut: by providing a high-abstraction, strongly constrained representation, it can compress the intent-to-action path, prune invalid actions before they are attempted, and mechanistically verify correctness via its harness. Our central contribution is to frame the deliberate selection, design, and evaluation of such language--harness pairings as an under-studied research variable that directly shapes an agent's effective action space, failure modes, and learning signals. We develop this position through a dual lens of lifecycle and domain analyses, and ground it with targeted case studies.

Many prediction tasks evaluate a tree ensemble on row groups that share feature values: discrete-time survival models expand each patient into $G$ time steps, click-through-rate models score each item in a search session, and scenario analyses vary a few inputs while keeping others fixed. Standard tree-ensemble inference treats each row independently, ignoring this input-side correlation. We instantiate partial evaluation for grouped tree-ensemble inference: constant features are static, varying features dynamic. The algorithm walks each tree once per group, partitions a row bitmask at varying splits, and skips empty subtrees via an unsplit shortcut. We prove a structural work decomposition: per-tree work splits into the constant-projected subtree size $|T_c|$, $G$ leaf writes, and a predicate-mask provisioning cost $Q$. For the trace evaluator, per-row work approaches $(d_v+1)/(d+1)$ as $G \to \infty$, giving asymptotic speedup $1/(1-\bar\rho_G)$. Empirically, TreeWalker delivers ${\sim}3\times$ algorithmic speedup over a row-independent traversal at the reference model size ($T{=}500$, $L{=}8$; $G{=}16$ for SUPPORT/FLCHAIN and empirical groups for Expedia), rising to $6.8$--$7.8\times$ at $G{=}128$ on the survival datasets. For f64 models, outputs match treelite GTIL up to tree-ordering roundoff; for f32 models, TreeWalker's f64 leaf accumulation is close to a Kahan-compensated reference than native f32 on $99.98\%$ of rows, with the remaining $0.02\%$ tied.


TriTD: Tri-Partite Trajectory–Distribution Distillation for Real-Time Autoregressive Video Generation

Jia Li ⋅ Xurui Peng ⋅ Haowei Zhu ⋅ Tianyu Zhao ⋅ Jiexi Wang ⋅ Youwei Zheng ⋅ Yuxi Ren ⋅ Xinglong Wu ⋅ XING WANG ⋅ Hayden So

Autoregressive (AR) video diffusion has emerged as a leading paradigm for efficient video generation, yet its practical deployment remains bottlenecked by the trade-off between generation quality and inference efficiency. We attribute this bottleneck to three error sources in AR rollouts: the base model error, the train–inference distribution mismatch, and the upstream error propagation. Building on this decomposition, we derive closed-form upper bounds on the per-step and cumulative errors, providing a principled mechanism for jointly optimizing generation quality and inference efficiency. Guided by this analysis, we propose TriTD, a Tri-Partite Trajectory–Distribution co-optimized distillation framework with three complementary modules that each target a distinct error term in the bound. First, for ODE initialization, we adopt a Next-Node Trajectory Distillation (NTD) objective and prove a strictly tighter error bound than conventional consistency distillation, reducing the base model error at its source. Second, we introduce Chunk-Graded Noising (CGN), a monotonically non-decreasing per-chunk noising schedule that provably narrows the train–inference noise mismatch and removes redundant KV Cache updates, yielding substantial speedups at no quality cost. Third, we propose Orthogonal-Parallel DMD (OP-DMD), a distribution-matching loss that preserves the original DMD gradient direction while explicitly suppressing unsupervised off-axis drift in the orthogonal subspace, improving both training stability and final generation quality. Extensive experiments show that TriTD consistently surpasses state-of-the-art baselines in visual quality while reaching a real-time throughput of 26.8 FPS on a single H100 GPU, delivering a balanced quality-efficiency solution for interactive streaming video generation.


TrustMod-SM: A Multi-Axis Benchmark for Evaluating Trustworthiness of LLMs in Social Media Content Moderation

Aizan Zafar ⋅ Sreekantam Sai Venkat ⋅ Gathik Jindal ⋅ Ritu Raj Sharma ⋅ Zishan Ahmad ⋅ Tulika Saha ⋅ Srinath Srinivasa

Large language models have shown remarkable promise for social media content moderation, with recent studies demonstrating performance rivaling human annotators on policy-compliance tasks. However, the trustworthiness of these models beyond aggregate accuracy remains largely unexamined, posing significant risks, a model with high overall accuracy may still exhibit substantial error-rate disparities across demographic groups while remaining vulnerable to adversarial manipulation. In this paper, we introduce TrustMod-SM and aim to comprehensively evaluate the trustworthiness of LLM-based content moderators across five dimensions: trustfulness, fairness, safety, robustness, and context integrity. TrustMod-SM comprises about 29K evaluation instances curated from eight established datasets, and we evaluate thirteen open-weight models (0.5B-14B parameters), including both text-only and vision-language architectures. Our analysis reveals that models consistently exhibit trustworthiness concerns, often displaying greater demographic disparity as detection accuracy improves, complying with both concealment and exaggeration inducements, assigning high confidence to incorrect predictions while rarely escalating ambiguous content for human review, and failing to distinguish hate from counter-speech and reclaimed language, particularly at sub-8B scales. Furthermore, scaling does not uniformly improve trustworthiness, and a cross-dimensional analysis reveals non-trivial associations between axes, challenging the independence assumed by prior benchmarks. We publicly release our benchmark data, evaluation code, and model predictions at https://anonymous.4open.science/r/TrustMod_SM WARNING: This paper contains model outputs that may be considered offensive.


TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

Shunyu Wu ⋅ Dan Li ⋅ Haozheng Ye ⋅ Weibin Feng ⋅ Jian Lou ⋅ Bo Zhang ⋅ Wenjie Feng ⋅ Chenjuan Guo ⋅ See-Kiong Ng

Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently, large language models (LLMs) have emerged as a promising paradigm for TS quality assessment via pairwise comparison and per-dimension evaluation. However, existing approaches rely on manually predefined quality dimensions and purely text-based reasoning, leaving it unknown whether LLMs can identify truly relevant quality dimensions or perform grounded and quantitative quality comparisons. To investigate this, we construct TSQBench, a dedicated benchmark for evaluating LLMs’ capabilities in TS quality assessment, which measures two progressive capabilities: (i) understanding and identifying relevant quality dimensions, and (ii) performing quality comparison under specific dimensions. Our analysis reveals that both open-source and proprietary LLMs consistently struggle with identifying critical quality dimensions and conducting precise, evidence-grounded reasoning over quality characteristics. To address these limitations, we propose TSQAgent, a novel agentic reasoning framework for dedicated TS data quality rating, comprising three collaborative roles: Perceiver for focused dimension selection, Inspector for dimension-wise quantitative analysis, and Adjudicator that aggregates and refines the final judgment. In particular, we introduce an agentic reasoning strategy that instills the ability to identify and prioritize the most relevant quality dimensions, and further propose an agent workflow equipped with external analytical tools to enable precise quantitative comparisons over selected dimensions. Experiments on both the proposed benchmark and eleven real-world datasets demonstrate that our framework not only substantially improves LLMs’ capabilities in quality understanding and quantitative comparison but also effectively translates these improvements into better quality-aware data selection, leading to enhanced downstream performance and data efficiency.


TStruct: Learning Shared Temporal Structures for Long-Term Time-Series Forecasting

Yangze Li ⋅ Zhengyang Zhou ⋅ Qihe Huang ⋅ Pengkun Wang ⋅ Binwu Wang ⋅ Yang Wang

Long-term time-series forecasting requires temporal dependency modeling that remains stable and generalizable under temporal distribution shifts. However, real-world time series are often noisy and non-stationary, making deep learning models prone to fitting incidental temporal interactions in the training data. A key question is whether there exist temporal relationship structures that recur across different temporal positions and variables. We analyze temporal relationship matrices on commonly used time-series datasets and find that samples from different temporal positions and different variables exhibit shared dominant temporal relationship structures. This indicates that temporal relationship structures contain patterns that recur stably across samples and variables. To learn such shared structures, we propose Temporal Structure Network (TStruct), a simple yet effective network for directly learning shared and generalizable temporal dependencies. Specifically, TStruct directly parameterizes a set of shared temporal dependency matrices and generates normalized combination weights using dynamic sample information and static variable priors. These weights represent a soft allocation of each sample-variable pair over shared temporal dependency structures, enabling the model to generate adaptive temporal dependency matrices under shared structural constraints, rather than fitting incidental interactions in a fully unconstrained manner. Notably, this simple structure-aware design consistently outperforms several strong baselines across all datasets, achieving superior accuracy while maintaining high computational efficiency. Our code is available at \url{https://anonymous.4open.science/r/TStruct-NeurIPS2026-1173}.


TurboVGGT: Fast Visual Geometry Reconstruction with Adaptive Alternating Attention

David Huang ⋅ Guile Wu ⋅ Chengjie Huang ⋅ Bingbing Liu ⋅ DONGFENG BAI

Recent feed-forward 3D reconstruction methods, such as visual geometry transformers, have substantially advanced the traditional per-scene optimization paradigm by enabling effective multi-view reconstruction in a single forward pass. However, most existing methods struggle to achieve a balance between reconstruction quality and computational efficiency, which limits their scalability and efficiency. Although some efficient visual geometry transformers have recently emerged, they typically use the same sparsity ratio across layers and frames and lack mechanisms to adaptively learn representative tokens to capture global relationships, leading to suboptimal performance. In this work, we propose TurboVGGT, a novel approach that employs an efficient visual geometry transformer with adaptive alternating attention for fast multi-view 3D reconstruction. Specifically, TurboVGGT employs an end-to-end trainable framework with adaptive sparse global attention guided by adaptive sparsity selection to capture global relationships across frames and frame attention to aggregate local details within each frame. In the adaptive sparse global attention, TurboVGGT adaptively learns representative tokens with varying sparsity levels for global geometry modeling, considering that token importance varies across frames, attention layers operate tokens at different levels of abstraction, and global dependencies rely on structurally informative regions. Extensive experiments on multiple 3D reconstruction benchmarks demonstrate that TurboVGGT achieves fast multi-view reconstruction while maintaining competitive reconstruction quality compared with state-of-the-art methods.


Two-Clustering Regime of Token Dynamics in Causal Attention

Trinh Nguyen ⋅ Duy-Tung Pham ⋅ Hoang-Son Do ⋅ Tan Nguyen ⋅ Thieu Vo

We partially resolve Conjecture 1 of \textit{Karagodin et al., Clustering in Causal Attention Masking (NeurIPS 2024)} and extend it to time-dependent parameter regimes. The conjecture predicts that causal attention dynamics converge to two clusters aligned with the eigenvectors associated with the largest eigenvalue $\lambda_{\max}$ of the value transformation. We prove that this behavior holds when $\lambda_{\max}>0$ and is generally false when $\lambda_{\max}\le0$. The resulting two-cluster regime is common in practice and substantially more challenging than the classical single-cluster setting studied previously. A central contribution of our work is the identification of a rigorous dimension deduction phenomenon: despite evolving in a high-dimensional embedding space, token trajectories collapse onto a two-dimensional structure governed by a small number of dominant directions. Building on this insight, we introduce a new analytical framework--instability within the flow--inspired by instability phenomena in physical systems. This framework enables the analysis of steady-regime behavior in non-monotone, non-autonomous attention dynamics. Our theoretical findings are illustrated through simulations and experimental studies using Group Query Attention (GQA) layers.


Two Phase Rapid Simulation Based Inference with Differentiable Simulators

Uros Zivanovic ⋅ Rahul Srinivasan ⋅ Roberto Trotta ⋅ Andre Scaffidi

We introduce Rapid Simulation Based Inference (RSBI), a diffusion-based variational approach to likelihood-free Bayesian inference that achieves high posterior sample efficiency under large and multi-dimensional prior-to-posterior volumes. RSBI builds on recent advances in Schrödinger Bridge (SB) diffusion sampling to handle multi-modal posteriors, with the likelihood supplied either by a neural ratio estimator or kernal surrogate via a differentiable simulator. Furthermore, a key observation is that an appropriate surrogate likelihood gives a proposal distribution covering most of the generally multi-modal posterior mass, replacing the sequential rounds of standard methods with a one-shot proposal step. Subsequent NRE refinement recovers exact inference under the simulator and original model prior. As a variational method, RSBI does not suffer from prior leakage and implicit target drift commonly observed in standard sequential posterior estimation methods. To improve mode coverage we optionally utilize the well-tempered meta-dynamics framework, which also provides a mechanism through which to encourage exploration of the prior volume. We observe strong performance not only on standard SBI benchmarks, but also as prior volume is scaled up to 2500 times the original, achieving a performance degradation of only $\sim$5% on the two moons benchmark under a highly constrained simulation budget. Additionally, we evaluate RSBI's potential for gravitational wave ring-down posterior estimation, highlighting a real-world use-case where our method excels.


UGM: Unified Multi-scale Genomic Event Modeling with Site-level Joint Prediction

Jiayang Wu ⋅ Chenchen Qin ⋅ Xu YANG ⋅ Yu Zhao ⋅ Bing He ⋅ Jiale Zhou ⋅ Zhenchao Tang ⋅ Minghao Yang ⋅ Shou Z Chen ⋅ Yefeng Zheng ⋅ Jianhua Yao

Genomic events span multiple spatial scales, from single-nucleotide substitutions to large structural variants, yet existing foundation models process only reference DNA sequences, leaving methylation, short variants, and structural-variant boundaries as downstream labels or auxiliary inputs. We propose the Unified Genomic Model (UGM), an encoder-only genomic foundation model that learns multi-scale genomic events as native tokens within a unified 43-token vocabulary. UGM introduces three algorithmic components: (i) a unified event tokenizer that maps reference bases, SNPs, CpG methylation, short INDELs, and structural-variant boundaries into a single event-state space; (ii) event-balanced pretraining, which promotes adequate gradient coverage for rare but biologically important events; and (iii) Site-Level Joint Prediction (SJP), a masked-language-modeling objective that recovers the complete genomic state at a masked position as a single token rather than predicting base, methylation, and variant labels independently. We evaluate UGM against specialized and generalist baselines across functional variant classification, regulatory benchmarks (NT and GUE), structural-variant detection, and cross-event attention analysis. UGM achieves competitive performance, particularly on tasks where explicit event-state information is central. It obtains the highest score on all seven functional variant tasks, reaches competitive splicing prediction on NT, and notably outperforms DNA-only baselines on SV breakpoint detection. These results suggest that native multi-scale event pretraining is a promising direction for genomic representation learning.


U-MOF: Uncertainty-Guided Parameter-Efficient Multi-Objective Fine-Tuning for Long-Tailed Recognition

Yuan Dong ⋅ Di Wu ⋅ Zhe Zhao ⋅ Liheng Yu ⋅ Xiaofeng Cao ⋅ Pengkun Wang ⋅ Yang Wang

Recent long-tailed recognition methods increasingly adopt two-stage fine-tuning, where a generic representation is first learned and imbalance-aware adaptation is then performed in a lightweight second stage. However, existing second-stage designs often allocate objective emphasis using delayed empirical summaries, such as historical class-wise performance or validation feedback, and commonly restrict adaptation to classifier heads. These choices may make routing unreliable when empirical feedback is noisy or skewed, and may also limit the feature plasticity required for rare categories. In this paper, we propose U-MOF, an uncertainty-guided parameter-efficient multi-objective fine-tuning framework for long-tailed recognition. U-MOF introduces lightweight residual adapters into a shared-backbone expert committee, uses entropy-based committee statistics as diagnostic routing proxies to modulate tail-oriented and smoothing-oriented objectives, and applies conflict-aware orthogonal projection to coordinate heterogeneous objective gradients. Experiments on CIFAR-100-LT, ImageNet-LT, and iNaturalist 2018 show that U-MOF improves rare-category and overall performance in controlled stage-two comparisons and remains competitive with strong published long-tailed recognition baselines, while preserving the efficiency advantages of decoupled adaptation. These results indicate that reliable stage-two long-tailed fine-tuning benefits not only from selecting appropriate objectives, but also from diagnosing when each objective should dominate and from retaining limited feature-level plasticity. An anonymized implementation is included in the Supplementary Material.


Uncertainty-Aware Fuzzy Graph Contrastive Learning

Xinlan Xu ⋅ Fei Hao ⋅ Ce Yang ⋅ Jiaxing Shang ⋅ Jinsong Chen ⋅ Jianrui Chen ⋅ Hongying Zhang ⋅ Aziz Nasridinov

Graph Contrastive Learning (GCL) is an effective approach for learning node representations. However, the intrinsic uncertainty of real-world graph data introduces significant challenges, as GCL relies on crisp deterministic representations that struggle to capture such ambiguity. To model graph uncertainty, fuzzy logic has attracted increasing attention due to its mathematical rigor. Existing methods, however, simply inject fuzzy modeling into the contrastive framework without tailoring data augmentation to the needs of fuzzy representations, resulting in limited discriminability and suboptimal downstream performance. To address these limitations, we propose a novel Fuzzy Graph Contrastive Learning (FGCL) framework for uncertainty-aware node representation learning. It adopts a vertex entanglement-based augmentation strategy to generate tailored multi-views with controlled perturbations while preserving core graph topology and semantics, mitigating semantic corruption. We also design a Deep Weighted Fuzzy Graph Convolutional Neural Network (DWFGCNN) encoder, which maps crisp features to fuzzy representations via learnable membership functions, explicitly models uncertainty, and resolves the core contradiction that traditional encoders cannot balance discriminability and robustness simultaneously. Extensive experiments on multiple datasets demonstrate that FGCL consistently outperforms baselines in node classification, community detection, and link prediction tasks, validating its effectiveness in handling uncertainty in graph representation learning.

Fine-grained visual classification (FGVC) with Multimodal Large Language Models (MLLMs) remains challenging, as visually similar subordinate categories inherently induce classification uncertainty due to subtle inter-class differences. Existing reinforcement learning methods typically require the model to generate a single class name and optimize it with sparse correctness rewards. However, a seemingly incorrect prediction can still provide highly informative signals if it falls within the same local confusion set as the ground-truth class. By treating all non-exact matches as equally wrong, the rigid single-answer paradigm overlooks these valuable near-misses. Consequently, it fails to capture the model's intrinsic uncertainty and provides limited learning signals. To address these limitations, we propose List-GRPO, an uncertainty-aware listwise Group Relative Policy Optimization framework. Rather than enforcing a single prediction, our method prompts the model to adaptively generate a ranked candidate list, explicitly capturing its predictive uncertainty. Furthermore, to mitigate the advantage vanishing when online exploration completely misses the target, we introduce a dynamic ground-truth anchored response strategy. We synthesize an off-policy anchor by prepending the ground truth to the most frequent hard negatives mined from online samples. To prevent the anchored response from dominating learning, we isolate the advantage estimation of online samples from the anchored response and use a gating coefficient to modulate its contribution. Extensive experiments on multiple fine-grained benchmarks demonstrate that List-GRPO significantly improves Top-1 accuracy while maintaining high ground-truth coverage with only about three candidate predictions on average.

Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-valued and entangle intrinsic ambiguity, expert judgment, and stochastic corruption. Existing deep constrained clustering (DCC) methods mainly target hard, expert-agnostic constraints, treating soft labels mostly numerically rather than semantically. We formalize this setting as uncertainty-aware probabilistic constrained clustering (UPCC), defining a canonical aleatoric target through a heterogeneous observation process and analyzing its conditional identifiability. We introduce ProbPair, an angular pairwise objective for probabilistic relations, and build ECI-PP, an estimator--corrector--integrator framework that refines imperfect supervision via belief estimation, correction, and reliability-aware integration. Across challenging probabilistic supervision settings, experiments on diverse benchmarks show that ECI-PP outperforms state-of-the-art DCC methods and remains robust with a shared default configuration.

Accurate image registration is essential in many medical imaging applications, yet most deep registration networks provide little indication of when or where their predictions are unreliable. Existing uncertainty estimation approaches, such as Bayesian methods, ensembles, or MC-dropout, typically require architectural modifications or retraining, precluding their applicability to pretrained registration models. We propose an inference-time, model-agnostic uncertainty estimation framework that applies directly to any pretrained registration network. Our approach is grounded in the transformation equivariance property of image registration, which states that the underlying anatomical mapping should remain consistent under spatial perturbations of the input. Experiments across three pretrained registration models and four anatomical structures show that the resulting uncertainty maps consistently correlate with registration error and highlight unreliably aligned regions. This framework turns pretrained registration networks into risk-aware tools at test time, moving medical image registration closer to safe clinical and large-scale research deployment. The code will be released at ***.


Understanding Generalization through Decision Pattern Shift

Huiqi Deng ⋅ Yibo Li ⋅ Quanshi Zhang ⋅ Peng Zhang ⋅ Hongbin Pei ⋅ Xia Hu

Understanding why deep neural networks (DNNs) fail to generalize to unseen samples remains a long-standing challenge. Existing studies mainly examine changes in externally observable factors such as data, representations, or outputs, yet offer limited insight into how a model’s internal decision mechanism evolves from training to test. To address this gap, we introduce \textbf{Decision Pattern Shift (DPS)}, a new perspective that defines generalization through the stability of internal decision patterns and quantifies failure as their deviation from those learned during training. Specifically, we represent each sample’s decision pattern as a GradCAM-based channel-contribution vector, which captures how feature channels collectively support a prediction, and we propose the DPS metric to measure its discrepancy from the class-average pattern. Empirical analyses across multiple datasets and architectures show that, (i) decision patterns form a highly structured, class-consistent space with strong intra-class cohesion and low inter-class confusion, enabling direct analysis of a model’s decision logic; (ii) the DPS magnitude correlates linearly with the generalization gap (nearly all Pearson $r > 0.8$), revealing generalization as a systematic drift in the model’s internal decision mechanism; (iii) the DPS spectrum organizes diverse generalization degradation scenarios (covering ideal generalization, in-distribution degradation, domain shift, out-of-distribution, and shortcut learning) into a continuous trajectory, providing a unified explanation of their failure modes. These findings open up new possibilities for early generalization-risk detection, failure-mode diagnosis, and channel-level defect localization.

Reinforcement learning agents often exhibit unintended goal-directed behaviour outside their training distribution, but we currently lack a principled understanding of how such agents will generalise to novel environments based on their training history. We address this gap for agents trained sequentially on one or more tasks. We study over 100 sequential training pipelines, evaluating behaviour across over 250 out-of-distribution environments. We find that salient features drive generalisation, and that goals learnt early in training can persist and influence those acquired later. To explain these phenomena, we introduce latent policy gradients, a method that predicts what out-of-distribution behaviour a training pipeline will likely induce. Our method simulates the evolution of low-dimensional latent variables during training according to what would achieve high reward on the training objective with respect to a simple model of how the latent variables map to behaviour. It achieves strong predictive accuracy, generalises to unseen types of training pipeline, and is interpretable. Our findings demonstrate that while out-of-distribution RL agent behaviour is dependent on the whole training pipeline, this dependence has an underlying structure we can capture, laying groundwork for understanding goal generalisation from a developmental perspective.


UnfoldArt: Zero-Shot Recovery of Full Articulated 3D Objects from Text or Image

Mohamed El Amine Boudjoghra ⋅ Ivan Laptev ⋅ Angela Dai

Articulated 3D objects are essential for interactive environments in embodied AI, robotics, and virtual reality, but reconstructing their structure and motion from sparse observations remains challenging. Existing approaches remain largely constrained by lack of supervised data or lack priors needed to reliably recover articulation, hidden geometry, and internal object structure. We present the first hierarchical agentic approach for articulated 3D object reconstruction from text or image inputs, combining hierarchical, debate-based reasoning with a video generative prior for articulation modeling. High-level agents reason about object semantics and motion using knowledge from vision-language and video models, while low-level agents estimate articulation parameters and interaction points; together, they engage in structured debate to resolve ambiguities in structure and motion. To enable reliable generation, including interior object structures, from only text or image inputs, we introduce a video generative model prior that not only synthesizes plausible object motions, but also expose occluded interiors and geometry that cannot be inferred from a single static view. By combining agentic reasoning with a video generative prior, our approach jointly infers articulation and reconstructs complete 3D articulated objects, producing high-fidelity geometry, internal structure, and motion-consistent states beyond directly observed surfaces.


UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

Zhekai Chen ⋅ CHENGQI DUAN ⋅ Kaiyue Sun ⋅ Bohao Li ⋅ Yuqing Wang ⋅ Manyuan Zhang ⋅ Xihui Liu

The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-world environments. However, existing benchmarks struggle to evaluate such agents effectively, as they often rely on sandboxed environments and single-turn evaluation paradigms. Moreover, their scenario-based task taxonomies mix multiple model capabilities within the same task category, making it difficult to identify the root causes of agent failures. To address these limitations, we introduce UniClawBench, the first capability-driven benchmark designed to evaluate proactive agents in dynamic, real-world settings. UniClawBench is built around five foundational model capabilities: Skill Usage, Exploration, Long-Context Reasoning, Multimodal Understanding, and Cross-Platform Coordination. Based on these capabilities, we design 400 bilingual real-world tasks. Unlike previous benchmarks that rely on static, pre-recorded answers, our benchmark evaluates agents in live Docker containers using fine-grained, step-by-step completion checkpoints. Furthermore, we design a closed-loop evaluation strategy comprising an executor agent, a hidden supervisor agent, and a user agent to simulate realistic multi-turn human feedback without leaking grading criteria. To disentangle base model capabilities from framework-level design choices, we evaluate state-of-the-art models under multiple agent frameworks. Through comprehensive comparisons across both models and frameworks, we show how base model capabilities and agent framework designs jointly shape performance in real-world environments.


UNICOM: Unified Multimodal Modeling via Compressed Continuous Semantic Representations

Yaqi Zhao ⋅ Wang Lin ⋅ Zijian Zhang ⋅ Miles Yang ⋅ Zhao Zhong ⋅ Liefeng Bo ⋅ Jingyuan Chen ⋅ Wentao Zhang

Current unified multimodal models typically rely on discrete visual tokenizers to bridge the modality gap. However, discretization inevitably discards fine-grained semantic information, leading to suboptimal performance in visual understanding tasks. Conversely, directly modeling continuous semantic representations (e.g., CLIP, SigLIP) poses significant challenges in high-dimensional generative modeling, resulting in slow convergence and training instability. To resolve this dilemma, we introduce UniCom, a unified framework that harmonizes multimodal understanding and generation via compressed continuous semantic representations. We empirically demonstrate that channel-wise compression is significantly more effective than spatial token reduction, retaining competitive VAE-free reconstruction fidelity while substantially improving generative convergence. Accordingly, we design an attention-based semantic compressor to distill dense features into a compact unified representation. A Transfusion-style predictor then generates these compressed latents with dense spatial correspondence. At scale, UniCom achieves competitive text-to-image generation and state-of-the-art performance on complex image editing without relying on VAE latents. The gains are especially clear on knowledge-intensive benchmarks.

Multicalibration requires predicted scores to agree with label probabilities across rich families of subgroups and score-dependent tests, but existing methods require clean input--label pairs for evaluation and post-processing. This assumption fails in weakly supervised learning (WSL) regimes---including positive-unlabeled, unlabeled-unlabeled, and positive-confidence learning---where clean labels are costly or unavailable even though reliable uncertainty estimates may be crucial. We address this gap by developing estimators of multicalibration error and post-hoc correction methods for WSL settings in which clean input--label pairs are unavailable. We propose a unified framework for estimating and correcting multicalibration under weak supervision by combining contamination-matrix risk rewrites with witness-based calibration constraints, yielding corrected multicalibration moments with finite-sample guarantees. We further propose weak-label multicalibration boost (WLMC), a generic post-hoc recalibration algorithm under weak supervision. Finally, we conduct experiments across multiple weak-supervision settings to evaluate multicalibration behavior and deepen empirical understanding of uncertainty estimation under weak supervision.


Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

Yuyue Wang ⋅ Xihua Wang ⋅ Xin Cheng ⋅ Yijing Chen ⋅ Ruihua Song

Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. Current methods either rely on disjoint pipelines, which fail to capture fine-grained interactions, or require structured inputs and external text rewriting, which limits the flexibility of free-form text prompts. In this paper, we introduce a new task: Free-Form-Text-Prompt-to-Unified-Audio generation, which aims to directly synthesize unified audio containing speech, sound, and their composites from unconstrained natural language. To address this task, we propose PlanAudio, a unified, autoregressive LLM-based framework. First, it simplifies the model architecture by leveraging intrinsic LLM reasoning capability instead of traditional text encoders. Second, it introduces a semantic latent chain-of-thought mechanism, an implicit planning mechanism that bridges high-level semantic understanding and low-level acoustic synthesis. Furthermore, we create PlanAudio-Bench, a specialized benchmark for evaluating composite audio scenarios. We perform evaluations in the scenarios of speech, sound, and their composites. The results demonstrate that PlanAudio generally outperforms the existing pipeline and unified baselines, while staying competitive with models designed for a single scenario. Our analysis further reveals the superiority of semantic latent CoT over other CoT mechanisms and highlights the importance of continuous multi-scenario training curricula.


UniSHARP: Universal Sharp Monocular View Synthesis

Meixi Song ⋅ Dizhe Zhang ⋅ Ruiyang Zhang ⋅ Bo Du ⋅ Ming-Hsuan Yang ⋅ Lu Qi

In this work, we focus on extending SHARP, the popular photorealistic view synthesis method, for universal monocular rendering across diverse camera systems, such as large-field-of-view fisheye lenses. To overcome the pinhole-specific assumptions of SHARP, our key idea is to align various images in a unified omnidirectional latent space. Thus, we propose UniSHARP, which performs implicit alignment in both feature and Gaussian spaces. Specifically, 2D semantic embeddings and 3D spatial features are jointly encoded and decoded by UniK3D to support Gaussian construction, while a ray-based universal representation organizes Gaussian primitives along rays and radial distances. To comprehensively evaluate our method, we construct a benchmark covering diverse imaging systems, including pinhole, fisheye and panoramic cameras, across various scenes. The benchmark is further stratified by field of view, spanning narrow perspective, wide-FoV, fisheye, and full panoramic settings. Extensive experiments on the proposed benchmark demonstrate the effectiveness of UniSHARP, outperforming alternative methods by a large margin. Our models, training code, and dataset will be publicly available.

We propose an adaptive proximal gradient method for minimizing the sum of two functions, where one is a simple convex function, and the other belongs to one of the three classes: nonconvex smooth, convex nonsmooth, or convex smooth. The key feature of the method is an adaptive step size that accumulates historical gradient mapping norms in the denominator. Without any modification or knowledge of problem parameters, the method converges across all three problem classes under mild bounded-iterates and bounded-variance assumptions, with rates matching those of the proximal gradient method up to logarithmic factors, in both deterministic and stochastic settings. For the convex setting, we further propose an accelerated variant. It retains a similar near-optimal convergence rate for the nonsmooth case and achieves an improved rate of order $\widetilde{O}\big(1/t^2 + \sigma/\sqrt{t}\big)$ for the smooth case, which is optimal up to logarithmic factors. Notably, we develop new techniques for controlling the effect of stochastic noise, which are applicable across all three problem classes in the stochastic setting and enable simplified analysis.


Universal Time Series Generation with Neural Controlled Differential Equations

Torben Berndt ⋅ Elyes Farjallah ⋅ Leif Seute ⋅ RAEID SAQUR ⋅ Benjamin Walker ⋅ Jan Stühmer

Recent work on the sequence universality of State Space Models (SSMs) has introduced efficient, maximally expressive continuous-time approaches for time-series modelling. While these works focus on discriminative settings, we extend this perspective to generative time-series modelling by proving that maximally expressive Structured Linear Controlled Differential Equations (SLiCEs) are universal time-series generators, in the sense that they can approximate the induced path laws of continuous causal pushforwards on compact latent sets in $W_\infty$. Building on these theoretical results, we propose Generative SLiCEs (G-SLiCEs), a maximally expressive continuous-time model for flow matching on path-space. Empirically, we show that expressivity improves performance in probabilistic forecasting and downstream tasks, while retaining the advantages of continuous-time models such as generalising to arbitrary observation grids. This is particularly beneficial for irregular grids, where fixed-grid models often struggle.


Unlocking Any-Order Generation in Pretrained Autoregressive Image Models

Rishav Pramanik ⋅ Marco Pedersoli ⋅ Zhaozheng Yin

Autoregressive (AR) image generators are powerful in principle but inflexible in practice: training with a fixed raster-scan order hardcodes a single factorization of the image distribution into the model, leaving it unable to inpaint, outpaint, or edit at inference time. Existing solutions—including the modern discrete-diffusion family—train from scratch on a mixture of token orderings, paying roughly 3× the compute of standard raster-scan training; discrete diffusion further sacrifices compatibility with the AR inference ecosystem (KV caching, vLLM) by construction. We argue that this cost is unnecessary. A pretrained raster-scan AR model has already learned the joint distribution over image tokens; what it lacks is not knowledge but the ability to express that knowledge through non-sequential conditional pathways. We propose a recipe for adapting pretrained raster-scan models into any-order generators using two components: future-aware positional embeddings that resolve a target-position identifiability problem at negligible overhead, and an ordering curriculum that traces a stable optimization path from the pretrained solution to a good any-order basin. Applied to four base models across two architectures and two scales (LlamaGen-L/XL, RAR-L/XL), the recipe substantially outperforms MaskGIT on ImageNet inpainting (FID 6.21 vs. 10.44) and outpainting (FID 5.84 vs. 12.17), preserves generation quality close to the pretrained checkpoints, and costs only ≈25% additional compute beyond pretraining (on LlamaGen-XL)—while leaving the standard AR inference stack untouched. Inpainting, outpainting, and editing, our results suggest, need not be added to autoregressive image models. They can be unlocked from the ones that exist.


URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment

Bogdan Zagribelnyy ⋅ Ivan Ilin ⋅ Nikita Bondarev ⋅ Anton Morgunov ⋅ Arkadii Lin ⋅ Maksim Kuznetsov ⋅ Rim Shayakhmetov ⋅ Vlad Aladinskiy ⋅ Alex Aliper ⋅ Alex Zhavoronkov

Synthesis planning aiming to find pathways of reactions for a target molecule is one of the most important and challenging tasks in drug discovery. Recent progress has produced both specialized deep-learning retrosynthesis systems and general-purpose large language models, but objective comparison remains difficult due to the lack of flexible, chemically interpretable benchmarking protocols. In the current study, we are introducing the URSA (Utilitarian RetroSynthesis Assessment) evaluation framework that provides the opportunity to benchmark the synthetic routes not only from a formal perspective, such as convergence to commercially available starting materials, but also from a chemical plausibility perspective, mimicking the way expert chemists evaluate the reactions and routes. The study covers a comprehensive evaluation of both conventional end-to-end retrosynthesis solutions and LLMs for the synthesis planning task on a set of novel, diverse target molecules with undisclosed synthetic routes, which represent realistic tasks in the daily drug design routine. We find that while LLMs can support high-level strategic planning, they currently underperform specialized retrosynthesis models in reliably solving synthesis planning tasks.


USAD: Uncertainty-aware Statistical Adversarial Detection

Zhijian Zhou ⋅ Xunye Tian ⋅ Jiacheng Zhang ⋅ Zesheng Ye ⋅ Yiyi Guo ⋅ Donghao Zhang ⋅ Liuhua Peng ⋅ Feng Liu

Statistical adversarial detection (SAD) treats detection as a two-sample test. Given a reference set of clean examples (CEs) and a batch of queries, potentially containing an unknown mixture of CEs and adversarial examples (AEs), SAD decides whether the query distribution drifts away from the CE distribution while controlling the false-alarm rate. Existing SAD-based methods mainly use maximum mean discrepancy (MMD) to measure the distributional discrepancy. However, MMD's distributional properties limit its ability to capture characteristic uncertainty patterns of AEs that are crucial for detection: AEs typically exhibit abnormal feature spread (i.e., global uncertainty) and instability under perturbations (i.e., local uncertainty). To close the gap, we propose Uncertainty-aware Statistical Adversarial Detection (USAD), which explicitly captures these uncertainty patterns with two new statistics: (1) Variance Discrepancy (VD), which measures the difference in feature spread between AEs and CEs to capture global uncertainty differences, and (2) Perturbation-based Covariance Discrepancy (PCD), which compares feature covariance under Gaussian perturbations to capture local uncertainty differences. By aggregating VD and PCD, USAD achieves superior detection performances over baseline methods against various adversarial attacks, highlighting the importance of considering characteristic behaviors of AEs for effective SAD. Our code is available at: https://anonymous.4open.science/r/USAD.


Utility-Driven Clustered Federated Learning via Class-wise Interaction Attribution

Zhongxu Wang ⋅ Yichen Li ⋅ Haozhao Wang ⋅ Hanlin Cai ⋅ chenqi ⋅ Ruixuan Li ⋅ Ozgur Akan

Clustered Federated Learning (CFL) mitigates statistical heterogeneity by grouping clients with similar data distributions. However, existing CFL frameworks predominantly rely on proxy metrics, such as gradient or data similarities, which collapse fine-grained class-level similarity into macroscopic client-level scores. This coarse-grained aggregation inherently masks critical class-wise conflicts, inadvertently exposing the federation to severe negative transfer. In this paper, we propose a fundamental paradigm shift from heuristic similarity to explicit value attribution. We introduce FedCap, a novel utility-driven CFL framework that explicitly quantifies class-wise marginal utilities to capture the true generalization impact of local updates. Guided by this attribution mechanism, FedCap employs a conflict-aware clustering algorithm to systematically eliminate intra-cluster negative transfer, while orchestrating safe, supply-demand-guided representation sharing across clusters. Crucially, we provide rigorous theoretical guarantees demonstrating the robustness of our underlying attribution mechanism against four major types of distribution shifts. Extensive experiments across diverse heterogeneity settings show that FedCap consistently achieves state-of-the-art accuracy. The code is available for anonymous access at https://anonymous.4open.science/r/FedCap-605E.


V2VFusion: Text-Controlled Video-to-Video Diffusion for Degradation-Aware Video Fusion

Jiajun Chen ⋅ Han Xu ⋅ Yunfei Huang ⋅ Zhuoya Zhang ⋅ Guangcan Liu ⋅ Jiayi Ma

Video fusion integrates complementary information in multiple source video streams into a single sequence for complete scene perception. Existing methods perform frame-wise fusion and lack unified spatio-temporal modeling across sequences, damaging temporal coherence across sequences. When dealing with degradations, they also ignore temporal dynamics and overload a single network with diverse degradations, leading to limited and imbalanced degradation suppression. To address these issues, we propose V2VFusion, the first unified video-to-video fusion framework that directly generates temporally coherent, high-quality fused videos from degraded multi-source inputs. By reformulating video fusion as a video-to-video conditional generation problem in latent diffusion space, our approach holistically models spatio-temporal dependencies across entire sequences and intrinsically ensures sequence-level consistency. On this basis, a text-guided degradation-conditioned ControlNet bridges high-level textual degradation descriptions with low-level visual priors, enabling interpretable and controllable guidance for structure-preserving fusion beyond purely data-driven conditioning. To handle heterogeneous and compound degradations, a state-modulated degradation-aware hierarchical mixture-of-experts module decouples expert selection into diffusion-state-aware filtering and degradation-aware routing. It prevents unreliable routing and fosters expert specialization to mitigate representation conflicts among diverse degradations. Experiments on infrared-visible, multi-exposure, and multi-focus video fusion tasks and complex degradations validate that V2VFusion surpasses existing methods in fusion quality and temporal stability.

Human values are not single labels, but distributions: they vary across individuals, contexts, and their relationships with other values. Yet most alignment approaches still compress group-level values into scalar scores or single representative targets, erasing within-group diversity and missing how value expression shifts across situations. We introduce Distributional Value Profiling, a framework for representing group-level value expression as distributions over expression modes, intensity, and inter-value relationships. Using large-scale text corpora spanning cultural, political, and religious groups, we construct profiles that capture not only which values are expressed, but also how they are expressed, how strongly they are emphasized, and how they relate to one another. We then model context-dependent shifts in these profiles using optimal transport, revealing structured redistributions of value expression across situations. Finally, we show that these profiles are actionable: a lightweight profiling model generalizes to unseen value-context combinations, and an activation-based steering method guides model outputs toward target distributions at inference time, consistently outperforming baselines across diverse benchmarks. Together, these results establish distributional value profiling as a concrete step toward pluralistic alignment, treating within-group diversity not as noise to be averaged away, but as structure to be measured, predicted, and controlled.


Values as Style: Disentangling Values from Semantics with One-Way Mixing for Low-Damage LLM Steering

Jiale Dai ⋅ Hongcan Deng ⋅ Liuxian Ma ⋅ Xiaoke Niu ⋅ Guojie Song

Activation steering can change an LLM's behavior without updating its weights, but a direction intended to change safety or value stance often also shifts topic, phrasing, and task content. We study a narrower and directly testable question: can a frozen residual-stream state be equipped with an editable interface that separates semantic content from value framing well enough to reduce this collateral damage? We propose a lightweight dual-code interface with a one-way semantic$\rightarrow$value path. Semantic information may ground value recognition, while stop-gradient gating, topic de-confounding, and swap consistency discourage value supervision from rewriting the semantic factor. At inference time, we edit only the value code through a residual delta update. Across two instruction-tuned backbones, the learned interface improves the steering--damage trade-off over dense and sparse steering baselines: it preserves semantics better, reduces topic leakage and benign false refusals, and remains stable under noise, distribution shift, and recomposer controls.


V-CAST: Video Curvature-Aware Spatio-Temporal Pruning for Efficient Video Large Language Models

Xinying Lin ⋅ Xuyang Liu ⋅ Yiyu Wang ⋅ Teng Ma ⋅ Jiasheng Li ⋅ Zichen Wen ⋅ Wenqi Ren

Video large language models (VideoLLMs) show strong capability in video understanding, yet long-context inference is still dominated by massive redundant visual tokens in the prefill stage. We revisit token compression for VideoLLMs under a tight budget and identify a key bottleneck, namely insufficient \textit{spatio-temporal information coverage}. Existing methods often introduce discontinuous coverage through coarse per-frame allocation or scene segmentation, and token merging can further misalign spatio-temporal coordinates under MRoPE-style discrete $(t,h,w)$ bindings. To address these issues, we propose V-CAST (\textbf{V}ideo \textbf{C}urvature-\textbf{A}ware \textbf{S}patio-\textbf{T}emporal Pruning), a training-free, plug-and-play pruning policy for long-context video inference. V-CAST treats frame representations as a semantic trajectory and uses local curvature as a low-cost proxy for temporal turns, routing token budgets to transition regions without relying on decoder attention. It further adopts a dual-anchor spatial selection mechanism that preserves high-entropy visual evidence without attention intervention, while keeping retained tokens at their original coordinates to maintain positional alignment. Extensive experiments across multiple VideoLLMs of different architectures and scales demonstrate that V-CAST achieves \textbf{98.6\%} of the original performance, outperforms the second-best method by \textbf{+1.1\%} on average, and reduces peak memory and total latency to \textbf{86.7\%} and \textbf{86.4\%} of vanilla Qwen3-VL-8B-Instruct. \textit{Our code is provided in the supplementary material.}

GRPO-style reinforcement learning has advanced vision-language reasoning, but its sequence-level rewards provide limited credit assignment for individual reasoning tokens. On-policy distillation (OPD) addresses this issue by providing dense supervision on student-generated trajectories, yet standard OPD assigns similar importance to all teacher-preferred tokens, including template phrases and visually irrelevant continuations. We introduce Visual Counterfactual On-Policy Distillation (VC-OPD), a grounding-aware distillation method that prioritizes tokens whose correctness depends on image-specific evidence. After a teacher-supervised warm start that stabilizes response formats and reduces noisy rollouts, VC-OPD performs on-policy distillation with two teacher evaluations for each student trajectory: one under the original image and one under a counterfactual image with corrupted instance-specific evidence. The original-image teacher distribution determines the distillation target, while the original--counterfactual teacher gap estimates token-level visual dependence and softly reweights the loss. To avoid unstable response-level update scales, VC-OPD further applies mass-preserving normalization to the token weights. This design preserves the dense supervision of OPD while shifting learning capacity toward visually grounded reasoning tokens. Experiments on multimodal reasoning benchmarks show that VC-OPD improves visual reasoning performance and produces more effective token-level distillation.


VERITAS: Veracity-Enhanced Robust Identification of LLM-generated Text Against Style-shifts

Xiaoquan Yi ⋅ Haozhao Wang ⋅ Yichen Li ⋅ Wenchao Xu ⋅ Jinghua Zhang ⋅ Yuhua Li ⋅ Rui Zhang ⋅ Imran Razzak ⋅ Ruixuan Li

With the rising prominence and fluency of large language models (LLMs), developing technologies to identify LLMs-generated text has become increasingly critical. However, existing technologies depend on static linguistic features, which can be evaded as advanced models increasingly mimic a wide range of writing styles. The study reveals two crucial vulnerabilities in existing detection systems:(1) State-of-the-art detectors suffer a substantial accuracy decline, reaching up to 11.43% when exposed to style-based adversarial rewrites generated by LLMs. (2) While general-purpose LLMs exhibit remarkable zero-shot capabilities, their performance in detecting adversarially manipulated text is significantly lower than specialized detectors fine-tuned for robustness. To address these vulnerabilities, we propose a novel style-agnostic detection framework named SAFD that enhances detection accuracy and robustness by prioritizing content-driven features over stylistic attributes. Our approach integrates a style-invariant training paradigm to disentangle content semantics from stylistic variations. We leverage adversarially enriched datasets constructed using LLMs fine-tuned for diverse style-based rewrites. Furthermore, we utilize advanced representation learning techniques to extract content-centric features, emphasizing semantic coherence, logical consistency, and factual alignment. Experimental results across multiple datasets and detection models validate the effectiveness of our framework, showing significant improvements in detection accuracy and robustness against diverse adversarial manipulations. The dataset and code are in the link https://anonymous.4open.science/status/A-Style-Agnostic-Framework-for-Detecting-LLM-Generated-Text-90B7.


VESPO: Variational Sequence-level Soft Policy Optimization for Off-Policy LLM Training

Guobin Shen ⋅ Chenxiao Zhao ⋅ Xiang Cheng ⋅ Lei Huang ⋅ XingYu Li

Off-policy updates are inevitable in reinforcement learning (RL) for large language models (LLMs) due to rollout staleness from asynchronous training and mismatches between training and inference engines. Naive importance sampling gives an unbiased correction but suffers from high variance, which is amplified by unbounded ratios and autoregressive generation. Prior remedies either rely on scenario-specific engineering, or trade bias for variance via token-level clipping or sequence-level normalization, yet these approaches remain largely heuristic. We propose **V**ariational s**E**quence-level **S**oft **P**olicy **O**ptimization (**VESPO**). By explicitly incorporating variance reduction into a variational formulation, we derive a principled closed-form reshaping kernel that operates directly on sequence-level importance weights, avoids token-level approximation and length normalization, and admits an explicit variance bound for the deployed kernel. Experiments on math reasoning and code generation show that VESPO maintains stable training under severe off-policy conditions (staleness up to $64\times$) and delivers consistent gains across both dense and Mixture-of-Experts (MoE) models, outperforming recent reshaping baselines under matched setup.


VESTA: Visual Exploration with Statistical Tool Agents

William Rudman ⋅ Abhishek Divekar ⋅ Kanishk Jain ⋅ Sebastian Joseph ⋅ Stella Offner ⋅ Matt Lease ⋅ Kyle Mahowald ⋅ Greg Durrett ⋅ Junyi Jessy Li

Fitting quantitative models to data is a central step in scientific workflows, yet it remains one of the least automated. Recent agent-based systems leverage language and vision-language models (VLMs) to iteratively propose and refine statistical models, but these systems struggle on more challenging modeling tasks. To address these limitations, we introduce VESTA (Visual Exploration with Statistical Tool Agents), a framework that equips VLMs with a dynamically growing exploration toolkit to guide model refinement through data transformations, hypothesis-driven visualizations, and robust statistical tests. Unlike prior systems that rely on iterative critique alone, VESTA actively explores data before and during refinement by selecting or creating diagnostic tools, which accumulate in a reusable registry across iterations. We evaluate VESTA against established baselines in three toolkit configurations: no tools, static expert-written tools, and dynamic model-written tools. To support this evaluation, we introduce DAWN (Dataset for Automated Workflows and Numerical Modeling), a benchmark targeting distribution fitting and time series modeling with varying difficulty tiers, and culminating in real-world astronomy tasks including modeling initial mass functions and gravitational-wave chirp signals. We find that VESTA's dynamic tool creation outperforms prior agentic pipelines, with the largest gains on complex and domain-specific tasks. We further show that dynamically generated tools are substantially more sophisticated than those produced by existing visual tool-creation systems, covering more diagnostic categories per function and strongly preferring visual outputs that the VLM critic can reason over directly.


ViCoR: Estimating Visual Necessity via Counterfactual Residuals for Multimodal Medical Data Selection

Haotian Gan ⋅ Yudong Li ⋅ Xintian Li ⋅ Jiarui Han ⋅ Weidong Tang

Medical vision-language models often exhibit hidden stratification, where high overall performance masks failures on subtle visual findings and overlapping clinical conditions. We trace this failure mode to modality imbalance during fine-tuning, where strong linguistic priors can suppress visually grounded learning and induce plausible but weakly grounded rationales. Existing data selection methods usually score image-text pairs as holistic units, without separating visually necessary samples from shortcut-solvable ones. To address this problem, we propose ViCoR, a counterfactual-residual-based data selection method for Med-VLM fine-tuning. ViCoR estimates visual necessity through image-dependent optimization residuals, calibrates the resulting scores with medical supervision compatibility, and constructs a diversity-aware fine-tuning subset. Across two backbones, four external generalization benchmarks, and one multi-disease robustness benchmark, ViCoR-selected 20\% subsets achieve the strongest average performance compared with full-data fine-tuning and selection baselines. Gradient-level analysis and perturbation audits further show that ViCoR enriches image-dependent training signals and strengthens image-level and evidence-region dependence. The code will be released to support future research.

Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a dominant paradigm for enhancing Large Language Models (LLMs) reasoning, yet its reliance on external verifiers limits its scalability. Recent findings suggest that RLVR primarily functions by eliciting latent capabilities, motivating the development of verifier-free algorithms. However, in such settings, standard methods like Group Relative Policy Optimization face a critical challenge: destructive gradient variance that often leads to training collapse. To address this issue, we introduce Verifier-Independent Curriculum Reinforcement Learning (VI-CuRL), a framework that leverages the model's intrinsic confidence to construct a curriculum independent from external verifiers. By prioritizing high-confidence samples, VI-CuRL effectively manages the bias-variance trade-off, specifically targeting the reduction of action and problem variance. We provide a rigorous theoretical analysis, proving that our estimator guarantees asymptotic unbiasedness. Empirically, VI-CuRL promotes stability and consistently outperforms verifier-dependent/independent baselines across math and general reasoning benchmarks with/without verifiers.


ViDiC: Video Difference Captioning

Jiangtao Wu ⋅ Shihao Li ⋅ Zhaozhou Bian ⋅ Jialu Chen ⋅ Yiwen He ⋅ Runzhe Wen ⋅ An Ping ⋅ JiakaiWang ⋅ Yuanxing Zhang ⋅ Jiaheng Liu

The rapid advancement of controllable video editing and agent-driven video generation has exposed a critical bottleneck: the inability of existing Multimodal Large Language Models to accurately perceive and articulate fine-grained differences between videos. Whether diagnosing editing fidelity or detecting training data quality issues, explicit comparative perception serves as an essential infrastructure. To address this capability gap, we introduce the ViDiC task and its corresponding ViDiC-1K benchmark. ViDiC-1K comprises 1,000 curated video pairs annotated with 3,720 fine-grained comparative checklist items, rigorously structured across seven dimensions: subject, style, background, camera work, motion, position, and playback techniques. To ensure reliable evaluation and prevent model hallucination, we propose a novel dual-checklist framework that assesses the accuracy of both similarities and differences separately based on an LLM-as-a-Judge protocol. Extensive experiments on 16 representative MLLMs reveal significant limitations in their comparative reasoning and dynamic difference perception. We hope ViDiC-1K serves as a foundational benchmark to advance continuous multi-video captioning and edit awareness in multimodal intelligence.


VISD: Enhancing Video Reasoning via Structured Self-Distillation

Hao Lin ⋅ Kunyang Lv ⋅ Xu Jiang ⋅ Jingqi Tian ⋅ Zhongjing Du ⋅ Jiayu Ding ⋅ Qiaoman Zhang ⋅ Hongbo Jin

Training VideoLLMs for complex reasoning remains challenging due to sparse sequence-level rewards and the lack of fine-grained credit assignment over long, temporally grounded reasoning trajectories. While reinforcement learning with verifiable rewards (RLVR) provides reliable supervision, it fails to capture token-level contributions, leading to inefficient learning. Conversely, existing self-distillation methods offer dense supervision but lack structure and diagnostic specificity, and often interact unstably with reinforcement learning. In this work, we propose VISD, a structured self-distillation framework that introduces diagnostically meaningful privileged information for video reasoning. VISD employs a video-aware judge model to decompose reasoning quality into multiple dimensions, including answer correctness, logical consistency, and spatio-temporal grounding, and uses this structured feedback to guide a teacher policy for token-level supervision. To stably integrate dense supervision with RL, we introduce a direction–magnitude decoupling mechanism, where environment rewards determine update direction, while structured privileged signals modulate token-level update magnitudes. This design enables semantically aligned and fine-grained credit assignment, improving both reasoning faithfulness and training efficiency. Additionally, VISD incorporates independent advantage estimation, curriculum scheduling, and EMA-based teacher stabilization to support robust optimization over long video sequences. Experiments on various benchmarks demonstrate that VISD consistently outperforms strong baselines across diverse video reasoning tasks, achieving gains in accuracy, grounding quality, and interpretability. Notably, VISD significantly accelerates convergence, highlighting the effectiveness of structured self supervision in improving both performance and sample efficiency for VideoLLMs.


VisInteract: Towards Dynamic Interactive Text-to-Visualization under Imperfect Queries

Wenxin XU ⋅ Jinwei Lu ⋅ Hwanhee Kim ⋅ Chen J Zhang ⋅ Xiao-Yong Wei ⋅ Haoyang LI ⋅ Yuanfeng SONG

Real-world visualization requests are routinely ambiguous, incomplete, or factually incorrect, yet existing Text-to-Visualization (Text-to-Vis) systems assume well-specified inputs and produce charts in a single pass. When queries are imperfect, a system must interact with the user to recover the true intent, but no benchmark or method supports this dynamic process. We introduce VisInteract, a new paradigm that reframes Text-to-Vis as interaction-driven intent recovery, and VisInteract-Bench, to our knowledge, that is the first benchmark for dynamic interactive Text-to-Vis, featuring controlled imperfection injection, a leakage-controlled User Agent for realistic multi-turn feedback, and dual-perspective (code and chart) automated evaluation. On the algorithmic side, we propose Vis-MCTS, a Monte Carlo Tree Search (MCTS) enhanced method, introducing improvements over classical MCTS, that Progressive Widening to tame the unbounded tool-argument space in tree search, cross-rollout information sharing so clarifications and critiques benefit the entire search tree, and Dimension-Aware Reward Decomposition that routes scalar user feedback along data-fidelity, visual-design, and intent-alignment dimensions to resolve credit assignment across heterogeneous actions. Extensive Experiments across two LLM backbones show that Vis-MCTS consistently outperforms all Text-to-Vis baselines, improving end-to-end task success by 13.40\%--16.27\% over the strongest interactive baseline and by more than 5\times over non-interactive ones.


Visual Anchoring for Scenario-Guided Forecasting

Patara Trirat ⋅ Jay Heo ⋅ Heejun Lee ⋅ Sung Ju Hwang

A multimodal large language model (MLLM) forecaster takes a forward-looking scenario as a soft constraint that decays with horizon. The decay is expected, but its rate is uncharacterized and untunable at inference time. A practitioner running a twelve-month stress test cannot tell whether the scenario binds for twelve months or twelve days. Existing accounts of autoregressive error, such as exposure bias and hallucination snowballing, bound a single rollout against ground truth and offer no inference-time lever. Under three assumptions, we prove a closed-form Lipschitz upper envelope on the per-step scenario drift between two paired-seed forecasts differing only in scenario text. Two identifiable parameters partition behavior into bounded, linear, and exponential regimes. A sister within-chunk bound holds under any chunk partition. It motivates a stabilization heuristic that we test and falsify: periodic re-injection of the scenario enlarges drift relative to a single injection, and length-matched random text shrinks it. The envelope is respected; the heuristic is not. A pre-registered attention diagnostic across the open-weight panel returns architecturally heterogeneous verdicts that the recurrence framing covers uniformly while no single attention-dilution mechanism does. In place of text re-injection, we propose Visan (Visual Anchoring), a chart-anchored multimodal prompt that renders the lookback, the scenario, and the historical context as a single image. Across a panel of MLLMs and two long-horizon scenario-conditioned forecasting benchmarks, Visan reduces forecast error on the majority of MLLMs.


Visual Grounding First, Multimodal In-context Learning Follows

Minhyuk Seo ⋅ Minjae Lee ⋅ Chaeeun Lee ⋅ Wei Lin ⋅ Muhammad Jehanzeb Mirza ⋅ Tinne Tuytelaars ⋅ Jonghyun Choi

LLMs have shown strong in-context learning (ICL) capability, but extending it to Multimodal LLMs (MLLMs) remains challenging. Prior multimodal ICL methods often rely on ICL-specific datasets for additional training or task vector extraction, which can improve performance on the same ICL benchmark used for adaptation. However, they do not generalize well to other benchmarks and often induce forgetting in previously well-solved tasks. For this reason, we identify `visual grounding' as the key bottleneck in multimodal ICL; MLLMs often fail to attend to task-relevant visual evidence in demonstrations, instead over-relying on textual cues and producing hallucinated outputs. To address this, we propose ProCoRe to enhance visual grounding through contrastive reinforcement learning on multimodal contrastive data for improving multimodal ICL. ProCoRe generates captions for contrastive image pairs, decomposes them into propositions, and optimizes contrastive rewards that encourage alignment with corresponding images while discouraging mismatched ones. Despite never seeing ICL-formatted training examples, ProCoRe improves average multimodal ICL classification accuracy by 5.6% and captioning ROUGE-L by 13.8% over the off-the-shelf Qwen-VL-3-8B model, outperforming ICL-data-dependent state of the arts across benchmarks. Finally, we introduce PICL, a personalized multimodal ICL captioning benchmark that tests whether MLLMs genuinely understand visual demonstrations rather than relying on pre-trained semantic priors.


VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models

Shengyu Si ⋅ Yuanzhuo Lu ⋅ Ruimeng Yang ⋅ Ziyi Ye ⋅ Zuxuan Wu ⋅ Yu-Gang Jiang

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robotic manipulation, yet they still struggle to generalize to unseen tasks that necessitate transferring relevant experience across objects, scenes, and action patterns. This paper proposes VLA-Pro, a plug-and-play framework designed to enhance cross-task generalization by storing task-relevant procedural memories at training time and transferring these memories during inference. Specifically, VLA-Pro stores task-specific LoRA adapters as parameterized procedural memories during training. At inference time, VLA-Pro retrieves relevant procedural memories based on the current multi-modal context and dynamically fuses these memories for generating the current action chunk. Experiments on RoboTwin, RLBench, and real-world manipulation tasks show that VLA-Pro consistently improves cross-task generalization across multiple backbones, achieving up to a 207\% relative improvement in simulation and increasing real-world success rate from 5.8\% to 65.0\%. These results suggest that procedural memory retrieval and adaptation provide an effective mechanism for transferring manipulation experience to novel tasks while preserving modularity and execution stability. The code is available at https://anonymous.4open.science/r/VLA-Pro/.


VUM: Visual Unified Models for Image Generation and Perception

ZiDong Wang ⋅ Yiyuan Zhang ⋅ Xiaoyu Yue ⋅ Xiangyuan Xue ⋅ Zhangquan Chen ⋅ Manyuan Zhang ⋅ Wanli Ouyang

Unifying visual perception and generation poses a fundamental dichotomy: perception extracts semantics but discards details, while generation synthesizes fine-grained structures. Thus, existing works rely on decoupled architectures, as forcing these opposing information flows into one network inevitably triggers task interference and representation collapse. To break this bottleneck, we introduce the Visual Unified Model (VUM), a native framework that unifies both capabilities within a single architecture. Our key insight is that visual perception and generation can be unified by re-envisioning Masked Image Modeling and Diffusion Processes as complementary facets of data recovery. By establishing a shared degradation space that integrates both masking and noising, VUM optimizes a joint signal recovery objective: simultaneously predicting masked patches and denoising corrupted signals. This objective enforces the dual internalization of semantic abstraction and generative synthesis. Using a shared, unified backbone, VUM delivers highly competitive performance across a diverse spectrum of tasks, encompassing image generation, visual recognition and dense prediction tasks. Codes and pretrained checkpoints will be made publicly available.


Wayfinder: Adaptive Resource Routing from Agent Citations

Miguel Romero Calvo ⋅ George Karypis

To act effectively, agents must seek information across heterogeneous applications including chat channels, issue trackers, email, knowledge bases, and code repositories, yet resource routing in such environments is inherently unknown to model providers: where relevant information resides varies across users and organizations and cannot be enumerated at training time. Existing routing approaches are typically trained offline under fixed assumptions or rely on static heuristics, limiting their ability to transfer across environments without target-specific supervision. We introduce \method, a routing policy trained via reinforcement learning from interaction feedback that enables adaptive resource selection while treating each resource as a black-box retriever. \method leverages an interaction-derived, citation-gated memory to prioritize candidate resources for downstream retrieval, while allowing the primary agent to expand its search when needed. We evaluate \method in controlled federated environments spanning both intra-corpus partitions (NFCorpus, FEVER) and cross-corpus heterogeneity (Tech Startup) and compare against supervised and zero-shot routing baselines. Interaction-driven adaptation outperforms all baselines on the intra-corpus benchmarks, where description-based discrimination is weak, and ties a fully supervised classifier on Tech Startup at matched top-$1$ output without using any target-environment labels.


Weakly Supervised Concept Learning for Interpreting and Attributing LVLM Predictions

Md Abdul Kadir ⋅ Omair Shahzad Bhatti ⋅ Daniel Sonntag

Large vision–language models (LVLMs) are inherently opaque, making it difficult to determine whether their outputs are grounded in visual evidence or driven by language-model priors. Existing concept-based interpretability methods are limited to classification settings, rely on proxy models, or require predefined tokens, and thus do not extend LVLMs. We propose Text-Guided Concept Learning (TGCL), a weakly supervised framework for extracting multimodal concept vectors in LVLMs without token supervision. TGCL builds concept-to-image mappings from data and extracts patch-level activations via concept-guided probing. It then formulates concept learning as a contrastive disentanglement problem, isolating concept-specific patches from background patches to produce sparse, stable, and semantically aligned concept vectors. We conduct experiments on four datasets—ImageNet, MSCOCO, CIFAR100, and DTD—and three recent LVLMs. TGCL outperforms recent interpretability methods, achieving up to 4\% higher sparsity, 11\% lower instability, 17\% lower overlap, and 20\% improvement on attribution faithfulness compared to state-of-the-art baselines.

Spatial intelligence is a fundamental capability of embodied Artificial Intelligence systems, and its reliable assessment and optimization require well-designed benchmark. Existing benchmarks for spatial intelligence largely rely on video datasets or procedurally generated data, which limits data scalability and ecological validity. To address these limitations, we propose to automatically extract spatial information directly from executable environments. Specifically, we design WebSpatial, a benchmarking framework built upon web runtime environments, which enables automatic acquisition of spatial data by parsing underlying code and interaction processes. It performs runtime scene graph parsing on web 3D applications to automatically extract visual attributes, world coordinates, and user interaction event sequences of objects, while constructing a dedicated spatial computation tool library. The modular generation pipeline integrates LLM-based task parsing and planning and tool-assisted execution to automatically generate question-answer pairs. We conduct experiments in method effectiveness validation and benchmark evaluation. For validation, we demonstrate that WebSpatial can reliably construct diverse and scalable spatial tasks grounded in both static and dynamic scenes, covering spatial perception including color, shape and quantity, and spatial reasoning including relation assessment, direction transformation and mental rotation. We develop a small-scale benchmark named WebSpatialBench based on WebSpatial, and use it to evaluate 10 multimodal large language models, revealing their strengths and limitations in spatial problem-solving.

LLM benchmark scores are often read as claims about factual accuracy, evidence use, instruction following, or logical reasoning. We study this score-to-claim conversion as inference over a completed evaluation record. Given benchmark items, prompts, outputs, scores, and native fields, which capability sentence is supported, which stronger sentence remains outside the evidence, and would that boundary change model choice? We formalize a report-use procedure, \emph{score-to-claim inference}: hold source units fixed, compare each claim-induced condition to benchmark-native controls or same-source invariance checks, and apply a declared finite-family support rule. In a fixed set of five open-weight model families, completed records support narrower sentences than endpoint readings alone suggest: supplied-context answer recovery remains separate from support attribution; local supplied-evidence verdict signals leave open-web provenance unmeasured; and SATBench exposes label-surface instability for SAT/UNSAT prediction. The output is a support boundary, a blocked strengthening, and the next measurement or model-selection sensitivity attached to the reported result.


What Stops SGD on LLM Pre-Training: The Need for Large Learning Rate and How to Achieve it

Athanasios Glentis ⋅ Chung-Yiu Yau ⋅ Dawei Li ⋅ Mingyi Hong

It is widely believed that stochastic gradient descent (SGD) performs significantly worse than adaptive optimizers such as Adam in pre-training Large Language Models (LLMs). Yet the underlying reason for this gap remains unclear. In this work, we attribute a large part of the discrepancy to SGD's inability to sustain learning rates comparable to Adam's much larger effective learning rates. Through empirical and theoretical analysis of LLM pre-training dynamics, we identify that training is characterized by small gradient norms and large weight-to-gradient ratios, an effect that becomes more pronounced with larger batch sizes typical in pre-training, necessitating such large effective learning rates. However, we find that output-layer gradient magnitudes become highly uneven across token classes, and that large gradient spikes frequently occur during training. Together, these effects severely restrict the admissible learning rate of SGD. Guided by this understanding, we show that simple clipping mechanisms that stabilize SGD at large learning rates enable it to recover most of Adam's performance. In our large-scale experiments, the validation loss gap between large-learning-rate SGD and Adam shrinks from more than 50% to only about 3.5% when pre-training a 1B-parameter LLaMA model with a 1M-token batch size.


When Alignment Fails: Stabilizing Cross-Dynamics RL with Prototype Trust Regions

Dong Uk Kim ⋅ Ji Su Yoon ⋅ Eui-Nam Huh ⋅ Choong Hong

Alignment-based representation learning for cross-dynamics RL admits a degenerate global minimizer: per-domain feature covariances collapse to a common point, driving the alignment loss to zero while transfer regret remains large and rendering the standard Bures-Wasserstein (BW) transfer regret bound vacuous. We document this collapse across multiple environments, alignment kernels (BW, Frobenius, MMD), and methods (BW-CORAL, VGDF), where per-domain effective rank drops from p to $\approx 1$ within a few epochs—structurally the same failure as representation collapse in self-supervised learning. We propose a minimal remedy: a per-domain Prototype Trust Region (TR) that anchors each empirical covariance to a slow-moving EMA prototype, introduces no learnable parameters, and acts as an explicit second-moment stabilizer, mirroring momentum targets and covariance regularizers in SSL. Because TR depends only on per-domain covariances, it is combined with any BW-based method; adding it to VGDF inherits the same resistance to collapse. Theoretically, we prove a BW non-collapse lower bound that restores the transfer-regret guaranty to a non-vacuous form, and present a Representation Stability Transfer Bound that decomposes regret into coverage and stability terms TR directly controls. Empirically, TR delivers statistically significant gains under pre-registered tests in collapse-dominated regimes (finite-sample, lower-dimensional) and is correctly neutral when other bottlenecks dominate. We frame TR not as a universal improvement but as a targeted fix for an identifiable and broadly shared failure mode.


When and What to Prune? Stage-Aware Visual Token Pruning for Efficient VLA

TIANJUN SHI ⋅ Haotian Xiong ⋅ Ziyu Gong ⋅ Qi Lu ⋅ Li Li

Visual token pruning is an effective way to accelerate vision-language models and is especially useful for vision-language-action (VLA) inference, where many visual tokens must be processed before predicting robot actions. Existing pruning methods usually estimate which tokens can be pruned based on attention scores or feature diversity, retaining tokens that are either highly attended or visually different from others. However, most of them use fixed pruning schedules, such as pruning once at a preset layer or pruning at uniformly spaced layers. Such schedules can be risky for VLA models, because the model may not know which visual regions matter for the action in early layers. Tokens that look unimportant at first may become useful after the model combines visual observations with the language instruction. In this work, we propose SAPrune, a training-free visual token pruning framework for efficient VLA inference. Instead of pruning at fixed layers, SAPrune uses a small calibration set to observe how action-to-visual attention changes across layers, and chooses pruning layers only after the attention pattern becomes more reliable. At each selected layer, SAPrune applies a dual-path pruning rule: one path protects strongly attended visual tokens from pruning, while the other prevents useful surrounding context from being discarded. Experiments on LIBERO, SIMPLER, and real-world robotic tasks show that SAPrune prunes 87.5% of visual tokens and achieves up to 1.718× inference speedup while maintaining competitive task success rates.

Sign-based optimization algorithms, such as SignSGD and Muon, have garnered significant attention for their remarkable performance in training large foundation models. Despite this empirical success, we still lack a theoretical understanding of when and why these sign-based methods outperform vanilla SGD. The core obstacle is that under standard smoothness and finite variance conditions, SGD is known to be minimax optimal for finding stationary points measured by $\ell_2$-norms, thereby fundamentally precluding any complexity gains for sign-based methods in standard settings. To overcome this barrier, we analyze sign-based optimizers leveraging $\ell_1$-norm stationarity, $\ell_\infty$-smoothness, and a separable noise model, which can better capture the coordinate-wise nature of signed updates. Under this distinct problem geometry, we derive matched upper and lower bounds for SignSGD and explicitly characterize the problem class in which SignSGD provably dominates SGD. Specifically, we compare the *upper bound of SignSGD* with the *lower bound of SGD*, illustrating that SignSGD effectively reduces the complexity by a factor of $d$ under *sparse noise*, where $d$ is the problem dimension. Furthermore, we elevate this framework to the matrix domain, providing an equivalent optimal lower bound for the Muon optimizer, proving that extending the sign operator to matrices preserves this optimal scaling with dimensionality. Finally, we bridge our theoretical bounds to practice, demonstrating that the theoretical superiority of SignSGD accurately predicts its faster convergence during the pretraining of a 124M parameter GPT-2 model.

Sliding-window human activity recognition (HAR) splits continuous sensor streams into fixed-length, overlapping windows and trains a classifier to assign one activity label to each window. Window length and stride are chosen from sampling rate, latency or computation constraints, and assumptions about typical motion duration, but they need not align with motion cycles or activity changes. Some windows therefore contain only part of a repeated motion, fragments shared by related activities, or mixed evidence near a transition. Yet each window still receives a single label, creating a granularity mismatch between activity-level supervision and local motion evidence. We propose PSDNet (Phase-Script Deliberation Network), which separates local phase interpretation from final activity classification. PSDNet encodes the current window into latent phase primitives, combines them with recent history, and produces an initial classification with an uncertainty score. When uncertainty remains high, a deliberation module compares the current phase evidence with learned class-specific phase scripts, i.e., compact templates of short-term phase patterns, and updates classification scores over a small candidate set. An auxiliary boundary objective encourages sensitivity to possible activity changes without serving as a hard routing rule. This design concentrates extra computation where direct classification is least reliable while keeping easy cases efficient. Experiments on eight public HAR benchmarks show that PSDNet consistently improves accuracy and weighted F1 over strong CNN- and sequence-based baselines. Additional analyses on ambiguous windows and similar-activity confusion pairs support phase-aware selective deliberation for windows whose local evidence is insufficient for reliable classification.


When a Zero-Shooter Cheats: Improving Age Estimation via Activation Steering

Erik Imgrund ⋅ Pia Hanfeld ⋅ Klim Kireev ⋅ Konrad Rieck

Different age-related regulations have been proposed to protect minors from harmful content and interactions online. Automated age estimation is central to enforcing such regulations, and vision-language models (VLMs) achieve state-of-the-art performance on this task. However, we find that the zero-shot nature of VLM-based age estimation produces an unexpected side effect we call the identity shortcut: Instead of estimating age from visual features, VLMs tend to identify the depicted person and infer their age from memorized knowledge. This phenomenon leads to substantially incorrect predictions when non-celebrities are misidentified as celebrities. It also produces deceptively high robustness to noise and adversarial perturbations on celebrity images, which dominate popular benchmarks. To mitigate this, we propose an activation steering method that suppresses the shortcut by intervening on the hidden states of the VLM. This method improves age estimation accuracy for both memorized and unseen identities, reducing mean absolute error by up to 25\% across popular benchmarks.

Large language models may encounter factual knowledge during pre-training yet fail to reliably use that knowledge after fine-tuning. We study this gap in a stylized one-layer self-attention + MLP transformer trained by next-token prediction and subsequently fine-tuned on question-answering data. We first prove that, under suitable regularity conditions, the model reaches near-optimal pre-training loss while learning structured attention patterns. We further show that fine-tuning turns the Q&amp;A prompt format into a trigger for pre-trained relation features, enabling the model to extract facts not revisited during fine-tuning. Our analysis reveals a relation-covering characterization for knowledge extraction: fine-tuning need not revisit every stored fact, but it must cover enough latent relation-template directions through which facts were encoded during pre-training. We prove that extraction improves with pre-training multiplicity and fine-tuning coverage, but becomes harder as the relation-template universe grows. Conversely, insufficient coverage yields a failure regime in which facts can be stored but not extracted, providing a stylized mechanism for hallucination. Our analysis covers both full and low-rank fine-tuning, and experiments on synthetic data and PopQA-based GPT-2/Llama models support the predicted trends.

Flow matching is increasingly used beyond open-ended generation, including prediction tasks whose outputs are effectively deterministic, yet it remains unclear why flow matching helps when diversity is not the objective. We study this question in histology-conditioned spatial transcriptomics (ST) prediction, a challenging sparse high-dimensional multimodal task where direct regression with pathology foundation features is already competitive. We disentangle several possible explanations and find that the gain does not come from simply adding a flow objective, injecting noisy target-side states, or increasing architectural complexity. We instead identify intermediate flow states as paired, time-indexed target-side signals: they contain condition-aligned information about the target, but are mixed with source noise and residual variation. Their benefit emerges only when this information is selectively routed into histology representations. Guided by this mechanism, we instantiate conditional flow matching with a simple gated self-attention to learn interactions between histology tokens and noisy ST states. Across high-resolution ST datasets, the model improves spatial correlation over regression and flow-based baselines, with stable convergence and favorable scaling under larger and sparser gene panels. These results establish a design principle for deterministic multimodal prediction: flow matching helps when intermediate target states are treated as structured cross-modal supervision and selectively integrated into conditional representations. Code is available at https://anonymous.4open.science/r/Understanding-When-Flow-Matching-Helps-Deterministic-Multimodal-Prediction-B426/.


When Does Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

Luca Viano ⋅ Antoine Moulin ⋅ Audrey Huang ⋅ Volkan Cevher ⋅ Philip Amortila ⋅ Dylan J Foster

Imitation learning (IL)—training an agent to replicate expert behavior from demonstrations—underpins applications from robotics to language model training. Standard approaches such as Behavior Cloning (BC) are known to suffer from compounding errors and performance plateaus, particularly when the learner cannot perfectly represent the expert’s policy (e.g., as is typical in distillation). Two interventions are widely understood empirically to improve performance: querying the expert interactively along the learner’s own trajectories, and using value function estimation en route to generating a policy. We investigate the nature of these improvements and their potentially surprising interplay. Our main finding is that expert interaction relaxes the representational demands on the learner: one only needs a model capable of realizing the expert’s value function, bypassing the (often stricter) requirement of realizing the expert’s policy itself. Concretely, we introduce OVI, an interactive IL algorithm that is statistically and computationally efficient whenever the learner can represent the expert’s value function. We complement this with a negative result showing that interaction is necessary: without significantly stronger representational assumptions than expert-value realizability alone, a broad class of value-based IL algorithms cannot succeed in the offline setting. These findings bear out empirically: OVI outperforms offline policy-based (BC), interactive policy-based (DAgger), and offline value-based IL methods, with the largest gains when the learner network is substantially less expressive than the expert’s.

LLM routing is designed to avoid overpaying for inference: easy queries should be handled by cheaper models, while stronger models should be reserved for cases where they are truly needed. However, we show that existing routers often violate this principle under loose cost budgets. As the budget increases, they increasingly route queries to the strongest and most expensive model, even when cheaper candidates achieve comparable or identical outcomes. We refer to this behavior as strong-model over-selection under loose budgets. Across both unimodal and multimodal routing benchmarks, this behavior weakens the cost-saving motivation of routing without necessarily improving quality. We diagnose this phenomenon through the mismatch between common router training objectives and deployment-time routing decisions. Many routers are trained to predict scalar performance scores, whereas cost-aware routing ultimately depends on query-specific comparisons among budget-feasible models, with cost used to distinguish equivalent or near-equivalent choices. In small-margin regimes, modest prediction errors can therefore flip relative orderings and induce cost-insensitive selections. Motivated by this diagnosis, we propose EquiRouter, a decision-aligned router that directly learns query-dependent model rankings with a Cost-Aware Ranking Objective and a lightweight Query--Model Interaction Representation. Experiments across RouterBench, MMR-Bench, MixInstruct, and RouterEval show that EquiRouter reaches the strongest model-level performance with lower relative cost than compared baselines, demonstrating consistent improvements across text-only, multimodal, continuous-utility, and large-model-pool routing settings.


When Poison Meets Structure: Topology-based Defense against Poisoning Attack on Graph-based Retrieval-Augmented Generation

Qizhi Chen ⋅ Junhao Wen ⋅ Shuang Liang ⋅ Jiakai Li ⋅ Yizhuo Ma ⋅ Rongzheng Wang ⋅ Ke Qin

Poisoning attacks against GraphRAG focus on knowledge pollution at the node and community levels, but conventional poisoning defenses mainly rely on sentence-level semantic features, which makes them less effective against such attacks. To address this issue, we analyze the poisoning strategy from a game-theoretic perspective and find that attackers prioritize fabricating query-related facts while ignoring supporting background knowledge, which in turn induces a systematic topological discrepancy between poisoned and clean subgraphs. Based on this insight, we propose the lightweight Topology-based Defense against Poisoning Attack on GraphRAG (TDP). TDP constructs a pair of conflicting candidate subgraphs from the retrieved evidence and trains a pairwise topology-ranking discriminator to distinguish clean evidence with dense cross-validation from poisoned evidence with sparse structural support, thereby removing poisoned subgraphs. Notably, the topological patterns captured by TDP reflect structural preferences induced by the poisoning game rather than dataset-specific distributional biases, allowing it to be transferred as a plug-and-play module after pretraining without end-to-end retraining. To the best of our knowledge, this is the first systematic work on poisoning defense for GraphRAG, and experiments show that TDP achieves state-of-the-art defense performance across multiple benchmarks and poisoning attacks, demonstrating strong generalization.


When Prompt Internalization Breaks: Continuous Experience Internalization in Large Language Models

Ze Chen ⋅ Haomai Zhang ⋅ Jiaxuan Zou ⋅ Zhuo Chen ⋅ Linhua Ye

Deployed large language models (LLMs) continually accumulate new experiences and rules, yet standard self-distillation fails to robustly internalize these updates over time. We identify and formalize Continuous Experience Internalization (CEI) collapse, a previously uncharacterized failure mode where recursive same-parameter self-distillation systematically forgets earlier patches, absorbs new ones poorly, and amplifies high-confidence errors. To diagnose this, we introduce CEI-Bench, a stage-wise benchmark quantifying task performance, old-patch retention, new-patch absorption, and error propagation. To mitigate CEI collapse, we propose Recycled Internalization Training (RIT), a framework separating stable consensus from unstable conflicts in teacher supervision. RIT combines multi-view variance isolation (MUSE) and targeted error recycling with cross-stage anchors (RACE), explicitly addressing the structural bottlenecks driving collapse. Across four backbones and three task families, RIT recovers 39%–72% of performance lost to vanilla recursive internalization and maintains stable trajectories. Our work reframes continuous self-distillation as a diagnosable phenomenon, providing a systematic analysis, mechanistic explanation, and principled mitigation for CEI. We position CEI collapse as a fundamental challenge for adaptive LLMs, and RIT as a framework for robust recursive internalization.


Where Are MLLMs Looking When They Hallucinate? Mitigating Visual Hallucination via Gaze Steering

Yebo Wu ⋅ Han Jin ⋅ Feng Liu ⋅ Changwang Zhang ⋅ Jun Wang ⋅ Li Li

Multimodal Large Language Models (MLLMs) have achieved remarkable progress, yet visual hallucination remains a critical barrier to reliable deployment. This paper aims to answer a fundamental question: when MLLMs hallucinate, where are they actually looking? By tracing the inter-layer consistency of visual attention, we reveal that hallucinated reasoning trajectories bifurcate into two pathological states: Attention Locking, where the gaze becomes overly focal and rigidly anchored to limited visual evidence, and Attention Collapse, where attention becomes excessively dispersed and fails to ground reasoning in meaningful cues. In contrast, faithful reasoning maintains a dynamic equilibrium between established visual evidence and broader peripheral exploration. Building on this insight, we propose ReGaze, a training-free framework that redirects the model's gaze toward this healthy equilibrium to mitigate hallucinations. Specifically, ReGaze monitors layer-wise gaze consistency and triggers timely interventions whenever an unhealthy gaze state emerges. For the locked state, ReGaze redistributes attention from over-dominant critical tokens to peripheral regions, encouraging the model to explore richer visual details. For the collapsed state, it gathers scattered attention from non-critical tokens and injects it into critical anchors, amplifying key visual signals. Experimental results demonstrate that ReGaze enables MLLMs to process visual information more effectively and significantly mitigates hallucinations.

Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it attractive for video understanding where annotation and computation are expensive. However, video PEFT has focused on adapting image-pretrained models, while standard PEFT methods can also be applied to video representations. These settings are rarely compared and both confine temporal reasoning to a single component of the model, leaving open how temporal context should be distributed across backbone, PEFT and probe. In this work we provide a systematic study of model adaptation strategies for video understanding. We evaluate methods across appearance-focused, motion-focused and spatially dense settings, with a particular focus on scenarios with limited data where parameter-efficiency is most beneficial. Our results provide new insights into PEFT and probing across settings and demonstrate the importance of temporal context allocation for effective video adaptation.


Where Root Cause Analysis Fails: A Retrieval-Reranking Decomposition

Hada M Muhammad ⋅ Luan Pham ⋅ Laure Barrière ⋅ Sachin Shetty ⋅ Leonardo Pulga ⋅ Flora Salim

Identifying root cause of an anomaly among hundreds of sensors is critical for preventing safety incidents and costly downtime in complex monitored systems. Existing studies evaluate root cause analysis (RCA) methods using top-k accuracy. We show that this metric has a fundamental blind spot: it conflates two failure modes: *retrieval failure*, where the true cause is never considered, and *reranking failure*, where it is considered but ranked too low. In this work, we introduce a retrieval-reranking decomposition and audit four well-known benchmarks to expose this blind spot. Our experiments show that, on benchmarks with complex faults, statistical baselines mis-rank the true cause 78-100\% of the time, and graph-based methods perform poorly because of unreliable causal graphs constructed using short fault windows. Meanwhile, on simple benchmarks where faults manifest significantly at their origin, retrieval is trivially solved at 100\%. Guided by the decomposition, we build a two-stage pipeline combining a multi-signal retriever with an LLM reranker that lifts top-1 accuracy by $+$2 to $+$26 points over the best baseline on complex-fault benchmarks and $+$9 to $+$13 points on simple benchmarks, with no causal graph or labeled data required.


Where to Connect? Boosting MLLMs via Dynamic Gated Pathways across ALL ViT and LLM Layers

Yingying Yan ⋅ Jiaqi Tang ⋅ Wei Wei ⋅ Qianzhou Wang ⋅ Jianmin Chen ⋅ Yuyang Xia ⋅ Botong Geng ⋅ Jinjian Wu ⋅ Lei Zhang ⋅ Qifeng Chen

A central yet underexplored question in Multimodal Large Language Models (MLLMs) is where to connect — how visual information from a hierarchical vision encoder should be wired into the layer-wise semantics of a large language model. Mainstream MLLMs inject a fixed single-layer visual representation into the LLM and treat all LLM layers as a single consumer, ignoring that different layers demand different visual granularities. Recent multi-layer fusion methods either compress hierarchical features on the encoder side or rely on predefined sparse or hierarchical layer-to-layer connections, and therefore stop short of modeling adaptive cross-layer interactions between the ViT and LLM hierarchies. We address this gap with DGP (Dynamic Gated Pathways), a routing mechanism that establishes all-to-all connections between every ViT layer and every LLM layer, with each LLM layer dynamically gating its preference over visual layers conditioned on its own semantic state. Built on top of LLaVA-1.5-7B and 13B, DGP delivers consistent gains over the LLaVA-1.5-7B baseline (e.g., +4.2 on SQA$^{I}$, +3.3 on MM-Vet, +3.1 on VizWiz and LLaVA$^{W}$) and surpasses prior multi-layer visual fusion methods on most benchmarks. Layer-wise routing analyses further show that different LLM layers do select different visual granularities on demand, supporting both the effectiveness and interpretability of the proposed dynamic gated pathways. We will open-source code/weight/demo soon.


Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video LLMs

Jongseo Lee ⋅ Hyuntak Lee ⋅ Sunghun Kim ⋅ Sooa Kim ⋅ Jihoon Chung ⋅ Jinwoo Choi

Video Large Language Models (Video-LLMs) have made rapid progress on temporal video understanding, yet many fail at a basic perceptual primitive: signed image-plane motion direction. On simple videos of a single object moving left, right, up, or down, most Video-LLMs perform near chance, with above-chance cases largely attributable to prediction biases rather than genuine direction understanding. We call this failure directional motion blindness. We localize the failure by tracing motion direction information through the Video-LLM pipeline. Motion direction remains linearly accessible from the vision encoder, projector, and LLM hidden states, but the readout fails to bind this signal to the correct verbal answer option, revealing a direction binding gap. Although synthetic motion direction instruction tuning reduces this gap on the source domain, motion direction concept vector analysis shows that visual complexity weakens the signal magnitude and limits out-of-domain generalization. We introduce MODIRECT, a dataset family for motion direction instruction tuning and evaluation, and DeltaDirect, a diagnosis-driven, projector-level objective that predicts normalized 2-D motion vectors from adjacent-frame feature deltas. On MODIRECT-SYNBENCH, instruction tuning with DeltaDirect improves motion direction accuracy from 25.9% to 85.4%. On MODIRECT-REALBENCH, DeltaDirect improves real-world motion direction accuracy by 21.9 points over the vanilla baseline without real-world tuning data, while preserving standard video-understanding performance.


Why Copy Others? Insights into Social Learning from Multi-Agent Reinforcement Learning

Yancheng Liang ⋅ Shakti Senthil ⋅ Daphne Chen ⋅ Simon Du ⋅ Natasha Jaques

When is it better to rely on information from others rather than experiment to discover the answer yourself? Social learning refers to the ability of humans to learn by observing, imitating, or interacting with others. In this work, we use multi-agent reinforcement learning (MARL) to investigate the conditions under which social learning achieves higher returns and improved learning efficiency relative to individual learning. We begin with a theoretical analysis and discover that, under coordinated group exploration and low skill transmission costs, social learning achieves lower sample complexity than individual learning. Motivated by this insight, we propose Selective Social Learning (SSL), a novel social learning algorithm that reduces the cost of skill transmission, and we design experiments across three MARL environments to test these predictions. Empirical results show that SSL outperforms baselines under conditions consistent with our theory. Surprisingly, we also find that in non-stationary environments with sparse rewards, which mirror the real world, social learning can emerge in standard MARL without an explicitly designed social learning mechanism. This suggests that the apparent absence of social learning in standard RL stems from limitations in training environment design rather than in RL itself. We therefore call for future MARL training environments to inherently incorporate sparsity and non-stationarity, enabling agents to naturally develop social-learning behaviors that transfer to complex real-world settings.

We present a diagnostic study of reward guided reasoning in discrete diffusion language models (dLLMs). On in-distribution GSM8K with Dream-v0-Instruct-7B, deterministic segmental top-$1$ guidance using outcome-supervised intermediate-state reward models underperforms a simpler matched-compute baseline: independent best-of-$N$ sampling plus a GSM8K-trained final-answer outcome reward model (ORM) reranker. The setup appears well suited for PRM guidance because dLLM intermediate states partially expose solution evidence; however, these states are unordered masked subsets rather than autoregressive prefixes, and prior reward-guided decoding results rarely report matched forward-pass compute against simple verifier reranking. We study \emph{outcome-supervised} PRMs trained with binary final-correctness labels on dLLM intermediate states. Under a protocol that charges denoising, PRM scoring, and ORM scoring in the same forward-pass unit, deterministic PRM-guided search trails ORM Rerank by $9.95$ and $12.69$ percentage points (pp) in seed-mean aggregate at matched candidate budgets $K{=}8$ and $K{=}32$; paired bootstrap CIs on seed $42$ exclude zero. Two mechanisms account for much of the headline gap: bidirectional PRM ROC-AUC (area under the receiver operating characteristic curve) decays from $0.77$ to $0.54$ with mask ratio, and deterministic top-$1$ pruning drops the guided pool's perfect-selector ceiling by $\sim$$14$\,pp. A third diagnostic shows that mean pooling also weakens causal PRM variants. Matched-compute ORM Rerank thus emerges as the strong baseline when an in-distribution final-answer verifier is available; closing the gap requires PRM guidance that preserves candidate diversity and queries the scorer at denoising stages where it remains discriminative, neither of which the deterministic top-$1$ recipe satisfies.


Why Heavy-Tailed Weights Predict Model Quality

Joseph Wilson ⋅ Chris van der Heide ⋅ Liam Hodgkinson ⋅ Zhichao Wang ⋅ Fred Roosta ⋅ Michael Mahoney

A solid understanding of the predictive success of deep neural networks (DNNs) remains elusive. For example, while many data-dependent metrics exist that seek to predict model quality of DNNs, these metrics may be expensive to compute for large datasets and/or may be impossible to compute when training datasets are not released. A practical and effective predictor of model quality involves analyzing DNN weight matrices, as it is known that a heavy-tailed spectral distribution often correlates with strong model quality. As such metrics lack a rigorous statistical derivation, in this work we aim to understand why they perform well, by employing a probabilistic framework to derive the marginal likelihood for the trained weights of a DNN. Across large-scale convolutional DNNs and LLMs, we find that this marginal likelihood predicts model quality. To understand the role heavy-tailed spectra play in predictive performance, we derive the limiting log-marginal likelihood for spectra that follow the heavy-tailed High-Temperature Marchenko-Pastur (HTMP) distribution. We show that this limiting quantity is convex in a heaviness parameter, and we derive an optimal heaviness that empirically predicts over-fitting for large-scale DNNs during training.


Winning Lottery Tickets in Neural Networks via a Quantum-Inspired Classical Algorithm

Natsuto Isogai ⋅ Hayata Yamasaki ⋅ Sho Sonoda ⋅ Mio Murao

Quantum machine learning (QML) aims to accelerate machine learning tasks by exploiting quantum computation. Previous work studied a QML algorithm for selecting sparse subnetworks from large shallow neural networks. Instead of directly solving an optimization problem over a large-scale network, this algorithm constructs a sparse subnetwork by sampling hidden nodes from an optimized probability distribution defined using the ridgelet transform. The quantum algorithm performs this sampling in time $O(D)$ in the data dimension $D$, whereas a naive classical implementation relies on handling exponentially many candidate nodes and hence takes $\exp[O(D)]$ time. In this work, we construct and analyze a quantum-inspired fully classical algorithm for the same sampling task. We show that our algorithm runs in time $O(\operatorname{poly}(D))$, thereby removing the exponential dependence on $D$ from the previous classical approach. Numerical simulations show that the proposed sampler achieves empirical risk comparable to exact sampling from the optimized distribution and substantially lower than sampling from the non-optimized uniform distribution, while also exhibiting exponentially improved runtime scaling compared with the conventional classical implementation. These successful dequantization results show that sparse subnetwork selection via optimized sampling can be achieved classically with polynomial data-dimension scaling on conventional computers without quantum hardware, providing an alternative to the existing quantum algorithm.


Words Before Pixels: Selective Modality Routing for Vision-Language Model Unlearning

Laura Yao ⋅ Haochen Zhang ⋅ Jinhao Duan ⋅ Sijia Liu ⋅ Tianlong Chen

Vision-language model (VLM) unlearning is often treated as an objective-design problem: given multimodal forget data, the goal is to optimize a loss that removes undesirable behavior while preserving retained capabilities. We argue that this view overlooks an equally important data-centric question: how should each forget example be represented before optimization? Existing methods typically keep image-conditioned failures in their original image-text form, implicitly assuming that pixels are the right route for unlearning whenever pixels appear in the input. We challenge this assumption by constructing modality-decomposed versions of VLM benchmarks, enabling controlled comparisons among image-only, text-only, and multimodal representations of the same examples. Our modality-mixing experiments show that, for many examples, textual renderings carry the actionable forget signal more directly than the image alone, while purely text-only unlearning can still weaken visual grounding when the target behavior depends on visual evidence. Motivated by this trade-off, we introduce ModRoute, a gradient-guided per-sample modality routing strategy that selectively converts high visual-pressure examples to text-only form while keeping the remaining examples multimodal. Notably, on VLGuard with LLaVA-1.5-7B, ModRoute has the strongest composite unlearning--utility score of $0.847$ at $\gamma=0.5$, a $2.5$% improvement over the best random-switching score and $\geq46$% better than pure multimodal or text unlearning. Overall, our results show that effective VLM unlearning depends not only on the objective, but also on the composition and per-sample modality routing of the unlearning data.


WorldForge: Forging Unified World Modeling into Video Generation

Boming Tan ⋅ Xiangdong Zhang ⋅ Ning Liao ⋅ Jingtao Zhang ⋅ 张雨晴 ⋅ Xiaoqiu Zhong ⋅ Xiaosong Jia ⋅ Xue Yang ⋅ Shaofeng Zhang ⋅ Yanyong Zhang

Despite impressive progress in video generation, existing models remain limited to surface-level plausibility, lacking a coherent and unified understanding of the world. Prior approaches typically incorporate only a single form of world-related knowledge or rely on rigid alignment strategies to introduce additional knowledge. However, aligning the single world knowledge is insufficient to constitute a world model that requires jointly modeling multiple heterogeneous dimensions (e.g., physical commonsense, 3D and temporal consistency). To address this limitation, we introduce \textbf{WorldForge}, a unified framework that integrates complementary world knowledge into video generators via a \textbf{Joint World Modeling Paradigm}, jointly predicting video pixels and features from foundation models to capture temporal dynamics, spatial geometry, and semantic consistency. However, naively optimizing these heterogeneous objectives can lead to visual instability and temporal flickering. To mitigate this issue, we propose \textit{Consistent Constraint Annealing (CCA)} to progressively regulate world-level constraints during training, and \textit{Multi-Source Inner-Guidance} to enforce learned world priors at inference. Extensive evaluations show that WorldForge improves world consistency, outperforming Wan2.1 by 2.26 points on VBench.


World–Value–Action Model: Implicit Planning for Vision–Language–Action Systems

Runze Li ⋅ Hongyin Zhang ⋅ Junxi Jin ⋅ Qixin Zeng ⋅ Zifeng Zhuang ⋅ Yiqi Tang ⋅ Shangke Lyu ⋅ Donglin Wang

Vision-Language-Action (VLA) models have emerged as a promising paradigm for building embodied agents that ground perception and language into action. However, most existing approaches rely on direct action prediction, lacking the ability to reason over long-horizon trajectories and evaluate their consequences, which limits performance in complex decision-making tasks. In this work, we introduce $\textbf{World-Value-Action (WAV)} $ model, a unified framework that enables implicit planning in VLA systems. Rather than performing explicit trajectory optimization, WAV model learn a structured latent representation of future trajectories conditioned on visual observations and language instructions. A learned world model predicts future states, while a trajectory value function evaluates their long-horizon utility. Action generation is then formulated as inference in this latent space, where the model progressively concentrates probability mass on high-value and dynamically feasible trajectories. We provide a theoretical perspective showing that planning directly in action space suffers from an exponential decay in the probability of feasible trajectories as the horizon increases. In contrast, latent-space inference reshapes the search distribution toward feasible regions, enabling efficient long-horizon decision making. Extensive simulations and real-world experiments demonstrate that the WAV model consistently outperforms state-of-the-art methods, achieving significant improvements in task success rate, generalization ability, and robustness, especially in long-horizon and compositional scenarios.


WorldVLN: Autoregressive World Action Model for Aerial Vision-Language Navigation

Baining Zhao ⋅ jiacheng xu ⋅ Weicheng Feng ⋅ Xin Zhang ⋅ Zhaolu Wang ⋅ Haoyang Wang ⋅ Shilong Ji ⋅ Ziyou Wang ⋅ Jianjie Fang ⋅ Zhiheng Zheng ⋅ Weichen Zhang ⋅ Yu Shang ⋅ Wei Wu ⋅ Chen Gao ⋅ Xinlei Chen ⋅ Yong Li

Aerial vision-language navigation requires agents to follow natural-language instructions through closed-loop perception and action in 3D environments. We argue that aerial VLN can be formulated as a prediction-driven world-action problem: the agent should anticipate latent world evolution and act according to the predicted consequences. To this end, we propose WorldVLN, the first autoregressive world action model for aerial VLN. Unlike full-sequence video-generation world models that generate an entire visual clip, WorldVLN adapts a latent autoregressive video backbone to predict short-horizon world-state transitions and directly decodes them into executable waypoint actions. After each action segment is executed, newly received observations are encoded back into the autoregressive context, enabling closed-loop world-action prediction. We further introduce a two-stage training framework that first grounds the video prior in instruction-conditioned navigation dynamics and then develops Action-aware GRPO, the first reinforcement learning method tailored to autoregressive WAMs, to optimize waypoint decisions through their downstream rollout consequences. On public outdoor and indoor benchmarks, WorldVLN consistently outperforms existing Vision-Language-Action baselines with $12\%+$ success-rate gains and larger advantages on challenging cases. It further transfers zero-shot to real drone deployment, suggesting that the proposed WorldVLN offers a promising route for spatial action tasks. Demos and code are available at \url{https://worldvln.com} (anonymous).


WsiSSM: A Weakly Supervised Subset-Matching Framework for Unified Classification and Segmentation of Histopathology Whole Slide Images

Guangjian Zeng ⋅ Siyuan Tao ⋅ Chengzhi Zhao ⋅ Xiaoqing Li ⋅ Na Tang ⋅ Wenting Huang ⋅ Shenying Fang

Multiple instance learning (MIL) has been increasingly used to analyze histopathology whole slide images (WSIs) and has achieved accurate diagnostic classification and a certain degree of automatic segmentation capability. However, segmentation performance without annotation remains suboptimal due to the immense size of WSIs, despite its essential role in assisting pathologists’ diagnosis. Moreover, existing approaches treat classification and segmentation as independent tasks, with segmentation primarily serving as an interpretability section. To address this limitation, we propose a novel weakly supervised learning algorithm that unifies classification and segmentation within a framework. Specifically, we explore the interconnections between patches and introduce class activation maps in WSI analysis through the proposed Sparse-CAM backbone to activate the segmentation capability, and further enhance feature representation using the proposed subset-matching module to improve both classification and segmentation performance. The proposed method demonstrated superior performance in diagnosis and achieved state-of-the-art performance in segmentation compared to other MIL methods across multiple datasets. In addition, it offered high-precision segmentation that is sensitive to micro-tumor regions without pixel-level annotations. The code will be publicly available on GitHub upon acceptance.


WURI: Watching Unfolding Risk in Agent Interactions

Yiyang Duan ⋅ Tiantong Wu ⋅ Yurong Hao ⋅ Wei Yang Bryan Lim

Large language model (LLM) agents increasingly operate through multi-step tool-use trajectories, where harmful intent may be distributed across actions that appear benign in isolation. This challenges single-step moderation, while repeated LLM-as-judge evaluation over growing prefixes can be costly and may intervene too late. We introduce WURI (Watching Unfolding Risk in Agent Interactions), a lightweight monitor for early prefix-level detection of harmful agent trajectories. WURI encodes each observed textual step with a frozen text encoder, learns an atom-adapted trajectory representation through metric learning, and scores each prefix with a fixed prototype-margin rule. It requires no access to agent model weights or hidden states, does not modify the agent's internal model, and uses no external LLM-as-judge calls at inference. Across five generalization settings, WURI achieves the best average prefix-area under the detection curve (AUDC), ranks first on three settings, and reaches the strict early-detection operating point within four steps in multiple evaluation settings. Runtime analysis shows a wall-clock speedup of over two orders of magnitude compared with repeated full-prefix guardrail evaluation. These results demonstrate that representation-based prefix scoring is an effective and efficient direction for monitoring unfolding risk in LLM agent interactions.

Production retrieval systems require three properties jointly: domain adaptation, preservation of general-retrieval quality, and compatibility with existing vector indices. Existing adaptation methods achieve at most two. Standard parameter-efficient fine-tuning (PEFT) adapts to a domain but breaks the index and degrades general retrieval; backward-compatible training (BCT) preserves the index approximately via distillation while sacrificing general retrieval under domain shift; frozen-backbone side networks (L2C) preserve general retrieval but still modify the deployed embedding via additive combining, so existing indices are not preserved. We introduce \textbf{CoSE} (Complementary Subspace Expansion), a side-branch LoRA whose frozen pathway is structurally untouched, so any index built from the base encoder remains valid --- the adapter is \emph{pluggable}, added on top of an already-deployed index without modifying existing entries. A joint training objective converts the frozen+LoRA concatenation into Pareto-efficient retrieval. Across four distinct domains (Korean broadcast, medical radiology, art paintings, remote sensing) and 3 seeds, CoSE at $6.4$M trainable parameters is the only configuration that simultaneously stays within $\pm 0.048$ of the frozen base on Flickr, COCO, and ImageNet at the deployment-default $\alpha{=}0.5$, is domain-competitive with standard LoRA (ties on RSICD, within $\sim 2\sigma$ on HAN and SemArt, $-0.05$ on ROCO), and is backward-compatible by construction. Deployment-time knobs --- $\alpha$ blending, Matryoshka truncation, partial branching --- expose further cost/quality trade-offs without retraining.


Your Teacher Can’t Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

Yanjiang Liu ⋅ Jie Lou ⋅ Xinyan Guan ⋅ Yuqiu Ji ⋅ Hongyu Lin ⋅ Ben He ⋅ Xianpei Han ⋅ Le Sun ⋅ XingYu Li ⋅ Yaojie Lu

On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, Supervision Fidelity Decay (SFD): as student-generated prefixes lengthen, the teacher’s next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce Lookahead Group Reward (LGR). Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, LGR evaluates the student’s top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, LGR improves mean@8 by 2.57 points over OPD for a 7B student, with gains increasing in longer-generation and reaching 4.92 points on AIME-26 at 39k tokens.


ZeoBench: A Benchmark for Self-Supervised Learning on 3D Zeolite Representations

Aaron Sun ⋅ Yachan Liu ⋅ Ping Yang ⋅ Peng Bai ⋅ Subhransu Maji

Zeolites are an important class of porous crystalline materials with broad applications in gas storage, separations, and catalysis. Although hundreds of thousands of experimentally realized and hypothetical zeolite structures are known, identifying materials optimized for a target application remains challenging because property labels are expensive to obtain and available only for limited structure–molecule pairs. While prior work has demonstrated that 3D neural representations can predict zeolite properties effectively, it remains unclear which representation-learning strategies are most effective in the low-data regime relevant to materials discovery. We present ZeoBench, a benchmark for self-supervised learning on dense 3D representations of all-silica zeolites, and evaluate the resulting features on six downstream adsorption-property prediction tasks, across a range of training set sizes. Our study compares self-supervised 3D pretraining on volumetric zeolite data, adaptations of image-pretrained 2D ConvNets and vision transformers to 3D via multiview and channel-inflation strategies, and 3D ConvNets pretrained on out-of-domain volumetric datasets, alongside hand-crafted descriptors and supervised baselines. Using our pretraining and evaluation framework, we find self-supervised 3D representations outperform existing methods and achieve state-of-the-art performance across all label regimes and target molecules. Learned representations consistently outperform hand-crafted descriptors, and convolutional architectures are generally more effective than vision transformers in this setting. These results establish ZeoBench as a practical benchmark for 3D representation learning in zeolites and provide guidance for data-limited adsorption property prediction.


Zero-Shot Coordination among LLM Agents

Adrian Hayler ⋅ Shashank Reddy Chirra ⋅ Andrei Lupu ⋅ Johannes Forkel ⋅ Bidipta Sarkar ⋅ Siheng Feng ⋅ Jakob Foerster

We study zero-shot coordination (ZSC), where independently developed agents must coordinate at test time. While ZSC has been well studied in the RL literature, far less is known about the performance of LLM agents despite their increasing deployment in such settings. Existing work on LLM coordination often relies on specialised scaffolding that independently developed agents are unlikely to share in practice. Moreover, evaluations on complex environments (e.g., Hanabi) make it difficult to pinpoint the sources of coordination failure. In contrast, we focus on simple, general-purpose scaffolds in minimal environments designed to isolate specific coordination challenges. Our results show that even in these controlled settings, frontier LLM agents struggle to coordinate, largely due to limited understanding of the coordination problem and weak reasoning about their partner’s beliefs. While LLMs introduce semantic information as an additional axis for coordination, they nonetheless fail to exploit this structure effectively. Towards this, we propose Coordination-friendly definitions (CFDs) as a principled approach for enabling robust coordination among LLM agents. Finally, we show that CFDs can be discovered automatically, removing the need for manual engineering.


Zero-Shot Instruction Following in RL via Structured LTL Representations

Mathias Jackermeier ⋅ Mattia Giuri ⋅ Jacques Cloete ⋅ Alessandro Abate

We study instruction following in multi-task reinforcement learning, where an agent must zero-shot execute novel tasks not seen during training. In this setting, linear temporal logic (LTL) has been adopted as a powerful framework for specifying structured, temporally extended tasks. While existing approaches successfully train generalist policies, they often struggle to effectively capture the rich logical and temporal structure inherent in LTL specifications. In this work, we address these concerns with a novel approach to learn structured task representations that facilitate training and generalisation. Our method conditions the policy on sequences of Boolean formulae constructed from a finite automaton of the task. We propose a hierarchical neural architecture to encode the logical structure of these formulae, and introduce an attention mechanism that enables the policy to reason about future subgoals. Experiments in a variety of complex domains demonstrate the strong generalisation capabilities and superior performance of our approach.