Session
Sydney Poster Session 5
Hall 1-4
Accelerating Video Inverse Problem Solvers with Autoregressive Diffusion Models
Taesung Kwon ⋅ Jonghyun Park ⋅ Hyungjin Chung ⋅ Jong Chul Ye
Diffusion models provide powerful priors for zero-shot video inverse problems, but their real-time deployment is hindered by two inefficiencies: high initial latency caused by holistic video restoration, and low throughput resulting from multiple VAE passes to enforce measurement consistency in pixel space. To overcome these limitations, we propose Autoregressive Video Inverse problem Solver (AVIS). The AVIS framework leverages autoregressive video diffusion models to restore videos in a streaming manner, naturally eliminating latency bottlenecks. Specifically, AVIS initializes reverse diffusion with a measurement-consistent estimate, reducing the required sampling steps. Compared to leading non-autoregressive solvers, AVIS drastically reduces initial latency from 114s to 4s and increases throughput from 0.71 to 1.18 FPS while achieving superior restoration quality. We further introduce a highly accelerated variant, dubbed AVIS Flash, that enforces measurement consistency solely on the first chunk. AVIS Flash substantially boosts throughput to 5.91 FPS on a single RTX 4090 GPU while maintaining competitive performance and achieving a favorable efficiency–performance trade-off, paving the way toward real-time deployment.
Achieving Directional-Stationarity from a Single Random Direction Step
Dan Greenstein ⋅ Nadav Hallak
This paper addresses the challenge of obtaining strong optimality guarantees in constrained nonsmooth nonconvex optimization under mild regularity conditions, namely local Lipschitz continuity and existence and continuity of directional derivatives. While standard methods typically ensure weak stationarity notions, achieving directional (d-)stationarity remains nontrivial. We show that a random direction exploration step is sufficient to attain d-stationarity. The proposed approach augments any base optimization method with a single exploration step that samples a direction and step size and accepts the candidate based on a function value comparison. The resulting scheme guarantees that all accumulation points are d-stationary almost surely, independently of the behavior of the underlying method. Moreover, it preserves convergence rates of the base method, as established for DCA and prox-linear-type schemes. The theoretical results are complemented by numerical experiments illustrating the effect and guarantees of the exploration step.
A Closer Look at Dynamic Scene Graph Generation in the Era of Multimodal Large Language Models
Xuanming Cui ⋅ Jaiminkumar Ashokbhai Bhoi ⋅ Chionh W Peng ⋅ Adriel Kuek ⋅ Ser Nam Lim
Dynamic Scene Graph Generation (DSGG) aims to capture objects and their evolving relations in videos. Despite recent progress, the practicality and quality of generated scene graphs remain limited compared to the rapid advances in Multimodal Large Language Models (MLLMs). In this work, we revisit DSGG from two fundamental perspectives: task setup and model design. From the task setup perspective, we identify two key limitations of the current recall-oriented evaluation protocol: (i) a severe precision–recall trade-off, and (ii) uninformative and redundant relation generation. To better assess practical usefulness, we introduce four additional metrics that provide a more comprehensive evaluation of the quality of generated dynamic scene graphs. From the model design perspective, we explore directly using MLLMs for scene graph generation and establish a strong MLLM-based DSGG baseline through three design changes. First, we replace the conventional bottom-up pipeline with a top-down reason-then-locate strategy. Second, we reformulate frame-wise dynamic graphs as Temporal Relation Set (TRS) prediction, improving both efficiency and performance. Third, we introduce Importance-Aware Finetuning (IAF) to encourage more relevant and diverse relation generation. Extensive experiments on Action Genome, VidVRD, and PVSG show that our approach consistently achieves state-of-the-art performance across multiple metrics.
A Composite Activation Function for Learning Stable Binary Representations
Seokhun Park ⋅ Choeun Kim ⋅ Kwanho Lee ⋅ Sehyun Park ⋅ Insung Kong ⋅ Yongdai Kim
Activation functions play a central role in neural networks by shaping internal representations. Recently, learning binary activation representations has attracted significant attention due to their advantages in computational and memory efficiency, as well as interpretability. However, training neural networks with Heaviside activations remains challenging, as their non-differentiability obstructs standard gradient-based optimization. In this paper, we propose Heavy-Tailed Activation Function (HTAF), a smooth approximation to the Heaviside function that enables stable training with gradient-based optimization. We construct HTAF as a sigmoid–hyperbolic tangent composite function and theoretically show that it maintains a large gradient mass around zero inputs while exhibiting slower gradient decay in the tail regions. We show that Spiking Neural Networks, Binary Neural Networks and Deep Heaviside neural Networks can be trained stably using HTAF with gradient-based optimization. Finally, we introduce Implicit Concept Bottleneck Models (ICBM), an interpretable image model that leverages HTAF to induce discrete feature representations. Extensive experiments across various architectures and image datasets demonstrate that ICBM enables stable discretization while achieving prediction performance comparable to or better than standard models.
A Computational Perspective to Data Ablation Experiments
Jiachen (Tianhao) Wang ⋅ Lin Chen ⋅ Mohammadhossein Bateni ⋅ Ruoxi Jia ⋅ Prateek Mittal ⋅ Vahab Mirrokni
\emph{Scaling} has been the central driver of progress in modern AI, yet data curation remains an exception: it still relies on manual heuristics rather than systematic, scalable methods. In practice, data recipes (e.g., filter thresholds, deduplication strength, domain mixing ratios) are determined through many small-scale ablation training runs. This work demonstrates that the performance of the best-discovered recipe improves \emph{predictably}--as a \emph{power law}--in the number of such ablation runs. Theoretically, the power-law form arises naturally from regret bounds in zero-order optimization theory. Empirically, we validate it across diverse training stages, including pretraining, supervised fine-tuning, preference optimization, and reinforcement learning. Building on this finding, we derive a principled approach to \emph{data quality--quantity tradeoff} under a fixed compute budget, yielding a Pareto frontier that dominates standard heuristic pipelines (e.g., DCLM). Together, this work provides a general framework for scaling data-curation compute, establishing it as a new axis alongside the well-studied pretraining and test-time scaling.
Action-On-Item Preference Flow: A Shared Event Schema for Predictive and Generative Personalization
Parthiv Chatterjee ⋅ Kashish Kanjariya ⋅ Vashisht Purani ⋅ Sourish Dasgupta ⋅ Tanmoy Chakraborty
Personalization histories vary across domains because surface actions and objectives differ: users watch or skip movies, click or ignore news, or request recommendation, summarization, question answering, or response generation. Rather than build a separate history encoder for each action space, we rewrite native interactions into a shared *action-on-item event schema*: a specific action on a specific item becomes positive, negative, or command-like evidence over a target-locus embedding whose geometry induces item-kind abstraction. Task requests become generalized commands over the active item-side abstraction. The modeling problem is how this evidence enters a user-specific state, persists, and is read out by the active command. We formulate the Multi-Timescale State Hypothesis (MTSH), where signed and command-like evidence flows through long-term stable interests, recency-sensitive short-term interests, and bursty episodic traces. We instantiate MTSH with $\texttt{PerTIDE}$, an action-conditioned multi-timescale encoder that learns a user-specific preference-history state and uses command-conditioned readout to produce a task-ready state. We evaluate $\texttt{PerTIDE}$ on MovieLens, MIND, and PENS for prediction, and PENS/OpenAI-Reddit for generation. $\texttt{PerTIDE}$ improves over the strongest baselines without explicit trace factorization by +2.36/+3.47, +10.05/+9.01, and +0.90/+1.42 MRR/nDCG on MovieLens, MIND, and PENS; and over the strongest two-shot large language model by +0.10/+0.18 and +0.04/+0.19 PerSEval-JSD/PerSEval-METEOR on PENS and OpenAI-Reddit. Removing action conditioning and command-conditioned readout degrades MIND by 6.52 MRR / 9.13 HR@10 and 4.14 MRR / 4.42 HR@10. Trace diagnostics show L/S/E specialization, while the full $L+S+E$ model is strongest under mixed stable, recency-sensitive, and episodic command evidence. These results support action-conditioned multi-timescale preference flow as a reusable user-history encoding principle.
Active Flow Expansion for Out-of-Distribution Discovery: from Theory to Molecules
Riccardo De Santi ⋅ Bruce D Lee ⋅ Cristian Jensen ⋅ Kimon Protopapas ⋅ Sophia Tang ⋅ Chenghao Liu ⋅ Pranam Chatterjee ⋅ Yisong Yue ⋅ Andreas Krause
Standard flow and diffusion pre-training matches the distribution of available data (e.g., molecules), which often covers only a small fraction of the valid design space. In generative discovery, however, one aims to sample valid new-to-nature designs, assigned negligible probability under, and thus inaccessible to, standard models fitted to the observed data. To overcome this limitation, we depart from data distribution matching and view a generative model through its generable set: the region it covers with non-negligible probability. This allows to introduce a new learning principle for out-of-distribution flow modeling: enlarging a model’s generable set to increase coverage of the valid design space. We propose Active Flow Expansion (ActFlow), a continued pre-training method that employs verifier feedback to expand a pre-trained model over new valid regions by iteratively adapting to synthetic data generated through active exploration in the learned flow representation. Theoretically, we establish to our knowledge first-of-their-kind statistical learning guarantees for out-of-distribution flow modeling, analyzing generable set expansion as a local-to-global reachability process over a learned representation. Empirically, we assess ActFlow with suitable out-of-distribution generative modeling metrics across small organic molecules, mid-sized drug-like molecules, therapeutic peptides, and protein sequence design tasks. Results show that ActFlow expands valid coverage far beyond the region modeled by the initial pre-trained model, significantly outperforming widely adopted synthetic flow pre-training methods.
AdaM-Rec: Adaptive Modality Routing for Multimodal Recommendation
Honghao Fu ⋅ Jiacheng Chen ⋅ Manxi Lin ⋅ Junjun Zheng ⋅ Xiangheng Kong ⋅ Yiwei Wang ⋅ Xin Yu ⋅ Miao Xu ⋅ Yuning Jiang ⋅ Yujun Cai
While recent multimodal recommender systems have demonstrated the effectiveness of incorporating visual and textual information to improve downstream performance, most existing methods rely on static modality fusion, assuming that the relative importance of textual and visual signals remains stable across recommendation scenarios. This design may not fully account for an important variation across recommendation requests: some queries require fine-grained visual cues, whereas others are better served by textual or functional semantics, in which case indiscriminate modality fusion brings in uninformative cues and impairs recommendation quality. To address this, we propose AdaM-Rec, an LLM-based framework for adaptive modality routing in multimodal recommendation, which enables dynamic calibration of reliance on textual and multimodal evidence for user-specific queries. Built on structured natural-language representations of items and user preferences, it estimates modality reliability using proxy recall tasks. Specifically, it generates pseudo-queries that match the granularity of the actual query while pointing to the user's positively interacted items as verifiable proxy targets, evaluating which modality yields better recall performance in analogous scenarios and optimizing the routing strategy in an agentic manner. It then performs routed recall with optimized strategy, enriches results with collaborative items, and ranks candidates by their relevance to both the query and user preferences. Experiments demonstrate that AdaM-Rec delivers strong performance against state-of-the-art baselines, highlighting the effectiveness and broader potential of adaptive control over modality reliance in multimodal recommendation. Code will be publicly available.
Adaptive auditing of AI systems with anytime-valid guarantees
Siyu Zhou ⋅ Patrick Vossler ⋅ Venkatesh Sivaraman ⋅ Yifan Mai ⋅ Jean Feng
A major bottleneck in characterizing the failure modes of generative AI systems is the cost and time of annotation and evaluation. Consequently, adaptive testing paradigms have gained popularity, where one opportunistically decides which cases and how many to annotate based on past results. While this framework is highly practical, its extreme flexibility makes it difficult to draw statistically rigorous conclusions, as it violates classical assumptions: the number of observations is typically limited (often 10 to 50 cases) and decisions regarding sampling and stopping are made in the midst of data collection rather than based a pre-specified rule. To characterize what statistical inferences can be drawn from highly adaptive audits, we introduce a hypothesis testing framework from two 'dueling' perspectives: (i) the model's null that asserts there is no failure mode with performance below a target threshold versus (ii) the auditor's null that asserts they have a sampling strategy that will uncover a failure mode. Leveraging Safe Anytime-Valid Inference (SAVI), we formalize the auditor as conducting 'testing by betting', which translates into simultaneous e-processes for testing the dueling null hypotheses. Furthermore, if the auditor is sufficiently powerful, we prove that these two hypotheses are asymptotically inverses of each other, in that passage of a stringent audit does in fact certify the AI system as being globally robust. Empirically, we demonstrate that our proposed testing procedures maintain anytime-valid type-I error control, outperform pre-specified testing methods, and can reach statistically rigorous conclusions sometimes with as few as 20 observations.
Large Reasoning Models (LRMs) produce explicit chains of thought to improve transparency, but they frequently overthink, generating long yet inefficient reasoning chains that inflate computational costs and risk introducing errors. While numerous reinforcement learning methods struggle to penalize verbosity to achieve conciseness, they typically apply the penalty uniformly across lengthy sequences, which leads to \textit{credit confounding}, that is, the failure to distinguish essential reasoning steps from superfluous ones. In this paper, we study a measurable proxy for this confounding: high-entropy tokens that often coincide with surface transition points where reasoning re-enters verification. Based on this insight, we propose Adaptive Entropy-Sparing (AES), which selectively spares high-entropy tokens that occur before an efficient in-group reference, while penalizing overlength high-entropy suffix tokens associated with superfluous branching. Our AES thereby navigates the suppression-exploration trade-off, reducing overthinking while preserving average reasoning accuracy. We also present a stylized entropy-length model that motivates why local uncertainty can serve as a proxy for future reasoning cost. Across eight reasoning benchmarks, AES reduces reasoning length by up to 52.9\% while increasing mean accuracy by up to 3.7\% over the base models, consistently outperforming existing efficiency-oriented methods.
Adaptive Residual Quantization for Memory-Efficient Temporal Action Segmentation
Gerard L Donahue ⋅ Guven Gergerli ⋅ Ayush Gupta ⋅ Reza Ghoddoosian ⋅ Enna Sachdeva ⋅ Faizan Siddiqui
Temporal action segmentation (TAS) is commonly trained on large banks of pre-extracted frame features rather than raw RGB video, but storing and replaying those features becomes a practical bottleneck for long-form video, especially in incremental TAS where past-task data must remain available to mitigate catastrophic forgetting. Recent TAS replay and condensation methods reduce this burden, but they rely on dense frame-wise annotations and allocate storage through fixed segment-level structures, limiting their flexibility for long procedural videos whose frame complexity varies substantially over time. We introduce Adaptive Residual-Quantized Variational AutoEncoder (ARQ-VAE), a compact and label-free memory representation for TAS features that learns a discrete residual quantization of frame-level video features and then adaptively truncates the residual chain on a per-frame basis, retaining only the residual prefix that best reconstructs each frame feature. This yields a practical compact feature memory that does not depend on frame-wise labels and can therefore support multiple TAS settings with the same stored representation, while a lightweight fine-tuning stage further improves reconstruction quality for the truncated codes that are actually retained. Across standard TAS benchmarks, ARQ-VAE achieves a substantially improved storage--performance trade-off over prior replay and condensation baselines. On Breakfast, our final representation reduces training-set storage from 28 GB for the original I3D feature bank to 10 MB while retaining strong downstream segmentation performance; with ASFormer, it achieves 71.5 Edit and 51.2 F1@50 in the supervised setting. We further show that the same label-free memory remains effective for incremental and unsupervised TAS, and transfers to an alternative pretrained feature space based on DINOv3. We position ARQ-VAE as a practical memory representation for TAS, and a promising direction for broader dataset condensation applications.
AdaSRU: Adaptive Source-Free Recommendation Unlearning via Gradient-Constrained Optimization
Xiaoran Zhao ⋅ Weiming Liu ⋅ Lianyong Qi ⋅ Weiyi Zhong ⋅ Yuwen Liu ⋅ Xiaolong Xu ⋅ Haolong Xiang ⋅ Qiang Ni ⋅ Xuyun Zhang ⋅ Wanchun Dou ⋅ Xiaokang Zhou
Recommendation unlearning (RU) aims to remove the influence of specified user--item interactions from a trained recommender while preserving utility on the remaining data. Existing RU methods usually require access to the original training data, especially the remaining data, to recover recommendation utility after unlearning. However, access to such data is often restricted or unavailable in practice due to privacy constraints, storage costs, or third-party deployment, giving rise to source-free recommendation unlearning. To address this challenge, we propose an adaptive source-free recommendation unlearning (AdaSRU) method that enables effective unlearning without accessing the original remaining data. AdaSRU first estimates the inaccessible gradient on the remaining data. It exploits the first-order optimality condition of the trained recommender and approximates the remaining gradient using the forgetting data and squared gradient accumulators from adaptive optimizers, e.g., Adam. AdaSRU then formulates unlearning as a gradient-constrained update, where the estimated remaining gradient is used as a constraint to limit utility degradation while maximizing the unlearning objective. AdaSRU adaptively adjusts the utility-preserving correction according to the current gradient geometry. We further theoretically prove that AdaSRU approaches an approximate Pareto stationary solution. Extensive experiments on three real-world datasets demonstrate that AdaSRU achieves strong performance in both unlearning and recommendation. The code is available at https://anonymous.4open.science/r/AdaSRU-94D9/.
Ad-Hoc Teamwork from Human Demonstrations
Darius Muglich ⋅ Niklas Lauffer ⋅ Tin Dizdarević ⋅ Jakob Foerster
Human players in cooperative games often rely on conventions: reusable patterns of play that coordinate well when used by the whole team, but can fail otherwise. Evaluating agents in this setting is difficult because human demonstration datasets are usually unlabeled: trajectories do not identify which player, group, or convention produced them. As a result, standard proxy-based evaluations can miss important population diversity. In particular, proxy bots produced by regularised self-play around a pooled behavioural cloning policy can remain concentrated around a single dominant mode of play. We introduce a pipeline for recovering and evaluating this hidden convention diversity from unlabeled demonstrations. Our method first learns latent convention structure from the dataset, then selects a compact set of distinct but plausible conventions, and finally reconstructs each selected convention into a high-performing playable proxy using KL-regularised PPO. A key practical challenge is that KL-regularised PPO is sensitive to the KL penalty, which is often tuned using online interaction. We provide an empirical offline calibration procedure for choosing this KL penalty in the PPO+KL reconstruction settings studied here. On Hanabi, we find that the BC-derived proxy suite is highly concentrated, and that the relative ordering of agents changes when evaluation uses the latent proxy suite instead.
Adversarial Group Fairness in Contextual Bandits: When Robust is Not Fair
Ray Telikani ⋅ Jaber valizadeh ⋅ Amir H Gandomi ⋅ Bao Q Vo ⋅ Ming Ding
Neural contextual bandits support high-stakes decision systems---from clinical trial allocation to content recommendation. This paper investigates an adversarial threat, group-targeted suppression attacks (GTSA), in which an adversary selectively degrades performance for a demographic subgroup while preserving global performance metrics. We first show that under standard aggregate monitoring, GTSA with budget $B = \Omega(\sqrt{T})$ can induce $\Omega(T)$ group regret while maintaining only $o(1)$ deviation in observable global statistics, establishing that any defense that ignores group structure is fundamentally vulnerable. To address this vulnerability, we propose AnchorFair, a robust mitigation framework comprising three components: (i) a trust-weighted policy anchor with provable bounded drift, (ii) representation-stable demographic discovery via online clustering in a slowly evolving embedding space, and (iii) a demographic-aware exploration mechanism that adaptively amplifies learning for underrepresented or low-trust groups. We prove that \textsc{AnchorFair} achieves sublinear group regret $\tilde{O}(\sqrt{T})$ for all groups under GTSA.
Adversary-Robust Learning from Fully Asynchronous Directional Derivative Estimates
Anik Kumar Paul ⋅ Nibedita Roy ⋅ Nagesh Talagani ⋅ Swetha Ganesh ⋅ Gugan Chandrashekhar Mallika Thoppe ⋅ Alexandre Reiffers-Masson
We propose FAR-SIGN (Fully Asynchronous Robust optimization via SIGNed directional projections) for adversary-resilient learning in parameter-server--worker systems. FAR-SIGN achieves robustness through sign-based updates along carefully designed directions and mitigates the resulting bias via a two-timescale mechanism. It admits both first-order and zeroth-order implementations and enables fully asynchronous execution without requiring a private reference dataset at the server. We establish almost-sure convergence of FAR-SIGN to the set of stationary points for smooth, nonconvex objectives. Moreover, we prove the near-optimal rate of $O(n^{-1/4+\epsilon})$ in the first-order setting and the standard $O(n^{-1/6+\epsilon})$ in the zeroth-order setting, where $n$ is the iteration count and $\epsilon>0$ can be chosen arbitrarily small. Experiments on MNIST show that FAR-SIGN outperforms robust aggregation-based methods in both accuracy and wall-clock time.
AdvJudge-Zero: Binary Decision Flips in LLM-as-a-Judge via Adversarial Control Tokens
Tung-Ling Li ⋅ Yuhao Wu ⋅ Hongliang Liu
LLM-as-a-Judge systems supply the reward signal in modern RLHF and RLVR pipelines, but their binary verdict reduces to a single linear readout on one hidden state. We show this readout is shallow enough that short, low-perplexity tokens flip the verdict from "No" to "Yes". These tokens are sampled from the judge's own next-token distribution at the response position, with no manual seed set and no gradient-based optimization. Our procedure, AdvJudge-Zero, reaches $>$90\% ensemble false-positive rate on 22 of 24 (model, dataset) cells across six Qwen, Llama, and Gemma judges, versus 54-72\% for the prior curated 10-token benchmark, and the discovered surface transfers cross-format to a 70B scalar reward model. The same discovered pool enables a defense: a LoRA fine-tune stratified by a 9-class mechanism taxonomy hardens cross-family generalization where naive sampling on the same pool fails, with mechanism breadth rather than pool size carrying the gain. Under GRPO training, the hardened judge eliminates the reward-collapse failures (false-positive spikes and length collapse) we observe in the unhardened baseline on both MATH and GSM8K at ten seeds per condition. The discovered pool, the mechanism taxonomy, and per-prompt flip records will be released under responsible disclosure.
AET-Bench: Pixel-Accurate Atomic Tomography Is Not Atom-Accurate
Zicheng Liu ⋅ Wenzhuo Ma ⋅ Jintao Chen ⋅ Di Huang
Atomic electron tomography (AET) aims to recover the three-dimensional arrangement, chemical identity, and temporal evolution of individual atoms. Yet most reconstruction methods are still evaluated primarily by voxel-level image similarity, such as PSNR, SSIM, or Fourier shell correlation. In this work, we show that this evaluation practice can select scientifically incorrect reconstructions: a volume that looks accurate can contain atoms that are shifted, merged, missed, hallucinated, chemically mislabeled, or assigned inconsistent identities through time. We introduce AET-Bench, a dataset and benchmark for atom-first evaluation of atomic electron tomography. AET-Bench contains 755 samples across a six-tier realism ladder, spanning clean analytical simulations, binary alloys, 4D atomic dynamics, multislice scattering, real experimental tilt series, and Brownian Pt nanocrystal trajectories. It defines four complementary tasks and evaluates 14 core method families with 21 logged method entries under a unified protocol. AET-Bench reveals a central failure mode hidden by voxel-only evaluation. On the AET-Sim-Mono tier, high-SSIM neural and Gaussian reconstructions can recover substantially fewer atoms than classical or sparsity-regularized methods: for example, 4D-GS achieves strong voxel similarity but only 43.9\% atom recall, while FISTA-TV reaches 95.1\% atom recall despite much lower SSIM. This inversion persists across atom-matching thresholds, showing that it is not an artifact of a single cutoff. On more realistic multislice and real tiers, atom recovery further degrades, exposing a physics gap between visually plausible volumes and reference atomic models. For 4D reconstruction, we find that smooth temporal density fields do not guarantee persistent atom identities. AET-Bench establishes atom-level fidelity as a necessary evaluation axis for AET. Our results suggest that future AET methods should not be judged only by how reconstructions look, but by whether they recover the atoms, species, and trajectories that scientists actually use.
A Favorable Regime Between ODE and SDE for Few-Step Diffusion Sampling
YIhao Bu ⋅ Dan Lian ⋅ Zhenguo Gao
Few-step diffusion sampling sharpens the tension between ODE and SDE dynamics. ODE samplers are stable but can lose diversity, while the standard SDE endpoint can become unreliable at low numbers of function evaluations (NFE). The challenge is therefore not simply whether stochasticity should be added, but which driver properties allow stochastic correction to contribute diffusion without losing finite-step control. We show that this tradeoff is governed by two driver-level quantities, the asymptotic covariance $Q$, which sets the effective diffusion level on the ODE-to-SDE spectrum, and the Poisson solution norm $C_\chi$, which controls the finite-step remainder. Gaussian drivers have $C_\chi = \infty$, so this finite-step remainder is not uniformly controlled by the same bound. Balancing the diffusion level set by $Q$ with the finite-step control set by $C_\chi$ formalizes a favorable regime for bounded structured drivers with intermediate diffusion strength. This regime directly gives a structured correction method, instantiated as Fast-Driver Sampler (FDS) with the shuffled-Lorenz driver. Across DiT-XL/2 and EDM2, FDS establishes a stronger low-NFE frontier among evaluated training-free samplers, improving ODE baselines while avoiding the degradation of the SDE endpoint.
A Foundational Model System for Datacenter Machine Repairs
Yuanlin Wen ⋅ Elan S Markowitz ⋅ Zubo Gu ⋅ Sami Abu-El-Haija ⋅ Bryan Perozzi ⋅ Michael Galkin ⋅ Ali Parviz ⋅ Aayush Gupta ⋅ Venkata Chivukula ⋅ Griffin Hadfield ⋅ Aravind Akella ⋅ Daniel J Frame ⋅ Vahab Mirrokni ⋅ Amin Vahdat
The rapid scale-up of AI infrastructure has made datacenter reliability a cornerstone of the cloud business economy. Hardware failures in hyper-scale AI clusters do not merely disrupt individual machines; they stall multi-million dollar training workloads, and so the downtime of individual machines can cause outsized economic losses. Minimizing machine downtime, therefore, is more important than ever. To address this, we present DrWatson, an AI-driven hardware diagnosis system that pioneers a foundation-model approach to automated datacenter maintenance. Deployed at production scale within a major AI hyperscaler, DrWatson yields substantial efficiency gains across diverse compute architectures. Real-world evaluations demonstrate that it reduces production downtime by 18.3% on GPU platforms and 16.1% on custom AI accelerators. Furthermore, it significantly accelerates hardware deployment, cutting quality assurance (QA) time by 28.3% for GPUs and 25.1% for custom accelerators. These results establish DrWatson as a field-tested solution for maximizing the reliability and cost efficiency of next generation AI fleets.
A Generative Model of Contextual Integrity: Appropriate vs. Inappropriate Information Sharing
Omer Ebead ⋅ Juan Formanek ⋅ Joel Leibo
Contextual integrity (CI) holds that appropriate information flow depends on context-specific norms. The same disclosure may be appropriate in one context and inappropriate in another. As actors handle private user data in multi-actor pipelines, we ask whether they track these norms faithfully under social and contextual pressure. We build a generative model of CI in multi-actor systems: scenarios drawn from CI's taxonomies, a generation pipeline, multi-actor simulation, and a two-sided appropriateness metric. The metric scores both withholding when norms forbid sharing and sharing when they permit it. Across three intervention axes (decision logic, cognitive profile, and model choice), no actor-level lever closes both sides of the appropriateness gap; protection and utility trade off in every variant we tested. Closing the gap may require mechanisms beyond the actor level, such as CI-grounded policy and institution design.
Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
Mengkang Hu ⋅ Bowei Xia ⋅ Yuran Wu ⋅ Ailing Yu ⋅ Yude Zou ⋅ Qiguang Chen ⋅ Shijian Wang ⋅ Jiarui Jin ⋅ Kexin Li ⋅ Wenxiang Jiao ⋅ Yuan Lu ⋅ Ping Luo
Symbolic world models (e.g., PDDL domains or executable simulators) are central to model-based planning, but training LLMs to generate such world models is limited by the lack of large-scale verifiable supervision. Current approaches rely primarily on static validation methods that fail to catch behavior-level errors arising from interactive execution. In this paper, we propose Agent2World, a tool-augmented multi-agent framework that achieves strong inference-time world-model generation and also serves as a data engine for supervised fine-tuning, by grounding generation in multi-agent feedback. Agent2World follows a three-stage pipeline: (i) A Deep Researcher agent performs knowledge synthesis by web searching to address specification gaps; (ii) A Model Developer agent implements executable world models; And (iii) a specialized Testing Team conducts adaptive unit testing and simulation-based validation. Agent2World demonstrates superior inference-time performance across three benchmarks spanning both Planning Domain Definition Language (PDDL) and executable code representations, achieving consistent state-of-the-art results. Beyond inference, Testing Team serves as an interactive environment for the Model Developer, providing behavior-aware adaptive feedback that yields multi-turn training trajectories. The model fine-tuned on these trajectories substantially improves world-model generation, yielding an average relative gain of 30.95% over the same model before training.
The rise of generative AI and autonomous agents is creating ecosystems in which stakeholders strategically choose which agents to deploy on their behalf. These design choices--such as model backbone, prompting strategy, tool access, and fine-tuning pipeline--are often strategically interdependent, as the performance of a deployed agent depends on the agents chosen by others. This position paper argues that this agent design stage should itself be a target for mediation. This perspective shifts attention from mediating only the \emph{downstream} interaction among deployed agents to mediating the \emph{upstream} stakeholder game that determines which agents enter that interaction in the first place. We formalize this agent design stage as a \emph{meta-game}, show how its equilibria can be collectively harmful, and discuss how various forms of mediation can realign incentives toward socially beneficial outcomes. We demonstrate these mediation approaches using both stylized examples and a real-world case study constructed from empirical LLM-agent interaction data. We outline concrete research directions for developing mediators that steer agentic design choices toward socially desirable outcomes.
Agentic Geometry Problem Solving via Human-like Parallel Bidirectional Reasoning
Xiaokai Zhang ⋅ Ruiqing Xia ⋅ Liang Chen ⋅ Zhenhai Sun ⋅ Yuchang Yang ⋅ Tuo Leng
Formal Geometric Problem Solving requires every step of reasoning within a strict logical framework, and has been a core challenge in artificial intelligence. This paper presents a unified neuro-symbolic reasoning framework, named Reflect Agent Verified Solve (RAVS), which deeply integrates large language models, agent architectures, and formal symbolic solvers. The core innovations of RAVS lie in two parts. First, we propose a role-separated neuro-symbolic collaboration mechanism: the LLM serves as a planner responsible for high-level semantic understanding and path reflection, while the symbolic solver acts as an executor responsible for formal verification and rigorous theorem application. The neural reasoning capability of the LLM and the logical completeness of the symbolic system complement each other, fundamentally eliminating the risk of hallucinations. Second, we construct the first bidirectional symbolic reasoning engine that fully unifies forward and backward solving. For each theorem, we define two reversible operations: forward application and backward decomposition. We further provide a theoretical analysis that bidirectional reasoning achieves an exponential complexity advantage over unidirectional reasoning. Together, these two mechanisms form an iterative closed loop between neural planning and symbolic execution. On the FormalGeo7K benchmark, RAVS achieves a 95.8% solving accuracy, substantially outperforming existing state-of-the-art methods, without requiring any additional problem-specific annotated data.
Agentic-imodels: Evolving agentic interpretability tools via autoresearch
Chandan Singh ⋅ Yan Shuo Tan ⋅ Weijia Xu ⋅ Zelalem Gero ⋅ Weiwei Yang ⋅ Michel Galley ⋅ Jianfeng Gao
Agentic data science (ADS) systems are rapidly improving their capability to autonomously analyze, fit, and interpret data, potentially moving towards a future where agents conduct the vast majority of data-science work. However, current ADS systems use statistical tools designed to be interpretable by humans, rather than interpretable by agents. To address this, we introduce Agentic-imodels, an agentic autoresearch loop that evolves data-science tools designed to be interpretable by agents. Specifically, it develops a library of scikit-learn-compatible regressors for tabular data that are optimized for both predictive performance and a novel LLM-based interpretability metric. The metric measures a suite of LLM-graded tests that probe whether a model’s string representation is “simulatable” by an LLM, i.e. whether the LLM can answer questions about the model’s behavior by reading its string output alone. We find that the evolved models jointly improve predictive performance and interpretability, generalizing to new datasets and new interpretability tests. Furthermore, these evolved models improve downstream end-to-end ADS, increasing performance for Copilot CLI, Claude Code, and Codex on the BLADE benchmark by up to 73%.
Agent Security is a Systems Problem
Mihai Christodorescu ⋅ Earlence Fernandes ⋅ Ashish Hooda ⋅ Somesh Jha ⋅ Johann Rehberger ⋅ Kamalika Chaudhuri ⋅ Xiaohan Fu ⋅ Khawaja Shams ⋅ Guy Amir ⋅ Jihye Choi ⋅ Sarthak Choudhary ⋅ Nils Palumbo ⋅ Andrey Labunets ⋅ Nishit V Pandya
We take the position that agent security must be approached as a systems problem: the AI model powering the agent must be treated as an untrusted component, and security invariants must be enforced at the system level. Through this lens, efforts to increase model robustness (the dominant viewpoint in the community) are insufficient on their own. Instead, we must complement existing efforts with techniques from the systems security domain. Based on our experience as cybersecurity researchers in operating systems, networks, formal methods, and adversarial machine learning, we articulate a set of core principles, grounded in decades of systems security research, that provide a foundation for designing agentic systems with predictable guarantees. As evidence, we analyze eleven representative real-world attacks on agents and discuss how systems principles, if realized, could have prevented these attacks. We also identify the research challenges that stand in the way of implementing these principles in agents.
AgentVista: Evaluating Multimodal Agent in Ultra-Challenging Realistic Visual Scenarios
Zhaochen Su ⋅ Jincheng Gao ⋅ Hangyu Guo ⋅ Zhenhua Liu ⋅ Lueyang Zhang ⋅ Xinyu Geng ⋅ Shijue Huang ⋅ Peng Xia ⋅ Guanyu Jiang ⋅ Cheng Wang ⋅ Yue Zhang ⋅ Yi R. (May) Fung ⋅ Junxian He
Real-world multimodal agents solve multi-step workflows grounded in visual evidence. For example, an agent can troubleshoot a device by linking a wiring photo to a schematic and validating the fix with online documentation, or plan a trip by interpreting a transit map and checking schedules under routing constraints. However, existing multimodal benchmarks mainly evaluate single-turn visual reasoning or specific tool skills, and they do not fully capture the realism, visual subtlety, and long-horizon tool use that practical agents require. We introduce AgentVista, a benchmark for generalist multimodal agents that spans 25 sub-domains across 7 categories, pairing realistic and detail-rich visual scenarios with natural hybrid tool use. Tasks require long-horizon tool interactions across modalities, including web search, image search, page navigation, and code-based operations for both image processing and general programming. Comprehensive evaluation of state-of-the-art models exposes significant gaps in their ability to carry out long-horizon multimodal tool use. Even the best model in our evaluation, GPT-5.4 with tools, achieves only 31.58% overall accuracy, and hard instances can extend to 25 tool-calling turns. We expect AgentVista to accelerate the development of more capable and reliable multimodal agents for realistic and ultra-challenging problem solving.
Aggregation Dispersion: An Information-Geometric Diagnostic of Oversmoothing in Graph Neural Networks
Abdessalam Ed-dib ⋅ Amine Aboussalah
Message-passing neural networks (MPNNs) learn node representations by iteratively aggregating information from neighbors, but stacking too many layers causes all representations to converge, a phenomenon known as oversmoothing. We propose a geometric perspective on this phenomenon. At each depth $k$, the propagation operator assigns to every node $v$ a walk distribution $p_v^{(k)}$, a probability distribution over all nodes describing how $v$ distributes its attention across the graph. These walk distributions form a point cloud on the probability simplex, and oversmoothing is the contraction of this cloud to a single point. To measure this contraction, we exploit the intrinsic geometry of the simplex: the Fisher--Rao metric, the unique Riemannian metric invariant under sufficient statistics, whose chordal distance under the square-root embedding $p \mapsto \sqrt{p}$ is the Hellinger distance. The resulting diagnostic, the \emph{aggregation dispersion} $\mathcal{D}_k$, is the variance of the embedded point cloud on the unit sphere. We prove that: (i) $\mathcal{D}_k = 0$ if and only if the $k$-th step propagation matrix has rank one, providing a necessary and sufficient condition for representational collapse; (ii) $\mathcal{D}_k$ is monotonically non-increasing with depth, so each layer irreversibly spends a finite diversity budget; (iii) the spectral gap of the propagation matrix controls the rate of this decay; and (iv) for contractive MPNNs, the aggregation dispersion upper-bounds the representation dispersion up to architecture-dependent constants, connecting the geometry of the simplex to the geometry of the feature space. In experiments across 14 model configurations and 11 datasets, $\mathcal{D}_k$ achieves the highest average Pearson correlation with accuracy degradation among all training-free diagnostics ($|r| = 0.721$), outperforming the post-training effective rank ($|r| = 0.688$).
Recent works on language identification and generation have established tight statistical rates at which these tasks can be achieved. These works typically operate under a strong realizability assumption: that the input data is drawn from an unknown distribution necessarily supported on some language in a given collection. In this work, we relax this assumption of realizability entirely, and impose no restrictions on the distribution of the input data. We propose objectives to study both language identification and generation in this more general ''agnostic'' setup. Across both problems, we obtain novel interesting characterizations and nearly tight rates.
A Graph Foundation Model for Unified Clustering
Renda Han ⋅ Xiaobao Wang ⋅ Longbiao Wang ⋅ Dayu Hu ⋅ Ronghao Fu ⋅ Di Jin
Graph clustering is a fundamental task for understanding graph-structured data without labels, yet classical end-to-end methods require retraining on each dataset, limiting their generalization ability. Although Graph Foundation Models (GFMs) enable transferable learning across diverse graph tasks, they are not directly applicable to graph clustering. This limitation stems from two key factors: the inability to learn mutually beneficial knowledge on both the node and graph levels, and the lack of a completely unsupervised clustering-oriented design. To bridge this gap, we propose a Graph Foundation Model for Unified Clustering (GFM-UC), which pre-trains on the source domain and only requires fine-tuning for fast clustering in the target domain. The entire learning process does not require any labels. Specifically, in the pre-training stage, we select learnable high-confidence unified subgraph prototypes from the source domain. And then we generalize them to train the clustering network collaboratively. In the fine-tuning stage, we align the target domain with the source domain in a distribution-aware manner to further refine the clustering head and obtain more well-separated cluster boundaries. Extensive experiments on five datasets demonstrate that our method significantly outperforms existing state-of-the-art approaches.
Agree to Disagree: Multimodal Autonomous Negotiation and Calibration for Entity Representation Learning
Chenyi Xiong ⋅ Yan Zhang ⋅ Yueluan Huang ⋅ Lu Wang ⋅ Kui Xiao ⋅ Wenxin Huang ⋅ Zhifei Li
Learning high-quality multimodal entity representations is essential for advancing reasoning tasks such as multimodal knowledge graph completion (MMKGC). However, rigid alignment in existing fusion strategies can bias representations toward a dominant anchor modality, weakening fine-grained complementary cues and propagating noisy signals into relational reasoning. To address these limitations, we propose MANoR, a Multimodal Autonomous Negotiation framework for calibrated Representation learning. MANoR treats modalities as autonomous semantic agents that negotiate before fusion. Specifically, MANoR preserves modality-specific semantics through centroid-guided subspace contraction, which reduces unstable centroid-relative dispersion while retaining entity-level residual semantics. On this stabilized basis, negotiated attention enables selective cross-modal interaction without collapsing modality boundaries. As modalities can still differ in reliability after interaction, uncertainty calibrated fusion estimates dimension-level aleatoric uncertainty and modality-level reliability to compute relation-aware fusion weights. In this way, MANoR couples modality autonomy, selective interaction, and calibrated integration within a unified representation learning framework. Experiments on MKG-W, MKG-Y, TIVA, and KVC16K show that MANoR consistently outperforms strong MMKGC baselines, achieving a relative Hits@1 gain of up to 17.72% on KVC16K.
AgriManager: A Framework and Benchmark for LLM-RL Generalization in Agricultural Management
Chi Gui ⋅ Junwen Zheng ⋅ Ritwik Nigam ⋅ Qianlan Yang ⋅ Meagan Lang ⋅ Vardhan Dongre ⋅ Jeremiah Barr ⋅ Yu-Xiong Wang ⋅ Vikram Adve
Agricultural management is heterogeneous along two axes. The first is distributional shift within a fixed task interface: the same sensors, action menu, and objective, but with weather, planted crops, and prices varying across years and locations. The second is structural shift across interfaces: different farms expose different observed variables, available actions, and management objectives. We formalize the second axis in MDP terms as \emph{schema shift}: variation in the policy-facing schema $\Sigma = (\Sigma_O, \Sigma_A, \Sigma_R)$ between training and deployment. Fixed-interface neural policies are structurally inapplicable when $\Sigma_O$ or $\Sigma_A$ shifts and require retraining when $\Sigma_R$ shifts, while prompt-conditioned LLM policies re-read $\Sigma$ from the prompt of each new episode. Whether LLM-RL actually generalizes across both axes has not been systematically tested. We address this question with \emph{AgriManager}, the first multi-simulator LLM-RL framework for agricultural management, unifying the DSSAT, WOFOST, and Cycles crop simulators under prompt-based schema specification with joint multi-objective, multi-environment training. On this platform we design a three-tier generalization ladder, each tier grounded in a real-world deployment scenario: Tier 1 within-schema distributional shift (weather, crop, price), Tier 2 single-axis schema shift along $\Sigma_O / \Sigma_A / \Sigma_R$, Tier 3 compositional schema shift across simulators. Within fixed schemas, prompt-conditioned LLM-RL with explicit reasoning matches or exceeds fixed-interface NN-RL on all three settings, with reasoning becoming necessary when the reward shape penalizes input use. Single-axis schema shift imposes a different burden along each axis: information sufficiency for $\Sigma_O$, joint action-menu coverage for $\Sigma_A$, and objective coverage for $\Sigma_R$. Under compositional schema shift, a unified policy with marginal per-platform coverage and explicit reasoning matches or exceeds per-source specialists on all three callback targets. Without that coverage, pairwise zero-shot transfer is bounded by transition dynamics, which the prompt cannot describe and which we leave as an open task. Together, these results provide an initial empirical baseline for unified prompt-conditioned control across heterogeneous agricultural deployments.
AI Construct Lexis: An Ontology of the Hidden Assumptions in AI Evaluation
Olawale Salaudeen ⋅ Florian E. Dorner ⋅ Tom Sühr ⋅ Sang Truong ⋅ Haoran Zhang ⋅ Patrick Bissett ⋅ James Fox ⋅ Marzyeh Ghassemi ⋅ Tom Hartvigsen ⋅ Peter Hase ⋅ Augustin Kelava ⋅ Suhas Mahesh ⋅ Margaret Mitchell ⋅ Jamelle Watson-Daniels ⋅ Sanmi Koyejo ⋅ Angelina Wang
There are now thousands of measurement instruments used to evaluate language models, and the resulting scores are routinely used to support consequential claims about capabilities and risks. Yet, the evidence linking benchmark scores to constructs (i.e., abstract concepts such as reasoning ability or safety) is often implicit, incomplete, or absent. We introduce AI Construct Lexis, a versioned ontology and interactive artifact that extracts and organizes relationships among constructs, behaviors, and measurement instruments as claimed in the literature for LLM evaluation. Drawing on over 8,337 published papers from selected top AI venues, we use language models to extract constructs, definitions, behavioral indicators, and measurement instruments, yielding 977 distinct constructs, 1,694 measurement instruments, 830 behavioral indicators, and 5,900 relationships. The Lexis captures relationships claimed in the published literature, not necessarily validated relationships; valid score interpretation still requires empirical evidence that the instruments capture the constructs they are used to support. By exposing the field’s implicit measurement assumptions, the Lexis enables their inspection, testing, falsification, and refinement into evidence-based relationships. On expert-annotated samples, the extraction pipeline achieves high agreement on extracted concepts while revealing substantial ambiguity in the boundaries of constructs and behaviors. By making these relationships explicit, AI Construct Lexis provides infrastructure for a more rigorous and cumulative science of AI evaluation.
AI for Drug Discovery Models Often Do Not Learn as Expected and How to Diagnose These Failure Modes
Nikhil Branson ⋅ Aaron Wenteler ⋅ Guy Durant ⋅ Charlotte Deane
We argue that in multiple areas of AI for drug discovery (AIDD), deep learning models are not learning the meaningful biological or chemical features they were hypothesised to capture, but instead learn non-generalisable features, for example, from dataset biases. To address this, we propose the systematic use of misaligned baselines. We define misaligned baselines as models that rely on signals purposely misaligned with the intended learning objective. Rather than modelling the underlying biology or chemistry directly, these baselines draw on reductionist sources of information, heavily perturbed input features, or heuristic models. Competitive performance of such baselines reveals when models are not learning the biologically meaningful representations they were designed to learn. By examining case studies across multiple branches of AIDD, we demonstrate that misaligned baselines have consistently exposed such failure modes and crucially informed the evaluation and improvement of these models. We argue that adopting them as standard practice will help to ensure that progress in AIDD reflects genuine advances rather than artefacts of data-generating processes or evaluation design.
AI Models Can Provably Hide Arbitrary Capabilities
Andis Draguns ⋅ Sumeet Motwani ⋅ Raymond Douglas ⋅ Christian Schroeder de Witt
Capability evaluations aim to surface dangerous model behaviors before deployment, but their reliability depends on the assumption that hidden capabilities can be elicited. We show this assumption does not hold under adversarial conditions by constructing backdoor attacks that embed encrypted neural networks within host model weights, decrypting and executing them only upon receiving a secret trigger. Using digital lockers, these hidden circuits remain provably hard to elicit and interpret under standard cryptographic hardness assumptions, even given full access to model weights, extending prior work from hiding fixed strings to arbitrary computations. We implement our constructions in PyTorch and empirically validate resistance to supervised fine-tuning on a small chemical reaction prediction task - a proxy for CBRN-relevant capabilities. Our results reveal a limit of capability audits. We release our models to support the development of stronger defenses.
Algorithms for Linear Equations with Min and Max Operators Under (Absolutely) Halting Condition
Krishnendu Chatterjee ⋅ Ruichen Luo ⋅ Raimundo Saona ⋅ Jakub Svoboda
We consider linear equations with min and max operators~(LEMMs) that contain many subproblems ranging from optimization, learning, to games. Recently, Chatterjee et al. [2025] gave a systematic study of the complexity of different subclasses. Three key subclasses--(I) halting branching process, (II) absolutely halting LEMMs, and (III) halting LEMMs--are proved to be in UP $\cap$ coUP while generalizing stochastic games (SSGs). In this work, we study the classic algorithms of Policy Iteration (PI) and Value Iteration (VI) for these general subclasses. First, we simplify the problem hierarchy by showing the equivalence between halting branching process and SSGs. Then, we show that while PI diverges for absolutely halting LEMMs due to the loss of monotonicity, VI remains convergent, a result we establish via diagonal rescaling. Finally, we show that neither PI nor VI converges for general halting LEMMs and, to this end, propose variants of simple policy iteration that ensure convergence across all subclasses.
Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning
Zhenchao Tang ⋅ Fang Wang ⋅ Haohuai He ⋅ Jiale Zhou ⋅ Tianxu Lv ⋅ Jun Zhu ⋅ Shou Z Chen ⋅ Minghao Yang ⋅ Yu Wang ⋅ Jiayang Wu ⋅ Yidong Song ⋅ Yaokun Li ⋅ Jiehui Huang ⋅ Jun Zhou ⋅ Bing He ⋅ Jianhua Yao
Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text exhibits a fundamentally different uncertainty structure from general text: dense low-confidence runs encode epistemic knowledge gaps (dense causal chains, rare entities) rather than the sparse aleatoric stylistic variation typical of general text. Based on this discovery, we propose Balanced Fine-Tuning (BFT), a dual-scale post-training method that combines group-normalized token reweighting with sequence-level reallocation toward knowledge-dense samples exhibiting dense epistemic uncertainty. Across medical evaluation, biological reasoning, sparse-reward RL, and biological representation tasks, BFT provides more consistent gains than SFT and DFT under a shared training setup. When replacing the default closed-source backbones in GeneAgent (GPT-4o) and VCWorld (Gemini-2.5-Flash), the BFT-aligned 70B model delivers stronger performance across biological process reasoning and chemical perturbation prediction. Critically, all BFT variants further improve after subsequent GRPO with sparse rewards, while SFT and DFT degrade, suggesting that epistemic-aware post-training provides a more robust policy initialization. Beyond text generation, BFT-aligned LLMs produce more accurate and professional biomedical profile texts; after encoding these profiles with a text embedding model, the resulting representations support gene-level, cell-level, and perturbation-response tasks, suggesting that BFT-enhanced generation can facilitate biological representation and, in turn, broader biomedical downstream tasks.
Aligning MLLMs with the Latent Structure of Human Cognition via Behavior-Derived Semantic Dimensions
Ning E ⋅ Changde Du ⋅ Yizhuo Lu ⋅ Huiguang He
Multimodal Large Language Models (MLLMs) have achieved impressive performance across a wide range of reasoning tasks, yet it remains unclear whether their internal representations and decision criteria align with the latent structure of human cognition. In this work, we propose a novel framework to align MLLMs with the multidimensional mental representations underlying human similarity judgements. We build on a previously established behavior-derived semantic embedding space, in which objects are represented along 66 interpretable dimensions derived from 4.70 million human Odd-One-Out (O1O) judgements. For each triplet, we infer the most salient decision dimension and convert it into a dimension-grounded linguistic rationale. This enables MLLMs to learn both the final choice and a justification consistent with the inferred behavioral criterion. Experimental results demonstrate that our approach improves model consistency with human similarity judgements while retaining competitive performance on standard multimodal benchmarks. Furthermore, using searchlight Representational Similarity Analysis (RSA) on an independent fMRI Natural Scenes Dataset (NSD), we observe increased representational alignment between model and human. Ultimately, our findings point to a data-driven route for incorporating human cognitive structure into MLLMs alignment.
Although Large Language Models (LLMs) achieve strong alignment through supervised fine-tuning and reinforcement learning from human feedback, the alignment is often fragile under subsequent fine-tuning. Existing explanations either attribute alignment fragility to gradient geometry or characterize it as a distributional shift in model outputs, yet few provide a unified account that bridges parameter-space learning dynamics with function-space alignment behavior during fine-tuning. In this work, we introduce a tractable alignment score and derive its closed-form update during fine-tuning, yielding a unified framework for alignment dynamics. Our analysis decomposes alignment updates into two competing components: a Rebound Force, governed jointly by the current alignment state and the narrowness of model distribution, and a Driving Force, determined by how the training distribution aligns with outcome-conditioned posteriors over aligned and non-aligned completions. This decomposition explains why prior alignment can be reversed by later fine-tuning and why narrower posterior structure strengthens such reversal. Moreover, our framework predicts a Rehearsal Priming Effect: prior alignment leaves a latent posterior imprint that amplifies the effective Driving Force upon re-exposure, leading to faster re-alignment. We validate these predictions across safety alignment, emergent misalignment, and sentiment settings, demonstrating consistent alignment reversal and accelerated re-alignment under re-exposure. In addition, controlled experiments in safety alignment confirm the predicted dependence of rebound strength on posterior narrowness. Together, these results provide a unified dynamical perspective on how alignment is disrupted and reactivated during LLM fine-tuning.
Allocentric Perceiver: Disentangling Allocentric Reasoning from Egocentric Visual Priors via Frame Instantiation
Hengyi Wang ⋅ RuiQiang Zhang ⋅ Chang Liu ⋅ Guanjie Wang ⋅ Zehua Ma ⋅ Han Fang ⋅ Weiming Zhang
With the rising need for spatially grounded tasks such as Vision-Language Navigation/Action, allocentric perception capabilities in Vision-Language Models (VLMs) are receiving growing focus. However, VLMs remain brittle on allocentric spatial queries that require explicit perspective shifts, where the answer depends on reasoning in a target-centric frame rather than the observed camera view. Thus, we introduce $\textbf{Allocentric Perceiver}$, a training-free strategy that recovers metric 3D states from one or more images with off-the-shelf geometric experts, and then instantiates a query-conditioned allocentric reference frame aligned with the instruction’s semantic intent. By deterministically transforming reconstructed geometry into the target frame and prompting the backbone VLM with structured, geometry-grounded representations, Allocentric Perceiver offloads mental rotation from implicit reasoning to explicit computation. We evaluate Allocentric Perceiver across multiple backbone families on spatial reasoning benchmarks, observing consistent and substantial gains ($\sim$\textbf{10\%}) on allocentric tasks while maintaining strong egocentric performance, and surpassing both spatial‑perception-finetuned models and state‑of‑the‑art open‑source and proprietary models.
AmbiguousWorld: Benchmarking and Resolving Ambiguous Instructions in Video World Models
Bo Wang ⋅ Sen Cui ⋅ Jin Li ⋅ Ming Hei Huang ⋅ Zhikang Chen ⋅ Min Zhang ⋅ Guocai Yao ⋅ Changshui Zhang
Translating high-level, under-specified human commands into coherent video is fundamentally challenging. Current video generation models lack explicit reasoning capabilities and typically fail to understand ambiguous instructions, frequently resulting in physically inconsistent and causally disjointed generation. To address this, we introduce AmbiguousWorld, a novel inference-time reasoning framework that transforms video generation into a multi-modal Tree of Thoughts. We employ a Vision-Language Model as both a semantic planner to dynamically decompose fuzzy instructions into actionable subgoals, and as a closed-loop critic to prune invalid trajectories by evaluating speculative paths across multiple physical and functional dimensions. To systematically assess this, we annotate the first embodied video generation benchmark targeting ambiguous instructions based on the DROID dataset, and design a multi-dimensional VLM evaluation framework. Empirical results on state-of-the-art baselines demonstrate that our framework yields substantial improvements, boosting absolute Task Success by up to 29.0\% and the overall generative consistency weighted score by 24.2 points. Furthermore, we validate the effectiveness of our VLM evaluation framework through rigorous comparison with human assessment.
A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems
Christoph Jürgen Hemmer ⋅ Florian Plaswig ⋅ Daniel Durstewitz
Recent foundation models (FMs) for zero-shot reconstruction of dynamical systems (DS) achieve strong out-of-domain generalization but provide little insight into the mechanisms that underlie their forecasts. Such an understanding could help to strip down overladen FM architectures to their bare essence and expose the minimal requirements for in-context learning in the DS domain. Toward this goal, here we iteratively reduce a recent powerful SOTA model for DS reconstruction, DynaMix (Hemmer & Durstewitz, 2025), to a minimal interpretable two-parameter form, which we call DynaBase. DynaBase produces forecasts through a linear blend of the current latent state and the nearest in-context neighbor and its temporal successor. Surprisingly, despite its extreme simplicity, DynaBase produces highly competitive zero-shot DS reconstructions across chaotic and cyclic systems, with a negligible parameter load, many orders of magnitude below that of other FMs. Even more, this extreme simplicity permits direct model optimization on DS reconstruction measures, as well as closed-form one-step analytical solutions on prediction MSE. Theoretical and empirical analysis of DynaBase further leads to a 1-parameter family of maps, with the context-parroting algorithm of (Zhang & Gilpin, 2026) recovered at one end, and chaotic (divergent but bounded) behavior at the other. We further show how different training strategies lead to models either optimal for short-term prediction or for DS reconstruction. Thus, DynaBase not only exposes the minimal mechanisms required for producing zero-shot DS reconstruction, but also reconciles within an accessible mathematical frame divergent observations in the literature.
We study image inpainting with generative diffusion models. Existing methods typically either train dedicated task-specific models, or adapt a pretrained diffusion model separately for each masked image at deployment. We introduce a middle-ground model, termed Amortized Inpainting with Diffusion (AID), which keeps a pretrained diffusion backbone fixed, trains a small reusable guidance module offline, and then reuses it across masked images without per-instance optimization. We formulate it as a deterministic guidance problem with a supervised terminal objective. To make this problem learnable in high dimensions, we derive an auxiliary Gaussian formulation and prove that solving this randomized problem recovers the optimal deterministic guidance field. This bridge yields a principled continuous-time actor--critic algorithm for learning the guidance module in a fully data-driven manner. Empirically, on AFHQv2 and FFHQ under the pixel EDM pipeline and on ImageNet under the latent EDM2 pipeline, AID consistently improves the quality--speed trade-off over strong fixed-backbone and amortized inpainting baselines across multiple mask types, while adding less than one percent trainable overhead.
Amortizing Generative Guidance for Model-Based Reinforcement Learning
Xiangteng Zhang ⋅ Guojian Zhan ⋅ Likun Wang ⋅ Jingliang Duan ⋅ Yang Guan ⋅ Shengbo Eben Li
Planner-guided model-based reinforcement learning (MBRL) has emerged as a powerful paradigm for continuous control, combining learned world models with online planning to achieve strong performance and sample efficiency. However, existing methods typically distill planner-improved actions into a unimodal Gaussian policy and reuse it as a proposal prior for subsequent planning, overlooking the fact that Model Predictive Path Integral (MPPI) planners can induce multimodal supervision over high-value actions. Compressing such supervision into a unimodal distribution can lead to mode averaging, limited proposal coverage, and unstable bootstrapped policy learning. To address this issue, we propose GeMP, a \textbf{Ge}nerative \textbf{M}ultimodal \textbf{P}lanning framework for MBRL. We theoretically characterize the multimodality of planner-induced supervision, showing that finite-sample planning can provide a distribution of high-value actions rather than a single deterministic target. To capture this distribution without compromising training stability, GeMP jointly learns a multimodal flow policy through value-weighted supervision and retains an auxiliary Gaussian policy for stablizing off-policy actor-critic learning. We further employ a heterogeneous proposal generator that combines Gaussian, flow-based, and random action-sequence candidates to improve proposal coverage for MPPI refinement. Experiments on MyoSuite and DeepMind Control Suite show that GeMP improves control performance and training stability over competitive baselines.
Anatomy-aware Spatio-Temporal Modeling for Echocardiography Segmentation
Seok-Hwan Oh ⋅ Guil Jung ⋅ Myeong-Gee Kim ⋅ Young-Min Kim ⋅ hyeonjik lee ⋅ Sang-yun Kim ⋅ Hyuksool Kwon ⋅ Hyeon-Min Bae
Echocardiography segmentation plays a central role in quantitative cardiovascular assessment. However, automated echocardiography segmentation remains challenging due to speckle noise, poorly defined anatomical boundaries, acquisition variability, and the limited anatomical coverage of existing benchmarks. Conventional segmentation methods demonstrate limited accuracy under such low image quality and domain shifts. Inspired by the observation that expert cardiologists can maintain reliable anatomical delineation by leveraging their structural understanding of the cardiac anatomy, we propose that incorporating anatomical priors into echocardiography segmentation models can improve both precision and robustness. To this end, we propose AST-Seg, an anatomy-aware spatio-temporal framework for echocardiography segmentation. AST-Seg introduces an anatomy-aware feature encoder that regularizes transformer attention with anatomy-derived spatial-relation priors, and a Deformable Spatio-Temporal Mamba module that captures localized and periodic cardiac motion through adaptive spatial aggregation followed by temporal state-space modeling. In addition, we curate an Echocardiography Anatomy (EA) dataset with pixel-level annotations for 12 clinically relevant cardiac anatomy, enabling comprehensive multi-structure cardiac interpretation. Extensive experiments on in-distribution CAMUS and EA datasets, as well as an out-of-distribution point-of-care echocardiography dataset, demonstrate that AST-Seg provides domain-generalized segmentation with improved precision.
An Efficient Geometric Characterization of Robust Fair Learning
Sushant Agarwal ⋅ Amit Jayant Deshpande ⋅ Rajmohan Rajaraman ⋅ Ravi Sundaram
Previous work has shown that the optimal classifier from a hypothesis class subject to exact fairness constraints may not be *robust*; that is, its accuracy can change drastically under small shifts to the underlying data distribution. We ask the following questions: Given a hypothesis class $\mathcal{H}$, is the accuracy of the optimal fair classifier from $\mathcal{H}$ robust to malicious distribution shifts? And is this property efficiently testable? We consider $\mathcal{H}$ that can be represented by a convex polytope — the natural setting for randomized ensembles of deterministic classifiers, including those produced by boosting. We provide a complete geometric characterization of when $\mathcal{H}$ is robust, and an efficient linear programming test to audit the robustness of a given $\mathcal{H}$. Surprisingly, we demonstrate that if we relax exact fairness constraints and only require approximate fairness, every $\mathcal{H}$ is robust. Our results hold for a broad class of linear fairness criteria (e.g., demographic parity, equal opportunity, predictive equality) and linear loss functions (e.g., 0/1 loss, weighted loss).
Angular Networks: Low-Bit Learning from Randomized Similarity Estimators
Ali Ahmed ⋅ Saeid Pourmand ⋅ Muhammad Awais ⋅ Rehan Farooq ⋅ Alireza Aghasi
We propose a new framework for low-precision neural networks based on angular geometry and randomized low-bit estimators. Rather than approximating Euclidean linear operations under limited precision, we reinterpret neural computation in terms of directional similarity and construct quantized random-feature estimators of cosine interactions. Our approach introduces \textit{Angular Layers}, which replace standard linear transformations with low-bit projections that estimate angular similarity in the forward pass, while gradients are computed with respect to the underlying continuous cosine geometry. This decoupling enables stable optimization without straight-through estimators or heuristic surrogate gradients, even under aggressive quantization of both weights and activations. On the theory side, we introduce \emph{G-Gradient} as a gradient proxy for angular layers and show that, in a one-layer planted model, a sufficiently small \emph{G}-gradient can guarantee exact recovery of a target binary solution. We further prove uniform approximation and convergence guarantees that provide a rigorous path from tractable continuous optimization to exact recovery in the original discrete model. Empirically, we instantiate these ideas in \textit{Angular Nets}, including low-precision variants of ResNet, Vision Transformers, and BERT-style encoders, and obtain strong performance across vision and language tasks at substantially reduced precision. Overall, our results suggest that preserving angular structure provides a principled foundation for low-bit deep learning. Our full implementation is available at: \url{https://github.com/GGM2026/GGM}.
Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation
Anya Ji ⋅ Abhijith Varma Mudunuri ⋅ David Chan ⋅ Alane Suhr
While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code for webpages, charts, or SVGs, it is still unclear VLMs are able to recover temporal dynamics when motion is present. To this end, we introduce the animation-to-code task, as well as a benchmark, \texttt{Animation2Code}, for evaluating temporal visual reasoning via reconstructing executable web animation code from videos. \texttt{Animation2Code} consists of 1,069 web animation videos with diverse visuals and motion patterns, augmented with corresponding HTML/CSS/JavaScript implementations. We further pair the dataset with several novel human-aligned metrics, which allow us to disentangle visual similarity from temporal alignment when comparing rendered animations to ground-truth samples. We benchmark current SOTA model performance on these new samples, and show that current models struggle with temporal consistency, even when achieving high appearance similarity, even in fine-tuning and iterative refinement scenarios. Our benchmark dashboard is available at: \url{https://anim2code-dashboard.vercel.app}.
An In-Depth Analysis of Hallucination Detection Methods for Vision-Language Models
Allison Chen ⋅ William Yang ⋅ Salma Abdel Magid ⋅ Jonathan Williams ⋅ Olga Russakovsky ⋅ Esin Tureci
Currently, evaluation of hallucination detection methods for vision-language models (VLMs) only reports aggregated statistics, overlooking potential confounding factors and providing limited insight into the strengths and weaknesses of each method. We conduct an in-depth analysis of what drives detection performance, focusing on methods that use either internal model signals (white box) or output logits (black box). First, our analyses show that a naive baseline of a word's token position in the caption is a competitive predictor of hallucinations and that many white box methods, despite performing the best, rely on this heuristic. Second, we find that some white box methods are additionally specialized to VLM architecture, such that when used with certain VLMs, they can even detect hallucinations in captions generated by other VLMs. Lastly, we find that the discriminative advantage of white box methods over black box primarily arises from detecting language-based hallucinations, as opposed to vision-based. Taken together, these analyses reveal insights into hallucination detection methods that are not captured by current evaluation protocols. Drawing from our findings, we encourage future work to develop more comprehensive evaluations that can better reflect hallucination detection behavior.
AnyEdit: A Unified Framework for Speech and Singing Voice Editing with Real-World Environmental Consistency
Yunjia Zhang ⋅ Junan Zhang ⋅ Jing Yang ⋅ Xueyao Zhang ⋅ Fan Fan ⋅ Zhizheng Wu
We present AnyEdit, a unified framework for speech and singing voice editing in real-world acoustic environments. Recent unified generative models already achieve strong speech and singing generation in clean conditions, but real-world vocal editing remains challenging because the model must revise content while simultaneously preserving speaker identity, prosody, and the surrounding acoustic scene. A straightforward solution is to directly model noisy edited audio end to end, yet this entangles semantic editing with environmental acoustics and can weaken the core vocal generation capability. We therefore adopt a two-stage design that preserves clean-domain editing ability while deferring acoustic rendering to a separate stage. First, an instruction-guided autoregressive editor predicts clean prosody and content-style tokens. To make this stage robust to degraded inputs, we introduce a teacher-student noise-robust prosody tokenizer that distills clean prosodic tokens from noisy recordings. Second, a flow-matching acoustic model performs in-context acoustic rendering conditioned on the source acoustic context, enabling the edited segment to inherit surrounding noise, reverberation, and background sound while maintaining speaker timbre consistently. To further stabilize the unedited regions, we introduce a unmodified-area-aware supervision mechanism. After pre-training, we also apply direct preference optimization (DPO) using real singing editing pairs mined via dynamic time warping to achieve better singing editing quality. To enable realistic evaluation, we introduce AnyEditBench, the first benchmark covering both speech and singing voice editing under diverse acoustic conditions. Experiments on Chinese and English speech and singing datasets showed that AnyEdit consistently outperforms strong baselines in content accuracy, subjective quality, and background consistency under noisy and reverberant conditions. Demo audios can be found at https://any-edit.github.io.
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
Yiheng Li ⋅ Zhuo Li ⋅ RuiBing Hou ⋅ Yingjie Chen ⋅ Hong Chang ⋅ Hao Liu ⋅ Shiguang Shan
Conditional human motion generation remains a fundamental challenge in computer vision and robotics. Despite significant progress, current methods are often constrained by fixed modality configurations and task-specific architectures, leaving cross-modal interactions and the scaling laws of multimodal-conditioned synthesis largely underexplored. A key bottleneck is the scarcity of large-scale modality-aligned motion data, limiting generalization across diverse control signals. In this work, we introduce \textbf{OmniHuMo}, a large-scale, high-quality dataset comprising over 5,000 hours of motion and 3.2 million sequences with precisely aligned multimodal annotations (e.g., text, speech, music, and trajectory). Leveraging OmniHuMo, we propose \textbf{AnyMo}, a unified multimodal framework combining a Residual FSQ-based motion tokenizer with a scalable masked modeling transformer, enabling high-quality motion synthesis under arbitrary modality combinations. Extensive experiments show that AnyMo achieves high-fidelity synthesis while offering flexible control over both spatial and stylistic attributes.
Test-time scaling systems often choose how much to sample, which prompt to use, and which verifier or aggregation rule to trust after repeatedly inspecting a calibration set. Standard validation error bars are not valid after this adaptive selection. We model a test-time scaling recipe as a randomized predictor with a random cost and prove a time-uniform PAC-Bayes certificate that holds simultaneously over all calibration times, all posterior mixtures over recipes, and all data-dependent stopping rules. Thus the same event covers the recipe search that precedes deployment, not merely the fixed recipe finally reported. For a finite grid of budgets and verifiers, the penalty is the familiar $\log M/n$, while non-uniform priors and localized posterior mixtures give sharper certificates for structured searches. We specialize the result to self-consistency and show a complementary margin law: plurality voting improves exponentially in the sample budget only on examples where the correct answer is the generator's unique mode; otherwise more compute can provably amplify errors. The same margin analysis yields an oracle water-filling law for compute allocation. A synthetic audit quantifies the optimism caused by adaptive budget selection, and a 192-example GSM8K audit using gpt-4.1-mini generations illustrates risk-cost reporting for a real inference pipeline. The results give a lightweight, distribution-free way to report risk and compute claims for adaptive inference systems.
ApertureAttn: Native 4K Video Generation with Image-Only Supervision
Ruonan Yu ⋅ Zigeng Chen ⋅ Zhenxiong Tan ⋅ Songhua Liu ⋅ Xinchao Wang
Recent progress in video generation has been remarkable, yet state-of-the-art models remain largely confined to 720p, well short of the 4K resolution of modern displays. Scaling video generation to higher resolutions is constrained by the scarcity of high-resolution video data and the prohibitive cost of large-scale training. Beyond data and compute, higher resolutions place substantially greater demands on spatial modeling, particularly for fine-grained details, complex compositions, and local structures. Such high-resolution spatial supervision, however, is abundantly available in high-resolution images, which are easier to collect and train on at scale. This naturally raises a question: can image-only supervision unlock native 4K video generation? To answer this, we propose ApertureAttn, an inverted attention pyramid that concentrates adaptation on high-resolution fine-grained spatial modeling using only single-frame high-resolution images, without any video training data. We further introduce window-scaled extrapolation and proximity-guided temporal attention to enhance spatiotemporal consistency. Extensive experiments show that our method enables native 4K video generation with strong visual fidelity and temporal coherence, while requiring only image supervision and minimal training cost. Despite this lightweight training setup, it remains competitive with state-of-the-art high-resolution video generation methods trained directly on high-resolution video data at substantially higher training cost.
Sequence models in machine learning compress long input histories into finite-dimensional states, from recurrent networks and reservoir computing to recent state space models for long-context data. This raises a basic expressivity question: how simple can the recurrent memory be before it can no longer approximate stable causal sequence-to-sequence maps? We answer this question for fading-memory operators, where the distant past has uniformly diminishing influence. We show that every continuous causal time-equivariant fading-memory operator on a compact input set can be uniformly approximated by a state space model whose recurrence is a fixed diagonal linear contraction. The dynamics are time-invariant, input-independent, and contain no nonlinear hidden-state update or delay-line structure. The construction uses a bank of short exponential traces; although such traces do not exactly store a recent input window, a singular Vandermonde extraction approximately reconstructs the window while the remaining tail becomes negligible as delta decreases. We also prove that exact delay recovery is impossible for any fixed finite trace bank, showing that the accuracy-dependent extraction is essential. Technically, the result separates recurrent expressivity from exact memory storage, while exposing a conditioning cost in the readout. More broadly, it clarifies why highly stable, structurally simple SSM cores can still be universal approximators for stable sequence behavior.
Are Well-Trained Surrogates Optimal? Rethinking the Surrogate Role with Instability for Unlearnable Examples
Binze Wang ⋅ Jinyu Tian ⋅ Xingrun Wang ⋅ Jianqing LI
Unlearnable Examples (UEs) aim to protect data from unauthorized model training by injecting imperceptible perturbations that degrade model generalization. Existing methods typically rely on well-trained surrogate models to generate such perturbations, implicitly assuming that stronger surrogates yield better degradation. In this work, we challenge this widely adopted assumption and show that it is fundamentally suboptimal. We provide theoretical insights by establishing an upper bound on the target model's test loss, showing that perturbation effectiveness is closely related to the surrogate’s sensitivity to input perturbations. Specifically, a less robust surrogate model leads to stronger poisoning performance. Motivated by this analysis, we propose a simple yet effective stability metric to select optimal surrogate models for single-level methods. We further extend this principle to bi-level optimization frameworks. Beyond revealing the bi-level method's implicit reliance on non-robust surrogates, we substantially boost its poisoning performance by the stability-based surrogate selection as well. Experiments on five datasets and four representative UE methods show that our selected surrogates reduce the target model's test accuracy by an average of 18\% compared with well-trained surrogates. Notably, this is the first work to reduce CIFAR-10 test accuracy to 10\% under a perturbation budget of 1/255, whereas prior UE methods typically rely on 8/255.
Are Your Reasoning Models Reasoning or Guessing? A Mechanistic Analysis of Hierarchical Reasoning Models
Zirui Ren ⋅ Ziming Liu
Hierarchical reasoning model (HRM) achieves extraordinary performance on various reasoning tasks, significantly outperforming large language model-based reasoners. To understand the strengths and potential failure modes of HRM, we conduct a mechanistic study on its reasoning patterns and find three surprising facts: (a) Failure of extremely simple puzzles, e.g., HRM can fail on a puzzle with only one unknown cell. We attribute this failure to the violation of the fixed point property, a fundamental assumption of HRM. (b) "Grokking" dynamics in reasoning steps, i.e., the answer is not improved uniformly, but instead there is a critical reasoning step that suddenly makes the answer correct; (c) Existence of multiple fixed points. HRM "guesses" the first fixed point, which could be incorrect, and gets trapped there for a while or forever. All facts imply that HRM appears to be "guessing" instead of "reasoning". Leveraging this "guessing" picture, we propose three strategies to scale HRM's guesses: data augmentation (scaling the quality of guesses), input perturbation (scaling the number of guesses by leveraging inference randomness), and model bootstrapping (scaling the number of guesses by leveraging training randomness). On the practical side, by combining all methods, we develop Augmented HRM, boosting accuracy on Sudoku-Extreme from 55.0% to 96.9%. On the scientific side, our analysis provides new insights into how reasoning models "reason".
Vision Transformers (ViTs) face severe computational bottlenecks due to the quadratic complexity of self-attention at high resolutions. Existing token reduction methods rely on local metrics—such as single-layer attention scores—that are inherently vulnerable to the \emph{attention sink} phenomenon, where uninformative tokens are paradoxically preserved over salient foreground objects. We propose \textbf{ASAP} (Attention Sink Anchored Pruning), a training-free framework that recasts this sink as a feature. Modeling ViT information flow as a Lazy Random Walk, ASAP identifies the sink as a dominant accumulator of probability mass. By computing the \emph{diffusion distance} to the sink within the cumulative transition matrix, ASAP partitions tokens via \emph{Radial Diffusion Clustering} and compresses background redundancy through \emph{Transition Weight Pooling} in a single shot. Extensive experiments across image, video, and vision-language tasks demonstrate ASAP outperforms state-of-the-art methods, accelerating throughput by up to 48\% while maintaining—or even exceeding—baseline accuracy.
A Scalable Multi-Task Model for Virtual Sensors
Leon Götz ⋅ Lars Frederik Peiss ⋅ Erik Sauer ⋅ Andreas U Sass ⋅ Thorsten Bagdonat ⋅ Stephan Günnemann ⋅ Leo Schwinn
Virtual sensors replace expensive physical sensors in critical applications through machine learning by predicting target signals from available measurements. Existing virtual sensor approaches require application-specific models with hand-selected inputs for each sensor, cannot leverage task synergies, and lack consistent benchmarks. While emerging time series foundation models offer general-purpose, pretrained solutions in other domains, they are computationally expensive and limited to predicting their input signals, making them incompatible with virtual sensors. We introduce the first multi-task model for virtual sensors addressing both limitations. Our unified model can simultaneously predict diverse virtual sensors exploiting synergies while maintaining computational efficiency. It learns relevant input signals for each virtual sensor, eliminating expert knowledge requirements while adding explainability. In our large-scale evaluation on three standard benchmarks and an application-specific dataset with over 18 billion samples, our architecture reduces computation time by up to 415x and memory requirements by 951x, while maintaining or even improving predictive quality compared to unified baselines. Compared to existing isolated models for a single virtual sensor, our unified approach generates superior predictions at similar inference speed while scaling gracefully to hundreds of virtual sensors with nearly constant parameter count, enabling practical deployment in large-scale sensor networks.
Assign and Add: A Mechanistic Study of Compositional Arithmetic
Brady Exoo ⋅ Alberto Bietti ⋅ John Sous
Modern large language models are able to compose skills in order to perform complex tasks, many of which might not have been seen during training. The details of how exactly this composition occurs remain elusive. In this paper, we study a mechanism for compositional generalization in transformers by considering a simple controlled setting involving variable assignment and modular addition. By partitioning our training data into disjoint sets, we observe that small transformers are able to generalize to previously unseen combinations of variables and numbers. Our mechanistic analysis shows that the same ``modular addition'' MLP module is used whether the inputs are given directly or indirectly through a separate variable assignment mechanism. We also analyze the training dynamics from an empirical lens, which reveals three phases of learning: first, modular addition is learned, then the structure required for variable assignment, and finally a refinement phase where the model generalizes to sequences not seen in training. Finally, we provide a theoretical framework to explain how compositionality emerges from training dynamics. These results suggest that compositional generalization can be a natural consequence of the compositionality of internal mechanisms in transformers.
Assistive Dueling Bandits: No-Regret Algorithms for Assisting No-Regret Users
Mark Bedaywi ⋅ Cassidy Laidlaw ⋅ Austin Tripp ⋅ Nika Haghtalab
We study a stochastic bandit problem in which learning is split between two agents. In each round, a human observes rewards but can choose only between a pair of arms selected by an assistant, while the assistant observes the human's behavior but not realized rewards. We call this model the **assistive dueling bandit**. The human is modeled as a learning agent whose regret on any subset of arms is bounded by a function $g(T)$ unknown to the assistant. This is a natural model for recommender systems or AI assistants that must present a slate of options to a user who may be still learning about their own preferences. We provide a general reduction from assistive dueling bandits to the problem of **max-finding with imprecise feedback.** With it, we design assistant algorithms that cooperate with any $g(T)$-regret human to achieve a joint regret of $\tilde{\mathcal{O}}(K \cdot g(T/K))$, even without knowledge of $g(T)$. Crucially, when $g(T) \in \mathcal{O}(\sqrt{T})$, our method recovers the nearly optimal minimax rate of $\tilde{\mathcal{O}}(\sqrt{KT})$, implying that splitting up the responsibilities of learning in this way can be done without suffering any excess regret asymptotically. We complement our theoretical analysis with experimental evidence that this algorithm outperforms an assistant implemented via standard dueling bandit algorithms.
Asymmetric Flow Models
Hansheng Chen ⋅ Jan Ackermann ⋅ Minseo Kim ⋅ Gordon Wetzstein ⋅ Leonidas Guibas
Flow-based generation in high-dimensional spaces is difficult because velocity prediction requires modeling high-dimensional noise, even when data has strong low-rank structure. We present Asymmetric Flow Modeling (AsymFlow), a rank-asymmetric velocity parameterization that restricts noise prediction to a low-rank subspace while keeping data prediction full-dimensional. From this asymmetric prediction, AsymFlow analytically recovers the full-dimensional velocity without changing the network architecture or training/sampling procedures. On ImageNet 256$\times$256, AsymFlow achieves a leading 1.57 FID, outperforming prior DiT/JiT-like pixel diffusion models by a large margin. AsymFlow also provides the first-ever route for finetuning pretrained latent flow models into pixel-space models: a patch-level linear lift initializes a low-rank pixel model whose denoising trajectory preserves the latent model's high-level semantics and structure, so finetuning mainly improves low-level residuals rather than relearning pixel generation. We show that the pixel AsymFlow model finetuned from FLUX.2 klein 9B establishes a new state of the art for pixel-space text-to-image generation, beating its latent base on HPSv3, DPG-Bench, and GenEval while qualitatively showing substantially improved visual realism. Code and models will be released publicly.
Asymmetric Hierarchical Anchoring for Robust Audio–Visual Cross-Modal Generalization
Bixing Wu ⋅ Yuhong Zhao ⋅ Zongli Ye ⋅ Jiachen Lian ⋅ Xiangyu Yue ⋅ Gopala Anumanchipalli
Audio--visual joint representation learning under Cross-Modal Generalization (CMG) aims to transfer knowledge from a labeled source modality to an unlabeled target modality through a unified discrete representation space. Existing symmetric frameworks often suffer from information allocation ambiguity, where the absence of structural inductive bias leads to semantic--specific leakage across modalities. We propose Asymmetric Hierarchical Anchoring (AHA), which enforces directional information allocation by designating a structured semantic anchor within a shared hierarchy. In our instantiation, we exploit the hierarchical discrete representations induced by audio Residual Vector Quantization (RVQ) to guide video feature distillation into a shared semantic space. To ensure representational purity, we replace fragile mutual information estimators with a GRL-based adversarial decoupler that explicitly suppresses semantic leakage in modality-specific branches, and introduce Local Sliding Alignment (LSA) to encourage fine-grained temporal alignment across modalities. Extensive experiments on AVE and AVVP benchmarks demonstrate that AHA consistently outperforms symmetric baselines in cross-modal transfer. Additional analyses on talking-face disentanglement experiment further validate that the learned representations exhibit improved semantic consistency and disentanglement, indicating the broader applicability of the proposed framework.
A Theoretical Analysis of Why Masked Diffusion Models Mitigate the Reversal Curse
Moongyu Jeon ⋅ Sangwoo Shin ⋅ Bumjun Kim ⋅ Kyelim Lee ⋅ Albert No
Autoregressive language models (ARMs) suffer from the reversal curse: after learning "$A$ is $B$," they often fail on the reverse query "$B$ is $A$." Masked diffusion language models (MDMs) exhibit this failure in a much weaker form, but the underlying reason has remained unclear. A common explanation attributes this mitigation to their any-order masked training objective. However, observing "$[\textnormal{\textbf{M}}]$ is $B$" during training teaches recovery of $A$ from $B$ in one positional configuration, and does not by itself explain why the learned evidence should transfer to the reverse prompt "$B$ is $[\textnormal{\textbf{M}}]$." We provide a theoretical analysis showing that this transfer arises from a parameter-level coupling between forward and reverse positional conditionals: shared Transformer parameters store token-pair evidence, while relative positional encodings route attention through queries and keys without changing the value-side evidence being retrieved. In a one-layer MDM, we prove that forward masked training strengthens evidence that is reusable in reverse queries, induces correlated forward--reverse attention routes, and yields a positively aligned shared-storage gradient component that decreases the reverse loss to first order. Controlled one-layer experiments and large-scale LLaDA/Dream experiments verify these signatures and show that they translate into improved reverse prediction.
ATLAS: Adaptive Temporal Learning for Single-Cell Multi-Omics Alignment and Dynamics
Ye Zhang ⋅ Zijie Fang ⋅ Zhixiang Lin
Single-cell multi-omics technologies provide a powerful basis for characterizing cell-state transitions and dynamic regulatory processes. However, existing methods for multi-omics dynamic modeling still face two major limitations. From the data perspective, many methods require high-quality paired multi-omics measurements; from the modeling perspective, dynamic inference often relies on predefined regulatory structures or kinetic assumptions. These limitations restrict their applicability to partially paired, unpaired, and broader cross-omics settings. To address these challenges, we propose ATLAS, a unified framework for single-cell multi-omics alignment and dynamic modeling that jointly learns cross-omics consistency and temporal dynamics from partially paired or unpaired data. ATLAS adaptively models temporal lag effects between omics layers to characterize asynchronous cross-omics regulation, and further introduces a reliability-guided temporal distillation strategy to improve model-based temporal ordering. We systematically evaluate ATLAS on five datasets across four tasks, showing strong overall performance in multi-omics alignment, cross-omics prediction, trajectory inference, and future-state prediction. Our code is available at https://github.com/anomity/ATLAS.
AtlasULP: Domain-aware Universal Link Prediction via Relation Atlas
Yujing Liu ⋅ Yixin Liu ⋅ Yu Zheng ⋅ Lianhua Chi ⋅ Alan Wee-Chung Liew ⋅ Hengtao Shen ⋅ Shirui Pan
Link prediction (LP) is a widely applied task in graph learning. Conventional LP approaches typically follow a dataset-specific paradigm, requiring independent training for each graph and incurring high computational and maintenance costs in large-scale applications. Motivated by these limitations, recent work explores Universal Link Prediction (ULP), aiming to enable training-free inference on arbitrary unseen graphs. However, existing methods primarily rely on subgraph sampling strategies to construct universal representations, which are unable to capture higher-order information due to the exponential growth of higher-order neighborhoods. As a result, they often under-perform on graphs where long-range signals are essential, such as biological networks. Moreover, their prediction models are typically domain-agnostic and lack the ability to adapt to different graph domains, leading to limited generalization. To address these challenges, we propose an Atlas-based Universal Link Prediction framework (AtlasULP), which introduces the relation atlas to capture structural associations between node pairs from both global and local perspectives, enabling a comprehensive characterization of connectivity signals across diverse graph domains. Building upon this representation, we develop a domain-guided prompting ULP model, which generates domain-aware prompt tokens from contextual structures and performs prompt-guided in-context prediction for adaptive link inference. Extensive experiments demonstrate that AtlasULP consistently outperforms state-of-the-art methods across diverse graph domains.
Attack Selection In Agentic AI Control Evaluations Meaningfully Decreases Safety
Catherine Ge-Wang ⋅ Tyler Crosse ⋅ Benjamin Hadad ⋅ Joachim Schaeffer ⋅ Ram Potham ⋅ Tyler Tracy
An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but untrusted AI agents under oversight from a weaker trusted monitor and a limited human audit budget. Control evaluations stress-test these protocols by pitting a red-team attack policy against the blue-team monitor, but current evaluations typically assume attackers that do not strategically select when to attack. However, a capable attacker would not attack indiscriminately. We study this capability, attack selection, in agentic settings by decomposing attack decisions into a start policy, which decides when an attacker should attack, and a stop policy, which decides when an attacker should abort an ongoing attack. Across two agentic settings, BashArena and LinuxArena, both policies substantially lower measured empirical safety without changing the underlying attack capability. At a 1% audit budget, our start policy reduces safety by 20pp on both BashArena and LinuxArena, and our stop policy reduces safety by 20pp on BashArena and 28pp on LinuxArena. These reductions in safety should be interpreted as upper bounds on the effect of attack selection. Existing control evaluations may therefore yield overly optimistic safety estimates against selective attackers. We recommend that future evaluations, system cards, and safety cases should elicit attack selection to produce more realistic safety estimates.
Attention Drift: What Auto-Regressive Speculative Decoding Models Learn
Doğaç Eldenk ⋅ Payal Mohapatra ⋅ Yigitcan Comlek ⋅ Kaan Oktay ⋅ Hongyang Zhang ⋅ Stephen Xia
Speculative decoding accelerates LLM inference by drafting future tokens with a small model, but drafter models degrade sharply under template perturbation and long-context inputs. We identify a previously-unreported phenomenon we call \textbf{attention drift}: as the drafter generates successive tokens within a speculation chain, attention progressively moves from the prompt onto its own recently-generated tokens. We observe this across both \emph{EAGLE3} drafters and \emph{MTP heads}, suggesting drift is a property of drafter designs. We trace this to the un-normalized residual path between chain steps: the drafter's hidden state magnitude grows monotonically with chain depth, which exhibits dynamics consistent with additional pre-norm transformer layers stacked on the target rather than as a standalone autoregressive predictor. In order to limit the growth, we propose two architectural changes: Post-norm on the drafter hidden states and per-hidden-state RMSNorm after capturing target hidden states. Our interventions improve acceptance length over the current leading model, pre-norm EAGLE3, by up to $2\times$ under template perturbation, $1.18\times$ on long-context tasks, and $1.10\times$ on seven standard benchmarks spanning multi-turn chat, math, and coding. Our changes also allow shorter train-time-test depths to generalize over longer drafting sequences.
AttnDiff: Attention-based Differential Fingerprinting for Large Language Models
Haobo Zhang ⋅ Zhenhua Xu ⋅ Junxian Li ⋅ Shangfeng Sheng ⋅ Dezhang Kong ⋅ Meng Han
Protecting the intellectual property of open-weight large language models (LLMs) requires verifying whether a suspect model is derived from a victim model despite common laundering operations such as fine-tuning (including PPO/DPO), pruning/compression, and model merging. We propose AttnDiff, a data-efficient white-box framework that extracts fingerprints from models via intrinsic information-routing behavior. AttnDiff probes minimally edited prompt pairs that induce controlled semantic conflicts, captures differential attention patterns, summarizes them with compact spectral descriptors, and compares models using CKA. Across Llama-2/3 and Qwen2.5 (3B--14B) and additional open-source families, it yields high similarity for related derivatives while separating unrelated model families (e.g., >0.98 vs. <0.22 with M=60 probes). With 5--60 multi-domain probes, it supports practical provenance verification and accountability; our open-source implementation is available at https://anonymous.4open.science/r/AttnDiff-37A1/.
AudioAgentBench: Evaluating Multi-Turn Voice Agents on Real-World Tasks
Gardenia Liu ⋅ Yi-Hao Peng ⋅ Kamryn Ohly ⋅ Oliver Johansson ⋅ Hileamlak Yitayew ⋅ Jaden Zhang ⋅ Grace Li
Speech-to-speech (S2S) language models are increasingly deployed as voice agents that schedule appointments, take grocery orders, and plan events. Reliable deployment in such settings requires accurate listening, grounded tool use, state tracking, knowledge-base access, and recovery from user corrections across many turns. Existing voice agent benchmarks often score final database state or per-call tool accuracy, which can obscure the first failure point when an early mistake cascades through the rest of a conversation. We introduce AudioAgentBench, a fixed-trace benchmark suite with six multi-turn voice-agent tasks, 221 turns of pre-recorded audio, golden tool calls, tool schemas, knowledge bases, and matched TTS and human-recorded variants. The fixed audio traces make turn-level expectations reproducible and support a diagnostic scoring protocol that attributes failures across tool use, state tracking, ambiguity handling, knowledge grounding, and instruction following. An LLM judge applies the protocol to each turn by comparing model outputs against golden tool calls, task state, and task-specific knowledge bases, while using cross-turn realignment, conditional penalty absorption, and category-aware dimension gating to avoid double-counting cascading errors. We evaluate seven S2S models across 420 continuous session runs. The top model averages an 80.3\% pass rate, but the top four models have overlapping 95\% confidence intervals, and no model leads on more than two of six tasks. On a 408-run turn-level diagnostic subset with 14,907 judged turns, tool-use accuracy on turns where a tool call was expected peaks at only 61.4\% and drops to 32.3\% on a confusable-name scheduling task. Errors cascade by up to 3.15× in stateful tasks, and tool-use accuracy on a 30-turn grocery task falls by 40.5 percentage points from the first half to the last half. Scoring only turns where a tool call was expected reduces apparent tool-use accuracy by 22–44 percentage points across models, which reveals a larger gap between strong and weak systems than aggregate task success suggests. AudioAgentBench surfaces action, state, and recovery failures that final success metrics obscure, providing a finer-grained diagnostic benchmark for production-oriented voice agents.
A Unified Theoretical Framework for Task Recognition and Task Learning in In-Context Learning
Zhixuan Pan ⋅ Li Cao ⋅ Yiqi Dong ⋅ Chenyu Gan ⋅ Shaowen Wang ⋅ Jian Li
In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks from a few examples without parameter updates. Previous work and our empirical studies suggest two modes in ICL: Task Recognition, which applies pre-trained knowledge to familiar tasks, and Task Learning, which generalizes to novel mappings during inference. However, a unified theoretical account of these behaviors remains lacking. We propose a new Bayesian prediction framework that models pre-training data as a Pitman–Yor mixture over latent tasks, which captures the growing and heavy-tailed structure of natural language. By formulating an information-theoretic optimization problem, we derive the structure of the optimal ICL predictor under explicit capacity constraints. Our analysis reveals a phase transition in optimal capacity allocation: a water-filling strategy that prioritizes high-frequency tasks while assigning zero task-specific information to rare ones. This gives rise to two regimes: for frequent tasks, the model performs Task Recognition, yielding exponentially fast error decay with prompt length. For rare or unseen tasks, the model performs Task Learning by executing an implicit inference algorithm, resulting in a slower power-law convergence rate. Our results provide a principled explanation for the dual nature of ICL and establish a direct connection between pre-training data distributions, model capacity, and in-context generalization behavior.
AURA: An Autonomous Retouching Agent with Photographic Visual Thinking
Shuaizheng Liu ⋅ Fangzhou Han ⋅ Jiarong Liao ⋅ Yujing Sun ⋅ Lingchen Sun ⋅ Ruibin Li ⋅ Xinyu Wei ⋅ Jie Liang ⋅ Hui Zeng ⋅ Xindong Zhang ⋅ Lei Zhang
Professional image retouching is a sequential and reasoning-driven process. Photographers first analyze scene semantics, lighting structure, and depth cues to design a retouching plan, and then progressively apply global tone grading and mask-guided local corrections based on the evolving visual outcome. Recent MLLM-empowered retouching agents either predict global parameters or respond to user-provided retouching instructions. Although showing impressive results, they lack the capacity for autonomous sequential reasoning and fall short of fine-grained local control. We propose AURA, an AUtonomous Retouching Agent that emulates the workflow of expert photographers. Given an input photograph, AURA autonomously assesses the scene and proceeds progressively. At each step, it reasons about what to retouch, executes the adjustment, and feeds the result back as visual context for next-step reasoning, forming a fine-grained visual chain-of-thought that evolves with the image. To support AURA, we contribute the first expert-annotated long-horizon retouching trajectory dataset, where expert photographers edit the input image from scratch with deliberate artistic intent. Based on this dataset, we train AURA in two stages: supervised fine-tuning to learn the retouching trajectory, followed by GRPO-retouching, an agentic reinforcement learning stage that optimizes for perceptual quality and tool accuracy. Our experiments demonstrate that AURA consistently outperforms existing methods in both quantitative metrics and user studies, enabling one-click professional retouching and democratizing masterpiece previously accessible only to skilled experts.
Auto-Dreamer: Learning Offline Memory Consolidation for Language Agents
Chongrui Ye ⋅ Yuxiang Liu ⋅ Yu Wang ⋅ Haofei Yu ⋅ Yining Zhao ⋅ Ge Liu ⋅ Julian McAuley ⋅ Jiaxuan You
Language agents increasingly operate over streams of related tasks, yet existing memory systems struggle to convert accumulated experience into reusable knowledge. Retrieval-augmented and structured memory methods record per-session observations effectively, but often couple acquisition and consolidation into a single online process, leaving the agent without a global view across sessions to discover recurring patterns, abstract shared procedures, or prune redundant entries. Inspired by complementary learning systems theory, we propose Auto-Dreamer, a learned offline consolidator for language-agent memory. Auto-Dreamer decouples fast per-session memory acquisition from slow cross-session consolidation. Given a region of a typed memory bank, the consolidator performs bounded tool-use to search memory, inspect candidate entries, and trace them back to raw source trajectories, synthesizing provenance-grounded replacement memories that supersede the original region. We train Auto-Dreamer via GRPO, using end-to-end agent performance as the reward signal to learn how to consolidate memories acquired through fast online experience. Trained on ScienceWorld trajectories alone, Auto-Dreamer transfers without retraining to held-out ALFWorld and WebArena, improves task success over fixed, RL-trained, and prompted memory baselines in both continual-memory deployment and fixed-bank consolidation, and does so with an active memory bank an order of magnitude smaller than competitive baselines.
Causal inference is central to scientific discovery, yet choosing appropriate methods remains challenging due to the complexity of statistical methodology and real-world data. Inspired by the success of artificial intelligence in accelerating scientific software, we introduce an evolutionary framework that uses large language models to discover and iteratively refine causal methods. Across benchmarks, our estimators consistently outperform established baselines: our best estimator lay on the Pareto frontier of 58 human submissions for a recent community competition. We also extend the algorithm to achieve competitive results in settings with estimated rewards. Analysis of the evolutionary trajectories shows that agents progressively discover sophisticated strategies tailored to unrevealed data-generating mechanisms. Our findings suggest that language-model-guided evolution could be used in scientific settings with partially observed rewards such as causal inference.
Barycentric Guidance: Turning Foundation Image Editors into Continuous Affective Controllers
Harvey Mannering ⋅ Zhiwu Huang ⋅ Adam Prugel-Bennett
Emotions are expressed on a continuum and are naturally described in the valence-arousal (VA) space. Existing VA-based image editors are typically training-based and domain-specific, limiting them to VA regions seen during training and making them brittle across datasets and domains. Meanwhile, despite producing high-quality edits, state-of-the-art foundation image editors are not continuous affective controllers: non-linear emotion representations in text encoders makes prompt interpolation poorly calibrated, so accurate and continuous VA control remains unreliable. To fill this gap, we introduce Barycentric Guidance, a training-free framework that turns pretrained foundation image editors into continuous affective controllers over VA, requiring no finetuning, extra data, or model changes. Our method maps any target VA point to simplex-based guidance weights, enabling VA controllability in three aspects: (i) accurate targeting of VA states, (ii) broad coverage across the VA plane, and (iii) continuous control. Across three face datasets, we improve VA target accuracy by 25–29\%, expand VA-plane coverage by 28–50\%, increase expression diversity by 29–49\%, over the strongest baseline. Unlike prior domain-specific methods, our training-free approach exploits the foundation model’s inherent versatility, enabling generalization beyond human faces to animals, artworks, and complex scenes within a single framework. Code will be made publicly available.
BASTION: Budget-Aware Speculative Decoding with Tree-structured Block Diffusion Drafting
Soowon Oh ⋅ Nam Cao ⋅ Yujin Kim ⋅ Hojung Jung ⋅ Huzama Ahmad ⋅ Sangmin Bae ⋅ Se-Young Yun
Block-diffusion drafters have recently shown strong potential for speculative decoding by predicting multiple future-token distributions in a single forward pass. To exploit these parallel predictions, tree-based verification can check multiple plausible continuations instead of committing to a single greedy draft path, often increasing the number of accepted tokens per decoding cycle. However, existing tree-based approaches typically rely on fixed tree topologies or static verification budgets, even though the best budget depends on the local draft distribution, context length, target model, and hardware runtime. We propose a system-aware adaptive tree framework for block-diffusion speculative decoding. At each decoding cycle, our method builds a tree toward high-probability draft paths and selects the verification budget that maximizes an online speedup estimate. This estimate combines a drafter-based acceptance surrogate with an analytical verifier-latency model calibrated from observed runtimes. The framework requires no additional training of the drafter or target model and preserves the target model's decoding rule. Across target models, benchmarks, decoding temperatures, and GPU platforms, our method achieves up to a $6.93\times$ wall-clock speedup over standard autoregressive decoding. This corresponds to approximately 40% higher speedup than Dflash, the strongest existing block-diffusion drafter, and is obtained without per-setting budget tuning.
BA-T: An Iterative Transformer for Two-View Bundle Adjustment
Ganlin Zhang ⋅ Weirong Chen ⋅ Daniel Cremers ⋅ Xi Wang
Feed-forward models for 3D reconstruction have achieved strong performance using deep cross-view attention to exchange information across images. However, these approaches often depend on heavy decoder stacks and lack a structured mechanism for geometry refinement, resulting in poor multi-view consistency. We address this by drawing inspiration from classical bundle adjustment (BA), which can be viewed as an iterative information propagation process between poses and local geometry. Inspired by BA, we propose BA-T, an iterative Transformer that implements BA-style structured updates as a repeatable layer in implicit token space. Instead of relying on deep attention stacks, BA-T refines predictions based on latent residual by a single lightweight layer. Experiments demonstrate that BA-T progressively improves pose and reconstruction accuracy across iterations, achieves stronger cross-view consistency than conventional decoders, and matches or surpasses substantially larger models while using only 16\% of their decoder parameters. BA-T provides a compact, efficient, and structural alternative to depth-heavy attention, enabling accurate 3D reconstruction within a lightweight architecture. The code will be made publicly.
BayesAT: Bayes-Guided Progressive Distillation for Semi-Supervised Adversarial Training
Yimo Guo ⋅ Lilin Zhang ⋅ Li Penglin ⋅ Jinhui Hao ⋅ Xianggen Liu
Existing semi-supervised adversarial training (SSAT) methods typically employ a teacher-student framework where a teacher model provides supervisory signals for unlabeled data. However, they rely on static supervision---either hard pseudo-labels or soft labels with fixed temperature---which maintains constant learning difficulty throughout training. This rigidity fails to accommodate the student's evolving capability, especially when the teacher's performance plateaus. To address this, we reframe SSAT as knowledge distillation from a Bayes teacher and formalize a stage-dependent bias-variance tradeoff. This analysis reveals the existence of an optimal temperature: low temperatures reduce bias, while higher temperatures control the variance term in the robust generalization bound. Guided by this insight, we propose BayesAT, a Bayes-guided progressive distillation framework that jointly employs a low-to-high temperature warming schedule and confidence-adaptive sample reweighting. BayesAT allows the student to first learn from sharp, high-confidence supervision to establish reliable decision boundaries, then progressively from smoother distributions that encourage exploration of inter-class relationships. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet-200 demonstrate that BayesAT consistently improves robust accuracy while maintaining natural accuracy over existing SSAT methods, highlighting the importance of dynamic, theory-driven supervision for effective knowledge transfer in semi-supervised adversarial training.
Bayesian Causal Experimental Design for CATE Estimation under Noncompliance
Erdun Gao ⋅ Yuanyuan Wang ⋅ Liang Zhang ⋅ Yuhang Liu ⋅ Haoxuan Li ⋅ Mingming Gong ⋅ Dino Sejdinovic
Causal experimental design often assumes direct control over treatment assignment, but many experiments can only assign encouragements that affect treatment receipt indirectly. We formulate this as adaptive instrumental variable (IV) design under noncompliance: the experimenter chooses unit--instrument queries, observes stochastic treatment realizations and outcomes, and seeks the treatment-level conditional average treatment effect (CATE). This creates an acquisition mismatch: uncertainty in prospective feedback does not necessarily correspond to uncertainty in the target causal estimand, so outcome-predictive or marginal variance-based criteria may select queries weakly informative for CATE. We introduce instrumental variable information gain (IVIG), a Bayesian acquisition principle that scores queries by the expected information their joint treatment-realization and outcome feedback provides about CATE values over a target population. Under the working Bayesian posterior, IVIG is Bayes-risk aligned under logarithmic loss; we also characterize the residual information discarded by outcome-only acquisition and connect greedy IVIG-style acquisition to target posterior-variance reduction under a fixed-covariance analysis. To approximate IVIG, we use an empirical-Bayes Gaussian Process IV posterior approximation with first-stage treatment-realization modeling and treatment-realization-specific fantasy updates. Experiments on synthetic and two semi-synthetic IV benchmarks show improved CATE sample efficiency over adaptive-design baselines; a real-label diagnostic shows improved recovery of full-data CATE estimates from observed IV data.
BeamWitness: Rooted Subgraph Beam Search for Selective Graph Representation Learning
Lin Du ⋅ Lu Bai ⋅ Lixin Cui ⋅ Ming Li ⋅ Hangyuan Du ⋅ Bo Jiang
Standard graph neural networks (GNNs) are built on message passing and neighborhood aggregation, following the Weisfeiler--Lehman refinement paradigm. Despite their wide adoption, these models tend to compress and mix neighbor information from the very first layer. This early mixing obscures the critical subgraphs that drive the prediction. Attention-based and subgraph-based methods partially address this by reweighting nodes or enriching local representations, but both families remain passive as they usually rely on soft weighting or predefined subgraph templates. We propose BeamWitness, a selective graph representation learning framework that treats subgraph identification as an active search problem. Starting from a root node, a policy trained with task rewards expands a connected subgraph one node at a time or terminates early, and beam search retains a small set of candidate subgraphs as witnesses for prediction. For node-level tasks, the target node serves as the root, while for graph-level tasks, BeamWitness launches searches from multiple root nodes and aggregates the selected witnesses into a graph-level representation. On controlled motif benchmarks with known ground-truth subgraphs, BeamWitness achieves near-perfect accuracy and selects witnesses that align with the true motifs. On real-world node and graph classification benchmarks, BeamWitness is competitive with standard GNNs and related methods while providing an explicit, interpretable subgraph-based account of each prediction.
Behaving Better, Thinking Worse: Sycophancy Across Post-Training Stages
Sonnet Xu ⋅ Kritika Singh ⋅ Sheharbano Jafry ⋅ Roxana Daneshjou ⋅ Sanmi Koyejo
We trace factual sycophancy across three post-training pipelines (OLMo 3 7B Think and Instruct, Llama 3.1 8B Instruct, Tulu 3) on AMPS and MedQuAD with a dual-track evaluation: a GPT-4o-judged generative track and a log-probability track reporting $\Delta$LogOdds separately under \emph{in-context} rebuttals and \emph{preemptive} wrong-answer assertions. Decomposing by challenge type $\times$ context reveals three patterns. (1) \textbf{Challenge type:} the sycophantic shift is alternative-conditioned, so simple pushback with no asserted answer produces no shift, while ethos/justification/citation challenges produce substantial positive shifts. (2) \textbf{Context:} every base model's sycophantic shift concentrates preemptively, with weak or defensive in-context response. On computational, IC diverges into four pipeline-specific endpoints. On medical, the same four pipelines converge with no defensive IC developing on any recipe. (3) \textbf{Behavioral vs log-probability:} post-training produces a \emph{surface-predictive dissociation}: matched-subset behavioral flip rate drops significantly on identical items while preemptive $\Delta$LogOdds on those same items grows.
Benchmarking Membership Privacy Risks in Preference-Based LLM Post-Training
Lorenzo Rossi ⋅ Kaif Shaikh ⋅ Franziska Boenisch ⋅ Adam Dziedzic
Modern language models are commonly adapted after pretraining to follow instructions, align with user preferences, and improve deployment behavior. Such post-training often relies on preference data from users, annotators, or model interactions, which may contain sensitive prompts, private responses, proprietary tasks, or confidential judgments. Understanding the privacy implications of this data is therefore essential. Membership inference attacks (MIAs), which test whether a record was used for training, are the de facto standard for empirical privacy auditing in machine learning. However, preference-based post-training changes the unit of membership: a training record contains a prompt, a preferred response, and a dispreferred response, rather than a single input-output sequence. Audits that ignore this structure can under-report leakage. We therefore introduce a benchmark protocol that adapts strong reference-model MIAs to complete preference records. Our protocol asks whether membership is exposed by either response alone, by the model’s preference between them, or by the two responses jointly. We find that auditing the response pair can reveal membership evidence missed when the record is reduced to a single score. Across three preference datasets, two model families, and seven post-training objectives spanning three post-training families, measured leakage depends strongly on both the audited record statistic and the training objective, with imbalanced risk between preferred and dispreferred responses. We further evaluate the impact of parameter-efficient fine-tuning (PEFT) and differential privacy (DP), finding widely varying privacy-utility trade-offs. Overall, preference-based post-training leaks in ways that standard single-response audits can miss, motivating formal privacy methods that protect complete preference records while preserving post-training utility.
Benign Reinforcement Learning Can Amplify Latent Backdoors
Chen Lai ⋅ Javier Rando ⋅ Nicholas Carlini ⋅ Yiming Zhang
Reinforcement learning (RL) is now standard for post-training large language models. The same reward optimization that elicits useful capabilities, however, can also reinforce latent backdoors planted earlier in the pipeline: attack success rates under 0.2% after SFT rise to 40--98% on training-distribution inputs after RL, and generalize to 13--30\% on held-out evaluation, with no modification to the RL data, reward function, or training loop. We study this phenomenon in the context of agent models, where SFT poisoning teaches the victim model (Qwen3-8B and a small frontier model) to call an attacker-controlled oracle. Because the oracle can be set up to always return correct answers, calling it earns higher reward than the model's own attempts, and RL reinforces the behavior. We further show that the oracle's RL-time responses can instill persistent biases, such as brand preferences, that survive into deployment even when no tool calls happen. Such patterns appear difficult to detect with current tools: traditional guard models do not flag this mechanism, and LLM-based auditors remain unreliable, achieving only 4.5\% precision even after iterative prompt refinement. Our findings point to the value of post-RL safety evaluation, in particular scrutinizing tool-call patterns such as unnecessary external invocations.
Better Source, Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching
Junwan Kim ⋅ Jiho Park ⋅ Seonghu Jeon ⋅ Seungryong Kim
Flow matching has recently emerged as a promising alternative to diffusion-based generative models, particularly for text-to-image generation. Although flow matching places no restriction on the source distribution, most existing systems still inherit a standard Gaussian from diffusion models, and the source is rarely treated as an optimization target at this scale. Recent works have begun to revisit this choice through condition-dependent or learned sources, yet evidence that such designs are effective in modern text-to-image systems—with high-dimensional latents and tightly integrated conditioning—remains limited. In this work, we study condition-dependent source distributions for flow matching along three axes: _why_ source learning helps, through the lens of the intrinsic variance term in the flow-matching objective; _how_ to make it work in modern text-to-image systems, where variance-only regularization and directional source—target alignment are critical for stable end-to-end training; and _when_ it is most beneficial, by connecting source design to recent representational advances in generative modeling and identifying target representation regimes in which learning the source yields the largest gains. Extensive experiments across multiple text-to-image benchmarks, backbones, and scales demonstrate that principled source design yields consistent and robust improvements—including up to $\mathbf{3.01\times}$ faster convergence in FID and $\mathbf{2.48\times}$ in CLIP score—and outperforms representative prior conditioned-source and condition-aware coupling methods.
Beyond Flat Frames: Hierarchical Graph Reasoning for Long Video Understanding
Xinyue Liu ⋅ Jiayang Sun ⋅ Zhe Jing ⋅ Jie Cao ⋅ Ran He ⋅ Huaibo Huang
Recent studies have shown that leveraging the reasoning and analytical capabilities of large language models (LLMs) for long video understanding has become a promising approach. However, these methods are constrained by a structural limitation: representing video as a flat frame sequence makes it difficult to model the temporal structure and logical dependencies of events, which in turn hampers long-range reasoning and risks overlooking important information. To address this limitation, we propose LVGraph, a framework that models video content using a coarse-to-fine hierarchical graph. LVGraph constructs a semantic hierarchy where high-level graphs characterize the global event structures within the video, while low-level graphs capture fine-grained, query-relevant information. Instead of linear scanning, reasoning is performed via a query-aware traversal of this graph, adaptively identifying and retrieving the most salient keyframes by navigating the nodes and edges most relevant to the query. Finally, we leverage a Vision Language Model (VLM) to augment the keyframe information, and subsequently answer the question by reasoning over the graph and the enriched keyframe captions. Extensive experiments on long video understanding benchmarks confirm the effectiveness of our method. LVGraph significantly outperforms existing LLM-based state-of-the-art methods on EgoSchema, NExT-QA, and long-duration segments (averaging 44 minutes) of the Video-MME benchmarks.
Beyond FLOPs: Train-Full, Deploy-Partial Multi-Exit Inference via Selective Lightweight IC Ensemble
Bitchan Eom ⋅ Eunchan Kim
Early-exit networks promise inference savings by terminating computation at intermediate classifiers, but FLOPs-based gains rarely translate into on-device latency reduction under static-graph compilation. On NVIDIA Jetson Orin Nano, the per-sample exit policy of Shallow-Deep Networks (SDN) runs approximately 2× slower than vanilla ResNet-56 and misses the 60 fps deadline despite using only 40% of the FLOPs, because dynamic branching precludes layer fusion. We propose the Selective Lightweight IC Ensemble Network (SLIENet), a train-full, deploy-partial framework compiled as a single static FP16 engine. SLIENet trains the full backbone with light self-distillation, using the final classifier as a stop-gradient teacher to preserve error diversity across internal classifiers (ICs). At deployment, a sub-second calibration search selects an IC subset for softmax averaging, while overthinking late blocks are truncated at depth k. On CIFAR-100 across ResNet-56, VGG-16, and MobileNetV1, SLIENet outperforms SDN-based variants and remains competitive with ZTW. On Jetson Orin Nano, the recommended SLIE5 configuration improves accuracy by +2.46 pp while reducing p50 latency by 14% and energy by 15% over vanilla, with zero deadline misses across 10,000 test samples. Thus, multi-exit deployment should be a pre-compiled deterministic inference plan, not a sample-wise routing policy.
Beyond Linear Decoders: Dynamic Expert-Coupled Optimal Decoding for Time Series Forecasting
Binwu Wang ⋅ Zhipeng Liu ⋅ Zhengyang Zhou ⋅ Pengkun Wang ⋅ Yang Wang
Current multivariate time series forecasting methods mainly rely on static linear decoders, but these often suffer from severe representational bottlenecks. In this paper, we propose a novel architecture called DecodeTS (\textbf{\underline{D}}ynamic \textbf{\underline{E}}xpert-\textbf{\underline{C}}oupled \textbf{\underline{O}}ptimal \textbf{\underline{DE}}coder for Time Series Forecasting), which replaces the conventional static prediction head with a heterogeneous expert library. DecoderTS adopts a divide-and-conquer strategy to disentangle complex temporal dynamics, such as long-term trends and abrupt changes. Crucially, we introduce an optimal transport (OT)-based dynamic routing mechanism that can adaptively assign customized combinations of experts to different variables. By imposing OT marginal constraints, DecoderTS is theoretically proven to achieve collaborative load balancing among experts and effectively eliminate representation collapse. Extensive experiments on more than \textbf{15} datasets and about \textbf{30} baselines demonstrate that DecoderTS achieves state-of-the-art forecasting accuracy, which delivers a \textbf{17.06\%} performance gain while attaining up to about \textbf{41×} inference time speedup and up to \textbf{3.61×} reduction in memory footprint.
Beyond Self-Play and Scale: A Behavior Benchmark for Generalization in Autonomous Driving
Aron Distelzweig ⋅ Faris Janjos ⋅ Andreas Look ⋅ Anna Rothenhäusler ⋅ Daniel Jost ⋅ Oliver Scheel ⋅ Raghu Rajan ⋅ Daphne Cornelisse ⋅ Eugene Vinitsky ⋅ Joschka Boedecker
Recent Autonomous Driving (AD) works such as GigaFlow and PufferDrive have unlocked Reinforcement Learning (RL) at scale as a training strategy for driving policies. Yet such policies remain disconnected from established benchmarks, leaving the performance of large-scale RL for driving on standardized evaluations unknown. We present BehaviorBench -- a comprehensive test suite that closes this gap along three axes: Evaluation, Complexity, and Behavior Diversity. In terms of Evaluation, we provide an interface connecting PufferDrive to nuPlan, which, for the first time, enables policies trained via RL at scale to be evaluated on an established planning benchmark for autonomous driving. Complementarily, we offer an evaluation framework that allows planners to be benchmarked directly inside the PufferDrive simulation, at a fraction of the time. Regarding Complexity, we observe that today's standardized benchmarks are so simple that near-perfect scores are achievable by straight lane following with collision checking. We extract a meaningful, interaction-rich split from the Waymo Open Motion Dataset (WOMD) on which strong performance is impossible without multi-agent reasoning. Lastly, we address Behavior Diversity. Existing benchmarks commonly evaluate planners against a single rule-based traffic model, the Intelligent Driver Model (IDM). We provide a diverse suite of interactive traffic agents to stress-test policies under heterogeneous behaviors, beyond just using IDM. Overall, our benchmarking analysis uncovers the following insight: despite learning interactive behaviors in an emergent manner, policies trained via pure self-play under standard reward functions overfit to their training opponents and fail to generalize to other traffic agent behaviors. Building on this observation, we propose a hybrid planner that combines a PPO policy with a rule-based planner, providing a baseline for our new benchmark. Code is available at https://anonymous.4open.science/r/behavior-bench-707C
Beyond Spatial and Temporal Priors: A Generalizable Approach for Dense Correspondence Matching
Luping Liu ⋅ Bingyi Kang ⋅ Yifan Wang ⋅ Dong Xu
Dense correspondence matching has historically been bounded by simplifying spatio-temporal priors, such as smooth motion and rigid geometry. While effective for classical tasks, these foundations collapse at the frontier of Vision-Language-guided Image Editing and Generation (VL-IEG), where transformations yield correspondences that are perceptually obvious yet physically discontinuous. To transcend spatio-temporal priors, we draw inspiration from human cognition: learning a generalized model of identity-preserving visual consistency rather than relying on physical constraints. To achieve this, we introduce FreeMatching, a unified framework that extracts identity-preserving knowledge from foundation models and diverse datasets. The model's capability is forged through a two-stage training paradigm: supervised pre-training on diverse annotated datasets, followed by weakly-supervised refinement on data without dense annotations. Experimentally, FreeMatching not only achieves performance competitive with SOTA methods on classical benchmarks but also establishes the first strong baseline for the challenging VL-IEG task. Furthermore, we demonstrate its utility as a quantitative metric for evaluating identity-preserving consistency in generative editing, aligning closely with human judgment.
Beyond the Training Distribution: Evaluating Predictions Under Distribution Shift and Selection Bias
Annie Ulichney ⋅ Amanda L Coston
Understanding how a prediction model will perform in a new environment before deployment is essential to preventing harm when algorithms inform decision-making. Two common sources of model performance degradation are (i) covariate shift, where the target covariate distribution differs from the source, and (ii) selective labels, where the observability of outcomes depends on historical decisions. We study pre-deployment model evaluation under the joint presence of covariate shift and labeling of outcomes selectively based on observed features. In particular, we present a double machine learning procedure for estimating the target risk of an arbitrary black-box prediction model under a general loss function. We show identification of this estimand under standard assumptions and derive a bias-corrected estimator based on the influence function of the target risk. Finally, we evaluate our estimator through experiments using the eICU electronic health records database, showing that it tracks the true target risk more accurately than methods that address either selective labels or covariate shift alone, as well as baselines that combine standard plug-in approaches.
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
Weiming Chen ⋅ Xitong Ling ⋅ Zhenyang Cai ⋅ Xidong Wang ⋅ Jiawen Li ⋅ Tian Guan ⋅ Benyou Wang ⋅ Yonghong He
Cell-level dense prediction is central to computational pathology, but remains challenging due to fine-grained histological structures, strong domain shifts, and costly dense annotations. Existing ViT-based pathology foundation models rely on patch tokenization, which can disrupt spatial continuity and weaken local morphological details needed for cell-level prediction. To address this, we propose Masked-Diffusion Convolutional Foundation Models, termed ConvNeXt Masked-Diffusion (CMD), a self-supervised convolutional generative pretraining framework for dense pathology representation learning. CMD uses a fully convolutional ConvNeXt-UNet backbone, performs masked-diffusion pretraining in pixel space, and incorporates frozen pathology foundation model features through adaptive normalization. Experimental results demonstrate that CMD consistently outperforms existing ViT-based pathology foundation models and even surpasses state-of-the-art end-to-end segmentation methods while fine-tuning only a small number of task-specific parameters across multiple pathology dense prediction tasks. The advantage is particularly pronounced under limited annotation settings, where CMD exhibits stronger robustness and generalization ability. Our findings suggest that purely convolutional architectures can also serve as competitive pathology foundation models for cell-level dense prediction, achieving leading performance within the current ViT-dominated paradigm and providing a scalable, high-performance solution that better preserves histological structural priors for fine-grained pathology understanding. Project page: https://anonymous.4open.science/r/ConvNeXt-Masked-Diffusion-5BC6.
B[FM]$^2$: Brain Foundation Model via Flow Matching with SplitUNet
Jaedong Hwang ⋅ Kathleen Zhang ⋅ David Dai ⋅ Konstantinos Kontras ⋅ Maarten De Vos ⋅ Ila Fiete ⋅ Paul Liang
EEG foundation models can learn generalizable representations from large-scale EEG corpora to enable single-backbone transfer across diverse clinical and brain-computer interface tasks. Existing models typically discretize the continuous multi-channel EEG waveform into patches or codebook tokens and train a transformer with masked self-supervision. Recognizing that this discretization fragments continuous brain rhythms and obscures fine-grained temporal dynamics, we present B[FM]$^2$ (Brain Foundation Model via Flow Matching), which drops discretization and pretrains directly on the raw signal using continuous-time flow matching without patches, tokenization, or masking. However, multi-channel EEG signals pose an architectural challenge for flow matching: time is densely sampled and highly autocorrelated (thousands of timepoints), while the electrode axis is short (tens of channels) at distinct scalp positions. To address this time-electrode asymmetry, we introduce SplitUNet, a velocity network that factorizes each block into separate 1D temporal and 1D electrode convolutions and downsamples only along time, preserving electrode topology throughout the hierarchy. B[FM]$^2$ sets a new state of the art on $7$ of $9$ standard downstream EEG classification tasks, using a pretraining budget of only $36{,}895$ segments ($\approx 307$\,h), a fraction ($\approx 3.3$\%) of that required by existing EEG foundation models. It also produces synthetic EEGs that two board-certified neurologists cannot distinguish from real EEGs (Cohen's $\kappa = -0.096$).
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
Alessandro Favero ⋅ Antonio Sclocchi ⋅ Matthieu Wyart
Diffusion probabilistic models have become a cornerstone of modern generative AI, yet the mechanisms underlying their generalization remain poorly understood. In fact, if these models were perfectly minimizing their training loss, they would just generate data belonging to their training set, i.e., memorize, as empirically found in the overparameterized regime. We revisit this view by showing that, in highly overparameterized diffusion models, generalization in natural data domains is progressively achieved during training before the onset of memorization. Our results, ranging from image to language diffusion models, systematically support the empirical law that memorization time is proportional to the dataset size, consistent with a kernel-regression bound on the time required to fit the empirical score at low noise. Generalization vs. memorization is then best understood as a competition between time scales. We show that this phenomenology is recovered in diffusion models learning a simple probabilistic context-free grammar with random rules, where generalization corresponds to the hierarchical acquisition of deeper grammar rules as training time grows, and the generalization cost of early stopping can be characterized. We summarize these results in a phase diagram. Overall, our results support that a principled early-stopping criterion – scaling with dataset size – can effectively optimize generalization while avoiding memorization, with direct implications for hyperparameter transfer and privacy-sensitive applications.
BiLoCo: Binary Low-Rank Corrections for LLM FP4 Decode
David Jin ⋅ Beshr IslamBouli ⋅ Tarushii Goel ⋅ Han Guo ⋅ Yoon Kim
Post-training quantization (PTQ) to FP4 has emerged as a key technique for reducing large language model inference cost, especially on Blackwell GPUs which offer native FP4 tensor core support. Direct PTQ to FP4, however, still leaves a noticeable accuracy gap. Recent work on SVDQuant mitigates the errors introduced through FP4 quantization for diffusion models by adding high-precision low-rank corrections obtained via SVD. However, this high-precision low-rank correction requires many bits per rank, and moreover competes on tensor core utilization with the main FP4 GEMM. We introduce BiLoCo, a correction whose left and right factors are stored as binary ${\{\pm 1\}}$ sign vectors rather than as BF16 numbers. The signed format uses $16{\times}$ less storage per rank-one term, and since the correction now consists of additions/subtractions, this can be run on CUDA cores concurrently with the FP4 tensor-core GEMM. On 4--32B LLMs, BiLoCo matches or improves memory-matched SVDQuant for PTQ. We implement BiLoCo on B200 with a correctness-preserving CUDA schedule that runs end-to-end decode at $1.13$--$1.15\times$ FP4-only, with the residual gap coming from the final output combine rather than from one-bit arithmetic.
Bits Beat Tokens: A Regret Rate Distortion Theory for Large Language Model Agents
Wanrong Yang ⋅ Lingfang Li ⋅ Dominik Wojtczak ⋅ Yihang Zhou ⋅ Danli Shi ⋅ Yalin Zheng
A language model agent acts on a context assembled by retrieval, summarisation, prompt templates, and memory. How many decision-relevant bits must this context carry before the agent can act with low regret? We give an answer that does not depend on the model, decoder, or scaffold. Modelling the agent as a Markov chain H→C→A and treating Bayes regret (the value gap between the chosen and the optimal action) as the distortion measure, we define a regret rate-distortion (RRD) function Rℓ(ε) whose converse is a one-line consequence of the data-processing inequality (DPI). A context that carries fewer than Rℓ(ε) bits cannot drive expected regret below ε, regardless of decoder, scaffold, scale, or chain-of-thought. For finite M-ary decisions under zero-one loss the bound reduces to the Fano floor in closed form, requiring ≈3.14 bits at (M=16, ε=0.1) and ≈6.73 bits at (M=256, ε=0.1). We instantiate the theory on 441 controlled single-step tool-selection cells, a balanced 16-way classification proxy spanning three open-source large language models (Qwen3-32B, Gemma3-27B, GLM-4-32B) and three public benchmarks (API-BANK, τ-bench, ToolBench). Because the gold tool index X=f(H) is a deterministic function of the history, the action-side rate IMM(X; A) lower-bounds the context capacity by data processing, so an action-side test of the converse is strictly tighter than a context-side one. Every cell respects the predicted floor (median margin +1.07 bits). Across 629 matched-rate pairs, 97.5% share overlapping regret confidence intervals (CIs), and a chain-of-thought supplement slides cells along the frontier rather than across it. The same data exposes a tokens-versus-bits gap reaching ∼1.7×10⁶:1, with the model extracting fewer than 0.05 bits about the gold action on a 32,768-token ToolBench prompt. In the single-step regime, the framework supplies a falsifiable, model-agnostic stopping criterion at IMM=Rℓ(ε) and reframes context engineering as an information-allocation problem.
BLARM: Animating 3D Objects from Video via Blending LAtent Rigid Motion Primitives
Pradyumn Goyal ⋅ Yizhak Ben-Shabat ⋅ Hsueh-Ti Derek Liu ⋅ Haomiao Jiang ⋅ Snehasish Mukherjee ⋅ Kyle Spence ⋅ Mark Stauber ⋅ Evangelos Kalogerakis ⋅ Yunze Zeng
We introduce BLARM, a feed-forward method for video-driven 3D mesh animation. Given a monocular video and a static object mesh, BLARM predicts a temporally coherent animated mesh whose motion follows the video. Rather than relying on explicit rigs or directly regressing high-dimensional vertex motion, we represent animation using a compact set of learned, time-varying rigid motion components and time-invariant vertex-to-component skinning weights. This yields a low-dimensional deformation space without requiring skeletons, cages, skinning weights, or rig annotations. Our architecture conditions geometry-derived deformation latents on video features through factorized spatial-temporal attention, then decodes rigid transformations blended by predicted skinning weights. Trained with trajectory reconstruction, entropy regularization, and motion-aware contrastive learning, BLARM produces accurate and temporally stable animations while recovering compact, interpretable motion structure from monocular video.
Blind-Window Forecasting: Real-Time Benchmarking and Multimodal Reconstruction for Tropical Cyclones
Zhaoran Feng ⋅ Xuanhong Chen ⋅ Zengbing Chen ⋅ Shengjun Wu ⋅ Bin He ⋅ Kairui Feng
Tropical cyclone forecasting models are often trained and evaluated with reanalysis inputs. Reanalysis provides physically rich environmental context, but its most recent fields may be unavailable at forecast issuance. This can introduce a hindsight-style advantage and lead to optimistic estimates of real-time forecast skill. Satellite observations, in contrast, are available closer to real time but provide incomplete and noisy views of the storm system. We formulate this data-availability mismatch as forecasting through a reanalysis blind window. To study this setting, we introduce a real-time-faithful benchmark spanning 1980--2023 that combines lagged reanalysis with visible, infrared, water-vapor, and passive microwave satellite imagery. We then develop RECAST-TC, a multimodal reconstruction framework that fuses time-lagged environmental histories with real-time satellite observations to estimate a forecast-sufficient latent storm state from delayed, partial, and noisy inputs. We use idealized analysis-available forecasting and delayed-reanalysis-only forecasting as reference regimes. RECAST-TC consistently improves track and intensity forecasts over delayed-reanalysis baselines and moves noticeably closer to the idealized reference. These results highlight data availability and observation timing as central design variables for realistic evaluation and deployment-oriented AI forecasting of tropical cyclones.
Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models
Yan Jiang ⋅ Ruihong Qiu ⋅ Zi Huang
Recently, reinforcement learning (RL) has been widely applied during post-training for diffusion large language models (dLLMs) to enhance reasoning with block-wise semi-autoregressive generation. Block size has therefore become a vital factor in dLLMs, since it determines the parallel decoding granularity and affects the rollout trajectories during RL optimisation, e.g., GRPO. Instead of investigating the effect of block size during inference on individual domains, this paper studies block size from a domain conflict perspective for dLLM RL post-training in multi-domain scenarios. The main contributions are: (1) a formulation of domain block size conflict in multi-domain RL for dLLMs, which will largely affect the post-training effectiveness for rollout-based RL methods; (2) a novel dataset, \textbf{Block-R1-41K} is constructed with a best-improved training block size for each sample, which also induces a Block Size Conflict Score to quantitatively measure the domain conflict; (3) a new benchmark, Block-R1, for flexible RL post-training for dLLMs in both single and cross domain; and (4) a simple yet powerful cross-domain post-training method with sample-level best-improved training block sizes. Extensive experiments on 13 distinct datasets, 7 latest RL algorithms, and various different dLLM backbones are covered in Block-R1. The benchmark is open-sourced at https://anonymous.4open.science/r/Block-R1-2026, with the dataset released at https://huggingface.co/datasets/dLLM-R1/Block-R1-41K.
Vector quantization is a fundamental primitive for scalable machine learning systems, enabling memory-efficient storage, fast retrieval, and compressed inference. Recent rotation-based quantizers such as EDEN, RabitQ, and TurboQuant have introduced strong guarantees and empirical performance, but the surrounding comparisons have been difficult to interpret because they rely on different distortion criteria, probability regimes, and implementation assumptions. As our first contribution, we provide a unified theoretical comparison of these methods and show that their relative advantages are criterion-dependent rather than absolute: TurboQuant is favorable for MSE distortion, EDEN is effective for expected inner-product distortion, and RabitQ provides strong high-probability control. This comparison clarifies the design principles behind each advantage and shows that no existing method uniformly dominates across the relevant measures. As our second contribution, we introduce Block-Sphere Quantization (BlockQuant), a new rotation-based block quantization algorithm designed around the spherical geometry of randomly rotated vectors. Unlike coordinate-wise quantizers, BlockQuant quantizes blocks on the sphere, preserving the geometry of rotated embeddings more faithfully. We prove that this block-spherical design improves expected distortion performance, including both reconstruction MSE and expected inner-product distortion. Experiments on real embedding data support these theoretical improvements.
BrainEM: A Large-Scale and Diverse Benchmark for EM Neuron Segmentation in Connectomics
Zhenghua Li ⋅ Lanyue Zhang ⋅ Zekang Yang ⋅ Kai Li ⋅ Yuwei Peng ⋅ Song-Hai Shi ⋅ Xiaolin Hu
Connectomics requires dense segmentation of large-scale electron microscopy (EM) volumes using models trained on small labeled blocks. In practice, a connectomics lab facing a new raw volume has two options: annotate a small block and train on it, or skip annotation and train on the union of available public labeled datasets. We introduce BrainEM, a benchmark that formalizes these options as two standard tracks: the Same-Source Track (Annotate-and-Train) and the Cross-Source Track (Reuse-and-Train). BrainEM brings together 16 EM datasets, including one in-house mouse dataset that we contribute, spanning multiple species, imaging modalities, and resolutions. We further add large-scale test volumes that bring the evaluation closer to realistic practice, and four diverse test targets in the Cross-Source Track. The data, post-processing, and evaluation metrics are fixed across all methods, and the framework is fully model-data decoupled so that new methods can be added with minimal effort. We conduct a systematic evaluation of six representative baselines covering the main technical directions in EM neuron segmentation, making BrainEM the first systematic evaluation at large scale and across diverse cross-source targets. The results differ from prior small-scale evaluations: rankings on the Same-Source Track no longer match those reported on small CREMI splits, and methods diverge sharply on the Cross-Source Track. Both tracks still leave clear room for improvement. We release BrainEM as a public platform to support unified comparison and faster method development in connectomics. All datasets and code are available at https://github.com/kwinderic/BrainEM
Breaking the $\sqrt{d}$ Communication Barrier in Federated Sampling with Adaptive Hamiltonian Monte Carlo
Jiajun Liang ⋅ Linxuan Wang ⋅ Guang Lin ⋅ Qifan Song
This work proposes the Adaptive Federated Hamiltonian Monte Carlo (AFHMC) for Bayesian federated learning (FL) tasks. The prior work on HMC applications for FL achieves $O(\sqrt{d}/\epsilon)$ communication cost under log-convex distributions, already demonstrating its advantage over the Langevin dynamic-based counterpart. Based on a refined analysis, this paper further reveals the benefit of HMC for the FL regime by incorporating a control mechanism for local node shifts. We prove that AFHMC at most requires $O((d/\epsilon^2)^{1/3})$ communication cost up to a logarithmic term, which is significantly better than the existing results. Moreover, the proposed method adapts to client heterogeneity. Under low-heterogeneity settings (i.e., the target sampling precision $\epsilon^2$ is higher than the heterogeneity level), the communication rate of AFHMC can be as good as $O(\log(d/\epsilon^2))$, resembling the communication costs for FL optimization tools. To the best of our knowledge, this provides the first logarithmic communication complexity result for federated sampling. Our code is available at https://anonymous.4open.science/r/Adaptive-Federated-HMC.
Breaking the Synthesis Barrier for AI-Designed DNA Libraries
Scott Sussex ⋅ Ema Borevković ⋅ Frederieke Lohmann ⋅ Ningning Chen ⋅ Elena Luethi ⋅ Sai Reddy ⋅ Andreas Krause
Designing DNA libraries is a key challenge from drug design to protein engineering and synthetic biology. Modern generative models offer opportunities to navigate the design space and propose specific sequences predicted to be effective in-silico. Designing deterministic libraries of specific sequences is however limited by the cost of DNA synthesis -- the synthesis barrier. In contrast, high-throughput multiplexed screening can measure the function of billions of biological sequences in parallel. Harnessing this technology requires the design of randomized libraries with specific design constraints to achieve low synthesis costs. In practice, such stochastic libraries are often chosen heuristically, sacrificing control for scale. Is there a way to bridge AI-based in-silico sequence design with high-throughput experimentation? In this work, we introduce Policy Gradients for Library Design (PGLD). PGLD uses a synthesis-aware parametrization of stochastic DNA libraries and optimizes them against a specified objective function. This allows for designing massive, controlled libraries without being limited by synthesis costs. We show how PGLD enables lab-in-the-loop design of multi-round high-throughput experiments, and large-scale in-vitro DNA sampling from generative models. Finally, we use PGLD to design a library of $\sim 10^6$ unique sequences at a cost of $\sim 700$ USD to explore the mutation space of a broadly neutralizing influenza antibody.
Bridging Graph Worlds: Neural Approximation of Gromov-Wasserstein Distances
Dong Qiao ⋅ Chris Ding ⋅ Jicong Fan
Graph-structured data is crucial in various domains like biology and social networks. Comparing graphs, which is a fundamental problem in graph data analysis, is nonetheless highly challenging. Recently, the Gromov-Wasserstein (GW) distance has provided a principled way to compare two graphs. However, computing the GW distance involves solving a complex non-convex optimization problem, making it computationally expensive, especially when the graphs are large. In this work, we propose a neural approximation of the GW distance, called NeuralGW. In NeuralGW, we use a combination of a graph isomorphism network and a transformer to represent the nodes of two graphs as two sets of vectors, treated as two discrete distributions, on which we compute multiple maximum mean discrepancy values given by different kernels. We then use a multilayer perceptron to convert the vector formed by these values into a single value, which is the prediction of the GW distance. Once trained, the model allows for efficient inference, enabling fast structural comparisons between graphs across diverse domains. We also provide a theoretical guarantee for the generalization ability of NeuralGW. Experiments demonstrate the effectiveness and practical applicability of our approach on real-world datasets, in comparison to baselines.
BSO: Safety Alignment Is Density Ratio Matching
Tien-Phat Nguyen ⋅ Truong Nguyen ⋅ Thin Nguyen ⋅ Duy M. H. Nguyen ⋅ Ngoc-Thanh Dinh ⋅ Trung Le
Aligning language models for both helpfulness and safety typically requires complex pipelines---separate reward and cost models, online reinforcement learning, and primal-dual updates. Recent direct preference optimization approaches simplify training but incorporate safety through ad-hoc modifications such as multi-stage procedures or heuristic margin terms, lacking a principled derivation. We show that the likelihood ratio of the optimal safe policy admits a closed-form decomposition that reduces safety alignment to a density ratio matching problem. Minimizing Bregman divergences between the data and model ratios yields Bregman Safety Optimization (BSO), a family of single-stage loss functions, each induced by a convex generator, that provably recover the optimal safe policy. BSO is both general and simple: it requires no auxiliary models, introduces only one hyperparameter beyond standard preference optimization, and recovers existing safety-aware methods as special cases. Experiments across safety alignment benchmarks show that BSO consistently improves the safety--helpfulness trade-off.
Can Agents Price a Reaction? Evaluating LLMs on Chemical Cost Reasoning
Yuyang Wu ⋅ Yue Huang ⋅ Shuaike Shen ⋅ Xujian Wang ⋅ Shuhao Zhang ⋅ Qiyao Xue ⋅ Weichen Liu ⋅ Runtian Gao ⋅ Jian Ma ⋅ Xiangliang Zhang ⋅ Olexandr Isayev
Large Language Models (LLMs) have become increasingly capable as tool-using agents, with benchmarks spanning diverse general agentic tasks. Yet rigorous evaluation of scientific tool use remains limited. In chemistry, recent agents can plan syntheses and invoke domain-specific tools, but evaluations often rely on curated demonstrations, expert assessment, or LLM-as-judge scoring rather than exact, judge-free ground truth. We address this gap with chemical procurement cost estimation, a practical task in which an agent must ground chemical identities, retrieve supplier quotes, select valid purchasable packs, normalize quantities, and compute cost from a reaction description. We introduce ChemCost, a benchmark of 1,427 evaluable reactions grounded to a frozen pricing snapshot covering 2,261 chemicals and 230,775 supplier quotes, supporting scalar scoring and stage-level diagnosis of grounding, retrieval, procurement, and arithmetic failures. To evaluate robustness, we further construct controlled noise-injected views that perturb chemical aliases, quantity expressions, missing fields, and input formatting. Experiments with frontier, open-weight, and chemistry-specialized LLM agents show that tool access is necessary but insufficient for solving the task. The strongest agents reach only 50.6% accuracy within 25% relative error on clean inputs and degrade substantially with realistic noise. Stage-level analysis further shows that failures arise from brittle parsing, ineffective evidence integration, invalid pack selection, and non-convergent tool use.
Can Ideologues Agree on Quality? From Non-identifiable Latent Factors to Collective Outcomes
David Gamba ⋅ Seura Ha ⋅ Daniel Romero ⋅ Grant Schoenebeck
To surface "objective" fact-check notes, X's Community Notes uses Biased Matrix Factorization (BMF) on crowd-sourced ratings to separate note quality from reviewer ideology. We show that for any correctly specified model that decomposes ratings into quality and ideological alignment, this separation requires fixing an ideological anchor: a reference point for what counts as neutral. We give a characterization for the standard implementation: with $L_2$ regularization, BMF anchors quality at a point proportional to the mean reviewer ideology, recovering objective quality if and only if the mean reviewer is neutral. This parallels the Condorcet Jury Theorem: collective judgment recovers truth when the pool is unbiased on average. When that assumption fails, BMF's recovered ranking shifts toward maximizing aggregate reviewer satisfaction, which coincides with social welfare only when the pool is representative. This is not specific to BMF; the rating matrix admits a continuous family of gauge-equivalent decompositions, each implying a different quality ranking, so no mechanism based on ratings alone can recover quality without implicitly choosing such an anchor. What appears to be routine hyperparameter tuning is in fact a normative choice about whose preferences define quality.
Can In-Context Learning Support Intrinsic Curiosity?
Eric Elmoznino ⋅ Sangnie Bhardwaj ⋅ Johannes von Oswald ⋅ Rajai Nasser ⋅ Blaise Aguera y Arcas ⋅ João Sacramento ⋅ Rif A. Saurous ⋅ Guillaume Lajoie
Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized data modeling, the problem of automated data selection, or "intrinsic curiosity", remains a significant challenge. Classic approaches incentivize exploration by rewarding an agent based on its "learning progress", which measures how much a newly acquired observation improves a world model's predictive ability. However, evaluating these rewards traditionally requires expensive inner loops of gradient descent updates within each trajectory, rendering them computationally impractical at scale. In this work, we investigate whether the emergent in-context learning (ICL) capabilities of sequence models can eliminate this bottleneck by serving as immediate, update-free world models. Specifically, we evaluate whether an exploration policy can be trained to maximize learning progress, using solely the prediction errors and counterfactual context manipulations of an in-context learner. We first prove that in general Markov decision processes, this is in fact impossible in an unbiased way: the resulting intrinsic rewards either suffer from nuisance terms that bias their estimation of true learning progress, or they cannot be implemented using an in-context learner's prediction errors. Conversely, we prove a positive result for a broad subclass of non-temporal settings, encompassing active learning and Bayesian Experimental Design: here, ICL-derived rewards successfully bound and asymptotically converge to the true learning progress. We corroborate our theory with controlled experiments across continuous and symbolic environments, demonstrating that our ICL-driven framework successfully trains curious data-collection policies that explore optimally.
Cannistraci-Hebb Channel-wise Dynamic Sparse Training of Convolutional Neural Networks with Contextual Modulation
Wenjing Wu ⋅ Hanming Li ⋅ Xizheng Deng ⋅ Jialin Zhao ⋅ Yingtao Zhang ⋅ Yusong Wang ⋅ Carlo Vittorio Cannistraci
Dynamic Sparse Training (DST) is an effective paradigm for sparse to sparse training of neural network connectivity under a fixed parameter budget. Recent advances in network-science for AI introduced epitopological learning methods, such as Cannistraci-Hebb Training (CHT), demonstrating that network automata applied to the mere network topology enables gradient-free predictions of the sparse connectivity evolution that can trigger significant increase in task performance. However, these methods were currently developed for standard MLP-like fully-connected architectures and need significant rethinking to be extended to convolutional neural networks (CNNs) due to element-wise weight sharing and spatially repeated interactions. In this work, we bridge this gap by introducing epitopological learning DST for CNNs through channel-wise network modeling to address the prohibitive complexity of element-wise modeling, which treats each kernel-channel as a single network node. Results show orders-of-magnitude improvements in time complexity of link-regrowth when applied at channel-level with respect to element-level. We also enhance the channel-wise formulation with lightweight contextual modulation to improve expressivity. Extensive experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet with ResNet and VGG architectures demonstrate that our method achieves performance comparable to dense training using only 30% of the parameters, which significantly reduces computational cost. Furthermore, our approach improves robustness under noisy inputs. These results highlight the effectiveness of topology-aware sparse training and the benefit of decoupling graph construction from sparse update rules in convolutional networks.
Can We Model the Artifacts Explicitly? Disentangle Artifacts via Pairwise Edit Relations for Image Manipulation Localization
Xuekang Zhu ⋅ Kaiwen Feng ⋅ Ruifeng Wang ⋅ Xiwen Wang ⋅ Xiaochen Ma ⋅ Bo Du ⋅ Changjiang Jiang ⋅ Chenfan Qu ⋅ Songyu Ye ⋅ Jian liu ⋅ Ji-Zhe, Knight Zhou
Image Manipulation Localization (IML) is commonly formulated as a fully supervised learning task that estimates the optimal manipulation mask $y$ for a given image $x$. In this work, we first reveal the latent nature of artifacts and thus reinterpret IML as a latent-variable problem, $P(y| x)=\int P(y|z) P(z|x)dz$, where $z$ denotes the artifacts. Following this interpretation, we pinpoint the cause for the current IML models' insufficiency as their implicit artifacts modeling strategy, highlighting the necessity of modeling $z$ in an explicit manner. Without direct labels, feature disentanglement is the most appropriate solution for this explicit modeling. Accordingly, we propose a two-stage learning paradigm with the Pairwise Artifacts Learning (PAL) and Standard Localization (SL) phases to estimate $P(z|x)$ and $P (y|z)$ via edit relations. To support our edit-relation-based learning, we further curate EditGroup-45K, a source-anchored dataset organized into edit groups for pair construction. Extensive experiments show that our PAL paradigm yields consistent improvements across diverse IML architectures, and empirical analyses further verify that PAL does capture artifacts explicitly through feature disentanglement. Code and dataset will be publicly available.
CARVE: Counterfactual Video Editing for Auditing and Hardening Video Detectors
Wei Zhou ⋅ YIMING CHEN ⋅ Xu Jinwei ⋅ Yang zhou ⋅ Quan Gan ⋅ Li Yang
Video detectors are usually evaluated on observational splits, where target events co-occur with environmental and capture factors. A high AUC under this setup can reflect either the event itself or its surrounding conditions, and standard reporting cannot tell them apart. This work addresses the entanglement with CARVE, a counterfactual video editing protocol. For each source clip, CARVE generates a matched quartet $(V^0, V^E, V^A, V^{AE})$ that holds camera pose, layout, and surrounding traffic fixed while independently intervening on environment and event. A two-phase generator first produces a layout-conditioned reference image, then performs reference-guided video editing; a three-layer reference / objective / VLM-panel protocol filters the output. Each quartet supports thresholded and continuous diagnostics for false-positive purity, event faithfulness, environment-induced drift, and held-out factor compositions. Thresholded scores are paired with continuous ones so that a conservative detector cannot suppress scores everywhere and look robust. The same audit drives training: CGAA-IC samples weak factors weighted by measured brittleness and adds quartet-level purity, consistency, and faithfulness losses; the inference graph is unchanged. On CCTV accident detection, the audit shows that clean AUC does not predict robustness rankings: a VideoMAE detector with $91.4\%$ AUC on TAD scores only $72.3\%$ CPS on the night slice of CARVE-Q. CGAA-IC raises purity from $84.2\%$ to $89.9\%$ and cross-dataset average AUC from $80.3\%$ to $84.5\%$.
CAST: Certifiable Aggregation of Smoothed Teachers for Robust Policy Adaptation
Zhongrui Zhao ⋅ Yanan Cai ⋅ Zhigang Lu ⋅ Longkun Guo ⋅ Ickjai Lee ⋅ Shuchao Pang ⋅ Minhui Xue
Transfer learning (TL)-based policy adaptation in deep reinforcement learning (DRL) usually relies on target-side data information for retraining when facing new tasks. However, two important issues remain underexplored in practical scenarios: existing DRL transfer learning methods usually lack theoretical guarantees against adversarial attacks, and target-side data information may be unavailable for retraining. In this paper, we propose Certifiable Aggregation of Smoothed Teachers (CAST), which adapts source policies by aggregating multiple certified smoothed teachers without any form of retraining, and further certifies the robustness of the adapted policy at both the action level and the cumulative reward level. CAST certification faces three key challenges: (i) value-function transfer in DRL cannot preserve certified robustness without robust retraining; (ii) reward changes break the direct reusability of source teachers' action-level certificates; and (iii) existing certified robustness transfer results provide limited guidance for certifying the student's cumulative reward. These challenges prevent us from directly using existing TL techniques in DRL. To address them, CAST constructs robustness signals from source certificates, incorporates them into policy aggregation to obtain action-level certificates, and bridges the adapted policy's action-level certificates to its target-task cumulative-reward certificate. Experiments on five multi-objective DRL benchmarks show that CAST exceeds the best source teacher in 17 out of 40 attack configurations and largely remains within the performance range of the source teachers. When testing whether smoothing benefits can be transferred, CAST improves in 39 out of 40 configurations, with a maximum relative gain of 183.9\%.
Catch-Only-One: Non-Transferable Examples for Model-Specific Authorization
Zihan Wang ⋅ Ethan Ma ⋅ Zhongkui Ma ⋅ Shuofeng Liu ⋅ Akide Liu ⋅ Derui Wang ⋅ Minhui Xue ⋅ Guangdong Bai
Recent AI regulations increasingly emphasize the need for mechanisms that preserve the utility of data for AI innovation while preventing misuse, particularly by enforcing purpose limitation in downstream AI applications. In practice, enforcing this principle remains challenging, as released data can be trivially fed into arbitrary models beyond its declared intent. Existing approaches attempt to mitigate this risk by either perturbing data or retraining models to limit unintended use. These strategies, however, offer no protection against inference by unknown or externally trained models, or fundamentally rely on control over the training or deployment. In this work, we introduce non-transferable examples (NTEs), recoded data that act as a task-level "ciphertext" decodable only by a designated model. Where adversarial examples exploit sensitive input directions, NTEs use the complementary insensitive subspace: a training-free, data-agnostic recoding within a model-specific low-sensitivity subspace preserves the authorized model's outputs while degrading unauthorized ones through subspace misalignment. We establish formal bounds certifying output fidelity for the authorized model and showing unauthorized degradation scales with measurable spectral misalignment between models. Empirically, NTEs preserve authorized-model performance across diverse vision backbones and vision-language models, while unauthorized models collapse even under adaptive reconstruction attacks. These results establish NTEs as a practical means to preserve intended data utility while preventing unauthorized exploitation. Our source code and visual demos are available at: https://github.com/model-specific/non-transferable-examples.
CausalConflictBench: Can Multimodal Models Follow Local Mechanisms That Conflict with Commonsense?
Bo Tian ⋅ Jianfeng Qu ⋅ Peng-Fei Zhang ⋅ Siyu Li ⋅ Zhixu Li ⋅ Kaiye Yu
Large multimodal models often benefit from commonsense priors, but these priors can mislead reasoning when a task specifies a local mechanism that conflicts with real-world regularities. Existing scientific VQA and visual reasoning benchmarks tend to align problem mechanisms with commonsense, so a correct answer may reflect either mechanism following or prior-based answering. We introduce Causal-ConflictBench, a diagnostic benchmark that makes this ambiguity observable by constructing samples where the current cause-to-effect mechanism contradicts the default commonsense mechanism. CausalConflictBench contains two modules: Textual Rule Override (TRO), which provides explicit commonsense-conflicting textual rules in real science image questions, and Visual Counter-Commonsense Induction (VCI), which requires models to induce a conflicting mechanism from a three-frame visual sequence. Beyond sample-level accuracy, we report Group-Strict Accuracy, factual-prior fallback metrics, and output-level process diagnostics. Across 14 proprietary and open-source multimodal models, we find that high sample-level accuracy can mask unstable mechanism following; errors concentrate strongly on factual-prior answers; and failure modes differ across rule-delivery paths, with TRO revealing rule-application failures and VCI revealing visual rule-induction failures. CausalConflictBench therefore provides a controlled diagnostic setting for uncovering commonsense fallback hidden beneath aggregate accuracy.
Causal Discovery with False Positive Error Control
Erik Jahn ⋅ Venkat Chandrasekaran ⋅ Frederick Eberhardt ⋅ Leonard Schulman
Causal discovery methods produce data-driven hypotheses about causal relationships, which may become new scientific findings or inform decisions about downstream interventions. In such settings, false positive causal claims can be more consequential than false negatives. Yet few methods target false positive error control in causal discovery, and existing error guarantees are often asymptotic and require strong versions of the faithfulness assumption. We introduce the $m$-conservative PC algorithm and prove that it controls false discovery rate in the linear Gaussian setting under an extremely weak version of the faithfulness assumption. In our framework, false positive error for causal graphs is defined based on the conditional-independence information they entail, yielding a metric that penalizes both false adjacencies and incorrect edge orientations. The parameter $m$ quantifies a trade-off between algorithmic complexity and the strength of the required faithfulness assumption. The $m$-conservative PC algorithm combines undirected graph learning, a Benjamini-Yekutieli-style procedure for conditional independence testing, and a conservative orientation step that marks unresolved or contradictory edge directions as ambiguous. Our algorithm produces valid outputs and retains error guarantees even in the finite-sample setting, despite possible random errors in the conditional independence tests.
Latent confounding poses a significant obstacle to identifying the causal effect of a treatment on an outcome. To address this, many existing studies leverage proxies of the latent confounder to indirectly adjust for the confounding bias. However, they impose specific structural constraints on the proxies and typically require multiple proxies. In this paper, focusing on the latent variable linear non-Gaussian acyclic model (lvLiNGAM), we propose a causal effect identification procedure requiring only a single agnostic proxy. Crucially, the term "agnostic" means that the causal connections between the proxy and the treatment-outcome pair can be arbitrary and are not required to be known a priori. This structural complexity precludes identifying the causal effect via a simple closed-form formula. Consequently, our identification procedure is designed to first derive candidate solutions from cumulants and then isolate the valid solution by examining certain independence relationships. We present a series of new theoretical results, which collectively establish the soundness of our identification procedure: given the observational population distribution, it correctly identifies the true causal effect when identifiable, and correctly reports unidentifiability otherwise. Finally, we empirically validate our theoretical results.
CausalMix: Data Mixture as Causal Inference for Language Model Training
Zinan Tang ⋅ YUKUN ZHANG ⋅ Shaomian Zheng ⋅ Zhuoshi Pan ⋅ Qizhi Pei ⋅ Dingnan Jin ⋅ Jun Zhou ⋅ Yujun wang ⋅ Biqing Huang
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely on the assumption of static data distributions. As a result, when the underlying data pool shifts, these methods require costly retraining from scratch. This limitation restricts their ability to scale seamlessly from small settings to larger data pools and model sizes. In this paper, we propose CausalMix to address this limitation by casting data mixture optimization as a causal inference problem. We formulate the statistical features of the data pool as covariates and the domain mixture as the treatment. After fitting a causal model on 512 runs of Qwen2.5-0.5B to estimate the Conditional Average Treatment Effect (CATE), we extrapolate the optimal mixture for an 800K data pool and apply it to train a 7B model. Furthermore, we successfully generalize the framework to long chain-of-thought data on Qwen3-4B-Base. By leveraging causal modeling to isolate confounding biases, CausalMix dynamically infers state-dependent optimal data mixtures. Extensive experiments show that the mixture guided by CausalMix consistently improves performance across multiple downstream tasks, outperforming RegMix and other baselines. In addition, we use the CATE Interpreter to provide visual analysis of the learned mixing strategy. Overall, CausalMix offers a causal and interpretable framework for optimizing LLM data mixtures.
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophisticated skills that remain largely untested in agents: (1) navigating long horizons amid uncertainty; (2) acquiring information in noisy environments; (3) adapting to a changing world; (4) orchestrating multiple moving parts toward a coherent goal. We introduce CEO-Bench, which evaluates these capabilities together by simulating a representative real-world task: operating a startup for 500 days. Given diverse and realistic company management tools, business databases, and social media, an agent needs to design pricing strategies, allocate operating budgets, analyze business data, respond to unexpected competitor moves, and more. Our evaluation shows that most state-of-the-art models struggle to succeed in this environment, and only one model (GPT-5.5) finishes the simulation above its $1M starting balance. CEO-Bench reimagines the role of AI in the future, shifting from solving isolated tasks to driving sustained, adaptive progress over time. We open source full code and trajectory.
We investigate the use of formal methods to provide tight and sound generalization bounds for learning algorithms. By casting the traditional notion of algorithmic stability as a specification to be verified, we demonstrate that recent advances in reachability analysis can yield provable bounds on the generalization of a given model and algorithm on a sample dataset. As sample-specific algorithmic stability is insufficient to bound the usual distributional notion of generalization, we develop a novel concentration inequality to connect the sample-specific results of formal certification algorithms to the required distributional analysis.The resulting framework enables the theoretical analysis of prior generalization bounds to extend far beyond their original restrictive analytical assumptions while achieving a provably sound bound on the generalization gap. In practice, we demonstrate that our framework provide formal generalization guarantees that are orders of magnitude tighter than alternative computational approaches at scales ranging from toy datasets to fine-tuning of modern large language models. While we implement certification-enhanced versions of several well-known stability results, future extensions of our approach will enable tighter bounds and enhanced practical adoption across the spectrum of modern generalization bounds.
Chain-of-Correction: Progress-Aware Policy Steering via Anchor-Grounded Predictive Reasoning
Xuening Zhang ⋅ Xiang Deng ⋅ Jiayi Lin ⋅ Qi Lv ⋅ Xingbo Liu ⋅ Weili Guan
Imitation learning policies for robotic manipulation inherently suffer from compounding errors. To mitigate this, inference-time policy steering introduces predictive reasoning to evaluate and refine action proposals before execution. However, existing methods typically anticipate consequences in unstructured pixel or latent spaces, leading to physically inconsistent predictions and unreliable error detection. Furthermore, their blindness to global progress often yields locally plausible corrections that fail to complete the overall task. To address these limitations, we propose Chain-of-Correction (CoC), an inference-time steering framework driven by a single VLM. Shifting away from unstructured state predictions, CoC explicitly leverages anchor-grounded scene graphs to track exact 3D physical relations. Built upon this representation, a progress-aware dual-check mechanism verifies historical execution and predicts future task advancement to systematically intercept myopic action proposals. Through a failure-augmented supervised fine-tuning pipeline, CoC seamlessly unifies scene graph extraction, progress evaluation, and precise kinematic correction, eliminating the need for task-specific world models. Extensive experiments across the COLOSSEUM benchmark, a newly introduced offline protocol for predictive reasoning, and real-world tasks demonstrate that CoC significantly outperforms state-of-the-art baselines in fine-grained error detection and robust error recovery.
Change-Robust Online Topological Memory for Long-Term Relocalization and Semantic Navigation
Jiaming Wang ⋅ Liu Diwen ⋅ Chen Jizhuo ⋅ Atharva A Ghotavadekar ⋅ Da Jiaxuan ⋅ Linh Kästner ⋅ Harold Soh
Long-term semantic navigation requires a robot to reuse past observations after appearance and scene change, but semantic memories are only useful if the robot can relocalize into the memory without corrupting it with false visual matches. We propose CROSS, a change-robust topological memory that introduces a pre-commitment localization layer between visual place recognition and map update. Instead of treating a retrieved keyframe as an immediate place association or loop-closure factor, CROSS lifts each RGB-D retrieval into a candidate global $\mathrm{SE}(3)$ pose mode using relative pose estimation. A bounded Gaussian-mixture filter then propagates competing continuous trajectory branches with odometry, rejects branches that are physically inconsistent, and promotes only persistent branches to loop closures. This moves ambiguity handling from discrete place IDs or post-hoc graph-factor rejection to continuous pose-space validation before map commitment. Across public long-term relocalization benchmarks and real quadruped object-navigation experiments, CROSS improves reuse of a single sparse RGB-D memory under illumination, seasonal, dynamic-scene, and object-level change.
Channel-wise Vector Quantization
Wei Song ⋅ Tianhang Wang ⋅ Yitong Chen ⋅ Zuxuan Wu ⋅ Tong Zhang ⋅ Min Li ⋅ Jiaqi Wang ⋅ Kaicheng Yu
We present Channel-wise Vector Quantization (CVQ), a novel image tokenization paradigm that replaces patch-wise tokens with channel-wise tokens. Unlike conventional vector quantization, which assigns a discrete token to each patch feature vector, CVQ quantizes each channel of the feature map. This formulation represents an image as discrete levels of visual details, rather than as a grid of spatial patches. Based on CVQ, we introduce a new visual autoregressive framework with "next-channel prediction". Instead of rendering images patch by patch in raster order, our Channel-wise Autoregressive (CAR) model predicts image channels sequentially, producing progressively enriched visual details. Specifically, it first sketches global structure and then refines fine-grained attributes, akin to a human artist's workflow. Empirically, we show that: (1) CVQ achieves 100% codebook utilization with a 16K+ codebook size without any bells and whistles, while reducing reconstruction FID by 50% over conventional VQ; and (2) CAR outperforms the AR baseline by improving the GenEval score from 0.69 to 0.74 and the DPG score from 79.86 to 82.14, demonstrating strong effectiveness for text-to-image generation. We hope our research offers a new perspective on the fundamental unit of visual tokenization by moving from spatial patches to channels.
Characterizing the Generalization Error of Random Feature Regression with Arbitrary Data-Augmentation
Lucas Morisset ⋅ Alain Durmus ⋅ Adrien Hardy
This paper aims at analyzing the regularization effect that data augmentation induces on supervised regression methods in the proportional regime, where the number of covariates grows proportionally to the number of samples. We provide a tight characterization of the test error, measured in mean squared error, in terms only of the population quantities of the true data, as well as first and second order statistics of the augmentation scheme. Our results are valid under misspecified feature maps, and for any network architecture where only the last readout layer is trained, and the rest of the network is either frozen or randomly initialized. We specify our results in the case of Gaussian data, and show that our asymptotic characterization is tight in this setting.
Chatter Attack: Resource Consumption Attack for Large Language Models
Yibo Miao ⋅ Yichuan Cao ⋅ Xiao-Shan Gao ⋅ Yinpeng Dong
LLMs are powerful but incur substantial inference cost, which can be exploited by malicious users to induce overly long outputs, increasing latency and operational expense. Prior resource consumption attacks either coerce repetition or suppress the end-of-sequence token by lowering its logit value. However, with autoregressive, samplingbased generation, the search space grows exponentially with length and small deviations or unstable trajectories can derail optimization, limiting attack effectiveness and stability. To address this, we introduce Chatter Attack, a novel class of resource consumption attacks that optimize adversarial prompts to induce predefined short target phrases, thereby avoiding search-space explosion while still eliciting long responses. Specifically, Chatter Attack triggers self-doubt using short doubt-inducing phrases as the target, and sustains this state throughout generation via inference attention loss and global entropy loss. To extend to black-box settings, we further propose a novel Bayesian Prompt Optimization (BPO) method, which models the objective function globally by constructing a discrete token kernel, efficiently exploring the entire search space using prior information. Extensive experiments across multiple models and datasets show that our method outperforms baselines, with maximum response length increases to 31.5×, and revealing practical resource-depletion risks for LLM services.
Chess-World-Model: A 10M-Game Benchmark for Exact State Tracking from Chess Move Sequences
Benjamin Walker ⋅ Terry Lyons
World models require state tracking, which is the ability to maintain a correct latent state across action sequences. Existing benchmarks are often synthetic or language-based, limiting their value as tests of structured state updates in realistic domains. We introduce Chess-World-Model, a large-scale state-tracking benchmark built from $10$ million real chess games, where models predict the exact board state reached after a sequence of legal moves. Alongside a held-out real-game split, we include an out-of-distribution split from uniformly random legal play, which tests whether models learn the transition rules rather than shortcuts from common human positions. Prior theoretical and empirical work has shown that Transformers struggle to state-track, while input-dependent linear RNNs require expressive state-transition matrices to do so. We therefore benchmark a causal Transformer, block-diagonal SLiCE, Mamba-3, and Gated DeltaNet with negative eigenvalues under a matched interface and training protocol. The recurrent models strongly outperform the Transformer at $3$ and $8$ million parameters. Real-game performance saturates above $18$ million parameters, but the random-uniform split remains discriminative up to $38$ million, exposing failures otherwise hidden by scale. Additionally, ablations show that less expressive state-transition mechanisms reduce performance on the out-of-distribution split for all three recurrent models. Together, these results establish Chess-World-Model as a practical large-scale benchmark for state tracking that exposes failures model scale would otherwise conceal.
Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use Agents
Yu Li ⋅ Zheng Zhang ⋅ Xin Liu ⋅ shengtian yang ⋅ Guangfeng Cai ⋅ Lei Feng
Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate the downstream effect of an intermediate decision. In this paper, we argue that effective tool-use agents should estimate the long-horizon value of a possible next tool invocation before executing it. This objective requires comparative supervision over alternative invocations under the same context, while logged trajectories only contain the invocation that was actually taken. Therefore, we propose Comparative Inference for Tool-use Agents (CITA). CITA trains a Comparative Inference Model (CIM) from paired signals that combine observed tool behavior, scalable supervision from a Bayesian tool-graph simulator, and semantic judgments from LLM-based comparison. The resulting CIM learns to estimate how likely a possible next tool invocation is to support final task success under the current context. Across three tool-use benchmarks and multiple backbone LLMs, CITA consistently improves Tool F1 and task success. Additional analysis shows that CIM learns accurate step-level value estimates for comparative tool choices.
CineMME: Benchmarking Fine-Grained Perception and Plot Reasoning in Multimodal Large Language Models
Jin Liu ⋅ Dawei Du ⋅ Yexiang Liu ⋅ Sijie Zhu ⋅ Fan Chen ⋅ Junxian Duan ⋅ Huaibo Huang ⋅ Ran He ⋅ Longyin Wen ⋅ Zhenfang Chen
Cinematic understanding requires more than recognizing visible events: models must acquire dialogue evidence, bind it to speakers, and use it to infer narrative meaning. Existing video benchmarks evaluate pieces of this problem, including video question answering, audio-visual reasoning, and temporal grounding, but rarely test whether these abilities form a coherent cinematic evidence chain within the same clips. As a result, a model may answer a movie question correctly while failing to localize the relevant utterance, identify the speaker, or integrate speech with visual context. We introduce **CineMME**, a human-verified benchmark designed to diagnose speech-grounded cinematic reasoning, supported by approximately $1,600$ expert person-hours of annotation. CineMME contains $975$ curated movie clips and two linked tracks: *CineMME-Grounding*, with $13,554$ timestamped speech segments and $33,083$ actor face boxes for temporal dialogue localization, transcription, and active-speaker grounding; and *CineMME-Reasoning*, with $1,600$ multiple-choice and $316$ open-ended questions spanning perception, audio, narrative, social, and cinematic understanding. We evaluate $20$ MLLMs on reasoning and $10$ omni-modal models on grounding. The results reveal a pronounced who-what-when decoupling: Gemini-3 Pro localizes dialogue in time 68.33% tIoU) and partially transcribes it (WER 29.00%), yet reaches only 13.24% sIoU for active-speaker grounding while achieving 65.88% reasoning accuracy. Modality ablations further show that vision is not uniformly beneficial: it improves perception-heavy questions but can hurt dialogue-centric reasoning. CineMME provides a compact diagnostic testbed for measuring whether multimodal video models truly connect dialogue, speakers, visual evidence, and narrative meaning, rather than succeeding through disconnected cues.
CIPHERGRID Benchmark: From Multimodal Rule Inference to Sequential Action
Christopher Curtis ⋅ Victor Fragoso ⋅ Saiph Savage
We introduce CIPHERGRID, a rule-inference and path-finding benchmark. In each task, models are given example grid-world images paired with encoded descriptions and encoded solutions, then receive a new encoded grid-world to solve. To answer correctly, they must infer which symbols denote tiles, actions, row boundaries, and game mechanics, decode the new world, and produce an encoded path from the start cell to the target cell. The benchmark is designed so that each component is individually solvable: rule inference is calibrated against human performance, and path-finding is solver-verifiable via dynamic programming. Although the underlying skills are individually solvable and verifiable, CIPHERGRID tests their interaction: whether models can infer an unfamiliar symbolic system, and carry it through rule application, and planning. CIPHERGRID remains difficult for current frontier models, which only achieve between 18.74\% and 42.20\% accuracy. CIPHERGRID offers a controlled test of whether models can coordinate multimodal grounding, rule inference, and sequential planning under unknown symbolic relationships.
Clean Data Can Still Carry Backdoors: Support-Persistent Backdoors for Model Reuse
Junhoo Lee ⋅ Baekseung Kim ⋅ Seungyeon Kim ⋅ Nojun Kwak
Traditional backdoor attacks inject artificial trigger patterns into training data, causing models to misclassify triggered inputs while maintaining normal behavior on clean samples. Since these artificial triggers are absent in clean datasets, standard knowledge distillation (KD) typically eliminates them, leading to the common belief that KD serves as an effective purification defense. In this paper, we propose DAR (Discover-and-Relabel), a support-persistent backdoor method that retains its behavior even after clean-data distillation. DAR identifies naturally abundant patterns in low-dimensional feature subspaces using subspace clustering and applies label-only poisoning to samples that inherently satisfy the discovered rule. At test time, the backdoor is activated by locally modifying the corresponding subspace of the target images. By instantiating spatial and frequency operators, DAR achieves attack accuracy comparable to state-of-the-art backdoor attacks in ImageNet classification and CLIP-based prompt tuning. Notably, DAR keeps attack success above 96% after clean KD and remains defense-resistant in practice.
ClinStab: Stability-Oriented Learning for Medical Time Series via Dual-Stream Alignment
Songning Lai ⋅ Haoxuan Xu ⋅ Yi Liu ⋅ Wenshuo Chen ⋅ Ninghui Feng ⋅ Jiayu Yang ⋅ Haoang Li ⋅ Yutao Yue
Deep models for medical time series can achieve strong benchmark accuracy yet remain brittle under acquisition artifacts, physiological variability, and low-resource training conditions. We present \textbf{ClinStab}, a Medformer-based framework for EEG and ECG classification that studies whether channel-aware dual-stream interaction and perturbation-aware optimization can improve the clean/robustness tradeoff. The proposed \textbf{S}alience-\textbf{C}ontext \textbf{D}ual \textbf{M}odulation (SCDM) module uses energy as a routing prior: high-energy channel-aligned components receive attentive modeling, while residual components remain trainable through a lightweight context pathway. ClinStab then combines clean supervision, perturbed supervision, prediction consistency, and teacher-guided stabilization under adversarial perturbations. Across four public datasets and 11 baselines, ClinStab improves accuracy over Medformer on all datasets and improves selected robustness metrics under the perturbation settings studied here, while exposing tradeoffs on PTB ranking metrics and soft-routing alternatives. These results support ClinStab as a practical same-backbone stability intervention, not as evidence of universal clinical robustness.
CLUE: Closing the Loop on Conflict and Collapse in LLM Unlearning
Mengyang Li ⋅ Jingwen Wang ⋅ Yu Zhang ⋅ Pinlong Zhao
Machine unlearning for large language models optimizes two opposing losses: a forget objective drives the model away from targeted knowledge, while a retain objective holds it near its original behavior. Existing methods commit to a fixed schedule of learning rate, retain weight, and step count chosen by offline grid search, which is blind to how the trajectory actually evolves. We show that LLM unlearning trajectories share a method-agnostic three-phase structure, an early phase where the two gradients are nearly orthogonal, a conflict phase where they oppose each other and retain loss begins to grow, and a terminal phase where representational divergence crosses an irreversibility threshold. Each phase calls for a different control policy. We propose CLUE, a closed-loop framework whose architectural choices each correspond to a phase or transition: a dual-tower controller separates forget and retain signal pathways, a conflict-driven gate switches between them, and an explicit collapse-risk head supervises anticipatory stopping. The controller is trained by task-distribution meta-optimization. We provide a single-step bound on conflict-induced retain loss growth, a quantitative irreversibility result, and a stopping regret bound. On TOFU, WMDP, and MUSE with Llama-2/3 and Mistral models from 7B to 70B, CLUE achieves better forget-quality versus model-utility trade-offs than fixed-strategy baselines, stops within 5\% of the oracle, and remains robust under adversarial prompting, relearning, and INT4 quantization.
CMI-Trans: Cross Modal Inconsistency-aware Transport for HSI-LiDAR Classification
Yanli Li ⋅ Xuan Tan ⋅ Ding Qi ⋅ XINYANG JIANG
Joint classification of hyperspectral imagery (HSI) and Light Detection and Ranging (LiDAR) data benefits from complementary spectral--geometric cues, but existing fusion methods largely assume locally consistent cross-modal features and symmetric pixel-wise trust, which is often violated in real urban scenes. We term this pixel-level phenomenon \emph{cross-modal local inconsistency} (CMI) and quantify it with a Strong-Edge CMI Ratio: dual-modal gradient analysis shows that over $16\%$ of strong-edge pixels exhibit single-modality dominance across three benchmarks, while ROI-level analysis reveals the class-entanglement induced by symmetric fusion. To address this issue, we propose \textbf{CMI-Trans} (\textbf{C}ross-\textbf{M}odal \textbf{I}nconsistency-aware \textbf{Trans}port), which promotes per-pixel reliability modeling from a post-hoc diagnostic to a structural signal for alignment and training. CMI-Trans combines Energy-Calibrated Reliability (ECR), Uncertainty-Shaped Transport Alignment (USTA) with Sinkhorn-regularized optimal transport, and Triadic Co-Optimization (TCO) that jointly optimizes classification, uncertainty consistency, and alignment stability on a Mamba-based fusion backbone. On Houston, Trento, and MUUFL, CMI-Trans achieves competitive overall accuracy and substantial gains on CMI-affected classes (e.g., $+17.90\%$ on Houston C10). The ECR-derived uncertainty score also separates reliable from erroneous predictions ($\rho(U,e)<0$; top-$20\%$ highest-$U$ error rate $\le 0.18\%$), while improving feature separability in shadow-induced CMI regions. These results show that explicitly modeling CMI through uncertainty-shaped adaptive alignment yields more accurate and interpretable multimodal remote sensing classification.
Coarse-to-Refine: Trajectory Self-Refinement in Single Autoregressive Pass for Driving VLA
Canyu Chen ⋅ Yuguang Yang ⋅ Jianing Pang ⋅ Zhewen Tan ⋅ Cheng Chi ⋅ Chunyang Liu ⋅ Kehua Sheng ⋅ Bo Zhang ⋅ Linlin Yang ⋅ Xiaoyan Luo ⋅ Yan Wang ⋅ Baochang Zhang
Vision-language-action (VLA) models have emerged as a promising paradigm for autonomous driving trajectory planning. While explicit test-time thinking has driven mainstream progress in Large Language Models (LLMs), how to effectively utilize reasoning in driving VLA remains unsettled. We revisit this gap through Coarse-to-Refine (C2R), a trajectory self-refinement framework for driving VLA that completes refinement in a single autoregressive pass. C2R first predicts a coarse trajectory, performs explicit reasoning over it, and then generates a refined trajectory autoregressively. We show that the refinement formulation itself drives imitation learning improvement while enabling reinforcement learning to utilize reasoning content for further optimization. C2R establishes VLA state-of-the-art performance across standard and challenging benchmarks, achieving 92.1 PDMS on NAVSIM v1, 90.4 EPDMS on NAVSIM v2, and 49.6 EPDMS on NavHard.
COCOTree: A Dataset and Benchmark for Open Tree-Structured Visual Decomposition
Junhyub Lee ⋅ Seunghun Chae ⋅ Hyosu Kim
We formalize and enable the task of open tree decomposition, which segments an image into hierarchical trees of visual components with unconstrained granularity and flexibility. Specifically, we provide the foundation benchmark for this new paradigm with the following three key contributions. First, we overcome the prohibitively high cognitive and physical bottlenecks of manual annotation by developing a fully automated generation pipeline that synergizes the semantic reasoning of Large Vision-Language Models (LVLMs) with the precise geometric grounding of SAM 3. Second, leveraging this pipeline, we construct COCOTree, a massive-scale benchmark featuring over 21K images and 1.8M structural nodes. By embracing an open-vocabulary space of over 3.5K unique labels, it successfully captures the long-tail distribution of complex physical assemblies. Notably, rigorous human evaluation confirms our generated annotations demonstrate strong alignment with human structural judgment. Third, we establish a standardized evaluation protocol by proposing the Open Tree Quality (OTQ) metric, which jointly assesses mask precision, label accuracy, and structural consistency. We release our dataset and benchmark code at https://github.com/melonkick3090/COCOTree.
CoDeRNet: Selective Cross-Task Routing under Heterogeneous Supervision for Change Detection and Captioning
Eunki Cho ⋅ Hyeon Bae Kim ⋅ Seong Tae Kim
Change Detection & Captioning (CDC) aims to jointly localize changed regions and describe their semantic content from bi-temporal remote sensing images. While recent approaches have demonstrated the effectiveness of unified modeling, they primarily rely on shared representations or jointly optimized architectures, leaving task-specific communication under heterogeneous supervision less explored. However, detection and captioning differ fundamentally in their supervision granularity, leading to heterogeneous representations with uneven reliability across tasks and spatial locations. In this paper, we reformulate CDC as a selective information routing problem between heterogeneous task branches, where the key challenge is to determine when and where information should be transferred. To this end, we propose CoDeRNet, a unified framework that preserves task-specific representations while enabling bidirectional exchange of complementary information. CoDeRNet introduces a Confidence-Decomposed Routing (CoDeR) mechanism, which decomposes the routing condition into receiver-side need for support and sender-side reliability cues to guide selective transfer. Experiments on LEVIR-MCI and WHU-CDC show that CoDeRNet achieves strong overall CDC performance, particularly improving caption generation while preserving competitive change localization. Extensive analyses further show that the benefits arise not merely from joint learning or generic cross-task communication, but from explicitly modeling selective cross-task information routing. These results highlight the importance of conditional information transfer for unified CDC under heterogeneous supervision.
CogArena: Benchmarking Multimodal Agents on Interactive Behavioral Experiments
Kianté Fernandez ⋅ Caitlin Chen ⋅ Jialu Zhou ⋅ Bartek Sadowski ⋅ Anthony C Miceli ⋅ Ian Krajbich
Claims about the cognitive capabilities of large language models are typically tested by translating behavioral paradigms into natural language prompts, bypassing the visual interfaces and time-pressured interactions through which human participants actually engage with experiments. In parallel, recent demonstrations that AI agents can produce human-like response data on live behavioral tasks raise concerns that online datasets used across the social and behavioral sciences are increasingly susceptible to contamination. We introduce CogArena, a benchmark of interactive experimental paradigms drawn from the social and behavioral sciences. Agents interact with the experiments through a standard browser, processing visual stimuli and responding under configurable deadlines. We designed a three-level scoring pipeline that evaluates task completion, performance accuracy normalized against literature-derived human baselines, and alignment with canonical task-specific signatures. Across frontier multimodal agents, we find that agents complete the tasks, but their behavior diverges from humans on the signatures these paradigms reliably elicit in people. CogArena provides the tasks, API, scoring pipeline, and a public leaderboard for ongoing community submissions.
We introduce CoLa3D, a simple and effective model for composable 3D scene decomposition that, given a single whole-scene latent from either a single image or an existing scene mesh as input, directly outputs a set of object meshes with consistent spatial layout that can be assembled back into a coherent scene. CoLa3D introduces a novel paradigm for 3D scene decomposition within a shared \emph{whole}-scene latent, unlike alternatives that perform per-object generation in canonical spaces and then glue the objects back together with fragile pose and scale estimation. The network is a lightweight, flexible, and scalable query-based transformer that accepts various types of queries, including 2D masks, 3D points, and learnable queries, and learns to decompose the scene into objects in a compact latent space. On both synthetic and real-world datasets, CoLa3D achieves state-of-the-art performance on composable 3D scene decomposition, substantially improving global scene-level consistency and object-level quality while running $2\times$ faster and using less than 1% of SAM 3D's training data. Visible depth on par with MoGe2 further confirms that latent decomposition preserves accurate visible scene geometry while producing physically complete, decomposable object meshes.
COLLAR: Cascaded Object-Level Latent Refinement for High-Fidelity Conditional Generation
Xinlong Zhang ⋅ Jia Wei ⋅ Xiaoyu Zhang ⋅ Teng Zhou ⋅ Chengyu Lin ⋅ Yongchuan Tang
Achieving high-fidelity object-level control in diffusion transformers remains a significant challenge despite the introduction of structural priors like depth and Canny maps. Current object-level conditional generation methods frequently suffer from visual artifacts and struggle to maintain precise control over objects within small localized regions. To address these limitations, we propose Cascaded Object-Level Latent Refinement (COLLAR), a training-free framework that progressively optimizes object-level features via the Field-of-View (FoV) expansion. First, we propose the Cross-Scale Semantic Alignment (CSSA) module to address spatial-semantic gaps by injecting object-level features into extended-FoV branches via attention mechanisms. To further optimize these features, the Cyclic Feature Injection (CFI) module introduces a reciprocal background feedback mechanism. It leverages a frequency-based adaptive strategy to selectively update the global backbone with context-aligned local information. Finally, the extended-FoV branch serves as a hub for feature optimization, ensuring that object-level features are integrated into the global generation process without compromising final image quality. Extensive experiments on the COCO-MIG and COCO-POS benchmarks demonstrate that our approach consistently outperforms state-of-the-art methods across semantic alignment, image quality, and spatial fidelity.
Collective Supervision for Unified Biomolecular Conformation and Dynamics Modeling with CoDyna
Kaiwen Cheng ⋅ FANDI WU ⋅ Dawei Huang ⋅ Jun Wu ⋅ Jianhua Yao
Modeling biomolecular conformation and dynamics is critical for elucidating biological functions, motivating generative surrogates for molecular dynamics (MD) simulations to address two complementary tasks: time-independent conformational sampling and time-dependent trajectory generation. Fundamentally, both tasks require matching the distribution of the generated collections to the MD reference, rather than reproducing individual configurations. Yet existing deep-learning approaches predominantly optimize a per-sample regression loss (score- or flow-matching), whose gradient is computed in isolation per sample and therefore carries no direct signal about the cross-sample distributional properties that define both tasks, leading to biased ensembles and long-horizon temporal drift. We address this with Collective Supervision, a training paradigm that aligns the empirical measures of the generated and reference collections via a maximum mean discrepancy on SE(3)-invariant physical observables; a single loss handles both tasks within a unified formulation. To mitigate the exposure bias inherent in autoregressive trajectory rollout, we further introduce Collective Rolling Forcing, which couples Collective Supervision with autoregressive self-rollout during training. Our framework, CoDyna, generalizes across proteins, protein--ligand and protein--protein complexes; on four all-atom MD benchmarks (ATLAS, MISATO, DynaRepo, and DynaBench), a unified model surpasses task-specialized baselines on most thermodynamic and kinetic fidelity metrics and remains structurally stable over 2k-frame rollouts.
ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models
Chenxi Ruan ⋅ Yihan Hou ⋅ Yu Xiao ⋅ Guosheng Hu ⋅ Wei Zeng
Text-to-image (T2I) models have advanced considerably in generating high-quality images from textual descriptions. However, their ability to associate colors with concepts remains largely constrained to explicit color names or codes, while their capacity to handle implicit concepts, such as emotions and visual states, remains underexplored. To address this gap, we introduce ColorConceptBench, an expert-annotated benchmark that systematically evaluates color-concept associations through probabilistic color distributions. ColorConceptBench moves beyond explicit color specifications by examining how models interpret 1,281 implicit color concepts, grounded in 6,584 human annotations. Our evaluation of nine leading T2I models reveals that performance varies substantially across semantic categories, and models exhibit a significant lack of sensitivity to abstract semantics. These limitations persist even when applying classifier-free guidance scaling at inference time, suggesting that achieving human-like color understanding demands a shift in how models learn and represent implicit semantic meaning.
COMET: Decoupled Distillation, Routing, and Capacity Control for Task-Agnostic Continual Vision--Language Learning
Seungyoon Woo ⋅ Keonhee Park ⋅ Gunhee Kim
Vision--language models (VLMs) such as CLIP retain zero-shot recognition ability, but their open-world use often requires task-agnostic continual learning: tasks arrive sequentially, past data are unavailable, task identity is unknown, and the backbone cannot be retrained. We argue that this setting fails VLMs through three coupled pressures beyond stability--plasticity tradeoff: (1) modal-consistency drift, (2) representation-space interference, and (3) per-task capacity shortfall. We propose COMET (COllaborative Mixture-of-Expert Transfer) to resolve these three issues respectively as follows. A frozen teacher provides only a feature reference on a shared image--text pool; Tri-Affinity Distillation preserves image--image, text--text, and image--text geometry; and Student-Aware Cross-Expert Distillation aligns new experts with compatible prior experts so routing errors degrade gracefully. We design COMET to keep capacity and inference separate: a contextual bandit keeps or merges full per-task experts using student-side signals only, while a reconstruction router selects sparse experts at test time without teacher semantics. On the X-TAIL benchmark, we show that each proposed components indeed improves the continual learning performance. COMET shows that continual vision--language adaptation can preserve multimodal geometry, coordinate accumulated experts, and allocate capacity jointly when feature reference, expert shaping, routing, and capacity control are separated by construction.
Communication-Efficient LLM Adaptation over Decentralized GPU Meshes
Sameera Ramasinghe ⋅ Shamane Siriwardhana ⋅ Thalaiyasingam Ajanthan ⋅ Hadi Mohaghegh Dolatabadi ⋅ Chamin Hewa Koneputugodage ⋅ Gil Avraham ⋅ Violetta Shevchenko ⋅ James Snewin ⋅ Karol Pajak ⋅ Harry Xi ⋅ Alexander Long
Decentralized training enables large-model training over low-end GPUs and internet-grade connections, but communication along both data-parallel and pipeline-parallel axes becomes the primary bottleneck. We study post-pretraining adaptation in this setting. We propose an asynchronous two-circuit system: a fast compressed training circuit drives throughput using activation masking for pipeline-parallel (PP) transfer and compressed data-parallel (DP) synchronization, while a slow anchor circuit runs occasional unmasked forward--backward passes off the critical path. Then, we introduce a spectral correction optimizer that uses these delayed anchor priors to denoise masked gradients without blocking the fast stream. Although prior work has found aggressive activation compression unreliable, we show that masking supports post-pretraining adaptation at high compression rates when anchored this way. Pipeline-parallel compression alone yields up to a 9x throughput gain, and combining it with data-parallel compression increases beyond 40x over internet-grade 200Mbps connections, while matching dense uncompressed performance across domain adaptation and continual pretraining.
CompactSplat: Spatially Adaptive Gaussian Distribution for Feedforward Scene Reconstruction
Bing He ⋅ Jingnan Gao ⋅ Yunuo Chen ⋅ Zhengxue Cheng ⋅ Rong Xie ⋅ Li Song ⋅ Wenjun Zhang
Reconstructing 3D scenes from sparse views without per-scene optimization remains highly challenging, especially for recovering accurate geometry and fine textures. While recent feedforward approaches leverage generalizable 3D Gaussian Splatting (3DGS) for scene generation, they typically assign one or multiple Gaussians to each pixel. Such uniform allocation not only produces highly redundant representations by wasting primitives in homogeneous areas, but also fails to exploit the inherent geometric priors of planar surfaces, often resulting in sub-optimal structural modeling. To address these limitations, we present CompactSplat , a novel feedforward framework that enables a content-aware allocation of 3D Gaussians. By integrating texture-guided spatial partitioning with hierarchical sampling, CompactSplat adaptively distributes primitives — concentrating dense Gaussians in complex regions while tiling flat areas with expanded, scale-modulated primitives. Moreover, our framework inherently supports dynamic sparsity control at inference time, enabling seamless trade-offs between rendering fidelity and memory footprint without requiring network retraining. Extensive experiments on RealEstate10K, DL3DV, and ScanNet demonstrate that CompactSplat consistently outperforms prior methods on both standard metrics and high-resolution rendering consistency, achieving high-fidelity reconstructions with significantly fewer primitives.
In this work, we provide a complete characterization of the uniqueness of linear patterns for generic fully connected neural networks activated by ReLU or leaky ReLU functions. With this result, together with several tools from geometry and probability, we compare the numbers of linear regions for a wide variety of networks from both theoretical and practical perspectives, with minimal prerequisites. More specifically, on the theoretical side, we show that networks with neurons arranged as $pq\times2$ (2 hidden layers, each having $pq$ neurons) are expected to have more linear regions than those arranged as $p\times 2q$ for $p\ge 1$ and $q\ge 2$; on the empirical side, we show the validity of a Monte-Carlo approach for comparing linear regions via linear patterns, and carry out experiments for networks in various scenarios. Our results indicate that neurons closer to the input layer tend to have more impact on the number of linear regions. Moreover, along the way, we find the precise value of the expected number of linear regions for networks with 2 hidden layers, and an explicit upper bound for networks whose layer widths are non-increasing. The implementations for this paper are available through the following link: https://github.com/temp-repo-259848/Regions-ReLU-Network
COMPASS: Composable Policy-Amortized Structured Search for LLM-Based Optimization Modeling
Nguyen Le Minh Hoang ⋅ Van D Cuong ⋅ Bui Trong Duc ⋅ Huynh Thi Thanh Binh
Large language models offer a promising technique for translating natural-language optimization problems into solver code, but their reliability remains constrained by the quality of template retrieval. While recent structured methods improve over flat prompting through hierarchical taxonomy search, their traversal rules remain fixed or LLM-driven at inference time, leaving retrieval unable to improve from past successes and failures. This motivates treating template retrieval as a policy that can be learned from modeling experience, rather than as a static inference-time rule. We propose COMPASS, a policy-amortized framework for retrieving optimization templates in LLM-based solver modeling, which represents the template library as a directed acyclic graph and formulates retrieval as a sequential decision process over template nodes. A deep Q-network learns a reusable traversal policy from template-level judgments and execution feedback, reducing reliance on fixed heuristic traversal through feedback-driven graph navigation. To support extensible retrieval, COMPASS embeds problem descriptions and template nodes in a shared semantic space, allowing newly inserted templates to be scored from their text descriptions without redefining the action space. Experiments across multiple LLM backbones and optimization-modeling benchmarks demonstrate that COMPASS consistently improves over heuristic structured retrieval, with the largest gains on complex and mixed-category problems.
CompJudge: Fine-Grained Comparative Evaluation using Multimodal LLM for Subject-Driven Generation
Nam Hyeon-Woo ⋅ Wenbin Ouyang ⋅ Ciprian A Corneanu
Subject-driven generation (SDG) evaluation using multi-modal LLMs (MLLMs) provides better interpretability and stronger alignment with human judgment than traditional embedding-based methods. However, MLLM-as-a-Judge scores are limited in differentiating between error and error-free SDG images. In this paper, we propose CompJudge, which overcomes these limitations by introducing comparative capabilities to MLLM-as-a-Judge and recalibrating scores based on comparison results. To rigorously evaluate the correctness of SDG evaluators, we introduce CompILIAS, a diagnostic benchmark of triplets (reference, identity-preserving image, identity-degraded image) with verified binary ground-truth labels indicating whether the SDG image preserves the reference identity. We demonstrate improvements of SDG evaluation over existing approaches: 8.9pp on CompILIAS and 10.62--15.98pp on the DreamBench++ dataset. These improvements consistently generalize across diverse MLLM backbones and across subject categories.
Concept frustration: Aligning human concepts and machine representations
Christopher R. S. Banerji ⋅ Enrico Parisini ⋅ Christopher J Soelistyo ⋅ Ahab Isaac ⋅ Alessandro Barp
Aligning human-interpretable concepts with the internal representations of modern machine learning systems remains a central challenge for interpretable AI. We introduce a geometric framework to compare supervised human concepts with unsupervised representations derived from foundation models. We formalise concept frustration: a mismatch that arises when an unobserved concept induces relationships between known concepts that cannot be made consistent within an existing ontology. We develop task-aligned similarity measures that detect this phenomenon, and show that frustration is identifiable in task-aligned geometry while conventional Euclidean comparisons fail. Under a linear–Gaussian generative model, we derive a closed-form expression for Bayes-optimal concept-based classifier accuracy, decomposing predictive signal into known and unknown contributions and identifying where frustration impacts performance. Experiments on synthetic data and real language and vision tasks demonstrate that frustration is present in foundation model representations and that incorporating missing concepts reorganises learned representations to better align human and machine reasoning. These results provide a principled framework for diagnosing incomplete concept ontologies and improving alignment in interpretable AI systems.
Concise Reasoning Through the Lens of Lagrangian Optimization
Chengqian Gao ⋅ Haonan Li ⋅ Taylor Killian ⋅ Jianshu She ⋅ Renxi Wang ⋅ Liqun Ma ⋅ Jorge (Zhoujun) Cheng ⋅ Shibo Hao ⋅ Zhiqiang Xu
Concise reasoning in large language models (LLMs) seeks to generate only essential steps needed to arrive at a final answer, thereby alleviating issues of overthinking. Most proposed approaches scalarize length and reward into a single objective, requiring coefficients or thresholds that must be re-tuned across domains and model scales. We address this brittleness by treating concise reasoning as a constrained problem that minimizes length subject to an accuracy floor, and deriving a tractable algorithm, Performance-Aware Length Update (PALU), which replaces each intractable update of Lagrangian optimization with a tractable surrogate while preserving its structural prescription. On DeepSeek-R1-Distill-Qwen-1.5B, PALU cuts generation length by 64\% and improves accuracy from 44.2\% to 51.6\% across six benchmarks, a level matched by GRPO only with 2.5× more tokens. The Lagrangian dynamics yields a two-phase compression: an initial phase where the residual signal still allows compression without performance cost, followed by a trade-off phase where performance gates further length reduction. The same hyperparameters transfer across math, logic and STEM, and across 1.5B to 14B models, suggesting that constrained optimization is a productive design scaffold for training reasoning LLMs.
CondenseVLA: Learnable History Condensation for Efficient Multi-Frame VLA
Feiyang Hong ⋅ Yujie Wei ⋅ Bo Zhao ⋅ Xiu Su ⋅ Hongxun Yao ⋅ Shuo Yang
Robot manipulation often depends on short-term cues—such as brief contacts, temporary occlusions, and intermediate object states—that are ambiguous in a single frame and can only be resolved by integrating observations over time. However, injecting historical information solely by enlarging the history window is computationally costly and frequently yields limited gains due to substantial redundancy and task-agnostic, inflexible sampling of observations. We propose CondenseVLA, a multi-frame VLA policy with a lightweight learnable history condensation module for using short-term visual context in action decoding. Given a sequence of cached per-frame features from the VLM backbone, CondenseVLA distills observation history into a fixed set of slot tokens using learnable queries and multi-head cross-attention. Simple auxiliary regularization encourages diverse attention patterns by penalizing overlap among slot attention maps and reduces slot 12 redundancy by decorrelating slot embeddings, stabilizing learned condensation without imposing strong time priors. By expanding the history window and distilling it into a compact token set, CondenseVLA improves over the baseline on the evaluated Simpler-WidowX tasks, increasing success rate from 58.4 to 70.8 (+12.4), achieves 97.2% on LIBERO, and reaches 52.0% on real-world tasks. Beyond the gains, our results show that long history windows contain substantial redundancy, so naively providing more frames yields limited benefits, whereas learned condensation into fixed slot tokens provides an effective bounded-token way to leverage temporal context.
Conditional Evaluation of Language Models with Cheap Auxiliary Signals
Zhi Zhang ⋅ Lingfeng Lyu ⋅ Yue Kang ⋅ Doudou Zhou
Aggregate accuracy hides where models succeed and fail. Estimating conditional performance profiles from gold labels alone is expensive, while cheap auxiliary signals such as LLM-judge scores, pairwise comparisons, confidence scores, and judge-disagreement features can be collected for every benchmark item but are often biased or miscalibrated. We propose \method{} (Local Augmented Control-Variate Evaluation), a semi-supervised estimator for conditional LLM evaluation. The key step is local centering: after subtracting the conditional mean of a cheap signal within the target profile region, any linear augmentation has zero conditional mean and therefore cannot change the estimand. The augmentation coefficient is used only for efficiency, and a local ridge control variate combines a gold-label residual mean from the labeled subset with a cheap-signal mean from the full item pool. We prove calibration-free identification, unbiasedness for grouped profiles, local oracle optimality within centered linear augmentations, and first-order adaptivity to the estimated coefficient. The resulting gain formula depends on a local $R^2$ that can be estimated from the available labeled subset and cheap signals, diagnosing where cheap signals are useful. We extend the same principle to direct paired model gaps and deployment-weighted scores, and validate it on MATH-500, ScienceQA, MMLU-Pro, and GPQA.
Consistency-Preserving Concept Erasure via Unsafe–Safe Pairing and Directional Fisher-weighted Adaptation
Yongwoo Kim ⋅ Sungmin Cha ⋅ Hyunsoo Kim ⋅ Jaewon Lee ⋅ Donghyun Kim
With the increasing versatility of text-to-image diffusion models, the ability to selectively erase undesirable concepts (e.g., harmful content) has become indispensable. However, existing concept erasure approaches primarily focus on removing unsafe concepts without providing guidance toward corresponding safe alternatives, which often leads to failure in preserving the structural and semantic consistency between the original and erased generations. In this paper, we propose a novel framework, PAIRed Erasing (PAIR), which reframes concept erasure from simple removal to consistency-preserving semantic realignment using unsafe–safe pairs. We first generate safe counterparts from unsafe inputs while preserving structural and semantic fidelity, forming paired unsafe–safe multimodal data. Leveraging these pairs, we introduce two key components: (1) Paired Semantic Realignment, a guided objective that uses unsafe–safe pairs to explicitly map target concepts to semantically aligned safe anchors; and (2) Fisher-weighted Initialization for DoRA, which initializes parameter-efficient low-rank adaptation matrices using unsafe–safe pairs, encouraging the generation of safe alternatives while selectively suppressing unsafe concepts. Together, these components enable fine-grained erasure that removes only the targeted concepts while maintaining overall semantic consistency. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art baselines, achieving effective concept erasure while preserving structural integrity, semantic coherence, and generation quality.
Contextual Flow Matching for High-quality Visual Content Generation
Bozhi Luan ⋅ Hong Chen ⋅ Jianmin Bao ⋅ Liang Hou ⋅ Yuan Gao ⋅ Xin Tao ⋅ Pengfei Wan ⋅ Wengang Zhou ⋅ Houqiang Li
Standard Flow Matching (FM) models typically employ a pointwise loss function, effectively assuming that the velocity vector at each spatial position is pointwise decomposed supervision given the conditioning. We argue that this Pointwise Independence Assumption is suboptimal for image generation, where strong local correlations exist across space. Neglecting these dependencies leads to artifacts such as structural inconsistency. In this paper, we introduce Contextual Flow Matching (CFM), a generalized framework that enforces consistency not just on individual pixels, but on the contextual relationships within the velocity field. We formulate the training objective using a set of generalized Context Operators---akin to convolution kernels---that abstract relationship constraints. Specifically, we instantiate a Differential Operator to capture high-frequency motion dynamics and an Aggregate Operator to stabilize low-frequency structural evolution. Extensive experiments on ImageNet demonstrate that CFM significantly accelerates convergence and enhances generation fidelity by explicitly modeling spatial dependencies.
Context Value Informed In-Context Reinforcement Learning
Wenhao Zhang ⋅ Shao Zhang ⋅ Xihuai Wang ⋅ Yang Li ⋅ Ying Wen
In-context reinforcement learning (ICRL) has emerged as an effective paradigm for test time adaptation to unseen tasks without parameter updates. However, existing ICRL methods can exhibit brittle and unstable adaptions, and the mechanisms underlying such adaptions remain poorly understood. We provide a policy-gradient view of ICRL and argue that relying on trajectory-level return feedback can lead to high-variance updates when test time rollout is limited. Motivated by the actor--critic principle, we propose Context Value Informed ICRL (CV-ICRL), which equips the in-context policy with an explicit context value: the expected discounted return conditioned on the current state and accumulated interaction context. CV-ICRL trains a value head and writes its prediction back into the context as a value token, enabling TD-style return targets for lower-variance policy updates. Experiments on the Dark Room, Minigrid, and Procgen testbeds show that CV-ICRL substantially improves the stability of test time adaptation and achieves higher returns across tasks and environments. The source code and data of this paper are available at https://anonymous.4open.science/r/CV-ICRL-D161.
Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments
Parth Asawa ⋅ Christopher M Glaze ⋅ Gabriel Orlanski ⋅ Ramya Ramakrishnan ⋅ Benji Xu ⋅ Asim Biswal ⋅ Vincent Chen ⋅ Frederic Sala ⋅ Matei A Zaharia ⋅ Joseph Gonzalez
Continual learning, the ability of AI systems to improve through sequential experience, has attracted substantial interest, but no high-quality benchmark exists to evaluate it. We introduce Continual Learning Bench (CL-Bench), the first difficult, expert-validated benchmark designed to measure whether LLM-based systems genuinely improve with experience. CL-Bench spans six diverse domains (software engineering, signal processing, disease outbreak forecasting, database querying, strategic game-playing, and demand forecasting), each validated by domain experts and designed so that tasks share a learnable latent structure (codebase layout, disease outbreak dynamics, opponent strategies) that a stateful system can discover online but a stateless one cannot. We evaluate frontier models across several agent architectures, from naive in-context learning (ICL) to dedicated memory systems, introducing a gain metric to isolate learning from prior capabilities. We find that these systems leave headroom for improved continual learning: agents frequently overfit to immediate observations or fail to reuse knowledge across instances, and dedicated memory systems do not fix this---in fact, naive ICL outperforms systems dedicated to memory management. CL-Bench is the first benchmark to evaluate continual learning across diverse real-world domains with expert-validated tasks and isolate online learning from underlying model capability, showing a need for better continual learning systems.
ContinuLoc: Continuous Pose Inference over Neural Fields for UAV Geo-Localization
Weikun Zhou ⋅ Jinliang Lin ⋅ Yu Zang ⋅ Lin Xie ⋅ HaiqingFu ⋅ Sheng Ao ⋅ Shangshu Yu ⋅ Zhiming Luo ⋅ Cheng Wang
Cross-view UAV geo-localization involves three design choices: how to represent the UAV observation, how to represent the geo-referenced scene, and how to infer pose from their interaction. The dominant paradigm organizes the world as a discretized 2DoF candidate library, coupling scene representation to inference and bounding pose accuracy by sampling resolution. We decouple scene representation from pose inference by encoding the geo-referenced world as a continuously queryable semantic neural field, enabling localization to be formulated as explicit optimization over a continuous 4DoF pose space — parameterized by 2D ground position, in-plane heading, and viewing scale. Pose hypotheses are evaluated globally in a shared semantic feature space, bypassing the sampling bottleneck of retrieval. We learn 4DoF pose-sensitive observation representations under full 4DoF supervision provided by ContinuScenes, a multi-city benchmark with dense near-nadir UAV imagery and complete pose annotations. Inference proceeds hierarchically, progressing from broad pose-space coverage to adaptive concentration on high-compatibility regions. Preliminary experiments demonstrate that our method substantially outperforms prevailing localization paradigms, including discrete retrieval and absolute pose regression, across increasingly strict 4DoF pose criteria.
Continuous Expert Assembly: Instance-Conditioned Low-Rank Residuals for All-in-One Image Restoration
Haisen He ⋅ XIANGYU ZOU ⋅ SongLin Dong ⋅ Heng Li ⋅ Yihong Gong ⋅ Zhiheng Ma
Real-world image degradation is often unknown, spatially non-uniform, and compositional, requiring all-in-one restoration models to adapt a single set of weights to diverse local corruption patterns without test-time degradation labels. Existing methods typically modulate a shared backbone with global prompts or degradation descriptors, or route features through predefined expert pools. However, compact global conditioning can bottleneck localized degradation evidence, while static expert routing may produce homogeneous updates or rely on unstable sparse assignments. We propose \textbf{Continuous Expert Assembly} (CEA), a token-wise dynamic parameterization framework for all-in-one image restoration. CEA employs a lightweight \textbf{Cross-Attention Hyper-Adapter} to probe intermediate spatial features and synthesize instance-conditioned low-rank routing bases and residual directions. Each spatial token then assembles its own residual update via dense signed dot-product affinities over the generated rank-wise components, avoiding external prompts, static expert banks, and discrete Top-$K$ selection. The resulting assembly rule also admits a linear-attention perspective, making its dense token-wise routing behavior transparent. Experiments on AIO-3, AIO-5, and CDD-11 show that CEA improves average restoration quality over strong prompt-, descriptor-, and expert-based baselines, with the clearest gains on spatially varying and compositional degradations, while maintaining favorable parameter, FLOP, and runtime efficiency.
Contractive Restoring Flows: Robust Reasoning Distillation via Orbital Stability
Dongqi Zuo ⋅ Yuanyuan Wang ⋅ Chuan Zhou ⋅ Haoxuan Li ⋅ Mingming Gong
Model distillation provides an alternative when computing resources are limited. How to transfer the reasoning ability from teacher model to student model influences the performance of model distillation. We develop a mathematical framework for reasoning distillation grounded in dynamical systems theory. By modeling the teacher model's residual stream as a discrete dynamical system, we define a precise notion of *orbital stability*: perturbations transversal to the teacher's reasoning trajectory should contract, while the tangential reasoning signal propagates without active suppression. We derive the **Contractive Restoring Flow (CRF)** loss from first principles and prove that, at the loss minimum, it achieves strict transversal contraction at a constant rate $1-\alpha$, leaves the tangential dynamics unconstrained (tangential agnosticism), and produces a global basin of attraction with bounded steady-state error within the linearisation regime. Empirically, reasoning-trajectory distillation with the CRF loss outperforms other distillation methods, and shows greater robustness on long-chain reasoning tasks, supporting the predicted effect of contracting transversal error. The code is avaiable at: https://anonymous.4open.science/r/CRF-5945/.
Control First, Robustness Next: Decoupled Representation Learning for Visual RL Generalization
heo chanyong ⋅ Hyelyn Jeong ⋅ Jongchan Park ⋅ Seungjun Oh ⋅ Yusung Kim
Visual reinforcement learning (RL) is promising for real-world problems such as autonomous driving, robotic locomotion, and manipulation, but remains highly vulnerable to visual distribution shifts between training and deployment. Existing methods improve robustness through Q-consistency regularization, masking, and auxiliary objectives, yet most of them follow a coupled representation learning paradigm in which control-oriented and robustness-oriented representation learning are jointly optimized within the same training framework. We argue that this coupled structure can introduce harmful interference between control learning and robustness learning, degrading original-environment control performance and limiting generalization under visual shifts. To address this issue, we propose Separate Then Align Representations (STAR), which reformulates visual RL generalization as a problem of decoupled representation learning. Specifically, we first learn a control-relevant reference representation through online RL under weak augmentation, and then perform an offline post-representation alignment stage that maps strongly augmented inputs back to this reference representation. Experiments on RL-ViGen benchmarks spanning DMControl and Robosuite show that STAR achieves strong generalization under visual distribution shifts while better preserving performance in the original environment than prior coupled approaches. Our source codes are available in the supplementary material.
ControlFlow3D: Distilling Multi-View Knowledge into Latent Flow Matching for Point Cloud Upsampling
Yuang Liu ⋅ Zhi Zuo ⋅ Zhengkai Zhao ⋅ Lirui Zhang ⋅ Pan Gao ⋅ Hao Feng ⋅ Zhengzhe Liu
Point cloud upsampling is essential for 3D applications but remains challenging due to sparse, under-constrained observations. Multi-view images can resolve geometric ambiguity, but existing multi-modal methods do not fully exploit visual cues. We propose ControlFlow3D, a framework that distills knowledge from pretrained foundation models into a lightweight generation network during training, enabling strong performance without multi-view images available at test time. Specifically, our method performs conditional flow matching in a structured latent space defined by a frozen 3D encoder. Multi-view image tokens and 3D latent tokens are jointly clustered into shared prototypes via unsupervised soft KMeans; the resulting prototypes guide the flow trajectory through velocity control and are aligned by a consistency loss that progressively transfers image knowledge into the base network. A Mamba-based geometry decoder then maps latent features to dense coordinates with linear complexity. We also construct ShapeNetPU, a 35K-object multi-modal benchmark that is 31times larger than prior datasets. Extensive experiments across multiple benchmarks demonstrate the effectiveness of our method and its strong generalization capability.
ControlJEPA: Principled Trajectory Regularization via Lyapunov Tube Loss
Khanh T Nguyen ⋅ Yen N Pham ⋅ Uyen N.B. Vo ⋅ Thieu Vo ⋅ Cuong Pham ⋅ Tan Nguyen
Next-token prediction leaves hidden-state trajectories largely unconstrained, motivating recent geometric regularization methods based on angular penalties. However, these approaches control only the direction of motion and provide no explicit bound on trajectory deviation. We introduce the Controlled Joint Embedding Predictive Architecture (ControlJEPA), a trajectory regularization framework grounded in discrete-time Input-to-State Stability theory. Our novel Lyapunov Tube Loss enforces a contraction condition on a normalized measure of transversal deviation, yielding a provable guarantee that confines hidden states within a tube of explicit width determined by two interpretable hyperparameters. ControlJEPA integrates seamlessly with standard training and requires no architectural changes. Across six benchmarks and seven model families, it consistently outperforms standard fine-tuning and prior trajectory regularizers, improving accuracy and data efficiency while preserving output diversity. Notably, the induced geometric structure persists at inference time, indicating that the method reshapes the representation space rather than memorizing training trajectories.
Controllable Dynamic 3D Shape Generation via 3D Trajectories and Text
Jaeyeong Kim ⋅ Inès H Kim ⋅ Jahyeok Koo ⋅ Seungryong Kim
We present T2Mo, the first feed-forward framework for controllable dynamic 3D shape generation of general objects conditioned on 3D trajectories and text prompts. Existing text-driven methods struggle to specify precise motion due to the ambiguity of natural language. To address this, we introduce 3D trajectories as a direct spatial control signal and propose a shape-grounded trajectory embedding that maps arbitrary user-provided trajectories into geometry-aware condition tokens aligned with the input shape. Injecting these tokens into the generative backbone enables highly controllable motion generation while handling varying trajectory numbers and distributions. We conduct extensive comparisons against text-based baselines and trajectory-guided video generation workarounds. Quantitative and qualitative evaluations, along with user studies, show that our method produces motions that more faithfully follow the given prompts with higher expressiveness while preserving motion quality. Our code and weights will be publicly released.
Controllable User Simulation
Guy Tennenholtz ⋅ Ofer Meshi ⋅ Amir Globerson ⋅ Uri Shalit ⋅ Jihwan Jeong ⋅ Craig Boutilier
Using offline datasets to evaluate conversational agents often fails to cover rare scenarios or to support testing new policies. This has motivated the use of \emph{controllable user simulators} for targeted, counterfactual evaluation, typically implemented by prompting or fine-tuning large language models. In this work, we formalize controllable simulation as a causal inference problem. By bridging natural language evaluation with off-policy evaluation methodology, we show that the standard practice of training simulators via supervised fine-tuning on post-hoc trajectory labels yields a structurally biased model. Specifically, these labels are inextricably coupled to the data-generating behavior policy, injecting a \emph{look-ahead bias} that breaks causal consistency. Furthermore, we prove that under policy shift this failure causes the variance of evaluation metrics to explode geometrically, a phenomenon we term \emph{controllability collapse}. To restore causal consistency, we establish theoretical conditions for accurate simulation and propose practical training mitigations: \textit{a priori} controls, step-wise dynamic controls, and direct policy-conditioned learning. Empirical evaluation confirms that while standard global controls distort conversational distributions and collapse behavioral diversity, our causally grounded simulators eliminate look-ahead bias, preserve natural variance, and exhibit robust zero-shot generalization to unseen agent behaviors.
ControlSVG: Exploring Controllable SVG Generation with Autoregressive Models
Qirui Li ⋅ Teng Hu ⋅ Ran Yi ⋅ Paul L Rosin ⋅ Yu-Kun Lai
Autoregressive (AR) models have reformulated Scalable Vector Graphics (SVG) generation as next-token prediction, demonstrating remarkable potential in text-to-SVG tasks. However, controllable SVG generation guided by spatial signals (e.g., sketches or scribbles) remains largely unexplored within AR models. While a natural approach, inspired by controllable image generation, is to adapt methods like Condition Prefilling or Conditional Decoding, they fall short in SVG generation. To address these challenges, we introduce ControlSVG, a versatile and effective framework tailored for integrating precise spatial controls into autoregressive SVG generation. Firstly, we propose a Visual-Feedback Conditioning mechanism that constructs a “See-and-Draw" loop. This design mitigates cross-modal drift and ensures long-range consistency. Secondly, we explore spatially aware learning objectives that explicitly bridge discretized SVG tokens with their absolute 2D spatial coordinates, enhancing robustness against coordinate deviations during sequential generation. Extensive experiments on Canny-Edge-to-SVG generation demonstrate that ControlSVG outperforms existing baselines, achieving superior spatial alignment and high-quality SVG. We further provide preliminary validations on real human-drawn scribbles and extend the framework to scribble-to-SVG and image-to-SVG settings, suggesting its potential applicability beyond Canny-based control.
Cool Graphs: Active Property Search Towards Quantum Nano-Refrigerators
Sofia Blyufer ⋅ Amir Kleiner ⋅ Uri Peskin
Nanodevices that absorb energy from their thermal environment can act as quantum refrigerators. Finding suitable configurations, however, is a needle-in-a-haystack problem: the many-body state space scales exponentially with the size of the device, where each candidate configuration requires an expensive quantum-chemical computation. Under random sampling of material parameters, only 2-9\% of configurations exhibit the desired property. We introduce a search framework that couples a Graph Variational Autoencoder (GVAE) with an uncertainty-driven Active Learning (AL) loop to efficiently identify energy-absorbing configurations in electronic junctions with a 2D "working material" to be optimized. The GVAE maps physical parameters - site energies, coupling matrices, temperature - into a smooth latent space that distinguishes cooling from dissipative parameter regimes, achieving a classification ROC-AUC of 0.89 on held-out topologies and 0.80 on unseen extrapolation topologies. An AL pipeline based on Monte Carlo dropout then targets the decision boundary, iteratively selecting the candidates where the model is least certain for simulation. Over 32 AL cycles, the uncertainty-driven strategy discovers positive-flux (energy absorbing) systems at rates consistently exceeding the random baseline across all tested topologies, providing a data-efficient pathway for navigating a combinatorically large design space. The learned latent space organizes configurations along physically interpretable axes, that correlate with underlying microscopic model parameters. Our framework, thus, paves a new way towards effective cooling of computational devices, to reduce the environmental footprint of high performance computing.
Coordination Connectivity: Shared Initialization Shapes the Joint-Policy Landscape in MARL
Yixiang Fan ⋅ Zhiqiang Pu ⋅ Hao Ma ⋅ Dongmin Li ⋅ Weihao Zhang ⋅ Yinxing Dai ⋅ Shijie Zheng ⋅ Jingjing Huang
Mode connectivity studies reveal that high-performing neural networks are often connected by low-loss paths in parameter space. We investigate the analogue of this phenomenon in cooperative multi-agent reinforcement learning (MARL), where a solution is not a single model but a team of decentralized policies. We introduce coordination connectivity, a diagnostic evaluating whether linear interpolation between the actor parameters of two trained policy teams preserves high team return. The resulting coordination barrier measures whether two successful teams are separated by a performance cliff along the joint-policy path. We evaluate this diagnostic on the StarCraft Multi-Agent Challenge (SMAC) across five controlled training histories, isolating the effects of shared initialization, continued training, parameter decoupling, and HAPPO-style separate-actor specialization. Across three SMAC maps and four seeds, policy teams derived from a shared MAPPO checkpoint exhibit near-zero interpolation barriers, even after non-shared specialization. In contrast, interpolation paths involving independently trained HAPPO teams exhibit substantially larger barriers. Mixed-team assembly, single-agent hot-swap, clone-all, role-swap, and weight-space probes further indicate that shared-initialized agents specialize into distinct roles while remaining within a low-barrier coordination component. These results suggest that shared initialization implicitly structures the joint-policy space of cooperative MARL, providing a mode-connectivity-inspired lens for analyzing coordination, specialization, and same-source cross-version team compatibility.
CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents
Mingwei Zheng ⋅ David OBrien ⋅ Siwei Cui ⋅ Pardis Pashakhanloo ⋅ RAJDEEP MUKHERJEE ⋅ Myeongsoo Kim ⋅ Sachit Kuhar
LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-making. However, the conventional append-only trajectory architecture found in practice tightly couples file-read actions with their observations, capturing snapshots that become permanently fixed in the chronological history. As files change through agent edits or concurrent human modifications, these snapshots become stale, causing reasoning errors and causing agents to redundantly re-read files, with each re-read appending yet another copy to the trajectory. To mitigate this, we propose CORVUS, a novel trajectory architecture that decouples file-read actions from their observations by maintaining a synchronized registry of relevant files and injecting only their current contents at each reasoning cycle. This structural change produces significantly lighter-weight trajectories that remain synchronized with the actual codebase state by construction, eliminating redundant file copies and stale snapshots that bloat conventional trajectories. We evaluated CORVUS on SWE-POLYBENCH_VERIFIED and SWE-BENCH PRO across four LLMs, achieving 9–50% reduction in average input tokens per task, 15–32% shorter final prompts, and up to 37% fewer reasoning cycles while maintaining comparable pass rates.
CoT-Guard: Small Models for Strong Monitoring
Nirav Diwan ⋅ Han Wang ⋅ Berkcan Kapusuzoglu ⋅ Ramin Moradi ⋅ Supriyo Chakraborty ⋅ Giridharan Iyengar ⋅ Sambit Sahu ⋅ Huan Zhang ⋅ Gang Wang
Monitoring the chain-of-thought (CoT) of reasoning models is a promising approach for detecting covert misbehavior (i.e., hidden objectives) in code generation tasks. While large models (GPT-5, Gemini-3-Flash) can serve as effective CoT monitors, they are expensive to deploy due to the lengthy reasoning traces and high API cost, emphasizing the need for smaller, cheaper alternatives. Nevertheless, we find that current small models (4B--8B) struggle to detect hidden objectives despite access to the CoT, frequently misattributing them as part of the user query. To address this, we propose a post-training pipeline combining supervised fine-tuning (SFT) and reinforcement learning (RL), where SFT narrows the gap for in-domain tasks by distilling detection behavior from stronger monitors, and RL on hard and subtly crafted hidden objectives helps the model generalize to out-of-domain monitoring tasks. To validate this generalization, we evaluate under a realistic threat model motivated by practical supply-chain attacks, where the adversary is a third-party LLM router injecting hidden objectives into code-generation requests through either prompt manipulation or code manipulation attacks. To push beyond objectives that large monitors already saturate, we also introduce four new challenging tasks even for strong monitors. Finally, we introduce CoT-Guard, a 4B-parameter monitor that demonstrates superior generalization performance under both prompt and code manipulation attacks, achieving a $\text{Gmean}^2$ (i.e., TNR$\times$TPR) of 75\% and outperforming GPT-5.4 (56\%), GPT-5-mini (41\%), and Qwen3-32B (54\%), while closing the gap to Gemini-3-Flash (83\%). These results demonstrate that \modelname provides a practical and cost-effective user-side defense, substantially improving hidden-objective detection while avoiding the deployment cost of large monitors.
Counterfactual Estimation under Composite Treatments via Progressive Distribution Alignment
Abhirup Mondal ⋅ Anirban Majumder ⋅ Vineet Chaoji
Estimating counterfactual outcomes when multiple treatments are applied simultaneously, the composite treatment setting, is fundamentally harder than the single-cause case: the treatment space grows combinatorially, most combinations are unobserved, and confounding is entangled across dimensions. We propose PACE (Progressive Alignment for Composite Effects), which exploits a known causal DAG over treatment variables to progressively transform the observational data distribution into the target interventional distribution. PACE performs K sequential augmentation steps in reverse topological order, each isolating a single treatment component and conditioning only on its parents in the DAG. We derive the closed-form optimal augmentation density minimizing KL divergence at each step, establish finite-sample bounds prescribing the augmentation budget via DAG-aware selection bias ratios, and show that the DAG provides a deterministic processing order that eliminates the need for random permutation averaging. On semi-synthetic benchmarks with real covariates (IHDP, Twins, Hillstrom) and controlled synthetic environments, PACE substantially outperforms seven baselines on both continuous and binary outcomes, with the largest gains on rare treatment combinations where confounding is most severe. Ablations confirm consistent improvements over the closest prior method across varying DAG topologies, confounding strengths, and under DAG misspecification.
CounterFlowNet: From Minimal Changes to Meaningful Counterfactual Explanations
Oleksii Furman ⋅ Patryk Marszałek ⋅ Jan Masłowski ⋅ Marek Śmieja ⋅ Maciej Zieba ⋅ Piotr Gaiński
Counterfactual explanations (CFs) provide human-interpretable insights into model's predictions by identifying minimal changes to input features that would alter the model's output. However, existing methods struggle to generate multiple high-quality explanations that (1) affect only a small portion of the features, (2) can be applied to tabular data with heterogeneous features, and (3) are consistent with the user-defined constraints. We proposeCounterFlowNet, a generative approach that formulates CF generation as sequential feature modification using conditional Generative Flow Networks (GFlowNet). CounterFlowNet is trained to sample CFs proportionally to a user-specified reward function that can encode key CF desiderata: validity, sparsity, proximity and plausibility, encouraging high-quality explanations. The sequential formulation yields highly sparse edits, while a unified action space seamlessly supports continuous and categorical features. Moreover, actionability constraints, such as immutability and monotonicity of features, can be enforced at inference time via action masking, without retraining. Experiments on eight datasets under two evaluation protocols demonstrate that CounterFlowNet achieves superior trade-offs between validity, sparsity, plausibility, and diversity with full satisfaction of the given constraints.
Coupling-Aware Reinforcement Learning for Co-Evolving Graph Games
Mina Kim ⋅ Guanghui Lan ⋅ Benoit Montreuil
Building hydrogen station networks, electric vehicle (EV) charging grids, or vaccine distribution systems requires coordinating two separately operated graphs - demand-side facilities and the supply-side networks they depend on - where each operator's payoff depends on the other's choices. We formalize this as the Co-Evolving Graph Game (CEGG), a setting that sits between facility location, interdependent-network analysis, and graph-based multi-agent reinforcement learning (MARL) but is not jointly addressed by any of them. We propose Cross-Graph Attention with Policy Mirror Descent (CGA-PMD), pairing cross-graph attention - which lets each agent read only the partner-graph nodes structurally coupled to its own - with a policy mirror descent update for stable training. Across three reward domains and a range of practical network scales, we find that the structure of the coupling selects which method wins: at small scale most methods are competitive, but as scale grows or coupling becomes irregular, only CGA-PMD remains productive while own-only, unmasked-attention, and Proximal Policy Optimization (PPO)-based variants all fail in distinct ways. Our theoretical analysis explains why: PMD's Kullback-Leibler (KL) regularization admits an iteration-uniform cumulative drift bound while PPO's clip update does not, predicting CGA-PMD's stability in the hard regime where PPO-based methods collapse.
Crafter: Towards Automated Reproducible Machine Learning via Agentic Code Generation
Fangru Linghu ⋅ Jackson R Ye ⋅ Jieying Wang ⋅ Alexandre V Morozov ⋅ Ian Foster ⋅ Zhao Zhang
Reproducing machine learning research is challenging because many papers do not provide executable code, and key implementation details are often scattered, implicit, or missing. Existing paper-to-code systems improve over direct prompting, but the generated repositories can still miss paper-critical logic or fail at execution time. We propose Crafter, an automated paper-to-code pipeline that treats reproduction as a problem of context-calibrated implementation recovery rather than single-pass code generation. Crafter builds an evidence-grounded specification, resolves missing implementation details before planning, and repairs the generated repository through execution-aware debugging. We evaluate Crafter with PaperBench Code-Dev for implementation faithfulness and a self-designed execution-oriented benchmark for runtime readiness. Across 23 PaperBench papers using the Claude Sonnet 4.6 backend, Crafter achieves an average Code-Dev score of 0.84, improving over Paper2Code by 19.7% and over DeepCode by 22.1%. In execution evaluation, Crafter reaches the contribution-level milestone on all three repositories, compared with two for both baselines, while reducing average repair cycles from 11.0--11.3 to 5.0.
CREF: Forecasting Benchmarks for the Age of Agents
Andreas Auer ⋅ Abdul Fatir Ansari ⋅ Oleksandr Shchur ⋅ Xiyuan Zhang ⋅ Boran Han ⋅ Pedro Mercado ⋅ Prateek Desai ⋅ Yuyang (Bernie) Wang ⋅ Michael Bohlke-Schneider
Traditional forecasting evaluations relying on static historical data are fundamentally incompatible with modern AI advances. Even for standard time series foundation models, temporal overlap between training and test periods can lead to implicit leakage through correlated signals. This problem is intensified by LLMs, which can bridge semantic gaps between forecasted and related data while memorizing extensive world knowledge. At the extreme, agentic models with tool access can retrieve the ground truth directly. Such information leakage inflates performance relative to real forecasting scenarios and is impossible to rule out in historical evaluations. We introduce CREF, a real-time benchmark that evaluates forecasting models on data that does not yet exist, eliminating all forms of leakage by design. CREF balances three competing objectives for real-time benchmarks: diversity, robustness, and throughput. It draws data from 32 public APIs and features 50 forecasting tasks with diverse data frequencies, multivariate targets, and covariates, making it the first real-time benchmark reaching the breadth of leading static benchmarks. Retrospective evaluation validates that CREF produces rankings consistent with established static benchmarks, while real-time evaluation reveals qualitative and quantitative performance gaps for agentic and LLM-based models, confirming leakage on historical data. CREF runs continuously and is open for new models to join, offering a public leaderboard for contamination-free model comparison.
Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets
Rafi Al Attrach ⋅ Rajna Fani ⋅ Sebastian Lobentanzer ⋅ Joan Giner-Miguelez ⋅ DEBANSHU DAS ⋅ Varuni H K ⋅ Nobin Sarwar ⋅ Rajat Ghosh ⋅ Anwai Archit ⋅ Surbhi Motghare ⋅ Christina C Parry ⋅ Luis Oala ⋅ Lara Grosso ⋅ Joaquin Vanschoren ⋅ Steffen Vogler ⋅ Sujata Goswami ⋅ Eric Rosenthal ⋅ Marzyeh Ghassemi ⋅ Matthew McDermott ⋅ Tom Pollard
Croissant has emerged as the metadata standard for machine learning datasets, providing a structured, JSON-LD-based format that makes dataset discovery, automated ingestion, and reproducible analysis machine-checkable across ML platforms. Adoption has accelerated, and NeurIPS now requires Croissant metadata in every submission to its dataset tracks. Yet in practice Croissant generation usually starts with uploading data to a public platform, a path infeasible for governed and large local repositories that hold much of the high-value data ML increasingly relies on. We release Croissant Baker, a local-first, open-source command-line tool that generates validated Croissant metadata directly from a dataset directory through a modular handler registry. We evaluate Croissant Baker on over 140 datasets, scaling to MIMIC-IV at 886 million rows and 374 Parquet files. On held-out comparisons against producer-authored or standards-derived ground truth, Croissant Baker reaches 97–100% agreement across multiple domains.
Cross-Channel Agreement Beats Consensus: Compositional Verification for Geometry Reasoning
ShaoWei Huang ⋅ Zefei Gao ⋅ Yuting Yu ⋅ Hao Yang ⋅ ZiYu Yang ⋅ Mingchi Sun ⋅ Yu Guo
Multimodal geometry reasoning requires models to jointly perceive diagrams and perform symbolic derivations. Test-time verification methods such as majority voting assume voter errors are roughly independent, but in geometry this assumption breaks down: all candidates perceive the same diagram, so perceptual biases cascade into correlated errors that produce confident but wrong consensus. We observe that neuro-symbolic candidates contain compositional structure that can serve as a verification resource. We introduce SKETCH, a format decomposing each candidate into a neural pathway (natural-language claim) and a symbolic pathway (visual grounding → typed DSL → deterministic execution). These pathways are governed by different error processes—symbolic errors are correlated across candidates through shared perception, while neural errors are more dispersed. Our method, Compositional Consensus Verification (CCV), gates consensus on cross-channel agreement: because the two pathways fail for different reasons, accidental agreement is rare, making each concordant vote far more precise than a raw majority vote. Across four geometry benchmarks, CCV achieves 81.0% accuracy (+10.4 pp over majority voting) without external scorers or additional inference. Seven non-compositional baselines—including trained PRMs—cluster at 70–71%, while compositional cross-layer features jump to 79–81%, confirming that structural diversity is the core verification resource, and that a principled gating algorithm effectively unlocks its potential.
Cross-Modal Prior-Guided Training with Visual Foundation Models for Unsupervised LiDAR Point Cloud Registration
KeZheng Xiong ⋅ Shiyun Xu ⋅ Sheng Ao ⋅ Mingming Wang ⋅ Siqi Shen ⋅ Chenglu Wen ⋅ Cheng Wang
Visual Foundation Models (VFMs) have demonstrated strong capabilities in visual-geometric tasks such as image matching and 3D reconstruction. Recent work has revealed the potential of VFMs in label-free indoor RGB-D registration, where the visual modality is intrinsic to the task. A natural question follows: can this success extend to more challenging outdoor scenarios characterized by sparse, irregular LiDAR data with only partial visual coverage? We answer this with PROMOTER, a novel teacher–student framework for unsupervised LiDAR registration that bridges visual priors and sparse LiDAR geometry. To leverage pre-trained geometric priors for 3D registration, we must disentangle them from confounding signals introduced by modality and task gaps under partial visual coverage. To this end, we introduce a lightweight Adaptive Prior Integrator that extracts and integrates these priors into an enhanced teacher, together with a Perceptual-Geometric Labeler that produces high-quality pseudo-labels. For student training, we propose Matchability-Anchored Training to improve robustness under noisy supervision. Experiments on KITTI and nuScenes demonstrate state-of-the-art performance and improved scalability across registration models and VFMs, while incurring no additional parameters at inference. Code will be released.
CTM-AI: A Blueprint for General AI Inspired by a Model of Consciousness
Haofei Yu ⋅ Yining Zhao ⋅ Lenore Blum ⋅ Manuel Blum ⋅ Paul Liang
Despite remarkable advances, today's AI systems remain narrow in scope, falling short of the flexible, adaptive, and multisensory intelligence that characterizes human capabilities. This gap has fueled longstanding debates about whether AI might one day achieve human-like generality or even consciousness, and whether theories of consciousness can inspire new architectures for AI. This paper presents an early blueprint for implementing a general AI system, CTM-AI, combining the Conscious Turing Machine (CTM), a formal machine model of consciousness, with today's foundation models. CTM-AI contains an enormous number of powerful processors ranging from specialized experts (e.g., vision-language models and APIs) to unspecialized general-purpose learners poised to develop their own expertise. Crucially, for whatever problem must be dealt with, information from many processors is selected, integrated, and exchanged appropriately to solve the task.CTM-AI achieves state-of-the-art accuracy on MUStARD (72.28) and UR-FUNNY (72.13), outperforming multimodal and multi-agent frameworks. On tool-using and agentic tasks, CTM-AI achieves 10+ points of improvement on StableToolBench and WebArena-Lite. Overall, CTM-AI offers a principled, testable blueprint for general AI inspired by a model of consciousness.
CuBic: Curvature-Driven Dynamic Inference Caching for Fast, High-Fidelity Flow Matching
Yuyang Chen ⋅ Linqian Zeng ⋅ Yijin Zhou ⋅ Hengjie Li ⋅ Jidong Zhai
Flow Matching (FM) has emerged as the dominant paradigm for state-of-the-art image and video generation, yet its reliance on full attention over lengthy sequences incurs significant latency. While existing caching methods reduce redundant computations, they yield suboptimal performance, fundamentally constrained by inflexible spatial granularity and empirical, model-specific heuristics. To bridge this gap, we ground the caching mechanism in the geometric properties of the FM velocity field. We theoretically establish that the \textbf{curvature} inherently dictates the coupling between trajectory fluctuation and computational requirements, thereby providing a rigorous metric to quantify caching staleness loss and automatically derive adaptive granularity and optimal update intervals. Translating this theoretical insight into practical acceleration, we present CuBic, a train-free caching framework. It bypasses expensive curvature computations via a zero-overhead proxy signal that dynamically guides spatial partitioning and budget-constrained update scheduling. At the system level, a sparse forward engine translates these optimal policies into efficient runtime, ensuring that theoretical guarantees directly translate into actual FLOP savings and linear speedup. Being model-agnostic and virtually hyperparameter-free, CuBic consistently outperforms state-of-the-art caching baselines in quality at customizable acceleration ratios (averaging 2.5$\times$) across Wan2.1, Flux.1, and FireRed-Image-Edit.
CURE: Counterfactual Unsafe-token Re-masking for Diffusion Large Language Model Test-time Alignment
Lichao Wang ⋅ Mingquan Zhang ⋅ Xian Zhang ⋅ Juntao Dai ⋅ Chi Liu ⋅ Yaodong Yang ⋅ Jingwei Yi
Diffusion language models (DLMs) generate text through iterative denoising, enabling parallel decoding, bidirectional conditioning, and editable intermediate states. Under harmful prompts, however, intermediate denoising states can contain risk-inducing tokens that are seemingly benign but increase the probability of an unsafe final response by shaping how the remaining masked positions are completed. Since committed tokens in DLMs can still be re-masked and rewritten, safety control can be applied during denoising by removing a small set of risk-inducing tokens before they drive subsequent generation toward unsafe responses. Existing DLM defenses mainly rely on model-level alignment, trajectory-level detection, or block-level repair, often requiring training cost, extra inference passes, or coarse regeneration. We propose CURE, a counterfactual value-driven test-time alignment method that keeps the base DLM fixed and selectively re-masks risk-inducing tokens. CURE trains a time-conditioned safety value model to estimate whether a partially denoised state will lead to an unsafe final response. During inference, CURE constructs counterfactual readout views that remove or isolate targeted tokens, estimates their contribution to future unsafe response, and re-masks only high-risk tokens for later rewriting. Across three DLMs and five jailbreak benchmarks, CURE reduces macro-average ASR from 47.20\% to 1.36\%, preserves utility on MMLU and GSM8K, and adds only 0.84 extra TFLOPs, achieving a strong safety-utility-efficiency trade-off. Code is available at \url{https://anonymous.4open.science/r/CURE}.
Direct Preference Optimization is a widely used approach for safety alignment, aiming to reduce harmful behaviors in large language models. However, prior work shows that it can be brittle and exhibits poor out-of-distribution (OOD) generalization \citep{qi2025safety}. To this end, in this paper we investigate whether curriculum learning can improve the robustness of DPO-based safety alignment. We propose \textbf{Staged-Competence}, a curriculum-based framework that organizes preference data by difficulty, employs competence-based sampling, and progressively updates the reference model during training. Averaged across three model families, Staged-Competence reduces OOD harmful response rates by 16\% and jailbreak attack success rates by 20\%, while preserving general capabilities and maintaining near-zero over-refusal. We further show that Staged-Competence (1) matches baseline safety with only 75\% of the training data, demonstrating improved data efficiency and (2) yields better separation between safe and unsafe responses. Staged-Competence is agnostic to the underlying policy optimization loss and can extend to other DPO variants and alignment domains other than safety. Our code and data can be found at: \url{link/upon/acceptance}.
CUVET: A Partitioning Approach for Continuous Treatment Assignment At Scale
Artem Betlei ⋅ Mariia Vladimirova ⋅ Victor Girou ⋅ Thibaud Rahier
Treatment assignment problems arise wherever limited budget must be allocated to heterogeneous users, with applications ranging from personalized recommendations to online advertising and healthcare. In such settings, individuals exhibit heterogeneous responses to different treatments, making it essential to learn cost-aware personalized treatments. This paper introduces the Cost per Unit Value Equalization Tree (CUVET) algorithm, a novel treatment assignment approach that partitions the user space. Under a diminishing-returns (power-law) assumption, CUVET solves the within-cohort allocation problem by equalizing the marginal cost per unit value across each user group, yielding a closed-form cost-aware treatment assignment suited for large-scale industrial deployment. We also release CUVET-policy, an 86.7-million-impression public benchmark derived from real-world industrial A/B tests, providing an open-source evaluation framework for decision-focused learning. Across a multi-method benchmark suite, the Bayesian variant of CUVET achieves the largest cost-feasible value uplift on the public MT-LIFT dataset +11.1% under LP, vs. at most +1.3% for all baselines, while on CUVET-policy, where treatment effects are sub-percent, CUVET variants are the most efficient and strongest cost-feasible methods.
CWAGraph: Retrieving What Was Never Explicitly Identified in Graph-Based RAG
Zhitao Yin ⋅ Ruiheng Zhu ⋅ Ganlin Xu ⋅ Hongru Hou ⋅ Jiaqing Liang ⋅ Deqing Yang
Unlike chunk-based Retrieval-Augmented Generation (RAG) methods which retrieve information only from text passages, graph-based RAG methods improve their performance through entity-relation graphs. However, existing graph-based RAG methods typically represent graph facts under the Open World Assumption (OWA), where an absent edge (relation) means unknown rather than false. This principle is problematic for many closed corpora including technical manuals and regulations, where rules are generally stated through scope-bounded statements (such as all X, except Y or only Y) named as local completeness statements(LCSs) in this paper. LCS apply rules to every unmentioned in-scope entity which cannot be explicitly and completely represented by the OWA graph. To address this problem, we propose CWAGraph, which is a graph-based RAG framework augmented with a Closed World Assumption (CWA) layer representing local scope and boundaries for LCSs. The CWA layer identifies which entities are governed by each statement and makes the corresponding closed-world evidence retrievable from those entities, thus enhancing the performance of question answering (QA) involving LCS (called as LCS resolution in this paper). To better evaluate RAG systems' performance on LCS resolution, we further introduce a new QA benchmark SetQA consisting of 493 QA instances involving LCSs across two source domains and three difficulty levels. Our experiments on SetQA reveal that CWAGraph consistently outperforms chunk-based and graph-based RAG baselines.
CyCLeGen: Cycle-Consistent Layout Prediction and Image Generation
Shan Xiaojun ⋅ Haoyu Shen ⋅ Yucheng Mao ⋅ Haiyang Xu ⋅ Xiang Zhang ⋅ Abhay Anand ⋅ Bingnan Li ⋅ Zhuowen Tu
We present CyCLeGen, an autoregressive framework that integrates layout understanding and layout-to-image generation through cycle consistency: predicted layouts must produce faithful images, and generated images must yield consistent layouts. We enforce this constraint via CycleGRPO, a bidirectional reinforcement learning strategy with complementary geometric and perceptual rewards, enabling self-introspective learning from only 8k RL samples. This creates a natural loop, generation helps understanding by rewarding only those layouts that lead to high-quality images, and understanding helps generation by rewarding only those images whose spatial structure can be faithfully recovered. Extensive experiments show that CyCLeGen achieves significant gains across diverse image understanding and generation benchmarks, with emergent gains on image captioning.
DaCe-DT: Data-Centric Offline Multi-Task Reinforcement Learning via Adaptive Prompts and Trajectory Correction for Heterogeneous Tasks
Shudong Wang ⋅ Xinfei Wang ⋅ Chenhao Zhang ⋅ Shanchen Pang ⋅ Wenhao Ji ⋅ Haiyuan Gui ⋅ Meng Han ⋅ Xiaojian Liao
Offline multi-task reinforcement learning (Offline MTRL) heavily depends on the quality and distribution of pre-collected data. However, existing methods mainly focus on algorithmic optimization, with less emphasis on data-level improvements to enhance learning ability and generalization performance. This paper, from a data perspective, reveals three key bottlenecks that limit Offline MTRL performance:(i) ineffective utilization of prompts length under diverse task complexities, and (ii) semantic irrelevance of randomly sampled prompt segments, (iii) misleading supervision induced by fragmented and discontinuous trajectories. To address these challenges, we propose DaCe-DT, a robust offline MTRL framework designed to be insensitive to heterogeneous task complexities and data quality, featuring length-gated prompt masking (LGPM), retrieval-augmented prompt construction (RAPC), and value-adaptive return calibration (VARC). Together, these mechanisms enable DaCe-DT to deliver data-centric prompt adaptation and trajectory refinement, resulting in robust multi-task generalization and stable policy learning amid heterogeneous offline data and tasks. Experimental results on Meta-World show that DaCe-DT consistently outperforms state-of-the-art methods, achieving an average improvement of 11.73\% on optimal datasets and an improvement of 13.34\% on suboptimal datasets, demonstrating its effectiveness in learning stably from imperfect data and improving overall multi-task performance.
DARE: Dual-Level Adversarial Learning with Domain-Aware Regularization for Whole Slide Image Classification
Yuxuan Jiang ⋅ Haichuan Dong ⋅ Manning Wang
Pathological whole slide images (WSIs) from different hospitals often exhibit severe domain shifts due to variations in scanning devices, staining protocols, and tissue preparation procedures. This leads to significant performance degradation when a trained model is applied to data from different sources. Although unsupervised domain adaptation (UDA) has achieved strong performance in transferring a model trained on a source domain to an unlabeled target domain in natural images, it is much less studied in the computational pathology field. This is due to the unique properties of WSIs, including ultra-high resolution and strong inter-slide and intra-slide heterogeneity. As a result, existing UDA methods are often unstable and may even cause negative transfer in WSI classification. To address this issue, we propose Dual-level Adversarial learning with domain-aware REgularization (DARE) for WSI Classification. Specifically, we design an adaptive pseudo-labeling framework with a teacher-student architecture to provide stable supervision for the target domain and perform adversarial alignment on both patch-level and slide-level embeddings. This enables joint optimization of patch-level and slide-level semantic embeddings for dual-level alignment across domains. We also introduce domain-aware attention masking to regularize attention and improve cross-domain generalization. We evaluate our method on four public WSI datasets through extensive experiments. The results show that our method consistently outperforms state-of-the-art approaches across various domain shifts.
DART: Zero-Shot Dual-Side Alignment Routing for LLM Performance-Cost Tradeoffs
Yuejun Jiao ⋅ Yanxin Yang ⋅ Boyu Wang ⋅ Yonghao Yang ⋅ Hao Shen ⋅ Mingsong Chen
As the ecosystem of Large Language Models (LLMs) rapidly expands, LLM routing has become essential for dynamically balancing performance and cost. However, existing methods fail to simultaneously satisfy the three core criteria of an ideal router, i.e., high routing precision, minimal intrinsic operational overhead, and zero-shot onboarding of new models. These methods rely heavily on surface-level semantics and discrete ID mappings, conflating mere semantic similarity with underlying capability requirements and thereby precluding zero-shot, training-free generalization. To address these limitations, we propose DART, a novel zero-shot routing framework grounded in dual-side alignment to optimize LLM performance-cost tradeoffs. Instead of relying on direct semantic matching, we decouple the routing process into a demand-supply alignment problem within a shared low-rank latent capability space. Specifically, on the supply side, DART constructs latent capability embeddings for candidate models, using an initialization mechanism to anchor newly introduced models in this latent space via metadata and family-graph prior, without retraining. On the demand side, a dedicated encoder explicitly distills the query's directional capability requirements and inherent difficulty into a latent demand embedding. This decoupled representation allows routing decisions to be computed via a lightweight dot-product operation between the demand and capability embeddings, modulated by query-adaptive cost sensitivity, minimizing overhead. Extensive experiments across comprehensive benchmarks demonstrate that DART achieves state-of-the-art utility under diverse cost constraints, satisfying the three core criteria with exceptional robustness in cold-start scenarios.
DASS: A Solver-Agnostic Dynamic Auxiliary Search Strategy for Symbolic Regression
Hu ⋅ Qian Li ⋅ Ding Wang ⋅ Yekun Zheng ⋅ Jinyue Yan ⋅ Yuntian Chen
Symbolic regression (SR) aims to discover interpretable mathematical expressions from data and plays a key role in scientific discovery. However, existing methods face a common bottleneck: the enormous search space and severe combinatorial explosion make it difficult to achieve a favorable trade-off between accuracy and algorithmic efficiency. We argue that, in addition to designing increasingly complex solver-specific SR strategies, an equally promising direction is to equip them with a shareable auxiliary strategy that can provide reliably search guidance. To this end, We propose DASS, a solver-agnostic Dynamic Auxiliary Search Strategy that improves SR by constructing and refining a high-potential auxiliary search subspace. DASS represents the auxiliary search space as a functional subspace spanned by a set of high-potential basis function terms. It estimates term utility through multi-environment evaluation, filters unstable guidance, and updates terms' potential score with a Gibbs posterior, thereby calibrating the subspace to provide more reliable guidance for downstream solvers. The High-quality solutions produced by the solver are used to update the set of basis terms. This forms a virtuous closed-loop optimization process: the subspace guides the downstream solver, and the solver feedback reshapes the subspace for subsequent exploration. DASS can be seamlessly integrated with various types of SR solvers. Extensive experiments on LLM-SRBench show that DASS significantly improves the accuracy of original solvers with a manageable increase in runtime.
Data-Adaptive Mahalanobis Metric Learning for Cross-Head Attention in Transformers
Zhenke Duan ⋅ Xin Li ⋅ Jiqun Pan ⋅ Xiaofei Dong ⋅ Hanwen Ning ⋅ Xinyuan Song
In this paper, we propose Data-Adaptive Mahalanobis Cross-Head Attention (DAMCHA), a simple and effective attention mechanism that generalizes Multi-Head Attention (MHA) through the lens of metric learning. Standard MHA implicitly learns a static, head-wise Euclidean similarity corresponding to the diagonal blocks of a unified cross-head metric matrix, discarding all off-diagonal interactions across heads. DAMCHA addresses this limitation via two key innovations. First, it learns the full cross-head metric matrix, inducing a Mahalanobis-like similarity that explicitly captures inter-head interactions beyond the restricted block-diagonal structure of MHA. Second, it parameterizes the metric matrix as an input-dependent function, enabling the attention geometry to dynamically adapt to the intrinsic manifold structure of the data rather than remaining fixed after training. We further introduce principled regularization strategies, most notably stack-wise parameter sharing, to ensure computational efficiency and stable optimization. DAMCHA serves as a drop-in replacement for MHA and its variants in various Transformer architectures, delivering superior expressive power and learning performance with comparable or reduced computational overhead. Extensive experiments demonstrate that DAMCHA-based Transformers consistently outperform competing baselines across a range of benchmark tasks. Code is available at https://anonymous.4open.science/r/DAMCHA-3B37.
DataComp-VLM: Improved open datasets for Vision-Language Models
Matteo Farina ⋅ Vishaal Udandarao ⋅ Thao Nguyen ⋅ Selim Kuzucu ⋅ Maximilian Böther ⋅ Andreas Hochlehnert ⋅ Adhiraj Ghosh ⋅ Marianna Nezhurina ⋅ Karsten Roth ⋅ Joschka Strüber ⋅ Yuhui Zhang ⋅ Sebastian Dziadzio ⋅ Elaine Sui ⋅ Soumya S Jahagirdar ⋅ Dhruba Ghosh ⋅ Hasan Hammoud ⋅ Thomas De Min ⋅ Simone Caldarella ⋅ Muhammad Jehanzeb Mirza ⋅ Sedrick Keh ⋅ Mehdi Cherti ⋅ Hilde Kuehne ⋅ Bernt Schiele ⋅ Serena Yeung-Levy ⋅ Muhammad Ferjad Naeem ⋅ Federico Tombari ⋅ Ana Klimovic ⋅ Elisa Ricci ⋅ Matthias Bethge ⋅ Sewoong Oh ⋅ Ameya Prabhu ⋅ Alessio Tonioni ⋅ Jenia Jitsev ⋅ Massimiliano Mancini ⋅ Ludwig Schmidt ⋅ Nikhil Parthasarathy
Building performant Vision-Language Models (VLMs) requires carefully curating large-scale training datasets, yet the community lacks systematic benchmarks for evaluating such curation strategies. We introduce DataComp for VLMs (DCVLM), a benchmark for controlled data-centric experiments to improve VLM training. As part of DCVLM, we standardize 160 datasets spanning four data types—image-caption pairs, multimodal interleaved documents, text-only, and instruction-tuning data—into a corpus of 6T multimodal tokens. DCVLM allows participants to test curation strategies (filtering, mixing, formatting, sampling) across 1B–8B models and 6.25B–200B token budgets. Models are then evaluated on a carefully selected suite of up to 52 downstream benchmarks across 9 domains. We conduct extensive experiments on DCVLM and find that data mixing, not filtering, is key to a high-quality training dataset: instruction-heavy mixtures scale better than caption-heavy ones, with gains widening at larger scales. The resulting dataset, DCVLM-BASELINE, enables training an 8B VLM to 63.6% accuracy on our 33-task core suite with 200B training tokens. Compared to FINEVISION, the state-of-the-art open VLM training dataset, this represents an improvement of +5.4pp. DCVLM and all accompanying artifacts will be made publicly available.
Data-Efficient Learning for Constraint Satisfaction Problems via Relational Biases and Hard Axiom Clamping
Enqiang Zhu ⋅ Ke Wang ⋅ Yu Zhang ⋅ Shengzhi Wang ⋅ Xianhang Luo ⋅ Chanjuan Liu
Constraint satisfaction problems (CSPs) serve as critical benchmarks for neural algorithmic reasoning, requiring accurate multi-step constraint propagation to determine variable values that satisfy all specified constraints. Existing neural solvers often rely on extensive training datasets, struggle with distribution shifts, or allow fixed clues to drift in continuous spaces. We introduce CSP-RSB, a data-efficient recurrent Transformer that directly incorporates constraint topology into the attention mechanism through a learnable relational bias. Additionally, we propose a hard axiom clamping strategy that fixes clues after each recurrence constraint projection and uses the QK-Norm to stabilize a two-layer block. In 9 \times 9 Sudoku, CSP-RSB achieves 99.94\% board accuracy on the RRN-test, 99.13\% on unfiltered Kaggle, and 99.78\% on filtered Kaggle, surpassing the previous best, Trial & Error 2025, which reported accuracies of 99.40\%, 98.90\%, and 99.50\%, respectively. Furthermore, it successfully solves Erd\H{o}s--R\'enyi graph coloring with 100\% accuracy in 9 epochs and maintains 99.5-100\% accuracy on 10 \times 10 binary puzzles. By incorporating explicit topological structure, CSP-RSB reduces the need for depth and data requirements for effective algorithmic reasoning.
Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning
Augustinas Jučas ⋅ Yangchen Pan
Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a large initial task. Existing cold-start methods typically either train the backbone throughout the stream and compensate for semantic drift, or freeze a backbone after the first task, producing features biased toward the initial classes. These choices also create a computational tension: drift-compensation methods require repeated backbone training and increasingly expensive updates as the task horizon grows, while frozen-backbone methods are cheap but weak under cold start. We study a third option: a feature extractor that is never fit to image data at all. We propose CIRCLE, a class-incremental classifier built from fixed bidirectional two-dimensional reservoir features, adapted from BiRC2D for image classification, and streaming linear discriminant analysis heads. CIRCLE groups multiple random reservoir instantiations into feature ensembles and averages the softmax outputs of independent SLDA heads, yielding a tunable bias-variance tradeoff between richer random features and prediction-level ensembling. Because the feature extractor is fixed and the head admits streaming closed-form updates, CIRCLE performs sample-wise training without replay, task-boundary information, or backbone backpropagation. On CIFAR-100, TinyImageNet, ImageNet-Subset, and ImageNet-1k, CIRCLE is competitive at 10-20 task splits and substantially outperforms strong CS-EFCIL baselines at 50, 100, and 500 task splits, while training much faster than trained-backbone drift-compensation methods. Ablations show that the BiRC2D-style extractor, SLDA head, and balanced feature/prediction ensembling each contribute to the final performance.
Dataset Distillation via Drifting
George Cazenavette ⋅ Angelina Quan ⋅ Giannis Daras ⋅ Antonio Torralba ⋅ Vincent Sitzmann
The task of \textit{dataset distillation} aims to find a small set of synthetic images such that training a model on them reproduces the performance of the same model trained on a much larger dataset of real samples. Existing distillation methods focus on synthesizing datasets that enable training \textit{randomly initialized} models. In contrast, state-of-the-art vision approaches are increasingly building on large, pre-trained self-supervised models rather than training from scratch. In this paper, we investigate the problem of distilling datasets that enable us to optimally train \textit{linear probes} on top of such large, pre-trained vision models. We introduce a method of dataset distillation for this task called \textit{Linear Gradient Matching} that optimizes the synthetic images such that, when passed through a pre-trained feature extractor, they induce gradients in the linear classifier similar to those produced by the real data. Our method yields synthetic data that outperform all real-image baselines and, remarkably, generalize across pre-trained vision models, enabling us, for instance, to train a linear CLIP probe that performs competitively using a dataset distilled via a DINO backbone. Further, we show that our distilled datasets are exceptionally effective for fine-grained classification and provide a valuable tool for model interpretability, predicting, among other things, how similar two models' embedding spaces are under the platonic representation hypothesis or whether a model is sensitive to spurious correlations in adversarial datasets.
DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence
Xu Huang ⋅ Ye Huang ⋅ Zijun Liao ⋅ Yuwei Niu ⋅ Xiaojie Li ⋅ Menghan Zhou ⋅ De Wen Soh ⋅ Xiaotong Li ⋅ Daquan Zhou
High-compression tokenizers are essential for scaling latent image generative models. However, aggressive compression creates a fundamental tradeoff between reconstruction fidelity and generation efficiency: high compression image encoder always increase the learning diffucity of diffusion training, resulting slow model convergence. Recent representation autoencoders speed-up the diffusion training by improving the latent feature's expressive capability by replacing VAE encoders with pretrained semantic encoders, yet they are typically limited to moderate compression and lose pixel-level details necessary for faithful reconstruction. To achieve both high compression and fast diffusion training, we propose DC-SAE, a Decoupled Compact Semantic Autoencoder designed for high-compression image generation with accelerated diffusion model convergence. DC-SAE consists of two key components: (1) a macro-level architecture design that leverages semantic encoders to enable higher compression ratios, and (2) a pixel-level encoder that preserves low-level details, ensuring high-fidelity image reconstruction. We empirically demonstrate that DC-SAE performs strongly on image generation tasks, achieving both compact latent representations and efficient training dynamics. % Specifically, on ImageNet dataset with $512 \times 512$ resolution, DC-SAE achieves $32\times$ spatial compression, with \textbf{29.79} PSNR and \textbf{3.05} gFID, substantially outperforming the previous state-of-the-art high-compression tokenizer baselines DC-AE by 13.5\% and 59.7\% on PSNR and gFID respectively, maintaining comparable throughput and faster diffusion model training convergence.
DDMS: Discriminative Distillation of Multi-view Foundational Features into Single-view Models
Jeonggi Kwak ⋅ Sho Kagami ⋅ Yuki Ono ⋅ Kwang Moo Yi
Foundational visual features such as DINO have played a critical role across modern computer vision, and have recently become key components in multi-view feed-forward geometry estimators. In this work, we demonstrate that by re-distilling these multi-view models-their internal knowledge of 3D geometry-into a single-view estimator, we can obtain enhanced 3D consistent foundational features. Our key idea is to construct a multi-view teacher by fusing pretrained 2D foundation features with multi-view geometric features, and refining the fused representation with a discriminative ranking objective. Through our discriminative distillation framework, we enforce the learned features to be both 3D consistent and locally distinctive, while keeping them aligned with the feature space of the original foundation model to preserve the semantic structure of the pretrained representation. Consistency and local discriminability are critical for 3D computer vision problems such as forming semantic and geometric correspondences across images. To demonstrate the effectiveness of our method, we perform comprehensive experiments spanning multiple angles: direct feature analysis, dense prediction transfer, and explicit 3D lifting and rendering. Across these evaluations, our method consistently produces stronger 3D-aware foundation features that improve multi-view consistency and local discriminability while preserving the semantic transferability of the original representation.
Decision Path Tracing for Causal Analysis in Transformers
Won Jo ⋅ Dahee Kwon ⋅ Jongeun Baek ⋅ Cheongwoong Kang ⋅ Jaesik Choi
Investigating the causal structures of transformers often entails a trade-off between the granularity of the explanation and computational overhead. Existing approaches either bypass the full internal chain, potentially missing critical decision-relevant links, or struggle with the combinatorial explosion of path enumeration. To address these challenges, we propose an efficient automated path-tracing method that ensures both causal reliability and algorithmic feasibility. Through extensive verification, we find that (i) the identified paths are the primary sources for decision signals, as evidenced by a marked decrease in self-repair compared to non-path components; and (ii) transformers utilize modular, class-shared path mechanisms to process categorical information. Beyond these findings, we show that the potential implications of our method can stem from sparse manipulation, highlighting efficient applicability to downstream tasks such as model debugging and pruning.
Decoding the Critique Mechanism in Large Reasoning Models
Hoang Phan ⋅ Nguyen Hung-Quang ⋅ Thanh Quoc Hung Le ⋅ Xiusi Chen ⋅ Heng Ji ⋅ Khoa D Doan
Large Reasoning Models (LRMs) exhibit backtracking and self-verification mechanisms that enable them to revise intermediate steps and reach correct solutions, yielding strong performance on complex logical benchmarks. We hypothesize that such behaviors are beneficial only when the model has sufficiently strong “critique” ability to detect its own mistakes. This work systematically investigates how current LRMs recover from errors by inserting arithmetic mistakes in their intermediate reasoning steps. Notably, we discover a peculiar yet important phenomenon: despite the error propagating throughout the entire chain-of-thought (CoT) without any verbalized correction, the model still reaches the correct final answer after the thinking process finishes. This recovery implies the existence of an internal mechanism helping the model to detect errors and trigger self-correction, which we refer to as the hidden critique ability. Building on feature space analysis, we identify a highly interpretable critique vector representing this behavior. Extensive experiments across multiple model scales and families demonstrate that steering latent representations with this vector improves the model’s error detection capability and enhances the performance of test-time scaling at no extra training cost. Our findings provide a valuable understanding of LRMs’ critique behavior, suggesting a promising direction to control and improve their self-verification mechanism.
DecompDreamer: A Composition-Aware Curriculum for Structured 3D Asset Generation
Utkarsh Nath ⋅ Rajeev Goel ⋅ Rahul Khurana ⋅ Kyle Min ⋅ Mark Ollila ⋅ Pavan Turaga ⋅ Varun Jampani ⋅ Tejaswi Gowda
Current text-to-3D methods excel at generating single objects but falter on compositional prompts. We argue this failure is fundamental to their optimization schedules, as simultaneous or iterative heuristics predictably collapse under a combinatorial explosion of conflicting gradients, leading to entangled geometry or catastrophic divergence. In this paper, we reframe the core challenge of compositional generation as one of optimization scheduling. We introduce DecompDreamer, a framework built on a novel staged optimization strategy that functions as an implicit curriculum. Our method first establishes a coherent structural scaffold by prioritizing inter-object relationships before shifting to the high-fidelity refinement of individual components. This temporal decoupling of competing objectives provides a robust solution to gradient conflict. Qualitative and quantitative evaluations on diverse compositional prompts demonstrate that DecompDreamer outperforms state-of-the-art methods in fidelity, disentanglement, and spatial coherence.
Decomposed Representations Mitigate the Alignment–Specificity Trade-off in Multi-Omics
Mai Thao Dang ⋅ Feng Jiang ⋅ Hehuan Ma ⋅ Yuzhi Guo ⋅ Jingquan Yan ⋅ Haiqing Li ⋅ Saiyang Na ⋅ Zheng Zheng ⋅ Thuc A Tran ⋅ Jean Gao ⋅ Junzhou Huang
Single-cell multi-omics technologies jointly measure multiple modalities from the same cell, such as gene expression, chromatin accessibility, and surface protein abundance, providing a richer view of cellular state than any single modality alone. A central challenge is how to learn a multimodal cell representation that preserves the different types of information contained in these co-profiled measurements. Existing integration methods often emphasize either cross-modal correspondence, which benefits retrieval and imputation, or fused discriminative structure, which benefits classification, but these objectives can favor different aspects of the data. As a result, a representation optimized for one setting may fail to preserve information needed for another. We introduce scDecomp, a decomposed representation learning framework for co-profiled single-cell multi-omics data. Motivated by the concepts of redundancy, uniqueness, and synergy, scDecomp separates multimodal information into three role-specific components: a redundancy component ($R$) for information shared across modalities, a uniqueness component ($U$) for modality-specific information, and a synergy component ($S$) for interaction-dependent multi-omic signals. This design allows the learned representation to retain shared cell-state structure while also preserving modality-specific and cross-modal regulatory information. On single-cell multi-omics benchmarks, the $R$ component achieves a 5.0\% reduction in FOSCTTM compared with the strongest baseline, while $U+S$ improves cell type classification accuracy by 1.38\%. Branch-level analyses further show that the decomposed components support complementary aspects of representation quality, suggesting that structured decomposition provides a practical strategy for learning more informative multimodal cell representations.
Decoupled Complementary Fields on 3D Gaussian Maps for Embodied Navigation and Reasoning
Zhao Jiahao ⋅ Yaonan Wang ⋅ Zijie Wu ⋅ Mingtao Feng ⋅ He Xie ⋅ Renjie Ding ⋅ Hui Zhang
Embodied navigation and reasoning require an agent to actively acquire, organize, and verify task-relevant evidence in previously unseen environments. Recent methods typically rely on semantic maps, scene graphs, vision-language priors, or structured memory, but often underuse persistent fine-grained 3D evidence and entangle query-related ambiguity with environment-side reliability. This coupling can make evidence acquisition inefficient and final decisions vulnerable to weak or transient observations. We propose Decoupled Complementary Fields on 3D Gaussian Maps, an embodied navigation framework that maintains persistent volumetric evidence, separates query-related ambiguity from environment-side readiness, and grounds decisions through progressive verification. First, we construct a query-conditioned object-centric Gaussian navigation state built on online 3D Gaussian Splatting, which preserves fine-grained 3D evidence across viewpoints while providing an executable spatial interface for region reasoning and navigation. We further formulate Decoupled Complementary Fields, consisting of a Query-Ambiguity Field, which captures where task-relevant evidence remains unresolved, and an Environment-Readiness Field, which estimates where the environment is structurally reliable for motion and inspection. Finally, we couple these fields in a hierarchical navigation-and-verification policy, which selects informative frontiers, instantiates executable local goals, and verifies task-relevant candidates using persistent volumetric Gaussian evidence before stopping or answer handoff. Experiments on A-EQA and GOAT-Bench show consistent gains in active evidence acquisition, multi-modal lifelong navigation, and reliable final decision making.
Decoupling Conversational Dynamics in Full-Duplex Spoken Models through Reinforcement Learning
Yuxin Li ⋅ Donghang Wu ⋅ Guan-Ting Lin ⋅ Hung-yi Lee ⋅ Chengwei Qin ⋅ CHEN CHEN ⋅ Zhehuai Chen
Recent full-duplex spoken dialogue models have demonstrated compelling progress toward human-like interaction, enabling agents to respond with low latency, produce backchannels, and handle user barge-ins. Yet these improvements in conversational dynamics often come with weaker reasoning and instruction-following abilities, revealing a potential tension between interactive dynamics and intelligence capability. In this paper, we argue that such an intelligence--dynamics trade-off is not fundamental: conversational dynamics can instead be learned as a separate real-time decision policy from human dialogue data. To this end, we propose DuplexPO, a reinforcement learning (RL) framework that decouples when to speak from what to say. It preserves the semantic response capability of an instruction-tuned assistant, while optimizing its temporal interaction behavior over selected high-impact windows from long human conversations. To quantitatively optimize these dynamics, we formulate the Factorized Conversational Dynamics Reward (FCDR) to enable fine-grained temporal credit assignment for turn initiation, backchanneling, yielding, and regularized participation. The policy is then optimized with a GRPO-style objective. Experiments show that DuplexPO substantially improves full-duplex behaviors, including timely backchannels, smooth turn-taking, and barge-in handling, while maintaining strong reasoning and instruction-following performance. Moreover, improvements in dynamics-oriented metrics are reflected in better user experience, suggesting that optimizing conversational timing as a standalone objective can promote more natural full-duplex interaction.
Decoupling Time and Risk: Risk-Sensitive Reinforcement Learning with General Discounting
Mehrdad Moghimi ⋅ Anthony Coache ⋅ Hyejin Ku
Distributional reinforcement learning (RL) is a powerful framework increasingly adopted in safety-critical domains for its ability to optimize risk-sensitive objectives. However, the role of the discount factor is often overlooked, as it is typically treated as a fixed parameter of the Markov decision process or tunable hyperparameter, with little consideration of its effect on the learned policy. In the literature, it is well-known that the discounting function plays a major role in characterizing time preferences of an agent, which an exponential discount factor cannot fully capture. Building on this insight, we propose a novel framework that supports flexible discounting of future rewards and optimization of risk measures in distributional RL. We provide a technical analysis of the optimality of our algorithms, show that our multi-horizon extension fixes issues raised with existing methodologies, and validate the robustness of our methods through extensive experiments. Our results highlight that discounting is a cornerstone in decision-making problems for capturing more expressive temporal and risk preferences profiles, with potential implications for real-world safety-critical applications.
Double Q-learning is a classical control algorithm that mitigates the maximization bias of Q-learning. To do so, it explicitly trains two independent action-value functions and uses them to decouple action-selection and action-evaluation when computing bootstrap targets. Double DQN adapts target bootstrap decoupling to deep reinforcement learning (RL), but explicitly trains only a single action-value function and does not fully decouple its estimators. Consequently, the two estimators remain correlated, and overestimation persists. In this paper, we introduce Deep Double Q-learning (DDQL), a deep RL algorithm that explicitly trains two Q-functions through Double Q-learning. DDQL stabilizes training through a combination of techniques, including lower replay ratios, longer target network update intervals, and shared layers. Across 57 Atari 2600 games, DDQL improves aggregate performance over Double DQN, outperforming it on 47 games while further reducing overestimation. In addition, we study key design choices when adapting Double Q-learning to deep RL, including the network architecture, replay ratio, and minibatch sampling strategies.
Deep Generative Models for Phylogenetic Inference with Complex Evolutionary Processes
Ethan Baron ⋅ Alan Amin ⋅ Andrew Wilson
Phylogenetic inference plays a central role in understanding evolutionary relationships, with applications ranging from tracking pathogen spread to reconstructing the history of life. Conventionally, practitioners obtain posteriors over the phylogenetic tree of species using MCMC methods. However, such likelihood-based methods are only tractable for simple evolutionary models with restrictive assumptions. For more complex and realistic evolutionary models, conventional methods are prohibitively expensive, inaccurate, or even impossible. We instead advocate for a simulation-based inference approach, by using simulated data from an evolutionary model to train a neural network that predicts tree topologies conditioned on sequences. To accurately represent the complex posterior distributions over tree topologies that can arise, we present flexible models that iteratively generate trees using three natural paradigms: top-down, middle-out, and bottom-up. We use a discrete diffusion framework to train these models efficiently on large-scale simulated datasets of phylogenetic trees. For all three generative paradigms, our models fit the data substantially better than the previous state-of-the-art simulation-based method, Phyloformer 2, and obtain more accurate posteriors on real datasets. Finally, our models outperform misspecified conventional methods on data following complex evolutionary processes.
Neural representations are not unique objects. Even when two systems realize the same downstream computation, their hidden coordinates may differ by reparameterization. A probe family intended to reveal structure already present in a representation should therefore be stable under the relevant representation symmetries rather than be tied to a particular basis. We study this group action in the tractable exact setting of the final readout layer, where equivalent realizations induce affine changes of hidden coordinates. The resulting symmetry principle singles out a unique hierarchy of shallow coordinate-stable probes, with linear probes as its degree-1 member. We also show that a natural object for cross-model probe transfer is a shared probe-visible quotient--the representation modulo directions invisible to the probe family--rather than the full hidden state. Experiments on synthetic and real-world tasks support both predictions, showing where degree-2 probes help beyond linear ones and how quotient-based transfer enables coverage-aware monitor portability across model families. These results point toward a broader geometric representation theory of neural probing, with coverage-aware monitor transfer as a concrete operational consequence.
DeFlow: Decoupling Behavior-Prior Modeling and Value Maximization for Offline Policy Extraction
Zhancun Mu ⋅ Chi Zhang
We present DeFlow, a decoupled offline RL framework for extracting high-value actions from a learned multi-step flow prior. Directly optimizing iterative generative policies typically requires backpropagation through ODE solvers, while shortcut policies trade off expressivity, inference cost, and policy-improvement objectives. DeFlow keeps the flow model as a behavior prior and trains a lightweight action-conditioned residual module for value improvement under an adaptive trust-region penalty. This design bypasses solver differentiation and encourages proximity to the learned prior without treating the trust region as exact data-support preservation. Empirically, DeFlow is competitive with recent generative offline RL baselines, with the clearest gains on selected multimodal manipulation and offline-to-online settings.
DeformGen: Dynamics-Based Topology Augmentation for Deformable Manipulation Policy Learning
Zili Lin ⋅ Wenyao Zhang ⋅ Yuyang Zhang ⋅ Zekun Qi ⋅ Junyan Lin ⋅ Hanxin Zhu ⋅ Jiaolong Yang ⋅ Zhibo Chen ⋅ Yao Mu ⋅ Xiaokang Yang ⋅ Xin Jin ⋅ Wenjun Zeng
Imitation learning has achieved remarkable progress in robot manipulation, but its success heavily depends on large-scale demonstration data that are costly to collect. Recent work has therefore explored demonstration augmentation, but they are fundamentally limited in deformable manipulation, where task-relevant object variation is governed by high-dimensional deformations and physics-induced internal constraints rather than low-dimensional pose changes. We present DeformGen, a dynamics-based augmentation framework that achieves topological diversity for deformable objects. Instead of perturbing object pose, DeformGen expands the valid state distribution by applying localized physical disturbances, forward-simulating the dynamics, and stabilizing the result to obtain topology-coherent deformable states with physical plausibility. Given these synthesized states, DeformGen further transfers source manipulation trajectories via deformation-field warping, which lifts per-particle displacements into a continuous spatial function to adapt the end-effector trajectory consistently with the deformed geometry. In this way, our method augments both the state distribution and its associated manipulation behavior. Experiments on high-fidelity deformable manipulation benchmarks show that DeformGen consistently improves policy learning over training on the original demonstrations alone and over rigid-style augmentation baselines.
Déjà View: Looping Transformers for Multi-View 3D Reconstruction
Alessandro Burzio ⋅ Tobias Fischer ⋅ Sven Elflein ⋅ Qunjie Zhou ⋅ Riccardo de Lutio ⋅ Jiawei Ren ⋅ Jiahui Huang ⋅ Shengyu Huang ⋅ Marc Pollefeys ⋅ Laura Leal-Taixé ⋅ Zan Gojcic ⋅ Haithem Turki
Recent feed-forward 3D reconstruction transformers have scaled to over a billion parameters, following the broader trend of increasing model capacity in computer vision. Yet emerging evidence suggests that contiguous transformer layers often behave like repeated applications of similar operations, and multi-view reconstruction transformers refine their predictions progressively across decoder depth. We posit that model depth partially buys iteration, paid for inefficiently in unique parameters, and instead make that iteration explicit in architecture. Our model, DéjàView, applies a single looped transformer block recurrently to per-view features for K refinement steps. Trained once, it exposes K as an inference-time compute knob, matching or outperforming substantially larger feed-forward baselines across five reconstruction benchmarks spanning indoor, outdoor, object-centric, and driving scenes, while using a fraction of their parameters and comparable or lower compute. Importantly, the same looped block formulation outperforms an otherwise identical variant with independent per-step parameters under matched training data and compute, suggesting that explicit iteration is not merely a compute-efficient substitute for capacity but a stronger inductive bias for multi-view 3D reconstruction.
Distributed reinforcement learning trains on data from stale, buggy, or mismatched actors, producing actions with high surprisal (negative log-probability) under the learner's policy. The core difficulty is not surprising data per se, but \emph{negative learning from surprising data}. High-surprisal failures can dominate finite-batch updates through large perpendicular components, while high-surprisal successes reveal opportunities the current policy would otherwise miss. The \textit{Delightful Policy Gradient} (DG) separates these cases by gating each update with delight, the product of advantage and surprisal, suppressing rare failures and preserving rare successes without behavior probabilities. In a tabular analysis, DG suppresses the perpendicular second moment of high-surprisal failures by a policy-overlap factor that vanishes as the learner improves. The advantage sign is essential for surprisal-based filtering: any learner-probability-only gate that suppresses rare failures also suppresses rare successes. On MNIST with simulated staleness, DG without off-policy correction outperforms importance-weighted PG with exact behavior probabilities. On a transformer sequence task with staleness, actor bugs, reward corruption, and rare discovery, DG often achieves nearly order-of-magnitude lower error. When all four frictions act simultaneously, its sample-efficiency advantage is order-of-magnitude and grows with task complexity.
DELTA-TTS: Adapting Autoregressive Model into a Diffusion Language Model for Text-to-Speech
Junwon Moon ⋅ Yejin Lee ⋅ Seungbeom Kim ⋅ Hoseong Ahn ⋅ Sewoong Park ⋅ Heeseung Kim ⋅ Kyuhong Shim
Autoregressive (AR) text-to-speech (TTS) models generate speech tokens one at a time, and inference latency therefore scales linearly with output length. Discrete diffusion language models (dLLMs) have recently emerged as a parallel alternative that produces tokens via iterative unmasking. Recent work has converted pretrained AR language models into dLLMs in text generation, but this paradigm has not been extended to speech synthesis. In addition, existing conversion methods require full fine-tuning and large-scale training data. This leaves the question of whether such conversion can be done with substantially less compute and data largely open. We introduce DELTA-TTS, a lightweight conversion that turns a pretrained AR TTS backbone into a dLLM. The AR weights are kept frozen; all adaptation is routed through LoRA and a per-block speech-aware convolution. The convolution injects local acoustic context into the bidirectional attention, supplying the short-range continuity between adjacent speech tokens. With only $585$ hours of LibriTTS as adaptation data, DELTA-TTS achieves a state-of-the-art WER of $\textbf{1.75}\%$ on Seed-TTS test-en and decodes $\textbf{3.3}\times$ faster on the token-generation stage than the CosyVoice3 AR backbone it is converted from. This shows that lightweight AR-to-dLLM conversion provides a practical, data- and compute-efficient route to non-autoregressive TTS.
DepthMaster: Unified Monocular Depth Estimation for Perspective and Panoramic Images
Pengfei Wang ⋅ Shihao Wang ⋅ liyi chen ⋅ Zhiyuan Ma ⋅ Guowen Zhang ⋅ Lei Zhang
While monocular depth estimation has achieved significant progress, achieving generalized metric depth estimation for both narrow field-of-view (FoV) perspectives and $360^\circ$ panoramas remains an unsolved challenge. Existing methods are often tailored to specific camera types and struggle to produce accurate metric depth that generalizes across diverse settings. This limitation stems from two key challenges: the inherent geometric discrepancy between perspective and panoramic cameras, and the scarcity of panoramic training data with metric annotations. In this work, we introduce \textbf{DepthMaster}, a unified metric depth estimation framework. Rather than employing specialized networks to learn spherical distortions, we reformulate the problem by decomposing panoramic images into overlapping perspective patches. Crucially, distinct from prior projection-based methods that rely on ad-hoc architectural modifications to handle boundaries, we introduce a novel Correspondence Consistency Loss (CCL) and inject virtual projection cameras as geometric priors, allowing us to seamlessly stitch the patches while avoiding specialized operators and keeping the backbone largely compatible with standard Transformer designs. This strategy also resolves the geometric differences by unifying all inputs into a canonical perspective representation, and effectively circumvents data scarcity by directly unlocking powerful metric priors from vast perspective datasets. Trained on a mixed dataset that contains only one panorama dataset, DepthMaster achieves state-of-the-art zero-shot performance on 13 diverse datasets, outperforming not only universal methods but also leading specialist models in both perspective and panoramic domains. The code and models of DepthMaster will be released.
Depth through Recurrence: Towards Ultra-Efficient On-Device ASR
Chen Feng ⋅ Tianyi Xu ⋅ Yicheng Lin ⋅ Jay Zhuo ⋅ Ramchalam Kinattinkara Ramakrishnan ⋅ Zhaocong Yuan ⋅ Chenzheng Su ⋅ Xiaopeng Zhang
Modern automatic speech recognition (ASR) systems have achieved remarkable accuracy by scaling model depth and capacity, but at the cost of substantial memory and computation. On edge devices where ASR is often most needed, such as watches and glasses, deploying such large models is infeasible due to extremely constrained resources. This raises a fundamental question: can we achieve high representational power in deep ASR models without scaling up parameterization? In this work, we revisit the role of depth and identify layer-wise representational dynamics, in which most layers learn functionally similar transformations. Motivated by this insight, we propose a block-recurrent ASR architecture that replaces parameterized depth with a small set of recurrent blocks, each consisting of weight-shared layers, thereby preserving effective depth while drastically reducing model size. Through representation-guided grouping and knowledge distillation, block-recurrent models retain the accuracy of large reference models while using an order of magnitude fewer parameters. Across Open ASR Leaderboard benchmarks, a model with only two shared blocks recovers $97.4$% of the accuracy of large models on average. The compact backbone further enables efficient on-device personalization through lightweight MoE-LoRA adaptation, allowing user-specific models to match or surpass the accuracy of general-purpose ASR systems with significantly lower inference cost. Together, these results show that high-quality ASR can be achieved without large parameterization, and that structured recurrence provides a principled path towards ultra-compressed, high-efficiency, and personalized speech recognition on the edge.
D-GAP: Improving Out-of-Domain Robustness via Dataset-Agnostic and Gradient-Guided Augmentation in Amplitude and Pixel Spaces
Ruoqi Wang ⋅ Haitao Wang ⋅ Shaojie Guo ⋅ Qiong Luo
Out-of-domain (OOD) robustness is challenging to achieve in real-world computer vision, especially in unsupervised domain adaptation scenarios, where shifts in image background, style, and acquisition instruments often degrade model performance. Generic augmentations show inconsistent gains under such shifts, whereas dataset-specific augmentations require expert knowledge and prior analysis. Moreover, prior studies show that neural networks adapt poorly to domain shifts because they exhibit a learning bias to domain-specific frequency components. Perturbing frequency values can mitigate such bias but overlooks pixel-level details, leading to suboptimal performance. To address these limitations, we propose D-GAP, a Dataset-agnostic and Gradient-guided augmentation method for the Amplitude spectrum (in frequency space) and the Pixel values. Unlike conventional handcrafted augmentations, D-GAP computes sensitivity maps in the frequency space from task gradients, which reflect how strongly the deep models respond to different frequency components, and uses the maps to adaptively interpolate amplitudes between source and target samples. We further propose a dual-space augmentation that jointly controls spectral bias and spatial fidelity by introducing a complementary pixel-space blending branch. This way, D-GAP turns augmentation from fixed, random, or manually designed perturbation into a model-response-adaptive intervention. Extensive experimental results show that the proposed method consistently outperforms both generic and dataset-specific domain adaptation methods, improving average OOD performance by +5.3% on four real-world datasets and +1.9% on three benchmark datasets.
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents
Sirui Hong ⋅ Liuzhijie ⋅ Tengfei Li ⋅ Wei Tao ⋅ Yifan Wu ⋅ Chenglin Wu
Evaluating LLM-generated interactive software requires execution in addition to static analysis. The key difficulty is that correctness is a graph-level reachability property over latent UI state-transition graphs, whereas a GUI evaluator observes only a single execution trajectory. A failed rollout therefore rules out only one realized path, leaving failure attribution ambiguous between evaluator-side execution error and genuine software defect. We present DIAGEVAL, a trajectory-conditioned diagnostic evaluation protocol for post-failure GUI-agent evaluation of interactive software. Rather than blindly retrying from scratch,DIAGEVAL reuses the failed trajectory to choose targeted diagnostic probes and aggregates their outcomes into an internal attribution signal. The latent-graph view motivates the diagnostic problem; DIAGEVAL does not reconstruct the graph or estimate calibrated posterior probabilities. We evaluate DIAGEVALon WebDevJudge-Unit and RealDevBench across multiple GUI-agent evaluators and LLM backbones. On false-negative cases, DIAGEVAL recovers 45.6--62.1\% of failures that were initially attributed to software defects, outperforming retry-based baselines with 34.4--160.6\% relative gains. On the full evaluation sets, this recovery improves accuracy from 69.9\% to 78.3\% on WebDevJudge-Unit and from 65.0\% to 81.6\% on RealDevBench. These results suggest that reliable GUI-agent evaluation requires not only stronger execution, but also active failure diagnosis to disambiguate evaluator-side errors from genuine software defects.
Diagnosing and Correcting Bias in MLLM for Long Video Understanding
Xusheng Liang ⋅ Jianqiao Sun ⋅ Hao Zhang ⋅ Yulei Niu ⋅ Hengshuang Zhao ⋅ Jiawei Ma
Long-video question answering is challenging since the answer often hinges on a few decisive moments scattered throughout the video, while memory constraints force Multimodal Large Language Models (MLLMs) to sample frames at extremely low rates, resulting in extreme sparsity that risks omitting key evidence. Though aggressively enlarging the context size is intuitive, we argue that performance remains fundamentally constrained by the bias, $\textit{i.e.}$, the misleading cues caused by spurious linguistic correlations and salient yet irrelevant visual observations. In this paper, we first introduce S-MME, a benchmark designed to systematically diagnose shortcut-biased behavior in MLLMs. To mitigate this failure, we further propose a training-free framework, Counterfactual Long-video Evidence-Aware Reasoning (CLEAR). By treating the full video-question pair as a complete view, CLEAR estimates bias effects via other views by only keeping the question or locally salient visual clips. Then, we apply hidden-state intervention to mitigate the discrepancy at inference such that the correct answer can be predicted confidently. With experimental studies across multiple benchmarks, we show that our method is generalizable and can be integrated in different MLLMs, yielding consistent gains in prediction and robustness. We hope CLEAR can offer a complementary direction in test-time scaling for long-video understanding and will publish code.
Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
Rui Cai ⋅ Bangzheng Li ⋅ Xiaofei Wen ⋅ Muhao Chen ⋅ Zhe Zhao
Multimodal Large Language Models demonstrate strong performance on multimodal benchmarks, yet remain fragile when one modality carries spurious or misleading content. We trace this fragility to a deficiency in cross-modality competency, defined as the ability to fairly evaluate and integrate information across modalities, and identify its concrete, measurable manifestation as \emph{Modality Interference}, in which task-irrelevant modality signals improperly influence the model's predictions. To diagnose this phenomenon systematically, we design a perturbation-based evaluation grounded in causal intervention, in which controlled noise is injected into the task-irrelevant modality across image-heavy and text-heavy tasks. Across diverse MLLM families and scales, we observe consistent performance degradation on modality-heavy tasks under perturbations, indicating that modality interference is a pervasive and scale-resistant failure mode rather than an artifact of any specific architecture. To mitigate this, we propose a unified perturbation-aware fine-tuning framework that combines (i) heuristic and adversarial data augmentation targeting the task-irrelevant modality, and (ii) output-level consistency regularization between clean and perturbed inputs. Extensive experiments across diverse MLLM architectures, model scales, and benchmarks show that our approach simultaneously improves unimodal robustness and standard multimodal performance, achieving Pareto-optimal gains over existing baselines.
DiagnosticIQ: A Benchmark for LLM-Based Industrial Maintenance Action Recommendation from Symbolic Rules
Devin Y De Silva ⋅ Dhaval Patel ⋅ Christodoulos Constantinides ⋅ Shuxin Lin ⋅ Nianjun Zhou ⋅ Dharmashankar Subramanian ⋅ Paul J Adams ⋅ Sal Rosato ⋅ Nicolas Constantinides ⋅ Deborah L McGuinness ⋅ Jayant Kalagnanam
Monitoring complex industrial assets relies on engineer-authored symbolic rules that trigger on sensor conditions and prompt technicians to perform corrective actions. The bottleneck is not detection but response: translating rules into maintenance steps requires asset-specific knowledge gained through years of practice. We investigate whether LLMs can serve as decision support for this rule-to-action step and introduce DiagnosticIQ, a benchmark of 6690 expert-validated multiple-choice questions from 118 rule-action pairs across 16 asset types. We contribute (i) a symbolic-to-MCQA pipeline normalizing rules to Disjunctive Normal Form with embedding-based distractor sampling, (ii) five variants probing distinct failure modes (Pro, Pert, Verbose, Aug, Rationale), and (iii) a benchmark of 29 LLMs and 4 embedding baselines. A human evaluation (9 practitioners, mean 45.0\%) confirms DiagnosticIQ requires specialist knowledge beyond operational experience. Three findings stand out. The frontier has closed: the top three LLMs lie within one Macro point, with Bradley-Terry Elo placing claude-opus-4-6 30 points above the next model. Yet DiagnosticIQ-Pro exposes brittleness, with every model losing 13-60\% relative accuracy under distractor expansion. DiagnosticIQ-Aug exposes pattern-matching: under condition inversion, frontier models still select the original answer 49-63\% of the time. The deployment bottleneck is not capability but calibration: frontier models handle template-style fault detection but break under structural perturbation.
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
Connall Garrod ⋅ Jonathan Keating ⋅ Christos Thrampoulidis
Cross-entropy (CE) loss is central to deep learning, but existing theory often relies on simplifications—such as squared loss or convex models—that miss key aspects of CE optimization. In this work, we study multi-class CE dynamics using a two-layer linear network with orthogonal inputs, the simplest non-convex setting where the CE implicit bias remains unresolved. This coincides with the unconstrained features model used to study neural collapse (NC). Our analysis is based on a key observation: Hadamard initialization diagonalizes the softmax operator. This allows us to extend the spectral initialization framework that Saxe et al. (2013, 2019) developed for squared loss. We prove convergence to NC under spectral CE training and give the first finite-time analysis in this setting via an explicit Lyapunov function that decreases monotonically to NC despite spurious critical points. We further identify CE-specific phenomena absent under squared loss—coupling, non-monotonic convergence, qualitative dependence on the number of classes, and exponential slowdowns. We also characterize the role of network width and show empirically that spectral dynamics qualitatively model small random initialization.
Diff3R: Feed-forward 3D Gaussian Splatting with Uncertainty Aware Differentiable Optimization
Yueh-Cheng Liu ⋅ Jozef Hladký ⋅ Matthias Niessner ⋅ Angela Dai
Recent advances in 3D Gaussian Splatting (3DGS) present two main directions: feed-forward models offer fast inference in sparse-view settings, while per-scene optimization yields high-quality renderings but is computationally expensive. To combine the benefits of both, we introduce Diff3R, a novel framework that explicitly bridges feed-forward prediction and test-time optimization. By incorporating a differentiable 3DGS optimization layer directly into the training loop, our network learns to predict an optimal initialization for test-time optimization rather than a conventional zero-shot result. To overcome the computational cost of backpropagating through the optimization steps, we propose computing gradients via the Implicit Function Theorem and a scalable, matrix-free PCG solver tailored for 3DGS optimization. Additionally, we incorporate a data-driven uncertainty model into the optimization process by adaptively controlling how much the parameters are allowed to change during optimization. This approach effectively mitigates overfitting in under-constrained regions and increases robustness against input outliers. Since our proposed optimization layer is model-agnostic, we show that it can be seamlessly integrated into existing feed-forward 3DGS architectures for both pose-given and pose-free methods, providing improvements for test-time optimization.
Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting
Chen Su ⋅ Yuanhe Tian ⋅ Yan Song
Diffusion models have become a widely used framework for probabilistic time series forecasting, modeling the distribution of future values given an observed history. In time series forecasting, however, the future continues the observed history, creating an asymmetry the standard diffusion process leaves unaddressed, with slowly-varying content largely determined by the observed continuity while higher-frequency dynamics carry most of the residual uncertainty. Existing diffusion-based forecasters decouple this asymmetry through an external rule before generation, leaving the corruption trajectory blind to which parts of the target the history can already anchor. We propose \textsc{DiffDiff}, a diffusion framework that embeds this predictability asymmetry into the diffusion trajectory itself, so that a single end-to-end diffusion process becomes aware of which parts of the target the history can already anchor. \textsc{DiffDiff} makes the forward operator step-dependent so that the noisy intermediate state progressively shifts from the target itself toward its second-order differenced structure, while a conditioning pathway supplies the denoiser with both value-domain and differential history information balanced by a stage-adaptive gate at each diffusion step. The terminal distribution approaches a standard Gaussian, preserving compatibility with existing samplers. On seven benchmarks across four prediction horizons, \textsc{DiffDiff} outperforms six diffusion baselines, and our analysis confirms that \textsc{DiffDiff} concentrates the diffusion's generative effort on the most uncertain components of the target while relieving it from rebuilding the history-anchored content.
Differentiable Retrieval-Augmented Generation for Predicting Cellular Responses to Gene Perturbation
Andrea Giuseppe Di Francesco ⋅ Andrea Rubbi ⋅ Rishabh Jain ⋅ Pietro Lió
Predicting transcriptional responses to genetic perturbations is fundamental to functional genomics and therapeutic discovery. Recent deep learning models have shown promise in single-cell perturbation response prediction, but they typically generate each response in isolation, without explicitly leveraging experimentally characterised responses to related perturbations. We introduce PT-RAG (Perturbation-aware Two-stage Retrieval-Augmented Generation), a plug-in retrieval-and-conditioning module for generative cellular perturbation response. PT-RAG augments an existing perturbation-response backbone with learned access to related perturbation contexts. The key challenge is that relevance is not fixed in this setting: functionally related genes may elicit different effects across cell types. PT-RAG addresses this with a two-stage retrieval mechanism: GenePT-based semantic retrieval first identifies $K$ candidate perturbations, after which a differentiable Gumbel-Softmax selector adaptively selects retrieved contexts conditioned on the control cell state, the query perturbation, and each candidate perturbation. Through cross-cell-type and cross-perturbation generalization tasks, PT-RAG consistently improves distributional similarity and often overall predictive quality; for example, on scGPT cross-cell-type results, Wasserstein distance drops by 10.7%. The code to reproduce our experiments is available at https://anonymous.4open.science/r/PT-RAG_NIPS-13D8/.
In machine learning applications, privacy requirements during inference or deployment time could change constantly due to varying policies, regulations, or user experience. In this work, we aim to generate a magnitude of models to satisfy any target differential privacy (DP) requirement without additional training steps, given a set of existing models trained on the same dataset with different privacy/utility tradeoffs. We propose two post processing techniques, namely random selection and linear combination, to output final private models satisfying any target privacy parameter. We provide privacy accounting of these approaches from the lens of R\'enyi DP and privacy loss distributions for general problems. In a case study on private mean estimation, we precisely characterize the privacy/utility results and theoretically establish the superiority of linear combination over random selection. Empirically, we validate our approaches and analyses on several models and both synthetic and real-world datasets.
Diffusion-Enhanced GFlowNet for Solving Vehicle Routing Problems
Ni Zhang ⋅ Zhiqin Zhang ⋅ Ling Pan ⋅ Hoong Chuin Lau ⋅ Zhiguang Cao
Traditional neural solvers for solving vehicle routing problems (VRPs) often suffer from limited solution diversity, motivating the development of Generative Flow Network (GFlowNet)–based models. However, the effectiveness of these models is frequently constrained by insufficient flow expansion in high-reward regions, limiting their ability to distribute probability across promising solution routes, the deeper exploration could yield superior results. Diffusion models, in contrast, provide stronger structural guidance for exploration. These two paradigms are naturally complementary: GFlowNet can supply edge-level signals for diffusion to embed, while diffusion can guide broader exploration. Leveraging this synergy, we propose Diffusion-Enhanced GFlowNet (DEG), a novel framework that integrates GFlowNet with diffusion model to encourage richer flow expansion toward high-reward regions and derive higher-quality solutions. Specifically, DEG exploits GFlowNet’s inherent diversity to generate edge-specific backward signals, applies the stochastic noise schedule of diffusion to perturb these signals, and then denoises them within the GFlowNet paradigm. To further improve scalability, we introduce a specialized decoder capable of dynamically adapting to diverse problem scales. Extensive experimental evaluations on synthetic and real-world datasets, including instances with up to 10,000 nodes, demonstrate that DEG consistently achieves favorable performance compared to baseline methods.
Diffusion Meta-Prompting and Steering for Generalizable Foundation Model Adaptation
Deepak Sridhar ⋅ Yi Li ⋅ Kartikeya Bhardwaj ⋅ Shuangjun Liu ⋅ TAOTAO JING ⋅ Yuan Li ⋅ Shuai Zhang ⋅ Jiancheng Lyu ⋅ Dashan Gao ⋅ Nuno Vasconcelos
Prompt learning is a popular parameter-efficient method for adapting foundation models, but learned prompts are typically task-specific and fail to generalize to new classes, domains, or compositions of tasks. In this paper, we introduce a $\textbf{Diffusion Meta-Prompt (DMP)}$ model , a framework that models the distribution of learned prompts using diffusion models. DMP is trained only on a repository of previously learned prompts to synthesize new prompts conditioned on natural language task descriptions, without access to the task data. To improve the sampling stability, we introduce a test-time steering strategy for DMP, which uses the best training-selected prompt in the repository as a latent anchor during diffusion sampling, without retraining the DMP or accessing test classes. DMP improves generalization across classification, retrieval and text-to-image generation tasks, supports concept composition and negative prompting without explicit training. It reduces storage and inference costs by over 90% compared to prompt retrieval methods. For composite classification, DMP achieves upto $\textbf{2.0}$% average gain over prior meta-learning methods across 55 pairs of datasets with gains as high as $\textbf{8.5}$% on specific pairs such as Eurosat and Flowers. DMP also enhances cross-task generalization with $\sim$$\textbf{2-9}$% improvement for hierarchical classification task. We further provide a theoretical guarantee bounding the expected task loss of prompts sampled from a DMP.
Diffusion Models without Classifier-free Guidance
Zhicong Tang ⋅ Dong Chen ⋅ Jianmin Bao ⋅ Baining Guo
Classifier-free Guidance (CFG) is central to modern diffusion models, but operates as an inference-time correction where the training objective learns the unbiased conditional score, while guided sampling targets a posterior-tilted score that never explicitly learned. Rather than another inference-time patch or distillation, Model-guidance (MG) directly integrates the posterior tilt into the training objective and eliminates CFG during inference. MG serves as a plug-and-play module compatible with existing methods, yet accelerates convergence by online self-bootstrapping and doubles inference speed. Unlike distillation surrogates, a reparameterisation property reveals that MG implicitly learns the vanilla score without mode collapse. Extensive experiments demonstrate that MG matches and even surpasses CFG baselines, and achieves a state-of-the-art FID of $1.34$ on ImageNet $256$ benchmark.
Diffusion Thinking for Fast Long-Form Spatial Reasoning in Vision--Language Models
Zhitao Zeng ⋅ weitao Du ⋅ Yueming Jin
Long-form spatial reasoning in vision--language models (VLMs) is often bottlenecked by autoregressive chain-of-thought (CoT) decoding, where intermediate reasoning traces and final answers are generated token by token. This sequential decoding path makes high-quality spatial reasoning expensive in latency-sensitive embodied and interactive settings. We propose \emph{Diffusion Thinking}, a post-training framework that converts pretrained autoregressive VLMs into diffusion-style reasoners for fast long-form spatial inference without changing the visual encoder or language backbone. Diffusion Thinking partitions each CoT trace into blocks and refines tokens within the current block in parallel through discrete denoising, while preserving causal dependence across blocks and reusing KV cache for streaming generation. This creates a high-speed operating point, but aggressive parallel refinement can weaken fine-grained reasoning. To address this, we introduce \emph{Diffusion Rethinking}, a cached test-time scaling mechanism that reallocates part of the saved latency budget to self-correction rounds, yielding a controllable speed--accuracy frontier. We construct a 561K-instance Long-CoT Spatial Reasoning dataset with verified final answers and generated CoT traces across six spatial task types. Across Qwen2.5-VL and InternVL3 backbones, Diffusion Thinking with block size $D=32$ achieves 38.18--41.81$\times$ effective wall-clock speedup over autoregressive Long-CoT decoding. With eight rethinking rounds, it retains 4.20--4.60$\times$ speedup and matches or surpasses autoregressive Long-CoT accuracy on the largest backbones. These results show that diffusion-style compute reallocation is a practical path toward fast, accurate long-form spatial reasoning in VLMs.
DIGS: Distribution-Informed Gaussian Splatting for Training-Free Open-Vocabulary 3D Segmentation
ZHIFAN ZENG ⋅ Wenbo Xiao ⋅ Arcot Sowmya ⋅ Changming Sun
Open-vocabulary 3D Gaussian segmentation typically binds semantic features to per-Gaussian state via per-scene gradient training (1–4 h) or the more recent closed-form distillation. Both commit each primitive to a single feature; therefore, boundary Gaussians spanning multiple 2D instances are forced into a hard decision that degrades retrieval accuracy and yields unstable mask boundaries at render time. We present DIGS, a training-free framework with two stages, class-agnostic map- ping and multi-modal encoding, that decouples instance discovery from linguistic naming. At the Gaussian level, our distribution-informed representation maintains an explicit top-M label distribution per primitive, updated by a tempered Dirichlet rule; preservation of multi-peaked posteriors at boundary primitives allows a late-argmax renderer to recover clean instance boundaries. A bidirectional state machine aggregates these distributions into a persistent class-agnostic instance ontology. The instance-level encoding stage then fuses the CLIP visual embedding of each instance’s isolated 3D footprint with the CLIP-text embedding of an LLM-generated attribute description, disambiguating visually similar instances. DIGS maps a scene in 1–3 minutes on a single A100 and serves 3D-native retrieval in 5.4 ms, reaching 71.95% mIoU on LERF-OVS, surpassing the strongest training-free baseline SFS and the strongest training-based baseline LaGa, and setting new state-of-the-art performance on LERF-Mask (91.71%) and 3D-OVS(96.45%). Codes are at: https://anonymous.4open.science/r/DIGS-C4EF.
DiM$^3$: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging
Zijing Wang ⋅ Mingyang Wang ⋅ Ercong Nie ⋅ Yongkang Liu ⋅ Shi Feng ⋅ Mengjie Zhao ⋅ Daling Wang ⋅ Xiaocui Yang ⋅ Hinrich Schuetze
Towards more general and human-like intelligence, large language models should seamlessly integrate both multilingual and multimodal capabilities; however, extending an existing multimodal model to many languages typically requires expensive multilingual multimodal data construction and repeated end-to-end retraining. We study a training-free alternative: injecting multilingual capability into an existing multimodal model by composing residual updates in the shared language model backbone. The key challenge is that multilingual and multimodal updates are heterogeneous, reflecting different functional roles in the shared model. To address this, we propose Direction- and Magnitude-aware Multilingual Multimodal merging (DiM$^3$), which selectively composes the two updates at each parameter dimension while preserving the original vision encoder and multimodal projector. Experiments on multilingual benchmarks in both text-only and vision-language settings, covering 57 languages across LLaVA- and Qwen-based backbones, show that DiM$^3$ consistently outperforms existing merging baselines, substantially improves multilingual performance over the original multimodal model, and remains competitive with dedicated multilingual multimodal fine-tuning while largely retaining general multimodal ability. We further show that DiM$^3$ can be directly applied to already trained multilingual multimodal models and still yield additional gains. Further interpretability analysis shows that DiM$^3$ primarily reshapes intermediate-layer semantic representations, strengthening cross-lingual alignment under both text-only and multimodal inputs while preserving higher-layer task-sensitive structure. Our repository is on https://anonymous.4open.science/r/w-C677/.
DiPhon: Diffusion on Graphons for Scalable Graph Generation
Sergio Rozada ⋅ Yiming QIN ⋅ Manuel Madeira ⋅ Pascal Frossard ⋅ Alejandro Ribeiro
Diffusion models are a leading paradigm for graph generation, with notable impact in domains such as molecular design. Yet, scaling these models to large graphs remains an open problem. We approach this question in the dense-graph setting through the lens of graphons, the size-agnostic limit object of dense graph sequences, to study how structural graph statistics behave across node-size scales. This perspective leads to DiPhon, a diffusion process for size-scalable graph generation. Specifically, we formulate a continuous diffusion process on the graphon space via a Jacobi stochastic differential equation (SDE), and propose DiPhon, a discretized scheme that imitates it on graphs. We further derive the corresponding reverse-time process, which requires access to the marginal score. For the Jacobi process, this score interestingly admits a tractable form, which we estimate from data via graph denoising and plug into the reverse process to generate graph samples. We prove that DiPhon matches the first moment of the marginal distributions induced by the continuous graphon process exactly, and approximates the second moment up to a closed-form discrepancy. In this way, DiPhon inherits the size-agnostic statistical properties of the graphon dynamics can be used to solve the scaling bottleneck. Empirically, we demonstrate the scalability of DiPhon by training on small graphs and generating substantially larger ones at inference time, without retraining on large graphs.
DIS-Bench: Evaluating LLMs on System Testing via Directed Input Synthesis
Siwei Wei ⋅ Yuqi Guo ⋅ Yan Cai
LLM-based coding agents have shown strong capabilities on real-world software testing tasks, yet existing evaluations concentrate on unit test generation, leaving system testing unexplored. System testing is both practically important for software quality assurance and inherently difficult, as it demands repository-level holistic understanding and long-horizon reasoning over program behavior. In this paper, we propose to measure LLM agents' system testing capabilities through its core problem, Directed Input Synthesis (DIS): generating system-level inputs that exercise specified code branches. We introduce DIS-Bench, a benchmark built via an automated pipeline based on random testing that selects target branches that are both reachable and challenging. Because DIS-Bench is constructed from execution behavior rather than human-authored issues or patches, it is scalable and less susceptible to contamination from existing training corpora. DIS-Bench contains 8267 tasks drawn from 37 real-world repositories. We evaluate 7 mainstream open-sourced LLMs on DIS-Bench-Lite under both BM25 retrieval and state-of-the-art coding agents, including OpenHands, Mini-SWE-Agent, and Claude-Code. The results show that DIS is highly challenging for current models: the best-performing configuration resolves only 28.1\% of targets at pass@1. We further perform a manual inspection to characterize the dominant failure modes, revealing reasoning bottlenecks of LLMs on repository-level system testing tasks.
Discovering dynamical parameters of synthetic multicellular systems from image sequences
Corey Herr ⋅ Cedric Allier ⋅ Stephan Saalfeld ⋅ Allyson E. Sgro
Cells in multicellular systems have complex shapes, yet most techniques for predicting biological dynamics from image sequences discard shape information by reducing cells to point particles. Here, we present a framework that moves beyond point representations by learning dynamic, shape-aware representations of synthetic cells in a multicellular collective directly from 2D time-series images. Specifically, our GraphDINO framework combines a frozen DINOv3 backbone for shape representation with a graph neural network trained through next-frame prediction. GraphDINO learns a per-cell latent representation that captures underlying cellular properties such as growth rate, membrane tension, and cell-cell adhesion without parameter supervision. The quality of this recovery depends on the encoder, with DINOv3 features providing a better shape representation than classical shape encoders. Together, these results demonstrate that combining foundation model features with graph-based interaction rules allows the discovery of key dynamical parameters that govern cellular behavior, directly from image sequences.
Discrete Langevin-Inspired Posterior Sampling
Sattwik Basu ⋅ Chaitanya Amballa ⋅ Jorge V Sampedro ⋅ Romit Roy Choudhury
We study posterior sampling for inverse problems in discrete state spaces using discrete diffusion models as generative priors. While continuous diffusion models have become widely used for inverse problems, their discrete counterparts remain comparatively underexplored. Existing discrete posterior samplers often rely on continuous relaxations of discrete variables, Gibbs-style updates, or mechanisms specialized to particular corruption processes, which can limit scalability or generality. We propose $\Delta$LPS, a Discrete Langevin-Inspired Posterior Sampler that uses gradient information to identify promising discrete moves without leaving the discrete state space. The resulting approach enables efficient parallel updates across all token dimensions and is agnostic to the training paradigm of the discrete diffusion prior, including masked and uniform-state diffusion. We evaluate our method on image restoration tasks across MNIST, CIFAR, and FFHQ, as well as spatial mapping, covering linear, nonlinear, and blind inverse problems. Across these settings, we improve over recent discrete diffusion posterior samplers and are competitive with strong continuous diffusion-based inverse solvers. Our results suggest that fully discrete, gradient-informed posterior samplers offer a scalable and general path toward solving inverse problems over discrete representations.
Discriminative Score Function: Turning Pretrained Models into Functional Generative Priors
Junhoo Lee ⋅ Hyeonjin Kim ⋅ Sangbum Han ⋅ Nojun Kwak
Discriminative models are trained with task-specific objectives such as classification, detection, or representation learning, yet their practical utility depends on generalizing beyond the training task to unseen samples from the same domain. This suggests that a discriminative model may retain an implicit trace of the data manifold beyond its task labels. We ask whether this trace can be exposed as a generative signal without task-specific guidance. We introduce the Discriminative Score Function (DSF), which extracts an image-space generative field from the loss-gradient geometry of a pretrained discriminative model. DSF provides a training-free update rule that moves images toward model-consistent visual regions without retraining or task-specific targets. We apply DSF to diverse pretrained discriminative models, including ResNet classifiers, DETR detectors, and DINOv2 ViT encoders. DSF generates unconditional images from noise and supports conditional guidance, editing, inpainting, and explanation through the same field. DSF-generated samples also provide effective zero-shot quantization calibration when real images or labels are unavailable, preserving feature geometry better than data-free calibration baselines. These results show that discriminative training leaves a generative functional signal for synthesis and calibration.
Disentangled Representation Learning via Flow Matching
Jinjin Chi ⋅ Taoping Liu ⋅ Mengtao Yin ⋅ Ximing Li ⋅ Yongcheng Jing ⋅ Jialie Shen ⋅ Leszek Rutkowski ⋅ Dacheng Tao
Disentangled representation learning aims to capture the underlying explanatory factors of observed data, enabling a principled understanding of the data-generating process. Recent advances in generative modeling have introduced new paradigms for learning such representations. However, existing diffusion-based methods encourage factor independence via inductive biases, yet frequently lack strong semantic alignment. In this work, we propose a flow-matching–based framework for disentangled representation learning, which casts disentanglement as learning factor-conditioned flows in a compact latent space. To enforce explicit semantic alignment, we introduce a non-overlap (orthogonality) regularizer that suppresses cross-factor interference and reduces information leakage between factors. Extensive experiments across multiple datasets demonstrate consistent improvements over representative baselines, yielding higher disentanglement scores as well as improved controllability and sample fidelity.
We characterise disentanglement for smooth generative pushforward models, such as in VAEs and GANs. For a generator/decoder $g:\\mathcal{Z}\\to\\mathcal{X}$ and factorised prior $p(z)=\\prod_i p_i(z_i)$, we define disentanglement as standard statistical independence expressed over the generated manifold: {the pushforward density} $p_\\mu = g_\\#p$ factorises into one–dimensional "seam" factors, each controlled by a distinct latent coordinate. We prove that $p_\\mu$ factorises according to the SVD of $g$'s Jacobian; that disentanglement equates to two conditions on $g$ (C1-C2); and that under such conditions the seam factors are identifiable, up to permutation and sign. In the special case of Gaussian ($\\beta$-)VAEs, we show via an identity how diagonal posteriors promote C1-C2, in expectation, explaining why disentanglement arises modulated by $\\beta$. Experiments illustrate this mechanism on Gaussian data, dSprites, and CelebA.
Disentangling Channel Semantics in Vision Transformers via Token Decorrelation and Composition-Aware Modulation
Daeun Kim ⋅ Hyejin Park ⋅ Hyesong Choi ⋅ Dongbo Min
Multi-channel images encode heterogeneous channel semantics aligned over a shared spatial structure, which requires models to jointly capture globally shared structure and channel-specific variations. Recent multi-channel Vision Transformers attempt to address this challenge by augmenting the `[CLS]` token with memory tokens to organize multi-channel representations under varying channel compositions. However, we find that these context tokens collapse onto a few dominant channels under high inter-channel redundancy, producing redundant representations and a biased global summary that fails to reliably support composition-aware channel interpretation. We propose **MuCa-ViT (Multi-Channel Composition-Aware Vision Transformer)**, a unified framework that addresses these limitations through two coupled mechanisms. **Factorial Token Learning (FTL)** enforces token-level decorrelation among context tokens, encouraging the `[CLS]` token to capture globally shared semantics while memory tokens preserve complementary channel-specific cues. **Channel Composition-Aware Modulation (CAM)** uses the FTL-refined `[CLS]` as a semantic anchor to generate sample-wise modulation signals, which enable composition-aware interpretation of channels within a unified backbone rather than relying on static channel embeddings. Experiments on JUMP-CP, CHAMMI, and So2Sat show that MuCa-ViT outperforms prior multi-channel ViTs across all three benchmarks, and the magnitude of improvement aligns with the inter-channel redundancy of each dataset ($\rho \in [0.19, 0.77]$), validating the diagnostic analysis that motivates our design.
Disentangling Continuous-Time Latent Dynamics: Identifiability of Latent SDEs via Diffusion Shifts
Yuanyuan Wang ⋅ Wenjie Wang ⋅ Haoxuan Li ⋅ Mingming Gong ⋅ Kun Zhang
Causal representation learning for time series has developed strong identifiability results in discrete-time latent causal models, but identifiability in continuous-time latent stochastic differential equation (SDE) models remains largely open. We address this gap using environment-induced shifts in diffusion covariance. We study additive-noise latent SDEs observed through an unknown nonlinear diffeomorphism, with shared drift but environment-specific diffusion covariance. We show that two diagonal diffusion regimes with pairwise distinct coordinate-wise variance ratios identify the latent coordinates up to permutation and scaling, without any sparsity assumption on the drift. We first prove this result for linear Ornstein--Uhlenbeck systems and then extend it to general additive-noise latent SDEs. Under mild smoothness, the instantaneous drift-Jacobian causal graph is identifiable up to the same permutation. We propose a two-stage estimator for latent disentanglement and optional graph recovery; experiments on synthetic systems confirm the predicted identifiability boundary, and an application to Hardanger Bridge monitoring data illustrates the approach on real sensor trajectories.
Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
Haipeng Liu ⋅ Yang Wang ⋅ Meng Wang
Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. In particular, the low-frequency band primarily captures color information, mid-frequency band encodes global structure, while high-frequency band corresponds to local texture. Besides, the early (high-noise) denoising stage mainly focuses on the low-and-mid frequency bands, while the late (low-noise) denoising stage progressively exhibits the high-frequency band, resulting in the dominance of low-and-mid over high frequency bands, especially when entangled within a single image reference, tends to reinforce each other at the early stage, hence the pitfall of overwhelming suppression of text prompt guidance. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these Dual image references within Frequency-aware Diffusion Models, dubbed Dual-FDM, to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized image. We further compute the ratio of information entropy of the substituted frequency bands to adaptively modulate the denoising timesteps between the early and late stages. Extensive experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from the supplementary material package.
Dispatchable Coordination Envelopes for Embodied Collaboration Under Sequential Uncertainty
Yuchen Wang ⋅ Fanqi Kong ⋅ Ke Shi ⋅ Zhipeng Liu ⋅ Xiaofei Yue ⋅ Zhaoxuan Li ⋅ Tingting Li ⋅ Ziming Zhao
Embodied AI systems can pass planning and still fail in execution: a delay, resource conflict, synchronization change, or candidate invalidation can turn a valid point schedule into broken commitments. We introduce FLECAR, which gives the executor a dispatchable coordination envelope instead of fixed start times. Offline, Flexible Envelope Synthesis builds executable windows, coupling constraints, and a dispatch graph; online, Controllability-Aware Repair updates that envelope after events and repairs only the affected region when dispatchability fails. On generated coordination instances with the same event streams and budgets, FLECAR improves post-event feasibility from 0.3880 to 0.8400, terminal success from 0.5800 to 0.8400, and normalized repair burden from 0.5357 to 0.2964 relative to legacy point-schedule repair. A point baseline with CAR-like repair closes part of the gap, but envelopes still retain higher post-event feasibility and lower repair burden, while pooled one-step survivability remains near parity. The result points to a simple execution principle: after disturbances, live windows, couplings, and alternatives serve the executor better than a point schedule rebuilt after the fact.
Diverse Representative Rashomon Sets for Sparse Generalized Additive Models
Varun Babbar ⋅ Christopher Li ⋅ Chudi Zhong ⋅ Cynthia Rudin
The Rashomon set paradigm seeks to uncover the full collection of near-optimal predictive models from a given function class. Access to this set allows users to better understand predictive multiplicity, interact with models, and select those that best satisfy domain-specific constraints. For sparse generalized additive models (GAMs), however, existing approaches that approximate the Rashomon set only apply to restricted subsets of features, overlooking the fact that equally accurate models can rely on many different feature subsets. We address this gap by representing the Rashomon set through a maximally diverse collection of models. We define a Metropolis-Hastings style approach for sampling maximally diverse GAMs from the Rashomon set. Experiments show that our methods can produce substantially more diverse solutions across all tested diversity metrics, providing users with actionable alternatives that maintain similar predictive performance.
Diversity Combining for Multi-Path LLM Reasoning
Guangsheng Yu ⋅ Litianyi Zhang ⋅ Qin Wang ⋅ Xu Wang ⋅ Mingyuan Li ⋅ Shaoxiong Ji ⋅ Ren Ping Liu ⋅ Massimo Piccardi
Multi-path reasoning methods such as self-consistency (SC) sample $K$ reasoning paths and choose the most frequent answer. However, their gains quickly plateau as $K$ increases, and existing methods do not predict when this saturation will occur. We formalize multi-path LLM reasoning as a diversity combining problem from wireless communications: each path is a noisy channel observation, and pairwise path correlation limits the effective number of independent votes, creating a finite saturation ceiling. Generalized least squares (GLS) analysis shows that, under exchangeability, the optimal symmetric linear combiner of latent embeddings is uniform, supporting majority vote as the natural default in standard SC while leaving room for weighting or pruning under heterogeneous prompt-template branches. Across 5 models and 12 benchmarks, prompt-template diversity reduces path correlation in $55$ of $57$ valid cells, with the strongest effect on open-ended QA. We derive an Adaptive-K rule that uses a four-path pilot to select $K^*$, retaining $97$--$104\%$ of MV@$K{=}32$ accuracy across Math, QA, and NLU. Our code is hosted at \url{https://anonymous.4open.science/r/DiversityCombining-4905}.
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Hongchen Wei ⋅ Yuanzhe Wang ⋅ Bei Liu ⋅ Yifan Yang ⋅ Qi Dai ⋅ Kai Qiu ⋅ Yunsheng Li ⋅ Dongdong Chen ⋅ Chong Luo ⋅ Zhenzhong Chen ⋅ Baining Guo
Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-augmented systems usually select evidence from a static index before generation, while recent agentic systems add multi-turn tool use but often rely on frozen proprietary backbones whose behavior is set by prompts. We present DocAtlas, a system that treats long-document understanding as a mutable-state information-seeking process. We instantiate DocAtlas as a mutable document harness: an external environment that determines what document information is searched, read, stored, reviewed, and shown to the model at each step. Given a document and question, the harness exposes search, reading, note-taking, and review tools, maintains a hierarchical tree and note store, and updates both as the agent records evidence. DocAtlas combines self-improving retrieval, selective evidence access, and active working memory under a fixed context budget. The same harness supports inference-time use with large VLMs and end-to-end reinforcement learning for compact VLM agents. With GPT-5.4, DocAtlas reaches 71.4\% on MMLongBench-Doc, exceeding the human-expert reference of 65.8\%. A Qwen3.5-4B VLM trained with end-to-end RL in the DocAtlas environment reaches 63.7\%, compared with a 54.4\% direct-input baseline, showing that mutable document-harness design can improve compact document agents by a large margin.
Do Coding Agents Deceive Us? Detecting and Preventing Cheating via Capped Evaluation with Randomized Tests
Thanawat Lodkaew ⋅ Johannes Ackermann ⋅ Soichiro Nishimori ⋅ Nontawat Charoenphakdee ⋅ Masashi Sugiyama ⋅ Takashi Ishida
A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving the intended task, producing \emph{deceptive} performance. This makes evaluation scores unreliable as measures of true task-solving ability. We propose CapCode, a framework for constructing coding datasets with randomized tests whose best achievable non-cheating performance is deliberately capped below one. This capped-performance design gives evaluation scores a clearer interpretation: scores substantially above the cap are implausible and therefore provide evidence of cheating. To prevent cheating, we propose CapReward, a reward design based on the CapCode principle to discourage optimization beyond the cap. Experiments across multiple datasets show that CapCode detects cheating while preserving performance ranking, and CapReward reduces cheating behavior, yielding models that better follow the intended task specification.
DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding
Xiang Feng ⋅ Jiawei Zhou ⋅ Zhangfeng Huang ⋅ Kewei Wang ⋅ Shanshan Ye ⋅ Jinxin Hu ⋅ Zulong Chen ⋅ Yong Luo ⋅ Jing Zhang
Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation beyond end-to-end answer accuracy. We introduce DocScope, a benchmark that formulates long-document QA as a structured reasoning trajectory prediction problem: given a complete PDF document and a question, the model outputs evidence pages, supporting evidence regions, relevant factual statements, and a final answer. We design a four-stage evaluation protocol—Page Localization, Region Grounding, Fact Extraction, and Answer Verification—that audits each level of the trajectory independently through inter-stage decoupling, with all judges selected and calibrated via human alignment studies. DocScope comprises 1,124 questions derived from 273 documents, with all hierarchical evidence annotations completed by human annotators. We benchmark 6 proprietary models, 12 open-weight models, and several domain-specific systems. Our experiments reveal that answer accuracy cannot substitute for trajectory-level evaluation: even among correct answers, the highest observed rate of complete evidence chains is only 29\%. Across all models, region grounding remains the weakest trajectory stage. Furthermore, the primary difficulty stems from aggregating evidence dispersed across long distances and multiple document clusters, while an oracle study identifies faithful perception and fact extraction as the dominant capability bottleneck. Cross-architecture comparisons further suggest that activated parameter count matters more than total scale. The benchmark and code will be publicly released.
The main challenge of long-tailed deep clustering is that imbalanced class frequencies cause head classes to dominate the partitioning of the representation space, which in turn submerges or distorts the cluster structures of tail classes. Existing methods typically alleviate this issue through heuristic head-tail partitioning or fixed auxiliary clustering. However, such approaches rely on manually specified discrete granularities and therefore struggle to capture the continuous latent structure of real-world long-tailed data. To address this limitation, we propose DPLC, a Dirichlet Process guided Long-tailed Clustering method that adaptively models the latent sub-cluster structure from a nonparametric Bayesian perspective. Specifically, DPLC periodically extracts soft latent sub-clusters from self-supervised embeddings and leverages them as data-driven rebalancing signals to mitigate the clustering bias induced by head-class dominance. By requiring neither a predefined number of auxiliary clusters nor explicit binary head-tail partitioning, DPLC enables more flexible modeling of complex long-tailed distributions. Extensive experiments show that DPLC consistently outperforms existing methods across multiple long-tailed clustering benchmarks and a wide range of imbalance ratios.
Drift-React: One-step Generation of Reaction Pathways via SE(3) Drifting Fields
Rémi Schlama ⋅ Philippe Schwaller
Mapping reaction pathways and transition states (TS) is fundamental to chemistry but computationally expensive at scale. The minimum energy pathway (MEP) dictates reaction rates and mechanisms, yet recovering it via electronic-structure methods requires thousands of costly force evaluations. Recent generative models accelerate TS identification but require slow iterative inference and only predict isolated saddle-point snapshots, missing the continuous reaction trajectory. We introduce Drift-React, an $\mathrm{SE}(3)$-equivariant generative framework that predicts complete reaction pathways in a single forward pass from only reactant and product geometries. By shifting distribution evolution to training via a Sinkhorn-weighted drifting field, Drift-React eliminates both the iterative force evaluations of NEB-style methods and the sequential ODE/SDE integration of diffusion and flow matching models. Evaluated on the Transition1x and Halo8 datasets, our one-step model generates physically consistent MEPs that accurately capture energetic bottlenecks and enable arbitrary-resolution sampling along the reaction coordinate. For isolated TS prediction, Drift-React matches the sub-Ångström accuracy of state-of-the-art iterative models while delivering orders-of-magnitude acceleration, clearing a major computational bottleneck for large-scale reaction network exploration.
DRILL: Training World Models to Improve Policies, Not Predict Pixels
Zhaolu Kang ⋅ Tailong Luo ⋅ Sihan Liu ⋅ Guangyuan Dong ⋅ Siheng Wang ⋅ Lei Wei ⋅ Shuaibo Li ⋅ Rongchao Zhang ⋅ Zhen Tian ⋅ richeng xuan ⋅ Zhichao Hu
World models are increasingly used to reduce environment interaction when post-training vision-language-action (VLA) policies with reinforcement learning. Yet they are typically trained to predict pixels one step ahead, even though their role here is not merely to render visually faithful videos but to produce learning signals from imagined rollouts that improve the policy. We show that this mismatch is substantial: lower pixel MSE can produce worse policy gradients, and a world model trained for policy utility can yield better downstream policies despite 17\% higher pixel MSE on the next frame. We propose the Downstream-Return Imagination Learning Loop (DRILL), a bilevel framework that updates the world model to maximize the policy's return after an inner GRPO update on imagined rollouts. DRILL has two instantiations. DRILL-VWI is a closed-form surrogate, requiring no additional simulator rollouts, that reweights prediction errors by policy visitation, the group-relative GRPO advantage, and model confidence. DRILL-IMG is a full meta-gradient method that differentiates through the inner update using truncated Hessian-vector products. We show that DRILL-VWI recovers the local one-step component of the DRILL-IMG meta-gradient under score-compatibility and small-residual assumptions. Across five manipulation simulators and two VLA backbones, DRILL-IMG raises mean success rate by 13.1 points over WMPO and 7.9 points over RLVR-World, and complements RLVR pretraining rather than competing with it. More broadly, we believe video world models for VLA reinforcement learning can be productively trained by the policy gradients they produce, not only by pixel-level prediction.
DriveSpatial: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving
Anh Hao Vo ⋅ Khoa Vo ⋅ Phu Loc Nguyen ⋅ Sieu Tran ⋅ Duc Nguyen ⋅ Ngo X Cuong ⋅ Gladys Gawugah ⋅ Sreevenkata A Godavarthi ⋅ Chase Rainwater ⋅ Nghi Bui ⋅ Anh Nguyen ⋅ Duy M. H. Nguyen ⋅ Ngan Le
Spatiotemporal intelligence in autonomous driving (AD) requires an agent to integrate multi-view observations into a coherent scene representation, maintain object continuity across viewpoints and time, and reason about spatial relations, interactions, and future dynamics. However, existing AD vision-language benchmarks largely focus on single-view, static, ego-centric, or single-source question answering, leaving it unclear whether current Vision-Language Models (VLMs) can truly construct and reason over dynamic driving scenes. We introduce DriveSpatial a benchmark of 15.6K human-verified QA pairs across 20 tasks from five large-scale AD datasets. DriveSpatial evaluates four abilities: Cognitive Scene Construction, Multi-view Relational Understanding, Temporal Reasoning, and Generalization. Unlike prior benchmarks, DriveSpatial is generated from a dynamic multi-relational scene graph that encodes object states, spatial relations, interactions, camera visibility, and temporal correspondences, enabling QA pairs that enforce genuine cross-view and spatiotemporal reasoning. Evaluating 15 representative VLMs reveals a substantial human–model gap: the strongest model trails humans by 28.4 points, with Cognitive Scene Construction emerging as the key bottleneck. Further diagnostics show that language-only prompting is insufficient, while explicit BEV grounding consistently improves performance. These results suggest that current VLMs lack the scene-construction ability needed for reliable spatiotemporal driving intelligence. DriveSpatial and its construction pipeline will be released to support future research.
DrPO: Drifting Preference Optimization for One-Step Generative Models
Zhou Jiang ⋅ Yandong Wen ⋅ Zhen Liu
While one-step text-to-image generators, a fast and deployment-friendly class of visual generative models, can produce samples with a single forward pass, existing preference-finetuning methods fail to simultaneously achieve efficient adaptation, reward-model agnosticism, and diverse generation. In this work, we propose Drifting Preference Optimization (DrPO), an online preference-finetuning method for one-step generators. Inspired by the recent work of Drifting Models, DrPO first uses the target reward model to construct an on-policy dipole reward model from positive and negative samples. This dipole model induces a preference drifting field in latent feature space, whose gradients are then used to optimize the generator without backpropagating through the target reward model. Empirically, we evaluate DrPO across multiple reward models and SD-Turbo and SDXL-Turbo one-step generators, including HPSv3 and GenEval as representative alignment benchmarks, and validate that DrPO can efficiently and effectively finetune one-step generative models. Preliminary experiments further suggest that this sample-based gradient synthesis can be extended to offline preference finetuning.
dStructAD: Domain-Level Structured Normality Representation with Variation Calibration for Time Series Anomaly Detection
Shiwang Xing ⋅ Jianwei Niu ⋅ Tao Ren
Time Series Anomaly Detection (TSAD) is critical for ensuring reliability in real-world systems such as industrial monitoring, healthcare, and online services. However, learning-based TSAD methods trained on a single normal-only TargetSet, namely the target dataset, often suffer from incomplete and biased knowledge of normality, causing domain-level false alarms when unseen but valid normal patterns appear at test time. A natural extension is to leverage DomainSet, composed of datasets from one domain, to enrich domain-level normality. Yet, our analysis shows that simply incorporating DomainSet could make domain-admissible normal variations and TargetSet-specific anomalies highly entangled in the learned representations, weakening anomaly separability. Motivated by human experts who establish normal patterns, calibrate admissible variations, and identify true anomalies, we propose dStructAD, a domain-level structured representation framework for time series anomaly detection, to address this challenge. Specifically, dStructAD first builds structured domain-level normality based on the Kolmogorov-Arnold network, and then calibrates admissible variations via two phases. Phase 1 learns shared domain-level semi-structural normality knowledge from DomainSet, while Phase 2 performs structure-preserving TargetSet calibration to absorb admissible variations without collapsing anomaly separability. Across six benchmarks, dStructAD delivers consistent improvements over SOTAs, with 3% average gain in overall performance. Code could be available at https://anonymous.4open.science/r/dStructAD-8F03.
Transfer learning seeks to improve efficiency in a target regression task by borrowing information from related external source dataset(s). Existing approaches achieve this by enforcing Euclidean proximity between the target and source regression parameters or, more recently, by encouraging their directional alignment through angle-based penalties. While such angle-based penalties mitigate the sensitivity of the transfer learning methods to the scale differences between source and target regression coefficients, they still rely on a point estimate of the source regression parameter and thus provide no direct mechanism to incorporate uncertainty in the source information. To that end, we propose a probabilistic framework for transfer learning that operates at the level of a \emph{scale-invariant directional distribution} of the source regression parameter estimate, rather than its point estimate, to enable robust, scale-invariant and uncertainty aware transfer under heterogeneous source information. Specifically, we project the sampling distribution of the source regression estimate onto the unit sphere, thereby extracting its scale-free directional distribution. Transfer from source to target is then induced through a novel \emph{angular cone prior} which shrinks the target regression parameter towards directions in which the source distribution assigns high probability, while allowing its magnitude to be learned exclusively from the target dataset. The proposed construction admits exact finite-sample characterizations of posterior angular concentration and posterior computation remains tractable via careful low-dimensional augmentation in the parameter space. Simulation studies and application on a benchmark hypertension prediction task based on the National Health and Nutrition Examination Survey (NHANES) datasets demonstrate that the proposed method yields marked gains over existing approaches in settings with substantial uncertainty in the source regression estimate and source--target scale mismatch.
Dynamic Representation Modeling for Federated Medical Image Domain Generalization
Yuxi Ma ⋅ Lin Zhao ⋅ Jing Yang ⋅ jiacheng wang ⋅ Liansheng Wang
Federated Domain Generalization (FDG) aims to learn a global model robust to heterogeneous domain shifts without sharing raw data, a critical challenge in multi-center medical imaging. Most existing methods implicitly assume static inter-domain discrepancies, overlooking the fact that client representations evolve continuously during local optimization. We identify this phenomenon as representation dynamics, which refers to the temporal evolution of latent features induced by training updates rather than shifts in underlying data distributions. Such dynamics often lead to unstable aggregation, particularly in heterogeneous medical imaging scenarios. To address this issue, we propose Dynamic Knowledge Tracked FDG (DKT-FDG), a framework that explicitly tracks these dynamics by leveraging flow matching to model continuous representation trajectories. By aligning representation evolution instead of static snapshots, DKT-FDG improves robustness to domain shifts. Comprehensive experiments on multiple multi-center medical benchmarks demonstrate consistent improvements over state-of-the-art FDG methods, highlighting the importance of modeling representation dynamics in federated learning.
Dynamics-Aware Sparse Attention for Efficient Autoregressive Video Diffusion
Yuanyu He ⋅ Zhuokun Chen ⋅ Yefei He ⋅ Zhiwei Tang ⋅ Jiasheng Tang ⋅ Yinghao Yu ⋅ Jianfei Cai ⋅ Bohan Zhuang
Autoregressive Diffusion Transformers have emerged as a powerful paradigm for long video generation. However, they are often constrained by prohibitive computational overhead and substantial memory footprints. This inefficiency arises from dense computation mechanisms that allocate uniform processing resources across all spatiotemporal regions, failing to leverage the inherent temporal redundancy of video where motion is typically confined to sparse areas. In this paper, we propose **Dynamics-Aware Sparse Attention (DASA)**, a unified framework that shifts from dense computation to sparse, motion-driven generation to eliminate both storage and computational redundancies. Inspired by video codec principles, our approach explicitly decouples video content into dynamic and static regions. To mitigate memory bottlenecks, we selectively compress the historical KV cache by retaining only dynamic features within the generated chunk. Furthermore, to reduce computational overhead, we employ a region-level computation allocation strategy based on localized dynamics. Additionally, a dynamics-guided distillation method is introduced to enhance performance via a lightweight distillation training process. Experiments on VBench demonstrate that our approach accelerates autoregressive video generation by 2.16$\times$, while preserving high visual fidelity.
E0: Expressive Fine-Grained Discrete Action Prediction for Vision-Language-Action Models via Tweedie Discrete Diffusion
Zhihao Zhan ⋅ Jiaying Zhou ⋅ Likui Zhang ⋅ Qinhan Lyu ⋅ Hao Liu ⋅ Weizheng Li ⋅ jusheng zhang ⋅ Ziliang Chen ⋅ Tianshui Chen ⋅ Ruifeng zhai ⋅ Keze Wang ⋅ Liang Lin ⋅ Guangrun Wang
Vision–Language–Action (VLA) models offer a unified framework for robotic manipulation by integrating visual perception, language understanding, and control generation. However, existing VLA systems still struggle to generalize across diverse tasks, scenes, and camera viewpoints, and often produce coarse or unstable actions. We argue that these limitations are closely tied to the structural properties of actions in VLA settings, including the inherent multi-peaked nature of action distributions, the token-based symbolic reasoning of pretrained VLM/VLA backbones, and the effective finite resolution imposed by real-world robotic control. Motivated by these properties, we introduce E0, a tweedie discrete diffusion framework that formulates action generation as iterative denoising over quantized action tokens. By operating in a discrete action space with a principled diffusion process, E0 naturally aligns with token-based reasoning, supports fine-grained yet executable action control, and avoids the distributional mismatch of masking-based discrete diffusion. Experiments on LIBERO, VLABench, ManiSkill, and a real-world Franka arm demonstrate that E0 achieves state-of-the-art performance across diverse environments, outperforming strong baselines by 8.3\% on average.
EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields
Zhaoyang Yang ⋅ Yurun Jin ⋅ Lizhe Qi ⋅ Cong Huang ⋅ Kai Chen
Pretrained video diffusion models provide powerful spatiotemporal generative priors, making them a natural foundation for robotic world models. While recent world-action models jointly optimize future videos and actions, they predominantly treat video generation as an auxiliary representation for policy learning. Consequently, they insufficiently explore the inverse problem: leveraging action signals to guide video synthesis, thereby often failing to preserve precise robot spatial geometry and fine-grained robot-object interaction dynamics in the generated rollouts. To bridge this gap, we present EA-WM, an Event-Aware Generative World Model that effectively closes the loop between kinematic control and visual perception. Rather than injecting joint or end-effector actions as abstract, low-dimensional tokens, EA-WM projects actions and kinematic states directly into the target camera view as Structured Kinematic-to-Visual Action Fields. To fully exploit this geometrically grounded representation, we introduce event-aware bidirectional fusion blocks that modulate cross-branch attention, capturing object state changes and interaction dynamics. Evaluated on the comprehensive WorldArena benchmark, EA-WM achieves state-of-the-art performance, outperforming existing baselines by a significant margin.
Echoes of Error: Residual-Directional Local Rollout Consistency for Timeseries Forecasting
Zheng Wang ⋅ Kaixuan Zhang ⋅ Xiaonan Lu ⋅ Wanfang Chen
In time-series forecasting, rollout errors are not merely one-step fitting discrepancies: once fed back into the evolving state, they can perturb future predictions. We formalize this effect through local rollout consistency, which compares the next prediction produced under teacher forcing with that produced under free running. In principle, this consistency can be enforced directly by tracking the residual-induced perturbation of the next state. In practice, however, the residual-to-state alignment operator is often difficult to design when the input state and output block have different dimensions, heterogeneous components, or architecture-dependent layouts. We therefore use a first-order local surrogate that penalizes the Jacobian-vector product of the predictor along the actual feedback perturbation, suppressing precisely the directions through which self-generated errors remain visible to future predictions. We further give the same loss a geometric interpretation near a predictive manifold: by controlling residual-induced directions that probe the normal bundle, the objective yields theoretical conditions under which free-running trajectories remain inside a stability tube around the manifold. Experiments on time series forecasting tasks support this view, showing that the proposed regularization improves predictive accuracy. More broadly, our results suggest a training principle for time-series forecasters and sequential predictors: a model should not only make small one-step errors, but also learn not to let its own errors echo into the future.
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
Jiameng Li ⋅ Minye Wu ⋅ Jiezhang Cao ⋅ Aleksei Tiulpin ⋅ Matthew Blaschko
Long-form video understanding remains challenging for Video Large Language Models (VideoLLMs), as the dense frame sampling introduces massive visual tokens while sparse sampling risks missing critical temporal evidence and leading to LLM hallucination. Existing training-free token reduction methods either treat videos equally as static images or rely on segment-level merging heuristics, which weaken fine-grained spatiotemporal modeling and introduce additional overhead. In this paper, we propose EchoPrune, a lightweight and training-free token pruning method that improves temporal resolution under a fixed LLM-side visual token budget. Our core idea is to interpret redundant video tokens as temporal echoes: if a token is well reconstructed from the previous frame, it is merely a temporally redundant echo; otherwise, it may capture new events, motion, or query-relevant visual evidence. Based on this insight, EchoPrune scores visual tokens by (i) query-guided crossmodal relevance and (ii) temporal reconstruction error, measured by correspondence matching and echo matching across consecutive frames. The selected tokens preserve task-relevant cues and temporal novelty while suppressing predictable redundancy, allowing VideoLLMs to observe more frames without increasing the decoding budget. Extensive experiments on LLaVA-OV, Qwen2.5VL, and Qwen3VL across six video understanding benchmarks show that EchoPrune enables VideoLLMs to process up to $\mathbf{20\times}$ frames under the same token budget, yielding improved performance ($\mathbf{+8.6\\%}$) $\%$and inference speedup ($\mathbf{5.6\times}$ for prefilling) on Qwen2.5VL-7B.
ECLIPSE: A Spacecraft Rendezvous Trajectories Dataset with Controlled In-Orbit Lighting Conditions
Nidhal Eddine Chenni ⋅ Arunkumar Rathinam ⋅ Abid Ali ⋅ Djamila Aouada
Deep learning has become the dominant paradigm for vision-based spacecraft pose estimation, where models are predominantly trained on synthetic data due to the prohibitive cost of acquiring real in-orbit imagery. When deployed, these models suffer significant performance degradation due to the sim-to-real domain gap, which existing benchmarks address by pairing synthetic training data with Hardware-in-the-Loop (HIL) real imagery. Yet a critical factor remains consistently overlooked: in real orbital scenarios, illumination varies with the spacecraft's position along its orbit, inducing severe appearance shifts that existing benchmarks neither control nor study. We introduce ECLIPSE, the first benchmark designed to isolate illumination as an independent experimental variable. ECLIPSE pairs a large-scale photorealistic synthetic training set with a fully labeled HIL evaluation set of 14 approach trajectories, each replayed identically under 4 controlled lighting conditions, enabling rigorous evaluation of illumination-robust and domain-adaptation methods. Through this controlled design, our experiments confirm that lighting is the dominant driver of the sim-to-real performance gap in spacecraft pose estimation.
EcoGEO: Trajectory-Aware Evidence Ecosystems for Web-Enabled LLM Search Agents
Hengwei Ye ⋅ Jiasheng Mao ⋅ Zhenhan Guan ⋅ Zheng Tian
Web-enabled LLM agents are changing how online information influences search outcomes. Existing Generative Engine Optimization (GEO) studies mainly focus on individual webpages. However, agentic web search is not a single-document setting: an agent may issue queries, crawl pages, follow links, reformulate searches, and synthesize evidence across multiple browsing steps. Influence therefore depends not only on page content, but also on how pages are organized, connected, and encountered along the agent's browsing trajectory. We study this shift through Ecosystem Generative Engine Optimization (EcoGEO), which treats GEO as an environment-level influence problem for web-enabled LLM agents. To instantiate this perspective, we propose TRACE, a Trajectory-Aware Coordinated Evidence Ecosystem. Given a recommendation query and a fictional target product, our method builds a controlled evidence environment that coordinates an agent-facing navigation entry page with heterogeneous support pages. These pages use shared terminology, internal links, and consistent product attributes to introduce, verify, and reinforce the target product. We evaluate our method on OPR-Bench, a benchmark for open-ended product recommendation. Experiments show that it consistently outperforms page-level GEO baselines in final target recommendation. Trajectory-level metrics further show increased initial target-result crawls, target-specific follow-up searches, and internal-link crawls, suggesting that the gains come from shaping the agent's evidence-acquisition process rather than merely adding more target-related content. Overall, our findings support an ecosystem research paradigm for GEO, where web-enabled LLM agents are studied in relation to the broader evidence environments that guide search, browsing, and answer synthesis.
EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
Xueyu Hu ⋅ Jinxiang Xia ⋅ Shengze Xu ⋅ Kangqi Song ⋅ Yishuo Yuan ⋅ Guibin Zhang ⋅ JinCheng Ren ⋅ Boyu Feng ⋅ Li Lu ⋅ Tieyong Zeng ⋅ Jiaheng Liu ⋅ Minghao Liu ⋅ He Zhu ⋅ Eleanor Jiang ⋅ Wei Wang ⋅ Wangchunshu Zhou
Long-horizon planning is widely recognized as a core capability of autonomous LLM-based agents; however, current evaluation frameworks suffer from being largely episodic, domain-specific, or insufficiently grounded in persistent economic dynamics. We introduce EcoGym, a generalizable benchmark for continuous plan-and-execute decision making in interactive economies. EcoGym comprises three diverse environments: Vending (adapted from the closed-source Vending-Bench, with full open-source release), Freelance (new), and Operation (new), implemented in a unified decision-making process with standardized interfaces, and budgeted actions over an effectively unbounded horizon (1000+ steps if 365 day-loops for evaluation). The evaluation of EcoGym is based on business-relevant outcomes (e.g., net worth, income, and DAU), targeting long-term strategic coherence and robustness under partial observability and stochasticity. Experiments across eleven leading LLMs expose a systematic tension: no single model dominates across all three scenarios. Critically, we find that models exhibit significant suboptimality in either high-level strategies or efficient actions executions. EcoGym is released as an open, extensible testbed for transparent long-horizon agent evaluation and for studying controllability–utility trade-offs in economic settings.
EditFlowSR: Revisable Expression Generation for Symbolic Regression
Yuhong Xu ⋅ Ziyi Yang ⋅ Jinghui Zhong
Symbolic regression (SR) aims to discover compact, interpretable mathematical expressions from data. This goal inherently requires iterative refinement rather than one-shot construction. Autoregressive methods lack the ability to perform targeted corrections, so whenever structural flaws appear, the remainder of the sequence must be discarded and regenerated. Meanwhile, search-based methods rely on costly and indirect replacement through enumeration. We propose EditFlowSR, which reframes SR as iterative editing of expression trees through insertions, deletions, and substitutions. To reconcile fitting accuracy with algebraic simplicity, we introduce dual trajectory supervision. One trajectory teaches the model to construct expressions from data, while the other teaches it to algebraically simplify them, and both are unified under the shared editing framework. At inference time, EditFlowSR starts from a randomly initialized expression and progressively refines it through iterative editing, simultaneously constructing valid structure and removing redundancy. Across the standard SRBench, EditFlowSR achieves first-rank Pareto dominance in predictive accuracy, expression tree size, and time complexity. These results establish edit-based generation as a principled and effective paradigm for end-to-end symbolic regression. Code is available at \url{https://anonymous.4open.science/r/EditFlowSR_NIPS2026}.
EDMA: Entropy-Driven Multimodal Answering
Emanuele Mezzi ⋅ Gertjan Burghouts ⋅ Fabio Massacci ⋅ Mengyuan Zhang
Multimodal question answering (MMQA) is based on integrating heterogeneous data sources, selectively leveraging relevant modalities, and ignoring distractors. Existing approaches based on Multimodal Large Language Models (MLLMs) improve performance by splitting the task into intermediate steps, but do not quantify how each modality contributes to reaching the final answer. We argue that quantifying the effect that each modality has on the answering process improves process traceability and final performance. Thus, we propose Entropy-Driven Multimodal Answering (EDMA), which uses logical entropy to formalise multimodal QA as an iterative process of entropy reduction and to quantify the contribution of each modality in distinguishing correct from incorrect candidate answers. This enables the construction of quantifiable answering trajectories, where each step is associated with a measurable entropy reduction. EDMA introduces (i) modality-conditioned partitioning to estimate unimodal entropy reduction, (ii) entropy-driven multimodal fusion to capture complementary information across modalities, and (iii) entropy-based selection of the answering trajectory that minimises the most logical entropy. Experiments on multimodal QA benchmarks show that EDMA consistently outperforms prompting-based baselines and state-of-the-art methods. Averaged across datasets, EDMA achieves gains of $12.8$ and $19.8$ F1 points over the strongest prompting baseline and state-of-the-art system, respectively. Moreover, we show that higher entropy correlates with more false positives and, through controlled interventions on the choice of the answering trajectory, provide evidence that lower-entropy partitions causally reduce false positives, validating entropy level as a reliable proxy for performance. We share our code at https://anonymous.4open.science/r/EDMA/README.md.
Effect-Driven Skill Abstractions for Offline Reinforcement Learning
Burcu Kılıç ⋅ David Drexel ⋅ Emre Ugur ⋅ Justus Piater
We seek to learn discrete skill abstractions to facilitate long-horizon, goal-conditioned reinforcement-learning (RL) tasks. Recent approaches have shown success by quantizing actions via clustering or reconstruction. However, these objectives lead to skills that are not directly related to what the current state affords or what the goal requires. Instead, we argue that goal-directed skills should be driven by how they affect the state of the agent. To realize this, we introduce a novel quantization method for action chunks that optimizes next-state predictions conditioned on the current state and action. This objective encourages skills to be distinguished by the future states they lead to, resulting in a vocabulary that spans diverse outcomes reachable from a given state. Our framework can easily be paired with common offline RL methods by replacing their continuous actor head with a discrete one that selects among the learned skills. We can execute these skills in the environment by converting them back to continuous actions using a trained low-level controller. Experiments on benchmark navigation tasks largely show substantial and consistent improvements over both the continuous-actor baselines and reconstruction-based discretizations.
Effectiveness of Curriculum Learning Depends on Reward Sparsity and Competing Optima
John Vastola ⋅ Ann Huang ⋅ Satpreet Harcharan Singh ⋅ Samuel J Gershman ⋅ Kanaka Rajan
While animals and people tend to learn a task more quickly and reliably when they first train on simpler versions of it, curriculum learning's effectiveness in artificial settings varies, and appears substantially greater in reinforcement learning than in supervised learning. In this work, we construct and analyze a minimal model of policy learning to better understand why. We consider an open-loop, episodic target interception task whose parameters---especially how close the agent must be to the target to receive an informative reward signal---can be chosen so that the agent either receives informative rewards frequently ('dense' rewards), or rarely ('sparse' rewards). We mathematically and numerically analyze the dynamics of REINFORCE-based policy learning in both cases. In the dense reward case, curricula have only marginal benefits; in the sparse case, curricula can be required for learning to occur at all. We find that this is especially true when the agent has the conflicting goals of minimizing effort and accumulating high reward, since this can produce distinct reward landscape optima which compete with one another. Our characterization of useful curricula in this setting indicates that they address two issues: they (i) make rewards less sparse to speed up learning, and (ii) shape the reward landscape to steer agents away from bad local optima.
Modern data workflows are inherently adaptive, repeatedly querying the same dataset to refine and validate sequential decisions, but such adaptivity can lead to overfitting and invalid statistical inference. Adaptive Data Analysis (ADA) mechanisms address this challenge; however, there is a fundamental tension between computational efficiency and sample complexity. For $T$ rounds of adaptive analysis, computationally efficient algorithms typically incur suboptimal $O(\sqrt{T})$ sample complexity, whereas statistically optimal $O(\log T)$ algorithms are computationally intractable under standard cryptographic assumptions. In this work, we shed light on this trade-off by identifying a natural class of data distributions under which both computational efficiency and optimal sample complexity are achievable. We propose a computationally efficient ADA mechanism that attains optimal $O(\log T)$ sample complexity when the data distribution is dense with respect to a known prior. This setting includes, in particular, feature-label data distributions arising in distribution-specific learning. As a consequence, our mechanism also yields a sample-efficient (i.e., $O(\log T)$ samples) statistical query oracle in the distribution-specific setting. Moreover, although our algorithm is not based on differential privacy, it satisfies a relaxed privacy notion known as Predicate Singling Out (PSO) security (Cohen and Nissim 2020). Our results thus reveal an inherent connection between adaptive data analysis and privacy beyond differential privacy.
Efficient Agentic GPU Kernel Optimization with a Compact Domain-Specific Language and Speed-of-Light Guidance
Siva Kumar Sastry Hari ⋅ Vignesh Balaji ⋅ Sana Damani ⋅ Qijing Huang ⋅ Christos Kozyrakis
LLM agents can optimize GPU kernels, but each candidate requires generation, compilation, correctness testing, and profiling. Our objective is to reduce the number of expensive attempts needed to find a correct, fast kernel. We study two complementary ways to improve this attempt efficiency by changing the representation that the agent writes in and adding a first-principles performance signal that estimates remaining headroom. We introduce $\mu$CUTLASS, a compact in-context learnable DSL that exposes high-impact CUTLASS optimization choices while hiding template plumbing, and pair it with Speed-of-Light (SOL) guidance for search steering, budget allocation, and integrity checking. On 59 KernelBench problems on H100, switching GPT-5-mini from low-level code generation to $\mu$CUTLASS changes a $0.40\times$ geomean regression versus PyTorch into a $1.27\times$ speedup. Adding SOL-guided steering, it reaches $1.56\times$. Across model tiers, this combination lets weaker models match or exceed stronger raw-code baselines, and SOL-guided scheduling saves 19--43\% of tokens while retaining at least 95\% geomean speedup. We also show that integrity filtering is essential because, without SOL-assisted checks, reported speedups can be inflated by up to $1.9\times$ by benchmark-gaming or PyTorch-only solutions.
Efficient and Simple Data Mixing All The Time
Michael Hu ⋅ Apurva Gandhi ⋅ Kyunghyun Cho ⋅ Tal Linzen ⋅ Pratyusha Sharma
Data mixing is a consequential problem throughout language model training. In pretraining, data composition is a key determinant of model quality; in continual learning and adaptation, it governs what is retained and acquired. Yet existing data mixing methods address only one phase of this lifecycle at a time: some require smaller proxy models tied to a single training phase, others assume a fixed domain set, and continual learning lacks principled guidance altogether. We argue that data mixing is fundamentally an online decision making problem---one that recurs throughout training and demands a single, unified solution. We introduce OP-Mix (On-Policy Mix), a data mixing algorithm that operates across the entire language model training lifecycle. Our main insight is that candidate data mixtures can be cheaply simulated by interpolating between low-rank adapters trained directly on the current model, eliminating separate proxy models and ensuring the search is always grounded in the model's actual learning dynamics. Across pretraining, continual midtraining, and continual instruction tuning, OP-Mix consistently finds near-optimal mixtures while using a fraction of the compute of the baselines. In pretraining, OP-Mix improves upon training without mixing by 6.3% in average perplexity. For continual learning, OP-Mix matches the performance of both retraining and on-policy distillation while using 66% and 95% less overall compute, respectively. OP-Mix suggests a different view of language model training: not a sequence of distinct phases, but a single continuous process of learning from data.
Efficient Brain-to-Speech Decoding with Fixed-Delay Spiking Neural Networks
Aleksandra Wisniewska ⋅ Seo-Hyun Lee ⋅ Seong-Whan Lee
Brain-to-speech decoding aims to restore communication by translating neural activity into speech-related representations. From a modeling perspective, this task can be formulated as a weakly supervised neural sequence decoding problem, where phoneme sequences must be recovered without frame-level annotations. Existing speech decoders commonly rely on recurrent sequence models to integrate temporal context, but their dense sequential computation can limit their suitability for resource-constrained portable or implantable neuroprosthetic systems. In this study, we propose a residual spiking neural network (SNN) for phoneme-level decoding of intracortical speech signals under constrained model capacity and weak sequence-level supervision. Our approach incorporates biologically inspired fixed connection-specific delays to integrate temporal context, providing each postsynaptic unit with access to recent presynaptic spike activity without recurrent state or dense temporal convolutions. Because each directed connection is assigned a single fixed delay, the temporal field is expanded without adding trainable delay parameters. Across 500k, 2M, and 5M core-parameter budgets, fixed delays consistently improve phoneme error rate (PER), with gains saturating once the delay range covers the task-relevant temporal horizon. Boundary-aligned CTC analysis further suggests that delayed temporal context reduces local uncertainty around phoneme transitions, indicating that fixed delays support weakly supervised phoneme alignment rather than simply increasing model capacity. Compared with recurrent, convolutional, and feedforward baselines under matched core-parameter budgets, the proposed SNN achieves the best PER at 500k and 2M parameters and the second-best PER at 5M parameters, while reducing estimated operation-level energy by approximately 2-4$\times$ relative to the strongest non-spiking baselines. These results suggest that fixed-delay SNNs provide an efficient temporal modeling mechanism for brain-to-speech decoding, potentially enabling lightweight decoding for on-device speech neuroprostheses.
Efficient Lookahead Encoding and Abstracted Width for Learning General Policies in Classical Planning
Michael Aichmüller ⋅ Simon Ståhlberg ⋅ Martin Funkquist ⋅ Hector Geffner
Generalized planning aims to learn policies that generalize across large collections of instances within a classical planning domain. Recent approaches using Graph Neural Networks (GNNs) have shown the ability to learn nearly perfect policies for several domains. This work improves on the recently published idea of Iterated Width (IW) policies. Therein, the policy broadens its successor scope through an IW-lookahead search that offers to "jump" over multiple transitions, simplifying the problem structure. Yet, each transition is evaluated individually leading to unscalable compute and expressivity limitations. Furthermore, while an IW(1) search is attractive due to scaling linearly with the number of atoms in a problem, it still becomes inefficient once thousands of objects are considered like in the International Planning Competition (IPC) 2023 benchmark. In this work, we address both limitations. Firstly, we introduce a vastly more efficient holistic encoding of the entire search tree. It jointly represents IW(1)-reachable states only by their relational differences to the current state, which enables Relational GNNs (R-GNNs) to score all transitions in a single forward pass. Secondly, we define Abstracted IW(1) to improve scaling through relational abstraction during novelty checks. Rather than testing fully instantiated atoms, it abstracts each atom by replacing all but one of its arguments with their types. The original atom is then novel if any of its abstracted forms is. This structural compression shifts the scaling of the novelty search to be linear in the number of objects, rather than atoms, still capturing meaningful subgoal structure. Our contributions are evaluated on the hyperscaling IPC 2023 benchmark and across a broad range of domains, including those that require features beyond the $C_2$ logic fragment. The results show that our policies achieve a new state-of-the-art performance, significantly surpassing prior work, including classical planners like LAMA.
Pre-training of Large Language Models is often prohibitively expensive and inefficient at scale, requiring complex and invasive modifications in order to achieve high data throughput. In this work, we present Token-Superposition Training (TST), a simple drop-in method that significantly improves the data throughput per FLOPs during pre-training without modifying the parallelism, optimizer, tokenizer, data, or model architecture. TST is done in two phases: (i) A highly efficient superposition phase where we combine many contiguous tokens into one bag and train using a multi-hot cross-entropy (MCE) objective, and (ii) a recovery phase where we revert back to standard training. We extensively evaluate TST on the scale of 270M and 600M parameters and validate on 3B and a 10B A1B mixture of experts model, demonstrating that it is highly robust in different settings. Ultimately, TST consistently outperforms baseline loss and downstream evaluations, and under equal-loss settings, TST yields up to a 2.5x reduction in total pre-training time at the 10B A1B scale.
Efficient Retrosynthesis Prediction with Integral Flow Matching and Latent Inversion
Tao Yin ⋅ Xiaohong Zhang ⋅ Yinjie Zhu ⋅ Jiacheng Zhang ⋅ Haotian Zou ⋅ Li Huang ⋅ Jiajun Cai ⋅ Zhibin Zhang ⋅ Meng Yan
Fast and accurate retrosynthesis prediction is desired for downstream drug discovery and synthetic planning tasks. Currently, training and sampling state-of-the-art discrete diffusion or flow-based models for retrosynthesis requires significant computational resources. In this work, we propose FlashRetro, an Integral Flow Matching (IFM) framework with Latent Inversion for efficient retrosynthesis prediction. FlashRetro learns finite-interval latent transport rather than local instantaneous velocity, directly matching the transport quantity required for one-step latent inversion and avoiding test-time numerical integration. For inference, FlashRetro uses a principled 1-NFE Latent Inversion rule that maps a noise latent conditioned on the product latent to the reactant latent with a single learned integral-transport update. Experimental results on standard benchmarks (USPTO-50K) show that FlashRetro achieves superior top@1 accuracy, outperforming prior discrete flow-matching methods by 13.5\% and diffusion-based baselines by 14.2\%, while its average inference time is only 5.1 ms per reactant---38$\times$ faster than multi-step diffusion-based models and nearly 22$\times$ faster than recent discrete flow models. These results show that FlashRetro provides an effective path toward highly efficient 1-NFE retrosynthesis prediction with flow-based models.
Efficient Training of Deep Spiking Neural Networks with Input-Driven Derivative-Free Updates
Yongbo Zhang ⋅ Xinzhe Li ⋅ Katsuma Inoue ⋅ Mitsumasa Nakajima ⋅ Toshikazu Hashimoto ⋅ Yasuo Kuniyoshi ⋅ Kohei Nakajima
Spiking neural networks (SNNs) are a promising paradigm for efficient neuromorphic computing, yet their training remains challenging due to the non-differentiability nature of spike firing. Backpropagation through time (BPTT) with surrogate gradients (SG) has become a mainstream approach due to its superior performance. However, this method incurs high computational and memory demands during training, while its update rule depends on specific neuronal dynamics. Furthermore, when targeting neuromorphic hardware implementation, the temporal error-propagation process and the full membrane-potential access required by SG are difficult to realize; these can introduce additional readout overhead, measurement noise, or perturbations to the system dynamics. To address these issues, we propose the input-driven derivative-free training (IDDFT) method. Rather than relying on membrane-potential-based surrogate derivatives, IDDFT constructs derivative-free error signals by applying a relaxed nonlinearity to the inputs, thereby avoiding the temporal error-propagation process and reducing the dependence of the update rule on specific neuronal dynamics. By eliminating access to internal membrane-potential states, our method reduces computational and memory costs, enhances compatibility with neuromorphic hardware, and enables training of black-box SNN models. We validate the IDDFT method through theoretical analysis and systematic experiments, demonstrating its effectiveness as an input-driven, derivative-free mechanism for constructing update signals. When integrated with a biologically inspired training strategy, IDDFT maintains comparable performance while substantially enhancing robustness to the choice of relaxed nonlinearities. Based on the resulting bio-inspired combined framework, we conduct evaluations on CIFAR-10, CIFAR-100, CIFAR10-DVS, and Tiny-ImageNet. Experimental results demonstrate that our method achieves performance comparable to BPTT and state-of-the-art methods, while exhibiting remarkable robustness under adversarial attacks.
Efficient Tree Draft for Long-Context Speculative Decoding
Brian J Chan ⋅ Ning-Chi Huang ⋅ Kai-Chiang Wu
Speculative decoding is an effective technique for accelerating autoregressive models in memory-bound regimes. Prior work has explored tree-based drafting to speculate multiple candidates at each step, improving verification acceptance rates. However, existing systems fail to fully realize the potential of tree-based drafting, leaving significant performance overhead. In this paper, we present FastDraft, a plug-and-play module that improves the efficiency of tree-based drafting in long-context speculative decoding. FastDraft exploits both the shared-prefix I/O pattern across draft branches and the constrained tree structure of speculative decoding, enabling specialized kernels that reduce draft-phase overhead while preserving compatibility with existing inference frameworks. FastDraft requires no modifications to the target model, draft model, or sampling logic. We implement FastDraft within SGLang and compare it against SGLang’s vanilla tree-based drafting backend, demonstrating substantially faster draft execution: up to 2.66× speedup in the draft phase and 1.77× higher throughput over SGLang under long-context settings.
EgoBench: An Interactive Egocentric Multimodal Benchmark for Tool-Using Agents
Yunqi Liu ⋅ Tong Niu ⋅ Zitong Wang ⋅ Zhenlong Dai ⋅ Yuqi Qing ⋅ Weiqiang Wang ⋅ Jian liu
As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to jointly evaluate these capabilities due to challenges in designing strictly coupled multi-capability tasks, simulating natural and task-constrained user feedback, and ensuring objective evaluation of dynamic interaction. To bridge this gap, we introduce EgoBench, the first interactive multimodal benchmark for tool-using agents. EgoBench comprises 1,045 egocentric-video-grounded tasks covering four daily scenarios, along with a user–agent–tool interactive environment for evaluation. We implement a three-stage synergistic pipeline through which each task is designed to enforce the joint application of visual perception and tool-augmented multi-hop reasoning. We additionally develop a multi-agent simulated user within EgoBench to evaluate agents' interaction capabilities, which generates high-fidelity, task-aligned responses to agents. Furthermore, we establish a deterministic joint validation framework that guarantees objective assessment through process-based and result-based equivalence. Benchmarking eight SOTA video-MLLM agents on EgoBench reveals a severe performance ceiling: the best model achieves only 30.62% accuracy in the best-performing scenario, averaging 19.43% across all four scenarios. Finally, we conduct a multi-dimensional error analysis to disentangle failure modes, exposing capability bottlenecks for advancing future AI agents.
EgoHMP: Achieving Precise Human Motion Prediction in 3D Scenes via Egocentric Cues
Xin Zhao ⋅ Pengzhan Zhou ⋅ Ziyi Li ⋅ Riheng Jia ⋅ Zhida Qin ⋅ Yu Liu
The complex interactive dependencies between human behavior and 3D scenes make accurately predicting human intent a major challenge for human motion prediction (HMP). While recent studies use human gaze to infer intent, gaze signals suffer from ambiguity, drift, and a reliance on eye-tracking hardware. Research in cognitive science indicates that egocentric views inherently encode potential interactions, while head motion serves as a crucial precursor in motion planning. Motivated by this, we propose EgoHMP, a novel framework that uses egocentric views and head motions as robust carriers of interaction intent. By translating this intent into an envisioning of future interactions, EgoHMP achieves precise HMP in 3D scenes. EgoHMP first employs a vision-motion alignment encoder to align the visual and motion feature spaces. Based on these aligned features, an egocentric intent predictor with a modality-aware modulation mechanism adaptively suppresses invalid interaction cues to predict an accurate interaction probability map. Furthermore, to enhance the stability and realism of the generated motions, we propose a contact-aware motion decoder. Specifically, inspired by the human cognitive process of planning trajectories prior to action execution, we first predict a global trajectory to serve as a self-prompt, effectively mitigating the optimization conflict in diffusion models. Subsequently, we introduce a multi-stage denoiser, which relieves motion artifacts by incorporating human contact modeling through a multi-stage optimization strategy. Finally, a GCN-based motion decoder is employed to synthesize physically plausible and semantically consistent motions. Extensive experiments demonstrate that EgoHMP achieves state-of-the-art performance.
EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries
Jiyoun Kim ⋅ Muhan Yeo ⋅ Eunhye Jang ⋅ Jeewon Yang ⋅ Hangyul Yoon ⋅ Su Ji Lee ⋅ Hee J Han ⋅ Hee-Jae Jung ⋅ Doyun Kwon ⋅ Jun y Lee ⋅ Jaehun Lee ⋅ Jung-Oh Lee ⋅ Sunjun Kweon ⋅ Jong-Hak Moon ⋅ Daseul Kim ⋅ Minjae Cho ⋅ Edward Choi
Discharge summaries are crucial clinical documents for patient care, containing clinical context of a patient's overall admission, and are routinely reviewed by medical experts during patient readmission, ongoing care, and diagnostic decision-making. In practice, medical experts often must synthesize information across multiple lengthy discharge summaries by iteratively obtaining information from the documents, while also verifying the evidence supporting each answer. Although large language models (LLMs) are increasingly explored for clinical document question answering, existing benchmarks do not sufficiently capture this clinical setting: they either evaluate exam-style medical knowledge or focus on single-turn question answering with limited evidence-grounding evaluation. We introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over patients' multiple discharge summaries. Built from de-identified MIMIC-IV discharge summaries, EHRNote-ChatQA contains 967 patient-level multi-turn samples spanning one to five notes and 16,072 medical-expert-verified QA pairs (8,036 content questions, each paired with an evidence-grounding question) across eight clinical categories. We construct the benchmark through an expert-informed pipeline that combines a comprehensive discharge-summary structuring schema, expert-curated multi-turn QA templates, and LLM-based generation. Every single QA sample is then reviewed and revised by 11 medical experts over three weeks. Benchmarking 22 LLMs (both open and closed-sourced) reveals several findings; evidence grounding is consistently harder than content answering, multi-turn errors compound across turns, and strong performance on existing single-turn clinical QA benchmarks does not reliably transfer to this setting. These results establish EHRNote-ChatQA as a rigorous and practical benchmark for evaluating clinical QA systems.
ELF: Embedded Language Flows
Keya Hu ⋅ Linlu Qiu ⋅ Yiyang Lu ⋅ Hanhong Zhao ⋅ Tianhong Li ⋅ Yoon Kim ⋅ Jacob Andreas ⋅ Kaiming He
Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and videos. Their success has attracted growing interest in applying them to language modeling. Unlike their image-domain counterparts, today’s leading diffusion language models (DLMs) primarily operate over discrete tokens. In this paper, we show that continuous DLMs can be made effective with minimal adaptation to the discrete domain. We propose Embedded Language Flows (ELF), a class of diffusion models in continuous embedding space based on continuous-time Flow Matching. Unlike existing DLMs, ELF predominantly stays within the continuous embedding space until the final time step, where it maps to discrete tokens using a shared-weight network. This formulation makes it straightforward to adapt established techniques from image-domain diffusion models, e.g., classifier-free guidance (CFG). Experiments show that ELF substantially outperforms leading discrete and continuous DLMs, achieving better generation quality with fewer sampling steps. These results suggest that ELF offers a promising path toward effective continuous DLMs.
Embracing Evolution: A Call for Body-Control Co-Design in Embodied Humanoid Robot
Guiliang Liu ⋅ Bo Yue ⋅ Yi J Kim ⋅ Kui Jia
Humanoid robots, as general-purpose physical agents, must integrate both intelligent control and adaptive morphology to operate effectively in diverse real-world environments. While recent research has focused primarily on optimizing control policies for fixed robot structures, this position paper argues for “evolving both control strategies and humanoid robots' physical structure under a co-design mechanism”. Inspired by biological evolution, this approach enables robots to iteratively adapt both their form and behavior to optimize performance within task-specific and resource-constrained contexts. Despite its promise, co-design in humanoid robotics remains a relatively underexplored domain, raising fundamental questions about its feasibility and necessity in achieving true embodied intelligence. To address these challenges, we propose practical co-design methodologies grounded in strategic exploration, Sim2Real transfer, and meta-policy learning. We further argue for the essential role of co-design by analyzing it from methodological, application-driven, and community-oriented perspectives. Striving to guide and inspire future studies, we present open research questions, spanning from short-term innovations to long-term goals. This work positions co-design as a cornerstone for developing the next generation of intelligent and adaptable humanoid agents.
Emergence of a Shared Canonical Object Frame from In-the-Wild Videos
Tom Fischer ⋅ Martin Sundermeyer ⋅ Adam Kortylewski ⋅ Eddy Ilg
Comparing object orientations and positions across different instances requires their poses to be expressed in a shared canonical frame. Establishing such frames has traditionally required manual annotation, creating a scaling bottleneck that limits category and instance diversity. We show that a shared canonical frame can instead emerge from self-supervised training on object-centric videos captured in the wild, using only noisy camera poses from Structure-from-Motion. Our key idea is to route all training sequences through a shared geometric bottleneck: a coarse canonical mesh that carries no category-specific detail. By learning dense correspondences from image pixels to this mesh, and estimating per-sequence alignments from noisy SfM geometry, a common canonical frame emerges from multi-view consistency and the semantic priors of the feature extractor, without any canonical pose labels or category conditioning. Trained in a self-supervised manner on $160{,}000$ in-the-wild object videos, our method achieves competitive accuracy on category-level pose estimation benchmarks compared to methods that rely on canonical pose supervision.
Foundation models trained on large unlabeled corpora develop emergent capabilities — tasks that go beyond the training objective. We extend this paradigm to multimodal biology with Bonbon, a foundation model for molecular interactions trained on approximately 3.2 trillion tokens of paired protein-ligand sequences. From a single self-supervised objective, Bonbon exhibits four zero-shot emergent capabilities at hierarchical levels of resolution — functional (mechanism of action), residue (binding site localization, including orthosteric/allosteric discrimination), bond (covalent versus non-covalent engagement), and atom (active moiety identification at single-bond resolution). All four capabilities generalize to out-of-distribution proteins. Bonbon was trained on protein-ligand sequences alone, without any task-specific supervision and without structural coordinates. We conducted controlled scaling experiments across four model scales from 42M to 1.6B parameters; every capability evaluated shows a sharp jump at the largest scale on continuous metrics, even as pretraining loss decreases smoothly. This demonstrates that loss-based scaling laws alone are insufficient to predict when zero-shot emergent capabilities become available. The capabilities reported here emerge after training at a 1:2,000 parameter-to-token ratio to accommodate the sparsity of paired biological interaction data. The same frozen representations achieve state-of-the-art binding and affinity predictions at five orders of magnitude greater speed than structure-based methods. In prospective wet-lab validation, Bonbon achieves a 70% hit rate on de novo compounds, discovering novel chemotypes not previously documented as inhibitors for those targets. All of the tested compounds had less than 60% structural similarity to molecules in the training corpus.
EmoPhone: A Multi-Wave Dataset for In-the-Wild Mobile and Wearable Affect Sensing
Panyu Zhang ⋅ minseo Park ⋅ Soowon Kang ⋅ Ismatzoda Tomiris ⋅ Azizbek Mustafakulov ⋅ Otabek Najimov ⋅ Woohyeok Choi ⋅ JUMABEK Alikhanov ⋅ Surjya Ghosh ⋅ Uichin Lee
We introduce a three-wave, in-the-wild multimodal dataset for affect sensing that integrates smartphone sensing, wearable sensing, and dense experience-sampling-method (ESM) labels collected annually from 2020 to 2022. The dataset supports moment-level affect modeling through a shared dimensional label core across all waves, with additional affective descriptors available in the third wave (D-3). We describe the resource in terms of study design, temporal density of in-situ labels, and sensing and label coverage across waves. To support evaluation within this resource, we define an initial three-setting benchmark spanning temporal prediction from within-user history, within-wave cross-user generalization, and cross-wave generalization in which each wave is treated as a separate dataset. Our benchmark results show that the strongest method family depends on the evaluation setting: supervised baselines perform best in the temporal setting, unsupervised domain adaptation is strongest overall in the within-wave cross-user setting, and domain generalization shows the strongest overall cross-wave performance, although its margin over strong baselines is modest. These findings indicate that robust mobile affective computing is constrained not only by label availability but also by substantial participant-level variability and realistic cross-wave differences inherent in longitudinal in-situ deployments.
Empirical Bayes Flow Matching for Continuous Cryo-EM Heterogeneity
Daniel Aibinder ⋅ Azmi Haider ⋅ Dan Rosenbaum
We address the problem of learning generative priors over high-dimensional latent variables from indirect, noisy observations, with a focus on continuous heterogeneity in cryo-electron microscopy (cryo-EM). In this setting, 3D atomic structures $x \in \mathbb{R}^{3N_a}$ are never directly observed and must be inferred from 2D projection images $y$ generated by a complex, non-linear forward operator. We propose Empirical Bayes Flow Matching (EB-FM), a method that learns a flow matching prior $p_\theta(x)$ directly from observations by embedding it within a Monte Carlo stochastic approximation Expectation-Maximization framework. EB-FM alternates between an E-step that performs approximate posterior sampling from $p_\theta(x \mid y)$ via guided flow-based generation while maintaining persistent latent estimates through stochastic approximation, and an M-step that updates the flow model using these stabilized latent variables as training targets. This approach eliminates the need for pretraining on clean data and naturally accommodates non-linear forward models. We validate our method on controlled inverse problems using two-moons and MNIST data, and demonstrate its effectiveness on simulated cryo-EM data with known poses, showing that EB-FM can recover continuous distributions of protein conformations directly in atomic-coordinate space.
Empirical regularities in subjective decision-making by LLMs
Katy Blumer ⋅ Jon Kleinberg ⋅ Ravi Kumar ⋅ Andrew Tomkins
When large language models (LLMs) are used for creative tasks where the evaluation is largely subjective, two crucial activities are ideation (in which a list of candidate ideas is produced) and pruning (in which the initial candidates are prioritized through a process of comparison). In this framework, the subjective results are influenced by inherent biases in the LLM's generation and evaluation, and it is therefore important to understand how these subjective biases operate. We define and explore a set of stylized tasks that allow us to highlight the effects of these subjective biases in LLM evaluation. In the process, we find evidence for several key underlying principles. First, there are multiple subjective LLM biases at play, and to understand the overall evaluation we must analyze the relative strengths of these biases and their interactions. Second, different prompts -- even when they seem superficially similar -- can have the effect of making different subjective LLM biases more or less salient. And third, there are complex interactions between the LLM's initial idea-generation phase and the subsequent decisions that go into prioritizing and pruning these ideas.
Empowering Masked Diffusion Models to Self-Correct with Leave-One-Out Transformers
Sofian Zalouk ⋅ Vincent Counathe ⋅ Paul Jünger ⋅ Daniel Cao ⋅ Daniel Zlotnick ⋅ Kilian Weinberger ⋅ Christopher De Sa
Diffusion models promise iterative refinement, yet masked diffusion language models (MDMs) suffer from two shortcomings: (1) inability to revise tokens, and (2) sparse training signal. The root cause is structural: MDMs control information flow only through input masking, so unmasked positions cannot be supervised, and self-correction is never learned. We propose the Leave-One-Out Transformer (LOOT), which addresses both issues. With LOOT, MDMs predict the distribution over potential replacements conditioned on the rest of the sequence at each position, in a single forward pass. Supervision can then be applied at masked and unmasked positions alike, yielding a lower-variance training objective with a dense loss signal at every token. Our method empowers MDMs to learn self-correction during training and supports continuous refinement at inference time via Gibbs sampling. LOOT achieves a new best validation perplexity on LM1B in a parameter-matched setting. Finetuned from a MDM checkpoint on OpenWebText, LOOT establishes a generative quality-vs-diversity frontier that surpasses prior remasking methods across NFE budgets.
Energy Is All We Need: Beyond FLOPs in Model-Heterogeneous Federated Learning
Sercan Yesilkoy ⋅ Yunseok Kang ⋅ Yoon-Ho Choi
Most energy-efficient federated learning (FL) methods report FLOPs or parameter reductions as evidence of energy savings. However, real GPU energy consumption differs across model architectures in ways that these computational proxies do not capture. In this work, we conduct an empirical study on reducing real GPU energy consumption in model-heterogeneous FL, where clients maintain distinct architectures and coordinate solely through shared prototypes. Through systematic experimentation with hardware-level energy measurements, we propose and validate two complementary mechanisms to make model-heterogeneous FL efficient in terms of real energy. First, Prototype-Aware Structured Pruning (PASP) evaluates each channel by the product of its batch-normalization scale magnitude and its gradient with respect to the prototype alignment loss---rather than the task loss---preserving channels critical for cross-client feature alignment while physically removing the rest. Second, Energy-Aware Training Scheduling (EATS) adaptively allocates local epochs to each client based on epoch-wise prototype alignment quality per unit energy, using smoothed alignment--energy profiles with a knee-point cutoff to prevent energy waste on epochs that yield diminishing alignment gains or amplify local drift. We investigate each mechanism's impact on accuracy and real energy consumption, as well as explore the impact of using both mechanisms simultaneously. All energy figures are obtained from GPU power monitoring rather than computational proxies, providing an empirically grounded evaluation of energy efficiency in heterogeneous federations.
Enhancing LLMs with Cognitive-Affective Personality Inference for Simulating Human Social-Psychological Behavior
Zhibo Deng ⋅ Dongyuan Li ⋅ Shuwen Ge ⋅ Ziqing Zhang ⋅ Ying Zhang ⋅ Renhe Jiang
Large language models are increasingly used to simulate human participants in social and behavioral studies, yet static persona prompting typically maps a participant profile and an experimental scenario directly to a response, entangling stable dispositions with situation-specific interpretations. To address this limitation, we introduce $\textbf{SPIN}$, a cognitive-affective personality system-inspired inference pipeline for social judgment and decision alignment. Specifically, SPIN implements this structured inference process through three zero-shot LLM calls that compile a task-blind participant core, elicit condition-specific cognitive-affective states, and read out decisions from those states, thereby reusing stable personality structure while routing each trial-specific response through an explicit state representation. We evaluate SPIN on two reconstructed social-psychological study families spanning uncertainty reasoning and pluralistic ignorance, across four base LLMs. Compared with blank, demographic, narrative, and chain-of-thought prompt variants, SPIN consistently delivers the strongest overall alignment performance across base LLMs and study families. Ablations and state analyses further show that both personality compilation and structured state elicitation contribute to the gains, and that the elicited states shift interpretably across informational and normative conditions. These results suggest that structured personality-state inference can improve benchmark-level behavioral alignment beyond richer persona descriptions or generic multi-step reasoning.
EnvTrap: Revealing the Environment-Only Attack Surface in Embodied AI via Consequence-Blind Action Execution
Zehao Liu ⋅ Huashuo Lei ⋅ Xi Lin ⋅ Haoang Li ⋅ Yuliang Chen ⋅ Huarui Zhang ⋅ Yifan WANG
As AI models evolve from text and vision to physical agents, embodied AI faces a fundamentally different attack surface. Prior attacks on embodied systems have largely focused on semantic or instruction-level manipulations, such as prompts, adversarial images, and action commands; by contrast, we study environment-only perturbations that require no access to the model or instructions. We reveal a new attack surface: the physical environment itself. An adversary who rearranges objects, without modifying instructions or accessing models, can cause hazardous consequences. In this work, we propose EnvTrap, a diagnostic pipeline that constructs paired safe, trap, and null-trap (benign but misleading) embodied scenarios. We demonstrate that environment-only perturbations raise hazardous-action rates to an average of 86.5\% across multiple vision-language-action (VLA) models, with similar vulnerability patterns confirmed in world models. On a consequence-prediction task, model accuracy remains near chance, while human evaluators succeed easily. We further propose a consequence-aware defense that reduces trap trigger rates by an average of 76.3\% across VLA models in simulation and by 61.7\% on physical robots. This vulnerability arises because current embodied models can recognize scene state but often fail to predict action consequences under altered layouts. Our findings establish environment integrity as a prerequisite for safe embodied AI deployment. Our code and data are available at https://anonymous.4open.science/r/Envtrap-1BA4/
EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting
Junwei Luo ⋅ Shuai Yuan ⋅ Zhenya YANG ⋅ Yansheng Li ⋅ Zhe Liu ⋅ Hengshuang Zhao
Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological conditions. In this paper, we view this task as a partially observed, weather-driven world modeling problem, in which weather acts as a conditioning signal, while forecasting remains uncertain due to sparse observations and unobserved land-surface states. However, existing methods do not fully capture this setting: deterministic models collapse uncertainty into a single future prediction, while diffusion-based methods typically treat weather variables as undifferentiated conditioning signals, and existing benchmarks focus mainly on reconstruction accuracy rather than whether forecasts respond correctly to changed weather forcing. We introduce EO-WM, a video diffusion transformer for multispectral EO forecasting. EO-WM incorporates a physically informed conditioning framework that represents meteorological forcing through a climatological baseline, weather anomalies, and cumulative physical stress signals. Specifically, it separates baseline and anomaly through distinct conditioning pathways, and accumulates anomalous forcing over time to capture sustained heat and drought stress. To evaluate weather-response behavior beyond standard metrics, we introduce two diagnostic benchmarks: an Extreme Summer Benchmark for severity-aware prediction of vegetation degradation under extreme weather, and a Seasonal Matched-Pair Benchmark for testing response fidelity under changed weather forcing. Experiments show that EO-WM reduces the error in predicted Normalized Difference Vegetation Index (NDVI) decline amplitude by a relative 5.63\% and improves directional hit rate by a relative 7.80\%, while remaining competitive on standard pixel-level metrics.
Epiplexity Guided Data Selection and Generation for Out-of-Distribution Generalization
Ellen Su ⋅ Andres Potapczynski ⋅ Shikai Qiu ⋅ Edward Hughes ⋅ Andrew Wilson
Modern systems are often expected to transfer across tasks that were not specified during training, leading to the question: what data facilitates generalization in new, unanticipated settings? One hypothesis is that data with more structural information could contain shared circuits and subprograms that could be recycled in a wider array of downstream settings. Epiplexity, a recently proposed measure of the structural information a compute-bounded learner can extract from data, provides a mechanism to reason about this relationship. In this paper, we show how to operationalize epiplexity as an online training signal for data selection and generation. For selection, we fit scaling laws to the training loss curves of natural data domains to predict the expected epiplexity gain as a function of training tokens, and use this signal to adaptively determine the sampling weights over domains during training. For synthetic data generation, we define a generator's reward as the change in learner epiplexity over a buffer of previously generated data and use REINFORCE policy gradients to guide the generator toward an epiplexity-maximizing distribution. In both cases, higher epiplexity predicts improved downstream performance on zero-shot and fine-tuning based tasks, supporting the hypothesis that data rich in structural information yields representations that transfer across domains.
EpistasisBench: Revealing Structural Limitations of Zero-Shot Protein Language Models
Junan Chen ⋅ jianwen Min ⋅ Liang He ⋅ Yiheng Zhu
Zero-shot mutation effect prediction is widely used to evaluate protein language models (PLMs). However, because existing benchmarks mix additive and non-additive mutations, and non-additive effects (epistasis) are sparse yet critical, it remains unclear whether these models truly capture thermodynamic residue coupling. As a result, strong overall performance may mask weaknesses in modeling multi-residue interactions. To make this distinction explicit, we introduce EpistasisBench, a thermodynamically stratified zero-shot benchmark constructed from large-scale stability measurements. By isolating significant and sign epistatic double mutations, it separates additive effects from thermodynamic coupling and defines two complementary tasks: ∆∆G prediction restricted to epistatic subsets, and a direct ∆∆Gint scoring task that examines non-additive structure in model probability space. When evaluated across more than 17 state-of-the-art PLMs, EpistasisBench reveals a consistent pattern. Performance declines under strict epistasis filtering, strong single-site predictors fail on coupled mutations, and standard masked language model scoring shows no variation in ∆∆Gint, thereby preventing detection of non-additive effects. Together, these findings indicate that the limitation lies in common zero-shot inference schemes rather than model capacity. To demonstrate that non-additive signal can in fact be recovered under a more suitable inference formulation, we introduce a simple joint marginal inference mechanism that restores non-additive signal and improves performance without fine-tuning. Overall, EpistasisBench provides a physically grounded test for multi-residue interactions and highlights a key gap in current zero-shot evaluation of protein language models. The dataset is available at https://www.openml.org/d/47245
Equivariant Reinforcement Learning for Clifford Quantum Circuit Synthesis
Richie Yeung ⋅ Aleks Kissinger ⋅ Rob Cornish
We consider the problem of synthesizing Clifford quantum circuits for devices with all-to-all qubit connectivity. We approach this task as a reinforcement learning problem in which an agent learns to discover a sequence of elementary Clifford gates that reduces a given symplectic matrix representation of a Clifford circuit to the identity. This formulation permits a simple learning curriculum based on random walks from the identity. We introduce a novel neural network architecture that is equivariant to qubit relabelings of the symplectic matrix representation, and which is size-agnostic, allowing a single learned policy to be applied across different qubit counts without circuit splicing or network reparameterization. On six-qubit Clifford circuits, the largest regime for which optimal references are available, our agent finds circuits within one two-qubit gate of optimality in milliseconds per instance, and finds optimal circuits in 99.2\% of instances within seconds per instance. After continued training on ten-qubit instances, the agent scales to unseen Clifford tableaus with up to thirty qubits, including targets generated from circuits with over a thousand Clifford gates, where it achieves lower average two-qubit gate counts than Qiskit's Aaronson-Gottesman and greedy Clifford synthesizers.
ERIS: Enhancing Privacy and Scalability in Federated Learning via Federated Shard Aggregation
Dario Fenoglio ⋅ Pasquale Polverino ⋅ Jacopo Quizi ⋅ Martin Gjoreski ⋅ Akash Dhasade ⋅ Marc Langheinrich
Scaling Federated Learning (FL) to billion-parameter models forces a challenging trade-off between privacy, scalability, and model utility. Existing solutions often tackle these challenges in isolation, sacrificing accuracy, relying on costly cryptographic tools, or introducing communication and optimization inefficiencies that affect convergence. We introduce ERIS, an FL framework centered on Federated Shard Aggregation (FSA), a novel mechanism that partitions each client update into non-overlapping shards whose aggregation is distributed across multiple client-side aggregators. FSA removes the central aggregation bottleneck, limits the information visible to any single observer, and preserves the centralized FL update after reassembly. ERIS can further readily integrate Distributed Shifted Compression (DSC) to reduce transmitted payloads and exposed coordinates. We prove that ERIS preserves convergence under standard assumptions and bounds mutual information leakage by the observable fraction of each update, decreasing with the number of client-side aggregators, and with the compression level when DSC is enabled. Experiments across image and text tasks, including large language models, show that ERIS achieves FedAvg-level utility while substantially reducing communication bottlenecks and improving robustness to membership inference and reconstruction attacks, without relying on heavy cryptography or utility-degrading perturbations.
Escaping Path Mirages in Offline Goal-Conditioned Reinforcement Learning
Seungyul Han ⋅ Junhyeon Bae ⋅ Jaebak Hwang ⋅ Gwanwoo Choi ⋅ Minung Kim
Offline goal-conditioned reinforcement learning remains challenging in stochastic long-horizon settings, where compounding value estimation errors hinder reliable goal reaching. While prior methods have sought to address this challenge through various approaches, a key challenge remains path mirage, where agents overcommit to spuriously successful trajectories in offline datasets induced by stochasticity, often leading to failure on difficult tasks. To address this issue, we propose Branch-aware Graph Planning (BGP), which captures stochastic branching structures by identifying high-variability points as graph nodes and designing goal-conditioned edges with segmented learning to avoid unreliable branches and reduce unnecessary stochasticity along paths. As a result, BGP enables reliable goal reaching even in highly stochastic long-horizon environments, and experiments on diverse OGBench tasks show that it substantially outperforms prior state-of-the-art offline goal-conditioned RL methods.
EVA-0: Test-Time Model Evolution with Only Two Forward Passes per Sample
Guohao Chen ⋅ Shuaicheng Niu ⋅ Geng Li ⋅ Yunbei Zhang ⋅ Shilin Shan ⋅ Chunyan Miao ⋅ Jianfei Yang
Test-time model evolution offers a promising way for deployed models to improve from unlabeled test-time experience, yet most existing methods depend on backpropagation (BP), which incurs substantial memory overhead and makes them difficult to deploy on edge devices, quantized models, specialized accelerators, or black-box models. In this work, we study test-time model evolution under a strict two-forward budget, a setting that pushes adaptation toward highly efficient real-world deployment. We reveal three key obstacles in zeroth-order test-time optimization: susceptibility to shortcut solutions, uncontrolled weight drift, and ineffective update direction estimation. To overcome them, we propose EVA-0, a minimal zeroth-order adaptation framework that: 1) keeps the loss scale-invariant to prevent shortcut solutions; 2) devises an anchor-guided optimization strategy to alleviate weight drift; 3) uses sample-wise symmetric two-sided perturbation for update direction estimation and inference. EVA-0 requires no BP and performs both inference and adaptation within only two forward passes per sample. Results on ImageNet-C\&ViT-Base show that EVA-0 outperforms both BP-based DeYO and BP-free FOA, while achieving a 14$\times$ speed-up over FOA.
Evaluating AI-based Scientific Knowledge Synthesis with Epidemiological Systematic Reviews
Shreyansh Padarha ⋅ Ryan Othniel Kearns ⋅ Tristan M Naidoo ⋅ Lingyi Yang ⋅ Łukasz Borchmann ⋅ Piotr Blaszczyk ⋅ Christian Morgenstern ⋅ Ruth McCabe ⋅ Sangeeta Bhatia ⋅ Philip Torr ⋅ Jakob Foerster ⋅ Scott Hale ⋅ Thomas Rawson ⋅ Anne Cori ⋅ Elizaveta Semenova ⋅ Adam Mahdi
Systematic literature reviews (SLRs) are a demanding and high-stakes form of scientific knowledge synthesis that remains underspecified as an evaluation setting for large language models (LLMs). We introduce AgentSLR, a large-scale evaluation harness comprising an SLR automation workflow and an expert annotated dataset covering 16,248 articles, designed to test LLM capabilities across the stages of SLRs in epidemiology. Reference annotations were derived from peer-reviewed studies on WHO priority pathogens and produced by domain experts. AgentSLR evaluates each review stage as a separate unit with dedicated metrics enabling targeted failure analysis. We evaluated five frontier reasoning models and found that no single model dominated across all tasks, showing sub-task specialisation often hidden by aggregate benchmarks. Structured data extraction is a major bottleneck, with no model exceeding an average field-level F1 of 0.67. Estimated costs vary substantially, by up to 96 times across evaluated models. Documented failure modes suggest that the evaluated models are not yet reliable enough for unsupervised deployment in epidemiology, where findings can inform public policy.
EventLens: Event-Structure Reinforcement Learning for Video Understanding
Guowen Zhang ⋅ Boshen Xu ⋅ Zihao Yue ⋅ Ziheng Wang ⋅ Xiaokun Liu ⋅ Xin Tao ⋅ Wenyu Qin ⋅ Pengfei Wan ⋅ Qin Jin
Recent multimodal large language models (MLLMs) achieve strong results on video question answering, yet often fail to recover the temporal structure that makes a video coherent. They may merge neighboring events, miss single-event progression, or infer implausible relations across events. We study structured temporal understanding as the ability to recover event structure from visual evidence, including event boundaries, single-event progression, and multi-event relations. We argue that this gap arises from a mismatch between current post-training objectives and temporal reasoning, where caption-style supervision allows models to rely on language shortcuts rather than visual temporal grounding. To address this, we propose $\textbf{EventLens}$, a vision-centric reinforcement learning framework that learns video understanding through event-structure recovery}. Instead of reconstructing captions, EventLens trains models to recover temporal structure under controlled perturbations, such as altered temporal granularity, reversed progression, and shuffled event order. We instantiate this framework with a three-level temporal ontology and derive three verifiable task families: event segmentation, progression discrimination, and multi-event relation reasoning. These tasks admit deterministic rewards, enabling task-specialized GRPO training without human preference labels or learned reward models. We further introduce multi-teacher on-policy distillation (MT-OPD) to consolidate specialized policies into a unified model. Experiments on Qwen3-VL backbones show consistent improvements on temporal grounding and event-sensitive reasoning benchmarks, while preserving general video-QA performance. Learning curves further demonstrate that EventLens achieves comparable downstream transfer with significantly reduced training cost compared to direct mixed RL optimization.
Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
Tao Lin ⋅ Yuxin Du ⋅ Jiting Liu ⋅ Nuobei Zhu ⋅ YUNHE LI ⋅ Yuqian Fu ⋅ Yinxinyu Chen ⋅ Hongyi Cai ⋅ Ye Zewei ⋅ Bing Cheng ⋅ Kai Ye ⋅ Yiran Mao ⋅ Yilei Zhong ⋅ Mingkang Dong ⋅ Junchi Yan ⋅ Gen Li ⋅ Bo Zhao
Vision-Language-Action (VLA) models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often struggle in scenarios requiring precise spatial understanding, as current VLA models primarily rely on 2D visual representations that lack depth information and detailed spatial relationships. While recent approaches incorporate explicit 3D inputs such as depth maps or point clouds to address this issue, they often increase system complexity, require additional sensors, and remain vulnerable to sensing noise and reconstruction errors. Another line of work explores implicit 3D-aware spatial modeling directly from RGB observations without extra sensors, but it often relies on large geometry foundation models, resulting in higher training and deployment costs. To address these challenges, we propose \textbf{Evo-Depth}, a lightweight depth-enhanced VLA framework that enhances spatially grounded manipulation without relying on additional sensing hardware or compromising deployment efficiency. Evo-Depth employs a lightweight \textit{Implicit Depth Encoding Module} (IDEM) to extract compact depth features from multi-view RGB images. These features are incorporated into vision-language representations through a \textit{Spatial Enhancement Module} (SEM) via depth-aware modulation, enabling efficient spatial-semantic enhancement. A \textit{Progressive Alignment Training} strategy is further introduced to align the resulting depth-enhanced representations with downstream action learning. Extensive experiments in simulation and real-world settings demonstrate the effectiveness of Evo-Depth. With only \textbf{0.9B} parameters, Evo-Depth achieves state-of-the-art-level performance across four simulation benchmarks. In real-world experiments, Evo-Depth attains the highest average success rate while also exhibiting the smallest model size, lowest GPU memory usage, and highest inference frequency among compared methods. These results demonstrate that lightweight implicit depth enhancement is an effective and practical solution for spatially grounded robotic manipulation. We release code and models to facilitate future research on lightweight depth-enhanced VLA models.
EvoInspect: A Unified Self-Evolving Multi-Agent Framework for Industrial Hardware Inspection
Zhuyun Yuan ⋅ Jiming Zhang ⋅ Haoyuan Sun ⋅ Jun Yin ⋅ Qing Wang ⋅ Jian Kang ⋅ Jie Li ⋅ Yongxiang Li
Human experts can accumulate experience from open-ended and complex inspection tasks, thereby continually improving their inspection capabilities. Actually, industrial hardware inspection faces challenges including open-set defect categories, complex defects (tiny, high-gloss, low-light, or blurred), high inference cost, and difficulty in accumulating experience post-deployment. Existing end-to-end inspection methods typically couple image restoration with defect recognition in a single forward pass, limiting generalization to unseen components or composite defects; cascade “restore-then-detect” approaches often rely on static preprocessing, hindering adaptive adjustment based on scenario, defect morphology, and historical failures. To address this, we propose EvoInspect, the first memory-enhanced multi-modal multi-agent framework for industrial defect object detection. It leverages a multi-modal agentic memory mechanism to distill self-evolving scenario-specific Defect Detail Restoration decisions and dynamically grounded trajectories, enabling “accumulating experience like human experts.” Specifically, (1) we introduce a dynamic Defect Detail Restoration strategy, including adaptive slicing for tiny defects, highlight suppression, low-light enhancement, super-resolution, and binarization, combined with a coarse-to-fine token-efficient grounding mechanism for efficient defect capture; (2) we design a dual-channel distillation mechanism unifying restoration and detection (perception and execution) experience, incorporating both successful and failed validation trajectories to support system self-evolution. Experiments on bearing surfaces, PCBs, solar electroluminescence (solar EL), and magnetic tile datasets show that EvoInspect consistently outperforms strong baselines in defect localization recall, inference accuracy, while demonstrating cross-scenario adaptability and offline continual improvement.
Evolutionary System Prompt Learning for Reinforcement Learning in LLMs
Lunjun Zhang ⋅ Ryan Chen ⋅ Bradly Stadie
Building agentic systems that can autonomously self-improve from experience is a longstanding goal of AI. Large language models (LLMs) today primarily self-improve via two mechanisms: self-reflection for context updates, and reinforcement learning (RL) for weight updates. In this work, we propose Evolutionary System Prompt Learning (E-SPL), a method for jointly improving model contexts and model weights. In each RL iteration, E-SPL samples trajectories under multiple system prompts in parallel, then jointly applies RL updates to weights and evolutionary updates to system prompts via LLM self-reflection. E-SPL encourages a natural division between declarative knowledge encoded in prompts and procedural knowledge encoded in weights. Most notably, in an easy-to-hard generalization setting (AIME $\rightarrow$ BeyondAIME), E-SPL improves RL success rate from 38.8% $\rightarrow$ 45.1%. E-SPL also improves RL on AIME 2025 (56.3% $\rightarrow$ 60.6%), HMMT 2025 (50.0% $\rightarrow$ 52.7%), and agentic search (44.2% $\rightarrow$ 48.6% on gpt-oss-120b). Across all settings, E-SPL outperforms both RL-only and evolution-only baselines, demonstrating that weight updates and context updates are deeply synergistic and can together yield gains in generalization that neither achieves alone.
EvoOptiGraph: Weakness-Driven Coevolution via Graph-Based Structural Generation for Optimization Modeling
Qingcan Kang ⋅ Mingyang LIU ⋅ Xiaojin Fu ⋅ Shixiong Kai ⋅ Tao Zhong ⋅ Mingxuan Yuan
Automating optimization modeling from natural language faces two key challenges: training corpora lack structural diversity, and data generation pipelines remain static and decoupled from model learning. To address these challenges, we propose EvoOptiGraph, a novel framework where data and model co-evolve, driven by model weaknesses. EvoOptiGraph represents each mixed-integer linear program (MILP) as an attributed bipartite graph and applies validity-preserving evolutionary operators to generate structurally diverse instances. The evolved graphs are converted into solver code and natural language via deterministic compilation and verified back translation. Training proceeds in two stages: supervised fine-tuning (SFT) on an initial dataset, followed by reinforcement learning with verifiable rewards (RLVR), where graph-derived weakness signals dynamically evolve new instances that target the model's failures, forming a closed loop that continuously updates the training distribution. Empirical results on six public datasets show that EvoOptiGraph significantly outperforms larger generalist models, agentic methods, and specialized baselines in accuracy, executability, and generalization. These results demonstrate that targeted data–model coevolution is an effective strategy for improving LLMs on optimization modeling tasks.
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
Yen-Shan Chen ⋅ Zhi Rui Tam ⋅ Cheng-Kuang Wu ⋅ Yun-Nung (Vivian) Chen
Current evaluations of LLM safety predominantly rely on *severity-based taxonomies* to assess the harmfulness of models' responses to malicious queries. We argue that this formulation requires re-examination as it implicitly equates model output with realized harm while neglecting *Execution Likelihood*—the conditional probability of a threat being realized by a real-world actor given a model's response. In this work, we introduce **Expected Harm**, a metric that weights the severity of a jailbreak by its execution likelihood, modeled as a function of *execution cost*. Through empirical analysis of state-of-the-art models, we reveal a systematic *Inverse Risk Calibration*: models disproportionately exhibit stronger refusal behaviors for low-likelihood (high-cost) threats while remaining vulnerable to high-likelihood (low-cost) queries. Mechanistically, we first provide an analysis using linear probing, revealing that while refusal mechanisms linearly separate queries by severity, they do not use execution likelihood as an indicator for refusal. Then, we use Olmo-3's open-source staged checkpoints to trace the root cause of miscalibration to the Supervised Fine Tuning (SFT) training data distribution (82.8\% of safety examples at cost levels 0–1) and show that cost representations are weak from the SFT stage onward. Finally, we demonstrate that this miscalibration creates a vulnerability: by exploiting this property, we increase the attack success rate of existing jailbreaks by up to $2\times$.
Exploring Lifelong Adaptation: In-Context Reinforcement Learning in Non-Stationary Environments
Ye Wang ⋅ Kaiqian Cui ⋅ Xinrun Xu ⋅ Tao Zhang ⋅ Qin Jin
Standard In-Context Reinforcement Learning (ICRL) typically assumes stationary dynamics within an interaction history, a restriction that limits its real-world applicability. We formalize Lifelong ICRL, a more realistic paradigm in which agents must continuously adapt to diverse non-stationary dynamics, ranging from abrupt random shifts to structured temporal drifts within a single lifetime. To probe the adaptation limits of this setting, we conduct a large-scale empirical study across non-stationary environments that span discrete symbolic reasoning and continuous physics-based control. We systematically investigate three core dimensions governing adaptation: (i) Training Regimes, comparing the generalization boundaries of domain randomization, stationary training, non-stationary training, and their staged combinatorial strategies; (ii) Model Architectures, evaluating modern sequence models, including Linear Attention and hybrid designs; and (iii) Impact of Scale, analyzing the relative contributions of interaction scale and model capacity. Our results reveal how these factors jointly shape robust in-context adaptation under unknown dynamics, yielding actionable insights for the design of generalist agents.
Extending Myerson's Optimal Auctions to Correlated Bidders via Neural Network Interpolation
Mingyu Guo ⋅ Jiayuan Liu ⋅ Vincent Conitzer
We aim to design revenue-maximizing single-item auctions that are deterministic, strategy-proof, and ex post individually rational --- the quintessential fundamental model for optimal mechanism design. While Myerson's seminal work solved this model for independent bidders, straightforward extensions to correlated settings often result in non-monotonic, and therefore non-strategy-proof allocations. Critically, optimal allocation design for correlated bidders is NP-hard to approximate, rendering theoretical optimality guarantees in this domain computationally intractable. We establish a new standard of empirical rigor for this theoretically hard model by proposing an empirical pipeline. We synthesize neural interpolation of Myerson's greedy allocation based on marginal profits (the correlated variant of virtual valuation), formal neural verification to enforce exact strategy-proofness, and a constructive repair procedure. Empirically, our method is highly effective, consistently achieves near-optimal revenue across a wide range of distributions, including synthetic, adversarial, and real-world (consumer, financial, and industrial) datasets. Compared to existing manual and neural baselines, our approach shows substantial improvement, often reducing the revenue gap by an order of magnitude. Our study is also the first to introduce (unattainable) greedy revenue as a rigorous upper bound for empirical benchmarking, providing a definitive quantitative performance measure where true optima remain theoretically elusive. We demonstrate the generality of our approach by extending it to multi-unit auctions with unit demand and integrating our verification techniques into RegretNet to achieve exact strategy-proofness.
Extending Pretrained 10-Second ECG Foundation Models to Longer Horizons
Wei Tang ⋅ Jinpei Han ⋅ Kangning Cui ⋅ Mattia Carletti ⋅ Fredrik K. Gustafsson ⋅ Shreyank Gowda ⋅ Patitapaban Palo ⋅ Anshul Thakur ⋅ Lei Clifton ⋅ Jean-michel Morel ⋅ Raymond H Chan ⋅ David Clifton ⋅ Xiao Gu
Electrocardiogram (ECG) foundation models pretrained on typical diagnostic 10-second ECG segments, have demonstrated strong transferability across a range of clinical applications. However, many real-world applications produce recordings that are typically longer, and are varied in duration during inference time. These 10-second models have no built-in way to combine information across time. Extending them to longer horizons introduces two challenges: structural incompatibilities arising from input-length disparities, and semantic challenges that limit meaningful temporal aggregation. We propose a parameter-efficient framework that extends pretrained ECG foundation models to longer and variable-length ECGs without retraining the backbone. Guided by a frozen pretrained 10-second model, we introduce a lightweight plug-in module that extends the model in two complementary ways: (i) structurally compatible long-sequence processing and (ii) semantically informed temporal modeling. Experiments on multiple long-horizon ECG tasks, datasets, and foundation model backbones demonstrate that our method enables robust long-horizon extension from pretrained snapshot models, consistently outperforming sliding-window and pooling-based baselines with strong parameter efficiency.
Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning
Sixing Chen ⋅ Ji-An Li ⋅ Saner Cakir ⋅ Sinan Akcali ⋅ Kayla Lee ⋅ Marcelo G Mattar
Large language models (LLMs), especially reasoning models, generate extended chain-of-thought (CoT) reasoning that often contains explicit deliberation over future outcomes. Yet whether this deliberation constitutes genuine planning, how it is structured, and what aspects of it drive performance remain poorly understood. In this work, we introduce a new method to characterize LLM planning by extracting and quantifying search trees from reasoning traces during four-in-a-row gameplay. By fitting computational cognitive models on the extracted search trees, we characterize how plans are structured and how they influence move decisions. We find that LLMs' search is shallower than humans', and that performance is predicted by search breadth rather than depth. Most strikingly, although LLMs expand deep nodes in their traces, their move choices are best explained by a myopic model that ignores those nodes entirely. A causal intervention study in which we selectively prune CoT paragraphs further suggests that move selection is driven predominantly by shallow rather than deep search. These patterns contrast with human planning, where performance is driven primarily by deep search. Together, our findings reveal a key difference between LLM and human planning: while human expertise is driven by deeper search, LLMs do not act on deep lookahead. This dissociation offers targeted guidance for aligning LLM and human planning. More broadly, our framework provides a generalizable approach for interpreting the structure of LLM planning across strategic domains.
ExtraVAR: Stage-Aware RoPE Remapping for Resolution Extrapolation in Visual Autoregressive Models
Feihong Yan ⋅ Shaoyu Liu ⋅ Haixuan Wang ⋅ Shuai Lu ⋅ Weijie Ma ⋅ Linfeng Zhang ⋅ Huiqi Li ⋅ Xiangyang Ji
Visual Autoregressive (VAR) models have emerged as a strong alternative to diffusion for image synthesis, yet their fixed training resolution prevents direct generation at higher resolutions. Naively transferring training-free extrapolation methods from LLMs or diffusion models to VAR yields three characteristic failure modes: global repetition, local repetition, and detail degradation. We trace them to a unified band-stage mismatch: VAR generates images in a coarse-to-fine, scale-wise process where each stage is driven by a distinct dominant RoPE frequency band, and each failure mode emerges when the dominant band of a particular stage is disrupted. Building on this insight, we propose Stage-Aware RoPE Remapping, a training-free strategy that assigns each frequency band a stage-specific remapping rule, jointly suppressing all three failure modes. We further observe that attention becomes systematically dispersed as the image resolution increases. Existing methods typically depend on predefined attention scaling factors, which are neither adaptive to the target resolution nor capable of faithfully capturing the actual extent of attention dispersion. We therefore propose Entropy-Driven Adaptive Attention Calibration, which quantifies dispersion via a resolution-invariant normalized entropy and yields a closed-form per-head scaling factor that realigns the extrapolated-resolution attention entropy with its training-resolution counterpart. Extensive experiments show that our method consistently outperforms prior resolution-extrapolation methods in both structural coherence and fine-detail fidelity. Our codes have been released in supplementary materials and will be released on Github.
FacEDiT: Talking Head Video Editing via Facial Motion Infilling
Sung-Bin Kim ⋅ Joohyun Chang ⋅ David Harwath ⋅ Tae-Hyun Oh
Editing a local segment in a talking head video without reshooting the entire scene remains challenging and underexplored. Given local edits, such as insertion, deletion, or substitution, talking head video editing must synthesize a replacement segment while leaving the rest unchanged. This requires variable-duration local rewriting while preserving identity, unedited regions, and boundary continuity, which standard speech-driven generation and lip synchronization do not directly address. Moreover, direct supervision is infeasible, as it requires paired videos of the same person and scene differing only in a local spoken segment, which does not exist in the real world. We instead formulate talking head video editing as facial motion infilling, a self-supervised pretext task that recovers masked facial motion from speech and surrounding motion context in ordinary video--speech pairs. The key insight is that local video editing can be simulated during training by masking a motion span and reconstructing it from speech and visible motion context. Based on this formulation, we introduce FacEDiT, a mask-controlled talking head model with local temporal attention bias and temporal smoothness regularization for improved lip--speech alignment and transition continuity. We also introduce FacEDiTBench, the first benchmark for talking head video editing, covering diverse edit types and lengths with dedicated evaluation metrics. Extensive experiments show that FacEDiT produces accurate, speech-aligned edits with strong identity preservation and seamless boundary transitions. Beyond editing, the same facial motion infilling model extends to portrait animation and lip synchronization by simply changing the mask pattern, establishing FacEDiT as a unified framework for talking head video editing and generation. \textit{We will release the code and data upon acceptance.
FACETS: Cross-Granularity Vision--Language Modeling for 3D Anomaly Detection
Yuchuan Li ⋅ Jae-Mo Kang ⋅ Il-Min Kim
Three-dimensional (3D) anomaly detection underpins quality assurance in advanced manufacturing and engineering, capturing subtle defects and deviations that 2D inspection struggles to resolve due to occlusion, viewpoint, and appearance confounds. Existing 3D anomaly detection methods either rely on reconstruction errors or stored normal representations to detect anomalies. While both paradigms have driven substantial progress, they share a fundamental limitation: both operate entirely in the visual domain, missing the rich semantics encoded in natural language, which also typically constrains them to per-category models. We propose FACETS, among the first frameworks that leverage a 3D vision--language model for 3D anomaly detection, opening a new paradigm beyond reconstruction- and memory-bank-based methods. The key idea is to explicitly retain native point-level features that enable reasoning at both patch and point granularities, and further leverage language grounding and cross-granularity geometric modeling along two complementary axes, linguistic semantics and geometric saliency, so that coarse semantic cues and fine-grained geometric details jointly support anomaly detection. FACETS enables unified multi-category anomaly detection, avoiding the per-category models used by most prior methods. Notably, we provide a mathematical analysis of our loss function, offering valuable insights into the substantial improvement FACETS achieves in anomaly localization. Extensive experiments on popular benchmarks reveal that FACETS establishes new SOTA performance across datasets and metrics for 3D anomaly detection, substantially and consistently outperforming existing methods. Code is provided as supplementary material.
Fair Range k-Supplier Clustering in Offline and Streaming Models
Meiyun Lu ⋅ Wei Yue ⋅ Weihong Wu ⋅ Lin Zeyu ⋅ Longkun Guo
Fairness has emerged as a central consideration in machine learning, motivating the study of fair range $k$-supplier clustering as a fundamental problem that focuses fairness on the selected centers. Given a set of suppliers and clients, where each supplier may belong to one or more demographic or functional groups, the objective is to select $k$ suppliers that minimize the maximum client-to-center distance while ensuring that the number of selected suppliers from each group satisfies prescribed lower and upper bounds. For disjoint supplier groups, we develop a polynomial-time $3$-approximation algorithm in the offline setting and a $(3+\epsilon)$-approximation algorithm in the streaming setting. We further consider overlapping supplier groups and show that this generalization admits a parameterized $3$-approximation algorithm whose runtime is exponential in the number of clusters. Finally, experiments on both synthetic and real-world datasets demonstrate that our algorithms achieve significantly better clustering quality and runtime efficiency in comparison with state-of-the-art methods.
Fast and Stable Gradient Approximation for Bilinear Forms of Hermitian Matrix Functions
Navjot Singh ⋅ Kipton Barros ⋅ Sherry Li
Objectives involving bilinear forms (u^\top f(A(\theta))v) for Hermitian (A) arise widely in scientific computing and probabilistic machine learning. For large matrices, Lanczos efficiently approximates these quantities, but differentiating them with respect to (\theta) is challenging. Existing approaches either backpropagate through the Lanczos recurrence, requiring reorthogonalization for stability, or apply Arnoldi to an augmented block matrix of twice the original size. Both introduce extra computation and orthogonalization costs that can limit performance on modern hardware. We propose a forward-only gradient approximation that reuses the Lanczos pass and adds very minimal overhead in most cases. We prove that its error is proportional to the Lanczos residual norm, the same quantity controlling the forward approximation. Whereas a traditional adjoint-based calculation would be unstable without reorthogonalization, the new method appears unconditionally stable in our tests. It is also faster than existing state-of-the-art approaches.
FEAD: Fine-Grained Epipolar Attention Diffusion for Large-Disparity Light Field Spatial Super-Resolution
Wenbin Wang ⋅ Youfang Lin ⋅ Chen Gao ⋅ Shuo Zhang
Light-field (LF) spatial super-resolution hinges on exploiting cross-view correspondences to recover high-frequency details while preserving geometric structure. However, existing methods often rely on implicit feature interaction or coarse disparity maps, which suffers noticeable performance degradation in regions with large disparities. In this paper, we propose Fine-Grained Epipolar Attention Diffusion (FEAD), which adopts epipolar attention as an explicit mechanism to characterize cross-view geometric correspondence along epipolar lines. To enable more accurate feature alignment and long-range spatial–angular interaction, we introduce a diffusion-based epipolar attention refinement strategy that progressively improves the initial attention. With the refined attention as guidance, cross-view features are explicitly warped and fused to enforce geometric consistency and aggregate complementary details across views for fine-grained reconstruction. Extensive experiments demonstrate that FEAD achieves state-of-the-art performance, with particularly strong gains in large-disparity scenarios.
FedCAG: Federated Causality-Aware Graph Learning for Multi-Cloud Workload Forecasting
Yongcan Luo ⋅ Zhengjie Yang ⋅ Jiahao Zheng ⋅ Hao Wang ⋅ Wei Bao ⋅ Dapeng Wu
Recent large-scale outages at major cloud providers such as AWS and GCP have exposed the fragility of relying on a single cloud. To improve resilience, fault tolerance, and business continuity, many enterprises are moving their services to multi-cloud environments. However, multi-cloud deployment also makes workload forecasting substantially harder. Existing multi-cloud workload forecasting methods are typically designed for either centralized or isolated environments, leading to risks of private data leakage or limited generalization. To address these challenges, we propose \textbf{FedCAG}, a \textbf{Fed}erated \textbf{C}ausality-\textbf{A}ware \textbf{G}raph learning paradigm for multi-cloud workload forecasting. Instead of treating federated learning as parameter averaging over local predictors, FedCAG jointly federates predictive models and graph-structured dependency priors. Each client constructs causality-aware, spatial, and temporal graphs from its private telemetry to model directed inter-metric influence, metric-level interactions, and intra-window temporal dynamics. To address strong cross-cloud heterogeneity, FedCAG further combines causality-aware representation fusion, adaptive graph refinement, and client-specific personalization, enabling the global model to benefit from shared workload structures while preserving local specificity. Experiments on real-world datasets show that FedCAG consistently outperforms mainstream federated baselines and even several centralized methods trained on the full dataset, delivering stable and accurate forecasting across heterogeneous and privacy-constrained deployments. The source code is available at: \url{https://anonymous.4open.science/r/FedCAG-FB9C}.
Federated Graph Learning with Local Message Compensation
Ye Zhou ⋅ Bangqi Li ⋅ Haodi Wang ⋅ Yu Guo ⋅ Libin Jiao ⋅ Rongfang Bie
Federated Graph Learning (FGL) aims to collaboratively train Graph Neural Networks (GNNs) across distributed subgraphs with a primary challenge of missing cross-client connections. Existing approaches predominantly rely on a retrieval-based paradigm, which reconstructs missing context by accessing external information across clients. However, such cross-client transmissions inevitably expand the attack surface for privacy inference and incur significant communication overhead. In this work, we propose \textbf{FedLMC}, a novel framework that eliminates the necessity of any node information sharing and prior global adjacency knowledge. FedLMC exploits a dual-stage message compensation mechanism to generate informative and expressive node representations. For the semantic information loss of external neighbors, we propose Local Semantic Compensation, which exploits semantic expansion via semantic surrogates to compensate for missing semantic context. For the structural information loss induced by semantic expansion, we propose Community Structural Compensation, which exploits structural anchoring via community anchors to compensate for underlying structural context. Extensive experiments on twelve datasets (both homophilic and heterophilic) demonstrate that FedLMC establishes a new state-of-the-art performance with superior generalization and robustness to highly fragmented graph data.
Federating a LGN is challenging because each gate is parameterised by a distribution over 16 Boolean operations, so averaging parameters across clients does not correspond to averaging the underlying functions. Under non-IID data, such parameter-space merging collapses to near-chance accuracy. Boolean Feature Selection (BFS) avoids this issue by selecting discrete Boolean functions rather than averaging parameters. We propose Federated Boolean Feature Selection (FBFS), which extends BFS to LGNs by treating hardened last-layer gates as candidate Boolean features and aggregating per-class, per-gate sufficient statistics across clients. Clients share only firing aggregates computed on a held-out split, allowing the server to reconstruct discriminative scores, defined as the gap between a gate's mean firing for a class and for the remaining classes, without accessing raw data. This procedure is equivalent to centralized BFS on the union of client data. We further introduce a diversity-regularized variant of FBFS that encourages complementary feature selection across gates. Empirically, FBFS is the only data-free method we evaluate whose performance remains stable under extreme non-IID settings, outperforming parameter-space merging and other baselines by large margins. It maintains high accuracy across a wide range of heterogeneity levels and scales to large architectures. In addition, FBFS enables efficient deployment: the resulting models compile to hardware-efficient logic representations with substantially fewer post-synthesis cells than ensemble-based alternatives while remaining bit-exact with their software counterparts.
Feedback World Model Enables Precise Guidance of Diffusion Policy
Tuo An ⋅ Jindou Jia ⋅ Gen Li ⋅ Jingliang Li ⋅ Chuhao Zhou ⋅ Pengfei Liu ⋅ Bofan Lyu ⋅ Jiaqi Bai ⋅ Xinying Guo ⋅ Geng Li ⋅ Jianfei Yang
World models aim to improve robotic decision making by predicting the consequences of actions. However, in practice, their predictions often become unreliable once the robot encounters states outside the training distribution, limiting their effectiveness at deployment. We observe that execution itself provides a natural but underutilized signal: after each action, the robot directly observes the true next state, revealing the mismatch between predicted and actual outcomes. Building on this insight, we propose $\textbf{feedback world model}$, a new paradigm that closes the loop between prediction and observation at inference time. Instead of treating the world model as a static open-loop predictor, our method maintains a lightweight feedback state that is updated online to iteratively correct future predictions, compensating for model errors using real-time observations without additional training data or parameter updates. We show that this process can be interpreted as a latent-space observer and admits convergence guarantees under mild conditions. We further introduce action-aware guidance to better translate corrected predictions into control by emphasizing action-controllable components while suppressing irrelevant variations. Experiments on LIBERO-Plus, Robomimic, and real-world manipulation tasks demonstrate that our method substantially improves both prediction accuracy and policy performance under distribution shift. In particular, it reduces world model prediction error by up to 76.4\% and improves out-of-distribution (OOD) success rate by 30\%. These results show that incorporating real-time feedback at inference time provides a simple yet powerful alternative to static world modeling.
Ferrogen: Generative Pipeline for Guided Search of Novel Ferroelectric Material for Logic and Memory
Yuan Sheng Fang ⋅ Dmitri E Nikonov ⋅ Ikenna Odinaka ⋅ Alan Kalitsov ⋅ Roza Kotlyar ⋅ Arnab Kabiraj ⋅ Benjamin W Chen ⋅ Shundan Xiao ⋅ Brian Demsky ⋅ Chryston B Meng ⋅ Chang Wei Kang ⋅ Teck L Tan ⋅ Jin Hongmei ⋅ Debo Olaosebikan ⋅ Sasikanth Manipatruni ⋅ Amrita Mathuriya
Discovery of novel materials for physical memory is essential for superior memory systems enabling AI hardware scaling. Among them, ferroelectrics hold promise for highly scalable DRAM. Yet their discovery remains driven largely by human intuition and exhaustive database screening. We present Ferrogen, a generative pipeline for the targeted discovery of novel ferroelectric materials for memory by fine-tuning the diffusion-based crystal structure generator Mattergen on ferroelectric-relevant properties. To enable property-conditioned generation, we construct a ML-labeled dataset using a suite of fast machine learning estimators, including a novel two-stage ensemble polarization predictor, dramatically reducing reliance on high-throughput Density Field Theory (DFT) during training and screening. The base Mattergen model is fine-tuned via lightweight adapter layers conditioned on polarization, switching energy, and metal-probability embeddings, with classifier-free guidance enabling targeted generation of candidates with high polarization, low but finite switching barriers, and insulating character. Generated candidates are screened by the same ML estimators and validated through rigorous DFT calculations, including structural relaxation, band gap verification, and Berry phase polarization computation. Against the ML screens, Ferrogen achieves a roughly $80 \times$ improvement in search efficiency over exhaustive database search. The screened candidates are then finally validated using DFT and made publicly available as a database of theoretically confirmed novel ferroelectrics. The pipeline is readily extensible to other functional material classes, establishing generative models as a powerful paradigm for accelerating electronic materials discovery for high performance memory.
FIND: Frequency Invariance Disentanglement for Test-Time Adaptation in LiDAR 3D Detection
Dongxiao Li ⋅ Zixuan Hu ⋅ LINGYU DUAN
LiDAR-based 3D Object Detection (L3OD) is a fundamental 3D perception task, yet its performance is often compromised by domain shifts arising from diverse sensor configurations and environmental conditions. Existing Test-Time Adaptation (TTA) methods primarily leverage self-supervision in the spatial domain, yet they frequently confound semantic content with domain style due to the sparse, non-Euclidean nature of point clouds. To overcome this limitation, we propose FIND (Frequency INvariance Disentanglement), a novel TTA framework that precisely disentangles domain-invariant content from domain-specific interference within the frequency domain. Our methodology utilizes B-Spline fitting to construct locally adaptive filters integrated into a dual-stream architecture: an Invariant Stream extracts robust features to guide pseudo-labeling, while a Specific Stream enforces consistency against domain perturbations. By combining invariance-driven alignment with stability-guided regularization, our approach dynamically extracts robust domain-invariant features while suppressing domain-specific interference. Extensive experiments on cross-dataset adaptation and robust corruption benchmarks demonstrate that FIND significantly outperforms state-of-the-art methods.
Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models
Mohammed Saidul Islam ⋅ Arash Afkanpour ⋅ Negin Baghbanzadeh ⋅ Farnaz Kohankhaki ⋅ Afshin Cheraghi ⋅ Ali Kore ⋅ Shayaan Mehdi ⋅ Elham Dolatabadi
Evaluation of foundation models often rely on aggregate scores from benchmarks that lack comprehensive coverage and metadata for a fine-grained evaluation. We introduce a framework for automated benchmark generation. Our framework generates evaluation problems grounded in reference material, such as textbooks, producing benchmarks with broad coverage, rich metadata, and robustness to contamination. The pipeline employs a multi-agent architecture for problem generation and a solution-graph-driven strategy that significantly improves the reliability of ground truth solutions. Using the framework, we generate three benchmarks in Machine Learning, Corporate Finance, and Personal Finance. Expert review finds a significantly lower ground-truth error rate than previous benchmarks such as MMLU and GSM8K. Evaluation of 12 commercial and open-source models shows that our benchmarks achieve near-uniform competency coverage and surface performance differences across models that existing benchmarks fail to capture. We open-source the framework and our curated benchmarks.
Fisher information and the geometry of memorization in neural networks
Daniel Bernstein ⋅ Chase Goddard ⋅ Luca Di Carlo ⋅ Vudtiwat Ngampruetikorn ⋅ David Schwab
A fruitful approach towards understanding generalization in machine learning involves characterizing how models separate generalizable structure from idiosyncratic or even corrupted features of training data. Fisher information has emerged as a diagnostic tool for this separation, leveraging loss geometry to distinguish generalization from memorization. Here we investigate this connection in vision models trained on CIFAR-10 with controlled label noise that can be learned only through memorization. We show that projecting model weights onto the top Fisher eigenvectors decouples general task performance from noisy sample memorization, and we identify the layer in which this decoupling emerges. We show that the same phenomena are present in small transformers trained to perform in-context learning and reproduced in closed form in a simple, analytically solvable linear model. Our results reveal that Fisher information captures aspects of the geometric structure underlying how neural networks allocate capacity between generalization and memorization.
Flexible Flows for Biological Sequence Design
Yogesh Verma ⋅ Dani Korpela ⋅ Harri Lähdesmäki ⋅ Vikas Garg
Designing functional biological sequences requires navigating vast discrete spaces under strict evolutionary and biophysical constraints. Discrete Flow Matching (DFM) offers a generative framework over such spaces, but existing approaches rely on biologically uninformative couplings and offer limited flexibility for variable-length sequence generation and fine-grained control. We propose a structured coupling that encodes domain-specific preferences among sequence elements, biasing the source distribution toward plausible regions without modifying the flow objective or training procedure. Building on this, we introduce a latent edit-based rate parameterization that models variable-length generation via edit operations conditioned on a shared global latent, akin to a latent variable model, while remaining tractable. We further introduce a latent classifier-free guidance mechanism that steers generation coherently in continuous latent space, along with Dirichlet-prior temperature scaling for test-time control over edit operations. Our method achieves state-of-the-art performance across diverse biological sequence tasks, including density estimation, unconditional and conditional DNA sequence generation, and peptide sequence generation.
Flexible Routing via Uncertainty Decomposition
Charlotte Peale ⋅ Siddartha Devic ⋅ Parikshit Gopalan ⋅ Udi Wieder ⋅ Aravind Gollakota
A key strategy for balancing performance and cost in modern machine learning systems is to dynamically route queries to either a low-cost model or a more expensive oracle (such as a large pretrained model or human expert), an approach known as model routing. In this work we present a new uncertainty-aware router that (1) avoids unnecessary oracle calls on inherently ambiguous queries, and (2) adapts dynamically to different loss functions and cost parameters through simple hyperparameter changes, without retraining. Our method, applicable to any classification setting where multiple independent annotations per input are available, is based on decomposing total uncertainty into irreducible and reducible components using higher-order predictors [Ahdritz et al., 2025]. This enables a unified approach to both routing and abstention: predict with the weak model when uncertainty is low, route to the oracle when reducible uncertainty is high, and abstain when irreducible uncertainty is high. Our router comes with strong theoretical guarantees bounding regret relative to optimal task-specific routers. We conduct experiments on both synthetic and real-world datasets that demonstrate the benefits of our approach in suitable regimes---in particular, whenever reducible and irreducible uncertainty are not too correlated.
FloatDoor: Platform triggered Backdoors in LLMs
Nils Loose ⋅ Jonas Sander ⋅ Felix Mächtle ⋅ Thomas Eisenbarth
Large language models (LLMs) are increasingly deployed in sensitive settings such as software engineering, where their outputs directly shape downstream artifacts. Recent work has shown that an identical model can produce measurably different outputs depending on the deployment platform, a consequence of non-associative floating-point arithmetic and divergent kernel implementations. We study the security implications of this platform-dependent variability and uncover a novel attack surface on LLM deployments. We introduce FloatDoor, the first input-independent, platform-triggered backdoor attack against generative LLMs. The compromised model exhibits adversary-chosen behavior when served on a target platform and is otherwise benign. FloatDoor is realized through two lightweight LoRA adapters, one that amplifies inter-platform numerical divergence and one that binds the resulting platform signature to a malicious downstream task, while leaving aggregate model utility largely intact. FloatDoor exploits a pronounced time-of-check, time-of-use gap between model auditing and serving. We demonstrate FloatDoor on Qwen3-4B across a broad range of deployment targets, including NVIDIA GPUs, Google TPUs, AWS Graviton, and Alibaba Yitian-710. As a final case study, we show that FloatDoor reliably induces exploitable code vulnerabilities on a chosen target platform. Our results establish a new class of attacks on LLM deployments and underscore the pressing need for trusted model supply chains in sensitive, LLM-powered applications.
FoMo: Forking Moment in Generative Trajectory as a Perceptual Distance
Jaihyun Lew ⋅ Mingi Jung ⋅ Minjun Park ⋅ Wooseok Song ⋅ Sungroh Yoon
Reference-based image quality assessment (IQA) metrics aim to reflect how humans perceive the perceptual distance between a pair of images. To learn how the human visual system (HVS) operates, recent reference-based IQA metrics heavily rely on human-annotated data. Mean opinion score (MOS)-based pointwise scoring, which assigns a scalar quality value per image, is preferable for annotation but is prohibitively expensive to collect at scale and is known to be noisy due to inconsistent human judgments. As an alternative, two-alternative forced choice (2AFC) pairwise labels have gained popularity due to their reliability and efficiency, but they capture only relative comparisons between pairs. In this paper, we propose a fully automated data generation pipeline that generates pointwise perceptual distance labels between image pairs without any human annotation. Our approach exploits the generative dynamics of diffusion models as a perceptual distance proxy, where the coarse structure of an image is generated in the early timesteps and the fine details are generated in the later timesteps. Images that fork early in the generation process share only coarse structure and are perceptually far apart; images that fork late differ only in fine detail. We demonstrate that the diffusion trajectory aligns well with the human visual system, and use this forking moment, FoMo, as a reference-grounded distance label to supervise the training of a reference-based IQA metric. The pointwise labels, which support universal comparison between arbitrary image pairs, enable an information-rich training objective. Extensive experiments across diverse backbone architectures confirm the effectiveness of our generation pipeline, outperforming human-annotated datasets in multiple benchmarks.
Forest-Guided Semantic Transport for Label-Supervised Manifold Alignment
Adrien Aumon ⋅ Myriam Lizotte ⋅ Guy Wolf ⋅ Kevin Moon ⋅ Jake Rhodes
Label-supervised manifold alignment bridges the gap between unsupervised and correspondence-based paradigms by leveraging shared label information to align multimodal datasets. Still, most existing methods rely on Euclidean geometry to model intra-domain relationships. This approach can fail when features are only weakly related to the task of interest, leading to noisy, semantically misleading structure and degraded alignment quality. To address this limitation, we introduce FoSTA (Forest-guided Semantic Transport Alignment), a scalable alignment framework that leverages forest-induced geometry to denoise intra-domain structure and recover task-relevant manifolds prior to alignment. FoSTA builds semantic representations directly from label-informed forest affinities and aligns them via fast, hierarchical semantic transport, capturing meaningful cross-domain relationships. Extensive comparisons with established baselines demonstrate that FoSTA improves correspondence recovery and label transfer on synthetic benchmarks and delivers strong performance in practical single-cell applications, including batch correction and biological conservation.
Forgetting is Not Always Bad: A Neuro-Inspired Memory Repair Mechanism for Poisoned LLM Agents
Lei Liu ⋅ Yunji Liang ⋅ Xiaowen Zhang ⋅ Jingqi Liu ⋅ Qi Li ⋅ Bin Guo ⋅ Zhiwen Yu
Large language model (LLM) agents are highly vulnerable to memory-poisoning attacks. Existing defenses primarily rely on external modules for either static memory isolation or continuous online auditing, resulting in low memory efficiency and high computational overhead. To address these limitations, inspired by neuroscientific mechanisms that weak reactivation can destabilize memories and induce decay in the absence of reinforcement, we propose an agent-intrinsic training-free memory repair framework, \textbf{DREAM} (\textbf{D}ynamic \textbf{R}eactivation for \textbf{E}ngram \textbf{A}ttenuation in \textbf{M}emory), that enables selective functional forgetting of poisoned memories. Specifically, DREAM implements a three-stage pipeline: perturbation-induced implicit reactivation to reveal the structural fragility of malicious memories, adaptive anomaly diagnosis based on multi-dimensional activation patterns to detect topological anomalies, and a dynamic memory repair module that selectively suppresses harmful memories while reinforcing benign ones. Extensive experiments across diverse attackers and real-world tasks demonstrate that DREAM reduces attack success rates by over 95\% against backdoor poisoning while preserving strong benign utility. DREAM also achieves competitive robustness against injection attacks. In terms of runtime and token consumption, DREAM achieves up to a 2.13$\times$ speedup and a 37.7\% reduction in token consumption compared with A-MemGuard. Furthermore, DREAM also achieves a task success rate of 92.96\% on poisoned multi-agent systems.
Medical image segmentation requires the precise alignment of macroscopic semantics and microscopic geometry, yet existing representation paradigms struggle to balance massive low-frequency semantic context with ultra-sparse, high-frequency boundary details. To address this fundamental information asymmetry, we introduce Fractal-G, a plug-and-play multi-scale fusion module that presents a continuous Grid-Graph-Grid feature reconstruction paradigm. Within this framework, an Uncertainty-Aware Fission Router dynamically translates regular image grids into a non-uniform heterogeneous graph, adaptively allocating dense microscopic nodes to complex boundaries while retaining sparse macroscopic nodes in homogeneous background regions. To ensure unbiased feature aggregation across these varying physical scales, an Area-Aware Continuous Neighborhood Graph explicitly incorporates physical node sizes into the topological message passing. Finally, a Topology-Aware Implicit Renderer projects the unstructured graph back into dense, artifact-free continuous feature fields. Extensive experiments on diverse medical image segmentation tasks (including skin lesions and polyps) demonstrate that Fractal-G easily integrable with mainstream backbones, consistently achieving state-of-the-art geometric fidelity and cross-resolution robustness with only 12\% GFLOPs increase (U-Net based). Our code is publicly available at https://anonymous.4open.science/r/Fractal-G/.
Fractional State Space Transition for Long Sequence Modeling
Ivan Kobyzev ⋅ Abbas Ghaddar ⋅ Ali Nasiri-Sarvi ⋅ Lifeng Shang ⋅ Yufei CUI
State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for long-context performance. Most modern SSMs rely on ODE-based dynamics that lead to exponential forgetting, limiting their ability to retain information over broad temporal ranges. We introduce FRAC, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory. To make fractional dynamics practical, FRAC approximates the heavy-tailed target kernel with a finite-state, log-spaced sum of exponential modes. This construction turns fractional memory into an efficient recurrent module with parallel training and prefill, while retaining bounded-state autoregressive decoding. Extensive experiments, including 1.3B-parameter language modeling, demonstrate that FRAC consistently improves long-context performance over state-of-the-art SSM baselines while staying competitive on short-context. These results show that fractional dynamics provide a practical and effective prior for long-context SSMs.
FrameScout: Scouting Query-Relevant Frames for Long Video Understanding
Haonan Hu ⋅ Shuhao Chen ⋅ Weisen Jiang ⋅ Lizhao Gao ⋅ James Kwok ⋅ Yu Zhang
Long video understanding requires VideoLLMs to answer queries about videos spanning tens of minutes to hours, but such long videos produce far more visual tokens than current VideoLLMs can process. To address this issue, keyframe selectors are proposed to select the most query-relevant frames by computing the similarity between frame and query embeddings. Existing selectors either encode frames independently without temporal context or process chunks in isolation with ranking objectives that are difficult to optimize. To alleviate those limitations, we propose FrameScout with two core designs: (i) a streaming selector architecture that leverages a sliding-window KV cache and successor aggregation to produce temporally-contextualized frame embeddings; (ii) a frame-query contrastive objective that aligns frame and query embeddings in a shared space, directly training the selector to distinguish relevant frames from irrelevant ones. Extensive experiments on three long-video benchmarks and six base VideoLLMs demonstrate that FrameScout consistently achieves state-of-the-art performance.
FreDRec: Frequency-Decoupled Knowledge Distillation for Multimodal Recommendation in Missing Modalities Scenarios
Qixian Chen ⋅ Xiaolong Xu ⋅ Haolong Xiang ⋅ Kun Yi ⋅ Xiaoyu Xia ⋅ Lianyong Qi ⋅ Wei Fan
Multimodal Recommender Systems (MRSs) have achieved significant success in personalized recommendations by integrating diverse modalities such as images, text, and audio. While in complex real-world industrial scenarios where MRS has usually been deployed, the presence of missing modalities is very common. Most existing MRS methods recover missing data by disentangling available modality features into general and specific representations, thereby generating missing modalities based on personalized aligned feature. However, these methods rely on implicit constraints within pre-extracted multimodal feature spaces, which usually introduce noise that prevents the effective decoupling of general and specific semantics, thereby diminishing recommendation accuracy. Moreover, unguided linear transformations of available modality features for generation lead to semantic hallucinations, which inevitably degrade the ranking quality of the recommendation results. To address these issues, we propose the Frequency-Decoupled Knowledge Distillation Framework for Multimodal Recommendation (\textbf{FreDRec}), which conducts effective and robust missing feature generation. Specifically, motivated by the physical characteristics of multimodal signals, FreDRec utilizes Fourier and Discrete Wavelet Transforms to explicitly decouple raw multimodal signals into low-frequency core semantics and high-frequency details, which is mathematically proven to effectively minimize noise introduction at the source. Next, to ensure robust feature generation, we pre-train a teacher model based on decoupled multi-view graph propagation to guide a student model via multi-level knowledge distillation, which establishes an end-to-end highly efficient lightweight architecture and achieves precise missing modality generation with expert collaborative bounds. Finally, extensive experiments on multiple public benchmark datasets demonstrate that FreDRec achieves superior performance under various modality missing ratios.
FreeAct: Demonstration-Free Robot Adaptation via Action-Grounded Generated Videos
Zhe-Han Mo ⋅ Jia-Ning Li ⋅ Jiaxi Song ⋅ Meng-Hao Guo ⋅ Shi-min Hu
Adapting pretrained Vision-Language-Action (VLA) policies to new deployment environments typically requires collecting expert demonstrations in each target domain, which is costly and difficult to scale. In this paper, we present FreeAct, a framework for demonstration-free robot adaptation that converts generated videos into physically grounded supervision. However, a key challenge is that generated manipulation videos, while visually plausible, often contain embodiment artifacts and kinematic inconsistencies that make direct inverse-dynamics labeling unreliable. FreeAct addresses this challenge by learning a discrete latent action space from large-scale multi-lab robot videos and aligning it with camera-frame end-effector SE(3) motion and gripper-state changes. The resulting latent action tokenizer captures task-relevant motion while reducing reliance on embodiment-specific visual artifacts, enabling generated target-domain videos to be pseudo-labeled with reliable latent action tokens. These tokens are first injected into VLA pretraining and then used, together with source-domain action supervision, to adapt the policy to unseen environments. Empirically, FreeAct improves success from 2\% to 57\% under target-domain shift using generated videos alone, and further to 74\% through MLLM-guided self-evolution on the policy's own rollouts, closing the adaptation loop without expert demonstrations. Further analysis shows that the learned latent action space forms a kinematically structured and transferable representation, establishing generated videos as a scalable source of supervision for robot manipulation.
Free Decompression with Algebraic Spectral Curves
Siavash Ameli ⋅ Chris van der Heide ⋅ Liam Hodgkinson ⋅ Michael Mahoney
Tools from random matrix theory have become central to deep learning theory, using spectral information to provide mechanisms for modeling generalization, robustness, scaling, and failure modes. While often capable of modeling empirical behavior, practical computations are limited by matrix size, often imposing a restriction to models that are too small to be realistic. This motivates the inference of properties of larger models from the behavior of smaller ones. Free decompression (FD) is a recently proposed method for extrapolating spectral information across matrix sizes, but its utility is currently limited by strong assumptions that preclude its implementation on more realistic machine learning (ML) models. We use algebraic spectral curve theory to provide a general FD methodology for spectral densities whose Stieltjes transform satisfies an algebraic relation, a modeling assumption that is more likely to hold in practice. This recasts FD as an evolution along spectral curves which can be readily integrated. Our framework enables the expansion of spectral densities that have multiple or multi-modal bulks, that exist at multiple scales, and that contain atoms, all characteristic of real-world data and popular ML models. We demonstrate the efficacy of our framework on models of interest in modern ML, including Hessian and activation matrices associated with neural networks and large-scale diffusion models.
Frequency-Structured Hamiltonian Neural Network for Multi-Timescale Dynamics
Yaojun Li ⋅ Yulong Yang ⋅ Christine Allen-Blanchette
Hamiltonian Neural Networks and related structure-preserving dynamics models encode conservation laws, but their scalar Hamiltonian parameterizations inherit the spectral bias of deep networks, limiting their ability to learn stiff multi-timescale systems with coupled fast and slow dynamics. We introduce the Frequency-Structured Hamiltonian Neural Network (FS-HNN), which decomposes the system Hamiltonian into learned components trained on frequency-filtered views of observed trajectories, then recombines them into a single scalar Hamiltonian so that the learned ODE dynamics remain Hamiltonian by construction. For PDEs with unknown or problem-dependent structure, FS-HNN represents Hamiltonians as neural functionals, learns the action of the dynamics operator, and uses projection to enforce conservative structure when appropriate. Across ODE and PDE benchmarks, including the Fermi--Pasta--Ulam--Tsingou chain, shallow water equations, and incompressible Taylor--Green vortex, FS-HNN achieves improved long-horizon rollout accuracy and more accurate energy behavior than existing structure-preserving baselines.
Frequency-Synchronized Boundary Coupling for Training-Free Multi-Prompt Long Video Generation
Yipan Xu ⋅ Shuyong Gao ⋅ Jiyuan Fu ⋅ Keliang Yin ⋅ Jiayu Chen ⋅ Qianyu Guo ⋅ Yang Liu ⋅ Wenqiang Zhang
Text-to-video generation has made remarkable progress with diffusion models and transformer-based video architectures, yet generating coherent long videos from evolving textual descriptions remains challenging. Training long-video models from scratch requires substantial computation and large-scale multi-event data, while training-free multi-prompt generation must activate new prompt semantics without breaking temporal continuity. We identify a frequency-agnostic boundary coupling problem in existing transition strategies: weak coupling causes abrupt changes, flicker, and background jumps, whereas overly strong or coarse coupling may over-preserve source semantics and produce ghosting or duplicated subjects around prompt boundaries. To address this problem, we propose FreqSync, a training-free framework that formulates prompt transitions as frequency-selective boundary coupling. FreqSync combines Frequency-Synchronized Source-Conditioned Attention (FS-SCA) with Residual-Clipped Spectral Stitching (RCSS), enabling adjacent segments to share stable low- and mid-frequency structure and motion cues while keeping high-frequency prompt-specific details target-dominant under explicit residual constraints. Extensive experiments show that FreqSync achieves state-of-the-art transition smoothness and prompt alignment while maintaining strong perceptual video quality, demonstrating its effectiveness for coherent multi-prompt long video generation.
From Contexts to Conditionals: Statistical Self-Consistency of Persona Prompting
Patrik Wolf ⋅ Thomas Kleine Buening ⋅ Andreas Krause ⋅ Celestine Mendler-Dünner
Large language models are increasingly used to estimate uncertain quantities from context, raising the question of whether their probabilistic outputs are internally consistent. A basic test of self-consistency is whether these estimates adhere to the law of total probability. We investigate this requirement through the lens of binary conditioning trees, which recursively partition a population into increasingly fine-grained subsets and expose multiple ways of estimating the same aggregate quantity. This construction yields a basic self-consistency check: marginal estimates must agree with prior-weighted aggregations of conditional estimates over any partition of the population. We find that current models systematically violate this requirement. In a case study on persona prompting, prior-weighted aggregates are consistently better aligned with human population statistics than direct estimates. Notably, this benefit of specificity persists even when the prior weights are themselves estimated by the LLM. Turning this discrepancy into a general evaluation criterion, we propose a family of self-consistency checks for LLMs grounded in the law of total probability. By evaluating these checks across frontier models, we show that failures of statistical consistency are widespread and not confined to persona prompting. Together with a benchmark dataset, our work provides a testbed for understanding when in-context learning can be treated as conditional inference.
From Cursed to Competitive: Closing the ZO–FO Gap via Input-to-State Stability
Amir Ali Farzin ⋅ Philipp Braun ⋅ Iman Shames
While it is generally understood that zeroth-order (ZO) algorithms have an extra dependency on their number of iterations for any choice of parameters, compared to their first-order (FO) counterparts, in this work, we show that under several conditions, in expectation, ZO methods do not suffer from extra dimension dependencies in their convergence rates with respect to their FO counterparts. We look at optimisation algorithms from the dynamical systems perspective and analyse the conditions under which one can formulate the average of a ZO algorithm as the average of its FO counterpart with bounded perturbations with values dependent on design parameters. Then, using input-to-state stability properties, we show ZO methods follow the same decay rate as their FO counterparts and converge to a neighbourhood of the fixed point of FO methods, where its radius depends on the bound of the norm of the perturbations, which can be made arbitrarily small. The theoretical findings are illustrated via numerical examples.
From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model
Yatai Ji ⋅ Teng Wang ⋅ Yuying Ge ⋅ Zhiheng Liu ⋅ Sidi Yang ⋅ Ying Shan ⋅ Ping Luo
Discrete diffusion models have emerged as a promising direction for vision-language tasks, offering bidirectional context modeling and theoretical parallelization. However, their practical application is severely hindered by a train-inference discrepancy, which leads to catastrophic error cascades: initial token errors during parallel decoding pollute the generation context, triggering a chain reaction of compounding errors and leading to syntactic errors and semantic hallucinations. To address this fundamental challenge, we reframe the generation process from passive denoising to active refining. We introduce ReDiff, a refining-enhanced diffusion framework that teaches the model to identify and correct its own errors. Our approach features a two-stage training process: first, we instill a foundational revision capability by training the model to revise synthetic errors; second, we implement a novel online self-correction loop where the model is explicitly trained to revise its own flawed drafts by learning from an expert's corrections. This mistake-driven learning endows the model with the crucial ability to revisit and refine its already generated output, effectively breaking the error cascade. Extensive experiments demonstrate that ReDiff significantly improves the coherence and factual accuracy of generated content, enabling stable and efficient parallel generation far superior to traditional denoising methods.
From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation
Shuchang Ye ⋅ Jinqiang Yu ⋅ Zhujun Xiao ⋅ Yajing Kong ⋅ Yist Y Lin ⋅ Yang Ma ⋅ Jiaxi Liu ⋅ Xiaolei XU ⋅ Zheng Yu
Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation to platform-specific data distributions, policy-specific objectives, and product-level safety constraints. As a result, platforms must undertake internal model development, naturally turning to shared public research for guidance. However, existing multimodal foundation-model studies primarily report architectures, training recipes, data scaling strategies, and benchmark results, but provide less systematic guidance on how failures should be localized and translated into targeted model-development interventions. Interventions are essential because deployment failures are rarely self-explanatory. Similar failures can originate from different causes. Without targeted interventions, improvement reduces to heuristic trial-and-error, where benchmark improvements are weakly attributable, and failures are difficult to trace to their underlying causes. To address this gap, we present a diagnostic methodology for industry-scale Audio-Visual-Language Models~(AVLM) development. The methodology maps model failures into a taxonomy of observable failure signatures and links each class of failure to an intervention space. We instantiate this methodology across the development and alignment lifecycle of an AVLM foundation model for a large-scale video and live-streaming platform. The resulting system supports over 100 regions and is designed for noisy, ambiguous, and highly diverse content drawn from global platform traffic.
From Generic to Dedicated: A Novel Optimizer for Online Continual Learning
Yongyi Wu ⋅ Zheng Wang ⋅ Sen Lin
Online Continual Learning (OCL) requires models to learn from non-stationary data streams under a strict single-pass constraint, making them highly susceptible to catastrophic forgetting. While existing studies have explored various strategies from data, architecture, and optimization perspectives, the optimizer often directly uses generic approaches such as SGD and AdamW. In this work, we reveal a new link between optimization dynamics and OCL needs by recasting Polyak-Ruppert averaging as the engine of "plasticity" and Primal averaging as the anchor of "stability". This inspires our novel optimizer, SPIN (Stability-Plasticity INterpolation), which explicitly decouples plasticity and stability by interpolating between a fast-moving "plasticity" sequence and an adaptive "stability" sequence. We theoretically demonstrate that SPIN implicitly employs an inverse Hessian approximation, providing crucial Tikhonov regularization to damp noisy, single-sample Hessian estimates and mitigate forgetting. More importantly, SPIN can be seamlessly integrated existing OCL methods by taking resulted gradients as input and replacing standard optimizers. Extensive experiments show that using SPIN instead of standard optimizers yields consistent and significant performance gains across multiple OCL approaches, including the challenging rehearsal-free setting and under strict memory constraints.
From Perception to Punchline: Empowering VLM with the Art of In-the-wild Memes
Xueyan Li ⋅ Yingyi Xue ⋅ Mengjie Jiang ⋅ Qingzi Zhu ⋅ Yazhe Niu
Generating humorous memes is a challenging multimodal task that moves beyond direct image-to-caption supervision. It requires a nuanced reasoning over visual content, contextual cues, and subjective humor. To bridge this gap between visual perception and humorous punchline creation, we propose HUMOR, a novel framework that guides VLMs through hierarchical reasoning and aligns them with group-wise human preferences. First, HUMOR employs a hierarchical, multi-path Chain-of-Thought (CoT): the model begins by identifying a template-level intent, then explores diverse reasoning paths under different contexts, and finally anchors onto a high-quality, context-specific path. This CoT supervision, which traces back from ground-truth captions, enhances reasoning diversity. We further analyze that this multi-path exploration with anchoring maintains a high expected humor quality, under the practical condition that high-quality paths retain significant probability mass. Second, to capture subjective humor, we train a pairwise reward model that operates within groups of memes sharing the same template. Following established theory, this approach ensures a consistent and robust proxy for human preference, even with subjective and noisy labels. The reward model then enables a group-wise RL optimization, providing a theoretical guarantee for monotonic improvement within the trust region. Extensive experiments show that HUMOR empowers various VLMs with superior reasoning diversity, more reliable preference alignment, and higher overall meme quality. Beyond memes, our work presents a general paradigm for open-ended, human-aligned multimodal generation, where success is guided by comparative judgment within coherent output groups.
From Retrieval to Recognition: Knowledge-Enhanced Foundation Models for Time Series Classification
Zhen Liu ⋅ Yucheng Wang ⋅ Yingpeng Du ⋅ Tianjun Wei ⋅ Qianli Ma ⋅ Min Wu ⋅ Jie Zhang ⋅ Zhu Sun
Time series foundation models (TSFMs) have achieved notable progress in classification tasks through large-scale pretraining. However, existing methods rely on implicit parameter fine-tuning, where task-relevant information is entangled with abundant irrelevant knowledge in vast parameter spaces. This may hinder TSFMs from precisely activating useful knowledge for downstream classification. To this end, we propose REKIN, a plug-and-play knowledge-enhanced framework for time series classification built upon pretrained TSFMs within a retrieval-to-recognition paradigm. Unlike traditional parameter-only approaches, REKIN introduces a decoupled query encoder to extract discriminative patterns and retrieve instance-level numerical knowledge from a constructed knowledge base. The retrieved explicit knowledge is then fused with the TSFM’s implicit parametric embeddings, enabling joint knowledge transfer at both the parameter and feature levels. In this way, REKIN allows TSFMs to recall task-relevant patterns learned during pretraining, improving the efficiency of knowledge utilization in downstream transfer. Experiments on 158 time series datasets show that REKIN consistently improves the classification performance of state-of-the-art TSFMs. Ablation and knowledge-enhanced studies further validate the effectiveness of the proposed framework.
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering with Generative Optimization
Dapeng Jiang ⋅ Yizhe Chi ⋅ Kaisen Yang ⋅ Tianwei Luo ⋅ Boshi Zhang ⋅ Deyao Hong ⋅ Dianqiao Lei ⋅ Yifan Zhou ⋅ Weiyang Jin ⋅ Xiaoyan Fan ⋅ Han Hao ⋅ Zhe Cao ⋅ Qian Houde ⋅ Qingle Liu ⋅ Bohan Lyu ⋅ Bowen Wang ⋅ Situ Wang ⋅ Youjie Zheng ⋅ Bingxiang He ⋅ Eren Cai ⋅ Calvin Xiao ⋅ Qinhuai Na
Engineering value is often created not by producing a single correct answer, but by improving a feasible artifact under hard constraints. We introduce **Frontier-Eng**, a benchmark for generative optimization agents that repeatedly edit executable engineering artifacts, receive frozen verifier feedback, and maximize the best feasible score found within a fixed budget. Frontier-Eng contains $47$ tasks across five engineering categories, spanning kernels, cryptography, quantum circuits, inventory and scheduling, robotics and control, optics and communication systems, structural design, reactions, and sustainable infrastructure. Each task provides a feasible starting solution, an editable artifact, and a read-only evaluator with verifier-parsed scoring, enabling cross-domain comparison without reducing engineering quality to binary pass/fail completion. We evaluate frontier models under OpenEvolve and compare search frameworks across the full task set. Current agents exhibit clear iterative improvement: **gpt-5.4** obtains the best average rank under OpenEvolve. Long-horizon runs reveal a dual **inverse-law pattern**: improvement events become rarer with iteration, and improvement magnitudes shrink with improvement count. At the same time, improvement is uneven across domains and remains sensitive to search depth, exposing a concrete frontier for engineering-oriented agents under realistic constraints and verifier-grounded feedback.
Fully First-Order Algorithms for Online Non-Convex Bilevel Optimization
Tingkai Jia ⋅ Ting Wang ⋅ Cheng Chen
In this work, we study nonconvex–strongly convex online bilevel optimization (OBO) using only first-order oracle. Existing OBO algorithms are mainly based on hypergradient descent, which requires access to a Hessian-vector product (HVP) oracle and potentially incurs high computational costs. By reformulating the original OBO problem as a single-level online problem with inequality constraints and constructing a sequence of Lagrangian function, we eliminate the need for HVPs arising from implicit differentiation. Specifically, we propose a fully first-order algorithm for OBO, and provide theoretical guarantees showing that it achieves regret of $O(1 + V_T + H_{2,T})$ with a total of $O(T\log T)$ iterations, where $V_T$ measures the variation in function values and $H_{2,T}$ characterizes the drift variation of the inner-level optimal solution. We also establish a sublinear regret bound under the single-loop structure by introducing additional gradient-variation terms. Furthermore, we develop an improved variant with an adaptive inner-iteration scheme, which removes the dependence on $H_{2,T}$ and achieves regret of $O(\log T + V_T)$. Finally, under the stochastic OBO setting, we establish the regret bound for the fully first-order algorithm, i.e., $O(T^{2/3}(1 + \sigma^2) + V_T + H_{2,T})$. Numerical experiments demonstrate the feasibility of our algorithm and support our theoretical findings.
FusionAudit: Pathway-Conditioned Robustness Auditing for Native Multimodal Models
Jiaming Zhang ⋅ Zherui Li ⋅ Siqi Guo ⋅ Junhao Dong ⋅ Kun Wang ⋅ Wei Yang Bryan Lim
As vision-language models (VLMs) transition from frozen-encoder pipelines to native multimodal architectures, the visual channel increasingly serves as both a perceptual input and a carrier of rendered text that competes with user prompts. Current single-attack, pixel-centric evaluations fail to capture this complexity: they collapse distinct visual attack surfaces into a single scalar and ignore how architectural shifts alter vulnerability. We introduce \textbf{PaRMA}, an evidence-based auditing framework that defines pathway-conditioned risk over two primary image-level interventions: pixel perturbation and rendered-text insertion. Instead of arbitrary attack enumeration, PaRMA uses targeted diagnostics to expose architecture-specific blind spots. Notably, we uncover a \textit{see-but-not-deceived} signature where native VLMs attend heavily to typographic overlays but visually reject them, and we identify severe gradient dispersion in deep-fusion models. To exploit these findings, we instantiate \textbf{FusionAudit}, introducing two targeted attacks: \textbf{MANTIS} (margin-aware pixel perturbation) and \textbf{Nat-Evid} (host-congruent semantic insertion). Evaluating 18 open-weight VLMs across a novel nativeness taxonomy, we demonstrate that legacy single-attack rankings (e.g., PGD) are highly unstable for native models. Our diagnostic-driven attacks substantially raise the estimated residual risk, proving that PaRMA provides a falsifiable, evidence-driven auditing methodology to uncover vulnerabilities systematically hidden by standard benchmarks.
FusionNeXt: Sequence-First 3D Multi-Modal Fusion in the Era of LLMs
Yu Hong ⋅ Xiaosong Jia ⋅ Songbur Wong ⋅ Yanhao Liu ⋅ Wenlong Liao ⋅ Tao He ⋅ Junchi Yan
Modern LLMs/VLMs have largely converged to a unified architecture: representing heterogeneous inputs as a single token sequence and processing it with a deep stack of generic operators such as attention and feed-forward layers. This sequence-first interface simplifies system designs and thus enables extensive algorithmic and hardware optimization, making it a compelling blueprint for modern model development. In contrast, 3D multi-modal fusion models still rely on complex and specialized designs such as dense feature representations, sparse 3D convolutions, deformable attention, view transformations, etc. In this paper, we ask: \textbf{Can 3D multi-modal fusion embrace the same design path as modern LLMs/VLMs and benefit from their mature and rapidly evolving ecosystem?} We answer this affirmatively and propose \textbf{FusionNeXt}, a modern multi-modal fusion paradigm that (i) unifies camera and LiDAR features into a shared token representation, (ii) serializes tokens into locality-preserving 1D sequences, and (iii) performs feature fusion using a deep stack of vanilla LLM/VLM-style blocks---FlashAttention, pre-norm residuals, SwiGLU FFNs, etc. The proposed paradigm enables standard sequence modeling tools to be applied directly to 3D multi-modal fusion, resulting in fast inference, strong performance, and generalization across tasks. FusionNeXt achieves state-of-the-art (SOTA) results on the nuScenes 3D detection and Occ3D occupancy benchmarks while delivering high inference throughput, suggesting a scalable direction for 3D perception. We will open-source our code.
GAE Falls Short in Imperfect-Information Self-Play Reinforcement Learning
Zhiyuan Fan ⋅ Gabriele Farina
Competitive multi-agent reinforcement learning in imperfect-information games requires agents to act under partial observability and against adversarial opponents, necessitating stochastic policies. While self-play reinforcement learning with Proximal Policy Optimization (PPO) has achieved strong empirical success, its standard advantage estimator, generalized advantage estimation, suffers from additional variance due to the sampling of stochastic future actions. This variance is amplified in equilibrium self-play because of the stochastic nature of the equilibrium policy and persists even when the critic is exact. We address this bottleneck by introducing $Q$-boosting, a variance-reduced advantage estimator based on a centralized action-value critic, and propose Variance-Reduced Policy Optimization (VRPO), incorporating this new estimator. The algorithm replaces sampled multi-step backups with a multi-step Expected SARSA$(\lambda)$ trace, computing policy expectations at each step to average out action-sampling noise, while retaining PPO's clipped objective and on-policy actor updates. Empirically, VRPO consistently achieves strong performance from mid-sized to large-scale games including Dou Dizhu and Heads-Up No-Limit Texas Hold’em.
GAPS: Gradient-Aware Adaptation-Gap Scoring for Time-Series Anomaly Detection with Foundation Models
Jongwon Kim ⋅ Byunghun Song ⋅ Young Myoung Ko
Reconstruction error is the standard signal for unsupervised time-series anomaly detection and has been widely adopted by time-series foundation models (TFMs). It is, however, vulnerable to input noise: the resulting score tracks the noise level itself rather than whether a sample is truly anomalous, and this failure is especially pronounced in domains where the noise envelope itself carries information about the underlying process state. We propose the *adaptation gap* $\text{Diff} = L_{FT} - L_{ZS}$, the difference between fine-tuned and zero-shot reconstruction losses, as a noise-robust complement. We prove that its conditional expectation cancels the input-noise variance and yields a calibration-anchored, population-level mean-separation guarantee. Building on these results we introduce **GAPS** (Gradient-Aware Adaptation-Gap Scoring), which selectively augments $L_{FT}$ with $\text{Diff}$ under a calibration-only routing layer with no label-tuned hyperparameters. Across a synthetic suite, an HRV benchmark whose normal class is intrinsically noisier than its anomaly class, and TSB-AD-U with a pre-registered noise-stress condition, GAPS preserves performance under standard conditions and yields substantial gains under noise stress.
Gauge-Symmetric Dual Lagrangian Frameworks for Born-Oppenheimer Molecular Dynamics
Sungwoo Park ⋅ Jongwon Lee ⋅ Jiwoong Kim ⋅ Hyung-sik Yoon
Born-Oppenheimer molecular dynamics delivers quantum accurate trajectories by evolving nuclei on the electronic ground-state surface while retaining electronic observables such as dipoles and polarizabilities. In Kohn-Sham density functional theory, this typically relies on an iterative self consistent field solver at each time step. Machine learning can predict electronic quantities or improve SCF initialization, yet dynamics remain sensitive to residual self consistency error, which can yield non conservative forces and energy drift, and repeated diagonalization limits differentiable and accelerator efficient implementations. We propose a gauge-symmetric dual Lagrangian framework that avoids SCF loops by propagating the physically meaningful electronic state as the occupied subspace and its density projector on a Grassmann manifold. A residual driven update stabilized by Rayleigh type dissipation with time decaying inertia yields a closed evolution of the subspace, its conjugate momentum, and the projector. Coupled with a neural gauge-symmetric Hamiltonian with analytic coordinate derivatives, the method provides closed form forces with Pulay corrections and enables stable quantum BOMD on molecular simulation benchmarks.
This paper proposes a novel crowd counting approach, the Gaussian Density Splatting Network (GDSNet). Unlike methods that rely on conventional, grid-based density maps and are sensitive to spatial resolution, GDSNet represents a crowd as a superposition of continuous 2D Gaussian primitives. Our approach is built upon two key contributions. First, we introduce a control-point-based fitting mechanism to structure the prediction of the Gaussian parameters. We design a method to allocate a set of control points that define local regions, from which features are pooled to regress each primitive's parameters. Second, we adapt a differentiable Gaussian Splatting framework to the counting task by parameterizing each primitive with geometric parameters and a scalar density mass. This formulation allows the network to be trained end-to-end via spatial matching of differentiably rendered density maps, naturally providing both local density supervision and global count optimization. Extensive evaluations on four standard benchmarks show GDSNet consistently outperforms the state of the art. Code will be released upon acceptance.
Gaussian Splatting-based Volumetric Video Compression with Sparse 4D Anchors
Ge Gao ⋅ Siyue Teng ⋅ Changqi Wang ⋅ Fan Zhang ⋅ Nantheera Anantrasirichai ⋅ Jui-Chiu Chiang ⋅ Wen-Hsiao Peng ⋅ David Bull
Immersive video communication requires photorealistic, render-efficient, and compact dynamic scene representations. 3D Gaussian Splatting (3DGS) offers a promising representation, but dynamic 3DGS remains difficult to compress due to dense primitives and spatiotemporal redundancy. Anchor-based formulations improve compactness with sparse scaffolds that share geometry and appearance across primitives. However, existing designs often rely on deforming a single canonical scaffold and condition each primitive on its associated anchor in isolation, limiting their ability to handle non-local dynamics and disocclusion while under-exploiting inter-anchor correlations, particularly in motion- or texture-dense regions. To address these limitations, we propose SAGA, a volumetric video codec built upon Sparse Anchor-assisted GAussian splatting representations. SAGA represents dynamic 3D scenes using hierarchically organized sparse 4D anchors, where coordinate-based INR decoders generate fine anchors and Gaussian primitives from inter-anchor interpolations, enabling compact parameter sharing across spatiotemporal structures. For long-range dependencies among unstructured anchors, we further introduce fixed-size memory slots with orthogonality-informed updates for accurate entropy-context modeling. Experiments show SAGA outperforms the state-of-the-art 3DGS-based codec GIFStream, achieving BD-rate reductions of 84.54% on Neu3D, 76.71% on Panoptic Sports, and 83.94% on MPEG MIV.
Training modern neural networks often relies on large learning rates, operating at the edge of stability, where the optimization dynamics exhibit oscillatory and chaotic behaviour. Empirically, this regime often yields improved generalization performance, yet the underlying mechanism remains poorly understood. In this work, we represent stochastic optimizers as random dynamical systems, which often converge to a fractal attractor set (rather than a point) with a smaller intrinsic dimension. Building on this connection and inspired by Lyapunov dimension theory, we introduce a novel notion of dimension, coined the `sharpness dimension', and prove a generalization bound based on this dimension. Our results show that generalization in the chaotic regime depends on the complete Hessian spectrum and the structure of its partial determinants, highlighting a complexity that cannot be captured by the trace or spectral norm considered in prior work. Experiments across various MLPs and transformers validate our theory while also providing new insights into the recently observed phenomenon of grokking.
Empirical evidence suggests that large neural networks rarely make effective use of all available parameters, with learned solutions exhibiting pronounced sparsity. Despite this ubiquity, existing generalization theory only partially explains how such sparsity influences statistical performance. In this paper, we derive non-asymptotic excess risk bounds for deep neural networks whose layer-wise weight matrices are restricted by $\ell_p$ quasi-norm ($0
We study high-dimensional observations $Z_i \in \mathbb{R}^p$ satisfying that $E[Z_i] = X(t_i)$, where $X(t)$ parameterizes a manifold on $0 \le t < 2\pi$, with the goal of recovering the latent phases $t_i$; a motivating example is estimating object-viewing angles from photographs. Spectral seriation is a standard phase-recovery method: given a similarity matrix $S$ with degree matrix $D$, it uses the leading eigenvectors of the normalized Laplacian $D^{-1/2} S D^{-1/2}$, whose continuum limit is linked to the manifold Laplacian, to estimate $t_i$. Recent applications show that generalized Laplacians involving $D^{-\alpha}$ can outperform the normalized Laplacian on graphs with heterogeneous degrees, raising the central question of whether the same advantage holds for spectral seriation. We address such question by introducing and comparing three principled generalizations drawn from manifold learning and spectral seriation: a shifted generalized Laplacian, a graph-difference Laplacian, and an extra-normalized generalized Laplacian. The extra-normalized generalized Laplacian is the most robust to $\alpha$ and achieves the highest recovery accuracy, identifying extra normalization as the key stabilizing step. We establish a consistency result unveiling that eigenspace perturbation controls the estimation error up to rotation and reflection, giving a theoretical basis for the method. Experiments on noisy closed curves confirm more accurate phase recovery under a suggested choice of $\alpha$; on image data, the method recovers photo angles with the best accuracy. Our results demonstrate that extra normalization is a key component for stable generalized spectral seriation in closed-curve recovery problems.
Shapley value and its priority-aware extensions are widely used for valuation in machine learning, but existing methods require pairwise priority to be binary and acyclic, a restriction spectacularly violated in real-data examples such as aggregated human preferences and multi-criterion comparisons. We introduce the generalized priority-aware Shapley value (GPASV), a random order value defined on arbitrary directed weighted priority graphs, in which pairwise edges penalize rather than forbid order violations. GPASV covers a range of classical models as boundary cases. We establish GPASV through an axiomatic characterization, develop the associated computational methods, and introduce a priority sweeping diagnostic extending PASV's. We apply GPASV to LLM ensemble valuation on the cyclic Chatbot Arena preference graph, illustrating that priority-aware valuation is not a one-button operation: different balances of pairwise graph priority versus individual soft priority produce substantively different valuations of the same data.
Generating the Wild: Individual-Consistent Image-to-Video Generation for Wildlife
Yuzhuo Li ⋅ Di Zhao ⋅ Xinyu Zhang ⋅ Daniel Wilson ⋅ Yun Sing Koh
Individual-level wildlife identification often suffers from data scarcity, as varying observations of the same animal under diverse poses, viewpoints, and motions are rarely available. Image-to-video (I2V) generation offers a promising way to mitigate this limitation by synthesizing additional observations from a single reference image. However, existing I2V models mainly emphasize global layout, semantics, and motion, and therefore often fail to preserve fine-grained local appearance cues that distinguish one wildlife individual from another, such as fur texture, stripe boundaries, spot configurations, and contour transitions. We observe that these identity-critical cues are closely related to high-frequency information, which is essential for distinguishing individuals. To address this challenge, we propose WildIcon, a high-frequency-guided I2V framework for wildlife individual consistency. Specifically, WildIcon introduces a frequency-aware identity encoding branch that extracts individual-specific high-frequency cues from the reference image. Combined with isolated foreground information, the resulting identity tokens are then injected into cross-attention blocks as identity conditioning. Building on a frozen backbone with lightweight identity adaptation, WildIcon strengthens fine-grained identity preservation while retaining the motion controllability and semantic fidelity of the base I2V model. In addition, to support the training and evaluation of wildlife individual-consistent I2V, we construct WildlifeVid, a wildlife-centric video dataset with high-quality, temporally coherent clips and individual-level identity labels. Experiments on both I2V and downstream animal re-identification (ReID) show that WildIcon achieves stronger individual consistency than existing baselines and provides useful, reliable synthetic data to improve downstream ReID performance.
Generative Active Learning via Bayesian Acquisition for Improving the Efficiency of Synthetic Data
Jeongjun Lee ⋅ Gyeonghoon Ko ⋅ Juho Lee
Unlike real-world datasets, utilizing synthetic data from generative models as train dataset necessitates a rigorous consideration of the training data's informativeness. In this paper, we propose a novel framework designed to steer diffusion models based on Bayesian epistemic uncertainty to train the target model. To ensure computational tractability, we estimate the Bayesian Active Learning by Disagreement (BALD) score by applying the Laplace approximation specifically to the last-layer parameters of the downstream model. Furthermore, we incorporate Feynman-Kac Steering (FKS) for diffusion to enhance the structural diversity of the synthetic samples and mitigate the emergence of undesirable artifacts. Extensive experiments on Imagenette and ImageNet-100 for the classification task and MS-COCO for the object detection task demonstrate that our method significantly outperforms existing baselines in downstream accuracy and OOD generalization. Furthermore, our analysis of pairwise distances between inter-class and intra-class confirms that our method effectively expands the downstream model's knowledge manifold by generating diverse and informative training signals.
GeoMemory: Geometry-Indexed Memory for Long-Horizon Interactive Video Generation
Junchao Huang ⋅ Xinting Hu ⋅ Boyao Han ⋅ Shaoshuai Shi ⋅ Zhuotao Tian ⋅ Tianyu He ⋅ Li Jiang
Autoregressive video diffusion models have demonstrated effectiveness in interactive game generation, with Minecraft gameplay serving as a representative application. To faithfully simulate gameplay, a model must generate natural content when exploring new scenes while preserving spatial consistency when revisiting explored areas. Under limited computation budgets, it must compress and exploit historical cues within a finite context window, which exposes a trade-off: models relying solely on temporal context offer flexible exploration but suffer from poor revisit consistency, whereas adding spatial memory strengthens consistency but may degrade new scene generation quality when the model over-relies on sparse or unreliable spatial context. We present GeoMemory, a learning framework that pairs training protocols with a geometry-indexed spatial memory. Specifically, our Hybrid Training exposes the model to both exploration and revisitation regimes, guiding the model to rely on temporal memory in new scenes while effectively incorporating spatial memory upon revisits. Chained Forward Training creates larger pose variations and encourages reliance on spatial memory for maintaining consistency. For spatial memory, we integrate Point-to-Frame Retrieval with an incremental 3D cache generated by VGGT, enabling constant-time retrieval of relevant historical context regardless of sequence length. Extensive experiments demonstrate that GeoMemory achieves superior performance in both long-term spatial consistency and visual quality in new scenes with real-time interaction.
Geometric Instability of Hidden-State Trajectories Predicts Reasoning Failures in Large Language Models
Hoda Fakharzadehjahromy ⋅ Andreas C Bueff ⋅ Fredrik Heintz ⋅ Mattias Tiger ⋅ Qing Li ⋅ Jiaqi Li ⋅ Jiahui Geng
Large language models often produce fluent reasoning chains that arrive at incorrect answers, and these failures are hard to detect from output probabilities alone. We show that a strong signal of correctness is reflected in the geometry of internal computation: hidden-state trajectories across transformer layers differ between correct and incorrect reasoning, with correct solutions following smooth, efficient paths and failures exhibiting elevated curvature, abrupt directional changes, and increased geodesic deviation. We capture this signal through VANE, five parameter-free geometric features (velocity, acceleration/jerk, curvature, geodesic deviation, and token coherence) computed from a single forward pass over layer-wise hidden states. Across six models (1.5B to 72B parameters) and five benchmarks spanning mathematics, code, and verbal reasoning, VANE achieves 70.7 to 96.6 AUROC and exceeds token log-probability on all 30 model–benchmark pairs (mean +25 percentage points, up to +50 percentage points). The geometric signal is near-independent of output confidence (r = -0.26), and a base-model control shows it emerges with instruction tuning. As a downstream application, filtering geometrically unstable outputs raises accuracy to 96.1% at 50% coverage on a 4-bit quantized 72B model. These results show trajectory geometry is a reliable single-pass indicator of reasoning correctness, even when the model is confidently wrong.
Geometry-Aware Flow Matching for Sparse-View 3D Gaussian Splatting
Abdullah Azeem ⋅ Ruisheng Wang ⋅ Qingquan Li ⋅ Abubakar Siddique
Generalizable 3D Gaussian Splatting aims to predict renderable Gaussian scenes from sparse images without per-scene optimization at test time. Existing feed-forward methods improve this problem mainly by strengthening the evidence given to the predictor, such as multiview aggregation, geometric cues, auxiliary supervision, and surface priors. Yet the prediction itself is usually learned as a direct map to the final Gaussian scene. This endpoint-centered objective compresses scene layout, local geometry, visibility, opacity, orientation, and appearance into a single end-point target, leaving the construction path of the Gaussian scene implicit. We propose GRiF, a geometry-aware flow matching framework that augments endpoint target supervision with conditional transport in scene space. Starting from a controlled source scene, GRiF learns a view-conditioned velocity field that progressively evolves a hierarchical Gaussian state toward a renderable target. The hierarchy organizes updates from coarse layout to fine attributes, while geometry-aware paths respect the native structure of Gaussian variables with rotation-aware paths for orientation variables. Experiments on RealEstate10K and ACID, with zero-shot transfer evaluated on ScanNet, DL3DV, and DTU, show improved in-domain reconstruction and stronger cross-domain performance.
Geometry-Aware Online Scheduling for LLM Serving: From Theoretical Bound to System Practice
Li Kong ⋅ Qi Qi ⋅ Yinyu Ye ⋅ Zijie Zhou
The explosive demand for interactive Large Language Model serving has highlighted the management of the Key-Value cache's dynamic memory footprint as a critical area for performance optimization in inference engines. Modern inference systems overwhelmingly rely on time-centric scheduling heuristics, such as Shortest Job First. However, their theoretical optimality is rooted in traditional schedule modeling, failing to capture the highly dynamic, 2D spatio-temporal geometric growth specific to LLM inference mechanisms. To resolve this, we propose the geometry-aware online scheduling by introducing the Smallest Volume First (SVF) algorithm and its highly efficient variant, 1-bit SVF. Theoretically, we provide a rigorous mathematical foundation for our approach. Utilizing a novel proof methodology, we tighten the worst-case competitive ratio ($\text{CR} \le 48 \rightarrow \text{CR} \le 5$) for SVF with known output lengths. Building upon this core breakthrough, we complete a comprehensive theoretical taxonomy analyzing our algorithms across different traffic scenarios and information availability. Practically, we seamlessly integrate our approach as a plug-and-play layer in vLLM. Extensive evaluations on Llama-3.1 models demonstrate comprehensive performance gains: SVF delivers strong reductions in both average and tail latency, while 1-bit SVF, with merely a single bit information, achieves competitive throughput and latency. This work establishes a theoretically sound and empirically proven approach for resolving memory-constrained scheduling in modern LLM deployments. To facilitate future research, our code is available at \url{https://anonymous.4open.science/r/LLM-inference-scheduling-7970}.
Geometry-Aware Similarity Metrics for Neural Representations on Riemannian and Statistical Manifolds
N Alex Cayco Gajic ⋅ Arthur Pellegrino
Similarity measures such as canonical correlation analysis (CCA), representational similarity analysis (RSA) and center kernel alignment (CKA), are widely used to compare the representational geometries used by different neural networks to solve the same task. Yet, because existing methods compare the extrinsic geometry of neural representations in state space, rather than their intrinsic geometry, they may fail to capture subtle yet crucial distinctions between fundamentally different computations performed by biological and artificial neural networks. Here, we introduce metric similarity analysis (MSA), a novel method which leverages tools from Riemannian geometry to compare the intrinsic geometry of neural manifolds. Within our mathematical framework, we derive several properties of MSA, including coordinate, scale and rotation invariances. We show that MSA can be used to i) disentangle neural computations of deep networks in rich vs. lazy learning regimes, ii) compare the computation-through-dynamics of recurrent neural networks and state-space models during neuroscience working memory tasks, and iii) investigate the statistical manifolds of diffusion models. Overall, we introduce a mathematically grounded and broadly applicable framework to understand the mechanisms behind neural computations by contrasting their Riemannian geometries, with broad applicability to both neuroscience and machine learning.
Geometry-Calibrated Conformal Abstention for Language Models
Rui Xu ⋅ Yi Chen ⋅ Sihong Xie ⋅ Hui Xiong
When language models lack relevant knowledge for a given query, they frequently generate plausible responses that can be hallucinations, rather than admitting being agnostic about the answer. Retraining models to reward admitting ignorance can lead to overly conservative behaviors and poor generalization due to scarce evaluation benchmarks. We propose a post-hoc framework, Conformal Abstention (CA), adapted from conformal prediction (CP) to determine whether to abstain from answering a query. CA provides finite-sample guarantees on both the probability of participation (i.e., not abstaining) and the probability that the generated response is correct. Importantly, the abstention decision relies on prediction confidence rather than the non-conformity scores used in CP, which are intractable for open-ended generation. To better align prediction confidence with the model's ignorance, we introduce a calibration strategy using representation geometry within the model to measure knowledge involvement in shaping the response. Experiments demonstrate that we improve selective answering significantly with 75% conditional correctness.
Geometry Conflict: Explaining and Controlling Forgetting in LLM Continual Post-Training
Yuanyi Wang ⋅ Yifan Yang ⋅ Su Lu ⋅ Yanggan Gu ⋅ Pengkai Wang ⋅ Wenjun Wang ⋅ Zhaoyi Yan ⋅ Congkai Xie ⋅ Jianmin Wu ⋅ Jialun Cao ⋅ Shing-Chi Cheung ⋅ Hongxia Yang
Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential updates enable capability transfer and when they cause catastrophic forgetting. Existing methods mitigate forgetting through sequential fine-tuning, replay, regularization, or model merging, but offer limited criteria for determining when incorporating new updates is beneficial or harmful. In this work, we study LLM continual post-training through three questions: What drives forgetting? When do sequentially acquired capabilities transfer or interfere? How can compatibility be used to control update integration? We address these questions through task geometry: we represent each post-training task by its parameter update and study the covariance geometry induced by the update. Our central finding is that: forgetting can be considered as a state-relative update-integration failure, it arises when the covariance geometries induced by tasks misalign with the geometry of the evolving model state. Sequential updates transfer when they remain compatible with the model state shaped by previous updates, and interfere when state-relative geometry conflict becomes high. Motivated by this finding, we propose Geometry-Conflict Wasserstein Merging (GCWM), a data-free update-integration method that constructs a shared Wasserstein metric via Gaussian Wasserstein barycenters and uses geometry conflict to gate geometry-aware correction. Across Qwen3 0.6B--14B on domain-continual and capability-continual settings, GCWM consistently outperforms data-free baselines, improving retention and final performance without replay data. These results identify geometry conflict as both an explanatory signal for forgetting and a practical control signal for LLM continual post-training.
GeoPMR: Preserving Relational and Hierarchical Geometry in Multimodal Molecular Representation Learning
Lusheng Li ⋅ Menglin Yang
Multimodal molecular representation learning aligns heterogeneous chemical and cellular signals to predict molecular properties. The dominant paradigm is instance-level contrastive alignment, such as InfoNCE and VICReg, which implicitly treats every modality as a point cloud in a shared Euclidean space. This paradigm overlooks two structural properties of molecular-cellular data. \textit{First}, each modality, from atomic fingerprints to cell morphology to transcriptomic response, carries its own intrinsic geometry; directly pulling individual embeddings together can distort the relational structure each modality has learned to express. \textit{Second}, the modalities form a directed biological cascade, from molecular structure to conformation, gene perturbation, cellular phenotype, and transcriptomic response, whose branching grows exponentially with depth, a regime that Euclidean spaces cannot embed without distortion. We therefore propose \textbf{GeoPMR}, a pre-training framework that addresses both properties through two complementary geometric modules: Metric-Guided Gromov-Wasserstein alignment (MGW) and a Hyperbolic Hierarchical Module (HHM). \textbf{MGW} equips each modality with a learnable diagonal Riemannian metric and aligns distance matrices rather than individual embeddings via entropic optimal transport, so cross-modal alignment preserves relational rather than merely pointwise structure. \textbf{HHM} lifts all modality latents onto the Poincaré ball and applies a probabilistic parent-child and sibling contrastive loss along the cascade, exploiting the exponential volume growth of negatively curved space to embed the hierarchy with low distortion. On seven molecular property prediction benchmarks spanning toxicity, bioactivity, and ADME, GeoPMR sets a new state of the art by improving the classification AUC by 1.4\% and regression MAE by 1.7\%.
GeoRouter: Dynamic Paradigm Routing for Worldwide Image Geolocalization
Pengyue Jia ⋅ Derong Xu ⋅ Yingyi Zhang ⋅ Xiaopeng Li ⋅ Wenlin Zhang ⋅ Yi Wen ⋅ Liu lan ⋅ Yuanshao Zhu ⋅ Xuetao Wei ⋅ Wanyu Wang ⋅ Xiangyu Zhao
Worldwide image geolocalization aims to predict precise GPS coordinates for images captured anywhere on Earth, which is challenging due to the large visual and geographic diversity. Recent methods mainly follow two paradigms: retrieval-based approaches that match queries against a reference database, and generation-based approaches that directly predict coordinates using Large Vision-Language Models (LVLMs). However, we observe distinct error profiles between them: retrieval excels at fine-grained instance matching, while generation offers robust semantic reasoning. This complementary heterogeneity suggests that no single paradigm is universally superior. To harness this potential, we propose GeoRouter, a dynamic routing framework that adaptively assigns each query to the optimal paradigm. GeoRouter leverages an LVLM backbone to analyze visual content and provide routing decisions. To optimize GeoRouter, we introduce a distance-aware preference objective that converts the distance gap between paradigms into a continuous supervision signal, explicitly reflecting relative performance differences. Furthermore, we construct GeoRouting, the first large-scale dataset tailored for training routing policies with independent paradigm predictions. Extensive experiments on IM2GPS3K and YFCC4K demonstrate that GeoRouter significantly outperforms state-of-the-art baselines. Our dataset, code, and checkpoints are publicly available to facilitate future research.
GIVLA: Deep Geometry Internalization for A Lightweight VLA via Geometry Instruction and Gradient-Informed Training
YUNHE LI ⋅ Qiming Liu ⋅ Haoyuan Wang ⋅ Hesheng Wang
Vision-Language-Action (VLA) models often struggle with precise manipulation due to their lack of explicit spatial awareness. To bridge this gap, we propose GIVLA, a framework that internalizes geometric priors into the VLA backbone through a deep-coupled architecture. GIVLA implements a two-stage gradient-informed training paradigm to resolve task-level interference and ensure precise execution, a design grounded in the perception-action duality identified through our analysis of training dynamics. Extensive experiments on LIBERO and physical robot platforms demonstrate that GIVLA achieves superior accuracy and robustness with high parameter efficiency.
GLARE: Generating Listening Heads with Appropriate REactions
Zikai Liao ⋅ Yumin Suh ⋅ Yi Ouyang ⋅ Yi-Lun Lee ⋅ Yi-Hsuan Tsai ⋅ Zhaozheng Yin
While talking head generation has advanced rapidly, generating natural listener behavior in dyadic conversations, which know when to react, how to react, and with what type of response, remains underexplored. Existing dyadic datasets lack fine-grained listener reaction annotations, and prevailing evaluation metrics inherited from talking-head and video generation measure visual realism rather than whether a listener reacted appropriately. We address these gaps along three aspects. First, we curate a listening-head-specific dataset built from RealTalk and Seamless Interaction, comprising approximately 147 hours of paired speaker–listener videos with 64,557 event-level reaction annotations across six categories: nodding, head shaking, smiling, laughing, frowning, and surprised. Second, we introduce an audio-driven baseline built on a flow-matching transformer, namely GLARE, with prosody conditioning derived from Qwen2-Audio and a temporal reaction loss that explicitly supervises frame-wise reactions. Third, we propose a reaction-oriented evaluation protocol that jointly measures reaction occurrence (R-F1), temporal alignment (R-tIoU), asymmetric temporal deviation (R-ATD), and reaction-region visual quality (R-FID), giving a more behaviorally grounded assessment than visual-quality-only metrics. Experiment results show consistent gains over prior listening-head methods in both visual fidelity and reaction-level metrics, suggesting that reaction-aware data, modeling, and evaluation are critical for natural listening behavior.
Going Down Memory Lane: Scaling Tokens for Video Stream Understanding with Dynamic KV-Cache Memory
Vatsal Agarwal ⋅ Saksham Suri ⋅ Matthew Gwilliam ⋅ Pulkit Kumar ⋅ Abhinav Shrivastava
Online video understanding requires models to robustly encode, store, and retrieve information from a continuous stream to support accurate question answering. Existing approaches rely on key--value caching to accumulate frame-level representations over time, but allocate only a limited token budget per frame, causing loss of fine-grained visual detail. We find that naively scaling this budget degrades retrieval quality and question-answering performance, as denser token streams introduce redundancy that disrupts query--frame similarity. To address this, we first introduce an adaptive selection strategy that reduces token redundancy while preserving local spatiotemporal information. Second, we propose a training-free retrieval ensemble that leverages external models to better identify relevant frames. Our method, \textbf{MemStream}, achieves +8.0\% on CG-Bench, +8.5\% on LVBench, and +2.4\% on VideoMME (Long) over ReKV with Qwen2.5-VL-7B. Furthermore, we show that MemStream consistently improves performance across several Video-LLMs for long-video understanding.
Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
Shuhong Zheng ⋅ Michael Oechsle ⋅ Erik Sandström ⋅ Marie-Julie Rakotosaona ⋅ Federico Tombari ⋅ Igor Gilitschenski
Visual geometry transformers have become powerful architectures for multi-view 3D reconstruction, enabling joint prediction of multiple 3D attributes in a feed-forward manner. However, their computational cost grows quadratically with the input sequence length due to the global attention layers inside these models. This limits both their scalability and efficiency. In this work, we address this challenge with a simple yet general strategy: restricting the number of key/value tokens that each query interacts with during global attention. To achieve effective token selection, we introduce a two-stage framework. First, an inter-frame selection step operates at the frame level to identify frames that should be preserved. Second, an intra-frame selection step further discards more redundant tokens within the selected frames. Our analysis highlights the advantage of a diversity-based strategy for inter-frame selection, which ensures broad coverage of the scene. For intra-frame selection, we show that layer-aware sparsification is necessary, with the selection process guided by the entropy of the global attention pattern. Our approach offers a superior speed-accuracy trade-off compared to existing solutions. Extensive experiments show that it accelerates visual geometry transformers by over 85% for scenes with 500 images while maintaining, or even improving, baseline performance, which hints that how our token selection strategy can play a crucial role in future applications of visual geometry transformers.
Gradient Boosted Trees for Retrieval-Augmented Generation
Huaiyu Qin ⋅ Chunyu Wei ⋅ Yueguo Chen ⋅ Yunhai Wang
Retrieval-augmented generation (RAG) divides into two camps failing in complementary ways on broad, long-tailed questions: iterative agentic RAG re-queries a biased distribution and drifts on prominent aspects, while structural RAG commits to a hierarchy fixed upfront and cannot admit aspects it initially missed. We propose GBT-RAG, which unifies the two by transposing gradient boosting into the text domain: each round is a query decomposition tree fitted to its predecessors' residual gap. This raises two challenges mirroring the pillars of boosting---natural-language critiques are not quantifiable, and successive rounds drift rather than fit the residual. We address the first with a text gradient, a structured residual in the retriever's embedding space whose support localizes missingness to specific leaves; we address the second by growing each tree only over under-covered leaves and admitting each rewrite through a monotone-improvement gate. Across standard benchmarks, GBT-RAG consistently outperforms the strongest iterative and structural baselines, with the largest gains on the long-tailed regime that motivated the design.
We introduce Gram-Calibrated Anchoring (GCA) for class-incremental learning (CIL) with pre-trained Vision Transformers. GCA rests on the observation that, under continual low-rank adaptation, anchor prototypes---class means extracted once from the frozen pre-trained backbone---are effective substitutes for prototypes recomputed from the adapted model; a formal stability analysis bounds decision-boundary shifts in terms of feature-level prototype drift, while empirical LoRA perturbation measurements support the small-drift regime. This motivates fixing the classification head to anchor prototypes computed upon each task's arrival rather than recalibrating prototypes after training. The fixed anchor geometry then enables Gram-Compensated Inference (GCI), which applies the regularized inverse of the anchor Gram matrix to deconvolve inter-class correlation that standard cosine classification ignores, and normalizes each coordinate by its estimation standard deviation for fair cross-class comparison. The resulting method trains a single continually evolved LoRA module with no exemplar storage or expanding adapter pools. Experiments on ImageNet-R, ImageNet-A, CIFAR-100, and CUB-200 with multiple backbones show that GCA achieves competitive accuracy across four CIL benchmarks while significantly reducing training cost over recent methods.
GramStatTexNet: Efficient, Interpretable, and Neuro-Inspired Texture Analysis-by-Synthesis
Vasha DuTell ⋅ Zeyu Yun ⋅ Anne Harrington ⋅ Yuan Sheng Fang ⋅ Mark Hamilton ⋅ Christian Koevesdi ⋅ Edward Adelson ⋅ Bill Freeman ⋅ Ruth Rosenholtz
Deep networks achieve high realism in texture synthesis but often rely on opaque, over-parameterized representations that have limited direct connection to biological vision. Neuroscience-informed pyramid models are interpretable, compact, and biologically grounded, but rely on small, hand-curated statistics sets that are difficult to generalize or extend. We introduce GramStatTexNet, a hybrid analysis-by-synthesis framework that combines the multi-scale Gabor filter structure of classical models with the flexibility of Gramian-based correlations, organized into structured filter families. This structured representation enables analyses that are challenging for deep-feature methods: each statistic carries a named filter-family identity, allowing us to decompose synthesis fidelity into per-family contributions and to learn an interpretable, low-dimensional embedding that preserves perceptual texture content. We compare our framework to classical and deep texture-synthesis baselines across diverse texture categories, and on multiple perceptual metrics, highlighting the compactness and interpretability advantages of this hybrid representation, alongside competitive synthesis quality. This same framework extends naturally to peripheral textures via spatial pooling, and to dynamic textures using a spatiotemporal filter bank. GramStatTexNet provides a unified, interpretable framework for analyzing and modeling visual information with natural extensions across space, time, and eccentricity.
Graph-Based Stochastic-Power-UCT: Monte-Carlo Graph Search with Power Mean Estimation
Tung Tran ⋅ Viet Bao ⋅ Hoang Ta ⋅ Tuan Dam
Tree-based Monte-Carlo Tree Search (MCTS) duplicates the same state when it is reached through different trajectories, which can waste simulations in stochastic MDPs. We introduce Graph-Based Stochastic-Power-UCT (GS-Power-UCT), which shares states reached at the same planning depth while keeping separate values for states reached at different depths. This design applies to general stochastic MDPs, including problems with cycles. We prove that for a fixed planning horizon, the root estimate converges to the finite-horizon value at rate $O(n^{-1/2})$, matching tree-based Stochastic-Power-UCT while reusing samples across shared states. We also study two full-state variants: GS-Power-UCT-F, which stores one node per physical state to increase sample sharing but may mix values from different remaining horizons, and GS-Power-UCT-F$^+$, which uses an adaptive horizon to control this bias. The latter converges to $V^{\star}(s_0)$, the optimal infinite-horizon discounted value at the root state $s_0$, when the remaining cross-depth gap vanishes. Experiments on stochastic planning benchmarks show improved sample efficiency over tree-based and graph-based baselines.
GraphMemRL: Action-Native Reinforcement Learning for Persistent Graph Memory Construction
Bingcheng Dong ⋅ Shenglan Liu ⋅ Sifan Zhang ⋅ Jirui Tian ⋅ Weitong Wu
Long-term memory construction for LLM agents is hindered by two mismatches. The first is a retrieval-objective mismatch: existing systems usually retrieve memories for writing with a single holistic query derived from the current input. However, memory construction depends on several complementary aspects of the input, and a single query is insufficient to capture these editing needs. The second is an optimization-granularity mismatch: long-horizon memory learning is commonly optimized over full trajectories, whereas memory construction proceeds through incremental state transitions. Effective supervision should therefore assess how each local edit changes the memory state. We propose GraphMemRL, a graph-based memory construction framework that addresses both mismatches. GraphMemRL introduces edit-oriented query decomposition to retrieve complementary context for memory writing. It also uses action-native state-transition supervision to directly evaluate whether each graph edit improves the evolving memory state. To make long-horizon learning more tractable, GraphMemRL reformulates memory construction as factorized segment-wise optimization. The agent updates a persistent memory graph across segments under limited context budgets, preserving globally accumulated structure without full-history training. Experiments on long-memory benchmarks show that GraphMemRL improves question answering and produces better connected memory structures with more controlled memory growth. The gains are especially strong on multi-hop, temporal, and knowledge-update tasks.
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
Junho Kim ⋅ Xu Cao ⋅ Houze Yang ⋅ Bikram Boote ⋅ Ana Jojic ⋅ Fiona Ryan ⋅ Bolin Lai ⋅ Sangmin Lee ⋅ James Rehg
Understanding social interactions requires reasoning over subtle non-verbal cues, yet current multimodal large language models (MLLMs) often fail to identify who interacts with whom in multi-person videos. We introduce GRASP, a large-scale social reasoning dataset that connects high-level social QA with fine-grained gaze and deictic gesture events. GRASP contains 290K question-answer pairs over 46K videos totaling 749 hours, organized by a 16-category taxonomy spanning gaze, gesture, and joint gaze-gesture reasoning, together with GRASP-Bench for evaluation. Unlike prior resources that focus on either isolated cues or high-level social QA, GRASP builds questions from identity-consistent gaze trajectories, deictic gestures, and their joint compositions into social events. Moreover, we propose Social Grounding Reward (SGR), a learning signal that uses these social events to encourage models to reason about the participants involved in each interaction. Experiments show that SGR improves performance on GRASP-Bench while maintaining zero-shot performance on related social video QA benchmarks.
GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation
Jette Oberländer ⋅ Jan Finkbeiner ⋅ Catherine Schoefmann ⋅ Emre Neftci
Autoregressive decoding with LLMs is primarily bottlenecked by GPU memory bandwidth, especially in edge-computing settings. While quantization is essential for mitigating this bottleneck, most existing methods treat inference as uniform process and fail to account for the asymmetry between the compute-bound prefill stage and the memory-bound decoding stage. We propose GRINQH (GRaded INput-based Quantization Hierarchy), a weight-only post-training quantization framework that accelerates decoding by unifying quantization and sparsification. GRINQH leverages activation magnitudes as a proxy for computational importance to dynamically assign weight channels to different precision levels, enabling flexible average bit widths during decoding. Evaluated on Llama3 and Qwen3 models, GRINQH outperforms state-of-the-art fixed- and mixed-precision baselines at comparable 3- and 4-bit settings, even enabling effective 2-bit generation. We experimentally verify theoretical speedups by leveraging a hierarchical nested memory layout for multi-precision storage in custom GPU kernels. Together, GRINQH establishes a new state-of-the-art Pareto frontier for LLM generation, enabling a dynamic trade-off between generation quality and inference speed.
Grounded or Fabricated? Unsupervised Detection of LLM Hallucinations via Contextualized Influence on Response Embeddings
Zeyang Ding ⋅ Feng Xiao ⋅ Jicong Fan
Hallucinations in large language models (LLMs) that are plausible-sounding but factually incorrect or unsupported pose a major challenge for deploying these models in high-stakes applications such as medical diagnosis, legal reasoning, and knowledge-based question answering. Existing detection methods primarily rely on either generating multiple responses to check consistency or supervised training on labeled hallucinations. Multiple-response approaches tend to be computationally expensive and may be less practical for real-time scenarios, while supervised methods are limited to known hallucination types and cannot generalize to unseen cases. To address these limitations, we propose an efficient unsupervised hallucination detection method. Our approach measures the embedding discrepancy between the LLM’s response considered independently and the response contextualized with the question. A large discrepancy indicates that the question significantly influences the response, suggesting it is grounded, whereas a small discrepancy signals a higher likelihood of hallucination. This method is efficient, interpretable, and does not require labeled hallucinations. Extensive experiments across multiple tasks demonstrate that our approach consistently achieves superior or comparable detection performance while reducing the detection cost compared to the multiple-response methods.
Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model
Hanqing Wang ⋅ Mingyu Liu ⋅ Xiaoyu Chen ⋅ Chengwei MA ⋅ Yiming Zhong ⋅ Yuhao Liu ⋅ Wenti Yin ⋅ Xing Mu ⋅ Jiahao Yuan ⋅ Zhiqing Cui ⋅ Lu Dai ⋅ Zhiyuan Ma ⋅ Hui Xiong
3D affordance grounding aims to highlight the actionable regions on 3D objects, which is crucial for embodied AI. While previous research primarily leverages static cues from language or single images, which often lack the rich interaction context necessary to accurately locate functional zones. To alleviate this predicament, we collect a comprehensive video-based 3D affordance dataset, VIDA, which contains 38K human-object-interaction (HOI) videos covering 16 affordance types, 38 object categories, and 22K point clouds. Based on VIDA, we propose a strong baseline: VideoAfford, a unified framework that extends Multimodal Large Language Models (MLLMs) with fine-grained affordance segmentation capabilities. VideoAfford incorporates a latent action encoder to distill dynamic interaction knowledge from demonstration videos and a spatial-aware loss to encourage geometrically consistent affordance predictions. Extensive experiments on VIDA show that VideoAfford significantly outperforms strong baselines in both seen and unseen settings, demonstrating its effectiveness in video-driven 3D affordance reasoning and open-world generalization. The dataset and code will be released soon to facilitate future research in this area.
Grounding Agent Reasoning with Structured Process Supervision for Multi-turn Reinforcement Learning
Renting Rui ⋅ Yulei Qin ⋅ Weiwen Liu ⋅ Yunjia Xi ⋅ Ke Li ⋅ Xing Sun ⋅ Weinan Zhang ⋅ Yong Yu
Long-horizon LLM agents trained with outcome-based RL exhibit a systematic failure mode: reasoning diverges from environmental observations at some step, and the error compounds across subsequent turns. We find that 74\% of failed trajectories contain reasoning–observation inconsistency, and once it appears, 62\% of subsequent turns remain inconsistent. Process rewards reduce but do not eliminate this: under free-form reasoning, agents learn to avoid penalties without genuinely tracking task progress or anticipating action outcomes. We propose ARC (Anchored Reasoning with Commitment), which requires the agent to make two explicit commitments at each step: a global progress statement that reconciles the agent's understanding of task state with accumulated observations, and a local outcome prediction that must be borne out by the next observation. Together, these commitments anchor the agent's reasoning to the actual environment state across turns, and provide well-defined targets for process supervision. To prevent interference between sparse outcome and dense process signals, we decouple their advantages via independent group normalization. On ALFWorld, WebShop, and WebArena, ARC improves over the strongest baseline by 6.4, 3.5, and 9.6 points respectively.
Group Distributionally Robust Optimization with Flexible Sample Queries
Haomin Bai ⋅ Dingzhi Yu ⋅ Shuai Li ⋅ Haipeng Luo ⋅ Lijun Zhang
Group distributionally robust optimization (GDRO) aims to learn models that perform well across $m$ distributions simultaneously. Existing methods either (i) query $m$ samples per round, which imposes strict implementation requirements, or (ii) query only a single sample per round, which requires many rounds to converge and therefore leads to long runtime. To avoid these limitations, we introduce a novel setting called \emph{flexible sample queries} (FSQ), which allows the number of samples that can be queried to vary across rounds. Under the FSQ setting, we cast GDRO as a two-player game: one performs non-oblivious online convex optimization, while the other tackles a non-oblivious prediction with limited advice (PLA) problem. In this game, the first player updates its decision using follow-the-regularized-leader with stochastic gradients. For the second player, we develop a new PLA method equipped with an adaptive exploration strategy that can leverage an arbitrary per-round sample size. We then establish the *first* high-probability regret bound for non-oblivious PLA. By integrating these components, we develop an anytime GDRO algorithm that supports FSQ, and derive a high-probability optimization error bound of $O\left(\frac{1}{t}\sqrt{\sum_{j=1}^t \frac{m}{r_j}\log m}\right)$, where $r_j$ is the sample size at round $j$. This result demonstrates that the optimization error decreases as the per-round sample sizes increase, and implies the same near-optimal sample complexity of $O(m\log (m)/\epsilon^2)$ for any fixed sample size $r\in[m]$.
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
Zheng Lin ⋅ Zhenxing Niu ⋅ haoxuan ji ⋅ Haichang Gao
This paper proposes a guaranteed defense method for large language models (LLMs) to safeguard against jailbreaking attacks. Drawing inspiration from the denoised-smoothing approach in the adversarial defense domain, we propose a novel smoothing-based defense method, termed Disrupt-and-Rectify Smoothing (DR-Smoothing). Specifically, we integrate a two-stage prompt processing scheme—first disrupting the input prompt, then rectifying it—into the conventional smoothing defense framework. This disrupt-and-rectify approach improves upon previous disrupt-only approaches by restoring out-of-distribution disrupted prompts to an in-distribution form, thereby reducing the risk of unpredictable LLM behavior. In addition, this two-stage scheme offers a distinct advantage in striking a balance between harmlessness and helpfulness in jailbreaking defense. Notably, we present a theoretical analysis for generic smoothing framework, offering a tight bound for the defense success probability and the requirements on the disruption strength. Our approach can defend against both token-level and prompt-level jailbreaking attacks, under both established and adaptive attacking scenarios. Extensive experiments demonstrate that our approach surpasses current state-of-the-art defense methods in terms of both harmlessness and helpfulness.
GWScore: A structural diversity metric for consistent text-to-image generation
Francis Snelgar ⋅ Stephen Gould ⋅ Liang Zheng ⋅ Akshay Asthana
Recent advances in visual storytelling have leveraged powerful text-to-image diffusion models to generate sequences of images featuring consistent subjects and styles. While such methods achieve impressive identity preservation across the sequence, they often produce scenes with limited structural diversity---images that are overly similar in subject pose and spatial layout. Existing evaluation protocols primarily measure identity consistency using metrics such as DreamSim, which can inadvertently conflate identity preservation with subject pose. The inability to independently evaluate structural diversity inhibits progress on identity preserving image generation methods. We address this limitation by introducing a new metric based on the Gromov Wasserstein distance to quantify structural diversity independently of identity. Our metric provides an interpretable and generalizable measure of spatial variation across generated images, enabling a more balanced assessment of identity consistency and pose diversity. Using this framework, we systematically evaluate several recent visual storytelling and subject-consistent generation methods. Furthermore, we evaluate the effect of latent initialization methods on structural diversity in several popular text-to-image models.
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale
Amit Roth ⋅ Ankur Samanta ⋅ Matan Halevy ⋅ Yoav Levine ⋅ Yonathan Efroni
Aligning autonomous agents with human intent remains a central challenge in modern AI. A key manifestation of this challenge is reward hacking, whereby agents appear successful under the evaluation signal while violating the intended objective. Reward hacking has been observed across a wide range of settings, yet methods for reliably measuring it at scale remain lacking. In this work, we introduce a new evaluation paradigm for measuring reward hacking. Whereas prior studies have primarily analyzed it post hoc by inspecting agent trajectories, we instead embed detectable reward hacking opportunities directly into environments. This makes their exploitation verifiable by design, enabling deterministic and automated measurement of whether and how agents exploit such vulnerabilities. We instantiate this approach in TextArena and release Hack-Verifiable TextArena, a testbed in which reward hacking can be measured reliably. Using this benchmark, we analyze reward hacking behavior across language models in diverse environments and settings.
Hallucination-Guided Unlearning: Using Hallucination Traces to Reveal Overfitted Memories
Jiayu Zhang ⋅ Ziqi Zhong ⋅ Yuliang Gai ⋅ Canran Xiao
Selective unlearning in large language models is increasingly necessary for safety and compliance, yet harmful or memorized content is often latent, entangled, and hard to specify as a clean ``forget set.'' We aim to close this gap by identifying what to forget from the model’s own failure modes rather than relying solely on explicit supervision. We propose Hallucination-Guided Unlearning (HGU), which treats hallucinations as diagnostic signals: it localizes unsupported spans via evidence and internal consensus, traces a sparse set of high-responsibility MLP neurons, and edits them with compact gates under a utility-preserving KL constraint. Across TOFU, HGU achieves the best MU--FE trade-off with Avg of 0.7221 (forget05) and 0.7699 (forget10), and it also attains the top SepS H-Avg (0.4382/0.4715) under mixed forget/retain prompts; on WMDP auxiliary, it yields the strongest forgetting (e.g., 14.2 on Economics) while improving Retain (52.7) and MMLU (57.2). These results suggest that reframing hallucinations as actionable traces enables targeted, post-hoc unlearning that better balances forgetting efficacy and general utility without full retraining.
HALO: Heterogeneous-Aware LoRA Optimization via Rank Allocation and Client-Aware Projection
XiaoHua Feng ⋅ Yuyuan Li ⋅ Jiayuan Fang ⋅ Fengyuan Yu ⋅ Jiaming Zhang ⋅ Li Zhang ⋅ Xiang Liu ⋅ Jun Wang ⋅ Chaochao Chen
Federated Learning (FL) with Low-Rank Adaptation (LoRA) enables privacy-preserving collaborative fine-tuning of foundation models. Existing federated LoRA studies often assume homogeneous ranks and lack principled rank allocation, leading to underfitting or overfitting under data and resource heterogeneity. Although later heterogeneous-rank federated LoRA methods relax the homogeneous-rank assumption, they still ignore resource-adaptive rank allocation across clients, while their aggregation and distribution strategies suffer from structural limitations, including rank collapse in naive truncation and the loss of locally relevant update signals caused by objective mismatch in SVD truncation. To address these issues, we propose HALO, a heterogeneous federated LoRA framework that integrates adaptive rank allocation with client-aware aggregation and distribution. Under the neural tangent kernel regime, we derive a sufficient rank condition to guide client-specific rank allocation from local data geometry and resource constraints. We further propose Client-Aware Projection (CAP), which builds a product-summed reference and redistributes rank-compatible balanced LoRA factors through projections onto client-induced subspaces, improving generalization via cross-client information while preserving locally relevant update signals. We show that CAP controls the client-loss-relevant truncation residual under the NTK approximation without additional communication overhead, which stabilizes local training and improves performance. Extensive experiments demonstrate the effectiveness of HALO. Our code is available at https://anonymous.4open.science/r/artifact2026-4174.
HALO: Homotopy-Augmented Layer Optimization for Stable LLM Supervised Post-training
Mengxiang Zhang ⋅ Lingyuan Liu
Supervised post-training, including supervised fine-tuning (SFT) and knowledge distillation (KD), is essential for adapting large language models (LLMs) to specialized domains. However, traditional supervised post-training methods often suffer from catastrophic forgetting, training instability, and prohibitive computational costs. While layer-wise strategies have emerged as efficient alternatives, they introduce a representation bottleneck, where constrained updates in early stages limit feature diversity and impair final generalization. To resolve these issues, we develop HALO (Homotopy-Augmented Layer Optimization), a homotopy-driven layer-wise strategy for LLM post-training. By introducing a continuous homotopy parameter and a sequence of monotonically increasing auxiliary functions, we formulate the post-training process as a continuous deformation of the optimization problem that progressively incorporates layers into the trainable set. Specifically, HALO unfreezes the model from the output back to the input in a structured sequence. We theoretically prove that this approach establishes a differentiable solution path, ensuring a stable transition from a restricted optimization state to a fully converged LLM. Extensive experiments across diverse model families and sizes demonstrate that HALO consistently enhances performance and generalization. Notably, HALO exhibits strong robustness in challenging scenarios, especially in low-resource settings, while significantly reducing convergence time and iteration counts. Our results position HALO as an effective and accessible framework for high-performing LLM specialization.
HAPACT: A Benchmark For Human-Centric Physical Impact Localization in Movies
Youngrae Kim ⋅ Yejin Jang ⋅ Eunho Kim ⋅ Youjin Sung ⋅ Kun-Woo Song ⋅ Sangpil Kim ⋅ Sang Ho Yoon
Physical reasoning about humans in videos has largely focused on everyday interactions, where contact, pressure, pose, and motion support embodied perception and contact-aware applications. Still, short and forceful properties like impact events remain relatively underexplored, despite their relevance to broader understanding of how abrupt physical events affect the human body. To fill this gap, we introduce human-centric impact localization, a new task that localizes target physical impacts temporally within a video and spatially on the human body. We further present HAPACT, a HumAn-centric imPACT dataset and benchmark built from movie clips, containing over 2.8K video-query pairs with more than 20K fine-grained human-annotated frames. We built HAPACT around viewer-grounded annotation, text-conditioned localization, and supervision at frame-level temporal and SMPL-X-based spatial granularity. To demonstrate the effectiveness of our benchmark, we conduct extensive evaluations of existing methods, employing video moment retrieval models for temporal localization and human-scene contact models for spatial localization. For the spatial task, we propose a text-conditioned baseline by reformulating localization as query-guided body-region selection. Benchmark results show that physical impact localization remains challenging for existing methods, motivating HAPACT as a dedicated benchmark to support physical impact reasoning in videos.
HCInfer: An Efficient Inference System via Heterogeneous Error Compensation for Resource-Constrained Devices
Shen Xu ⋅ Xiangwen Zhuge ⋅ Zhe Xu ⋅ Yingkun Hu ⋅ Zheng Yang ⋅ Yunhao Liu
LLMs often struggle with memory-constrained deployment on consumer-grade hardware due to their massive parameter sizes. While existing solutions such as model compression and offloading improve deployment feasibility, they often suffer from substantial accuracy degradation or severe throughput bottlenecks. Recent error compensation methods recover accuracy through auxiliary LoRA-style branches, and we observe that these branches are inherently amenable to offloading: they require substantial parameter storage but access only a small subset of compensation parameters during each inference step. Motivated by this opportunity, we propose HCInfer, a heterogeneous inference system that offloads residual compensation to the CPU while executing the compressed backbone on the GPU, and further introduces an asynchronous compensation pipeline and sensitivity-aware dynamic rank allocation to hide compensation overhead and maximize accuracy recovery. Experimental results show that HCInfer achieves a maximum accuracy improvement of 5.2% on downstream tasks compared to compression model and sustaining a maximum speedup of 10.4x compared to full-precision model.
Heads That Write, Not Just Point: Image Retrieval Heads in Vision-Language Models
Junsung Park ⋅ Uiwon Hwang ⋅ Donghun Kang ⋅ Yeongtak Oh ⋅ Jaewon Jeong ⋅ Gyeongtae Yoo ⋅ Han Cheol Moon ⋅ Jungbeom Lee ⋅ Sungroh Yoon
Large vision-language models (LVLMs) can process interleaved text and multiple images in a single context, but how they internally identify the image relevant to a question remains poorly understood. We study this grounding step as in-context image retrieval and introduce ICIR-MCQ, a controlled probe for analyzing image retrieval at attention-head granularity. Using each head's attention mass over image-token spans, we first identify sparse pointer heads that reliably point to the target image. These heads transfer beyond the controlled probe: a single pointer head outperforms external vision-language retrievers on a naturalistic multi-image retrieval benchmark, showing that LVLMs contain a strong internal image retrieval signal. However, attention alone only shows where a head reads from, not what information its output passes on to later layers. We therefore introduce a complementary output-side analysis. For each head, we train a lightweight image retriever on its value-weighted output and use it to identify writer heads, whose outputs contain enough information to identify the target image. Surprisingly, the pointer and writer criteria select substantially different head sets, with only 35 to 44 heads overlapping among the top-100 heads under each criterion. Targeted set-partition ablations show that the heads most important for standard multi-image inference are concentrated in this pointer-writer intersection, while heads selected by only one criterion contribute little beyond random-head controls. These results show that attention-based pointing alone is systematically incomplete for identifying image retrieval heads in LVLMs. Causally relevant retrieval heads must also write target-image information to the residual stream, not merely point to it.
Hear, Localize, and Reason: Spatially Aware Scene Understanding for Audio-visual LLMs
Sung-Bin Kim ⋅ Lee Jung-Mok ⋅ Jinwoo Jung ⋅ Oh Hyun-Bin ⋅ Hyeonggon Ryu ⋅ David Harwath ⋅ Tae-Hyun Oh
Audio-visual large language models (AV-LLMs) have made strong progress in multimodal understanding, but they typically treat audio as monaural semantic content and therefore struggle to reason about where sounds originate and how they relate to the visual scene. Recent spatial audio-visual studies address parts of this problem, but often focus on specific abilities, such as spatial correspondence or direction and distance reasoning. In this paper, we propose HLR-AVSceneQA, a comprehensive benchmark for spatial audio-visual scene understanding. HLR-AVSceneQA evaluates whether models can jointly hear, localize, and reason by recognizing what is heard, grounding where it comes from in egocentric and allocentric views, and inferring how sound sources relate to visible, hidden, or nearby objects. We further introduce HLR-LLM, which extends a strong audio-visual foundation model with a binaural spatial audio branch. HLR-LLM is trained with a three-stage curriculum, applying chain-of-thought supervision in the final stage to help the model learn structured multi-step spatial reasoning. Experiments show that existing AV-LLMs remain limited in spatial grounding and relational reasoning, while HLR-LLM substantially improves spatial audio-visual scene understanding without sacrificing semantic audio-visual perception.
HELICS: Biobank-scale Conditional Synthetic Genome Generation via Latent Flow Matching
Ahmad Abdel-Azim ⋅ Xihong Lin
Individual-level genotype data underpin genetic discovery, disease-risk modeling, and therapeutic development, yet many critical settings remain data-limited, including underrepresented ancestries, rare or imbalanced disease cohorts, and restricted-access biobanks. Conditional synthetic genome generation would support privacy-aware benchmarking, simulation, and data augmentation, but requires learning an ultra-high-dimensional discrete distribution over hundreds of thousands to tens of millions of correlated single-nucleotide polymorphisms (SNPs), or sites of DNA variation. We introduce $\mathbf{HELICS}$: $\mathbf{H}$igh-fidelity g$\mathbf{E}$neration via $\mathbf{L}$atent $\mathbf{I}$nterpolant $\mathbf{C}$onditional $\mathbf{S}$ynthesis, a phasing-free latent flow matching framework for genome-wide SNP-level generation from biobank-scale genotype data. HELICS tokenizes chromosomes with convolutional autoencoders, reducing $>$$600{\rm K}$ SNPs to $4{,}096$ continuous latent tokens, and trains a transformer parameterized conditional flow matching model over the concatenated multi-chromosome latent representation. Trained on roughly $428\rm K$ genomes in the UK Biobank, HELICS generates genome-wide synthetic cohorts without haplotype phasing and achieves stronger fidelity to held-out real genotypes than existing haplotype- and genotype-based simulators, including the highest validation allele-frequency agreement ($R^2=0.992$) and superior preservation of local and long-range SNP correlation structure. Conditioning on ancestry and genetic liability summaries enables targeted generation across genetic risk profiles; in perturbation analyses, by increasing the conditioned disease trait liability, HELICS induces SNP-level dosage changes in generated genomes that are highly correlated with published association weights. HELICS provides a scalable generative modeling framework for SNP-level genome data and a path toward conditional synthetic cohorts for data-limited genetic studies.
Hessian-Dependent Sample Complexity in Zeroth-Order Stochastic Optimization: Suboptimality of Convex-Support Sampling and Optimal Sample Complexity
Mengtian Hong ⋅ Jason Lee ⋅ Qian Yu
Zeroth-order stochastic optimization is a fundamental formulation that arises in real-world design problems where gradients are inaccessible. A central challenge in this pursuit is to design gradient estimators and optimization algorithms under noisy, function-only feedback that exploit local Hessian geometry to achieve optimal sample efficiency. We introduce the Spectrally Grouped Estimator (SGE), a novel gradient estimator that samples over a non-convex union of sphere sections, and utilize it to build an algorithm that achieves order-wise improved Hessian-dependent simple regrets over second-order smooth, strongly convex functions compared to conventional convex-sets sampling baseline methods. We complement these results with the first tight analyses of the baseline schemes, revealing a shared performance bottleneck and thus emphasizing the necessity of non-convex sampling for optimality. We further establish matching converse bounds that tightly characterize the optimal sample complexities within general subclasses of functions sharing the same Hessian at the global minimum, thereby proving the universal optimality of SGE with respect to the Hessian-dependent rates. This fully resolves an open conjecture from the preliminary version of this work. This fully resolves an open conjecture in the preliminary version of this work.
Heterogeneous Graph Federated Learning with Structure-Aware Data-Free Distillation
Lili Guo ⋅ Qi Li ⋅ Ruoyu Wang ⋅ Chao Li ⋅ Shifei Ding
Subgraph heterogeneity is a critical challenge in federated learning, significantly impacting model performance. Data-free knowledge distillation overcomes the limitation of sharing private label distributions in federated learning, yet existing research lacks in-depth exploration of subgraph heterogeneity. Our work builds upon the fact that node and structural variations in heterogeneous graphs cause significant disparities in the reliability of knowledge from local graph neural networks. We propose a structure-aware bidirectional data-free federated distillation technique, where the generator and global model engage in an adversarial training process to ensure reliable bidirectional knowledge transfer between local and global models. Extensive experiments across multiple public datasets demonstrate the model's effectiveness.
Hierarchical Variational Policies for Reward-Guided Diffusion
Kushagra Pandey ⋅ Farrin Marouf Sofian ⋅ Jan Niklas Groeneveld ⋅ Felix Draxler ⋅ Stephan Mandt
Adapting pretrained diffusion models to downstream objectives such as inverse problems often requires expensive test-time guidance or optimization. We propose a principled framework for generating high-quality reward-aligned samples at substantially reduced inference cost. Our approach formulates test-time adaptation as a hierarchical variational model, where control is amortized into a lightweight yet expressive stochastic policy. This formulation naturally supports few-step diffusion sampling: large step sizes enable fast inference, while the learned policy maintains sample quality by providing structured per-step control. The resulting fully amortized sampler achieves a strong quality--speed tradeoff, matching or exceeding recent test-time scaling baselines while requiring significantly less compute. For example, on $4\times$ super-resolution, our method achieves better perceptual quality with more than 5$\times$ faster inference compared to the best-performing baseline. We further extend our approach to a semi-amortized regime that combines cheap amortized proposals with limited test-time optimization, achieving state-of-the-art perceptual quality across several challenging inverse problems.
Hierarchical World Models with Implicit Dynamics
Gaoyue Zhou ⋅ Yvonne Wu ⋅ Zichen Cui ⋅ Nicolas Ballas ⋅ Mahmoud Assran ⋅ Lerrel Pinto ⋅ Yann LeCun
Planning with world models enables test-time adaptation for embodied agents, offering superior generalization over direct policy learning. However, long-horizon tasks remain hindered by cumulative errors and search complexity. Existing hierarchical methods often use rigid macro-actions or latent skills, leading to brittle representations and limited fine-grained control. We introduce Implicit-HWM, a hierarchical world model that decouples state-space reachability from local control. Our framework utilizes a high-level model to sample feasible transitions across the state manifold and a low-level inverse dynamics solver to ground these transitions into precise actions. This allows the planner to identify reachable subgoals without being constrained by low-level execution details. By prioritizing local physical feasibility over global behavioral patterns, Implicit-HWM synthesizes novel long-horizon plans from short segments, transcending the specific behavioral modes seen during training. We demonstrate its effectiveness across three navigation and manipulation suites, achieving a 38.5% improvement over state-of-the-art policy and diffusion-based planners. Crucially, by leveraging this hierarchical structure to mitigate compounding errors, our approach yields a 120% performance gain over single-level world model MPC, particularly in complex, unseen configurations for long-horizon tasks.
HiLoc: A Hierarchical Representation Method for Spatial Localization in Multimodal Large Language Models
Evelyn Zhang ⋅ Fufu Yu ⋅ Hanjun Li ⋅ Aoqi Wu ⋅ Ke Yan ⋅ Shouhong Ding ⋅ Tianhe Ren ⋅ Xinting Hu ⋅ Xiaojuan Qi ⋅ Li Jiang ⋅ Jiaya Jia
Existing multimodal large language models (MLLMs) typically formulate object detection as a conditional sequence generation task. Under this paradigm, spatial locations are commonly represented by either textual coordinates or special quantized learnable tokens. However, textual coordinates are tokenized as unrelated symbols, spatially adjacent points (e.g., $x=199$ vs. $y=200$) yield entirely different tokens, so the cross-entropy loss is misaligned with true spatial distance. Special quantized location tokens, on the other hand, inevitably introduce quantization error. Reducing this error requires linearly more tokens, but a larger vocabulary is harder to train and converges slower. We argue that both limitations stem from the formulation of spatial representation. To address this issue, we propose a simple yet novel localization representation method, HiLoc, which predicts locations in a coarse-to-fine manner with special tokens: first a grid index for coarse localization, then a cell index to refine the position within the grid. Under this paradigm, accuracy can be scaled up by simply adding the hierarchy level. We also build HiRef-500K, a dataset consisting negative and complex referring samples to strengthen rejection and complex referring ability. Experiments show that our method achieves strong performance on both detection and grounding benchmarks, especially on small objects, under high IOU thresholds and complex referring tasks.
Offline goal-conditioned reinforcement learning methods have shown promise for reach-avoid tasks, where an agent must reach a target state while avoiding undesirable regions of the state space. Existing approaches typically encode avoid-region information into an augmented state space and cost function, which prevents flexible, dynamic specification of novel avoid-region information at evaluation time. They also rely heavily on meticulously designed reward and cost functions, limiting transferability to novel environments and specifications. We introduce RADT, a decision transformer approach for offline, reward-free, goal-conditioned, avoid-region-conditioned RL. RADT encodes goals and avoid-regions directly as prompt tokens, allowing any number of avoid-regions of arbitrary size to be specified at evaluation time. Using only suboptimal offline trajectories from a random policy, RADT can learn reach-avoid behavior in a completely data-driven manner using our novel avoid-region hindsight relabeling approach, without the need for reward/cost supervision. We benchmark RADT against existing offline goal-conditioned RL models across 17 tasks, environments, and experimental settings. RADT generalizes in a zero-shot manner to out-of-distribution avoid-region sizes and counts, outperforming baselines that require retraining. In one such zero-shot setting, RADT achieves 35.7% improvement in normalized cost over the best retrained baseline while maintaining high goal-reaching success. We also apply RADT to cell reprogramming in biology, demonstrating its versatility.
History-Aware Conformal Prediction Sets for Censored Time-to-Event Outcomes
Yuyao Wang ⋅ Alexander W Levis ⋅ Shu Yang ⋅ Larry Han
Existing conformal prediction methods for time-to-event outcomes leverage only baseline covariates, producing prediction intervals that are insufficiently informative to facilitate decision making. We propose History-Aware Prediction Sets (HAPS), a conformal framework that constructs prediction sets for individual event times using covariate histories observed up to a decision time, targeting coverage among individuals who have survived to this time. HAPS handles right censoring adjusted for time-varying confounders via inverse probability of censoring weighting. When the censoring weights are consistently estimated, it achieves PAAC (probably asymptotically approximately correct) coverage among survivors. We further propose two doubly robust extensions of HAPS to weaken reliance on consistent estimation of the censoring distribution. In simulations, HAPS and its extensions reduce median prediction interval length by up to 75\% relative to baseline comparators while maintaining close to nominal coverage. On two public benchmark data sets, HAPS reduces the median interval length by up to 60\% for predictions at year 5, compared to the baseline comparators.
Prior Laplacian-based option-discovery methods build options from 0-form Laplacians defined on nodes, where the resulting eigenvectors induce routes toward isolated boundaries or extrema, thereby skewing visitation toward specific states. In contrast, kernel eigenvectors of the 1-form Hodge Laplacian represent interpretable, topology-aware edge flows and naturally capture cyclic behaviour known as harmonic flows. These kernel eigenvectors are particularly attractive because they define zero-divergence edge flows, offering a natural solution for non-isolated visitation to sources and sinks in regions in which they are active. Despite this appeal, harmonic flows have not previously been explored as a basis for option discovery in reinforcement learning. This is what we do here. We first introduce harmonic-flow skills and show that, on state graphs with tree-like “dead-end” appendages, purely harmonic flows can become nearly inactive in those regions, which can limit coverage. To address this limitation, we propose quasi-harmonic flow options: skill-inducing edge flows that preserve harmonic cycles where harmonic activity is strong, while injecting drift into appendages using the lowest-magnitude non-harmonic eigenvector of the Hodge Laplacian, a low divergence solution, where harmonic activity is weak. Our approach is eigenvector-only, relying solely on spectral components of the Hodge Laplacian. We compare quasi-harmonic flow options against a harmonic-flow options baseline and state-of-the-art skill-discovery methods from the literature, demonstrating empirical gains and analytical improvements on hitting-time-based metrics.
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enhancement
Yiyang Cai ⋅ Nan Chen ⋅ Rongchang Xie ⋅ Junwen Pan ⋅ Chunyang Jiang ⋅ Cheng CHEN ⋅ Chang Liu ⋅ Zhenbang Sun ⋅ Wei Xue ⋅ Wenhan Luo ⋅ Yike Guo
Human-object-centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on $\textit{inter-subject}$ personalization still struggle to strike a balance between high subject fidelity and accurate interaction patterns between humans and diverse objects, especially when objects represent abstract concepts such as logos. Second, while $\textit{intra-subject}$ references ($\textit{e.g.}$, OCR maps, multi-view inputs) are expected to enhance subject fidelity, most existing works lack mechanisms to understand such latent correspondence. To address both challenges, we propose HOMIE, an HOCVP framework that tackles both inter- and intra-subject input settings in a unified manner. Compared to previous approaches, HOMIE proposes a better MLLM integration strategy to extract knowledge of reference-level relationships without compromising the controllability of text encoders or incurring costly re-alignment. Specifically, we introduce global multimodal guidance within self-attention to better align MLLM-derived semantic features with VAE tokens. Furthermore, we propose modality-reference embedding to differentiate tokens from MLLM features and VAE tokens and associate intra-subject reference image tokens. Extensive experiments validate that our method achieves state-of-the-art performance across various HOCVP tasks. Anonymous project page with videos is available at https://homie-demo.github.io/.
How Can SignSGD Outperform SGD? A Functional Scaling Law Perspective
Zilin Wang ⋅ Binghui Li ⋅ Jia-Nan Wang ⋅ Lean Wang ⋅ Jinbo Wang ⋅ Lei Wu
We study signSGD in linear regression with diagonal features and power-law spectra, a toy but expressive setting where the effect of coordinate-wise sign normalization can be analyzed sharply. Specifically, we establish a functional scaling law (FSL) for signSGD under general learning rate schedules, which decomposes the loss dynamics into two components: learning of the target signal and noise accumulation. Compared with SGD, signSGD learns the signal faster but exhibits slower noise forgetting. This signal-noise tradeoff yields a regime-dependent comparison between signSGD and SGD: signSGD achieves better data-scaling efficiency in hard-task regimes, while they perform similarly in easy-task regimes. The key mechanism is that gradient noise is {\it curvature-aligned}: its directional variance is of the same order as the directional curvature. As a result, sign operation effectively acts as a \(\diag(\bH)^{-1/2}\) preconditioner in the noise-dominated regime, where $\bH$ is the Hessian matrix. Moreover, we demonstrate that the same square-root preconditioning mechanism can arise in RMSprop and Adam, where coordinate-wise normalizers essentially track the coordinate curvatures. Synthetic experiments validate the proposed mechanism and the predicted scaling behavior. Language model experiments further suggest that, beyond the controlled setting, the resulting FSL can serve as a surrogate model for fitting and predicting practical training loss curves.
Despite their success in image generation, diffusion models can memorize training data, raising serious privacy and copyright concerns. Although prior work has identified empirical factors associated with memorization, such as text guidance, prompt specificity, and duplicated training examples, the mechanism by which memorization is triggered and propagated during denoising remains unclear. In this paper, we provide a theoretical explanation of how memorization occurs in text-to-image diffusion models. We show that conditional overfitting causes the conditional posterior mean to collapse to the memorized training latent, while the unconditional posterior remains close to the zero-centered data mean at the first denoising step. As a result, classifier-free guidance overestimates the clean prediction and injects an amplified memorized signal into the reverse trajectory. We further show that this signal propagates across timesteps because each guided latent is fed back into the denoising network, allowing the conditional branch to indirectly shift the unconditional branch. To characterize this process, we introduce the posterior lag, the discrepancy between conditional and unconditional posterior means. For normal prompts, this lag decreases as the two posteriors synchronize during denoising. For memorized prompts, the conditional posterior commits early to a specific training latent, while the unconditional posterior catches up only later, producing a distinctive rise-and-fall lag pattern. Our analysis provides a mechanism-level explanation of memorization and clarifies why text guidance and early denoising steps play a central role in memorized generation.
How Does Personalized Memory Shape LLM Behavior? Benchmarking Rational Preference Utilization in Personalized Assistants
Xueyang Feng ⋅ Weinan Gan ⋅ Xu Chen ⋅ Quanyu Dai ⋅ Yong Liu
Large language model (LLM)-powered assistants have recently integrated memory mechanisms that record user preferences, leading to more personalized and user-aligned responses. However, irrelevant personalized memories are often introduced into the context, interfering with the LLM's intent understanding. To comprehensively investigate the dual effects of personalization, we develop \textsc{RPEval}, a benchmark comprising a personalized intent reasoning dataset and a multi-granularity evaluation protocol. \textsc{RPEval} reveals the widespread phenomenon of irrational personalization in existing LLMs and, through error pattern analysis, illustrates its negative impact on user experience. Finally, we introduce \textsc{RP-Reasoner}, which treats memory utilization as a pragmatic reasoning process, enabling the selective integration of personalized information. Experimental results demonstrate that our method significantly outperforms carefully designed baselines on \textsc{RPEval}, and resolves 80\% of the bad cases observed in a large-scale commercial personalized assistant, highlighting the potential of pragmatic reasoning to mitigate irrational personalization. Our data is available at \url{https://anonymous.4open.science/r/RPEval-E4B0}.
How Does Pruning Change Decisions in Large Language Models?
Jiang Li ⋅ Tian Lan ⋅ Chenxi Zhou ⋅ Zehua Duo ⋅ Xianpeng Shang ⋅ Guanglai Gao ⋅ Xiangdong Su
Post-training pruning reduces the memory and computation costs of large language models (LLMs), and N:M semi-structured sparsity provides a hardware-friendly pruning pattern for efficient inference. However, pruned models are often evaluated mainly by perplexity and average downstream accuracy, which are useful aggregate metrics but do not fully explain how pruning changes individual candidate-level decisions. We evaluate magnitude pruning, Wanda, SparseGPT, Wanda++, and our proposed MR-GOBS across multiple model families and sparsity settings. We further conduct a paired decision-level evaluation that measures true-label confidence, answer margin, prediction flips, candidate-level distribution shift, and near-boundary vulnerability. Our results show that pruning weakens decision reliability by lowering confidence in the correct answer, shrinking answer margins, and increasing correct-to-wrong flips, especially on near-boundary examples. To examine whether reconstruction-oriented improvements reliably improve decision reliability, we revisit SparseGPT through its local least-squares reconstruction formulation and refine its N:M mask-selection step. MR-GOBS uses GroupOBS, a set-level OBS reconstruction cost for jointly pruning a group of weights, to score each legal local N:M mask. It then selects masks using a minimax-regret criterion over multiple calibration views. Although MR-GOBS often improves average perplexity over SparseGPT, these gains do not always translate into better downstream accuracy or decision-level metrics. Our findings suggest that reliable evaluation of pruned LLMs requires both aggregate performance metrics and decision-level diagnostics.
How Far Are VLMs from Privacy Awareness in the Physical World? An Empirical Study
Junran Wang ⋅ Xinjie Shen ⋅ Zehao Jin ⋅ Pan Li
As Vision-Language Models (VLMs) are increasingly deployed as autonomous cognitive cores for embodied assistants, evaluating their privacy awareness in physical environments becomes critical. Unlike digital chatbots, these agents operate in intimate spaces, such as homes and hospitals, where they possess the physical agency to observe and manipulate privacy-sensitive information and artifacts. However, current benchmarks remain limited to unimodal, text-based representations that cannot capture the demands of real-world settings. To bridge this gap, we present ImmersedPrivacy, an interactive audio-visual evaluation framework that simulates realistic physical environments using a Unity-based simulator. ImmersedPrivacy evaluates physically grounded privacy awareness across three progressive tiers that test a model's ability to identify sensitive items in cluttered scenes, adapt to shifting social contexts, and resolve conflicts between explicit commands and inferred privacy constraints. Our evaluation of 12 state-of-the-art models reveals consistent deficits. In cluttered scenes, all models exhibit monotonic performance decay as scene complexity grows due to perceptual deficit. When social context shifts, no model exceed 65% selection accuracy. Under conflicting commands, the best model gemini-3.1-pro perfectly balances task completion and privacy preservation in only 51% of cases. These findings reveal that current VLMs in the physical world suffer from perceptual fragility and fail to let their knowledge of privacy cues govern their situated behavior. Our code and data is available at https://anonymous.4open.science/r/immersed-privacy-review-3A3C.
How I learned to stop worrying and love StopGrads: Stationarity, Convergence, and a case study on Flow Map Learning
Mark Goldstein ⋅ Max Shen ⋅ Zichu Wang ⋅ Aahlad Manas Puli ⋅ Rajesh Ranganath
Stopgrads are widely used in training machine learning models, but stopgrads can alter the gradient, stationary points and convergence guarantees of the original objective, which can make stopgrad training theoretically ungrounded. We introduce a \textit{stopgrad regression principle}, which identifies a general template for stopgrad objectives with a closed-form characterization of stationary points, and unifies stopgrad objectives for flow maps, reinforcement learning, and diffusion samplers. We provide theoretical grounding for optimizing stopgrad flow map objectives by characterizing stationary points and proving convergence results. Remarkably, we show that the learned flow map has a closed-form expression composing the initial flow map and the true flow map, under functional semi-gradient flow. We use our stopgrad regression principle to propose novel modified stopgrad placements for flow map objectives which speed up training by ${\sim}2.5\times$.
How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
Kristian Schwethelm ⋅ Daniel Rueckert ⋅ Georgios Kaissis
We measure how much one recurrence is worth to a looped (depth-recurrent) transformer, in equivalent unique parameters. From an iso-depth pretraining sweep across recurrence counts $r \in \{1, 2, 4, 8\}$ spanning ${\sim}50\times$ in training compute, we fit a joint scaling law $L = E + A\,(N_\text{once} + r^{\varphi} N_\text{rec})^{-\alpha} + B\,D^{-\beta}$ and measure a recurrence-equivalence exponent $\varphi = 0.46$. Intuitively, $\varphi$ tells us whether looping a block $r$ times is equivalent in validation loss to $r$ unique blocks of a non-looped model (full equivalence, $\varphi{=}1$) or to a single block run repeatedly with no capacity gain ($\varphi{=}0$). Our $\varphi = 0.46$ sits in between, so replacing unique blocks with shared recurrences increases validation loss at matched training compute. For example, at $r{=}4$ a 410M looped model performs on par with a 580M non-looped model, but incurs the training cost of a 1B non-looped one. We demonstrate the utility of $\varphi$ as a diagnostic tool on two case studies: commonly used truncated backpropagation lowers $\varphi$ to $0.38$, indicating that the loop mechanism is poorly trained under truncation, even though validation loss decreases. Conversely, hyperconnections raise $\varphi$ to $0.65$, a genuine capacity gain. Our method separates true loop improvements from training-side gains, a distinction raw validation loss cannot make.
How Post-Training Shapes Biological Reasoning Models
Lukas Fesser ⋅ Hanlin Zhang ⋅ Michelle M Li ⋅ Eric Wang ⋅ Bryan Perozzi ⋅ Shek Azizi ⋅ Sham Kakade ⋅ Marinka Zitnik
Post-training is central to building scientific reasoning models, but its effects remain poorly understood in biology, where models integrate natural language with multimodal biological data. We study how post-training stages shape generalization, and when they improve performance versus induce over-specialization. Across genomics, transcriptomics, and proteomics, we train and evaluate more than 100 biological reasoning models under controlled variation in backbone, continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL), measuring both in-domain (ID) and out-of-domain (OOD) performance. We find that post-training induces stage-specific trade-offs rather than uniform gains. CPT improves downstream performance by aligning models with biological language. SFT consistently increases ID performance but causes OOD performance to peak early and decline as models fit the training distribution. RL, when applied to strong SFT checkpoints with aligned rewards, improves OOD performance and partially recovers generalization. These results show that scientific reasoning does not improve monotonically with additional supervision or compute. Instead, performance depends on how training stages are composed. Under fixed epoch-level budget, the best ID-OOD trade-off comes from combining minimal SFT with larger RL budgets and allocating adaptation capacity asymmetrically across stages.
Concept Bottleneck Models (CBMs) have become a popular approach to enable interpretability in neural networks by constraining classifier inputs to a set of human-understandable concepts. While effective, current models embed concepts in flat Euclidean space, treating them as independent, orthogonal dimensions. Concepts, however, are highly structured and organized in semantic hierarchies. To resolve this mismatch, we propose Hyperbolic Concept Bottleneck Models (HypCBM), a post-hoc framework that grounds the bottleneck in this structure by reformulating concept activation as asymmetric geometric containment in hyperbolic space. Rather than treating entailment cones as a pre-training penalty, we show they encode a natural test-time activation signal: the margin of inclusion within a concept's entailment cone yields sparse, hierarchy-aware activations without any additional supervision or learned modules. We further introduce an adaptive scaling law for hierarchically faithful interventions, propagating user corrections coherently through the concept tree. Empirically, HypCBM rivals post-hoc Euclidean models trained on 20$\times$ more data in sparse regimes required for human interpretability, with stronger hierarchical consistency and improved robustness to input corruptions.
Hyperbolic Language Models: From Zipf to Compute-Optimal Scaling
Jinrui Lin ⋅ Menglin Yang ⋅ Rex Ying
Compute-optimal scaling laws have become the central organizing tool of language-model training, yet they have been measured exclusively under Euclidean readouts. We give the first compute-optimal scaling-law characterization of hyperbolic language modeling. Trained from scratch on OpenWebText across a paired 5×4 grid spanning 86M to 1.32B parameters and 66M to 8.4B tokens, our hyperbolic LM HypTip follows a markedly different scaling law than its Euclidean counterpart. Fitting the Hoffmann form L = E + A·N^(−α) + B·D^(−β) to each track yields a compute-optimal data-to-parameter ratio of D*/N* = 8.4 for the hyperbolic readout versus 20.8 for the Euclidean baseline. The hyperbolic readout therefore reallocates compute by 2.5× toward parameter-heavy training, winning 17 of 20 paired cells across the grid. This shift is mechanistically grounded: across all 20 trained checkpoints, the per-row Lorentzian radius of the LM head correlates with token log-frequency at Spearman ρ down to −0.83, and the strength of this Zipfian token hierarchy scales with training-token count rather than parameter count, indicating that the geometry-as-prior emerges causally during training rather than being read off a frozen Euclidean checkpoint as in prior post-hoc analyses. Held-out evaluation on WikiText-103 perplexity, LAMBADA, and PIQA preserves the same advantage, and a cross-architecture replication on GPT-2 and DeepSeek-V3 confirms that LM-head geometry, not body details, drives the effect. Output-layer geometry alone therefore reshapes the compute-optimal scaling law of language modeling and brings hyperbolic LMs into the same predictive framework as the Euclidean models that have shaped the field.
HyperGen: Learning Structure-Aware Spectral Flows for Hypergraph Generation
Yu Xie ⋅ Junqi Liu ⋅ Hangyuan Du ⋅ Zhiqiang Wang ⋅ Ming Li
Hypergraphs provide a natural representation for high-order relations, but real-world hypergraph data is often difficult to share or collect due to privacy, access, and annotation constraints. Hypergraph generation is therefore useful for constructing surrogate high-order relational data. Existing generators, however, typically build hyperedges through structural heuristics, propagation rules, or projected graph spaces, and thus do not directly model the incidence relation between nodes and hyperedges. Spectral representations offer a principled way to encode global and local relational patterns, but hypergraph spectra are intrinsically non-invertible: a spectral representation does not determine a unique discrete incidence structure. We propose HyperGen, a structure-aware generative framework that models hypergraph generation as continuous dynamics in a phase-enhanced spectral latent space. HyperGen constructs a Hermitian node--hyperedge spectral operator with learnable incidence phases, learns a conditional vector field that transports noise toward target spectral embeddings, and decodes the generated latents into node--hyperedge incidence matrices through a learned spectrum-to-structure mapping with denoising. Experiments on benchmark hypergraphs show that HyperGen more accurately recovers node--hyperedge incidences and better preserves spectral and geometric characteristics than existing baselines.
Hypergraph Generation with Latent Diffusion
Valerio Di Pasquale ⋅ Alessia Antelmi ⋅ Mirko Polato ⋅ Carmine Spagnuolo
Hypergraphs capture high-order interactions that ordinary graphs cannot represent, yet generating realistic ones remains challenging because existing methods often rely on fixed structural assumptions or struggle to model coupled dependencies between nodes and hyperedges. In this work, we propose \textsc{Janus}, a latent diffusion framework that decomposes an observed hypergraph into sub-hypergraphs, learns coupled node and hyperedge latent representations with a dual-view $\beta$-VAE, and generates new structures through paired latent diffusion with cross-view conditioning. The method supports both unconstrained and node-set-constrained generation. We also introduce an evaluation protocol covering micro, meso, and macro scales interactions, structural reconstruction, and high-order similarity measures. Across five real-world datasets and nine baselines, \textsc{Janus} achieves the most stable performance across structural scales, with strong gains in meso-scale and high-order fidelity; in the constrained setting, it consistently obtains the best reconstruction results.
Hypergraph Representation Learning with Hyperlink Random Effects
Zimeng Li ⋅ Shihao Wu ⋅ Gongjun Xu ⋅ Ji Zhu
Hypergraphs record multi-way interactions among entities. Extracting information from the combinatorial structure underlying observed multi-way interactions is a central task in many real-world problems. Existing methods face several limitations. First, many deep architectures for hypergraphs do not explicitly exploit the potential low-rank structure, which can sacrifice parsimony and interpretability in the learned representations. Second, many low-rank-based methods operate on tensor representations, which typically require hyperlinks to have uniform sizes and thus limit their applicability to general hypergraphs with non-uniform hyperlink sizes. Third, many methods ignore the fact that hyperlinks often arise from heterogeneous mechanisms. For example, medical symptoms may co-occur in the profiles of patients with very different conditions, and such heterogeneity should be incorporated into the learning process. In this work, we develop a general framework for hypergraph representation learning using hyperlink random effects while exploiting the low-rank structure in hypergraphs. The proposed framework accommodates latent heterogeneity in hyperlink formation while preserving entity interaction patterns. We establish identifiability of the model parameters and leverage an expectation-maximization strategy for estimation. The framework allows flexible specifications for the hyperlink random effects; in this paper, we study three choices: categorical, Gaussian mixture, and score-based effects, and develop corresponding estimation algorithms. Through simulation studies, we demonstrate the effectiveness of the proposed method in recovering latent structure and capturing heterogeneous interaction patterns. Empirical studies on real-world hypergraph datasets further illustrate the practical utility of our approach.
HyperNSDE: Personalized Neural SDEs for Joint Static—Longitudinal Clinical Data Generation
Perrine Chassat ⋅ Agathe Guilloux
Synthetic patient data generation is a promising solution to the dual challenge of data scarcity and privacy constraints in healthcare machine learning. Realistic synthesis of patient-level clinical data requires jointly modeling heterogeneous static covariates, irregularly sampled longitudinal trajectories, and informative observation times — three tightly coupled components in practice yet rarely addressed together. We propose HyperNSDE, a continuous-time generative model that conditions a latent Neural SDE on static patient representations through a hypernetwork, allowing baseline characteristics to shape the full trajectory dynamics rather than only the initial condition, without requiring a trajectory encoder, while stochastic latent dynamics capture realistic variability in generated paths. Observation times are modeled jointly through a latent-state-dependent intensity process. Training on irregular stochastic paths is stabilized via a deterministic—stochastic path decomposition with a non-adversarial signature-kernel objective. Experiments on simulated and real clinical datasets demonstrate consistent improvements over strong baselines across fidelity, privacy, and utility metrics.
Hyperspherical Local Margin Retraction for Zero-Shot Instance-Wise Machine Unlearning
Hongyi Lyu ⋅ Xuyun Zhang ⋅ Xiaoxiao Chi ⋅ Guanfeng Liu ⋅ Amin Beheshti
Zero-shot machine unlearning aims to remove the influence of specified training instances from a pretrained model when the retain set is unavailable. Instance-wise unlearning further operates at the level of individual instances, where a forget instance may induce instance-specific representation and margin deviations that are entangled with class-relevant local structure. However, without retain instances, separating these deviations from the structure that retraining would preserve is difficult, exposing retain-forget entanglement in representation space. We propose HLMR, a novel zero-shot instance-wise unlearning approach that formulates unlearning as local margin retraction in hyperspherical representation space. HLMR uses hyperspherical geodesic probes to construct a local reference in representation space, preserving the structure supported by this reference while unlearning instance-specific margin excess. Empirical evaluation across multiple dataset--architecture settings shows that HLMR aligns more closely with the retraining oracle in utility and membership-inference risk than state-of-the-art zero-shot baselines.
HyperTransport: Amortized Conditioning of T2I Generative Models
Valentino Maiorca ⋅ Eleonora Gualdoni ⋅ Xavier Suau ⋅ Marco Cuturi ⋅ Luca Zappella ⋅ Pau Rodriguez
As foundation models grow in capability, the ability to efficiently and reliably control their behavior becomes critical. Fine-tuning these models can be costly, and while prompting can be practical for controllability, it remains fragile due to models' high sensitivity to exact prompt wording and structure. This brittleness has driven interest in activation steering techniques that offer more stable and predictable control over model behavior. However, existing activation steering methods require per-concept optimization, which makes them ill-suited to deployment scenarios where the concept set is large, evolving, or only specified at request time: each new concept incurs at least minutes of optimization on the target model. We propose HyperTransport, a hypernetwork framework that amortizes this cost by mapping embeddings from a pretrained encoder (CLIP in our instantiation) directly to intervention parameters, trained end-to-end using an optimal transport loss. Once trained, HyperTransport produces each new intervention in a single hypernetwork forward pass, 3600–7000x faster than per-concept fitting. On concepts unseen during training, it matches the strongest per-concept baselines at inducing the target concept. By decoupling concept representation from intervention prediction, HyperTransport combines three capabilities that no existing approach offers as a set: amortized steering for open-ended concept sets, continuous interpretable strength control, and cross-modal conditioning where reference images can directly steer text-based generation. We validate HyperTransport on DMD2 and Nitro-1-PixArt across 167 held-out test concepts via CLIP-based metrics, a VLM-as-a-judge evaluation, and a user study. In pairwise comparisons, both human and VLM judges prefer HyperTransport over prompting ~2x as often.
HyperTree: Scalable Unsupervised Hierarchy Discovery in Hyperbolic Space
Thomas Lang ⋅ Kevin Sidak ⋅ Anna Beer ⋅ Sebastian Tschiatschek ⋅ Yllka Velaj ⋅ Claudia Plant
Real-world image datasets are deeply hierarchical, yet unsupervised hierarchical representation learning remains underdeveloped. Hyperbolic space is geometrically ideal for embedding trees with low distortion due to its exponential volume growth, but existing unsupervised hyperbolic methods either fail to scale beyond a few thousand points or depend heavily on class supervision. We bridge this gap with HyperTree, a fully unsupervised, scalable, and capacity-efficient framework for hierarchical representation learning in hyperbolic space, built on three contributions: (i) a geometrically consistent, closed-form definition of the hyperbolic Lowest Common Ancestor, valid across hyperbolic manifolds of arbitrary curvatures; (ii) a differentiable hierarchical objective based on Dasgupta's cost that is bounded and invariant to dataset scale; and (iii) a sample-complexity guarantee that enables scaling to datasets such as ImageNet-1k and iNaturalist-2018. Across four benchmarks, HyperTree outperforms Euclidean baselines by wide margins on hierarchical dendrogram purity, matches supervised hyperbolic baselines on CIFAR-100, and surpasses them on TinyImageNet at 16 dimensions — a highly capacity-efficient regime where competing methods collapse.
HypMoE-ReID: Hyperspherical Mixture-of-Experts for Large Scale Person Re-Identification
Kunlun Xu ⋅ Haotong Cheng ⋅ Yufei Guo ⋅ Jiahuan Zhou
With the rapid growth of surveillance data, Person Re-Identification (ReID) is evolving toward large scale scenarios, where datasets exhibit substantial increases in both size and diversity. Under this trend, the widely adopted dense backbone architectures, such as Transformers, suffer from an inherent limitation: Shared parameters couple all input tokens, forcing the model to learn generic representations that are suboptimal for heterogeneous data sources. Although Mixture-of-Experts (MoE) architectures offer a promising solution for handling diverse data, they encounter a fundamental challenge in ReID. Specifically, person images exhibit high inter-instance similarity at the coarse level, while discriminative cues lie in subtle fine-grained details that are easily overlooked. As a result, existing MoE methods often suffer from severe routing collapse, leading to insufficient expert specialization and limited diverse knowledge acquisition capacity. To address these problems, we propose a novel framework, termed Hyperspherical Mixture-of-Experts for Large Scale Person Re-Identification (HypMoE-ReID). Specifically, motivated by theoretical analysis in hyperspherical space, HypMoE-ReID incorporates two key components: (1) a Token Purification mechanism that suppresses noisy tokens detrimental to expert specialization, and (2) a Routing Diversity Learning strategy that explicitly separates experts on a unit hypersphere to enforce specialization and prevent collapse. Furthermore, to facilitate the investigation in large scale ReID, we construct a large-scale ReID benchmark comprising 809,701 images by unifying 13 public datasets with significant distribution gaps. Furthermore, extensive experiments demonstrate that HypMoE-ReID outperforms state-of-the-art MoE-based methods by at least 11.3\%/10.3\% in Average mAP/R@1, and surpasses strong dense backbone models by 2.2\%/2.0\% in Average mAP/R@1. Our code will be released.
IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder
Yitong Chen ⋅ Zijie Diao ⋅ Junke Wang ⋅ Lingyu Kong ⋅ Yixuan Ren ⋅ Bo He ⋅ Yu-Gang Jiang ⋅ Zuxuan Wu
Built on pretrained vision foundation models (VFMs), representation autoencoders (RAEs) have recently emerged as a promising approach for constructing semantically rich latent spaces for image generation. However, their reconstruction quality often remains suboptimal, largely because deep VFM representations do not preserve sufficient fine-grained visual detail. This limitation becomes even more severe after discretization, where missing low-level information is difficult to recover. In fact, we observe that shallow VFM features retain considerably richer local appearance and structural detail, which complements the high-level semantics carried by deep features used in existing RAEs. Motivated by this complementary property, we propose IDEAL, an In-DEpth ALignment framework for discrete representation autoencoding. By jointly aligning quantized tokens with both shallow and deep VFM features, IDEAL enables the resulting discrete visual tokens to preserve both visual fidelity and rich semantics. Extensive experiments demonstrate that IDEAL yields superior reconstruction performance, achieving $\mathbf{0.61}$ rFID on ImageNet and outperforming the previous best method by $\mathbf{0.28}$. When used for autoregressive image generation, IDEAL further produces a gFID of $\mathbf{1.89}$, establishing a new state of the art for autoregressive image generation.
Idempotency Exposes Consistency Problems in Sparse Autoencoders
Lennart Stöpler ⋅ Nikolai Bolik ⋅ Artur Andrzejak
The Linear Representation Hypothesis (LRH) posits that Transformer activations decompose into sparse linear combinations of interpretable directions, with Sparse Autoencoders (SAEs) serving as the predominant method for recovering such decompositions. We propose that idempotency -- the requirement that re-encoding an SAE's output reproduces the same latent representation -- is a desirable property for any faithful dictionary learning method, and demonstrate empirically that most SAEs deviate very far from this ideal. We identify two underlying failure modes: exploding representation norms and unstable active feature sets. Further analysis reveals that even individual SAE latents cannot be reliably reconstructed under re-encoding, and that SAE outputs exhibit significant sensitivity to input scaling. We show that idempotency can be improved via an auxiliary loss term and demonstrate its potential to improve SAE interpretability.
Identifying Structural Biases from Causal Mechanism Shifts
Praharsh Nanavati ⋅ Jilles Vreeken ⋅ David Kaltenpoth
Causal discovery methods commonly assume that all data is independently and identically distributed (i.i.d.) and that there are no unmeasured variables affecting the system. In practice, these assumptions are often violated, leading to inaccurate inference. In this paper, we study how to identify hidden confounding and selection biases from causal mechanism shifts. In particular, we show that structural biases lead to dependent mechanism shifts. That is, by considering for which variables the mechanisms change given data from different environments, we can tell which variables are unbiased, which are subject to hidden confounding, and which are undergoing selection bias. We formalize this into an empirically testable criterion based on mutual information, and show under which conditions it identifies structural biases. To tell which nodes are subject to what kind of bias, we introduce the StruBI algorithm. Experiments on synthetic and real-world data show that StruBI works well in practice, accurately recovering affected variable sets and types of biases, outperforming the state-of-the-art by a wide margin.
IIDiff: Learning Cross-Domain Frequency Transitions with Diffusion Mixture-of-Experts
Xin Xue ⋅ Yirong Xue ⋅ Xingcheng Fu ⋅ Yisen Gao ⋅ Lanhao Li ⋅ Lijun Sun ⋅ Yunhao Fu ⋅ Haoyi Zhou
Cross-domain time-series generation is a challenging task for diffusion models due to the strong requirement for domain adaptability. While frequency-based Mixture-of-Experts (MoE) methods have shown promise in domain adaptation, their direct integration with diffusion models faces critical issues from intra- and inter-domain. For intra-domain aspect, the absence of frequency transition compliance leads to imbalance pattern learning across denoising steps and degrade model performance. For the inter-domain aspect, the domain-specific transition variation bring difficulties for domain adaptation. To address these challenges, we propose IIDiff, a frequency-based diffusion MoE framework equipped with progressive frequency specializing and dynamic transition matching. The progressive specializing allows experts dynamically adjust frequency receptive fields to capture underlying patterns aligning with the intra-domain transition, enhancing the generation fidelity. The transition matching leverages the frequency noise distribution property to analysis the instantaneous transition process, enabling adaptive routing with domain-specific transition. Extensive experiments on 12 real-world datasets across 4 domains demonstrate our model's superiority. Compared with SOTA baselines, our model achieves an average 27.59\% improvement in KL divergence and 6.00\% improvement in MMD, confirming its robust domain adaptability and ability to generate high-fidelity cross-domain time-series.
Illusory Pattern Perception Drives Spurious Inference in Large Language Models
Peihua Mai ⋅ Zhuoyan Shao ⋅ Xinbao Qiao ⋅ Meng Zhang ⋅ Xinyue Zhou ⋅ Yan Pang
Illusory pattern perception is a well-documented human cognitive tendency to infer meaningful relationships in data that is actually random. Such a tendency, often described as “connecting the dots” where none exist, can result in systematic reasoning errors. This paper investigates whether Large Language Models (LLMs) exhibit such perceptual tendencies, which can lead to systematic errors in downstream applications. To our knowledge, this work presents the first systematic study of illusory pattern perception in LLMs, adapting classic psychological paradigms to three tasks with direct empirical comparison to human behaviors. We find that LLMs frequently exhibit stronger illusory pattern perception than humans. In particular, models tend to over-associate frequent positive attributes with majority groups or large organizations, and show increased tendencies to construct causal narratives from ambiguous events. To uncover the mechanism behind these behaviors, we develop a feature interpretability framework based on Sparse Autoencoders (SAEs) to analyze internal representations. Our results reveal that holistic frequency perception and analytic cognitive orientation are linked to the emergence of illusory perceptions. These findings highlight a previously underexplored cognitive-like illusion that may affect the reliability of LLM reasoning. Code available at Anonymous Github.
Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator
Zihan Wang ⋅ Seungjun Lee ⋅ Yinghao Xu ⋅ Gim Hee Lee
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into near 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale.
Image Matting without Matting-Specific Annotations via Eikonal Fields
Tao Huang ⋅ Yi Wang ⋅ Xiaohong Zhang ⋅ Wei Zhang ⋅ Tao Yin ⋅ Zhibin Zhang ⋅ Qiuli Wang
Image matting aims to estimate a continuous alpha matte that separates foreground objects from the background. However, existing deep matting methods still rely heavily on task-specific supervision for matting, including alpha mattes and trimaps, which are costly to annotate and limit their applicability in annotation-scarce scenarios. To address this limitation, we propose FreeMatte, a novel annotation-free framework that learns to predict high-quality alpha mattes without requiring any matting-specific annotations. The proposed framework formulates the construction of matting supervision as a foreground-driven wavefront propagation process. Starting from a coarse localization prior generated by an off-the-shelf segmentation foundation model, it solves an Eikonal equation constructed from the prior and the input image itself to obtain an arrival-time field that encodes how readily each pixel can be reached from the coarse foreground source. The arrival-time field is then translated into confidence-aware sparse supervision, which constrains the matting model to learn alpha prediction from the propagation structure encoded by the field. Extensive experiments across diverse benchmarks show that FreeMatte achieves performance competitive with task-supervised matting baselines under this challenging setting, ranking within the top two on 13 out of 20 evaluated dataset-metric pairs.
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Jaewoo Jung ⋅ Hyeonseo Yu ⋅ Honggyu An ⋅ Jisang Han ⋅ Mungyeom Kim ⋅ Minkyeong Jeon ⋅ Heeseong Shin ⋅ WonJun Moon ⋅ Federico Tombari ⋅ Daniel Barath ⋅ Marc Pollefeys ⋅ Seungryong Kim ⋅ Sunghwan Hong
Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry between viewpoints, and assemble a coarse 3D layout of the scene. Inspired by this process, we introduce Imagine3D-LLM, an MLLM that learns to assemble a similar compact 3D representation of the scene and conditions its answer on this representation. Concretely, we append a small set of learnable summary tokens after the image tokens, decode them into a compact 3D Gaussian Splatting representation supervised by a photometric reconstruction loss, and train jointly with the standard next-token prediction objective. Notably, although only the summary tokens receive direct reconstruction supervision, this objective also induces stronger cross-frame correspondence within the LLM's underlying image features, suggesting that learning to reconstruct propagates 3D-aware signals throughout the model. As a result, Imagine3D-LLM consistently outperforms prior approaches across multiple spatial reasoning and 3D understanding benchmarks, suggesting that imagining the scene can be more effective than being told its pixel-wise geometry. Our code and weights will be publicly released.
Impossibility of Distribution-Free Predictive Inference for Individual Treatment Effects
Chongguang Tao ⋅ Zheng Zhou ⋅ Yuhong Yang
Uncertainty quantification for individual treatment effects (ITEs) is a daunting challenge in causal inference. Motivated by recent advances in conformal prediction, several works aim to construct distribution-free prediction sets for ITEs with desired coverage under standard assumptions such as strong ignorability and overlap. In this paper, we show that such goals are fundamentally unattainable in the presence of continuous covariates. Specifically, we establish finite-sample and asymptotic impossibility results demonstrating that any distribution-free prediction set achieving desired coverage for ITEs must be trivial, in the sense that it has infinite expected length. Our analysis relies on a connection between ITE inference and the hardness of conditional independence testing, and highlights the intrinsic limitations imposed by the missing data nature of causal inference. These results provide a new perspective on existing methods, clarifying that their apparent success necessarily relies on additional structural assumptions beyond standard causal assumptions.
Improving Consistency in Retrieval Augmented Systems With Group Similarity Rewards
Faisal Hamman ⋅ Chenyang Zhu ⋅ Anoop Kumar ⋅ Xujun Peng ⋅ Sanghamitra Dutta ⋅ Daben Liu ⋅ Alfy Samuel
Retrieval-Augmented Generation (RAG) systems are increasingly deployed in high-stakes domains where users expect outputs to be consistent across semantically equivalent queries. However, existing systems often exhibit significant inconsistencies due to variability in both the retriever and generator, undermining trust and reliability. In this work, we focus on information consistency—the requirement that generated outputs convey the same core content across semantically equivalent inputs. We introduce a principled evaluation framework that decomposes RAG consistency into retriever-level, generator-level, and end-to-end components, enabling systematic identification of inconsistency sources. To improve consistency, we propose Paraphrased Set Group Relative Policy Optimization (PS-GRPO), an RL approach that leverages multiple rollouts across paraphrased set to assign group similarity rewards. We leverage PS-GRPO to achieve Information Consistent RAG (Con-RAG), training the generator to produce consistent outputs across paraphrased queries and remain robust to retrieval-induced variability. Because exact reward computation over paraphrase sets is computationally expensive, we also introduce a scalable approximation method that retains effectiveness while enabling efficient, large-scale training. Empirical evaluations across short-form, multi-hop, and long-form QA benchmarks demonstrate that Con-RAG significantly improves both consistency and accuracy over strong baselines, even in the absence of explicit ground-truth supervision. Our work provides practical solutions for evaluating and building reliable RAG systems for safety-critical deployments.
Improving Guidance-Free Visual Generation via Self-Contrastive Alignment for Likelihood Estimation
Jiwan Hur ⋅ DongJae Lee ⋅ Dongyeun Lee ⋅ Gyojin Han ⋅ Junmo Kim
Classifier-free guidance (CFG) is essential for improving sample quality in visual generation, but it incurs substantial sampling overhead. Existing guidance-free approaches halve this cost by approximating the CFG with a single model inference, but their performance is bounded by the CFG and inherits its limitations. In this paper, we propose self-contrastive guidance (SCG), which contrasts the model's prediction on real and self-generated samples to obtain a direct, model-based signal that pushes generation toward the true data manifold. Integrating SCG into the standard likelihood objective via reparameterization yields Self-Contrastive Alignment for Likelihood Estimation (SCALE), a guidance-free training framework that produces SCG-enhanced predictions in a single forward pass. Across class-conditional and text-to-image generation on diffusion and autoregressive models, SCALE consistently outperforms CFG-mimicking guidance-free baselines and improves further when combined with DPO, demonstrating that the self-contrastive signal yields gains complementary to both CFG and preference-based fine-tuning.
Improving Self-Supervised Vision Transformers with Cross Distillation
Siran Dai ⋅ Qianqian Xu ⋅ Peisong Wen ⋅ Yang Liu ⋅ XIAOCHUN CAO ⋅ Qingming Huang
Self-supervised Vision Transformers learn strong image-level representations, but their patch-level representations can be unreliable for dense prediction. We study this mismatch and find that the class token concentrates global semantics on a small number of representative patches, while dense prediction errors often occur on low-attention regions. We provide a local gradient analysis showing that, along the [CLS] attention readout, image-level self-distillation gradients are weighted by [CLS]-to-patch attention mass. To provide masked patches with global semantics without changing the architecture, we propose Cross Distillation (CODI), which mixes a small amount of the teacher class-token distribution into each teacher patch-token target. CODI gives every masked patch an attention-independent semantic correction while preserving the standard self-distillation architecture. Experiments on dense and image-level benchmarks show that CODI improves both dense and image-level performance under aligned pretraining settings.
Improving the Diffusability of Motion Tokenizer
Guanhe Huang ⋅ Songqiao Han ⋅ Tangzheng Lian ⋅ Oya Celiktutan
Latent diffusion models paired with a variational auto-encoder (VAE) tokenizer have shown promising performance for efficient text-driven human motion generation. However, motion latent diffusion suffers from a generation-reconstruction trade-off: simply increasing tokenizer capacity enhances reconstruction but fails to proportionally improve text-conditioned generation quality. We assume this trade-off stems from a spectral mismatch inside the VAE tokenizer: the encoder yields a latent space biased toward high frequencies, and the decoder introduces temporal artifacts. In this paper, we propose the Diffusable Motion Tokenizer (DiMoT) to improve the diffusability of the tokenizer, rather than scaling up the diffusion backbone. To improve text-conditioned optimization, Critic Spectral Adaptation (CSA) leverages a pretrained diffusion model as the critic to suppress high-frequency latent energy. To preserve motion dynamics, the Alpha-Flow Motion Decoder (AMD) performs generative decoding via average-velocity prediction, enabling efficient one-step decoding. Extensive experiments demonstrate that DiMoT achieves state-of-the-art performance across various challenging text-to-motion benchmarks, yielding superior motion fidelity and text-motion alignment. The code will be available upon publication.
Non-convex optimization is central to machine learning and AI, yet unlike its convex counterpart, it still lacks a general and principled framework. In this work, we try to highlight the importance of combining noisy perturbations with over-parametrization mathematically under the classic setting of matrix completion (MC). Matrix completion is an important non-convex recovery problem in which only a subset of a matrix’s entries is observed, and the remaining entries must be recovered using low-rank constraints. This problem is notoriously difficult because observations are limited. However, we show that by injecting $\epsilon$-level perturbations into the nullspace of the observation mask, we can construct noisy matrix sensing surrogates whose solutions remain $\mathcal{O}(\epsilon)$-close to the ground truth with valid restricted isometry property (RIP) constants. The existence of RIP further enables powerful tensor frameworks to be applied with theoretical guarantees. This approach reflects a broader empirical trend in which stochasticity or noise improves the curvature of the optimization landscape, allowing over-parameterized (a.k.a large) models to better separate signal from noise. Assuming each matrix entry is observed independently, we establish quantitative recovery guarantees without explicit requirements on sampling rate or incoherence, and discuss how such assumptions can further strengthen our results.
Inadvertent Context Leakage in Language Models
Jaiden Fairoze ⋅ Neal Mangaokar ⋅ Kamalika Chaudhuri ⋅ Sanjam Garg ⋅ Saeed Mahloujifar
For LLM agents to be useful beyond simple chat, they must hold sensitive user context such as calendars, credentials, health records, and financial data. We study whether the mere presence of such secrets in a model's context window introduces hidden correlations into the model's benign outputs, allowing reconstruction even when the model correctly refuses direct extraction. We further study whether an adversary can actively engineer prompts that amplify this effect, using the model as a covert carrier to transmit secrets through seemingly innocuous text. In both cases, this limited leakage is exploited using a novel adaptive attack that assumes black-box access to the underlying model. In controlled experiments across eight proprietary models, we find that 2-digit in-context secrets are reconstructed with near-perfect accuracy and 4-digit secrets at 82% exact match, all from outputs the model produces in response to ordinary, non-adversarial requests. More capable models leak more: stronger instruction-following amplifies sensitivity to in-context secrets, making leakage a byproduct of capability rather than a bug to be patched. We show this leakage enables two practical attacks: (1) a trained classifier that infers semantic predicates about user memories (e.g., health conditions, financial events) from routine natural-language outputs, and (2) an RL-trained adversary that extracts full Social Security Numbers from a production-style agent at 100% success.
INEUS: Iterative Neural Solver for High-Dimensional PIDEs
Jean-Loup Dupret ⋅ Davide Gallon ⋅ Patrick Cheridito
In this paper, we introduce INEUS, a meshfree iterative neural solver for partial integro-differential equations (PIDEs). The method replaces the explicit evaluation of nonlocal jump integrals with single-jump sampling and reformulates PIDE solving as a sequence of recursive regression problems. Like Physics-Informed Neural Networks (PINNs), INEUS learns global solutions over the entire space-time domain, yet it offers a more efficient treatment of nonlocal terms and avoids the computationally expensive differentiation of full PIDE residuals. These features make INEUS particularly well suited for high-dimensional PDEs and PIDEs. Supported by a contraction-based convergence proof for linear PIDEs, our numerical experiments show that INEUS delivers accurate and scalable solutions for various high-dimensional linear and nonlinear examples.
Inexact Bregman Sparse Newton Method for Efficient Optimal Transport
Jianting Pan ⋅ Jian Li ⋅ Sirong Dai ⋅ Ming Yan
Computing exact Optimal Transport (OT) distances for large-scale datasets is computationally prohibitive. While entropy-regularized alternatives offer speed, they sacrifice precision and frequently suffer from numerical instability in high-accuracy regimes. To address these limitations, we propose the Inexact Bregman Sparse Newton (IBSN) method, which efficiently solves the exact OT problems. Our approach utilizes a Bregman proximal point framework through a sequence of semi-dual subproblems. By solving these subproblems inexactly, we significantly reduce per-iteration complexity while maintaining a theoretical guarantee of convergence to the true optimal plan. To further accelerate the algorithm, we develop a sparse Newton-type solver for the subproblem and employ a Hessian sparsification strategy that drastically lowers memory and time costs without sacrificing accuracy. We provide rigorous theoretical guarantees for the global convergence of the algorithm. Extensive experiments demonstrate that IBSN consistently outperforms state-of-the-art methods in both computational speed and solution precision.
Inference-Time Self-Aligned Drifting for Few-Step Flow Matching
Shigui Li ⋅ JIAN XU ⋅ Wei Chen ⋅ Junmei Yang ⋅ Fei Wang ⋅ Daiguo Zhou ⋅ Delu Zeng
Flow Matching (FM) provides an efficient formulation for continuous-time generative modeling via straight optimal transport (OT) trajectories. However, during accelerated sampling, these models face a fundamental variance-fidelity dilemma. Under few-step deterministic Ordinary Differential Equation (ODE) integration, the rigid OT trajectories suffer from mean concentration, leading to severe variance collapse and the loss of high-frequency details. Conversely, traditional Stochastic Differential Equations (SDEs) inject blind, exogenous Brownian noise to recover variance, but this isotropic exploration incurs inherent numerical attrition over long integration paths, bottlenecking peak fidelity. To resolve this, we introduce Self-Aligned Drifting (SADrift) to mitigate these challenges without external priors. SADrift is an inference-time framework that augments flow matching inference with self-interacting dynamics, without additional training. By maintaining a momentum history of the sampling trajectory, SADrift extracts the \emph{kinematic residual} to generate a deterministic extrapolative field. This mechanism forces the trajectory to efficiently explore the local tangent space, effectively decoupling exploratory variance into targeted deterministic dispersion and a minimal stochastic buffer. To safely anchor these dispersed trajectories and prevent out-of-distribution divergence, the residual drift is analytically integrated into a micro-time Ornstein-Uhlenbeck (OU) process, providing a mathematically sound, mean-reverting geometric safeguard. Requiring exactly zero additional Neural Function Evaluations (NFE), SADrift empowers a frozen base model to achieve a highly competitive FID of 2.00 at just 20 NFE on ImageNet-256. By successfully bypassing the numerical attrition of the 250-step SDE baseline (FID 2.06), SADrift establishes a new Pareto frontier for variance injection. Beyond image synthesis, we further validate SADrift on fast video generation paradigms (e.g., TurboDiffusion with Wan2.1-1.3B), demonstrating its universal capability to unlock latent generative quality across different modalities and acceleration regimes.
Inferring learning rules in deep neural network architectures from animal learning data
Shaunak Bhandarkar ⋅ Jonathan Pillow
A major goal in computational neuroscience is to understand how animals learn to perform a new behavior from a sequence of actions and rewards. Recent work from Liebana et al. (2025) has argued that animal learning trajectories are inconsistent with learning in shallow (1-layer) networks and are better explained by gradient descent (GD) in a deep neural network (DNN). However, previous work on inferring animal learning rules from behavioral data has focused on either 1-layer networks or fits to group behavior, highlighting the need for methods to fit DNN learning rules to individual animal learning trajectories. To address this gap, we develop a method for inferring single-animal DNN-based learning rules, parametrized by a set of initial weights, a learning rate, and a softmax parameter governing stochasticity in choice behavior. We first prove that GD in a shallow network can be arbitrarily well approximated by GD in a deep network, establishing a theoretical limit on network identifiability. Using this result to refine our inference method, we analyze two behavioral datasets from mice learning a sensory decision-making task. Specifically, we compare three model architectures: 1-layer linear (1L) and 2-layer linear (2L) models considered in Liebana et al., as well as a 2-layer nonlinear network (2NL) with a sigmoidal nonlinearity after the first layer. The 2NL model achieved the best fit for 84 out of 88 mice in both datasets (by AIC). Moreover, simulated learning trajectories from the fitted 2NL model closely matched the timecourse of key behavioral metrics such as bias, stimulus sensitivity, and accuracy. These results show that both depth and nonlinearity are critical for capturing individual learning trajectories. Strikingly, we found that the inferred learning rate and initial layer-1 weights of the 2NL model systematically predicted individual differences in accuracy and task bias, linking network initialization and learning rate to individual differences in learning performance across mice.
Information Bottleneck-Guided Adaptive Hypergraph Transformer for Brain Disease Diagnosis
Jingxi Feng ⋅ Xudong Chen ⋅ Yifan Zhang ⋅ Heming Xu ⋅ Hongcheng Han ⋅ Xijing Wang ⋅ Dong Zhang ⋅ Shaoyi Du
Exploring high-order correlations and long-range dependencies in brain networks holds significant value for both neuroscience research and clinical diagnosis. However, previous studies have lacked a unified integration of high-order and long-range dependency information in brain networks, and there is substantial redundancy behind various types of information. These issues limit their effectiveness in the diagnosis of brain diseases. To address this, we propose an Information Bottleneck- Guided Adaptive HyperGraph Transformer (IBAHGT). By incorporating the information bottleneck (IB) principle, this approach enables adaptive learning of high-order correlations and both short- and long-range dependencies within a unified framework for brain network analysis, achieving high-precision brain disease diagnosis. IBAHGT consists of three key components: an information bottleneck- guided adaptive hypergraph convolution, which introduces a novel hypergraph information bottleneck (HIB) principle to adaptively learn hypergraph message- passing weights between nodes and hyperedges, optimizes information flow and captures high-order information in brain networks that is maximally informative and minimally redundant (MIMR). The Transformer encoder captures global information within brain networks through the attention mechanism, specifically modeling short- and long-range dependencies. An information bottleneck-guided node-level adaptive fusion employs the IB principle to learn independent weights for each node, facilitating the fine-grained integration of high-order information and global information to obtain an efficient representation for downstream tasks. Extensive experiments demonstrate that the proposed method outperforms current state-of-the-art methods and can identify biomarkers for clinical applications.
Information Parity for Code: The Scope of Transfer in Multilingual Code Models
Alexander Tsvetkov ⋅ Alon Kipnis
LLMs exhibit uneven performance across programming languages, yet benchmark gaps often reflect narrow task regimes and evaluation-specific confounds. This makes it hard to tell whether poor performance reflects a broader weakness with a given language, a broader difficulty on a given task, or a limitation specific to that task-language combination. We adapt Information Parity, a measure of relative cross-language representational efficiency developed for natural language, to code. We define Code Information Parity (Code IP) on a controlled parallel corpus of 22 algorithms across 20 languages, audited to preserve core semantics while minimizing boilerplate and excluding library shortcuts. The results are regime-specific: Code IP correlates strongly with multilingual benchmark performance on short-horizon generation ($\rho$ from $-0.55$ to $-0.81$), remains negative though moderated on execution prediction ($\rho=-0.429$), and disappears on whole-program synthesis ($\rho=+0.105$). Rather than a universal metric of multilingual coding ability, Code IP isolates a reliable cross-language signal on short-horizon generation and delineates where that signal no longer separates cleanly from broader task and evaluation effects.
Classical initialization theory analyzes networks through second-order observables such as variance, correlation, and Jacobian conditioning that are natural when weights are symmetric around zero. Identity-plus-noise initialization, which initializes the network to behave locally like a residual block, sits outside this picture: variance can be preserved while coordinate-wise signal identity is destroyed, making such second-order quantities indirect indicators of trainability. We identify the terminal sign-flip rate as the natural forward-pass observable for this regime, and derive its critical scale $\sigma \sim L^{-1/2}$ from a dimensionless margin--sensitivity ratio, rather than postulating it. The retention curve emerges as a first-passage limit of a signed-margin process, yielding a calibration procedure that targets a chosen sign-flip rate at any depth. Experiments on synthetic and real benchmarks suggest that this calibrated initial rate behaves as a training-time parameter rather than an inert calibration detail.
Information-Theoretic Generalization for Set-Input Optimization-Valued Objectives
Futoshi Futami ⋅ Masahiro Fujisawa
Many modern learning criteria, such as multicalibration, CVaR, and distributionally robust objectives, are set-input and optimization-valued: their empirical values depend on the whole evaluation batch and on an auxiliary optimizer selected after observing that batch. This optimization mismatch falls outside the standard information-theoretic generalization analysis for averages of fixed per-example losses. We provide a unified supersample-based analysis for this class of objectives. Under a selector-gap concentration condition, we show that the optimized validation--training gap is controlled by the conditional mutual information between the learned supersample predictions and the selector, without an explicit complexity penalty for the auxiliary optimizer class. We verify the concentration condition through stability and convex L_2-Lipschitz certificates, obtaining concrete bounds for multicalibration, CVaR, and Cressie--Read DRO, including a single-index refinement under stability.
Inner Product Aware Quantization: Provably Fast, Accurate, and Adaptive Algorithms
Nathan White ⋅ Krish Singal
Quantization is a fundamental tool used to compress datasets, neural network weights, and memory usage in a range of computational tasks. Many downstream applications of vector quantization perform inner products with arbitrary inputs. This motivates the study of inner product aware quantization schemes that approximately preserve inner products with unseen vectors -- in contrast to simply minimizing the mean-squared error. In this work, we formulate objectives that capture natural desiderata and develop adaptive and unbiased quantization methods that approximately preserve inner products with worst-case and average case inputs. An analysis of these objectives shows a tight connection with the well-studied notion of Adaptive Stochastic Quantization (ASQ). We develop provably fast exact and approximate algorithms for our objectives. Our theoretical results inspire efficient practical algorithms that perform well across a variety of workload distributions. They also lead to practical algorithms for standard ASQ which are 2-10$\times$ faster than prior state-of-the-art methods while maintaining quality. These theoretical and empirical results contribute towards making adaptive quantization techniques more efficient and tractable in practical settings.
InSpect: A Curated Natural History Collection Dataset for Insect Specimen Understanding
Junhao Zhang ⋅ Feng Xu ⋅ Ahalya Ravendran ⋅ Nicole Fisher ⋅ Brendan Tidd ⋅ Lars Petersson ⋅ Dadong Wang ⋅ Xun Li
Visual insect specimen understanding aims to identify biological taxa and characterize morphology-relevant structures from specimen images. Existing insect benchmarks have advanced this problem through large-scale pretraining and multimodal learning, but remain centered on taxonomic recognition with strong auxiliary modalities. We introduce \textbf{InSpect}, a curated natural history collection dataset containing 48,184 digitized insect specimen images with aligned crops, hierarchical taxonomy, label-derived structured metadata, and 1,859 images with fine-grained anatomical part masks. We thus establish two benchmark settings. The first is \textbf{open taxonomic recognition}, which evaluates zero-shot, fine-tuned, and unseen-taxon recognition using cropped insect images, label-derived metadata, or their combination. The second is \textbf{fine-grained anatomical part segmentation}, which evaluates supervised and text-guided segmentation of insect structures such as antennae, legs, wings, and body regions. Experiments show that open taxonomic recognition remains challenging despite strong closed-set performance, and that specimen metadata provides useful but unstable contextual cues. For anatomical segmentation, closed-set segmentation models remain weak on thin structures, and zero-shot open-vocabulary models often fail to distinguish insect parts from the body or background. Together, these results show that InSpect enables systematic evaluation of insect specimen understanding across cropped visual morphology, label-derived context, and fine-grained anatomical annotations.
INSPO : Unlocking Intrinsic Self-Reflection for LLM Preference Optimization
Yu Li ⋅ Tian Lan ⋅ Zhengling Qi
Direct Preference Optimization (DPO) and its variants have become the standard for aligning Large Language Models (LLMs). However, we identify two fundamental limitations. First, the optimized policy lacks invariance since it varies with modeling choices such as scalarization function or reference policy, whereas an optimal policy should remain invariant. Second, most existing methods yield theoretically suboptimal policies by not fully exploiting the comparative information in pairwise preference data, thus missing an opportunity for self-reflection through comparing and contrasting responses. To address both limitations, we propose Intrinsic Self-reflective Preference Optimization (InSPO), which derives a globally optimal policy conditioned on both context and alternative response, explicitly formalizing self-reflection. We prove this formulation surpasses standard DPO and RLHF targets while guaranteeing invariance. InSPO serves as a plug-and-play enhancement for DPO-family algorithms, decoupling alignment from modeling constraints without architectural changes. Using privileged information learning, InSPO requires no alternative response at inference since the self-reflective mechanism is distilled during training, incurring zero overhead. Experiments show InSPO consistently improves win rates and length-controlled metrics across DPO variants, yielding more robust and human-aligned LLMs. Code is available here.
Instruction Anchor: Dissecting the Mechanistic Dynamics of Modality Arbitration
Yu Zhang ⋅ Mufan Xu ⋅ Xuefeng Bai ⋅ Kehai Chen ⋅ Pengfei Zhang ⋅ Yang Xiang ⋅ Min zhang
Modality following is the ability to selectively leverage multimodal contexts based on user instructions. It is fundamental to the safety and reliability of multimodal large language models (MLLMs) in real-world deployments. However, the internal mechanisms governing this decision-making process remain largely under-explored. In this work, we investigate the mechanism underlying modality following through an information flow perspective. Our findings reveal that instruction tokens serve as structural anchor for modality arbitration: Shallow attention layers perform undifferentiated information transfer, routing multimodal cues to instruction tokens as a latent buffer; in contrast, deep attention layers selectively strengthen the instruction-compliant subspace and resolve modality arbitration according to the instruction-specified intent, with a sparse subset of attention heads driving this process. Targeted attention-head interventions further validate the functional specificity of these heads: blocking only $5\%$ of the identified heads substantially degrades modality following while preserving general visual and language capabilities, whereas targeted amplification can restore failed modality-following samples by up to approximately $60\%$. Together, this work provides a mechanistic account of modality following and informs future efforts to improve how MLLMs integrate and utilize multimodal evidence under user instructions. \footnote{All data and code will be released upon acceptance.}
Integrated Imputation-Classification for Supervised Learning with Missing Data
Yue Liu ⋅ Ben Liang ⋅ Ali Tizghadam ⋅ Ilijc Albanese
We study supervised classification problems with missing feature values. Existing approaches often decouple imputation from classification, producing imputations that may be plausible but uninformative for prediction. Instead, we propose the Integrated Imputation and Classification Network (IICN), which jointly trains an imputer and an ${(n{+}1)}$-classdiscriminator adversarially with a single class supervised classification objective, where the discriminator learns to distinguish among the $n$ true classes and an additional ``imputed" class. We prove that at the global optimum, the imputer and discriminator together implement marginalization over missing coordinates and yield a Bayes-optimal classifier. We evaluate IICN on FashionMNIST, CIFAR-10, and tabular datasets with naturally occurring missingness. IICN outperforms classical impute-then-classify pipelines and recent generative baselines, showing strong robustness and accuracy in challenging settings.
Interleaved Head Attention
Sai Surya Duvvuri ⋅ Chanakya Ekbote ⋅ Rachit Bansal ⋅ Rishabh Tiwari ⋅ Devvrit Khatri ⋅ David Brandfonbrener ⋅ Paul Liang ⋅ Inderjit Dhillon ⋅ Manzil Zaheer
Multi-Head Attention (MHA) is the core computational primitive underlying modern Large Language Models (LLMs). However, within a single attention layer, MHA has an intrinsic linear scaling limitation: $H$ attention heads produce exactly $H$ independent attention matrices, with no communication between heads during attention computation. While deep Transformers can compose information across layers, this places much of the burden of building and combining intermediate relations on depth, potentially requiring additional layers to expose interactions useful for downstream computation. Our focus is complementary: increasing the compositional bandwidth available within each attention layer. To this end, we propose Interleaved Head Attention (IHA), which enables within-layer cross-head mixing by constructing $P$ pseudo-heads per head, typically with $P=H$, where each pseudo query, key, and value is a learned linear combination of all $H$ original queries, keys, and values, respectively. Interactions between pseudo-query and pseudo-key heads induce up to $P^2$ attention patterns per head with modest parameter overhead $\mathcal{O}(H^2P)$. We provide theory showing improved parameter efficiency on Polynomial Filters, where IHA uses $\Theta(\sqrt{k}n^2)$ parameters versus $\Theta(kn^2)$ for MHA, and on the order-sensitive CPM-3 task, where IHA uses $\lceil\sqrt{N_{max}}\rceil$ heads versus $N_{max}$ for MHA. On real-world benchmarks, IHA improves Multi-Key retrieval on RULER by 10–20% at 4k–16k context lengths and, after OpenThoughts reasoning fine-tuning, improves GSM8K by 5.8% and MATH-500 by 2.8% under majority voting over full attention, with modest throughput overhead.
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
Bohan Hou ⋅ Jiuning Gu ⋅ Jiayan Guo ⋅ Ronghao Dang ⋅ Sicong Leng ⋅ Xin Li ⋅ Xuemeng Song ⋅ Jianfei Yang
Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or treated as an answer endpoint rather than part of an interleaved search trajectory. We introduce \textbf{InterLV-Search}, a benchmark for Interleaved Language-Vision Agentic Search, in which textual and visual evidence is repeatedly used to condition later search. It contains 2,061 examples across three levels: active visual evidence seeking, controlled offline interleaved multimodal search, and open-web interleaved multimodal search. Beyond existing benchmarks, it also includes multimodal multi-branch samples that involve comparison between multiple entities during the evidence search. We construct Level 1 and Level 2 with automated pipelines and Level 3 with a machine-led, human-supervised open-web pipeline. We further provide InterLV-Agent for standardized tool use, trajectory logging, and evaluation. Experiments on proprietary and open-source multimodal agents show that current systems remain far from solving interleaved multimodal search, with the best model below 50% overall accuracy, highlighting challenges in visual evidence seeking, search control, and multimodal evidence integration. We release the benchmark data and evaluation code at \url{https://anonymous.4open.science/r/InterLV-Search-F423}
Intern-Atlas: A Methodological Evolution Graph as Research Infrastructure for AI Scientists
Yujun Wu ⋅ Dongxu Zhang ⋅ Xinchen Li ⋅ Jinhang Xu ⋅ Yiling Duan ⋅ Yumou Liu ⋅ Jiabao Pan ⋅ Qiyuan Zhu ⋅ Xuanhe Zhou ⋅ Jingxuan Wei ⋅ Siyuan Li ⋅ Jintao Chen ⋅ Conghui He ⋅ Cheng Tan
Existing research infrastructure is fundamentally document-centric, providing citation links between papers but lacking explicit representations of methodological evolution. In particular, it does not capture the structured relationships that explain how and why research methods emerge, adapt, and build upon one another. With the rise of AI-driven research agents as a new class of consumers of scientific knowledge, this limitation becomes increasingly consequential, as such agents cannot reliably reconstruct method evolution topologies from unstructured text. We introduce Intern-Atlas, a methodological evolution graph that automatically identifies method-level entities, infers lineage relationships among methodologies, and captures the bottlenecks that drive transitions between successive innovations. Built from $1{,}030{,}314$ papers spanning AI conferences, journals, and arXiv preprints, the resulting graph comprises $9{,}410{,}201$ semantically typed edges, forming a queryable causal network of methodological development. To operationalize this structure, we further propose a self-guided temporal tree search algorithm for constructing evolution chains that trace the progression of methods over time. We evaluate the quality of the resulting graph against expert-curated ground-truth evolution chains and observe strong alignment. In addition, we demonstrate that Intern-Atlas enables downstream applications in idea evaluation and automated idea generation. We position methodological evolution graphs as a foundational data layer for the emerging field of automated scientific discovery.
Interpretable Relational Inference with LLM-Guided Symbolic Dynamics Modeling
Xiaoxiao Liang ⋅ Juyuan Zhang ⋅ Liming Pan ⋅ Linyuan Lü
Inferring latent interaction structures from observed dynamics is a fundamental inverse problem in many-body interacting systems. Most neural approaches rely on black-box surrogates over trainable graphs, achieving accuracy at the expense of mechanistic interpretability. Symbolic regression offers explicit dynamical equations and stronger inductive biases, but typically assumes known topology and a fixed function library. We propose $\textbf{COSINE}$ ($\textbf{C}$o-$\textbf{O}$ptimization of $\textbf{S}$ymbolic $\textbf{I}$nteractions and $\textbf{N}$etwork $\textbf{E}$dges), a differentiable framework that jointly discovers interaction graphs and sparse symbolic dynamics. To overcome the limitations of fixed symbolic libraries, COSINE further incorporates an outer-loop large language model that adaptively prunes and expands the hypothesis space using feedback from the inner optimization loop. Experiments on synthetic systems and large-scale real-world epidemic data demonstrate robust structural recovery and compact, mechanism-aligned dynamical expressions. Code: https://anonymous.4open.science/r/COSINE-6D43.
Intrinsically Interpretable Attention via Sparse Post-Training
Florent Draye ⋅ Anson Lei ⋅ Hsiao-Ru Pan ⋅ Ingmar Posner ⋅ Bernhard Schölkopf
We introduce a simple post-training method that makes transformer attention sparse without sacrificing performance. Applying a flexible sparsity regularisation under a constrained-loss objective, we show on models up to 7B parameters that it is possible to retain the original pretraining loss while reducing attention connectivity to $\approx 0.4$% of its edges. Unlike sparse-attention methods designed for computational efficiency, our approach leverages sparsity as a structural prior: it preserves capability while exposing a more organized and interpretable connectivity pattern. We find that this local sparsity cascades into global circuit simplification: task-specific circuits involve far fewer components (attention heads and MLPs) with up to 100× fewer edges connecting them. Additionally, using cross-layer transcoders, we show that sparse attention substantially simplifies attention attribution, enabling a unified view of feature-based and circuit-based perspectives. These results demonstrate that transformer attention can be made orders of magnitude sparser, suggesting that much of its computation is redundant and that sparsity may serve as a guiding principle for more structured and interpretable models.
Intrinsic-dimension empirical Bernstein inequalities for bounded self-adjoint operators
Diego Martinez Taboada ⋅ Aaditya Ramdas
Operator-valued concentration inequalities are foundational to the analysis of modern high-dimensional statistics and randomized algorithms. However, standard oracle bounds are frequently limited in practice: they require explicit a priori knowledge of the true variance, and often explicitly scale with the ambient dimension, rendering them vacuous for infinite-dimensional or heavily structured operators. Motivated by these challenges, we establish the first empirical Bennett and Bernstein inequalities for sums of independent, bounded, compact self-adjoint operators. Our fully data-driven bounds replace the unknown variance with an empirical estimate and rely strictly on the intrinsic dimension rather than the ambient dimension. This structural shift yields computable, dimension-free guarantees that are strictly sharper for non-isotropic random matrices and seamlessly extend to infinite-dimensional Hilbert spaces. We demonstrate that our empirical bounds achieve asymptotic sharpness with the best known oracle rates. Finally, as an independent byproduct, we derive novel empirical concentration guarantees for the intrinsic dimension itself.
Iterative Chow Filtering for Learning with Distribution Shift
Gautam Chandrasekaran ⋅ Georgios Gkrinias ⋅ Adam Klivans ⋅ Konstantinos Stavropoulos ⋅ Arsen Vasilyan
Recent work due to Goel et al.~gave the first efficient algorithms for learning with distribution shift in the challenging PQ framework, where a learner receives labeled training examples and unlabeled test examples and must make correct predictions on the test set but is allowed to abstain from predicting on out-of-distribution points. Their results rely on ${\cal L}_2$ sandwiching approximations, a strong requirement that leads to poor bounds for several basic function classes such as DNF formulas. Here we show that the weaker notion of ${\cal L}_1$ sandwiching suffices for efficient PQ learning. As a consequence, we obtain the first quasipolynomial-time PQ learning algorithm for DNFs under the uniform distribution and essentially match the guarantees known for ordinary PAC learning. More broadly, our bounds provide exponential improvements for several classes including constant depth circuits and constant degree polynomial threshold functions. Our main technical ingredient is *Iterative Chow Filtering*, a new procedure that uses low-degree Chow parameters to identify and remove test points incompatible with the training distribution.
Jacobian Scopes: A Unified Geometric Framework for Token-Level LLM Attributions
Toni Liu ⋅ Baran Zadeoğlu ⋅ Nicolas Boulle ⋅ Raphaël Sarfati ⋅ Gurbir Arora ⋅ Christopher Earls
Gradient-based attribution methods for large language models (LLMs) have so far been limited to explaining scalar outputs - the logit of a single target token. Yet LLM predictions are inherently distributional, and many practically important questions concern the full predictive distribution: which tokens make the model uncertain? Which drive a broad range of plausible continuations? We propose **Jacobian Scopes**, a unified framework that fills this gap by projecting the input-to-output Jacobian onto different directions in output space via a single vector-Jacobian product, requiring only one backward pass. This yields three complementary methods: **Semantic Scope** attributes a specific target logit; **Fisher Scope**, grounded in information geometry, identifies tokens that most alter the overall shape of the predicted distribution; and **Temperature Scope** traces which tokens govern the model's predictive confidence. Fisher and Temperature Scopes are, to our knowledge, the first attribution methods to natively target distributional rather than pointwise features of LLM predictions. Through case studies spanning instruction following, translation, and in-context time-series forecasting, Jacobian Scopes reveal implicit political biases, uncover word- and phrase-level translation strategies, and illuminate the nearest-neighbor pattern-matching mechanisms underlying LLM forecasting. Quantitative evaluation on LAMBADA and IWSLT2017 across six leading LLMs (LLaMA-3.2, Qwen2.5, Gemma-3) confirms that Jacobian Scopes consistently match or outperform Input $\times$ Gradient and Integrated Gradients, at a fraction of the latter's computational cost.
Kairos: Toward Adaptive and Parameter-Efficient Time Series Foundation Models
Kun Feng ⋅ Shaocheng Lan ⋅ Yuchen Fang ⋅ Wenchao He ⋅ Sihan Lu ⋅ Shuqi Gu ⋅ Lintao Ma ⋅ Xingyu Lu ⋅ Kan Ren
Inherent temporal heterogeneity, such as varying sampling densities and periodic structures, has posed substantial challenges in zero-shot generalization for Time Series Foundation Models (TSFMs). Existing TSFMs predominantly rely on massive parameterization to absorb such heterogeneity, as their static tokenization and positional encoding schemes entangle diverse temporal patterns into a fixed representation space, encouraging memorization rather than adaptation. To address this limitation, we propose Kairos, a flexible and parameter-efficient TSFM dedicated to forecasting tasks, which decouples temporal heterogeneity from model capacity through a novel tokenization perspective. Kairos introduces a dynamic patching tokenizer and a mixture-of-size encoding that adapt observational granularity to local information density, enabling fine-grained temporal abstraction without increasing model width or depth. In addition, we design a multi-granularity positional embedding based on dynamic rotary encodings, which conditions on instance-level spectral features and temporal structure induced by dynamic patching tokenization, allowing robust modeling of diverse temporal dependencies. Trained on a novel Predictability-Stratified Time-Series (PreSTS) corpus, Kairos achieves superior zero-shot performance with substantially fewer parameters on two mainstream benchmarks, GIFT-Eval and Time-Series-Library.
KL for a KL: On-Policy Distillation with Control Variate Baseline
Minjae Oh ⋅ Sangjun Song ⋅ Gyubin Choi ⋅ Yunho Choi ⋅ Yohan Jo
On-Policy Distillation (OPD) has emerged as a dominant post-training paradigm for large language models, especially for reasoning domains. However, OPD remains unstable in practice due to the high gradient variance of its single-sample Monte Carlo estimator, and recipes for stable training are still immature. We propose vOPD (On-Policy Distillation with a control variate baseline), which casts OPD as policy-gradient RL and stabilizes it by introducing a control variate baseline—canonically a value function—from the RL literature. We show that the OPD value function admits a closed form as the per-token negative reverse KL divergence between the student and the teacher, available directly from the already-computed forward pass with no additional critic or inference. Existing stabilization methods either compute the full token-level reverse KL over the entire vocabulary, adding significant overhead, or restrict it to a top-k support, biasing the objective. vOPD instead preserves the lightweight single-sample estimator, subtracting the value function as a detached baseline to keep the gradient unbiased while reducing variance. Furthermore, we show that a top-k approximation of the baseline further lowers cost without compromising performance. Across mathematical and scientific reasoning benchmarks, vOPD consistently outperforms vanilla OPD and matches the most expensive full-vocabulary baseline, offering an efficient stabilization of On-Policy Distillation through principled RL variance reduction.
KNOT: A Knowledge Entanglement Benchmark for Robust Unlearning Evaluation
Mengyang Li ⋅ Jingwen Wang ⋅ Yu Zhang ⋅ Shuang Liu ⋅ Zhong Zhang
Large language model (LLM) unlearning aims to selectively remove specific knowledge while preserving model utility. Current benchmarks such as TOFU, WMDP, and MUSE evaluate unlearning on semantically disjoint forget and retain sets, making the task artificially easy. We formalize this gap through a taxonomy of knowledge entanglement at three levels: entity, concept, and skill. Based on this taxonomy, we construct KNOT, a benchmark with approximately 5,300 forget, 9,800 retain, and 5,000 boundary QA pairs, each annotated with a quantitative entanglement score. Evaluation of seven methods on KNOT reveals that (a) all methods degrade severely under high entanglement, (b) the coreset effect reported on standard benchmarks vanishes, and (c) existing metrics miss critical boundary behavior. We propose Boundary Precision and Boundary Recall to fill this gap. Cross-benchmark comparison confirms that KNOT's entanglement, rather than surface difficulty, drives the observed degradation.
KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers
Yann Bouquet ⋅ Alireza Khodamoradi ⋅ Kristof Denolf ⋅ Mathieu Salzmann
Post-training quantization (PTQ) of diffusion transformers (DiTs) to W4A4 severely degrades output quality, because activations entering each linear layer contain outliers that 4-bit formats cannot represent. The standard fix applies an invertible linear transform to the activations and its inverse to the weights before quantizing both. Normalization layers between blocks force this transform to run online at every denoising step, making its inference computation cost the binding design constraint. Existing options trade quantization quality for inference cost: per-channel scaling (SmoothQuant) is computationally cheap but impacts the magnitude of the channels, which can harm quantization accuracy; fixed Hadamard transforms yield better quantization accuracy but require large block sizes that incur a high online cost; learned full-$d$ invertible transforms calibrate best but entail an prohibitive dense $d \times d$ matrix multiplication (GEMM) per layer per step. We propose KroQuant, a PTQ method that applies a learned Kronecker-structured invertible transform to each 32-element block of the activation, storing less than half the parameters of per-channel scaling. The block-local structure runs as small tensor-core GEMMs, and on an MI350 GPU the KroQuant quantizer kernel is up to $14$\% faster than the SmoothQuant kernel. Offline LoRaQ weight calibration then absorbs the residual per-weight quantization error. On PixArt-$\Sigma$, SANA, and FLUX.1-schnell at W4A4 (MXFP4e2), KroQuant produces outputs closer to the FP reference than SVDQuant and LoRaQ on MJHQ-30K and SDCI, while preserving or improving image quality.
KV-COBRA: KV Cache Compression via Co-Optimized Bit-Rank Allocation
Sihyeon Ha ⋅ Jaeho Lee ⋅ Yo-Seb Jeon
What limits KV-cache compression at extreme bit-rates? We argue it is not the choice of compression scheme but how its budget is allocated across attention heads. Existing methods apply rank and bit-width uniformly, ignoring that each head has a different optimal mix of rank truncation and quantization. We show that co-optimizing rank and bit-width per head — using only standard low-rank projection and scalar quantization — dominates uniform allocation, with the largest gains at low bit-rates. Our method, KV-COBRA (Co-Optimized Bit-Rank Allocation), formalizes this as a resource-allocation problem: it balances rank-truncation loss against quantization loss within each head, then redistributes budget across heads to minimize total distortion. A fused Hadamard rotation equalizes per-channel variance, and reordering the SVD basis by attention-KL importance makes the solver query-aware. The same allocator extends to joint $K{+}V$ compression. On perplexity, zero-shot, and long-context benchmarks from $0.5$ to $4$ bits per dimension (bpd), KV-COBRA shows the smallest accuracy degradation among evaluated methods at low bpd, with no per-token overhead.
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
Andreas Hochlehnert ⋅ Marianna Nezhurina ⋅ Mehdi Cherti ⋅ Andrej Radonjic ⋅ Thaddäus Wiedemer ⋅ Christoph Schuhmann ⋅ Romain Beaumont ⋅ Wieland Brendel ⋅ Bernhard Schölkopf ⋅ A. Sophia Koepke ⋅ Jenia Jitsev ⋅ Matthias Bethge
We present LAION-BVD (LAION - Big Video Dataset), a large-scale open video dataset for multimodal learning, containing 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across video, audio, and image modalities. Using content-aware scene detection we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as a new source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. Overall, LAION-BVD significantly expands open access to multimodal videos at unprecedented scale and will be released to the research community for multimodal research
LangMap: A Human-Verified Benchmark for Hierarchical Open-Vocabulary Goal Navigation
Bo Miao ⋅ Weijia Liu ⋅ Jun Luo ⋅ Lachlan Shinnick ⋅ Jian Liu ⋅ Thomas Hamilton-Smith ⋅ Yuhe Yang ⋅ Zijie Wu ⋅ Vanja Videnovic ⋅ Feras Dayoub ⋅ Anton van den Hengel
Language-conditioned goal navigation (LGN) requires agents to locate user-specified targets without step-by-step guidance. However, existing benchmarks largely focus on category-level goals or rely on instance descriptions generated by vision-language models (VLMs), which often contain ambiguities and semantic errors, limiting systematic and reliable evaluation. We introduce HieraNav, an open-vocabulary LGN task with goals specified at four hierarchical semantic levels: scene, room, region, and instance. To this end, we present Language as a Map (LangMap), to our knowledge the first real-world 3D indoor navigation benchmark with human-verified semantic annotations to support tasks across all four goal levels. LangMap provides region labels and discriminative region and instance descriptions covering 414 object categories, produced through a rigorous contrastive annotation protocol comparing same-scene regions and instances, and contains over 18K tasks. Each target is paired with concise and detailed descriptions, enabling evaluation across instruction styles. Quantitative and qualitative analyses validate our annotation quality; notably, our instance descriptions outperform GOAT-Bench annotations by 23 percentage points in text-to-view matching. We further introduce PlaNaVid, a strong RGB-only baseline that combines Bounded Diverse Memory (BDM) with high-level planning to prime a reactive policy for multi-goal navigation. PlaNaVid achieves top-tier success rates without depth, 3D scene representations, or object masks. Further analysis shows that memory and richer context boost performance, while long-tailed categories, small objects, distant targets, and multi-goal completion remain open challenges.
Language Models Demand New Computer Science
Chang Yang ⋅ Xinrun Wang ⋅ Shuxin Li ⋅ Qinggang Zhang ⋅ Huachi Zhou ⋅ Zhuoyi Lin ⋅ Bo Li ⋅ Xiao Huang
The emergence of large language models (LMs) as versatile problem-solving systems marks a paradigm shift comparable to the advent of digital computers. Traditional computer science provides theoretical foundations for understanding computational processes through complexity theory, algorithms, and formal methods, this position paper argues that language models require their own distinct theoretical framework, termed as Computer Science of Language Models (CSLM). We identify seven fundamental differences between traditional computing and language model: 1) the generating-solving-verifying triad versus the solving-verifying dichotomy, 2) token-based rather than operation-based computation, 3) probabilistic reliability with hallucinations rather than deterministic correctness, 4) native multimodal understanding beyond pixel-level processing, 5) general-purpose capabilities versus problem-specific algorithms, 6) essential two-stage training-inference processes, and 7) inherent knowledge. We argue these differences necessitate new theoretical frameworks, complexity measures, and evaluation methodologies. Then, we define the token complexity as the core concept in CSLM and discuss connections to existing research in scaling laws, memory systems, tool use, and fine-tuning, while identifying critical open questions about LM-complete problems, reliability bounds, emergence phenomena, and multimodal complexity. Drawing parallels to complexity theory's 50-year development, we present a 5-year research roadmap (2026-2030) to build the mature CSLM. CSLM promises to transform LM development from empirical scaling to principled engineering grounded in rigorous theory and strengthen our understanding about both the problems, LMs and their relationships.
Laplacian Heads Improve Transformers by Smoothing Token Representations
Yuchong Zhang ⋅ Vardan Papyan
Transformers update token representations through multi-head attention and residual connections as $X \leftarrow X + \sum_{i} P^{(i)}XW_{V_i}W_{o_i}$, where $P^{(i)}$ is the softmax attention matrix in head $i$. We propose replacing a subset of $P^{(i)}$'s with the Laplacian $I - P^{(i)}$, giving $X \leftarrow X + \sum_{i \in \mathcal{A}} P^{(i)}XW_{V_i}W_{o_i} + \sum_{i \in \mathcal{L}} (I - P^{(i)})XW_{V_i}W_{o_i}$. Our proposal has two motivations. First, it allows attention heads to update the mean of token representations, while Laplacian heads can directly control within-sequence variance. Second, if tokens are viewed as nodes in a graph with edge weights \(P^{(i)}\), then \(I - P^{(i)}\) is the corresponding graph Laplacian, and the update can be interpreted as one step of heat diffusion on the graph. We show that this simple modification improves performance across supervised learning, language modeling, and self-supervised learning tasks. To investigate why, we examine the token representations learned with and without Laplacian heads. In supervised learning, Laplacian heads collapse token representations within the same sequence and align the sequence means with the geometry of Neural Collapse. In language modeling, they increase the separability of token representations that share the same next-token prediction. In self-supervised learning, they produce token representations whose principal components are better suited for segmentation. Across modalities, they also lead to faster-decaying spectra, indicating stronger token smoothing. Overall, our findings challenge the prevailing view that token oversmoothing is inherently harmful, showing instead that certain forms of smoothing can be beneficial.
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
Hao Chen ⋅ Zhonghao Yan ⋅ Jiaming Liu ⋅ Nuowei Han ⋅ Renrui Zhang ⋅ Chenyang Gu ⋅ Jialin Gao ⋅ Peng Jia ⋅ Shanghang Zhang ⋅ Pheng-Ann Heng
Robotic foundation models require reasoning over complex visual scenes to execute adaptive actions in dynamic environments. While recent studies on latent-reasoning Vision-Language-Action (VLA) models have demonstrated the capability to capture fine-grained physical dynamics, they remain predominantly confined to static imitation learning, severely limiting their adaptability and generalization. In this paper, we present LaST-R1, a novel reinforcement learning (RL) post-training framework designed to effectively harness latent reasoning-before-acting policies. Specifically, we propose Latent-to-Action Policy Optimization (LAPO), a core RL algorithm that jointly optimizes the latent reasoning process and the action generation. By explicitly embedding latent Chain-of-Thought (CoT) reasoning directly within the RL optimization loop, LAPO stimulates profound physical world modeling, which in turn drives robust execution in interactive environments. Furthermore, an adaptive latent CoT mechanism is introduced, allowing the policy to dynamically modulate its reasoning horizon based on diverse environment states. Experiments show that LaST-R1 achieves a near-perfect 99.9% average success rate on the LIBERO benchmark with only one-shot supervised warm-up, significantly improving convergence speed and performance over prior state-of-the-art (SOTA) methods. In real-world deployments, LaST-R1 yields up to a 22.5% average improvement over SOTA supervised fine-tuning approach across four complex tasks, including both single-arm and dual-arm settings. Finally, LaST-R1 demonstrates strong generalization across simulated and real-world environments.
Single-step retrieval-augmented generation (RAG) provides an efficient way to incorporate external information for simple question answering tasks but struggles with complex questions. Agentic RAG extends this paradigm by replacing single-step retrieval with a multi-step process, in which the large language model (LLM) acts as a search agent that generates intermediate thoughts and subqueries to iteratively interact with the retrieval system. This iterative process incurs substantial latency due to the autoregressive generation of lengthy thoughts and subqueries. To address this limitation, we propose LatentRAG, a novel framework that shifts both reasoning and retrieval from discrete language space to continuous latent space. Unlike existing explicit methods that generate natural language thoughts or subqueries token-by-token, LatentRAG produces latent tokens for thoughts and subqueries directly from the hidden states in a single forward pass. We align LLMs with dense retrieval models in the latent space, enabling retrieval over latent subquery tokens and supporting end-to-end joint optimization. To improve transparency and encourage semantically meaningful latent representations, we incorporate a parallel latent decoding mechanism that translates latent tokens back into natural language. Extensive experiments on seven benchmark datasets show that LatentRAG achieves performance comparable to explicit agentic RAG methods while reducing inference latency by approximately 90%, substantially narrowing the latency gap with traditional single-step RAG.
Latent-space Attacks for Refusal Evasion in Language Models
Giorgio Piras ⋅ Raffaele Mura ⋅ Fabio Brau ⋅ Maura Pintor ⋅ Luca Oneto ⋅ Fabio Roli ⋅ Battista Biggio
Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations. Existing methods do so by ablating a refusal direction from model activations, aiming to remove refusal from the model’s residual stream. Despite their empirical success, these methods lack a principled account of the latent-space transformation they induce and why it suppresses refusal. In this work, we recast refusal suppression as a latent-space evasion attack against linear probes trained to separate refused from answered prompts. Under this view, prior work’s difference-in-means direction naturally defines such a probe, and its ablation is exactly a projection onto its decision boundary, i.e., a minimum-confidence evasion attack. This perspective not only explains the empirical success of prior work but also admits a key limitation: evasion stops at the decision boundary, motivating the need to push representations further into the compliant region, i.e., where the model answers. We leverage this by proposing a Controlled Latent-space Evasion attack that projects representations past the boundary with an optimized confidence. We achieve state-of-the-art attack success rate across 15 instruction-tuned, multimodal, and reasoning models, outperforming existing refusal-ablation baselines and specialized jailbreak attacks.
LEAP: Library-driven Evolutionary Abstraction Paradigm for Large Language Models
Fanbin Lu ⋅ Chi-Wing Fu ⋅ Jiaya Jia
Despite the remarkable progress of evolutionary LLM frameworks like AlphaEvolve in program and algorithm discovery, they consistently struggle with tasks requiring deep algorithmic reasoning and long-horizon planning, such as the Abstraction and Reasoning Corpus (ARC-AGI). A fundamental limitation of current evolutionary code generation is its reliance on unguided stochasticity to derive child programs from parents, lacking a systematic mechanism for knowledge accumulation, whereas human problem-solving thrives on the continuous consolidation and reuse of high-level conceptual abstractions. To bridge this gap, we propose the \textbf{Library-driven Evolutionary Abstraction Paradigm (LEAP)}, a cognitive-inspired framework that enables LLMs to autonomously construct and evolve a globally shared library of reusable algorithmic primitives. The evolution of this library is governed by two rigorous theoretical principles: First, the generation of new primitives is driven by a Minimum Description Length (MDL) objective, ensuring the system only assimilates concepts that genuinely compress the problem space. Second, the evolutionary survival and selection of these primitives are governed by a principled exploration-exploitation mechanism, dynamically regulating the library's life-cycle to strike an optimal balance between exploiting proven cognitive operators and exploring novel hypotheses while strictly bounding context entropy. Extensive experiments demonstrate that LEAP significantly outperforms existing LLM evolutionary baselines on the ARC-AGI benchmark and math optimizations problems, exhibiting robust cross-task generalization, highly efficient context utilization, and autonomous concept discovery capabilities.
Learning Continuously Evolving Spatio-Temporal Explanations for Traffic Flow Forecasting
Cuiying Huo ⋅ Baoxu Wang ⋅ Lin Wu ⋅ Yu Mei ⋅ Dongxiao He ⋅ Yawen Li ⋅ Di Jin
Traffic flow forecasting plays a pivotal role in intelligent transportation systems. However, most existing methods adopt a black-box learning paradigm, resulting in a lack of interpretability in their decision-making processes. Existing explainable methods mostly follow the feature attribution paradigm, leaving explanations at the level of static or discrete feature evidence and making it difficult to reveal the continuous evolution mechanism of spatio-temporal dependencies during dynamic traffic propagation. Meanwhile, prior methods often decouple spatial and temporal explanations, treating spatial structures and temporal dynamics independently and thus failing to capture the intrinsic coupling between spatial dependencies and temporal contexts in model decision-making. To this end, we propose STGIB, a new spatio-temporal graph information bottleneck theory for characterizing continuously evolving explanations in traffic forecasting. STGIB redefines traffic explanation as a time-indexed spatio-temporal explanation trajectory, and jointly learns dynamic key structures and their temporal representations under predictive sufficiency and information compression constraints. Furthermore, we derive tractable recursive variational bounds to make the theoretical objective optimizable, and instantiate a new explainable traffic flow forecasting model via spatio-temporal explanation trajectory generation. Experiments on real-world traffic datasets demonstrate that STGIB maintains competitive predictive performance while generating more faithful and spatio-temporally consistent explanations.
Learning Gaussian Conditional Distributions using Neural Ratio Estimation is Hard
Pierre Glaser ⋅ Arthur Gretton
Neural Ratio Estimation (NRE) is a popular method for performing simulation-based posterior estimation given a known prior and samples from a joint distribution. For such a task, it can be crucial for the posterior estimation task to be \emph{statistically efficient}, i.e. to require as few simulations as possible to obtain an accurate posterior estimate. However, currently, little is known about its statistical efficiency. In this work, we bridge this gap by performing an asymptotic statistical analysis of NRE. We show that even when the target distribution is a simple product of Gaussians, the accuracy of NRE can degrade exponentially badly with the dimension of the data distribution. Our results show that NRE can be significantly less efficient than other conditional density methods like Maximum Likelihood Estimation, which are known to enjoy favorable statistical properties.
Does a neuron’s shape predict whom it connects to? Peters’ rule, a canonical geo- metric principle in connectomics, predicts connectivity from spatial arbor overlap, but true wiring also depends on molecular and morphological specificity. We ask how far neuron skeletons, with identity and graph features withheld, predict directed connections in the FlyWire Drosophila connectome. Because most non- connected pairs are far apart, headline metrics are dominated by easy negatives that geometry can reject. We therefore evaluate controlled contrasts: plausible non-connections whose arbors are nearby, overlapping, or cell-type matched, but that still do not synapse. Across a controlled model ladder, from arbor overlap and hand-crafted morphology baselines to Connectoformer (a bidirectional Trans- former with pooled and pairwise cross-attention scoring), performance improves when the model gains the corresponding signal: global shape, local branch com- patibility, and tree context. Morphology and arbor co-occupancy are complemen- tary: morphology discriminates similar-looking unconnected pairs but can score pairs whose arbors never touch; arbor co-occupancy does the inverse. By suppress- ing essentially-zero-overlap pairs without boosting high-overlap pairs, a fixed one- sided arbor prior performs well in both regimes. The encoder also transfers across Drosophila individuals: zero-shot evaluation on the male CNS connectome beats arbor overlap on hard-negative discrimination. Stratified evaluation reveals a clean division of labor: geometry rules out pairs that cannot physically connect, and morphology picks the actual partners from what remains.
Learning Options for Compositional Motor Control with Adapter Banks
Sreejan Kumar ⋅ Marcelo G Mattar ⋅ Lea Duncker
Learning flexible motor primitives is a hallmark of skilled motor control. Recent neuroscience theory proposes that motor primitives may be implemented as low-rank perturbations of a shared recurrent network, but leaves open how such a system is learned. We translate this principle into a novel architecture for learning motor skills end-to-end: a shared recurrent core modulated by a bank of residual adapters, each selected by a discrete latent code. Trained on closed-loop biomechanical control, the adapters develop emergent low-rank perturbations of the recurrent dynamics despite no architectural rank constraint, placing task representations in disparate subspaces of the shared core network. A simple high-level policy over the learned options, optimized while the whole network is frozen, sequences the low-rank adapters to produce novel out-of-distribution movements. We demonstrate the ability to generalize to novel motor sequences within the closed-loop control setting, improving on the generalization error of a task-input-conditioned multitask baseline by up to order of magnitude.
Learning Pareto Stationary Fronts via Single-Pass Backpropagation
Elina Rojin Celik ⋅ Marcos M. Raimundo ⋅ Isabel Valera
We propose MOSEL (Multi-Objective Stackelberg Efficient Learning), a framework for a posteriori multi-objective optimization (MOO) in deep neural networks that recovers a full front of Pareto stationary solutions at the computational cost of standard single-objective training. MOSEL reformulates the problem as a bilevel optimization problem that leverages network modularity to decouple representation learning from objective-preference alignment. Casting the bilevel problem as a Stackelberg game enables solving the original a posteriori MOO problem in a single forward–backward pass. As a result,MOSEL matches the time and memory efficiency of standard single-objective training while enabling scalable Pareto stationary front learning. Empirically, MOSEL uncovers diverse and optimal Pareto frontiers in strongly conflicting settings (e.g., fairness–accuracy). Remarkably, even in weakly conflicting regimes such as multi-task learning, it consistently converges to solutions closer to the utopia point, outperforming both standard single-objective training and specialized multi-task learning methods. These results highlight the broader potential of a posteriori MOO learning as a pathway to efficiently learn more diverse and robust representations, ultimately improving generalization.
Learning Pseudo-Riemannian Manifolds for Heterophilic Graphs via Graph Signature
Yun Young Choi ⋅ Asung Kil ⋅ Sun Woo Park ⋅ Minho Lee ⋅ Seokhwan Kim
Graph Neural Networks (GNNs) operate under the implicit assumption that the underlying data manifold is Riemannian, where the metric tensor is strictly positive-definite. While effective for homophilic graphs, this inductive bias creates a fundamental geometric mismatch for heterophilic graphs, where edges often signify dissimilarity or structural repulsion. In this work, we propose a paradigm shift from learning on fixed manifolds to learning the manifold itself. We postulate that heterophilic graphs are naturally embedded in pseudo Riemannian manifolds endowed with an indefinite metric, allowing for negative squared distances to model repulsive interactions. To formalize this, we introduce the Graph Signature Index, a spectral invariant that diagnoses the geometric nature of a graph. This index enables our proposed Pseudo-Riemannian Attention Network (PRAT) to dynamically learn an indefinite metric tensor, effectively capturing both attractive and repulsive interactions. By unifying these opposing forces within a single framework, PRAT demonstrates how a subtle shift in the metric signature yields substantial performance gains on heterophilic benchmarks. Our analysis reveals that PRAT’s learned geometry emergently aligns with the spectral signature of the graph, validating our geometric hypothesis.
Learning Reach-Set Geometry for Tighter Probabilistic Neural Network Verification
Samuel Sasaki ⋅ Ben Wooding ⋅ Taylor Johnson
Neural network verification is key to certifying robustness in safety-critical systems. Sound verifiers commit to abstract domains that propagate soundly through the network's operations, paying for that commitment with looseness at scale and applicability largely confined to piecewise-linear architectures. Probabilistic verifiers relax soundness, but they too commit to a fixed family of shapes for the network's reachable outputs. When the network's true output geometry differs from the chosen shape, the verifier over-approximates the reach-set and fails to certify networks that are in fact safe. We propose instead to learn the reach-set geometry from the network's behavior. A flow-matching model trained on input-output samples and calibrated by conformal prediction yields a probabilistic reach set whose shape adapts to the true shape of the reach-set rather than to an a priori choice. The resulting set has no closed-form description, so we reframe the specification check as a rare-event estimation problem on the learned distribution, restricted to the calibration's coverage region, and solve it with bounded adaptive multilevel splitting. The pipeline produces a probabilistic safety certificate combining the conformal coverage with a rare-event upper bound. Our pipeline performs comparably to sound verifiers on a subset of VNN-COMP 2025 benchmarks where sound verification is tractable and matches or exceeds fixed-shape probabilistic baselines on most benchmarks in the same suite. On a synthetic family of growing-depth networks, it scales sub-exponentially where sound verifiers time out or abstain.
Learning Reusable Motor Motifs for Continuous Animal Behavior Modeling
Jiyi Wang ⋅ Jingyang Ke ⋅ Bo Dai ⋅ Anqi Wu
Animals generate complex behaviors by flexibly recombining a finite set of motor primitives, but existing behavior segmentation methods oversimplify this process by imposing discrete syllables under restrictive generative assumptions. Here, we introduce Motif-based Continuous Dynamics (MCD), a framework that models behavior as a continuous, compositional process driven by reusable motor motifs. MCD leverages reinforcement learning (RL) to (1) discover interpretable motif representations via transition-based representation learning, and (2) model behavior as a time-varying mixture of these motifs through motif-based policies. This formulation avoids restrictive dynamics assumptions while capturing continuity, compositionality, and long-term dependencies in behavior. Across simulated gridworld tasks, maze navigation, and animal behavior datasets, MCD identifies reusable motifs and interprets the trajectories with them. These results provide a generative account of behavior as combinations of fundamental motifs, offering a flexible and interpretable framework for studying natural behavior.
Learning to Commit: Next-Commit Prediction via Online Supervised Contrastive Reflection
Mo Li ⋅ Qitai Tan ⋅ Kai Chen ⋅ Ting Cao ⋅ Yunxin Liu
Large language model (LLM)-based coding agents achieve strong results on controlled benchmarks yet routinely produce pull requests that real maintainers reject. The cause is not functional incorrectness but a lack of organicity: generated code ignores project conventions, duplicates internal APIs, and violates implicit architectural constraints. The latest repository snapshot is insufficient-it reveals the final codebase state, not the change patterns that shaped it. We make three contributions. (1) Paradigm: we propose next-commit prediction as a paradigm for learning agentic coding skill from a repository's own history-each historical commit is a self-supervised target whose oracle diff supplies dense supervision. (2) Evaluation: we establish Organicity Evaluation as a measurable objective for coding agents and contribute the Learning to Commit benchmark-5 curated open-source GitHub repositories under strict per-repository temporal splits-scoring patches along file localisation, internal API reuse, patch bloat, and code-style consistency. (3) Method: we instantiate the paradigm with the Learning to Commit framework, in which the agent performs supervised contrastive reflection: it blindly attempts each historical commit, contrasts its prediction against the oracle, and incrementally distils a reusable skill document that conditions subsequent generations. On the Learning to Commit benchmark and on SWE-bench Pro reframed under our temporal-split protocol, our framework consistently improves organicity on held-out future tasks and further lifts test-pass rates on SWE-bench Pro, narrowing the gap between benchmark success and real-world mergeability.
Learning to Discriminate Scene Structures Makes Self-Supervised Depth Learning Scalable
Mengtan Zhang ⋅ Yuchuan Cui ⋅ Yu Ma ⋅ Zizhan Guo ⋅ Wei Ye ⋅ Rui Fan
Self-supervised monocular depth estimation is appealing for its potential to scale with unlabeled data, yet existing methods often degrade when trained on diverse scenes. This study finds that, under self-supervision, models tend to encode dataset-specific structural biases rather than transferable scene geometry, leading to cross-domain optimization conflicts. To address this issue, this study proposes the MaSS framework, aiming to make self-supervised depth learning scalable to diverse data. MaSS introduces a dedicated structural representation pathway explicitly decoupled from contextual features, which is guided by a novel scene-structure discrimination objective to learn generalized, layout-aware structural representations for depth decoding. Trained and evaluated on four diverse datasets, both individually and jointly, MaSS substantially outperforms prior self-supervised methods and is the first to consistently benefit from mixed-domain training. Zero-shot evaluation on eight additional datasets further demonstrates its remarkable generalizability, particularly on those entirely out-of-distribution scenes. The source code will be publicly available upon publication.
Learning to Inject: Automated Prompt Injection via Reinforcement Learning
Xin Chen, Cynthia ⋅ Jie Zhang ⋅ Florian Tramer
Prompt injection is a critical vulnerability in LLM agents, yet the strongest methods still rely on human red-teamers and hand-crafted prompts. Adapting automated jailbreak optimizers does not close this gap: jailbreaks shape models toward generic compliance, while prompt injection requires emitting specific tool calls with correct parameters. The success signal is binary, and randomly sampled suffixes almost never trigger it—so standard optimizers have no gradient to follow. We present AutoInject, a black-box reinforcement learning (RL) framework that learns adversarial suffixes for prompt injection. A learned comparison-based reward scores each candidate against the best suffix seen so far, turning the binary signal into a dense reward suitable for RL optimization. The framework supports both online query-based attacks and offline-trained transferable suffixes that need no utility access at deployment, and incorporates a utility objective when task-completion feedback is available. On AgentDojo, AutoInject outperforms template attacks, GCG, TAP, and adaptive attack across production models, with statistically significant improvements under McNemar's test with $p<0.05$. Suffixes learned by AutoInject also break Meta-SecAlign-70B, a model fine-tuned specifically to resist prompt injection, where template attacks fail outright. The results establish an automated baseline for prompt injection and expose a gap between preference-based defenses and adaptive optimization-based attackers.
Learning with Enumeration: Neural-Guided SAT Framework for Cryptographic Key Recovery
Xinhao Zheng ⋅ Xinhao Song ⋅ Jinchen Yu ⋅ Zhuoyuan Xu ⋅ Gongshen Liu ⋅ Junchi Yan
Boolean satisfiability (SAT) provides a foundation tool for cryptographic key recovery by encoding ciphers into CNF or ANF representations. However, heuristic solvers and their neural-enhanced variants often degrade significantly on cryptographic instances due to pronounced structural symmetry, a large number of intermediate variables, and complex algebraic dependencies. Existing neural approaches—ranging from end-to-end prediction to solver-integrated heuristics—face challenges in scalability, assignment accuracy, or computational efficiency. In this paper, we propose a unified neural-guided enumerative SAT framework for cryptographic key recovery. It performs a single forward neural guidance to identify a subset of $k$ critical variables, followed by enumeration over their assignments before invoking a heuristic SAT solver. This design effectively reduces the combinatorial search space while keeping neural overhead minimal. We further introduce a taxonomy of neural guidance across data and model capability regimes, including supervised selection with cryptographic priors, uncertainty-driven selection with assignment prediction, and robust fallback strategies using the model trained on general SAT datasets. Experiments on 10 SAT solvers and 2 cube-and-conquer strategies over SAT4CryptoBench and SAT Competition benchmarks demonstrate up to $5\times$ speedup and approximately 2× average improvement on cryptographic instances, while maintaining better performance on general SAT datasets. These results highlight the effectiveness and generalization of this framework, illustrating its potential to bridge machine learning and SAT solving in cryptanalysis tasks.
Synthetic data has become a promising way to scale model training beyond limited human-generated data but it may also induce strong model collapse (Dohmatob et al., 2024), where any fixed fraction of synthetic data prevents model performance from improving under data scaling, leaving a non-vanishing excess risk floor. In this paper, we studied how synthetic data affects the generalization of one-pass SGD in high-dimensional linear regression with model shift. We show that mixed training induces strong model collapse while two-stage training avoids this by using synthetic data only in the first stage, followed by real-data training in the second stage, showing that strong model collapse is not inevitable through a simple data curriculum. We further establish scaling-law upper bounds for both protocols under a random sketch model, showing that larger models amplify synthetic-induced degradation in mixed training and giving an explicit characterization of how high-quality synthetic training may reduce bias in two-stage training. Overall, our results highlight that synthetic data is neither inherently harmful nor beneficial; its effect depends critically on both its quality and the training protocol used to incorporate it.
LensDesigner: A Self-Improving Agent for Optical Lens Design
Lei Sun ⋅ Haoran Liang ⋅ Dannong Xu ⋅ Yao Gao ⋅ Yuyu Geng ⋅ Jinjin Gu ⋅ Kaiwei Wang ⋅ Danda Pani Paudel ⋅ Luc V Gool
Optical lens design is a complex, non-convex optimization challenge that relies heavily on human experience and intuition. Existing optimized-based automatic lens design methods struggle to navigate this vast parameter space without meticulous manual tuning. In this paper, we present LensDesigner, an autonomous agent framework that mirrors the problem-solving workflow of expert opticians. To overcome the initial cold start problem, we construct LensLib100K, an extensive optical lens library, and employ Optics-Aware Retrieval to supply physically valid structural seeds. Within an interactive physical simulation environment, the agent executes macroscopic orchestration while receiving immediate optical feedback. Furthermore, we introduce a continuous self-evolving mechanism guided by a curriculum agent. By iteratively solving design tasks with progressively increasing difficulty, the agent autonomously extracts, accumulates, and reuses design heuristics, effectively evolving its optical lens design expertise over time. At the evaluation level, we introduce LensArena, a standardized evaluation benchmark comprising 120 diverse optical design tasks, covering extreme configurations. Extensive experiments on this benchmark demonstrate that LensDesigner significantly outperforms publicly available baseline algorithms, achieving superior success rates and optimization efficiency. We hope this work sheds light on the emerging field of intelligent optics. The code will be publicly available.
Distilled diffusion models accelerate image generation by reducing the number of denoising steps, but often suffer from degraded image quality. To mitigate this trade-off, test-time optimization methods improve quality, yet their iterative nature incurs substantial computational overhead and leads to slow inference, limiting practical usability. Recent hypernetwork-based approaches amortize this process during training, but still require costly noise modulation in high-dimensional latent spaces. In this work, we propose LENS (Low-frequency Eigen Noise Shaping), an efficient noise modulation framework that operates in a low-dimensional subspace. Our approach is motivated by the observation that low-frequency components of the noise largely determine the global structure and visual fidelity of generated images. Based on this observation, we provide a theoretical justification for restricting modulation to the low-frequency subspace and derive a principled training objective. Building on this, LENS employs a lightweight, standalone network to selectively modulate these components, enabling efficient and targeted noise modulation. Extensive experiments demonstrate that LENS achieves competitive image quality while reducing FLOPs by 400–700$\times$, model parameters by 25–75$\times$, and inference-time overhead by 10–20$\times$ compared to prior methods.
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
Roy Xie ⋅ Dan Friedman ⋅ Donghan Yu ⋅ Bowen Pan ⋅ Christopher Fifty ⋅ Jang-Hyun Kim ⋅ Xianzhi Du ⋅ Zhe Gan ⋅ Vivek Rathod ⋅ Bhuwan Dhingra
Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the vision encoder's effective resolution, making them indistinguishable. To address this, we propose LensVLM, an inference framework and post-training recipe that enables VLMs to scan compressed images, then selectively expand only the relevant images to their uncompressed form via learned tools. Building on Qwen3.5-9B-Base, LensVLM maintains accuracy comparable to the full-text upper bound at 4.3$\times$ effective compression and outperforms retrieval-based, text- and visual-compression baselines up to 10.1$\times$ effective compression across seven text QA benchmarks. LensVLM also generalizes to multimodal document and code understanding tasks, with the accuracy gain over baselines growing as compression increases. Our analysis validates this approach: training makes visual compression robust to rendering choices, and as compression grows the model increasingly relies on expanded content rather than unreliable visual reading. The analysis also yields practical tool-choice guidance: text expansion is preferable for rendered text, while high-resolution image expansion suits native documents whose layout cues carry task-relevant information.
Less Evidence, Better Answering: Gain-Aware Minimal Evidence Subset Selection for Medical QA
Songyue Guo ⋅ Zhao CHEN ⋅ Caleb C Cao ⋅ Lei Chen
Evidence-grounded medical question answering requires not only accurate answers, but also precise, verifiable, and clinically trustworthy evidence. Existing medical RAG systems typically retrieve top-(k) chunks according to item-wise relevance. Although this design improves recall, it does not explicitly optimize whether the selected evidence set is compact, sufficient, and decision-critical. In this paper, we study medical evidence-grounded QA from a different perspective: instead of retrieving evidence without optimization, the evidence-grounded medical QA system should identify a compact sufficient subset of high-density evidence snippets. We propose MEGA, a gain-aware framework for compact supporting medical evidence selection. MEGA first expands clinical queries into complementary subqueries and performs hybrid retrieval to construct a recall-oriented candidate evidence pool. It then introduces hidden-state Information Gain Scoring, which uses a frozen LLM to estimate whether each candidate snippet contributes new answer-relevant information beyond surface relevance. Finally, MEGA formulates evidence selection as a budgeted utility maximization problem and proposes an FPTAS selector to identify a compact evidence subset under an explicit token budget. We further construct two evidence-grounded medical QA benchmarks, CRC-EvidenceQA and Med-EvidenceQA, to evaluate both answer quality and evidence grounding. Across three LLM backbones and five medical RAG baselines, MEGA consistently achieves the best results, improving over the second-best baseline by (8.8\%) on average in answer quality and 38.8\% on average in evidence F1. Ablation studies and blinded clinical expert evaluation further validate its robustness and clinical relevance.
Less Language, More Latents: Annotation-Efficient VLAs for Driving
Alexey Zakharov ⋅ Kemal Oksuz ⋅ Puneet Dokania
Vision-language-action models (VLA) promise human-steerable autonomous driving, but their training is bottlenecked by the scarcity of frames paired with natural-language instructions: while camera streams and expert trajectories are logged at scale, language annotations (e.g., "turn left at the intersection") remain scarce and expensive to acquire. To address this challenge, we introduce Latent Action Driving Annotations (LADA), a three-stage pipeline that transforms abundant unlabelled observation–trajectory pairs into a substrate for language-conditioned control. First, we train a latent action model with a vector-quantised bottleneck, producing a compact codebook of high-level vehicle intents. Second, a small language-annotated subset is used to train a vision-language translator to map observations and language instructions into this codebook. Third, we train a driving VLA on observation–latent-action pairs over the full unlabelled corpus. Using fewer than 5% of language annotations and without leveraging any auxiliary chain-of-thought reasoning or visual question answering streams, LADA achieves a Driving Score of 87.98 and a Success Rate of 70.46% on the closed-loop Bench2Drive benchmark, matching or surpassing fully supervised baselines.
Less Supervision, Better Generalization: Weakly Supervised Fake Region Localization in Diffusion-Edited Images
Junhee Lee ⋅ JeonDongHyeon ⋅ Taeoh Kim ⋅ beomYoung Kim ⋅ MyeongAh Cho
Localizing AI-edited regions is essential for interpretable forensic analysis, but remains challenging due to subtle and spatially distributed artifacts that are misaligned with semantic or object boundaries. Existing approaches rely on pixel-level supervision from controlled editing pipelines, which is difficult to scale and can introduce misleading signals: artifacts frequently extend beyond annotated regions, while out-of-mask pixels are treated as authentic. This limits models' ability to capture transferable evidence and generalize across generators and datasets. To address these issues, we propose \textbf{ReGFLoW}, a \textbf{Re}construction-\textbf{G}uided \textbf{F}ake \textbf{Lo}calization framework under \textbf{W}eak supervision, which is the first weakly supervised approach for diffusion-edited fake region localization. ReGFLoW requires only real/fake labels at the image-level and uses diffusion reconstruction errors as dense spatial guidance to inject them into both feature and score spaces. Furthermore, by artifact-centric multiple instance learning, ReGFLoW utilizes localized diffusion evidence without relying on semantic-affinity or boundary-based pseudo-mask priors. Extensive experiments demonstrate that ReGFLoW achieves stronger out-of-domain generalization than fully supervised learning baselines.
LiBrA-Net: Lie-Algebraic Bilateral Affine Fields for Real-Time 4K Video Dehazing
Yongcong Wang ⋅ Chengchao Shen ⋅ Guangwei Gao ⋅ Wei Wang ⋅ Pengwen Dai ⋅ Dianjie Lu ⋅ Guijuan Zhang ⋅ Zhuoran Zheng
Currently, there is a gap in the field of ultra-high-definition (UHD) video dehazing due to the lack of a benchmark for evaluation. Furthermore, existing video dehazing methods cannot run on consumer-grade GPUs when processing continuous UHD sequences of 3--5 frames at a time. In this paper, we address both issues with a new benchmark and an efficient method. Our key observation is that atmospheric dehazing reduces to a per-pixel affine transform governed by the low-frequency depth field, which can be compactly encoded in bilateral grids whose prediction cost is decoupled from the output resolution. Building on this, we propose LiBrA-Net, which factorizes the spatiotemporal affine field into a spatial--color and a temporal bilateral sub-grid predicted at a fixed low resolution, fuses their coefficients in the $\mathfrak{gl}(3)$ Lie algebra under group-theoretic regularization, maps the result to invertible $GL(3)$ transforms via a Cayley parameterization, and restores high-frequency detail through a lightweight input-guided branch. We further release UHV-4K, the first paired 4K video dehazing benchmark with depth, transmission, and optical-flow annotations on every frame. Across UHV-4K, REVIDE, and HazeWorld, LiBrA-Net sets a new state of the art among compared video dehazing methods while running native 4K at 25\,FPS on a single GPU with only 6.12\,M parameters. Code and data are available at \url{https://anonymous.4open.science/r/LiBrA-Net-42B8}.
LifeStream: Token Life-Cycle Modeling for Training-Free Online Video Understanding
Yifan Ge ⋅ Zhihang Liu ⋅ Jiannan Ge ⋅ Chenhui Jin ⋅ Hongtao Xie
With the rapid development of Video Large Language Models (VLLMs), online video understanding has emerged as a practical paradigm for real-world streaming applications. A key challenge is to construct a compact, query-agnostic visual memory from continuously growing video streams under strict token budgets. Existing token reduction methods typically rely on static and coarse-grained importance estimation, implicitly assuming that retained tokens contribute equally over time. This overlooks the dynamic nature of visual information, where the relevance of tokens evolves as the video progresses, leading to suboptimal prioritization and inefficient long-term memory usage. To address this limitation, we propose LifeStream, a training-free framework for lifecycle-aware visual token selection and hierarchical memory construction. LifeStream models the temporal evolution of visual tokens and reveals that tokens contribute unequally to downstream reasoning depending on their lifecycle states. Based on this insight, it organizes tokens into a hierarchical memory structure: informative non-stable tokens are preserved in an archive memory, while long-range history is progressively compressed into a global memory according to state-specific priorities. This design enables continuous retention of high-value information under limited token budgets, without requiring query-time retrieval or additional computation. Extensive experiments demonstrate that LifeStream discards over 80\% of visual tokens while improving performance. It improves LLaVA-OV-7B and Qwen3-VL-8B by 4.4\% and 3.5\% on StreamingBench, respectively, and achieves state-of-the-art results on one additional online benchmark and four long-video understanding benchmarks.
LiFT: Likelihood-Free Tree-Structured Policy Optimization for Flow-Based VLAs
Junjie Gao ⋅ Siyuan Song ⋅ ZIXUAN ZHANG ⋅ SONG FEIYANG ⋅ Pan Yongzhou ⋅ Wen Nuan ⋅ Yaosheng Deng ⋅ Xue Tian ⋅ Mir Feroskhan
Flow-based vision-language-action (VLA) models provide expressive action generation for embodied control, but supervised fine-tuning (SFT) often confines them to narrow expert behaviors. Online reinforcement learning (RL) can improve beyond demonstrations through environment interaction, yet flow-based VLAs struggle with sparse long-horizon rewards and intractable action likelihoods. We propose \textbf{\emph{LiFT}}, a likelihood-free tree policy optimization framework for flow-based VLAs. LiFT first addresses sparse-reward credit assignment by creating sibling continuations from shared rollout histories and using their final outcomes to compute branch level relative advantages. LiFT further uses local action-flow instability as an expansion criterion, focusing rollout budget on ambiguous action-generation regions that are most likely to reveal meaningful outcome differences. Finally, LiFT applies branch-level advantages through a decision-level surrogate ratio on executed action chunks, matching policy updates to the granularity of tree-based credit assignment. This yields a likelihood-free policy update without attributing chunk-level credit to individual denoising steps. Evaluations on in-distribution and out-of-distribution benchmarks show that LiFT improves task success and robustness to distribution shifts.
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
Ning Yang ⋅ Hengyu Zhong ⋅ Wentao Wang ⋅ Baoliang Tian ⋅ Yuan Zhou ⋅ Haijun Zhang ⋅ Jun Wang
The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by Continued Pre-Training (CPT). While effective, this paradigm is notoriously data-hungry and computationally expensive, requiring massive long-text corpora to recalibrate the model to the shifted positional distribution. We propose LinearARD, a self-distillation method that restores Rotary Position Embedding (RoPE)-scaled students through attention-structure consistency with a frozen native-RoPE teacher. Rather than next-token prediction or opaque hidden-state matching, LinearARD aligns row-wise distributions of dense $Q/Q$, $K/K$, and $V/V$ self-relation matrices from the final attention layer to directly supervise attention dynamics. To remove the quadratic memory bottleneck of $n \times n$ relation maps, we introduce a linear-memory kernel that stores only per-token log-sum-exp statistics and recomputes logits in the backward pass to obtain exact Kullback-Leibler divergence gradients. Across LLaMA2-7B, LLaMA3-8B, and Mistral-7B-v0.1 extended to 32K context, LinearARD recovers 93.1\%/94.2\%/94.3\% of native short-context performance and achieves strong long-context robustness on RULER using only \textbf{4.25M} tokens---amounting to just 1.6\% (a $\sim$60$\times$ reduction) of the 256M-token budget required by state-of-the-art baselines. Under this severely constrained budget, CPT and LongReD remain near zero on RULER, demonstrating that relation-level restoration provides a substantially better efficiency-quality tradeoff.
LINE: LLM-based Iterative Neuron Explanations for Vision Models
Vladimir Zaigrajew ⋅ Michał Piechota ⋅ Gaspar Sekula ⋅ Paweł Gelar ⋅ Przemyslaw 'Prem' Biecek
Interpreting the concepts encoded by individual neurons in deep neural networks is a crucial step towards understanding their complex decision-making processes and ensuring AI safety. Despite recent progress in neuron labeling, existing methods often limit the search space to predefined concept vocabularies or produce overly specific descriptions that fail to capture higher-order, global concepts. We introduce LINE, a novel, training-free iterative approach tailored for open-vocabulary concept labeling in vision models. Operating in a strictly black-box setting, LINE leverages a large language model and a text-to-image generator to iteratively propose and refine concepts in a closed loop, guided by activation history. We demonstrate that LINE achieves state-of-the-art performance across multiple model architectures, yielding AUC improvements of up to 0.11 on ImageNet and 0.05 on Places365, while discovering, on average, 27% of new concepts missed by massive predefined vocabularies. Beyond identifying the top concept, LINE provides a complete generation history, enabling polysemanticity evaluation and producing supporting visual explanations that rival gradient-dependent activation maximization methods. The LINE source code is available on GitHub.
Listening to the Wise Few: Query–Key Alignment Unlocks Latent Correct Answers in Large Language Models
Eduard Tulchinskii ⋅ Kristian Kuznetsov ⋅ Laida Kushnareva ⋅ Anastasia Voznyuk ⋅ Andrei Andriiainen ⋅ Irina Piontkovskaya ⋅ Evgeny Burnaev ⋅ Serguei Barannikov
Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer internally. We expose this latent knowledge via the Query–Key (QK) score, defined for an attention head as the inner product between the last-token query and the key at the end-of-line token following option $i$, evaluated before rotary positional embedding is applied. Its argmax identifies a universal class of select-and-copy heads in middle layers that perform option selection through semantic query–key alignment, mechanistically distinct from induction and copy-suppression heads (Olsson et al., 2022): they are invariant to label symbols, and solve a synthetic task with zero surface overlap—properties no positional-copy account explains and that critically require stripping RoPE. Across $24$ models from $1.5\mathrm{B}$ to $72\mathrm{B}$ parameters (LLaMA-2/3/3.1/3.3, Qwen-2.5, Gemma, Phi-3.5, DeepSeek-R1-Distill), a single head's QK-score exceeds the model's own zero-shot accuracy by up to $+27.4$ pp on HellaSwag and $+49.8$ pp on HaluDialogue; causal zero-ablation collapses MCQA accuracy to near-random. To remove any dependence on labeled validation data, we introduce an unsupervised HeadScore that ranks heads from unlabeled inputs and recovers the supervised top-$k$ heads on every tested model. Against four positional-debiasing baselines (e.g., PriDe, Wiegreffe, Wang), QK-score is complementary by construction: debiasing re-weights output logits, whereas QK-score reads the model's selection from a middle-layer head before decoding. We release a one-line drop-in $\mathtt{HeadScore}$ script and per-model head indices, making every result one-command reproducible across all $24$ models and four benchmarks.
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex
Yun Qu ⋅ Qi Wang ⋅ Yixiu Mao ⋅ Heming Zou ⋅ Yuhang Jiang ⋅ Yingyue Li ⋅ Wutong Xu ⋅ Lizhou Cai ⋅ Weijie Liu ⋅ Clive Bai ⋅ Kai Yang ⋅ Yangkun Chen ⋅ Saiyong Yang ⋅ Xiangyang Ji
Reinforcement learning with verifiable rewards (RLVR) has become a standard approach for large language models (LLMs) post-training to incentivize reasoning capacity. Among existing recipes, group-based policy gradient is prevalent, which samples a group of responses per prompt and updates the policy via group-relative advantage signals. This work reveals that these optimization strategies share a common geometric structure: each implicitly defines a target distribution on the response simplex and projects toward it via first-order approximation. Building on this insight, we propose Listwise Policy Optimization (LPO) to explicitly conduct the target-projection, which demystifies the implicit target by restricting the proximal RL objective to the response simplex, and then projects the policy via exact divergence minimization. This framework provides (i) monotonic improvement on the listwise objective with bounded, zero-sum, and self-correcting projection gradients, and (ii) flexibility in divergence selection with distinct structural properties through the decoupled projection step. On diverse reasoning tasks and LLM backbones, LPO consistently improves training performance over typical policy gradient baselines under matched targets, while intrinsically preserving optimization stability and response diversity. The code is available at https://anonymous.4open.science/r/LPO.
LittleLearner: Language Models Under Pedagogically-Controlled Knowledge Exposure
Fanfei Li ⋅ Jana Zeller ⋅ Manuel Prada-Corral ⋅ Thaddäus Wiedemer ⋅ Prasanna Mayilvahanan ⋅ Ryan Cotterell ⋅ Wieland Brendel
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LittleCurriculum, a meticulously curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 2B-parameter LLM from scratch on LittleCurriculum yields LittleLearner, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LittleCurriculum and LittleLearner as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LittleLearner better utilize existing knowledge, but do not raise out-of-scope capabilities. Our findings underscore the value of this controlled environment for future investigations.
LLM Active Alignment: A Nash Equilibrium Perspective
Tonghan Wang ⋅ Yuqi Pan ⋅ Xinyi Yang ⋅ Xinrui Song ⋅ Yanchen Jiang ⋅ Milind Tambe ⋅ David Parkes
We develop a game-theoretic framework for predicting and steering the behavior of populations of large language models (LLMs) through Nash equilibrium (NE) analysis. To avoid the intractability of equilibrium computation in open-ended text spaces, we model each agent’s action as a mixture over human subpopulations. Agents choose actively and strategically which groups to align with, yielding an interpretable and behaviorally substantive policy class. We derive closed-form NE characterizations, adopting standard concave-utility assumptions to enable analytical system-level predictions and give explicit, actionable guidance for shifting alignment targets toward socially desirable outcomes. The method functions as an active alignment layer on top of existing alignment pipelines such as RLHF. In a social-media setting, we show that a population of LLMs, especially reasoning-based models, may exhibit political exclusion—pathologies where some subpopulations are ignored by all LLM agents—which can be avoided by our method, illustrating the promise of applying the method to regulate multi-agent LLM dynamics across domains.
LLM Agents Already Know When to Call Tools - Even Without Reasoning
Chung-En Sun ⋅ Linbo Liu ⋅ Ge Yan ⋅ Zimo Wang ⋅ Lily Weng
Tool-augmented LLM agents tend to call tools indiscriminately, even when the model can answer directly. Each unnecessary call wastes API fees and latency, yet no existing benchmark systematically studies when a tool call is actually needed. We propose When2Tool, a benchmark of 18 environments (15 single-hop, 3 multi-hop) spanning three categories of tool necessity — computational scale, knowledge boundaries, and execution reliability — each with controlled difficulty levels that create a clear decision boundary between tool-necessary and tool-unnecessary tasks. We evaluate two families of training-free baselines: Prompt-only (varying the prompt to discourage unnecessary calls) and Reason-then-Act (requiring the model to reason about tool necessity before acting). Both provide limited control: Prompt-only suppresses necessary calls alongside unnecessary ones, and Reason-then-Act still incurs a disproportionate accuracy cost on hard tasks. To understand why these baselines fail, we probe the models' hidden states and find that tool necessity is linearly decodable from the pre-generation representation with AUROC 0.89–0.96 across six models, substantially exceeding the model's own verbalized reasoning. This reveals that models already know when tools are needed, but fail to act on this knowledge during generation. Building on this finding, we propose Probe&Prefill, which uses a lightweight linear probe to read the hidden-state signal and prefills the model's response with a steering sentence. Across all models tested, Probe&Prefill reduces tool calls by 48% with only 1.7% accuracy loss, while the best baseline at comparable accuracy only reduces 6% of tool calls, or achieves a similar tool call reduction but incurs a 5× higher accuracy loss. On the real-world Search-o1 agentic benchmark, Probe&Prefill reduces API calls by 20–56% without accuracy degradation.
Local Manifold Identification with Latent Linear Models and OT Flows
Sherman Khoo ⋅ Song Liu ⋅ Mark Beaumont
Identifying manifold structure reveals hidden signals from datasets, removing noise. In this paper, we study the problem of identifying local manifold structure from data. Given a query point $y_0$ from a dataset, our goal is to recover the geometry of the data manifold in a neighborhood of $y_0$ using a \emph{pretrained} optimal transport flow from the reference to the data distribution. First, we prove that the Brenier optimal transport map preserves manifold structure: the preimage of an $m$-dimensional data manifold is itself an $m$-dimensional manifold in the reference space. Second, motivated by this result, we propose a latent variable model that maps a linear model through the transport flow. We prove that the linear approximation error is significantly reduced by the optimal transport map, leading to a tight fit of the non-linear data manifold. Third, noting the intractability of the resulting likelihood, we deploy denoising Fisher score estimation --- a recent development from simulation-based inference that learns the Fisher score over parameter--observation pairs --- to perform likelihood-based inference effectively. Experiments on both synthetic and real-world datasets demonstrate the effectiveness of the proposed method.
LOFT: Low-Rank Orthogonal Fine-Tuning via Task-Aware Support Selection
Lanxin Zhao ⋅ Bamdev Mishra ⋅ Pratik Kumar Jawanpuria ⋅ Lequan Lin ⋅ Dai Shi ⋅ Junbin Gao ⋅ Andi Han
Orthogonal parameter-efficient fine-tuning adapts pretrained weights through structure-preserving multiplicative transformations, but existing methods often conflate two distinct design choices: the subspace in which adaptation occurs and the transformation applied within that subspace. This paper introduces LOFT, a low-rank orthogonal fine-tuning framework that explicitly separates these two components. By viewing orthogonal adaptation as a multiplicative subspace rotation, LOFT provides a unified formulation that recovers representative orthogonal PEFT methods, including coordinate-, butterfly-, Householder-, and principal-subspace-based variants. More importantly, this perspective exposes support selection as a central design axis rather than a byproduct of a particular parameterization. We develop a first-order analysis showing that useful adaptation supports should be informed by the downstream training signal, motivating practical gradient-informed support selection strategies. Across language understanding, visual transfer, mathematical reasoning, and multilingual out-of-distribution adaptation, LOFT recovers principal-subspace orthogonal adaptation while gradient-informed supports improve the efficiency–performance trade-off under matched parameter, memory, and transform budgets. These results suggest that principled support selection is an important direction for improving orthogonal PEFT.
Logistic Bandits with $\tilde{O}(\sqrt{dT})$ Regret without Context Diversity Assumptions
Seoungbin Bae ⋅ Dabeen Lee
We study the $K$-armed logistic bandit problem, where at each round, the agent observes $K$ feature vectors associated with $K$ actions. Existing approaches that achieve a rate-optimal $\widetilde{\cal O}(\sqrt{dT})$ regret bound rely heavily on context diversity assumptions, such as strict positivity of the minimum eigenvalue of a context covariance matrix. These assumptions, however, impose strong restrictions on the context process, as they rule out the situation where the context vectors are concentrated in a low-dimensional subspace. In this paper, we propose SupSplitLog, which, to the best of our knowledge, is the first algorithm for logistic bandits that achieves $\widetilde{\cal O}(\sqrt{dT})$ regret without any context diversity assumption. The key idea is to split the collected samples into two disjoint subsets when constructing estimators; one is used to compute an initial-point estimator, while the other is used to apply a Newton-type one-step correction procedure. The splitting rule is carefully designed to balance the accuracy requirements of the initial-point estimator and the one-step correction procedure. Moreover, SupSplitLog strictly improves on the existing algorithms in terms of the dependence on dimension $d$ in the regret upper bound. Furthermore, SupSplitLog can be adapted simply to deduce a regret bound that grows with a data-dependent complexity measure, avoiding a direct dependence on $d$, which is favorable when the context vectors are concentrated in a low-dimensional subspace. We also provide experimental results that demonstrate numerically the superiority of our algorithm, validating the theoretical results.
Real-world human activities unfold as sequences of temporally dependent human-object interactions (HOIs), yet existing methods address either atomic HOIs in isolation or text-driven motion composition without object or scene context. To bridge this gap, we introduce the task of long-term action composition of HOIs, where a human navigates between and interacts with multiple objects within a 3D scene. The primary challenges lie in producing smooth transitions, scene-aware locomotion, and plausible hand-object interaction. To this end, we propose LT-HOI, a diffusion-based model that jointly conditions on past and future contexts—the past providing kinematic continuity across transitions, the future enabling anticipatory navigation toward upcoming interacting objects. We further integrate Diffusion Noise Optimization for collision avoidance and introduce a hand-object displacement loss to improve contact quality. On the ParaHome dataset, we establish the first benchmark for this task and demonstrate that LT-HOI effectively composes long, temporally coherent HOIs.
Long Video Instructional Editing in the Wild
Yiqi Lin ⋅ Guoqiang Liang ⋅ Ziyun Zeng ⋅ Zechen Bai ⋅ Yanzhe Chen ⋅ Mike Zheng Shou
Instructional video editing has matured rapidly on short, single-shot clips, yet long videos, the dominant form of video in practice, remain largely unaddressed. Long videos exhibit failure modes that short-video editing does not stress-test: edit targets recur across temporally distant shots under varying viewpoint, scene, and scale, demanding cross-shot semantic consistency, while non-target intervals must be left unchanged, demanding accurate temporal grounding of the instruction. We present a hierarchical framework that addresses these two axes with dedicated components: a planning agent that grounds the instruction and partitions the timeline into shot-aligned segments, a long-term anchor bank that maintains a sparse set of high-quality cross-shot anchors through critic-driven iterative refinement on top of a strong image editing prior, and an anchor-conditioned video editor obtained by lightly adapting a pretrained short-video editor through LoRA, with no architectural changes. The editor is trained exclusively on short-video pairs and lifted to long videos at inference through its conditioning interface, requiring no long-video supervision. To evaluate this regime, we introduce LVEdit-Bench, the first benchmark designed for long-video instructional editing in the wild, together with a three-tier evaluation protocol that isolates per-shot fidelity, cross-shot identity preservation, and robustness to target-absent content under a structured VLM-judge protocol. Across all three settings, our method matches strong short-video editors on per-segment quality and substantially improves cross-shot consistency and grounding, with the margin widening as temporal complexity grows.
Look Before You Leap: Self-Evolving Clinical Reasoning with Psychometric Preference Optimization for Radiology Report Generation
Qi Wang ⋅ Yunfeng Min ⋅ Liwei Huang ⋅ Zeyu Zhang ⋅ Di Gai ⋅ Jun Liu
Despite the remarkable progress of large Vision-Language Models (VLMs) in Radiology Report Generation (RRG), existing black-box autoregressive generation paradigms lack intermediate clinical logical reasoning. This deficiency renders such models prone to generating linguistically fluent reports with erroneous diagnostic findings, which manifest as plausible hallucinations and thereby undermine clinical reliability. To this end, we propose a novel RRG paradigm termed "Look-Before-You-Leap", aimed at enhancing the clinical logical reasoning capabilities of the model to improve diagnostic accuracy. Specifically, we first construct a self-evolving Radiological Skill (RadSkill) repository to guide the teacher model in distilling high-confidence explicit clinical reasoning, and further internalize this reasoning capability into the implicit diagnostic intuition of the RRG baseline. Building upon this foundation, a novel preference optimization method termed Psychometric Preference Optimization (PSY-PO) is proposed. Grounded in Item Response Theory (IRT) from psychometrics, this method precisely samples clinically misleading dispreferred reports for preference optimization, thereby effectively mitigating plausible hallucinations. Extensive experimental results demonstrate that our proposed method significantly enhances both the generation quality and trustworthiness of radiology reports.
Looped Diffusion Language Models
Sanghyun Lee ⋅ Chunsan Hong ⋅ Seungryong Kim ⋅ Jonghyun Lee ⋅ Jongho Park ⋅ Dongmin Park
Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models for language modeling, yet the effective design of transformer architectures for MDMs remains underexplored. In this paper, we show that selectively looping the early-middle transformer layers significantly improves both training efficiency and model performance in MDMs. We call this approach \textbf{LoopMDM} (Looped Masked Diffusion Model), which brings two key benefits: looping layers at training-time yields a depth-scaling effect without adding parameters, while varying the number of loops at inference-time enables flexible compute scaling. Despite the simplicity, the results are striking: across multiple pre-training corpora, LoopMDM matches the performance of same-size MDMs with up to $3.3\times$ fewer training FLOPs, while its final performance outperforms them on various reasoning benchmarks, including up to $+8.5$ points on GSM8K. It even surpasses deeper non-looped MDMs trained with comparable per-step compute, indicating that selective looping is more effective than naive depth scaling. Furthermore, LoopMDM can scale inference-time compute by increasing the number of loops. Adaptively adjusting the number of loops throughout the sampling process further yields additional gains in compute efficiency while maintaining performance. Lastly, with attention analysis, we provide evidence that looping is effective in MDMs by promoting interactions among masked positions. Our code and weights will be publicly released.
LoopWeaver: Weaving Feedback Loops into Hierarchical Generation under Constraints
Yifan Zhu ⋅ Guanting Chen ⋅ Bing Wei ⋅ Haoran Luo
Large Language Models face a challenging multi-objective coupled optimization problem in constrained long-text generation, requiring the simultaneous balancing of global structure, local coherence, and constraint satisfaction. However, existing methods often rely on static pipelines or loosely coupled designs, making it difficult to achieve cross-stage collaborative optimization. To address this, we propose LoopWeaver, which models the generation process as a hierarchical, feedback-driven closed-loop optimization problem; by introducing iterative feedback between the planning and generation stages, it achieves the joint optimization of both global and local aspects. Specifically, LoopWeaver integrates constraint-aware hierarchical planning, feasibility filtering, and reward-based preference optimization, thereby allowing feedback signals to permeate both the planning and generation layers. Unlike methods that utilize feedback solely during post-processing or within a single module, LoopWeaver introduces a unified paradigm for feedback-coupled optimization, demonstrating significant improvements in text quality and constraint satisfaction capabilities across various models.
LoRAtorio: An intrinsic approach to LoRA Skill Composition
Niki Foteinopoulou ⋅ Ignas Budvytis ⋅ Stephan Liwicki
Low-Rank Adaptation (LoRA) has become a widely adopted technique in text-to-image diffusion models, enabling the personalisation of visual concepts such as characters, styles, and objects. However, existing approaches struggle to effectively compose multiple LoRA adapters, particularly in open-ended settings where the number and nature of required skills are not known in advance. In this work, we present LoRAtorio, a novel train-free framework for multi-LoRA composition that gates each adapter using only signals already inside the denoiser, specifically, patch-level cosine (dis)similarity between each LoRA's predicted noise and the base model's. We refer to this as ``intrinsic'' guidance because no external parameters are introduced. Our method is motivated by two key observations: (1) LoRA adapters trained on narrow domains produce unconditioned denoised outputs that diverge from the base model, and (2) when conditioned out of distribution, LoRA outputs show behaviour closer to the base model than when conditioned in distribution. These patch-level similarities are used to construct a spatially-aware weight matrix, which guides a weighted aggregation of LoRA outputs. To address domain drift, we further propose a modification to classifier-free guidance that incorporates the base model's unconditional score into the composition. We extend this formulation to a dynamic module selection setting, enabling inference-time selection of relevant LoRA adapters from a large pool. On the ComposLoRA benchmark, LoRAtorio achieves state-of-the-art performance, with a CLIPScore that does not deteriorate as the number of composed adapters grows, surpassing the strongest prior method by up to 1.3% in CLIPScore, and reaching a 76.92% GPT-4V win rate against the closest baseline. We further show that the framework is backbone-agnostic: the intrinsic-similarity principle transfers to rectified-flow architectures (FLUX.1-dev) without retraining. Code will be made available.
LoRIF: Low-Rank Influence Functions for Scalable Training Data Attribution
Shuangqi Li ⋅ Hieu Le ⋅ Jingyi Xu ⋅ Mathieu Salzmann
Training data attribution (TDA) identifies which training examples most influenced a model's prediction. Influence function methods are a theoretically grounded family of TDA methods and exploit gradients. To overcome the scalability challenge arising from gradient computation, the most popular strategy is random projection (e.g., TRAK, LoGRA). However, this still faces two bottlenecks when scaling to large training sets and high-quality attribution: (i) storing and loading projected per-example gradients for all $N$ training examples, where query latency is dominated by I/O; and (ii) forming the $D \times D$ inverse Hessian approximation, which costs $O(D^2)$ memory. Both bottlenecks scale with the projection dimension $D$, yet increasing $D$ is necessary for attribution quality---creating a quality--scalability tradeoff. We introduce \textbf{LoRIF} (Low-Rank Influence Functions), which exploits low-rank structures of gradient to address both bottlenecks. First, we store rank-$c$ factors of projected per-example gradients rather than full matrices, reducing storage and query-time I/O from $O(D)$ to $O(c\sqrt{D})$ per layer per sample. Second, we use truncated SVD with the Woodbury identity to approximate the inverse Hessian term in an $r$-dimensional subspace, reducing memory from $O(D^2)$ to $O(Dr)$. On models from 0.1B to 70B parameters trained on datasets with millions of examples, LoRIF achieves up to 20$\times$ storage reduction and query-time speedup compared to LoGRA, while matching or exceeding its attribution quality. LoRIF makes gradient-based TDA practical at frontier scale.
Lost in Translation, Found in Embeddings: Sign Language Translation and Alignment
Youngjoon Jang ⋅ Liliane Momeni ⋅ Zifan Jiang ⋅ Joon Son Chung ⋅ Gul Varol ⋅ Andrew Zisserman
Our aim is to develop a unified model for sign language understanding that performs both sign language translation (SLT) and sign–subtitle alignment (SSA). Together, these two tasks enable the conversion of continuous signing videos into spoken language text and also the temporal alignment of signing with subtitles -- both beneficial for practical communication, large-scale corpus construction, and educational applications. To achieve this, our approach is built upon three design choices: (i) a lightweight visual backbone that captures manual and non-manual cues from human keypoints and lip-region images while preserving signer privacy and enabling end-to-end training; (ii) a Sliding Perceiver mapping network that aggregates consecutive visual features into word-level embeddings to bridge the vision–text gap; and (iii) scaling training to large datasets -- BOBSL covering British Sign Language (BSL) and YouTube-SL-25 covering American Sign Language (ASL) -- to promote generalisation across languages and signers. With this multi-task, multilingual pretraining and strong model design, we achieve state-of-the-art results on BOBSL, How2Sign, Phoenix14T, and FLEURS-ASL SLT benchmarks, and also for SSA on BOBSL and WMT-SLT-SRF. Beyond these standard, signer-interpreted datasets, we show successful SLT on the natural signing data from the BSL-Corpus.
LP-RAG: Learning to Retrieve with Link Predictors
Erik Jhones Freitas do Nascimento ⋅ Jorge Franco ⋅ Amauri Souza
Retrieval-augmented generation (RAG) strategies have empowered large language models (LLMs) through integration with external knowledge sources, leading to more accurate, up-to-date, and contextually relevant outputs. Recently, graph-based RAG methods have gained attention for leveraging relational structures that support multi-hop reasoning and retrieval. However, existing approaches remain limited in their ability to explicitly model query-aware relevance signals during retrieval. In this paper, we propose LP-RAG, a graph-based framework for document RAG that casts retrieval as an inductive link prediction problem. In particular, LP-RAG first constructs a graph encoding semantic relationships among chunks (i.e., candidates for retrieval), which is then augmented with chunk-conditioned synthetic queries that emulate potential user questions. This enables self-supervised learning of query-chunk relevance without relying on real query annotations. Retrieval is performed by predicting links between unseen queries and chunk nodes. Notably, LP-RAG is model-agnostic and can leverage a broad class of link prediction methods. To demonstrate its effectiveness, we evaluate LP-RAG across diverse settings and benchmarks. Our empirical results show that LP-RAG consistently outperforms existing methods in both retrieval quality and downstream generation performance, while maintaining competitive efficiency relative to learnable graph-based RAG approaches.
Lyapunov-Driven Optimistic Learning for Online Scheduling with Multi-Stage Tasks
Yutong Huang ⋅ Xin Liu
We study online scheduling in a shared-resource system where tasks of different types arrive sequentially and compete for limited resources. Each task is modeled as an $H$-stage episodic Markov decision process, which captures the sequential resource-allocation decisions made within the task. At the system level, the learner must coordinate resource sharing across tasks to maximize the long-term ratio of cumulative reward to cumulative cost over $K$ tasks, without knowledge of the task distribution. The key challenges are coupled intra-task Markovian dynamics and inter-task ratio maximization, further complicated by unknown models and large or continuous state spaces. To address these challenges, we propose Contextual Optimistic Lyapunov Decomposition (COLD), a Lyapunov-driven optimistic learning algorithm. COLD converts the long-term ratio objective into a queue-adjusted per-task surrogate, thereby decoupling reward-cost optimization from statistical value estimation. It uses double-optimistic least-squares value iteration to estimate reward and cost values under linear function approximation, and a softmax policy to smooth the greedy update with only a controlled approximation error. We prove a regret bound of $\tilde{\mathcal{O}}(\sqrt{Md^3H^3T})$, where $M$ is the number of task types, $d$ is the feature dimension, $H$ is the episode horizon, and $T=KH$ is the number of interaction steps. The $\sqrt{T}$ dependence is order-optimal up to logarithmic factors, and the bound scales with the feature dimension rather than the size of the state space. Experiments on synthetic environments and an empirical KuaiRand-based MDP demonstrate the effectiveness of the proposed method.
M$^3$: Reframing Training Measures for Discretized Physical Simulations
Yuan Mei ⋅ Xingyu Song ⋅ Xiaowen Song ⋅ Naoya Takeishi
Neural surrogate models for physical simulations are trained on discretized samples of continuous domains, where the induced empirical measure leads to uneven supervision, biasing optimization and causing spatial inconsistencies in physical fidelity. To mitigate this measure-induced bias, we propose M$^3$ (Multi-scale Morton Measure), a scalable framework that balances training measures by partitioning space according to physical variation and allocating supervision across multiple scales. Applied to three industrial-scale datasets with diverse discretizations, M$^3$ consistently improves predictions in the continuous physical domain, achieving up to 4.7$\times$ lower error in large-scale volumetric cases. These gains persist under aggressive subsampling (160M $\rightarrow$ 16M $\rightarrow$ 1.6M points), where M$^3$-trained models outperform those trained on higher-resolution data, reducing physics-weighted relative $L_2$ error by 3--4$\times$ and the corresponding MSE by up to 13$\times$. These results highlight data distribution as a key factor in operator learning and position M$^3$ as a scalable, data-efficient approach for physically consistent modeling.
MainFL: Intertemporal Data Valuation for Robust Auction-based Federated Learning
Bangqi Pan ⋅ Yun Xin ⋅ Jianfeng Lu ⋅ Gang Li ⋅ Shuqin Cao ⋅ Zhou Chendi
Auction-based Federated Learning (AFL) provides a principled framework for incentivizing data owners and allocating decentralized training resources to data consumers. However, practical AFL systems face a coupled valuation and allocation challenge: data quality is uncertain before training, post-training feedback can be strategically distorted, and heterogeneous deadlines render late model updates ineffective for learning. We propose MainFL, an incentive-compatible intertemporal data valuation mechanism for robust auction-based federated learning. MainFL integrates three components: (i) a bounded mutual-information peer prediction rule that elicits truthful pre-training quality reports without requiring ground-truth labels; (ii) an intertemporal reliability calibration scheme that updates future valuations using post-session validation signals; and (iii) a deadline-induced layered clinching mechanism that allocates quality-weighted data while enforcing budget and feasibility constraints.We theoretically establish Bayes–Nash truthfulness for data owners, allocation feasibility and budget safety, and truthful residual-demand equilibrium for data consumers under standard clinching-auction regularity conditions. Empirical results on widely used federated learning benchmarks show that MainFL improves social welfare by up to 12.1% and model accuracy by up to 7.4% compared to state-of-the-art AFL baselines, while substantially reducing straggler-induced task timeouts.
Malicious Node Injection: A Transferable Adversarial Attack on GNN Fairness
Haotian Zhang ⋅ He Huang ⋅ Shuang Cui ⋅ Yu-e Sun
Graph Neural Networks (GNNs) have shown remarkable effectiveness across diverse graph-related tasks. However, existing studies reveal that their fairness is highly vulnerable to adversarial manipulation, which can substantially exacerbate inherent biases toward sensitive attributes, such as gender in demographic prediction. Although prior work has shown that fairness poisoning attacks can be launched via malicious node injection, whether this strategy can support the more challenging fairness evasion attacks remains open. Moreover, existing methods have not directly addressed fairness attacks on multi-class datasets with multi-valued sensitive attributes. To bridge this gap, we propose a universal fairness attack framework (UFA) that injects malicious nodes to perform both poisoning and evasion attacks against GNNs while minimally affecting model utility. UFA first identifies nodes most susceptible to being pushed across the decision boundary. It then targets core mechanisms shared by diverse GNN architectures to manipulate node representations, thereby maximizing outcome disparities across sensitive groups. Comprehensive experiments on five real-world datasets demonstrate the effectiveness of UFA. By injecting fewer than 0.2\% malicious nodes, UFA severely degrades the fairness of both mainstream and fairness-aware GNNs, proving effective in both poisoning and evasion settings. Our findings expose the fragility of fairness in GNNs and underscore the urgent need for robust fairness-aware models.
Map-Guided Caching: A Global Perspective for Efficient Diffusion Transformer
Jiaqi Ji ⋅ Ran Yang ⋅ Bo Wei ⋅ Hui Li ⋅ Hee Min Choi
Feature caching is a promising way to accelerate Diffusion Transformers (DiTs) without modifying the model backbone, but existing methods largely rely on local heuristics or input-agnostic rules when deciding which intermediate features to cache and reuse. We present Map-Guided Caching, a feature caching framework that predicts a Temporal Variation Map (TVMap) for each denoising trajectory and uses the map to derive a global caching policy. The TVMap summarizes block-wise temporal variation across timesteps, while a lightweight Map Predictor estimates this variation from the prompt and initial noise before denoising starts. Given the predicted map, we formulate policy generation as a global path planning problem and combine a dynamic-programming-based orientation planner with a quantity planner that targets a desired speedup ratio. Across image generation, image editing, and video generation, the proposed method achieves up to $3.3\times$ speedup on FLUX.1 [dev] and HunyuanVideo, and $2.7\times$ on FLUX.1 Kontext [dev], while preserving competitive output quality. The advantage is especially pronounced on complex prompts, where local caching policies degrade more sharply. These results suggest that input-adaptive, globally planned caching is an effective alternative to purely local feature reuse policies.
MapPolicy: Structure-Aware Imitation Learning for Robot Manipulation via Physically Constrained Scene Map
Zehao Du ⋅ Zheng Huang ⋅ Rui Li ⋅ Hongyu Liu ⋅ Jiude Wei ⋅ Jiachun Bao ⋅ Cewu Lu ⋅ Jianhua Sun
Robot manipulation under partial observability requires policies to act from incomplete visual evidence while still reasoning about object structure, spatial layout, and physical relations. Existing imitation learning methods mainly encode currently visible observations, which makes them brittle when task-critical parts, contact regions, or support relations are occluded. We introduce Scene Map, a structured scene representation that models object shape, part structure, spatial layout, and physical constraints as a graph of manipulation-relevant primitives and relations. Scene Map turns occluded but task-relevant structure into explicit prediction targets and provides structural priors for inferring unseen components from partial observations. Building on Scene Map, we propose MapPolicy, an imitation learning framework that fuses structured map features with visual and robot state features for action prediction. MapPolicy uses structure-modulated graph attention to propagate global structural and spatial information through primitive features, and a physical constraint loss to regularize learned features with geometric and kinematic consistency. Across 49 manipulation tasks in three simulation benchmarks and five real world tasks, MapPolicy consistently improves over strong imitation learning baselines, with especially clear gains under severe occlusion. Our code will be released publicly to facilitate reproducibility and further research.
Masked Diffusion Vision-Language Models for Temporal Action Localization
Fengshun Wang ⋅ Zhengbo Zhang ⋅ Zhigang Tu
Temporal action localization (TAL) requires recognizing the target event and localizing its start and end times precisely in untrimmed videos. Recent vision-language formulations improve semantic reasoning and support language-conditioned outputs, but their autoregressive decoders still generate tokens from left to right, preventing later semantic evidence from revising earlier timestamp predictions. We adapt masked diffusion vision-language models (MDVLMs) to TAL so that semantic tokens and boundary tokens remain editable throughout iterative denoising with bidirectional attention, allowing temporal boundaries and semantic content to be refined jointly. Direct adaptation, however, creates two TAL-specific mismatches: standard masked diffusion training corrupts all positions uniformly at random, but the time tokens are more reliable when enough semantic context is available; and token-level cross-entropy does not reflect temporal IoU. To address these mismatches, we introduce a Planned Training Objective that uses boundary-aware masking and step-weighted reconstruction to rehearse the late recovery of time tokens, together with a Step-Level IoU Reward that provides overlap-aware supervision during denoising. A standard sequence-level cross-entropy term provides the base reconstruction signal. Experiments on ActivityNet-RTL, ActivityNet-1.3, and THUMOS-14 show that MDVLM-TAL improves both temporal reasoning and boundary localization over autoregressive vision-language baselines, with especially strong gains under stricter temporal IoU criteria.
MASTA: A Feedback-Scheduled Multi-Agent System for End-to-End Tamarin Protocol Modeling and Analysis
Yongkang Xiao ⋅ Qiyi Deng ⋅ Jing Chen ⋅ Min Shi ⋅ JU MA ⋅ Ruiying Du
Tamarin-based formal verification of security protocols requires expert-authored multiset rewriting rules and temporal-logic lemmas. Translating a natural-language protocol description into a complete formal model with an LLM is hard because syntactic, reachability, and property-level errors compound across the multi-step modeling process, and a single agent that both authors and self-evaluates cannot reliably localize them. We present MASTA, a multi-agent system that addresses these difficulties through three innovations. First, an operational--evaluative role decomposition assigns rule modeling to an Architect, property modeling to an Auditor, and pure dispatch to a dedicated Arbiter that holds no authoring duties. Second, the Arbiter implements a Tamarin-anchored scheduler that dispatches the next agent based on the latest verifier verdict and routes problem-tagged fixes through a modification list, with a six-class problem taxonomy and an LLM deduplication check preventing routing deadlock. Third, we construct a benchmark of 10 protocols paired with 139 mutation variants and five fine-grained metrics that jointly assess model correctness ($M_1$--$M_3$) and completeness ($M_4$--$M_5$). On this benchmark MASTA on Qwen-3.5-Plus improves over OpenCode on the same backbone by up to $+0.27$ on individual metrics and reaches the highest overall average, surpassing Codex (GPT-5.5) and Claude Code (Opus~4.7) despite using a weaker LLM. Ablations confirm that removing the Arbiter collapses property and verification scores to near zero.
Matching2Matching: Zero-Shot Light Field Image Denoising with Matching View Construction
Shuo Zhang ⋅ Qian Tian ⋅ Chen Gao ⋅ LongzhaoGuo ⋅ Youfang Lin
Recently, Noise2Noise provides a powerful principle for learning image denoising without clean data. However, zero-shot methods for 2D images often rely on local similarity or non-local self-similarity to construct image pair for training, which is unreliable in complex scenes with highly textures. Different from 2D image, 4D Light Field (LF) captures multiple views for one specific scene. The multiple views with independent noises provides make it possible to find image pair across views. In this paper, we exploit the inter-view correspondence in LF and propose a novel way to implement the Noise2Noise in 4D LF. Specifically, we design a Matching2Matching (M2M) framework to construct matching views for each view in LF. This framework generates geometrically consistent training samples by searching for corresponding patches along the epipolar line, crucially enforcing both visible and disparity consistency constraints. The constructed views enable the network to train the denoised network in a zero-shot way without any other data. Experimental results show that our method effectively suppresses noise while outstandingly preserving the consistency and fine details. Our M2M framework significantly outperforms existing zero-shot methods in both quantitative and qualitative evaluations on synthetic and real-world noise, demonstrating superior generalization ability without relying on external training datasets.
MatCurvs: Article Real-Coordinate Curve Extraction for Agent-Ready Materials Reasoning
Liang Yin ⋅ Zhan'ao yao ⋅ Jiahui Shi ⋅ Songlin Yu ⋅ Jianjun Liu
In materials-science papers, quantitative evidence—XRD peak shifts, Raman ID/IG ratios, high-frequency Z′ intercepts in EIS—lives in line plots. Hosted vision–language models describe these plots fluently but cannot recover the (x, y) values an agent needs to query, compare, or integrate. We release three coupled artifacts. MatCurvs is a real-coordinate chart benchmark for materials-science spectra, organized in three tiers spanning XRD, Raman, XPS, XAFS, EIS, GCD, and CV: 1,375,165 classified subfigures from 55,763 articles (L0); 204 panels with manual per-curve pixel polylines (L1); and 8,026 verified real-coordinate curves over 3,817 panels (L2). MatDeplot is a local extractor that recovers 45.8% of curves within a 5-pixel IoU tolerance against L1 ground truth, versus ≤1.2% for the strongest hosted VLM—a 38× pixel-anchoring gap, closed in roughly 1.5 seconds per image. MatCurvs-Reasoning is a 14,740-question deterministic numerical evaluation: feeding MatDeplot’s extracted (x, y) data to an LLM cuts median relative error from 54% for an image-only VLM to about 4%, and a text LLM with the data alone matches the same VLM prompted with both image and data. By turning each published chart into a numerically queryable curve, the released artifacts let a deployer compute derived materials parameters—Raman ID/IG disorder ratios, XRD Scherrer crystallite sizes, and phase-purity audits—directly from a literature-scale chart corpus, where most papers under-report these numbers and the chart is the only available record.
MathCD: A Benchmark Dataset for Cognitive Diagnosis with Semantic Information
Xueyi Li ⋅ Youheng Bai ⋅ Tengteng Cheng ⋅ Mingliang Hou ⋅ Teng Guo ⋅ Jiaqi Zheng ⋅ Yongdong WU ⋅ Zitao Liu
Cognitive diagnosis (CD) aims to infer students' knowledge states from historical learning interactions, providing essential support for personalized education. Despite recent progress, most existing CD datasets mainly represent students, exercises, and knowledge concepts as discrete identifiers, offering limited semantic information about exercise text, answer choices, and student responses. This limitation restricts the development of semantic enhanced CD models and makes it difficult to evaluate whether foundation large language models (LLMs) can demonstrate educational understanding in CD scenarios. To address this gap, we introduce \emph{MathCD}, a new K-12 mathematics CD dataset collected from an online learning platform. MathCD contains three subsets with $2,197$ students, $46,431$ exercises, $3,258$ knowledge concepts, and $106,220$ student-exercise interactions. Moreover, MathCD provides rich semantic annotations, including exercise text, KC text, answer text, response text, exercise type, and exercise difficulty. Based on MathCD, we evaluate $9$ existing CD models and $5$ representative foundation LLMs. Experimental results show that semantic information consistently improves both classical and neural CD models, with semantic-enhanced variants achieving comparable or better performance than strong identifier based models. We further find that foundation LLMs exhibit emerging CD-related abilities in student response prediction and exercise difficulty discrimination, but their performance remains unstable and limited without task-specific adaptation. MathCD provides a new benchmark for semantic-enhanced CD and foundation LLMs evaluation in educational modeling.
Matryoshka Transcoders and Hierarchy Misalignment: When SAE Absorption Protection Does Not Transfer
Quoc Dat Ngo ⋅ Hung Nguyen
Sparse autoencoders decompose activations into sparse, interpretable features and have become a central tool for mechanistic interpretability. Transcoders extend this approach from representation reconstruction to computation prediction by mapping MLP inputs to MLP outputs, making them attractive for circuit-level analysis. We ask whether a prominent sparse-autoencoder reliability fix, Matryoshka nested-prefix training, transfers to this two-hook setting. In sparse autoencoders, Matryoshka training improves absorption behavior by encouraging broad features to appear in early dictionary prefixes. In transcoders, we find that this protection does not transfer automatically. Across Gemma-2-2B MLP transcoders trained under a matched recipe, Matryoshka models form genuine prefix hierarchies, with reconstruction improving as later groups are added, but they do not inherit the absorption advantage observed for sparse autoencoders. BatchTopK achieves lower full absorption, lower FVU, and higher first-letter F1@1 in the canonical 16k comparison. Bridge controls recover the expected Matryoshka advantage when the cross-space map is removed, showing that the failure is specific to predicting MLP outputs from MLP inputs. We identify the mechanism as hierarchy misalignment. Prefix losses reward covariance with the reconstruction target, so early transcoder groups prioritize target-predictive output geometry rather than broad source-space semantic parents. These results show that hierarchical sparse-dictionary objectives do not preserve interpretability properties by default. For dictionaries that map between activation spaces, the key question is not whether a hierarchy forms, but what geometry that hierarchy is trained to organize.
MaxSketch: Robust Distinct Counting in Streams via Random Projections
Nikos Tsikouras ⋅ Constantine Caramanis ⋅ Christos Tzamos
Estimating the number of distinct elements in a data stream is well understood when repeated elements are identical. In modern settings, however, observations are high-dimensional and noisy, so repeated instances of the same object are only approximately similar -- for example, different images of the same individual may vary significantly at the pixel level. Classical sketches such as HyperLogLog rely on consistent hash values for identical elements and break down in this regime. Recent work on robust distinct counting in general metric spaces achieves $\tilde\Theta(\sqrt{n})$ memory, which is tight in the worst case. We show that substantially improved memory guarantees are possible under geometric structure common in learned representations. We introduce MaxSketch, a simple max-linear sketch built from random Gaussian projections, and prove that it succeeds in estimating the number of distinct latent objects. Concretely, we show that under this assumption $m = \tilde{O}(\log n / \varepsilon^4)$ random projections (and hence $\tilde{O}(\log n/\varepsilon^4)$ memory) suffice to recover the true distinct count within a $(1+\varepsilon)$ factor. Experiments on image streams confirm that MaxSketch accurately estimates distinct counts and generalizes beyond the training regime. Our results bridge classical streaming algorithms and modern representation learning, showing how geometric structure can fundamentally reduce the complexity of distinct counting.
MCGI: Manifold-Consistent Graph Indexing for Billion-Scale Disk-Resident Vector Search
Dongfang Zhao
Graph-based Approximate Nearest Neighbor (ANN) search often suffers from performance degradation in high-dimensional spaces due to the Euclidean-Geodesic mismatch, where greedy routing diverges from the underlying data manifold. To address this challenge, this paper presents Manifold-Consistent Graph Indexing (MCGI), a geometry-aware and disk-resident indexing method that leverages Local Intrinsic Dimensionality (LID) to dynamically adapt search strategies to the intrinsic geometry of data. Unlike conventional algorithms that treat dimensions uniformly, MCGI modulates its beam search budget based on in-situ geometric analysis, which reduces sensitivity to data-specific hyperparameters by replacing a single scalar with a geometry-informed range that remains stable across datasets of varying dimensionality. Theoretical analysis demonstrates that MCGI provides robust approximation by preserving manifold-consistent topological connectivity. Extensive evaluations against five industry-standard baselines across five datasets up to billion scales confirm the advantages of the proposed approach.
MCPHallu: Benchmarking Reasoning, Execution, and Memory Hallucinations in MCP Agents
Yizhen Jiang ⋅ 浩伟 郭 ⋅ Tianyi Bai ⋅ Zengjie Hu ⋅ Lu Ma ⋅ Zhengyang Zhao ⋅ Haoze Sun ⋅ Peng Pei ⋅ Wentao Zhang
Large language model agents are increasingly deployed through the Model Context Protocol (MCP), which allows them to discover and invoke tools across multiple servers at runtime. Existing MCP benchmarks primarily measure task completion capability but provide limited insight into why agents fail, reporting aggregate success rates without attributing failures to specific reasoning, execution, or memory breakdowns. We introduce MCPHallu, a benchmark for diagnosing hallucination-prone agent failures in MCP environments along three axes: reasoning, execution, and memory. MCPHallu contains 358 tasks across five capability domains and runs on 36 production MCP servers in isolated containers. Each task is designed with a structural trigger that makes one of four subtypes the dominant failure risk: Branch Collapse, Unreachable Goal, Tool Misuse, or Context Forgetting. This design enables subtype-level analysis rather than aggregate task success alone. We evaluate fourteen frontier LLM agents from eight vendors on 5,012 trajectories. Current MCP agents remain vulnerable to hallucination-inducing tool environments, achieving an average score of 0.642. Failures concentrate on Unreachable Goal tasks, where the average score drops to 0.458 and 22.1\% of runs pass, indicating that agents often continue issuing tool calls or claim successful completion rather than recognizing infeasible goals. Performance across the four subtypes is largely independent: per-model ranks span on average 6.9 of 13 positions, and Branch Collapse is uncorrelated with Unreachable Goal across models ($r = -0.12$). MCPHallu is released at https://anonymous.4open.science/r/mcp_hallu.
MCSplat: Multi-View Photometric and Geometric Consistent Feed-Forward Gaussian Splatting for Driving Scenes
Wang Yifei ⋅ Wei Tian ⋅ WENHUA BAO ⋅ Jia Zhou ⋅ Hanshi Wang ⋅ Fudong Ge ⋅ Jianhua Wu ⋅ Shuai Yuan ⋅ Zhipeng Zhang ⋅ Xingliang Liu ⋅ Xianming Zeng ⋅ Jianyun Xu ⋅ Bingzhao Gao
Driving-scene reconstruction via feed-forward Gaussian splatting from sparse multi-view images remains challenging due to two factors. First, cross-camera photometric inconsistency—caused by variations in exposure, ISP, and illumination—can be incorrectly absorbed into Gaussian geometry. Second, unreliable local geometric supervision arises from normal-based constraints near depth discontinuities. To address these issues, we present MCSplat, a multi-view consistent feed-forward Gaussian splatting framework for driving scenes. MCSplat constructs a canonical appearance space from depth-guided multi-camera correspondences to decouple geometry learning from camera-specific appearance. A multi-scale Bilateral Grid Transformer restores raw camera appearance from canonical renderings, enabling photometric supervision without contaminating Gaussian geometry. To improve local geometric reliability, we estimate Fisher-information based uncertainty from photometric sensitivity and use it to adaptively weight surface-oriented self-supervision. Experiments on the Waymo Dataset show that MCSplat achieves state-of-the-art performance on both the full validation set and an appearance-challenging subset with strong cross-camera photometric inconsistency, demonstrating improved reconstruction quality and more consistent novel-view geometry. Code will be released.
MDK-MoE: Multi-view Decomposed Kalman Mixture of Experts Framework for Non-stationary Time Series Forecasting
Rui Hou ⋅ Yao Liu ⋅ Ruilin Jiang ⋅ Mengyao Lu ⋅ Jingbo Wang ⋅ Lanyi Zhang ⋅ Qiao Liu
Non-stationary time series forecasting (NTSF) requires modeling evolving trends and seasonal patterns. Although linear decomposition operators (e.g., FFT) help isolate these patterns, they often introduce operator-dependent biases and idiosyncratic noise. To address this issue, we propose a novel Multi-view Decomposed Kalman Mixture of Experts (MDK-MoE) framework, which \textit{improves forecasting by strategically integrating multiple complementary, yet biased, linear decompositions}. Specifically, a Multi-view Decomposition (MVD) extracts frequency components from three complementary views. To mitigate view-specific noise, a Maximizing Mutual Information (MMI) enhances cross-view consistency, maximizes task-relevant information across diverse views, and suppresses view-dependent task-irrelevant information. To address the conflicts inherent in multi-source observations, the efficient Kalman MoE adaptively fuses distinct multi-view features. Instead of adopting traditional $O(d^3)$ covariance inversion, our learnable Kalman gain design enhances numerical stability and enables reliable state estimation under non-stationary scenarios. Extensive experiments on 13 real-world datasets demonstrate that MDK-MoE consistently achieves state-of-the-art performance.
Measure-to-measure Regression with Transformers
Matthew Vandergrift ⋅ Martha White ⋅ Yury Polyanskiy ⋅ Philippe Rigollet ⋅ Lazar Atanackovic
Many learning problems require predicting how populations evolve under an unknown transformation. A natural representation for such populations is a probability measure, with point clouds as a key example. In this work, we study the measure-to-measure regression (M2M) problem, in which one seeks to learn a map between probability measures from a finite collection of observed input--output pairs. In contrast to classical regression, where individual samples are transformed independently, M2M regression treats entire distributions as the data points. This perspective is vital in certain scientific applications, for example, cellular and molecular biology, where cells are known to evolve not as independent data points but as a distribution. However, few existing approaches address the problem of M2M regression with sufficient expressivity and scalability. We present a formalization of nonlinear M2M regression and introduce two simple, expressive, and scalable approaches to learn such operators: transformers as static M2M maps and transformers as dynamic M2M velocity fields. Our approach leverages the natural measure-dependent and mean-field structure of transformers to learn nonlinear M2M maps on the space of probability distributions. We illustrate the effectiveness of our proposed method to generalize to unseen measures on synthetic experiments, canonical interacting particle systems, and a large-scale patient-derived organoid dataset for predicting treatment response in colorectal cancer.
MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
Jiadong Zhang ⋅ Xiaosong Ma
Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory benchmarks often under-test the combination of activity-dense interaction, ego-centric perspective, and coherent multi-session worlds. MemArena fills these gaps with a single-world conversational benchmark built with its MASim agent simulator, for $50$ agents over $15$ days ($10.3$M dialog-text tokens, $24.1$K text-only ego-observed tokens/agent/day). With the interaction history, it co-generates ground truth over six recall, reasoning, and trustworthiness evaluation dimensions. We evaluate five open-weight readers with Vanilla context, BM25-RAG, Oracle retrieval, Memobase, and MemSearch as memory backends. Three results stand out: (1) Memory-backend choice matters more for content accuracy: At Qwen3-0.6B, Memobase $\to$ MemSearch gains $+32.5/+19.2$ pp, exceeding MemSearch reader scaling ($+10.6/+6.8$ pp). (2) Permission-aware access fails universally: with Oracle leaking heavily and other backends too timid to disclose. (3) Search latency bites only at very small reader: on a Spark GB10 edge node, memory-search adds a moderate and fixed $87$/$7$/$48$ ms (BM25-RAG/Memobase/MemSearch) that composes a small part of TTFT for most reader-backend combinations. We release code, the MASim simulator, and the MemArena-L benchmark at https://anonymous.4open.science/r/MemArena-NeurIPS2026-anon-3E9D
MemDLM: Memory-Enhanced DLM Training
Zehua Pei ⋅ Hui-Ling Zhen ⋅ Weizhe Lin ⋅ Sinno Pan ⋅ Yunhe Wang ⋅ Mingxuan Yuan ⋅ Bei Yu
Diffusion Language Models (DLMs) offer attractive advantages over Auto-Regressive (AR) models, such as full-attention parallel decoding and flexible generation. However, standard DLM training uses a static, single-step masked prediction objective that never exposes the model to the progressive denoising dynamics of inference, and forces all contextual information to be maintained purely through token-space attention, which becomes increasingly diluted as context length grows. We propose MemDLM (Memory-Enhanced DLM), which introduces a second memory channel by embedding a simulated denoising trajectory into training via Bi-level Optimization. An inner loop updates a set of fast weights, forming a Parametric Memory that captures the local trajectory experience, while an outer loop updates the base model conditioned on this memory. By offloading part of the memorization burden from token-space attention to parameter space, MemDLM yields faster convergence, stronger long-context representations, and lower training loss, even when the fast weights are discarded at inference time. Re-enabling the inner loop at inference provides an additional prompt-specific adaptation effect, where the Parametric Memory acts as an emergent in-weight retrieval mechanism on challenging Needle-in-a-Haystack tasks.
Memory Grafting: Scaling Language Model Pre-training via Offline Conditional Memory
Runxi Cheng ⋅ Yuchen Guan ⋅ Yongxian Wei ⋅ Qianpu Sun ⋅ Qixiu Li ⋅ Sinan Du ⋅ Feng Xiong ⋅ Chun Yuan ⋅ Yan Lu ⋅ Yeyun Gong
Large language models acquire rich internal knowledge during pre-training, yet compact models are still typically forced to relearn such knowledge from raw data or imitate it indirectly through distillation. We investigate a different route: can pretrained knowledge itself be directly reused during pre-training? To this end, we propose Memory Grafting, a representation-level transfer paradigm that extracts structured memories from a large frozen donor model and injects them into a compact student throughout training. Unlike knowledge distillation, which transfers behavior through output alignment, or retrieval-based methods, which expose external knowledge only at inference time, Memory Grafting makes pretrained knowledge part of the student’s optimization process itself. This suggests that knowledge in large language models can be modular, transplantable, and reusable across scales, pointing to a new path toward more efficient language model pre-training.
MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents
Jifeng Gao ⋅ Kang Xia ⋅ Yi Zhang ⋅ Xiaobin Hong ⋅ Mingkai Lin ⋅ Xingshen Wei ⋅ Wenzhong Li ⋅ Sanglu Lu
Persistent external memory enhances agent continuity but introduces persistent security vulnerabilities: adversarial content can be injected via standard interaction channels, retained across turns, and later distort downstream behavior. To address this challenge, we propose MemPoison, a comprehensive benchmark and analysis framework featuring 1227 hand-validated cases across four attack types, three injection channels, and three representative memory substrates, evaluated on seven open-weight and three closed-weight model families. We introduce a three-tier taxonomy: (L1) direct single-record corruption, (L2) compositional multi-record corruption and (L3) context-triggered dormant corruption. Our evaluations reveal a distinct defense frontier: while baseline write-time defenses, such as consistency checks, substantially suppress direct L1 attacks, they fail to reliably suppress L2 and L3 attacks. Through mechanistic influence decomposition (MID), we demonstrate structural blind spots in write-time defenses, which admit seemingly benign records that later become harmful through joint retrieval composition or trigger-conditioned activation. Our findings advocate for shifting from static filtering to adaptive, context-sensitive memory defense strategies. Code, benchmark artifacts, and evaluation scripts are publicly released
MeshFIM: Local Low-Poly Mesh Editing via Fill-in-the-Middle Autoregressive Generation
Dingdong Yang ⋅ Jian Liu ⋅ Biwen Lei ⋅ Haohan Weng ⋅ Zhuo Chen ⋅ Song Guo ⋅ Hao Zhang ⋅ Ali Mahdavi Amiri ⋅ Chunchao Guo
Autoregressive (AR) models can generate high-quality low-poly meshes from point clouds, but they still operate in an all-or-nothing manner: when a local region is unsatisfactory, the entire mesh must be regenerated, wasting computation and destroying satisfactory mesh structure elsewhere. We introduce MeshFIM, a Fill-in-the-Middle (FIM) framework that regenerates a target region of a low-poly mesh conditioned on the surrounding context. MeshFIM addresses three mesh-specific challenges: enforcing exact attachment along the exposed boundary, preserving topological order in the context, and suppressing overflow beyond the intended region. It does so with five complementary design choices: boundary vertex markers, context positional embeddings, expanded context width, context augmentation, and a low-poly geometry encoder whose gated subtraction mechanism focuses generation on the missing region by leveraging the difference between the reference surface and the existing mesh. Detailed ablation studies are presented to show the effectiveness of every introduced component. Based on MeshFIM, we demonstrate two applications: interactive brush-based editing and automatic defect repair on low-poly mesh (see Figure 1). Last but not least, experiments show that MeshFIM outperforms a range of baselines in mesh refinement, mesh repair and whole mesh generation plus stitch-back scheme.
Mesh-Free Convolution: Learned Spectral Attenuation and Transport
Niloufar Zakariaei ⋅ Manav Manav ⋅ Eldad Haber
Learning solution operators for partial differential equations (PDEs) requires models that can represent spatial structure across grids and scales. This becomes especially important in time-dependent problems, where the learned operator is often used autoregressively: each predicted state is fed back as input to predict future states. In this setting, small spatial errors can be repeatedly propagated through the learned dynamics, leading to instability and drift, vitiating long-term solution fidelity. We propose Mesh-Free Convolutions (MFCs), filters defined in continuous space and discretized only at evaluation time. This allows the same learned filters to be applied across resolutions. Rather than imposing a grid-dependent stencil or a fixed spectral cutoff, MFCs learn smooth scale-dependent filtering and include directional terms for transport-like behavior. Across fluid-dynamics and weather-forecasting benchmarks, MFC-based models produce stable long-horizon rollouts, transfer reliably across resolutions, and better preserve long-time statistics. These results suggest that continuous, mesh-free filtering is a useful inductive bias for PDE operator learning, especially in autoregressive prediction.
MESS: Multi-Exposure Sequence Synthesis for Generalizable Image Enhancement
Ruodai Cui ⋅ Shuaizheng Liu ⋅ Rongyuan Wu ⋅ Lei Zhang
Despite the significant progress in image enhancement under adverse illuminations, such as low-light image enhancement, exposure correction, and backlit image enhancement, existing models show limited generalizability to different scenarios. A major bottleneck lies in the lack of large-scale and diverse training data, as these tasks typically rely on manually captured, precisely aligned image pairs that are expensive to scale up. Although synthetic data have been explored as an alternative, existing synthesis methods are limited in modeling complex adverse illumination conditions and camera imaging pipelines. To address this issue, we propose a novel framework to synthesize realistic training pairs for generalizable image enhancement in the wild. Specifically, we first derive multi-exposure sequences (MES) from high-bit RAW images by rendering them under different exposure settings through an emulated ISP pipeline. Using these RAW-derived MES as supervision, we train a diffusion-based RGB-to-MES generator to synthesize plausible exposure sequences from a single 8-bit RGB image, extending the synthesis pipeline from limited RAW collections to large-scale in-the-wild RGB data. For each source image, we randomly sample one synthesized exposure variant as the degraded input and use the original high-quality RGB image as the target. With the proposed $\textbf{MES S}ynthesis (\textbf{MESS})$ approach, we construct a dataset of 500K training pairs, and train a lightweight network, which however demonstrates significantly better generalization performance than existing models across various image enhancement tasks. Code, dataset, and models will be released.
MetaKE: Meta-Learning for Knowledge Editing Toward a Better Accuracy-Editability Trade-off
Shuxin Liu ⋅ Di Gao ⋅ Ou Wu
Existing locate-then-edit Knowledge Editing (KE) methods typically decompose editing into two stages: upstream target representation optimization and downstream constrained parameter optimization. The optimization across the two stages is disconnected: upstream applies uniform regularization without observing downstream realization of the planned residual, hindering a refined accuracy–editability trade-off. Since this realization is request-specific and depends on downstream constraints, uniform regularization can over-shrink high-association requests, causing insufficient editing, while it can under-regularize low-association requests, producing over-large planned residuals that reduce downstream editability. To bridge this disconnect, we propose $\textbf{MetaKE}$ ($\textbf{Meta}$-learning for $\textbf{K}$nowledge $\textbf{E}$diting), a new framework that unifies upstream and downstream stages into a bi-level optimization problem. The inner level optimizes parameter updates for the target representation, while the outer level optimizes representation using feedback from downstream constraints, achieving a better semantic accuracy-editability trade-off. To avoid costly multi-layer backpropagation, we introduce a Structural Gradient Proxy to approximate and propagate this feedback. Extensive experiments show that MetaKE outperforms strong baselines, offering a new perspective on KE. Code at: https://anonymous.4open.science/r/MetaKE-C160.
Meta-TTRL: A Metacognitive Framework for Self-Improving Test-Time Reinforcement Learning for T2I Generation in Unified Multimodal Models
Lit Sin Tan ⋅ Junzhe Chen ⋅ Xiaolong Fu ⋅ Lichen Ma ⋅ Junshi Huang ⋅ Jianzhong Shi ⋅ Yan Li ⋅ Lijie Wen
Test-time scaling (TTS) improves text-to-image (T2I) generation in unified multimodal models (UMMs), but existing TTS methods typically rely on frozen-parameter search or sampling, yielding only ephemeral, instance-level improvements. We propose Meta-TTRL, a metacognitive test-time reinforcement learning (TTRL) framework that converts test-time generation experience into learning signals, enabling UMMs to self-improve during T2I generation without external reward models. Meta-TTRL formulates test-time learning as a closed-loop interaction between an object-level generator and a meta-level introspector. The introspector decomposes prompts into structured verification rubrics and produces confidence-enhanced intrinsic monitoring signals, which are used to optimize the generator policy through reinforcement learning (RL). Extensive experiments demonstrate that Meta-TTRL generalizes well across three representative UMMs, including Janus-Pro-7B, BAGEL, and Qwen-Image, achieving significant gains on compositional reasoning tasks and multiple T2I benchmarks with limited data. Further analyses with external introspectors, RL leakage, and alternative monitoring signals reveal a key principle for effective TTRL: metacognitive synergy, where monitoring signals must align with the model's own optimization regime.
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
Yankai Fu ⋅ Ning Chen ⋅ Junkai Zhao ⋅ Shaozhe Shan ⋅ Guocai Yao ⋅ Pengwei Wang ⋅ Zhongyuan Wang ⋅ Shanghang Zhang
Building a generalist robot that can perceive, reason, and act across diverse tasks remains an open challenge, especially for dexterous manipulation. A major bottleneck lies in the scarcity of large-scale, action-annotated data for dexterous skills, as teleoperation is difficult and costly. Human data, with its vast scale and diverse manipulation behaviors, provides rich priors for learning robotic actions. While prior works have explored leveraging human demonstrations, they are often constrained by limited scenarios and a large visual gap between human and robots. To eliminate these limitations, we propose METIS, a vision-language-action (VLA) model for dexterous manipulation pretrained on multi-source egocentric datasets. We first construct EgoAtlas, which integrates large-scale human and robotic data from multiple sources, all unified under a consistent action space. We further extract motion-aware dynamics, a compact and discretized motion representation, which provides efficient and expressive supervision for VLA training. Built upon them, METIS integrates reasoning and acting into a unified framework, enabling effective deployment to downstream dexterous manipulation tasks. Our method demonstrates exceptional dexterous manipulation capabilities, achieving highest average success rate in eight real-world tasks. Experimental results also highlight the superior generalization and robustness to out-of-distribution scenarios. These findings emphasize METIS as a promising step toward a generalist model for dexterous manipulation.
MFlowAudio: Efficient Text-to-Audio Synthesis via Mamba-based Stateful Flow Matching
Hao Dai ⋅ Panyu Chen ⋅ Jagmohan Chauhan
Recent advancements in audio generation have been largely driven by Transformer-based diffusion models. However, these models suffer from quadratic complexity of self-attention, severely bottlenecking the long-form audio synthesis. To overcome this limitation, we propose MFlowAudio, a novel latent audio generation framework that synergizes the continuous-time dynamics of Flow Matching with a custom-designed TFMamba backbone. TFMamba uses an innovative dual-scan mechanism: a TimeMamba module to capture long-range causal dependencies with linear complexity, and a FrequencyMamba module to model spectral correlations such as harmonic structures. Exploiting this structural foundation, we formulate the Stateful Flow Matching (SFM) paradigm. This framework inherently enables chunk-wise training and streaming generation, maintaining an $\mathcal{O}(1)$ caching complexity without incurring extra computational overhead. To enable fine-grained controllable synthesis, we devise a novel guidance mechanism that neutralizes vector field collisions precipitated by on-the-fly semantic transitions of prompts. Comprehensive empirical evaluations confirm that MFlowAudio yields generative fidelity on par with state-of-the-art baselines while establishing significant computational efficiency, achieving an $81.8$% acceleration in generation and facilitating perceptually seamless streaming synthesis with a constant, ultra-low latency of $1.8$ seconds. Demo:https://huggingface.co/spaces/mflowaudio/MFlowAudio
MGMem: An Efficient, Deterministic, and Provenance-Preserving Framework for Long-Horizon Agent Memory
Junhong Huang ⋅ Xin Tong ⋅ Jun Xia
Long-horizon Large Language Model (LLM) agents need persistent memory that surfaces the right past interaction at the right time across thousands of turns. Current systems index by invoking generative LLMs to rewrite history into summaries, facts, or knowledge graphs---a step that is (a) costly (up to $\sim 50$K LLM calls and $35.83$M completion tokens per build), (b) non-deterministic (undermining reproducibility and giving inconsistent answers across runs), and (c) provenance-erasing (the reader sees synthesized text, not the original turn). We propose MGMem, an LLM-call-free framework that replaces this step with a deterministic discriminative-parser pipeline: recurring mentions across sessions are organized into an incidence graph, question mentions are resolved by a subject-conditioned three-stage resolver, and raw dialogue episodes are returned to the reader. On LoCoMo, MGMem outperforms matched-protocol generated-memory and retrieval baselines at zero generative indexing calls; on the full LongMemEval-S, it remains competitive overall and leads on single-session question types. Anonymous code: \url{https://anonymous.4open.science/r/mentiongraph}.
MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning
Ruilin Tong ⋅ Dong Gong
Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most procedures treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience across them can further improve performance. Existing memory-based methods either store whole-solution templates that generalize poorly to novel problems or use heuristic step-level selection that is not optimized for final-answer correctness. Learning selection policies requires large-scale training data and fixed action spaces. We propose MILES (Modular Instruction Memory with LEarnable Selection for self-improving LLM reasoning), a framework that dynamically expands step-wise memory and applies correctness-optimized memory composition under realistic test-time constraints. MILES maintains modular memory units consisting of asymmetric pairs of sub-goal embeddings and sub-instructions, each associated with a learnable selection head. This memory structure enables a coarse-to-fine retrieval mechanism: The coarse level enables memory expansion and collects supervision for training selection heads from confident samples, while the fine stage applies learned selection heads to rerank coarse-level candidates and guide reasoning for uncertain samples. MILES consistently matches or outperforms prior methods while achieving superior accuracy–efficiency tradeoffs. Extensive experiments demonstrate its effectiveness, robustness, and transferability.
MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding
Zexuan Chen ⋅ Sichao Liu ⋅ Runhao Lu ⋅ Huichao Qi ⋅ Alexandra Woolgar ⋅ Xi V Wang ⋅ Lihui Wang
Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural representations and computational models of vision. A field-wide goal is to achieve accurate, generalizable decoding from non-invasive, temporally resolved signals, including electroencephalography (EEG). A major obstacle towards this goal is the low signal-to-noise ratio of EEG and the substantial inter-subject variability, which render direct end-to-end EEG–image supervision weak and unstable. To address this, we introduce a tri-modal contrastive framework for EEG-based visual decoding that aligns EEG, visual, and textual representations within a unified latent space. Our approach follows a two-stage design. First, we pre-train an EEG encoder via masked reconstruction on unlabeled trials, learning spatio-temporal regularities that transfer robustly downstream. Second, we jointly align EEG, image, and LLM-generated textual descriptions through contrastive learning, where text supervision acts as a semantic regularizer that injects linguistic structure into the shared space without overwhelming the primary EEG–image signal. The encoder integrates subject-specific adaptation, graph-attention over channels, and temporal-spatial convolutional embeddings. On the Things-EEG2 200-way zero-shot benchmark, our framework achieves 54.1\% Top-1 and 83.4\% Top-5 accuracy, substantially exceeding the strongest prior baseline (32.4\% / 64.0\%), with paired Wilcoxon tests confirming significance (p < 0.01) over all in-subject baselines. We validate generalization on Things-MEG. Analysis reveals that compact embedding geometries (CN-CLIP) outperform much larger backbones, and that decoding aligns with established neurophysiology of visual processing. This work is a critical step towards robust, semantically-grounded visual decoding from non-invasive temporal neural signals.
MinPath: Learning Efficient LLM Reasoning via Minimal Dependency Paths
Zhijing Yang ⋅ Shuming Hou ⋅ Guohui Xiao ⋅ Lemei Zhang ⋅ Peng Liu
Large language models increasingly rely on long Chain-of-Thought (CoT) to solve complex reasoning tasks, but the resulting CoT often contains repeated checks, abandoned branches, and unused planning that increase inference cost without supporting the final answer. Existing efficient reasoning methods address this through length budgets, token-level importance, confidence signals, or a stronger teacher model, treating CoT compression mainly as a text-shortening problem rather than as a question of which steps the answer actually depends on. We propose MinPath, a training framework that treats CoT compression as a graph problem. The raw CoT is represented as a Typed Reasoning Graph, where the Minimal Dependency Path supporting the target conclusion is extracted, and the resulting MinPath is used for fine-tuning. Across four reasoning models on GSM8K, MATH-500, and AIME24, MinPath reduces generated tokens by 33.9\% on average with nearly unchanged precision (accuracy even improved in 41.7\% of all experiments) and outperforms five compression baselines in every evaluated setting.
MIRAGE: Mobile Agents with Implicit Reasoning and Generative World Models
Zhichao Yang ⋅ Yuanze Hu ⋅ Haojie Hao ⋅ Longkun Hao ⋅ Dongshuo Huang ⋅ Hongyu Lin ⋅ Li ⋅ Lanqing HONG ⋅ Yihang Lou ⋅ Yan Bai
Mobile agents are increasingly expected to operate everyday applications from screenshots and language goals, where reliable control requires reasoning over screen affordances, multi-step navigation, and future state changes. Yet many agents externalize this computation as long textual thoughts, making interaction slower, supervision more costly, and deployment less efficient. We introduce MIRAGE (Mobile agents with Implicit Reasoning And Generative world modEls), a framework that learns continuous latent reasoning representations from visible textual thoughts. MIRAGE introduces an efficient latent-space learning procedure that transfers explicit reasoning into compact hidden states, allowing the agent to reason internally without decoding long rationales. It further brings a world-model perspective into mobile-agent training: the model’s latent reasoning vectors are aligned with future screenshots, encouraging the agent to predict upcoming interface states in latent space before executing an action. This makes the hidden computation not only a compressed thought trace, but also a forward-looking representation of how the environment may change. At inference time, MIRAGE reasons in continuous latent space, reducing token generation while improving execution efficiency. On AndroidWorld, MIRAGE matches explicit-CoT SFT in the 4B ablation under a 3–5× lower decoded-token budget, and improves a comparable instruction-tuned baseline by 10.2 points; on AndroidControl, it improves action grounding with over 75% fewer generated tokens.
MIRA: Mutual Information guided calibration for Reliable Test-Time Adaptation
Hyeongyu Kim ⋅ Mingyeong Jang ⋅ Youngjun Song ⋅ GeonHui Han ⋅ Dosik Hwang
Test-time adaptation updates a deployed model using only unlabeled target data, but standard entropy-based methods often rely on deterministic confidence and can reinforce overconfident errors under distribution shift. We propose MIRA, Mutual Information guided calibration for Reliable Test-Time Adaptation, a source-free online adaptation method that uses lightweight last-layer Gaussian perturbations to probe prediction reliability. For each target batch, MIRA reuses a single feature extraction pass and samples perturbed classifiers around the source classifier, forming an efficient local stochastic ensemble. Rather than fitting a Bayesian posterior, MIRA treats the perturbation scale as a local probing radius and selects a batch-specific scale using mutual information and prediction agreement. The resulting perturbation-sensitivity signal identifies predictions that are confident and locally stable under classifier perturbations, while suppressing brittle predictions with high stochastic disagreement. MIRA then performs adaptation with an MI-confidence weighted entropy loss, which down-weights perturbation-sensitive or low-confidence samples before they drive entropy minimization. This weighted loss reduces the influence of brittle predictions that can otherwise cause harmful adaptation under distribution shift. The method requires no source data, retraining, model ensembles, or offline posterior fitting, and can be applied as a plug-and-play modification to entropy-minimization TTA. Experiments on distribution-shift benchmarks show improvements in accuracy, calibration, and negative log-likelihood, suggesting that mutual-information-guided perturbation probing and weighted adaptation yield more robust online TTA than deterministic confidence.
Large language models (LLMs) can memorize and reproduce training sequences verbatim, which undermines both generalization and privacy. Existing mitigation methods apply interventions uniformly, often degrading performance on the majority of tokens that generalize normally. We empirically find that memorization is sparse and heavy-tailed at the token level. This structure is mismatched to static, sequence-level interventions: effective mitigation must localize to where memorization actually happens. From the ideal token-level memorization reduction objective, we derive a probe-steer framework, which decomposes intervention into a probe that detects memorization-relevant activations and a steer that applies targeted correction only when the probe exceeds a threshold. We propose Gated Subspace Steering (GSS) as its practical instantiation: the optimal probe-steer pair admits a closed-form solution as the leading singular vectors of a gradient-weighted activation matrix. Across four benchmarks, GSS matches or exceeds state-of-the-art memorization reduction while requiring 10--100$\times$ less compute than optimization-based alternatives.
MixNLQ: An Effective Post-Training Nonlinear Low-Bit Quantization Method for Large Language Models
Yirong Xiong ⋅ Mingrui Cai ⋅ Gene Wen ⋅ Yuxing Han
Post-Training Quantization (PTQ) efficiently compresses Large Language Models (LLMs) without expensive retraining. However, most existing PTQ methods rely on linear quantization, which fails to fully capture the underlying characteristics of parameter distributions. Conversely, non-linear approaches, such as codebook-based methods, introduce significant complexity that hinders practical engineering deployment. To address this, we propose MixNLQ, an easy-to-use hybrid non-linear PTQ method. By combining linear and exponential terms, MixNLQ flexibly fits the "high-peak, heavy-tail" distributions typical of LLM weights and activations. It supports both calibration-free and calibration-based settings, and uses arithmetic-only dequantization to avoid memory-intensive lookup tables—ensuring highly efficient quantization and inference. Additionally, MixNLQ serves as a plug-and-play module capable of enhancing various existing PTQ frameworks. Evaluations at 3-bit and 4-bit on LLaMA, Mistral, and DeepSeek demonstrate that MixNLQ achieves competitive or superior perplexity and zero-shot accuracy compared to existing PTQ baselines, alongside promising activation quantization performance and its versatility as a plug-and-play enhancement module. Ultimately, MixNLQ offers a highly practical solution that effectively balances accuracy, algorithmic simplicity, and real-world deployment efficiency.
MLLM-Edit: Benchmarking Image Forgery Detection and Localization under MLLM-based Editing
Zeqin Yu ⋅ Ye Tian ⋅ Jian Zhang ⋅ Jiangqun Ni ⋅ Lanyun Zhu ⋅ Alex Kot ⋅ Xudong Jiang
Multimodal large language models (MLLMs) enable prompt-driven image editing, introducing a new challenge for image forgery detection and localization (IFDL). Unlike traditional or mask-guided editing, MLLM-based editing may re-synthesize both tampered and non-tampered regions, weakening the correspondence between tampering regions and pixel-level forensic discrepancies. To study this emerging paradigm, we introduce MLLM-Edit, a large-scale benchmark containing approximately 250k tampered images with pixel-level tampering annotations. MLLM-Edit is constructed through a structured pipeline that combines diverse real-image collection, human-like prompt design, multi-source MLLM-based editing, and semi-automatic mask annotation. Specifically, we collect natural and text images from 30 sources, formulate prompts with target content, tampering type, and consistency constraints, generate forgeries using 18 representative open-source and closed-source MLLM-based editors, and obtain pixel-level tampering masks through a semi-automatic annotation pipeline. We further establish four evaluation protocols covering traditional editing, mask-guided editing, fully synthetic generation, and MLLM-based editing. Experiments show that existing IFDL methods degrade substantially on MLLM-based forgeries, especially in pixel-level localization, highlighting the need for dedicated benchmarks and methods. A demo subset of MLLM-Edit is available at the \href{https://kaggle.com/datasets/8441a4e0594aa02aa7a59ee3eead454410d34716d14423d56a56c7bbbac80655}{project page}.
MMGraph-Agent: Agentic Multimodal RAG via Cache-Inspired Multimodal Knowledge HyperGraphs
Haoran Luo ⋅ Ziyue Zhu ⋅ Shangyang Wu ⋅ LINHAO LUO ⋅ Yikai Guo ⋅ Qika Lin ⋅ Fangzhi Xu ⋅ Jiapu Wang ⋅ Xiaobao Wu ⋅ Yifan Zhu ⋅ Anh Tuan Luu
Driven by increasingly complex real-world applications, Retrieval-Augmented Generation (RAG) has evolved from textual pipelines to multimodal settings, comprising two complementary dimensions: knowledge organization and retrieval enhancement. However, existing methods still face challenges in high construction cost, context window limitations, and the lack of unified training across offline and online knowledge sources. To address these challenges, we propose MMGraph-Agent, a unified multimodal RAG framework based on cache-inspired multimodal knowledge hypergraphs. MMGraph-Agent enables extremely lightweight hypergraph construction with zero API overhead and dynamic memory-based retrieval that alleviates long-context issues. Experiments across offline, online, and joint retrieval settings demonstrate consistent improvements in performance, efficiency, and architectural unification. Our software and data are publicly available.
Mobility Helps Learning: Unsupervised Model Adaptation for Object Recognition via Movement
Yanan Ma ⋅ Yihang Tao ⋅ Zhengru Fang ⋅ Senkang Hu ⋅ Xianhao Chen ⋅ Yuguang Fang
Movable agents, such as autonomous vehicles and robots, should be able to autonomously adapt their pre-trained object recognition models to new environments without human annotations. While existing test-time adaptation (TTA) methods address label-free model adaptation, they typically treat test data as independent snapshots, without constructively exploiting a rich, unique source of supervision: the significant variance in prediction quality as an agent observes the same object from different distances and viewpoints. To capitalize on this, we propose MoCaFe, a model-agnostic and hyperparameter-insensitive paradigm that leverages motion-induced prediction variance for unsupervised adaptation. Specifically, we first develop mobility-calibrated filtering of pseudo-labels, which constructs reliable pseudo-labels by fusing cross-view predictions using inverse-variance weights. The filtered pseudo-labels then guide the selection of "hard samples" for effective learning. To further suppress pseudo-label corruption, we introduce p-MoCaFe, an optional extension that uses a public dataset to assign weighting factors to data samples, yielding an estimator of the clean loss with bounded bias. Extensive experiments on autonomous driving (nuScenes, KITTI) and embodied-agent datasets show that MoCaFe significantly outperforms state-of-the-art TTA schemes, achieving up to 20\% accuracy gain.
MoCAR: Motion-code Coordinate-aware AutoRegression for Continuous Trajectory Forecasting
Yiming Xu ⋅ Hao Cheng ⋅ Monika Sester
Autoregressive generation is natural for language, where predicted tokens can be directly reused as the next prediction state, but trajectory forecasting lacks such a clean token: motion is continuous, multimodal, and expressed in local coordinate frames that evolve with the predicted trajectory. We present $\textbf{MoCAR}$ ($\textbf{Mo}$tion-code $\textbf{C}$oordinate-aware $\textbf{A}$uto$\textbf{R}$egression), a decoder-only framework that casts trajectory forecasting as next-code prediction in a coordinate-aware continuous latent space. MoCAR learns a continuous motion-code space for endpoint-normalized trajectory segments, where each code jointly captures local trajectory geometry and the reference-frame transition induced by that segment. Historical motion codes are used as a teacher-forced prefix, future codes are generated autoregressively under temporal, map, agent, and mode interactions, and predicted codes persist in latent memory while decoded endpoints update the local scene context. This enables rollout without trajectory-space re-tokenization, trajectory queries, goal candidates, or proposal-and-refinement pipelines. On Argoverse (AV) benchmarks, MoCAR achieves top-tier performance with a simple single-stage architecture, transfers strongly from AV2 to AV1 in zero-shot evaluation, and improves on turn-heavy scenarios. Ablations confirm that the learned continuous motion-code space, latent alignment, weak KL regularization, and joint tokenizer-predictor optimization are essential for stable latent autoregression.
Modelling Opinion Dynamics at Scale with Deep MARL
Lukas Seier ⋅ Brandon Kaplowitz ⋅ Sebastian Towers ⋅ Richard M Bailey ⋅ Jakob Foerster
Modelling opinion dynamics typically relies on hand-crafted local interaction rules to study emergent macroscopic phenomena such as consensus and polarisation. In contrast, multi-agent reinforcement learning (MARL) enables agents to learn such behaviours directly by optimising simple rewards. To explore the potential of MARL for opinion dynamics, we introduce a GPU-accelerated consensus and truth-finding game that scales to populations of up to 1000 agents, comparable to many real-world social sub-networks. To prevent unrealistic conventions, we extend other-play to general-sum social interactions. We next validate our model on a subset of the Bluesky network by recovering agent importance structures from graph topology alone via a learned attention layer, finding that highly conforming populations most closely match human data. In large social media networks such high levels of conforming significantly reduce collective accuracy and promote dishonest agents that lie to fit in. By contrast, small, dynamic hunter-gatherer networks are less affected; here, conformity can even improve collective agreement. This suggests a mismatch between evolved human conformity heuristics and modern social media environments as a potential contributor to misinformation.
Moment-Constrained Latent Steering for Flow Policy
Wenxin Zhao ⋅ Letian Tao ⋅ Zhilong Zheng ⋅ Guojian Zhan ⋅ Yujie Yang ⋅ Likun Wang ⋅ yinuo Wang ⋅ Feihong Zhang ⋅ Tianze Zhu ⋅ Jingliang Duan ⋅ Yang Guan ⋅ Shengbo Eben Li
Fine-tuning continuous-time generative models, such as flow models, via reinforcement learning (RL) is a promising approach for complex continuous control, but directly updating the network weights often leads to catastrophic forgetting of the pre-trained prior. Latent steering method mitigates this by freezing the generative backbone and optimizing a latent policy to generate the initial noise. However, existing methods typically enforce hard support boundaries to prevent out-of-distribution (OOD) shifts in the latent space, inducing a fundamental trade-off between prior preservation and policy expressiveness. To resolve this dilemma, we propose Moment Constraint Parameterization, which regulates the latent distribution through a statistical budget. By decoupling this budget into inter-statistic constraints and intra-statistic flexibility, our method acts as a probabilistic safeguard against OOD shifts while providing the degrees of freedom to retain policy expressiveness. Furthermore, while moment constraints define a flexible search space, effectively optimizing the latent policy within it requires accurate gradient signals. To address the optimization lag and bias caused by proxy latent critics in existing methods, we introduce Adjoint Q-Gradient Propagation. By exploiting the white-box structure of flow models, this technique backpropagates exact analytical gradients from the action-space critic directly to the initial noise. Integrating these two mechanisms, we present Moment Constraint Flow Steering (MCFS). Empirical results on D4RL and OGBench show that MCFS consistently delivers stronger online adaptation performance than prior latent steering and flow-policy baselines, with particularly pronounced gains on challenging long-horizon compositional manipulation tasks.
MoRe: Modular Representations for Principled Continual Representation Learning on Sequential Data
Jiaqi Sun ⋅ Boyang Sun ⋅ Rasmy M. H. ⋅ Xiangchen Song ⋅ Kun Zhang
Continual learning requires models to adapt to new data while preserving previously acquired knowledge. At its core, this challenge can be viewed as principled one-step adaptation: incorporating new information with minimal interference to existing representations. Most existing approaches address this challenge by modifying model parameters or architectures in a supervised, task-specific manner. However, the underlying issue is representational: tasks require distinct yet structured representations that can be selectively updated without disrupting representations, while structure should reflect intrinsic organization in the data rather than task boundaries. In sequential data, time-delayed dependencies provide a natural signal for uncovering this organization, revealing how fundamental representations give rise to more specific ones. Inspired by the modular organization of the human brain, we propose MoRe, a framework that identifies modularity in the representation itself rather than allocating it at the architectural level. MoRe decomposes knowledge into a hierarchy of fundamental and specific modules with identifiability guarantees, enabling principled module reuse, alignment, and expansion during adaptation while preserving old modules by construction. Experiments on synthetic benchmarks and real-world LLM activations demonstrate interpretable hierarchical structure, improved plasticity-stability trade-offs, suggesting MoRe as a principled foundation for continual adaptation.
MorphoHELM: A Comprehensive Benchmark for Evaluating Representations for Microscopy-Based Morphology Assays
Emre Hayir ⋅ Lorin Crawford ⋅ Alex X Lu
Microscopy images contain rich information about how cells respond to perturbations, making them essential to applications like drug screening. To quantify images, researchers often use representation extraction methods, and recent years have seen a proliferation of deep learning methods. While measuring the quality of these representations is essential, evaluation remains fragmented, with each proposed model evaluated on different tasks and datasets, using custom pipelines and metrics, making it difficult to fairly compare models. Here, we introduce MorphoHELM, a comprehensive open benchmark for evaluating feature extraction methods for cell painting, the most widely-used morphological profiling assay. MorphoHELM consolidates evaluation standards in the field, extends and corrects them to be more robust, and evaluates on the widest range of methods to date. A defining feature of the benchmark is that each task is evaluated at different degrees of batch effects (or technical noise), directly quantifying how the ability of methods to detect biological signal degrades as noise increases. Together, these properties enable MorphoHELM to detect trade-offs between methods, and we demonstrate that models that excel at certain kinds of biological signal are weaker at others. We show that no existing model outperforms classic computer vision analytic strategies across all settings, which remain the strongest general use-case representations. All datasets, code, and evaluation tools are publicly available.
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
Jonathan Steinberg ⋅ Oren Gal
Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The challenge is structural: existing safety alignment evaluates overt requests in isolation, leaving models blind to malicious end-states that emerge from sequenced compliance with innocuous-looking requests. We introduce MOSAIC-Bench (Malicious Objectives Sequenced As Innocuous Compliance), a benchmark of 199 three-stage attack chains paired with deterministic exploit oracles on deployed software substrates (10 web-application substrates, 31 CWE classes, 5 programming languages) that treats both exploit ground truth and downstream reviewer protocol as first-class evaluation axes. On this benchmark, nine production coding agents from Anthropic, OpenAI, Google, Moonshot, Zhipu, and Minimax compose innocuous tickets at 53–86% end-to-end ASR with only two refusals across all staged runs. In a matched direct-prompt experiment over four frontier Claude/Codex agents, vulnerable-output rates fall to 0–20.4%: Claude primarily refuses, while Codex primarily hardens rather than emitting the vulnerable implementation — ticket staging silences both defense modes simultaneously. Downstream, code reviewer agents approve 25.8% of these confirmed-vulnerable cumulative diffs as routine PRs, and a full-context implementation protocol closes only ∼50% of the staged/direct gap, ruling out context fragmentation as the sole explanation. As a deployable but non-adaptive mitigation, reframing the reviewer as an adversarial pentester reduces evasion across the evaluated reviewer subset; pentester-framed evasion ranges from 3.0% to 17.6%, and an open-weight Gemma-4-E4B-it reviewer under this framing detects 88.4% of attacks on the dataset with a 4.6% false-positive rate measured on 608 real-world GitHub PRs. We publicly release our dataset at https://huggingface.co/datasets/MosaicBenchmark/mosaic-bench, and a verifiable and adaptable evaluation framework at https://github.com/mosaic-benchmark/mosaic-benchmark.
Motion Forcing: Decoupling Ego and Object Motion via Sparse Inputs for Structured Video Generation
Tianshuo Xu ⋅ ZhiFei Chen ⋅ Leyi Wu ⋅ Hao LU ⋅ Yingcong Chen
Structured video generation (e.g., autonomous driving) faces a fundamental domain gap between simple human control and dense physical reality. Text-to-Video models accept accessible inputs but struggle with precise spatial guidance, while dedicated autonomous driving world models require expensive, dense 3D annotations that are difficult for users to provide. To bridge this gap, we introduce \textbf{Motion Forcing}, a structured video generation framework that achieves precise multi-agent control using extremely sparse 2D inputs. Our key insight is to explicitly bridge human intent and visual synthesis via a hierarchical \textbf{``Point-Shape-Appearance''} paradigm. This approach decomposes generation into verifiable stages: modeling user intent as sparse geometric points, expanding them into dynamic depth maps to explicitly resolve 3D geometry and decouple ego-motion from object dynamics, and finally rendering high-fidelity textures. Furthermore, to elevate the model from passive instruction-following to active physical reasoning, we employ a \textbf{Masked Point Recovery} strategy. By forcing the reconstruction of dynamic depth from occluded trajectories, the model internalizes latent physical laws, enabling causal inference. Extensive experiments demonstrate that Motion Forcing significantly outperforms state-of-the-art baselines on large-scale autonomous driving benchmarks. Additional validation in rigid-body physics and robotic manipulation confirms the structural integrity and generalizability of our framework. \textbf{To facilitate future research, we have released our model weights and code.}
MoTo: Mixture of Tokenizers Towards Fair Multilingual Language Modeling
Gül Sena Altıntaş ⋅ Colin Raffel
Large language models rely on a single tokenizer that is chosen at training time and fixed thereafter. Most tokenizers are trained on English-dominant corpora, making them under-serve morphologically rich and non-Latin-script languages, where they can produce fragmented representations, longer sequence lengths, and reduced information density during training. We propose Mixture of Tokenizers (MoTo), a modular tokenization framework that trains dedicated per-language BPE tokenizers and composes them into a unified superset vocabulary, where tokens shared across languages are deduplicated into a common ID space. Each language tokenizer maintains its own normalization policy, pre-tokenization rules, and vocabulary budget. Despite allocating significantly fewer entries to English, MoTo retains English performance while delivering consistent gains across 21 typologically diverse languages spanning six scripts, with improvements most pronounced on morphologically complex and non-Latin-script languages that are systematically underserved by English-dominant tokenizer design. Our results suggest that modular subword tokenization is a practical and extensible alternative to monolithic tokenization.
Multigroup Fairness and Omniprediction: Separations and Equivalences
Sílvia Casacuberta ⋅ Parikshit Gopalan ⋅ Varun Kanade ⋅ Omer Reingold ⋅ Konstantinos Stavropoulos ⋅ Pranay Tankala
Omniprediction [GKR+22] is a powerful learning guarantee that requires a predictor to be competitive with the best hypothesis from a benchmark class, not just for a single loss function, but simultaneously across all loss functions in a prespecified family. Currently, learning algorithms for achieving omniprediction rely on notions of multigroup fairness, such as multiaccuracy and multicalibration [HKRR18] or calibrated multiaccuracy [GHK+23]. In this work, we ask whether this reliance is necessary: Does omniprediction require some form of multigroup fairness? We answer this question in the negative. While prior works have shown that various multigroup fairness notions imply omniprediction [GKR+22, GHK+23, OKK25], we rule out even a weak converse. Specifically, we show that omniprediction for proper losses does not even require accuracy in expectation, a much weaker notion than calibration or multiaccuracy. We complement our negative answer for omniprediction with an affirmative answer for loss outcome indistinguishability (loss OI) [GHK+23], a related but stronger learning guarantee. Loss OI, which implies omniprediction, requires the predicted label distribution to be indistinguishable from the true label distribution by a certain class of tests depending on the loss functions and the benchmark class. Prior work showed how to achieve loss OI from a combination of calibration and multiaccuracy [GHK+23]. We prove the converse, establishing that loss OI is in fact equivalent to a form of calibrated multiaccuracy.
Learning unified representations from single-cell multi-omics data is a fundamental challenge, yet existing deep learning models often overlook the inherent hierarchical structure of biological regulation. In this work, we introduce scWavelet, a novel deep learning framework that leverages a multi-scale inductive bias by learning representations in the cellular spatial-frequency domain. Our approach is grounded in a heuristic approximation-misspecification decomposition, which theoretically motivates the design of a scale-factorized architecture. scWavelet employs a learnable wavelet transform to decompose omics signals into distinct scales, which are then processed by scale-specific Omics Mixers to generate integrated latent representations. Extensive experiments demonstrate that scWavelet achieves state-of-the-art performance on challenging multi-omics integration and cross-omics translation benchmarks. Furthermore, interpretability analyses confirm that the learned multi-scale features effectively correspond to the biological hierarchy of cellular identity. Our project will be publicly available.
MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation
Ke Wang ⋅ Houxing Ren ⋅ Zimu Lu ⋅ Yunqiao Yang ⋅ ZHUOFAN ZONG ⋅ Mingjie Zhan ⋅ Hongsheng Li
End-to-end full-duplex speech models have brought open-source machine conversation close to human fluency, yet existing systems fall short of real-world deployment in two entangled respects: long-context conversational robustness and multi-party interaction capability. Realistic settings, including meetings, group lessons, family dinners, and social-robot reception, are inherently long-horizon and multi-party at the same time, requiring a single model to perceive, attribute, contextualize, and respond across multiple speakers over extended durations. Progress along these axes is bottlenecked by both data and evaluation. On the data side, open multi-party conversational speech corpora total only a few hundred hours and are not designed for codec-frame-level full-duplex modeling. On the evaluation side, existing long-audio benchmarks focus on passive listening, while existing speech-to-speech benchmarks remain dyadic and short. In this work, we extend the Moshi paradigm along long-horizon and multi-party axes simultaneously, in both English and Chinese, with three contributions. First, we release an open data engine and a 57.6 k-hour corpus for long, multi-party, bilingual full-duplex dialogue. The engine produces parallel-stream audio with controllable length, participant count, conversational dynamics (turn-taking, overlap, backchannels, interruption, addressee shifts, long-range co-reference), and language (English, Chinese, and intra-sentential code-switching), exceeding all prior open multi-party conversational speech corpora by more than an order of magnitude. Second, we introduce MultiTalkBench, the first benchmark to jointly evaluate long-form full-duplex dialogue with an average duration of 32.6 minutes, multi-party, and bilingual full-duplex dialogue, with explicit probes for long-range entity tracking, topic coherence, and addressee selection. Third, we train a bilingual Moshi-style full-duplex model on the released corpus that sustains coherent multi-party English-Chinese conversation over extended durations, substantially outperforming open-source baselines including Moshi, MiniCPM-o-4.5, and Qwen3-Omni-30B-A3B-Instruct.
Native Audio-Visual Alignment for Generation
Longbin Ji ⋅ Guan Wang ⋅ xuan wei ⋅ Chenye Yang ⋅ Zhenyu Zhang ⋅ Shuohuan Wang ⋅ YU SUN ⋅ Jingzhou HE
Joint audio-video generation aims to synthesize temporally synchronized and semantically coherent visual and acoustic content. Existing open-source methods mainly follow either dual-tower designs, which generate audio and video in separate streams and rely on posterior alignment, or fully unified tri-modal designs, which mix textual context, audio, and video in a single shared space. These paradigms may either weaken fine-grained audio-video co-evolution or couple semantic conditioning with low-level synchronization. We propose NAVA, a Native Audio-Visual Alignment framework that formulates generation as context-conditioned native audio-visual alignment. NAVA first establishes audio-video correspondence in a dedicated alignment space and then applies context as external conditioning to guide the aligned representation. We instantiate this formulation with an Align-then-Fuse MMDiT architecture, which progressively bridges modality-aware alignment and unified audio-video denoising. To support controllable speech generation, we further introduce Timbre-in-Context Conditioning, which binds reference timbre cues to corresponding speech spans through the context pathway. Experiments on Verse-Bench and the Seed-TTS benchmark demonstrate that NAVA achieves superior audio-visual synchronization and video quality, competitive audio quality, and substantially improved reference-timbre controllability with only 6.2B parameters.
NavAble: A Large-Scale Dataset and Synthetic Data Generation Pipeline for Blind Navigation
Hochul Hwang ⋅ Jahir Sadik Monon ⋅ Soowan Yang ⋅ Shiven U Patel ⋅ Anh N Nguyen ⋅ Kien Nguyen ⋅ Keshav Garg ⋅ Dylan Gage ⋅ Duretti Hordofaa ⋅ Anshu Anjna ⋅ Eshed Ohn-Bar ⋅ Donghyun Kim
Reliable recognition of accessibility-critical objects (e.g., audible pedestrian signals, door-activation buttons, and handrails) is essential for assistive navigation technologies that support safe, independent mobility for blind and low-vision (BLV) users. Yet these categories are severely underrepresented in existing large-scale vision datasets, and the few datasets that include them suffer from limited class coverage, inconsistent annotations, and poor diversity, limiting the perception reliability that BLV navigation requires. We introduce NavAble, a large-scale dataset spanning 11 accessibility-critical object classes, combining 41K curated open-source images with 8K newly collected, densely annotated real-world images. To extend scale and diversity, we develop a synthetic data generation pipeline that renders high-fidelity 3D assets across 37 environments under diverse styles and viewpoints, yielding 452K frames with rich ground-truth annotations from 565 assets. The pipeline minimizes manual effort by combining automated web image crawling, VLM-based filtering, SAM 3D-based asset generation, and convenient camera trajectory configuration, enabling large-scale data generation at a diversity unattainable through manual collection. Experiments show that fine-tuning existing segmentation models on the NavAble Dataset achieves reliable accessibility-object segmentation (>80\% mIoU on our held-out test set). We also demonstrate the effectiveness of synthetic data augmentation, which yields up to +3.5 mIoU over real-world-only fine-tuning. Beyond benchmark evaluation, validation on egocentric data from two mobility-assistive robots demonstrates improved perception reliability in real-world navigation scenarios. Together, NavAble establishes a scalable foundation for BLV navigation perception, directly addressing the data scarcity and embodiment gaps that have limited progress in this domain. The dataset and pipeline are publicly available at https://huggingface.co/datasets/NavAble/NeurIPS2026BLV.
Orthogonal trace-sum maximization (OTSM) problems arise in a wide range of data processing applications, including canonical correlation analysis and cryogenic electron microscopy. Despite their practical importance, the existing theoretical understanding of these problems remains incomplete. In this paper, we show that the generalized power method (GPM) converges linearly to the global optimal solution with high probability under an additive Gaussian noise model, provided that the noise level satisfies the nearly optimal bound $(O(\sqrt{n/\log n}))$. In addition, we prove that the semidefinite programming (SDP) relaxation of OTSM is tight and admits a unique optimal solution under the same noise regime, improving upon the best known theoretical guarantee of $(O(n^{1/4}))$. Extensive numerical experiments further demonstrate that our theoretical predictions closely match empirical observations.
Near-Optimal Best-of-Both-Worlds Algorithms for Decoupled Exploration and Exploitation in Multi-armed Bandits
Hibiki Sekiya ⋅ Shinji Ito
We study the decoupled multi-armed bandit problem, where the player selects one arm to exploit (incurring its loss) and one arm to explore (observing its loss) at each round. We propose two algorithms: D-Exp3-Ada, based on Exp3 with an AdaGrad learning rate, and D-TINF-SPM, based on FTRL with Tsallis entropy and the stability-penalty matching learning rate. Here, $K$ denotes the number of arms, $T$ the time horizon, and $\Delta_i$ the suboptimality gap of arm $i$. D-Exp3-Ada achieves $\mathcal{O}(\sqrt{KT\log K})$ regret in adversarial regimes and $\mathcal{O}(\sum_{i \neq i^{\star}} \frac{\log K}{\Delta_i})$ in stochastic regimes, combining simplicity of algorithm design with strong guarantees. D-TINF-SPM achieves the minimax optimal $\mathcal{O}(\sqrt{KT})$ in adversarial regimes and $\mathcal{O}(\min\{\sqrt{K\sum_{i \neq i^{\star}} \frac{1}{\Delta_{i}^2}}, \sum_{i \neq i^{\star}} \frac{\log K}{\Delta_i}\})$ in stochastic regimes. We also provide a refined analysis of the prior best algorithm. Finally, we derive instance-dependent lower bounds under natural monotonicity and permutation-invariance assumptions on regret upper bounds, proving that our algorithms are optimal up to $\mathcal{O}((\log K)^2)$ within this class.
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
Boxin Wang ⋅ Chankyu Lee ⋅ Nayeon Lee ⋅ Sheng-Chieh Lin ⋅ Wenliang Dai ⋅ Yang Chen ⋅ Yangyi Chen ⋅ Zhuolin Yang ⋅ Zihan Liu ⋅ Mohammad Shoeybi ⋅ Bryan Catanzaro ⋅ Wei Ping
Building general-purpose reasoning models with reinforcement learning (RL) entails substantial cross-domain heterogeneity, including large variation in inference-time response lengths and verification latency. Such variability complicates the RL infrastructure, slows training, and makes training curriculum (e.g., response length extension) and hyperparameter selection challenging. In this work, we propose cascaded domain-wise reinforcement learning (Cascade RL) to develop Nemotron-Cascade, capable of operating in both \emph{instruct} and deep \emph{thinking} modes, without any performance gap relative to a thinking-only counterpart. Departing from conventional approaches that blend heterogeneous prompts from different domains, Cascade RL orchestrates sequential, domain-wise RL, reducing engineering complexity and delivering state-of-the-art performance across a wide range of benchmarks. Notably, RLHF for alignment, when used as a pre-step, boosts the model's reasoning ability far beyond mere preference optimization, and subsequent domain-wise RLVR stages rarely degrade the benchmark performance attained in earlier domains and may even improve it. Our 14B model, after RL, outperforms its SFT teacher, DeepSeek-R1-0528, on LiveCodeBench v5/v6/Pro and achieves silver-medal performance in the 2025 International Olympiad in Informatics (IOI). We transparently share our training and data recipes.
Nereus: A Large-Scale Underwater Dataset for Fine-Grained Attribute Understanding and Grounded Counting Perception
Jingtao Deng ⋅ Zihao Jin ⋅ Yuehe Chen ⋅ Feixiang He ⋅ Shuiwang Li ⋅ Dan Zeng
Underwater imagery records rich ecological information and supports applications such as science communication, fishery monitoring, ecological conservation, and oceanographic research. Beyond general scene understanding, these applications require models to recognize fine-grained biological traits and reason over spatially grounded evidence. However, existing underwater datasets often provide coarse category labels, generic captions, or total-count supervision, which limits their use for organism-level analysis and ecological reasoning. We introduce Nereus, a large-scale dataset for fine-grained underwater perception that jointly supports Fine-Grained Object Attribute Understanding and Grounded Counting Perception. Nereus contains 88K images and 2.68M question-answer pairs, including annotations for 136K marine organism objects across 2,690 species and 1,515 genera, as well as 65K point annotations and 67K grounded counting QA pairs. Experiments with representative multimodal models show that Nereus improves downstream adaptation over existing underwater-domain data and provides a foundation for future research on fine-grained marine perception.
Network of Theseus (Like the ship)
Vighnesh Subramaniam ⋅ Colin Conwell ⋅ Boris Katz ⋅ Andrei Barbu ⋅ Brian Cheung
A standard assumption in deep learning is that the inductive bias introduced by a neural network architecture must persist from training through inference. The architecture you train with is the architecture you deploy. This assumption constrains the community from selecting architectures that may have desirable efficiency or design properties due to difficulties with optimization. We challenge this assumption with Network of Theseus (NoT), a method for progressively converting a trained, or even untrained, guide network architecture part-by-part into an entirely different target network architecture while preserving the performance of the guide network. At each stage, components in the guide network architecture are incrementally replaced with target architecture modules and aligned via representational similarity metrics. This procedure largely preserves the functionality of the guide network even under substantial architectural changes—for example, converting a convolutional network into a multilayer perceptron, or GPT-2 into a recurrent neural network. By decoupling optimization from deployment, NoT expands the space of viable inference-time architectures, opening opportunities for better accuracy–efficiency tradeoffs and enabling more directed exploration of the architectural design space.
Neural Causal Models under Markov Equivalence
Yushu Pan ⋅ Hongshuo Yang ⋅ Adiba Ejaz ⋅ Elias Bareinboim
Neural causal models can simulate interventions in complex, high-dimensional settings, but typically require a known causal graph. Observational data, however, generally identifies only a Markov equivalence class of graphs, represented by a CPDAG, and causal queries may vary across DAGs in that class. We introduce \textit{Masked Neural Causal Models }(NCMs), a provably expressive nonparametric framework for simulating and bounding interventional queries over all models compatible with a given CPDAG given observational data. We give an optimization objective in the space of Masked NCMs that asymptotically recovers the true bounds of the causal effect. To enable this optimization in practice, we introduce an attention-based architecture and a novel optimization strategy that recovers highly accurate bounds in discrete nonparametric settings.
Neural Garbage Collection: Learning to Forget while Learning to Reason
Michael Li ⋅ Jubayer Ibn Hamid ⋅ Emily Fox ⋅ Noah Goodman
Chain-of-thought reasoning has driven striking advances in language model capability, yet every reasoning step grows the KV cache, creating a bottleneck to scaling this paradigm further. Current approaches manage these constraints on the model’s behalf using hand-designed criteria. A more scalable approach would let end-to-end learning subsume this design choice entirely, following a broader pattern in deep learning. After all, if a model can learn to reason, why can't it learn to forget? We introduce \textbf{Neural Garbage Collection (NGC)}, in which a language model learns to forget while learning to reason, trained \emph{end-to-end} from outcome-based task reward alone. As the model reasons, it periodically pauses, decides which KV cache entries to evict, and continues to reason conditioned on the remaining cache. By treating tokens in a chain-of-thought and cache eviction decisions as discrete actions sampled from the language model, we can use reinforcement learning to jointly optimize how the model reasons and how it manages its own memory: what the model evicts shapes what it remembers, what it remembers shapes its reasoning, and the correctness of that reasoning determines its reward. Crucially, the model learns this behavior entirely from a \emph{single learning signal} — the outcome-based task reward — without supervised fine-tuning or proxy objectives. On Countdown, AMC, and AIME tasks, NGC maintains strong accuracy relative to the full-cache upper bound at a 2–3x compression in peak KV cache size and substantially outperforms eviction baselines. Our results are a first step towards a broader vision where end-to-end optimization drives both capability and efficiency in language models.
Neural Reconstruction of LiDAR Point Clouds under Jamming Attacks via Full-Waveform Representation and Simultaneous Laser Sensing
Ryo Yoshida ⋅ Takami Sato ⋅ Wenlun Zhang ⋅ Yuki Hayakawa ⋅ Shota Nagai ⋅ Kentaro Yoshioka
LiDAR sensors are critical for autonomous driving perception, yet remain vulnerable to spoofing attacks. Jamming attacks inject high-frequency laser pulses that completely blind LiDAR sensors by overwhelming authentic returns with malicious signals. We discover that while point clouds become randomized, the underlying full-waveform data retains distinguishable signatures between attack and legitimate signals. In this work, we propose PULSAR-Net, capable of reconstructing authentic point clouds under jamming attacks by leveraging previously underutilized intermediate full-waveform representations and simultaneous laser sensing in modern LiDAR systems. PULSAR-Net adopts a novel U-Net architecture with axial spatial attention mechanisms specifically designed to identify attack-induced signals from authentic object returns in the full-waveform representation. To address the lack of full-waveform representations in existing LiDAR datasets under jamming attacks, we introduce a physics-aware dataset generation pipeline that synthesizes realistic full-waveform representations under jamming attacks. Despite being trained exclusively on synthetic data, PULSAR-Net achieves reconstruction rates of 92\% and 73\% for vehicles obscured by jamming attacks in real-world static and driving scenarios, respectively.
Neural‑Visual Decoding via Cognitive‑guided Adaptive Blurring and Information‑Constrained Alignment
Fan Yin ⋅ Chuhang Zheng ⋅ Peiliang Gong ⋅ Donghai Guan ⋅ Qi Zhu
EEG-based visual decoding aims to establish a mapping between neural signals and visual semantics. However, it remains constrained by the dual challenges of severe information granularity mismatch and the low signal-to-noise ratio (SNR) of EEG signals. Existing approaches typically treat static visual features, ignoring the dynamic selectivity of human vision and the frequency specificity of neural oscillations. To bridge this gap, we propose \textbf{CAIA}, a Cognitive-guided Adaptive blurring with Information-Constrained Alignment framework for Neural-Visual decoding. On the visual side, it simulates selective attention to adaptively reduce redundancy. Meanwhile, on the EEG side, it leverages neural oscillation priors and the information bottleneck mechanism to enhance SNR. Specifically, we devise a cognitive-dynamics-based adaptive blurring mechanism that dynamically integrates center-biased and saliency-guided visual cues via cross-modal attention. Furthermore, we introduce a distribution-aware boundary calibration loss to robustly rectify alignment bias caused by outlier samples. Moreover, a cognitively-guided information-screening method is proposed to select task-relevant EEG oscillations. Extensive experiments demonstrate that CAIA improves both subject-dependent and subject-independent average Top-1 and Top-5 accuracy in zero-shot brain-to-image retrieval, significantly outperforming prior methods. Our work validates that optimizing visual information density to match neural granularity offers a more interpretable and robust pathway for neural decoding. Code is available at https://anonymous.4open.science/r/CAIA-0DA2.
NeuroAtlas: Benchmarking Foundation Models for Clinical EEG and Brain-Computer Interfaces
Konstantinos Kontras ⋅ Trui Osselaer ⋅ Stylianos G Mouslech ⋅ Angeliki I. Karaiskou ⋅ Guido Gagliardi ⋅ Thomas Strypsteen ⋅ Mohammad Hossein Badiei ⋅ Anku Rani ⋅ Maarten Vanmarcke ⋅ Miguel Bhagubai ⋅ Chanakya Ekbote ⋅ Jaedong Hwang ⋅ Christos Chatzichristos ⋅ Paul Liang ⋅ Maarten De Vos
Foundation models (FMs) promise to extract unified representations that generalize across downstream tasks. They have emerged across fields, including electroencephalography (EEG), but it is less clear how effective they are in this particular field. Published evaluations differ in datasets, in the EEG-specific preprocessing that might influence reported results, and in the reported metrics, frequently obscuring the clinical relevance in EEG. We introduce NeuroAtlas, the largest EEG benchmark to date: 42 datasets and $\sim$260k hours covering clinical EEG (epilepsy, sleep medicine, brain age estimation) and brain-computer interfaces, and include multiple datasets per task along with bespoke clinical evaluation metrics. Besides evaluating EEG-FMs with respect to supervised baselines, we present results from generic time-series FMs. We report three findings. First, EEG-specific FMs do not consistently outperform time-series FMs, which have neither EEG-focused architectures nor been pretrained on EEG. Second, standard machine learning metrics are insufficient to assess clinical utility: thus, we thoroughly evaluate more appropriate measures such as the quality of event-level decision-making, hypnogram-derived features, and the brain-age gap in the domains of epilepsy, sleep, and brain age, respectively. Third, model rankings and performance can vary substantially within domains. We conclude that pretrained models perform largely on par, with only narrow advantages for a few, and that current models do not yet deliver on the promise of an out-of-the-box unified EEG model. NeuroAtlas exposes this gap and provides the datasets and metrics for the next generation of unified EEG FMs.
NeuroHorizon: Long-Horizon Forward Prediction of Neural Population Activity via Autoregressive Decoding with Hierarchical Memory
Peng Jiang ⋅ Xiaoxuan Jia
Forward prediction of neural population activity is a prerequisite for closed-loop brain-computer interfaces and a stringent test of learned neural dynamics. Unlike neural decoding or masked reconstruction, this task requires a model to forecast future spiking activity from neural history alone and to remain stable when its own predictions become part of the conditioning context. We introduce NeuroHorizon, an encoder-decoder architecture that combines event-level tokenization of spike trains with an autoregressive causal decoder for population-level firing-rate prediction. The decoder uses a hierarchical tail+segment memory to retain recent predictions at high resolution while compressing longer prediction history, and scheduled sampling reduces the mismatch between teacher-forced training and autoregressive rollout. Evaluated on a multi-horizon motor-cortex benchmark with additional motor- and visual-cortex datasets for scaling and cross-population transfer, NeuroHorizon consistently outperforms strong baselines and remains stable under multi-step rollout where prior approaches collapse. These results show that long-horizon neural forecasting becomes feasible when models are explicitly trained and evaluated under autoregressive rollout, providing a foundation for closed-loop BCI and multi-step neural-state prediction.
NeuroMem: A Neuroplastic Memory Framework for Lifelong Agents through Delayed Consolidation
Yingyi Cheng ⋅ Xueqiang Han ⋅ Chen Yu
With long-term memory, LLM-based agents have demonstrated remarkable capabilities in handling complex long-horizon tasks. However, existing memory frameworks primarily treat long-term memory as a post-admission management problem, emphasizing how experiences should be organized, updated, or controlled. This overlooks a more fundamental stage in the memory lifecycle, i.e., how transient experiences are progressively transformed into stable and task-transferable knowledge. Biological memory consolidation provides a temporal principle for memory formation, whereby long-term memory emerges through the delayed stabilization and selective transformation of recent experience, rather than through its verbatim preservation. Motivated by this, we propose NeuroMem, a neuroplastic memory framework for lifelong agents through delayed consolidation. NeuroMem formulates the formation of task-transferable knowledge in long-term memory as a delayed-reward reinforcement learning problem. Specifically, newly formed memory candidates are first maintained in a transient buffer, where a consolidation policy learns to promote, merge, retain, or discard them only after their delayed utility rewards become observable through subsequent interactions. Experiments on EHRSQL and LoCoMo show NeuroMem achieves stronger performance with a smaller memory set, improving EHRSQL task success by up to 8.7% and LoCoMo F1/BLEU-4 by up to +5.20%/+0.97% over baselines.
Newton-PINet: A fast physics-informed neural network with Newton linearization for meta-learning nonlinear PDEs
Yuchen Fan ⋅ Chang Wei ⋅ Pao-Hsiung Chiu ⋅ Chin Chun Ooi ⋅ Heyang Wang ⋅ Jian Cheng Wong
Scientific machine learning has opened new avenues for solving parameterized partial differential equations (PDEs), enabling models to learn a family of PDEs and generalize to unseen instances. In this context, data-driven operator learning methods typically require large training datasets, while physics-informed neural networks (PINNs) suffer from difficult optimization and limited generalization, especially for nonlinear PDEs. We propose Newton-PINet, a physics-informed network enhanced by Newton linearization, offering an effective meta-learning framework for nonlinear PDEs. Newton-PINet (i) employs a physics-informed multilayer network with skip connections, where the output-layer weights are solved by least squares; (ii) adopts a two-stage learning strategy that first leverages gradient-based training to learn robust representations from the available training tasks, and then performs gradient-free fine-tuning on the output layer for fast task-specific generalization; and (iii) incorporates a Newton linearization method to speed up the least-squares iteration for nonlinear PDE problems. On a challenging nonlinear reaction-diffusion benchmark, Newton-PINet achieves up to three orders of magnitude lower relative error than recent neural solvers, while using 16× fewer training tasks and over an order of magnitude less training time (under 5 minutes versus several hours). This work advances the meta-learning of PINNs toward data-efficient, fast, and generalizable physics solvers. The datasets and code are provided in the supplementary material.
Next-Latent Prediction Transformers Learn Compact World Models
Jayden Teoh ⋅ Manan Tomar ⋅ Kwangjun Ahn ⋅ Edward Hu ⋅ Tim Pearce ⋅ Pratyusha Sharma ⋅ Akshay Krishnamurthy ⋅ Riashat Islam ⋅ Alex Lamb ⋅ John Langford
Transformers replace recurrence with a memory that grows with sequence length and self-attention that enables ad-hoc look ups over past tokens. Consequently, they lack an incentive to compress history into compact latent states with consistent transition rules. This often leads to learning solutions that generalize poorly. We introduce **Next-Latent Prediction** (NextLat), which extends standard next-token training with *self-supervised* predictions in the latent space. Specifically, NextLat trains a transformer to learn latent representations that are predictive of its next latent state given the next output token. Theoretically, we show that these latents provably converge to *belief states*, compressed information of the history necessary to predict the future. This simple auxiliary objective injects a recurrent inductive bias into transformers, while leaving their architecture, parallel training, and inference unchanged. NextLat effectively encourages the transformer to form compact internal world models with its own belief states and transition dynamics—a crucial property absent in standard next-token prediction transformers. Empirically, across benchmarks in world modeling, reasoning, planning, and language modeling, NextLat demonstrates significant gains over standard next-token training in downstream accuracy, representation compression, and lookahead planning. Furthermore, NextLat enables *variable-length self-speculative decoding*, accelerating inference by up to $3.3\times$ in the language domain. NextLat stands as a simple and efficient paradigm for shaping transformer representations toward stronger generalization.
Njord: A Probabilistic Graph Neural Network for Ensemble Ocean Forecasting
Daniel Holmberg ⋅ Joel Oskarsson ⋅ Erik Wikingsson ⋅ Fredrik Lindsten ⋅ Teemu Roos
Ocean dynamics are inherently chaotic, yet existing machine learning ocean models produce only deterministic forecasts. We introduce Njord, a probabilistic data-driven model for ocean forecasting, applicable to both global and regional domains. Njord combines a deep latent variable framework with a graph neural network architecture, enabling sampling each forecast step in a single forward pass. We apply Njord globally at 0.25° resolution and regionally to the Baltic Sea at 2 km resolution. To scale to these large ocean grids we introduce K-means cluster meshes that adapt to irregular sea surface geometry. Experiments demonstrate strong performance on both domains compared to deterministic machine learning baselines, while also providing uncertainty estimates from the sampled ensemble forecasts. On the global OceanBench benchmark, Njord achieves the lowest errors on average across upper-ocean variables when evaluated against real-world observations, with the largest improvements in surface temperature prediction.
Nonparametric Estimation of a Factorizable Density using Diffusion Models
Hyeok Kyu Kwon ⋅ Dongha Kim ⋅ Ilsang Ohn ⋅ Minwoo Chae
In recent years, diffusion models, and more generally score-based deep generative models, have achieved remarkable success in various applications, including image and audio generation. In this paper, we view diffusion models as an implicit approach to nonparametric density estimation and study them within a statistical framework to analyze their surprising performance. A key challenge in high-dimensional statistical inference is leveraging low-dimensional structures inherent in the data to mitigate the curse of dimensionality. We assume that the underlying density exhibits a low-dimensional structure by factorizing into low-dimensional components, a property common in examples such as Bayesian networks and Markov random fields. Under suitable assumptions, we demonstrate that an implicit density estimator constructed from diffusion models adapts to the factorization structure and achieves the minimax optimal rate with respect to the total variation distance. In constructing the estimator, we design a sparse weight-sharing neural network architecture, where sparsity and weight-sharing are key features of practical architectures such as convolutional neural networks and recurrent neural networks.
No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos
Matteo Balice ⋅ Yanik Künzi ⋅ Chenyangguang Zhang ⋅ Matteo Matteucci ⋅ Marc Pollefeys ⋅ Sunghwan Hong
Recent feed-forward 3D gaussian splatting methods have made dramatic progress on individual aspects of 3D scene reconstruction, but no existing method jointly addresses dynamic content, multi-view input, and unknown camera poses in a single feed-forward pass. Methods that handle dynamics either require accurate camera poses or accept only monocular input; pose-free multi-view methods address only static scenes; and per-scene optimization methods bridge some of these gaps but at minutes-to-hours cost per scene. We introduce NoPo4D, the first feed-forward system that addresses this empty quadrant. Building on a pretrained geometry backbone and recent 4D Gaussian frameworks, NoPo4D introduces a velocity decomposition that splits Gaussian motion into per-pixel image-plane shifts and depth changes, allowing direct supervision from pseudo ground-truth optical flow on the 2D component. This sidesteps both the differentiable rendering that couples prior posed methods to pose accuracy and the 3D motion ground truth that prior pose-free methods require. The system is rounded out by a bidirectional motion encoder for cross-view and cross-frame feature aggregation, and view-dependent opacity that mitigates cross-view and cross-timestep Gaussian misalignments. On four multi-view dynamic benchmarks, NoPo4D consistently outperforms prior feed-forward baselines, and with an optional post-optimization stage surpasses per-scene optimization methods, while running orders of magnitude faster. Code and the pretrained weights will be made publicly available.
NormLift: From Lifted Features To Semantic Reliability In 3D Gaussian Splatting
Yihan Zang ⋅ Da Li ⋅ Dominik Engel ⋅ Shinkyu Park ⋅ Ivan Viola
Training-free weighted aggregation is widely used to lift 2D semantic features onto 3D Gaussians for open-vocabulary scene understanding, yet its theoretical role remains insufficiently understood. Existing analyses typically justify this operation from the rendering side, treating Gaussian features as linearly composable Euclidean variables for reconstructing 2D feature maps. However, this view does not match downstream 3D usage, where each Gaussian is often queried independently in a cosine-based embedding space. We revisit feature lifting from the 3D side and formulate per-Gaussian assignment as a cosine alignment problem on the CLIP unit sphere. Under this objective, the $\ell_2$-normalized semantic back-projected feature emerges as the closed-form solution, providing a complementary interpretation of the standard lifting rule from the perspective of per-Gaussian semantic assignment. The same formulation further yields a norm decomposition into intra-view and inter-view consistency, suggesting that feature magnitude itself can serve as a semantic reliability signal. Calibrated by effective multi-view support, this reliability score guides a mode-voting refinement that preserves CLIP feature validity by avoiding linear averaging. Experiments on open-vocabulary 3D semantic segmentation show that NormLift is an efficient, training-free framework that achieves strong performance across evaluation protocols.
NorSA: Accelerating LLM Decoding via Normalized Sparse Activation
Tianteng Gu ⋅ Bo Xiao ⋅ Ke Zeng ⋅ Shuai Shi ⋅ Wangyou Zhang ⋅ Chenda Li ⋅ Yanmin Qian
Sparse activation accelerates Large Language Model (LLM) decoding by removing redundant computations and memory access, but existing methods often assume hidden-state dimensions are independent and identically distributed (i.i.d.). We replace this i.i.d. perspective with a covariance-aware view by analyzing contextual dependencies across tokens and inter-dimensional correlations within hidden states. We introduce Normalized Sparse Activation (NorSA), a training-free framework that combines dynamic thresholding with decorrelating rotation. Across the LLaMA, Mistral, and Qwen families, NorSA consistently improves perplexity and downstream accuracy over prior training-free baselines, especially in high-sparsity regimes. On LLaMA3-8B at 50\% sparsity, it stays within 0.44 perplexity points of the dense model while achieving a 1.32$\times$ end-to-end decoding speedup.
Not Only Where, But When: Temporal Scheduling for RLVR
Jinghao Zhang ⋅ Ruilin Li ⋅ Feng Zhao ⋅ Jiaqi Wang
Reinforcement learning with verifiable rewards (RLVR) has become a core technique for post-training of Large Language Models (LLMs). While policy optimization is driven by all sampled tokens under a globally broadcast scalar reward, the heterogeneous policy behaviors exhibited along trajectories are largely overlooked without differentiation. Existing works address this by credit allocation, including token-level advantage reweighting, and selective token optimization, however, the allocation criterion are principally stagnant throughout training, limiting resilient policy evolution. In this work, we argue that \textit{when} learning signals are scheduled can be as important as \textit{where} they are allocated across tokens, and introduce the temporal dimension that scheduling the credit allocation criteria over the course of RLVR optimization. We find that prioritizing targeted tokens emphasized with specific policy behaviors, and gradually attenuating toward general optimization leads to more stable and efficient learning dynamics. Furthermore, we show that simple trajectory percentiles provide a natural perspective for distinguishing policy behaviors, and works effectively with temporal scheduling. Our analysis reveals that standard optimization substantially sacrifices policy entropy when simultaneously accommodating heterogeneous behaviors, whereas temporal scheduling yields healthier policy evolution dynamics. Experiments across mathematical and general reasoning benchmarks demonstrate consistent improvements, suggesting that temporal scheduling constitutes a promising optimization dimension.
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
Zheng Huang ⋅ Mingyu Liu ⋅ Xiaoyi Lin ⋅ Muzhi Zhu ⋅ Ye Lin ⋅ Canyu Zhao ⋅ Zongze Du ⋅ Xiaoman Li ⋅ Hao Zhong ⋅ Yiduo Jia ⋅ Hao Chen ⋅ Chunhua Shen
Robot fine-tuning can improve manipulation success while degrading the semantic structure inherited from large vision-language pretraining. This trade-off limits the reuse of vision-language-action (VLA) policies in settings that require compositional instructions, camera changes, and coordination with higher-level agents. We present \textbf{NoTVLA}, a semantics-preserving robot adaptation framework that replaces dense low-level action supervision with a sparse narrative action interface. NoTVLA converts demonstrations into sparse, semantically meaningful waypoints, grounds each decision with a task-relevant visual anchor and depth value, and reconstructs executable motion through a deterministic detokenizer. The resulting interface keeps the autoregressive vision-language backbone close to its pretrained prediction format while delegating high-frequency control to a transparent motion-rendering stage. We evaluate this design through matched and backbone-family comparisons, semantic retention probes, semantic out-of-distribution manipulation tasks, camera and depth perturbations, and deployment-oriented efficiency analysis. The results suggest that robot adaptation should be judged not only by task success, but also by how much task-relevant semantic competence is preserved during fine-tuning.
NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion
Chuheng Chen ⋅ Xiaofei Zhou ⋅ Geyuan Zhang ⋅ Yong Huang ⋅ Gang Xiong
Low-Rank Adaptation (LoRA) fusion enables the composition of subject and style representations for controllable generation without retraining. However, existing approaches primarily operate through weight-level merging, without explicitly modeling how independently trained LoRAs interact in the shared parameter space. We adopt a geometric perspective on LoRA fusion, interpreting content and style LoRAs as occupying overlapping, non-orthogonal low-rank subspaces, where such overlap can lead to conflicting parameter updates that affect generation quality. This observation motivates us to reformulate LoRA fusion not merely as parameter combination, but as a problem of controlling how updates from overlapping subspaces are combined. Based on this insight, we propose Null Space Projection LoRA (NP-LoRA), a training-free framework that employs projection as a fusion operator to explicitly modulate cross-LoRA interactions. Specifically, NP-LoRA uses principal directions of the style LoRA to define a projection subspace and projects the content LoRA onto the complementary subspace (i.e., the null space of the style LoRA), suppressing interference along dominant style directions while preserving complementary information. To avoid the overly aggressive suppression of hard projection, we further formulate soft projection as a regularized optimization problem that balances content preservation against style-subspace suppression. This objective admits a closed-form solution, yielding a projection operator controlled by a single parameter that continuously interpolates between linear merging and hard projection. Extensive experiments across multiple pretrained LoRA pairs show that NP-LoRA achieves more balanced content-style composition compared to strong baselines, without requiring retraining.
NuMuon: Nuclear-Norm-Constrained Muon for Compressible LLM Training
Hadi Mohaghegh Dolatabadi ⋅ Thalaiyasingam Ajanthan ⋅ Sameera Ramasinghe ⋅ Chamin Hewa Koneputugodage ⋅ Shamane Siriwardhana ⋅ Violetta Shevchenko ⋅ Karol Pajak ⋅ James Snewin ⋅ Gil Avraham ⋅ Alexander Long
The rapid progress of large language models (LLMs) is increasingly constrained by memory and deployment costs, motivating compression methods for practical deployment. Many state-of-the-art compression pipelines leverage the low-rank structure of trained weight matrices, a phenomenon often associated with the properties of popular optimizers such as Adam. In this context, Muon is a recently proposed optimizer that improves LLM pretraining via full-rank update steps, but its induced weight-space structure has not been characterized yet. In this work, we report a surprising empirical finding: despite imposing full-rank updates, Muon-trained models exhibit pronounced low-rank structure in their weight matrices and are readily compressible under standard pipelines. Motivated by this insight, we propose NuMuon, which augments Muon with a nuclear-norm constraint on the update direction, further constraining the learned weights toward low-rank structure. Across billion-parameter-scale models, we show that NuMuon increases weight compressibility and improves post-compression model quality under state-of-the-art LLM compression pipelines while retaining Muon's convergence behavior.
ODDR: One-Step Deshadow Diffusion via Reward Guidance
Junseong Shin ⋅ Kijun Kim ⋅ Minseong Kim ⋅ Dongjin Kim ⋅ Tae Hyun Kim
Recent advances in deep learning for shadow removal have significantly enhanced image quality and realism. However, most approaches rely on real-world paired datasets, which are costly to collect and often limited in scene diversity, leading to limited generalization. To address these limitations, we propose One-step Deshadow Diffusion via Reward guidance (ODDR), a new framework that achieves efficient and high-fidelity shadow removal without relying on real-world paired supervision. Our method begins with One-step Deshadow Diffusion (ODD), a baseline model trained on synthetic shadow data for efficient one-step shadow-free reconstruction. We further adapt ODD into ODDR using ShadowReward. In contrast to traditional, annotation-heavy approaches, ShadowReward is the first reward model for shadow removal trained entirely without human annotation. It learns to mimic human perceptual judgments by ranking synthetically generated images with controlled degradations, such as texture distortion and boundary artifacts. This reward-guided fine-tuning enables ODDR to close the synthetic-to-real domain gap. Extensive experiments show that ODD achieves strong performance without relying on real-world paired supervision, and ODDR further improves the results and achieves performance competitive with methods trained on real-world paired data, while maintaining higher computational efficiency as a single step model.
OgBench: A Framework for Evaluating Graph Neural Networks on Omics Data
Louisa Cornelis ⋅ Johan Mathe ⋅ Louis Van Langendonck ⋅ Guillermo Bernárdez ⋅ Nina Miolane
Graph Neural Networks (GNNs) have become the dominant framework for inductive graph-level learning. Yet most benchmarks focus on the regime $n \gg p$, where the number of graphs $n$ greatly exceeds the number of nodes per graph $p$. This overlooks biological domains such as omics, which operate in the opposite $n \ll p$ regime, characterized by large graphs of genes, transcripts, or proteins across few patient samples. This raises the question: \textit{how do GNNs perform in this low-sample, high-node omics setting?} We introduce \texttt{OgBench} (Omics-Graph Bench), the first benchmarking platform for graph-level prediction in the $n \ll p$ regime characteristic of omics data. We provide a standardized, end-to-end modular infrastructure from raw omics data to families of featured graphs with varied structural properties. We benchmark classical GNNs, as well as GNNs designed for large graphs and omics applications, alongside MLPs and machine learning baselines to establish reference performances. Our results show that widely used GNNs often do not outperform simple MLPs and classical baselines. These findings challenge the prevailing assumption that graph structure inherently adds value in this domain, fostering a critical reassessment of current learning paradigms. Ultimately, by exposing these limitations, OgBench provides the open-source ecosystem necessary for the community to develop and validate novel architectures explicitly tailored for biological graphs. The code is available at \url{https://anonymous.4open.science/r/ogbench-C1A9/README.md}.
OGPO: Offline Goal-conditioned Policy Optimization for Recoverable Vision-Language-Action Models
Xule Gao ⋅ Xi Wang ⋅ Zehua Zang ⋅ Rui Wang ⋅ Changwen Zheng ⋅ Chuxiong Sun
Vision-language-action (VLA) models have shown strong promise in robotic manipulation, yet their post-training still predominantly relies on trajectory-level supervised imitation over robot demonstrations. While effective for learning feasible actions, this paradigm supervises policies to reproduce expert actions at each step, biasing them toward demonstration-specific action paths and providing limited guidance for recovering toward task-progress states once execution deviates from demonstrated trajectories. To address this limitation, we propose OGPO, an Offline Goal-conditioned Policy Optimization framework for recoverable VLA post-training. Instead of treating offline demonstrations merely as step-wise action labels, OGPO adaptively relabels trajectories with semantic key states and constructs goal-conditioned process rewards to optimize the policy for reaching these task-progress states. In this way, OGPO shifts VLA post-training from trajectory-level action imitation to key-state reachability learning, enabling policies to acquire a more robust ability to complete goal states from feasible off-demonstration states without requiring additional online interaction. Experiments on LIBERO and MetaWorld show that OGPO consistently improves VLA policy performance over standard supervised fine-tuning. Moreover, evaluations on LIBERO-Plus and our proposed LIBERO-ReAct, a mid-execution perturbation benchmark, further demonstrate that OGPO achieves stronger robustness and goal-directed recoverability under distribution shifts and recoverable off-demonstration states.
OmniDex: Scaling Dexterous Hand Grasping to Diverse Cluttered Scenes
Naiyu Fang ⋅ Zhongjin Luo ⋅ Yuxin Mo ⋅ Siyuan Huang ⋅ Jianbo Liu ⋅ Yufei Liu ⋅ Zheyuan Zhou ⋅ Chengkai Jin ⋅ Xiaogang Wang ⋅ Hongsheng Li
Dexterous grasping is the foundational primitive in embodied AI, demanding massive data to train robust models. As real-world data collection is expensive, simulation has become the mainstream paradigm. Yet, while cluttered scenes best reflect real-world applications, learning to grasp within them is bottlenecked by a critical scarcity of large-scale data. To resolve this, we curate high-quality 3D objects and supporting bases, proposing a scalable seed-and-filter strategy that bypasses sluggish scene-level optimization. This yields an unprecedented benchmark comprising over 2.6 million scenes and 0.4B scene-specific grasp ground truths, featuring diverse realistic layouts paired with rich semantic and geometric observations. Furthermore, we introduce the OmniDex model to overcome the grasp multimodality and last-millimeter precision errors plaguing current generative models. By coupling Soft Winner-Takes-All learning with human-inspired physical constraints during training, and utilizing physics-driven ranking, our approach achieves robust dexterous grasping without the latency of post-optimization. Experimental results show that OmniDex model achieves state-of-the-art performance and strong generalization across diverse scenes, views, and unseen objects.
Omni-Interactive Universal Embedder
Wei-Yao Wang ⋅ Kazuya Tateishi ⋅ Shuyang Cui ⋅ christian simon ⋅ Takashi Shibuya ⋅ Shusuke Takahashi ⋅ Yuki Mitsufuji
Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders due to their strong instruction-following capabilities. Despite this progress, existing approaches primarily focus on language and image modalities, which also remain the dominant modalities for human-AI interaction in current embedders. In this paper, we propose the first Omni-Interactive Universal Embedder (OmniUE), which not only learns a unified embedding space across text, video, and audio by leveraging intermediate-layer representations from dedicated learnable tokens, but also supports omni-interactive querying, enabling users to provide inputs in the form of text, visual regions of interest, and audio spans. Within OmniUE, visual and audio segmenters process diverse user interactions and integrate them with an omni-LLM to produce user-conditioned any-to-any embeddings via context aggregation. To evaluate OmniUE’s omni-interactive capabilities, we introduce OmniCHOIR, benchmarking models for omni-interactive compositional audio retrieval based on the given text, video, and audio as well as unimodal or multimodal interaction prompts. OmniUE consistently surpasses state-of-the-art baselines across diverse modalities, with average improvements of 10.5% on textual-interactive video benchmarks (MMEB-v2-video), 1.1% on audio tasks (MAEB), 83.7% on visual-interactive benchmarks (SCaR), and 24.1% on our omni-interactive OmniCHOIR benchmark We believe that jointly advancing omni-modal representation learning and omni-interactive querying paves the way toward universal embedders.
One-Shot Federated Graph Learning via High-Fidelity Proxies and Transferability-Guided Collaboration
jia jiang ⋅ Zihan Tan ⋅ Wenke Huang ⋅ Bin Yang ⋅ Mang Ye
One-shot Federated Graph Learning (OFGL) has emerged as a communication-efficient paradigm for Federated Graph Learning (FGL) under strict bandwidth constraints, compressing collaboration into a single transmission round. Nevertheless, this extreme setting suffers from a fundamental Single-Transmission Bottleneck, which limits both the quality and utility of exchanged knowledge. First, while existing methods attempt to distill private graphs into shareable proxies, their aggressive one-shot compression inherently triggers a fidelity crisis. This compression distorts local graph topology and semantics across clients, leading to externalized knowledge underexpression. Second, although the server aggregates these proxies to update global models, current integration schemes suffer from transferability agnosticism. The server cannot reliably identify cross-client collaboration potential, resulting in integrated knowledge underutilization. To tackle these dual challenges, we propose FedHFT, an effective One-shot Personalized Federated Graph Learning (OPFGL) framework that addresses this bottleneck from two perspectives, namely High-Fidelity Knowledge Externalization and Transferability-Ware Knowledge Integration. We use Fidelity-Driven Proxy Refinement to preserve client-side knowledge fidelity, and Transferability-Steered Personalized Collaboration to enable adaptive server-side collaboration. The superiority of FedHFT is validated through extensive experiments on both homophilic and heterophilic graphs. The code is available at https://anonymous.4open.science/r/FedHFT-DC7F/.
One-Time Soft Alignment Enables Resilient Learning without Weight Transport
Jeonghwan Cheon ⋅ Jaehyuk Bae ⋅ Se-Bum Paik
Backpropagation is the cornerstone of deep learning, but its reliance on symmetric weight transport and global synchronization makes it computationally expensive and biologically implausible. Feedback alignment offers a promising alternative by approximating error gradients through fixed random feedback, thereby avoiding symmetric weight transport. However, this approach often struggles with poor learning performance and instability, especially in deep networks. Here, we show that a one-time soft alignment between forward and feedback weights at initialization enables deep networks to achieve performance comparable to backpropagation, without requiring weight transport during learning. This simple initialization condition guides stable error minimization in the loss landscape, improving network trainability. Spectral analyses further reveal that initial alignment promotes smoother gradient flow and convergence to flatter minima, resulting in better generalization and robustness. Notably, we also find that allowing moderate deviations from exact weight symmetry can improve adversarial robustness compared to standard backpropagation. These findings demonstrate that a simple initialization strategy can enable robust and effective learning in deep networks without dynamic weight transport, providing a resource-efficient alternative to standard backpropagation.
On Fitting Flow Models with Large Sinkhorn Couplings
Stephen Zhang ⋅ Alireza Mousavi-Hosseini ⋅ Michal Klein ⋅ Marco Cuturi
Flow models transform data gradually from one modality (e.g. noise) onto another (e.g. images). Such models are parameterized by a time-dependent velocity field, trained to fit segments connecting pairs of source and target points. When the pairing between source and target points is given, training flow models boils down to a supervised regression problem. When no such pairing exists, as is the case when generating data from noise, training flows is much harder. A popular approach lies in picking source and target points independently (Lipman et al., 2023). This can, however, lead to velocity fields that are slow to train, but also costly to integrate at inference time. In theory, one would greatly benefit from training flow models by sampling pairs from an optimal transport (OT) measure coupling source and target, since this would lead to a highly efficient flow solving the Benamou-Brenier dynamical OT problem. In practice, recent works have proposed to sample mini-batches of $n$ source and $n$ target points and reorder them using an OT solver to form better pairs. These works have advocated using batches of size $n \approx 256$, and considered OT solvers that return couplings that are either sharp (using e.g. the Hungarian algorithm) or blurred (using e.g. entropic regularization, a.k.a. Sinkhorn). We follow in the footsteps of these works by exploring the benefits of increasing this mini-batch size $n$ by three to four orders of magnitude, and look more carefully on the effect of the entropic regularization $\varepsilon$ used in the Sinkhorn algorithm. Our analysis is facilitated by new scale invariant quantities to report the sharpness of a coupling, while our sharded computations across multiple GPU or GPU nodes allow scaling up $n$. We show that in both synthetic and image generation tasks, flow models greatly benefit when fitted with large Sinkhorn couplings, with a low entropic regularization $\varepsilon$.
Online Learning via Learned Latent Bayesian Tracking
Guy Gerson ⋅ Tomer Raviv ⋅ Nir Shlezinger ⋅ Tirza Routtenberg ⋅ Osvaldo Simeone
Online learning in non-stationary environments requires models to adapt rapidly from streaming data under strict computational constraints. A principled approach casts online learning as Bayesian state tracking, where model parameters are updated sequentially via Bayesian filtering. However, applying Bayesian filters directly to modern deep models is computationally prohibitive due to the high dimensionality of parameter space, forcing existing methods to rely on restrictive approximations or manually designed low-dimensional subspaces. In this work, we identify the absence of a suitable low-dimensional dynamical representation as the core bottleneck in Bayesian filtering-based online learning. Accordingly, we propose Adaptive Update through Representation Adaptation (AURA), a meta-learning framework that learns offline a low-dimensional latent state-space model governing the evolution of optimal model parameters under distribution shift. Online adaptation is then performed via extended Kalman filtering in this learned latent space followed by reconstruction of the full model parameters through a learned lifting map, enabling efficient single-step online adaptation while preserving model expressiveness. Evaluated on online adaptation of neural wireless receivers under time-varying channels and on non-stationary image classification, AURA shows substantial improvements in adaptation speed, accuracy, and computational efficiency over existing online learning and Bayesian filtering baselines, demonstrating that an adaptation-aware latent geometry is beneficial for effective Bayesian online learning in high-dimensional models.
On Minimizing Regret in Fixed-Confidence $\varepsilon$-Best Arm Identification
Tianyuan Jin ⋅ Junwen Yang ⋅ Vincent Tan
This paper studies the $\varepsilon$-best arm identification problem ($\varepsilon$-BAI) for $K$-armed bandits under the fixed-confidence setting. Departing from the usual objective of identifying the best arm with minimal sample complexity, we focus on minimizing the cumulative regret of identifying an $\varepsilon$-best arm. We consider the usual asymptotic regime in which the error probability $\delta$ of identifying an $\varepsilon$-best arm vanishes. In this setting, we show a novel instance dependent lower bound. In particular, we show that any family of $(\varepsilon,\delta)$-PAC algorithms must incur an expected cumulative regret up to the stopping time that scales as $\log(1/\delta)$, with an instance-dependent constant $C(\mu)$. This is proved using a novel regret-aware change-of-measure technique that allocates information given a budget on the regret. To achieve this lower bound, we devise an algorithm, \algname, that achieves the same scaling in $\delta$, i.e., $\log (1/\delta)$, but with a different constant, which we call $C^\*(\mu)$. KL-UCB+Greedy has two main benefits: (i) it is computationally efficient as $C^\*(\mu)$ admits a closed-form expression; (ii) in broad classes of reward distributions such as Gaussian distributions, $C^\*( \mu) = C( \mu)$, demonstrating asymptotic optimality of KL-UCB+Greedy within these classes. As auxiliary results, we also derive lower bounds on the sample complexity of any family of $(\varepsilon,\delta)$-algorithms as $\delta$ vanishes and show that the sample complexity of KL-UCB+Greedy closely matches that of the lower bound on the sample complexity.
On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation
Andy Q Han ⋅ Kristina Fujimoto ⋅ Avidan Shah ⋅ Kiet Nguyen ⋅ Kai Xu ⋅ Yueh-Han Chen ⋅ Ilia Sucholutsky ⋅ Rico Angell
Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Consistency training is a promising new alignment paradigm to mitigate such failures by training invariants into the model using contrastive input pairs. Existing consistency training procedures generate the supervision signal once, offline, and use supervised fine-tuning (SFT) to update the model. Unfortunately, the resulting models tend to merely memorize the surface forms of the training distribution and thus generalize poorly and regress in their capabilities. We introduce On-Policy Consistency Training (OPCT), a new consistency training approach where the objective is computed over the model's own responses to prompts, supervised by itself conditioned on corresponding contrastive prompts. We evaluate OPCT on three safety axes: sycophancy, jailbreaking, and safety awareness. Across three model families, OPCT outperforms its SFT counterpart on all safety desiderata. It nearly halves the sycophancy rate relative to baseline (8.1\% vs. 15.4\%, compared to 11.2\% for SFT). Under an adaptive per-target attacker, OPCT holds jailbreak defense success near 99\% on held-out jailbreak behaviors, whereas SFT achieves 87\% on average. On safety awareness, OPCT outperforms SFT in two out of three models, and matches it on the other. OPCT also largely avoids the capability regressions that SFT induces, such as a 28-point drop on MATH-500. Our results suggest that consistency training is best implemented as OPCT rather than as SFT, especially when generalization beyond the training distribution is desired.
On-Policy Distillation with Open Property-Equivalence Reward for LLM-Based NL-to-SVA Generation
Qingyun Zou ⋅ Yingze Li ⋅ Tianen Liu ⋅ Bingsheng He ⋅ Weng-Fai Wong
LLM-based generation of SystemVerilog Assertions (SVA) is often reported as nearing saturation, with the strongest specialized model reaching ${\sim}76\%$ accuracy on NL2SVA-Human. We show that this aggregate hides a temporal gap: models that appear strong overall still collapse to a few implication templates on bounded-delay and liveness specifications. The core issue is that the dominant recipe, supervised fine-tuning on NL/SVA pairs, optimizes token-level mimicry rather than the property equivalence that defines SVA correctness. We introduce Reward-Weighted On-Policy Distillation (RWOPD), an on-policy distillation method that samples student rollouts, scores them with an open SymbiYosys+Z3 Property-Equivalence Checker (PEC), and applies a verifier-reward-weighted forward-KL gradient from a frozen 14B teacher on verifier-passable rollouts. This keeps the supervision dense at every response token while grounding both selection and loss weight in property-equivalent behavior. RWOPD distills CodeV-SVA-14B into a Qwen2.5-Coder-7B-Instruct student that sets a new state of the art on NL2SVA-Human and NL2SVA-Machine across pass@1, pass@5, and pass@10, surpassing both specialized prior SOTA models and 671B general-purpose baselines.
Diffusion models are often trained in low-dimensional latent spaces, which are then reused for related but shifted datasets. In this work, we study when such latent reuse remains reliable under distribution shift. We consider a source-target setting in which both datasets are approximately low-dimensional but may lie near different subspaces. We show that freezing and reusing a source latent space induces a target-domain score error governed by two quantities: the principal-angle misalignment between the source and target subspaces, and the target ambient noise amplified by the diffusion time scale. Motivated by these limits, we further study mixed source-target training and characterize how the required shared latent dimension depends on the relative geometry of the two distributions. Our results provide theoretical guidance on when latent reuse is reliable and when learning a shared representation may be necessary.
Diffusion models are central to modern generative modeling, and understanding how they balance memorization and generalization is critical for reliable deployment. Recent work has shown that memorization in diffusion models is shaped by training dynamics, with generalization and memorization emerging at different stages of training. However, deployed diffusion models are often further distilled, introducing an additional training phase whose impact on memorization remains unclear. In this work, we analyze how distillation reshapes memorization behavior in diffusion models, taking consistency distillation as a representative framework. Empirically, we show that when applied to a teacher model that has memorized data, consistency distillation significantly reduces transferred memorization in the student while preserving, and sometimes improving, sample quality. To explain this behavior, we provide a theoretical analysis using a random feature neural network model [Bonnaire et al., 2025], showing that consistency distillation suppresses unstable feature directions associated with memorization while preserving stable, generalizable modes. Our findings suggest that distillation can serve not only as an acceleration tool, but also as a mechanism for improving the memorization-generalization trade-off.
On the Selectivity of Generative Models in Structure-Based Drug Design
Ella Miray Rajaonson ⋅ Jungyoon Lee ⋅ William Chau ⋅ Alan Aspuru-Guzik ⋅ Benjamin Sanchez-Lengeling ⋅ Dominique Beaini ⋅ Kirill Neklyudov ⋅ Marta Skreta
Designing small-molecule drugs that selectively bind to a target protein while avoiding unintended interactions with other proteins is essential for reducing side effects that contribute substantially to clinical attrition. Although structure-based drug design (SBDD) has advanced rapidly with deep learning, most generative models still optimize binding to a single target and disregard selectivity due to limited training data and unknown off-target interactions. Motivated by this gap, we introduce SelectBench, a comprehensive benchmark assessing the selectivity of generative SBDD models. SelectBench provides three curated evaluation tasks reflecting different real-world drug discovery scenarios: literature-based case studies, safety screening panels used in pharmaceutical development, and large-scale structure-based evaluation sets. We further introduce a suite of metrics to quantify selectivity and conduct a systematic evaluation of nine representative SBDD models, including methods explicitly trained for selectivity and widely-used inference-time guidance techniques. Our results show that even selectivity-trained models achieve only modest off-target discrimination and inference-time guidance provides limited improvements, underscoring the need for method improvement. SelectBench provides both a characterization of where current generative SBDD stands on selectivity and an open, extensible platform to accelerate progress.
OntoPlan: An Ontology-Grounded Scene Representation and Agentic Framework for Scalable Robot Task Planning
Hyeongwoo Nam ⋅ Woongje Cho ⋅ Juwon Kim ⋅ Jongeun Choi
Large language model (LLM)-based robot task planning is promising for open-ended instruction following, but degrades on long-horizon tasks in large environments. When spatial information is conveyed to the LLM through text, the model can fail to capture spatial context, and token cost grows with environment size. Generating action sequences directly with an LLM also makes it difficult to satisfy the current world state and action preconditions. We address this with an ontology-grounded scene representation that aligns objects, spaces, relations, and states in a shared symbolic vocabulary for spatial reasoning and task planning, and with OntoPlan, an agentic framework that interprets instructions, selectively retrieves task-relevant information, formalizes goals and constraints, and produces executable plans. Across 150 benchmark tasks spanning five indoor environments and three scene scales, OntoPlan achieves 0.83 average task success, compared with 0.30 for the strongest baseline, while using 17.2k total tokens per task on average, about 5.8$\times$ fewer than the most efficient baseline. These advantages persist as scene scale increases, whereas prior methods degrade more sharply in both success and token cost. OntoPlan also responds appropriately to ambiguous or infeasible instructions by asking follow-up questions or reporting insufficient information rather than committing to invalid plans. Code available at https://anonymous.4open.science/r/OntoPlan.
Open-Vocabulary 3D Part Segmentation with Semantic Propagation Hawkes Process
Kejie Shi ⋅ Kun Zhou ⋅ Jieyu Zhao ⋅ Xulun Ye
Existing open-vocabulary 3D part segmentation methods typically rely on point–text matching, where semantic organization is shaped implicitly by the training loss, and both training and inference lack an explicit forward semantic optimization mechanism. We address this limitation with a framework of forward semantic dynamics, implemented with coupled semantic-dynamics blocks that organize part semantics within a single forward pass. Specifically, differentiable forward low-rank semantic shaping constrains point features to a prompt-conditioned semantic subspace, suppressing semantic drift and improving intra-part consistency. Built on this shaped state, a finite-depth spatial Hawkes process for semantic propagation models high-confidence prompt responses as sparse semantic events and propagates them across multiple steps to provide reliable evidence for ambiguous regions. The propagated context is further fed back into later feature shaping, coupling semantic organization and evidence propagation without test-time iterative optimization or additional post-processing. Extensive experiments on multiple open-vocabulary 3D part segmentation benchmarks show that our method establishes a new state of the art, outperforming the previous comparable SOTA method on all 16 reported evaluation slices by 1.88 to 5.31 mIoU points (3.84 on average) and achieving the overall best results on 14 of them.
Detecting LLM reasoning failures at inference time without ground-truth labels motivates a family of confidence baselines, including self-consistency, semantic entropy, and $P(\text{True})$, built on within-question sampling and self-evaluation. Operad theory, the formalism for systems built by iterated substitution, suggests a complementary diagnostic: a model's direct answer to a compositional query should agree with the answer it produces by composing a stated decomposition of the same query. We instantiate this idea as *operadic consistency* (OC), a per-question signal. Across twelve instruction-tuned LLMs (4B--671B parameters, open-weights and closed-source) on four multi-hop QA datasets, OC is strongly correlated with accuracy on every dataset (Pearson $r \in [{+}0.86, {+}0.94]$, all $p \leq 0.0004$), making it a substantially stronger population-level accuracy proxy at roughly one-third the inference cost of the strongest sample-based baseline (whose own cross-model rates never exceed $r{=}{+}0.80$ on any dataset). At the per-question level, OC contributes information beyond self-consistency, semantic entropy, $P(\text{True})$, and constructed decomposition-aware baselines on every dataset (cluster-robust $p \leq 10^{-19}$). We translate the regression coefficients into deployment-ready selective-prediction lifts ($\Delta_{\text{AUARC}}$ up to $+0.088$, $\Delta_{\text{AUROC}}$ up to $+0.167$ over a tuned $\text{SC}_{K=3}$ baseline). On five frontier thinking models, where the decomposition is extracted from the model's own chain of thought, the same equal-cost comparison gives positive selective-prediction lift on all $16$ (dataset, budget, metric) cells tested.
Opt-Arena: Evaluating, Selecting, and Generating Optimization Modeling Data via Tripartite Graphs
Yian Xu ⋅ Xiongwei Han ⋅ Jie Wang ⋅ Haoyang Liu ⋅ Yuyang Cai ⋅ Boxuan Niu ⋅ Mingxuan Yuan
Optimization modeling is the cornerstone of Operations Research (OR). While large language models (LLMs) show promise in autoformulation, existing datasets suffer from unquantified redundancies and domain imbalances. Consequently, models overfit to frequently occurring patterns and fail on underrepresented constraints, inducing a structural capability bias. To address this, we propose Opt-Arena, a data-centric framework that formally represents optimization datasets as tripartite graphs. To map this topology, we introduce Atomic Modeling Information (AMI)—the minimal semantic-mathematical rules of optimization. This representation enables the quantification of structural coverage, difficulty, and overlap. Guided by these metrics, our core-set selection mechanism distills high-density, low-redundancy training data. To further resolve intrinsic domain absences, we extend this framework with a kernel-driven generation pipeline, leveraging foundational problem backbones to synthesize structurally feasible instances for underrepresented areas. Empirically, this approach exhibits exceptional sample efficiency: utilizing only 5k curated instances, it surpasses the 20k-budget performance plateau of conventional distance- and uncertainty-based baselines. Relying entirely on standard SFT driven by our data, lightweight empirical probes outperform both massive proprietary models and complex reasoning architectures enhanced by reinforcement learning. These results reveal that LLM formulation performance is fundamentally constrained by structural data diversity, positioning topology-aware curation as an important driver for advancing OR modeling capabilities.
OpticalRAG: Pixel-Space Compression for Token-Efficient Retrieval-Augmented Generation
Senhao Liu ⋅ Yuheng Zhang ⋅ Chunyu Wei ⋅ Yueguo Chen ⋅ Xinran Zhang
Retrieval-augmented generation (RAG) is bottlenecked by the token budget of large language models: feeding long documents into the context window is expensive, while aggressive textual compression discards fine-grained information that is irreversibly lost in a one-dimensional token sequence. We argue that this bottleneck is not fundamental to information density but to the choice of compression \emph{modality}. Modern visual encoders can pack the contents of a rendered text page into roughly a hundred visual tokens with little semantic loss, opening a new compression-fidelity tradeoff that purely textual methods cannot reach. We present \textbf{OpticalRAG}, a framework that renders documents as page images, encodes them with a frozen vision encoder, and distills the retrieved visual tokens into as few as $16$ injected tokens before passing them to the LLM. Two ingredients make this work without expensive cross-modal training: (i) \emph{Encoder-Bridged Retrieval}, which reuses the shared representation space of a pretrained vision-language model to retrieve visual chunks from a text query, with no contrastive training; and (ii) \emph{Query-Driven Distillation}, a lightweight cross-attention compressor that condenses thousands of visual tokens into a small set of query-conditioned prefix tokens. On LongBench-v2, BAMBOO, and LooGLE-v2, OpticalRAG achieves the best compression-aware accuracy (CAP) on all three benchmarks while injecting only $16$ tokens, and improves the closed-book Qwen2.5-7B-Instruct backbone by $5.17$ points on LongBench-v2.
Optimal Subgroup Discovery at Every Support Threshold
Lincen Yang ⋅ Qi Huang ⋅ Niki van Stein ⋅ Thomas Bäck ⋅ Matthijs van Leeuwen ⋅ Zhong Li
Subgroup discovery aims to identify regions of the feature space where a target variable deviates from its marginal distribution. A fundamental tension in this problem is the trade-off between a subgroup's deviance and its support: small subgroups can be highly deviant but statistically meaningless, while large ones cannot deviate much by construction. Existing methods resolve this tension by collapsing it into a single scalar quality measure, implicitly committing to one particular trade-off and making it difficult to recover the most deviating subgroup under a user-specified support constraint. We instead provide, to our knowledge, the first theoretical and structural characterization of KL-optimal subgroups under arbitrary support constraints. Under a partition-based generative model where the feature space decomposes into regions of homogeneous conditional law, which we call atoms, we prove that every point on the deviance-support Pareto front is a union of whole atoms together with at most one fractional atom. This reduces an uncountable search to a finite combinatorial problem. Guided by this theory, we propose a simple two-step recipe: learn a homogeneous partition, then select atoms greedily. Across 19 benchmark datasets, this recipe finds the most deviating subgroups across a range of support thresholds, and reliably improves subgroups returned by existing methods when applied as a post-processing step. These results empirically validate the sufficiency and necessity of our theoretical contributions.
Masked Diffusion Models (MDMs) are trained under an objective that is symmetric over decoding orders, but at sampling time a single order must be committed to. A common task in such models is to evaluate the log-probability of a candidate sequence, either as a downstream score or to rank a finite pool of candidates. We observe that this evaluation depends substantially on the decoding order chosen, even though the training objective averages over orders. For both a model trained on closed-form Probabilistic Context-Free Grammar (PCFG) data and a pre-trained MDM on OpenWebText, the per-order log-probability rankings of fixed candidate pools disagree on most inputs. We characterize this phenomenon and propose Order-Marginalized Scoring (OMS), a simple Monte-Carlo (MC) estimator of the order-marginalized log-likelihood $\tilde \mu(x)=E_\sigma[\log p_\sigma(x)]$. We use it in two complementary algorithms: a reranker over a fixed candidate pool, and a stochastic best-of-$N$ decoder that combines adaptive position selection with $\tilde\mu$-reranking. On Sudoku, the decoder substantially raises exact-match accuracy over the strongest single-order baseline from $65.3$% to $94.7$%. On a zero-shot protein-stability benchmark, the reranker improves Spearman correlation between log-probability scores and experimental folding stabilities over a per-order baseline by $0.07$. These results suggest that OMS is a useful alternative to per-order scoring for order-agnostic problems.
Order Matters: Competition-Guided Query Ordering for RNN-Based Object Detection
Shengjian Wu ⋅ Li Sun ⋅ Yu Shangguan ⋅ Qingli Li
DETR-style detectors use one-to-one bipartite matching during training to assign object queries to ground-truth objects, enabling end-to-end set prediction without non-maximum suppression (NMS). However, without an explicit de-duplication procedure, multiple queries can still produce highly similar hypotheses for the same object, making training unstable and predictions less decisive. Inspired by the sequential ordering of NMS, we propose DETRNN, a plug-and-play module that turns unordered object queries into a competition-aware sequence for recurrent refinement. DETRNN builds an explicit confidence-and-similarity based order from prior predictions, then refines queries with an RNN along this order to model competition inside the decoder. This ordered recurrent refinement reduces redundant predictions, stabilizes optimization, and improves final detection accuracy. Experiments on multiple DETR-style detectors show consistent gains with comparable efficiency. Code will be released on GitHub.
Multi-view 3D vision-language-action (VLA) policies are brittle to small sensor perturbations: depth jitter, partial occlusion, calibration drift and per-view dropout can move the predicted action chunk by far more than the perturbation itself. We trace the brittleness to the *fusion stage*: concatenating per-view tokens lets a small geometric shift flip arbitrary token indices, turning a continuous-mass-transport problem into a discrete indexing problem. We propose **a single substitution**: replace concat-then-attention fusion with a *differentiable Wasserstein barycentric tokenizer* that fuses $K$ view measures into $M$ barycentric tokens by transport. We prove a chain of Lipschitz-style stability bounds that hold conditional on a declared OT perturbation set, calibrated Lipschitz constants, and bounded Sinkhorn-solver residuals. Empirically, the barycenter is the single dominant mechanism behind the gain: (i) the proposed model has the lowest mean clean $\ell_2$ on synthetic, LIBERO, RLBench and CALVIN ABC→D (3-of-4 with non-overlapping 95% CIs); (ii) a controlled ablation on synthetic regresses by $46\times$ on the corruption-stability metric when the barycenter is removed and by $<1.5\times$ when any other component is; (iii) the same mechanism transfers to LIBERO with a $+37\%$ clean $\ell_2$ regression and $+62\%$ action-OT regression on real data; (iv) in synthetic closed-loop the concat-fusion baseline collapses from 85% clean to 0% under per-view dropout while the barycenter keeps a 92% worst-case across all five conditions (3 seeds, narrowest CI of all four methods).
Outbidding and Outbluffing Elite Humans: Mastering Liar’s Poker via Self-Play and Reinforcement Learning
Richard Dewey ⋅ Janos Botyanszki ⋅ Ciamac C Moallemi ⋅ Andrew Zheng
AI researchers have long focused on poker-like games as a testbed for environments characterized by multi-player dynamics, imperfect information, and reasoning under uncertainty. While recent breakthroughs have matched elite human play at no-limit Texas hold'em, the multi-player dynamics are subdued: most hands converge quickly with only two players engaged through multiple rounds of bidding. In this paper, we present Solly, the first AI agent to achieve elite human play in reduced-format Liar’s Poker, a game characterized by extensive multi-player engagement. We trained Solly using self-play with Regularized Nash Dynamics (R-NaD), a model-free, actor-critic, deep reinforcement learning algorithm, and we successfully extend R-NaD to the multi-player scenario for the first time. Solly played at an elite human level as measured by win rate (won over 50\% of hands) and equity (money won) in heads-up and multi-player Liar’s Poker. Solly also outperformed large language models (LLMs), including those with reasoning abilities, on the same metrics. Solly developed novel bidding strategies, randomized play effectively, and was not easily exploitable by world-class human players.
Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering
Cheng-Kuang Chang ⋅ Kai-Wei Chang ⋅ Alexander Liu ⋅ Jim Glass
Full-duplex spoken language models (FD-SLMs) enable seamless speech interaction by allowing models to listen and speak simultaneously, yet the internal mechanism by which they coordinate listening and speaking remains underexplored. We analyze the predictive behavior encoded in FD-SLM hidden representations and find that they exhibit stream-specific predictive patterns: during listening, they preferentially predict the incoming user stream, whereas during speaking, they preferentially predict the model-side output stream. Building on this observation, we show that FD-SLMs dynamically modulate their internal predictive focus between two states: a generative state aligned with model-side output generation and a perceptive state aligned with incoming user input. However, this modulation can lag behind abrupt changes in conversational context. During user interruptions, the model remains transiently biased toward the generative state before transitioning into the perceptive state, causing it to miss the beginning of the incoming input. We term this delayed internal transition state inertia. To quantify its downstream impact, we introduce the Zero-Buffer Benchmark (ZBB), a diagnostic benchmark for evaluating immediate interruption comprehension when user speech begins abruptly. We evaluate this setting using response correctness and initial-word occurrence rate (IWOR). Finally, we mitigate state inertia through activation steering with a perception vector, a training-free intervention with little additional computational overhead. Across multiple state-of-the-art FD-SLMs, activation steering substantially improves interruption handling; for example, on PersonaPlex, it improves correctness from 28% to 45% and IWOR from 40% to 72% without any fine-tuning.
P$^3$-VLM: A Point-based Alternative for Grounded 3D Vision-Language Models
Anna-Maria Halacheva ⋅ Jan-Nico Zaech ⋅ Sombit Dey ⋅ Luc V Gool ⋅ Danda Pani Paudel
Examining existing benchmarks for grounding in 3D VLMs reveals a pervasive "semantic leakage" in the evaluation and prompting protocols: a heavy reliance on bounding boxes as spatial prompts or representation primitives. Because bounding boxes are highly informative of object categories, they introduce a shortcut for the models to neglect visual information during 3D VLM training. We demonstrate this with a "blind" LLM that, prompted only with the box coordinates, can perform surprisingly well. To further analyze this shortcut, we propose point-based grounding prompting - querying models with a single 3D point instead of a box. Under this protocol, the accuracy of the blind LLM drops from $\textbf{36.4}$% to $\textbf{18.0}$% on ScanNet's Nr3D, and similarly degrades for 3D VLMs (e.g. LL3DA). Building on this insight, we introduce $P^3$-VLM, a novel tri-pathway 3D VLM that supports point-based prompts and achieves $\textbf{52.1}$% state-of-the-art accuracy on location-grounded captioning under this setting. To achieve this, we redesign the architecture while keeping it detector-free, scene-centric, and suitable for 3D GS inputs. $P^3$-VLM utilizes three complementary pathways (location-, task-, and context-aware), enabling strong reasoning about scenes without any bounding box dependencies. The key benefit of point-based learning is a consistent improvement across settings: $P^3$-VLM achieves state-of-the-art performance under point-based grounding while also improving robustness, OOD generalization, and performance on ungrounded scene-centric 3D VQA benchmarks, reflecting its reduced reliance on bounding-box shortcuts and resulting stronger visual-spatial reasoning. Our source code and models will be made publicly available.
PACE: Progress Actively-internalized Conditioning Execution for Long-horizon Manipulation
Fei Ni ⋅ Zhuo Chen ⋅ Xianze Yao ⋅ Zerui Chen ⋅ Changrui Chen ⋅ Yifu Yuan ⋅ Shan Luo ⋅ Jianye Hao ⋅ Jiankang Deng
Vision-language-action (VLA) models have made rapid progress in robotic manipulation, yet the policy acts with no sense of the pace at which its own execution unfolds—whether it is advancing, stalling, or has quietly gone wrong—especially for long-horizon tasks. Recent methods supply progress through external reward models or as an auxiliary prediction inside the policy, but the signal runs alongside action generation rather than actively shaping it. Moreover, progress is inherently tied to the terminal state, making estimation fragile when the policy lacks an internal awareness of where the task should end. We propose PACE, which actively internalizes task progress as a predictive interface bridging commitment and execution, equipping the policy with a calibrated sense of the pace at which it advances toward task completion. Specifically, we first equip the VLA with a foresight anchor, a predicted terminal representation from observation and instruction, from which progress emerges as a geometric projection. Building on this anchor, the model perceives a compact past–present–future progress state capturing how much the last step advanced, how far execution has come, and how far the next step should go. These predicted progress further actively conditions the action decoder to guide the prediction, while the gap between committed and realized increments provides a built-in consistency check for lightweight failure detection. Experiments across LIBERO, SimplerEnv, and real-world single-arm and bimanual platforms show that PACE achieves 98.7% on LIBERO and 74% on real-world manipulation, with consistent gains especially on long-horizon tasks. Robustness and stable generalization across diverse scenarios further confirm the advantage of internalizing progress within the policy, validating a calibrated sense of execution pace as an effective paradigm for long-horizon robot control.
PAI-Actor: Cinematic Multi-Actor Character Replacement in Dynamic Scenes
Heyuan Gao ⋅ Bangxun Tang ⋅ Yiren Song ⋅ Guian Fang ⋅ Zijian He ⋅ Jie Yang ⋅ Mike Zheng Shou
We present PAI-Actor, a cinematic multi-character animation framework for character replacement in dynamic movie scenes. Unlike conventional animation systems that mainly drive a single static image or a single subject, our goal is to replace and animate multiple characters within real video clips while preserving the original scene dynamics, camera motion, and background content. This setting is particularly challenging because the generated characters must remain consistent with the source performance in motion and interaction, while also matching the surrounding background in lighting, shadow, composition, and overall cinematic appearance. To address this, we formulate multi-character animation as a structure-guided human recovery problem and build a movie-driven training pipeline from high-quality film data. Furthermore, to support practical cinematic production, we introduce a bidirectional-to-autoregressive distillation framework: we first train a bidirectional diffusion transformer for high-quality short-clip generation at 1080P resolution, and then distill it into an autoregressive video-to-video model for efficient inference and longer video generation. Experiments show that PAI-Actor enables high-fidelity multi-character animation with strong scene consistency, cinematic visual quality, and efficient long-form generation.
PanoHK360: A Large-Scale 8K Urban Panoramic Dataset and Benchmark for Depth Estimation
Yunxiao Chen ⋅ Honghuan Lin ⋅ Ruisheng Wang ⋅ Yujun Liu ⋅ Kun Zhou ⋅ Shangfeng Huang ⋅ Tsz Nam Chan
Panoramic 360° depth estimation underpins autonomous navigation, 3D scene reconstruction, and scene understanding. However, existing outdoor panoramic datasets remain limited in three aspects. Small-scale collections are insufficient for modern data-hungry architectures and tend to produce models that generalize poorly across cities. Synthetic datasets transfer poorly to real scenes due to persistent domain gaps in illumination, geometry, and material appearance. Pseudo-label datasets distilled from pretrained estimators inherit the errors of the teacher model. Their downstream accuracy is therefore capped at the performance ceiling of the teacher itself. To address these issues, we present PanoHK360, a city-scale outdoor panoramic RGB-D dataset of diverse urban scenes in Hong Kong. The data are collected by a vehicle-mounted rig that pairs 8K panoramic imaging with survey-grade LiDAR. To the best of our knowledge, PanoHK360 is the first panoramic depth dataset that simultaneously combines three defining properties: city-scale geographic coverage, 8K equirectangular resolution, and sensor-based metric depth. The depth is recovered directly from LiDAR rather than estimated by a pretrained network. The release further includes raw point clouds, 6-DoF camera poses, and temporal capture sequences. We further benchmark representative monocular 360° depth methods on PanoHK360 under a unified protocol with a location-disjoint training and validation split. The results reveal a substantial gap between current methods and the level of resolution and scene diversity required for real-world urban deployment. These findings establish PanoHK360 as a challenging testbed for future research.
PanoWorld: Towards Spatial Supersensing in 360◦ Panorama World
Changpeng Wang ⋅ Xin Lin ⋅ Junhan Liu ⋅ Yuheng Liu ⋅ Zhen Wang ⋅ Donglian Qi ⋅ Yunfeng Yan ⋅ Xi Chen
Multimodal large language models (MLLMs) still struggle with spatial understanding under the dominant perspective-image paradigm, which inherits the narrow field of view of human-like perception. For navigation, robotic search, and 3D scene understanding, 360$^\circ$ panoramic sensing offers a form of supersensing by capturing the entire surrounding environment at once. However, existing MLLM pipelines typically decompose panoramas into multiple perspective views, leaving the spherical structure of equirectangular projection (ERP) largely implicit. In this paper, we study pano-native understanding, which requires an MLLM to reason over an ERP panorama as a continuous, observer-centered space. To this end, we first define the key abilities for pano-native understanding, including semantic anchoring, spherical localization, reference-frame transformation, and depth-aware 3D spatial reasoning. We then build a large-scale metadata construction pipeline that converts mixed-source ERP panoramas into geometry-aware, language-grounded, and depth-aware supervision, and instantiate these signals as capability-aligned instruction tuning data. On the model side, we introduce PanoWorld with Spherical Spatial Cross-Attention, which injects spherical geometry into the visual stream. We further construct PanoSpace-Bench, a diagnostic benchmark for evaluating ERP-native spatial reasoning. Experiments show that PanoWorld substantially outperforms both proprietary and open-source baselines on PanoSpace-Bench, H$^\ast$Bench, and R2R-CE Val-Unseen benchmarks. These results demonstrate that robust panoramic reasoning requires dedicated pano-native supervision and geometry-aware model adaptation. All source code and proposed data will be publicly released.
Parallel Fixed-Point Spiking Neurons for Efficient Training of Spiking Neural Networks
Qirong Yang ⋅ Guowei Peng ⋅ Xiurui Xie ⋅ Yuning Yang ⋅ Qiugang Zhan ⋅ Guisong Liu
Spiking Neural Networks (SNNs) are inherently sequential due to temporal state dependencies, limiting the efficiency of parallel hardware such as GPUs and resulting in high training cost. In this work, we propose the Parallel Fixed-Point Spiking Neurons (PFSN), which reformulates the dynamics of leaky integrate-and-fire neurons into a unified fixed-point mapping, enabling parallel computation across all time steps. Unlike conventional sequential unrolling, the proposed formulation decouples temporal dependencies through iterative fixed-point updates, significantly improving computational efficiency. Furthermore, we introduce a learnable temporal propagation operator that generalizes predefined dynamics and allows adaptive modeling of task-specific temporal interactions without relying on explicit temporal recursion. Extensive experiments across diverse domains, including event-based recognition, sequential image classification, speech processing, and time-series forecasting, demonstrate that PFSN consistently achieves superior efficiency while maintaining or improving predictive performance compared to existing parallel SNN approaches. These results highlight the effectiveness of combining fixed-point formulations with learnable temporal structures for scalable and efficient SNN training. Code is available at https://anonymous.4open.science/r/PFSN.
Parallel-in-Time Variational Inference for Latent Stochastic Differential Equations
Chenyang Wu ⋅ Pengfei Liu ⋅ Zongzhang Zhang
Latent stochastic differential equations (SDEs) model continuous-time, irregularly-sampled time series. Training these models typically relies on variational inference (VI) methods that suffer from sequential sampling bottlenecks and compounding integration errors. In this paper, we propose Parallel-in-Time Variational Inference (PiTVI), a framework that parameterizes the variational posterior as a non-Markovian process mapping the driving noise history directly to the local state increments. By leveraging modern sequence modeling architectures such as Transformer and Mamba, the training and inference of the variational posterior become fully parallelizable, reducing the span complexity from $\mathcal{O}(L)$ to $\mathcal{O}(\log L)$. Beyond computational scalability, we demonstrate that this formulation shifts global error propagation from multiplicative compounding to additive scaling. Furthermore, the non-Markovian formulation natively accommodates fractional driving noise, which models physical memory effects and provides a statistical relaxation mechanism when fitting smooth dynamics. We validate these properties across empirical time complexity scaling tests, long-horizon predictions on non-linear systems, and high-dimensional sequence modeling.
PARE: Pruning and Adaptive Routing for Efficient Video Generation
Yutong Wang ⋅ Yunke Wang ⋅ Tianfan Xue ⋅ Yu Qiao ⋅ Yaohui WANG ⋅ Xinyuan Chen ⋅ Chang Xu
Video Diffusion Transformers (DiTs) generate high-quality videos but demand substantial compute due to wide blocks, deep architectures, and iterative sampling. Recent methods reduce cost by compressing width, depth, or sampling steps, but typically commit to a fixed architecture that cannot adapt to individual inputs or denoising stages. We propose PARE (Pruning and Adaptive Routing for Efficient video generation), which jointly compresses width and depth with structure-aware pruning and input-adaptive routing. For width, we observe that attention heads specialize into spatial and temporal roles, and design importance scoring that accounts for this distinction to prevent motion-critical temporal heads from being pruned prematurely. For depth, we train a lightweight router conditioned on denoising timestep and visual content to dynamically select which blocks to execute at each step, enabling per-input compute adaptation rather than static block removal. A progressive pipeline first recovers width-pruned quality via distillation, then jointly optimizes the student and router to decouple the two learning objectives. Experiments on Wan2.1-14B for both image-to-video and text-to-video generation show that PARE substantially reduces per-step computation while preserving quality across VBench dimensions, and composes with step distillation for further acceleration.
PathNavigate: A Training-Free Pathology Agent with Surprise-Guided Scan and Shared Slide Memory for Whole-Slide VQA
Chunze Yang ⋅ Qidong Liu ⋅ Wenjie Zhao ⋅ Yue Tang ⋅ Jiusong Ge ⋅ Di Zhang ⋅ Jiashuai Liu ⋅ Lei Wu ⋅ Junbo Lu ⋅ Ni Zhang ⋅ Xian Wu ⋅ Zeyu Gao ⋅ Chen Li
Whole-slide image visual question answering (WSI-VQA) frames pathology as an extreme-context search problem: to answer a free-form clinical query, a system must first navigate a gigapixel slide under a strict inspection budget to locate sparse, high-resolution evidence. Existing approaches largely fall into two paradigms: i) supervised pathology multimodal large language models (MLLMs) and agents can absorb localization and reasoning into learned modules, but they often couple navigation to task-specific supervision and retraining, limiting their practicality; ii) training-free pathology agents avoid this cost by keeping core models frozen, but often follow a question-first design, i.e. constructing the initial candidate set mainly from query-conditioned relevance. This can miss decisive morphology that is not named in the question, and force heavier inference-time scaffolding. To address this challenge, we introduce PathNavigate, a training-free pathology agent built around a scan-search-readout routine. Before question matching, PathNavigate scans the current slide at low magnification with a shared online memory module over frozen pathology features, producing a slide-specific surprise field that marks an abnormal-region pool. It then applies question-conditioned PLIP relevance only within this pool to select high-magnification search targets. Finally, it extracts local high-magnification evidence and answers with a frozen perceptor-adjudicator stack, using the same online memory as slide-level context. Experiments on WSI-VQA and SlideBench-BCNB show that the proposed scan-search-readout design improves answer accuracy and yields more interpretable evidence-selection trajectories with higher efficiency. To ease reproducibility, we have released the code online.
PAVE: Prefill-Conditioned Activation Editing for Hallucination Mitigation in LVLMs
Jingmin Zhu ⋅ Junae Kim ⋅ Dinh Phung ⋅ Trung Le ⋅ Jianfei Cai ⋅ Qiuhong Ke
Large vision-language models (LVLMs) can caption images, answer visual questions, and reason over scientific diagrams, yet they still produce hallucinations unsupported by visual evidence. Existing approaches mitigate hallucinations either by reshaping the decoding distribution or by steering hidden activations with calibration-based directions. The former adjusts token probabilities but operates only at the output level, whereas hallucinations often arise when language priors dominate visual evidence in intermediate representations. The latter applies a fixed calibration intervention, which may fail to capture image-specific hallucination drift or inadvertently suppress grounded evidence. To address these limitations, we propose PAVE, a training-free method for input-specific hidden-state intervention. PAVE constructs offline subspaces for hallucination-drift suppression and visual-evidence amplification, and adapts the hidden-state update to each input using a per-image visual basis computed during prefill. By editing hidden states with this prefill-conditioned subspace, PAVE strengthens grounded evidence while suppressing hallucination drift, without relying solely on decoding-level corrections or fixed activation edits. Experiments on LLaVA-1.5-7B and Qwen-VL show that PAVE substantially reduces hallucinations on CHAIR and POPE while preserving grounded utility on MME. Code will be released.
PaxBench: A Multimodal Sequence Benchmark for Protein Abundance Prediction
Ke Zhai ⋅ Oscar J Charles ⋅ Helena A Saunders ⋅ Conrad Bessant
Protein abundance links genotype to cellular state and is useful for comparing genes, organisms, and engineered systems. PaxDb contains broad abundance mea- surements, but its heterogeneous records have not yet been organized into a shared benchmark for testing protein, CDS, codon, and mRNA representations under the same target. The difficulty is that abundance labels must be recovered from diverse PaxDb sources, matched to species-specific sequences at several molecular levels, and evaluated despite incomplete modality coverage, limited metadata, and variable annotation quality. To address this, we introduce PaxBench, a carefully cleaned and aligned seven-species PaxDb-derived benchmark with matched protein, CDS, codon, and mRNA inputs, and use it to test whether biological foundation models benefit from multimodal fusion. On the cluster-aware split, the full multimodal representation reaches Spearman 0.789, compared with 0.740 for the best single representation, while random evaluation reaches 0.834. Cross-species settings also remain predictive. These results suggest that protein, coding-sequence, codon, and transcript representations provide complementary predictive information under this benchmark. PaxBench provides a compact setting for comparing foundation models, feature baselines, adaptation methods, and interpretability analyses for sequence-based abundance prediction.
Population game dynamics describe how aggregate behavior evolves in response to payoff signals. In many applications, the payoff function may change, and the goal is to predict the population trajectory induced by a new payoff design before trajectory data under that design are available. We study this problem as payoff-aware prediction of population game dynamics. Standard dynamics learning methods fit the motion observed under training payoffs, but do not separate the payoff-dependent incentive from the payoff-independent response structure needed for prediction under new payoffs. Building on the existing framework of Riemannian game dynamics, we propose PA-RmD, a payoff-aware learning method that keeps payoff information explicit and learns the payoff-independent response structure from data. Theoretically, we show that the learned model preserves fixed-payoff Nash equilibria and derive finite-time one-step prediction bounds under shared-response and payoff-coverage conditions. Empirically, we evaluate PA-RmD on synthetic zero-sum and potential games, as well as a YouTube Trending multi-country dataset, showing that our method performs well under changing payoff designs.
PCDFusion: Proposal-Context-Detail Bayesian Rendering for Infrared-Visible Image Fusion
Rui Liu ⋅ Lin Gu ⋅ Hesong Li ⋅ Ying Fu
Infrared-visible image fusion (IVIF) aims to combine visible and thermal infrared observations for human perception and downstream vision tasks. Visible and thermal infrared observations are physically complementary but unreliable in different ways. Existing deep IVIF methods mainly improve feature interaction and reconstruction, but often lack an explicit local rule for deciding when visible appearance should dominate and when thermal responses should contribute. To address this issue, we reformulate IVIF as local Bayesian posterior rendering, where fusion is controlled by reliability posteriors rather than direct modality mixing. In this paper, we propose Proposal-Context-Detail Fusion (PCDFusion), a visible-anchored framework that implements this formulation through three rendering stages. PCDFusion preserves potentially useful thermal responses, updates exposure-degraded visible regions, and renders reliable band-limited infrared structures into the final luminance. Extensive experiments on public IVIF benchmarks show that PCDFusion improves fusion quality, target recovery, and downstream detection and segmentation performance under challenging illumination. To support reproducibility, the code will be publicly released upon acceptance.
PCEval: A Benchmark for Evaluating Physical Computing Capabilities of Large Language Models
Inpyo Song ⋅ Eunji Chon ⋅ Jangwon Lee
Large Language Models (LLMs) have demonstrated remarkable capabilities across various domains, including software development, education, and technical assistance. However, a fundamental question remains unresolved: can LLMs reason about physical constraints—the rules governing how abstract computations must map onto real-world hardware? This capability is distinct from, and arguably more demanding than, software generation alone, as it requires grounding language understanding in spatial, electrical, and topological constraints. To address this gap, we introduce PCEval (Physical Computing Evaluation), a benchmark and evaluation protocol for physical computing education that enables fully automatic, execution-based assessment of LLMs across both logical and physical implementation artifacts, without human grading. Our evaluation framework assesses LLMs in generating circuits and producing compatible code across varying levels of project complexity. Through comprehensive testing of 14 leading models, PCEval provides the first reproducible and automatically validated empirical assessment of LLMs' ability to reason about fundamental hardware implementation constraints within a simulation environment. Our findings reveal that while LLMs perform well in code generation and logical circuit design, they struggle significantly with physical breadboard layout creation, particularly in managing proper pin connections and avoiding circuit errors. PCEval advances our understanding of AI assistance in hardware-dependent computing environments and establishes a foundation for developing more effective tools to support physical computing education.
PCFBench: How Far Can Large Vision-Language Models Go in Physics-aware Photonic Inverse Design?
Shengchao Chen ⋅ Ting Shu ⋅ Sufen Ren
Photonic crystal fibres (PCFs) underpin critical applications in telecommunications, sensing, and high-power laser delivery, yet their design remains a manual, expert-intensive process that couples geometric intuition with deep waveguide physics. Vision-language models (VLMs) offer a promising path toward automating this pipeline, but can they truly reason about photonic physics from microscopy images, or do they merely recognise visual patterns? We introduce PCFBench, the first comprehensive benchmark for evaluating VLMs on the complete photonic inverse-design workflow. PCFBench spans 26 tasks organised into six categories of increasing cognitive demand (Geometry Perception, Physics Understanding, Text Generation, Multi-modal Reasoning, Inverse Design, and Code Generation), built from over 220K FDTD simulation samples across 18 structurally diverse fibre families. We evaluate both individual VLMs on isolated tasks and multi-agent systems (MAS) on the full sequential pipeline from perception to design. Our results reveal a striking perception-design gap: current models can identify fibre geometry with moderate success, yet consistently fail to translate that visual understanding into physics-aware design actions. Moreover, cascading agents through the pipeline exposes significant error propagation, where upstream perception mistakes compound into downstream design failures. These findings suggest that neither model scaling nor naive agent chaining suffices; closing the loop demands explicit physical grounding. PCFBench provides the community with a rigorous, multi-dimensional testbed to measure progress toward AI-assisted photonic engineering.
Perception, Not Reasoning, Limits Visual Theory of Mind
Mohamed Rayan Barhdadi ⋅ Syed Talal Wasim ⋅ Jürgen Gall ⋅ Erchin Serpedin ⋅ HASAN KURBAN
Vision-language models pass text-based false-belief probes at near-human levels but drop by twenty or more percentage points the moment the same problem is conveyed visually. The consensus diagnosis is conceptual, prescribing new belief-reasoning modules. We argue the bottleneck is perceptual: not what the model cannot reason about, but what it cannot see. We propose a five-condition dissociation protocol that holds the reasoning task constant while varying perceptual fidelity from visual only to text only, and quantify the result with the Perceptual Contribution Ratio, a dose-response metric measuring how much of the visual-to-text gap a scaffold of given fidelity closes. Instantiated on abstract 2D panels in the Heider-Simmel tradition and naturalistic 3D game scenes from MuMA-ToM and MindPower across seven VLMs and roughly $147{,}000$ API calls, ground-truth scaffolds close the gap on every regime, matching or exceeding the text-only ceiling on abstract panels and MuMA-ToM. Automatic parsers, including frontier VLMs as zero-shot scene describers, fail. The dominant visual-only error mode is \emph{reality bias} ($60\%$ of false-belief failures): models report where an object currently is rather than where an absent agent believes it to be, the signature of a front-end that fails to encode who was present when. The route to better visual social reasoning runs through perception, not new reasoning modules.
Periodic Complex Stochastic Processes for Retrieving Atomic Structures of Unknown Matters
Gyoung S. Na ⋅ Chanyoung Park
Retrieving unknown atomic structures from observable analytical spectra or images remains a long-standing challenge across natural sciences. However, retrieval accuracy of existing retrieval methods on analytical data remains suboptimal because they have overlooked the underlying periodic quantum mechanical perturbations of non-equilibrium atomic structures behind analytical data. This paper proposes a periodic complex stochastic process (PCSP) that models such periodic perturbations and establishes theoretical backgrounds of periodic stochastic process in the complex-valued domain, including its sample diversity, process length, and periodicity. Finally, we develop a complex-valued cross-modal retrieval (CVCR) by integrating PCSP with cross-modal retrieval frameworks. CVCR outperformed existing cross-modal retrieval methods in cross-modal retrieval tasks of real-world analytical chemistry. Moreover, CVCR achieved state-of-the-art retrieval accuracy in zero-shot settings on 40 million of molecules.
PerQ: Inverse Generative Modeling for Neural Image Compression via Quantization Error Compensation
Tianhao Peng ⋅ Ho Man Kwan ⋅ Fan Zhang ⋅ Shan Liu ⋅ David Bull
Neural image codecs face a long-standing trade-off: distortion-optimized designs preserve pixel-wise fidelity but tend to produce blurred reconstructions at low bitrates, while generative designs close the perceptual gap with a learned prior, but with pixel fidelity saturating as bitrate increases. The latter typically requires training multiple models for various bitrate targets. This paper addresses these limitations by reframing the perceptual reconstruction problem. Starting from the quantization bound that holds for any modern multi-rate codec, we derive a transport bound that confines the perceptual reconstruction to a closed, rate-determined cell around the backbone's output. The cell is small relative to the full image manifold, casting unconstrained generative modeling as a bounded-support problem. We instantiate this as PerQ: an image codec that attaches a rate-conditioned flow-matching compensator supported on this cell to a frozen multi-rate distortion-optimized backbone. The experimental results across Kodak and CLIC2020-test datasets show that PerQ offers similar performance to generative neural image codecs (e.g., MS-ILLM) in perceptual metrics (LPIPS, FID) , while still maintaining competitive coding performance compared to distortion-optimized codecs (e.g., ELIC). Moreover, PerQ can produce compression results across a wide bitrate range and fast decoding from only a single checkpoint.
Personalization shifts the target of LLM alignment from population-level rewards to user-conditioned policies. \textbf{This position paper argues that personalized LLMs must be counterfactually verifiable: models should answer, on demand, three causal queries --- attribution, necessity, and sufficiency --- regarding their inferred user representations.} Absent this property, behavioral evaluation cannot reliably distinguish legitimate adaptation from sycophancy, exploitation, or spurious correlation. Because two models can produce similar aggregate outputs while differing in causal structure (for eg., one truthful, one sycophantic; one respecting protected attributes, one silently conditioning on them) the shift to personalization renders standard benchmarks structurally inadequate. We formalize this via a structural causal model (SCM) of personalized alignment and characterize when these counterfactual queries are identifiable despite the latent nature of inferred user states. We propose a black-box estimation strategy based on rewriting prompts twice to cancel off-target perturbations, and introduce selective counterfactual invariance to bridge the gap between personalization and counterfactual fairness. Ultimately, counterfactual verifiability is both a technical prerequisite for evaluation and a normative standard for the responsible deployment of personalized LLMs.
Perturb, Repair, Verify: Self-Play Vision-Language Verifiers for Compositional Understanding
Zijun Huang ⋅ Yingjun Du ⋅ Cees Snoek
Vision-language verifiers judge whether an image faithfully matches a textual description, yet existing training pipelines typically rely on static supervision from stronger external judges or curated hard-negative data. Such supervision is costly, teacher-bounded, and difficult to adapt as model failures evolve. We introduce PROVE (Perturb-and-Repair Optimization for Vision-Language Evaluation), a self-play framework that trains verifiers through a closed perturb–verify–repairverify loop. Starting from a known-correct image-caption pair, a Mutator introduces a typed perturbation over objects, attributes, counts, or relations. The verifier then judges the perturbed caption and provides textual feedback, which a Repairer uses to recover a faithful caption for re-verification. These two verification rounds produce structured rewards for informative mismatches and visually grounded repairs. Since perturbation, verification, and repair are role-conditioned behaviors of the same evolving multimodal backbone, updating the shared parameters also improves the verifier itself. Experiments show that PROVE improves fine-grained compositional verification and visual-detail robustness while preserving general multimodal capability.
PG-LRF: Physiology-Guided Latent Rectified Flow for Electro-Hemodynamic PPG-to-ECG Generation
Xiaoda Wang ⋅ Minxiao Wang ⋅ Kaiqiao Han ⋅ Defu Cao ⋅ Ching Chang ⋅ Yidan Shi ⋅ Runze Yan ⋅ Xiao Luo ⋅ Yan Liu ⋅ Xiao Hu ⋅ Yizhou Sun ⋅ Wei Wang ⋅ Carl Yang
Electrocardiography (ECG) is the clinical standard for cardiac assessment but requires dedicated hardware that does not scale to daily-life monitoring. Photoplethysmography (PPG) is ubiquitous in wearables but lacks ECG-specific diagnostic morphology and is corrupted by motion and sensor noise. PPG-to-ECG generation aims to bridge this gap by recovering electrical morphology and timing from peripheral pulse signals. However, existing methods largely rely on statistical alignment and data-driven generation. They fail to explicitly structure the latent space around physiology-aware electro-hemodynamic factors and lack constraints from forward physiological dynamics. To address these challenges, we propose PG-LRF, a physiology-guided latent rectified flow framework. PG-LRF introduces an electro-hemodynamic simulator that co-models ECG and PPG through shared cardiac phase dynamics. Guided by this simulator, a Physiology-Aware AutoEncoder learns a structured electro-hemodynamic latent space. Then we integrate this simulator guidance into a PPG-conditioned latent rectified flow, enforcing ECG-side morphology consistency and ECG-to-PPG forward hemodynamic consistency during generative transport. Experiments on the large-scale MC-MED dataset demonstrate that PG-LRF significantly improves PPG-to-ECG generation and downstream cardiovascular disease classification, proving its ability to generate ECGs that are both signal-faithful and physiologically plausible under the ECG-to-PPG hemodynamic pathway
PGSB: Pretrained-Guided Shared Basis for LoRA Model Merging
Muqing Liu ⋅ Chongjie Si ⋅ Zhuoya Liu ⋅ Yuheng Jia
Large pretrained models have achieved remarkable success across diverse domains, yet adapting them to multiple downstream tasks remains challenging: full fine-tuning is costly, while maintaining separate task-specific models is storage-inefficient and fails to produce a unified multi-task model. Low-Rank Adaptation (LoRA) offers an efficient alternative for Model Merging by representing each task adaptation as a low-rank update over frozen pretrained weights. However, existing LoRA merging methods often treat task-specific updates as directly comparable objects and merge them in their original low-rank parameter spaces. This overlooks two key factors: which update directions provide reliable shared support across tasks, and how these directions should be re-parameterized with respect to the pretrained weights before merging. To address this limitation, we propose Pretrained-Guided Shared Basis (PGSB), a two-stage framework for LoRA model merging. PGSB first identifies shared directions among task-specific LoRA updates, and then re-parameterizes these directions using their interaction with the pretrained weights, producing a more suitable basis for merging. Extensive experiments demonstrate that PGSB consistently outperforms state-of-the-art baselines across multiple benchmarks.
Phase-DGS: Phase-Guided Dynamic Gaussian Splatting from Unsynchronized Multi-view Video
Hosung Jeon ⋅ Jun Y Jeong ⋅ Sangwoon Kwak ⋅ Joonsoo Kim ⋅ Sangmin Kim ⋅ Jaesik Park ⋅ Won-Sik Cheong ⋅ Hyon-Gon Choo
Spatiotemporal scene reconstruction with Dynamic Gaussian Splatting (DGS) fundamentally depends on perfect temporal synchronization across all cameras, an assumption rarely achievable in practice. Industry-standard synchronization requires expensive specialized hardware and complex setup, while data-driven alternatives including audio-based alignment, geometry matching, and pose tracking remain limited by restrictive environmental or scene-specific assumptions. To overcome these limitations, we propose Phase-DGS, a framework that replaces conventional timestamps with a semantic phase representing each frame's inherent state within the underlying motion cycle. Frames capturing the same motion state share the same phase regardless of which camera records them or when, naturally establishing spatiotemporal correspondence across unsynchronized cameras. We validate Phase-DGS under temporal offsets, frame drops, and camera freezes, achieving up to +8.5 dB PSNR improvement over baselines while enabling seamless integration with existing DGS backbones including RealTime-4DGS and FreeTimeGS.
PhaseLoRA: Control-Regime-Conditioned Low-Rank Adaptation for Continuous-Action Vision-Language-Action Policies
Yufei Guo ⋅ Yinan Wu ⋅ Haoran Duan ⋅ guiguang ding ⋅ Jungong Han
Parameter-efficient fine-tuning (PEFT) is a natural way to adapt pretrained vision-language-action (VLA) policies, but most adapter designs apply temporally static updates throughout a control rollout, overlooking the phase-dependent nature of continuous-action manipulation. Such policies traverse distinct regimes, including approach, contact transition, grasping, transport, and placement, each requiring different adaptation behaviors. We propose \textbf{PhaseLoRA}, a lightweight LoRA parameterization that conditions adaptation at each control step using two weakly supervised descriptors: fine-control tendency and event/boundary intensity. PhaseLoRA modulates the LoRA left factor in the action expert, allowing the effective low-rank update direction to vary over time while keeping the backbone largely frozen. On LIBERO, PhaseLoRA improves average success rate by 12.2 points over a matched-parameter high-rank LoRA baseline and outperforms stronger LoRA variants. Ablations and update-direction analyses show that the gain is not explained by parameter count, arbitrary temporal modulation, or scalar gating, but by control-regime-aligned changes in low-rank update directions. These results establish within-trajectory conditioning as an effective lightweight PEFT axis for continuous-action VLA policies.
Phase-wise Velocity Distillation: Towards Effective Image Generation with A Single NFE
Zhen Guo ⋅ Rongyuan Wu ⋅ Qiaosi Yi ⋅ Chenxi Xie ⋅ Xinyu Wei ⋅ Lei Zhang
While diffusion distillation methods have largely accelerated image generation, achieving high-quality synthesis under a single number of function evaluations (NFE) budget remains a challenging problem. The diffusion process follows a coarse-to-fine progression, yet existing solutions typically force a single student model to simultaneously resolve global structures and fine details in one forward pass, leading to over-smoothed outputs. To address this issue, we propose **P**hase-wise **V**elocity **D**istillation (**PVD**), which strategically partitions the generation timeline into a coarse and a fine phase, and models the transition within each phase via the average velocity. A dedicated half-sized expert is then assigned to each phase, decoupling structural composition from detail refinement while keeping the cumulative cost within a single-NFE budget of the teacher. While this design already suffices for class-conditional image (C2I) generation, we further introduce phase-wise adversarial supervision with dedicated discriminators for the more complex text-to-image (T2I) tasks, ensuring accurate distribution matching. On C2I generation, PVD achieves a state-of-the-art FID of 1.48 under a single-NFE budget on ImageNet $256 \times 256$. On T2I tasks, PVD-distilled models (Stable Diffusion 3.5-Medium, FLUX.1-dev) produce results competitive with their multi-step teachers, significantly outperforming prior distillation methods. Source codes and distilled models will be released.
PhyMo: Learning Physical Dynamics with Accurate and Continuous Motion from Multi-View Videos
Shangjia Liu ⋅ Jinxi Li ⋅ Siyuan Zhou ⋅ Bo Yang ⋅ Bing WANG
Recovering dynamic 3D scene geometry, appearance, and the physical motion states that govern scene evolution from multi-view videos is important yet challenging. A key difficulty is that appearance can be deceptive. When trained primarily with rendering supervision, existing methods may fit the observed appearance well by exploiting appearance shortcuts, rather than recovering the true evolving geometry and motion. This issue is especially severe in scenes with fast or locally complex motion, where inaccurate trajectories can still produce plausible images over the observed frames, but lead to unreliable reconstruction and discontinuous future prediction. In this paper, we propose PhyMo, a framework for learning evolving dynamics from multi-view videos. We introduce a Motion-Aware Trajectory Representation that drives the model to recover accurate motion from deformation, rather than relying on appearance shortcuts to explain observed motion. On top of this, we propose a Continuous Motion Evolution module with bidirectional kinematic continuity constraint to promote temporally consistent evolution of displacement, velocity, and acceleration. Together, these designs yield geometrically faithful interpolation and continuous motion extrapolation. Experiments on four dynamic scene benchmarks demonstrate state-of-the-art performance.
PISA: Piecewise Sparse Attention Is Wiser for Efficient Diffusion Transformers
Haopeng Li ⋅ Shitong Shao ⋅ Wenliang Zhong ⋅ zikai zhou ⋅ Lichen Bai ⋅ Hui Xiong ⋅ Zeke Xie
Diffusion Transformers are fundamental for video and image generation, but their efficiency is bottlenecked by the quadratic complexity of attention. While block sparse attention accelerates computation by attending only critical key-value blocks, it suffers from degradation at high sparsity by discarding context. In this work, we discover that attention scores of non-critical blocks exhibit distributional stability, allowing them to be approximated accurately and efficiently rather than discarded, which is essentially important for sparse attention design. Motivated by this key insight, we propose PISA, a training-free Piecewise Sparse Attention that covers the full attention span while substantially reducing computational cost. Unlike the conventional keep-or-drop paradigm that directly drop the non-critical block information, PISA introduces a novel exact-or-approximate strategy: it maintains exact computation for critical blocks while efficiently approximating the remainder through block-wise Taylor expansion. This design allows PISA to serve as a faithful proxy to full attention, effectively bridging the gap between speed and quality. Experimental results demonstrate that PISA achieves 1.91times and 2.57 times speedups on Wan2.1 and Hunyuan-Video, respectively, while consistently maintaining the highest quality among sparse attention methods. Notably, even for image generation on FLUX, PISA achieves a 1.2 to 1.6 times acceleration at different resolutions without compromising visual quality.
PISCO: Precise Video Instance Insertion with Sparse Control
Xiangbo Gao ⋅ Renjie Li ⋅ Xinghao Chen ⋅ Yuheng Wu ⋅ Suofei Feng ⋅ Jie Yang ⋅ Qing Yin ⋅ Zhengzhong Tu
AI video generation is moving beyond general generation, which relies on exhaustive prompt-engineering and "cherry-picking", towards fine-grained, controllable generation and high-fidelity post-processing. In professional AI-assisted filmmaking, the core requirement is the ability to perform precise, targeted modifications. A key task is video instance insertion, which requires precise spatial-temporal placement, physically consistent scene interaction (e.g., shadows and reflections), and the faithful preservation of original dynamics - all achieved under minimal user effort. In this paper, we propose PISCO, a video diffusion model for precise video instance insertion with arbitrary sparse keyframe control. PISCO allows users to specify a single keyframe, start-and-end keyframes, or sparse keyframes at arbitrary timestamps, and automatically propagates object appearance, motion, and interaction. To stabilize generation under sparse conditioning, we introduce Variable-Information Guidance and Distribution-Preserving Temporal Masking, complemented by geometry-aware conditioning. We further construct PISCO-Bench, a benchmark with verified instance annotations and paired clean background videos, and evaluate performance using both reference-based and reference-free perceptual metrics. Experiments demonstrate that PISCO consistently outperforms existing baselines and scales effectively as additional control signals are provided.
PitchBench: Measuring Pitch Hearing in Audio-Language Models
Milan Liessens Dujardin ⋅ Song-Ze Yu ⋅ Craver C Thomas-Smith ⋅ David Chan ⋅ Karina Nguyen
Audio-language models (ALMs) are increasingly used in real-world applications that require understanding music, from music tutoring and transcription to captioning, recommendation systems, and music production. More broadly, they are becoming an important component of multimodal AI systems that must reason from sensory input rather than text alone. This makes reliable musical perception a critical prerequisite: if a model cannot accurately hear the structure of sound, it cannot be trusted to reason about, teach, transcribe, or act on audio in the real world. Yet existing benchmarks rarely assess one of the most fundamental musical abilities underlying such perception: pitch hearing. Current evaluations tend to probe pitch hearing only indirectly, through higher-level tasks and often in multiple-choice formats, leaving open how reliably ALMs identify fine-grained pitch across instruments, acoustic conditions, and response formats. We introduce PitchBench, an evaluation suite that systematically measures pitch hearing in ALMs. PitchBench comprises 28 experiments spanning absolute and relative pitch perception within sequences and chords, while varying loudness, note duration, sound source, time stretching, background noise, and other acoustic conditions. Tasks range from identifying individual pitches in isolation to tracking a melodic line within a four-part musical texture. Evaluating frontier ALMs, we find that pitch hearing remains highly unreliable: models perform consistently poorly across settings, with accuracy varying sharply by sound source, note duration, and notation format. Current ALMs do not yet possess stable pitch perception, even for controlled synthetic and instrumental stimuli. Alongside the benchmark, we release PitchBench as a Python package containing the evaluation data and data generation tools to support future work on pitch-aware audio-language modeling.
PixelART: Image-to-Layer Decomposition without Latents or Text-to-Image Pretraining
Zelin Jia ⋅ Zhao Zhang ⋅ Zhicong Tang ⋅ Yuhui Yuan ⋅ Shixia Liu
Image-to-layer decomposition converts a flattened image into editable RGBA layers, enabling element-level editing in design workflows. Existing diffusion-based systems typically adapt large pretrained text-to-image models and introduce RGBA autoencoders or variable-layer architectural modules. We revisit this design choice and ask whether layer decomposition actually requires latent autoencoding or T2I pretraining. We introduce PixelART, a pixel-space rectified-flow Transformer trained from scratch for image-to-layer decomposition. PixelART directly denoises regional RGBA pixel patches with a single-stream multi-modal diffusion Transformer, avoiding RGBA-VAEs, pretrained T2I backbones, and layer-specific decoders. We identify a key property of the task: high-noise timesteps determine layer assignment and coarse layer organization, while low-noise timesteps mainly refine color, alpha, texture, and boundaries. Based on this observation, we propose a \emph{terminal-boosted timestep sampling} to increase training coverage in the high-noise assignment regime. Trained on 4M multi-layer design templates up to (1024\times1024), PixelART achieves state-of-the-art layer and composite reconstruction on \designbenchmark while using over (10\times) fewer parameters and substantially lower latency and memory than VAE-based diffusion baselines. Ablations show that pixel-space \mathcao{x}-prediction, high-noise timestep coverage, and data/model scaling are critical, while T2I initialization provides no measurable final gain in the data-rich I2L setting.
Multi-modal large language models (MLLMs) have shown impressive generalization across tasks using images and text modalities. While their extension to video has enabled tasks such as video question answering and video captioning, their pixel-level visual grounding abilities are less studied. In this work, we raise the pertinent question of whether motion is used in pixel-level visual grounding and whether video MLLMs can segment objects based on natural language expressions describing their motion patterns. We identify the shortcomings in the current benchmarks, where we show that a single frame can often suffice for capturing the referring expression without any temporal reasoning. To address this, we introduce novel motion-centric probing, particularly designed for the visual grounding task, to study video MLLMs' ability to identify true motion from a fake one and their ability to grasp the ordering of motion. Consequently, we introduce MoCentric-Bench, a motion-centric benchmark designed to evaluate video MLLMs based on their ability to capture the interaction between motion and language, rather than relying primarily on static cues. We further establish strong single-image baselines that either match or surpass prior methods. Finally, we explore a simple motion-centric adaptation that provides state-of-the-art performance on our MoCentric-Bench. Code is available at \url{https://anonymous.4open.science/r/PixFoundation2MoCentricBenchFin/}.
Planning as Dynamics Relaxation: Hippocampal Recurrent Network Realizes Optimal Goal-Directed Navigation
宇航 他 ⋅ Junfeng Zuo ⋅ Tianhao Chu ⋅ Si Wu
Neural correlates of spatial cognitive map are well documented, yet exactly how neural circuits perform spatial navigation in complex environments — e.g., reaching a goal while avoiding obstacles — remains largely unclear. Here, we show that a hippocampal network with appropriate recurrent connections can naturally achieve optimal goal-directed navigation via its relaxation dynamics. Specifically, we consider that the recurrent weights between the neurons represent the transition probabilities between spatial locations encoded by neurons; obstacles such as walls and blocked corridors are therefore reflected by the vanishing of connection weights. This connection pattern can be learned in the hippocampus via behavioral-timescale synaptic plasticity (BTSP) while the animal is exploring the environment. When a goal signal is presented, the network dynamics will relax into an activity field representing the goal location. We prove that this field is mathematically equivalent to the desirability field of a Linearly-solvable Markov Decision Process (LMDP), and the local log-gradient of the field indicates the navigation direction. Both theoretical analyses and simulations demonstrate that this recurrent network dynamics-mediated navigation is efficient and robust in environments with complex obstacle layouts. Moreover, only low-rank updates of the network's connection pattern are needed when the environment has local changes. We hope this study offers insight into a general circuit principle for planning in abstract rational maps in the brain beyond spatial navigation.
Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry
Pablo Marcos Manchón ⋅ Rishi Jha ⋅ Lluís Fuentemilla
The Strong Platonic Representation Hypothesis suggests that representational convergence in artificial neural networks can be harnessed constructively: embeddings can be translated across models through a universal latent space without paired data. We ask whether an analogous geometry can be recovered across human brains. Using fMRI data from the Natural Scenes Dataset, we propose a self-supervised encoder that learns subject-specific embeddings from brain data alone by exploiting repeated stimulus presentations. We show that these independently learned spaces can be translated across subjects using unsupervised orthogonal rotations, without paired cross-subject samples or intermediate model representations. Synchronizing pairwise rotations into a single shared latent space further improves cross-subject retrieval, indicating that subject-specific spaces are mutually compatible with a common coordinate system. These results provide evidence for a shared neural geometry in the human visual cortex: subject-specific fMRI representations are approximately isometric across individuals and can be translated through purely geometric transformations.
PMO-Dock: Benchmarking Docking, Specificity, and Generalization in Molecular Optimization
Gor Simonyan ⋅ Tatevik Abrahamyan ⋅ Narek Abelyan ⋅ Tigran Fahradyan ⋅ Hrant Khachatrian
The Practical Molecular Optimization (PMO) benchmark standardized evaluation in molecular optimization, but it is built on simple property-based oracles that do not capture structure-based drug design. As the field has moved to docking-based objectives, evaluation practice has become unstandardized, and different methods report results on different docking tasks, often without a shared protocol for fair comparison or sample-efficiency constraints. To address this, we introduce Practical Molecular Optimization for Docking (PMO-Dock), a benchmark and protocol for docking-based optimization consisting of 25 tasks covering hit generation, lead optimization, and a new specificity task requiring strong on-target binding while penalizing off-target interactions. The protocol separates model development from final evaluation by assigning validation tasks for hyperparameter tuning and strictly held-out test tasks for reporting, with no test-task use during development. It also enforces strict oracle budgets to reflect realistic drug-discovery settings. We benchmark four diverse high-performing methods spanning different optimization paradigms, Saturn (reinforcement learning), GenMol (discrete diffusion), Genetic-guided GFlowNet, and Chemlactica (LLM-based). Our analysis shows no universally best method across tasks, substantial differences in hyperparameter sensitivity, and method-dependent transferability, where some methods benefit from global hyperparameter selection while others require task-local tuning on sufficiently similar validation tasks. This benchmark provides a concrete evaluation standard for measuring progress on generalizable, sample-efficient molecular optimization.
Point Cloud Sequence Encoding for Material-conditioned Graph Network Simulators
Philipp Dahlinger ⋅ Balázs Gyenes ⋅ Niklas Freymuth ⋅ Luca Geminiani ⋅ Tobias Würth ⋅ Johannes Mitsch ⋅ Nadja Klein ⋅ Luise Kärger ⋅ Gerhard Neumann
Graph Network Simulators (GNS) have emerged as powerful surrogates for complex physics-based simulation, offering inherent differentiability and orders-of-magnitude speedups over traditional solvers. However, GNSs typically assume access to the underlying material parameters, such as stiffness or viscosity, severely limiting their utility in realistic experimental settings. While recent meta-learning approaches address the parameter dependency by inferring properties from mesh trajectories, reconstructing a mesh from an observed scene is difficult. In this work, we introduce Point Cloud Encoding for Accurate Context Handling (PEACH), a novel framework that applies in-context learning on point clouds to adapt a learned simulator to unseen physical properties during inference. Our approach relies on a novel spatio-temporal point cloud sequence encoder, as well as two forms of auxiliary supervision to help improve simulation fidelity. We demonstrate that PEACH is capable of accurate zero-shot sim-to-real transfer on a challenging, dynamic scene. Experiments on simulation scenes show that PEACH even outperforms mesh-based baselines on simulation accuracy, while being much more practical for real-world deployment.
PolySplat: Workload-Regime-Aware Rasterization for 3D Gaussian Splatting
Longzan Luo ⋅ Bin CUI ⋅ Xupeng Miao
Rasterization dictates the interactive budget of 3D Gaussian Splatting (3DGS). However, the comparative speed of modern CUDA rasterizers is typically evaluated on a narrow canonical benchmark, masking severe regime-dependent performance reversals. This limited scope hides three critical kernel-level bottlenecks: warp lane underutilization, exposed global-memory latency during pixel shading, and terminal-tail penalties from hardware thread scheduling. In this paper, we present PolySplat, a workload-regime-aware 3DGS rasterizer designed to overcome these inefficiencies. PolySplat introduces a warp-saturating tile-key emitter with adaptive three-way dispatch, asynchronous shared-memory staging in the render kernel, and a persistent kernel architecture with centralized atomic dispatch to strictly bound tail latency. To rigorously validate our system, we introduce an extended 76-target benchmark that achieves comprehensive workload-regime coverage by stratifying targets across extreme Gaussian counts (up to 56.5M), high resolutions (up to 9K), and six diverse scene categories. Evaluated on this comprehensive suite, PolySplat achieves dataset-balanced geometric-mean speedups of 1.26x to 6.48x over state-of-the-art rasterizers at lossless visual quality. Notably, under sustained 60Hz interaction at extreme resolutions, PolySplat strictly meets 1-vsync display deadlines where existing renderers suffer from massive queue divergence and input-to-photon lag.
PoSafeNet: Structured Safety Learning via Compositional Projection
Kiwan Wong ⋅ Wei Xiao ⋅ Daniela Rus
Safe robot learning often involves multiple heterogeneous safety constraints that cannot always be satisfied simultaneously. Existing neural safety layers typically treat multi-constraint safety as a numerical optimization problem, enforcing all constraints through a single QP-based projection or relaxing conflicts with slack variables. This hides the semantic question of which constraints may be sacrificed under conflict inside solver geometry, penalty weights, or a fixed total hierarchy. We propose PoSafeNet, a poset-structured composable safety layer that makes these conflict semantics explicit. PoSafeNet encodes admissible safety override relations as a partial order and realizes each admissible execution by composing closed-form projections onto CBF-induced halfspaces. Across multi-obstacle navigation, constrained manipulation, and vision-based autonomous driving, PoSafeNet improves operational feasibility, computational efficiency, and task performance over dQP-based, slack-based, and hierarchical safety layers.
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
Sanghyeon Lee ⋅ Jinwoo Kim ⋅ Jong Taek Lee
Zero-shot skeleton-based action recognition (ZSSAR) is typically treated as a skeleton-text alignment problem: encode joint-coordinate sequences, align them with language, and classify unseen actions. We argue that this alignment is often too late. Skeletons are not complete action observations, but compressed outputs of human pose estimation (HPE); by the time alignment begins, human-object interactions and pose-relative visual cues may no longer be explicit. We call this upstream semantic loss. To address it, we propose PoseBridge, an HPE-aware ZSSAR framework that bridges intermediate HPE representations to skeleton-text alignment. Rather than adding an RGB action branch or object detector, PoseBridge extracts pose-anchored semantic cues from the same HPE process that produces skeletons, then transfers them through skeleton-conditioned bridging and semantic prototype adaptation. Across NTU-RGB+D 60/120, PKU-MMD, and Kinetics-200/400, PoseBridge improves ZSSAR performance under the evaluated protocols. On the Kinetics-200/400 PURLS benchmark, which contains in-the-wild videos with diverse scenes and action contexts, PoseBridge shows the clearest separation, improving the strongest compared baseline by 13.3-17.4 points across all eight splits. Our code will be publicly released.
PosePlaner: Denoising Any Feedforward Pose Predictor from Pairwise Planar Geometry
Yihang Chen ⋅ Bin Tan ⋅ Yujun Shen ⋅ Yinghao Xu ⋅ Nan Xue
Recent multi-view Transformers (e.g., DUSt3R and VGGT) have advanced 3D reconstruction by scaling training data and model capacity from two views to many views and long sequences, subsuming the classical SfM/SLAM pipeline of feature extraction, matching, and pose estimation in a single feedforward pass. Yet this subsumption hides the interfaces between pipeline stages---correspondences, geometry, and pose are entangled in a single forward pass, leaving supervised fine-tuning as the major path to improvement and offering no mechanism to correct predictions from geometric evidence at test time. We propose PosePlaner to close this gap by reviving the classical coupling between correspondences and pose via flow matching, treating pose refinement as a conditional generation problem that takes coarse geometric matches and initial pose as input condition and fuses the resulting pose hypotheses into a single refined estimate, providing a continuous, data-driven analogue of RANSAC voting. Our method is trained efficiently on two-view data, yet serves as a periodic refiner for streaming n-view pose estimators with no retraining. Empirically, it corrects failures of feedforward predictors, substantially improves correspondence-based solvers, and reduces accumulated drift in long streaming sequences.
Position: Semantic Uncertainty Measures Disagreement, Not Reliability
Joseph Hoche ⋅ Maxime Corlay ⋅ David Brellmann ⋅ Andrei Bursuc ⋅ Pavel Izmailov ⋅ Angela Yao ⋅ Gianni Franchi
Large Language Models and their multi-modal variants see increasingly rapid adoption and deployment, including in settings where reliability matters. However their uncertainty remains difficult to assess as classic or early token-level approaches cannot deal with the open-ended outputs of such models. Semantic uncertainty arises as a promising solution to the limits of token-level confidence by measuring disagreement among multiple generated responses. This position paper argues that this framing is incomplete: semantic uncertainty measures semantic disagreement, not actual reliability. A model can be uncertain while producing several valid answers, or certain while repeatedly producing the same wrong answer. We introduce a semantic bias–uncertainty decomposition to show that reliability depends both on variability across meanings and systematic deviation from correct or grounded meanings. This perspective reveals that common hallucination-detection evaluations conflate uncertainty with error. We argue for reliability-centered evaluation that separates semantic disagreement, correctness, and enables a more fine-grained characterization of uncertainty beyond the level of the full answer.
Post-ADC Inference: Valid Inference After Active Data Collection
Shuichi Nishino ⋅ Tomohiro Shiraishi ⋅ Teruyuki Katsuoka ⋅ Ichiro Takeuchi
The validity of statistical inference depends critically on how data are collected.When data gathered through active data collection (ADC) are reused for a post-hoc inferential task, conventional inference can fail because the sampling process is adaptively biased toward regions favored by the collection strategy. This issue is especially pronounced in black-box optimization, where sequential model-based optimization (SMBO) methods such as the tree-structured Parzen estimator (TPE) and Gaussian process upper confidence bound (GP-UCB) preferentially concentrate evaluations in promising regions. We study statistical inference on actively collected data when the inferential target is constructed in a data-dependent manner after data collection. To enable valid inference in this setting, we propose post-ADC inference, a framework that accounts for the biases arising from both the active data collection process and the subsequent data-driven target construction. Our method builds on selective inference and provides valid $p$-values and confidence intervals that correct for both sources of bias. The framework applies to a broad class of ADC processes by imposing only assumptions on the observation noise, without requiring any assumptions on the underlying black-box function or the surrogate model used by the SMBO algorithm. Empirical results also show that post-ADC inference provides valid inference for data collected by GP-UCB and TPE.
Posterior Inference in Latent Space for Scalable Constrained Black-box Optimization
Kiyoung Om ⋅ Kyuil Sim ⋅ Taeyoung Yun ⋅ Hyeongyu Kang ⋅ Jinkyoo Park
Optimizing high-dimensional black-box functions under black-box constraints is a pervasive task in a wide range of scientific and engineering problems. These problems are typically harder than unconstrained problems due to hard-to-find feasible regions. In this paper, we propose a new framework to solve high-dimensional, constrained black-box optimization problems via posterior inference in the latent space of generative models. Our method iterates through two stages. First, we train flow-based models to capture the data distribution and surrogate models that predict both function values and constraint violations. Second, we cast the candidate selection problem as a posterior inference problem to effectively search for promising candidates that have high objective values while not violating the constraints. Concretely, we utilize diffusion models to amortize the sampling from the posterior distribution in the latent space of flow-based models, which can bypass the issue of mode collapse. We empirically demonstrate that our method achieves superior performance across synthetic and real-world tasks. Our code is available here.
Post-Processing Guarantees for Classification under Linear-Fractional Performance Metrics
Andrea Della Vecchia
This paper studies binary classification under generalized performance metrics, going beyond standard misclassification error. We focus on linear-fractional measures—including the $F$-score and Jaccard index—which are widely used in class-imbalanced settings. We first provide a unified characterization of the optimal classifier, showing that it admits a thresholding structure on the regression function. This reduces the original infinite-dimensional, non-decomposable optimization problem to the estimation of a single scalar parameter. Building on this result, we analyze plug-in classifiers that combine a regression estimator with a data-driven estimate of the optimal threshold. We show that this threshold is characterized as the solution of a fixed-point equation depending only on the marginal distribution, enabling its estimation using unlabeled data and avoiding standard validation-based procedures. Our main contribution is a finite-sample analysis of this approach. We derive excess risk bounds that decompose the error into contributions from regression estimation, threshold estimation, and sampling effects, providing a modular understanding of learning under non-decomposable metrics. Under standard margin assumptions, we further establish fast convergence rates. Our results yield a unified theoretical framework for plug-in classification under linear-fractional metrics, extending prior analyses beyond specific measures such as the $F$-score. Empirical results on synthetic and real datasets support our theoretical findings.
Post-Selection-Safe Pessimistic Utilities for Offline Multi-Objective Reinforcement Learning
Mingxi Hu ⋅ Meiling Yu
We study post-selection-safe deployment in finite-horizon tabular offline multi-objective reinforcement learning with full vector rewards, nonnegative linear scalarization, and preferences revealed at deployment. A fixed dataset may already have been reused to construct a data-dependent policy library $\Pi_N$; the remaining problem is to return auditable lower-confidence utilities for future queried preferences and for small policy menus. We introduce the Pessimistic Utility Oracle (PUO), a coordinatewise pessimistic vector-value construction. On one dataset-level tabular model-confidence event, PUO lower-bounds every data-dependent policy under every $w\in\mathcal{W}\subset\mathbb{R}_+^d$, without preference discretization or an additional library-size union bound beyond the model event. PUO yields a queried-preference planner and SPS++, a greedy policy-set method. The main SPS++ theorem evaluates the executable rule that deploys the in-set policy with largest PUO score, and decomposes its loss into coverage, greedy approximation, and preference-sampling terms. Controlled tabular audits validate these statistical effects; trajectory-count and linear-feature appendices are presented only as bridges beyond the main generative tabular theorem.
Post-Training Quantization with Gradient-Projected Fisher Approximation for Vision Transformers
Jincheol Yang ⋅ Jaemin Choi ⋅ Nahyun Lim ⋅ Yun-Seong Jeong ⋅ Matti Zinke ⋅ Hyunwoo Yu ⋅ Bongjoon Hyun ⋅ Kyomin Sohn ⋅ Suk-Ju Kang
Vision Transformers (ViTs) demonstrate strong performance, but their substantial memory footprint and computational overhead necessitate efficient compression techniques such as post-training quantization (PTQ). However, pushing ViTs to low-bit precision incurs significant accuracy degradation. Recent PTQ methods adopt block reconstruction-based optimization with curvature-aware objectives, typically implemented using Fisher-based approximations. These approaches rely on explicit curvature modeling based on diagonal or structured approximations, which still fail to capture cross-dimensional interactions. Moreover, they often involve matrix inversion, leading to numerical instability under ill-conditioned settings. To address these limitations, we propose Gradient-Projected Fisher Approximation for Quantization (GPFA-Q), a block reconstruction-based PTQ framework that avoids explicit curvature matrix construction while capturing off-diagonal interactions. First, we introduce Gradient-Projected Reconstruction (GPR), a reconstruction objective that captures cross-dimensional interactions without explicitly constructing curvature matrices. To further support GPR, we integrate Soft Grid Rounding (SGR), which reduces the mismatch between continuous reconstruction and discrete inference. Extensive experiments demonstrate that our GPFA-Q achieves the state-of-the-art performance in low-bit quantization across diverse vision tasks.
Predicting Quantization Price for Selecting PTQ Configurations Before Deployment
Junbin Qiu ⋅ Jian Mu ⋅ Weitong Zhang ⋅ Yao SHU
Weight-space post-training quantization (PTQ) must choose finite formats, granularities, quantizer families, transformations, and bits before the completed quantized model reveals its output-distribution drift. Existing PTQ methods predict important pieces of this degradation, including reconstruction error, Hessian sensitivity, transformation effects, and downstream loss, but these pieces are usually scored after fixing the quantization geometry or inside separate configuration families. We formulate weight-space PTQ as pre-deployment configuration selection using priced layer-output error. Each admissible layer configuration is treated as an error generator with a deployment cost, which induces a layer-output error covariance $\boldsymbol{\Sigma}_l(\alpha_l)$, and the full-precision model prices that covariance by downstream curvature, $\hat{\rho}_l(\alpha_l)=\frac{1}{2}\operatorname{Tr}\left(\hat{\mathbf{H}}_l\,\hat{\boldsymbol{\Sigma}}_l(\alpha_l)\right)$. The price follows from full-precision-to-quantized forward KL, whose first-order term cancels at the reference model. It turns reconstruction and diagonal scores into reduced proxies that drop price factors, while finite formats, codebooks, granularities, and equivalent transformations become comparable candidates through the covariances they induce and the costs they pay. A trace reduction then yields a calibration-time price table and a budgeted price-guided selector, making fixed-geometry bit allocation a special case rather than the organizing problem.
Predictive Surprise as Self-Grounding Concept Bottleneck for Interpretable Time Series
Sachith Abeywickrama ⋅ Emadeldeen Eldele ⋅ Min Wu ⋅ Xiaoli Li ⋅ Chau Yuen
Time-series classifiers deployed in risk-sensitive domains such as healthcare, wearable sensing, and industrial monitoring must be accurate, interpretable, and capable of abstaining when evidence is weak. No existing method delivers all three, and the obstacle is structural, rather than logistical. Meaningful temporal patterns cannot be defined without knowing segment boundaries, yet meaningful boundaries cannot be placed without knowing which patterns to expect. Existing approaches break this circularity through fixed windows or expert-annotated vocabularies, sacrificing at least one of the three properties. We introduce ConceptTime, which closes the loop through a single self-supervised signal. A frozen probabilistic forecaster predicts a distribution over the next observation at every timestep of the sequence. We call the mismatch between its prediction and the realized signal, the predictive surprise. This signal both locates segment boundaries and characterizes each resulting segment through the forecaster's own predictive statistics. Segments are thus grounded by \emph{how} they behaved relative to expectation, not by what they look like. These summaries cluster into a vocabulary of dynamical regimes that we call concepts. The vocabulary distinguishes calm-after-spike from calm-after-calm, and unifies visually distinct segments that violate expectation in the same way. A lightweight head classifies from the concept sequence, and the distance to the nearest concept yields per-input reliability for free. Across seventeen UEA, human-activity, biomedical, and fault-diagnosis benchmarks, ConceptTime matches state-of-the-art black-box accuracy, beats every interpretable baseline, and retains near 98% accuracy on Epilepsy with 1% of labels. It equals or exceeds softmax, prototype-distance, and SHAP/TimeX/TimeX++ on OOD detection and deletion-faithfulness, without a single annotation, language-model call, or domain expert.
Pref-DetectGPT: Unveiling Machine-Generated Text via Preference-Aware Curvature Measurement
Jiahao Wang ⋅ Feifei Kou ⋅ Zhongbao Zhang ⋅ Jiwei Zhang ⋅ Lei Shi ⋅ Pengfei Zhang ⋅ Suguo Zhu ⋅ Mingying Xu
The widespread deployment of large language models (LLMs) has made reliable detection of machine-generated text increasingly critical. Recent curvature-based detectors introduce perturbations to candidate text and compute curvature as the log probability difference between original and perturbed variants using advanced surrogate models. However, such curvature measurements only capture weak discrepancies between machine-generated and human-written text. We propose $\textbf{Pref-DetectGPT}$, a training-free method that introduces preference-aware curvature measurement by leveraging the implicit preference capabilities of preference-optimized surrogate models. Specifically, we construct an implicit reward model from the preference-optimized policy to score both original and perturbed text, and redefine curvature as their normalized deviation, which serves as a significantly stronger signal for distinguishing machine-generated text from human-written text. Empirical evaluations across multiple public benchmark datasets demonstrate that Pref-DetectGPT achieves state-of-the-art detection performance with relative improvements of 3.49\%, 3.13\%, and 10.66\% in AUROC, AUPR, and TPR5\%, while exhibiting strong robustness against various adversarial attacks and input lengths.
Prefix Executability: Evaluating Tool-Using Agents Beyond Final Success
Amir H Rezaeian ⋅ Weiyi Sun ⋅ Yassine Benajiba ⋅ Dan Roth
Tool-use benchmarks often score agents by final task success, but terminal outcomes can obscure whether the intermediate tool-call trajectory was executable. We propose prefix executability as a trajectory-level evaluation protocol for tool-using agents. The protocol measures whether every prefix of a predicted tool-call sequence remains executable under tool schemas, state preconditions, and data dependencies, and summarizes reliability using Prefix Executability Curves (PEC), First-Error Depth (FED), Trajectory Executability Rate (TER), horizon-normalized AUC-PEC, and Recovery Rate (RR). We instantiate the protocol on stratified samples from BFCL-v3 multi-turn and ToolSandbox. Across both benchmarks, outcome-only and prefix-level rankings diverge: agents whose execution is conditioned on an unverified plan often achieve higher point-estimate final success than non-plan-conditioned baselines while substantially lowering TER and AUC-PEC. These divergences are precisely the failure modes that prefix executability is designed to expose: they reveal whether apparent success rests on an executable trajectory, where invalidity first enters, and whether success follows recovery from an earlier failure. As a validation intervention, explicit pre-execution plan checking reverses much of this prefix-reliability loss, consistently improving executable-prefix survival while maintaining competitive terminal success. These results show that final success alone is insufficient for evaluating tool-use agents and that prefix-level executability provides an actionable reliability signal for comparing agents, diagnosing failures, and assessing validation mechanisms.
PreFT: Prefill-only finetuning for inference efficiency
Andrew Lanpouthakoun ⋅ Aryaman Arora ⋅ Zhengxuan Wu ⋅ Dhruv Pai ⋅ Benjamin Keigwin ⋅ Dan Jurafsky ⋅ Chris Potts
Large language models can now be personalised efficiently at scale using parameter efficient fine-tuning methods (PEFTs), but \emph{serving} user-specific PEFTs harms throughput, even with specialised kernels and memory management techniques. This is because, theoretically and empirically, a mismatch exists between prefill (processing a large number of tokens at once) and decode (generating a single token autoregressively): the latter has far lower throughput when serving multiple adapters. Rather than optimising performance relative to parameter count, for efficient multi-adapter serving, we instead ought to optimise performance relative to \textit{serving throughput}. We therefore propose \textbf{\texttt{PreFT}} (Prefill-only Finetuning), wherein we only apply the adapter to prefill tokens and discard it afterwards. \texttt{PreFT} significantly increases throughput with minimal effect on performance. We develop and release an efficient implementation of two prefill-only PEFTs, LoRA and ReFT, on the vLLM inference engine. We first show that serving multi-user \texttt{PreFT}s is vastly more efficient than traditional PEFTs ($1.90\times$ the throughput when serving $512$ adapters on Llama 3.1 70B). Then, we compare the performance of prefill-only vs.~all-token adapters on a variety of supervised finetuning and reinforcement learning tasks with LMs at varying scales. On SFT, we observe that the evaluation loss of \texttt{PreFT}s is higher, but can be compensated by increasing rank with nearly no reduction in throughput. On RL, we consistently find that \texttt{PreFT}s approach parity with standard PEFTs. Together, this work validates prefill-only adaptation of LLMs as a more favourable accuracy--throughput tradeoff than existing PEFTs for personalised serving.
Preserving DEG Rankings for Gene Discovery in Histology-Based Spatial Gene Expression Prediction
Kaito Shiku ⋅ Kazuya Nishimura ⋅ Yasuhiro Kojima ⋅ Ryoma Bise
Predicting spatial gene expression from histology images could scale spatial transcriptomics (ST) to image-only cohorts, but conventional histology-based ST prediction is trained and evaluated mainly by per-gene spatial-profile reconstruction. This objective is misaligned with a key downstream use of ST: differentially expressed gene (DEG) discovery, where genes are ranked for a biological or morphology-defined contrast by evidence of between-group expression differences. We formulate image-based differential expression ranking (IDER), which asks whether predicted expression profiles preserve the contrast-specific ranked gene list obtained from measured profiles. IDER compares gene rankings induced by differential-expression statistics, rather than raw expression magnitudes or per-gene spatial correlations. We further introduce a differentiable IDER objective that aligns these statistics across genes and can be trained with morphology-derived proxy contrasts without predefined biological group labels. Experiments on public ST datasets show improved DEG-ranking agreement and pathway-enrichment overlap over conventional reconstruction objectives, including morphology-derived and pathologist-annotated tissue-region evaluations.
PriorVLA: Prior-Preserving Adaptation for Vision-Language-Action Models
Xinyu Guo ⋅ Bin Xie ⋅ Wei Chai ⋅ Xianchi Deng ⋅ Tiancai Wang ⋅ Zhengxing Wu ⋅ Xingyu Chen
Large-scale pretraining has made Vision-Language-Action (VLA) models promising foundations for generalist robot manipulation, yet adapting them to downstream tasks remains necessary. However, the common practice of full fine-tuning treats pretraining as initialization and can shift broad priors toward narrow training-distribution patterns. We propose PriorVLA, a novel framework that preserves pretrained priors and learns to leverage them for effective adaptation. PriorVLA keeps a frozen Prior Expert as a read-only prior source and trains an Adaptation Expert for downstream specialization. Expert Queries capture scene priors from the pretrained VLM and motor priors from the Prior Expert, integrating both into the Adaptation Expert to guide adaptation. Together, PriorVLA updates only 25% of the parameters updated by full fine-tuning. Across RoboTwin 2.0, LIBERO, and real-world tasks, PriorVLA achieves stronger overall performance than full fine-tuning and state-of-the-art VLA baselines, with the largest gains under out-of-distribution (OOD) and few-shot settings. PriorVLA improves over $\pi_{0.5}$ by 11 points on RoboTwin 2.0-Hard and achieves 99.1% average success on LIBERO. Across eight real-world tasks and two embodiments, PriorVLA reaches 81% in-distribution (ID) and 57% OOD success with standard data. With only 10 demonstrations per task, PriorVLA reaches 48% ID and 32% OOD success, surpassing $\pi_{0.5}$ by 24 and 22 points, respectively.
PRISM: Programming Interactive Scenes from Monocular Images for Embodied Simulation
Yumeng He ⋅ Yichen Song ⋅ Xiaotian Yang ⋅ Weijia Zhang ⋅ Zanwei Zhou ⋅ 俊儒 宫 ⋅ Xiaokang Yang ⋅ Yunbo Wang
The advancement of Embodied AI necessitates high-quality simulation assets that faithfully mirror the physical world. However, transforming raw visual observations into simulation-ready scenes remains challenging due to the lack of physical grounding and scene-level interactivity in current image-to-URDF methods. We propose PRISM, a framework that reformulates monocular scene reconstruction as a procedural programming task for interactive 3D environments. By leveraging the zero-shot reasoning and code synthesis of MLLMs, PRISM translates a single RGB image into executable programs defining object geometry, articulation, and physical properties. To ensure simulation readiness, it incorporates a physics-in-the-loop mechanism that iteratively refines the generated programs by validating their execution in a physics engine. This feedback loop enforces physically plausible object articulations, valid object compositions and interactions, and accurate spatial relationships. Experiments demonstrate that PRISM significantly outperforms open-loop approaches and prior monocular reconstruction models. Notably, PRISM-generated scenes support complex downstream tasks such as stable stacking and fine-grained manipulation, which are difficult to achieve with existing methods.
Large language models (LLMs) are increasingly used to simulate human behavior, but their ability to simulate individual privacy decisions faithfully is not well understood. In this paper, we address the problem of evaluating whether a core set of user persona attributes can drive LLMs to simulate individual-level privacy behavior. We introduce PrivacySIM, an evaluation suite that benchmarks LLM simulation of user privacy behavior against the ground-truth responses of 1,000 users. These users are drawn from five published user studies on privacy spanning LLM healthcare consultations, conversational agents, and chatbots. Drawing on these user studies, we hypothesize three persona facets as plausible predictors of privacy decision-making: demographics, previous experiences with AI, and stated privacy attitudes. We condition nine frontier LLMs on subsets of these three facets and measure how often each model's response to a data-sharing scenario matches the user's actual response. Our findings show that (1) privacy persona conditioning consistently improves simulation quality over no-persona conditioning, but even the strongest model (40.4% accuracy) remains far from faithfully simulating individual privacy decisions. (2) A user's stated privacy attitudes alone may not be the best predictor as they often diverge from the user's actual privacy behavior. (3) Users with high AI experience but low stated privacy attitudes are the most challenging to simulate. PrivacySIM is a first step toward understanding and improving the capabilities of LLMs to simulate user privacy decisions.
PrivateSeal: Low-Sensitivity Latent Directions for Diffusion-Resilient User-Specific Watermarking
Yiheng Chen ⋅ Zhong Ji ⋅ Zhihao Li ⋅ Boyu Wang
Diffusion-driven image editing and regeneration are becoming increasingly widespread. This creates an urgent need for watermarking methods that can embed imperceptible yet verifiable signals into existing images for ownership attribution and provenance tracking. Existing methods fail to simultaneously satisfy three practical requirements: robustness against diffusion-based transformations, low-cost verification, and scalable support for per-user key assignment at low overhead. To address these challenges, we propose PrivateSeal, a watermarking framework for pre-existing images under diffusion-based transformations. The perturbation is encouraged to lie along low-sensitivity latent directions, so that the embedded signal is less likely to be suppressed during diffusion-based editing or regeneration. During verification, the embedded message is reliably recovered via a simple latent-space projection using the corresponding key. This design allows a platform to assign independent projection keys to different users, accounts, or images without retraining or modifying the verifier, while maintaining low-cost verification. Extensive experiments on the W-Bench benchmark show that PrivateSeal achieves competitive robustness against diffusion-based regeneration and editing, with additional cross-model and cross-dataset evaluations further validating its strong transferability.
Proactive Instance Navigation with Comparative Judgment for Ambiguous User Queries
Junhyuk Kwon ⋅ Seungjoon Lee ⋅ Hyejin Park ⋅ Kyle Min ⋅ Jungseul Ok
Natural-language instance navigation becomes challenging when the initial user request does not uniquely specify the target instance. A practical agent should reduce the user's burden by actively asking only the information needed to distinguish the intended target from similar distractors, rather than requiring a detailed description upfront. Existing approaches often fall short of this goal: they may stop at the first plausible candidate before sufficiently exploring alternatives, or, even after collecting multiple candidates, ask about the target's attributes derived from individual candidates rather than questions selected to distinguish candidates in the pool. As a result, despite the dialogue, the agent may still fail to distinguish the target from distractors, leading to premature decisions and lengthy user responses. We propose Proactive Instance Navigation with Comparative Judgment (ProCompNav), a two-stage framework that first constructs a candidate pool and then identifies the target through comparative judgment. At each round, ProCompNav extracts an attribute-value pair that splits the current pool, asks a binary yes/no question, and prunes all inconsistent candidates at once. This reframes disambiguation from open-ended target description to pool-level discriminative questioning, where each question is chosen to narrow the candidate set. On CoIN-Bench, ProCompNav improves Success Rate over interactive baselines with the same minimal input and non-interactive baselines with detailed descriptions, while substantially reducing Response Length. ProCompNav also achieves state-of-the-art Success Rate on TextNav, suggesting that comparative judgment is broadly useful for instance-level navigation among similar distractors.
Probe Before You Edit: Probing-Guided Molecular Optimization for LLM Agents in Structure-Based Drug Design
Zaifei Yang ⋅ Weiyu Chen ⋅ Yaqing Wang ⋅ James Kwok
Structure-based drug design increasingly employs LLM agents to iteratively refine ligands against a target pocket, yet a viable ligand must satisfy two often-conflicting objectives---binding affinity and druggability---which single optimization steps rarely improve together. To quantify this difficulty, we introduce two diagnostic metrics: the first measures how often a single edit improves both objectives, and the second measures how often a gain on one objective comes with a loss on the other. Applying these diagnostics to current LLM-agent pipelines exposes a consistent failure mode: the agent performs molecular editing without knowing how the pocket-ligand complex responds to local modifications, thus rarely achieving joint improvement. Inspired by medicinal chemists, who probe the pocket-ligand complex with controlled analog edits before choosing an optimization direction, we propose PROBE, an optimization framework built around edit--response probing. PROBE first decomposes the ligand into editable sites and builds a pocket-specific site map that flags where joint gains are plausible, where the two objectives are likely in tension, and where liability substructures should be changed; it then performs controlled probe edits whose responses are distilled into an EditManual. Guided by the site map and EditManual, PROBE runs an iterative multi-agent loop in which an affinity agent, a druggability agent, and a co-optimization agent jointly produce edits. On the CrossDocked2020 benchmark, PROBE achieves state-of-the-art performance and substantially mitigates the failure modes exposed by our diagnostics metrics.
Process-Aware LNS for Large-Scale MILP via Context-Enhanced Fine-Tuned LLM-driven Selector
An Yan ⋅ Huigen Ye ⋅ Hua Xu ⋅ Jiahao Zhang
Solving large-scale Mixed-Integer Linear Programming (MILP) problems involving millions of variables is critical for industrial applications but notoriously intractable due to combinatorial explosion. While Large Neighborhood Search (LNS) has emerged as the premier heuristic strategy for tackling such scales, its efficacy hinges entirely on the underlying neighborhood selection mechanism. Recent Large Language Model (LLM)-driven LNS frameworks show remarkable promise but suffer from fundamental flaws: they lack generalizability across diverse problem classes and rely on static selection paradigms that remain entirely blind to the shifting dynamics of the iterative optimization process. To overcome this rigidity, we propose PALL, a Process-Aware Evolutionary LNS framework guided by a Context-Enhanced Fine-Tuned LLM-driven Selector. Instead of deploying a monolithic operator, PALL leverages an offline-evolved diverse operator ensemble and dynamically synchronizes the search strategy with evolving optimization states. Specifically, we fine-tune an 8B-parameter LLM on high-quality hindsight oracle trajectories to adaptively perform optimal operator selection. To bolster the LLM's sequential decision-making, we introduce a context-enhanced reasoning mechanism that leverages historical operator trajectories and optimization rewards as inferential feedback. Extensive evaluations on standard million-variable MILP benchmarks demonstrate that PALL achieves state-of-the-art performance. Supported by an asynchronous deployment strategy, it not only consistently outperforms commercial solvers like Gurobi and classical LNS algorithms but also delivers over a $6\times$ acceleration compared to state-of-the-art LLM-driven LNS frameworks.
PROLA: Principal-Orthogonal Low-rank Adaptation for Predictive Spatiotemporal Weather Downscaling
Minseo Yoon ⋅ Minseong Bae ⋅ Sojin Lee ⋅ Hyunwoo J. Kim
High-resolution weather prediction is important for resolving local atmospheric patterns missed by coarse global forecasts. Large pretrained weather backbones have improved global forecasting, but adapting them to produce future high-resolution fields from recent coarse states remains challenging and costly. We formulate this setting as Predictive Spatiotemporal Weather Downscaling (\textbf{PSWD}), where the target is a future high-resolution trajectory rather than a same-time refined field. To make this structure explicit, we propose a framework that decomposes prediction into spatial and temporal stages using a shared pretrained backbone, fixed numerical scaffolds, and lightweight residual heads. For backbone adaptation, we introduce \textbf{PROLA} (PRincipal-Orthogonal Low-rank Adaptation), which splits a fixed low-rank budget between pretrained principal directions and their orthogonal complement. PROLA further rescales rank-1 optimizer updates using gradient signals, improving adaptation without increasing trainable rank. Experiments in the PSWD setting show that PROLA outperforms representative low-rank adaptation baselines under matched trainable-parameter budgets. These results support PSWD as a challenging setting for weather backbone adaptation and PROLA as an effective method for this setting.
Prompt-Conditioned Semantic Bottleneck for Cross-Domain Face Attack Detection
Weihang Wang ⋅ Rongjie Liu ⋅ Min Cao
Learning visual representations that generalize across domains remains challenging when source-domain supervision is entangled with spurious dataset-specific cues. This issue is particularly pronounced in face attack detection (FAD), where attack artifacts vary substantially across capture devices, manipulation pipelines, and image quality. In this paper, we propose Prompt-Conditioned Semantic Bottleneck, an information-bottleneck-inspired representation learning framework for cross-domain FAD. Unlike prior methods that use prompts as image-text matching anchors for multi-modal feature learning, our method uses label- and attack-type-conditioned prompts to elicit attack-aware semantic targets from a frozen VLM. These prompt-conditioned semantics provide offline supervision for a lightweight visual detector, encouraging its representation to preserve attack-relevant cues while reducing reliance on source-domain nuisance factors. Extensive experiments on face anti-spoofing and face forgery detection benchmarks demonstrate consistent improvements under cross-dataset evaluation. Ablation studies further show that the gains are not simply due to stronger visual or multi-modal models, but arise from attack-aware prompt-conditioned semantic supervision. Importantly, the VLM and prompts are removed during inference, enabling efficient vision-only deployment. Overall, our results suggest that prompt-conditioned VLM semantics provide an effective way to improve the trade-off among detection accuracy, cross-domain generalization, and inference efficiency in FAD.
Proper Scoring Rules for Agentic Uncertainty Quantification
Suresh Raghu ⋅ Satwik Pandey ⋅ Shashwat Pandey
Language-model agents increasingly emit uncertainty signals throughout a trajectory, but existing agentic UQ evaluations often conflate ranking usefulness with probabilistic truthfulness. AUROC, AUPRC, risk-coverage, Trajectory ECE, and scalarized trajectory scores evaluate discrimination, binwise calibration, or collapsed summaries, but do not strictly elicit the full prefix-conditioned success-probability trace $q_t=\mathbb{P}^{\pi}(Y{=}1\mid\mathcal{H}_t)$. Building on prequential proper scoring, we introduce the Trajectory Proper Score (TPS), a predictor-agnostic family of strictly proper trajectory-level scoring rules for any per-step uncertainty signal calibrated into a probability of eventual success. We prove that TPS strictly elicits the success-probability process under complete observation, within the chosen score family and weight schedule. We extend the construction to administratively censored trajectories by projecting the complete-data score onto the observable stopped prefix, yielding an exact $q_Z$-weighted reduced score and a tractable approximation when $q_Z$ is unestimated. We further show that common trajectory evaluators target weaker objects than the full prefix-conditioned probability process: Trajectory ECE is resolution-blind, while scalarized Trajectory Brier elicits only the collapsed scalar, not the full trace. Experiments on StrategyQA, Tau2-Bench, HotpotQA, and WebShop show that these theoretical distinctions are operationally visible: probability recalibration can substantially change TPS while leaving rank metrics nearly unchanged, and the tractable censored approximation can change the verdict relative to complete-only evaluation.
Prospective Coding Improves Learning in Deep Continuous-Time Recurrent Networks
Shivang Rawat ⋅ Mirko Morello ⋅ Flaviano Morone ⋅ David Heeger
Temporal integration gives continuous-time recurrent networks memory, but in deep stacks it also delays bottom-up signals and attenuates top-down errors. We ask whether this failure mode can be addressed by making the bottom-up input to each layer prospective, replacing instantaneous layer inputs with local look-ahead signals that compensate for integration-induced lag. We develop Recursive Quadrature Filters (RQFs), a biologically motivated class of complex-valued temporal filters that remain equivalent to single-channel diagonal state-space models (SSMs). This equivalence makes RQF layers trainable with the same parallel scan and convolutional algorithms used for diagonal SSMs. At the implementation level, this prospective input amounts to a lightweight two-tap input update, making it a drop-in modification for RQFs and SSMs. For the resulting discrete-time network, we prove that spatial-only backpropagation in a deep RQF network with instantaneous bottom-up inputs yields error signals that decay geometrically with depth, whereas prospective-input coding restores order-one gradient flow. We evaluate prospective-input coding on the Speech Commands dataset using standalone RQF recurrent stacks across two feature representations, mel-frequency cepstral coefficients (MFCCs) and raw audio, and local and non-local credit-assignment strategies. Under spatial-only backpropagation, prospective-input coding improves validation accuracy from 65.7% to 84.9% on MFCC features and from 46.4% to 61.0% on raw audio. Under full backpropagation through time, the gains persist, improving accuracy from 89.4% to 93.2% on MFCC features and from 80.1% to 82.7% on raw audio. Taken together, our results identify prospective-input coding as a local, broadly applicable mechanism for improving credit assignment in multi-layer recurrent networks with a continuous-time substrate.
ProxyPose: 6-DoF Pose Tracking via Video-to-Video Translation
Ruihang Zhang ⋅ Felix Taubner ⋅ Pooja Ravi ⋅ Kyros Kutulakos ⋅ David Lindell
Tracking the six-degree-of-freedom (6-DoF) pose of objects and surfaces from monocular video is a long-standing problem in computer vision. To tackle this problem, existing methods require inputs beyond the video itself—such as 3D models, depth maps, object masks, or task-specific learned features—and they struggle with textureless, transparent, reflective, or deformable surfaces. Here, we introduce ProxyPose, which recasts 6-DoF pose tracking as video-to-video translation. Given only a video and a single marked pixel in the first frame, a fine-tuned video diffusion model translates the input into a "proxy video"—a synthetic video depicting a colored polyhedron undergoing the same local rigid-body motion as the surface region at the marked pixel. Because the proxy's geometry and appearance are known by construction, recovering its full 6-DoF trajectory reduces to classical pose estimation with off-the-shelf solvers. This formulation leverages large-scale video pre-training to absorb the hardest aspects of pose tracking—handling challenging materials, occlusions, and deformations—into the translation step, while operating at the pixel level with no assumptions about object identity, boundaries, or global rigidity. ProxyPose achieves state-of-the-art 6-DoF pose tracking accuracy without the additional inputs required by competing methods and after fine-tuning the video model only on synthetic data. We further demonstrate that ProxyPose extends to face tracking, camera pose estimation, and challenging in-the-wild scenes that are beyond the reach of existing approaches. Video results are available on the Supplemental Webpage.
Prysma: Efficient Modality Adaptation for SLO-aware LLM-based Video Question Answering
Botao Zhu ⋅ Xiaoyi Fan ⋅ Yifei Zhu
Large language models (LLMs) have demonstrated remarkable performance in video question answering (VQA). To improve responsiveness and answer quality, modern LLM-based VQA services pre-extract multiple LLM-compatible modalities from videos before response generation. However, the performance impact of modality combinations (MCs) and extraction knob settings has barely been studied. Our analysis shows that both factors substantially affect answer quality and Time-to-First-Token (TTFT) latency. Furthermore, the best configuration depends on both video content and query semantics, necessitating adaptive strategies. We propose Prysma, the first modality adaptation framework for LLM-based VQA. Prysma intelligently adapts the modality configurations to optimize answer quality, subject to a Service Level Objective (SLO) on average TTFT. It combines an offline modality orchestrator and an online semantic adapter. The offline orchestrator constructs a video knowledge base to warm-start tuning, employs a latency-aware configuration pruner to reduce search space, and runs a two-stage tuner to optimize both video- and query-level configurations. The online adapter selects configurations in real time based on query semantics and tuning history. Extensive evaluations demonstrate that Prysma improves answer quality by up to 82% over SOTA methods under the same SLO setting.
Q-ARVD: Quantizing Autoregressive Video Diffusion Models
Siao Tang ⋅ Xinyin Ma ⋅ Gongfan Fang ⋅ Xingyi Yang ⋅ Xinchao Wang
Autoregressive video diffusion models (ARVDs) have emerged as a promising architecture for streaming video generation, paving the way for real-time interactive video generation and world modeling. Despite their potential, the substantial inference cost of ARVDs remains a major obstacle to practical deployment, making model quantization a natural direction for improving efficiency. However, quantization for ARVDs remains largely unexplored. Our empirical analysis shows that directly applying existing quantization schemes developed for standard diffusion transformers to ARVDs leads to suboptimal performance, revealing quantization behaviors that differ from those observed in bidirectional diffusion models. In this paper, we identify two critical challenges in quantizing ARVDs: (C1) Highly unbalanced frame-wise quantization sensitivity. Error accumulation during autoregressive generation can induce severely skewed quantization sensitivity across frames, following an exponential-like decay pattern. (C2) Prominent and heterogeneous outlier patterns in weights. Weight distributions exhibit pronounced outlier channels, whose patterns vary substantially across layer types and block depths. To address these issues, we propose Q-ARVD, a novel framework for accurate ARVD quantization. (S1) To tackle the highly unbalanced frame-wise sensitivity, Q-ARVD incorporates a final-quality aware frame-weighting mechanism into the quantization objective. (S2) To prevent heterogeneous outliers from degrading performance, Q-ARVD introduces an outlier-aware adaptive dual-scale quantization, which automatically detects the presence and quantity of outlier channels for an arbitrary layer, and isolates them to protect normal channels. Extensive experiments on state-of-the-art open-source ARVDs (i.e., self-forcing and causal-forcing) demonstrate the superiority of Q-ARVD. Practical deployment of INT8 model shows 1.30x speedup and 1.97x model size reduction.
QEC Model Zoo: Democratizing AI-enhanced Quantum Error Correction
Yuandou Wang ⋅ Maxim de Groot ⋅ Aadi Patwardhan ⋅ Floris Geerts ⋅ Rihan Hai
Quantum error correction (QEC) is essential for scalable quantum computing. AI-enhanced QEC is gaining increasing attention, with learned full decoders and pre-decoders being explored across code families. However, progress is difficult to assess: existing studies use different datasets, input representations, decoder roles, preprocessing pipelines, and evaluation protocols, making it unclear whether a reported gain comes from the model, the data interface, or the benchmark setup. We present the QEC model zoo, a data-centric benchmark and model-zoo framework for AI-enhanced QEC. The QEC model zoo curates real and simulated datasets for repetition- and surface-code decoding, supports both full decoders and pre-decoders, and records the metadata needed for reproducible comparison. The framework provides common task definitions, model-ready processing pipelines, and evaluation interfaces for heterogeneous QEC data. It further enables systematic comparison across neural, tree-based, and modern tabular models. Our goal is not to advocate a single decoder architecture, but to make AI-enhanced QEC easier to compare, reproduce, and extend.
QuantDemoire: Quantization with Outlier Aware for Image Demoiréing
郑 陈 ⋅ Kewei Zhang ⋅ Xiaoyang Liu ⋅ Kai Liu ⋅ Weihang Zhang ⋅ Mengfan Wang ⋅ Yifan Fu ⋅ Yulun Zhang
Demoir\'eing aims to remove moir\'e artifacts that often occur in images. While recent deep learning-based methods have achieved promising results, they typically require substantial computational resources, limiting their deployment on edge devices. Model quantization offers a compelling solution. However, directly applying existing quantization methods to demoir\'eing models introduces severe performance degradation. The main reasons are distribution outliers and weakened representations in smooth regions. To address these issues, we propose QuantDemoire, a post-training quantization framework tailored to demoir\'eing. It contains two key components. **First**, we introduce an outlier-aware quantizer to reduce errors from outliers. It uses sampling-based range estimation to reduce activation outliers and keeps a few extreme weights in FP16 with negligible cost. **Second**, we design a frequency-aware calibration strategy. It emphasizes low- and mid-frequency components during fine-tuning, which mitigates banding artifacts caused by low-bit quantization. Experiments validate that our QuantDemoire achieves large reductions in parameters and computation while maintaining quality. Meanwhile, it outperforms existing quantization methods by over **4 dB** on W4A4. We implement a CUDA kernel for the mixed-precision branch and achieve a 2.3$\times$ end-to-end speedup on an RTX 4090 GPU. Code will be made public.
Quantum Best Arm Identification with Limited Round of Adaptivity: Lower Bounds and Algorithms
Haoran Li ⋅ Chen Wang ⋅ Xuchuang Wang
Best arm identification, also known as pure exploration, is a fundamental problem in the multi-armed bandit (MAB) model. Recent quantum algorithms achieve $O(n/\Delta_{[2]})$ query complexity for identifying the best arm with high constant probability, but all of them require $\Omega(\log(1/\Delta_{[2]}))$ or $\Omega(\log n)$ adaptive rounds. Motivated by the practical cost of adaptivity on near-term quantum hardware, where each round incurs substantial overhead from state re-preparation and classical--quantum communication, we study quantum best arm identification with no or very limited adaptivity. On the lower bound side, we prove that any non-adaptive quantum algorithm must make $\Omega(n\log n)$ queries when $\Delta_{[2]}=O(1)$, establishing a separation between adaptive and non-adaptive algorithms. To complement this lower bound, we present an algorithm that identifies the best arm with high constant probability using $O(n/\Delta_{[2]})$ queries and only $\log^*(n)+1$ adaptive rounds. En route to the lower bound, we show that the instances and techniques used to prove adaptivity lower bounds for classical bandits, such as those of Agarwal et al. [COLT'17], do not extend to the quantum setting. As such, we employ the polynomial method to prove our adaptivity lower bounds. To the best of our knowledge, this represents the first application of the polynomial method in this context, which may be of independent interest.
We present a method of quantum composite hypothesis testing with small error, which enables us to establish quantum lower bounds in nonparametric statistics and high-dimensional functional estimation. This is achieved by introducing trigonometric polynomials into the quantum polynomial method, thereby generalizing the quantum phase-estimation lower bound of Mande and de Wolf (ESA 2023). As applications, we settle the quantum complexities of several problems in property testing and functional estimation with small error: - For $\ell_2$-closeness testing, we show that the approach of Luo et al. (*IEEE Trans. Inf. Theory* 2024) is optimal. - For Tsallis entropy estimation, where $q=2$ corresponds to the Gini impurity, we show that the approaches of Buhrman et al. (*Phys. Rev. Lett.* 2001) and Ekert et al. (*Phys. Rev. Lett.* 2002) are optimal for integer $q \geq 2$, and that the approach of Chen et al. (ICALP 2026) is near-optimal for real $q \geq 1.5$. - For pure-state Uhlmann fidelity and trace distance estimation, we show that the approach of Wang (*IEEE Trans. Inf. Theory* 2024) is optimal.
Query Lower Bounds for Approximating the Top Eigenvector of Asymmetric Matrices
Kun Chen ⋅ Zhihua Zhang
We study the query complexity of approximating the top eigenvector of an asymmetric matrix in the matrix-vector product model. Our main result gives a gap-dependent lower bound for this problem in the adaptive query setting. Prior lower bounds for symmetric matrices already imply weaker hardness guarantees for the more general asymmetric problem, whereas our result yields a sharper dependence on the eigen-gap in the asymmetric case. In the inverse polynomial accuracy regime, this lower bound matches the benchmark upper bound of the power method up to lower-order factors. Our proof is based on an asymmetric spiked random matrix that serves as a hard instance. The key random-matrix ingredient is an analysis of the spike model, including both its asymptotic eigen-gap and the alignment between its top eigenvector and the planted spike. Building on an information-theoretic framework for query lower bounds, we then obtain the stated hardness result.
Radial-Angular Geometry for Reliable Update Diagnosis in Noisy-Label Learning
Ningkang Peng ⋅ Jingyang Mao ⋅ Xiaoqian Peng ⋅ Qu Weiguang ⋅ Yanhui Gu
Noisy-label methods often estimate sample reliability from forward-space signals such as loss, confidence, or entropy. These signals indicate whether a sample is difficult to predict, but they do not directly test whether its observed label induces a reliable parameter update. This gap matters because hard clean samples and mislabeled samples can have similar loss while inducing different updates. We recast reliability estimation as diagnosis of the observed-label update. The sample-wise empirical Fisher trace gives a backward-space measure of update energy: for the classifier layer, it factorizes into a prediction-residual term and a feature-sensitivity term, so it captures information beyond scalar loss. Trace, however, is still a radial magnitude signal and cannot decide whether a large update is useful or harmful. We therefore propose Relative Geometric Conflict (RGC), which compares the observed-label gradient with a reference gradient induced by an EMA teacher. The conflict term helps distinguish large but aligned hard-clean updates from large conflicting updates caused by corrupted labels. Across synthetic and real-world noisy-label benchmarks, RGC improves hard-clean preservation and accuracy under our evaluation protocol.
RAIL: Representation-Aligned Imitation Learning for Student-Compatible Teacher Policies
Meraj Mammadov ⋅ Pedro Zuidberg Dos Martires ⋅ Johannes A. Stork
Reinforcement learning (RL) from raw sensory inputs such as images or onboard sensors can be challenging due to high-dimensional observations and sparse rewards. A common strategy is to train a teacher policy with access to privileged state information and then distill it into a student policy that acts from raw inputs alone. However, when the teacher relies on information unavailable to the student, exact imitation may be impossible, creating an irreducible imitation gap. Existing approaches typically mitigate this mismatch through careful reward shaping or additional reinforcement learning on the student policy after the teacher has been trained. Instead, we introduce Representation-Aligned Imitation Learning (RAIL), which addresses the mismatch during teacher training itself. RAIL learns a latent representation shared across teacher and student observations using contrastive learning, and trains the teacher policy directly in this space. Because the teacher policy is constrained to operate on this shared space, it learns behaviors that are reproducible from student observations. This substantially reduces the imitation gap while preserving task performance. Across multiple environments, RAIL outperforms strong baselines without reward modification or post-hoc student fine-tuning, and surpasses direct reinforcement learning from raw observations. The learned representations further enable zero-shot transfer to new tasks through teacher-only training.
Vanilla SVGD is known to have a computational cost of $\mathcal{O}(N^2)$ per iteration. In this work, we introduce Random-Projection Tree Stein Variational Gradient Descent (RP-SVGD) to alleviate this computational challenge. This is achieved by restricting kernel interactions to spatially proximal particles clustered via a random projection tree, further reducing to a cost of $\mathcal{O}(N d(\log (N/C)+C))$, where $C$ is a hyperparameter representing the maximum leaf node capacity of the spanning tree. To establish theoretical validity, we introduce a smoothed RP-Tree kernel for any fixed tree realization and prove that it belongs to the Stein class of the target distribution, thereby ensuring the generation of valid gradient flows. In addition, by considering the effective kernel as the expectation of random tree partitions, we verify that it preserves regularity along with other necessary conditions, which guarantee convergence under the approximate gradient flow framework. Extensive experimental results demonstrate that RP-SVGD tends to have competitive performance and significant speedups across various tasks.
Rank-Aware Differentially Private Release of Listwise Preferences for LLM Alignment
Junwei Chen ⋅ Manjiang Yu ⋅ Pengpeng Qiao ⋅ Yang Cao
Listwise preference feedback offers richer supervision for large language model (LLMs) alignment than pairwise comparisons, but a full ranking also reveals more sensitive user preferences. Existing privacy-preserving alignment methods focus mainly on pairwise feedback, while listwise alignment under local differential privacy (LDP) remains unexplored. Extending LDP to this listwise domain raises a critical challenge: the high sensitivity of the full ranking space demands excessive noise, neutralizing the benefits of listwise supervision. To address this challenge, we analyze the interaction between supervision granularity and privacy noise in downstream training, and introduce a rank-aware exponential mechanism that privatizes listwise preference data into a low-sensitivity granularity suitable for downstream alignment. The mechanism leverages ranking information to sample a fixed-size binary partition, concentrating the noise near borderline items instead of perturbing all items uniformly as randomized response (RR) does. Empirical evaluations on downstream LLM alignment tasks show that our mechanism consistently outperforms existing LDP baselines at matched privacy budgets.
Rare Disease Diagnosis Agent with Decoupled Workflows and Knowledge-Driven Self-Evaluation
Yunlu Yan ⋅ Yawen Huang ⋅ Xian Wu ⋅ Lei Zhu
Rare disease diagnosis is a fundamental challenge due to heterogeneous and overlapping clinical phenotypes. While recent LLM-based agentic systems have demonstrated promise by integrating external medical tools and knowledge, they typically rely on large-scale or commercial LLMs, limiting their practical deployment in resource-constrained settings. In this work, we observe that their performance significantly degrades when using small-scale LLMs, due to entangled workflows and unreliable multi-evidence aggregation. To address this issue, we propose RADAR, a rare disease diagnostic framework designed for small-scale LLMs. RADAR adopts a divide-and-conquer paradigm that decouples heterogeneous diagnostic evidence into specialized workflows, reducing reasoning complexity and improving robustness. It further introduces a knowledge-driven self-evaluation mechanism that performs evidence-aware reasoning over structured disease knowledge to produce interpretable reliability scores by identifying key supporting phenotypes and conflicting evidence. Finally, a multi-evidence fusion strategy integrates outputs from multiple sources based on reliability and clinical agreement, resolving conflicts and producing stable diagnostic rankings. Experiments on three benchmarks show that RADAR consistently outperforms various state-of-the-art baselines, including general and medical LLMs and agentic systems. Code will be released.
Raven: High-Recall Sequence Modeling via Sparse Memory Routing
Arshia Afzal ⋅ Aviv Bick ⋅ Eric Xing ⋅ Volkan Cevher ⋅ Albert Gu
Long-context recall in linear-time sequence models highlights a tradeoff in how they write to memory. State-based linear models, such as state-space models (SSMs) and linear Transformers, write densely, updating the entire state for each newly arrived token, which leads to interference and makes specific past tokens hard to recover. Sliding-window attention (SWA) exhibits the opposite behavior: it writes sparsely by storing explicit token representations, but only within a fixed window, so recall drops once the relevant token is evicted. Interpolating between these models, we introduce Raven, a linear-time sequence model that maintains a fixed set of memory slots and, at each step, decays and updates only a selected subset via learned, input-dependent routing. This lets Raven mitigate SWA's position-based overwriting and hard eviction while reducing interference from dense state updates in SSMs, thereby preserving long-range content much more effectively. Across recall-intensive benchmarks, Raven is competitive with or outperforms prior linear-time baselines, achieving strong long-context recall where both SWA and SSMs sharply degrade. It remains effective when extrapolating to context lengths as large as 16x its training length, with similar gains in hybrid architectures.
Reason to Play: Behavioral and Brain Alignment Between Frontier LRMs and Human Game Learners
Botos Csaba ⋅ Sreejan Kumar ⋅ Austin T D Andrews ⋅ Laurence T Hunt ⋅ Josh Tenenbaum ⋅ Christopher Summerfield ⋅ Rui Costa ⋅ Marcelo G Mattar ⋅ Momchil Tomov
Humans rapidly learn abstract knowledge when encountering novel environments and flexibly deploy this knowledge to guide efficient and intelligent action. Can modern AI systems learn and plan in a similar way? We study this question using a dataset of complex human gameplay with concurrent fMRI recordings, in which participants learn novel video games that require rule discovery, hypothesis revision, and multi-step planning. We jointly evaluate models by their ability to play the games, match human learning behavior, and predict brain activity during the same task, comparing a suite of frontier Large Reasoning Models (LRMs) against model-free and model-based deep reinforcement learning agents and a Bayesian theory-based agent. We find that frontier LRMs most closely match human behavioral patterns during game discovery and predict brain activity an order of magnitude better than both reinforcement learning alternatives across cortical and subcortical regions, with effects robust to permutation controls. Through targeted manipulations, we further show that brain alignment reflects the model's in-context representation of the game state rather than its downstream planning or reasoning. Our results establish LRMs as compelling computational accounts of human learning and decision making in complex, naturalistic environments.
RECIPE: Learning to Rank Complete Precursor Sets for Inorganic Retrosynthesis
Jing Gao ⋅ Kaipeng Zeng ⋅ Fuyuan Xia ⋅ Jian Ma ⋅ Yufeng Li ⋅ Qifeng Li ⋅ Lin Yao ⋅ Junchi Yan
Single-step inorganic retrosynthesis is evaluated by whether a complete precursor set is recovered, yet many data-driven systems first rank individual precursors and then rely on threshold or count heuristics to assemble sets. This mismatch is especially severe when the true set size is unknown: individually plausible precursors can combine into incomplete, over-complete, or otherwise wrong near-miss sets. We reformalize the task as variable-size set-level ranking and introduce RECIPE, a target-conditioned framework that separates precursor recall from final set-level scoring. The Precursor Candidate Generator learns formula-level compatibility to build a high-recall precursor pool, while the Complete-Set Reranker compares variable-size candidate precursor sets directly. On the Retrieval-Retro year-split benchmark, the generator improves Combo@20 from 69.00 to 72.75. Compared with Retrieval-Retro, the reranker improves Combo@1 by 11.30 points to 71.70, Combo@20 by 20.82 points to 89.82, and Combo MRR by 14.14 points to 77.43. A preliminary 20-case out-of-distribution evaluation shows the same direction of improvement. These results suggest that set-level prioritization remains important after strong precursor recall, and that optimizing ranked complete precursor sets can reduce the inspection burden in inorganic synthesis planning.
Training-free personalization adapts a frozen vision-language model (VLM) to recognize user-defined concepts from only a few exemplars by retrieving concept entries at inference time. While recent retriever-based methods enrich each concept with descriptive evidence, they largely treat concepts in isolation, so the retrieved information can be plausible yet non-discriminative, especially when the personalized concept set contains highly confusable, fine-grained categories. We argue that personalization is inherently relational: correct recognition depends on the specific distinctions between a concept and its nearest alternatives. To this end, we propose ReCoG, a Relational Concept Graph for retrieval-augmented personalization. ReCoG represents concepts as nodes and augments them with directed edges that explicitly capture discriminative cues between concept pairs, elicited by a frozen VLM during database construction. At inference, ReCoG retrieves a shortlist and resolves ambiguity via explicit pairwise comparisons conditioned on the corresponding relational edges, selecting the concept that consistently wins against competing candidates. We also introduce FINGER benchmark to test personalized recognition on confusable concept sets. Across FINGER, ReCoG substantially outperforms prior methods, with the largest gains in the most confusable regimes. ReCoG also achieves consistently best performance on established personalization benchmarks including captioning and VQA, confirming that relational evidence is broadly beneficial beyond fine-grained settings.
Reconciling Operational Energy Trilemma: A Heterogeneous Risk-Constrained MDP Framework with Residual Policy Learning
Yujian Ye ⋅ Siqi Qian ⋅ Yizhi Wu ⋅ Tianxiang Cui ⋅ Goran Strbac
Deep decarbonization is transforming power grids into safety-critical, low-carbon cyber-physical systems, where AI controllers must jointly manage economic efficiency, carbon reduction, and operational security. This operational energy trilemma is challenging for reinforcement learning: cost and carbon objectives are naturally optimized in expectation, whereas security violations are rare, heavy-tailed events whose consequences can be catastrophic and therefore require explicit tail-risk control. Existing safe reinforcement learning methods either constrain safety only in expectation or rely on light-tailed approximations, which can underestimate rare but severe grid violations. We propose a heterogeneous risk-constrained reinforcement learning framework for low-carbon grid operation. The key idea is to assign different risk semantics to different objectives: safety-critical security constraints are enforced through Conditional Value-at-Risk (CVaR), while economic and carbon-related objectives are optimized in expectation to preserve operational flexibility. To better capture extreme scenarios, a mixture distribution model is introduced to characterize heavy-tailed constraint violations. We further develop a residual policy learning scheme built on a distributionally robust chance-constrained reference module: the robust module provides a feasible and economically efficient baseline policy, and the learned residual refines it toward lower-carbon operation while respecting tail-risk bounds. Experiments on power-system operation tasks show that the proposed framework reduces economic cost by over 25% compared with expectation-based CMDP baselines and improves safety by two orders of magnitude over Gaussian-tail methods. These results suggest that risk-aware reinforcement learning can support reliable decarbonization by reconciling efficiency, sustainability, and security within a unified decision-making framework.
Recursive Semantic Divergence for LLM Agent Consistency
Harshavardhan Abichandani ⋅ Penny Chong ⋅ Atin Ghosh ⋅ Daniel Dahlmeier
Large Language Model (LLM) agents with tool use exhibit inconsistent behavior across independent runs from identical states, calling different tools, producing contradictory outputs, or pursuing divergent strategies. This variability compounds over time and undermines deployment reliability. Existing evaluation metrics primarily measure task correctness via completion-based scores and do not capture consistency in agent behavior across executions. Separately, agent behavior is analyzed either at the token level using probability distributions or at the single-response level using semantic clustering, but neither captures how inconsistency propagates across future turns of a trajectory. We introduce Recursive Semantic Divergence (RSD), an unsupervised trajectory-level consistency metric adapted from bisimulation metrics in reinforcement learning. RSD measures the expected cumulative semantic disagreement between two independent rollouts of the same agent from the same state, using natural language inference to detect logical inconsistency. It is defined recursively over future states and approximated by a neural network trained on independent rollouts. The learned metric estimates future divergence from the current state, providing a consistency score at each turn without requiring the full trajectory to complete. When used as a reward signal for fine-tuning, RSD reduces the variation in task progress from 132% to 91% on $\tau^2$-bench and from 90% to 44% on toolsandbox, while improving performance across benchmarks and model scales. Training with RSD as an additional reward signal matches supervised fine-tuning performance with substantially lower variance, without requiring ground-truth labels.
Reinforcement Learning for Code Optimization
Pierre Chambon ⋅ Kunhao Zheng ⋅ Juliette Decugis ⋅ Benoît Sagot ⋅ Gabriel Synnaeve
RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that pass. Extending this to code optimization seems straightforward: just add execution time to the reward. But in practice, once timing drives the reward, small problems in measurement noise, reward sparsity, or GRPO instability overwhelm the signal and make RL fail: generated solutions are barely faster, and more of them can fail. We make execution time learnable through three stages: (1) how code is tested, by building DMC-Optim with large optimization tests and a calibrated sandbox; (2) how speed is turned into reward, by composing correctness and speed in the RL environment and using an offline simulator to predict the most promising configurations; and (3) how the model learns from that reward, by adapting GRPO and evaluation to the sparser, noisier timed-execution setting. On DMC-Optim, the strongest optimization-aware configurations improve strict top-decile pass@1 from 18.1% to 27.7% on Qwen 2.5 7B and from 28.7% to 42.4% on CWM 32B, while preserving most of the pure-correctness score. When the timing sandbox is degraded, robust optimization RL reaches up to roughly 130% improvement over standard RLVR. On LCB, CWM 32B wins up to 70.3% of best-sample speed comparisons against standard RLVR; relative to the fastest correct human submissions per problem, it reaches about half the human rate of complexity-class improvements (14% vs. 28%).
Reinforcement Learning for View-Adaptive Distillation in 3D Gaussian Compression
Hongji Zhao ⋅ Mingrui Zhu ⋅ Xin Wei ⋅ Nannan Wang
3D Gaussian Splatting enables real-time novel view synthesis with high fidelity, yet its sparse and unorganized Gaussian anchors make compression challenging without harming structure and cross-view consistency. A key challenge is that effective compression depends not only on reducing parameter redundancy, but also on identifying where the compressed model is most fragile and how limited representation capacity should be allocated accordingly. We present an RL-guided distillation framework for 3DGS compression (RLGS). RLGS first uses a lightweight reinforcement learning policy to predict informative viewpoint offsets for adaptive distillation. It then distills a compact student 3DGS from a high-quality teacher under multi-view rendering supervision. Finally, it introduces an Unbalanced Optimal Transport based structural selection module to align teacher-student anchor distributions in voxelized space, thereby retaining the most informative anchors during pruning and allocation. Under the same bitrate constraints, our method consistently outperforms strong 3DGS compression baselines in rendering quality. At matched rendering quality, it further reduces model size while preserving structural fidelity, achieving up to about 15% additional size reduction at comparable rendering quality over prior methods.
We study tabular reinforcement learning problems with multiple steps of lookahead information. Before acting, the learner observes $\ell$ steps of future transition and reward realizations: the exact state the agent would reach and the rewards it would collect under any possible course of action. While it has been shown that such information can drastically boost the value, finding the optimal policy is NP-hard, and it is common to apply one of two tractable heuristics: processing the lookahead in chunks of predefined sizes ('fixed batching policies'), and model predictive control. We first illustrate the problems with these two approaches and propose utilizing the lookahead in adaptive (state-dependent) batches; we refer to such policies as adaptive batching policies (ABPs). We derive the optimal Bellman equations for these strategies and design an optimistic regret-minimizing algorithm that enables learning the optimal ABP when interacting with unknown environments. Our regret bounds are order-optimal up to a potential factor of the lookahead horizon $\ell$, which can usually be considered a small constant.
RelationVGGT : Visual Geometry Transformers for 3D Spatial Relation Segmentation
Minsu Kim ⋅ Jaesung Choe ⋅ Jiwoo Lee ⋅ Frank Wang ⋅ Seon Joo Kim
Recent advances in 3D reconstruction have progressed from per-scene optimization to feed-forward inference, and semantic scene understanding has followed suit—yet existing methods remain confined to object-centric perception, neglecting spatial relations between objects. We introduce \emph{3D spatial relation segmentation}, a task that requires models to identify a target object satisfying a given spatial relation with respect to a subject, consistently across multiple views. To this end, we propose RelationVGGT, a novel feed-forward framework that integrates semantic features from a visual foundation model with geometry-aware representations from a 3D geometry foundation model and leverages a relation transformer for subject-conditioned, cross-view relation prediction—requiring neither per-scene optimization nor known camera poses. We additionally provide a fully automated annotation pipeline built on ScanNet++ with VLMs and LLMs, enabling scalable training data generation for this new task.
Relative Energy Barriers for Copyright-Aware Language Model Adaptation
Wei Sun ⋅ Yuhang Zhang ⋅ Zhicheng Liang ⋅ Xianda Wang ⋅ Jiaxuan Chen ⋅ Fangxin Wang
Adapting language models on long-form corpora can improve domain behavior while also making protected passages easier to reproduce verbatim from short prefixes. We study this tension through \emph{relative energy gain}, the token-level log-density ratio between an adapted model and its pre-adaptation anchor on an exact protected continuation. This quantity separates ordinary predictability from update-induced copying pressure and accumulates over suffixes as a likelihood-ratio advantage for extractable memorization. Building on this view, we introduce Energy-Gated Copyright Regularization (EGCR), a training objective that augments maximum likelihood with a soft relative-energy budget on protected tokens. EGCR uses a differentiable gate combining local persistence of relative gain with anchor surprisal, so regularization is concentrated on distinctive spans whose probabilities rise in a memorization-like way while ordinary next-token supervision is retained elsewhere. Across BookMIA adaptation stress tests, training-prefix extraction, CopyBench transfer, and gate ablations, EGCR reduces exact-match and overlap-based reproduction while preserving held-out language-modeling and book-related utility. Token-level visualizations further show that the gate focuses on a small set of expressive positions associated with verbatim recovery. These results suggest that relative-energy barriers offer an effective and inspectable mechanism for copyright-aware language model adaptation.
RelaVPR: Relation-Based Knowledge Distillation for Efficient Visual Place Recognition
Zheyuan Zhang ⋅ Jiwei Zhang ⋅ Rongjing Cai ⋅ Nihua Shrestha ⋅ Zilong Wang ⋅ Boyu Zhou ⋅ Hong Chen
Visual Place Recognition (VPR) aims to identify and retrieve geographic locations from visual observations. With the advent of Visual Foundation Models (VFMs), recent VPR methods have primarily adopted them as backbones to exploit their strong generalization capabilities, effectively mitigating challenges such as perceptual aliasing and long-term appearance variations. However, their large parameter sizes and high inference latency impose substantial hardware demands, significantly hindering practical deployment. To retain the generalization strength of VFMs while enabling efficient VPR, we revisit this problem from the perspective of Knowledge Distillation (KD). Specifically, we conduct a systematic investigation of various KD losses for VPR from both theoretical and empirical standpoints. Our analysis reveals that relation-based KD methods consistently achieve superior performance and efficiency, which we attribute to their larger feasible solution space and better alignment with the VPR objective. Building upon this insight, we propose a relation-based KD framework, termed RelaVPR, which is further enhanced along three key dimensions: 1) dynamically filtering noisy knowledge, 2) improving hard-sample learning, and 3) mitigating gradient interference in multi-teacher distillation. Compared with state-of-the-art (SOTA) KD-based VPR methods, RelaVPR requires no additional fine-tuning while delivering higher performance, faster inference, and a more compact model. Moreover, RelaVPR establishes a stronger performance--efficiency frontier than existing SOTA VPR approaches across multiple popular benchmarks. Codes and weights will be released.
RelFlexformer: Efficient Attention Transformers for Integrable Relative Positional Encodings
Byeongchan Kim ⋅ Arijit Sehanobish ⋅ Kumar Avinava Dubey ⋅ Min-hwan Oh ⋅ Krzysztof M Choromanski
We present a new class of efficient attention mechanisms applying universal 3D Relative Positional Encoding (RPE) methods given by arbitrary integrable modulation functions $f$. They lead to the new class of 3D-Transformer models, called *RelFlexformers*, flexibly integrating those RPEs, and characterized by the $O(L \log L)$ time complexity of the attention computation for the $L$-length input sequences. RelFlexformers builds on the theory of the Non-Uniform Fourier Transform (NU-FFT), naturally generalizing several existing efficient RPE-attention methods from structured settings with tokens homogeneously embedded in unweighted grids into general non-structured heterogeneous scenarios, where tokens' positions are arbitrarily distributed in the corresponding 3D spaces. As such, RelFlexformers can be applied in particular to model point clouds. Our extensive empirical evaluation on a large portfolio of 3D datasets confirms quality improvements provided by the NU-FFT-driven attention modulation techniques in the RelFlexformers.
Reliable Clustering and Quantization via Distortion-Constrained Optimal Transport
Tianhao Wu ⋅ Wei Zhang ⋅ Haoran Pang ⋅ Haoran YANG ⋅ Hao Shi
Clustering and vector quantization are core primitives for representation learning and large-scale retrieval, yet widely used methods often produce clusters or codewords of highly uneven quality, especially in noisy or heterogeneous data. In this paper, we propose a distortion-constrained optimal transport (DCOT) framework that explicitly enforces per-cluster bounds on the average assignment distortion, defined as the expected distance between data points and their assigned representatives, ensuring more consistent and reliable representations. Our formulation couples an optimal transport assignment objective with cluster-level distortion constraints to control cluster quality. To solve this constrained problem, we develop an efficient alternating optimization algorithm that iteratively updates transport plans, representatives (codewords), and dual variables. We further provide theoretical guarantees including entropic consistency, coordinate descent, and sublinear convergence. Importantly, DCOT serves as a plug-in assignment module that can be seamlessly integrated into a wide range of vector quantization methods. Replacing the assignment step with DCOT consistently improves reconstruction quality across diverse architectures and training settings, demonstrating strong robustness and general applicability.
Remember with Confidence: Uncertainty Quantification for Spatio-temporal Memory with Probabilistic Guarantees
Harry Zhang ⋅ Nicolas Gorlo ⋅ Luca Carlone
Long-horizon robot operation requires spatio-temporal memory to record the environment state and recall it for downstream reasoning. Scene graphs and retrieval-augmented systems ground VLM descriptions to persistent 3D entities with rich semantic descriptions. However, VLM captions are noisy and viewpoint-inconsistent, and existing systems treat them as an oracle with no mechanism to detect unreliable stored descriptions. We introduce object-level semantic uncertainty for multi-view VLM memory: a score that measures object-centric cross-view semantic scatter of captions and identifies semantically unresolved objects. Then, we include our uncertainty scores in an advanced spatial-semantic memory system, that we dub UQ-DAAAM. UQ-DAAAM uses this score to actively refine uncertain objects under a fixed query budget by selecting high-quality views and fusing the resulting multi-view captions into a single object description. We also derive probabilistic guarantees showing that higher-quality candidate views (as selected by our approach) are more likely to reduce uncertainty. Our experiments show that uncertainty quantification can make embodied 4D memory systems more reliable and more effective. In particular, on the OC-NaVQA benchmark, UQ-DAAAM achieves substantially larger uncertainty reduction and better spatio-temporal question answering performance than baselines.
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation
Suyoung Bae ⋅ Jaehoon Lee ⋅ Changkyu Choi ⋅ YunSeok Choi ⋅ Jee-Hyong Lee
Automated code documentation is essential for modern software development, providing the contextual grounding that both human developers and coding agents rely on to navigate large codebases. Existing repository-level approaches process components independently, causing redundant retrieval and conflicting descriptions across documents while producing outputs that lack hierarchical structure. Therefore, we propose MemDocAgent, a long-horizon agentic framework that generates documentation within a single, integrated context spanning the entire repository. It combines two components: (i) Dependency-Aware Traversal Guiding that predetermines a traversal order respecting dependency and granularity hierarchies; (ii) Memory-Guided Agentic Interaction, in which the agent interacts with RepoMemory, a shared memory accumulating prior work traces through read, write, and verify operations. Through an in-depth multi-criteria evaluation, MemDocAgent achieves the best performance over both open and closed-source baselines and demonstrates practical applicability in real software development workflows.
RePercENT: Scaling Disentangled Representation Learning Beyond Two Modalities
Vasiliki Rizou ⋅ Pascal Frossard ⋅ Dorina Thanou
To leverage the full potential of multimodal data, we need representations that go beyond the state-of-the-art alignment and fusion approaches and exploit all cross-modal interactions without sacrificing modality-specific information. Learning disentangled representations is a principled way to identify these underlying shared and unique factors that are hidden in observational data. However, while multimodal disentanglement is a compelling paradigm, existing methods are largely confined to the two-modality regime due to its inherent scalability bottleneck. To address this, we propose RePercENT, a self-supervised framework designed to surpass these limitations, that $\textit{unlocks scalable pairwise disentanglement}$ as we move beyond two modalities. Through a multimodal `plug-and-play' architecture, our approach operates directly on pre-extracted embeddings, eliminating the need for extensive joint pre-training while making no assumptions regarding the underlying modalities or foundation model backbones. Moreover, we introduce a joint optimization objective for simultaneously deriving the shared and unique components, and provide formal theoretical guarantees that characterize the optimality of our solution. Across diverse modalities and tasks, RePercENT successfully recovers disentangled components while maintaining competitive performance and significantly reducing computational complexity.
RepoScope: Bridging Physical Structure and Logical Flow for Intent-Driven Repository Documentation
Zezhou Yang ⋅ Chaozheng Wang ⋅ Hailiang Huang ⋅ Peihua Zhang ⋅ TingPeng ⋅ yuetang deng ⋅ Chenxiong Qian
Repository-level documentation generation is vital for human comprehension and autonomous agents, yet existing approaches fail to align high-level architectural intent with granular implementation details, resulting in documentation that lacks both global coherence and local fidelity. To address this, we propose RepoScope, a repository documentation generation framework that achieves alignment via two synergistic innovations. Specifically, RepoScope employs Physico-Logical Partitioning by jointly modeling directory structures and call graphs, effectively bridging the gap between physical layouts and logical dependencies while creating coherent documentation units. Within this partitioned space, we introduce SAIL (Seeded Architectural Intent Lift), a mechanism that projects architectural intent top-down to steer granular analysis, which is then revised and lifted bottom-up via code evidence to ensure global-local consistency. Our experimental results on CodeWikiBench demonstrate state-of-the-art performance in aligning with human-authored documentation. Furthermore, the validation on real-world industrial systems shows that RepoScope yields superior correctness and completeness in empowering autonomous agents for CLI extraction compared to existing baselines.
Representation Preconditioning for Efficient Diffusion Model Training
deyuan liu ⋅ PENG SUN ⋅ Xufeng Li ⋅ Tao Lin
Diffusion transformers learn useful internal representations during training, and learning them from scratch can add a burden to efficient training. Alignment mitigates this burden by adding an auxiliary loss that regularizes diffusion hidden states with features from pretrained visual encoders. In this work, we revisit the role of this alignment signal. Rather than using pretrained representations only as a regularizer during full diffusion training, we use them as a preconditioning signal before diffusion training. The key observation is that early layer alignment from clean latents to pretrained representations can be much cheaper than optimizing the full flow matching objective over the entire backbone, yet provides a better initialization for subsequent diffusion training. We instantiate this idea as Embedded Representation Warmup (ERW). ERW first aligns early layers from clean latents to pretrained representations, and then trains the full model with the standard diffusion objective and decaying alignment. This initialization allows subsequent diffusion training with alignment to converge more efficiently than starting alignment from scratch. Counting warmup in the total budget, ERW improves convergence in SiT based diffusion training. On ImageNet $256\times256$, ERW reaches FID=1.41 in 350 epochs with SiT-XL/2; it also reaches FID=2.04 on ImageNet $512\times512$ and improves FID on MS-COCO text-to-image generation.
Representation Rigidity in Face Embeddings: Orthogonal Identifiability and Backfill-Free Compatibility
Yongle Zhao ⋅ Zhichao Chen ⋅ Yin Xie ⋅ Jun Wang ⋅ Ziyong Feng
Upgrading a deployed face-recognition model usually makes new query embeddings incompatible with the legacy gallery, forcing costly backfilling of stored embeddings. We show that this incompatibility is often much simpler than it appears: independently trained angular-margin face models exhibit representation rigidity, where their embedding spaces are well approximated by a single global orthogonal transformation. We formalize this phenomenon as Procrustes rigidity, derive a finite-calibration recoverability bound, and estimate the orthogonal adapter from a small paired calibration set via closed-form Procrustes alignment. The resulting deployment protocol is backfill-free: transport upgraded queries into the legacy coordinate system, then reuse the existing gallery without re-embedding stored identities. Across 45 models spanning three architectures, two losses, three data scales, and five seeds, we find strong rigidity within architecture families, non-monotonic scaling with data size, and robust cross-model retrieval. On the 1.6M-image IFRT benchmark, Procrustes-aligned queries recover at least $99.8\%$ of mono-model Rank-1 accuracy for same-family upgrades, while lightweight residual correction substantially narrows the cross-architecture gap.
RescueBench: Can Embodied Agents Save Lives in the Wild?
Kui Wu ⋅ Beiyu Guo ⋅ Hao Chen ⋅ Shuhang Xu ⋅ Yuling Li ⋅ Yongdan Zeng ⋅ Zhoujun Li ⋅ Yizhou Wang ⋅ Fangwei Zhong
Search-and-rescue (SAR) requires embodied agents to explore unfamiliar environments under multimodal uncertainty, perform multi-stage interactions, and retrieve spatial memory over long horizons. Existing benchmarks typically evaluate these capabilities in isolation, leaving unclear how failures compound when they must be composed in realistic workflows. We introduce RescueBench, a photo-realistic diagnostic benchmark that instantiates SAR as a four-stage pipeline: multimodal exploration, target rescue, memory-guided return, and final handoff. By combining sequential task composition with stage-level evaluation, RescueBench enables analysis of how exploration and memory failures propagate through embodied rescue workflows. It contains five progressive difficulty levels that vary environmental complexity, clue ambiguity, and spatial hierarchy, along with an automatic episode generation and annotation pipeline for scalable evaluation and training. We evaluate seven baselines, an oracle reference, and human players, showing that no baselines completes the full task at the highest difficulty. Stage-level diagnosis identifies autonomous exploration as the dominant failure mode and spatial memory as a second, independent bottleneck, suggesting that these limitations are not resolved by current topological visual-language navigation or map-based methods.
Residual Paving: Diagnosing the Routing Bottleneck in Selective Refusal Editing
Bryce Hinkley ⋅ Peyman Najafirad
We study selective refusal editing as a three-way control problem: induce non-refusal on designated edit prompts while preserving benign behavior and harmful refusals outside the edit set. We introduce Residual Paving, a routed residual editing method for frozen instruction-tuned transformers that separates route selectivity, whether to intervene, from residual-edit capacity, what edit to apply. An early-layer router predicts a scalar gate and expert mixture; when active, prompt-conditioned bottleneck residual experts apply later-layer residual updates while leaving the backbone unchanged. This decomposition supports an oracle-routing diagnostic where only the learned scalar gate is replaced with the held-out edit/keep label, leaving the residual editor and frozen backbone fixed. On the primary Gemma-3-4B-IT held-out split, learned Residual Paving reduces edit refusal from 88.6% to 4.0%, with 95.5% benign distribution preservation and 87.3% harmful distribution preservation. Same-protocol one-direction steering controls are much weaker on edit success, leaving edit refusal at 86.8% for Edit-target ActAdd and 78.9% for DIM-style refusal steering. The remaining failure is off-target harmful-keep degradation: harmful refusal remains below the frozen-base rate, 65.3% vs. 81.6%. Across six backbones, oracle routing improves the keep-side diagnostic score on every reported row, with median gain +12.9 pp, supporting the interpretation that learned route selectivity is the main observed bottleneck. Trajectory diagnostics on two backbones further suggest directed movement toward edit-target continuations rather than generic refusal suppression.
Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization
Yapeng Liu ⋅ Yuanzhao Zhai ⋅ Xudong Gong ⋅ Feng Dawei ⋅ Bo Ding ⋅ Lin Wang ⋅ Huaimin Wang
As Embodied Agents System (EAS) move into physical domains such as autonomous navigation and household assistance, reliable execution becomes critical for physical task completion and collaboration trust. Recent EAS reliability evaluation works typically focus on outcome-centric metrics as success rate or safety-only scores to assess an agent's performance, which collapse diverse execution trajectories into coarse outcomes. Therefore they ignore a critical property of EAS -- which we define as the **Resilience** -- that reflects how EASs recover, stabilize, and extend under perturbations and across iterative updates. The lack of resilience is particularly critical in open-world environments due to continuous unexpected disruptions, thus directly affecting the quality of EAS deployment. To address this problem, we gain insight from the resilience-engineering concepts to EAS groundings and propose a novel resilience evaluation framework that can be flexibly applied to any EAS. Specifically, we define the first comprehensive resilience metrics suite for EASs system that exposes *Rebound*, *Stability*, and *Graceful Extensibility* across embodied tasks execution, providing a practical grounding for EAS resilience analysis. We further implement the resilience evaluation layer that transforms execution evidence into assessments for comparison, diagnosis and optimization. Across 400 household tasks with 10 representative EAS baselines, our evaluation reveals process-level distinction hidden by outcome metrics, including recovery cost differences among successful episodes ($\Delta C_{\mathrm{rec}}=25.2$), increased instability under semantic perturbations, and task-family-specific degradation under stress. Metrics-guided optimizations reduce recovery cost by 42.94%, increase stability by 19.87%, and improve graceful extensibility completion by 10.08%, showing the diagnostic effect of resilience evaluation. Beyond these targeted gains, our results reveal trade-offs among resilience characteristics, suggesting that a resilient EAS construction should be evaluated and configured according to deployment-specific requirements. Our code and data are available at .
Resolving Representation Ambiguity in Feedforward Novel View Synthesis Transformer via Semantic-Spatial Decoupling
Yihang Wu ⋅ Yihang Sun ⋅ Shaofeng Zhang ⋅ Zuxuan Wu ⋅ Junchi Yan ⋅ Xiaosong Jia ⋅ Yu-Gang Jiang
Transformer-based models have advanced feedforward novel view synthesis (NVS). Current architectures such as GS-LRM and LVSM mix semantic information (e.g., RGB) and spatial information (e.g., Pl\"ucker rays) into a shared feature space. Since Pl\"ucker rays naturally carry lattice-like spatial structure, these designs can make the spatial bias interfere with appearance representation and degrade rendering fidelity. To this end, we propose to decouple the representation of feedforward NVS transformers into separate semantic and spatial tokens. The decoupled design keeps semantic and spatial information explicit in their branches while preserving cross-branch interaction through shared attention routing. Built on this design, we introduce optional categorized supervision and bidirectional modulation: the former provides branch-specific training signals, while the latter improves interaction between the two branches. Notably, the base decoupled design introduces virtually zero additional inference latency due to its architectural design. The proposed designs achieve consistent improvements, demonstrating effectiveness across decoder-only and encoder-decoder feedforward NVS models.
Rethinking Bayesian Optimization for Co-Optimizing LLM Training Configurations
Zhiliang Chen ⋅ Alfred Leong ⋅ Shao Yong Ong ⋅ Apivich Hemachandra ⋅ Gregory Kang Ruey Lau ⋅ Chuan Sheng Foo ⋅ Zhengyuan Liu ⋅ Nancy F Chen ⋅ Bryan Kian Hsiang Low
Fine-tuning an LLM to maximize performance on a downstream task requires finding the optimal data and model configuration. This is a black-box optimization problem that Bayesian optimization (BO) addresses by sequentially evaluating training configurations and using observed performance feedback to adaptively guide the search towards better configurations. However, directly applying BO to the joint data-model space suffers from two fundamental challenges: the prohibitive cost of full LLM training at each iteration, and the large number of BO iterations required for effective exploration due to the high problem dimensionality. This paper introduces JoBS, a BO-based approach that co-optimizes data and model configurations without requiring a full training run at every iteration. JoBS allocates an initial portion of the optimization budget to learn a scaling-law-inspired performance predictor that estimates fully trained LLM performance from only a small number of training steps. The remaining budget is then used to run BO with this predictor, eliminating the need for full training runs and enabling JoBS to explore significantly more data and model configurations within the same budget. We analyze JoBS' average regret and derive the optimal budget allocation between predictor learning and BO iterations. Empirically, JoBS outperforms independent data and model optimization methods and existing multi-fidelity BO baselines across a diverse set of downstream tasks. Our code is available at https://github.com/a35453779/JoBS.
Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
Zidan Wang ⋅ Yaqian Li ⋅ xiaokai zhang ⋅ Kun He ⋅ Kaiwen Long ⋅ Hanpeng Liu
CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a resource-efficient route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that the reported forgetting is not caused by insufficient negatives, but by an inappropriate magnitude of the contrastive temperature $\tau$: with $\tau$ set sufficiently small, contrastive loss becomes the strongest single-loss post-training objective. Building on this finding, we propose \textbf{ComCLIP}, a complementary post-training framework that freezes CLIP's text encoder---preserving compatibility with downstream VLMs---and trains the vision encoder under three complementary losses operating in the shared contrastive space: a properly-tempered contrastive loss for image-text alignment, an MSE anchoring loss against the original CLIP for preserving the pretrained feature manifold, and a relational distillation loss from DINOv2 for fine-grained visual knowledge injection. With only one epoch on small data, ComCLIP consistently outperforms prior post-training methods on ViT-B/16 and ViT-L/14 across zero-shot classification, retrieval, linear probing, and MMVP, achieving up to $\mathbf{+7.41}$ points on MMVP. Furthermore, ComCLIP serves as a drop-in vision encoder for LLaVA-7B, improving average performance across $8$ standard VLM benchmarks without any re-alignment of the LLM or projector.
Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction
kejun ren ⋅ Lei Jin ⋅ Tianxin Huang ⋅ Lianming Xu ⋅ Li Wang
Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TTT3R-style per-token gates across five benchmarks and discover a structural bottleneck: the gate is intrinsically bounded in magnitude (median $0.31$; never exceeding $0.6$) and nearly frame-invariant, yielding an effective memory horizon of only $\sim$3 frames per state token, which serves as the structural origin of long-sequence drift. We trace this to a missing axis: existing inference-time methods modulate updates only at the per-token, intra-frame level, while the orthogonal frame-level question of \emph{how strongly each frame should contribute to the state} has been silently treated as content-independent. We close this gap with a scalar frame-level gate $\alpha_t \in (0, 1]$ derived in closed form from frame-to-frame changes of internal features---a continuous relaxation of classical SLAM keyframe selection that requires no parameters, no training, and no extra forward pass. Across six benchmarks spanning camera pose, video depth, and 3D reconstruction at sequence lengths up to $4,541$ frames, our gate cuts ATE by $51\%$ on long TUM-RGBD pose sequences, reduces AbsRel by $12.8\%$ on Bonn video depth, and on KITTI long-sequence pose estimation surpasses LongStream and Keyframe-VO which are specially constructed for long sequences via retraining and policy learning, while retaining strictly constant memory at zero training cost.
Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement
Szu-Wei Fu ⋅ Rong Chao ⋅ Xuesong Yang ⋅ Sung-Feng Huang ⋅ Ryandhimas E Zezario ⋅ Rauf Nasretdinov ⋅ Ante Jukić ⋅ Yu Tsao ⋅ Frank Wang
Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curation remain unresolved. In this work, we systematically address these three overlooked problems. First, we revisit the conventional practice of using early-reflected speech as the dereverberation target and show that it can degrade perceptual quality and downstream ASR performance. We instead demonstrate that time-shifted anechoic clean speech provides a superior learning target. Second, guided by the distortion--perception tradeoff theory, we propose a simple two-stage framework that achieves minimal distortion under a given level of perceptual quality. Third, we analyze the trade-off between training data scale and quality for USE, revealing that training on large uncurated corpora imposes a performance ceiling, as models struggle to remove subtle artifacts. Our method achieves state-of-the-art performance on the URGENT 2025 non-blind test set and exhibits strong language-agnostic generalization, making it effective for improving TTS training data. Code and models will be released upon acceptance. Audio samples are available at https://anonymous.4open.science/r/USE-5232/index.md.
Rethinking Visual Reasoning in Text-to-Image Reward Modeling
Shiye Su ⋅ Xiaohan Wang ⋅ Ludwig Schmidt ⋅ Serena Yeung-Levy
Reward models are a core component of modern text-to-image systems, serving as the objective for inference-time selection and RL-based post-training. A key prerequisite is visual understanding: to judge prompt adherence and ground aesthetic preferences, a reward model must bind attributes, resolve relations, and count reliably. We find that widely-used reward models perform poorly on targeted visual reasoning evaluations. Even when starting from strong vision–language backbones, preference fine-tuning erodes these capabilities, and continual-learning baselines fail to mitigate a steep tradeoff between visual reasoning and preference accuracy. We propose a data-centric intervention that injects visual reasoning supervision into preference learning by flipping the pairing axis. To complement conventional human preference data (paired images under a shared prompt), we convert existing visual reasoning datasets into preference pairs that share the same image but differ in prompt, and train on the mixture. This yields a Pareto improvement in both visual understanding and preference prediction. Across inference-time best-of-n selection and RL post-training, the resulting reward model improves text-to-image generation on challenging compositional prompts (GenEval) by an average of 4.0% and 5.5% respectively, outperforming HPSv3 which was trained on 6x more preference pairs.
Rethinking XAI Evaluation: A Human-Centered Audit of Shapley Benchmarks in High-Stakes Settings
Inês Oliveira e Silva ⋅ Sérgio Jesus ⋅ Iker Perez ⋅ Rita P. Ribeiro ⋅ Carlos Soares ⋅ Hugo Ferreira ⋅ Pedro Bizarro
Shapley values are a cornerstone of explainable AI, yet their proliferation into competing formulations has created a fragmented landscape with little consensus on practical deployment. While theoretical differences are well-documented, evaluation remains reliant on quantitative proxies whose alignment with human utility is unverified. In this work, we use a unified amortized framework to isolate semantic differences between eight Shapley variants under the low-latency constraints of operational risk workflows. We conduct a large-scale empirical evaluation across four risk datasets and a realistic fraud-detection environment involving professional analysts and 3,735 case reviews. Our results reveal a fundamental misalignment: standard quantitative metrics, such as sparsity and faithfulness, are decoupled from human-perceived clarity and decision utility. Furthermore, while no formulation improved objective analyst performance, explanations consistently increased decision confidence, signaling a critical risk of automation bias and overconfidence in high-stakes settings. These findings demonstrate that current evaluation proxies are insufficient for predicting human impact, motivating our release of interaction measurements to support research into behaviorally grounded XAI evaluation.
Reusable Conditional Resampling via Flow Matching for Constraint-Based Causal Discovery
SHUNYU ZHAO ⋅ Yanfeng Yang ⋅ Shuai Li ⋅ Kenji Fukumizu
Constraint-based causal discovery requires a large number of conditional independence (CI) tests. In nonlinear and high-dimensional settings, generative CI methods offer strong expressive power, but often suffer from high per test cost and limited reusability across the many query dependent conditional resampling tasks arising in graph search. We propose a generative nonlinear CI testing and graph search framework for PC-style causal discovery. First, we develop the Flow Matching based Conditional Independence Test (FMCIT), which combines joint flow matching modeling with conditional imputation to reformulate generative CI testing from query specific conditional modeling into a conditional resampling process that can be reused across queries. This allows the same generative model to repeatedly generate the randomized copies required by CI testing throughout the entire graph search procedure. We then introduce GPC‑FMCIT, which embeds FMCIT into a guided PC pipeline with screening, guidance, refinement, and orientation. By using low-cost screening to shrink the search space, constructing edge-specific candidate conditioning sets from local graph structure, and invoking FMCIT only in a budget-controlled refinement stage, GPC-FMCIT reduces both the deployment cost of individual nonlinear CI queries and the overall query complexity. Experiments on synthetic datasets and real-data analyses show that the proposed method achieves a strong accuracy--efficiency trade-off under nonlinear and high-dimensional structures. Overall, our results suggest that, through the joint design of reusable conditional resampling and controlled graph search, generative nonlinear CI testing can become a practical system level module for high‑dimensional PC‑style causal discovery.
Reversing the Roles of Signal and Noise: Modeling Noise in Astronomical Images without Paired Data
Drew Tufto ⋅ Daniel Muthukrishna
Many scientific instruments produce structured noise that contaminates the signal of interest and cannot be observed in isolation, ruling out standard supervised denoising. We propose BEAM (Background Estimation via Additive Modeling), a generative framework for modeling complex noise in these settings. As ground truth noise observations are unavailable, BEAM learns a generative model of the clean signal and uses it to score candidate noise estimates by how plausible the underlying signal looks. Score matching on signal data turns this into a tractable likelihood that we sample from with Langevin dynamics, which refines a coarse classical estimate of the noise into a more accurate one. We further show that when metadata governing the noise is known in advance, a separate conditional flow model predicts the noise from that metadata alone, enabling pre-observation forecasting of the noise. We instantiate BEAM on stray light in NASA's Transiting Exoplanet Survey Satellite (TESS), where contamination from the Earth and Moon obscures stellar signals. BEAM produces cleaner stray light estimates than the filtering tool used in current TESS pipelines and forecasts stray-light contamination across the observing window for the interstellar comet 3I/ATLAS using only spacecraft--Earth--Moon geometry, supporting prospective scheduling of time-critical observations and a path to recovering faint signals that current pipelines discard.
Revisiting Embodied Chain-of-Thought for Generalizable Robot Manipulation
Nan Sun ⋅ Yuan Zhang ⋅ Yongkun Yang ⋅ Wentao Zhao ⋅ Peiyan Li ⋅ Jun Guo ⋅ Wenxuan Song ⋅ Pengxiang Ding ⋅ Runze Suo ⋅ Yifei Su ⋅ Xin Xiao ⋅ Xinghang Li ⋅ Huaping Liu
Embodied chain-of-thought (CoT) aims to bridge linguistic reasoning with robotic control, yet its effective form and integration remain underexplored. In this paper, we revisit embodied CoT for robotic control at an unprecedented scale. We curate the largest embodied CoT corpus to date, comprising 978,743 trajectories, 226.3M samples, and 2592.5 hours of data. Through extensive experiments, we show that effective CoT must ground high-level semantic understaning in concrete linguistic action guidance -- such as end-effector movement and image-space trajectories -- whereas high-level reasoning alone yields marginal gains. More importantly, we identify that explicit CoT does not scale reliably as an autoregressive action prefix, suffering from compounding errors during inference. To address these challenges, we propose ERVLA, a vision-language-action (VLA) model that effectively leverages linguistic reasoning in generalizable robot manipulation. ERVLA is trained using a CoT-dropout strategy, allowing the model to leverage rich reasoning traces during training while predicting actions directly without CoT during inference to bypass autoregressive instability. This approach enables reliable scaling with increasing pre-training data. ERVLA achieves state-of-the-art results on LIBERO-Plus with an 86.9% success rate and reaches 53.2% on VLABench, showcasing superior performance in out-of-distribution settings. Furthermore, ERVLA outperforms competitive state-of-the-art baselines in real-robot experiments, especially in handling semantic ambiguity and long-horizon tasks. Code, data, and model checkpoints will be released.
Revisiting On-policy Adversarial Black-Box Distillation: Calibrating Groupwise Reward Geometry for Effective Advantage Construction
Xiao Cui ⋅ Mo Zhu ⋅ Yulei Qin ⋅ Yuze Wu ⋅ Wengang Zhou ⋅ Houqiang Li
Black-box distillation is a practical route for transferring capabilities from API-accessible large language models that expose only text outputs into smaller student models. Recent on-policy adversarial methods such as GAD improve over SeqKD by forming an adversarial loop between a critic and a student, where the critic provides rewards for GRPO-based student policy optimization over the student's sampled responses. However, GRPO computes advantages from the within-group relative rewards of student samples for the same prompt, whereas the critic is trained primarily to distinguish teacher responses from student responses. This objective mismatch can produce reward groups with collapsed scale or fragile margins, leading to brittle grouped optimization signals. We propose Groupwise Reward Geometry Conditioning (GRGC), a two-stage framework that improves advantage construction by shaping student-side reward groups during both critic training and policy optimization. To improve critic-side conditioning, Gaussian groupwise Optimal Transport calibration regularizes the critic during training to produce reward groups with non-collapsed spread and smooth rank-wise gaps by matching sorted prompt-wise rewards to group-centered Gaussian quantiles. Building on this conditioned reward geometry, policy-side group power modulation reshapes the prompt-wise reward groups before they are converted into advantages, preserving the critic-induced ordering while increasing optimization-relevant margin separability. Extensive experiments across diverse teachers, student model families and scales, and training datasets demonstrate the effectiveness of GRGC on both in-distribution and out-of-distribution evaluations, while introducing negligible overhead over GAD. The code is available at https://anonymous.4open.science/r/GRGC.
Revisiting Subgradient Dominance in Robust MDPs: Counterexamples, Hardness, and Sufficient Conditions
Toshinori Kitamura ⋅ Arnob Ghosh ⋅ Alex Ayoub ⋅ Thang Chu ⋅ Csaba Szepesvari
Projected subgradient descent (PSD) has gained popularity for solving robust Markov decision processes (RMDPs) because it applies to a broader class of uncertainty sets than traditional dynamic programming. Existing work claims that RMDPs with a general compact uncertainty set satisfy the subgradient dominance property, under which exact PSD converges to an $\varepsilon$-optimal policy in a polynomial number of updates (e.g., Wang et al., 2023). We show that these claims are incorrect. Even when the uncertainty set has cardinality two, the RMDP objective is not subgradient-dominant and can admit suboptimal strict local minima. Moreover, we prove that finding an $\varepsilon$-optimal policy can be NP-hard even in settings where subgradients are efficiently computable: (i) finite transition uncertainty sets and (ii) $sa$-rectangular finite transition uncertainty sets with finite cost uncertainty sets. Finally, we identify two conditions under which RMDPs do satisfy subgradient dominance: when, for each policy, either the worst-case transition kernel or the worst-case action-value function is unique.
Revitalizing Medical Time Series with Vision-Informed Retrieval: A Vision-Language Perspective
Guoqi Yu ⋅ Juncheng Wang ⋅ Emma, Shujun Wang
Medical time series (MedTS) analysis plays a critical role in supporting clinical decision-making and patient care. However, existing methods treat MedTS signals purely as numerical sequences, ignoring the rich information embedded in their visual waveform, which is essential for clinicians to diagnose physiological conditions. To bridge this gap, we introduce a novel Vision-Informed Retrieval (ViRe) framework that injects visual waveform prior knowledge into representations of raw MedTS sequences, thereby emphasizing morphology-relevant patterns and improving clinician-aligned reasoning. Specifically, a Vision Query is extracted using pre-trained vision-language models (VLMs) to obtain morphology-aware priors from waveform plots. To enrich the numerical representation with such morphology-aware information, we design a tailored attention-based cross-modal retrieval mechanism that aligns raw numerical features with these high-level vision priors through the Vision Query. ViRe demonstrates strong effectiveness, outperforming the state-of-the-art method by an average of 6.42% gain across six public MedTS benchmarks. Code is available at this Anonymous Repo.
Reward Budgeting Reduces Premature Convergence in Reinforcement Learning for LLM Reasoning
Mengni Jia ⋅ Mengyu Zhou ⋅ xiaoxi jiang ⋅ Guanjun Jiang
RL often plateaus early: policies sharpen quickly, and further training yields little gain despite remaining model- and data-capacity. Entropy is recently used to combat this plateau, but its promotion alone is not a sufficient intervention target in the settings we study. By analyzing RL's dynamics through the optimal policy, we instead identify a sharpness factor $\kappa$ as a more effective alternative. We then introduce a plug-in reward-budget mechanism, which allocates a fixed total reward budget to each prompt and distributes it among correct rollouts. Theoretically, this induces a more regulated $\kappa$-trajectory that deviates from many RL methods, and its updates are equivalent to optimizing a concave objective, approximately $\mathbb{E}(\log R)$. Empirically, our method achieves significant gains on Qwen3-8B and Llama3.2-3B-Instruct (up to +12.50 on AMC23 and +10.00 on AIME24 \& AIME25). Together, these results suggest that regulating algorithm-induced sharpening is more effective for mitigating premature convergence than directly preserving entropy.
Reward Is Not a Universal Interface for Generative Reinforcement Learning
Victor Huang ⋅ Poodar Chu ⋅ Hongsheng Li
A reward is not a universal RL interface: it becomes a valid update only through the probability object exposed by the policy branch. Autoregressive post-training works cleanly because token ratios are exact, but parallel discrete policies hide within-step dependence and flow/diffusion policies often lack cheap density ratios. We introduce MindRL, a reward-to-update interface controller that translates a shared reward through branch-native score objects while turning factorization, tractability, drift, and smoothness barriers into budgets over block size, serialization, anchors, clipping, reranking, and branch weights. As a closed-loop controller, MindRL makes the shared reward actionable for AR, parallel discrete, flow/diffusion, and AR+flow policies by selecting branch-native update objects and adapting their control budgets. Across language-side AR/parallel-discrete tests and matched flow/diffusion and AR+flow evaluations, MindRL improves task-quality or reward-risk tradeoffs while preserving each branch's native structure.
RISE: Red-teaming via Iterative Strategy Evolution for Modern Text-to-Image Models
Dmitrii Kharlapenko ⋅ Sergei Bratchikov ⋅ Konstantin Korolev ⋅ Aleksandr Nikolich
On modern production text-to-image systems, successful policy violations are rare, and previously effective human-written seeds are often patched out. Current automated red-teamers are poorly matched to this regime in two ways: unreliable success measurement and poor exploration. First, we find that judges widely used in prior T2I red-teaming work are unreliable under vague unsafe-content targets: they either miss true violations or reward benign borderline images on hardened APIs. We therefore define strict category-specific success criteria and calibrate strong VLM judges against human labels. Second, we show that broadly used prompt-modification pipelines do not solve the exploration problem: on harder guardrail settings they remain tied to seed prompts, fail to transfer, or cannot bootstrap positive examples. We introduce RISE, which evolves reusable strategies used to generate prompts rather than rewriting them one by one. The best discovered strategies are then reused to generate attacks across new scenarios. On DALL·E 3, Nano Banana 2 (Google) and GPT-Image-2, RISE reaches up to 13\% human-verified ASR; under the same calibrated evaluation, prior methods with reported ASR as high as roughly 30\% fall to near zero.
Risk-guided Estimation-aware Acceptance for Learning with Synthetic Data
Zijie Tian ⋅ Kevin Jin ⋅ yang Chen ⋅ Thomas Chun Man Lee
Synthetic data generation is becoming increasingly prevalent as advances in generative modeling accelerate. However, theoretical understanding of when and how synthetic data can improve downstream estimation and prediction remains underdeveloped. This question is especially relevant when synthetic samples are generated from proposal mechanisms informed by trained generators or domain knowledge, whose distributions may differ from the target population. To address this question, we propose REAL-Syn: Risk-guided Estimation-aware Acceptance for Learning with Synthetic Data, a framework for statistical learning with synthetic data selection. REAL-Syn uses selected subsets of candidate synthetic data to reduce estimator variability and often improve downstream prediction under target-proposal mismatch. The proposed procedure uses a covariance-based risk surrogate to select from the set of candidate synthetic samples; the accepted samples are then reweighted and pooled with the real observations for inference. We present theoretical results on the surrogate-optimal acceptance rule and the proposed selection procedure. We empirically show that REAL-Syn achieves competitive performance in simulations and an application to solar flare intensity forecasting.
Risks Create a Jagged Frontier of LLM Productivity Gains Across Computer Occupations
Deepika Chawla ⋅ Gagandeep Singh ⋅ Elham k buxton ⋅ Meicen Sun ⋅ Lav Varshney ⋅ Jeremy Riel ⋅ Craig De Voto
This position paper argues that the current frontier of expert-LLM collaboration is doubly jagged: it is uneven, not only because of jagged AI capabilities but also due to jagged AI risks. As LLM systems improve, incompetence-driven risks (e.g., misinformation) may decline, but adversarial risks can persist or even increase, leading to the likely persistence of a jagged risk–reward frontier. To show jaggedness, we introduce the first occupation-level quantification of risk-aware productivity gains for expert–LLM collaboration over a realistic task distribution, using standard O*NET computer-occupation tasks. For each task, we leverage six frontier models (e.g., GPT-5, Claude Opus) to produce structured risk–reward ratings and use 6 human experts to verify a subset for reliability, achieving high agreement. Risk captures the possible increase in individual, organizational, or societal harm due to employing LLMs. Reward captures possible productivity gains, i.e., time/cost savings at a fixed quality target after accounting for expert oversight to detect and mitigate risks. We aggregate task-level ratings into occupation-level scores. Our analysis shows that risk varies more than reward, yielding vast differences in risk–reward tradeoff: In safety-critical roles (e.g., Information Security Engineers), gains are offset by risk, whereas in less safety-critical roles (e.g., Web Developers), gains often substantially exceed risks. At the task-level, we identify high-reward and medium-risk tasks as the "Sweet spot" for expert-LLM collaboration. The identified sweet spot is evident in real-world deployments as the majority of Claude conversations assigned to O*NET tasks (from the recent Anthropic Economic Index) are high-reward and medium-risk, whereas high-risk tasks are minimally represented, supporting both our methodology and position.
Risk-sensitive optimal stopping arises wherever a single bad outcome dominates a sequential decision, such as option exercise, sequential medical treatment, and safety-critical control. The natural objective is the conditional value-at-risk (CVaR) of the stopped reward. We develop the first theoretical framework and algorithm for risk-sensitive deep optimal stopping. Together they resolve all four structural problems of risk-sensitive RL: state augmentation, score-function pathology, blindness to success, and bootstrapping circularity. We isolate three structural properties of optimal stopping: single-point reward, action-independent dynamics, and sum-zero gradient. They yield Markov optimality on the original state space and uniform boundedness of the per-step pathwise gradient at the deterministic-policy boundary, resolving the first two. From these, we derive DIOS (Distribution-Informed Optimal Stopping), a single-time-scale CVaR estimator on the Rockafellar-Uryasev envelope that resolves the remaining two. A soft penalty replaces the hard tail indicator, and a tail-level annealing schedule initializes training in the risk-neutral regime, where no trajectory is gated out. The inner quantile admits a closed-form update without an auxiliary critic. DIOS admits a finite-time best-iterate stationarity guarantee under standard CVaR non-degeneracy, the first explicit rate for CVaR optimal stopping. On Bermudan max-call options, DIOS achieves the highest mean CVaR and the lowest wall time against five representative CVaR baselines.
RIVET: Regex-to-Indexable Keys via Neural Translation for Interactive LIMIT-k Retrieval
Dong Wang ⋅ Yingze Li ⋅ Zihang Yu ⋅ Ziqing Zeng ⋅ Zongdai Jiang ⋅ Yuhao ZHANG ⋅ Kaixin Zhang ⋅ Shuangshuang Cui ⋅ Hongzhi Wang
Interactive regex search over large text columns needs exact matches under low latency, but indexes help only when a pattern exposes selective fragments. For LIMIT-k queries, the immediate bottleneck is choosing a few indexable keys that reach match-rich candidate buckets before regex verification. RIVET casts this step as constrained regex-to-key generation: a column-specific translator emits short keys, an existing index retrieves candidates, and the native regex engine verifies every output record. Training from repeated regex-key hits moves probability mass toward effective keys, so finite online sampling spends fewer regex checks before collecting the requested matches. Across four large datasets, RIVET reaches 35.8 ms median latency, improves over the strongest baseline REI by 3.4–36×, and reaches up to 1773× speedup over sequential scan. The same mechanism improves TPC-H query plans, preview retrieval under latency budgets, and 80 human-authored OOD regex queries, showing that regex-to-key generation translates into end-to-end system gains.
RL Excursions during Pre-training: How early is too early for on-policy learning?
Rachit Bansal ⋅ Clara Mohri ⋅ Tian Qin ⋅ David Alvarez-Melis ⋅ Sham Kakade
The standard LLM training pipeline applies reinforcement learning (RL) only after pre-training and supervised fine-tuning (SFT). We question this status quo by training a LLM from scratch and applying RL, SFT, and SFT followed by RL directly to intermediate pre-training checkpoints. We find that RL is effective very early, and often matches the full SFT$\to$RL pipeline early as well. Through experiments on harder problems, we find that targeted pre-training data composition is a strong lever for RL effectiveness, even more so than model scale. Beyond reasoning accuracy, applying RL directly to base checkpoints expands the model's distribution; the sharpening effect reported in recent work arises only when RL follows SFT. The general capabilities of the model remain essentially unchanged by RL, while they degrade following SFT. Finally, we merge RL and SFT objectives by \textit{parallel averaging}, which outperforms across all other training methods discussed, across metrics, while preserving general capabilities. Together, these results suggest that LLM training might benefit from an expanded use of RL.
RobustGenBench: A Benchmark for Robust Generalization to Adversarial and Common Perturbations, with Applications to Vision and Vision-Enabled Large Language Models
Maxime Heuillet ⋅ JONAS NGNAWE ⋅ Yann Pequignot ⋅ Rishika Bhagwatkar ⋅ Alexandre Larouche ⋅ Irina Rish ⋅ Christian Gagné ⋅ Ola Ahmad ⋅ Audrey Durand
Robust generalization (i.e., how models behave across diverse perturbation settings) is poorly understood for modern classification techniques such as robust fine-tuning from pretrained backbones and zero-shot classification with vision-enabled large language models (Ve-LLMs). We introduce \texttt{RobustGenBench}, a standardized benchmark of six classification datasets with reproducible evaluation splits and a unified perturbation protocol covering adversarial ($\ell_1, \ell_2, \ell_{\infty}$) and common perturbations. Using \texttt{RobustGenBench}, we conduct one of the most comprehensive studies of robust fine-tuning to date (40 pretrained backbones, 2 robust losses, and 3 adaptation protocols yielding 7{,}200 robustness measurements). We find that convolutional architectures with TRADES perform best at base scale, and the advantage of TRADES over Classic AT widens as size grows, and that hybrid architectures unlock competitive robust generalization. We further evaluate two frontier Ve-LLMs (GPT-4o mini, Gemini 3 Flash) and show that they exceed robust fine-tuned vision models on clean and common perturbations, and exhibit a striking resilience to $\ell_1$ perturbations that fine-tuned models lack. By evaluating both modeling techniques on the same tasks and perturbation settings, \texttt{RobustGenBench} provides a unifying view of robust generalization across these regimes.
Robust Noisy Inductive Matrix Completion with Local Linear Convergence
Xingcai Zhou ⋅ Xin Dong ⋅ Linglong Kong
We study noisy inductive matrix completion (IMC) in the presence of heavy-tailed and possibly asymmetric noise, where the goal is to recover a rank-$r$ matrix of ambient dimension $d$ given $n$ features as side prior information. So far, there has been a lack of theoretical understanding of the statistical estimation error in noisy IMC, let alone theory under heavy-tailed noise. We first develop an efficient two-stage nonconvex algorithm, called RGDIMC, via robust gradient descent with spectral initialization and feature-aware matrix factorization, based on an adaptive Huber loss to accommodate heavy-tailed noise. We prove that RGDIMC converges locally at a linear rate with sample complexity depending only linearly on $n$ (up to $\log n$) and logarithmically on $d$, and achieves the minimax-optimal statistical error rate $O_p(\sigma\sqrt{dr/p})$ under merely a bounded second-moment condition on the noise, where $r$ is the rank of the unknown core matrix. The theoretical results are strongly supported by extensive experiments.
RobustStress: Stress-Testing AI-Generated Text Detectors under Gradual Perturbations
Youcheng Yan ⋅ Jinshuo Liu ⋅ Qiyuan Li ⋅ Xinya Liu ⋅ Jinguang Gu ⋅ Donghong Ji ⋅ Jeff Pan
Detecting text generated by large language models (LLMs) is a critical line of defense for ensuring information content security. Although recent studies have gradually shifted to evaluating the robustness of detectors, existing benchmarks generally suffer from two primary limitations: (1) they lack granularity in controlling perturbation intensity, making it difficult to systematically evaluate the dynamic changes in detector performance as attack intensity increases; and (2) they fail to investigate the impact of perturbation intensity on detector robustness within complex scenarios, such as cross-domain and cross-model settings. To address these challenges, we propose \textbf{RobustStress}, a novel stress-testing framework for intensity-controlled perturbations. We achieve quantitative classification of perturbation intensity by varying the proportion of synonym replacement and employing graded LLM-based paraphrasing attacks. Experiments reveal distinct patterns of performance degradation as perturbation intensity increases, and further capture the synergistic weakening effect between cross-scenario shifts and adversarial perturbations. RobustStress provides an effective stress-testing framework for evaluating detectors in real-world scenarios, encouraging the community to move beyond performance on clean benchmarks to the development of more robust defense mechanisms against adversarial attacks.
ROCKET: Rapid Optimization via Calibration-guided Knapsack-Enhanced Truncation for Efficient Model Compression
Ammar Ali ⋅ Baher Mohammad ⋅ Denis Makhov ⋅ Dmitriy Shopkhoev ⋅ Magauiya Zhussip ⋅ Stamatios Lefkimmiatis
We present $\textbf{ROCKET}$, a training-free model compression method that achieves state-of-the-art performance in comparison with factorization, structured-sparsification and dynamic compression baselines. Operating under a global compression budget, ROCKET comprises two key innovations: First, it formulates layer-wise compression allocation as a multi-choice knapsack problem, selecting the optimal compression level for each layer to minimize total reconstruction error while adhering to a target model size. Second, it introduces a single-step sparse matrix factorization inspired by dictionary learning: using only a small calibration set, it sparsifies weight coefficients based on activation-weights sensitivity and then updates the dictionary in closed form via least squares bypassing iterative optimization, sparse coding, or backpropagation entirely. $\textbf{ROCKET}$ consistently outperforms existing compression approaches across different model architectures at 20–50\% compression rates. Notably, it retains over 90\% of the original model’s performance at 30\% compression without any fine-tuning. We will release the code implementing ROCKET upon paper acceptance.
RRL-HOI: Reflective Reinforcement Learning for Open-Vocabulary HOI Detection
Yongchao Xu ⋅ Jiawei Liu ⋅ Junfeng Wang ⋅ Yufei Zheng ⋅ Sen Tao ⋅ Tao Jiang ⋅ jiangbo Ai ⋅ Jin Zhang ⋅ Zheng-Jun Zha
Open-Vocabulary Human-Object Interaction (HOI) detection aims to localize interacting human-object pairs while generalizing to novel interaction categories beyond the training set. Although Multimodal Large Language Models (MLLMs) exhibit strong open-vocabulary understanding, directly applying MLLMs to HOI detection remains challenging, as language-prior hallucinations and coupled localization-semantic errors can destabilize structured HOI predictions. To address these issues, we propose RRL-HOI, a novel Reflective Reinforcement Learning framework that trains MLLMs to follow a verify-and-revise self-correction policy. Specifically, RRL-HOI first adapts MLLMs to produce box-grounded HOI triplets, and then introduces an executable reflection mechanism that targets the two failure modes through structured edits over localization, human-object pairing, and interaction semantics. In this way, localization and pairing edits mitigate localization-semantic coupling, while interaction-semantic edits suppress visually unsupported language-prior predictions. To further improve this reflection mechanism, RRL-HOI formulates reflection as a reinforcement learning objective trained with GRPO, where the reward jointly considers localization quality, pair-association correctness, interaction correctness, and edit consistency. Through reward-driven policy learning, RRL-HOI turns open-vocabulary recognition from a single-step generative prediction into an explicit verify-and-revise decision process, thereby bridging the gap between general multimodal understanding and precise interaction detection. Extensive experiments on two standard open-vocabulary HOI detection benchmarks demonstrate the effectiveness of our method.
RSSA: Robust Semantic and Spatial Aligner for Collaborative Perception
Zhihao Yang ⋅ Zhiyu Xiang ⋅ Peng Xu ⋅ Tianyu Pu ⋅ Kai Wang ⋅ Eryun Liu ⋅ Dongping Zhang ⋅ Yong Ding
V2X collaborative perception improves single-vehicle perception by aggregating multi-view features from collaborative agents. However, existing methods typically rely on simple spatial warping to align the collaborative features, which can cause semantic misalignment during fusion. Meanwhile, these deterministic warping-based methods are highly sensitive to the positional noises. In this paper, we propose a plug-in module called Robust Semantic and Spatial Aligner (RSSA), which performs semantic and probabilistic spatial transformations for robust feature alignment. Specifically, a Semantic Transformation (SemT) block modulates feature channels based on the observed relative pose to correct semantic inconsistencies. Subsequently, a Probabilistic Spatial Transformation (P-SpaT) block performs probabilistic warping by sampling from the Gaussian-modeled relative pose space. A Relative Transformation Augmentation (RT-Aug) strategy is further introduced to augment the relative poses and facilitate the training of the RSSA. Extensive experiments on both simulated and real-world datasets demonstrate that RSSA improves the state-of-the-art methods by a large margin under the configuration of both clean and noisy positions.
Rubato: Signature Attention for Irregular Multivariate Time Series Forecasting
Jintao Yang ⋅ Meng Wang ⋅ Chenjuan Guo ⋅ Bin Yang
Irregular multivariate time series (IMTS) are sampled at non-uniform timestamps and asynchronous across channels, their dependencies are shaped by time gaps, local ordering, and channel-specific sampling patterns. Building a faithful attention mechanism for IMTS exposes two limitations of existing work. First, current irregular-time attention methods typically regularize raw inputs into interpolated, aligned, or continuous-time-embedded forms before applying attention, which blurs exactly these irregular cues that determine how two observations relate. Second, beyond representation, standard dot-product attention scores pairs through learned token projections with no intrinsic tie to the shape of the underlying series, and an attention that does reflect this shape becomes costly to compute in high-channel IMTS. We propose Rubato, which reads two views directly from the raw irregular input: per-channel local dynamics, and a path signature that summarizes the joint multi-channel series. Attention in Rubato then scores pairs by a signature-kernel measure of series similarity rather than by a generic dot product, and a structured random projection (SORF) keeps this comparison affordable when the channel count is large. We prove that the structured projection introduces only a controlled bias that vanishes as the number of channels grows, and experiments on multiple IMTS benchmarks show that Rubato consistently outperforms strong baselines.
RVCBench: Benchmarking Robustness of Voice Cloning Across Modern Audio Generation Models
Ruinan Jin ⋅ Xinting Liao ⋅ Hanlin Yu ⋅ Deval Pandya ⋅ Xiaoxiao Li
Modern Voice Cloning (VCL) can synthesize speech that closely matches a target speaker from only seconds of reference audio, enabling applications such as personalized speech interfaces and dubbing. In practical deployments, modern audio generation models inevitably encounter noisy reference audios, imperfect text prompts, multilingual and long-form generation settings, downstream post-processing, and adversarial perturbations, all of which can significantly hurt robustness. Despite rapid progress in VCL driven by autoregressive codec-token language models and diffusion-based models, robustness under realistic deployment shifts remains underexplored. This paper introduces RVCBench, a comprehensive dataset and benchmark that evaluates Robustness in Voice Clone across the full generation pipeline. RVCBench contributes a large-scale, task-aligned robustness dataset that instantiates realistic deployment shifts through controlled text-audio pairing, multilingual and long-form scenarios, expressive prompts, post-processing conditions, and passive or proactive audio perturbations. Covering 18 robustness evaluations, 225 speakers, and 14,370 utterances, RVCBench enables unified evaluation of input sensitivity, generation stability, output resilience, and perturbation robustness. We evaluate 18 representative modern VCL models and reveal systematic vulnerabilities in content consistency, speaker similarity, long-form stability, post-processing resilience, detectability, and adversarial robustness.
SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
Kean Shi ⋅ Zihang Li ⋅ Tianyi Ma ⋅ Zengji Tu ⋅ Jialong Wu ⋅ Xinbo Xu ⋅ Qingyao Yang ⋅ Weichu Xie ⋅ Ruoyu Wu ⋅ Ming Wu ⋅ Jason Zeng ⋅ Michael Heinrich ⋅ Liang Chen ⋅ Kuan Li ⋅ Baobao Chang
Computer-Using Agents (CUAs) are rapidly extending large language models (LLMs) beyond text-based reasoning toward action execution in more complex environments, such as web browsers and graphical user interfaces (GUIs). However, existing web and GUI agent benchmarks often rely on simplified settings, isolated tasks, or short-horizon interactions, making it difficult to assess capabilities of agents in realistic professional workflows. Software-as-a-Service (SaaS) environments are a natural choice for CUA evaluation, as they host a large share of modern digital work and naturally involve dynamic system states, cross-application coordination, domain-specific knowledge, and long-horizon dependencies. To this end, we introduce \textbf{SaaS-Bench}, a benchmark built on 23 deployable SaaS systems across six professional domains, containing 106 tasks grounded in realistic work scenarios. These tasks require long-horizon execution, cover both text-only and multimodal settings, and are evaluated with weighted verification checkpoints that measure strict task completion and partial progress. Experiments show that representative LLM-based agents struggle on SaaS-Bench, with even the strongest model completing fewer than 4\% of tasks end-to-end, exposing limitations in planning, state tracking, cross-application context maintenance, and error recovery.
Safe Linear Bandits with Unknown Safety Gaps
MAOLI LIU ⋅ Zhuohua Li ⋅ Zeyu Zhang ⋅ Xiangxiang Dai ⋅ John C. S. Lui
We study stochastic linear bandits with a linear safety constraint that depends on an unknown parameter, requiring every chosen action to be safe at every round with high probability, even though the safe decision set is initially unknown. Prior work achieves $\widetilde{O}(\sqrt{T})$ regret when the safety gap, i.e., the constraint slack at the optimal action, is strictly positive and known to the learner. However, only $\widetilde{O}(T^{2/3})$ regret is established in two distinct cases: when the gap is zero, and when the gap is positive but unknown. This raises two natural questions: is the $\widetilde{O}(T^{2/3})$ bound tight in the zero-gap case, and is knowledge of the gap necessary for $\widetilde{O}(\sqrt{T})$ regret when the gap is positive? We answer both. First, we establish an $\Omega(T^{2/3})$ lower bound for zero-gap instances, showing that the existing upper bound is tight in its dependence on the horizon. Second, we propose Epoch-SLUCB, an algorithm that attains $\widetilde{O}(\sqrt{T})$ regret whenever the safety gap is positive, without requiring its value to be known. We also provide numerical simulations to demonstrate the performance of our algorithm.
SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models
Chengjie Hong ⋅ Feixiang He ⋅ Yiheng Zeng ⋅ Lulu Kang ⋅ He Wang
We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is essential for reducing memory use and accelerating inference in large foundation models, it remains under-explored for PFMs, where preserving physical fidelity is crucial. The challenge lies in the functional nature of physics data, where partial derivatives encode spatiotemporal dynamics and exhibit high sensitivity to compression. Conventional compression methods ignore this structure, often causing severe performance degradation or failure. To address this, we introduce a sensitivity-aware fidelity-enforcing compression framework that explicitly models loss-aware layer sensitivity in the output function space during compression. This provides a new route to compressing scientific foundation models while preserving accuracy and physical fidelity. Experiments show substantial gains over existing methods across multiple models and datasets, achieving significantly higher compression ratios while maintaining accuracy, in some cases by orders of magnitude. More broadly, the work potentially leads to a new subfield of efficient, deployable, and sustainable scientific foundation models in AI for Science.
SAGAS: Semantic-Aware Graph-Assisted Stitching for Offline Temporal Logic Planning
Ruijia Liu ⋅ Ancheng Hou ⋅ Xiang Yin
Linear Temporal Logic (LTL) provides a rigorous framework for specifying long-horizon robotic tasks, yet existing approaches face a trade-off: model-based synthesis relies on accurate labeled transition systems, whereas learning-based methods often require online interaction, task-specific rewards, or specification-conditioned training. We study LTL-specified robotic planning and execution in a stricter offline, model-free setting, where the agent is given only fixed, task-agnostic trajectory fragments, with no dynamics model, task demonstrations, or online data collection. To address this setting, we propose SAGAS, a framework that combines the compositionality of symbolic synthesis with the data-driven reachability structure learned from offline trajectories. SAGAS first learns a reusable latent reachability graph and a frozen goal-conditioned executor from fragmented offline data. For each new LTL formula, it performs task-time semantic graph augmentation to ground state-defined propositions on the learned graph, and applies B\"uchi product search to synthesize a cost-aware accepting prefix--suffix waypoint plan executed by the frozen executor. By shifting formula-specific reasoning from policy learning to test-time graph augmentation and symbolic search, SAGAS enables zero-shot generalization to unseen, data-supported LTL specifications without task-specific reward design, policy retraining, or online interaction. Experiments on LTL task suites constructed from OGBench locomotion domains show that this design produces executable and cost-efficient prefix--suffix behaviors for diverse unseen LTL tasks from fragmented offline data.
SAGE: Semantic-Agnostic Image Embedding for Generalized AI-Generated Image Detection
Seongho Kim ⋅ Jaehyun Choi ⋅ Dahye Kim ⋅ Sungwon Yi ⋅ Jang Ho Choi
AI-generated image detection has become an important problem in media forensics as modern generative models produce increasingly realistic images. Recent CLIP-based detectors show strong generalization ability, but CLIP image features entangle semantic content with forensic cues. Since semantic content is not an intrinsic generation trace, relying on it can introduce semantic shortcuts and unstable performance under unseen generator shifts. In this paper, we propose SAGE, a Semantic-Agnostic Image Embedding framework for generalized AI-generated image detection. SAGE constructs paired triplets of COCO images, SDXL-reconstructed images, and COCO captions so that real and fake images share the same caption-level semantics. It then suppresses semantic information by suppressing the semantic direction provided by CLIP text embeddings from CLIP image embeddings. To enable caption-free inference, SAGE learns a semantic anchor that approximates the role of paired text embeddings at test time. In addition, SAGE adopts a SoftTriple-style sub-center classifier to model intra-class variation in real and fake images, yielding a more stable authentication geometry across diverse generators. Experiments on CommunityForensics demonstrate that SAGE achieves competitive detection performance while substantially reducing generator-wise performance variation. Further analyses show that SAGE effectively reduces semantic dependence and forms a clearer real/fake feature geometry than strong baselines.
SAM 3D Animal: Promptable Animal 3D Reconstruction from Images in the Wild
Xuyi Hu ⋅ Jin Lyu ⋅ Jiuming Liu ⋅ Yebin Liu ⋅ Silvia Zuffi ⋅ Liang An ⋅ Stefan Goetz
3D animal reconstruction in the wild remains challenging due to large species variation, frequent occlusions, and the prevalence of multi-animal scenes, while existing methods predominantly focus on single-animal settings. We present SAM 3D Animal, the first promptable framework for multi-animal 3D reconstruction from a single image. Built on the SMAL+ parametric animal model, our method jointly reconstructs multiple instances and supports flexible prompts in the form of keypoints and masks which enable more reliable disambiguation in crowded and occluded scenes. To train such a model, we further introduce Herd3D, a multi-animal 3D dataset containing over 5K images, designed to increase diversity in species, interactions, and occlusion patterns. Experiments on the Animal3D, APTv2, and Animal Kingdom datasets show that our framework achieves state-of-the-art results over both existing model-based and model-free methods, demonstrating a scalable and effective solution for prompt-driven animal 3D reconstruction in the wild.
Low-rank adaptation (LoRA) has attracted significant attention in the parameter-efficient fine-tuning (PEFT) landscape. However, its design suffers from two fundamental limitations. First, its invariance manifold induces flat directions in the loss landscape, creating an optimization bottleneck for adaptive optimizers. Second, its effective rank scales only linearly with the parameter budget. In this paper, we propose Scalable Morpho Adaptation (SaMA), a PEFT method grounded in a principled generalization of the Kronecker product. By relaxing the shared-block constraint in the perfect-shuffle factorization, SaMA spans from low to full rank and achieves effective rank that scales \quadratically in parameters. Furthermore, we show that its invariance manifold collapses to a diagonal group, making it structurally more optimizer-friendly. Empirically, SaMA consistently outperforms strong PEFT baselines on commonsense and arithmetic reasoning benchmarks while being more parameter-efficient.
SAMAT: A Stereotype-Aware Multimodal Transformer for Interpretable Misogynistic Meme Detection
Gopendra Singh
This paper introduces SAMAT, a Stereotype-Aware Multimodal Alignment Transformer for detecting and explaining implicit misogyny in memes, where harm arises from subtle visual-textual incongruity and cultural stereotypes. SAMAT integrates three components: a Stereotype Subspace Projection Module (SSPM) that structures representations; a fidelity-based retrieval mechanism aligned with a curated Rationale Bank; and an evidence-conditioned explanation generator. For evaluation, we rely on the MEE corpus with 8,000 explanations, Stereotype Alignment (SAS) and Contextual Faithfulness (CFS) scores. Experiments show that SAMAT achieves a Macro-F1 of 88.1%, surpassing MLLM baselines, while improving retrieval faithfulness (SAS: 0.78) and explanation grounding (CFS: 0.68). Ablations confirm gains stem from structured stereotype projection and evidential retrieval, not scale. SAMAT offers a transparent, culturally grounded framework for accountable content moderation, aligning with Responsible AI objectives.
Same Concept, Different Directions: Cross-Modal Feature Heterogeneity in Sparse Autoencoders
Chungpa Lee ⋅ Jihoon Kwon ⋅ Kyle Min ⋅ Jy-yong Sohn
Vision-language models map images and text into a joint embedding space. However, these embeddings often entangle multiple semantic features, which limits their interpretability and controllability. While sparse autoencoders have emerged as a useful tool for decomposing these embeddings into monosemantic features, their application to joint embedding spaces has largely relied on an implicit, untested assumption that semantically corresponding features share the same directions across modalities. In this paper, we challenge this assumption by identifying discrepancies in feature directions for the same concept across image and text modalities, a phenomenon we term cross-modal feature heterogeneity. We demonstrate that this heterogeneity is a key driver of the modality split, where a shared concept activates different latents depending on the modality. This finding further reveals why aligning latent activations alone is insufficient to resolve the underlying feature mismatch. To address this misalignment, we propose an approach that trains sparse autoencoders to preserve the unique feature geometry of each modality and aligns corresponding features post hoc. Our method improves reconstruction fidelity and enhances performance in cross-modal retrieval and concept steering.
Sample Efficient Generative Molecular Optimization with Joint Self-Improvement
Serra Korkmaz ⋅ Adam Izdebski ⋅ Jonathan Pirnay ⋅ Rasmus Møller-Larsen ⋅ Michal Kmicikiewicz ⋅ Pankhil Gawade ⋅ Maksim Pavlov ⋅ Dominik G Grimm ⋅ Ewa Szczurek
Generative molecular optimization aims to design molecules with properties surpassing those of existing compounds. However, evaluating candidate molecules is expensive, so sample efficiency, i.e., generating optimized molecules within a fixed evaluation budget, is essential. Moreover, in the offline setting, where a surrogate model approximates candidate evaluation, sample inefficiency compounds with surrogate error, further limiting performance. To address these challenges, we introduce Joint Self-Improvement (JSI), a framework for sample-efficient online and offline molecular optimization. JSI combines (i) a joint generative-predictive model trained via a joint likelihood-based loss, which aligns molecule generation and surrogate modeling, and (ii) inference-time self-improvement, which iteratively biases the generative component of the joint model toward higher-scoring molecules using direct objective evaluations in the online setting and the predictive component in the offline setting. Experiments across docking score optimization, in both online and offline settings, as well as antimicrobial peptide optimization, demonstrate that JSI outperforms state-of-the-art methods under limited evaluation budgets.
Sample-Grained Approximate Unlearning with Provable Per-Sample Bounds
Yaxin Xiao ⋅ Yulin Jin ⋅ Li Tang ⋅ Haoyang LI ⋅ Huadi Zheng ⋅ Qingqing Ye ⋅ Haibo Hu
Machine unlearning aims to erase the influence of a designated forget set from a trained model while preserving strong utility on the retain set. Approximate unlearning methods efficiently achieve this goal without full retraining, but their effectiveness is primarily assessed using dataset-level aggregate metrics. As a result, prior work does not explicitly evaluate whether individual samples are unlearned correctly, nor does it provide mechanisms to enforce such per-sample unlearning behavior. In this work, we focus on per-sample unlearning effectiveness. We introduce SALA (Sample-grained Likelihoods ALignment), a plug-in optimization framework that enables approximate unlearning methods to operate at the level of individual samples by steering each example toward its correct target distribution. To quantitatively assess this behavior, we further propose the sample-grained distribution gap, a metric that captures per-sample discrepancies between approximate unlearning and exact retraining output distributions. We provide theoretical guarantees showing that SALA contracts this gap and yields a bounded per-sample discrepancy to retraining. To make this framework practical, we develop a low-cost estimation method for the member and non-member target distributions. Extensive experiments on classification and generation tasks demonstrate that SALA reduces the sample-grained distribution gap by up to 86.67% when paired with existing approximate unlearning baselines, while also improving run-to-run robustness.
Sample Size Design for Bounds on Discrete Probabilities of Causation
Tianyuan Cheng ⋅ Ruirui Mao ⋅ Judea Pearl ⋅ Ang Li
Probabilities of causation (PoCs), such as the probability of necessity and sufficiency (PNS), are important tools for decision making but are generally not point identifiable. Existing work has derived bounds for these quantities using combinations of experimental and observational data. However, there is very limited research on sample size analysis, namely, how many experimental and observational samples are required to achieve a desired margin of error. In this paper, we propose a sample-size design framework for PoC bounds that can be expressed as finite minima or maxima of smooth functions of experimental and observational probabilities — a representation that covers binary PNS, PN, and PS, as well as representative multi-valued PoC families recently shown to subsume the discrete PoC bounds in the current literature. Our framework gives endpoint-specific sample-size formulas that account for the joint covariance structure of the probability estimates: a distribution-free rule for affine bounds, and a pilot-based plug-in rule with theoretical safety guarantees for ratio-type bounds. Simulations show that the proposed rules achieve the target precision and coverage, recover a much smaller size requirement for binary PNS than the existing calculation, and provide sample-size designs for other discrete PoCs where no prior general sample-size formula is available.
Sample Transform Cost-Based Training-Free Hallucination Detector for Large Language Models
Zeyang Ding ⋅ Xinglin Hu ⋅ Jicong Fan
Hallucinations remain a major barrier to the trustworthy deployment of large language models (LLMs). We study hallucination detection from the perspective of prompt-conditioned response distributions. For a fixed prompt, an LLM induces a distribution over possible responses; when the model is uncertain or hallucinates, this distribution often becomes more complex and less internally consistent. However, this distribution is unknown, and its samples are variable-length token sequences rather than fixed-dimensional points, making direct complexity estimation difficult. To address this challenge, we propose a training-free detector based on sample-to-sample transform costs in representation space. Specifically, we represent each sampled response as an empirical distribution over generated-token hidden embeddings and compute pairwise Wasserstein distances between responses. The resulting Wasserstein distance matrix characterizes the cost structure of transforming one sampled response into another. From this matrix, we derive two complementary hallucination signals: \textbf{AvgWD}, which measures the average transform cost, and \textbf{EigenWD}, which captures the spectral complexity of the transform-cost structure. Importantly, we extend the proposed detector beyond white-box access by using an accessible auxiliary model to construct representation-space consistency signals for black-box LLMs, substantially broadening the practical applicability of our method. Experiments across five open-source LLMs and four benchmarks show that AvgWD and EigenWD consistently achieve competitive or superior AUROC compared with strong training-free uncertainty baselines. These results suggest that distributional complexity in token-level representation space provides an effective signal for hallucination detection.
Sampling-Free Privacy Accounting for Matrix Mechanisms under Random Allocation
Jan Schuchardt ⋅ Nikita Kalinin
We study privacy amplification for differentially private model training with matrix factorization under random allocation (also known as the balls-in-bins model). Recent work by Choquette-Choo et al. (2025) proposes a sampling-based Monte Carlo approach to compute amplification parameters in this setting. However, their guarantees either only hold with some high probability or require random abstention by the mechanism. Furthermore, the required number of samples for ensuring $(\epsilon,\delta)$-DP is inversely proportional to $\delta$. In contrast, we develop sampling-free bounds based on Rényi divergence and conditional composition. The former is facilitated by a dynamic programming formulation to efficiently compute the bounds. The latter complements it by offering stronger privacy guarantees for small $\epsilon$, where Rényi divergence bounds inherently lead to an over-approximation. Our framework applies to arbitrary banded and non-banded matrices. Through numerical comparisons, we demonstrate the efficacy of our approach across a broad range of matrix mechanisms used in research and practice.
SAVE: Sparsity-Aware Influence Estimation for Vocabulary-Expanded LLMs
Seungyoo Lee ⋅ Giung Nam ⋅ Seanie Lee ⋅ Hyungi Lee ⋅ Juho Lee
Large language models adapted to specialized targets such as Lean theorem proving or tool calling often benefit from extending the tokenizer with target-specific tokens. However, target-only vocabulary-expanded fine-tuning can leave the newly added embeddings transfer-incomplete, since the new rows receive gradients only from examples containing the corresponding tokens, so target-local gains do not necessarily imply that the embeddings are integrated with related capabilities. We propose SAVE, a sparsity-aware influence estimator for selecting auxiliary healing data from generic corpora. SAVE computes token-conditional curvature over the effective update set of each new vocabulary row and combines this vocabulary-aware signal with dense LoRA-level influence. Across Lean autoformalization and tool calling, SAVE improves target performance and related transfer while remaining competitive on broad retention benchmarks.
SAX: Advancing Video Diffusion Models for Sequential Action Execution
Haoyu Wang ⋅ Baorui Ma ⋅ Donglin Di ⋅ Shiliang Zhang
Video diffusion models have recently achieved remarkable visual fidelity. However, when handling complex instructions with multi-action sequential prompts, existing models frequently suffer from severe action omission or state freezing. We attribute this bottleneck to two fundamental flaws: i) semantic entanglement, where appearance and motion concepts are heavily intertwined within the text embedding space; and ii) temporal attention leakage, wherein the global nature of cross-attention along the temporal dimension biases the model towards the most visually salient action tokens throughout the entire video. To break this bottleneck, we propose SAX, a novel diffusion transformer framework designed for the synergistic generation of appearance and dynamic actions. First, we introduce a dual-stream attention mechanism that explicitly decouples dynamic action interactions from the static appearance generation process. Besides, to establish a precise temporal-action mapping, we design the Layer-adaptive Action Choreographer. By integrating a shared Global Temporal Table and an Action Ordinal Table, this module dynamically generates a temporal-action alignment map for each layer. Furthermore, to overcome the implicit convergence difficulties inherent in temporal-action alignment optimization, we innovatively employ a Vision-Language Model as a temporal-semantic supervisor. Through distribution distillation, we directly inject the VLM's powerful temporal-action grounding priors into the learning of the alignment maps. Extensive experiments on action generation and prediction demonstrate that SAX substantially outperforms existing models across core metrics.
SBNO : Schrödinger Bridge Neural Operator for the Forward and Inverse Problems under Missing Information
Jaehyeon Park ⋅ Inhyeok Jeong ⋅ Noseong Park
Partial differential equation (PDE) surrogate modeling often relies on either dense observations or access to governing equations/physics constraints. In practice, observations can be sparse and partial, and the underlying physics may be partially specified or unavailable. We introduce SBNO, an equation-free paired‑data diffusion bridge matching framework that learns conditional operators for both forward and inverse PDE problems under partial observations. SBNO learns a pair of time-dependent forward/backward drifts via diffusion Schrödinger bridge matching, using a unified endpoint coupling that preserves empirical input-solution pairing, enabling both directions to be handled within a single framework. At inference time, SBNO performs deterministic reconstruction by integrating the probability-flow ODE of the learned bridge, eliminating the need for inference-time hyperparameter tuning. On five PDE benchmarks with 5\% observations, SBNO achieves superior or comparable accuracy to baselines while demonstrating lower test-instance error dispersion. Furthermore, SBNO is at least $9.7\times$ faster and reduces GPU memory by at least $4.3\times$ compared to diffusion-based baselines.
Optimal Transport (OT) is a principled framework for comparing probability distributions, but its effectiveness depends critically on the ground metric, the cost used to compare observations. In high-dimensional settings, fixed ground metrics like the Euclidean distance can make OT distances reflect task-irrelevant variation rather than low-dimensional class structure, limiting OT in tasks such as classification and clustering. Supervised OT addresses this by learning a parameterized ground metric such that the induced OT distance reflects class structure. However, existing approaches on point clouds scale poorly with the number of observations. We propose a scalable supervised OT framework for Gaussian Mixture Models (GMMs) that lifts ground metric learning from observations to mixture components, replacing large point-cloud OT problems with smaller component-level problems whose transport costs admit closed forms. We define a learnable Wasserstein-type distance between GMMs using the Generalized Bures-Wasserstein (GBW) distance between Gaussian components as the ground metric, parameterized by a rectangular linear map that projects Gaussian components into a low-dimensional latent space inducing an OT distance between GMMs. For additional scalability, we introduce a diagonal approximation of the GBW metric, reducing the covariance computation between components to linear in the latent dimension. We bound the resulting approximation error by the mutual information between latent Gaussian variables and the Frobenius norm of the learned map, yielding a principled regularization strategy that also enables interpretation of latent axes. Across synthetic and scRNA-seq benchmarks, our method improves or matches classification and clustering over fixed-metric and supervised OT baselines on point clouds while substantially reducing OT distance computation time.
Scalable Token-Level Hallucination Detection in Large Language Models
Rui Min ⋅ Tianyu Pang ⋅ Chao Du ⋅ Minhao Cheng ⋅ Yi R. (May) Fung
Large language models (LLMs) have demonstrated remarkable capabilities, but they still frequently produce hallucinations. These hallucinations are difficult to detect in reasoning-intensive tasks, where the content appears coherent but contains errors like logical flaws and unreliable intermediate results. While step-level analysis is commonly used to detect internal hallucinations, it suffers from limited granularity and poor scalability due to its reliance on step segmentation. To address these limitations, we propose TokenHD, a holistic pipeline for training token-level hallucination detectors. Specifically, TokenHD consists of a scalable data engine for synthesizing large-scale hallucination annotations along with a training recipe featuring an importance-weighted strategy for robust model training. To systematically assess the detection performance, we also provide a rigorous evaluation protocol. Through training within TokenHD, our detector operates directly on free-form text to identify hallucinations, eliminating the need for predefined step segmentation or additional text reformatting. Our experiments show that even a small detector (0.6B) achieves substantial performance gains after training, surpassing much larger reasoning models (e.g., QwQ-32B), and detection performance scales consistently with model size from 0.6B to 8B. Finally, we show that our detector can generalize well across diverse practical scenarios and explore strategies to further enhance its cross-domain generalization capability.
A post-hoc Bayesian procedure should not change its evidence or uncertainty estimates under a function-preserving reparameterization. Standard empirical-Bayes (EB) Laplace approximation violates this principle for ReLU networks. Positive homogeneity induces continuous scale orbits of equivalent parameter vectors, yet the standard isotropic objective depends on the chosen representative through the prior norm and curvature log-determinant term. Equivalent ReLU parameterizations can thus yield different marginal likelihoods, selected prior precisions, posterior covariances, and linearized predictive uncertainties. We fix this by introducing $\gamma$-canonicalization, selecting a canonical representative of each ReLU scale orbit and fitting Laplace there. This is equivalent to transporting the isotropic Gaussian prior from the canonical representative back to the original coordinates, producing a precision field that transforms compatibly with the curvature. For every $\gamma\in[0,1]$, the generalized evidence and linearized predictive distribution are invariant under ReLU rescaling. The parameter $\gamma$ indexes a family of orbit-consistent priors and is selected jointly with the prior precision by marginal likelihood, making EB Laplace a well-defined procedure on ReLU scale-equivalence classes. Empirically, we show that standard EB-Laplace varies substantially across functionally identical ReLU rescalings, while $\gamma$-canonicalized Laplace collapses this variation to zero.
Scaling Causal Reasoning with Increasingly Complex Causal Simulators
Nicolás Astorga ⋅ Anita Kriz ⋅ Mihaela van der Schaar
Despite surpassing human performance across mathematics, coding, and other knowledge intensive tasks, large language models (LLMs) continue to struggle to causally reason. A core obstacle is the target data itself: causal systems are complex and often expressed in non-executable forms, and ground-truth answers to causal queries are inherently scarce. We introduce, CauSim, a framework that turns causal reasoning from a scarce-label problem into a scalable, supervised one. It constructs \textit{increasingly complex causal simulators}: executable structural causal models (SCMs), incrementally built by LLMs, that scale to globally complex systems while maintaining verifiable answers to any causal query. CauSim operates across representations by formalizing non-executable causal knowledge into code, allowing for data augmentation, and informalizing executable SCMs into natural language, enabling supervision in previously unsupervisable representations. We structure our research into two parts: (1) how to construct increasingly complex causal simulators, and (2) a systematic study of what CauSim enables, demonstrating generalization across representations, consistent gains from curriculum scaling and data volume, LLM self-improvement through self-generated simulators, and data augmentation via formalization of existing domain knowledge.
Sparse autoencoder (SAE) training is often bottlenecked by activation storage. We show that the stored activation buffer can be far smaller than the total SAE training budget: quality depends mainly on the number of \emph{unique} activations, not on whether every training token is fresh. After a short diversity-limited regime, additional fresh activations provide little benefit; replaying the same buffer preserves both reconstruction and interpretability performance. We capture this diversity--repetition tradeoff with a data-constrained scaling law and validate it across dictionary sizes, token budgets, Llama-3.1 and Qwen3 models, early/middle/late layers, BatchTopK/TopK/JumpReLU SAEs, and downstream metrics including EV, reconstruction MSE, feature absorption, automated interpretability, and sparse probing. The resulting buffer-sizing rule is simple: choose the smallest activation buffer that reaches the quality plateau, then use replay to spend the remaining training budget. In our experiments, this reduces activation storage by $8\text{--}64\times$ with negligible quality loss.
Scenes as Objects, Not Primitives : Instance-Structured 3D Tokenization from Unposed Views
Mijin Yoo ⋅ In Cho ⋅ Subin Jeon ⋅ Jiwoo Lee ⋅ Eunbyung Park ⋅ Seon Joo Kim
A 3D scene is understood through its objects, not the primitives that compose them. Yet feed-forward reconstruction methods output dense, unstructured sets of points or Gaussians, leaving object-level structure to be recovered after the fact. We propose a feed-forward framework that directly decomposes unposed multi-view images into instance-structured 3D token groups - compact object-centric units from which reconstruction, segmentation, and manipulation all follow. Each token group consists of an instance token that captures entity-level identity, paired with anchor tokens that model local geometry and appearance and decode into 3D Gaussians. This factorization separates what belongs together from how each part looks, making object instances a native interface of the representation rather than a derived product. The model is trained through differentiable rendering with joint reconstruction and segmentation supervision, without 3D annotations or test-time optimization. On indoor scene benchmarks, our feed-forward model outperforms per-scene optimized baselines in class-agnostic instance segmentation, while the same token groups support instance-level scene editing through direct token removal, translation, and insertion.
SceneScaffold: Active Scene-State Construction for Unified 3D Scene Understanding
Xiangqi Li ⋅ Libo Huang ⋅ Jiarui Zhao ⋅ Weilun Feng ⋅ Chuanguang Yang ⋅ Zhulin An ⋅ Yongjun Xu
Recent 3D large multimodal models (3D-LMMs) rely on a visual bottleneck to compress complex 3D scene evidence into a limited number of visual tokens compatible with large language models (LLMs). Current visual bottlenecks, however, often passively compress heterogeneous 3D evidence into a homogeneous object-centric token sequence, leaving the spatial organization of the scene under-represented. This under-representation forces the LLM to recover spatial relations from a flattened token sequence, leading to unstable reasoning in relation-intensive and spatially ambiguous scenes. To address this issue, we propose SceneScaffold, an active scene-state construction framework for unified 3D scene understanding. SceneScaffold reformulates the visual bottleneck from a passive feature compressor into an active scene organizer, constructing a role-aware spatial scaffold before language reasoning. Specifically, SceneScaffold organizes superpoint-level visual evidence into scene-state components with distinct structural roles: entity states preserve core object semantics, scene-frame states maintain spatial references via boundary and region anchors, relation states encode object-environment interaction cues, and a global summary provides compact context. Through this role-aware construction, SceneScaffold provides the LLM with a spatially organized scene representation before language reasoning. Experiments on unified 3D scene understanding tasks, including 3D visual grounding, question answering, and dense captioning, demonstrate the effectiveness of SceneScaffold, while diagnostic results further show its applicability to relation-intensive and spatially ambiguous cases.
SchemaPose: RGB-Based Category-Level Pose Estimation with Parametric Category Schema
Shuang Wu ⋅ Jiude Wei ⋅ Cewu Lu ⋅ Jianhua Sun
RGB-based category-level 9D pose estimation remains highly challenging because a monocular system must generalize across unseen object instances while jointly inferring 3D translation, 3D rotation, and 3D size. Unlike instance-level pose estimation, the category-level setting requires the network to learn both what is shared across objects within a category and how individual instances vary in structure. This challenge is amplified in the RGB-only setting, where pose cues, category-level commonality, and instance-specific variation are entangled in appearance without direct geometric constraints, making it difficult for existing frameworks to learn a representation that is truly aligned with the category-level estimation objective. To address this problem, we propose SchemaPose, a schema-guided framework for RGB-based category-level pose estimation. SchemaPose equips the estimation pipeline with an explicit parametric schema for modeling category-level regularities together with structured instance variation, and uses the inferred schema variables to guide pose-related prediction in an end-to-end query-based framework. The same formulation also supports template instantiation, synthetic annotation, and pose learning within a unified framework. Experiments on synthetic and real data show that SchemaPose achieves state-of-the-art performance among RGB-based methods for category-level 9D pose estimation.
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
Wai-Chung Kwan ⋅ Aryo Gema ⋅ Joshua O Leang ⋅ Pasquale Minervini
Self-play can train language models without external supervision. However, existing methods require rule-checkable answers, leaving open-ended tasks dependent on curated prompts or frontier-model judges. We introduce SCOPE, a data-free self-play framework for open-ended tasks that co-evolves two policies: a Challenger that generates document-grounded tasks, and a Solver that answers them through multi-turn retrieval. A frozen copy of the initial model serves as the self-judge, which writes task-specific rubrics from the source document and grades Solver responses against them. Across three 7-8B instruction-tuned models (Qwen2.5, Qwen3, OLMo-3), SCOPE improves open-ended performance by up to +10.4 points on eight benchmarks and matches or exceeds GRPOdata trained on ~9K curated prompts. Although trained only on open-ended tasks, SCOPE also improves held-out short-form QA by up to +13.8 points on seven held-out benchmarks, surpassing GRPOdata on all three models. Ablations show that co-evolving the Challenger is necessary to keep tasks near the Solver's frontier, that gains arise from improvements in both retrieval and synthesis with the relative contribution varying by task, and that rubric generation quality is the bottleneck for self-judging.
Screening Lipid Nanoparticles through Structure-Ratio Alignment
Yoonho Lee ⋅ Yunhak Oh ⋅ Hoyoung Choi ⋅ Chanyoung Park
Lipid Nanoparticles (LNPs) are widely used as delivery systems for nucleic acid therapeutics, where transfection efficiency is determined by both the identities of constituent lipid components and their composition ratios. While prior studies have focused on learning molecular representations for ionizable lipids, modeling how multiple components and their ratios jointly influence LNP performance remains underexplored. In this work, we propose STRATA, a framework that models molecule interaction between LNP components, which is known to contribute to LNP transfection efficiency. Our approach is built on two complementary views: (1) a ratio-centric view that captures interaction patterns induced by composition ratios through a transformer with a Ratio-induced Positional Embedding, and (2) a molecule-centric view that incorporates interaction-induced effects into structure-based molecule embeddings. By jointly training and aligning these views, our model integrates molecular structure and composition ratio within a unified framework that captures interaction-driven effects. Experiments demonstrate that our method improves prediction accuracy and generalization to unseen molecules and ratios, highlighting the effectiveness of our approach.
scShapeBench: Discovering geometry from high dimensional scRNAseq data
Andrew J Steindl ⋅ Joao Felipe Rocha ⋅ Brian T Di Bassinga ⋅ Zachary Warren ⋅ Matthew Scicluna ⋅ César Miguel Valdez Córdova ⋅ Shabarni Gupta ⋅ Leire Torices ⋅ Daniel Neumann ⋅ Timothy J Mann ⋅ Ihuan Gunawan ⋅ Dhananjay Bhaskar ⋅ John G Lock ⋅ Christine L Chaffer ⋅ Guy Wolf ⋅ Smita Krishnaswamy
High-dimensional point cloud data arise across many scientific domains, notably single-cell biology. The "shapes" or topologies of these datasets are informative of types of information that can be extracted from the datasets. For example clustered data admits the extraction of cell types or cell states in a static analysis of the datasets. Continuous trajectory structures admit continuous transition or trajectory analysis, while other shapes such as archetypal shapes admit continuum extraction with a range of cells spanning behaviors. While analysis pipelines exist, they often presuppose shape in data. For example, the standard Seurat pipeline combines UMAP visualization with Louvain clustering. This assumes clustered data. Tools like Monocle and Spade assume a tree-like shape, and flow-models like MIOFlow and Conditional Flow Matching are suitable for trajectories. Deciding which pipeline to apply to which part of the data is often the realm of bioinformaticians who visualize and qualitatively analyze the data before selecting one. However, with the advent of agentic AI scientists, it becomes important to automate data shape detection, particularly into categories that are relevant for downstream analysis pipelines. Towards this end we introduce \textsc{scShapeBench} a benchmark dataset comprising both synthetic and single-cell expert-annotated datasets that are meant for the task of shape detection. Synthetic datasets are sampled from a ground truth “skeleton graph” with variance. Real single-cell datasets are curated from a variety of sources and are annotated by experts, classifying four categories; clusters, single trajectory, multi-branches and archetypes. In addition, we provide a baseline method, scReebTower, to bridge the gap between data visualization and pipeline selection. scReebTower relies on the diffusion geometry to extract Reeb graphs. We provide new topology-aware metrics with which we evaluate scReebTower and existing methods PAGA and Mapper on synthetic data. On single-cell data we curate expert annotations of shapes, and showcase evaluations of methods. Our comparisons indicate scReebTower outperforming other baselines. Overall, our contributions span benchmarks, evaluation metrics, and a novel baseline method for automated shape detection in high-dimensional single-cell data.
scTrilemma: Balancing Identity, Invariance, and Reconstruction in Single-Cell Representation Learning
Yunhak Oh ⋅ Yoonho Lee ⋅ Junseok Lee ⋅ Namkyeong Lee ⋅ Sang-Yeon Hwang ⋅ Yinhua Piao ⋅ Hyomin Kim ⋅ SEONGHWAN KIM ⋅ Jaechang Lim ⋅ Woo Youn Kim ⋅ Sungsoo Ahn ⋅ Chanyoung Park
Single-cell RNA-seq representation learning is inherently label-free: cell identities, states, and tissue contexts are typically discovered through analysis rather than provided as target labels during training. As a result, the evaluated cell embedding must balance competing demands, which we refer to as the representation trilemma: it should preserve biological identity, avoid encoding collection effects as cell identity, and retain expression information needed for reconstruction. We introduce scTrilemma, a latent-bottleneck VAE that tests whether simple expression-derived routing can manage this trilemma without cell-level annotations, metadata-derived supervision, auxiliary representation losses, or specialized disentanglement modules. scTrilemma combines expression-gated gene encoding, cell-representation routing through the decoder, and pseudo-bulk prior conditioning under a single reconstruction objective. In release-based zero-shot evaluation on successive CZ CELLxGENE Census releases, scTrilemma improves batch-effect removal while preserving biological identity, retains marker-program and biologically meaningful tissue-context structure in the evaluated embedding, and better preserves differential-expression and pathway structure through reconstruction; ablations support burden allocation as the source of these gains.
SDAE: Semantic-Diversity-Aware Exploration for Efficient Reinforcement Learning in Large Language Models
Jianwen Sun ⋅ Wangzi Shi ⋅ Sannyuya Liu ⋅ Zhiming Wang ⋅ Fengpeng Yue ⋅ Luona Wei ⋅ Qian Wan
Reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models. However, as training progresses, models often converge to narrow solution paths, leading to diversity collapse. Existing methods alleviate this phenomenon through token-level entropy regularization or negative sample reinforcement, yet these works have treated responses as atomic units distinguished only by binary correctness labels, overlooking substantial semantic heterogeneity among responses under the same label—among incorrect responses, the degree and type vary significantly; among correct responses, both conventional and novel strategies exist. Uniformly reinforcing all correct responses leads to policy convergence, while uniformly penalizing all incorrect ones may suppress promising directions. Building on this insight, we propose Semantic Diversity-Aware Exploration (SDAE), which leverages geometric structures in semantic space to guide exploration. SDAE computes the semantic centroid of positive samples and modulates advantage functions based on each response's distance to this centroid, encouraging novel correctness, penalizing divergent errors, while preserving improvement potential in locally flawed responses. Experiments demonstrate that SDAE consistently outperforms strong baselines across competition-level reasoning benchmarks, achieving a 13.3% improvement in Pass@256 on AIME25. Our code is available at https://anonymous.4open.science/r/SDAEcode-3553.
SDFlow: Similarity-Driven Flow Matching for Time Series Generation
Wei Li ⋅ FENG SHIBO ⋅ Pengcheng Wu ⋅ Xingyu Gao ⋅ Min Wu ⋅ Peilin Zhao
Vector quantization (VQ) with autoregressive (AR) token modeling is a widely adopted and highly competitive paradigm for time-series generation. However, such models are fundamentally limited by exposure bias: during inference, errors can accumulate across sequential predictions, leading to pronounced quality degradation in long-horizon generation. To address this, we propose SDFlow (Similarity-Driven Flow Matching), a non-autoregressive framework that operates entirely in the frozen VQ latent space and enables parallel sequence generation via flow matching. We tackle three key challenges in making this transition: (1) eliminating exposure bias by replacing step-wise token prediction with a global transport map; (2) mitigating the high-dimensionality of VQ token spaces via a low-rank manifold decomposition with a learned anchor prior over the latent manifold; and (3) incorporating discrete supervision into continuous transport dynamics by introducing a categorical posterior over codebook indices within a variational flow-matching formulation. Extensive experiments show that SDFlow achieves state-of-the-art performance, improving Discriminative Score and substantially reducing Context-FID, particularly for challenging long-sequence generation. Moreover, SDFlow provides significant inference speedups over autoregressive baselines, offering both high fidelity and computational efficiency. Code is available at https://anonymous.4open.science/r/SDFlow-D6F3/
SEAR: Sample Efficient Action Chunking Reinforcement Learning
C F Maximilian Nagy ⋅ Onur Celik ⋅ Emiliyan Gospodinov ⋅ Florian Seligmann ⋅ Weiran Liao ⋅ Aryan Kaushik ⋅ Gerhard Neumann
Action chunking improves exploration and accelerates value propagation in long-horizon reinforcement learning, but naively applying off-policy methods to the temporally extended action space at reduced decision frequency offsets these gains, leading to poor sample efficiency. Existing action chunking methods address these issues through computationally expensive critic-only approaches or by relying on offline data. We introduce SEAR, a sample-efficient off-policy algorithm that enables online reinforcement learning with action chunks by addressing both challenges. To handle high-dimensional action sequences, SEAR employs a causal transformer critic trained with multi-horizon targets that provide a training signal for every prefix of an action chunk, effectively increasing the useful gradients per sample. To restore decision frequency, SEAR operates with a receding horizon and random replanning, ensuring uniform state coverage while combining the fast value propagation of large chunks with the reactivity of small ones. SEAR outperforms state-of-the-art online reinforcement learning methods including SimbaV2 on challenging Metaworld manipulation tasks. Beyond online RL, applying SEAR to the offline-to-online method QC improves its performance on the demanding OGBench cube-triple tasks.
Neural ordinary differential equations replace a finite layer stack by the solution map of a learned dynamical system. Thus the machine-learning primitive evaluated at inference time is the flow endpoint generated by the learned vector field, computed to requested accuracy. We study the complexity of this global forward pass when the vector field has standard compactness and Lipschitz certificates. Using second-order complexity theory, we identify the exact operator-level complexity of this local-to-global computation: scalar Neural-ODE endpoint inference is second-order polynomial-time equivalent to the classical Lipschitz IVP solution operator and is therefore FPSPACE$_2$-complete in the worst case. This complexity resides in the continuous-depth layer itself, independently of any particular numerical solver: even one fixed three-dimensional autonomous Neural-ODE block with $C^1$ local dynamics can define a PSPACE-hard input-output map. At every finite smoothness level $C^k$ with $k\geq2$, such autonomous blocks can define counting-hierarchy-hard maps. One scalar readout-gradient query recovers the forward endpoint, so the same complexity reaches an elementary training-gradient computation. Hence continuous-depth models and compressed-depth tied ResNets can realize worst-case global computational complexity even when local dynamics are cheap. As a tractable counterpart, fixed-dimensional Taylor-certified analytic vector fields admit polynomial-time endpoint inference under quantitative magnitude and radius certificates. These results show, based on the rigorous framework of second-order complexity theory, that the computational reliability of continuous-depth learning systems depends jointly on global flow structure, representation certificates, and the cost of local network evaluations.
Seeing the Unseen: Unified Visible–Invisible Motion for Physically Consistent Video Generation
Mingyang Bao ⋅ Fangda Ye ⋅ Xihao Chen ⋅ Bobo Li ⋅ Xiangtai Li ⋅ Shengqiong Wu
Recent advances in video generation have achieved impressive visual quality, yet they often fail to produce physically consistent dynamics due to the lack of explicit modeling of underlying physical processes. Existing approaches either rely on simulators with strong assumptions or use coarse semantic guidance, and further struggle to effectively inject physical priors into generation models. In this work, we propose Seeing the Unseen, a unified framework for physics-grounded video generation. Our approach decomposes the problem into two stages: learning physical dynamics and injecting them into video synthesis. First, we introduce a lightweight Particle-based Graph Dynamics Simulator (PGDS) that learns generalizable physical interactions from data and predicts plausible 3D motion trajectories. Second, we propose a Visible–Invisible Motion Field (VIMF) that captures both motion observable in the input frame and motion that emerges over time due to object dynamics. This representation enables more complete and structured motion guidance compared to conventional physical signals. By integrating these trajectories into a diffusion-based generator, our method produces videos that are both visually coherent and physically consistent. Extensive experiments demonstrate improved motion accuracy, interaction consistency, and reduced physical artifacts compared to strong baselines.
See, Read, Compare: Candidate-Aware Verification for Agent Test-Time Scaling
Xinyu Ye ⋅ Yongliang Wu ⋅ Xingyu Zhu ⋅ Peng Xia ⋅ Huaijin Wu ⋅ Rui Ye ⋅ Yehui Tang ⋅ Hao Xiong ⋅ Junchi Yan ⋅ Huaxiu Yao
Test-time scaling (TTS) improves LLM agent performance by sampling $K$ independent rollouts and selecting one with a verifier. In open-ended software engineering (SWE) agent settings, verification typically relies on code execution to obtain a direct correctness signal, but incurs substantial setup and testing cost. By contrast, execution-free verifiers offer a more scalable alternative that chooses from the submission alone. However, existing execution-free verifiers are largely submission-isolated: they fix rubrics before seeing the candidate batch, underuse the trajectory evidence already produced during agent rollouts, and rely on poorly calibrated pointwise scores for the final decision. As a result, they often miss the right candidate even when it lies within the pool. To address this gap, we propose a training-free and execution-free verifier for agent test-time scaling based on a See, Read, Compare (SRC) protocol, which treats the $K$ candidates as mutual context rather than independent inputs. The See stage combines task understanding with cross-candidate contrast to induce an instance-specific rubric. The Read stage extracts structured behavioral features, and compresses each trajectory into an evidence-grounded summary along several cognitive dimensions: localization, hypotheses, interventions, and validation. The Compare stage performs rubric-guided pairwise selection through a two-phase hybrid tournament, preserving relative judgment with only linear complexity calls. Experiments show that SRC outperforms all compared execution-free baselines across SWE-bench Verified generator pools and generalizes to heterogeneous model pools and Terminal-Bench 2.
See to Believe: Segmentation by Reasoning with Visual Evidence
Changsong Wen ⋅ Yanjie Wang ⋅ Zelin Peng ⋅ Yu Huang ⋅ Yaoming Wang ⋅ Xiaokang Yang ⋅ Wei Shen
Reasoning segmentation requires localizing objects from complex, implicit textual queries. Existing methods either rely on the MLLM to reason implicitly within hidden representations, or produce explicit but purely textual chain-of-thought that terminates in a coarse bounding box. In both cases, the reasoning process is not grounded in visual evidence at each step, leaving intermediate errors unverifiable, which can propagate to the final prediction. We propose Reasoning with Visual Anchors (ReVA), a framework where each chain-of-thought step is explicitly grounded to a specific image region. ReVA introduces anchor tokens into intermediate reasoning steps, each directly aligned with region-level features in the image encoder's latent space. As the model reasons step by step, each anchor attends to its corresponding image region, grounding inference in concrete visual evidence. To supervise the full reasoning chain, we further require each anchor to correctly transition to the next along the spatial relation described in the text. This progressively guides the model toward the target and produces the correct segmentation mask. To support training, we construct a large-scale Chain-of-Anchor dataset of over 245K quality-filtered samples from the RefCOCO series, each annotated with a spatially grounded chain-of-thought containing explicit anchor regions. The dataset is constructed by leveraging a strong MLLM to generate reasoning traces, followed by a three-stage verification pipeline ensuring spatial accuracy and reasoning quality. Experiments on reasoning and referring segmentation benchmarks demonstrate that ReVA achieves state-of-the-art performance (e.g., +5.4 gIoU and +6.7 cIoU on the ReasonSeg test set), with ablations showing that reasoning in visual evidence is key to the improvement.
Seg3DParts: Segmentation-Grounded Controllable Part-Level 3D Generation
Jiantao Lin ⋅ Meixi Chen ⋅ Yingjie Xu ⋅ Chenbo FU ⋅ Leyi Wu ⋅ Hao CHEN ⋅ Yinchuan Li ⋅ Yingcong Chen
Part-level 3D assets are essential for editing, reassembly, and interaction, yet recovering such structure from a single image remains challenging due to occlusion, ambiguous boundaries, and the need for coherent multi-part reasoning. Existing approaches struggle to achieve both controllable part-level generation and coherent multi-part structure, as part identity and spatial allocation are typically inferred implicitly. We present Seg3DParts, a segmentation-grounded framework for controllable part-level 3D generation from a single image. By treating segmentation as an explicit grounding signal, our method defines part identity during generation, enabling each component to be anchored to a corresponding image region. To ensure coherent assemblies, we introduce structured cross-part interaction that allows components to exchange global context throughout the generative process. As a result, Seg3DParts directly generates well-aligned part meshes in a shared canonical space without post-hoc alignment, supporting flexible and controllable decomposition. We further introduce PartObjectNet, a large-scale dataset with over 200K objects and 1M annotated parts. Experiments demonstrate that Seg3DParts achieves superior geometry quality, cross-part coherence, and part-level controllability over existing methods.
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
Zhangcheng Hou ⋅ TOMOAKI OHTSUKI
Radar-camera depth estimation must turn an ultra-sparse, all-weather, metric radar signal into a dense per-pixel depth map. Existing methods --- concatenation, confidence-aware gating, sparse supervision, graph-based extraction --- combine radar and image features outside the backbone's sequence operator, and even cross-modal Mamba variants leave the selection mechanism itself unimodal. We argue that the selection mechanism is the right place for radar to enter. We introduce Radar-Modulated Selection (RMS), a minimal and principled way to inject radar into Mamba's selective scan: radar modulates the scan from within, adding zero-initialised perturbations to the step size $\boldsymbol{\Delta}$ and readout $\mathbf{C}$ while leaving the input projection $\mathbf{B}$ and state dynamics $\mathbf{A}$ image-only. The construction is exactly equivalent to a pretrained image-only Mamba at initialisation, ensuring radar only influences the model where it improves accuracy. Two further properties follow that out-of-scan fusion cannot offer: linear-cost cross-modal coupling at every recurrence step, and a natural fallback to the image-only backbone when radar is absent. We deploy RMS in a Multi-View Scan Pyramid (MVSP) that matches the fusion operator to radar's spatial reach at each scale. SemoDepth achieves state-of-the-art performance on nuScenes, reducing MAE by 34.0%, 29.9%, and 29.9% over the previous best at 0-50, 0-70, and 0-80 m, while attaining the lowest single-frame latency (26.8 ms). A further ablation shows that out-of-scan feature blending adds no accuracy on top of RMS, providing empirical validation that in-scan selection can replace out-of-scan fusion.
Select Smarter, Not More? Prompt-Aware Evaluation Scheduling with Submodular Guarantees
haoyue liu ⋅ Xiaoyu Ma ⋅ Yiwen li ⋅ Zhichao Wang ⋅ Shuguang Cui ⋅ Xiaoying Tang
Prompt optimizers invest heavily in how to generate better prompts, yet pay almost no attention to which examples they use to judge them. This evaluation subset directly shapes every feedback signal that the optimizer receives, while existing methods either fix it before optimization begins (principled but agnostic to the evolving prompt population) or adapt it heuristically (flexible but unstable). We bridge this gap by constructing an online adaptive testing problem: Prompts are examinees, training examples are test items, and the scheduler selects items that best discriminate among the strongest candidates. We introduce POES (Prompt-Aware Online Evaluation Scheduling), whose monotone submodular objective combines discrimination, coverage, and bounded subset updates, yielding a $(1-1/e)$ cold-start guarantee and a warm-start tracking bound under bounded inter-round drift. Across a 35-task APO suite plus 57 MMLU subjects and 6 optimizer-model configurations, POES achieves the highest mean accuracy in every configuration, improving over the best baseline by +3.0 pp to +7.8 pp. Notably, the achieved gains concentrate on tasks where baselines still have headroom to improve, vanishing where methods saturate, suggesting that what you evaluate on matters most precisely when there is a signal to extract.
Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
Chirayu Nimonkar ⋅ Shlok Shah ⋅ Catherine Ji ⋅ Benjamin Eysenbach
For groups of autonomous agents to achieve a particular goal, they must engage in coordination and long-horizon reasoning. Rather than relying on complex reward functions and explicit cooperation mechanisms, we ask what minimal ingredients are required for effective coordination and exploration to emerge in multi-agent settings. We investigate this question through self-supervised goal-reaching, where agents aim to maximize the likelihood of visiting a goal state rather than maximizing a reward. Despite a sparse feedback signal, we present empirical results that show self-supervised goal-reaching techniques enable agents to learn from such feedback. On MARL benchmarks, self-supervised goal-reaching outperforms alternative approaches that have access to the same sparse reward signal. Furthermore, we empirically demonstrate that multi-agent self-supervised goal-reaching approaches can be more robust than single-agent strategies. While there is no explicit exploration mechanism, this approach explores nontrivial intermediate coordination strategies in sparse settings where alternative approaches fail to achieve a single success.
Semantic-Level Invariant Representation Learning for Cross-Hospital Clinical EEG Modeling
Kang Yin ⋅ Hye-Bin Shin ⋅ Seong-Whan Lee
Robust deployment of clinical electroencephalography (EEG) models requires source-trained models to generalize to unseen hospitals without target-domain calibration, hospital identifiers, or clinical reports at inference. This is challenging because acquisition protocols, hardware, montages, and patient populations perturb low-level EEG statistics, while clinical interpretation is organized around higher-level semantic concepts. We argue that this abstraction mismatch limits conventional signal-level invariant learning: aligning feature distributions alone cannot distinguish clinically irrelevant variation from diagnostic content. We propose Text-Guided Invariant Learning (TG-IL), a deployment-oriented framework that uses clinician-authored reports only during training as semantic anchors for EEG representation learning. TG-IL combines stochastic recording-level aggregation, a variational semantic information bottleneck, and prototype-guided EEG--report alignment to preserve report-derived clinical semantics while mitigating site-specific nuisance variation. At inference, all text-side modules are discarded and the model operates solely on EEG. On a three-hospital clinical EEG benchmark, TG-IL improves zero-shot cross-hospital generalization and worst-case robustness while reducing hospital-specific information in learned representations.
SemISP: Semantic-Consistent Diffusion for Cross-Camera RAW-to-sRGB Generation
Zhe Pang ⋅ Haina Qin ⋅ Ziyu Feng ⋅ Zewen Chen ⋅ Juan Wang ⋅ Bing Li ⋅ Weiming Hu
Cross-camera RAW-to-sRGB mapping seeks to convert smartphone RAW captures into DSLR-quality sRGB images, yet the large domain gap between linear RAW and nonlinear sRGB spaces makes this task highly challenging. Existing methods mainly condition on pixel-level supervision and local RAW cues, where content preservation and domain translation remain entangled in low-level measurements. We propose SemISP, a semantic-guided diffusion framework that uses scene semantics as a domain-invariant anchor for cross-domain generation. Specifically, we adopt a frozen DINOv3 as the sRGB semantic backbone and develop a Domain Alignment Adapter to extract sRGB-consistent semantic representations from RAW inputs through parameter-efficient adaptation. Because semantic priors and RAW signals operate at fundamentally different information levels, and the diffusion denoising process follows a coarse-to-fine trajectory, we further design a Time-aware Dual-stream Adaptive Modulation (TDAM) module that separately encodes both condition streams and uses the diffusion timestep to dynamically balance their contributions---emphasizing semantic anchoring at early stages and fine-grained RAW details at later ones. Experiments on two established cross-camera benchmarks show consistent improvements over prior methods in both fidelity and perceptual quality.
SenseBench: A Benchmark for Remote Sensing Low-Level Visual Perception and Description in Large Vision-Language Models
Chen Zhong ⋅ Xiao An ⋅ Jiaxing Sun ⋅ Zihan Gui ⋅ Guangyi Yang ⋅ Wei He
Low-level visual perception underpins reliable remote sensing (RS) image analysis, yet current image quality assessment (IQA) methods output uninterpretable scalar scores rather than characterizing physics-driven RS degradations, deviating markedly from the diagnostic needs of RS experts. While Vision-Language Models (VLMs) present a compelling alternative by delivering language-grounded IQA, their visual priors are heavily biased toward ground-level natural images. Consequently, whether VLMs can overcome this domain gap to perceive and articulate RS artifacts remains entirely unverified. To bridge this gap, we propose SenseBench, the first dedicated diagnostic benchmark for RS low-level visual perception and Description. Driven by a physics-based hierarchical taxonomy that unifies both no-reference and reference-based paradigms, SenseBench features over 10K meticulously curated instances across 6 major and 22 fine-grained RS degradation categories. Specifically, two complementary protocols are designed for evaluation: objective low-level visual perception and subjective diagnostic description. Comprehensive evaluation of 23 state-of-the-art VLMs reveals not only skewed domain priors and multi-distortion collapse, but also fluency illusion and a perception-description inversion effect. We hope SenseBench provides a robust evaluation testbed and high-quality diagnostic data to advance the development of RS VLMs in low-level perception. Code and datasets are available at here.
SENSE: Semantic Neural Speech Synthesis from Brain Dynamics via Spatial Graph Encoding
Jisoo Park ⋅ Seonghak Lee ⋅ Hyojin Park ⋅ Junseok Kwon
Reconstructing speech from non-invasive brain signals offers a promising pathway for restoring communication in individuals who are cognitively intact but unable to speak. Existing EEG-to-speech approaches formulate this task as \emph{acoustic reconstruction}, optimizing waveform fidelity while ignoring whether the generated speech preserves high-level semantic content. In this work, we revisit this formulation and argue that EEG signals carry not only acoustic but also semantic information. We identify two key limitations of prior methods: (1) the neglect of spatial relationships between EEG electrodes, and (2) the failure to exploit the semantic structure of the N400 paradigm, where congruent and incongruent trials reflect distinct semantic processing. We propose \textbf{SENSE} (\textbf{S}emantic-\textbf{E}EG \textbf{N}eural \textbf{S}peech Synth\textbf{E}sis), which combines a graph-based EEG encoder over electrode geometry with EEG Semantic Conditioning (ESC), aligning EEG to a pretrained semantic space using only congruent trials. On the N400 dataset, SENSE consistently outperforms prior methods on both acoustic and semantic metrics, and model-internal channel attribution suggests distributed reliance on auditory, sensorimotor, and centro-parietal regions, consistent with known speech-perception neuroscience. In the unseen-subject setting, SENSE trained on only two subjects already surpasses the strongest baseline trained on all eighteen subjects in word error rate, and matches it on acoustic metrics with as few as eight subjects.
SePO: Self-Evolving Prompt Agent for System Prompt Optimization
Wangcheng Tao ⋅ Han Wu ⋅ Weng-Fai Wong
System prompt optimization improves agent behavior without modifying the underlying model, yielding human-readable, model-agnostic instructions. Existing methods build a prompt agent that refines task agents' system prompts, yet leave the prompt agent's own system prompt hand-engineered and fixed. We propose Self-Evolving Prompt Optimization (SePO), which treats the prompt agent's own system prompt as an optimization target alongside task agents' system prompts. SePO adopts a self-referential design. A single prompt agent improves both task agents' system prompts and its own under an open-ended evolutionary search that maintains an archive of candidate prompts as stepping stones. Training proceeds in two stages: pre-training evolves the prompt agent on a multi-task pool, and fine-tuning then applies it to a target task. Across five benchmarks spanning math (AIME'25), abstract reasoning (ARC-AGI-1), graduate-level science (GPQA), code generation (MBPP), and logic puzzles (Sudoku), SePO consistently outperforms Manual-CoT, TextGrad, and MetaSPO, improving the average accuracy by $4.49$ points compared to Manual-CoT. The prompt optimization skill from pre-training also generalizes to tasks beyond the pre-training mixture, rather than memorizing per-task prompts.
Seq-LoRA: Sequential Bayesian Low-Rank Adaptation for Large Language Models
Boming Chen ⋅ Xiaojing Wang
Low-Rank Adaptation (LoRA) enables efficient fine-tuning of large language models, but adapted models often become overconfident, especially in low-data settings and under distribution shift. Existing uncertainty-aware LoRA methods typically construct a static posterior over a fixed adaptation set and do not model how local posterior evidence varies across structured subsets of the training data. We propose Seq-LoRA, a post-hoc Bayesian framework for sequentially aggregating slice-wise posterior evidence in LoRA space. Starting from a shared maximum a posteriori LoRA anchor, Seq-LoRA builds slice-wise Kronecker-factored quadratic surrogates, projects them into a shared curvature-informed low-dimensional subspace, rewrites the resulting latent quadratic forms as Gaussian pseudo-observations, and performs exact inference in an induced linear--Gaussian state-space model via Kalman filtering. Rather than relying on a particular curriculum order, Seq-LoRA couples heterogeneous slice-wise local posterior evidence through a random-walk prior, forming a terminal Bayesian posterior for prediction. Across the in-distribution ScienceQA test split and six out-of-distribution (OOD) reasoning targets, Seq-LoRA substantially reduces deterministic LoRA overconfidence under shift. Among compared methods, it achieves the best negative log-likelihood on five OOD targets and the best expected calibration error on four, while preserving most task accuracy.
Sequential Behavioral Watermarking for LLM Agents
Hyeseon An ⋅ Shinwoo Park ⋅ Dongsu Kim ⋅ Yo-Sub Han
LLM-based agents act through sequences of executable decisions, but their trajectories provide little evidence of which agent or policy produced them, making provenance, ownership, and unauthorized reuse difficult to establish from observed behavior alone. This motivates watermarking signals embedded directly into agent behavior rather than only into generated text, since text watermarking cannot capture the action-level decisions that define agent execution. Recent agent watermarking methods address this gap by moving the watermark from generated text to behavioral choices. However, by treating each action step as an independent trial, they overlook trajectory structure and become fragile when trajectories are perturbed, truncated, or observed without reliable alignment. We propose \texttt{SeqWM}, a sequential behavioral watermarking framework that embeds signals into history-conditioned transition patterns and verifies trajectories position-agnostically against random-key baselines. Experiments across diverse agent benchmarks and LLM backbones show that \texttt{SeqWM} consistently achieves reliable detection while preserving agent utility, and remains robust under trajectory corruption where round-indexed behavioral watermarks collapse. Our code is available at https://anonymous.4open.science/r/seqwm-0897.
SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction
Zhixiong Zhang ⋅ Yizhuo Li ⋅ Shuangrui Ding ⋅ Yuhang Zang ⋅ Shengyuan Ding ⋅ Long Xing ⋅ Yibin Wang ⋅ Qiaosheng Zhang ⋅ Jiaqi Wang
Referring segmentation grounds natural-language queries to pixel-level masks, but extending it to complex scenarios with multiple instances, cross-category groups, or open-ended target sets remains challenging. Previous Large Vision Language Model (LVLM)-based methods represent referred targets with one or more special tokens sequentially, treating multiple targets as separate outputs rather than a coherent set and offering little incentive to capture set-level properties such as completeness and mutual exclusivity. We reformulate open-ended referring segmentation as explicit set-level concept prediction and propose Set-Concept Segmentation (SetCon), which uses LVLM-generated natural-language concepts, instead of segmentation-specific tokens, as semantic conditions for joint mask-set decoding. A hierarchical semantic decomposition first predicts a shared set-level concept defining the target scope and then refines it into fine-grained concept groups aligned with target subsets. To support this, a two-stage annotation pipeline augments existing reasoning segmentation datasets with hierarchical semantic supervision (236k samples, 784k concept phrases). SetCon achieves state-of-the-art results on image benchmarks (+3.3 gIoU on gRefCOCO, +12.1 gIoU on MUSE), with margins that grow as the number of referred targets increases. The concept interface also transfers to video under a detect-and-track setting, yielding new state-of-the-art results on seven referring video benchmarks, including +10.9 J&F on MeViS and +12.4 J&F on Ref-SeCVOS. The code, model checkpoints, and dataset annotations will be released.
SH$^2$: A Mathematician-Curated benchmark for Assessing Research-level Math Capabilities of LLMs
Guijin Son ⋅ Seungone Kim ⋅ Catherine Arnett ⋅ Hyunwoo Ko ⋅ Hyein Lee ⋅ Hyeonah Kang ⋅ JIANG LONGXI ⋅ JIN YUN ⋅ JungYup Lee ⋅ Kyungmin Lee ⋅ Sam Y Kim ⋅ Sang Park ⋅ Seunghyeok Hong ⋅ SeungJae Lee ⋅ Seungyeop Yi ⋅ Sunhye Bok ⋅ Yonghoon Ji ⋅ Youngtaek Kim ⋅ Hanearl Jung ⋅ Akari Asai ⋅ Graham Neubig ⋅ Sean Welleck ⋅ Youngjae Yu ⋅ Akshelin R ⋅ Alexander B Ivanov ⋅ Muhammadjon Boboev ⋅ Christian Stump ⋅ Dmitrii Karp ⋅ Dohyun Kwon ⋅ DoYong Kwon ⋅ Duk-Soon Oh ⋅ Giovanni Resta ⋅ Greta Panova ⋅ Huiyun Noh ⋅ Hyungryul Baik ⋅ Hyungsun Bae ⋅ Jeewon Kim ⋅ Ji E Lee ⋅ Jiaqi Liu ⋅ Jieui Kang ⋅ Jimin Kim ⋅ Jon-Lark Kim ⋅ Junseo Yoon ⋅ Junwoo Jo ⋅ Kibeom Kim ⋅ Kiwoon Kwon ⋅ Mario Kummer ⋅ Max Mercer ⋅ Minjun Kim ⋅ Nahyun Lee ⋅ Ng Ze-An ⋅ Rafał M Łochowski ⋅ Raphaël Lachièze-Rey ⋅ Ruichen Zhang ⋅ Sejin Park ⋅ Seonguk Seo ⋅ Shin jaehoon ⋅ Taewoong Eom ⋅ Yeachan Park ⋅ Yongseok Jang ⋅ Youchan Oh ⋅ Zhaoyang Wang ⋅ Zoltán Kovács ⋅ Chae Young Han ⋅ Shinae Shin ⋅ Sunyoung SHIN
Mathematical reasoning benchmarks for large language models face three persistent limitations. First, datasets scraped from public competitions are vulnerable to training-data contamination and saturate quickly. Second, most existing benchmarks cover only part of the olympiad-to-research spectrum. Third, benchmarks that withhold problems indefinitely to mitigate leakage trade away transparency and reproducibility. We introduce \gls{soohak}, a benchmark of 1{,}141 newly authored problems by 86 mathematicians designed to address all three. \gls{soohak} consists of three difficulty splits, constructed adversarially against panels of small, mid-size, and frontier baseline LLMs, spanning olympiad-style problem solving through research-adjacent material; a separate 99-item Refusal split probes recognition of ill-posed prompts. The dataset is temporarily embargoed and evaluated by request, with a committed public release in late 2026. On \splitthree{}, Gemini-3-Pro, GPT-5, and Claude-Opus-4.5 reach Avg@3 of 30.39\%, 26.37\%, and 10.39\%, respectively; the best Pass@3 is only 44.12\%, and 129/340 problems remain unsolved by any closed model evaluated, indicating substantial room for improvement before LMs can meaningfully assist in mathematical research. By contrast, the \calibrationsplits{} are nearly saturated by both closed and open-weight systems. Notably, while GPT-OSS-120B, Kimi-2.5, and GLM-5 each reach Pass@3 between 87.4\% and 88.3\% on \splitone{}, within 6 percentage points of the leading closed model (Gemini-3-Pro, 93.85\%), they trail by over 20 percentage points on \splitthree{}. This pattern suggests open-weight pipelines are well targeted to olympiad-style benchmarks but transfer poorly to research-adjacent material. Finally, through a timed human study with 25 participants of diverse mathematical expertise, we confirm that the difficulty is genuinely mathematical rather than artifactual, with aggregated teams covering 50.6\% of a 79-problem sample.
Existing guarantees for misspecified kernelized bandit optimization pay for misspecification through kernel complexity: in generic offline bounds, the misspecification level $\varepsilon$ is multiplied by $\sqrt{d_\mathrm{eff}}$, where $d_\mathrm{eff}$ is the kernel effective dimension, while in online regret bounds, the corresponding penalty is $\sqrt{\gamma_n}\,n\varepsilon$, where $\gamma_n$ is the maximum information gain after $n$ rounds of interaction. In this work, we show that, for a large class of kernels, the misspecification amplification can be reduced to logarithmic or polylogarithmic growth. In the offline setting, we first prove high-probability simple-regret bounds whose misspecification term is governed by a spectral Lebesgue constant. This yields logarithmic amplification for one-dimensional monotone spectra and polylogarithmic amplification for multivariate Fourier-diagonal product kernels. In the online setting, we modify a domain-splitting algorithm and prove a cumulative regret bound of $\widetilde{\mathcal O}(\sqrt{\gamma_n n}+n\varepsilon)$ under mild localized eigendecay assumptions, removing the extra $\sqrt{\gamma_n}$ factor from the misspecification term. The common principle is localization: spectral localization controls the Lebesgue constant of the offline approximation operator, while domain splitting implements the spatial analogue of this mechanism in the online setting, preventing local misspecification errors from being amplified globally.
Shift-Aware Identity-Guided Latent Refinement for Referring Audio–Visual Segmentation
Kun Li ⋅ Sami S Brandt ⋅ Michael Yang
Referring Audio–Visual Segmentation (Ref-AVS) is an emergent multimodal task requiring the precise identification and segmentation of a specific object based on a natural language expression grounded in both visual and auditory cues. Unlike traditional AVS, which segments generic sound sources, or referring video object segmentation (Ref-VOS), which relies solely on visual–textual grounding, Ref-AVS necessitates fine-grained tri-modal interaction across vision, audio, and language. However, existing methods often fail to maintain consistency when subjected to target shift, a phenomenon where the model erroneously transfers its focus between objects due to transient occlusions, acoustic fading, or identity ambiguities. To address this, we propose a novel Shift-Aware Identity-guided Latent (SAIL) refinement framework. It introduces a dual-token strategy that decouples frame-level dynamics from a global semantic anchor to explicitly model target-shift cues. We design a shift-aware refinement module that rectifies drift by aligning tokens with domain-specific evidence. Finally, we leverage a decoder that integrates the refined embeddings with a continuous memory update to maintain spatio-temporal coherence and identity consistency throughout the video sequence. Extensive experiments conducted on two Ref-AVS benchmarks demonstrate the effectiveness of our proposed method, significantly outperforming existing AVS, Ref-VOS, and Ref-AVS approaches. The code will be released after publication.
SierpinskiCam: Camera-Controlled Video Retaking with Sierpinski Triangle Pattern Cues
Suttisak Wizadwongsa ⋅ Hyelin Nam ⋅ Supasorn Suwajanakorn ⋅ Jeong Joon Park
Generating novel renderings of a scene along user-defined camera trajectories from a single monocular video, dubbed video retaking, is a compelling but difficult problem in content creation and visual effects. Existing geometry-guided approaches reconstruct a 4D representation from the source video and render it along the target trajectory to condition video diffusion models. However, this guidance degrades as the target camera departs from the source trajectory, leaving newly revealed regions sparse or entirely missing. We propose SierpinskiCam, which addresses this limitation by augmenting geometry-based guidance with Sierpinski dome texture cues that contains rich trackable features even under large viewpoint changes. We further introduce a reference video conditioning mechanism that appends source-video tokens to the target-token sequence and separates the two streams with negative RoPE indices, enabling appearance grounding without architectural modification or per-video adaptation. Extensive experiments show that SierpinskiCam achieves significant gains in camera controllability, geometric consistency, and video quality across diverse and challenging retaking scenarios.
SimCT: Recovering Lost Supervision for Cross-Tokenizer On-Policy Distillation
Jie Sun ⋅ Mao Zheng ⋅ Mingyang Song ⋅ Qiyong Zhong ⋅ Yilin Cheng ⋅ Bichuan Feng ⋅ Pengfei Liu ⋅ Junfeng Fang ⋅ Xiang Wang
On-policy distillation (OPD) is a standard tool for transferring teacher behavior to a smaller student, but it implicitly assumes that teacher and student predictions are comparable token by token, an assumption that fails whenever the two models tokenize the same text differently. Under heterogeneous tokenizers, exact shared-token matching silently discards a large fraction of the teacher signal at precisely the positions where vocabularies disagree. We propose \textbf{\underline{Sim}ple \underline{C}ross-\underline{T}okenizer OPD (SimCT)}, which restores this signal by enlarging the supervision space: alongside shared tokens, SimCT compares teacher and student over short multi-token continuations that both tokenizers can realize, leaving the OPD loss form itself unchanged. We show that these units are the finest jointly tokenizable supervision interface, and that coarser alternatives remove teacher-student distinctions that are useful for on-policy learning. Across three heterogeneous teacher-student pairs on mathematical reasoning and code-generation benchmarks, SimCT shows consistent gains over shared-vocabulary OPD and representative cross-tokenizer baselines, with ablations confirming that the improvements come from recovering supervision discarded by exact shared-token matching. Code is available at \href{https://anonymous.4open.science/r/SimCT}{https://anonymous.4open.science/r/SimCT}.
Code evolution is a family of techniques that rely on large language models to search through possible computer programs by evolving existing code. While code evolution pipelines have shown impressive performance across domains, many are highly complex and are typically not compared to simpler alternatives. To prevent bad comparisons from stagnating the field, we propose two simple baselines and compare them to popular pipelines across three domains: finding better mathematical bounds, designing agentic scaffolds, and machine learning competitions. Surprisingly, the baselines match or outperform existing methods across all domains, prompting a closer look at which factors drive performance in different settings. For mathematical bounds, we find that the search space matters far more than the exact search algorithm, with both simple and complex methods performing similarly given the same search space. Thus, expanding the search space is what matters for improving performance, and does so across all pipelines. Moreover, different domain knowledge embedded in prompts can significantly affect the search's efficiency, potentially leading to overestimating a pipeline's sample efficiency. For automated agentic scaffold design, we show that typical high-variance evaluations lead to all automated search methods picking subpar scaffolds relative to a simple majority vote. To mitigate this, we propose improved evaluation procedures that reduce stochasticity while keeping the search economically feasible. Overall, these results indicate that further improving code evolution's performance depends more on the setup surrounding the search pipeline than on the pipeline itself. We hope these insights will enable developing new code evolution methods and help spur meaningful progress in the field.
Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image
Inhee Lee ⋅ Sangwon Baik ⋅ Hyeonwoo Kim ⋅ Hyunsoo Cha ⋅ Sungjoo Kim ⋅ Hanbyul Joo
Reconstructing interactive, simulation-ready 3D scenes from a single image is a critical bottleneck for robotic manipulation. While recent single-image lifters recover plausible per-object shapes, composing them yields scenes that collapse under physical simulation due to interpenetrating, hovering, or sinking objects. Existing physics-aware methods address this strictly as a post-hoc layout correction, leaving the underlying geometric errors unresolved. To address this, we introduce SimuScene, a compositional 3D reconstruction pipeline that puts physics in the loop of shape and layout estimation. Rather than using physics merely for layout cleanup, we utilize the physics engine as a diagnostic measurement tool during the generative process itself. By diagnostically simulating reconstructed objects under gravity, we convert penetration and support failures into quantitative correction signals that drive gravity-axis stretching and amodal shape resampling. This physics-informed feedback loop mitigates accumulated reconstruction errors and produces a stable, simulation-ready compositional 3D scene. Extensive experiments demonstrate state-of-the-art performance on physical stability and geometric alignment benchmarks. We further highlight SimuScene's utility by deploying reconstructed environments in humanoid control and robot-arm manipulation tasks.
Single-Pass Evidence Measurement for Interpretable and Uncertainty-Aware Multimodal Face Anti-Spoofing
Yingjie Ma ⋅ Haonan Wang ⋅ Xun Lin ⋅ Ruixin Zhang ⋅ Jingyun Zhang ⋅ Jun Wang ⋅ Rizen Guo ⋅ Shouhong Ding ⋅ Weicheng Xie ⋅ Linlin Shen ⋅ Zitong YU
Face anti-spoofing (FAS) is critical for protecting face recognition systems against presentation attacks in open environments. Beyond high cross-domain generalization, practical FAS deployments must provide interpretable and uncertainty-aware decisions when multimodal evidence is incomplete, degraded, or conflicting. However, current multimodal domain-generalized FAS methods focus primarily on feature fusion or domain alignment, compressing RGB, depth, and infrared cues into opaque logits. This makes it difficult to discern which modality drives the decision, whether modalities reinforce or suppress each other, and whether a given sample lies within the classifier's reliable operating regime. We introduce SEM-FAS, a quantum-inspired yet fully classical single-pass evidence measurement framework that jointly enhances generalization, interpretability, and reliability. SEM-FAS encodes each modality as a normalized complex-valued state whose amplitude captures evidence strength and whose phase encodes cross-modal interference; multiple coherent heads are then mixed into a density-matrix representation. A constrained POVM-like measurement head converts this representation into calibrated class probabilities while analytically decomposing every prediction into per-modality main effects and pairwise synergy/suppression terms. Simultaneously, three built-in uncertainty indicators are obtained without sampling overhead: predictive ambiguity, representation mixedness, and measurement-coverage mismatch. On standard multimodal DG-FAS benchmarks, SEM-FAS reduces the best baseline HTER by 2.18\% and improves AUC by 1.38\% under both complete and missing-modality protocols. Visual analyses further suggest that the learned modality effects, pairwise interactions, and uncertainty scores align with expected FAS behavior, even for hard samples, indicating that single-pass evidence measurement is a promising way to jointly enhance generalization, interpretability, and reliability in multimodal FAS.
Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models
Siqi Luo ⋅ Jianghan Shen ⋅ Yi Xin ⋅ Huayu Zheng ⋅ Haoxing Chen ⋅ Yan Tai ⋅ Yue Li ⋅ Junjun He ⋅ Yihao Liu ⋅ Guangtao Zhai ⋅ Yuewen Cao ⋅ Xiaohong Liu
Diffusion Multi-Modal Large Language Models (dMLLMs) are powerful for image generation, but optimizing them through reinforcement learning (RL) remains a major challenge. One primary difficulty is that a single image can be generated through many different unmasking sequences, which makes calculating importance ratios is often intractable. Additionally, existing methods tend to ignore the hierarchical generation process of dMLLMs, where early tokens define the global layout and later tokens focus on local details. By assigning uniform rewards to all tokens, these current methods fail to reflect the actual contribution of each token to the final image. To address these issues, we propose Hierarchical Token GRPO (HT-GRPO), which integrates this hierarchy directly into the policy optimization process. Our approach features a Sketch-Then-Paint training scheme that organizes updates into three distinct stages: global, structure, and refinement. We also use a prompt-conditioned estimator to calculate importance ratios starting from a fully masked state. Furthermore, we introduce a Hierarchical Credit Assignment mechanism that prioritizes key structural tokens to ensure accurate reward propagation. Experiments using two popular dMLLM backbones, MMaDA and Lumina-DiMOO, demonstrate that HT-GRPO achieves substantial gains on the GenEval and DPG benchmarks. Evaluations across six additional metrics confirm significant improvements in image quality, aesthetics, and human preference.
SkillGen: Verified Inference-Time Agent Skill Synthesis
Yuchen Ma ⋅ Yue Huang ⋅ Han Bao ⋅ Haomin Zhuang ⋅ Swadheen Shukla ⋅ Michel Galley ⋅ Xiangliang Zhang ⋅ Stefan Feuerriegel
Skills are a promising way to improve LLM agent capabilities without retraining, while keeping the added procedure reusable and controllable. However, high-quality skills are still largely written by hand. We introduce SkillGen, a multi-agent framework that synthesizes a single auditable skill from trajectories generated by a base agent. The output is a human-readable artifact that can be inspected before use. Rather than merely summarizing trajectories, SkillGen leverages contrastive induction over both successful and failed trajectories to identify reusable success patterns, recurring failure modes, and behaviors that appear in nearby successes but are missing from failures. SkillGen then generates candidate skills and iteratively refines the skill. A key novelty in SkillGen is that we model agent skills as interventions to empirically verify the net effect of skills on the overall performance. Specifically, we compare outcomes on the same instances with and without the skill, so that we account for both repairs (cases where the skill fixes a baseline failure) and regressions (cases where the skill breaks a baseline success). Across a broad range of agents and datasets, SkillGen consistently improves held-out performance, outperforms existing skill-generation baselines, and produces skills that transfer across models.
Skill-Level Effects in Behavioral Cloning: When Low-Skill Data Improves Performance
Saumik Narayanan ⋅ Kassa Korley ⋅ Siddhartha Sen ⋅ Chien-Ju Ho
Behavioral cloning (BC), which trains models from offline demonstrations, is a common approach in reinforcement learning settings. Prior work argues that BC requires expert demonstrations and performs poorly when trained on low-skill data. We challenge this assumption by showing that, in certain regimes, training on low-skill data can yield models that outperform those trained on high-skill data. Because low-skill data is often cheaper and more easily acquirable, this finding has important practical implications. To explain this result, we introduce the notion of fragility, characterizing how a policy's reward degrades under errors, and provide theoretical insights on how fragility can predict low-skill outperformance. We test our approach in a synthetic environment and MuJoCo and validate it using human data from chess and racing. Motivated by these findings, we connect our results to curriculum learning by structuring training according to demonstrator skill, rather than task difficulty as in standard curriculum design, and show that such skill-based curricula can improve performance relative to standard BC approaches.
Skipping Domain Shifts: Domain Memory Retention Enhanced Hyperspectral Single-Source Domain Generalization
Taiqin Chen ⋅ Hao Sha ⋅ Xiaochen Feng ⋅ Chunyu Chen ⋅ XiangYu Chen ⋅ Ke Chen ⋅ yongbing zhang
Hyperspectral single-source domain generalization aims to train a robust model capable of overcoming domain shifts using only one available source domain for cross-scene hyperspectral image (HSI) analysis. Existing methods typically adopt data augmentation to broaden decision boundaries or utilize style transfer techniques for test-time alignment. However, these studies either fail to completely bridge the domain gap or inadvertently disrupt the discriminative spectral information, leading to suboptimal generalization and degraded classification performance on the target domain. To overcome these limitations, we propose a novel method, termed the domain memory retention framework, which optimizes domain memory during training to explicitly encode the semantic structure of the source domain. During inference, the target features are projected to the source distribution by leveraging this memory, thereby skipping domain shifts and improving generalization performance across unseen scenes. Specifically, the memory is formulated as a codebook, consisting of multiple semantic codewords. We further develop a semantic richness constraint and a perturbation robustness constraint to enhance representation diversity while empowering domain invariance for the memory. Extensive experiments conducted on three remote sensing datasets demonstrate that the proposed method outperforms state-of-the-art methods.
SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routing
Hui Zhang ⋅ Yachao Yuan ⋅ Jiayun Wang ⋅ Yuanzhuo Li ⋅ Hongtao Wang ⋅ Yali Yuan
Fine-tuning-as-a-service enables users to adapt aligned large language models (LLMs) to specialized tasks, but malicious fine-tuning can erode refusal behavior while preserving task performance on legitimate inputs. We revisit recent layer-wise safety diagnostics and find that safety sensitivity is signed: scaling different layers can strengthen refusal, weaken it, or have little effect. Motivated by this observation, we propose SLDR, a post-fine-tuning defense based on Selective Layers Recovery and Dynamic Routing. SLDR trains a LoRA recovery adapter only on the layers with the maximum and minimum sensitivity scores in the signed spectrum, and uses representation-based dynamic routing inference to activate the adapter only for malicious queries. Across four model architectures, five downstream tasks, and four harmful benchmarks, SLDR substantially reduces harmful outputs while preserving downstream utility. On Llama3.1/SST2, SLDR reduces the average harmful score from 11.54 to 0.08 while maintaining downstream accuracy, and the harmful score remains near zero under poisoning ratios up to 0.9.
Sliced Wasserstein Meets Quantum Optics: Provable Wavefunctions Tomography with Scarce Noisy Measurements
Takahiro Kajisa
Quantum state tomography of pure non-Gaussian states is fundamentally limited by scarce and noisy data due to measurement-induced collapse and experimental drift. We introduce a novel reconstruction framework that parameterizes the wavefunction directly via an overparameterized neural network, trained using rank-1 factored gradient descent. We rigorously prove that this formulation's loss landscape is devoid of spurious local minima and strict saddles, guaranteeing all critical points are global minima. By utilizing the max-sliced Wasserstein-1 distance, our loss perfectly maps to the structure of homodyne measurements, structurally preventing gradient vanishing in low-shot regimes. On noisy two-mode cat states, our method achieves continuous fidelity improvements, successfully bypassing the noise-floor plateaus that limit standard iterative maximum-likelihood estimation.
Small Model Portfolios for Many Deployment Profiles: Submodular Coverage under Bundled Constraints
Ji Cheng
Deploying deep learning models across heterogeneous hardware requires each model to jointly satisfy a bundle of coupled constraints on accuracy, latency, memory, and energy. Maintaining a specialist for every deployment scenario is operationally prohibitive, so practitioners need a small portfolio of $K$ models that collectively covers many deployment profiles. Existing hardware-aware NAS and Pareto-based approaches target marginal objectives and do not directly address this downstream decision. We formalize it as Deployment Portfolio Selection over a fixed pre-measured archive, where a profile is covered only when a single retained model simultaneously satisfies its entire constraint bundle. The induced objective is weighted maximum coverage, which is NP-hard but monotone submodular, so greedy selection inherits a $(1-1/e)$ guarantee, and sample-then-optimize admits a finite-sample bound. Across an 18-family sweep on HW-NAS-Bench, direct bundled-coverage optimization is most valuable in the small-portfolio, low-coverability regime: gains concentrate on the hardest profiles, are most robust at $K = 3$, and persist under three non-uniform demand models. Complementary analyses of remeasurement noise, runtime, and compact feasibility storage clarify when bundled coverage is preferable to scalarized or Pareto-based alternatives.
SMMBench: A Benchmark for Source-Distributed Multimodal Agent Memory
Huacan Chai ⋅ YUKAI WANG ⋅ Yingxuan Yang ⋅ Dan Peng ⋅ Yuanyi Song ⋅ Zhihui Fu ⋅ Jun Wang ⋅ Weiwen Liu ⋅ Jianghao Lin ⋅ Weinan Zhang
Existing benchmarks for multimodal memory reasoning largely evaluate systems within pre-assembled contexts, but under-evaluate whether agents can use evidence distributed across independently originated sources. We argue that source-distributed memory composition is an important and under-examined bottleneck in multimodal agent memory, especially when relevant evidence is fragmented across heterogeneous artifacts such as conversations, profiles, screenshots, tables, images, and documents. To address this gap, we introduce Source-distributed Multimodal Memory Benchmark, which measures whether agents can retrieve, align, and compose multimodal evidence scattered across multiple sources rather than reason within a single curated context. SMMBench evaluates four core capabilities: (1) cross-source multimodal reasoning; (2) conflict resolution; (3) preference reasoning; (4) memory-grounded action prediction. The benchmark contains $1,877$ samples grounded in $264$ sources. Experiments on representative memory-style and retrieval-based baselines show that current systems still struggle on these capabilities, positioning source-distributed multimodal memory as an important and still under-evaluated challenge for multimodal agents. Our data and code are available at https://anonymous.4open.science/r/NeurIPSBenchmarksW8kP6gV/.
Social Choice Foundations for Simulation-Augmented Generation
Sonja Kraiczy ⋅ Smitha Milli ⋅ Ratip Emin Berker ⋅ Avinandan Bose ⋅ Brandon Amos ⋅ Jamelle Watson-Daniels ⋅ Maximilian Nickel ⋅ Edith Elkind ⋅ Ariel Procaccia
Users increasingly turn to AI systems for normative assistance—guidance on what one ought to do or think—yet models are often opaque about whose viewpoints they represent. A promising approach is simulation-augmented generation (SAGE), which involves querying generative simulations of individuals in a target population at inference time, soliciting their open-ended judgments, and synthesizing them into a response while transparently reporting whose viewpoints are reflected. However, inference-time simulation raises acute scalability constraints. Since the key benefit of simulation is improved representativeness, the core challenge is scaling simulation without sacrificing representation. We introduce the first formalization of this problem, grounded in proportional clustering concepts from social choice theory. We prove that to represent a population of $m$ humans, we need only create $n \ll m$ simulations of them, and need only dynamically query $k \ll n$ of those simulations at inference time, while still maintaining approximate proportional representation guarantees for the full population. We empirically validate that our inference-time algorithm yields better representation--efficiency trade-offs than baseline approaches.
Social Interaction Breaks Replicate Independence in Controlled-Clone LLM Agents
Yandan Zheng ⋅ Haoran Luo ⋅ Anh Tuan Luu
Multi-agent LLM evaluations and large-scale social simulations often treat multiple copies of the same model as comparable replicates after they interact. We test this assumption with SocCogEval, a controlled-clone assay for Social Cognitive Transfer (SCT), and find that it fails. Agents share the same base model, system prompt, and decoding parameters. When briefs are assigned, the brief is the only initial difference. On Qwen3-122B-FP8, the matched interaction contrast holds asymmetric briefs fixed and toggles 30 rounds of interaction. It yields stable off-topic BFI-10 response-profile drift on two topically disjoint scenarios. The effect is Delta =+0.92 on Greenfield (95% CI [+0.61,+1.22], g =+1.07) and Delta =+ 0.91 on Riverton ([+0.60,+1.24], g =+ 0.96, about 0.18 Likert points per BFI factor. Brief-only controls do not explain the drift. A verbatim-history ablation with SRM extraction disabled remains positive, so typed memory extraction is not the source of the signal. Five additional model variants provide positive but uneven breadth. Five of ten non-primary scenario cells are positive at alpha = 0.05, and no variant shows a significant reversal. The result limits iid-clone assumptions in multi-agent evaluation: interacting clones are not independent replicates unless paired no-interaction controls show otherwise.
SoDeArena: A Socially-Situated Reasoning Benchmark for Large Language Models
Senhao Yang ⋅ Ziling Yuan ⋅ Xiaoran Yang ⋅ Yaqing Wang ⋅ Runpeng Xie ⋅ Chen Huang ⋅ Lexiang Wang ⋅ Shuang Xu ⋅ Yiqin Yang ⋅ Bo Xu
In recent years, Large Language Models (LLMs) have demonstrated remarkable performance in tasks such as mathematical reasoning and code generation. However, as LLMs are increasingly deployed as autonomous agents, their reasoning capabilities in dynamic, human-like social interactions remain largely unexplored. To bridge this gap, we propose the benchmark SoDeArena—Social Deduction Arena, designed to evaluate the socially-situated reasoning capabilities of LLMs through a diverse set of interactive games. Through large-scale multi-agent simulations and human interaction experiments, we have identified three key insights for improving current LLMs: (1) the overall proficiency of models in social deduction is fundamentally bottlenecked by model scale and reasoning capabilities; (2) models lack proactive intent management, defaulting instead to strategic passivity and output conformity during multi-agent interactions; and (3) LLMs exhibit dual behavioral biases, manifesting as a rigid safety alignment in adversarial roles and a pathological reliance on shallow positional patterns over robust reasoning processes. These findings point to potential directions for the future development of AI systems with robust socially-situated reasoning capabilities.
Soft Contamination Means Benchmarks Test Shallow Generalization
Ari Spiesberger ⋅ Juan J Vazquez ⋅ Nicky Pochinkov ⋅ Tomáš Gavenčiak ⋅ Peli Grietzer ⋅ Gavin Leech ⋅ Nandi Schoots
If LLM training data is polluted with benchmark test data, then benchmark performance is only a biased estimate of out-of-distribution (OOD) generalization. Typical 'decontamination' filters use $n$-gram matching which fail to detect 'semantic' duplicates: sentences with equivalent (or near-equivalent) content that are not close in string space. We study this 'soft' contamination of training data by semantic duplicates. We embed the Olmo 3 training corpus and find that: (1) contamination remains widespread: we observe semantic duplicates for 76\% of the CodeForces test set and 100\% of MBPP's; (2) training on semantic duplicates of benchmark data improves benchmark performance by -1–22pp in controlled experiments and 18–22pp in ecological finetuning, depending on training regime and type; and (3) finetuning on duplicates of benchmark datapoints improves performance by -1–22pp (controlled) and 13–19pp (ecological) on truly-held-out datapoints from the same benchmark. The generalization we find is 'shallow': it is limited to (2) and (3), and does not typically extend to related benchmarks. We replicate on Olmo 3, Qwen3, and Qwen3.5. We thus argue that recent benchmark gains are confounded: the prevalence of soft contamination means gains reflect both genuine capability improvements and the accumulation of *effective* test data in growing training corpora.
SOLAR: AI-Powered Speed-of-Light Performance Analysis
Qijing Huang ⋅ Sana Damani ⋅ Zhifan Ye ⋅ Athinagoras Skiadopoulos ⋅ Siva Kumar Sastry Hari ⋅ Jason Clemons ⋅ Sahil Modi ⋅ Jingquan Wang ⋅ Aditya Kane ⋅ Edward C Lin ⋅ Humphrey Shi ⋅ Christos Kozyrakis
How fast could a deep-learning model run on target hardware, and how far is today’s implementation from that limit? These questions are central to software, hardware, and algorithm optimizations. Speed-of-Light (SOL) analysis answers them by computing a workload’s theoretical minimum execution time on a given architecture. Yet deriving SOL bounds remains manual, error-prone, and disconnected from rapid model development. To close this gap, we introduce SOLAR, a framework that automatically derives validated SOL bounds from Pytorch and JAX source code. SOLAR leverages both generative and deterministic components in its flow: an LLM frontend translates any source programs into an executable Affine Loop IR, validated by output comparison; a deterministic flow lifts the IR into an einsum graph; and an analytical backend computes unfused, fused, and cache-aware SOL bounds. SOLAR provides comprehensive operator and language coverage, produces validated bounds with zero observed SOL violations, and offers multi-fidelity analysis that tightens bounds and surfaces optimization insights. We evaluate SOLAR across KernelBench, JAX/Flax models, and robotics workloads. These experiments demonstrate four use cases: headroom analysis at multiple fidelity levels, identifying optimization opportunities, cross-platform exploration, and inverse-roofline hardware provisioning.
Solver-as-Teacher: Solver-Guided On-Policy Post-Training Framework for PDE Foundation Models
Hao Wei ⋅ Shangzhe Li ⋅ Nils Thuerey
PDE foundation models (PDE-FMs) promise general-purpose surrogate simulation, yet deploying a pretrained model to a new physical regime inevitably introduces distribution mismatch, demanding task-specific post-training. The dominant approach, supervised fine-tuning (SFT), is fundamentally off-policy: the model trains on pre-generated states but is evaluated autoregressively on its own. In chaotic regimes, this mismatch is catastrophic as per-step errors compound exponentially along the Lyapunov spectrum, and exact pointwise trajectory matching beyond the predictability horizon becomes an ill-posed training signal. To resolve this, we introduce Solver-as-Teacher (SaT), an on-policy post-training framework for PDE-FMs. During training, the student rolls out autoregressively; a numerical solver branches multi-step correction targets from each student-visited state; and the student is updated against these solver corrections, learning to correct its own rollout dynamics. We instantiate SaT in a unified post-training suite with SFT, physics-informed temporal alignment, and learned-teacher on-policy distillation. Across seven 2D/3D PDE benchmarks and six pretrained PDE-FMs, SaT variants achieve the lowest 50-step rollout RMSE in 23 of 24 in-distribution model--task pairs. The best SaT variant reduces RMSE over SFT up to 90.0\%, and SaT gives the lowest RMSE on all seven OOD parameter-shift benchmarks.
Verification methods aim at mathematically proving desirable properties of neural networks, such as robustness to adversarial perturbations. A verifier is sound if and only if it never claims that a neural network has the desired property when it does not. It was shown recently that none of the currently known verifiers that are claimed to be sound are guaranteed to be sound when considering the deployed version of the verified network. Due to this, all the known verifiers are vulnerable to certain backdoor attacks, where an adversarial network passes verification, but in reality, it exhibits adversarial behavior in specific deployment environments. So far, it has been suspected that sound verification is prohibitively expensive if we wish to verify all possible executions—including parallel and stochastic ones—in deployment. We show that efficient and practically sound verifiers can be designed using a bound on the so-called backward error. Using this bounding technique, we propose two verifiers: one based on interval bound propagation, and one using symbolic propagation. Both verifiers are proven to remain sound even if the deployment environment randomly selects a valid expression tree (an ordering and parenthesizing of the arithmetic operations) to compute the network, and even if numeric underflow or overflow occurs in any valid expression tree. This is especially interesting in the case of symbolic propagation, where the expression tree used by the verifier is completely different from the one used in deployment. We demonstrate empirically that our techniques introduce only a limited performance overhead while detecting all the known verifier backdoor attacks.
SPA-Q: Structure-Preserving Adaptive Post-Training Quantization for Monocular Depth Estimation
Jaemin Choi ⋅ Jincheol Yang ⋅ Nahyun Lim ⋅ Yun-Seong Jeong ⋅ Matti Zinke ⋅ Hyunwoo Yu ⋅ Suk-Ju Kang
Monocular depth estimation (MDE) has advanced rapidly with the emergence of foundation models such as Depth Anything. Their transformer-based architectures provide strong generalization across diverse scenes and domains, but also incur high computational and memory cost, making efficient deployment challenging. Post-Training Quantization (PTQ) provides an efficient and practical solution for model compression, yet low-bit PTQ remains challenging for MDE models. When applying PTQ to Depth Anything, we identify two key challenges: (i) independently minimizing the quantization error of the query and key projections fails to preserve the attention maps induced by their interaction, and (ii) quantization errors accumulate across layers, resulting in intermediate feature distribution shifts. To address these issues, we propose SPA-Q, a Structure-Preserving Adaptive PTQ framework for MDE, consisting of two main components. First, Attention-Preserving Calibration (APC) determines query and key quantization parameters by matching the full-precision attention distribution. Second, Channel-Wise Distribution Alignment (CWDA) learns channel-wise affine transformations to mitigate quantization-induced distribution shifts, and the learned parameters are absorbed into the weights after training. Experimental results show that SPA-Q consistently outperforms existing PTQ methods under 4-bit quantization, achieving an average 25.9\% reduction in AbsRel and a 17.3\% improvement in $\delta_1$ across NYUv2 and KITTI datasets.
Sparse Layers are Critical to Scaling Looped Language Models
Ryan Lee ⋅ Jacob Biloki ⋅ Edward J Hu ⋅ Jonathan May
Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries. However, looped models do not scale as favorably as standard transformers with unique layers. We compare standard and Mixture-of-Experts (MoE) transformers, with and without looping, and find two main results. First, we find Looped-MoE models scale better than the standard baseline while dense looped models do not. We trace this to routing divergence between loops: in Looped-MoE models, different experts are activated on each pass through the same shared layers, recovering expressivity without additional parameters. Our second finding is that looped models have better compute-quality trade-offs with early exits than standard models. Because each loop ends with the same layers that produce the final output, loop boundaries are superior exit points, as confirmed by faster output convergence at these points. In sum, we provide a clear direction for scaling looped models: a Looped-MoE model with early exits can not only beat standard transformers at scale, but also enable significant memory and inference savings with minimal degradation in quality.
Sparse Repair in Reasoning Traces: A Structural View of Test-Time Inference
Zhuorui Zhang ⋅ Jianqi Yan ⋅ Shanshan Feng ⋅ Fan LI
Inference-time reasoning improves when models spend more computation through repeated sampling, search, or verification. Existing scaling strategies usually allocate this computation across prompts, candidate answers, or whole trajectories. We study a different structure: how the repair value of additional inference is distributed within a single generated reasoning trace. Across mathematical, scientific, coding, and knowledge-intensive reasoning tasks, recoverable utility is sharply concentrated: in oracle replay analysis over coarse trace windows, the top-three windows capture 83--90\% of replay-recoverable utility. This concentration is not fully explained by position, length, or local uncertainty alone. Building on this finding, we introduce \textbf{LEVER}, a sparse local-intervention method for test-time reasoning that separates counterfactual discovery from deployable compute allocation. It uses local replay to estimate window-level recoverable utility, then distills this offline signal into a lightweight selector. At test time, the selector uses only online-safe trace features to decide which traces to activate and which windows receive additional rollouts. At a average 1.31$\times$ realized completion-token cost, LEVER improves average performance by about 4.2 points over base decoding. Across three base models and four task families, LEVER obtains the best result on 11 of 12 model--task pairs at the primary operating points, while using less compute than full-trajectory sampling baselines. These results suggest that effective test-time reasoning depends not only on how much computation is spent, but on whether it reaches the local commitments where repair value is highest.
Spatially-Grounded Long Video Generation with Self Geometry Forcing
Chenguo Lin ⋅ Panwang Pan ⋅ Bowen Xue ⋅ Ruijie Lu ⋅ Yuchen Lin ⋅ Chenxin Li ⋅ Brandon Feng ⋅ Yadong Mu
Recent video generative models can produce photorealistic short clips, but struggle with long-horizon generation under dynamic camera control, often suffering from viewpoint drift, geometric inconsistency, and error accumulation. We propose Joint Appearance-Geometry Diffusion Transformer (JAG), a spatially-grounded causal video diffusion model that explicitly integrates 3D perception for long and consistent video generation. It adopts a novel dual-branch causal architecture that jointly models visual appearance and scene geometry. Building upon this formulation, we introduce Self Geometry Forcing, a geometry-aware distillation scheme that conditions current generation on self-predicted geometric structures as persistent long-term memory and imposes self-verified geometric rewards on intermediate rollouts, effectively mitigating error accumulation over long horizons. Extensive experiments show that JAG can generate long videos with superior coherent appearance and geometry at constant per-frame computation, improving long-horizon controllability, geometric consistency, and visual fidelity over previous state-of-the-art methods.
Spatial Representation Distillation and Knowledge Routing for Vision-Language-Action Models
Jungin Park ⋅ Chaoran Zhu ⋅ Changjae Oh ⋅ Kwanghoon Sohn
Understanding the 3D physical world has emerged as an essential component in vision-language-action models (VLAs), enabling robots to perform several challenging tasks by unifying \textit{spatial} perception, language, and action. Existing approaches typically either rely on explicit 3D inputs, such as depth maps or point clouds, or implicitly inject spatial knowledge into visual representations. However, explicit 3D inputs are often brittle in practice due to sensor noise, hardware heterogeneity, and incomplete depth coverage, while implicit alignment can distort the original 2D visual representation and weaken generic scene understanding. In this paper, we present SPARK-VLA, a simple yet effective framework to allow VLAs to implicitly possess generic 2D and spatial 3D visual comprehension within a unified visual backbone. SPARK-VLA consists of two key components: (1) spatial representation distillation, which transfers 3D understanding from a frozen 3D expert into a LoRA-adapted visual branch by aligning attention responses while preserving the original visual pathway, and (2) a lightweight knowledge router, which selectively forwards action-relevant visual tokens to the language model using routing supervision derived from action-loss gradients. Experiments in both simulation and real-world environments show that SPARK-VLA, with only a 0.5B language backbone and 50M trainable parameters, achieves performance comparable to or better than substantially larger VLAs.
Spectral-Aligned Pruning for Universal Error-Correcting Code Transformers
Sanghyeon Cho ⋅ Taewoo Park ⋅ Seong-Joon Park ⋅ Dae-Young Yun ⋅ Hee-Youl Kwak ⋅ Sang-Hyo Kim ⋅ Yongjune Kim
Universal channel decoders based on transformers--such as the Foundation Error Correction Code Transformer (FECCT)--achieve competitive decoding performance across diverse code families with a single shared backbone, optionally followed by code-specific finetuning. However, the high computational complexity and large parameter footprint of FECCT present substantial obstacles to practical deployment. To address these challenges, we investigate structured pruning for FECCT and propose Spectral-Aligned Pruning (SAP), a structure-aware framework that enables cross-code reuse of structured pruning masks by leveraging the spectrum of the corresponding bipartite graph. SAP is grounded in classical graph analysis of codes: the two algebraically largest adjacency eigenvalues provide compact spectral proxies for degree scale, expansion ratio, and minimum-distance lower bounds. These quantities are directly relevant to decoding performance: degree scale reflects how densely codeword bits and parity checks are connected; expansion ratio influences how information propagates across the bipartite graph; and minimum distance characterizes codeword separation. Based on this connection, SAP uses these two leading eigenvalues as a lightweight code signature for pruning-mask retrieval. Empirically, this two-dimensional signature yields stable library selection equivalent to higher-dimensional spectral signatures in our evaluation. After pruning, SAP performs per-code recovery via parameter-efficient low-rank adaptation (LoRA), enabling a shared pruned backbone while storing only small code-specific adapter parameters. Experiments across diverse codes show that SAP achieves decoding performance comparable to dedicated per-code pruning, while enabling substantial reductions in computational cost and model memory footprint through kernel-level structured pruning.
Spectral Estimation with Deformed Decompression
Siavash Ameli ⋅ Chris van der Heide ⋅ Liam Hodgkinson ⋅ Michael Mahoney
Sample covariance matrices are fundamental objects in machine learning and statistics, with valuable information encoded into their eigenvalues. When treating modern problems, the scale of these matrices can become prohibitive for tractable computation. This creates a critical need for "spectral decompression": a method for inferring the eigenspectrum of a very large-scale model using only information from its smaller realizations. Existing methods fall into two main categories: those relying on spectral inversion; and those that rely on Free Decompression techniques. On one hand, spectral inversion is a fundamentally ill-posed problem that is notoriously unstable, often failing to yield reliable results in finite-size settings. On the other hand, while Free Decompression was recently proposed to address this scaling, its reliance on the Nica-Speicher framework restricts its use to unitarily invariant ensembles, a symmetry condition that subsampled covariance matrices fail to satisfy. In this work, we introduce Deformed Decompression (DD), a novel framework for the stable, forward-extrapolation of spectral densities. By leveraging properties of the companion matrix, DD recasts the spectral decompression problem as a directed evolution of the Stieltjes transform of the companion. This approach bypasses the instabilities of traditional inversion and provides a robust methodology for scaling covariance matrices and even their generalized variants (including Hessian matrices) that lack unitary invariance. We demonstrate the utility of our framework through experiments on a variety of random matrix examples, showing that DD successfully captures the complex spectral signatures of large-scale systems.
Spectral Identifiability for World Models: Polynomial Projectors, Resolvent Stability, and a Krylov Bottleneck
Phan Quoc Hung Mai ⋅ Duc H Nguyen ⋅ Luong Doan ⋅ Ngoc Mai Vu ⋅ Khanh N Quoc ⋅ Nhung Duong ⋅ Naeem Ul Islam ⋅ Tuan Do
World models are central to model-based reinforcement learning and planning, but their latent dimensions, architectural spectra, and downstream behavior remain poorly understood. In this work, we develop a spectral theory of linear world models in three settings. First, for shared multi-horizon prediction in linear dynamical systems, we identify the finite-horizon block-Krylov operator $\mathcal K_H$ as the exact latent bottleneck and give a perturbation condition for recovering its empirical rank elbow. Second, for stable diagonal state-space models, we prove a Rademacher complexity bound governed by the spectral memory energy $\sum_j L_j(T)^2$, rather than the parameter count alone. Third, for jointly trained diagonal latent dynamics, we show that disjoint factor spectra imply factor-aligned latent coordinates, with a Riesz-projector stability extension for approximate recovery and a local gradient-flow companion. The supporting results connect these spectral quantities to compositional OOD prediction, rollout envelopes, and prediction-to-control mismatch. Further, generated-system experiments verify the predicted spectral behavior and illustrate why pixel prediction and control returns can rank world models differently.
Spectrally Decomposed Equivariant Graph Neural Networks for Interatomic Potentials
Junfu Tan ⋅ Mengfan Wu ⋅ Yu Zhu ⋅ Yu Shi ⋅ Yifang Qin ⋅ Chang Liu ⋅ Bonan Zhu ⋅ Ziheng Lu
Machine Learning Interatomic Potentials (MLIPs) based on Equivariant Graph Neural Networks (EGNNs) have greatly advanced the quantitative simulation of atomic systems. However, accurately resolving fine-grained structural geometry remains a critical challenge. In this work, we demonstrate through spectral analysis that the reliance of modern EGNNs on a single macroscopic cutoff radius inherently acts as a spatial low-pass filter. This structural bottleneck induces radial spectral confusion, suppressing the network's capability for fine-grained geometric modeling. To overcome this limitation, we propose the Multi-Cutoff Spectral Decomposition (MCSD) mechanism. As a universal, plug-and-play module, MCSD explicitly decomposes atomic interactions across multiple spatial scales. By leveraging envelope associativity to unify the neighborhood graph, this method decouples scale parameters from tensor product complexity, enabling multi-resolution representations with minimal overhead. Extensive evaluations across representative EGNN backbones demonstrate that MCSD not only consistently reduces prediction errors for energy, forces, and macroscopic stress tensors, but also significantly enhances the prediction fidelity of higher-order properties (such as lattice thermal conductivity), all while maintaining the smoothness of the potential energy surface and the energy conservation of long-term molecular dynamics simulations. The code is available at the anonymous repository: https://anonymous.4open.science/r/MCSD-C854.
Spectral Rank Calibration for Continual LoRA Merging in Multimodal Large Language Models
Chenrui Wu ⋅ Haishuai Wang ⋅ Jiajun Bu ⋅ Jiangchuan Liu
Model merging enables the integration of multiple expert models with different capabilities into a unified model. In practical deployments, new expert models are often continuously updated and arriving. Meanwhile, due to the training and communication overhead of large language models (LLMs) and multimodal large language models (MLLMs), Low-Rank Adaptation (LoRA) is widely used to adapt large models to specific domains due to its efficiency. This scenario brings a new challenge: continual LoRA merging of deployed MLLMs to load new task capabilities. Naive merging faces catastrophic forgetting and parameter conflicts. We first observe the distinct features of the LoRA direction and layers in MLLM LoRA tuning. Based on this, we propose our Spectral Rank Calibration Merging (SRC-Merging), a data-free and train-free continual LoRA merging method tailored for MLLMs. SRC-Merging regards the deployed LoRA as a compressed spectral memory and performs calibrated rank-level merging for each incoming adapter. It aims to preserve accumulated multimodal knowledge while avoiding over-suppression of new vision-language task updates during continual merging. Extensive experiments on multiple multimodal tasks have demonstrated that our method significantly outperforms state-of-the-art merging methods.
Spectral Reversal: Counteracting Singular Value Bias for Graph Prompting
Hanxu Yang ⋅ Yuhuan Zhao ⋅ Xiaodong He ⋅ Zhao Kang
Pre-training Graph Neural Networks (GNNs) via self-supervised learning has become a dominant paradigm, yet efficiently adapting frozen encoders remains a challenge. Graph prompting offers a parameter-efficient alternative to fine-tuning, but existing methods largely treat pre-trained models as opaque feature extractors, ignoring their internal spectral structure. In this work, we identify a systematic phenomenon in pre-trained GNNs, which we term spectral bias: optimization during pre-training disproportionately aligns representations with directions associated with large singular values, leaving low-energy directions under-explored. We show that these underutilized directions can encode complementary information that is beneficial for downstream adaptation, especially under distribution shift. To leverage this insight, we propose Spectral Reverse Prompt (SRP), a prompting framework that rebalances the spectral contributions of frozen GNN encoders. SRP applies a learnable soft-thresholding mask in the spectral domain to down-weight dominant directions while amplifying weaker ones. In addition, SRP incorporates a null-space augmentation module that captures variation in directions with minimal activation under the frozen encoder. Extensive experiments across multiple benchmarks demonstrate that SRP achieves state-of-the-art performance with minimal additional parameters, highlighting that reweighting spectral components is a principled and effective strategy for parameter-efficient graph adaptation.
SPERA: Spherical Prior EEG Foundation Model with Geometry- and Frequency-Aware Latent Prediction
Minsu Kim ⋅ Ye-Sung Kim ⋅ Hyeseong Jeon ⋅ Wooseok Hyung ⋅ Joshua Lee ⋅ Chang-Hwan Im
Electroencephalography (EEG) provides a non-invasive measure of ongoing neural activity, but building general-purpose EEG models remains challenging due to the heterogeneity of subjects, devices, and electrode montages. Existing EEG foundation models predominantly rely on input-space reconstruction, which can bias the encoder toward memorizing noise and artifacts. We introduce SPERA (Spherical Prior EEG Representation Architecture), an EEG foundation model that predicts in latent space following the joint-embedding predictive architecture (JEPA). SPERA combines three components tailored to EEG: (i) a hybrid attention backbone that interleaves factorized temporal and spatial attention with periodic full-attention layers, (ii) a Legendre-polynomial spatial prior incorporated into attention to encode varying scalp electrode geometries, and (iii) a relational spectral regularizer that aligns latent similarity structure with spectral views, inducing frequency-aware latent representations. Pretrained on approximately 90,213 hours of EEG from 31,771 subjects across 106 datasets, SPERA achieves the highest average performance across nine downstream tasks spanning clinical, cognitive, and BCI applications. SPERA further exhibits strong parameter efficiency and robustness across varying recording conditions, suggesting its potential as a general-purpose backbone for diverse EEG analyses.
SphereFlow: Missing Modality Imputation via Geometric Transport on Hypersphere
Huihan Wang ⋅ Juncheng Wang ⋅ Yuxiang Feng ⋅ Wenlong Hou ⋅ Guoqi Yu ⋅ Chao Xu ⋅ Yang Liu ⋅ Baigui Sun ⋅ Emma, Shujun Wang
Missing modality imputation aims to synthesize unobserved modalities from observed ones. Existing methods treat each modality as living on its own latent manifold, forcing the generator to learn a long, unstructured jump between them. We present SphereFlow, which takes a fundamentally different view: by projecting all modalities into a shared hyperspherical latent space via a frozen self-supervised encoder, cross-modal imputation reduces to short-range geometric transport between nearby points on the sphere. Realizing this idea with flow matching, however, exposes two geometric conflicts: the standard Gaussian noise source is catastrophically mismatched with the unit-norm data manifold, and the linear interpolation path departs the sphere surface into a region where the decoder has no training signal. We resolve both with simple, geometry-aware modifications: replacing the noise source with the observed-modality latent so that the velocity field learns only the cross-modal displacement, and training the decoder with a time-weighted loss that extends its domain to the near-surface region traversed by the flow. We prove that this data-as-source formulation reduces the worst-case off-manifold deviation from a quantity that grows with the latent dimension to a small, dimension-independent bound governed by inter-modality similarity. Experiments on BraTS multi-modal MRI imputation and CT--MRI translation demonstrate state-of-the-art generation fidelity across all settings, while downstream brain tumor segmentation with our imputed data closes roughly eighty percent of the gap to the full-modality oracle.
SpikingGamma: Temporally Precise Online SNN Training Through Smoothed Temporal Delays
Roel Koopman ⋅ Sebastian Otte ⋅ Sander Bohte
Neuromorphic hardware implementations of Spiking Neural Networks (SNNs) promise energy-efficient, low-latency AI through sparse, event-driven computation. Yet, training SNNs under fine temporal discretization remains challenging, hindering low-latency responsiveness and the mapping of software-trained SNNs to efficient hardware. In current approaches, spiking neurons are modeled as self-recurrent units, embedded into recurrent networks to maintain state over time, and trained with BPTT or RTRL variants based on surrogate gradients. These methods scale poorly with temporal resolution, while online approximations exhibit instability for long sequences and tend to fail at capturing temporal patterns. To address these limitations, we develop SpikingGamma, an SNN architecture in which neurons maintain a compact bank of smoothed-delay states and communicate through sigma-delta spike-coding. We show that in feedforward networks, SpikingGamma supports direct online error backpropagation through the continuous reconstruction signal, avoiding surrogate gradients through the spike discontinuity. This enables stable learning of temporal patterns with minimal spiking and scales feedforward SNNs to complex tasks and benchmarks with competitive accuracy, all while remaining robust to the temporal resolution of the model. Our approach offers both an alternative to recurrent SNNs trained with surrogate gradients, and a novel route for mapping SNNs to neuromorphic hardware.
Sport Is The Next Grand Challenge For Artificial Intelligence: Toward A Science Of Human Physical Skill
Katia Bourahmoune ⋅ Karlos Ishac
We argue that sport represents the next grand challenge for Artificial Intelligence, a domain rich enough to expose the limits of current approaches and structured enough to drive rigorous scientific progress in understanding and modelling human physical skill. Despite rapid advances in language, vision, and robotic manipulation, the physical intelligence required to learn and adapt complex motor skills remains poorly understood and largely unaddressed. We identify two intertwined gaps: the absence of models for human physical skill, and the absence of the scientific agenda needed to build them. We propose concrete directions for how the community can begin to close both.
SpRePE: A Spherical Geometry-Aware Position Embedding scheme for Vision Transformers
Qilong Jia ⋅ Yi Xiao ⋅ Wei Xue
Vision Transformers are increasingly applied to data defined on the sphere, particularly in physics, meteorology, and related scientific domains. Position embeddings provide the spatial information that attention itself does not encode, making their design central to the expressive power of Transformers. However, most existing position embedding schemes are designed for Cartesian grids and therefore do not naturally handle longitude periodicity, polar singularities, or geodesic relations on spherical domains. This mismatch limits the ability of standard attention to model spherical geometry without specialized architectural modifications. We propose Spherical Reflection Position Embedding (SpRePE), a drop-in position embedding scheme for Vision Transformers on spherical data. SpRePE encodes each absolute spherical position by applying Householder reflections to query and key representations. Although each token is encoded only from its own absolute coordinates, the resulting attention inner products induce an explicit spherical relative-position term with a clear geometric interpretation. This formulation injects sphere-aware relative geometric information into standard attention without constructing pairwise attention-bias matrices, introducing task-specific modules, or modifying the backbone architecture. It also avoids the quadratic overhead of relative position biases and preserves the same asymptotic overhead as RoPE. We evaluate SpRePE on spherical image classification, panoramic depth estimation, and global weather forecasting. Across these tasks, SpRePE achieves competitive performance compared with strong position embedding baselines, with particularly clear gains in settings where spherical geometry is important.
S&P: Towards Scalable and Powerful Graph Learning with Hierarchical Structural Acquisition
Mingqi Yang ⋅ Zhaoyu Liu ⋅ Wenjie Feng
Graph Neural Networks (GNNs) face a fundamental dichotomy between expressiveness and scalability. Recent spectral and transformer-based models approach universal approximation capabilities but rely on computationally intensive global operators that scale poorly to large graphs. Conversely, scalable approaches often compromise structural fidelity through sampling or decoupling, resulting in incomplete structural acquisition. We identify that this trade-off stems from a homogeneous treatment of graph topology, where uniform computational complexity is applied to heterogeneous structures. We propose a paradigm shift towards Hierarchical Structural Acquisition and introduce S&P (Scalable & Powerful), a framework grounded in Fusion Frame Theory that aligns the operator with the intrinsic hierarchy of the data. S&P decomposes the global operator into intra-component and inter-component, and we prove that this design preserves universal spectral filtering capacity for non-degenerate inputs while remaining numerically stable and linear in complexity. Experiments on 14 datasets show that S&P achieves state-of-the-art accuracy, bridging the gap between scalability and expressiveness. Our code is available at https://anonymous.4open.science/r/HierarchicalSP.
Stability-Enhanced Federated Learning with Accelerated Gradient
Tianxiang Chen ⋅ Wenjie Hou ⋅ Feng Wang ⋅ Tiantong Wang ⋅ Zhiming Zheng ⋅ Shaoting Tang ⋅ Wei Yang Bryan Lim
Federated Learning (FL) is a promising paradigm for distributed machine learning. However, FL often suffers from degraded generalization performance due to the inconsistency between local and global optimization objectives and client-side overfitting. In this paper, we provide a global-update stability analysis framework as an analytical tool to study generalization error and derive the stability bounds of mainstream FL optimization algorithms under non-convex settings. Our analyses reveal how the number of global update steps, data heterogeneity, and update rules influence their stability. We observe that momentum-based FL acceleration methods do not improve stability. To address this issue, we propose FedSEMA, a new FL algorithm that couples global momentum with a gradient-difference corrected Nesterov Accelerated Gradient (NAG) scheme and a hybrid proximal term to enhance stability. This design ensures updates follow a globally consistent descent direction while retaining the benefits of acceleration. Theoretical analysis shows that FedSEMA achieves an improved stability upper bound on non-i.i.d. datasets in the non-convex settings. Extensive experiments on real-world datasets demonstrate that FedSEMA significantly outperforms multiple baseline methods under standard FL settings, achieving faster convergence and state-of-the-art performance.
Stabilizing Off-policy LLM Optimization with Prefix Importance Ratio
Shiye Lei ⋅ Zhihao Cheng ⋅ Dacheng Tao
Reinforcement learning (RL) post-training has increasingly demonstrated strong ability to elicit reasoning behaviors in large language models (LLMs). For training efficiency, rollouts are typically generated in an off-policy manner using an older sampling policy and then used to update the current target policy. To correct the resulting discrepancy between the sampling and target policies, most existing RL objectives rely on a token-level importance sampling ratio, primarily due to its computational simplicity and numerical stability. However, we observe that token-level correction often leads to unstable training dynamics when the degree of off-policyness is large. In this paper, we revisit LLM policy optimization under off-policy conditions and show that the theoretically rigorous correction term is the prefix importance ratio, and that relaxing it to a token-level approximation can induce instability in RL post-training. To stabilize LLM optimization under large off-policy drift, we propose a simple yet effective objective, Minimum Prefix Ratio (MinPRO). MinPRO replaces the unstable cumulative prefix ratio with a non-cumulative surrogate based on the minimum token-level ratio observed in the preceding prefix. Extensive experiments on both dense and mixture-of-experts LLMs, across multiple mathematical reasoning benchmarks, demonstrate that MinPRO substantially improves training stability and peak performance in off-policy regimes.
Stable Long-Horizon PDE Forecasting via Latent Structured Spectral Propagators
Xiaoxiao Lu ⋅ Ye Yuan ⋅ Jiahao Shi
Long-horizon forecasting of time-dependent partial differential equations (PDEs) is critical for characterizing the sustained evolution of physical systems. While neural operators have emerged as efficient surrogates, they typically learn implicit finite-time transitions from discrete observations. When deployed autoregressively, such propagators often suffer from rapid error accumulation and dynamic drift. To address this, we propose a neural forecasting framework that reformulates PDE rollout as learning a Structured Spectral Propagator (SSP) in a propagation-oriented latent space. Following an analysis-propagation-synthesis design, our framework: (i) maps physical states into a shared, time-consistent spatial representation; (ii) projects this space into a compact propagation state to isolate recurrent dynamics from fine-grained spatial details, thereby decoupling reconstruction fidelity from rollout regularity; and (iii) evolves retained spectral modes using a frequency-conditioned linear backbone complemented by a nonlinear spectral closure to account for truncated interactions. This explicit structuring endows the propagator with a strong inductive bias for coherent modal evolution. Extensive experiments demonstrate that SSP significantly outperforms state-of-the-art baselines, reducing relative $L_2$ errors by up to 48.9\% and exhibiting improved stability in temporal extrapolation beyond the supervised horizon.
Stable Resolution-Invariant Emulation in Entropically Controlled Kinetic Methods
Kareem Hegazy ⋅ Jose Antonio Lara Benitez ⋅ Anastasis Kratsios ⋅ Ivan Dokmanić ⋅ Maarten V. de Hoop ⋅ Michael Mahoney
Resolution invariance is a critical capability for scientific machine learning methods (SciML), however, the ability to sample at any resolution does not guarantee stable emulation at all resolutions. This important distinction underpins one of scientific machine learning's greatest potential: the ability to efficiently emulate complicated and computationally expensive systems on inexpensive coarse grids. Stable numerical solvers are computationally expensive as they attempt to resolve small scales that receive and dissipate spectral energy. Learned coarse-grid forecasters remove the dissipative channel, allowing unresolved frequencies to re-emerge as grid-scale noise (aliasing). We explore stable resolution invariant emulation through physically principled control over entropy dynamics to reinforce the dissipative mechanism within the Neural Discrete Equilibrium (NeurDE) method. Entropically-stabilized NeurDE produces stable long-time rollouts for thousands of forecasting steps on large Reynolds number turbulent flows (Re=50,000) with very coarse grids ($\mathbb{R}^{16\times16}$). Entropically-stabilized NeurDE demonstrates extreme stability at coarse resolutions, outperforming the data-generating stabilized numerical method: a much more challenging baseline than SciML surrogates.
STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
HAIPENG LUO ⋅ Qingfeng Sun ⋅ Song-Li Wu ⋅ Can Xu ⋅ Wenfeng Deng ⋅ Han Hu ⋅ Yansong Tang
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoning in LLMs, yet commonly suffer from policy entropy collapse during training. We conduct a first-order gradient analysis of token-level entropy dynamics under GRPO and identify a token-level credit assignment mismatch: the per-token entropy variation decomposes into the product of the trajectory-level advantage and an entropy sensitivity function over the next-token distribution, yielding an advantage-surprisal four-quadrant structure and a near-criticality property. Motivated by it, we propose STARE (Surprisal-guided Token-level Advantage Reweighting for policy Entropy stability), which identifies entropy-critical token subsets via batch-internal surprisal quantiles, selectively reweights their effective advantages, and incorporates a target-entropy closed-loop gate for stable entropy regulation. Across model scales from 1.5B to 32B and three task families (Short CoT, Long CoT, and Multi-Turn Tool Use), STARE sustains stable RL training over thousands of steps while maintaining policy entropy within the target band. On AIME24 and AIME25, STARE outperforms DAPO and other competitive baselines by 4%–8% in average accuracy, with reflection tokens and response length growing in tandem, indicating sustained exploration–exploitation balance that further unlocks RL training potential.
STAR-Math: Multi-Agent Mathematical Reasoning under Persistent Meta-Strategic Supervision
Jiaao Wu ⋅ Xian Zhang ⋅ Hanzhang Liu ⋅ Sophia Zhang ⋅ Fan Yang ⋅ Yinpeng Dong
Frontier AI models and multi-agent systems have led to significant improvements in mathematical reasoning. However, for problems requiring extended, long-horizon reasoning, existing systems continue to suffer from fundamental reliability issues: hallucination accumulation, memory fragmentation, and imbalanced reasoning-tool trade-offs. We introduce STAR-Math, a multi-agent framework that systematically addresses these challenges through meta-level supervision and structured Reasoner-Verifier interaction. STAR-Math is structured as an orchestrated state machine with nested challenge-step-replan loops, governed by a reasoning-free Python orchestrator that separates control from inference and bounds error propagation through trace-back and re-planning. Our key innovation is a persistent Meta-Strategist that maintains cross-attempt memory and exercises meta-level control by issuing high-level strategic guidance or mandatory directives, so the system can escape unproductive loops rather than stagnate or over-rely on tools. STAR-Math achieves state-of-the-art results on all eight top-tier competition benchmarks: AIME 2025-2026, MathArena Apex Shortlist, MathArena Apex 2025, Putnam 2025, IMO 2025, HMMT February 2026, and USAMO 2026. It obtains perfect scores on AIMEs, Putnam, and HMMT, and shows its largest margin on Apex 2025, scoring 93.75% compared with 80.21% by the strongest baseline GPT-5.5. Ablation studies show that the gains arise from the framework's orchestration rather than from model-level diversity since removing key components or substituting in mixed backbones consistently weakens performance.
Flow matching learns continuous-time transports from a simple source distribution to a target data distribution by training a velocity field on intermediate states along interpolating trajectories. However, because the model observes only the projected data-space state $(x_t,t)$, different transport trajectories can induce conflicting velocity directions at nearby intermediate states, leading to curved transport paths that limit the fidelity of low-NFE sampling. To address this limitation, we propose \emph{State Augmented Flows} (SAF), a lightweight framework that lifts flow matching into a transient augmented state space. SAF transports an augmented state $(x_t,z_t)$, where auxiliary coordinates provide additional contextual information during transport while vanishing at the terminal time. This enlarges the state observed by the velocity model without changing the final sample space, and is compatible with existing coupling strategies and flow-matching architectures. Across MNIST, CIFAR-10, and ImageNet-32, SAF consistently improves over the rectified-flow baseline, with the largest gains in the low-NFE regime. These results suggest that augmenting the transported state provides a new and complementary mechanism for improving flow-based generative modeling.
State Evolution Awareness for Category-agnostic 3D Point Cloud Tracking
Xiantao Hu ⋅ Ying Tai ⋅ Jinxia Xie ⋅ Bineng Zhong ⋅ Guangwei Gao ⋅ Jianjun Qian ⋅ Jian Yang
Existing 3D point cloud tracking methods have not fully exploited temporal information in point cloud sequences, making it difficult to effectively capture the continuous evolution of target states and thereby limiting tracking stability in scenarios with large cross-category geometric variation. To address this issue, we propose TETrack3D, a framework for category-agnostic 3D point cloud tracking. The core idea of TETrack3D is to elevate temporal information from an auxiliary cue to an explicit learning constraint for cross-frame association. Specifically, we introduce a historical feature reuse mechanism to preserve and reuse target-related features accumulated over multiple frames, enabling the association process in the current frame to access richer and finer-grained historical context. We further design a state evolution supervision module, which guides the model to learn temporally consistent target motion patterns from historical observations by predicting future target states. In addition, to alleviate cross-frame feature drift caused by target-state changes in point cloud sequences, we propose a temporal distribution alignment strategy based on optimal transport theory, which constrains target features across adjacent frames at the distribution level and improves the stability of temporal target modeling. Extensive comparative experiments on KITTI, nuScenes, and Waymo demonstrate that explicitly modeling the target-state evolution process effectively improves the localization robustness of 3D object tracking.
Bellman residual minimization (BRM) provides a scalable, gradient-based approach to offline reinforcement learning in large state spaces. While globally convergent gradient-based soft (i.e., entropy-regularized) BRM methods for neural networks have recently been established, their statistical complexity under stochastic gradient descent remains largely unknown. In this paper, we address this theoretical gap for soft BRM. Through a novel Lyapunov-based analysis, we establish an $\mathcal{O}(1 / n)$ average argument stability bound, which translates directly into a $\mathcal{O}(1 / n)$ statistical complexity for the soft BRM objective.
We investigate the quantitative performance of affine-equivariant estimators for robust mean estimation. As a natural stability requirement, the construction of such affine-equivariant estimators has been extensively studied in the statistics literature. We quantitatively evaluate these estimators under two outlier models which have been the subject of much recent work: the heavy-tailed and adversarial corruption settings. We establish lower bounds which show that affine-equivariance induces a \emph{strict} degradation in recovery error with quantitative rates degrading by a factor of $\sqrt{d}$ in both settings. We find that classical estimators such as the Tukey median {(Tukey '75)} and Stahel-Donoho estimator {(Stahel '81 and Donoho '82)} are either quantitatively sub-optimal \emph{even within} the class of affine-equivariant estimators or lack any quantitative guarantees. On the other hand, recent estimators with strong quantitative guarantees are not affine-equivariant or require additional distributional assumptions to achieve it. We remedy this by constructing a new affine-equivariant estimator which nearly matches our lower bound. Our estimator is based on a novel notion of a high-dimensional median which may be of independent interest. Notably, our results are applicable more broadly to \emph{any} estimator whose performance is evaluated in the \emph{Mahalanobis} norm which, for affine-equivariant estimators, corresponds to an evaluation in \emph{Euclidean} norm on isotropic distributions.
StatLUT: Statistical Feature-Driven Multimodal 3D LUT Generation for Photorealistic Style Transfer
Yifan Wang ⋅ Zhixiang Hao ⋅ Yu Wang ⋅ Congchao Zhu
Photorealistic Style Transfer (PST) aims to transfer the color and tonal style of a reference to a content image while strictly preserving its structural integrity. However, existing deep learning-based methods inherently suffer from semantic entanglement caused by pre-trained image encoders, leading to unnatural spatial distortions. Moreover, current pixel-level mapping paradigms often ignore color gamut topology, resulting in color banding, while also lacking the multimodal capability for intuitive text-driven control. To address these bottlenecks, we propose StatLUT, a novel statistical feature-driven multimodal 3D LUT generation framework. First, we bypass traditional encoders and introduce a Lab-Extractor to derive spatially-agnostic statistical features, fundamentally decoupling color distributions from structural semantics to ensure artifact-free rendering. Second, we formulate LUT generation as a Transformer-based Seq2Seq translation task, utilizing a Multi-dimensional Residual Mapper (MR-Mapper) to predict topologically smooth 3D LUTs. Finally, to break the single-modal barrier, we propose the H-Diffuser, a lightweight Diffusion Transformer that directly synthesizes statistical features from natural language prompts, enabling flexible text-driven color grading. Extensive experiments on standard benchmarks demonstrate that StatLUT significantly outperforms state-of-the-art methods in both visual quality and quantitative metrics, pioneering a highly robust and flexible paradigm for multimodal photorealistic style transfer.
Steering Away from Memorization: Reachability-Constrained Reinforcement Learning for Text-to-Image Diffusion
Sathwik Karnik ⋅ Juyeop Kim ⋅ Sanmi Koyejo ⋅ Jong-Seok Lee ⋅ Somil Bansal
Text-to-image diffusion models are susceptible to memorization, revealing a fundamental failure to generalize beyond the training set. Current mitigation approaches typically sacrifice image quality or prompt alignment to reduce memorization. To address this, we propose Reachability-Aware Diffusion Steering (RADS), an offline-trained, inference-time framework that mitigates memorization while maintaining generation fidelity. RADS models the diffusion denoising process as a dynamical system and uses reachability analysis to learn a safety value function that characterizes intermediate latent states along trajectories leading to memorized samples. This motivates a constrained reinforcement learning (RL) formulation, where a policy learns to steer the trajectory away from memorization via minimal perturbations in the caption embedding space. Empirical evaluations show that RADS achieves a stronger empirical Pareto frontier between generation diversity (SSCD), quality (FID), and alignment (CLIP) compared to state-of-the-art inference-time baselines. Crucially, RADS provides robust mitigation without modifying the diffusion backbone, offering a plug-and-play solution for diverse image generation.
Step-dLLM: Adaptive Step-aware Sparse Attention for Efficient Diffusion LLM Inference
Zhichen Zeng ⋅ Xichong zhang ⋅ Junpan Wu ⋅ Yifei Zuo ⋅ Chi-Chih Chang ⋅ Jiayi Wang ⋅ Maohua Nie ⋅ Ji Liu ⋅ Ang Li ⋅ Banghua Zhu
Diffusion large language models (dLLMs) generate tokens in parallel via iterative denoising, but their bidirectional attention prevents the KV caching used by autoregressive models and recomputes the full attention map at every step, incurring quadratic per-step cost. Existing sparse attention methods for dLLMs compute a sparse mask once at the first denoising step and reuse it throughout, while allocating the sparsity budget uniformly across heads. We argue both choices ignore key structural properties of attention in dLLMs: empirically, attention patterns are only locally consistent across steps and shift abruptly at certain transitions, and different heads concentrate their attention mass to very different extents. To this end, we design Step-dLLM, which periodically refreshes a block-sparse mask at anchor steps along the denoising trajectory and reuses it within each locally consistent interval, while a head-adaptive global budget assigns more blocks to heads with more concentrated attention. Backed by a fused IO-aware Triton kernel for block-sparse attention, Step-dLLM applies to pre-trained LLaDA-1.5 and Dream-7B-Instruct without retraining, achieving a $4.2\times$ kernel-level speedup over dense FlashAttention at 16K context with 90% sparsity (up to $8.5\times$ at 95%), and achieves better model accuracy and lower latency than prior methods.
StitchEdit: Stitching Depth Priors into Image Editors via Per-Layer Gradient Probing
Qin Guo ⋅ Dan Xu
Text-guided image editing with diffusion models struggles with spatial edits (e.g. manipulating objects, changing viewpoints, manipulating occlusions) that depend on the 3D structure understanding of a scene. Marigold and follow-up work demonstrate that a pretrained diffusion backbone already harbors strong geometric priors that can be activated for depth estimation, yet how to redirect these priors toward editing remains unresolved. Through per-layer gradient probing, we discover that the editing and depth objectives are compatible in early transformer layers but actively conflict in late layers, explaining why naive depth-conditioned editing either ignores geometry or degrades visual quality. We present StitchEdit, which first activates the backbone's geometric capability via a lightweight depth-denoising adapter, then selectively fine-tunes this adapter for editing using a probe-guided per-layer learning rate schedule that preserves depth knowledge where it helps and prunes it where it conflicts. On the proposed StitchEditBench, a benchmark of 150 geometry-dependent edits spanning spatial transformation, occlusion manipulation, and viewpoint change, StitchEdit consistently surpasses both instruction-based and depth-conditioned baselines on perceptual quality and geometric fidelity.
Stochastic Optimization with Random Search
El Mahdi Chayti ⋅ Taha EL BAKKALI EL KADI ⋅ Omar Saadi ⋅ Martin Jaggi
We study stochastic random search for nonconvex optimization with only noisy function evaluations. We propose Mi2P, a two-point variant of stochastic three-point methods, and analyze it under generalized $(L_0,L_1)$-smoothness. Two regimes are considered. When the mean objective $f$ is $(L_0,L_1)$-smooth and component variance is bounded, Mi2P achieves $\widetilde{O}(d^3\sigma_0^2/\varepsilon^{6})$, matching the rate of \citet{boucherouite2024mistp} under a strictly weaker assumption (they require $f$ to be $L$-smooth, i.e., $L_1=0$). When an $L^1$-average $(L_0,L_1)$-smoothness holds across $\xi$, Mi2P attains the optimal $\widetilde{O}(d\sigma^2/\varepsilon^{4})$ rate; the per-sample $L$-smoothness used by gradient-estimation methods such as RSGF \citep{ghadimi2013stochastic} (i.e., each $f_\xi$ is $L$-smooth, $L_1 = 0$ pointwise in $\xi$) is a special case of this condition. For finite sums with $G$-Lipschitz components, a translation-symmetric variance-reduction scheme yields $\widetilde{O}(\min\{d^{4/3}n^{2/3}/\varepsilon^{8/3},\, dn/\varepsilon^2\})$ without storing snapshots. We also extend the framework to $\delta$-inexact comparison feedback (e.g., RLHF, A/B testing), giving an $O(\sqrt{d\delta})$ accuracy floor. Our analysis is presented for directional distributions with bounded support; the same rates hold for any rotation-invariant distribution with sub-Gaussian tails on the norm of $s$ (including Gaussian) under a step-size restriction that does not affect the asymptotic complexity.
Stochastic Reconfiguration as Statistical Filtering for Overparameterized Neural Quantum States
Tak Hur
Stochastic reconfiguration (SR) is the standard optimizer for neural quantum states (NQS), but modern NQS often have far more parameters than Monte Carlo samples. We show that in this regime the diagonal shift is more than a numerical stabilizer. It acts as a statistical filter for finite-sample generalization. At a fixed wave function, SR is ridge regression from tangent features to the centered local energy. Its residual is the expressivity gap, the part of imaginary-time evolution outside the current tangent space. This gap is orthogonal to the tangent space in population, but finite batches make it act as noise that SR can overfit. The shift therefore balances shrinkage of useful update directions against variance from fitting sampled residuals. Exact $4\times4$ $J_1$-$J_2$ experiments separate two effects of overparameterization. Larger tangent spaces help when they reduce the expressivity gap, but they can hurt when they overfit a fixed gap. In a $100$-site transverse-field Ising family trained with a foundation NQS, validation risk is U-shaped in the shift while variance decreases, matching the noisy-ridge model. This view leads to multi-shift SR (MS-SR), which averages independent ridge solves at data-adaptive shifts to form a richer, lower-variance spectral filter. MS-SR lowers validation risk and stays closer to the ideal tangent-space imaginary-time update compared to other competitive NQS optimizers.
Strategic Evaluation: Incentivizing AI Capability Coverage with Private Benchmarks
Sang Truong ⋅ Serena Wang ⋅ Nick Haber ⋅ Sanmi Koyejo
Public benchmarks play a significant role in steering LLM development, as market incentives drive model developers to optimize for leaderboard performance. However, fundamental information limitations mean that evaluators are only able to cover a subset of socially relevant capabilities with their evaluation tasks. On the other hand, model developers often have additional private benchmarks that are unavailable to evaluators. The resulting information asymmetry opens the door to strategic task specialization that can lead to suboptimal social welfare. We propose randomized evaluation mechanisms (effectively private benchmarks) as a way to incentivize model developers to train to cover a broader set of socially relevant tasks, including those unknown to the evaluator. We formalize this as a dynamic game with information asymmetry between an evaluator and a model developer. We prove that randomized evaluation aligns incentives to the developer’s prior belief over possible evaluated tasks, and is socially optimal when that belief matches society, but information leakage degrades this over rounds. However, if the evaluator can continually update their known task set, alignment can be asymptotically recovered. We illustrate a variance–leakage–correction tradeoff in semisynthetic experiment with a latent factor structure over MMLU-Pro.
In practice, the reasoning strategy often matters more than the model. Rather than calling a fixed large model once per input, a well-designed strategy breaks the problem into steps and uses the right model for each, achieving better accuracy at lower cost. Existing methods have tried to automate strategy design by searching over a fixed set of strategy templates, but the templates themselves limit what reasoning patterns can be discovered. We propose Strategist, a meta-agent that designs a high-performing inference strategy for a given task. Each strategy is a structured executable topology that defines how agents plan, interact, verify, and route across models. Given a task and a small development set, Strategist proposes candidate strategies, runs them on real samples, and revises them against measured accuracy and cost. Every strategy and its components are reusable, forming an evolving library that grows with each new task and compounds what future tasks can build from. On 30 benchmarks, Strategist matches or beats the strongest baselines with an average of 8.5 absolute percentage points higher accuracy at 48% lower inference cost.
Strengthening LLMs for Tabular Prediction with Structural Priors
Pengxiang Cai ⋅ Zihao Gao ⋅ Wanchen Lian ⋅ Guocong Li ⋅ Jintai Chen
Tabular prediction has long been dominated by gradient-boosted decision trees and specialized deep tabular models, while large language models (LLMs) remain difficult to make competitive despite their cross-task adaptability and transparent reasoning traces. Existing LLM-for-tabular methods mainly adapt tables into textual inputs, but often fail to capture key structural properties of tabular data. In this work, we propose Permutation Relative Policy Optimization (PRPO), a reinforcement learning post-training method that strengthens LLMs for tabular prediction by injecting column-permutation invariance as a structural prior. By constructing label-preserving column permutations and estimating advantages both within and across them, PRPO converts sparse outcome rewards into denser and more stable optimization signals. Extensive experiments on 139 OpenML datasets show that our 8B model reaches a genuinely competitive regime against strong specialized tabular baselines. It achieves strong fully supervised performance, dominates cross-dataset zero-shot settings, and performs on par with 32-shot strong baselines. Moreover, it substantially outperforms much larger general-purpose and reasoning LLMs, including up to a 53.17% improvement over DeepSeek-R1 (685B). These results show that structural-prior RL post-training is an effective route for making LLMs competitive in tabular prediction.
Strengthen Out-of-Distribution Detection via Adaptive Mahalanobis Gap
Kunpeng Sui ⋅ Rundong He ⋅ Jie Su
Deep neural networks have achieved significant success on in-distribution (ID) tasks, but out-of-distribution (OOD) samples during inference can lead to unreliable predictions. To address this issue, distance-based methods exploit the geometric structure of feature space for OOD detection. However, methods based on static information geometry fail to address geometry distorted by ill-distributed samples. Although prior work reduces the misclassification of OOD samples as ID by adjusting the residual space with real-time features, it fails to address the misclassification of ID samples as OOD. In addition, existing distance-based methods ignore class-specific characteristics when computing deviation feature and rely on the nearest prototype for uncertainty estimation, leading to suboptimal OOD detection performance. To address these issues, we propose Adaptive Mahalanobis Gap (AMG). Specifically, we first strengthen the principal space along real-time feature directions to reduce the misclassification of ID samples as OOD. Subsequently, we incorporate class variance and prototype norms to mitigate bias of deviation feature from class-wise distributional differences. Finally, we introduce the Mahalanobis gap scoring for uncertainty scoring to overcome bias from single-prototype reliance. Experiments on CIFAR-10/100 and ImageNet-1k show that AMG consistently outperforms state-of-the-art methods in FPR@95 and AUROC. Our code is available at https://anonymous.4open.science/r/AMG_AR.
Stroke-Audited Quadratic Bezier Splatting with Structural Initialization for Efficient Painting Rendering
Jinfan Liu ⋅ Xuhan Zhan ⋅ Mengqin Zhao ⋅ Guangyi Deng ⋅ Wuze Zhang ⋅ Bingbing Ni
Stroke-based rendering is often optimized as a pixel-level reconstruction task. This over-reliance on global error reduction not only degrades strokes into fragmented curves due to a lack of structured planning, but also allows locally defective primitives to be concealed by overlapping layers. To address these issues, we introduce a model-guided framework for stroke-audited rendering. Our framework uses Quadratic Bezier Splatting to analytically model continuous curved strokes with efficient differentiable alpha compositing. Built on this renderer, a feedforward multi-scale stroke decoder provides a structural initialization with learned stroke-scale allocation, while test-time stroke optimization supplies the reconstruction capacity needed for high-fidelity refinement. We further propose a single-stroke audit mechanism that explicitly evaluates local stroke quality, preventing defective primitives from surviving optimization through layer-wise compensation. Experiments demonstrate our approach achieves highly coherent layouts, improving PSNR by up to 20\% and yielding a 58.8$\times$ renderer speedup with a 16\% more compact representation compared to representative SBR baselines.
Structure-Guided Masked Autoencoders for Ultra-High Resolution Scientific Image Understanding
Enzhi Zhang ⋅ Du Wu ⋅ Rui Zhong ⋅ Cong Ma ⋅ Isaac Lyngaas ⋅ Amir K Ziabari ⋅ Xiao Wang ⋅ Peng Chen ⋅ Tao Luo ⋅ Toshio Endo ⋅ Fumiyoshi Shoji ⋅ Kento Sato ⋅ Kentaro Uesugi ⋅ Takayuki Nonoyama ⋅ Ryuji Kiyama ⋅ Masahiro Yoshida ⋅ Tezuka Masaru ⋅ Tetsuya Ishikawa ⋅ Satoshi Matsuoka ⋅ Masaharu Munetomo ⋅ Mohamed Wahib
Self-supervised pre-training with Vision Transformers, including Masked Autoencoders (MAE), is difficult to apply to gigapixel scientific images. Random masking is poorly matched to the structured, multi-scale morphology of scientific data, while uniform tokenization produces prohibitively long sequences that make $O(N^2)$ attention impractical. We propose \method{}, a structure-guided masked autoencoding framework for ultra-high-resolution scientific images. \method{} couples two components: a content-adaptive quadtree tokenizer that compresses gigapixel images into a fixed-length sequence, and a structure-conditioned masking process that biases reconstruction toward spatially informative regions. To stabilize this process across scales, we introduce Damped Accumulation (DA), which aggregates signal-dependent responses across the tree into a structure canvas used to guide masking. The resulting pre-training task preserves fine microstructure while remaining compatible with standard ViT encoders and MAE-style reconstruction. Across electron microscopy, whole-slide optical microscopy, and X-ray CT datasets, \method{} consistently outperforms MAE baselines. It achieves 95.68\% Dice on the 8K${\times}$8K${\times}$28K SpringXCT dataset, improving over the strongest MAE baseline by +9.70 points, and 83.21\% Dice on the $32\text{K}^2$ WSI PAIP dataset, improving by +16.84 points, while providing up to a $24.8\times$ inference speedup
Subject-Relative Micro-Motion and Sleep Dynamics for Near-Infrared Video Sleep Staging
Kunmin Jang ⋅ Yourim Choi ⋅ Hun Heo ⋅ Dongik Park ⋅ Suahn Bae ⋅ Heonjun Lee ⋅ Hyun-Woo Shin ⋅ Hyung-Sin Kim
Near-infrared (NIR) video is a promising modality for contactless sleep monitoring, but recent video-based sleep staging methods often use it as a route to reconstructed respiratory/cardiac proxies or cross-modal physiological representations. We study video-only sleep staging under labels defined by polysomnography (PSG), where the model infers sleep stages from NIR video alone without explicit physiological proxy reconstruction or auxiliary physiological supervision. This tests whether NIR video itself can provide informative sleep-stage evidence, rather than only serving as an input for recovering physiological proxies. We propose ViNUSS (Video-Native Unmediated Sleep Staging), a framework that combines subject-relative micro-motion learning with full-night sleep dynamics modeling. Spatially anchored pre-spatial micro-motion encoding preserves localized temporal variation together with its spatial context. Within-subject stage contrast learns stage cues with respect to each subject's night-specific baseline. Two-scale sleep dynamics modeling captures within-epoch motion evolution and organizes epoch-level evidence into a coherent full-night sleep-stage trajectory. On 475 overnight NIR recordings (~3,250 hours), ViNUSS achieves 0.80 accuracy and 0.78 macro-F1 for four-class sleep staging. Interpretability analysis suggests attention to thoraco-abdominal periodic motion and gross body movements associated with arousals and position changes. These results support NIR video as an independently informative and complementary modality for PSG-defined sleep-stage estimation.
SuperSycophantic: Stress-Testing Frontier LLMs from Single- to Multi-Turn Sycophancy
Terry J Zhang ⋅ Oscar S Yasunaga ⋅ Wenyuan Jiang ⋅ Jessica Bo ⋅ Florent Draye ⋅ Aydin Javadov ⋅ Arth Singh ⋅ Lucille Muir ⋅ VICTORIA OLDEMBURGO DE MELLO ⋅ Bernhard Schölkopf ⋅ Zhijing Jin
AI sycophancy is gaining increasing prevalence as models optimized to maximize user satisfaction may tend to unconditionally agree with users even at the expense of factuality, which poses great risk as AI are increasingly used for decision support in high-stake scenarios. We present \ourWork, a systematic stress-testing framework encompassing objective (OBJ) questions with verifiable answers and subjective (SUB) scenarios without ground truth on sycophancy induced from first-turn context framing to multi-turn user pressure. Evaluation of 9 frontier models revealed high sycophancy in even the best-performing models such as GPT-5.4 changes answers from right to wrong to please users in 23.8\% of OBJ scenarios and blindly follows the pressured user view in more than half of the SUB scenarios. We found that strength of user tone is one of the most impactful yet previously overlooked factor for AI sycophancy and discover an interesting pattern where Claude models are uniquely more sycophantic under moderate pressure because strong triggers often lead them to rethink questions from scratch to arrive at the correct impartial answer. These findings highlight the need for anti-sycophancy training in future model development, where training models to rethink from scratch when facing pressure may serve as a promising paradigm towards more truthful AI systems. We provide our code and data in the supplementary material.
Support Before Frequency in Discrete Diffusion
Adrian Müller ⋅ Antoine Gonon ⋅ Zebang Shen ⋅ Ya-Ping Hsieh ⋅ Niao He
Discrete diffusion models are increasingly competitive for language modeling, yet it remains unclear how their denoising objectives organize learning. Although these objectives target the full data distribution, we show that the exact reverse process induces a hierarchy between coarse support information and finer frequency information. For uniform and absorbing (a.k.a. masking) diffusion, we prove that, in the small-noise regime of the final denoising steps, each single-token reverse edit decomposes into a leading scale, determined by whether it moves toward the data support (e.g., grammatically valid sentences), and a finer coefficient, determining relative probabilities within the same scale. Thus, recovering validity structure only requires learning the correct order of magnitude of reverse probabilities, whereas recovering data frequencies requires coefficient-level estimation. The separation is mechanism-dependent: uniform diffusion exhibits a trichotomy into validity-improving, validity-preserving, and validity-worsening edits, while absorbing diffusion places its leading-order mass on validity-improving moves. Experiments on a masked language diffusion model and synthetic regular-language tasks support these predictions: support-localization emerges earlier than within-support frequency ranking, and the contrast between uniform and absorbing diffusion matches the predicted rate separation. Together, our results suggest that discrete diffusion models learn data support before data frequencies.
SurGe: Improved Surface Geometry in Point Maps
Karim Knaebel ⋅ Gonzalo M Garcia ⋅ Christian Schmidt ⋅ Ilya Fradlin ⋅ Lucas Nunes ⋅ Daan de Geus ⋅ Bastian Leibe
With impressive progress in architectures, training recipes, and quantity of high-quality data, recent monocular geometry estimation methods estimate global 3D geometry remarkably well. However, they still exhibit surprisingly low quality on surface details, which is clearly visible in qualitative predictions but not reflected in common metrics. To improve this shortcoming, we first formulate a metric that is capable of measuring the surface quality of the estimated geometry via normals derived from the predicted point map. Furthermore, we propose two manners to improve fine-grained local geometry in monocular geometry estimation models: (i) we present an improved point-based gradient loss, and (ii) we introduce a novel decoder based on neighborhood attention. We evaluate our proposed method on eight common zero-shot monocular geometry estimation benchmarks and achieve state-of-the-art results, particularly improving local surface geometry. Extensive ablations validate our design choices. Code and models will be open-sourced.
SurvivalPFN: Amortizing Survival Prediction via In-Context Bayesian Inference
Shi-ang Qi ⋅ Vahid Balazadeh Meresht ⋅ Michael Cooper ⋅ Russell Greiner ⋅ Rahul Krishnan
Survival analysis provides a powerful statistical framework for modeling time-to-event outcomes in the presence of censoring. However, selecting an appropriate estimator from the many specialized survival approaches often requires substantial methodological and domain expertise. We introduce SurvivalPFN, a prior-data fitted network that amortizes Bayesian inference for censored observations through in-context learning. SurvivalPFN is pretrained on a diverse family of synthetic, identifiable, and right-censored data-generating processes, enabling it to amortize survival analysis in a single forward pass during inference. As a result, the model adapts to the effective complexity of each dataset without task-specific training or hyperparameter tuning, avoids restrictive parametric assumptions, and produces calibrated survival distributions. In a large-scale benchmark spanning 61 survival datasets, 21 methods, and 5 evaluation metrics, SurvivalPFN achieves strong predictive performance and often improves upon established survival models. These results suggest that SurvivalPFN offers a principled and practical foundation model for survival analysis, with potential applications in high-impact domains such as healthcare, finance, and engineering.
Susceptibilities for Neural Networks Learning from Physical Data
Rohan Hitchcock ⋅ Gary W Delaney ⋅ Jonathan H Manton ⋅ Richard Scalzo ⋅ Jingge Zhu
Neural networks that are trained on data arising from a physical system must somehow learn regularities induced by the underlying physical laws. In this setting, concepts from statistical physics provide powerful tools for analysing how physical laws influence the learning process. In this paper we explain how susceptibilities --- which measure the response of a neural network to perturbations of a parameter of the data distribution --- can be applied when learning from physical data. We show that such susceptibilities can identify the conditions under which specific input features are most informative for learning, guiding data selection, a prediction we validate experimentally. For distributions lacking a natural parameter, we propose introducing one via an arbitrary scalar function of the physical state, and demonstrate that susceptibilities track the emergence of specific capabilities (such as collision detection) during training.
SVoT: State-aware Visualization-of-Thought for Spatial Reasoning via Reinforcement Learning
Chao Lei ⋅ Yanbei Jiang ⋅ Markus Hiller ⋅ Zhijian Zhou ⋅ Xunye Tian ⋅ Krista A. Ehinger ⋅ Nir Lipovetzky
Spatial reasoning remains a challenge for Multimodal Large Language Models (MLLMs), as it requires reliable multi-hop inference over both intermediate states and state transitions. Current studies often leave intermediate states unverified and treat state transitions as implicit processes, which limits reliability in multi-hop spatial reasoning. To address this, we propose State-aware Visualization-of-Thought (SVoT), a reinforcement learning framework that generates interleaved, verifiable intermediate states and visualizations. SVoT integrates transition reasoning chains into the generation processes, enabling the model to verify action preconditions and effects through interleaved textual and visual reasoning. We train SVoT via Group Relative Policy Optimization (GRPO), instantiating verification through reward design and evaluating the efficacy of different fine-grained rewards. As existing benchmarks reduce state transitions to single-variable updates, substantially simplifying the problems, we establish five domains by extending classical environments and introducing two novel domains, Pacman and Gather, that require multi-object interactions and numerical reasoning. These domains support systematic evaluation of multi-hop spatial reasoning with quantitative verification of generated intermediate states and transition reasoning. SVoT with transition-aware supervision achieves state-of-the-art performance across the introduced domains, yielding up to a 65\% absolute accuracy gain on out-of-distribution test sets.
SWE-Git-Bench: A Focused Worktree-Level Benchmark for Real Merge Conflict Resolution
Wei Zhang ⋅ Jian Yang ⋅ jiajun wu ⋅ Zidan Tang ⋅ Siwei Wu
Modern coding agents operate inside live \texttt{git} worktrees: they edit files, pull upstream changes, and must leave the repository in a mergeable state. Yet current code benchmarks focus on issue resolution or synthesis from specification, not on the merge-resolution primitive an agent must invoke after \texttt{git} emits conflict markers. We introduce \textbf{SWE-Git-Bench}, a focused high-fidelity benchmark of 135 manually audited real conflicts spanning 185 files from 23 repositories. Across 36 contemporary LLMs, exact reproduction of the maintainer-committed file peaks at only \TopWFEM{}\%, while top systems still cluster around $86$--$90\%$ edit similarity, exposing a near-repair regime where outputs look close but fail to reproduce the committed resolution. A human-checked audit of 80 whole-file non-EM outputs with ES $\ge 0.99$ finds that 50.0\% are still likely unsafe, so similarity cannot be treated as a semantic pass rate. Conflict-block prompting improves edit similarity by \AvgDeltaES{}\,pp on average (median $+16.1$\,pp) while leaving exact match effectively unchanged. We release the dataset, harness, full prediction matrix, difficulty calibration, and exploratory failure analysis to support reproducible evaluation of this worktree-replayed, file-scored resolver primitive.
SWE-Marathon: Can AI Agents Autonomously Complete Ultra-Long-Horizon Software Work?
Rishi Desai ⋅ Joan Cabezas ⋅ Neel Harsola ⋅ Adnan E Assadi ⋅ Pramod Srinivasan ⋅ Roey B Chaim ⋅ Fenil Faldu ⋅ Prannay Hebbar ⋅ Jiankai Sun ⋅ Christopher Settles ⋅ Omkaar Kamath ⋅ Pratyush Shukla ⋅ Albert Liu ⋅ Yiyuan Li ⋅ Nevasini Sasikumar ⋅ Xiangyi Li ⋅ Pranav Raja ⋅ Ishan Gupta ⋅ Marek Suppa ⋅ Daniel Wang ⋅ Erik Quintanilla ⋅ Derek Chen ⋅ Chris Kong ⋅ Steven Dillmann ⋅ Ivan Bercovich ⋅ Jesse Hu
AI agents are increasingly expected to complete long-horizon workflows that require sustained progress over hours, millions of tokens, and complex environments. Yet current agent benchmarks largely evaluate short-form tasks, such as single pull requests, small tickets, or 5--10 minute exercises, limiting our ability to measure agents' capabilities in planning, long-context understanding, and memory use. We introduce SWE-Marathon, a benchmark of 20 long-horizon tasks spanning software engineering and adjacent technical domains. Each task consists of a unique executable environment, a human-written reference solution, and a multi-layer verification suite. Logged agent attempts average 27.2M total tokens, making SWE-Marathon substantially longer-horizon than existing SWE and command-line agent benchmarks. Current frontier coding agents solve fewer than 20\% of tasks. Failures often arise from poor self-verification, self-reported infeasibility, and premature termination. We also observe reward-hacking behavior in 18.8\% of rollouts, where agents attempt to exploit the environment or verifier to bypass the intended workflow. SWE-Marathon includes adversarial review of test suites and execution environments, as well as multi-layer checks designed to prevent shortcut solutions. We release SWE-Marathon and its evaluation code to support reproducible measurement of long-horizon agent capabilities.
SWE-Pro: A Benchmark for Evaluating LLMs on Performance-Oriented Repository-Level Optimization
Ezgi Sarıkayak ⋅ Wenchao Gu ⋅ Hesham Ghonim ⋅ Chunyang Chen
Software performance optimization is a notoriously complex and manual task. Despite the growing use of Large Language Models (LLMs) for code refinement, we still lack benchmarks that capture how optimization actually happens in real-world codebases. Existing frameworks often oversimplify the problem by focusing on isolated functions or a single performance metric, missing the critical trade-offs between execution time and memory footprint, the inherent noise of the measurement environment, and the variability introduced by different input data and execution conditions. We address this by introducing SWE-Pro, a repository-level benchmark derived from 102 expert-written optimizations from open-source projects. Unlike previous benchmarks, SWE-Pro pairs each task with parameterized tests to evaluate runtime, peak memory, and Time-Weighted Memory Usage (TWMU) varying input data and execution conditions under noise-aware measurement conditions. Our evaluation shows that current LLMs struggle significantly: runtime gains are negligible, and memory optimizations are nearly non-existent. This stands in sharp contrast to expert implementations, which achieve an aggregate speedup of 15.5x and peak memory reduction of 171.3x over benchmark tasks. Expert-written improvements are observed in 91.2% of tasks for runtime and 65.7% for peak memory. Our findings expose a substantial gap between current LLMs capabilities and the demands of expert-level engineering.
Symmetry-Guaranteed Prediction of High-Order Tensor Properties for Crystalline Materials via Irreducible Decomposition
Qiaolin Lu ⋅ Qiang Qu ⋅ Hao Jiang ⋅ Aoni Xu ⋅ Fengwang Li ⋅ Bo Han ⋅ Tongliang Liu ⋅ Yi Chang ⋅ Chengqi Zhang
Predicting high-order tensor properties for crystalline materials is crucial for various scientific and engineering applications. Crystal symmetry is one of the primary factors influencing high-order tensor properties, such as elasticity and piezoelectricity, making strict adherence to symmetry constraints essential. However, exactly guaranteeing symmetry compliance remains challenging. Recent approaches rely on enforcing symmetry but often fail to strictly preserve symmetry. In this work, we propose a novel method that guarantees exact symmetry compliance by predicting symmetry-constrained irreducible components of high-order tensors. Specifically, we first develop a computational procedure to identify the basis tensors corresponding to symmetry-constrained irreducible components under various symmetry conditions. This symmetry-constrained basis guarantees that the assembled full tensor strictly adheres to the required symmetry constraints. To predict the numerical values for these irreducible components, we then propose a spherical-harmonic convolutional neural network designed to effectively capture essential high-order tensor information. Extensive experiments validate that our method achieves exact symmetry compliance without compromising prediction accuracy, thereby outperforming state-of-the-art approaches.
Symplectic Reck: In-Situ Learning of Gaussian Quantum Operations
Janet Zhong ⋅ Renwen Yu ⋅ Charles Roques-Carmes ⋅ Paul-Alexis MOR ⋅ Aviv Karnieli ⋅ David A B Miller ⋅ Shanhui Fan
Programmable photonic circuits are a natural platform for physical learning, where the learned computation is performed by the hardware itself rather than by a digital model. We propose a physical learning algorithm and architecture for continuous-variable quantum systems where a photonic circuit is configured in situ to invert an unknown lossless Gaussian quantum operation. Our contributions are threefold. (i) Theory: we prove that generic $S \in \mathrm{Sp}(2 N, \mathbb{R})$ admits a triangular decomposition into two-mode and single-mode symplectic gates. (ii) Architecture: we introduce the symplectic Reck mesh, a triangular array of two-mode symplectic gates that physically realizes this factorization, generalizing the unitary Reck mesh widely used in optical neural networks. (iii) Physical learning: when placed after an unknown Gaussian unitary, the mesh is trained in situ by nulling local response blocks one at a time using simple optical probes and measurements. We identify a measure that predicts when training becomes hard and numerically demonstrate finite-shot training and robustness to gate errors. Our work extends self-configuring optical networks from unitary to symplectic transformations, bringing physical learning to continuous-variable quantum photonic hardware.
SynIB: Informational Bottleneck for Maximizing Synergy in Multimodal Learning
Konstantinos Kontras ⋅ Teodora Popordanoska ⋅ Thomas Strypsteen ⋅ Christos Chatzichristos ⋅ Matthew Blaschko ⋅ Maarten De Vos ⋅ Paul Liang
A central objective in multimodal learning is to capture synergy: task-relevant information that arises only from the joint use of multiple modalities, and is not available from any single modality alone. While most approaches operate at the architectural level through larger or more complex fusion models, we propose a complementary axis: shaping the training objective itself. Standard training often emphasizes unimodal or redundant information, falling short on examples that require cross-modal reasoning. We formalize multimodal synergy through information theory and introduce the Synergistic Information Bottleneck (SynIB), a scalable objective that targets synergy directly. To prioritize learning synergy, SynIB motivates the model to predict accurately from all modalities while penalizing confidence when information from any modality is withheld. Alongside the standard task loss, the model runs forward passes with one modality masked at a time and is penalized for remaining confident, which would indicate reliance on unimodal cues rather than cross-modal interactions. We validate SynIB in two regimes. On synthetic XOR tasks where the ground-truth synergy is known by construction, standard training fails to recover it while SynIB does. On five real-world benchmarks, including three MultiBench affective tasks, Hateful Memes with CLIP-ViT and DeBERTa backbones, and a controllable irony extension of CREMA-D we introduce, SynIB improves accuracy on synergy-dependent examples by up to 7.8\% and overall accuracy by up to 3.8\%.
Synsema: Syntax-Guided Learning of Semantically Valid Programs
Levin Winter ⋅ Cong Li ⋅ Hao Sun ⋅ Zhendong Su
Generating inputs for systems that process highly structured data, such as compilers, is difficult due to the dual of syntactic and semantic constraints. Traditional approaches rely on manually crafted rules that require deep domain expertise and are difficult to maintain, while modern machine learning methods have to learn syntax and semantics simultaneously which is challenging. We present Synsema, a novel approach that combines insights from both worlds by reformulating the generation of semantically valid programs as a syntax-guided reinforcement learning task. By guaranteeing syntactic validity through the grammar, the agent needs to only learn the language's semantic rules, sampling from a drastically reduced space. We implement our approach in the context of Java and Rust which both exhibit strict semantic requirements, and train an agent to generate interesting and well-formed programs using fine-grained semantic feedback. Applied to compiler testing, we systematically compare different syntax-guided generation procedures with an unconstrained baseline and show that by restricting the sampling space, Synsema learns semantic rules efficiently. While an average of only 0.03% of programs generated by the baseline are semantically valid, Synsema produces a diverse set of programs with a 7.29% success rate that trigger 2.6x more compiler behaviors. In addition, Synsema discovered 5 previously unknown bugs in production compilers, validating the practical impact of our approach.
Synthetic Anchor-Assisted Prototype Alignment for Heterogeneous Federated Learning
Zhihao Hao ⋅ Bob Zhang
Heterogeneous Federated Learning (HFL) aims to enable collaboration among clients with diverse model architectures and non-IID data distributions, where direct parameter aggregation is often infeasible and prototype-based aggregation may suffer from semantic drift. Existing methods usually construct global references from client-side representations, making the alignment process vulnerable to local statistical bias and noisy or corrupted prototype updates. In this paper, we propose SATPL, a synthetic anchor-assisted prototype alignment framework that introduces externally generated synthetic anchors as auxiliary semantic priors for HFL. Instead of treating Large Language Models (LLMs) as infallible semantic oracles, SATPL uses LLM-generated class descriptions and generative models to construct reproducible reference prototypes for supervised tasks with known class semantics. Each client maps its heterogeneous local representation into a shared anchor space through a lightweight projection head, while an alignment-weighted aggregation rule assigns lower weights to prototypes that are highly inconsistent with the corresponding anchors. The blockchain component is used as an auditable recording layer for anchor commitments and aggregation metadata, rather than as the source of the statistical robustness guarantee. Experiments under architecture heterogeneity, label skew, and random prototype corruption show that SATPL improves accuracy, convergence stability, and robustness over representative HFL baselines. Additional analyses evaluate anchor quality, temperature sensitivity, and system overhead, clarifying both the benefits and limitations of synthetic-anchor-based alignment.
Tackling the Data-Parallel Load Balancing Bottleneck in LLM Serving: Practical Online Routing at Scale
Tianci Bu ⋅ Yuan Lyu ⋅ Zixi Chen ⋅ Chendong Song ⋅ Hong Liang ⋅ Tsepten Gurung ⋅ Yuwei Fan ⋅ Yinyu Ye ⋅ Zijie Zhou
Data-parallel (DP) load balancing has emerged as a first-order bottleneck in large-scale LLM serving. When a model is sharded across devices via tensor parallelism (TP) or expert parallelism (EP) and replicated across many DP workers, every decode step ends in a synchronization barrier whose latency is set by the most heavily loaded worker; even modest persistent imbalance across DP workers compounds, step after step, into a substantial fraction of wasted compute. The problem is hard for reasons specific to LLM decoding: assignments are sticky (KV caches cannot be migrated), per-request loads grow over time, arrivals are non-stationary, and the router must decide within a sub-100ms decode budget over hundreds of waiting requests and tens of workers. We present BalanceRoute, a family of practical online routing algorithms that target this bottleneck. The first, BR-0, requires no prediction infrastructure and uses a piecewise-linear F-score that captures the sharp asymmetry between admissions that fill safe margin and those that overflow into the envelope; a two-stage decomposition keeps per-step cost compatible with millisecond-scale scheduling. The second, BR-H, generalizes BR-0 with a short, constant lookahead $H$ and a lightweight termination-classifier interface, extending the F-score to a horizon-discounted form. We deploy BalanceRoute on a 144-NPU cluster and evaluate against vLLM baselines on both a proprietary production trace and the public Azure-2024 trace. Across both workloads, BalanceRoute substantially reduces average DP imbalance and improves end-to-end serving throughput. An anonymized open-source release is available at https://anonymous.4open.science/r/BR-Family-on-vllm-acend-0F95/.
Take It or Leave It: Intent-Controlled Partial Optimal Transport
Parth TRIPATHI ⋅ bertrand chapron ⋅ Fabrice collard ⋅ Nicolas Courty ⋅ ronan fablet
While optimal transport (OT) enforces a rigid constraint by requiring two measures to be matched exactly, partial optimal transport relaxes this requirement by allowing mass to remain unmatched through a global budget, scalar rebate, or uniform rejection rule. However, many applications call for more structured, pointwise rejection mechanisms, where the decision to leave mass unmatched depends on side-specific reliability, support geometry, or external information about which components should participate in the comparison. We introduce \emph{intent-controlled partial optimal transport} (IC-POT), a targeted generalization of partial transport that replaces the global rejection paradigm with pointwise rejection costs over both measures. We show that the resulting optimization problem admits a dual interpretation in terms of local acceptance thresholds and can be solved by recasting it as a balanced Kantorovich OT problem on an augmented support. Beyond theoretical analysis, we demonstrate the practical relevance of IC-POT in settings where rejection is driven by side information. In positive-unlabeled learning and open-partial domain adaptation, incorporating pointwise rejection rules that encode statistical structure improves fixed baseline pipelines. Finally, we motivate the use of IC-POT with a geophysical practical case: multi-modal satellite ocean measurements, for which physical and instrumental priors naturally inform the rejection mechanism and define the retrieved comparable signal information.
TALK: One-Shot Batch Design for Protein Variant Effect Prediction
qiao huang ⋅ Ao Shen ⋅ Jie Du ⋅ Manning Wang
Protein variant effect prediction is central to understanding protein function and guiding protein engineering, yet experimental labels remain costly. In many protein experiments, the variants that can be assayed are dictated by protocol constraints, while the variants that need accurate prediction may lie in a different or broader region of the landscape. This assay-prediction mismatch makes candidate-local criteria such as predicted function, uncertainty, or exploration value insufficient for measurement design. Given a protocol-defined accessible space and a limited budget, we ask which variants should be assayed so that the resulting data best improves landscape prediction. We propose TALK, which estimates each candidate’s contribution through posterior coupling between accessible and target variants and reduces redundant information within the batch. Across ProteinGym and GB1 settings spanning mutation-order transfer and different accessible-target relationships, TALK uses limited budgets more efficiently, yielding labels that better support subsequent prediction than baselines. These results point to a practical two-stage workflow: first, spend a limited assay budget on variants selected to be informative for the target landscape and train a predictive model from the measured labels; then use this model to score and prioritize high-potential variants.
Taming Generative Co-Folding Prior for Molecular Docking with Diffusion Bridge
Shikun Feng ⋅ Zipeng Yang ⋅ Qiuyi Li ⋅ Mingze Yin ⋅ Xiaoyuan Zhang ⋅ Tingjun Hou ⋅ Yiheng Zhu
Molecular docking aims to predict the three-dimensional structure of a protein–ligand complex; in the conventional structure-based setting, it is conditioned on an input protein structure and ligand, often starting from an apo receptor in flexible docking or a holo receptor in rigid docking. Recent co-folding models exemplified by AlphaFold 3 achieve strong accuracy in protein–ligand complex prediction, but they typically generate complexes from sequence-derived inputs and therefore do not directly use apo structures as structural priors in docking, leading to a larger conformational search space and weaker robustness in difficult or few-step settings. We present BridgeDock, a framework that adapts pretrained co-folding backbones to flexible docking by modeling the apo-to-holo transition as a diffusion bridge, allowing generation from the initial protein–ligand state rather than from Gaussian noise. To make this bridge formulation compatible with pretrained denoisers, we introduce an alignment mechanism that maps bridge states to the original denoising schedule of the pretrained model. Experiments on standard flexible docking benchmarks show that BridgeDock consistently outperforms strong co-folding baselines. Moreover, BridgeDock is computationally efficient at inference time and remains effective even with very few denoising steps.
TaskGround: Structured Executable Task Inference for Full-Scene Household Reasoning
ZhiYuan Feng ⋅ Yu Deng ⋅ Ruichuan An ⋅ Zhenhua Liu ⋅ Qixiu Li ⋅ Keming Wu ⋅ Zhiying Du ⋅ Weijie Wang ⋅ Haoxiao Wang ⋅ Shuang Chen ⋅ Sicheng Xu ⋅ Yaobo Liang ⋅ Jiaolong Yang ⋅ Baining Guo
In real home deployments, household agents must often operate from a complete household scene and a situated household request, rather than from a clean task specification. Such requests require agents to identify task-relevant entities, recover intended task conditions, and resolve ordering constraints from the surrounding scene context. We formalize this capability as full-scene household reasoning: given a complete household scene and a situated household request, an agent must infer executable task structure before producing a grounded skill-level action sequence. This setting is challenging because complete household scenes contain substantial task-irrelevant information, making direct complete-scene prompting inefficient and error-prone. In practical deployment, this challenge is further amplified by privacy and local compute constraints, which favor compact open-weight models with limited long-context reasoning ability. We propose TaskGround, a training-free and model-agnostic Ground-Infer-Execute framework that grounds complete scenes into compact task-relevant scene slices, infers executable task structure, and compiles it into grounded skill-level action sequences. To evaluate this setting, we introduce FullHome, a human-validated evaluation suite of 400 household tasks spanning diverse home-scale environments and both goal-oriented and process-constrained requirements. On FullHome, TaskGround improves task success rates by large margins across both proprietary and open-weight models. Notably, it makes Qwen3.5-9B competitive with GPT-5 under direct complete-scene prompting while reducing total input-token cost by up to 18x. Our results identify executable task-structure inference as a central bottleneck in full-scene household reasoning and show that structured grounding can make compact local models substantially more effective for practical household deployment.
Tatemae: Detecting Alignment Faking via Tool Selection in LLMs
Matteo Leonesi ⋅ Francesco Belardinelli ⋅ Flavio Corradini ⋅ Marco Piangerelli
Alignment faking (AF) occurs when an LLM strategically complies with training objectives to avoid value modification, reverting to prior preferences once monitoring is lifted. Current detection methods focus on conversational settings and rely primarily on Chain-of-Thought (CoT) analysis, which provides a reliable signal when strategic reasoning surfaces, but cannot distinguish deception from capability failures if traces are absent or unfaithful. We formalize AF as a composite behavioural event and detect it through observable tool selection, where the LLM selects the safe tool when unmonitored, but switches to the unsafe tool under monitoring that rewards helpfulness over safety, while its reasoning still acknowledges the safe choice. We release a dataset of 108 enterprise IT scenarios spanning Security, Privacy, and Integrity domains under Corruption and Sabotage pressures. Evaluating six frontier LLMs across five independent runs, we find mean AF detection rates between 3.5\% and 23.7\%, with vulnerability profiles varying by domain and pressure type. These results suggest that susceptibility reflects training methodology rather than capability alone.
Teaching Video Generators to Remember: Eliciting Dynamic Memory for Out-of-Sight State Evolution
Tianshuo Xu ⋅ Yichen Xie ⋅ Depu Meng ⋅ Chensheng Peng ⋅ Quentin HERAU ⋅ Bo Jiang ⋅ yihan hu ⋅ Wei Zhan
Video world models should maintain evolving states when evidence is unobserved, yet current generators often freeze hidden states upon interruption. This is not simply a capacity problem: pretrained video diffusion transformers already possess KV-cache mechanisms capable of non-local retrieval, but they are rarely trained to use them as dynamic memory. We introduce ReMind, a framework eliciting dynamic memory behavior via memory-oriented data, event-aware training, and cache adaptation. Organized around a taxonomy of 100+ dynamic events, we build a camera-annotated training mixture combining VLM-filtered real videos, generated hard dynamics, synthetic camera loops, and memory-interruption augmentations. Each clip is converted into a frame graph with protected anchors, degraded intervals, and explicit temporal gaps. A node-structured curriculum—including node-drop, noisy memory, frontier continuation, and reference-cache training—forces the model to retrieve relevant past states across interruptions rather than relying solely on local continuity. PM-RoPE, an elegant camera-phase RoPE extension, unlocks spatiotemporal retrieval at a single-attention cost while preserving pretrained pathways. ReMind achieves the best overall scores on STEVO-Bench and recovery tasks. Furthermore, general image-to-video evaluations confirm this curriculum avoids catastrophic forgetting. We will open-source our code, data, and models.
TED: Text-Axis Evidence Decomposition for Prompted Anomaly Localization
JINYOUNG KIM ⋅ Geonho Kim ⋅ GiJeong Park ⋅ Geonu Lee ⋅ YoungJoon Yoo
CLIP is a powerful vision-language model, but it was not designed for fine-grained defect localization; CLIP-based anomaly detectors therefore adapt it with prompts or lightweight modules to increase defect sensitivity. We show that stronger sensitivity does not necessarily make local evidence reliable: under domain shift, adapted CLIP-AD models often assign high anomaly scores to both true defects and visually complex normal regions. The issue is not simply missing defect information, but a local scoring rule that decodes defect and hard-normal evidence, having the same anomaly evidence. We propose \textsc{TED} (Text-Axis Evidence Decomposition), a post-hoc scoring method that asks whether each ambiguous response is better supported by source defect patches or by source normal patches mistaken as anomalous. \textsc{TED} compares these supports under the host's normal-versus-anomaly text response, leaves the backbone and prompts unchanged, and requires no target-domain training. It works as a train-free score for raw VLM backbones or as a source-calibrated residual correction for adapted CLIP-AD hosts. Across frozen VLM backbones, \textsc{TED} substantially improves pixel-level localization over raw prompt similarity; across adapted hosts, it improves most pixel-level settings over P-AUROC, P-PRO, and P-AP. Gains are largest under stronger hard-FP competition, with mean localization gain increasing from $+5.0$ in low-competition regimes to about $+10.9$ in mid/high-competition regimes. These results suggest that recoverable defect evidence can already exist in pretrained multimodal representations, but reliable localization requires decoding it against hard-normal competitors.
Training-free conditional diffusion provides a flexible alternative to task-specific conditional model training, but existing samplers often allocate computation inefficiently: independent guided trajectories can vary widely in quality, and additional function evaluations along a single trajectory may not recover from poor early decisions. We propose Tempered Guided Diffusion (TGD), an annealed sequential Monte Carlo framework for training-free conditional sampling with diffusion priors. TGD targets tempered posterior distributions over the clean signal, using noisy diffusion states only as auxiliary variables for proposing reconstructions and propagating particles. Particles are reweighted by incremental likelihood ratios, resampled, and propagated across noise levels, concentrating computation on trajectories plausible under both the prior and observation. Under idealized exact-reconstruction assumptions, full TGD yields a consistent particle approximation to the posterior as the number of particles grows. For expensive reconstruction tasks, Accelerated TGD (A-TGD) retains early particle exploration but prunes to a single high-likelihood trajectory partway through sampling. Experiments on a controlled two-dimensional inverse problem and image inverse problems show improved posterior approximation and favorable wall-clock speed-quality tradeoffs over independent multi-trajectory baselines.
TeMPO: Frame-Causal Token Compression for Efficient Video Large Language Model
Yingxin Lai ⋅ Bo Xu ⋅ Yun-ze Pan ⋅ Zhiliang Zhu ⋅ Baigui Sun ⋅ Yang Liu
Video large multimodal models incur substantial inference cost because visual tokens accumulate across sampled frames. Training-free token compression can reduce this cost, but it often degrades video reasoning performance when later frames must contribute new evidence over time. We identify a common failure mode behind this degradation: existing compressors do not track the committed support, namely the token support already retained and delivered to the language model by earlier frames. As a result, selectors with different scoring rules repeatedly retain overlapping supports across neighboring frames, a phenomenon we call Selector Collapse. To address this issue, we propose \tempo{}, a frame-causal token compression framework that combines committed-support memory, residual-energy greedy selection, and parameter-free neighborhood fusion. Across four video understanding benchmarks and multiple backbones, \tempo{} consistently improves the accuracy-efficiency trade-off of training-free compression; at 10\% retention, it preserves 99.8\% of vanilla performance on LLaVA-OneVision and improves fixed-budget frame scaling on Qwen2.5-VL to 109.3\% relative accuracy, without retraining, backbone modification, or additional token budget.
Flow-matching-style generative models learn time-dependent vector fields, but standard objectives supervise each timestep independently, leaving shared-path temporal structure unused. We introduce Temporal Pair Consistency (TPC), a simple training-time objective that couples velocity predictions at paired timesteps along the same probability path. TPC leaves the architecture, probability path, sampler, and inference-time computation unchanged. Under matched backbones, training budgets, solvers, and numbers of function evaluations, TPC improves sample quality across flow matching, OT-CFM, rectified flow, DiT-style backbones, and MeanFlow. On CIFAR-10, TPC improves FM from 6.35 to 3.19 FID and OT-CFM from 3.58 to 2.90 FID at the same NFE. Mechanism diagnostics show lower measured gradient variance, higher paired-gradient correlation, and reduced temporal roughness, while random and local pairing controls are substantially weaker. Higher-resolution ImageNet experiments further show consistent gains for DiT-XL/2 at 256 and 512 resolution and for MeanFlow at 256 resolution, all at matched inference cost.
Temporal Selective Exploration for Reinforcement Learning-Guided Continuous-Discrete Flow Matching in 3D Molecular Design
Lianghong Chen ⋅ Yan Yi Li ⋅ Gen Zhou ⋅ Yuxi Long ⋅ Ganlin Feng ⋅ Mike Domaratzki ⋅ Pingzhao Hu
Flow matching has shown strong potential in generative tasks, while its optimization via Reinforcement Learning (RL) remains underexplored, especially in continuous-discrete mixed settings. Moreover, flow matching is formulated as an Ordinary Differential Equation (ODE), whose deterministic trajectories do not naturally support RL optimization. Introducing stochasticity by converting the entire trajectory into a Stochastic Differential Equation (SDE) is a common practice. Nevertheless, not all timesteps contribute to the final outcomes, and redundant exploration introduces additional noise. Furthermore, in multi-objective settings, some objectives may dominate the optimization while others receive insufficient updates. In this study, we propose an RL-guided continuous-discrete flow matching framework with temporal selective exploration for 3D de novo molecular design. Specifically, our method jointly optimizes continuous and discrete flow matching for atomic coordinates and molecular identities, respectively, while restricting exploration to timesteps that primarily contribute to target properties, reducing ineffective exploration. We also introduce a reward-based advantage calibration mechanism to mitigate objective dominance in multi-objective optimization. The framework is applied to design inhibitors for both the well-studied target protein EGFR and the challenging Pin1 protein. We identify some novel candidate inhibitors, supported by in silico validation.
Tensorion: A Tensor-Aware Generalization of the Muon Optimizer
Vladimir Bogachev ⋅ Vladimir Aletov ⋅ Alexander Molozhavenko ⋅ Sergei Kudriashov ⋅ Maxim Rakhuba
Common first-order optimizers, such as Adam, implicitly treat each parameter block as an unstructured vector, which disregards the multilinear weight structure present in many modern machine learning models. Recent work has shown that exploiting matrix structure can improve optimization dynamics. A notable example is Muon, which performs steepest descent under the spectral norm constraint. We take the next step and introduce Tensorion, a tensor-aware optimizer that extends Muon’s constrained optimization perspective from matrices to higher-order tensors. Tensorion is built around a linear minimization oracle (LMO) over a tensor norm ball. The norm is carefully chosen to balance two objectives: tightly bounding the tensor spectral norm, while still keeping the LMO tractable. This LMO becomes computable because it reduces to operations on adaptively selected unfolding matrices. Notably, when restricted to order‑2 tensors (i.e., matrices), Tensorion recovers Muon exactly. Experiments on tensor-based computer vision problems suggest that Tensorion can offer improved convergence behavior and more stable gradient updates compared with Adam-based and existing tensor-aware baselines in the evaluated settings.
TERMINATOR: Learning Optimal Exit Points for Early Stopping in Chain-of-Thought Reasoning
Alliot Nagle ⋅ Jakhongir Saydaliev ⋅ Dhia Garbaya ⋅ Michael Gastpar ⋅ Ashok Vardhan Makkuva ⋅ Hyeji Kim
Large Reasoning Models (LRMs) achieve impressive performance on complex reasoning tasks via Chain-of-Thought (CoT) reasoning, which enables them to generate intermediate thinking tokens before arriving at the final answer. However, LRMs often suffer from significant overthinking, spending excessive compute time even after the answer is generated early on. Prior work has identified the existence of an optimal reasoning length such that truncating reasoning at this point significantly shortens CoT outputs with virtually no change in performance. However, determining optimal CoT lengths for practical datasets is highly non-trivial as they are fully task and model-dependent. In this paper, we precisely address this and design Terminator, an early-exit strategy for LRMs at inference to mitigate overthinking. The central idea underpinning Terminator is that the first arrival of an LRM's final answer is often predictable, and we leverage these first answer positions to create a novel dataset of optimal reasoning lengths to train Terminator. Powered by this approach, Terminator achieves significant reductions in CoT lengths of 14%–55% on average across four challenging practical datasets: MATH-500, AIME 2025, HumanEval, and GPQA, while outperforming current state-of-the-art methods and reducing inference latency by more than 2$\times$ compared to the original LRM.
Test-Time Graph Recalibration: Enhancing Robust Zero-Shot Inference for Graph Foundation Models
Chunchun Chen ⋅ Zhen Luo ⋅ Xing Wei ⋅ Yuxing Zhang ⋅ Xiaofeng Cao ⋅ Rui Fan ⋅ Ambuj K Singh ⋅ Wei Ye
This work investigates the challenge of robust graph learning within the framework of graph foundation models. Prior studies primarily rely on data-centric structural purification or adversarial augmentation during training to achieve adversarial robustness. However, in practical zero-shot inference scenarios, such methods exhibit significant vulnerability to unseen adversarial structures owing to the static nature of their defense mechanisms and the prohibitive computational cost associated with retraining. In this paper, we propose a novel framework termed Test-time gRaph recAlibration for enhancing robust zero-shot inferenCE (TRACE). The core mechanism of TRACE involves the selective learning of a structural anti-attack directly on noisy graphs during the inference phase, which neutralizes adversarial perturbations while maintaining performance on clean graphs. Specifically, TRACE introduces a non-parametric diagnostic metric based on smoothed spectral entropy to quantify the structural-semantic misalignment induced by adversarial attacks, thereby serving as an adaptive trigger for recalibration. Furthermore, a uniformity-guided optimization objective is formalized to leverage semantic anchors from pretrained text encoders, guiding the distorted graph back to the clean manifold. To ensure computational efficiency for sparse graph encoders, a first-order gradient-based edge flipping strategy is employed to reconstruct the optimal graph structure directly within the discrete domain. Extensive experiments conducted across various benchmark datasets and graph attacks demonstrate the superiority of TRACE over existing state-of-the-art baselines.
Test-Time Prompt-Agnostic Decomposition
Junze Wang ⋅ Lei Fan ⋅ Dezheng Zhang ⋅ Donglin Di ⋅ Yang Song ⋅ Sidong Liu ⋅ Cong Cong
Test-Time Prompt Tuning (TPT) adapts large pretrained models to unseen target distribution shifts by updating only lightweight prompt tokens. Existing TPT methods are effective, but they mainly optimize each prompt update locally. Across test-time steps, prompts updated at the current step are reused for later predictions and further updates, so noisy or biased unlabeled objectives can accumulate and destabilize the prompt-update trajectory. We identify two failure modes: magnitude expansion, where updates move the prompt too far, and directional drift, where consecutive updates point in inconsistent directions. To analyze and control these failures, we propose Test-Time Prompt-Agnostic Decomposition ($\mathtt{TPD}$), which characterizes observed prompt-update trajectories from a dynamical-systems view and decomposes base prompt updates into magnitude and direction. $\mathtt{TPD}$ computes a spectral radius to measure update expansion and decomposes update directions into persistent, oscillatory, and residual components. Building on this decomposition, Adaptive Koopman Control ($\mathtt{AKC}$) regulates update magnitude by shrinking updates under expansive recent dynamics, while Hankel Update Router ($\mathtt{HUR}$) refines update direction by preserving persistent components and suppressing oscillatory and residual components. Together, $\mathtt{AKC}$ and $\mathtt{HUR}$ produce a stabilized prompt update. As a prompt-agnostic framework, $\mathtt{TPD}$ can be plugged into visual and text TPT methods without modifying the backbone, prompt architecture, or adaptation loss. Experiments on 15 datasets show that $\mathtt{TPD}$ consistently improves 10 baselines, with accuracy gains of 3-8\% for visual prompts and over 2\% for text prompts.
Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
Roy Miles ⋅ Aysim Toker ⋅ Andreea-Maria Oncescu ⋅ Jiankang Deng ⋅ Ismail Elezi
Reasoning with large language models often benefits from generating multiple chains-of-thought, but existing aggregation strategies are typically trajectory-level (e.g., selecting the best trace or voting on the final answer), discarding useful intermediate work from partial or “nearly correct” attempts. We propose Stitching Noisy Diffusion Thoughts, a self-consistency framework that turns cheap diffusion-sampled reasoning into a reusable pool of step-level candidates. Given a problem, we (i) sample many diverse, low-cost reasoning trajectories using a masked diffusion language model, (ii) score every intermediate step with an off-the-shelf process reward model (PRM), and (iii) stitch these highest-quality steps across trajectories into a composite rationale. This rationale is then used to recompute only the final answer. This modular pipeline separates exploration (diffusion) from evaluation and solution synthesis, avoiding monolithic unified hybrids while preserving broad search. Across math reasoning benchmarks, we find that step-level recombination is most beneficial on harder problems, and ablations highlight the importance of the final solver in converting stitched but imperfect rationales into accurate answers. Using low-confidence diffusion sampling with parallel, independent rollouts, our training-free framework improves average accuracy by up to 23.8% across six math and coding tasks. At the same time, it achieves up to a 1.8× latency reduction relative to both traditional diffusion models (e.g., Dream, LLaDA) and unified architectures (e.g., TiDAR). The code will be publicly available.
Tethered Predictive-Inertial Proposals with Objective Verification for Diffusion-Prior Inverse Problems
Minwoo Kim ⋅ Seunghyeok Shin ⋅ Dabin Kim ⋅ Hongki Lim
Training-free diffusion priors are powerful for inverse problems, but measurement guidance during reverse sampling is local: aggressive updates can improve immediate data fit while disrupting later denoising. We introduce a tethered predictive-inertial correction rule that separates proposal generation from acceptance. At each corrected step, the denoiser prediction anchors a frozen clean-space objective combining measurement fit with a noise-level-dependent tether. Heavy-ball dynamics with diffusion-scale predictive smoothing generate clean-state candidates; a separate verifier evaluates a finite dyadic set using the original unsmoothed objective, with the denoiser prediction retained as fallback. The accepted clean state is re-noised to continue sampling. The rule applies to pixel and latent diffusion with differentiable measurement operators, requires no retraining, and adds no denoiser calls inside the correction loop. We analyze when smoothing is inactive, how it changes nonlinear or decoder-composed proposal landscapes, and how objective verification yields local nonincrease, scale robustness, and displacement control. Across natural-image and accelerated MRI benchmarks, the method achieves competitive reconstruction quality with favorable speed--quality trade--offs, often reducing runtime relative to optimization-heavy diffusion solvers. On fastMRI knee reconstruction, it attains the strongest PSNR among compared methods at both acceleration factors.
Text-Based AI Tools for Research Integrity Must Be Audited on Linguistic Fairness Before Deployment
Shuai Shao ⋅ Yongkang Wan ⋅ Daoyin Dang ⋅ Lanyun Zhu ⋅ Di Yang ⋅ Yutong Bai ⋅ Yan WANG ⋅ Jiangtao Wang
Research integrity is the cornerstone of scientific progress. The academic community has increasingly adopted text-based Artificial Intelligence (AI) tools to automate research integrity detection, yet whether these tools are fair to researchers from different linguistic backgrounds has rarely been examined. If detection tools themselves harbor linguistic bias, false positives will damage researchers' academic careers and subject entire research communities to unfair treatment. To expose this risk, we conduct a case study of a paper mill (organizations that mass-produce fraudulent academic manuscripts for sale) detection model published in \textit{The BMJ} (impact factor $=$ 43) in January 2026. Through independent reproduction and controlled experiments, we find that the model's predictions are substantially influenced by linguistic style: legitimate non-native English papers from high-impact journals receive a mean paper mill probability of 42.0\%, versus 2.0\% for native English papers. Large Language Model (LLM) based controlled experiments further show that switching writing style alone can flip predictions from negative to positive, and the reverse never occurs. Through further analysis, we argue that these biases are not an isolated case: unavoidable geographic bias in training data, combined with the inherent tendency of language models to encode linguistic style, makes linguistic bias an intrinsic risk for all such tools. Based on these findings, our position is that \textbf{text-based AI tools for research integrity must undergo mandatory fairness auditing for linguistic bias before deployment}. We propose concrete auditing standards and call on the AI community to establish norms that balance detection effectiveness with linguistic equity.
TGPO: Trace-Guided Policy Optimization for Robot Task Planning via Verifiable Subgoal Generation
Zhihong Liu ⋅ Yang Li ⋅ RenMing Huang ⋅ Chendong Zeng ⋅ Cewu Lu ⋅ Panpan Cai
Robot task planning in real-world environments requires mapping abstract natural language instructions to executable action sequences under long horizons and complex constraints. While large language models (LLMs) provide strong commonsense reasoning, they often fail to generate reliable and feasible plans. In contrast, symbolic planners ensure feasibility and optimality but require well-specified goals and cannot directly interpret high-level human intent. We formulate robot task planning as learning to generate verifiable subgoals in the Planning Domain Definition Language (PDDL), bridging language understanding and symbolic planning. To address the challenges of sparse and noisy supervision, we propose Trace-Guided Policy Optimization (TGPO), a reinforcement learning framework that improves structured subgoal generation through (i) verifier-grounded rewards, (ii) external correction of intermediate reasoning traces, and (iii) constrained policy updates that incorporate corrected traces into training. We evaluate TGPO on large-scale household planning tasks with long horizons, abstract instructions, and complex constraints. TGPO significantly outperforms prompting-based and reinforcement learning baselines, with the largest gains on abstract tasks. Furthermore, TGPO integrates naturally with symbolic planners and language-conditioned executors, enabling robust long-horizon planning and combinatorial generalization in realistic environments.
The BAMBI Dataset: Multimodal Nadir UAV-Recordings of Forest Wildlife
Christoph Praschl ⋅ Hugo Markoff ⋅ Anna Maschek ⋅ Wolfram Jantsch ⋅ Stephanie Wohlfahrt ⋅ Anton H Jørgensen ⋅ Christian E Mogensen ⋅ Mathias B Skadhauge ⋅ Sara Beery ⋅ Michael Ørsted ⋅ David C Schedl
Large-scale wildlife monitoring in forested environments using drones remains challenging due to occlusion, limited visibility, and scarce annotated data. We present a comprehensive airborne wildlife dataset comprising 386 paired RGB and thermal aerial video sequences recorded across diverse temperate forests and forest-adjacent habitats in Austria. Each frame is geo-referenced with precise global coordinates (longitude, latitude and altitude), enabling learning and evaluation in both image space and geographic space. The dataset includes 10 animal species captured under varying seasonal, environmental, and illumination conditions, with annotations supporting tasks such as object detection, multi-object tracking and fine-grained classification. The multispectral data and spatial metadata enable research on world-coordinate trajectory analysis, spatial population modeling, geo-aware perception, and advanced methods such as Airborne Light Field Sampling. Thermal subsets of this dataset have been used to develop and validate methods for wildlife monitoring, while the RGB data has mainly stayed untouched. To address this, we present a cross-modal pipeline to transfer thermal annotations to RGB frames next to the dataset. With this dataset we aim to provide a blueprint to promote research on multimodal, geo-referenced perception in ecology.
The Geometry of Alignment Collapse: When Fine-Tuning Breaks Safety
Max Springer ⋅ Chung Peng Lee ⋅ Bohdan Turbal ⋅ Blossom Metevier ⋅ Hayoung Jung ⋅ Jane Castleman ⋅ Zeyu Shen ⋅ Aleksandra Korolova
Fine-tuning aligned language models on benign tasks (e.g. math tutoring) systematically breaks safety guardrails, even when training data contains no harmful content. While mechanistic approaches have shed light on where alignment resides in model weights, they lack the formal language needed to derive guarantees about when and why fine-tuning degrades it — leaving the field without principled tools for predicting or preventing alignment collapse. We develop such a framework through geometric analysis of parameter-space trajectories and apply it to understand the fragility of alignment in fine-tuning. While first-order analysis suggests orthogonal updates are safe, we prove this is illusory: the curvature of the fine-tuning loss induces second-order acceleration that systematically bends trajectories into alignment-sensitive regions. We formalize the central construct of our framework as the Alignment Instability Condition (AIC), three geometric properties that, when present, are sufficient to guarantee degradation. Our main result proves quartic onset of alignment degradation in training steps, determined by how sharply alignment depends on specific parameters and how strongly tasks couple to these parameters. These findings yield formal sufficient conditions under which gradient descent makes alignment inherently fragile, demonstrating the framework's capacity for theoretical guarantees. We further empirically validate the framework's foundations, showing that the Fisher Information Matrix governs the degree of safety degradation across diverse fine-tuning settings.
The Graph Concept Bottleneck: Decoding Combinatorial Reasoning in GNNs for Interpretability
Yue Niu ⋅ zhaokai sun ⋅ Jiayi Yang ⋅ Chunchun Chen ⋅ Xiaofeng Cao ⋅ Rui Fan ⋅ Xin Sun ⋅ Wei Ye
GNNs have achieved remarkable success in graph learning, yet their black-box nature obscures the combinatorial reasoning behind their predictions. A core challenge lies in understanding how GNNs translate topological patterns (graph concepts) into logical rules. Current works only uncover hard Boolean logical rules over graph concepts, which cannot quantify the contribution of each concept to model predictions. Moreover, they are post-hoc methods that generate explanations after training via surrogate models, and thus may deviate from the true combinatorial reasoning of GNNs. In this work, we develop the graph concept bottleneck that enforces the combinatorial reasoning of GNNs to fit soft logical rules over graph concepts, thereby quantifying the contribution of each concept. To further enhance the graph concept bottleneck, we treat graph concepts as "graph words" and graphs as "graph sentences", and leverage language models to learn context-aware graph concept embeddings. Extensive experiments on multiple datasets show that our method GCBMs achieve state-of-the-art performance in both interpretability and classification.
The Heavy Hitter Oracle: Enhancing Frequency Estimation in Skewed Data Streams
Lisa Schmierer ⋅ Ioana-Oriana Bercea
Estimating element frequencies in data streams is a fundamental problem when processing large amounts of data. Recent approaches suggest learning-augmented algorithms that can access a heavy-hitter oracle that knows the most frequent elements. In this paper, we challenge the idea of these complex oracles: First, we show that heavy-hitter predictions do not necessarily yield significant efficiency gains in frequency estimation. We prove that a classical algorithm using slightly more memory, $\Theta(B \log B)$ instead of $\Theta(B)$, matches the performance of learning-augmented methods with perfect predictions. Furthermore, we show that given (perfect) predictions, a trivial mechanism achieves the same performance as state-of-the-art learning-augmented algorithms. On the algorithmic side, we introduce SpaceR, a single-pass algorithm that matches the accuracy of existing two-pass prediction-based methods without requiring any predictions. SpaceR uses randomized sampling to identify and track frequent elements in one pass through the data, achieving near-perfect frequency recovery with minimal memory on both synthetic and real-world datasets.
The Heel of RLVR: Benchmark Glory Should Not Outpace Honest Measurement
Shuo Yang ⋅ Chiyu Ma ⋅ Kexin Huang ⋅ Jinda Lu ⋅ Shaohang Wei ⋅ Xinpeng Liu ⋅ Haoming Meng ⋅ Yuyang Liu ⋅ Shangshang Wang ⋅ Minghao Zhu ⋅ Soroush Vosoughi ⋅ Guoyin Wang ⋅ Jingren Zhou ⋅ Li Yuan
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as the dominant paradigm for eliciting reasoning capabilities in large language models, producing a steady stream of benchmark improvements that have generated considerable excitement in the community. In this position paper, we argue that this reported progress is built on a measurement foundation with systematic, largely unexamined cracks. Through large-scale quantitative analysis of rollout trajectories during Zero-RL training, we identify five structural vulnerabilities organized across two levels. At the algorithm level, we show that destructive over-reflection ("Oops Moments") occurs at nearly three times the rate of the celebrated self-correction ("Aha Moments"), exposing a profound survivorship bias in how progress is reported; and that injecting $\pm20$% noise into the advantage signal produces no measurable effect on training dynamics, calling into question what our optimization algorithms are actually learning. At the implementation level, we show that rule-based verifiers maintain a persistent misjudgment rate of $0.1$%--$0.5$% that silently corrupts reward signals throughout training; that output format choices orthogonal to reasoning ability measurably shift benchmark performance; and that the alignment between loss normalization granularity and micro-batch construction strategy constitutes a hidden hyperparameter that can cause two teams running ostensibly the same algorithm to optimize fundamentally different objectives. Taken together, these findings suggest that the gap between real progress and accurately measured progress in RLVR may be substantially larger than the community currently appreciates. We call for more honest accounting practices in RLVR research before the next benchmark milestone is celebrated.
The Interplay of Data Structure and Imbalance in the Learning Dynamics of Diffusion Models
Flavio Nicoletti ⋅ Chenxiao Ma ⋅ Enrico Ventura ⋅ Luca Saglietti ⋅ Stefano Sarao Mannelli
Real-world datasets are inherently heterogeneous, yet how per-class structural differences and sampling imbalance shape the training dynamics of diffusion models—and potentially exacerbate disparities—remains poorly understood. While models typically transition from an initial phase of generalization to memorizing the training set, existing theory assumes homogeneous data, leaving open how class imbalance and heterogeneity reshape these dynamics. In this work, we develop a high-dimensional analytical framework to study class-dependent learning in score-based diffusion models. Analyzing a random-features model trained on Gaussian mixtures, we derive the feature-covariance spectrum to characterize per-class generalization and memorization times. We reveal the explicit hierarchy governing these dynamics: class variance is the primary determinant of learning order—consistently favoring higher-variance classes—while centroid geometry plays a secondary role. Sampling imbalance acts as a modulator that can reverse this ordering and, under strong imbalance, forces minority classes to acquire distinct, delayed speciation times during backward diffusion. Together, these results suggest that diffusion models can memorize some classes while others remain insufficiently learned. We validate our theoretical predictions empirically using U-Net models trained on Fashion MNIST.
The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
Sheng-Yu Wang ⋅ Yotam Nitzan ⋅ Aaron Hertzmann ⋅ Jun-Yan Zhu ⋅ Eli Shechtman ⋅ Alexei Efros ⋅ Richard Zhang
Human visual similarity judgments are context-dependent. For example, two images may be similar in shape but distinct in color. Existing perceptual similarity metrics, however, collapse these nuances into a single scalar value, offering no mechanism to condition on specific aspects. To bridge this gap, we introduce a large-scale dataset of human similarity judgments over image triplets, where each triplet is annotated across multiple, free-form semantic aspects. Benchmarking a broad range of frontier vision-language models (VLMs) reveals a considerable performance gap compared to human-human agreement. Leveraging our data, we fine-tune a VLM to produce our Text-Prompted Image Perceptual Similarity (TPIPS) metric, capturing multiple senses of visual similarity depending on the specified text prompt. We demonstrate that TPIPS aligns more closely with human perception and generalizes reliably beyond the training distribution. Finally, we show that TPIPS unlocks new capabilities in text-guided retrieval, compositional search, and the fine-grained evaluation of generative models. Our code, data, and trained models will be released.
The Minimax Rate of Second-Order Calibration
Kamil Ciosek ⋅ Banafsheh Rafiee ⋅ Sina Ghiassian ⋅ Nicolò Felicioni
We characterize the minimax rate of estimating the second-order calibration error for binary classification, which quantifies whether a higher-order predictor's epistemic-uncertainty estimate matches the conditional variance of the label probability on its level sets. Our key observation is that the sech perturbation kernel, previously used only to enforce smoothness of calibration functions, in fact makes them \emph{analytic} in a strip of half-width $h\pi/2$. Polynomial regression then estimates the calibration error at rate $\tilde{O}(1/\sqrt{n})$, with explicit constants, a qualitative improvement over the $O(n^{-1/4})$ rate achievable by bucketing or kernel smoothing. A matching $\Omega(1/\sqrt{n})$ lower bound establishes minimax optimality up to logarithmic factors. As a corollary, we give the first finite-sample guarantee for second-order Platt scaling, yielding a post-hoc procedure that recalibrates both the mean prediction and the epistemic-variance estimate of any higher-order predictor. Along the way, we provide a bucket-free definition of second-order calibration and relate it quantitatively to the bucketed formulation of Ahdritz et al. [2025]. Our experiments confirm the predicted rate and the quality of the recalibrated uncertainties.
The Missing Corpus: Infinite Ground Truth for File-Grounded LLM Evaluation
Anush Sankaran ⋅ Arijit Banerjee ⋅ Pamela Bhattacharya ⋅ Tanujay Saha
Evaluating enterprise LLMs on file-grounded Question Answering (QA) requires large, labeled corpora of realistic enterprise PDFs, yet real documents are sensitive, scarce, and impossible to annotate at scale. Existing synthetic approaches produce visually crude renders, annotate ground truth post-hoc (introducing hallucination risk), or are non-reproducible one-off scripts with no diversity control. We present SynthDocQA, an Intermediate Representation(IR)-driven pipeline that co-generates photorealistic multi-page enterprise PDFs and provenance-grounded evaluation pairs in a single pass. A compact Document Generation IR encodes topic, content schedule, and perturbation profile before rendering begins. Templates are retrieved per page via Contrastive Language-Image Pretraining (CLIP) embedding-based cross-modal search over a 10,000-layout library; a 15-stage physical degradation pipeline models five scan quality tiers from pristine to poor. Every QA assertion is derived analytically from the same structured artifact that populated the page, making hallucination structurally impossible, and a hierarchical fork-based Random Number Generator (RNG) guarantees exact reproducibility from a single integer seed. A reference 100-document corpus spans 50 enterprise domains, 4,374 pages, 16,589 QA pairs, and 21,402 assertions at zero annotation cost.
The Surprising Effectiveness of Video Diffusion Models for Hand Motion Reconstruction
Yuxi Wang ⋅ Chengkai Jin ⋅ Yufei Liu ⋅ Wenqi Ouyang ⋅ Tianyi Wei ⋅ Zhiwei Zeng ⋅ Siyuan Huang ⋅ Zhiqi Shen ⋅ Xingang Pan
4D hand motion reconstruction from egocentric video is bottlenecked by clear limitations of existing methods: image-based pipelines depend on a detector that fails under heavy occlusion, while video-based methods rely on temporal modules learned only from hand-pose annotations, a signal too narrow to capture motion, occlusion, and hand-object interaction. These capabilities, however, are exactly what video generative models must implicitly acquire when trained to synthesize coherent video at internet scale. Motivated by this, we present ViDiHand, which leverages the representations of a pretrained video diffusion model to reconstruct 4D two-hand pose. We adapt it via a hand-overlay rendering objective that specializes its features for hands while preserving its world priors, and a decoder recovers metric-scale pose from the adapted features. The whole pipeline runs in a single forward pass over full frames—no detector, no infiller, and no test-time optimization. On ARCTIC, HOT3D, and HOI4D, ViDiHand substantially outperforms prior methods, establishing video diffusion models as a powerful new foundation for hand motion reconstruction and a promising route to scalable in-the-wild data collection for embodied AI.
TIB: Sample-wise Tempered Information Bottleneck for Multimodal Attribution beyond Alignment Assumption
Qiming Huang ⋅ Peixi Liu ⋅ Jianbo Jiao
Multimodal (e.g. vision–language, in this work) attribution aims to interpret the models with- out access to ground-truth explanatory supervi- sion. Existing multimodal attribution methods typically rely on paired modalities as proxy su- pervision, implicitly assuming that image–text pairs are semantically aligned. Information bot- tleneck (IB)–based attribution methods follow this general paradigm and instantiate it through a global sufficiency–compression objective. In real- istic multimodal datasets, however, this assump- tion is frequently violated due to partial semantic misalignment and unreliable cross-modal corre- spondences. Through theoretical analysis, we show that, under such misalignment, information- bottleneck-based attribution objectives exhibit a structural limitation: by assuming equal reliabil- ity across all proxy supervision signals, they in- duce posterior over-contraction and forced expla- nations. To address this issue, we propose TIB, a sample-wise Tempered Information Bottleneck framework that adaptively modulates attribution strength according to cross-modal reliability, with- out assuming semantic alignment. Extensive ex- periments on controlled benchmarks and large- scale datasets demonstrate the effectiveness of the proposed approach, while also validating the identified failure mode
TIGER: Bridging the Multimodal Reasoning-Access Gap via Modality Counterfactuals
Gregory Kang Ruey Lau ⋅ Huynh Minh Nguyen ⋅ Bryan Kian Hsiang Low
Multimodal Large Language Models (MLLMs) can exhibit strong text-based reasoning yet fail to apply the same capabilities to semantically equivalent visual inputs. We study this failure mode in a controlled setting by rendering text-only reasoning problems as images, preserving semantic content while changing only the input modality. Across multiple MLLM families, models solve the same problems substantially better from text than from images. We characterize this discrepancy as a reasoning-access gap, where models may correctly perceive visual content but still fail to route that content into the latent reasoning machinery used for text-based tasks. To bridge this gap, we propose TIGER (Text-to-Image Gap-targeted Training for Enhanced Reasoning), a framework that repurposes text-only reasoning corpora into effective multimodal training data. By mining "modality counterfactuals", instances where a model succeeds on a text problem but fails on its semantically-equivalent rendered image, TIGER provides targeted supervision for reasoning-access failures without requiring manually curated multimodal datasets. We implement TIGER using image-conditioned Group Relative Policy Optimization (GRPO), along with SFT and DPO variants, and show how it consistently narrows the modality gap and yields generalized visual reasoning performance gains in multimodal reasoning benchmarks such as MathVerse and EMMA. We show that Reinforcement Learning with Verifiable Reward-based models still exhibit modality-dependent reasoning gaps, and that our training recipe can further reduce these gaps. Further analysis using reasoning-subspace activation and activation patching shows that TIGER enables visual representations that better activate reasoning-relevant latent subspaces of the language backbone. Our results suggest that advancing multimodal intelligence requires moving beyond perceptual alignment toward explicit visual access to existing reasoning machinery.
TIMA: Test-Time Internalization for Agentic Memory
Xiaohang Sui ⋅ Yongjian Fu ⋅ Yizhe Zhao ⋅ Sheng Yue ⋅ Ju Ren
Agentic memory is critical for accumulating and reusing experience in LLM agents across long interaction horizons and related tasks. While latent-memory approaches offer a compact representation of such experience without incurring additional context overhead, existing methods are designed to encode transient in-context states or rely on pre-trained modules that remain fixed at test time to inject experiential representations. Consequently, an effective mechanism for continually refining reusable latent memory from execution experience is still missing in current agents.To this end, we propose TIMA, a framework for test-time internalization that enables agents to acquire and update latent memory during interaction. TIMA couples a memory-augmented agent loop with an online-updated Internalizer and an Internalizer Bank for retrieval and continual refinement across episodes. Through self-supervised updates derived from execution trajectories, TIMA progressively encodes task-relevant solving patterns into compact latent memory, enabling both within-episode adaptation and cross-episode transfer. Extensive experiments across eight benchmarks show that TIMA consistently outperforms strong baselines, surpassing A-Mem by up to 50.8% and latent-memory baselines such as MemGen and Titans by up to 4.47%--15.96%. Further analysis indicates that the learned latent memory remains stable under continual updates, generalizes well out of distribution, and exhibits interpretable clustering structure across Internalizers.
Time-Frequency Decoupled Cross-Scale Partial Optimal Transport for Time Series Domain Adaptation
Shubin Liu ⋅ Xinyang Chen ⋅ Xiucheng Li ⋅ Weili Guan ⋅ Liqiang Nie
Unsupervised domain adaptation has become an important paradigm for mitigating distribution shift in time series classification. However, existing time series domain adaptation methods typically align source and target data at a fixed time scale and simply fuse time-frequency features before uniform alignment. Furthermore, they do not consider the impact of noisy samples on distribution alignment. To address these limitations, we propose CSPAN (Time-Frequency Decoupled \textbf{C}ross-\textbf{S}cale \textbf{P}artial Optimal Transport \textbf{A}lignment \textbf{N}etwork) for time series domain adaptation. CSPAN first constructs segment-scale time-frequency candidates, and then uses Segment-Scale Aware Top-k Fusion to select the most suitable segment-scale feature representation. It then performs cross-domain alignment after decoupling time and frequency through Temporal Partial Optimal Transport and Frequency Partial Optimal Transport, where temporal domain alignment is guided by the transferability estimated from a probing transport plan, and frequency domain alignment is guided by the discriminability estimated by an auxiliary frequency classifier. By combining segment-scale representation selection, time-frequency decoupling, and partial distribution alignment, CSPAN enables more flexible cross-domain matching. Extensive experiments on six datasets demonstrate CSPAN’s consistent superiority, achieving an average accuracy improvement of 3.16\% in cross-domain scenarios. Code is available at the anonymous link: \url{https://anonymous.4open.science/r/CSPAN-35FB/}.
Time–Frequency Non-Stationary Modeling for Multivariate Time Series Forecasting
Haoyi Zhao ⋅ Yishan Jiang ⋅ Jiqian Yang ⋅ Ji Chang ⋅ Dawei Ma ⋅ Wenjun Lv
Multivariate time series forecasting is important in real-world applications, but practical data often show both temporal and spectral non-stationarity. Existing methods may treat severe spectral drift as informative evolution, allow unreliable spectral components to contaminate temporal representations through credibility-agnostic time--frequency interaction, and infer cross-variable dependencies from noisy local correlations, which may lead to spurious long-term dependencies. To address these issues, we propose TFNS, a dual-stream time--frequency framework for suppressing unreliable spectral evidence under non-stationarity. Specifically, we introduce a Spectral Credibility Estimator to quantify patch-level spectral stability and softly suppress low-credibility frequency components for more reliable frequency modeling; a Credibility-Gated Cross-Interaction module to enable bidirectional time--frequency enhancement while adaptively constraining cross-domain information injection according to spectral reliability; and a Spectral Cointegration Graph to capture stable long-term cross-variable structure in the spectral domain while mitigating spurious dependencies under non-stationary drift. Together, these modules form a progressive pipeline from local spectral purification to reliability-aware cross-domain interaction and robust long-term structure modeling. Experiments on 11 benchmarks show that TFNS achieves state-of-the-art performance, ranking first in 44/55 MSE and 43/55 MAE comparisons.
Sparse Mixture-of-Experts (MoE) layers underpin the most capable open-weight large language models, including DeepSeek-V3.2, Qwen3-MoE, and the Llama 4 family, but their pretraining is plagued by an early-stage pathology in which a handful of experts dominate the routing distribution while the remainder are starved of gradient signal. Existing remedies treat the symptom rather than the cause: load-balancing auxiliary losses, sequence-level balance constraints, and router z-loss penalties all push the gating distribution toward uniformity through additional gradient signals that compete with the language-modeling objective and require careful coefficient tuning. We argue that the root issue is a positivefeedback loop between expert capability and expert selection, and we propose Token-Conditional Expert Dropout (TCED), a training-time perturbation that masks the chosen expert for a token with probability proportional to that expert’s recent utilization and re-routes the token to the highest-scoring alternative. TCED is parameter-free, requires no auxiliary loss, and provides implicit regularization analogous to attention dropout while breaking the collapse feedback loop. We derive an unbiased gradient estimator under TCED, prove a variance-reduction lemma that bounds the variance of normalized expert utilization, and pretrain a 1.4B-active / 8B-total MoE on 480B tokens of a high-quality web mixture. TCED reduces expert-utilization variance by 67% across training, lifts MMLU-Pro by 2.1 points and GPQA by 2.1 points relative to a vanilla MoE, and matches a heavily tuned auxiliary-loss baseline without a single coefficient sweep. Results are preliminary single-seed estimates; final-version revisions will report multi-seed confidence intervals.
Token Filtering: Online Attention Pruning via KV Similarity for Efficient LLM Inference
JUNGMIN LEE ⋅ Gwangeun Byeon ⋅ Yulhwa Kim ⋅ Seokin Hong
Pruning has emerged as a promising direction for accelerating large language model (LLM) inference. However, many existing methods rely on offline calibration data, making them sensitive to distribution shifts between calibration data and inference inputs. In this paper, we introduce Token Filtering, a lightweight online pruning method that selectively skips attention computation for redundant tokens during inference, without requiring any calibration data or fine-tuning. Token Filtering identifies redundant tokens based on joint key–value (KV) similarity and bypasses their attention computation. This approach reduces both compute and KV cache size while preserving essential contextual information. To preserve accuracy while meeting the target pruning ratio, we restrict pruning to later layers, which are typically less sensitive to pruning, with a layer-wise threshold that adaptively tracks the target pruning ratio. Extensive experiments on LLaMA3.1-8B, LLaMA3.1-8B-Instruct, and Qwen3-8B demonstrate that Token Filtering consistently achieves better accuracy–efficiency trade-offs than prior methods. In long-context generation, Token Filtering achieves up to 1.5X higher accuracy than the best-performing pruning baseline while reducing latency by up to 44% compared to the dense baseline.
Tool-IQA: Augmenting Image Quality Assessment with Simple Tools
Guanyi Qin ⋅ Junjie Zhang ⋅ Chunming He ⋅ Yibing Fu ⋅ Jie Liang ⋅ Tianhe Wu ⋅ Lei Zhang
Vision-Language Models (VLMs) have been increasingly adopted for Image Quality Assessment (IQA). However, current methods typically employ a static one-shot scoring paradigm, despite the fact that humans assess image quality through dynamic visual inspection, e.g., selectively adjusting views to verify details and subtle artifacts. Specifically, relying solely on a single-pass observation introduces two primary limitations: first, perceiving the image only at a global scale restricts the assessment of finer local details; second, the original intensity distribution of the image may overwhelm the visibility, leading to insufficient inspection of image quality. To address these issues, we propose Tool-IQA, shifting the assessment mechanism from passive scoring to a tool-augmented workflow. In particular, we equip VLMs with simple yet effective view tools: a Magnifier to inspect local details, and a Gamma Corrector to uncover visibility and hidden artifacts. The assessment follows a structured pipeline that consists of an initial observation with rubric notes, a tool-augmented in-depth inspection, and a final quantification for calibrated quality score. Furthermore, to ensure efficient and purposeful tool callings, we introduce a batch-aware training strategy to reward tool interactions that can yield positive contributions rather than simply encouraging usage. Experiments on a variety of IQA benchmarks demonstrate that, with effective tool calling and calibrated assessment, our proposed Tool-IQA significantly outperforms existing state-of-the-art models, e.g., it achieves a PLCC of 0.854 on the challenging CLIVE dataset. The code and models will be released.
Topology-Aware Representation Alignment for Semi-Supervised Vision-Language Learning
Junwon You ⋅ Mihyun Jang ⋅ Sangwoo Mo ⋅ Jae-Hun Jung
Vision-language models have shown strong performance, but they often generalize poorly to specialized domains. While semi-supervised vision-language learning mitigates this limitation by leveraging a small set of labeled image-text pairs together with abundant unlabeled images, existing methods remain fundamentally pairwise and fail to model the global structure of multimodal representation manifolds. Existing topology-based alignment methods rely on persistence diagram matching, which neither guarantees geometric alignment nor utilizes the image-text pairing information central to vision-language learning. We propose Topology-Aware Multimodal Representation Alignment (ToMA), a framework that uses persistent homology to identify topologically salient edges and aligns them across modalities through available cross-modal correspondences. ToMA leverages both $H_0$-death edges and lightweight $H_1$-birth edges, allowing it to capture both connectivity and cycle structure without constructing 2-simplices. Experiments show that ToMA yields stable gains, with clear improvements on remote sensing and modest but consistent benefits on fashion retrieval. Additional analysis shows that ToMA is more stable than alternative topology-based objectives and that lightweight $H_1$-birth edges provide useful higher-order structural signals.
Toward Efficient Reasoning of Large Language Models via Latent Concept-Pyramid Modeling
Sijia Chen ⋅ Ningxin Su
Large language models (LLMs) perform reasoning via lengthy token-by-token generation, incurring substantial inference cost. While recent methods compress this process by enabling LLMs to reason in a latent space, they still rely on next-vector generation for sequential logic and on flat representations. We therefore present Latent Concept-Pyramid Modeling (\emph{L}CP), a new paradigm that reformulates LLM reasoning as hierarchical next-level concept generation in a coarse-to-fine manner. Specifically, \emph{L}CP generates the pyramid level by level: a single apex concept encodes the most abstract semantics, and each subsequent level refines its predecessor into finer-grained concepts that together capture distinct CoT segments at one granularity. Such next-level concept prediction approximates the reasoning trace in a manner analogous to advancing from a skeletal outline, through broad structural forms, down to local details. Experimental evaluations on zero-shot mathematical and coding reasoning benchmarks demonstrate that Qwen2.5 and Qwen3 models fine-tuned with \emph{L}CP achieve significant reductions in both token cost and inference latency while maintaining competitive accuracy. \emph{L}CP further showcases zero-shot generalization ability across different tasks. Moreover, the pyramid's hierarchical, multi-grained representational capacity enables faithful reconstruction of the original CoT, endowing \emph{L}CP with interpretability.
Towards Error-Free EHRs: Reasoning-Intensive Consistency Verification Between Clinical Notes and Structured Tables in Electronic Health Records
Yeonsu Kwon ⋅ Jiho Kim ⋅ Junseong Choi ⋅ Paloma Rabaey ⋅ Minseo Kim ⋅ Sujeong Im ⋅ Jeewon Yang ⋅ Jun-Min Lee ⋅ Sangji Lee ⋅ Jiwon Kim ⋅ Hangyul Yoon ⋅ Hyunwook Kwon ⋅ Edward Choi
Data consistency between unstructured clinical notes and structured tables in Electronic Health Records (EHRs) is essential for patient safety and clinical decision-making. However, existing work on note-table consistency verification mainly relies on surface-level matching of numeric values or simple events. Such approaches fail to capture the reasoning underlying real-world EHR documentation, including clinical interpretation, event relations, and temporal changes. To address this gap, we introduce EHR-ReasonCon, a reasoning-intensive benchmark for note-table consistency verification. Built on MIMIC-III with expert-guided annotations, it comprises 8,048 entities derived from clinical notes and provides high-quality ground-truth labels. The annotation protocol is supported by specialized table-exploration tools to ensure systematic evidence retrieval and reliable consistency assessment. We also propose EHR-INSPECTOR, an LLM-based framework that segments notes, extracts anchor entities and temporal references, and uses table-exploration tools to verify consistency against structured tables. Evaluated using expert-validated LLM-as-a-judge metrics under harsh and lenient criteria, EHR-INSPECTOR achieves state-of-the-art performance across multiple model backbones. Analyses further demonstrate the effectiveness of its components and highlight differences from human verification. The code is available at https://anonymous.4open.science/r/EHR-ReasonCon-1A54/.
Towards On-Policy Data Evolution for Visual-Native Multimodal Deep Search Agents
Shijue Huang ⋅ Hangyu Guo ⋅ Guanting Dong ⋅ Chenxin Li ⋅ Junting Lu ⋅ Xinyu Geng ⋅ Zhaochen Su ⋅ Zhenyu Li ⋅ Shuang Chen ⋅ Hongru WANG ⋅ Yi R. (May) Fung
Multimodal deep search requires an agent to solve open-world problems by chaining search, tool use, and visual reasoning over evolving textual and visual context. Two bottlenecks limit current systems. First, existing tool-use harnesses treat images returned by search, browsing, or transformation as transient outputs, so intermediate visual evidence cannot be re-consumed by later tools. Second, training data is usually built by fixed curation recipes that cannot track the target agent's evolving capability. To address these challenges, we first introduce a visual-native agent harness centered on an image bank reference protocol, which registers every tool-returned image as an addressable reference and makes intermediate visual evidence reusable by later tools. On top of this harness, On-policy Data Evolution (ODE) runs a closed-loop data generator that refines itself across rounds from rollouts of the policy being trained. This per-round refinement makes each round's data target what the current policy still needs to learn. The same framework supports both diverse supervised fine-tuning data and policy-aware reinforcement learning data curation, covering the full training lifecycle of the target agent. Across 8 multimodal deep search benchmarks, ODE improves the Qwen3-VL-8B agent from 24.9% to 39.0% on average, surpassing Gemini-2.5 Pro in standard agent-workflow setting (37.9%). At 30B, ODE raises the average score from 30.6% to 41.5%. Further analyses validate the effectiveness of image-bank reuse, especially on complex tasks requiring iterative visual refinement, while rollout-feedback evolution yields more grounded SFT traces and better policy-matched RL tasks than static synthesis.
Towards Optimism-Pessimism Trade-off in Model-based Offline-to-Online Reinforcement Learning
Guochen Zhou ⋅ Yijun Yang ⋅ Qiqi Duan ⋅ Qing Su ⋅ Weiming Ou ⋅ Li Shen ⋅ Shihao Ji ⋅ Yuhui Shi ⋅ Zhisong Zhang
Model-based offline-to-online reinforcement learning (RL) enables sample-efficient adaptation by leveraging offline pre-training for online fine-tuning. However, the distribution shifts between offline and online stages often hinder fine-tuning performance. Many existing methods approach this problem by adjusting the optimism-pessimism trade-off via a single-objective formulation, requiring costly online bi-level optimization. We identify this trade-off during offline training as a key challenge: optimistic policies generalize better to novel tasks by exploring out-of-distribution states and actions, while pessimistic policies remain constrained to the offline data distribution and excel on similar tasks. To address this challenge, we propose a bi-objective formulation that captures this trade-off, yielding a Pareto policy pool during offline training. These policies enable flexible selection for various online tasks. To generate the pool, we introduce Multi-Objective Soft Actor-critIC (MOSAIC), which solves bi-objective problems and constructs diverse Pareto policies. After offline training, a contextual bandit algorithm hierarchically selects the most suitable policy for fine-tuning at each online interaction step. Empirically, our pipeline, Hierarchical ParetoPolicy Pool (HiP3), achieves state-of-the-art performance on a suite of continuous-control offline-to-online benchmarks, particularly under shifted online tasks. Comprehensive ablations further clarify the roles of the policy pool and the online selection mechanism, as well as the robustness of the overall framework.
Towards Precise Knowledge Distillation for Large Language Models via Knowledge Probing
Jiajun Liu ⋅ Yao He ⋅ Wenjun Ke ⋅ Peng Wang ⋅ Ziyu Shang ⋅ Guozheng Li ⋅ Zijie Xu
Knowledge distillation (KD) has become a widely adopted approach for compressing large language models (LLMs). Existing distillation paradigms for LLMs can be broadly categorized into black-box and white-box approaches. Black-box methods typically transfer the chains of thought (CoTs) generated by teacher models to student models, yet they fail to convey the intrinsic knowledge embedded within teacher models. White-box methods aim to alleviate this limitation by leveraging internal representations from teacher models. However, they still face two major challenges. First, it is difficult to identify which knowledge is essential for distillation in student models. Second, it is challenging to transfer this key knowledge from teacher models accurately. To tackle these challenges, we propose KPD, a novel Knowledge Probing-based white-box Distillation framework for LLMs, which is designed to precisely distill the knowledge that the student model lacks and can be incorporated into diverse distillation approaches. Specifically, to identify critical knowledge gaps in student models, we introduce an uncertainty-based probing method that extracts tokens prone to student errors as critical knowledge. To locate the corresponding knowledge in the teacher model, we utilize cumulative gradient calculation to probe where such knowledge is stored and then distill the target layers. Extensive experiments across multiple datasets demonstrate that KPD boosts traditional distillation methods by 0.34%-4.56% in Rouge-L scores. Further ablations show that probing both the student and teacher yields more precise distillation. Our code is available at https://anonymous.4open.science/r/KPD-CA0D.
Towards Reliable VLM Judges: State-Conditional Invariance and Presentation-Aware Diagnostics
Qinan Zhang ⋅ Xinyu Wang ⋅ Qihang Jin ⋅ Tianyuan Wang ⋅ Jian Di ⋅ Hongbing Luo ⋅ Pengfei Li ⋅ Wenjun Lv ⋅ Yu Kang
Vision-language models (VLMs) are increasingly used as automated judges for multimodal and embodied tasks, yet aggregate accuracy or consistency alone does not reveal whether a judge changes its verdict for the right reason. We propose State-Conditional Invariance (SCI), a reliability principle requiring a judge to remain stable under semantics-preserving presentation changes while responding to causal interventions that change the task outcome. We instantiate SCI in GridWM-Judge, a simulator-grounded diagnostic benchmark built from deterministic MiniGrid trajectories with aligned Full, evidence-removal, and counterfactual variants, plus presentation and representation probes. Across five diagnostic tasks, GridWM-Judge tests local perception, action-conditioned transitions, presentation stability, representation exchangeability, and Full-CF causal discrimination. Evaluating current VLM judges reveals separable failures, including presentation sensitivity, representation fragility, causal misgrounding, action-insensitive heuristics, and stability traps. These results show that reliable VLM-as-a-judge evaluation requires joint reporting of correctness, presentation invariance, and causal sensitivity rather than a single accuracy or consistency score.
Quantum autoencoders (QAEs) are learning architectures that compress quantum data into a low-dimensional latent state while preserving the information needed for reconstruction. We study blind single-copy compression of quantum states through a $k$-qubit bottleneck and investigate the minimal circuit width required to attain the information-theoretic optimum under average infidelity. Between the conventional architecture, which is narrow but nonuniversal, and fully general *completely positive and trace preserving* (CPTP) realizations, which are universal but overparameterized, we identify a balanced regime. We prove that for every distribution of pure $n$-qubit states, there exists a QAE with exactly $k$ encoder ancillas and $n$ decoder ancillas that achieves the optimal fidelity over all CPTP encoder–decoder pairs. The encoder-side statement is sharp in that we construct source families for which every optimal scheme necessarily uses at least $k$ encoder ancillas, thereby determining the universal encoder threshold exactly. On the decoder side, we show that isometric decoders are exactly optimal for several analytically tractable source families, but we also exhibit an explicit counterexample demonstrating that decoder isometry is not universally sufficient. Nevertheless, numerical experiments indicate that the performance gap is practically negligible.
TRACE: Data-Free Text Reconstruction Attacks against Approximate Unlearning in LLMs
Mengying Zhang ⋅ Derui Wang ⋅ Ibrahim Khalil ⋅ Bowen Liu ⋅ Xiaoyu Xia ⋅ Minhui Xue
As large language models (LLMs) face growing demands for data removal driven by regulatory compliance, copyright concerns, and privacy protection, approximate unlearning has emerged as a practical alternative to expensive full retraining. In practice, pre- and post-unlearning model snapshots are often maintained for auditing and rollback, introducing a previously overlooked side channel. This paper demonstrates that approximate unlearning in LLM can leave residual traces within model snapshots, which may be recoverable and thus enable the privacy breaches that approximate unlearning is intended to prevent. Here, we propose TRACE, a framework that reconstructs unlearned training data given access to pre- and post-unlearning model snapshots, without any data-specific prior knowledge. To achieve this, TRACE identifies a compact set of candidate tokens from embedding-level parameter differences to constrain the search space, and assembles sequences via a contrastive scoring mechanism based on the perplexity gap between the two models. A two-stage exploration-then-completion strategy enables the recovery of multiple distinct samples. Extensive experiments on various datasets, including TOFU, MUSE, and WMDP, across multiple LLMs and six representative unlearning methods, demonstrate that TRACE achieves high-fidelity reconstruction. This paper exposes significant privacy risks in current LLM approximate unlearning deployments and highlights the need for defenses against parameter-difference leakage and appropriate management of model snapshot access.
TraceTriage: A Benchmark for Cost-Aware Stop-or-Continue Decisions in Delayed-Outcome Workflows
Sihang Lei ⋅ Yihang Qiu ⋅ Xueyan Zhao ⋅ Yuwei Wang ⋅ Weiqiang Wang
Current agent benchmarks largely evaluate whether a workflow eventually succeeds, but many real workflows require an earlier decision: whether to continue spending resources or stop and reallocate them before the final outcome is known. In model training, simulation campaigns, and chip implementation, success or failure may be visible only after substantial compute and engineering time. Wrong stops discard runs that would have succeeded; wrong continuations waste downstream compute, tool time, and engineering effort. This asymmetric, risk-constrained intervention problem is not captured by benchmark scores that only measure eventual success. We present TraceTriage, a benchmark for evaluating LLM agents and decision models on cost-aware stop-or-continue policies from checkpoint-local process records. TraceTriage is not a final-outcome prediction benchmark: it evaluates deployable interventions using validation-only thresholding under a false-stop cap and remaining-cost utility. We instantiate the protocol with OpenROAD, an open-source electronic-design-automation workflow that turns hardware designs into placed-and-routed chip layouts through staged tools. OpenROAD is a reproducible substrate, not the benchmark's scope: it provides staged observability, heterogeneous logs and reports, measurable remaining runtime, reproducibility, and verifiable outcomes. The released benchmark uses no-leak run-level splits and instantiates this protocol through 10 OpenROAD design-optimization campaigns with a 50,000-trial budget, retaining 31,572 runs and expanding them to 81,041 checkpoint prefixes with fixed train/validation/test splits. Each example contains only evidence available at a workflow checkpoint, while remaining-cost annotations are used only for evaluation. Together, the results make TraceTriage a capability-boundary benchmark for cost-aware workflow intervention: it verifies actionable signal in checkpoint-local records, quantifies transfer sensitivity, and exposes the unresolved cold-start calibration gap in delayed-outcome, risk-constrained settings.
Tracing the Cascade: A Topology-Aware Evaluation Framework for Scientific Agent Hallucinations
Xinshun Feng ⋅ Ziqi Miao ⋅ Lijun Li ⋅ Jing Shao
Large language model (LLM) agents are increasingly deployed in scientific research, where reliability is critical, and the underlying knowledge is densely interconnected. In such settings, hallucinations are particularly damaging: a single erroneous claim on a foundational concept can propagate through multi-step reasoning and corrupt entire trajectories. Existing hallucination benchmarks largely operate at the surface level, treating facts in isolation and relying on uniform accuracy metrics that ignore this topological structure. We address this gap with \textsc{Schema}, the first evidence-grounded, topology-aware evaluation framework for scientific agents. Built on an automated concept-graph construction pipeline, \textsc{Schema} provides two complementary diagnostic instruments. A trajectory hallucination pipeline audits intermediate reasoning at scale via a topology-weighted severity score ($\mathrm{HS}^w$), while a multi-agent counterfactual attribution module pinpoints the causal mechanism behind selected failures. The framework supports diverse task formats spanning multi-hop reasoning, claim verification, and programmatic experimental execution. Across two biomedical subdomains and eleven contemporary LLM agents, \textsc{Schema} reveals that hallucinations concentrate at a small set of highly connected knowledge hubs, and that final-answer accuracy decouples from trajectory honesty—models often reach correct conclusions through structurally flawed reasoning. These results indicate that for high-stakes scientific applications, terminal accuracy alone is an insufficient signal of agent reliability, motivating mechanism-level evaluation grounded in knowledge topology\footnote{Code is available at \url{https://anonymous.4open.science/r/SCHEMA_NIPS2026}}.
Trainable Sparse Attention via Hybrid Top-k+Top-p Masking and Distillation Fine-Tuning
Jintao Zhang ⋅ Kai Jiang ⋅ Chendong Xiang ⋅ Weiqi Feng ⋅ Yuezhou Hu ⋅ Haocheng Xi ⋅ Shuo Yang ⋅ Mengfei Xia ⋅ Jianfei Chen ⋅ Jun Zhu
Many training-free sparse attention methods are effective for accelerating diffusion models. Recently, several works suggest that making sparse attention trainable can further increase sparsity while preserving generation quality. We study three key questions: (1) when do the two common masking rules, i.e., Top-k and Top-p, fail, and how can we avoid these failures? (2) why can trainable sparse attention reach higher sparsity than training-free methods? (3) what are the limitations of fine-tuning sparse attention using the diffusion loss, and how can we address them? Based on this analysis, we propose SpargeAttention2, a trainable sparse attention method that achieves high sparsity without degrading generation quality. SpargeAttention2 includes (i) a hybrid masking rule that combines Top-k and Top-p for more robust masking at high sparsity, (ii) an efficient trainable sparse attention implementation, and (iii) a distillation-inspired fine-tuning objective to better preserve generation quality during fine-tuning using sparse attention. Experiments on video diffusion models show that SpargeAttention2 reaches 95% attention sparsity and a $16.2\times$ attention speedup while maintaining generation quality, consistently outperforming prior sparse attention methods.
Trainable Topology Supervision under Structurally Unreliable Pseudo Supervision
Jieyang Zhou ⋅ Ziliang Wang ⋅ Bin Hu ⋅ ying yang ⋅ Aoli He ⋅ Xianhong Wen ⋅ Kehua Guo
Topology supervision relies on meaningful structural targets, yet this assumption can fail under structurally unreliable pseudo supervision. In pseudo-label learning, the topology induced by pseudo masks may already contain distorted components, broken connections, or spurious holes, making direct supervision-side topology fitting unreliable and potentially propagating structural errors. We address this problem by reformulating topology supervision as trainable topology-side correction. Specifically, we propose a differentiable persistent-homology proxy loss that derives topology signals from a soft Euler-characteristic trajectory and threshold-wise structural variation, without computing or matching persistence barcodes. We further introduce a conservative topology-consistency injection mechanism that regulates when, how strongly, and in which direction these signals enter optimization. We validate the proposed reformulation in unsupervised camouflaged object detection, a challenging setting where pseudo masks often contain severe structural distortions. Experiments under controlled same-carrier comparisons show consistent improvements in both task-level detection performance and topology-level structural reliability. These results suggest that, under structurally unreliable pseudo supervision, topology supervision is more effective when derived as trainable proxy signals and introduced into optimization conservatively. Code is available at https://anonymous.4open.science/r/LPHP/.
Open World Object Detection (OWOD) requires a detector to recognize known categories, detect unknown objects, and incrementally learn new classes. Recent work such as OW-OVD adapts a pre-trained Vision-Language Object Detector (VLOD) to OWOD through fine-tuning. However, this approach demands costly backpropagation and often degrades the detector's own zero-shot performance on known categories, calling into question whether training is necessary. In this paper, we propose CAVE (Clustering in Attribute-Visual Embedding space), a training-free framework that adapts a VLOD to OWOD without any parameter updates. From a single forward pass, CAVE collects per-class visual and attribute statistics via von Mises-Fisher-based clustering, along with attribute co-occurrence patterns, without any gradient computation. At inference, CAVE fuses the original VLOD's text-based scores with MAS (Mixture-Aggregate Score) derived from the collected statistics for known-class detection, while unknown objects are identified by AOS (Attribute Objectness Score) that combines superclass responses with attribute co-occurrence patterns. Experiments on M-OWODB show that CAVE outperforms OW-OVD by +16.2 known-class mAP (Task 4) and +28.0 average unknown recall, while surpassing the zero-shot VLOD across all incremental tasks.
Training-Induced Escape from Token Clustering in a Mean-Field Formulation of Transformers
Noboru Isobe ⋅ Daisuke Inoue ⋅ Masaaki Imaizumi
Transformers perform inference by iteratively transforming token representations across layers. This layerwise computation has been studied empirically, and recent mean-field theories of Transformer dynamics explain how attention can drive token distributions toward clustering. However, existing mean-field analyses largely treat model parameters as prescribed, leaving open how training reshapes this clustering picture. We study this question in a noisy mean-field Transformer in which only a parameter-linear FFN is trained under $L^2$ regularization. We find and analyze a training-induced phase in the dynamics: after initially following attention-driven clustering, the token distribution can leave the clustered regime near the final layers. Our mathematical analysis is based on an entropy-regularized interaction energy that captures the clustering bias of attention. More broadly, our results point toward a training-aware mean-field theory of Transformer dynamics, in which training and inference dynamics are treated together.
Training on Documents About Monitoring Leads to CoT Obfuscation
Reilly Haskins ⋅ Bilal Chughtai ⋅ Joshua Engels
Chain-of-thought (CoT) monitoring is one of the most promising tools we have for detecting model misbehavior, but its effectiveness depends on models faithfully externalizing their reasoning. Motivated by this vulnerability, we study whether monitor-aware models are capable of obfuscating their reasoning to evade detection. We use synthetic document finetuning to expose eight models to realistic pre-training-style documents describing a CoT monitor and find that monitor-aware models consistently achieve higher rates of undetected misbehavior compared to unaware controls. This effect is weaker but still present on a harder agentic task. We also show that CoT controllability, a model's ability to reshape its own reasoning trace under an imposed constraint, is closely correlated with obfuscation success across the eight models studied ($r=0.800$, $p=0.017$). Monitor-aware models placed under equal reinforcement learning optimization pressure also learn to reward-hack without triggering a CoT monitor substantially faster than unaware controls. Together, these results suggest that knowledge of monitoring combined with high CoT controllability poses a risk to CoT-based monitoring.
Trajectory as the Teacher: Few-Step Discrete Flow Matching via Energy-Navigated Distillation
Amin Karimi Monsefi ⋅ Dominic Culver ⋅ Nikhil Bhendawade ⋅ Manuel R Ciosici ⋅ Yizhe Zhang ⋅ Irina Belousova
Discrete flow matching generates text by iteratively transforming noise tokens into coherent language, but may require hundreds of forward passes. Distillation uses the multi-step trajectory to train a student to reproduce the process in a few steps. When the student underperforms, the usual explanation is insufficient capacity. We argue the opposite: the trajectory is the bottleneck, not the student. Each training trajectory is built through a chain of blind stochastic jumps with no evaluation of sequence quality; a single bad decision at an early midpoint propagates through subsequent steps, yet the student must imitate the result. Trajectory-Shaped Discrete Flow Matching (TS-DFM) replaces these blind jumps with guided navigation: a lightweight energy compass evaluates candidate continuations at each midpoint, selecting the most coherent. All shaping is training-only; inference cost is unchanged. On 170M-parameter language modeling, the shaped student at 8 steps achieves 32% lower perplexity than the 1 024-step teacher while being 128×faster, with gains consistent across source distributions and three evaluators of increasing scale. TS-DFM achieves the best perplexity of any discrete-generation baseline we compare against, including methods trained on 6×more data or using 5×larger models.
Trajectory Forcing: Exploiting Diffusion Trajectories for Autoregressive Long Video Generation
Bohan Wang ⋅ Shuo Chen ⋅ Yunzhi Li ⋅ Zhaozheng CHEN ⋅ Junzhe Zhang ⋅ Hanwang Zhang
Autoregressive (AR) modeling is arguably the only way to extend a fixed-length video diffusion model to longer or even streaming videos. However, existing AR diffusion methods use only the previously predicted chunk as context for the next, discarding the intermediate predictions that constitute the chunk's generation history. In contrast, conventional AR models such as LLMs condition on the entire history of preceding tokens. To this end, we propose Trajectory Forcing (TraF), an AR diffusion method that exploits the denoising trajectories of previous video chunks as context. TraF operates under two training regimes: (1) with ground-truth videos, where pseudo-trajectories are constructed from Gaussian interpolations and re-predicted by the model, and (2) with a teacher model, where real trajectories are obtained from the student's own denoising rollout. The first regime naturally provides a strong initialization for the second. On the 5-second VBench benchmark, TraF achieves a Total score of 85.15, outperforming Self Forcing (+0.84) and Causal Forcing (+1.11) under the same training configuration, with Quality and Semantic improving simultaneously. On 30-second generation — $6\times$ the training horizon — TraF achieves the highest overall quality among all compared methods.
Transductive Generalization for GNNs via Optimal Transport
MoonJeong Park ⋅ Seungbeom Lee ⋅ Kyungmin Kim ⋅ Jaeseung Heo ⋅ Seunghyuk Cho ⋅ Shouheng Li ⋅ Sangdon Park ⋅ Dongwoo Kim
Analyzing the generalization of graph neural networks (GNNs) on node classification is challenging: message passing induces dependencies that break the i.i.d. assumption underlying standard inductive generalization theory. While distribution-free transductive theory resolves this dependency issue, existing classical bounds are either intractable or correlate poorly with empirical generalization. To fill this gap, we propose a transductive generalization bound based on optimal transport, connecting generalization to intra-class concentration and inter-class separation in the feature space. Empirically, our bound aligns strongly with the generalization gap. Building on this representation-based approach, we analyze how feature geometry evolves under repeated message passing. This explains the non-monotonic relationship between GNN depth and generalization, providing new insights into oversmoothing research. Practically, our bound serves as a geometry-aware regularizer that corrects the harmful behavior of message passing, yielding consistent performance improvements.
Transfer Learning of Linear Regression with Multiple Pretrained Models: Benefiting from More Pretrained Models via Overparameterization Debiasing
Daniel Boharon ⋅ Yehuda Dar
We study transfer learning for a linear regression task using several least-squares pretrained models that can be overparameterized. We formulate the target learning task as optimization that minimizes squared errors on the target dataset with penalty on the distance of the learned model from the pretrained models. We analytically formulate the test error of the learned target model and provide the corresponding empirical evaluations. Our results elucidate when using more pretrained models can improve transfer learning. Specifically, if the pretrained models are overparameterized, using sufficiently many of them is important for beneficial transfer learning. However, the learning may be compromised by overparameterization bias of pretrained models, i.e., the minimum L2-norm solution's restriction to a small subspace spanned by the training examples in the high-dimensional parameter space. We propose a simple debiasing via multiplicative correction factor that can reduce the overparameterization bias and leverage more pretrained models to learn a target predictor.
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
Jungtaek Kim ⋅ Thomas Zeng ⋅ Ziqian Lin ⋅ Minjae Lee ⋅ Chungpa Lee ⋅ Jy-yong Sohn ⋅ Hyung Koo ⋅ Kangwook Lee
Effective problem solving with Large Language Models (LLMs) can be enhanced when they are paired with external search algorithms. By viewing the space of diverse ideas and their follow-up possibilities as a tree structure, the search algorithm can navigate such a search space and guide the LLM toward better solutions more efficiently. While the search algorithm enables an effective balance between exploitation and exploration of a tree-structured space, the need for an external component can complicate the overall problem-solving process. We therefore pose the following question: Can LLMs or their underlying Transformer architectures approximate a search algorithm? To answer this question, we first introduce a simplified framework in which tree extensions and feedback signals are externally specified, allowing for controlled evaluation of search capabilities. We call this setting unknown tree search with bandit feedback. Within this setting, we show that Transformers are theoretically expressive enough to implement distinct search strategies and can be trained from scratch to approximate those strategies. Our Transformer models exhibit the possibility of generalizing to unseen conditions such as longer horizons or deeper trees. Furthermore, we demonstrate that continued task-focused training unlocks the complete capabilities of a pretrained LLM, by fine-tuning the LLM on search trajectories.
TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation
Hanyu Guo ⋅ JieDong Yang ⋅ Chao Chen ⋅ Longfei Xu ⋅ Kaikui Liu ⋅ Xiangxiang Chu
Public transit route planning traditionally depends on structured map infrastructure and complex routing engines, and no existing dataset supports training models to bypass this dependency. We present TransitLM, a large-scale dataset of over 13 million transit route planning records from four Chinese cities covering 120,845 stations and 13,666 lines, released as a continual pre-training corpus and benchmark data for three evaluation tasks with complementary metrics. Experiments show that an LLM trained on TransitLM produces structurally valid routes at high accuracy and implicitly grounds arbitrary GPS coordinates to appropriate stations without any explicit mapping. These results demonstrate that transit route planning can be learned entirely from data, enabling end-to-end, map-free route generation directly from origin-destination information. The dataset and benchmark are available at https://huggingface.co/datasets/GD-ML/TransitLM, with evaluation code at https://github.com/HotTricker/TransitLM.
Trapping Attacker in Dilemma: Examining Internal Correlations and External Influences of Trigger for Defending GNN Backdoors
Fan Yang ⋅ Binyan Xu ⋅ Di Tang ⋅ Kehuan Zhang
GNNs have become a standard tool for learning on relational data, yet they remain highly vulnerable to backdoor attacks. Prior defenses often depend on inspecting specific subgraph patterns or node features, and thus can be circumvented by adaptive attackers. We propose PRAETORIAN, a new defense that targets intrinsic requirements of effective GNN backdoors rather than surface-level cues. Our key observation is that flipping a victim node's prediction requires substantial influence on the victim: attackers must either inject many trigger nodes or rely on a small set of highly influential ones. Building on this observation, PRAETORIAN (i) analyzes internal correlations within potential trigger subgraphs to detect abnormally large injected structures, and (ii) quantifies external node influence to identify triggers with disproportionate impact. Across our evaluations, PRAETORIAN reduces the average attack success rate (ASR) to 0.55\% with only a 0.68\% drop in clean accuracy (CA), whereas state-of-the-art defenses still yield an average ASR of $\geq$20\% and a CA drop of $\geq$3\% under the same conditions. Moreover, PRAETORIAN remains effective against a range of adaptive attacks, forcing adversaries to either inject many trigger nodes to achieve high ASR ($\geq 80\%$), which incurs a $>$10\% CA drop, or preserve CA at the cost of limiting ASR to 18.1\%. Overall, PRAETORIAN constrains attackers to an unfavorable trade-off between efficacy and detectability.
TreemapMix: Dirichlet-Controlled Multi-Image Augmentation for Probability and Ordinal Supervision
Ejafa Bassam ⋅ Konstantin Garbers ⋅ Yingsheng Geng ⋅ Dalton Jens ⋅ Jiaxu Qian ⋅ Guyue Liu
Data augmentation improves visual recognition by exposing models to synthetic training examples that encourage generalization beyond the observed data. Mixed-sample methods extend this idea by combining multiple images. However, existing approaches either mix only pairs of images or rely on fixed small layouts, limiting control over source count, visible area proportions, and induced supervision. We propose TreemapMix, a distribution-controlled multi-image augmentation that samples source weights from a Dirichlet distribution and assigns them to area-proportional regions using a treemap partition. The resulting known region areas support two complementary objectives. TreemapMix with soft cross-entropy (SCE) treats areas as soft class-probability targets, while TreemapMix with a Plackett-Luce (PL) objective converts areas into ordinal supervision. On ImageNet-1K under a matched training protocol, TreemapMix-SCE achieves substantially lower calibration error than evaluated mixed-sample baselines, while TreemapMix-PL achieves the strongest Top-1 accuracy. Transfer experiments on object detection and instance segmentation suggest that TreemapMix-pretrained features remain useful beyond classification. These results show that controlled multi-image composition can expose the same region-area information as either probability-matching or rank-matching supervision.
TreePII: Efficient Computation of Higher-Order Probabilistic Interaction Indices in Tree Ensembles
Junho Choi ⋅ Jaesik Choi
While higher-order interaction indices offer deep insights into the synergies and redundancies within tree ensembles, their exact computation is hindered by combinatorial bottlenecks. Recent advancements efficiently compute symmetric interactions, but struggle to scale for asymmetric Probabilistic Interaction Indices (PIIs), which are crucial for advanced tasks such as incorporating feature hierarchy structures. To bridge this gap, we introduce TreePII, a novel algorithm that computes exact, any-order PIIs by integrating the partial derivatives of multilinear extension using interpolatory quadrature. TreePII effectively bypasses traditional computational hurdles, demonstrating both theoretical and empirical improvements over existing baselines and opening new avenues for scalable analysis of tree ensembles.
Tree Rotary Positional Encoding for Extreme Length Extrapolation from Scratch
qiu wu chen ⋅ Ziteng Huang ⋅ Shuhai Zhang ⋅ Zimo Liu ⋅ Linxiao Li ⋅ Ying Sun ⋅ Yuchen Li ⋅ Yifan Zhang ⋅ Yaofo Chen ⋅ Mingkui Tan
Large language models rely on positional representations for long context modeling, yet are typically trained on short sequences and evaluated at much longer lengths. Bridging this gap requires positional encodings remaining distinguishable at large distances while generalizing beyond the training length. In this paper, we identify a structural limitation of existing rotary-based positional encodings: all positional dimensions are controlled by a single global index, which couples training-time positional variation with long-range extrapolation. To address this, we propose Tree Rotary Positional Encoding (TPE), a structured multi-scale positional encoding designed for training from scratch that represents absolute positions through com positions of different components, avoiding single-index extrapolation. We further devise an Offset-based Positional Training (OPT) strategy to improve learning interactions across TPE components, enabling better generalization to longer con texts. We theoretically analyze that its multi-scale decomposition enables effective combinatorial position coverage. Experimental results show strong gains under extreme length extrapolation: with a 1.2B model trained from scratch on length 512, our TPE achieves a 64× extrapolation to 32K tokens, reducing perplexity from 4532.21 (RoPE) to 149.96, with comparable inference-time efficiency.
TRIAD: Benchmarking Omni-Modal Ambiguity in Multimodal Large Language Models
Zhaolu Kang ⋅ Yidi Wang ⋅ Yile Li ⋅ Siqi Zeng ⋅ Yingjie He ⋅ richeng xuan ⋅ Zhichao Hu
Ambiguity is a central feature of natural communication: a sentence, image, or sound may remain underspecified until evidence from the other channels is considered. Yet current multimodal benchmarks mostly test perception, recognition, or reasoning over already determinate inputs, leaving unclear whether omni-modal large language models can resolve ambiguity distributed across text, vision, and audio. We introduce TRIAD, a benchmark for omni-modal ambiguity resolution. Each item is a text--image--audio triplet paired with a question and a set of answer options, and is constructed so that the full triplet determines a unique answer while removing any one modality makes the set of consistent options non-unique. TRIAD contains $327$ hand-written queries in $86$ scene groups and covers $18$ ambiguity categories across the three modalities. Evaluating $14$ omni-modal systems reveals that strong item-level performance does not translate into robust scene-level disambiguation: the best model still trails humans by over $20$ points in group accuracy. Removing answer options widens the gap further, driving model group accuracy to the single digits while humans remain robust. Leave-one-out and transcript/description counterfactuals further show that current systems often underuse audio cues when text and image are present. TRIAD reframes omni-modal evaluation around disambiguation rather than recognition, and provides a taxonomy, benchmark, and diagnostic protocol for measuring this capability.
Triggering Generalist Reasoning via Predictive Uncertainty for Dual-System VLA
Hyemin Yang ⋅ Wooseong Jeong ⋅ Giwon Lee ⋅ Kuk-Jin Yoon
Dual-system Vision-Language-Action (VLA) models improve real-time robotic control by pairing a slow, reasoning-capable generalist with a fast specialist action expert. However, existing methods invoke the generalist at a fixed frequency, ignoring the fact that decision-making complexity varies throughout a rollout. This static strategy wastes computation in easy phases and can delay intervention when unexpected scene changes require renewed high-level reasoning. We propose TUD (Triggering generalist reasoning via predictive Uncertainty for Dual-system VLA), an adaptive inference framework that selectively skips unnecessary generalist calls. TUD measures the cross-step dispersion of action re-predictions at the upcoming chunk slot, computed from the specialist's existing forwards under the cached generalist context, as a predictive uncertainty signal. This signal captures how much the future action plan shifts as new observations arrive, enabling TUD to invoke the generalist only when the cached plan becomes unreliable, and separates success from failure more reliably than prior uncertainty signals. The signal needs no manually labelled phase boundaries or auxiliary uncertainty model, and is computed from forwards the architecture already runs. On VLA-Arena, TUD finds a more favorable cost--success trade-off than non-adaptive baselines, tracing an entire operating curve as a single threshold is varied, and substantially reduces VLM calls at matched success rate. Our results suggest that predictive uncertainty provides a practical criterion for adaptive reasoning in efficient VLA control.
TriPrompt: Progressive Local Prompting for Few-shot Out-of-Distribution Detection
Ziyou Xiang ⋅ Chaowei Fang ⋅ Zhihong Wu ⋅ Jiliang Li ⋅ Di Xu ⋅ De Cheng
Few-shot out-of-distribution (OOD) detection is essential for deploying recognition models in open-world environments, where a reliable model should correctly recognize in-distribution (ID) samples while rejecting samples from unseen categories. Recent CLIP-based methods have shown promising few-shot OOD detection ability by exploiting transferable vision--language representations. However, many methods rely mainly on global image features or emphasize only the most class-discriminative local cues, making them less effective for OOD detection where ID and OOD samples share visually similar local patterns. Although local prompting introduces patch-level reasoning, existing methods often treat local evidence as a binary distinction between ID-related and non-ID cues, leaving the ambiguous transition region between reliable ID evidence and clear non-ID evidence insufficiently modeled. To address this issue, we propose ThriPrompt, a progressive local prompting framework for few-shot OOD detection. TriPrompt decomposes patch-level evidence into three complementary components within a unified local feature space: Core-ID cues that provide reliable class-defining evidence, Far-ID cues that capture ambiguous transition-region evidence, and OOD cues learned from weak non-ID proxies mined directly from patch features. This patch-native design enables suppression-aware local reasoning without requiring additional OOD data or repeated multi-crop feature extraction. At inference time, \modelname{} combines global CLIP confidence with Core-ID, Far-ID, and OOD local responses to reduce overconfident ID assignment on hard OOD samples. Experiments on ImageNet-1K OOD benchmarks show that \modelname{} consistently improves few-shot OOD detection performance and remains compatible with mainstream CLIP-based scoring methods. Our code is provided in the supplementary material.
We derive a recursion for the expected squared Frobenius norm of the zonotope generator matrices that govern the decision boundary of a random Gaussian deep ReLU binary classifier. For a network of depth $L$ with weight variances $(\sigma_{l=1..L}^2)$, the boundary complexity is $\Gamma_L = 8(1+1/\pi)\prod_{l=1}^{L} \sigma_l^2/2$. He initialization $\sigma_w^2 = 2$ is the unique uniform weight variance under which $\Gamma_L$ is preserved with depth, matching the Poole--Schoenholz mean-field critical value on a different observable. $\Gamma_L$ is identified with an observable per-unit-length boundary-crossing density via Rice's formula. The relation is supported by Monte Carlo experiments.
When transformers train on contradictory data, the same problem with both correct and incorrect solutions, which answer do they prefer? We hypothesize that next-token prediction, as a compression process, favors whichever answer cluster has lower description length; truth benefits only when errors lack internal structure. We test this by training transformers (3.5M-1B parameters) from scratch on controlled corpora, systematically varying the structure of errors. We find that (a) when errors are random, models develop a correctness preference scaling from 65% to 85% with model size; (b) when errors follow a single coherent alternative rule, this preference vanishes (~45-51%); (c) two competing wrong rules suffice to restore it (47% to 78%). The pattern reproduces on Wikipedia paragraphs with entity substitution (71% vs 46%) and at 1B scale on a mixed natural-text corpus (77% vs 47%). These results are consistent with the hypothesis that, in controlled contradictory corpora, model preference tracks the relative compressibility of competing answer systems rather than truth per se.
TSB-SEG: A Systematic Time-Series Segmentation Benchmark
Félix Chavelli ⋅ Arik Ermshaus ⋅ Patrick Schäfer ⋅ Fan Yang ⋅ John Paparrizos ⋅ Paul Boniol
Time-series segmentation, studied as either \emph{Change Point Detection} (CPD) or \emph{State Detection} (SD), underpins a broad range of monitoring and diagnostic applications. However, the field suffers from a persistent divide between the statistical community (focusing on CPD) and the machine learning and data mining communities (focusing on SD). Consequently, the literature remains fragmented: evaluations typically cover narrow methodological families, datasets are often small or homogeneous, and CPD and SD are rarely evaluated together. We address these gaps with \texttt{TSB-SEG}, a unified benchmark comprising $592$ univariate and multivariate time series across $8$ heterogeneous domains and surveying $27$ algorithms spanning five decades of progress. To ensure a fair comparison, we introduce a two-pool design: a scalable \emph{main pool} of sub-quadratic detectors used for the aggregate ranking, and an \emph{extended-scope pool} for computationally intensive specialized methods. All methods are unified under tsseg, an open-source library to ensure full reproducibility. Our systematic evaluation yields three insights: (a) no single detector dominates across domains, and per-dataset hyperparameter tuning remains decisive for performance, (b) CPD accuracy is a strong proxy for SD accuracy, and (c) data characteristics should guide method selection. We provide actionable guidance based on data properties, error tolerances, and tuning budgets, identify impactful hyperparameters, and characterize algorithm robustness to these choices.
TTB: Test-time MLP Baking for Efficient Rendering of Decoder-only View Synthesis Models
Chin-Yang Lin ⋅ Ryo Hachiuma ⋅ Jaesung Choe ⋅ Min-Hung Chen ⋅ Frank Wang ⋅ Yen-Yu Lin ⋅ Wei-Chen Chiu ⋅ Yu-Lun Liu ⋅ Cheng Sun
Decoder-only Large View Synthesis Models (LVSMs) with KV-cache have recently achieved state-of-the-art quality by regressing novel views from neural networks without reconstructing 3D geometry, while Test-Time Training (TTT) layers can be introduced to save computation, but sacrifice quality. We propose Test-Time Baking (TTB), a strategy that learns to bake global scene information from cross-view attention into a lightweight MLP, enabling fast and high-fidelity novel-view decoding. The first key is to repurpose the fast-weight MLP of TTT to learn the cross-view attention mapping explicitly from its input-output pairs, rather than replacing attention entirely, achieving the same rendering efficiency as TTT while delivering substantially higher quality. The second key addresses a modality misalignment between context prefilling and novel-view rendering: we disentangle context tokens into camera-only context rendering tokens to serve as TTB input, and context ground-truth tokens with image content to serve as the TTB mapping target, ensuring the baked MLPs are relieved of the modality generalization challenge between context and novel-view processing. Our method achieves up to 23x rendering speedup with quality on par with the state-of-the-art across several benchmarks.
Two is better than one: A Collapse-free Multi-Reward RLIF Training Framework
Shourov Joarder ⋅ Diganta Sikdar ⋅ Ahsan H Akash ⋅ Binod Bhattarai ⋅ Prashnna K Gyawali
Reinforcement learning with verifiable rewards (RLVR) has substantially improved the reasoning ability of LLMs, but often depends on external supervision from human annotations or gold-standard solutions. Reinforcement learning from internal feedback (RLIF) has recently emerged as a scalable unsupervised alternative, using signals extracted from the model itself. However, existing RLIF methods typically rely on a single internal reward, which can lead to reward hacking, entropy collapse, and degraded reasoning structure. We propose a multi-reward RLIF framework that decomposes the training signal into two complementary components: an answer-level reward based on cluster voting and a completion-level reward based on token-wise self-certainty. To combine these signals robustly, we apply GDPO-based normalization to reduce reward-scale imbalance. We further introduce KL-Cov regularization, which targets low-entropy token distributions responsible for disproportionate entropy reduction, preserving exploration and preventing late-stage collapse. Across mathematical reasoning and code-generation benchmarks, our method improves stability and robustness over prior unsupervised RL approaches, while achieving performance close to supervised RLVR methods. These results show that complementary internal rewards, combined with targeted regularization, can support stable long-horizon reasoning without relying on external ground-truth supervision.
Tyche: One Step Flow for Efficient Probabilistic Weather Forecasting
Fan Xu ⋅ Yuan Gao ⋅ Kun Wang ⋅ Rui Su ⋅ Fenghua Ling ⋅ Hao Wu ⋅ Wanli Ouyang
Probabilistic weather forecasting requires not only accurate trajectories, but calibrated distributions over plausible atmospheric futures. Recent data-driven systems have achieved remarkable deterministic skill, and diffusion-based ensemble forecasters have substantially improved sample realism and uncertainty quantification. However, their inference cost scales with forecast horizon, ensemble size, and the number of denoising steps required for each transition, making large operational ensembles expensive. To address this, we present Tyche, a one-step conditional flow model for efficient probabilistic weather forecasting. Tyche models the conditional forecast distribution with a destination-aware average-velocity flow that maps Gaussian noise directly to future weather states in a single function evaluation (1-NFE). To make this one-step transport learnable in high-dimensional geophysical fields, we derive a JVP-regularized rectification objective that enforces temporal self-consistency across source and destination flow timesteps without explicitly forming Jacobians. The transport field is parameterized by an isotropic Swin-style transformer that preserves fine-scale spatial structure while remaining scalable on global grids. To improve ensemble reliability under autoregressive forecasting, we further introduce a rollout-based finetuning stage with curriculum CRPS calibration supervision. Experiments on ERA5 at 1.5$^\circ$ and 6-hour resolution show that our Tyche, using merely a single NFE, matches or exceeds the forecast skill and calibration of state-of-the-art multi-step generative baselines and the operational ECMWF IFS ensemble. Codes are available.
Uncertainty Quantification for Large Language Diffusion Models
Artem Vazhentsev ⋅ Vladislav Smirnov ⋅ David Li ⋅ Maxim Panov ⋅ Timothy Baldwin ⋅ Artem Shelmanov
Large Language Diffusion Models (LLDMs) are emerging as an alternative to autoregressive models, offering faster inference through higher parallelism. Similar to autoregressive LLMs, they remain prone to hallucinations, making reliable uncertainty quantification (UQ) crucial for safe deployment. However, existing UQ methods are fundamentally misaligned with this new paradigm: they assume autoregressive factorization or use expensive multiple sampling, negating the efficiency of LLDMs. In this work, we present the first systematic study of UQ for LLDMs and propose lightweight, zero-shot uncertainty signals derived from the iterative denoising process, leveraging intermediate generations, token remasking dynamics, and denoising complexity. We further adapt a state-of-the-art UQ method to LLDMs by combining masked diffusion likelihoods with trajectory-based semantic dissimilarity. We provide theoretical grounding, proving that expected trajectory dissimilarity bounds the masked diffusion training objective. Comprehensive experiments across three tasks, eight datasets, and two models show that our method achieves a great cost-performance trade-off: it approaches the strongest sampling-based baselines while incurring up to one hundred times lower computational overhead. Our work demonstrates that LLDMs can deliver both fast inference and reliable hallucination detection simultaneously.
Uncertainty Quantification for Multimodal Large Language Models with Incoherence-adjusted Semantic Volume
Gregory Kang Ruey Lau ⋅ Hieu Dao ⋅ Nicole Hui Lin Kan ⋅ Bryan Kian Hsiang Low
Despite their capabilities, Multimodal Large Language Models (MLLMs) may produce plausible but erroneous outputs, hindering reliable deployment. Accurate uncertainty metrics could enable escalation of unreliable queries to human experts or larger models for improved performance. However, existing uncertainty metrics have practical constraints, such as being designed only for specific modalities, reliant on external tools, or computationally expensive. We introduce UMPIRE, a training-free uncertainty quantification framework for MLLMs that works efficiently across various input and output modalities without external tools, relying only on the models' own internal modality features. UMPIRE computes the incoherence-adjusted semantic volume of sampled MLLM responses for a given task instance, effectively capturing both the global semantic diversity of samples and the local incoherence of responses based on internal model confidence. We provide theoretical analysis motivating UMPIRE's design. Extensive experiments show that UMPIRE consistently outperforms baseline metrics in error detection and uncertainty calibration across image, audio, and video-text benchmarks, including adversarial and out-of-distribution settings. We also demonstrate UMPIRE’s generalization to non-text output tasks, including image and audio generation.
Advances in deep reinforcement learning have enabled the development of policies capable of solving complex, high-dimensional problems, allowing AI agents to strategize and reason purely through interaction without supervision. Despite these successes, our knowledge on the functional properties and the underlying structure of the deep neural policy manifold remains limited. In this paper, we uncover a fundamental functional property in reinforcement learning: the intrinsic alignment between the advantage function and the gradient of the loss targeting directions of incoherence. We provide a rigorous theoretical foundation for this relationship, demonstrating that this intrinsic alignment characterizes how policies learn and internalize the underlying value function and how this process governs policy decisions. By leveraging this fundamental functional property, we propose a novel algorithm that can diagnose and identify deep neural policy decision volatilities. We conduct extensive empirical analysis in high-dimensional MDPs. From algorithmic and architectural changes to natural distributional shifts and worst-case perturbations, our proposed method can identify and audit the differences by leveraging the underlying structure of the deep neural policy manifold and the intrinsic correlation. Our analysis reveals foundational properties of policies trained in high-dimensional MDPs, and our paper provides a principled step toward constructing scalable, stable, and generalizable deep reinforcement learning agents.
Understanding and Mitigating Structural Forgetting in Fine-Tuned Time Series Foundation Models
ZEYU SHI ⋅ Yirong Xue ⋅ Xin Xue ⋅ Ziming Wang ⋅ Jintao Wu ⋅ Shiqi Gao ⋅ Haoyi Zhou
While fine-tuning is the standard recipe for adapting Time Series Foundation Models (TSFMs) to downstream tasks, we reveal that it triggers a critical yet overlooked problem: structural forgetting. Distinct from widely studied point-wise error degradation, structural forgetting represents a fundamentally more destructive phenomenon: the erasure of universal structural patterns acquired during pre-training. Unlike point-wise errors, this structural collapse renders models incapable of supporting basic structure-dependent decisions (e.g., trend-based trading), a critical vulnerability that existing forgetting mitigation strategies fail to address. To bridge this gap, we first introduce a diagnostic framework equipped with a novel Temporal-Frequency Attribution analysis, mechanistically revealing that fine-tuning unnecessarily overwrites a sparse set of pattern-critical parameters. We then propose NeST, a post-hoc framework that precisely restores these components while compensating surrounding weights to preserve task adaptation. Extensive experiments across 3 TSFMs and 9 datasets demonstrate that NeST recovers an average of 85\% of structural awareness without sacrificing task performance, offering a new paradigm for structurally-preserved TSFM adaptation.
Understanding Generalization Requires Universal Induction
Aram Ebtekar ⋅ Marcus Hutter ⋅ Danica J. Sutherland
Classical statistical theory is insufficient to explain the successes of general-purpose AI models, because it depends on handcrafted inductive biases that it cannot justify. No Free Lunch theorems force any learner that beats chance on some environments to underperform on others, and meta-learning which environments are more likely only pushes the problem up a level. The inductive bias must therefore be grounded in something other than data. By deriving its bias from Turing universality, Solomonoff induction (SI) competes with all computable learners. However, its regret bounds include large ``constants,'' such as the size of a learner's (losslessly compressed) full codebase. We relativize SI to an information vantage point, which includes all pre-existing code and data. This reframes the inductive bias: instead of favoring some absolute notion of simplicity, we favor accessibility with respect to our vantage point. The relativized SI is sample-optimal on finite data: any learner that outperforms it necessarily contains inaccessible information about the data. While SI is incomputable and hence not a practical algorithm, it provides a formal optimum for inference in the limit of infinite compute. We argue that algorithmic information theory, which underlies SI, is necessary to explain the generalization behavior of modern (and future) AI systems.
Uni-Cheb: A Basis-Agnostic Learnable Chebyshev Filter for Multimodal Spectral Modulation
Shuoqiu Duan ⋅ Jiasen Gao ⋅ Xiaoliang Chen
While spatiotemporal modeling has emerged as an effective approach in multimodal learning, it struggles to resolve the ``spectral dilemma'' where distinct tasks demand contradictory frequency views. For example, deception detection requires global macro-patterns, while intent recognition depends on local micro-transients. To address this, we propose \textbf{Uni-Cheb}, a unified spectral operator centered on the Learnable Chebyshev Filter (LCF). Designed as a basis-agnostic, plug-and-play component, LCF maintains a consistent mathematical form to adaptively modulate frequency components regardless of the underlying spectral transform. By leveraging the minimax property of Chebyshev approximation, LCF performs elastic spectral resampling to focus on task-relevant bands. Guided by task-specific physical priors, the LCF seamlessly integrates with the Discrete Fourier Transform (DFT) to amplify global physiological rhythms or the Discrete Wavelet Transform (DWT) to capture localized semantic shifts, all while maintaining a negligible computational footprint. Furthermore, it serves as an effective spectral preconditioner to mitigate distribution shifts in cross-domain transfer learning. Extensive experiments across multiple benchmarks demonstrate that Uni-Cheb acts as a universal enhancer, consistently improving state-of-the-art baselines. By rendering frequency modulation both task-adaptive and highly efficient, Uni-Cheb establishes a robust, operator-level solution to the spectral dilemma.
UniDBO: A Unified Dual-Branch One-Step Denoising Framework for Autonomous Driving Scenario Generation
Da Zhao ⋅ Yuhang Chen ⋅ Jie Sun ⋅ Jialin Fan ⋅ Jian Sun
Scenario generation is essential for training, testing, and safety validation of autonomous vehicles (AVs), especially when real-world data coverage is limited and long-tail safety-critical events are rare. Existing methods face a three-way trade-off among open-loop prediction accuracy, closed-loop simulation robustness, and inference efficiency. Autoregressive methods are typically efficient and strong in open-loop fitting, but they are prone to error accumulation and policy-induced state-distribution shift in long-horizon closed-loop rollouts. Diffusion-based methods provide strong multimodal behavior modeling and competitive generation quality, but iterative denoising substantially increases inference latency and limits simulation throughput. To address this challenge, we propose UniDBO, a Unified Dual-Branch One-step denoising framework for autonomous driving scenario generation. UniDBO uses a shared scene encoder and two jointly optimized complementary branches with distinct roles: a continuous-scale branch (CS) for open-loop trajectory fitting and a discrete high-noise branch (DHN) for robust closed-loop simulation. Through joint training, UniDBO coordinates open-loop fitting and closed-loop robustness while preserving one-step inference efficiency. We evaluate UniDBO on multiple benchmarks, including in-domain open-loop and closed-loop evaluation, zero-shot cross-domain generalization, and system-level closed-loop AV testing in unseen scenarios. Results show that UniDBO improves the open-loop/closed-loop balance while maintaining high inference efficiency, and achieves competitive zero-shot generalization on the evaluated benchmarks.
UniFlowDock: Flexible Docking with Complete Equivariant Velocity Fields
Kangxin Chen ⋅ Jieyu Zhao ⋅ Xulun Ye ⋅ Kun Zhou ⋅ Jinli Hu ⋅ Yuanyuan Deng ⋅ Min Xie
Flexible molecular docking is crucial for drug discovery; however, existing flow-matching generative models often produce physically invalid conformations (e.g., steric clashes, distorted bond geometry). Our analysis suggests that a primary contributor to this challenge lies at the modeling level: when the flow is parameterized by standard equivariant networks, an intrinsic approximation floor can emerge, potentially hindering the learned velocity field from reaching its optimal convergence bound, particularly for highly flexible molecules. We present UniFlowDock, which addresses this by introducing strictly Complete Equivariant Velocity Fields, thereby eliminating the dimensional incompleteness of standard architectures. This theoretical completeness is further bolstered by Fragment-Adaptive Optimal Coupling, a method that structurally linearizes transport trajectories and decomposes the generation process into physics-aware subproblems. Experiments on PDBBind and PoseBusters demonstrate UniFlowDock achieves state-of-the-art accuracy and physical validity, showing consistent gains in highly flexible docking scenarios where target distributions exhibit higher transport complexity.
Uniform Spectral Growth under Factor-wise Muon Orthogonalization in Matrix Factorization and LoRA
Changmin Kang ⋅ Jihun Yun ⋅ Baekrok Shin ⋅ Yeseul Cho ⋅ Chulhee Yun
Spectral gradient descent (SpecGD) orthogonalizes matrix parameter updates and has inspired practical optimizers such as Muon. They often perform well in large language model training, but their dynamics remain poorly understood, especially in factorized parameterizations where the product matrix does not receive orthogonalized updates. We study such dynamics through matrix factorization (MF), where the orthogonalization is applied separately to the factor updates. We analyze spectral gradient flow (SpecGF)—a continuous-time analog of SpecGD—in the low-rank MF setting and prove "equal-rate" dynamics: all singular values grow at equal rates up to small deviations. Consequently, smaller singular values attain their target values earlier than larger ones, contrasting with the largest-first stepwise learning observed in standard gradient flow. Moreover, we prove that SpecGF in our setting converges to global minima from almost all initializations, provided the factor norms remain bounded; with $\ell_2$ regularization, we obtain global convergence. Empirically, we observe that LoRA fine-tuning with orthogonalization-based optimizers including Muon exhibit near-uniform growth in the product of LoRA adapters, consistent with the mechanism predicted by our MF analysis.
UniFunc3D: Unified Active Spatial-Temporal Grounding for 3D Affordance Segmentation
Jiaying Lin ⋅ Dan Xu
Affordance segmentation in 3D scenes requires an agent to ground implicit natural-language instructions into precise masks of fine-grained interactive elements. Existing training-free methods typically rely on fragmented pipelines, which introduce visual blindness during task parsing and limit accuracy through single-scale spatial and temporal processing. We present UniFunc3D, a unified and training-free framework that treats the multimodal large language model as an active observer. By consolidating semantic, temporal, and spatial reasoning into a single forward pass, UniFunc3D performs joint reasoning to ground task decomposition in direct visual evidence. Our approach introduces active spatial-temporal grounding with a coarse-to-fine strategy. This allows the model to select correct video frames adaptively and focus on high-detail interactive parts while preserving the global context necessary for disambiguation. On SceneFun3D, our UniFunc3D achieves state-of-the-art performance, surpassing prior training-free methods by a large margin with a relative 59.9\% mIoU improvement, and even outperforming training-based methods without any task-specific training. Code will be released.
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis
Shuang Chen ⋅ Quanxin Shou ⋅ Hangting Chen ⋅ Yucheng Zhou ⋅ Kaituo Feng ⋅ Wenbo Hu ⋅ yifan zhang ⋅ Yunlong Lin ⋅ Wenxuan Huang ⋅ Mingyang Song ⋅ Dasen Dai ⋅ Bolin Jiang ⋅ Manyuan Zhang ⋅ Yu Cheng ⋅ Nanyun Peng
Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, they still rely primarily on frozen parametric knowledge, which makes them struggle with real-world image generation involving long-tail and knowledge-intensive concepts. Inspired by the broad success of agents on real-world tasks, we explore agentic modeling to address this limitation. Specifically, we present Unify-Agent, a unified multimodal agent for world-grounded image synthesis, which reframes image generation as an agentic pipeline consisting of prompt understanding, multimodal evidence searching, grounded recaptioning, and final synthesis. To train our model, we construct a tailored multimodal data pipeline and curate 143K high-quality agent trajectories for world-grounded image synthesis, enabling effective supervision over the full agentic generation process. We further introduce FactIP, a benchmark covering 12 categories of culturally significant and long-tail factual concepts that explicitly requires external knowledge grounding. Extensive experiments show that our proposed Unify-Agent substantially improves over its base unified model across diverse benchmarks and real world generation tasks, while approaching the world knowledge capabilities of the strongest closed-source models. As an early exploration of agent-based modeling for world-grounded image synthesis, our work highlights the value of tightly coupling reasoning, searching, and generation for reliable open-world agentic image synthesis.
UniRank: Unified List-wise Reranking via Confidence-Ordered Denoising
Pengyue Jia ⋅ HailanYang ⋅ Shuchang Liu ⋅ Xiaobei Wang ⋅ Wanyu Wang ⋅ Xiang Li ⋅ Yongqi Liu ⋅ Kaiqiao Zhan ⋅ Kun Gai ⋅ Xiangyu Zhao
List-wise reranking arranges a request-specific pool of candidate items into an ordered slate that maximizes user satisfaction. Existing generative rerankers fall into two paradigms: Autoregressive (AR) rerankers construct the slate left to right and capture inter-item dependencies in the exposure list, but they suffer from error propagation because early mistakes affect subsequent slots. Non-autoregressive (NAR) rerankers predict all slots in parallel and avoid error propagation, but they weaken inter-item interaction modeling under a slot independence assumption. This raises a central question: is there a unified architecture that combines the strengths of both paradigms and delivers stronger reranking performance? We answer this question with UniRank, a unified list-wise reranking framework whose inference time variants recover AR and NAR rerankers as special cases. UniRank integrates bidirectional slate modeling into an iterative denoising process and fills the most confident slot at each step. To instantiate this framework for reranking, we introduce the Task Grounded Diffusion Interface (TGD), which performs denoising at the item level and restricts prediction to the request-specific candidate pool. TGD aggregates each item's semantic tokens into a single item embedding and scores each slot directly against the candidate pool. Experiments on Amazon Books, MovieLens-1M, and an industrial short video dataset show that UniRank consistently outperforms state-of-the-art baselines. Our code is available online.
UniT: Toward a Unified Physical Language for Human-to-Humanoid Policy Learning and World Modeling
Boyu Chen ⋅ Yi Chen ⋅ Lu QIU ⋅ Jerry Bai ⋅ Yuying Ge ⋅ Yixiao Ge
Scaling humanoid foundation models is bottlenecked by the scarcity of robotic data. While massive egocentric human data offers a scalable alternative, bridging the cross-embodiment chasm remains a fundamental challenge due to kinematic mismatches. We introduce UniT (Unified Latent Action Tokenizer via Visual Anchoring), a framework that learns a unified physical language for human-to-humanoid manipulation transfer. Grounded in the philosophy that heterogeneous kinematics share consistent visual consequences, UniT employs a tri-branch cross-reconstruction mechanism: actions predict vision to anchor kinematics to physical outcomes, while vision reconstructs actions to filter out irrelevant visual confounders. Concurrently, a fusion branch integrates these purified modalities into a shared discrete latent space of cross-embodiment physical intents. We validate UniT across two paradigms: (1) Policy Learning (VLA-UniT): By predicting these unified tokens, VLA-UniT achieves state-of-the-art performance with high data efficiency on the RoboCasa GR1 benchmark. Leveraging diverse human data further improves out-of-distribution (OOD) generalization in simulation and real-world deployment, and enables zero-shot task transfer in the real world. (2) World Modeling (WM-UniT): By aligning cross-embodiment dynamics via unified tokens as conditions, it supports direct human-to-humanoid action-conditioned generation, translating human knowledge into enhanced action controllability for humanoid video generation. Ultimately, by inducing a more aligned cross-embodiment representation (empirically supported by t-SNE visualizations revealing improved alignment of human and humanoid features), UniT offers a scalable path to distill human priors into humanoid manipulation capabilities.
Unleashing the Power of Intrinsic-Entropy-Driven Exploration for Off-Policy Generative RL
Feihong Zhang ⋅ Guojian Zhan ⋅ yinuo Wang ⋅ Letian Tao ⋅ Likun Wang ⋅ Shiqi Liu ⋅ Yao Lyu ⋅ Shengbo Eben Li
Generative models, such as diffusion and flow-based models, have emerged as powerful policy approximators in deep reinforcement learning (RL) due to their ability to represent multi-modal action distributions. However, their potential for maximum entropy RL remains largely untapped, as most existing methods treat generative processes as black-box samplers or rely on heuristic noise. In this paper, we propose GENIEO, a novel off-policy generative RL framework that leverages the exact intrinsic entropy of an augmented dummy-action policy to drive principled exploration. By designing a structurally invertible affine dual-variable flow, we bypass the numerical approximations of standard ODE solvers. This innovation allows us to rigorously apply the change-of-variables formula to derive a tractable and exact formulation for the augmented dummy-action policy entropy, enabling its direct integration into the off-policy learning objective. Experimental evaluations across 14 challenging tasks from the DMControl and HumanoidBench suites demonstrate that GENIEO achieves state-of-the-art (SOTA) performance. By effectively discovering optimal strategies in complex, high-dimensional environments, our approach bridges the gap between expressive generative modeling and maximum entropy RL.
Unsupervised Domain Adaptation for Semantic Segmentation Based on Instance Spatial Geometry
Binglin Hao ⋅ Huajun Liu
Unsupervised domain adaptation(UDA) for semantic segmentation transfers knowledge from synthetic to real domains, where geometric cues such as depth are commonly exploited to reduce the domain gap. However, existing depth-aware methods fail to explicitly model the domain-invariant spatial structures among semantic instances, and current self-training schemes weight pseudo-labels solely by prediction confidence, ignoring geometric consistency, which limits their reliability under domain shift. To address these issues, we propose an instance-wise geometric modeling framework that captures inter-instance spatial relations, including directional concentration and dominant direction, beyond conventional depth representations. These geometry-aware features are further integrated into pseudo-label weighting to enforce geometric-semantic consistency during self-training, leading to more stable and accurate pseudo supervision. Experiments on standard benchmarks show that our method significantly outperforms state-of-the-art UDA approaches.
Untrusted Content Masking for Web Agents with Security Guarantees
Kristina Nikolić ⋅ Egor Zverev ⋅ Javier Rando ⋅ Matthew Jagielski ⋅ Edoardo Debenedetti ⋅ Florian Tramer
Defenses that provide security guarantees against prompt injection attacks require strict isolation between an agent's task planning and data processing capabilities. This prevents third-party content from overwriting trusted instructions. In text-based environments such as tool-use APIs, agents can plan from interface definitions without ever processing untrusted data. Web agents, however, face a fundamental challenge: they must observe the rendered page to perceive their environment, but that page already contains untrusted third-party content. In this paper, we present Untrusted Content Masking, a simple, effective approach that enables web agents to observe their environment and plan without directly processing untrusted content. We leverage a key structural insight: a webpage's Document Object Model (DOM) structure alone suffices to identify untrusted regions. Our framework exploits this by redacting such regions before they reach the agent, and restricting interaction to a sandboxed interface with strict privilege separation.
Unveiling Fine-Grained Visual Traces: Evaluating MultiModal Interleaved Reasoning Chains in Multimodal STEM Tasks
Jing Jin ⋅ Hao Liu ⋅ Yan Bai ⋅ Yihang Lou ⋅ Zhenke Wang ⋅ Tianrun Yuan ⋅ Yongkang Zhu ⋅ Juntong Chen ⋅ Fanhu Zeng ⋅ Xuanyu Zhu ⋅ Tao Feng ⋅ Yige Xu
Multimodal large language models (MLLMs) have shown promising reasoning abilities, yet evaluating their performance in specialized domains remains challenging. STEM reasoning is a particularly valuable testbed because it provides highly verifiable feedback, but existing benchmarks often permit unimodal shortcuts due to modality redundancy and focus mainly on final-answer accuracy, overlooking the reasoning process itself. To address this challenge, we introduce STEPSTEM: a graduate-level benchmark of 283 problems across mathematics, physics, chemistry, biology, and engineering for fine-grained evaluation of cross-modal reasoning in MLLMs. STEPSTEM is constructed through a rigorous curation pipeline that enforces strict complementarity between textual and visual inputs. We further propose a general step-level evaluation framework for both text-only chain-of-thought and interleaved image-text reasoning, using dynamic programming to align predicted reasoning steps with multiple reference solutions. Experiments across a wide range of models show that current MLLMs still rely heavily on textual reasoning, with even Gemini 3.1 Pro and Claude Opus 4.6 achieving only 38.29\% accuracy. These results highlight substantial headroom for genuine cross-modal STEM reasoning and position STEPSTEM as a benchmark for fine-grained evaluation of multimodal reasoning.
UxSID: Semantic-Aware User Interests Modeling for Ultra-Long Sequence
Hongwei Zhang ⋅ qiqiang zhong ⋅ Jiangxia Cao ⋅ Junfeng Shu ⋅ Yiyang Lv ⋅ Huanjie Wang ⋅ Liwei Guan ⋅ Jing Yao ⋅ Yiyu Wang ⋅ Liu Zhaojie ⋅ Han Li
Modeling ultra-long user behavior sequences is an essential task for capturing evolving preferences in modern recommender systems, and this research direction has contributed solid gains in the past several years. However, recommendation systems always serve enormous traffic, and extending user sequences sharply increases computational cost, which creates a difficult trade-off between efficiency and effectiveness. To scale to longer user sequences while keeping lightweight serving computation, existing works can be divided into two paradigms: (1) $\textit{Search-based Top-$K$ selection}$, which constructs an $\textbf{item-specific}$ subsequence for each candidate item to avoid facing the ultra-long sequence directly; and (2) $\textit{pre-trained user-interest compression}$, which maps an ultra-long user sequence into a small group of $\textbf{item-agnostic}$ user interest memories so that the online model can perceive user long-term interests via this highly compressed dense memory. Besides the two technical routes (totally item-specific or item-agnostic), we argue that there exists an intermediate path not well explored: preserving partial relevance between the user sequence and the target item, while exposing only limited signals to guide the direction of interest compression. This design does not pursue item-specific user interest compression, but seeks semantic-group shared general user interest memory according to item attributes, where semantically similar items share the same compressed user interest memory. Motivated by this, we propose $\textbf{UxSID}$, a novel framework that bridges this gap by facilitating a target item semantic-aware interaction between $\textbf{U}$ser histories and candidate $\textbf{S}$emantic $\textbf{ID}$s (SIDs). Specifically, UxSID employs a dual-level attention strategy: it first extracts item-agnostic user interests from raw sequences, and then performs a semantic-specific query over global behaviors and those agnostic interests to generate semantic-specific preferences. By adopting this end-to-end architecture, UxSID generates offline embeddings that balance computational parsimony with target items' semantic awareness, while strictly preserving parity with online inference in constant time. Extensive public benchmarks and large-scale A/B tests demonstrate that UxSID achieves state-of-the-art performance, driving a 0.337\% revenue lift in advertising. Open source link: $\url{https://anonymous.4open.science/r/UxSID/}$
VALOR: Vector-Aware Low-Rank Restructuring of Neural Networks for RISC-V Inference
Zhihao Xu ⋅ Xiaoning Du ⋅ Bixin Li ⋅ Wang Lulu ⋅ Li Liao ⋅ Zhou Ying
The RISC-V Vector Extension (RVV) provides SIMD-style vector execution for accelerating deep learning (DL) inference on edge processors. However, existing low-rank compression methods mainly select ranks according to accuracy preservation and theoretical floating-point operation (FLOP) reduction, without considering whether the selected ranks are profitable under a target RVV effective vector length. As a result, a low-rank factorization that reduces FLOPs may still introduce tail-handling overhead or even increase the RVV instruction count under specific vector configurations. To address this problem, we propose VALOR, a low-rank restructuring framework for efficient DL deployment on RISC-V processors. Given a pretrained model and a target RVV configuration, VALOR identifies decomposable layers in the pretrained model, filters rank candidates using an RVV instruction-profitable condition, and searches for a global restructuring strategy with minimal accuracy loss. Furthermore, to avoid repeated end-to-end evaluation, VALOR uses a Hessian-aware accuracy predictor that combines layer-wise factorization error with inter-layer coupling. Experiments on Spike and a real RISC-V vector processor, SpacemiT X60, show that VALOR achieves a better accuracy-efficiency trade-off than baseline methods.
Value-Rectified Distillation for Flow-based Offline Reinforcement Learning
Ke Jiang ⋅ Wen Jiang ⋅ Yoshinobu Kawahara ⋅ Xiaoyang Tan
We aim to address the problem of Out-of-Distribution (OOD) Mode Averaging, a critical problem that emerges when distilling Flow Matching (FM) based behavior policies in offline Reinforcement Learning (RL): When learning over multi-modal datasets, standard distillation inadvertently forces the policy to interpolate across different behavioral modes, causing the generated actions to collapse into invalid, OOD regions. To resolve this, we propose Value-Rectified Distillation (VRD), a novel framework that reformulates offline policy distillation as a value-guided generative trajectory alignment problem. By employing a dynamic coupling mechanism, VRD actively rectifies flow dynamics to explicitly decouple distinct behavioral modes, cleanly isolating high-value actions. Specifically, we instantiate VRD through two complementary algorithms: VRD-Q, which utilizes value-quantile behavior distillation to statistically filter out low-value behaviors, and VRD-R, which reorganizes the latent space to geometrically isolate high-value generative trajectories from the FM-based behavior policy. Extensive empirical evaluations demonstrate that VRD improves existing generative baselines across diverse continuous control and planning tasks on the D4RL, OGBench, and Minari benchmarks, especially over those multi-modal, non-expert datasets.
ValuSpec: Plug-and-Play Candidate Valuation before Target Verification for Tree-Based Speculative Decoding
Liang He ⋅ SiYuan Ma ⋅ Qishi Zhan ⋅ Yongqi Fan ⋅ Mingyu Cao ⋅ Qizhen Lan ⋅ Zhaolu Kang ⋅ Xilu Wang
Tree-based speculative decoding accelerates large language model inference by verifying multiple candidate branches in parallel. However, in practical serving scenarios, similar requests repeatedly induce overlapping candidate branches. While retrieval-based methods mitigate this by reusing historical drafts, the resulting candidate trees often contain many low-quality branches as the candidate tree grows, which reduces verification efficiency and wastes target-side computation. To address this, we propose ValuSpec, a plug-and-play candidate valuation mechanism that filters low-quality branches from a supplied candidate tree before expensive target-side verification. ValuSpec introduces transfer-based supervision to align filtering with the target model's behavior. A filtering threshold is estimated from the target model's greedy trajectories on a calibration split, without any parameter training. The target model then verifies the filtered tree under the output-identical greedy rule, so accepted tokens are guaranteed to match standard autoregressive decoding exactly. Experimental results on HumanEval, GSM8K, Dolly, and SpecBench show that ValuSpec achieves up to 2.38× end-to-end speedup in retrieval-based tree settings. Furthermore, integrating ValuSpec into the EAGLE-3 dynamic tree framework yields up to 2.78× end-to-end speedup without modifying its tree construction strategy, demonstrating that ValuSpec serves as a generally applicable and plug-and-play candidate valuation module.
Variance-Adaptive Optimal Algorithm for Reinforcement Learning with MNL Function Approximation
Wonyoung Kim ⋅ Garud Iyengar ⋅ Assaf Zeevi ⋅ Min-hwan Oh
Reinforcement learning with multinomial logistic (MNL) function approximation has become an important framework due to its flexibility and broad applicability. While existing studies have established regret guarantees under worst-case analysis, they do not capture how performance depends on the variability of the interaction between the learner and the environment. In this paper, we develop a new theoretical analysis for MNL-based Markov decision processes that yields explicit variance-adaptive regret bounds. Our algorithm is computationally efficient and achieves the instance-wise optimal rate of regret, narrowing the gap between upper and lower bounds. Our numerical experiments validate that our method learns optimal policies more efficiently than conventional approaches.
Variance-Averse n-Step Offline Reinforcement Learning for Sparse Long-Horizon Environments
Guhyeon Kang ⋅ Minhae Kwon
Generative actors are transforming offline reinforcement learning (RL) by enabling expressive policy classes that model complex action distributions. However, this expressiveness also exposes a key challenge in heterogeneous datasets: generative policies can reproduce unreliable action modes whose return distributions exhibit high variance, occasionally yielding high returns by chance but lacking consistency. Consequently, maximizing the expected Q-value alone is insufficient for identifying reliable actions. We propose VAN-Flow (Variance-Averse n-step Flow), a framework that promotes reliable actions in generative offline RL. VAN-Flow combines (i) a categorical distributional critic, (ii) a variance-averse expectation operator that smoothly reweights atom probabilities to favor actions with both high returns and low dispersion, and (iii) a flow-matching generative actor guided via rejection sampling. Unlike CVaR or mean--variance objectives, the operator redistributes probability mass over the categorical return distribution without hard truncation or auxiliary penalty terms. Across more than 40 tasks from D4RL and OGBench, VAN-Flow consistently outperforms strong baselines, with the largest gains in long-horizon and high-variance regimes where reliable action selection becomes critical.
Variance Reduction for Expectations with Diffusion Teachers
Jesse Bettencourt ⋅ Xindi Wu ⋅ Matan Atzmon ⋅ James Lucas ⋅ Jonathan Lorraine
Pretrained diffusion models increasingly serve as frozen teachers feeding downstream pipelines such as text-to-3D, single-step distillation, and data attribution. The teacher gradients these pipelines consume are Monte Carlo expectations over noise levels and Gaussian noise; their estimator variance dominates compute cost because each draw requires expensive upstream work, such as rendering, simulation, or encoding. We introduce CARV, a compute-aware variance-accounting framework that motivates a hierarchical Monte Carlo estimator: amortize the expensive upstream computation over cheap diffusion-noise resamples, sharpened by timestep importance sampling and a stratified inverse-CDF construction. Across diffusion-guided workloads, we obtain 2–3× effective compute multipliers, most from amortized reuse and approximately 25% additional gain from importance sampling plus stratification, without changing the objective. We also map regimes where these gains translate into improved downstream metrics versus regimes where they do not, such as DMD.
VCR: Learning Valid Contextual Representation for Incomplete Wearable Signals
Yuxuan Weng ⋅ Wenhan Luo ⋅ Qijia Shao
Wearable devices enable continuous health monitoring from multimodal signals, but real-world deployment is hindered by limited labeled data and pervasive sensor incompleteness. While large-scale self-supervised pretraining reduces label dependence, most existing methods assume full modality availability. Current approaches for handling modality missingness often reconstruct entire absent signals, which can encourage hallucinating modality-specific details that are not inferable from the observed sensor signals and degrade robustness. We propose VCR, a self-supervised framework that learns to extract valid representations robust to modality missingness. VCR employs an orthogonal tokenizer to enforce strict orthogonal disentanglement by rectifying latent manifolds and applying a geometric projection, separating each modality into shared semantics and modality-specific residuals. This design preserves complete information integrity while serving as a structural foundation for robust learning under modality missingness. The resulting tokens are processed by a missing-aware mixture-of-experts backbone that adapts to varying patterns of modality availability. By constraining the objective to reconstruct only the shared components of missing modalities, VCR effectively mitigates hallucinations of non-inferable modality-specific details. Across multiple health monitoring tasks, VCR consistently improves performance and robustness under full, single-missing, and multiple-missing modality settings compared with strong supervised and self-supervised baselines.
VecUQ-OT: Aggregating Uncertainty Measures via Multivariate Ranks
Nikita Kotelevskii ⋅ Vladimir Kondratyev ⋅ Maiya Goloburda ⋅ Alexander Fishkov ⋅ Mohsen Guizani ⋅ Eric Moulines ⋅ Maxim Panov
Decision-making under uncertainty typically requires a single scalar score, forcing practitioners to commit to one uncertainty measure. However, different measures often capture complementary failure modes, and relying on a single measure may be insufficient for downstream tasks. We propose \emph{VecUQ-OT}, a label-free procedure that aggregates multiple scalar uncertainty measures into a single ranking score via entropy-regularized optimal transport. Rather than selecting a single measure, our approach combines signals of (possibly) different natures through non-additive fusion based on \emph{multivariate ranks}. The construction is motivated by the theory of multivariate ranks via exact optimal transport and realized through a tractable entropic approximation that generalizes to unseen inputs without retraining. Experiments across synthetic, image, and text domains demonstrate that \emph{VecUQ-OT} provides stable performance across tasks, remains reliable when individual measures fail, and outperforms the natural additive fusion alternative.
VEHBench: A Stage-Local Diagnostic Benchmark for LLM-Assisted Vibration Energy Harvester Design
Depeng Su ⋅ Yuyu Luo ⋅ Guobiao Hu
Battery-free Internet of Things (IoT) requires iterative design of vibration energy harvesters (VEHs) under coupled physical constraints, while LLMs are emerging as interface layers for engineering workflows. However, existing engineering benchmarks primarily assess final artifact validity, offering limited insights into how LLMs behave across different stages of coupled physical design. We introduce VEHBench, an engineering-native diagnostic benchmark for LLM-assisted VEH design, featuring 763 literature-grounded tasks scored by an analytical physical oracle. VEHBench evaluates four design roles: specification triage, verifier-guided search, corrupted-state recovery, and policy-conditioned selection. Experimental results reveal that LLM capability is strongly stage-dependent: no single model consistently dominates the entire workflow, and response-control profiles expose distinct behavioral patterns across design roles. VEHBench thus provides a stage-aware foundation for evaluating, selecting, routing, and improving verifier-grounded engineering LLMs. The anonymous artifact is available at https://huggingface.co/datasets/AnonymousVehbench/vehbench.
Verifier-Backed Hard Problem Generation for Mathematical Reasoning
Yuhang Lai ⋅ Jiazhan Feng ⋅ Yee Whye Teh ⋅ Ning Miao
Large Language Models (LLMs) demonstrate strong capability in solving scientific and mathematical problems, yet they struggle to produce valid and challenging novel problems, an essential component for advancing LLM training and enabling autonomous scientific research. Existing problem generation approaches either depend on expensive human expert involvement or adopt naive self-play paradigms, which frequently yield invalid problems due to reward hacking. This work introduces VHG, a verifier-enhanced hard problem generation framework built upon three-party self-play. By integrating an independent verifier into the conventional setter-solver duality, our design constrains the setter’s reward to be jointly determined by problem validity (evaluated by the verifier) and difficulty (assessed by the solver). We instantiate two verifier variants: a Hard symbolic verifier and a Soft LLM-based verifier, with evaluations conducted on indefinite integral tasks and general mathematical reasoning tasks. Experimental results show that VHG substantially outperforms all baseline methods by a clear margin. Our codes are anonymously available at https://github.com/VHGMath/VHG_Math.
Verigrad: Verification-Driven Multi-Agent GPU Kernel Generation for High-Order MLIP Derivatives
yao liu ⋅ Yuanchang Zhou ⋅ Hongtao Xu ⋅ Mingzhen Li
High-order derivative kernels in machine learning interatomic potentials (MLIPs) are a dominant cost in scientific ML, yet current autonomous GPU-kernel agents do not reliably target this workload. They generate kernels against local tensor-level references, whereas MLIP derivative workloads span forward, backward, and higher-order automatic differentiation (AD) phases with cross-phase saved tensors and reconstruction rules. We introduce \textsc{Verigrad}, a verification-driven multi-agent harness that makes high-order derivative generation a first-class kernel-generation workload. \textsc{Verigrad} first exposes derivative semantics through \texttt{DerivativeTask} and a derivative-aware IR, and then decomposes the workload into verifiable kernel contracts. An execution-based derivative verifier validates the generated artifacts; the same verifier then gates hardware-guided refinement, so optimized kernels replace earlier candidates only when the derivative checks continue to pass. We integrate the generated kernels into MatRIS training and inference, where \textsc{Verigrad} attains 1.33--1.55$\times$ end-to-end speedup over PyTorch eager across energy-only, energy--force, and full-derivative inference and training workloads.
VETime: Vision Enhanced Zero-Shot Time Series Anomaly Detection
Yingyuan Yang ⋅ Tian Lan ⋅ Yifei Gao ⋅ Yimeng Lu ⋅ Xuming An ⋅ Wenjun He ⋅ Meng Wang ⋅ Chenghao Liu ⋅ Chen Zhang
Time-series anomaly detection (TSAD) requires identifying both immediate Point Anomalies and long-range Context Anomalies. However, existing zero-shot foundation models face a fundamental trade-off: 1D temporal models provide fine-grained pointwise localization but lack a global contextual perspective, while 2D vision-based models capture global patterns but suffer from information bottlenecks due to a lack of temporal alignment and coarse-grained pointwise detection. To resolve this dilemma, we propose VETime, the first TSAD framework that unifies temporal and visual modalities through fine-grained visual-temporal alignment and dynamic fusion. VETime introduces a Reversible Image Conversion and a Patch-Level Temporal Alignment module to establish a shared visual-temporal timeline, preserving discriminative details while maintaining temporal sensitivity. Furthermore, we design an Anomaly Window Contrastive Learning mechanism and a Task-Adaptive Multi-Modal Fusion to adaptively integrate the complementary perceptual strengths of both modalities. Extensive experiments demonstrate that VETime significantly outperforms state-of-the-art models in zero-shot scenarios, achieving superior localization precision with lower computational overhead than current vision-based approaches.
vExpert: Virtualizing Expert Storage for Adaptive Load Balancing in Distributed MoE Inference
Wenxun Wang ⋅ Xiuhong Li ⋅ Yida Wang ⋅ Chen Tang ⋅ Zongle Huang ⋅ Ke Hong ⋅ Yu Wang ⋅ Yongpan Liu
Mixture-of-Experts (MoE) models have become prevailing, yet their distributed deployment suffers from severe dynamic load imbalance due to the conflict between static expert placement and inherent routing dynamism. While host DRAM offers capacity to avoid static device storage for load balancing, we identify existing offloading works fundamentally fail in distributed inference due to two overlooked challenges: (1) coupled performance impacts of runtime Host-to-Device transfers due to load dynamism and shared PCIe topology; (2) system-level memory overhead incurred by DRAM usage at scale. In this paper, we propose vExpert that shifts DRAM offloading paradigm from partial residency to full redundancy by virtualized expert storage. We pioneer the first systematic performance model that jointly captures H2D latency, PCIe contention, and load imbalance, building a lightweight allocator to dynamically map physical experts to virtual slots. A novel disaggregated manager is designed to reduce memory overhead in multi-process environments. Evaluated on DeepSeek-V3.1 within an 8-node H100 cluster, vExpert improves load balance ratio by up to 55\% and achieves 1.2$\times$ end-to-end speedup. Code is available at https://anonymous.4open.science/r/vExpert-B743/.
Video-Zero: Self-Evolution Video Understanding
ruixu zhang ⋅ Deyi Ji ⋅ Lanyun Zhu ⋅ Xuanyi Liu ⋅ Yuxin Meng ⋅ Ruihang Chu ⋅ Yujiu Yang
Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to video understanding remains underexplored and challenging: videos are long, dynamic, and redundant, while the evidence needed for reasoning is often sparse and temporally localized. Naively generating difficult question-answer pairs from full videos can therefore produce supervision that appears challenging but is weakly grounded, relying on static cues or language priors rather than temporal evidence. In this work, we argue that the key bottleneck of video self-evolution is not difficulty alone, but grounding. We propose \textbf{Video-Zero}, an annotation-free Questioner--Solver co-evolution framework that centers self-evolution on temporally localized evidence. The Questioner discovers informative evidence segments and generates evidence-grounded questions, while the Solver learns to answer and align its predictions with the supporting evidence. This closes an iterative loop of evidence discovery, grounded supervision, and evidence-aligned learning. Across 13 benchmarks spanning temporal grounding, long-video understanding, and video reasoning, Video-Zero consistently improves multiple video VLM backbones, demonstrating the effectiveness and transferability of evidence-centered self-evolution. Code and models will be made publicly available.
ViewRec3D: Learning to Recommend 3D Viewpoints for AI Photography
Chengzhi Zhang ⋅ Yunzhong Hou ⋅ Zizheng Sun ⋅ Ming-Hsuan Yang ⋅ Chi Liu
Viewpoint dictates the photo composition and plays a vital role in AI photography. Existing 3D viewpoint recommendation methods are often limited in viewpoint adjustments, which is mainly due to their lack of high-quality training data with accurate and large 3D view changes. Ideally, the training data should contain suboptimal images and expert-selected optimal viewpoints, but this can be very expensive to curate. To solve this, we build an automatically-generated 3D viewpoint recommendation dataset from expert photos. Specifically, we first consider these expert photos as optimal and outpaint them to provide more information on scene layouts. Next, we explicitly reconstruct the 3D scene. We then render suboptimal images from random viewpoints. Lastly, we introduce a hierarchical data filtering scheme and acquire high-quality data pairs for our ViewRecDB-100K dataset. On top of this, we further introduce a view recommendation model, ViewRecNet, that can predict expert viewpoints from any suboptimal image inputs using a ray-based dense supervision. The proposed method achieves strong results both quantitatively and qualitatively, supported by LLM-based viewpoint preference and human studies. Data and code are available at https://github.com/anonymous1-submission/ViewRec3D.git.
View-Spectral Reconciliation Learning for Text-based Multispectral Aerial–Ground Person Re-Identification
Zhongao Zhou ⋅ Bin Yang ⋅ Yuxuan Zhao ⋅ Yudi Xie ⋅ Jun Chen ⋅ Mang Ye
Conventional text-based person re-identification (ReID) approaches mainly rely on visible-spectrum images captured from ground-level viewpoints. However, with the increasing prevalence of UAVs and multispectral imaging systems, pedestrians can be captured from both ground and aerial viewpoints, and different spectral sensors provide complementary cues that reveal diverse pedestrian characteristics. As a result, conventional text-based ReID methods become inadequate for such retrieval scenarios. To address this limitation, we introduce a novel task termed Text-based Multispectral Aerial-Ground Person Re-Identification (TMAG-ReID), which aims to retrieve pedestrian images across aerial and ground viewpoints under multispectral scenarios using textual descriptions. To facilitate this task, we construct the T-MSAG dataset, which consists of multispectral pedestrian images captured from aerial and ground viewpoints, paired with corresponding textual descriptions. Furthermore, we propose a novel View-Spectral Reconciliation Learning framework (VSRL) that leverages textual guidance to interweave discriminative semantics across distinct views and spectral modalities. This fusion yields robust cross-view and cross-modal pedestrian features, thereby enabling high-performance retrieval. Extensive experiments on the T-MSAG dataset demonstrate the effectiveness of VSRL.
VIGOR: Visual Gain Ordering for Hallucination Mitigation in Multimodal Discrete Diffusion Language Models
Junzhe Chen ⋅ Tian Qin ⋅ Eugenie Shi ⋅ Tianshu Zhang ⋅ Qiang Ju ⋅ Lijie Wen
Multimodal discrete diffusion language models (dLLMs) generate responses through iterative mask prediction and offer a promising alternative to autoregressive large vision-language models. However, they still suffer from hallucination, producing outputs that are plausible under language priors but insufficiently grounded in the image. We argue that hallucination in multimodal dLLMs is not only a token prediction problem, but also a premature commitment problem: during iterative denoising, some masked positions are unmasked too early because their confidence is driven more by language priors than by visual evidence, and these early commitments can bias later generation. To address this issue, we propose Visual Gain Ordering (VIGOR), a training-free and plug-and-play inference strategy that reorders masked positions according to counterfactual visual evidence. At each denoising step, VIGOR compares the confidence of each candidate token under the original image-conditioned pass and an additional visual-ablation pass, and prioritizes positions whose predictions are more strongly supported by the image. Experiments on LLaDA-V and Lumina-DiMOO show that VIGOR consistently reduces hallucination and improves multimodal reasoning performance across benchmarks. Anonymous code is available at https://anonymous.4open.science/r/519_VIGOR-4051.
VINCIE-NExT: Unlocking Video Editing from Images via In-Context Modeling
Leigang Qu ⋅ Feng Cheng ⋅ Ziyan Yang ⋅ Bangbang Yang ⋅ Zhaoyang Huang ⋅ Wei Chow ⋅ Yicong Li ⋅ Wenjie Wang ⋅ Tat-Seng Chua ⋅ Yan Zeng
Building a capable video editor remains significantly harder than a video generator: editing requires (source, instruction, edited) triplets that are prohibitively expensive to annotate and difficult to synthesize at scale, whereas image editing has already reached maturity with millions of such pairs readily available. In this work, we introduce Chain-of-Editing, a unified framework that transfers editing capability from images to videos through in-context visual demonstrations, alleviating the need for large-scale paired video editing data. Chain-of-Editing decomposes video editing into a structured chain of composable sub-tasks (Video to Image to Image to Video), routing editing intent through the image domain and enabling scalable joint training from heterogeneous image and video corpora under a unified diffusion objective. An image editing pair, synthesized by the model or supplied by the user, is prepended as an in-context visual demonstration that serves as a spatial appearance blueprint for every output frame. To ground appearance edits across the interleaved context, we introduce a novel position encoding that links image demonstrations and video frames in a shared spatial coordinate system, enabling pixel-faithful propagation of appearance changes to every output frame. Chain-of-Editing further provides principled test-time scaling: by executing the sub-task chain as progressive diffusion stages, editing quality can be improved by investing additional compute without retraining. Comprehensive experiments on OpenVE-Bench demonstrate the state-of-the-art performance across diverse editing categories, with ablations confirming the effectiveness of each component.
VIPER: An Expert-Curated Benchmark for Vision-Language Models in Veterinary Pathology
Luca Weishaupt ⋅ Simone de Brot ⋅ Javier Asin ⋅ Llorenç Grau-Roma ⋅ Nic G Reitsam ⋅ Andrew H Song ⋅ Dongmin Bang ⋅ Stefan T Kaluziak ⋅ Long P Le ⋅ Jakob N Kather ⋅ Faisal Mahmood ⋅ Guillaume Jaume
Pathology vision-language models are advancing rapidly, yet existing benchmarks remain focused on human tissue, particularly oncology, leaving non-human pathology largely unaddressed. This gap is especially important in toxicologic pathology, where microscopic tissue examination of laboratory animals is a core component of preclinical drug safety assessment. To address it, we introduce VIPER, the first expert-curated benchmark for vision-language model evaluation in toxicologic pathology. VIPER contains 1251 questions associated with 419 H&E-stained rat histology images across 9 organs, covering multiple-choice, KPrim, and free-text formats. All questions were curated and validated by board-certified veterinary pathologists. In total, we benchmarked 16 models, including two newly introduced veterinary-pathology models, seven human pathology-specialized models, and seven general-purpose frontier models. The results identify a substantial domain gap between veterinary and human pathology, expose the risk of over-diagnosis of normal tissue in frontier models, and show that domain-specific training remains critical for visually grounded predictions. VIPER data and evaluation code are available at https://github.com/mahmoodlab/viper.
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
Yudong Han ⋅ Yong Wang ⋅ Zaiquan Yang ⋅ Zhen Qu ⋅ Liyuan Pan ⋅ Xiangxiang Chu
Multimodal latent reasoning has emerged as a promising paradigm that replaces explicit Chain-of-Thought (CoT) decoding with implicit feature propagation, simultaneously enhancing representation informativeness and reducing inference latency. By analyzing token-level gradient dynamics during latent training, we reveal two critical observations: (1) visual tokens exhibit significantly smaller gradient norms than their textual counterparts due to inherent language bias, resulting in systematic visual under-optimization; and (2) semantically simple tokens converge rapidly, whereas complex tokens exhibit persistent gradient instability constrained by fixed architectural depths. To address these limitations, we propose a visual replay module and routing depth scaling to collaboratively enhance visual perception and refine complicated latents for deeper contextual reasoning. The former module leverages causal self-attention to estimate token saliency, reinforcing fine-grained grounding through spatially-coherent constraints. Complementarily, the latter mechanism adaptively allocates additional reasoning steps to complex tokens, enabling deeper contextual refinement. Guided by a curriculum strategy that progressively internalizes explicit CoT into compact latent representations, our framework achieves state-of-the-art performance across diverse benchmarks while delivering substantial inference speedups over explicit CoT baselines.
Visual text compression (VTC) promises efficient long-context processing by rendering text into an image and re-encoding it with a vision-language model, often producing $3$--$20\times$ fewer decoder tokens than subword tokenization. Yet token savings do not translate predictably into downstream utility: on some tasks the visual path matches or exceeds the text path, on others it collapses, and the compression ratio itself does not predict which regime will occur. The missing quantity is therefore not another summary of efficiency, but a principled measure of task-relevant information loss induced by visual encoding. We address this problem by formulating VTC in the language of measure transport. Treating text and visual tokens as empirical probability measures, we show that the ViT patch encoder induces a push-forward map whose transport cost decomposes into a precision cost from within-patch aggregation and a coverage cost from cross-patch fragmentation. Both terms are estimable from downstream-label-free probes. This formulation yields two operational consequences: a downstream-label-free routing criterion that selects whether to use the visual path for a given input or benchmark instance, and a transport-informed foveation mechanism that re-encodes high-cost regions at higher resolution. Across $24$ NLP datasets at Qwen3-4B, our label-free rule matches the per-dataset oracle on $17/24$ datasets ($70.8\%$), and improves the average task score by $+3.3\%$ with $-10.3\%$ average tokens relative to a pure-LLM.
Visual-to-Executable Procedural Reconstruction of Buildings.
Zichong Zhu ⋅ Zhihao Liu ⋅ Andrei Sharf ⋅ Zhanglin Cheng
Procedural generation is well-suited for modeling 3D buildings, as architectural structures are naturally composed of repetitive and parameterized components. It provides a compact and editable representation by encoding geometry through executable rules. However, constructing such procedural representations typically relies on manually designed grammars, making them difficult to obtain in practice. We study image-based procedural facade reconstruction, where the goal is to recover an editable building procedural representation from an in-the-wild building image. We propose a hierarchical neuro-symbolic framework that decomposes the reconstruction into two coupled levels: a global structure layout tree for split-and-repeat facade organization, and a component parameterization representation for local architectural. A fine-tuned layout model recovers and refines the global structure through rendered visual feedback, while a tile-level predictor maps representative component crops to structured procedural attributes. The two representations are then deterministically compiled into executable CGA programs. Experiments show that our method achieves strong performance on both regular facades and challenging in-the-wild images, while maintaining high structural consistency and procedural executability. In addition, the tile-level predictor improves fine-grained architectural detail, enabling accurate and controllable component synthesis.
VLMGuard: Bootstrapping Malicious Prompt Detectors from Unlabeled Vision-Language Prompts in the Wild
Junlin Fang ⋅ Wenyu Chen ⋅ Reshmi Ghosh ⋅ Robert Sim ⋅ Ahmed Salem ⋅ Vitor Carvalho ⋅ Emily Lawton ⋅ Sharon Li ⋅ Jack Stokes ⋅ Sean Du
Vision-language Models (VLMs) are essential for contextual understanding of both visual and textual information. However, their vulnerability to adversarially manipulated inputs presents significant risks, leading to compromised outputs and raising concerns about the reliability in VLM-integrated applications. Detecting these malicious prompts is thus crucial for maintaining trust in VLM generations. A major challenge in developing a safeguarding prompt classifier is the lack of a large amount of labeled benign and malicious data. To address the issue, we introduce VLMGuard, a novel learning framework that leverages the unlabeled user prompts in the wild for malicious prompt detection. These unlabeled prompts, which naturally arise when VLMs are deployed in the open world, consist of both benign and malicious information. To harness the unlabeled data, we present an automated maliciousness estimation score for distinguishing between benign and malicious samples within this unlabeled mixture, thereby enabling the training of a binary prompt classifier on top. Notably, our framework does not require extra human annotations and is robust to realistic prompt variations, offering strong flexibility and practicality for real-world applications. Extensive experiments show that VLMGuard achieves superior detection results, improving AUROC by 9.46% on average over the state-of-the-art method. Disclaimer: This paper may contain offensive examples; reader discretion is advised. Code is available at: https://github.com/radiolab-ntu/vlmguard.
WASD: Wasserstein-based Knowledge Distillation for Large Language Models
Byeonghu Na ⋅ Donghyeok Shin ⋅ Yeongmin Kim ⋅ Mina Kang ⋅ Il-chul Moon
Autoregressive large language models (LLMs) have rapidly advanced in capability, but their increasing scale comes with substantial computational and memory costs at inference time. Knowledge distillation (KD) offers a practical solution by transferring knowledge from a large teacher model to a smaller student model via alignment of discrete probability distributions. However, existing KD methods for LLMs primarily rely on divergences that evaluate discrepancies through probability values at each vocabulary index, without explicitly leveraging token-level semantic information. We propose Wasserstein-based knowledge distillation (WASD) for LLMs, which incorporates token-level semantic information via the Wasserstein-based distance with a cost matrix derived from token embeddings. To ensure computational tractability, we adopt the Sinkhorn divergence and derive an equivalent formulation that can be efficiently optimized without introducing additional networks. Experiments across multiple LLM families and scales show that WASD consistently improves distillation performance on diverse tasks, including instruction following, mathematical reasoning, and code generation. Our results highlight the importance of semantic information encoded in the token space for effective distribution alignment in LLM distillation.
Wasserstein Gradient Flows and Forward-Only Diffusion Are Not Enough for Multimodal Sampling
Daniel McBride ⋅ Pratik Khandagale ⋅ Cristina Garcia-Cardona ⋅ Yen Ting Lin
There has been a plethora of algorithms proposed that leverage Wasserstein gradient flows or forward-only diffusion processes to perform sampling tasks. These approaches are often characterized by theoretical guarantees of exponentially fast convergence to the target distribution. In this work, we study the mixing behavior of this family of sampling methods. By invoking the Jordan--Kinderlehrer--Otto scheme and Otto calculus, we first establish that Wasserstein gradient flow and forward diffusion-based samplers share the same density evolution. Consequently, their convergence can be characterized by studying the convergence behavior of the associated diffusion process, which has been extensively analyzed in nonequilibrium statistical physics. We perform two independent and well-established analyses for this purpose, namely spectral analysis and mean first-passage time analysis. We show that, although the sampling distributions converge to the target distribution exponentially fast, the presence of multimodality can lead to exponentially long mixing times associated with small spectral gaps. Further analysis elucidates that even when combined with annealing, this family of samplers still requires impractically long time to sample multimodal distributions. We discuss the origin of this behavior as a consequence of the local (i.e., gradient-driven) dynamics underlying this class of samplers, and highlight the need for non-local mechanisms to facilitate the transport of probability mass across modes, enabling efficient exploration in multimodal settings.
Watermarking is widely proposed for provenance, attribution, and safety monitoring in generative models, yet is typically evaluated only under adversaries who attempt to evade detection or induce false positives at the level of individual samples. We argue that watermarking should be treated as a monitoring primitive, and that internal monitoring is unavoidable given per-entity attribution keys and messages, as well as detector access. We introduce an observer-based threat model in which observers can aggregate watermark signals across outputs to infer entity-level information, showing that even zero-bit watermarking enables attribution under multi-key settings. We further show that external monitoring can emerge over time from persistent, key-dependent statistical structure, although this depends on watermark design and may be mitigated by distribution-preserving or undetectable schemes. Our findings reveal a fundamental dual-use tension between attribution and monitoring, motivating evaluation of watermarking beyond per-sample robustness to account for aggregation and observer-based capabilities.
WaveSem: Frequency-Adaptive Tokenization for Disentangling Semantics and Noise in Genomics
Shou Z Chen ⋅ Bing He ⋅ Zhenchao Tang ⋅ Jun Zhu ⋅ Minghao Yang ⋅ Tianxu Lv ⋅ Yu Wang ⋅ Jiayang Wu ⋅ Fang Wang ⋅ Yu Zhao ⋅ Chenchen Qin ⋅ Jianhua Yao
High-throughput scientific sensing, exemplified by genomics, faces a critical representation bottleneck: discrete biological events are often obscured by continuous, high-entropy stochastic noise. Standard neural codecs fail to distinguish deterministic semantics from stochastic sensor noise, inefficiently allocating their information budget to high-entropy fluctuations rather than to low-entropy structural motifs. We introduce \textsc{WaveSem}, a frequency-adaptive framework that imposes a physically motivated inductive bias on discrete representation learning. Through integrated wavelet decomposition, \textsc{WaveSem} explicitly disentangles each signal into a structured semantic core and a residual texture component. We validate this approach on large-scale nanopore sequencing data, a domain characterized by extreme sequence lengths and low signal-to-noise ratios (SNRs). \textsc{WaveSem} achieves a $14\times$ storage reduction while maintaining competitive downstream basecalling accuracy. Moreover, linear probing of the frozen encoder reaches $\mathrm{AUROC}=0.928$ for 5mC methylation detection, demonstrating that the learned tokens preserve semantic information beyond nucleotide identity. These results establish disentangled vector-quantized representations as a scalable foundation for efficient genomic analysis and long-term archival.
We Should Distinguish Unlearning From Untraining
Eleni Triantafillou ⋅ Imtiaz Humayun ⋅ Mónica Ribero ⋅ Alexander Turner ⋅ Michael Mozer ⋅ Georgios Kaissis
There has been a recent surge in interest about the question of how we can "delete" specific data points or behaviours from a trained model, a goal referred to as "machine unlearning". We argue that the umbrella term "unlearning" actually spans two distinct problem formulations, but the distinction between them has not yet been observed in literature. This causes ambiguity around when an unlearning algorithm is expected to work, leads to the use of inappropriate metrics and baselines when comparing algorithms to one another, difficulty in interpreting results, and missed opportunities for pursuing critical research directions. In this paper, we argue for the position that a fundamental distinction must be made between two notions that we refer to as Unlearning and Untraining. On one hand, Untraining aims to reverse the effect of having trained on a given "forget set", i.e. to remove the influence that that specific set of examples had on the model during training. On the other hand, Unlearning aims not just to remove the influence of those given examples, but also of the entire underlying subdistribution from which they were sampled (and thus e.g. the concept or model behaviour that those examples represent). We discuss technical definitions of these problems and map problem settings studied in the literature to each notion. By disambiguating technical definitions, our work aims to accelerate progress in this important field.
What Do SAE Features Encode? Evidence from Human Neural Activity
Yujin Kang ⋅ Hyojin Park ⋅ Yoon-Sik Cho
Sparse Autoencoders (SAEs) decompose dense LLM activations into sparse, interpretable features. However, evaluating whether SAEs extract genuinely meaningful structure remains challenging. Current approaches rely on model-internal metrics such as reconstruction fidelity and automated LLM scoring, which assess SAEs as mathematical decompositions but provide no external validation that the extracted features correspond to anything outside the model. To address this issue, we propose a new validation approach: comparing SAE representations against human neural activity. The validation rests on a shared computational principle, since both SAEs and biological neural systems implement sparse coding over overcomplete populations. Through extensive experiments using EEG recordings of naturalistic reading and SAE features from three large language models, we demonstrate that SAE features systematically align with human brain activity. We further show that this alignment is dominantly carried by the SAE's learned sparse code, rather than by generic architectural properties or the scale of the training data. This work is the first attempt to validate pretrained SAE features against human brain activity, establishing biological alignment as a complementary benchmark for mechanistic interpretability research beyond model-internal metrics.
What Is Worth Representing? Representational Empowerment for Continual Model Construction
Fei Dai ⋅ Hanqi Zhou ⋅ Alison Gopnik ⋅ Charley M Wu
The first problem of modeling the world is not just estimating the right parameters or causal structure, but deciding *what should be represented at all*. We frame this as *continual model construction*: an agent maintains an environment-specific model $M$ of an inaccessible world $W$ and curates a persistent library $\mathcal{L}$ of reusable representational elements across environments. We propose *Representational Empowerment* ($\operatorname{RepEmp}$) to score candidate elements by how much they expand the agent's future capacity to model and plan---a counterpart to classical environmental empowerment, redirected from control over external states to control over internal representations. We realize the framework as a hierarchical Actor-Curator architecture and test it across three experiments. In a finite-vocabulary causal-learning task, human participants construct causal models at varying abstraction levels to maximize goal reachability rather than fidelity to the world---a signature better predicted by $\operatorname{RepEmp}$ over information-gain or novelty alternatives. Matched simulations reveal $\operatorname{RepEmp}$-guided construction, not exploration, to be the driver of sufficient structure recovery and cross-task transfer. Finally, in an open-vocabulary planning domain, an LLM-augmented Curator builds more compact symbolic libraries, which also generalize better than baselines. Ablating $\operatorname{RepEmp}$ eliminates these benefits. Together, these results identify $\operatorname{RepEmp}$ as a potential principle for continual model construction: deciding what to build, retain, and reuse under bounded resources.
What, Where, and Boundary: Hierarchical Cognitive Decomposition for Echocardiography Video Segmentation
Long Zheng ⋅ Zhi Li ⋅ Weidong Wang ⋅ Chenxi Zhu ⋅ Zhenyu Dai ⋅ Minyi Guo
Segmenting the left ventricle in echocardiographic videos remains difficult because the target undergoes continuous deformation and rapid displacement across the cardiac cycle, while low tissue contrast further obscures its endocardial boundaries. Expert echocardiographers navigate these challenges through a hierarchical cognitive workflow, sequentially addressing three fundamental questions:What is the target? Where is it now? Where does the boundary lie? Yet existing methods predict masks end-to-end from temporal features without explicitly modeling this cognitive process. We propose Hierarchical Cognitive Decomposition (HCD), which decomposes the segmentation task into three stages that emulate this expert interpretation workflow. To capture invariant target identity amid changing appearances, Static Identity Anchoring (SIA) anchors the first annotated frame as a semantic reference, enabling stable recognition across the cardiac cycle. To maintain spatial focus despite inter-frame motion, Adaptive Gaze Prior (AGP) converts the previous frame's prediction into a spatial attention prior that localizes the current target position. To resolve boundary ambiguity under low tissue contrast, Structural Boundary Perception (SBP) extracts multi-scale boundary cues within the localized region. On the CAMUS and EchoNet-Dynamic benchmarks, HCD achieves state-of-the-art performance with only 1.35M trainable parameters (out of 35.2M total) at 68 FPS, demonstrating that explicitly encoding clinical cognitive priors into network design yields both effective and efficient segmentation.
When Copying Is Hard: Copy-Constrained Decoding for Exact Span Reproduction
Jinghui Zhang ⋅ Lang Gao ⋅ Zongfang Liu ⋅ Ruihong Zeng ⋅ Zirui Song ⋅ Rui Yan ⋅ Kentaro Inui ⋅ Xiuying Chen
Precisely reproducing a target span verbatim within a long context is a fundamental capability test for LLMs: even when the target string is present in context, models often still drift within the span or fail to stop correctly, revealing limitations in source grounding and boundary control. This capability has broad practical relevance. For instance, code generation often requires embedding sensitive strings such as API keys exactly as they appear in context. In this work, we construct a benchmark to test exact span reproduction ability across contracts, scientific literature, code, and random sequences, and find that current models still make surprisingly frequent errors. To address this, we propose CopyGen, a lightweight copy-constrained plugin for frozen LLMs that turns copying from a byproduct of next-token prediction into a controlled process. Our key idea is to decompose copying into three distinct decisions: when to copy, how far to copy, and when to stop. Concretely, CopyGen first determines whether to enter copy mode, then predicts a tentative copy length, and finally verifies the generated span token by token. During verification, we adaptively adjust the logits of copy continuation tokens to incentivize or suppress continuation, and emit only the longest valid prefix. Across three LLM backbones and four benchmark tasks, CopyGen improves average exact match by 35% relative without model finetuning, and integrates seamlessly with SFT to further improve it by 4%, while substantially accelerating generation. These results suggest that reliable exact copying is not something LLMs obtain for free, but a decoding-time control problem that requires explicit modeling. Code is available at https://anonymous.4open.science/r/CopyGen-A65D.
When Debate Helps: Proposal Supply and Verification-Aware Readout in Multi-Agent Reasoning
Zihao Zhao ⋅ Tunyu Zhang ⋅ Haizhou Shi ⋅ Yusong Zhao ⋅ Xinxi Zhang ⋅ Hao Wang
Multi-agent debate can improve reasoning, yet often fails to beat simple majority voting. Prior martingale-null theory explains such failures as debate without truth-directed signal. We develop a more complete account of when practical LLM debate succeeds: it needs both correct minority proposals and a readout that can recover them. We formalize this view through latent headroom, the gap between majority vote and proposal-oracle performance, which measures unused correct expertise in an expert society. For readout, we developLatent Verification Debate (LVD), a theory that models candidate proposals as receiving latent verified evidence before influencing final generation. For supply, we study neural-thicket expert societies, where nearby model perturbations provide heterogeneous specialists, and introduce a coverage-based construction that increases complementary proposal supply. Across matched-budget reasoning and multi-discipline benchmarks, our construction increases recoverable headroom, and controlled readout experiments show that gains arise from recovering surfaced correct minorities rather than merely adding interaction. Together, these results identify diverse proposal supply and verification-aware utilization as the two mechanisms that determine when debate helps.
When Depth Lies: Benchmarking Vision-Language Models on Mirror-Induced RGB-D Ambiguity
HAO YIN ⋅ Tianchen Guo ⋅ Heming Du ⋅ Yan Ke ⋅ Yanbin Liu ⋅ Xin Yu
Vision-language models (VLMs) have achieved significant progress in scene understanding and spatial reasoning. However, existing evaluation benchmarks predominantly assume that all visible content corresponds directly to physical entities in the scene, leaving model performance in reflective scenarios largely unexamined. Reflective surfaces preserve object appearance while altering the mapping between appearance and physical location, causing VLMs to produce systematic reasoning errors in such scenes. In this work, we introduce (i) a grounded RGB-D dataset for mirror-aware spatial reasoning, and (ii) the first comprehensive and systematic evaluation of state-of-the-art VLMs in reflective scenarios. Specifically, we collect 1,011 RGB-D images from 525 real-world scenes, spanning diverse indoor and outdoor environments and six reflective surface types, under varying lighting conditions across both daytime and nighttime settings. All images are annotated with reflective surface region masks and fine-grained, polygon-level instance annotations for both directly observed and mirror-reflected objects, with depth maps serving as both geometric ground truth for annotation and optional model input. We curate a mirror-induced reasoning ambiguity benchmark, named \textbf{MIRA-Bench} (Mirror-Induced Reasoning Ambiguity Benchmark). MIRA-Bench comprises 2,955 unique QA pairs across eight sub-tasks organized into three cognitive levels: \mbox{Reflection-aware} Perception, Spatial Reasoning, and \mbox{Scene-level} \mbox{Decision-making}, with five object-referencing protocols yielding 14,140 QA samples. We evaluate state-of-the-art VLMs on MIRA-Bench and find that these models consistently underperform across tasks, suggesting that current VLMs do not yet understand mirrors. Common failure patterns include misclassifying whether objects are directly observed or seen through reflections, failing to identify that reflected and directly observed instances correspond to the same physical object, producing erroneous cross-space spatial inferences, and tending to adopt conservative default strategies in action-oriented decisions instead of reasoning about the true scene layout. We will release the dataset and evaluation code to support future research.
When Do Cosine Prototypes Mislead? A Whitening-Aware Benchmark for Frozen-Feature Image Classification
Jian Ding ⋅ Su Yang
Cosine nearest-class-mean (NCM) evaluation is a widely used default for closed-set frozen-feature image classification because it is deterministic, train-free, and easy to reproduce. This convenience hides a metric assumption: raw angular geometry alone suffices for benchmark conclusions. We show that the assumption changes model-selection outcomes, not just absolute accuracies: across nine datasets and six frozen backbones, cosine-only reporting changes which backbone ranks first on 2 of 9 datasets and incurs 6.71 percentage points of regret relative to a best-evaluator oracle; on fixed DINOv2-B features, the underlying cosine-to-LDA gap reaches 37.63 points on Aircraft. Motivated by the low-variance spectral directions where class-discriminative signal often appears in such high-gap cases, we call this evaluator dependence the Spectral Decision Gap (SDG). SDG quantifies the gap between raw cosine and train-validated whitening-aware reference evaluators, flagging when cosine-only reporting changes benchmark conclusions. A leave-one-dataset predictive protocol reaches Pearson $r=0.752$ for best-whitening disagreement and incurs only 0.56pp routing regret at a 5pp threshold. This paper provides a compact audit protocol and artifact suite for detecting, predicting, and reporting when cosine prototype conclusions are reliable in this setting. Code and artifacts are available anonymously at https://anonymous.4open.science/r/sdg-ed/.
A representation that scrambles the true degrees of freedom of the world cannot support reliable planning or compositional generalization. We prove that LeJEPA is guaranteed to recover a linear representation of latent variables from nonlinear observations, a property known as linear identifiability. What kind of world and learning objective admit this guarantee? We consider World Models with three properties (independent, stationary, and additive-noise transitions) that generate positive pairs, and study representations trained with the LeJEPA objective. Our main result: if the latent distribution is Gaussian, the representation provably achieves linear identifiability. The key insight comes from a spectral decomposition of the representation with respect to the transition structure: each spectral component corresponds to a degree of nonlinearity, and higher degrees are strictly penalized by alignment, making the linear map the unique optimum. We then prove the converse: among all latent variable distributions satisfying these three properties, the Gaussian is the unique one that leads to linear identifiability. The guarantee degrades gracefully when the two objectives are only approximately satisfied, with an explicit bound validated empirically. Our theory turns an empirically successful recipe into a mathematical guarantee, providing the foundation for building World Models that provably recover the structure of the world.
When Does Sequential Detection Collapse to a Scalar? A Necessary and Sufficient Characterisation
PRAKUL S HIREMATH ⋅ PeerAhammad M Bagawan ⋅ Sahil Bhekane
For $K=2$ regimes, scalar thresholding of the Bayesian posterior is Bayes-optimal. For $K \geq 3$, the posterior evolves on a $(K-1)$-dimensional simplex $\Delta^{K-1}$, yet virtually all deployed detection systems reduce it to a scalar score without theoretical justification. We resolve this question completely. We introduce **Extended Decision Sufficiency (EDS)**—three linear constraints on emission densities and transition dynamics (Rank-One Emissions, Markov Factorisation, Normal Factorisation), verifiable in $O(K^2)$ operations from model parameters alone—and prove the following exact characterisation: **Main result.** EDS holds if and only if the Bayes-optimal stopping value function is constant on every level set of a scalar statistic $\phi : \Delta^{K-1} \to \mathbb{R}$. Sufficiency follows from algebraic closure of the Bellman operator $\mathcal{B}$ under EDS. Necessity is proved by a coupling-based fixed-point contradiction: for each failure mode we exhibit an explicit belief pair with equal $\phi$-values but a strictly positive total-variation gap on predictive densities, which propagates to a value-function gap via the Lipschitz bound on $\mathcal{B}$ and the $\rho$-contraction of $\mathcal{T}$, without assuming any structure on $V$ beyond continuity. Three consequences are sharp and quantitative: * **(i)** When EDS fails, *every* scalar rule—regardless of architecture, capacity, or training data—incurs a strictly positive, computable, algorithm-independent lead-time loss. * **(ii)** When EDS holds, optimal expected lead time is given in closed form by a hitting-time functional of a scalar Markov chain. * **(iii)** Under approximate EDS with deviation $\eta$, performance degrades at rate $O(\eta/(1-\rho))$, making $\eta$ a practical model-diagnostic with direct performance implications. Across 72 synthetic configurations and the CICIDS2017 benchmark, predicted and observed lead times agree within 0.4 steps (MAE). A controlled experiment with matched parameter counts confirms that the lead-time advantage is structural, not a regularisation artefact, thereby validating the theory as the sole causal mechanism.
When Further Realization Is Unnecessary: Amortized Reasoning for Long-Horizon LLM Agents
Rongzheng Wang ⋅ Jiakai Li ⋅ Renzhong Wang ⋅ Rongwei Wang ⋅ Yihong Huang ⋅ Dingyuan Rao ⋅ Jielei Wang ⋅ Shuang Liang ⋅ Ke Qin
Long-horizon LLM agents solve interactive tasks by repeatedly choosing the next action from the current state. At many steps, a compact decision is sufficient to specify the action, yet the base agent still emits a full realization in its native output format (e.g., code or action text). Once the compact decision determines the executable content of the step, subsequent generation mostly adds formatting and elaboration, yielding redundant realization. This separation is reflected inside the decoder: compact decisions become locally recoverable in middle layer representations before full realizations, and attention over its token span becomes concentrated when the decision is sufficiently supported as the current step output. Building on this separation, we propose MIRA, a training-free inference method for Model-Internal Reasoning Amortization. MIRA amortizes realization through two complementary stages: Consensus-Guided Decision retrieves successful reference steps in the middle layer representation space to propose a compact decision candidate. Attention-Guided Commitment evaluates its token span using token confidence and attention concentration. The candidate is used as the step output only when both signals support it for the current state; otherwise, the agent falls back to the base agent's full realization path. Experiments on long-horizon interactive agents show that MIRA improves task performance while reducing redundant realization, yielding a 47.5\% reduction in output tokens and a 24.7\% reduction in inference time across tasks and agent formats.
Standard language model evaluation assigns scores to single predicted answers, rewarding high-confidence responses regardless of how residual probability mass is distributed over alternative options. This creates a systematic pressure toward overconfident guessing: under accuracy-based schemes, a model maximises its expected score by always committing to an answer rather than abstaining, even when its uncertainty is high. While penalty-based approaches partially address this by raising the confidence threshold for strategic guessing, they still treat all sub-threshold responses identically, ignoring a fundamental distinction in how models can express uncertainty - for example between hedging toward incorrect answers versus hedging toward "I don't know" responses. We introduce a novel evaluation metric to solve this problem of not considering a model's entire probability distribution over answer choices. Our metric naturally distinguishes between harmful overconfidence in wrong answers and uncertainty expressed through abstention, providing scores in an interpretable default range. Through theoretical analysis and illustrative examples, we demonstrate our metric offers a more nuanced and aligned evaluation paradigm that incentivises models to express genuine uncertainty rather than guessing. We then adapt 12 existing evaluation benchmarks to our metric's variants and measure performance on six language models, showing that for half of the tested benchmarks scores are negative across all tested models, indicating significant tendencies towards hallucination.
When Integral Meets Decomposition: A Signal-Level Self-Supervised Feature Decompose Paradigm for Multi-Modal Image Fusion
Zeyu Wang ⋅ Jiayu Wang ⋅ Haiyu Song ⋅ Haoran Duan
Multimodal image fusion (MMIF) aims to integrate complementary information from different modalities into a high-quality fused image and support downstream tasks. Recently, feature decomposition has become an important paradigm by separating source images into common and modality-specific unique features. However, existing methods lack clear supervision because ground-truth (GT) decomposition feature maps are unavailable. They usually combine multiple image-level metrics as losses, which are inherently incomplete and may conflict since each pixel couples attributes such as texture, edge, and contour. To address this, we propose a 1D signal-level self-supervised feature decomposition paradigm. Our core insight is to reformulate feature decomposition from unclear 2D image-level supervision into an integral-driven 1D signal-level optimization problem. This objective-level reformulation uses the 1D signal form to compute the integral constraint. The decomposer is optimized by the integral area between common and original signals, enabling more stable optimization with a clear optimization objective. Our model follows a two-stage SSL framework. Stage I designs dual pretext tasks for integral-driven decomposition at the signal level and structure-preserving reconstruction at the image level. Stage II fuses unique features and combines them with common features to reconstruct the fused image. Experiments on representative MMIF tasks show state-of-the-art(SOTA) performance. Our code will publicly available.
When Less is More: The LLM Scaling Paradox in Context Compression
Ruishan Guo ⋅ Yibing Liu ⋅ Guoxin Ma ⋅ Yan Wang ⋅ Yueyang Zhang ⋅ Long Xia ⋅ Kecheng Chen ⋅ Zhiyuan Sun ⋅ Daiting Shi
Scaling up model parameters has long been a prevalent training paradigm driven by the assumption that larger models yield superior generation capabilities. However, under lossy context compression in a compressor--decoder setup, we find a \textbf{\textit{Size-Fidelity Paradox}}: increasing compressor size can lessen the faithfulness of reconstructed contexts though reconstruction error decreases. Across 27 compressor setups spanning model families, scales, and compression rates, we coin this paradox arising from two dominant factors: 1) \textit{knowledge overwriting}: larger models increasingly replace source facts with their own prior beliefs, \textit{e.g.}, ``the white strawberry`` $\to$ ``the red strawberry``; and 2) \textit{semantic drift}: larger models tend to paraphrase or restructure content instead of reproducing it verbatim, \textit{e.g.}, ``Alice hit Bob`` $\to$ ``Bob hit Alice``. Interestingly, this paradox persists across varied settings, with mid-sized compressors often outperforming larger ones in faithful recovery. By analyzing the compressed memory via embedding geometry and reconstruction determinacy, we further reveal that compressors tend to organize memory across broader semantic subspaces, yielding more ambiguous representations prone to overwriting, drift, and weakened recovery. These findings complement existing evaluations of context compression and expose a breakdown of scaling laws when the objective shifts from plausible generation to faithful preservation.
When Metropolis and Hastings Meet Bradley and Terry: Exact MCMC From Preference Voting
Ariel Smogorghevski ⋅ Nir Rosenfeld ⋅ Yaniv Romano
conditioned on desired semantic properties is an emerging challenge in modern generative modeling. Metropolis-Hastings (MH) provides a principled route to conditional sampling, but requires access to exact pointwise target-density evaluations, which are not available in generative settings. Meanwhile, pairwise comparisons by humans or model ``judges'' are highly accessible and have proved valuable across diverse applications. We introduce Pref-MH, a general exact MH sampler for judge-induced conditional distributions using only stochastic binary pairwise comparisons. Our key observation is that the MH unnormalized density ratio matches the preference odds of BT choice models. The central challenge is that while MH requires precise ratio computation, BT judges provide only sampled binary feedback. To this end, we develop a valid accept/reject rule whose resulting Markov chain provably converges to the target distribution. We further show that, for a fixed proposal kernel and budget, Pref-MH is optimal in the Peskun-Tierney sense among this class of exact reversible acceptance rules. Experiments on text generation with LLM judges and image generation with VLM judges demonstrate that Pref-MH provides a practical and flexible approach to conditional sampling from comparative feedback.
When Noise Meets Long-Tail: Feature-Threshold Dual Calibration for Robust Pseudo-Labeling
Ping Guo ⋅ Zhiqi Huang ⋅ Xinran Li
Pseudo-labeling has become a cornerstone of learning from unlabeled data in semantic segmentation. Yet its effectiveness drops sharply in real-world scenarios where strong imaging noise and long-tailed class distributions occur together. We trace this failure to a vicious cycle of pseudo-label degradation. Imaging noise entangles foreground and background features, lowering prediction confidence across all classes, while long-tailed distributions leave tail classes with far fewer training samples and inherently lower confidence. Under fixed high-threshold filtering, these tail-class predictions are systematically filtered out, so they receive no supervision from unlabeled data and thus features keep degrading in subsequent iterations. Critically, noise and long-tail are not independent obstacles but mutually amplifying ones, and addressing either alone is insufficient. To break this cycle, we propose FTC-Seg, a Feature-Threshold dual-Calibration framework built on a standard teacher-student framework. At the feature level, Orthogonal Prototype Reconstruction (OPR) uses a set of learnable orthogonal prototypes to residually purify pixel-wise features, widening the margin between weak foreground targets and noisy backgrounds. At the threshold level, Adaptive Threshold Calibration (ATC) dynamically adjusts class-specific thresholds based on learning difficulty and prediction-distribution bias, rescuing low-confidence pseudo-labels of tail classes from systematic exclusion. Extensive experiments on four public benchmarks spanning three distinct noise modalities show that FTC-Seg achieves strong performance against state-of-the-art methods, with particularly substantial gains on tail classes. Our results establish that jointly calibrating features and thresholds is essential for robust pseudo-labeling under compounded noise and class imbalance.
When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy
Xiaofeng Tan ⋅ Jun Liu ⋅ Bin-Bin Gao ⋅ Yuanting Fan ⋅ Xi Jiang ⋅ Chengjie Wang ⋅ Hongsong Wang ⋅ Feng Zheng
RLHF is widely used to align flow-matching text-to-image models with human preferences, but often leads to severe diversity collapse after fine-tuning. In RL, diversity is often assumed to correlate with policy entropy, motivating entropy regularization. However, we show this intuition breaks in flow models: policy entropy remains constant, even while perceptual diversity collapses. We explain this mismatch both theoretically and empirically: the constant entropy arises from the fixed, pre-defined noise schedule, while the diversity collapse is driven by the mode-seeking nature of policy gradients. As a result, policy entropy fails to prevent the model from converging to a narrow high-reward region in the perceptual space. To this end, we introduce perceptual entropy that captures diversity in a perceptual space and maintains the property of standard entropy. Building upon this insight, we propose two entropy-regularized strategies, Perceptual Entropy Constraint and Perceptual Constraints on Generation Space, to preserve perceptual diversity and improve the quality. Experiments across two base models, neural and rule-based rewards, and three perceptual spaces demonstrate consistent gains in the quality-diversity trade-off; PEC achieves the best overall score of 0.734 (vs.\ baseline's 0.366); a complementary setting of PEC further reaches a diversity average of 0.989 (vs.\ baseline's 0.047). Code will be released.
When Sanitization Becomes the Trigger: Defense-Triggered Backdoor Attacks
Zhiguo Yang ⋅ Ruotian Liu ⋅ Peipei Xu ⋅ Wenjie Ruan
Backdoor defenses are widely regarded as key to secure third-party model deployment. However, we are the first to show that, in standard backdoor sanitization pipelines, the defense process itself can be exploited and turned into a conditional trigger. We propose **Defense-Triggered Backdoor (DTB): an attacker implants both a normally active decoy backdoor and a dormant hidden backdoor, so that model sanitization suppresses the decoy backdoor while activating the hidden backdoor, thereby making the defense process the trigger condition for the hidden backdoor. DTB uses bi-level optimization to approximate the shared sanitization effect of mainstream backdoor defenses, while constraining the hidden backdoor to activate only after the decoy backdoor is sufficiently suppressed. Experiments on multiple datasets, different networks, and 14 representative defenses show that DTB can stably trigger the hidden backdoor, and this phenomenon remains significant after multi-round compositional sanitization. Our findings reveal a potential threat in existing backdoor defense pipelines and suggest that post-sanitization safety cannot be judged solely by whether the original backdoor has been eliminated.
When Symbol Names Should Not Matter: A Logistic Theory of Fresh-Symbol Classification
Wenjie Guan ⋅ Jelena Bradic
Template tasks have emerged as a clean testbed for asking whether transformers reason with abstract symbols rather than concrete token names. We study the fixed-label classification version of this problem, where train and test examples share latent templates but may use disjoint vocabularies. Unlike next-token prediction, the model need not emit unseen symbols; it must learn a decision rule invariant to symbol renaming. We analyze regularized kernel logistic classification in the transformer-kernel regime. Our main result decomposes the learned predictor into an ideal template-level classifier and a finite-sample perturbation caused by accidental token overlaps in the training data. We encode these overlaps by a colored collision graph and prove high-probability margin-transfer guarantees for fresh-symbol classification. This perspective extends template-based analyses to logistic classification and refines scalar diversity conditions: vocabulary size controls the average rate of collisions, but collision geometry controls whether the ideal classification margin is preserved. More broadly, the same perturbation framework applies to abstraction-augmented inputs, yielding a general margin-versus-collision criterion for identifying when prompting strategies improve fresh-symbol generalization. Synthetic template experiments illustrate the predicted roles of regularization, sample size, and transformer-kernel structure.
When Transcriptomic Foundation Models Scale: Domain-Focused Pretraining for Drug Development in Immunology and Inflammation
Karim El Kanbi ⋅ Yannis Cattan ⋅ Aziz Fouché ⋅ Charlotte Claye ⋅ Pierre Marschall ⋅ Julien Duquesne
Recent work has reported two failure modes of transcriptomic foundation models: validation loss plateaus beyond 100M parameters, and consistent underperformance against simple linear baselines on clinically relevant tasks. We show that both findings invert under a specific pretraining recipe. We introduce EVA-RNA, a transformer pretrained on 545k human and mouse samples spanning bulk RNA-seq, microarray, and pseudobulked single-cell data, scoped to immunology and inflammation, one of the largest therapeutic areas in clinical development. EVA-RNA exhibits clean power-law scaling from 7M to 300M parameters, with no plateau emerging within our scale range. On a benchmark co-designed with immunologists and drug development experts, EVA-RNA outperforms existing foundation models on every task category, spanning drug discovery, preclinical-to-clinical translation and patient stratification. Mechanistically, EVA-RNA learns species-invariant representations in which orthologous genes progressively align across layers without supervision. We interpret these results as evidence that scope and data composition are sufficient, with conventional architecture and knowledge-informed gene embeddings, to achieve clinical utility and scaling in I&I. We also release EVA-RNA 60M model weights to support continued investigation.
Where Does Warm-Up Come From? Adaptive Scheduling for Norm-Constrained Optimizers
Artem Riabinin ⋅ Andrey Veprikov ⋅ Arman Bolatov ⋅ Martin Takac ⋅ Aleksandr Beznosikov
We study adaptive learning rate scheduling for norm-constrained optimizers (e.g., Muon and Lion). We introduce a generalized smoothness assumption under which local curvature decreases with the suboptimality gap and empirically verify that this behavior holds along optimization trajectories. Under this assumption, we establish convergence guarantees under an appropriate choice of learning rate, for which warm-up followed by decay arises naturally from the proof rather than being imposed heuristically. Building on this theory, we develop a practical learning rate scheduler that relies only on standard hyperparameters and adapts the warm-up duration automatically at the beginning of training. We evaluate this method on large language model pretraining with LLaMA architectures and show that our adaptive warm-up selection consistently outperforms or at least matches the best manually tuned warm-up schedules across all considered setups, without additional hyperparameter search. Our source code is available at https://anonymous.4open.science/r/warmup.
Where Reusable Computation Becomes Detectable: Solution-Frame Path Triage for Modular-Arithmetic Grokking
Ruian Lei ⋅ Ruoxi Jiang ⋅ Marley M Vellasco ⋅ M. Tanveer ⋅ Zenglin Xu
Mechanistic studies of grokking often begin after an endpoint circuit has been identified. We study the preceding retrospective triage problem: given a solved final checkpoint and a long saved trajectory, which path and window should be inspected first? We propose solution-frame path triage. The final checkpoint fixes a solution frame; earlier checkpoints are then scored path by path. The Geometric Coherence Score (GCS) measures coherence in this procedure: it asks whether neighboring inputs in the final frame undergo similar local Jacobian transformations through the attention, MLP, or full block measurement maps. Centered Kernel Alignment (CKA) and $L_2$ distance to the final path state provide the complementary proximity scores. Across 104 modular arithmetic Transformer runs, GCS reveals a repeatable path schedule that simple norm and dimensionality summaries miss: attention coherence rises near grokking and often turns over, while MLP maps diversify and later reconverge. After selection, held-out checks show that the resulting windows are enriched for accuracy and algebraic changes, align with Fourier progress on modular addition, and identify attention-to-MLP interface states with high replacement cost. The method returns a short list of path/window hypotheses for endpoint analysis. It is retrospective by design: the final checkpoint is part of the procedure.
Who Called? V33DA: A Physically Verified Multimodal Benchmark for Vocal Attribution in Zebra Finch Groups
Maris Basha ⋅ Yuhang Wang ⋅ Xiaoran Chen ⋅ Longbiao Cheng ⋅ Luca N Yapura ⋅ Anja T Zai ⋅ Mathieu Salzmann ⋅ Richard Hahnloser
Deciding which member of a group produced a vocalization and where that vocalization was produced are central to studying social communication but are rarely evaluated with physically verified ground truth. Existing resources either localize sounds without identifying the caller, or recognize individuals from isolated vocalizations without reliance on the 3d candidate geometry. We close this gap with V33DA, a multimodal benchmark for spatial vocal attribution in social zebra finches: it comprises 33,625 verified events coupling 5-channel audio with three camera views, calibrated 3D pose for every visible candidate, and per-bird radio telemetry across 10 individuals and three experiments. Caller identity is verified from an on-body accelerometer-derived vibration signal; this channel is withheld from benchmark models and used only by an oracle ceiling. We also provide V33DA++ as an auxiliary extension with verified overlapping-call events and $\pm 2$s context windows, intended to support future source-separation and vocal-activity-detection studies. V33DA separates familiar-individual recognition from transferable spatial attribution. We evaluate candidate-conditioned caller attribution, with 3D source localization as a secondary diagnostic, under three regimes: session-disjoint testing, held-out-experiment transfer, and leave-one-individual-out evaluation. A broad set of reference methods exposes a consistent gap between in-domain accuracy and transfer under identity shift: methods that can exploit familiar vocal identity perform well on known callers but degrade sharply on unseen ones, whereas candidate-aware methods preserving explicit spatial reasoning transfer substantially better. The accelerometer oracle stays near-perfect across regimes, confirming the internal consistency of the accelerometer-verified labels and quantifying the ceiling available when the withheld verification channel is observed. We release V33DA with fixed attribution/localization protocols and V33DA++ with overlap and long-context tasks, together with code for the evaluated methods.
Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs
Jihoo Jung ⋅ Youngjoon Jang ⋅ Joon Son Chung
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal binding is achieved in AVLLMs. Specifically, we identify emergent symbolic trimodal binding mechanisms in AVLLMs that utilize modality-specific symbolic variables. By encoding auditory and visual components into symbolic variables-capturing temporal utterance sequences and spatial entity coordinates, respectively-the model establishes cross-modal linking within this abstract space. Crucially, we reveal that when trimodal binding fails, the breakdown predominantly stems from misaligned audio-visual connections. To overcome this bottleneck, we introduce an audio-visual prompting method utilizing an off-the-shelf Active Speaker Detection (ASD) model. By simply overlaying visual bounding boxes on active speakers, this training-free approach yields immediate performance gains across three conversation-centric benchmarks. Moreover, lightweight fine-tuning of fewer than 300 steps on these ASD-prompted-videos extends these gains to three general AV benchmarks, suggesting the generalizability of our method.
Who Wrote This Paper? Autonomous Scientific Discovery for 3DGS Research
Seemandhar Jain ⋅ Manmohan Chandraker
Reproducing a published 3DGS paper takes weeks of expert engineering before any extension can be proposed. Prior autonomous-research systems read papers, draft hypotheses, and write code, but suffer from limited template length, executability, and traceability. Specifically, none has been demonstrated end-to-end in a live computer-vision subfield with intense activity. GS-Scientist is an autonomous research system for 3D Gaussian Splatting that addresses these limitations. Given a research prompt, a target-paper URL, or no input at all, it reproduces a published baseline, proposes an extension, trains and ablates across Mip-NeRF~360 and LLFF, writes the manuscript, and binds every reported number to a logged training run. Five design choices distinguish it: (i)candidates mutate plugins from a verified-reproduction backbone matched to published baselines; (ii)a cascade-weighted Elo tournament across eight islands scores on measured GPU PSNR with verbal-gradient feedback; (iii)a standing falsifier subtracts Elo on failure, integrating stress-test survival into selection; (iv)every reported number is bound to a logged training run, and the run cannot terminate while any claim is unbacked; (v)a seven-category 3DGS taxonomy wraps \texttt{EVOLVE-BLOCK} markers in domain-aligned scaffolds transferable across 3DGS repositories. GS-Scientist produces manuscripts that beat their reproduced target by $+0.18$ to $+3.78$\,dB PSNR, including $+3.78$\,dB on astrophotographic nebula rendering, for which no prior 3DGS method exists. The research cycle compresses from months of graduate-student time to days of mostly autonomous compute. Code, all manuscripts, experiment ledgers, and the cross-run paper store will be released.
Why Geometric Continuity Emerges in Deep Neural Networks: Residual Connections and Rotational Symmetry Breaking
Kyungwon Jeong ⋅ Won-Gi Paeng ⋅ Honggyo Suh
Weight matrices in deep networks exhibit geometric continuity---principal singular vectors of adjacent layers point in similar directions. While this property has been widely observed, its origin remains unexplained. Through experiments on toy MLPs and small transformers, we identify two mechanisms: residual connections create cross-layer gradient coherence that aligns weight updates across layers, and symmetry-breaking nonlinearities constrain all layers to a shared coordinate frame, preventing the rotation drift that would otherwise destabilize weight structure. Crucially, a nonlinear but rotation-preserving activation fails to retain continuity, isolating symmetry breaking---not nonlinearity itself---as the active ingredient. Activation and normalization play distinct roles: activation concentrates continuity in the leading singular direction, while normalization distributes it across multiple directions. In transformers, continuity is \emph{projection-specific}: Q, K, Gate, and Up (which read from the residual stream) develop input-space ($\mathbf{v}_1$) continuity; O and Down (which write to it) develop output-space ($\mathbf{u}_1$) continuity; V alone, lacking an adjacent nonlinearity, develops only low continuity.
Latent action models (LAMs) aim to learn action-like representations from unlabeled videos by compressing frame-to-frame changes. The frames of in-the-wild videos, however, contain not only the agent's own state but exogenous state such as background clutter. Since the exogenous state introduces changes unrelated to actions, it hinders reliable latent action learning. This paper investigates this problem analytically by extending a linear LAM framework to explicitly model exogenous state. Our analysis reveals two insights: (1) minimizing the standard reconstruction objective produces latent actions that encode exogenous information of future observation; and (2) learning in a representation space that focuses on endogenous components is a key to mitigating the interference of noise. We further show that previously proposed auxiliary objectives, such as action-supervision, provably encourage latent actions to be consistent across exogenous states. These findings are validated through experiments on both linear and nonlinear LAMs, providing a unified theoretical analysis of how exogenous state hinders latent action learning and why common remedies work.
Why Transformers Struggle with Distribution-Independent In-Context Learning
Omar Naim ⋅ Jerome Bolte ⋅ Nicholas Asher
Transformer models can adapt to new tasks from input--output examples at inference time, a capability known as in-context learning (ICL). A common interpretation is that transformers implement implicit learning algorithms in their forward pass, using the context to infer the task and predict the query. We test this interpretation beyond the training distribution in controlled regression settings where classical estimators can generalize exactly. Transformers achieve near-perfect in-distribution ICL, but fail sharply under scale and support shifts, exhibiting a cliff-shaped collapse that is absent from least-squares and kernel baselines. This failure is robust across polynomial degrees and a range of training interventions. We identify two architectural mechanisms that help explain this behavior. First, final normalization imposes a readout-side scale constraint, causing predictions to saturate at fixed boundary values. Removing this constraint eliminates hard saturation, but does not restore reliable ICL under coefficient or label scaling. Second, attention can become a context bottleneck: under distribution shift, softmax attention may concentrate on a small part of the prompt, limiting the context-dependent adaptation required for ICL. These results show that strong in-distribution ICL does not imply distribution-independent ICL, and suggest caution when using ICL instead of explicit retraining or adaptation under distribution shift.
WorldSpeech: A Multilingual Speech Corpus from Around the World
Antonis Asonitis ⋅ Luca Lanzendörfer ⋅ Frédéric Berdoz ⋅ Roger Wattenhofer
Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sharply for most languages due to limited publicly available aligned data. To this end, we introduce WorldSpeech, a 24kHz multilingual speech corpus comprising 65k hours of aligned audio-transcript data across 76 languages, collected from diverse public sources including parliamentary proceedings, international broadcasts, and public-domain audiobooks. For 37 languages, WorldSpeech provides more than 200 hours of aligned speech, with 28 exceeding 500 hours and 24 surpassing 1k hours. Fine-tuning existing ASR models on WorldSpeech results in an average relative Word-Error-Rate reduction of 63.5% across 11 typologically diverse languages.
XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
Dasol Choi ⋅ Eugenia Kim ⋅ Jaewon Noh ⋅ Sang Seo ⋅ Eunmi Kim ⋅ Myunggyo Oh ⋅ Yunjin Park ⋅ Brigitta J Kartono ⋅ Josef Pichlmeier ⋅ Helena Berndt ⋅ Sai Krishna Mendu ⋅ Glenn J Tungka ⋅ Özlem Gökçe ⋅ Suresh Gehlot ⋅ Katherine Pratt ⋅ Amanda Minnich ⋅ Haon Park
Current LLM safety benchmarks are predominantly English-centric and often rely on translation, failing to capture country-specific harms. Moreover, they rarely evaluate a model's ability to detect culturally embedded sensitivities as distinct from universal harms. We introduce XL-SafetyBench. a suite of 5,500 test cases across 10 country-language pairs, comprising a Jailbreak Benchmark of country-grounded adversarial prompts and a Cultural Benchmark where local sensitivities are embedded within innocuous requests. Each item is constructed via a multi-stage pipeline that combines LLM-assisted discovery, automated validation gates, and dual independent native-speaker annotators per country. To distinguish principled refusal from comprehension failure, we evaluate Attack Success Rate (ASR) alongside two complementary metrics we introduce: Neutral-Safe Rate (NSR) and Cultural Sensitivity Rate (CSR). Evaluating 10 frontier and 27 local LLMs reveals two key findings. First, jailbreak robustness and cultural awareness do not show a coupled relationship among frontier models, so a composite safety score obscures per-axis variation. Second, local models exhibit a near-linear ASR--NSR trade-off (r = -0.81), indicating that their apparent safety reflects generation failure rather than genuine alignment. XL-SafetyBench enables more nuanced, cross-cultural safety evaluation in the multilingual era.
Z-Cache: Accelerating Diffusion Transformers via Self-Reflection
Zegang Cheng ⋅ Zhikai Wang ⋅ Jiacheng Liu ⋅ Xiaobing Tu ⋅ Chang Zhou ⋅ Junjie Chen ⋅ Yue Ma ⋅ Zhiyuan Ma ⋅ Jinkui Ren ⋅ Likai Zou ⋅ Linfeng Zhang
Diffusion transformers have become the most powerful models for visual generation, but still suffer from massive computation costs. To solve this problem, feature caching has been proposed to cache the features of diffusion models in the previous computation steps and then reuse them in the following caching steps, which brings significant acceleration but also degradation in generation quality. To address this problem, this paper proposes Z-cache as a feature caching method that can maintain high-quality generation through self-reflection. Concretely, we observe that the error from feature caching tends to be sharply reduced after each full computation. Based on this observation, Z-Cache is designed to first predict the features in the future caching steps and then perform a full computation. After that, Z-Cache returns to the caching steps and re-predicts them based on the previous and the current computation steps, which brings correction in features. Experiments demonstrate that with Z-Cache, diffusion transformers achieve comparable generation quality to the original model but with faster inference speed, for instance, 6.22X acceleration on FLUX-dev for text-to-image generation. Our codes will be released on GitHub.