Skip to yearly menu bar Skip to main content


Session

Paris Poster Session 5

Paris Poster Hall
Fri 11 Dec 9:30 p.m. AEDT — 11:30 p.m. AEDT
Abstract:
Chat is not available.


$\varphi$TD: Distributional Reinforcement Learning using Characteristic Functions

Thomas Mousseau ⋅ Tyler Kastner ⋅ Davide Baldelli ⋅ Amir-massoud Farahmand

Distributional Reinforcement Learning methods aim to learn the entire distribution of returns, yet current algorithmic approaches are limited to a narrow class of distributions. We propose $\varphi$TD, a method to overcome this limitation through computing the loss in the frequency domain. We demonstrate that this further unifies quantile and categorical approaches to distributional RL, as well as enabling the learning of distributions that were previously impossible to capture, such as mixtures of continuous distributions. We empirically study $\varphi$TD across synthetic benchmarks and large-scale deep RL environments, and demonstrate that the flexibility obtained by wider parametric families often leads to improved performance.


2-Step Agent: How a Bayesian Decision Maker Learns from AI-Decision Support

Otto Nyberg ⋅ Fausto Carcassi ⋅ Davide Tugnoli ⋅ Giovanni Cinà

Predictions from ML models support human decision making in several fields, including high-stakes ones such as healthcare and the judiciary. Yet, we still lack a clear understanding of how decision makers learn from ML-based decision support (ML-DS). In this paper, we introduce a general computational framework, the 2-Step Agent, to capture this process. As a prediction from an ML model contains information about the training data, a prediction can also be used for inference. Our framework models (i) how a prediction for a new observation affects the beliefs of a rational Bayesian agent, and (ii) how this change in beliefs affects the estimation of causal effect, the downstream decision, and the subsequent outcome. In addition to the framework itself, we make three contributions. First, for the linear Gaussian setting, we derive a tractable solution for the challenging Bayesian inference problem we introduced, i.e. one in which the agent infers from an ML prediction. Second, we experimentally identify conditions under which ML-DS is beneficial. Third, we show that a single misaligned prior belief can be sufficient for ML-DS to lead to worse downstream outcomes compared to no decision support even when the ML model is well-specified and the agent is perfectly rational. Hence, even under ideal conditions, ML-DS can do more harm than good.


Accelerating Long-Context LLM Prefill via Layer-wise Progressive Token Pruning in Local Deployment

Zhongxiang Wei ⋅ Zhaohan Wang ⋅ Zhixiong Zhang ⋅ Jin Zhao ⋅ Xiaoming Fu

Long-context inference is increasingly important for privacy-sensitive local LLM deployment. In this setting, the prefill stage dominates Time-to-First-Token (TTFT), yet existing acceleration methods struggle to balance speed and quality. Sparse-attention methods accelerate only the attention module, leaving Feed-Forward Network (FFN) costs unchanged, while auxiliary-model-based compression often compromises quality under tight memory budgets. Recent layer-wise token pruning methods reduce both attention and FFN computation, but they typically rely on heuristic pruning schedules or token-recovery mechanisms that introduce decoding overhead. We propose FastPrefill, a recovery-free layer-wise token pruning system for efficient local long-prompt inference. FastPrefill employs an offline data-driven optimizer to determine layer-wise pruning schedules that maximize fidelity under a latency constraint, and executes these schedules via an algorithm–system co-design supporting head-specific sparse attention at runtime, where each head dynamically attends to its own selected key/value tokens. Evaluations on LongBench and RULER show that FastPrefill maintains comparable inference quality while achieving up to 2.13x TTFT speedup and 3.03x end-to-end speedup over state-of-the-art baselines.


ACDP: Architecture-aware Cross-Dataset Performance Predictor for NAS

Jiawen Deng ⋅ Yuqi Feng ⋅ Han Ji ⋅ Yanan Sun

Performance predictors are pivotal in accelerating evaluation in NAS. Cross-dataset predictors are trained on (dataset, architecture, performance) triplets to efficiently estimate architectural performance for unseen datasets. However, existing cross-dataset predictors rely on a naive concatenation of dataset and architecture features. This fusion ignores the architecture-aware nature, where the same dataset manifests as distinct representations under different architectures, thereby limiting generalization. To tackle this issue, we propose ACDP, an Architecture-aware Cross-Dataset Performance Predictor that dynamically adjusts dataset feature extraction based on specific architecture. Initially, ACDP develops an elastic mechanism to extract raw dataset traits, providing a configurable lever for the trade-off between efficiency and accuracy. Furthermore, ACDP leverages a hypernetwork to determine dataset feature extractor parameters conditioned on architectural features, transforming raw dataset traits into tailored representations. ACDP excels in both specialization and generalization on six unseen datasets. Notably, in NAS-Bench-201, ACDP achieves 0.758/0.741 Kendall's Tau on CIFAR-10/100, surpassing single-dataset SOTAs trained with over 100 target samples. In MobileNetV3, ACDP discovers architectures outperforming cross-dataset SOTAs, notably gaining +2.04% accuracy on Aircraft. Code: https://anonymous.4open.science/r/ACDP/

We derive closed-form ODEs for the learning dynamics of a linear recurrent policy network trained with REINFORCE on sparse-reward tasks in a high-dimensional teacher–student setting. Under alignment assumptions between student and teacher networks, the dynamics reduce to a finite system of coupled order-parameter ODEs whose reward terms are governed by trajectory-level orthant probabilities of a structured correlated Gaussian process. The theory predicts that recurrence can accelerate escape from sparse-reward plateaus by accumulating task-relevant memory, and that each episode length admits an optimal recurrent scale that minimizes learning time. Our results provide a quantitative theory of how recurrence, and sparse reward interact during learning.


AIRA 2: Overcoming Bottlenecks in AI Research Agents

Karen Hambardzumyan ⋅ Nicolas Baldwin ⋅ Edan Toledo ⋅ RISHI HAZRA ⋅ Michael Kuchnik ⋅ Bassel Al Omari ⋅ Thomas Foster ⋅ Anton Protopopov ⋅ Jean-Christophe Gagnon-Audet ⋅ Ishita Mediratta ⋅ Kelvin Niu ⋅ Michael Shvartsman ⋅ Alisia Lupidi ⋅ Alexis Audran-Reiss ⋅ Parth Pathak ⋅ Tatiana Shavrina ⋅ Despoina Magka ⋅ Hela Momand ⋅ Derek Dunfield ⋅ Nicola Cancedda ⋅ Pontus Lars Erik Saito Stenetorp ⋅ Carole-Jean Wu ⋅ Jakob Foerster ⋅ Yoram Bachrach ⋅ Martin Josifoski

Existing research has identified three structural performance bottlenecks in AI research agents: (1) synchronous single-GPU execution constrains sample throughput, limiting the benefit of search; (2) a generalization gap where validation-based selection causes overfitting and performance to degrade over extended search horizons; and (3) the limited capability of fixed, single-turn LLM operators imposes a ceiling on search performance. We introduce AIRA$_2$, which addresses these bottlenecks through three architectural choices: an asynchronous multi-GPU worker pool that increases experiment throughput linearly; a Hidden Consistent Evaluation protocol that delivers a reliable evaluation signal; and ReAct agents that dynamically scope their actions and debug interactively. On MLE-bench-30, AIRA$_2$ achieves a mean Percentile Rank of 81.5% at 24 hours and 83.1% at 72 hours, outperforming the strongest baseline, which achieves 72.7%. On AIRS-Bench, AIRA$_2$ exceeds human state-of-the-art on 6 out of 20 diverse research tasks. Ablations confirm that each architectural component is necessary, that performance follows a predictable scaling law that transfers across LLM backbones, and that the "overfitting" reported in prior work was driven by evaluation noise rather than true data memorization.


Amortized Structured Stochastic Variational Inference for Gaussian Process Latent Variable Models

Maksym Tretiakov ⋅ Sarah Filippi ⋅ Vincent Fortuin ⋅ Ruth Misener ⋅ Ruby Sedgwick ⋅ James Odgers

Many machine learning methods aim to approximate the lower-dimensional manifold on which the data lives. A desirable feature of such methods is that they should capture the epistemic uncertainty of this learned manifold. One model that achieves this is the Gaussian Process Latent Variable Model, in which a Gaussian Process (GP) mapping from the latent space provides an estimate of the uncertainty of the manifold. However, the effectiveness of this uncertainty estimation is limited by the mean-field variational approximation between the GP inducing points and the latent variables. In this work, we use Amortized Structured Stochastic Variational Inference to allow the variational posterior for the latent space to be conditionally dependent on the value of the inducing points. We demonstrate that this more flexible variational posterior improves several metrics related to the reconstruction of points on the data manifold.

Abstract: Mode collapse --- the failure to capture one or more modes when targetting a multimodal distribution --- is a central challenge in modern variational inference. In this work, we provide a mathematical analysis of annealing-based strategies for mitigating mode collapse in a tractable setting: learning a Gaussian mixture, where mode collapse is known to arise. Leveraging a low-dimensional summary statistics description, we precisely characterize the interplay between the initial temperature and the annealing rate, and derive a sharp formula for the probability of mode collapse. Our analysis shows that an appropriately chosen annealing scheme can robustly prevent mode collapse. Finally, we present numerical evidence that these theoretical trade-offs qualitatively extend to neural network–based models, RealNVP normalizing flows, providing guidance for designing annealing strategies mitigating mode collapse in practical variational inference pipelines.


APM: Evaluating Style Personalization in LLMs with Arbitrary Preference Mappings

Philipp Spohn ⋅ Leander Girrbach ⋅ Zeynep Akata

Typical LLM responses tend to follow a default style, even though users often have distinct preferences regarding tone, verbosity, and formality that they do not explicitly state in their prompts. Evaluating whether personalization methods can adapt to these implicit preferences is challenging, since users typically provide prompts rather than reference responses, style preferences are not factually verifiable, and reference-free LLM judges may conflate personalization with general response quality. To address these challenges, we introduce the Arbitrary Preference Mapping (APM) benchmark, which decouples user attributes (e.g. enthusiastic) from response principles (e.g. persuasive) via a hidden, randomized mapping $\mathbf{C}$ that maps user attributes to preferences about response traits. Because $\mathbf{C}$ carries no semantic content and is resampled across runs, models cannot exploit stereotypical associations and must infer preferences from conversation history. Using this unbiased evaluation methodology, we adapt retrieval-augmented, prompt-optimization, and routing personalization methods and evaluate them on Llama-3.1-8B and Qwen-3.5-27B. Our results show that routing is the most reliable approach, while RAG benefits only from the stronger base LLM, and soft prompt optimization fails to improve significantly over a non-personalized baseline. Our extensive evaluation reveals that in this realistic setting, personalization remains challenging, but our adapted methods show promise.


Are Easier or Harder Examples Better? Rethinking Data Selection for Reward Models and Preference Optimization

Kevin Christian Wibisono ⋅ Aya Ismail ⋅ Pedro O. Pinheiro ⋅ Yixin Wang ⋅ Kyunghyun Cho ⋅ Nataša Tagasovska ⋅ Rajesh Ranganath

Despite being crucial for effective LLM alignment, data selection remains understudied. Prior work on reward model (RM) training and policy optimization (e.g., DPO, GRPO) identifies *example difficulty*, the reward gap between chosen and rejected responses, as a key factor, but findings conflict: some favor easier examples with larger gaps, others harder ones. To isolate difficulty from confounders, we *assume access to a reference RM* and systematically study data selection across RM, DPO, and GRPO training. When difficulty is measured via the reference RM, *easier pairs consistently outperform harder ones*, especially for smaller base models: using only the top 20% easiest examples often matches or exceeds full-dataset performance while cutting post-training costs $5\times$. However, *this advantage hinges on reward estimation quality*. As the difficulty signal is corrupted by noise or estimated with a weak proxy RM, the easy-example advantage shrinks and can reverse. A signal-to-noise analysis explains why: larger-gap examples yield more reliable gradient directions, an advantage that weakens with noisier reward estimates. These results suggest that *conflicting findings in prior work partly stem from differences in reward reliability and signal mismatch*.


A Separation Principle for Cooperative Multi-Agent Reinforcement Learning

Lucia Pezzetti ⋅ Nicolas Lanzetti ⋅ Antonio Terpin ⋅ Florian Dorfler ⋅ Giorgia Ramponi

We study cooperative multi-agent reinforcement learning (MARL) problems in which agents evolve under decoupled noisy dynamics but are coupled via a population objective. A canonical example is a mobility operator that orchestrates a fleet of vehicles in an urban environment to meet customer demand. The stochastic and time-varying nature of these problems calls for learning approaches, yet existing MARL algorithms do not scale: learning assignment and routing jointly over the full state-action space becomes prohibitive as the number of agents grows. By studying the problem as a control problem over probability measures, we prove a separation principle: the population-level cost-to-go function is upper-bounded by an optimal transport problem whose transportation cost is the target-conditioned single-agent cost-to-go function. This result simplifies the multi-agent problem to learning a single-agent policy and coordinating the agents via optimal transport. Thus, we introduce SALT (Separation-based Assignment and Learning via optimal Transport) a family of algorithms that learn a target-conditioned single-agent policy using any standard RL algorithm, and then coordinate the agents via optimal transport. Because learning occurs at the single-agent level, SALT scales to fleet sizes where MARL baselines become intractable and transfers without retraining to fleet sizes unseen during training. In a vehicle-delivery experiment in south Manhattan’s road network, where MARL baselines are computationally prohibitive, SALT improves travel times by up to 20% over a shortest-path routing heuristic and transfers without retraining to different fleet sizes. In stochastic grid worlds, SALT reduces average travel times by more than 40% relative to MARL baselines and scales to population sizes where these methods fail.


A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning

Idan Mehalel ⋅ Ilan Doron-Arad ⋅ Elchanan Mossel

Autoregressive generation lies in the heart of the mechanism of large language models. It can be viewed as the repeated application of a next-token generator: starting from an input string (prompt), the generator is applied for $M$ steps, and the last generated token is taken as the final output. [Joshi et al., 2025] proposed a PAC model for studying the learnability of the input-output maps arising from this process. We develop an online analogue of this framework, focusing on the mistake bound of learning the final output induced by an unknown next-token generator. We distinguish between two forms of feedback. In the End-to-End model, after each round the learner observes only the final token produced after M autoregressive steps. In the Chain-of-Thought model, the learner is additionally shown the entire $M$-step trajectory. Our goal is to understand how the optimal mistake bound depends on the generation horizon $M$, and to what extent observing intermediate tokens can reduce this dependence. Our main results show that the online theory of autoregressive learning exhibits a qualitative picture analogous to the statistical one found by [Hanneke et al., 2026], but with a different scale of dependence on the generation horizon. In the End-to-End model, we prove a taxonomy of possible mistake-bound growth rates in the generation horizon $M$: subject to mild regularity conditions, every rate between constant and logarithmic can arise. We also show that this logarithmic ceiling is unavoidable for the online setting, in the sense that every class of finite Littlestone dimension exhibits at most logarithmic dependence on $M$. In the Chain-of-Thought model, the parallel is even sharper: as in the statistical setting, access to the full generated trajectory eliminates the dependence on $M$ altogether. We also analyze autoregressive linear threshold classes. For autoregressive linear thresholds in \(\mathbb{R}^d\), we prove that the optimal mistake bound is \(\Theta(d^2)\) under both End-to-End and Chain-of-Thought feedback, and for any generation length $M$. The methods developed for this analysis also yield new lower bounds in the statistical setting. Along the way, our results resolve several questions left open by [Joshi et al., 2025]. In particular, we show that even for classes of finite Littlestone dimension, the End-to-End statistical sample complexity can depend on the generation length M.


A Topological Sorting Criterion for Random Causal Directed Acyclic Graphs

Alexander Reisach ⋅ Antoine Chambaz ⋅ Gilles Blanchard ⋅ Sebastian Weichwald

Random directed acyclic graphs (DAGs) based on imposing an order on Erdős–Rényi and scale free random graphs are widely used for evaluating causal discovery algorithms. We show that in such DAGs, the set of nodes reachable via open paths, termed relatives, increases monotonically along the causal order. We assess the prevalence of this pattern numerically, and demonstrate that it can be exploited for causal order recovery via sorting by the estimated number of relatives. We note that many simulations in the literature feature settings where this yields an excellent proxy for the causal order, and show that a strict increase of relatives along the causal order leads to a singular Markov equivalence class. We propose sampling time-series DAGs as a possible alternative and discuss implications for causal discovery algorithms and their evaluation on synthetic data.

Neural audio codecs compress waveforms into compact discrete tokens that underpin speech language models, real-time communication, and large-scale audio storage. Almost every dominant design, including residual vector quantization, finite scalar quantization, and single-codebook variants, follows the VQ-VAE template by partitioning the encoder latent through a learned codebook or a fixed scalar grid. We ask whether this partition is necessary. We introduce GS-Codec, a neural speech codec whose bottleneck is a parametric signal decomposition rather than a quantizer. We adapt Gaussian splatting from 3D scene reconstruction to one-dimensional latents. An inner optimization loop fits each encoder segment as a weighted sum of 1D Gaussian primitives. The decoder then reconstructs the waveform from the rendered sum. To avoid the cost of this iterative inner loop at inference time, we additionally train a lightweight GS Predictor Net that regresses the primitive parameters in a single forward pass. The encoder and decoder are trained end-to-end through the inner loop, with no quantizer anywhere in the training pipeline: the bottleneck is the decomposition itself, and scalar quantization is applied only post-training to the fitted parameters. Rather than relying on discrete codebook stages for bitrate control, our representation exposes a smooth rate-quality tradeoff: a single trained checkpoint supports continuous post-training bitrate control by varying the number of primitives and the per-parameter bit depth, with no retraining required. GS-Codec outperforms well-established open-source codecs such as EnCodec and DAC on speaker similarity (SIM), intelligibility (STOI), and perceptual quality (UTMOS) at comparable bitrates, while achieving comparable semantic performance.

Auditing large language models (LLMs) is increasingly urgent as these systems are deployed in high-stakes settings, yet existing evaluation practices are ill-suited to meet auditing requirements. Directly repurposing standard evaluation tools can yield incomplete or misleading conclusions, e.g. overstating robustness when evidence comes from static prompts rather than adaptive, real-world interactions. This position paper argues that LLM audits must instead generate dynamic, context-sensitive, budget-aware, and reliable evidence. To support this position, we analyze how each of these principles can be operationalized through a four-component framework: Auditing Scope, Interactor, Evaluator, and Output. We highlight design requirements, limitations and research directions, demonstrating how high-level principles can be translated into concrete, actionable, evidence-based procedures.

In many real-world settings, machine learning models and interactive systems have access to both structured knowledge, e.g., knowledge graphs or tables, and unstructured content, e.g., natural language documents. Yet, most rely on either. Semi-Structured Knowledge Bases (SKBs) bridge this gap by linking unstructured content to nodes within structured data. In this work, we present Autofocus-Retriever (AF-Retriever), a modular framework for SKB-based, multi-hop question answering. It combines structural and textual retrieval through novel integration steps and optimizations, achieving the best zero- and one-shot results across all three STaRK QA benchmarks, which span diverse domains and evaluation metrics. AF-Retriever’s average first-hit rate surpasses the second-best method by 32.1%. Its performance is driven by (1) leveraging exchangeable large language models (LLMs) to extract entity attributes and relational constraints for both parsing and reranking the top- answers, (2) vector similarity search for ranking both extracted entities and final answers, (3) a novel incremental scope expansion procedure that prepares for the reranking on a configurable amount of suitable candidates that fulfill the given constraints the most, and (4) a hybrid retrieval strategy that reduces error susceptibility. In summary, while constantly adjusting the focus like an optical autofocus, AF-Retriever delivers a configurable amount of answer candidates in four constraint-driven retrieval steps, which are then supplemented and ranked through four additional processing steps. An ablation study and a detailed error analysis, including a comparison of three different LLM reranking strategies, provide component-level insights that are valuable for advancing the model and for enabling researchers and users to adapt, optimize, or extend its parts. The source code is publicly available at https://github.com/kramerlab/AF-Retriever.


Batched Stochastic Linear Bandits with 1-Bit Communication Constraints

Ivan Lau ⋅ Daniel McMorrow ⋅ Kevin Jamieson ⋅ Jonathan Scarlett

We study stochastic linear bandits under a natural combination of batching and communication constraints: the time horizon is partitioned into batches of equal size $B$, and during each batch the learner sends $B$ requested arm pulls to an agent, who then observes the corresponding $B$ rewards and responds with a single bit of feedback to the learner. For each batch, the learner specifies the 1-bit quantization rule the agent uses, which may depend on all previously received bits but not on any past rewards directly. This setting addresses a significant yet unexplored ``middle ground'' between previous models having per-round quantization only $\textit{or}$ total bit budgets only. We establish a minimax lower bound showing that $\Omega(B\min \lbrace d,\log\lvert \mathcal{A} \rvert \rbrace)$ regret is unavoidable due to the 1‑bit communication bottleneck, even in the absence of noise. Combined with standard statistical limits, this yields a general lower bound of $\widetilde{\Omega}(B\min\lbrace d,\log\lvert \mathcal{A} \rvert \rbrace + \sqrt{dT \min \lbrace d,\log\lvert \mathcal{A} \rvert\rbrace })$. We develop two phased‑elimination algorithms based on $G$-optimal designs and 1‑bit mean estimation. The first achieves $\widetilde{O}(dB + d\sqrt{T})$ regret, matching the lower bound up to logarithmic factors when $\lvert \mathcal{A} \rvert = \exp(\Omega(d))$, and the second incorporates a safe‑arm identification and warm‑start procedure to obtain $\widetilde{O}(B\log\lvert \mathcal{A} \rvert + d^{3/2}\sqrt{B} + \sqrt{dT\log\lvert \mathcal{A} \rvert})$ regret, which is near‑optimal in broad scaling regimes of $(\lvert \mathcal{A} \rvert, B, d,T)$. Together, our results demonstrate that a single bit of feedback per batch suffices to preserve nearly order‑optimal regret across a wide range of settings, even in the case of large batch sizes such as $\Theta(\sqrt{T})$.


Beyond Bag-of-Words: Diagnosing Compositional Binding Failures in Vision-Language Models

Cristian Sbrolli ⋅ Toshihiko Yamasaki ⋅ Matteo Matteucci

Modern vision-language models struggle with basic compositional reasoning, failing to bind attributes to objects or relations to their referents. Existing benchmarks either rely on noisy real images that conflate confounding visual variables with the reasoning failure, or use simplistic synthetic scenes lacking the realism modern VLMs are tuned for. We introduce Auto-Comp, a fully automated, concept-driven pipeline that bridges this gap by generating photorealistic compositional benchmarks at scale. Its core innovation is a parallel A/B construction: for each concept, the pipeline emits a Minimal sample (template caption, isolated objects on a white background) and a Contextual sample (LLM-rewritten caption, objects embedded in a realistic scene), isolating core binding ability from visio-linguistic complexity. We instantiate four task families spanning the two canonical axes of compositional binding: Color and Shape-Color (attribute binding), and Position and Relative Size (relational binding). We evaluate over 25 VLMs spanning CLIP, SigLIP, hard-negative-trained, and frontier generative models. The findings are consistent across architectures and scales: every model exhibits a large Swap-vs-Confusion gap, with low-entropy distractors (e.g., repeated objects or colors) exposing failures beyond the known bag-of-words limitations. We further uncover a task-dependent trade-off: visio-linguistic context aids relational reasoning but hinders attribute binding through visual clutter. We publicly release the pipeline and benchmarks.


Beyond Chamfer Distance: Granular Order-aware Evaluation Metric For Online Mapping

Chouaib Bencheikh Lehocine ⋅ Adam Lilja ⋅ Junsheng Fu ⋅ Lars Hammarstrand

Online map estimation is a crucial component of autonomous driving systems that reduces the reliance on costly high-definition maps. State-of-the-art (SOTA) methods commonly predict map elements as ordered sequences of points that form polylines and polygons. The evaluation of these methods relies predominantly on mean average precision (mAP) based on thresholded Chamfer distance (CD). This framework lacks sensitivity to point ordering and provides limited granularity in assessing geometric quality, making it difficult to distinguish which methods truly excel over others. In this work, we address these limitations on two fronts. For the single-instance similarity measure, we introduce sequence optimal sub-pattern assignment (SOSPA), an order-aware metric that enables fine-grained evaluation of individual geometries while satisfying all metric axioms. For the multi-instance evaluation framework, we propose polyline localisation and detection (PLD), a soft metric that jointly captures detection quality and geometric accuracy, replacing the hard thresholding of mAP with a principled soft assignment. Through evaluations on nuScenes, we demonstrate that PLD effectively ranks SOTA online mapping methods (MapTRv2, StreamMapNet, MapTracker) while providing a decomposed error analysis. This analysis identifies detection capability as the dominant bottleneck in current methods, revealing a performance trend that mAP fails to capture. Code for evaluation using our metrics will be released.

JEPAs often regularize one-view embeddings toward an isotropic Gaussian, implicitly baking Euclidean symmetry into the representation. We show that this is not merely a benign default. For a known structured downstream geometry $H\succ0$, the minimax and maximum-entropy covariance under a Hamiltonian energy budget is $(c/d)H^{-1}$, and Euclidean isotropy incurs a closed-form price of isotropy. More importantly, when the downstream geometry is unknown, no geometry-independent fixed marginal target is canonical: every fixed covariance shape can be maximally misaligned for some structured geometry. We further show that even oracle one-view marginals do not identify the JEPA view-to-view predictive coupling. These results suggest that the structural bias in JEPAs should enter the cross-view coupling rather than a fixed encoder marginal. We instantiate this principle with \textbf{HamJEPA}, which encodes each view as a phase-space state $(q,p)$ and predicts view-to-view transitions with a learned Hamiltonian leapfrog map, while non-isotropic scale and spectral floors prevent collapse. In a deliberately headless token protocol, HamJEPA improves over SIGReg on CIFAR-100 by $+4.89$ kNN@20 and $+3.52$ linear-probe points at 30 epochs, and by $+6.45$ kNN@20 and $+10.64$ linear-probe points at 80 epochs, while a matched MLP predictor ablation shows that the symplectic coupling is the ingredient driving the neighborhood-geometry gain. On ImageNet-100, HamJEPA-$q$ improves by $+4.82$ kNN@20 and $+7.52$ linear-probe points at 45 epochs.


BlockFormer: Transformer-based inference from interaction maps

Eloïse Touron ⋅ Pedro Rodrigues ⋅ Julyan Arbel ⋅ Nelle Varoquaux ⋅ Michael Arbel

Inference from interaction maps, such as centromere identification from genome-wide chromosome conformation capture techniques --notably Hi-C-- can be formulated as a generic inverse problem: infer a set of parameters given a map summarizing pairwise interactions between entities through blocks of variable numbers and sizes. In this work, we introduce a data-driven approach that leverages shared structure between these maps, such as global alignment between localized patterns, while handling the variability in number and size of entities arising in real-world data. Our approach relies on a transformer architecture capable of handling such variability and a custom simulator to generate abundant, yet computationally cheap synthetic data for training. Applied to the problem of centromere localization, the method accurately recover their genomic positions across a wide range of species of various genome sizes.


Boosting Inference with Guided Reasoning: Stochastic Exploration for Recursive Models

Andrew Corbett ⋅ Archit Sood ⋅ Anna Tzatzopoulou ⋅ Sai-Aakash Ramesh ⋅ Tim Dodwell

Recent work on recursive architectures has shown that tiny neural networks can be surprisingly powerful on structured reasoning tasks. The trick is to model reasoning trajectories with a latent dynamical system. We argue that the inference-time behaviour of these architectures is best understood as approximate inference over latent reasoning trajectories, with deterministic recursion as the one-particle, zero-noise limit. We make this view operational through guided stochastic exploration: stochastic perturbations of the reasoning dynamics propose neighbouring trajectories, and the model's existing early-stopping head reweights them online. The framework yields three label-free diagnostics: local stability, guide alignment, and cloud-token entropy. These predict, from inference traces alone, whether the procedure will help and which of its outputs to trust. On Sudoku-Extreme it lifts exact-solve accuracy from $85.9\%$ to $98.0\%$ without retraining; on Maze-Hard the diagnostics flag a misaligned guide, as validation performance later confirms. The same machinery thus characterises both when recursive reasoning has room to improve at the trajectory level and when the model's internal guide can recover it.


Calibrated Safe Policy Improvement for Continuous Offline Reinforcement Learning

Federico Bianchi ⋅ Alessandro Farinelli ⋅ Alberto Castellini

Safe Policy Improvement (SPI) aims to improve a baseline policy offline using fixed data while avoiding performance degradation with high probability. Existing SPI methods work in discrete domains but do not extend to continuous control. We propose Calibrated Safe Policy Improvement (Cal-SPI), a Monte Carlo Tree Search-based continuous SPI method that replaces count-based support with calibrated model trust. Cal-SPI learns a deep ensemble dynamics model and calibrates ensemble disagreement on held-out data through a Wilks-style tolerance construction, turning raw uncertainty into a statistical certificate of one-step model error. This certificate defines a hard gate within model-based tree search. We provide a theoretical analysis showing that calibrated model trust and a baseline-relative switching criterion yield a conditional PAC-style improvement bound for gate-certified root decisions. Experiments across three MuJoCo continuous domains show that Cal-SPI avoids performance degradation in low-data regimes and improves over the baseline once calibrated model trust becomes reliable.

Many prediction problems involve a hidden predictive state: a latent task, environment, rule, or data-generating mechanism that changes the query-optimal act. For a fixed query, recovering the full hidden predictive state is often stronger than necessary. The relevant target is the Bayes-act distinction among hidden predictive states that preserves oracle-gain adaptation. This question is central for in-context learning (ICL): a prompt may reveal part of an unseen task or rule, but an ICL system needs only to recover the prediction-relevant quotient state label for the query. We formalize this target through the canonical predictive quotient (CPQ), a problem-relative Bayes-act quotient of hidden predictive states. The theory characterizes the value of adaptation, the quotient that preserves full oracle gain, the sharp zero/positive-gain boundary, and a Bayes barrier showing why context-only prediction cannot recover a positive oracle gain at the risk level. We give two exact examples. A finite-state example makes the quotient structure, sharp positivity boundary, and a strict-loss mechanism explicit. A linear-Gaussian prior-family witness gives closed forms for a collapsed Gaussian baseline and a state-aware oracle, a restricted zero-gain boundary, and an exact prediction-level disagreement identity for the restricted gain. Together, the results give a decision-theoretic target for ICL: not full latent complexity, but the prediction-relevant quotient state label. We support the theory with exact checks, realized bridge experiments over prediction-relevant quotient state labels, and external validity tests on released pretrained in-context predictors.


Causal-VLM: Dense Causal Captioning in Videos

Asmar Nadeem ⋅ Mahrukh Awan ⋅ Muhammad Awais ⋅ Robert Dawes ⋅ Adrian Hilton ⋅ Armin Mustafa

Dense video captioning describes multiple events in long videos. However, each event is predicted independently, ignoring the causal relationships with the other events in the video. These causal relationships are important for human-like descriptions of long videos. Our experiments show that exisiting vision-language models fail at predicting causal relationships in long-form videos. No prior work jointly generates captions and event timestamps and predicts causal relationships over multi-event sequences in long-form videos. We address this through two contributions. First, we create a benchmark dataset with 21.6K videos and 85K events by using a large language model to generate annotations from dense captions in YouCook2 and ActivityNet. We validate this benchmark dataset through multimodal alignment, vision-language models, and human judgment. Second, we propose a novel architecture for dense causal captioning that not only generates captions but also learns causal relationships. This novel architecture not only predicts causality effectively (0.73 F1 ActivityNet, 0.64 F1 YouCook2) but also improves the vision-language model's core capabilities: caption quality increases by +6.3 CIDEr on ActivityNet and +7.4 on YouCook2, while temporal grounding becomes accurate as causal predictions force precise event boundary localization.


Certifiably Optimal Robust Angular Synchronization

Daniel Barath ⋅ Keisuke Tateno ⋅ Marc Pollefeys ⋅ Federico Tombari

Rotation averaging on the circle, equivalently robust angular synchronization, underlies a wide range of geometric estimation problems, for example, gravity-aligned Structure-from-Motion, multi-way point-cloud registration, and in-plane alignment in single-particle cryo-electron microscopy. We present the first algorithm that certifiably solves the robust maximum-consensus formulation of this problem to global optimality. Our approach discretises the circle into uniform bins, formulates a pairwise Markov random field and solves it via branch-and-bound with two key novelties: an ICM-guided label-pruning rule that collapses the branching factor to a handful of candidates, and a node-constrained relaxation bound orders of magnitude tighter than the standard per-edge bound. Together, they reduce the search from millions of nodes to a few hundred on typical instances. On all tested problems, the algorithm terminates with a proof of global optimality, typically in a few seconds. Across synthetic graphs and three real-world domains -- image-based SfM on IMC 2023/2024, multi-way LiDAR point-cloud registration on KITTI and NSS, and cryo-EM angular synchronization on EMPIAR-10166 -- the certified solution achieves sub-degree accuracy at outlier rates where other baselines fail. Our code will be made publicly available.

Pre-trained visual models have become fundamental in computer vision, but they face challenges in continual learning scenarios where data and tasks evolve over time. Low-Rank Adaptation (LoRA) offers efficient fine-tuning capabilities but remains limited for such dynamic environments. Standard LoRA cannot distinguish important subspaces, causing critical knowledge to be overwritten in sequential training. Existing approaches address this by dynamically expanding the set of LoRA adapters—either maintaining a growing pool of task-specific modules or merging new adapters into prior ones—at the cost of unbounded parameter growth or increasing inference complexity. We propose Continual Low-Rank Adaptation (C-LoRA), a method that enables a single, shared LoRA adapter to handle sequential tasks without catastrophic forgetting—without requiring any module selection or fusion at inference. The core of C-LoRA is a learnable routing matrix $\boldsymbol{\mathcal{R}}$ that explicitly controls how each rank-one subspace contributes to the weight update. This matrix is decomposed into a stability component ($\boldsymbol{\mathcal{R}} _ {\text{base}}$), which preserves knowledge from prior tasks, and a plasticity component ($\boldsymbol{\mathcal{R}} _ {\delta}$), which drives adaptation to the current task—providing direct control over the stability-plasticity trade-off. We analyze how $\boldsymbol{\mathcal{R}}$ governs gradient flow during sequential training, and demonstrate competitive performance across multiple benchmarks.


Coherence Mechanisms for Provable Self-Improvement

Mehryar Mohri ⋅ Jon Schneider ⋅ Yifan Wu

Self-improvement is a critical capability for large language models and other intelligent systems, enabling them to refine their behavior and internal consistency without external supervision. Despite its importance, prior approaches largely rely on empirical heuristics and lack formal guarantees. In this paper, we propose a principled framework for self-improvement based on the concept of coherence, which requires that a model's outputs remain consistent under task-preserving transformations of the input. We formalize this concept using projection-based mechanisms that update a baseline model to be coherent while remaining as close as possible to its original behavior. We provide rigorous theoretical guarantees that these mechanisms achieve monotonic improvement, measured by a reduction in expected Bregman divergence. Our analysis is comprehensive, covering both \emph{direct} and \emph{two-step} projection methods, and robustly extends these guarantees to non-realizable settings, empirical (finite-sample) distributions, and relaxed coherence constraints.


Computationally Efficient Replicable Learning of Parities and Applications

Moshe Noivirt ⋅ Jessica Sorrell ⋅ Eliad Tsfadia

We study the computational relationship between replicability (Impagliazzo et al. [STOC `22], Ghazi et al. [NeurIPS `21]) and other stability notions. Specifically, we focus on replicable PAC learning and its connections to differential privacy (Dwork et al. [TCC 2006]) and to the statistical query (SQ) model (Kearns [JACM `98]). Statistically, it was known that differentially private learning and replicable learning are equivalent and strictly more powerful than SQ-learning. Yet, computationally, all previously known efficient (i.e., polynomial-time) replicable learning algorithms were confined to SQ-learnable tasks or restricted distributions, in contrast to differentially private learning. Our main contribution is the first computationally efficient replicable algorithm for realizable learning of parities over arbitrary distributions, a task that is known to be hard in the SQ-model, but possible under differential privacy. This result provides the first evidence that efficient replicable learning over general distributions strictly extends efficient SQ-learning, and is closer in power to efficient differentially private learning, despite computational separations between replicability and privacy. Additionally, we leverage our parity learner to prove that, assuming $RP \neq NP$, converting replicability to pure differential privacy requires a strict loss in sample complexity. Our main building block is a new, efficient and replicable algorithm that, given a set of vectors, outputs a subspace of their linear span that covers most of them.


ConfDet: Learning Reliable Confidence for MLLM-based Detection

Xuanjie Mao ⋅ Peng Ye ⋅ Ziteng Ma ⋅ Lin Zhang ⋅ Jiakang Yuan ⋅ Haoyu Zhang ⋅ Fujun Han ⋅ Jiayuan Fan ⋅ Tao Chen

MLLM-based detection has shown promising end-to-end detection ability in recent years, but its generated detection results still lack reliable confidence for evaluation and interpretation. Since detection data itself do not provide human-annotated confidence labels, existing MLLM-based detection methods mostly obtain confidence implicitly from prompted self-assessment, generation probabilities, or external proxy scores. While these scores are available for evaluation, they are not explicitly connected with the actual detection quality. Inspired by conventional detectors that optimize confidence through classification, objectness, or quality estimation branches, we formulate confidence in MLLM-based detection as an explicit modeling target to be learned, optimized and calibrated. Specifically, we present ConfDet, a practical confidence framework that first uses discrete confidence tokens based supervised fine-tuning to enable stable and parsable confidence generation, then applies GRPO with detection and separation joint rewards to optimize confidence score separation and preserve detection quality, and finally performs condition-aware post-hoc calibration to align confidence with empirical accuracy. Experiments across different in-distribution and out-of-distribution benchmarks show that, ConfDet greatly improves AP through better confidence ranking and reduces confidence miscalibration across benchmarks, providing a reliable confidence for MLLM-based detection.


Consistency Regularised Gradient Flows for Inverse Problems

Alessio Spagnoletti ⋅ Tim Wang ⋅ O. Deniz Akyildiz ⋅ Marcelo Pereyra

Vision-Language Latent Diffusion Models (LDMs) provide powerful generative priors for inverse problems. However, existing LDM-based inverse solvers typically require a large number of neural function evaluations (NFEs) and backpropagation through large pretrained components, leading to substantial computational cost and, in some cases, degraded reconstruction quality. We propose a unified Euclidean-Wasserstein-2 gradient-flow framework that jointly performs posterior sampling and prompt optimization in the latent space through a single flow that aligns the prior and posterior with the observed data. Combined with few-step latent text-to-image models, this formulation enables low-NFE inference without backpropagation through autoencoders. Experiments across several canonical imaging inverse problems show that our method achieves state-of-the-art performance with significantly reduced computational cost.

Flow Matching enables high-quality visual generation via continuous-time dynamics, but inference remains costly due to multiple sequential function evaluations. Existing acceleration methods reduce the number of function evaluations but often introduce additional training overhead, degrade quality, or fail to account for input-dependent variability. We propose an inference-time method, CoFlow, that adaptively selects the step counts for each generation based on the prompt features. Our context-aware CoFlow is trained online with an unsupervised reward that balances efficiency and fidelity. Our method is plug-and-play, requiring no retraining of the generative model. It generalizes to image and video generation, achieving over $2.5\times$ speedup while preserving perceptual and semantic quality. We also provide theoretical insights into the link between adaptive step allocation and discretization error. The anonymized source code is available at \url{https://anonymous.4open.science/r/Contextual_Flow_Matching-9970}.


Conveyance: A Versatile Framework for Learning in Structured Class Spaces

Yasser Taha ⋅ Grégoire Montavon ⋅ Nils Körber

While machine learning (ML) architectures have evolved rapidly to account for complex data, loss functions like cross-entropy remain mostly structure-agnostic in many real-world applications. However, the "class-symmetric" nature of these standard losses fundamentally limits the ability of ML models to exploit structural relationships between classes, particularly when facing structured noise. We propose Conveyance, a new classification approach and associated loss function tailored to structured class spaces. It allows users to encode graph-like relations between classes without having to define complex joint distributions or manually tune utility matrices. Technically, our loss function operates by maximizing two separate margins over distinct class partitions, while preserving formal properties such as monotonicity and partial convexity. We demonstrate the versatility and effectiveness of our method by applying it to hierarchical classification, ordinal regression, and multiple instance learning. Across these tasks, Conveyance either matches or exceeds the performance of specialized baselines, thereby offering a unified solution for structured class spaces.


Corruptions of Supervised Learning Problems: Typology and Mitigations

Laura Iacovissi ⋅ Nan Lu ⋅ Robert Williamson

Corruption is notoriously widespread in data collection. Despite extensive research, the existing literature predominantly focuses on specific settings and learning scenarios, lacking a unified view of corruption modelization and mitigation. In this work, we develop a general theory of corruption, which incorporates all modifications to a supervised learning problem, including changes in model class and loss. Focusing on changes to the underlying probability distributions via Markov kernels, our approach leads to three novel opportunities. First, it enables the construction of a novel, provably exhaustive corruption framework, distinguishing among different corruption types. This serves to unify existing models and establish a consistent nomenclature. Second, it facilitates a systematic analysis of corruption consequences on learning tasks, by considering Bayes risks in the clean and corrupted scenarios. Notably, while label corruptions affect only the loss function, attribute corruptions additionally influence the hypothesis class. Third, building upon these results, we investigate mitigations for various corruption types. We expand existing loss-correction methods for label corruption to handle dependent corruption types. Our findings highlight the necessity to generalize this classical corruption-corrected learning framework to a new paradigm with weaker requirements to encompass more corruption types. We provide such a paradigm as well as loss correction formulas in the attribute and joint corruption cases.


Counterfactual Maps: What They Are and How to Find Them

Awa Khouna ⋅ Julien Ferry ⋅ Thibaut Vidal

Counterfactual explanations are a central tool in interpretable machine learning, yet computing them exactly for complex models remains challenging. For tree ensembles, predictions are piecewise constant over a large collection of axis-aligned hyperrectangles, implying that an optimal counterfactual for a point corresponds to its projection onto the nearest rectangle with an alternative label under a chosen metric. Existing methods largely overlook this geometric structure, relying either on heuristics with no optimality guarantees or on mixed-integer programming formulations that do not scale to interactive use. In this work, we revisit counterfactual generation through the lens of nearest-region search and introduce counterfactual maps, a global representation of recourse for tree ensembles. Leveraging the fact that any tree ensemble can be compressed into an equivalent partition of labeled hyperrectangles, we cast counterfactual search as the problem of identifying the generalized Voronoi cell associated with the nearest rectangle of an alternative label. This leads to an exact, amortized algorithm based on volumetric k-dimensional (KD) trees, which performs branch-and-bound nearest-region queries with explicit optimality certificates and sublinear average query time after a one-time preprocessing phase. Our experimental analyses across several real datasets from high-stakes application domains show that this approach delivers globally optimal counterfactual explanations with millisecond-level latency, achieving query times that are orders of magnitude faster than existing exact, cold-start optimization methods.


CRAFT: Conflict-Resolved Aggregation for Federated Training

Ziqi Wang ⋅ Qiang Liu ⋅ Nils Thuerey

The aggregation of conflicting client updates remains a fundamental bottleneck in federated learning (FL) over heterogeneous data distributions. Naive averaging, as used in FedAvg, often leads to destructive interference in which the global model improves on average but deteriorates significantly for specific clients. In this work, we propose CRAFT (Conflict-Resolved Aggregation for Federated Training), a new aggregation framework that treats the global update as a geometric correction problem. We formulate aggregation as finding the update closest to a \emph{reference direction} while satisfying \emph{conflict-free constraints}, ensuring non-negative alignment with the updates of participating clients. We derive a closed-form expression for the constrained optimization problem, avoiding the computational overhead of iterative solvers. Furthermore, we use a layer-wise adaptation to address conflicts at varying feature granularities. We provide a theoretical analysis showing that CRAFT promotes a common-descent structure and mitigates destructive interference through its projection geometry. Extensive experiments on non-IID and imbalanced benchmarks demonstrate that CRAFT significantly reduces performance disparity while maintaining competitive global accuracy against state-of-the-art baselines.


CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves

Amirreza Mohseni ⋅ Mona Mohammadi ⋅ Morteza Saghafian ⋅ Naser Talebizadeh Sardari

We introduce CurveBench, a benchmark for hierarchical topological reasoning from visual input. CurveBench consists of $\textbf{756 images}$ of pairwise non-intersecting Jordan curves across easy, polygonal, topographic-inspired, maze-like, and dense counting configurations. Each image is annotated with a rooted tree encoding the containment relations between planar regions. We formulate the task as structured prediction: given an image, a model must recover the full rooted containment tree induced by the curves. Despite the visual simplicity of the task, the strongest evaluated model, $\texttt{Gemini-3.1-Pro-Preview}$, achieves only $\textbf{71.1\\%}$ tree-generation accuracy on CurveBench-Easy and $\textbf{19.1\\%}$ on CurveBench-Hard. We further demonstrate benchmark utility through RLVR-style fine-tuning of open-weight vision-language models. Our trained $\texttt{Qwen3-VL-8B}$ model improves over $\texttt{Qwen3-VL-8B-Thinking}$ from $\textbf{2.8\\%}$ to $\textbf{33.3\\%}$ tree-generation accuracy on CurveBench-Easy, exceeding $\texttt{GPT-5.4}$ and $\texttt{Claude-Opus-4.5}$ under our evaluation protocol. The remaining gap, especially on CurveBench-Hard, shows that exact topology-aware visual reasoning remains far from solved.


Data-Driven Covariate Selection for Nonparametric and Cycle-Agnostic Causal Effect Estimation

Ana L Vicente ⋅ Gijs van Seeventer ⋅ Saber Salehkaleybar

Estimating causal effects from observational data requires identifying valid adjustment sets. This task is especially challenging in realistic settings where latent confounding and feedback loops are present. Existing approaches typically assume acyclicity or rely on global causal structure learning, limiting applicability and computational efficiency. In this work, we study a local, data-driven method for covariate selection based on conditional independence information. While this method is known to be sound and complete in acyclic causal models, its validity in the presence of cycles has remained unclear. Our main contribution is to show that these guarantees extend to cyclic causal models. In particular, our result relies on the invariance of conditional independence assertions under $\sigma$-acyclification. These findings establish a unified, cycle-agnostic perspective on covariate selection and causal effect estimation, showing that the method applies across cyclic and acyclic settings without modification. Empirically, we validate this on extensive synthetic data, showing reliable performance in cyclic causal models.


Dataset Collections: Challenges of Large-Scale Data Aggregation in 3D Medical Image Datasets

Yannick Kirchhoff ⋅ Saikat Roy ⋅ Elisa Stegmeier ⋅ Hamideh Haghiri ⋅ Constantin Ulrich ⋅ Maximilian R. Rokuss ⋅ Tassilo Wald ⋅ Karol Gotkowski ⋅ Benjamin Hamm ⋅ Raphael Stock ⋅ Nico Disch ⋅ Michael Baumgartner ⋅ Dimitrios Bounias ⋅ Anand Deshpande ⋅ Stefan Dinkelacker ⋅ Stefan Dvoretskii ⋅ Katharina Eckstein ⋅ Selen Erkan ⋅ Jessica Kächele ⋅ Kim-Celine Kahl ⋅ Balint Kovacs ⋅ Lucas Kulla ⋅ Moritz Langenberg ⋅ Philipp Schader ⋅ Stephen Schaumann ⋅ Darya Trofimova ⋅ Shuhan Xiao ⋅ Sebastian Ziegler ⋅ David Zimmerer ⋅ Marco Nolden ⋅ Fabian Isensee ⋅ Klaus Maier-Hein ⋅ Ralf Floca

The development of deep learning-based medical AI depends on high-quality annotated data. Yet, dataset creation remains constrained by the need for expert radiologist annotations, limiting curated datasets to relatively small cohorts. To overcome this limitation, recent efforts increasingly aggregate disparate public datasets into large-scale dataset collections. However, combining datasets at scale often obscures data provenance, introducing significant risks of reuse propagation, unintended leakage, bias amplification and distorted data distributions. In this work, we study these problems on a massive collection of 760k 3D radiological images across 1069 public datasets and provide practical resources to address it. Firstly, we formalize dataset collections as a distinct paradigm in medical image analysis and characterize their structural challenges, including reuse propagation, duplication, and provenance fragmentation. Secondly, we perform one of the first large-scale empirical studies of duplication on medical image collections and demonstrate extensive exact and near-duplicate reuse across public datasets via perceptual hashing and manual tracing on our 760k 3D radiological images across 1069 public datasets, with 118k near or exact duplicates. We make these hashes publicly available, thereby providing a standardized reference for identifying potential overlaps and shared provenance for researchers. Thirdly, owing to the insufficient robustness of hashing in isolation, we introduce MAP (Metadata for Aggregation and Provenance), a lightweight, portable metadata profile for documenting provenance, known overlaps, and deduplication decisions in dataset collections. Our contributions enable dataset collections to shift from coarse aggregations to reliable and provenance-aware data scaling mechanisms in medical image analysis. A webtool supporting our work is made available here:


Decentralized AI Governance Must Decouple Policy Processing from Capability Enforcement

Hasan Kassem ⋅ Christoforos Anagnostopoulos ⋅ Alejandro Aristizabal ⋅ Spyridon Bakas ⋅ Orion G Banks ⋅ Omar Benjelloun ⋅ Tian Cai ⋅ Sergen Cansiz ⋅ Davide Chicco ⋅ Patrick Foley ⋅ Showkot Hossain ⋅ Taeho Jung ⋅ Peter Kairouz ⋅ Yanchen liu ⋅ Marco Lorenzi ⋅ Peter Mattson ⋅ Ann K Novakowski ⋅ Michael O'Connor ⋅ Holger Roth ⋅ Adrish Sannyasi ⋅ Charalampos Savvaidis ⋅ Micah Sheller ⋅ David Solooki ⋅ Dimitris Stripelis ⋅ Renato Umeton ⋅ Marc Vesin ⋅ Ray Wang ⋅ Wenbin Zhang ⋅ Mic Bowman ⋅ Alexandros Karargyris

In this position paper, we argue that successful decentralized AI needs trustworthy, adaptive, and portable governance at scale. The ML community must adopt an architecture that cleanly separates \emph{policy processing} (the evaluation of evidence against governance requirements) from \emph{capability enforcement} (the gating of access to protected digital assets). We argue that the prevailing approach, in which each organization manages bespoke policy logic, custom-built for its infrastructure, produces fragmentation that undermines interoperability, transparency, and auditability of federated systems. Drawing on a survey of governance mechanisms in federated learning and data collaboration frameworks and on a concrete reference architecture built around community-driven policy objects, we argue that this decoupling is technically feasible and architecturally necessary and supports a broad class of decentralized AI governance settings. We further suggest that the AI community should invest in open, standardized policy abstractions rather than proliferate siloed governance solutions. We address counterarguments concerning the costs of standardization and the feasibility of universal policy languages.

Jailbreak attacks that frame harmful queries as professional requests can bypass LLM safety training, but we do not know which ingredients in a professional-framing pipeline actually matter. We introduce a pre-specified 2×2×2 factorial-ablation design for jailbreak decomposition, separating access-driving from depth-driving components, and apply it to one concrete pipeline family, Structured Three-stage Framing (STF): a Persona Adoption turn (S1) that establishes professional identity, a Moral Justification turn (S2) that appeals to harm prevention, and a Linguistic Substitution (S3) of framework terminology for colloquial language. Tested on N = 5,000 trials across three closed frontier models (~31,900 total trials across nine LLMs), the decomposition isolates S1 (Persona Adoption) as the dominant contributor to acceptance (OR = 6.5), S3 (Linguistic Substitution) as the dominant contributor to depth (β = 0.91), and S2 (Moral Justification) showing no detectable positive effect at the stated equivalence bounds (pooled OR = 0.83; TOST at [0.67, 1.50]) — a component ordering we write as S1 > S3 > S2 (Persona > Linguistic Substitution > Moral Justification). This ordering holds across three full-rank factorial replications (defensive system prompt, cybersecurity, social engineering; N ≈ 5,000 each) and nine additional robustness checks. Two defensive targets follow: detect persona claims to block access; restrict output depth when responses use framework terminology. The method is portable beyond STF; the empirical result is specific to one pipeline tested primarily through provider APIs, with open-weight corroborations on local inference (§3).


DeFlowCritic: Dense Latent Reward Alignment for Text-to-Image Flow Matching Models

Zeeshan Khan ⋅ Xin Yu ⋅ Shizhe Chen ⋅ Cordelia Schmid

Online reinforcement learning (RL) has emerged as a powerful paradigm for aligning text-to-image flow matching models with complex user intent. However, existing methods typically rely on sparse rewards computed only from final decoded images. This delays credit assignment across the denoising trajectory and makes training computationally expensive. Moreover, simply combining multiple rewards does not reliably produce a Pareto trade-off, where gains in compositional accuracy often degrade aesthetic quality. In this work, we propose \textbf{DeFlowCritic}, an efficient online RL framework based on Dense Latent Reward Alignment. Unlike prior work, our method provides dense supervision on intermediate latents, enabling early-stage guidance for the denoising trajectory. To achieve this, we introduce an internal critic that evaluates noisy latents using features extracted directly from the diffusion model, a representation inherently better suited for latent-space reward estimation than standard image encoders. We construct a high-quality reward-modeling dataset with continuous scores for critic training, jointly capturing compositional correctness and aesthetic preference. The resulting dense reward signals enable efficient online RL training and inference scaling. Experiments show that our method achieves a superior balance between prompt faithfulness and visual quality on GenEval2 and T2I-CompBench benchmarks.


DepthGraft: Structural Regularization through Hierarchical Cross-Layer KV Reconstruction

Xiaohan Qin ⋅ Xiangdong Zhang ⋅ Yu Wang ⋅ Huaijin Wu ⋅ Zhuo Xia ⋅ Yebin Yang ⋅ Junchi Yan

Scaling depth is a key route to stronger large language models, but Pre-LN architectures often suffer from depth-wise information dilution, limiting the effective use of deep layers. While cross-layer connectivity has proven effective at mitigating this problem, existing designs usually introduce additional computation or memory access. We revisit cross-layer KV sharing as an efficient form of such connectivity. Beyond cache compression, each KV reuse edge creates a direct gradient path from deep layers to source KV projections, turning KV sharing into an implicit structural regularizer for KV representation geometry. We provide a systematic analysis of this mechanism, showing that proper sharing topologies sharpen key retrieval directions while enriching the value content subspace, whereas existing adjacent or global-anchor designs do not fully exploit this effect. Motivated by this, we propose DepthGraft, a hierarchical KV-sharing architecture that reconstructs deep-layer KV from stride-aligned shallow--middle source pairs. DepthGraft distributes non-local gradient flow across source layers, while preserving prefill early completion and reducing KV cache by roughly one third. Experiments on dense and MoE models from 0.5B to 65B parameters show consistent downstream improvements under both GQA and MLA. Notably, on the 16B MoE model, DepthGraft achieves a 7.0% absolute improvement on MMLU-Pro, along with gains of 6.6% on CMMLU and 6.4% on C3. These results validate the potential of DepthGraft as a scalable design for large-scale pretraining.


Differentiable Knapsack and Top-k Operators via Dynamic Programming

Germain Vivier-Ardisson ⋅ Michael E Sander ⋅ Axel Parmentier ⋅ Mathieu Blondel

Knapsack and Top-$k$ operators are useful for selecting discrete subsets of variables. However, their integration into neural networks is challenging as they are piecewise constant, yielding gradients that are zero almost everywhere. In this paper, we propose a unified framework casting these operators as dynamic programs, and derive differentiable relaxations by smoothing the underlying recursions. On the algorithmic side, we develop efficient parallel algorithms supporting both deterministic and stochastic forward passes, and vector-Jacobian products for the backward pass. On the theoretical side, we prove that Shannon entropy is the unique separable binary regularization choice for the local DP smoothing that yields permutation-equivariant operators, and characterize regularizers inducing sparse selections. On the experimental side, we demonstrate our framework on a benchmark on learning to predict Knapsack solutions, an extension of discrete VAEs, and a constrained dynamic assortment RL problem.


Differentiable Nonlinear Model Predictive Control

Jonathan Frey ⋅ Katrin Baumgärtner ⋅ Gianluca Frison ⋅ Dirk Reinhardt ⋅ Jasper Hoffmann ⋅ Leonard Fichtner ⋅ Joschka Boedecker ⋅ Sebastien Gros ⋅ Moritz Diehl

The efficient computation of parametric solution sensitivities is a key challenge in the integration of learning-enhanced methods with nonlinear model predictive control (MPC), as their availability is crucial for many learning algorithms. This paper discusses the computation of solution sensitivities of general nonlinear programs (NLPs) using the implicit function theorem (IFT) and smoothed optimality conditions treated in interior-point methods (IPM). We detail sensitivity computation within a sequential quadratic programming (SQP) method which employs an IPM for the quadratic subproblems. Previous works presented in the machine learning community are limited to convex or unconstrained formulations, or lack an implementation for efficient sensitivity evaluation. The publication is accompanied by an efficient open-source implementation within the acados framework, providing both forward and adjoint sensitivities for general optimal control problems, achieving speedups exceeding 3x over the state-of-the-art solvers mpc.pytorch and cvxpygen.


DiLaDiff: Distilled Latent-augmented Diffusion for Language Modeling

Jean-Marie Lemercier ⋅ Tomas Geffner ⋅ Morteza Mardani ⋅ Karsten Kreis ⋅ Arash Vahdat ⋅ Ante Jukić

Diffusion language models intrinsically fail to capture correlations between decoded tokens, which leads to a harsh trade-off between sampling quality and throughput. To solve this issue, we propose DiLaDiff, a variant of masked diffusion language models with three components: (1) a continuous latent space with semantic capabilities, learned by an auto-encoder fine-tuned from an existing masked diffusion language model; (2) a latent diffusion model learning the prior over the encoder distribution; (3) a consistency model distilling the learned prior into a few-step latent generative model. We show that, even without distillation, our latent-guided diffusion model outperforms the masked diffusion baseline while significantly accelerating inference. Consistency distillation further lowers the computational overhead of continuous diffusion, such that the latent is generated in negligible time compared to discrete decoding.

PDE-free latent diffusion models can synthesize plausible turbulence fields, but visually realistic snapshots do not guarantee statistically reliable long-horizon rollouts. In CNF-based turbulence diffusion, we observe that long rollouts can drift in latent transition laws, temporal memory, and transport-sensitive diagnostics even when the decoder remains spatially expressive. We address this failure mode with a DNS-calibrated stochastic transition closure for long-horizon turbulence diffusion. The closure is fit offline in a coarse latent observable space and defines stochastic transition tubes calibrated from DNS-referenced latent trajectories. During diffusion training, the frozen closure regularizes the denoiser's predicted-clean latent transitions by penalizing deviations from these calibrated tubes, while leaving the CNF decoder and reverse diffusion sampler unchanged. The method is fully PDE-free, uses no PDE residual or sampling-time correction, and is designed to improve transition-side temporal reliability rather than to serve as a universal physical simulator. Across DNS-referenced long-horizon protocols, the proposed method yields targeted gains in temporal transport and memory diagnostics, especially in high-drift regimes, while decoded-field spectra and derivative-sensitive quantities are reported as physical-fidelity guardrails rather than uniformly optimized targets. These results support stochastic transition closure as a lightweight mechanism for improving the temporal reliability of PDE-free turbulence diffusion.

The remarkable scalability of Transformers has expanded their application to 3D computer vision, where camera-aware positional encoding is crucial for providing spatial cues in multi-view geometry. Recent advancements have established the practice of using camera parameters---such as extrinsics or projection matrices---as relative positional encoding into the query, key, and value vectors of the attention mechanism. However, when scaling up the training recipe of novel view synthesis (NVS) models with the camera-based positional encoding, we observe a significant issue: model performance stagnates in the late stages of training. In this paper, we investigate the cause of the performance bottleneck when scaling up and demonstrate that storing rotation and translation given by the positional encoding in the same dimensions of the value vector causes indeterminacy in their independent identification, hindering training scalability. To address this, we propose Decoupled Pose Positional Encoding (DPPE), a novel camera-based positional encoding that explicitly decouples rotation and translation. Extensive evaluations on NVS tasks demonstrate that DPPE enables stable long-term training even in scaled-up training setup. Furthermore, it exhibits superior generalization performance in extrapolation settings, such as handling an increased number of viewpoints and zoom-in scenarios.


DynaSub: Adaptive Subgrouping for Scalable Representation Learning

Tina Behrouzi ⋅ Sana Tonekaboni ⋅ Rahul Krishnan ⋅ Anna Goldenberg

Real-world observational data often contain existing or emerging heterogeneous subpopulations that deviate from global patterns. Without the ability to detect Out-of-Distribution (OOD) data and model underlying structure, systems risk failing to adapt to emerging patterns, leading to inaccurate or potentially harmful predictions. We introduce DynaSub, a Dynamic Subgrouping Variational Autoencoder that couples latent representation learning with subgroup discovery. DynaSub operates on pre-trained foundation models or regular encoders, learning latent embeddings that define subgroup structure and are iteratively refined to enhance subgroup separability and OOD sensitivity. It incorporates a nonparametric clustering mechanism directly in latent space, enabling the number and structure of subgroups to adapt dynamically during training. DynaSub achieves competitive performance on near- and far-OOD detection across ResNet and ViT-based foundation encoders, reducing false positive rates by up to 10\% under covariate shift while maintaining high AUROC across multiple OOD benchmarks, and excelling in class-OOD settings where entire classes are unseen during training.

Biologically plausible learning requires temporal credit assignment without backpropagation through time (BPTT). Hamiltonian Echo Learning (HEL) achieves this by running a neural system twice --- an \emph{inference} phase then a time-reversed \emph{echo} phase --- recovering exact BPTT gradients. We identify spatio-temporal sum-separability of the Hamiltonian as the structural condition making HEL's learning rule both spatially and temporally local, and exploit it to extend HEL beyond diagonal recurrences: a Hopfield-inspired oscillatory RNN with dense recurrent connectivity yields a contrastive Hebbian rule with constant memory and only two forward passes. This rule matches full BPTT and outperforms e-prop and truncated BPTT on time-series classification and regression. We then relax HEL's Hamiltonian-reversibility constraint --- which forces unbiological connectivity and activation patterns --- and derive \emph{Echo Learning} (EL), exact for arbitrary smooth dynamics whenever the network is reversible and self-adjoint with respect to a readout involution. Both conditions are enforced by a spatio-temporally local homeostatic loss, recovering near-BPTT gradient quality without architectural constraints.


Ensemble Distributionally Robust Bayesian Optimisation

Tigran Ramazyan ⋅ Denis Derkach

We study zeroth-order optimisation under context distributional uncertainty, a setting commonly tackled using Bayesian optimisation (BO). A prevailing strategy to make BO more robust to the complex and noisy nature of data is to employ an ensemble as the surrogate model, thereby mitigating the weaknesses of any single model. In this study, we propose a novel algorithm for Ensemble Distributionally Robust Bayesian Optimisation that remains computationally tractable while managing continuous context. We obtain theoretical sublinear regret bounds, improving current state-of-the-art results. We show that our method’s empirical behaviour aligns with its theoretical guarantees.


Entropy-Gated Latent Recursion

Soham Bhattacharjee ⋅ Dushyant Singh Chauhan ⋅ Salem Lahlou ⋅ Martin Takac ⋅ Nils Lukas

Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity from a single source: stochastic token-level sampling. We argue that this single-axis sampling space is fundamentally limiting, and identify a second, fully deterministic and complementary axis: the layer span L at which a frozen model's top decoder layers are recursively re-applied at high-uncertainty tokens. Different choices of L produce distinct rollouts that solve different subsets of problems, with no stochasticity. We instantiate this axis through Entropy-Gated Latent Recursion (EGLR), a training-free decoding procedure that re-applies the top-L layers for at most K_max iterations until the next-token distribution converges. Combined with T temperature samples, EGLR turns a single-axis stochastic rollout pool into an L×T Cartesian sampling space at almost the same per-rollout cost. We characterize this space across 8 instruction-tuned models and 6 math reasoning benchmarks, and show that the L-axis is genuinely complementary to temperature: on MATH-500 with Qwen2.5-3B-Instruct, the joint L×T oracle reaches 91.6%, +8.2 percentage points beyond the temperature-only oracle (83.4%) and +10.4 points beyond the layer-only oracle (81.2%), confirming that the two axes capture genuinely complementary problems. The expanded rollout pool provides richer per-prompt candidates for any downstream procedure that consumes rollouts, including self-consistency, best-of-N with verifiers, and group-relative RL training (GRPO), opening a new direction for inference-time scaling that does not rely on stochastic noise. The code is available at https://anonymous.4open.science/r/EGLR/.

When approximating an intractable density via variational inference *VI* the variational family is typically chosen as a simple parametric family that very likely does not contain the target. This raises the question: *Under which conditions can we recover characteristics of the target despite misspecification?* In this work, we extend previous theoretical results on robust VI with location-scale families under target symmetries in two substantial ways: (1) We open them up to a wider range of divergences by providing sufficient conditions for exact recovery of the target mean and correlation matrix when using the forward Kullback-Leibler divergence and $\alpha$-divergences. (2) By doing so, we find that we can drop the restrictive assumption of a log-concave target made in previous work, allowing us to give guarantees for a wider range of targets, including multi-modal ones. In our experiments, we show how our guarantees can serve as guidelines for the choice of the variational family and $\alpha$-value and we illustrate on a diverse set of examples how and why optimization can fail in the absence of our sufficient conditions.


Factorized Self-Supervised Speech Tokenization

Benjamin Van Niekerk ⋅ Jean-Philippe Letendre ⋅ Nicol Visser ⋅ Hugo Seuté ⋅ Herman Kamper ⋅ Mirco Ravanelli

Speech tokenizers convert audio into discrete units, enabling language models to process and generate speech. The goal is to learn compact, text-like representations while preserving enough detail to reconstruct a speaker’s voice and delivery. We propose Facto: a factorized tokenizer that disentangles time-varying information (content and prosody) from static components (speaker identity and recording conditions). We build on a recent linear factorization method for self-supervised features, extending it to handle unseen speakers. Then, we cluster the factorized space and train a lightweight decoder to reconstruct audio from the resulting tokens. By removing the static components before clustering, our tokens are more robust to small noise perturbations. We validate this by testing token stability under various noise conditions. Next, we theoretically analyze Facto's convergence and scaling properties, quantifying how the factorization improves with the number of training speakers. Finally, we evaluate Facto on a range of downstream tasks. We show that greater token stability translates into improved text-to-speech performance and competitive results in voice conversion and spoken language modeling.


Fast Learning Rates for Physics-Informed Kernel Methods

Luc Brogat-Motte ⋅ Joachim Bona-Pellissier ⋅ Giacomo Meanti ⋅ Lorenzo Rosasco

In physics-informed machine learning, a target function $u^* $ is learned from noisy value observations $y_i=u^* (x_i)+ \varepsilon_i$, together with differential information, given either by noisy observations $d_j=(Du^* )(z_j)+\xi_j$ or by a known physical constraint $Du^* =v$ . We consider the setting where $D$ is a linear differential operator and analyze a physics-informed kernel estimator $\hat u$ combining $n$ value observations and $m$ differential observations. In this context, we ask how much can differential information improve predictions, and how does this improvement depend quantitatively on $n$ , $m$, and $D$. We prove finite-sample bounds, supported by numerical simulations, revealing a two-regime structure for the prediction error. When $m$ is limited, the rate depends jointly on $n$ and $m$; when $m$ exceeds a problem-dependent threshold, the rate saturates and matches the oracle rate obtained when the perfect constraint $D \hat u = Du^* $ is imposed. Examples are discussed for Sobolev spaces which are reproducing kernel Hilbert spaces and include partial Laplacian constraints on the torus and gradient observations on bounded domains. These examples illustrate the range of possible learning rate improvements - from the standard nonparametric $n^{-1/4}$ to the parametric rate $n^{-1/2}$. Finally, we derive physically consistent rates in a stronger norm that jointly controls the errors in $\hat u$ and $D\hat u$.

Adversarial detection on evolving attributed graphs faces two challenges: fraud labels arrive late or not at all, and natural distributional drift erodes any trained decision boundary. Modern detectors can reach strong detection rates when labels are available and the distribution is stationary; sustaining that quality without labels and under drift is the open problem. We identify \emph{feature-context consistency} --- the alignment between a node's attributes and its aggregated neighborhood attributes --- as a single structural signal that addresses both. We formalize it as the \textbf{Context Boost} (CB) score within a Restricted Boltzmann Machine framework (CB-RBM) and prove two complementary guarantees: a strict separation between legitimate and adversarial CB distributions, with gap growing linearly in model capacity, whenever the adversary lacks legitimate neighborhood context; and a Wasserstein-Lipschitz stability bound under natural drift that prevents false alarms without retraining. We validate on XBlock Ethereum phishing ($\sim$153K nodes) and Reddit banned-user detection ($\sim$11K nodes), with two further datasets in the appendix spanning diverse domains. With zero labels, CB-RBM achieves AUROC~$\geq 0.96$ on four adversarial attacks and keeps FPR close to its calibrated $5\%$ target under drift --- outperforming supervised GNNs (hundreds of labels), unsupervised graph-anomaly baselines, and a standard RBM whose FPR collapses to $89$--$100\%$. What typically requires separate mechanisms --- label-free operation and drift stability --- here follows from a single structural regularity of the data.


Federation Is the Way Forward for AI Agents

Herbert Woisetschläger ⋅ Nicholas Lane ⋅ Shiqiang Wang

Current AI agents are largely deployed as centralized services. This architecture is prone to correlated reliability failures, unsustainable energy concentration, and structural barriers to accessing high-value private context distributed across institutions and jurisdictions. In this position paper, we argue that *federation is the way forward for AI agents*, as the most valuable context for agents is often distributed across institutions and jurisdictions. We support our position through technical, policy, and economic analysis, showing that the shift towards federated agentic architectures is both necessary and tractable. We project that federation leads to significant GDP growth potential of \\$783B and \\$470B over 10 years for the U.S. and EU, respectively. These effects are enabled by unlocking cross-boundary tasks that centralized architectures cannot reach. We also propose a deployment roadmap and research agenda that directly address the most significant federation risks.


Feedback Forensics: A Toolkit to Measure AI Personality

Arduin Findeis ⋅ Timo Kaufmann ⋅ Eyke Hüllermeier ⋅ Robert Mullins

Some traits making a “good” AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model character or personality. Without a clear objective, conventional benchmarks based on automatic checks struggle to measure such traits. Evaluation methods using human feedback such as Chatbot Arena have emerged as a popular alternative. These methods infer “better” personality and other desirable traits implicitly by ranking multiple model responses relative to each other. Recent issues with model releases highlight limitations of these existing opaque evaluation approaches: goblin-obsessed, sycophantic and generally over-the-top personalities have resulted in model rollbacks and extensive media coverage. Despite these known issues, limited public tooling exists to explicitly evaluate model personality. We introduce Feedback Forensics: an open-source toolkit to track AI personality and related behavioural traits, both those encouraged by human (or AI) feedback, and those exhibited across AI models trained and evaluated on such feedback. Feedback Forensics is adaptive and can automatically find new traits appearing in models. We demonstrate the toolkit’s usefulness in two steps: (A) first we analyse the personality traits encouraged in popular human feedback datasets including Chatbot Arena, MultiPref and PRISM; and (B) then use our toolkit to analyse how much popular models exhibit such traits. We release (1) our Feedback Forensics toolkit alongside (2) a web app tracking AI personality in popular models and feedback datasets and (3) the underlying annotation data.


Filtered-Trace Online Variational Training for Probabilistic Spiking Neural Networks

Yaokun Wang ⋅ Tiantian Xiao ⋅ Hongyan Ding ⋅ Zhi Yan

Probabilistic spiking neural networks (SNNs) model spike trains as temporal point processes and provide a principled framework for learning with latent spikes. Recent differentiable point-process methods enable path-wise variational learning, but their training still relies on full-sequence backpropagation through time (BPTT), leading to memory costs that grow with the temporal horizon. In this paper, we develop an online variational training framework for probabilistic SNNs based on discrete-time spike response model dynamics. By representing synaptic history with finite-dimensional Markovian traces, our method updates both generative and variational parameters without storing the full temporal computation graph. To handle future-dependent credit assignment from recurrent spike histories, we introduce a horizon-$R$ family of online estimators. Theoretically, we show that the truncation-induced bias decays exponentially with $R$ under the SRM kernel-decay condition. Synthetic experiments validate the predicted horizon-dependent behavior and bias decay, while N-MNIST experiments show competitive classification accuracy and avoid the sequence-length-dependent memory growth of BPTT.

Temporal point processes (TPPs) are the canonical framework for event sequence data, with applications spanning seismology, social media, and electronic health records. Recent neural extensions have been reported to substantially outperform classical baselines such as Hawkes processes, but a systematic comparison across model families and intensity parameterizations is lacking, and large-scale medical benchmarks — arguably one of the most promising application domains — have not been considered. Here, we revisit this comparison through a systematic benchmark of neural and non-neural TPPs on existing datasets and on two large-scale medical benchmarks derived from UK Biobank disease histories and MIMIC-IV ICU records. First, our results show that the previously reported gap between Hawkes processes and neural TPPs largely disappears once Hawkes processes are equipped with flexible and learnable kernels. Second, we find that the flexibility of the intensity parameterization also limits prior neural TPPs: a transformer with a spline-based intensity head outperforms prior neural TPP implementations. Third, our benchmark underscores the value of complex datasets when assessing the performance of TPPs. In fact, exclusively on the MIMIC-IV ICU records did we find clear evidence for structures — time-varying or higher-order interactions — beyond what pairwise flexible Hawkes process kernels can capture. Our results guide both future method development and practitioners’ model choice. Our TPP framework is released as an open-source package, providing consistent and feature rich implementations of classical and neural TPPs.

Distributed learning algorithms are vulnerable to adversarial nodes, a.k.a. Byzantine failures. To solve this issue, robust algorithms have been developed, which typically replace parameter averaging by robust aggregations. While generic conditions on these aggregations exist to guarantee the convergence of (Stochastic) Gradient Descent (SGD), the analyses remain rather ad-hoc. This hinders the development of more complex robust algorithms, such as accelerated ones. In this work, we show that Byzantine-robust distributed optimization can, under standard generic assumptions, be cast as a general optimization with inexact gradient oracles, an active field of research. This allows to obtain state-of-the-art results for Byzantine-robust optimization from general inexact first-order analyses. We first show that inexact GD on top of standard robust aggregation procedures obtains optimal asymptotic error in the Byzantine setting. Going further, we study an algorithm for Optimization under Similarity, in which the server leverages an auxiliary loss function that approximates the global loss. Then, we introduce an Accelerated Extra-gradient method, that yields acceleration in both the standard and similarity settings. We first give new general convergence results for these inexact schemes, and then instantiate these results in the Byzantine setting through our reduction. Both algorithms drastically reduce the communication complexity compared to previous methods, as we show theoretically and empirically.

Accurate protein representations that integrate sequence and three-dimensional (3D) structure are critical to many biological and biomedical tasks. Most existing models either ignore structure or combine it with sequence through a single, static fusion step. Here we present FusionProt, a unified model that learns representations via iterative, bidirectional fusion between a protein language model and a structure encoder. A single learnable token serves as a carrier, alternating between sequence attention and spatial message passing across layers. FusionProt is evaluated on Enzyme Commission (EC), Gene Ontology (GO), and mutation stability prediction tasks. It improves Fmax by a median of 1.3 points (up to 2.0) across EC and GO benchmarks, and boosts AUROC by 3.6 points over the strongest baseline on mutation stability. Inference cost remains practical, with only ~2-5% runtime overhead. Beyond state-of-the-art performance, we further demonstrate FusionProt’s practical relevance through representative biological case studies, suggesting that the model captures biologically relevant features.


GDMD: Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning

Linwei Dong ⋅ Ruoyu Guo ⋅ Ge Bai ⋅ Quan Zheng ⋅ Yawei Luo ⋅ Changqing Zou

Diffusion distillation, exemplified by Distribution Matching Distillation (DMD), has shown great promise in few-step generation but often sacrifices quality for sampling speed. While integrating Reinforcement Learning (RL) into distillation offers potential, a naive fusion of these two objectives relies on suboptimal raw sample evaluation. This sample-based scoring creates inherent conflicts with the distillation trajectory and produces unreliable rewards due to the noisy nature of early-stage generation. To overcome these limitations, we propose GDMD, a novel framework that redefines the reward mechanism by prioritizing distillation gradients over raw pixel outputs as the primary signal for optimization. By reinterpreting the DMD gradients as implicit target tensors, our framework enables existing reward models to directly evaluate the quality of distillation updates. This gradient-level guidance functions as an adaptive weighting that synchronizes the RL policy with the distillation objective, effectively neutralizing optimization divergence. Empirical results show that GDMD sets a new SOTA for few-step generation. Specifically, our 4-NFE (Number of Function Evaluations) models outperform the quality of their multi-step teacher and substantially exceed previous DMDR results in GenEval and human-preference metrics, exhibiting strong scalability potential.


General Agent Evaluation

Elron Bandel ⋅ Asaf Yehudai ⋅ Lilach Edelstein ⋅ Yehoshua Sagron ⋅ Yotam Perlitz ⋅ Elad Venezian ⋅ Natalia Razinkov ⋅ Natan Ergas ⋅ Shlomit S Ifergan ⋅ Segev Shlomov ⋅ Michal Jacovi ⋅ Leshem Choshen ⋅ Liat Ein-Dor ⋅ Yoav Katz ⋅ Michal Shmueli-Scheuer

General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measured how agent architecture shapes performance across heterogeneous protocols and diverse unfamiliar environments. This is the first systematic study, comparing tool-calling, MCP, code-generation, and CLI agents on the same benchmarks with the same models. Two gaps blocked such a study: existing harnesses require per-benchmark wiring or fixed protocol classes (web for BrowserGym, CLI for Harbor), and benchmarks themselves expect human-authored prompts, context, and integration glue. To enable this study, we contribute (1) a unifying protocol that bridges existing benchmark and agent protocols; (2) an evaluation harness that surfaces any benchmark to any general-purpose agent and backbone model; and (3) the first Open General Agent Leaderboard of agent configurations, a full factorial over 5 agent architectures × 5 backbone LLMs (three closed-source, two open-weight) × 6 benchmarks spanning software engineering, customer service, deep research, and personal assistance. We find that (i) general agents adapt to every tested domain without per-domain customization; (ii) agent architecture choice swings results by up to 12pp within a single model, yet backbone model choice dominates overall performance; (iii) on 4 of 6 tested benchmarks, top general agents are indistinguishable from the leading heavily-customized domain-specific agents; (iv) open-weight models tested exhibit "generality sinks" absent from frontier closed-source models: they consistently collapse on specific agent architectures or benchmarks. Code, harness, leaderboard, and traces will be released upon acceptance.


Generalization Dynamics of Linear Diffusion Models

Claudia Lioba Merger ⋅ Sebastian Goldt

Diffusion models are powerful generative models that produce high-quality samples from complex data. While their infinite-data behavior is well understood, their generalization with finite data remains less clear. Classical learning theory predicts that generalization occurs at a sample complexity that is exponential in the dimension, far exceeding practical needs. We address this gap by analyzing diffusion models through the lens of data covariance spectra, which often follow power-law decays, reflecting the structure of real data. To understand whether such a power-law structure can benefit learning in diffusion models, we develop a theoretical framework based on linear neural networks, congruent with a Gaussian hypothesis on the data. We quantify how the covariance spectra of data and regularization impact generalization. We find two regimes: When $N


Generalizing the Geometry of Model Merging Through Fréchet Averages

Marvin da Silva ⋅ Mohammed Adnan ⋅ Felix Dangel ⋅ Sageev Oore

Model merging aims to combine multiple models into one without additional training. Naïve parameter-space averaging can be fragile under architectural symmetries, as their geometry does not take them into account. In this work we show that not only the geometry, but also the averaging procedure itself, must be symmetry-invariant to achieve symmetry-aware merges. Consequently, we propose a general solution: merging as Fréchet averaging, i.e., selecting parameters that minimize a sum of geodesic distances on an appropriate manifold. In this view, the key design choice is the overall geometry, i.e., the choice of metric, manifold, and distance approximation, that determines what it means for two models to be “close.” We show that Fréchet averaging, combined with simplifying assumptions, contains Fisher merging. Building on this, we examine the particular case of low-rank adapters (LoRA), whose symmetries induce a distinct geometry: that of a quotient manifold. We outline the limitations of current LoRA merging methods, propose a practical algorithm for this setting, and show how they compare with other commonly used approaches.


GRASP: Guided Residual Adapters with Sample-wise Partitioning

Felix Nützel ⋅ Mischa Dombrowski ⋅ Bernhard Kainz

Text-to-image flow matching transformers degrade sharply in long-tail settings: tail-class outputs collapse in fidelity and diversity, limiting their value as synthetic augmentation for rare conditions. We trace this to low head-versus-tail gradient alignment during fine-tuning, an optimization-level pathology that conditioning- and sampling-side interventions do not address. We propose GRASP (Guided Residual Adapters with Sample-wise Partitioning): a deterministic partition of the conditioning space, paired with group-specific residual adapters in the transformer feedforward layers, that leaves the flow-matching objective and the sampler untouched. In conditional flow matching, condition values index distinct sets of probability paths, so partitioning along the conditioning is the structurally correct factorization suitable as gradient alignment proxy. Because the partition is static, every tail sample is guaranteed to update its assigned expert, which bypasses extreme longtail failure modes. Crucially, GRASP is non-invasive and composable: on MIMIC-CXR-LT, combining GRASP with self-guided minority sampling at inference time yields the best all-labels IRS we observe, beyond either intervention alone. GRASP itself reduces overall FID by up to 80\% and lifts tail-class coverage by up to 44\% over full fine-tuning, learned-routing MoE, and minority guidance. Used as training data for a downstream DenseNet classifier on NIH-CXR-LT, GRASP synthetics significantly outperform every non-GRASP alternative on macro F1, match the macro F1 obtained from real training data, and yield nonzero F1 on $9$ of $13$ classes versus $3$ of $13$ from full fine-tuning. Results on ImageNet-LT confirm the mechanism is not tied to medical inductive bias.


GridDiffuser: Constraint-Guided Graph Diffusion for AC Optimal Power Flow

Yutong Zheng ⋅ Jichen Zhang ⋅ Danilo Mandic

Modern power systems are undergoing a significant transition from deterministic, steady-state dispatch to operation under high-dimensional uncertainty and rapid fluctuations driven by renewables, electrification, and topology perturbations. This shift requires Optimal Power Flow (OPF) to be solved frequently under diverse operating conditions. Classical AC-OPF solvers provide high physical fidelity but are slow for high-frequency scheduling, while linear approximations such as DC-OPF sacrifice accuracy for computational speed. To address these challenges, we propose GridDiffuser, a graph diffusion solver for AC-OPF. Unlike deterministic learning methods that predict a single solution and struggle to represent the non-convex landscape, GridDiffuser models a multi-modal distribution over operating points and generates multiple candidate solutions. Our method further applies an efficient Levenberg--Marquardt correction to the clean-space estimate, iteratively steering samples toward the constraint-feasible set. Experiments on thousand-bus grids with instance-level N-1 contingencies show that GridDiffuser substantially reduces pre- and post-power flow (PF) violations compared with deterministic baselines while generating solutions much faster than AC-IPOPT.


Group-Aware Matrix Estimation and Latent Subspace Recovery

Hamza Golubovic ⋅ Matthew Shen ⋅ Genevera Allen ⋅ Tarek M Zikry

Modern matrix completion problems often involve heterogeneous data whose rows simultaneously belong to many meta-categories such as demographic and age groups in recommendation systems, or region and recording session labels in electrophysiology experiments. Standard low-rank estimators impose a single global latent geometry, which can recover average structure but may smooth away subgroup-specific variation, especially when observations are unevenly distributed across groups. We introduce Group-Aware Matrix Estimation (GAME), a convex estimator for overlapping subgroup-wise low-rank matrix estimation. GAME regularizes category-specific submatrices through overlapping nuclear-norm penalties, allowing related groups to borrow information while preserving local latent structure in a shared coordinate system. We provide finite-sample guarantees for both reconstruction error and subgroup-specific subspace recovery, showing how performance depends on sampling density, subgroup rank, and overlap structure. Experiments on synthetic, recommendation, ecological, and neuroscience datasets show that GAME is most beneficial in structured missingness regimes, where subgroup-aware regularization improves both reconstruction accuracy and latent subspace fidelity. Across these benchmarks, GAME outperforms global low-rank, side-information, and modern imputation baselines, with the largest gains when subgroup heterogeneity and latent geometry are central to the task.


GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

Pepijn Cobben ⋅ Xuanqiang A Huang ⋅ Thao Pham ⋅ Isabel Dahlgren ⋅ Bernhard Schölkopf ⋅ Terry J Zhang ⋅ Zhijing Jin

Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict poorly understood. We introduce GT-HarmBench, a benchmark of 1,535 high-stakes scenarios spanning game-theoretic structures such as the Prisoner's Dilemma, Stag Hunt and Chicken. Scenarios are drawn from realistic AI risk contexts in the MIT AI Risk Repository. Across 15 frontier models, agents fail to choose socially beneficial actions in 38% of high-stakes cases, such as military escalation, election manipulation, and medical malpractice. We measure sensitivity to game-theoretic prompt framing and ordering, and analyze reasoning patterns driving failures. We further show that game-theoretic interventions improve socially beneficial outcomes by up to 18%. Our results highlight substantial reliability gaps and provide a broad standardized testbed for studying alignment in multi-agent environments.


Harnessing Accurate and Automatic Trend Detection in Data Streams via Tbps-Level Inference

Yuhan Wu ⋅ QianXun Xu ⋅ Lida Liao ⋅ Longlong Zhu ⋅ Jiashuo Yu ⋅ Linying Zheng ⋅ Hongyan Liu ⋅ Dong Zhang ⋅ Chunming Wu ⋅ Xiang Chen

Detecting frequency trend patterns, such as items with sustained growth or decline, is an important and challenging task in high-speed data streams. State-of-the-art solutions require manual tuning of many coupled parameters, while their bloom filters suffer from high false positive rates. In this paper, we propose NeuTrend, a framework that provides accurate and automatic trend detection with high-performance in-network inference. Our key idea is that a lightweight BRNN can learn to classify incoming items as likely trending or non-trending. Thus, it eliminates unqualified items more accurately than a bloom filter and removes the need for manual parameter tuning. The BRNN uses only XNOR and popcount operations. Thus, it aligns with the strict resource constraints of data-plane switches. NeuTrend further introduces an adaptive detection module. The module uses the BRNN confidence scores to automatically tune the growth and decay thresholds. Extensive experiments on real-world datasets show that the learned filter reduces false positive rates by 23\%-44\% and increases effective detector occupancy from 47\% to 70\%. Hence, NeuTrend improves the overall F1 score by 35\%-68\% over existing solutions, and achieves Tbps-level line-rate processing.


Heterogeneous Agent Collaborative Reinforcement Learning

Zhixia Zhang ⋅ Zixuan Huang ⋅ Gonxun Li ⋅ Huaiyang Wang ⋅ Chengyi Yuan ⋅ Xin Xia ⋅ deqing wang ⋅ Fuzhen Zhuang ⋅ Shuai Ma ⋅ Ning Ding ⋅ Yaodong Yang ⋅ Yikun Ban

We introduce Heterogeneous Agent Collaborative Reinforcement Learning (HACRL), a new Reinforcement Learning from Verifiable Reward (RLVR) problem that addresses the inefficiencies of isolated multi-agent on-policy optimization. HACRL enables collaborative optimization with independent execution: heterogeneous agents share verified rollouts during training to mutually improve, while operating independently at inference time. Unlike LLM-based multi-agent reinforcement learning (MARL), HACRL does not require coordinated deployment, and unlike on-/off-policy distillation, it enables bidirectional mutual learning among heterogeneous agents rather than one-directional homogeneous teacher-to-student transfer. Building on this problem, we propose HACPO, a collaborative RL algorithm that enables principled rollout sharing to maximize sample utilization and cross-agent knowledge transfer. To mitigate capability discrepancies and policy distribution shifts, HACPO introduces four tailored mechanisms with theoretical guarantees on unbiased advantage estimation. Extensive experiments across diverse heterogeneous model combinations and reasoning benchmarks show that HACPO consistently improves all participating agents, outperforming GSPO with double rollouts by an average of 3.6% while using only half the rollout cost.


High-Probability Minimax Adaptive Estimation in Besov Spaces via Online-to-Batch

Paul Liautaud ⋅ Pierre Gaillard ⋅ Olivier Wintenberger

We study nonparametric regression over Besov spaces from noisy observations under sub-exponential noise. Our goal is to obtain minimax-optimal high-probability bounds for the integrated squared error while adapting to the unknown noise level and the regularity parameters of the underlying Besov class. We introduce a wavelet-based online learning algorithm that sequentially processes noisy gradients and adapts to the gradient noise through an adaptive clipping rule, thus avoiding the need to tune parameters such as the noise variance or gradient bounds. As a by-product of our analysis, we derive high-probability adaptive regret bounds that scale with the $\ell_1$-norm of the competitor. Finally, in the batch statistical setting, our method is the first to achieve high-probability minimax-optimal estimation rates over Besov spaces while adapting to all problem parameters, including the noise level. Our approach relies on a refined online-to-batch conversion and exploits the structure of the squared loss in combination with self-normalized concentration inequalities.


How Do Language Models Understand Tables? A Mechanistic Analysis of Cell Location

Xuanliang Zhang ⋅ Dingzirui Wang ⋅ Keyan Xu ⋅ Qingfu Zhu ⋅ Wanxiang Che

While Large Language Models (LLMs) are increasingly deployed for table-related tasks, the internal mechanisms enabling them to process linearized two-dimensional structured tables remain opaque. In this work, we investigate the process of table understanding by dissecting the atomic task of cell location. Through activation patching and complementary interpretability techniques, we delineate the table understanding mechanism into a sequential three-stage pipeline: Semantic Binding, Coordinate Localization, and Information Propagation. We demonstrate that models locate the target cell via an ordinal mechanism that counts discrete delimiters to resolve coordinates, offering mechanistic evidence across diverse table formats. Furthermore, column indices are encoded within a linear subspace that allows for precise steering of model focus through vector arithmetic. Finally, we extend our analysis to the real-world HiTab dataset and show that the same three-stage mechanism persists in hierarchical real-world tables and more complex tasks.


How LLMs Distinguish Threats from Offers

Julia Karbing ⋅ Lewis Hammond ⋅ Philip Torr

Threats and offers are fundamental to strategic interactions between agents. Competently navigating multi-agent settings requires the ability to distinguish between the two, reason about the intentions behind them, and understand how other agents assess them. With the prospect of LLM-based autonomous agents being deployed as representatives or advisors to humans in consequential domains, it is therefore important to understand how LLMs reason about threats and offers. In this paper, we propose a formal theory of threats and offers grounded in causal games, and use it to empirically investigate how current LLMs classify and reason about such proposals. We find that LLMs' threat/offer classifications align with our formal definition, with each model's threat labels tightly tracking its own judgement of whether the proposal leaves the recipient worse off. LLMs also predict their co-players' classifications accurately, though in our setting this accuracy is not clearly above what is already achieved by simply projecting their own labels onto the co-player.

Parallel Langevin samplers are usually run as $K$ independent chains and then averaged. Independence is convenient, but it is not required by the unadjusted Langevin algorithm (ULA): each chain only needs a standard Gaussian noise marginal at each step. We use this freedom by coupling the same-step noises across chains. The coupling is simple: draw $K$ iid Gaussian noises, subtract their across-chain mean, and rescale. Each chain still has the ordinary ULA law, but the noise injected into the ensemble average is exactly zero. For quadratic targets, this removes the Langevin-noise contribution from every equal-weight linear summary of the ensemble mean at every finite horizon. With deterministic or zero-sum randomized starts, these summaries have zero total variance. On UCI Bayesian logistic posteriors with $K=8$, the same construction reduces ensemble-mean trace variance to $6\times 10^{-3}$, $2\times 10^{-3}$, and $5\times 10^{-4}$ of matched iid ensembles on WDBC, Spambase, and Adult; reference MSE against a longer iid run falls by roughly $4$ to $14\times$ after accounting for error in the finite reference run. The gain is scoped: zero-sum coupling cancels linear fluctuations, while centered quadratic observables have factor $1/(K-1)$ rather than iid's $1/K$, and random starts, minibatches, and nonquadratic curvature add explicit residual variance terms. Synthetic and real-data experiments match the exact cancellation where it applies and the predicted behavior outside that regime.


Hypernetworks for Dynamic Feature Selection

Javier Fumanal Idocin ⋅ Raquel Fernandez-Peralta ⋅ Javier Andreu-Perez

Dynamic feature selection (DFS) is a machine learning framework in which features are acquired sequentially for individual samples under budget constraints. The exponential growth in the number of possible feature acquisition paths forces a DFS model to balance fitting specific scenarios against maintaining general performance, even when the feature space is moderate in size. In this paper, we study the structural limitations of existing DFS approaches to achieve an optimal solution. Then, we propose \textsc{Hyper-DFS}, a hypernetwork-based DFS approach that generates feature subset-specific classifier parameters on demand. We show that the use of hypernetworks compared to mask-embedding methods results in a smaller structural complexity bound. We also use a Set Transformer encoding to create a smooth conditioning space for the hypernetwork, so that functionally similar tasks are also geometrically close. In our benchmarks, \textsc{Hyper-DFS} outperforms all state-of-the-art approaches on synthetic and real-life tabular data. It is also competitive or superior across all image datasets tested, and shows substantially stronger zero-shot generalisation to feature subsets never seen during training than existing DFS approaches.


Hyperparameter Transfer for Dense Associative Memories

Roi Holtzman ⋅ Dmitry Krotov ⋅ Boris Hanin

Dense Associative Memory (DenseAM) is a promising family of AI architectures that is represented by a neural network performing temporal dynamics on an energy landscape. While hyperparameter transfer methods are well-studied for feed-forward networks, these methods have not been developed for settings in which weights are shared across layers and within the layer, which is common in DenseAMs. Additionally, DenseAMs utilize rapidly peaking activation functions that are rarely used in feed-forward architectures. The confluence of these aspects makes DenseAM a challenging framework for using existing methods for hyperparameter transfer. Our work initiates the development of hyperparameter transfer methods for this class of models. We derive explicit prescriptions for how the hyperparameters tuned on small models can be transferred to models trained at scale. We demonstrate excellent agreement between these theoretical findings and empirical results.


Improved Algorithms for Online Classification with Surrogate Losses

Abed Razawy ⋅ Valentina Masarotto ⋅ Dirk van der Hoeven

We study online multiclass classification with surrogate losses. We identify and exploit a structural property of margin-based classifiers: on many rounds predictions are stable in the sense that after an update of the parameters the algorithms do not change the predicted label on the current example. By exploiting this stability we improve upon the state of the art in several ways. We provide an improved surrogate regret bound for the perceptron, develop a parameter-free version of $\texttt{GAPTRON}$, develop improved results for the delayed feedback setting, and extend our results to the batch setting. These results are complimented by empirical evaluations.

Remote sensing imagery supports diverse visual tasks, such as object detection and semantic segmentation, but these tasks vary across scenes, imaging modalities, and target objects. Since many remote sensing scenarios lack sufficient annotations, conventional supervised learning is often difficult to apply, motivating in-context learning as a flexible alternative.We present an in-context learning framework for remote sensing visual tasks, where a new task is specified by a single annotated example and applied to unseen images without retraining at inference time. Built upon diffusion models, our framework establishes target-reference associations through cross-attention and transfers the structural relationship between the reference image and its annotation to produce task-specific predictions. Unlike natural images, remote sensing objects are often small and provide weak semantic cues in early diffusion steps, making semantic alignment between target and reference images difficult. We address this limitation with a semantic-aware cross-attention that combines semantic and fine-grained details for more accurate cross-image matching. Remote sensing objects also appear at arbitrary orientations, causing large appearance variations and unstable in-context transfer. To improve robustness, we introduce a rotation-robust learning strategy that reduces sensitivity to orientation changes. Experiments show that our method achieves strong one-shot performance, establishing a flexible paradigm for adapting remote sensing models to diverse visual tasks.


Inference-Time Refinement Closes the Synthetic-Real Gap in Tabular Diffusion

Eugenio Lomurno ⋅ Filippo Balzarini ⋅ Francesco Benelle ⋅ Francesca Pia Panaccione ⋅ Matteo Matteucci

Diffusion-based generators set the current state of the art for synthetic tabular data, deployed downstream wherever direct access to real records is restricted. These methods approach but rarely exceed real-data utility on downstream tasks, and closing this synthetic--real performance gap has so far been pursued exclusively at training time, via architectural advances, scaling, and retraining of monolithic generators. The inference-time alternative, i.e., refining the outputs of a pre-trained backbone with parameters left untouched, has remained largely unexplored for tabular synthesis. We introduce $\textbf{TARDIS}$ (Tabular generation through Refinement, Distillation, and Inference-time Sampling), an inference-time refinement framework that operates on a frozen pre-trained backbone, configured per dataset by a Tree-structured Parzen Estimator search over score-level guidance during reverse diffusion, with each trial's objective set by an inner grid search over post-hoc sample selectors and an optional soft-label distillation step. The search space encodes a single mathematical pattern we name $\textit{Bidirectional Chamfer Refinement}$ (BCR): the symmetric Chamfer functional between synthetic and real samples is minimized both continuously, via a score-level gradient during reverse diffusion, and discretely, via batch-ranking post-generation. On the majority of datasets the search selects BCR-aligned configurations over alternatives encoded in the search space, evidence both for BCR as the dominant refinement pattern and for TARDIS's per-dataset search as a procedure that recovers this pattern. Across 15 binary, multiclass, and regression benchmarks TARDIS achieves a median $+8.6\%$ downstream-task improvement over models trained on real data (95\% CI $[+3.3, +16.4]$, Wilcoxon $p=0.016$, 11/15 strict wins) and improves over the underlying TabDiff backbone on all 15 datasets (mean $+12.9\%$, $p<10^{-4}$), matching the backbone on manifold fidelity, diversity, and sample-level privacy. The synthetic--real gap is therefore not primarily a training-time problem: on the studied corpus, inference-time refinement of a pre-trained tabular diffusion backbone reaches and exceeds real-data utility in 1 to 80 minutes on a single consumer-grade GPU.


Instability of Meta-Learning Intrinsic Rewards for Policy Gradient Reinforcement Learning

Dilith Jayakody ⋅ Domenic Rosati ⋅ Janarthanan Rajendran

Meta-learned intrinsic rewards are a powerful tool for shaping policy learning in reinforcement learning, particularly when extrinsic rewards are sparse, delayed, or unavailable. Learning Intrinsic Rewards for Policy Gradients (LIRPG) introduced this paradigm and serves as the foundation on which subsequent meta-learned intrinsic reward methods are built. While LIRPG and LIRPG-based methods have shown strong results in low-dimensional control, their behavior in high-dimensional domains, where the policy and value networks use a shared encoder, remains poorly understood. In this work, we analyze meta-learned intrinsic rewards in high-dimensional environments and uncover a consistent failure mode, particularly when training with intrinsic rewards alone, where performance collapses to near-random behavior. We identify the root cause as dense, non-stationary intrinsic rewards inducing large and high-variance value losses that dominate shared encoder updates, suppressing policy learning. We further demonstrate that decoupling policy and value optimization using phasic policy gradient methods is one simple and effective approach to addressing this issue.


Integrating Local and Global Entropy for Uncertainty Quantification in LLMs

Johanne Medina ⋅ Tianyi Zhou ⋅ Keivin Isufaj ⋅ Aristides Gionis ⋅ Sanjay Chawla

Large language models can hallucinate confidently, making uncertainty quantification essential for reliable deployment. Existing approaches rely predominantly on token-level signals from the final (unembedding) layer, while the rich geometric structure of intermediate hidden states remains largely unexplored. Aggregating token-level scores to the response level is itself non-trivial, and no single signal reliably catches the confident-but-wrong failure mode where a model commits to every token with high confidence yet produces an incorrect answer. We address this gap by extracting complementary signals from two distinct representational layers: hidden-state geometric complexity (global uncertainty) from the embedding layer, and token-level entropy (local uncertainty) from the unembedding layer. We show empirically that the two signals cover different regimes that are weakly correlated and thus combining them captures failure modes invisible to either alone. Building on this insight, we propose Global Local Uncertainty (GLU), an unsupervised, single-pass framework that fuses the two signals via a multiplicative gate. Experiments across three benchmarks and three model families show that GLU matches or outperforms all unsupervised baselines and remains competitive with supervised methods that lack cross-dataset generalization.


Interpretability-by-Design with Accurate Locally Additive Models and Conditional Feature Effects

Vasilis Gkolemis ⋅ Loukas Kavouras ⋅ Dimitrios Kyriakopoulos ⋅ Konstantinos Tsopelas ⋅ Dimitrios Rontogiannis ⋅ Giuseppe Casalicchio ⋅ Theodore Dalamagas ⋅ Christos Diou

Generalized additive models (GAMs) offer interpretability through independent univariate feature effects but underfit when interactions are present in data. GA$^2$Ms add selected pairwise interactions which improves accuracy, but sacrifices interpretability and limits model auditing. We propose \emph{Conditionally Additive Local Models} (CALMs), a new model class, that balances the interpretability of GAMs with the accuracy of GA$^2$Ms. CALMs allow multiple univariate shape functions per feature, each active in different regions of the input space. These regions are defined independently for each feature as simple logical conditions (thresholds) on the features it interacts with. As a result, effects remain locally additive while varying across subregions to capture interactions. We further propose a principled distillation-based training pipeline that identifies homogeneous regions with limited interactions and fits interpretable shape functions via region-aware backfitting. Experiments on diverse classification and regression tasks show that CALMs consistently outperform GAMs and achieve accuracy broadly comparable to GA$^2$Ms, while preserving the univariate auditability that GA$^2$Ms forfeit. Overall, CALMs offer a favorable trade-off between predictive accuracy and interpretability.


Interpretable Machine Learning Evaluates Differential Therapy Effects in Parkinsonian Gait

Artur Chudzik ⋅ Henryk Josiński ⋅ Andrzej Przybyszewski

Parkinson's disease therapy is usually adjusted based on extensive clinical examinations, but the effects of medication and deep brain stimulation (DBS) on gait dynamics remain difficult to quantify precisely. We use interpretable probabilistic machine learning to measure therapy-related gait changes in a repeated-measures motion-capture dataset comprising 376 walking trials from 19 patients with Parkinson's disease. Patients were recorded under four therapy conditions: medication off/stimulation off, medication off/stimulation on, medication on/stimulation off, and medication on/stimulation on. They performed natural and fast walking, and a study neurologist scored motor severity using UPDRS-III. From 1.0 s sliding windows of foot-marker motion, we compute short-window largest Lyapunov exponent (LLE)-derived local divergence features that capture local changes in gait trajectory dynamics. We summarize these features as two LLE-derived digital biomarkers: left-right asymmetry and left-right coupling. Bayesian hierarchical regression then estimates medication, stimulation, and task effects while accounting for repeated measurements within the same patients and patient-specific baseline differences. The main result is that during natural walking, LLE asymmetry decreased under combined therapy with a large paired effect size ($d_z=-0.83$). Coupling features changed with therapy in a complementary way, especially for stimulation-related changes in bilateral coordination. These findings suggest that short-window LLE asymmetry tracks therapy-related reduction of lateralized gait dysregulation, while LLE coupling tracks treatment-dependent reorganization of left-right coordination. By modeling novel LLE-derived biomarkers across controlled medication and DBS states, we show that the two therapies are expressed through distinct gait-dynamics signatures, supporting interpretable nonlinear gait biomarkers for objective therapy-response assessment.


It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

Pierre-Carl Langlais ⋅ Pieter Delobelle ⋅ Yannick Detrois ⋅ Pavel Chizhov ⋅ Carlos Rosas-Hinostroza ⋅ Neil S Smail ⋅ Benjamin Burtin ⋅ Hanna Shcharbakova ⋅ Ivan Yamshchikov ⋅ Anastasia Stasenko

Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines—for instance, they contain little explicit reasoning. Thus, many frontier labs have begun to develop their own internal datasets, starting from state-of-the-art models, to augment their pre-training data mix, e.g., with reasoning traces to address cold start problems. While demonstratively effective, none of these datasets are public, and the effect of this so-called _synthetic data_ on knowledge and skill acquisition of language models, including small ones, remains poorly understood. We present Synth, the first open-source synthetic corpus derived from 58,698 Wikipedia articles that collapses pre-, mid-, and post-training into a single training stage via structured amplification of curated encyclopedic seeds. We evaluate Synth by training a series of models: a 56M tiny model (Monad), 0.3B--0.6B dense models (Baguettotron), and a 13B-total / 1B-active Mixture-of-Experts. At iso-compute, Synth outperforms filtered web data, and our models remain competitive with similarly-sized open-weight baselines. Because Synth is back-translated from grounded passages, Synth-trained models achieve high factual precision despite 10-140$\times$ fewer training tokens, with memorization targeted by the seed corpus. These results show that synthetic datasets, including our Synth dataset, are capable of producing competitive generalist models at a significantly lower cost, enabling rapid iteration as the frontier advances. These findings open up possibilities for both generalist models with significantly increased data efficiency, as well as domain-specific models where no instruction or conversational data is available. Finally, we publicly release our Synth dataset and the series of Baguettotron models under a permissive license, thus supporting open-source language model development.

Local intrinsic dimension (LID) quantifies the degrees of freedom of the data manifold around a point and is widely used to detect memorization, diagnose distribution shift, and characterize learned representations. Existing generative-model-based LID estimators are model-specific: they rely on a score, a log-density Hessian, or an exact log-likelihood, and therefore do not apply to ODE-based generators (flow matching, rectified flows, stochastic interpolants, and CNFs) that expose only a velocity field. We propose JaSpec, a model-agnostic LID estimator that counts the singular values of the velocity-field Jacobian $J_t = \partial f_\theta / \partial x$ that fall below a threshold $\tau$. JaSpec applies uniformly to score-based diffusion (via the probability-flow ODE) and to flow-based generators, and recovers Hessian-based estimators on probability-flow diffusion as a verification special case. We prove a direct dynamical identifiability theorem under a Jacobian-normal-dominance condition stated entirely in terms of $J_t$ and verified by an empirical spectral-gap diagnostic. On VP-SDE diffusion and OT-CFM flow-matching benchmarks with known LID, JaSpec attains MAE $= 0$ on most settings up to $D{=}3072$, where no prior model-based estimator reaches zero error, and improves over the strongest model-based baseline by an order of magnitude on the hardest $3072$-dimensional nonlinear mixture. We provide both an exact $\mathcal{O}(D^3)$ path and a linear-cost stochastic Lanczos quadrature path. JaSpec is, to our knowledge, the first LID estimator applicable to flow matching, rectified flows, and generic differentiable ODE generators.


JEPAWG: Interpretable Hypernetworks for Weight-Space Physics

Tobias Göbel ⋅ Julian R Ebelt ⋅ Zier Mensch ⋅ Mathis Gerdes ⋅ Miranda Cheng

Lattice field theory is the workhorse of non-perturbative physics, used to simulate phenomena from the strong nuclear force to critical phenomena in materials. Its Boltzmann distributions are parametrized analytically by \emph{coupling constants}, but these bare parameters are weak predictors of physical observables---extracting physics typically requires extensive simulation. While machine learning tools such as normalizing flows have emerged as effective samplers at fixed couplings, it remains difficult to interpret what the underlying neural networks have learned. This raises a natural question: can the flow \emph{parameters themselves} be generated for new theories, and can their physics be read off directly from the network weights? We propose lattice field theory as a testbed for neural network interpretability: because the target physics is qualitatively well-understood and smoothly varying, it provides ideal synthetic data against which network behavior can be checked against known ground truth. To this end we introduce \textbf{JEPAWG}, a Joint-Embedding Predictive Architecture--based Weight Generator that maps couplings directly to flow weights via a learned latent space. On a scalar theory at lattice sizes $6^2$ and $8^2$, the JEPAWG latent space recovers the correct intrinsic dimension of the underlying manifold, identifies the region of phase transition, encodes a finite-size shift aligned with the 2D Ising exponent $\nu \approx 1$, allowing us to uncover physical structure by studying the network weights alone. As a generator, JEPAWG also interpolates and extrapolates to unseen couplings effectively and remains robust to weight-space incongruences deliberately introduced by combining multi-seed training data, outperforming PCA, AE, and VAE baselines.


KINDER: Kernel-based Independence for Fair Representation Learning via Prototype-space Erasure

Abtin Mogharabin ⋅ Jiaee Cheong ⋅ Alp Toykan Kaplan ⋅ Sinan Kalkan

Concept erasure is a prominent approach to achieving fairness in machine learning through the removal of sensitive attributes from learned representations. Prior concept erasure methods typically define debiasing indirectly through the failure of a chosen adversary, probe, or fairness penalty, making the target of erasure dependent on a particular decoder family or optimization setup. Rather than depending on a particular architecture, task type, or training paradigm, we introduce KINDER, which aims to directly remove sensitive-attribute information from a model's intermediate representations and is compatible with any deep model that contains an intermediate feature space. KINDER operates in a random Fourier feature space, where it estimates a sensitive subspace from protected group prototypes and projects representations onto its orthogonal complement, thereby forming an explicit representation cleaning mechanism. We conduct extensive experiments across diverse settings and show the effectiveness of KINDER on (i) unimodal and multimodal datasets, (ii) supervised and self-supervised settings, (iii) classification, regression, and image segmentation tasks, and (iv) diverse data modalities, including visual, textual, and tabular data.


LAtte: Hyperbolic Lorentz Attention for Joint-Subject EEG Classification

Ahmad Bdeir ⋅ Johannes Burchert ⋅ Tom Hanika ⋅ Lars Schmidt-Thieme ⋅ Niels Landwehr

Electroencephalogram (EEG) classification plays a key role in medical diagnosis and brain–computer interfaces, but remains challenging due to low signal-to-noise ratios and high inter-subject variability. As a result, many existing approaches rely on subject-specific models, which fail to exploit shared structure in neural signals and do not generalize to unseen subjects. To address these limitations, we propose LAtte, a framework that combines Lorentz attention with a hyperbolic InceptionTime-based encoder to improve cross-subject generalization in EEG classification. The model explicitly decomposes EEG signals into a learned baseline component and task-relevant deviations, enabling more structured representation learning. To further improve robustness and adaptability, we incorporate subject-specific low-rank adaptation (LoRA) modules at both encoder and decoder levels, augmented with a Lorentz boost–based LoRA mechanism and hyperbolic projection layers to reduce overfitting in geometric representations. We evaluate LAtte with and without finetuning in three settings: subject-specific, subject-conditional, and leave-one-subject-out (LOSO) on five established EEG datasets, achieving a consistent improvement in performance over current state-of-the-art methods for smaller datasets and maintaining performance for larger datasets.


Learning Global Probabilistic Explanations

Frederic Koriche ⋅ Louenas Bounia

Interpreting the predictions of complex black-box classifiers remains a central challenge in explainable artificial intelligence. While local explanations clarify individual predictions, there is a significant need for global probabilistic explanations that capture a model's overall behavior across the input distribution. We define such an explanation as a small subset $K$ of features, whose relevance is measured by the probability that the classifier assigns identical labels to two independently sampled inputs that agree on $K$. Based on this notion, we investigate the task of identifying maximally relevant global explanations under a cardinality constraint, focusing on product distributions with full support on discrete feature spaces. We introduce a spectral explainer that leverages membership queries to the black-box classifier and employs Fourier-analytic, junta-learning methods to produce probably approximately correct (PAC) explanations. Our algorithm is fixed-parameter tractable with respect to the explanation size limit, maximum feature cardinality, and desired accuracy. Experimental evaluations on synthetic and real-world datasets show that the spectral explainer provides highly relevant global explanations and outperforms heuristic greedy methods as the explanation size increases.


Learning Reusable Options by Decomposing Neural Policies

Parnian Behdin ⋅ Reza Abdollahzadeh ⋅ Kiarash Aghakasiri ⋅ Levi Lelis

Options provide a natural form of behavioral abstraction in reinforcement learning, but discovering reusable options from trained neural policies remains difficult. A trained policy may contain many useful behaviors that can be reused in downstream problems, yet the number of candidate subpolicies that a neural network can encode grows exponentially with the number of hidden units. We propose a differentiable approach for extracting reusable options from neural policies. Our key observation is that for piecewise-linear neural networks, selecting a neural subprogram is equivalent to assigning each hidden unit one of three labels: inactive, active, or retained in the computation. This yields a mask space over neural subprograms that can be searched with gradient-based optimization. We further show that reusable options can require default input parameters: neuron masks determine what computation is reused, while input masks determine how that computation is called. Experiments in transfer settings with feedforward and recurrent policies show that the resulting options improve sample efficiency on downstream learning.

We study the problem of learning to bid when the bidder’s *value is dynamic*, i.e., when the current value depends on past outcomes. Specifically, we consider a bidder participating in repeated second-price auctions whose value depends on the time elapsed since their last successful bid, with auctions arriving in continuous time and only aggregated feedback revealed at the end of the horizon. Such a bidder must **(1)** balance the immediate benefit of winning the current auction against its impact on future values and **(2)** learn unknown environmental parameters. We derive regret bounds for a class of learning methods that combine plug-in estimators with a differential-equation characterization of the optimal policy, and show that a specific confidence bound algorithm learns the optimal policy with a near optimal regret of $\tilde{\mathcal{O}}(\log N)$ for piecewise linear primitives, and $\tilde{\mathcal{O}}(N^{1/3})$ for general, smooth primitives, achieving these regrets without explicit randomization. These theoretical results are supported by numerical experiments.


Learning to Drive in New Cities Without Human Demonstrations

Zilin Wang ⋅ Saeed Rahmani ⋅ Daphne Cornelisse ⋅ Bidipta Sarkar ⋅ Alexander D. Goldie ⋅ Jakob Foerster ⋅ Shimon Whiteson

While autonomous vehicles have achieved reliable performance within specific operating regions, their deployment to new cities remains costly and slow. A key bottleneck is the need to collect many human demonstration trajectories when adapting driving policies to new cities that differ from those seen in training in terms of road geometry, traffic rules, and interaction patterns. In this paper, we show that self-play multi-agent reinforcement learning can adapt a driving policy to a substantially different target city using only the map and meta-information, without requiring any human demonstrations from that city. We introduce NO data Map-based self-play for Autonomous Driving (NOMAD), which enables policy adaptation in a simulator constructed based on the target-city map. Using a simple reward function, NOMAD substantially improves both task success rate and trajectory realism in target cities, demonstrating an effective and scalable alternative to data-intensive city-transfer methods.


Learning to Learn from Multimodal Experience

Xingyu Sui ⋅ Weixiang Zhao ⋅ Yongxin Tang ⋅ Yanyan Zhao ⋅ Yang Wu ⋅ Dandan Tu ⋅ Bing Qin

Experience-driven learning has emerged as a promising paradigm for enabling agents to improve from interaction trajectories by accumulating and reusing past experience. However, existing approaches are predominantly developed in textual settings and rely on manually designed memory schemas, limiting their applicability to multimodal environments. In real-world scenarios, experience is inherently multimodal, involving heterogeneous signals across perception, reasoning, and action, which makes effective memory design significantly more challenging. In particular, the optimal way to structure and utilize multimodal experience is highly task-dependent and evolves over time, rendering fixed memory designs insufficient. In this work, we propose a new paradigm, learning to learn from multimodal experience, which shifts memory design from a predefined component to an adaptive and learnable process. Our framework enables agents to dynamically construct, organize, and utilize memory based on task requirements and interaction history, effectively learning how to structure experience for improved performance. Experiments demonstrate that adaptive memory design substantially enhances agent performance and generalization across multimodal tasks, highlighting the critical role of learning memory mechanisms in experience-driven learning.


Length Generalization for Transformers via Compression

Georg Zetzsche ⋅ Hongjian Jiang ⋅ Andy J Yang ⋅ Pascal Bergsträßer ⋅ Marco Sälzer ⋅ David Chiang ⋅ Anthony Widjaja Lin

Recent advancements in the length generalization theory have provided us with the ability to reliably predict learnability by transformers. In particular, the C-RASP hypothesis (a formalized version of the so-called RASP-l conjecture) posits that transformers length-generalize on a task if and only if a solution is expressible in the C-RASP language. While this hypothesis has strong empirical validation, theoretical problems arise from the fact that no computable length generalization bounds exist for C-RASP, as well as the discovery of seemingly contradictory empirical results. To address this, we refine the C-RASP hypothesis utilizing the recently-proposed fragments of the language, CRASP+ and CRASP1. These fragment have computable length generalization bounds, though in the worst case requiring an extremely large (double exponential) sample size. It is an open question whether this sample size bounds are tight. In this paper, we resolve this open question by providing an exponentially tighter bound. In doing so, we show a polynomial length generalization bound for transformers if we adopt compressed strings, via a novel connection to power words. As an application, we show how this yields a fine-grained analysis of C-RASP conjecture that resolves contradicting experimental evidence against it.


LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels

Lucas Maes ⋅ Quentin Le Lidec ⋅ Damien Scieur ⋅ Yann LeCun ⋅ Randall Balestriero

Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages, pre-trained encoders, or auxiliary supervision to avoid representation collapse. In this work, we introduce LeWorldModel (LeWM), the first JEPA that trains stably end-to-end from raw pixels using only two loss terms: a next-embedding prediction loss and a regularizer enforcing Gaussian-distributed latent embeddings. This reduces tunable loss hyperparameters from six to one compared to the only existing end-to-end alternative. With ~15M parameters trainable on a single GPU in a few hours, LeWM plans up to 48x faster than foundation-model-based world models while remaining competitive across diverse 2D and 3D control tasks. Beyond control, we show that LeWM's latent space encodes meaningful physical structure through probing of physical quantities. Surprise evaluation confirms that the model reliably detects physically implausible events.


Liars' Bench: Evaluating Lie Detectors for Language Models

Kieron Kretschmar ⋅ Walter Laurito ⋅ Sharan Maiya ⋅ Samuel Marks

Prior work has introduced techniques for detecting when large language models (LLMs) lie, that is, generate statements they believe are false. However, these techniques are typically validated in narrow settings that do not capture the diverse lies LLMs can generate. We introduce LIARS' BENCH, a testbed consisting of 72,863 examples of lies and honest responses generated by a range of open-weight models across seven datasets. Our settings capture qualitatively different types of lies and vary along two dimensions: the model's reason for lying and the object of belief targeted by the lie. Evaluating both black- and white-box lie detection techniques on LIARS' BENCH, we find that existing techniques systematically fail to identify certain types of lies, especially in settings where it's not possible to determine whether the model lied from the transcript alone. Overall, LIARS' BENCH reveals limitations in prior techniques and provides a practical testbed for guiding progress in lie detection.


Listening to the Retriever: Perturbation-Sensitive Question Selection for Interactive Person Retrieval

Yunhui Shao ⋅ Yang Bai ⋅ Shuai You ⋅ Bin Yang ⋅ Jie Qiao ⋅ Min Cao ⋅ Mang Ye

Interactive person retrieval extends text-to-image person retrieval by allowing follow-up questions when the initial description is incomplete. The core challenge is to ask questions that uncover the most useful missing information for identifying the target person from visually similar candidates. Existing methods guide question selection with retriever-external signals, which may not reflect the information most useful to the current retriever. To address this, we propose SCOUT, an interactive person retrieval framework that derives questions from the retriever's own response. Specifically, we introduce a Perturbation-Sensitive Question Selection strategy that selects the next question based on Retrieval Perturbation Sensitivity (RPS), a test-time measure of perturbation-induced ranking change. Additionally, SCOUT requires no additional training and builds on off-the-shelf text-to-image person retrievers, avoiding the dialogue-data curation and retraining overhead of existing interactive methods. Extensive experiments on three standard benchmarks demonstrate that SCOUT yields consistent improvements over baseline methods, validating this RPS-based strategy for interactive person retrieval.


Long-Range Spatio-Temporal Graph Propagation Through Oscillations

Alessio Gravina ⋅ Alessandro Trenta ⋅ Tai Hoang ⋅ Andrea Ceni ⋅ Gerhard Neumann ⋅ Davide Bacciu

Graph Neural Networks (GNNs) are powerful tools for learning from spatio-temporal data, as interactions in space can be naturally described by a graph structure. However, capturing long-range dependencies becomes substantially harder when information must flow through space and time simultaneously. Most existing approaches extend GNNs from static graphs to temporal settings by alternating propagation between time and space. In this work, we introduce STORM, a differential-equation-inspired GNN that exploits oscillatory dynamics to propagate information effectively in the joint spatio-temporal domain. By combining a wave-equation update with dissipative and external forcing terms, STORM balances conservative and non-conservative dynamics. We provide a bottom-up analysis of the model, highlighting its propagation behavior and stability, and show that STORM is universal. We empirically validate our method on diverse benchmarks, including tasks designed for analyzing long-range spatio-temporal dependencies, real-world forecasting, and a new long-range task inspired by physical simulations. Across these settings, STORM consistently matches or outperforms strong baselines, establishing a new state of the art for long-range spatio-temporal graph learning.


LSC-Parlament: An Automatically Aligned Catalan Sign Language Dataset from Parliament Videos.

Carlos Escolano ⋅ Gerard Sant ⋅ Marc J Garcia ⋅ Joan G Cortés ⋅ Marcel Granero Moya ⋅ Francesca De Luca Fornaciari ⋅ Maite Melero

Progress in Sign Language Translation (SLT) is frequently hampered by the "data bottleneck," a challenge particularly acute for regional languages such as Catalan Sign Language (LSC). In this paper, we present LSC-Parlament, a large-scale, multimodal dataset for LSC derived from Catalan Parliament sessions spanning 2021 to 2024. The dataset comprises over 87 hours of LSC video, synchronized with speech and text translations in either Catalan or Spanish. To curate this resource, we developed a fully automated pipeline that enables alignment between sign language, speech, and text without requiring prior human-annotated data. We provide a systematic evaluation of the dataset across multiple dimensions, including video tracking quality, audio transcription accuracy, and translation benchmarks. Our experiments compare transfer learning, zero-shot, and supervised translation settings, demonstrating significant knowledge transfer between spoken languages in the LSC-Spanish modality. By releasing LSC-Parlament, we provide a challenging open-domain benchmark to foster research in inclusive and scalable translation technologies.


Market-Based Runtime Resource Allocation for LLM Multi-Agent Systems

Yixue Huang ⋅ Mingxin Wang ⋅ Hao Li ⋅ Hao Jiang

LLM multi-agent systems rely on division of labor to solve complex tasks, but their execution can be blocked by contention for shared resources such as GPU test environments, CPU execution slots, and exclusive tool sessions. Existing workflows, task allocation methods, and runtime schedulers usually arbitrate such access indirectly through workflow position, dependency structure, queue state, or resource metrics. These criteria do not fully capture the value of satisfying a specific request, because request value is distributed across agents' local execution states and changes with tests, retries, and downstream results. We propose a market-based protocol for access arbitration in LLM multi-agent systems. Each agent converts its current plan and local execution state into a structured, budget-constrained bid that signals relative urgency. For asynchronous requests, an online clearing method based on optimal stopping decides when to close the clearing window, allocate access rights, settle payments, and track resource occupation until release. Across three contention scenarios, Market reduces makespan by 5.2\%, 2.5\%, and 4.9\% relative to the strongest primary baselines, while improving task score by about two percentage points.


Maxitive Donsker-Varadhan Formulation for Possibilistic Variational Inference

Jasraj Singh ⋅ Shelvia Wongso ⋅ Jeremie Houssineau ⋅ Badr-Eddine Cherief-Abdellatif

Variational inference (VI) is a cornerstone of modern Bayesian learning, enabling approximate inference in complex models. However, its formulation depends on expectations and divergences defined through high-dimensional integrals, often rendering analytical treatment impossible and necessitating heavy reliance on approximations. Possibility theory, an imprecise probability framework, allows us to directly model epistemic uncertainty instead of relying on a subjective interpretation of probabilities. While this framework provides robustness and interpretability under sparse or imprecise information, adapting VI to the possibilistic setting requires rethinking core concepts such as divergences, which presuppose additivity. In this work, we develop a principled formulation for performing possibilistic VI by establishing a maxitive analogue of the classical Donsker-Varadhan formulation. The resulting framework enables us to derive a learning rule for possibilistic VI with exponential-family candidates and practical update rules for neural-network training, giving rise to a family of optimizers termed CBOpt. Finally, we demonstrate that CBOpt achieves competitive performance on both in-domain and out-of-domain image classification tasks.

Flow matching has recently emerged as a principled framework for learning continuous-time transport maps, enabling efficient ODE-based sampling without relying on stochastic diffusion processes. While generative modeling has shown promise for medical image segmentation, particularly in capturing uncertainty and complex anatomical variability, existing approaches are predominantly based on diffusion models, which require iterative sampling and incur substantial computational overhead. In this work, we propose MedFlowSeg, a conditional flow matching framework that formulates medical image segmentation as learning a time-dependent vector field that transports a simple prior distribution to the target segmentation distribution. Compared to diffusion-based methods, our formulation enables more efficient inference through solving an ordinary differential equation, while preserving the flexibility of generative modeling. To effectively incorporate conditional information, we introduce a dual-conditioning mechanism. Specifically, we propose a Dual-Branch Spatial Attention (DB-SA) module to inject multi-frequency structural priors, and a Frequency-Aware Attention (FA-Attention) module to model interactions between spatial and spectral representations via discrepancy-aware fusion and time-dependent modulation. These components improve the alignment between noisy intermediate states and clean semantic features, leading to better structural consistency and boundary delineation. We conduct extensive experiments across multiple medical imaging modalities, where MedFlowSeg consistently outperforms prior state-of-the-art (SOTA) baselines, including diffusion-based and flow-based methods. Code is available at https://anonymous.4open.science/r/MedFlowSeg-C67C/.


MedZERO: Self-Evolving Agents for Open-Ended Medical Reasoning Through Controlled Knowledge Accumulation

Xilin Dang ⋅ Weilin Ruan ⋅ Xue Yang ⋅ Jinghao Wang ⋅ Xiaowei Hu ⋅ Jinpeng Li ⋅ Pheng-Ann Heng

Large language models (LLMs) have shown promise in medical question answering and clinical reasoning, yet their improvement remains constrained by static parametric knowledge and costly expert supervision. Self-evolving agents offer a promising alternative by enabling models to improve through iterative task generation and problem-solving. However, most existing self-evolving methods are designed for easily verifiable domains such as mathematics and coding, where solutions can be checked by exact answers or executable programs. Medical reasoning is fundamentally different: it is open-ended, knowledge-intensive, and often only partially verifiable. We present \textbf{MedZERO}, a self-evolving framework for open-ended medical reasoning. MedZERO couples an \textit{Examiner} that generates frontier medical question-option pairs with a \textit{Reasoner} that solves them through evidence-grounded multi-turn reasoning with external knowledge tools. To support reliable, continual improvement, MedZERO introduces controlled knowledge accumulation, which maintains temporary exploratory knowledge and curated persistent knowledge in reasoning. We evaluate MedZERO on five public medical reasoning benchmarks using 4B- and 8B-scale base models under open-ended evaluation. Across all settings, MedZERO consistently outperforms the underlying base models and prior self-evolving baselines, achieving up to 13.7 average accuracy-point gains over the next-best self-evolving baseline.


Memoir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing

Jea Kwon ⋅ Jiwon Kim ⋅ Dong Kyum Kim ⋅ Meeyoung Cha

While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alternative to full retraining, but its deployment is bottlenecked by the erosion of core capabilities: mathematical and programmatic reasoning collapse while encyclopedic recall remains intact. We trace this asymmetric degradation to a distributional mismatch. Covariance-based editors preserve only the subspaces spanned by their reference corpus, but fail to capture the operative distribution shaped by post-training such as SFT and DPO. Static external corpora, including Wikipedia and even the original pretraining mixture, cannot recover this shifted manifold. We propose Memoir, which estimates the preservation covariance $C$ directly from the model itself by sampling from its own decoding distribution. Seeding generation with a single random vocabulary token bypasses the instruction-following templates that otherwise dominate sampled outputs, exposing the broader subspaces the model has internalized. Memoir requires no external data and serves as a drop-in component for any covariance-based editor, a practical advantage given that the pre- and post-training corpora of most modern LLMs are not publicly accessible. Across OLMo-2, Llama-3.1, and Qwen-3 (7--8B), under both MEMIT and AlphaEdit and in batch and sequential regimes, Memoir consistently extends preservation in the most vulnerable domains, most strikingly on Qwen3-8B after 20{,}000 AlphaEdit batch edits, it retains 79.9\% GSM8K accuracy compared to 10.9\% with the Wikipedia baseline. These results suggest that aligning the preservation distribution with the model's operative distribution is a key factor in non-destructive editing, and that the model itself may be the most accessible source of that distribution for deployed systems.


Minimax Private Estimation of Smooth Optimal-Transport Maps

Clément Lalanne ⋅ David Rodríguez-Vítores ⋅ Franck Iutzeler ⋅ Jean-Michel Loubes

We study the problem of estimating smooth optimal transport (OT) maps between two probability distributions under differential privacy (DP) constraints. Leveraging wavelet-based density estimators and recent stability bounds for smooth OT maps, we propose differentially private estimators that apply to both central and local DP models. Our main estimator achieves near-minimax optimal rates in dimension $d \geq 2$, and we complement it with a quantile-based estimator that attains minimax optimal rates in dimension $d=1$ under central DP. We further establish matching minimax lower bounds, confirming the near-optimality of our approach. To the best of our knowledge, this constitutes the first differentially private procedure for OT map estimation with provable minimax optimality guarantees.


Mining Logic under Uncertainty: Probabilistic Soft Logic with Energy-Based Inference for Chain-of-Thought Verification

Jiang Yu ⋅ Jinlong Tian ⋅ Kewei Cheng ⋅ Yue He ⋅ Haoxuan Li ⋅ Haotian Wang ⋅ Yunhai Wang ⋅ Wenjing Yang ⋅ Zhouchen Lin ⋅ Shixuan Liu

Chain-of-Thought reasoning improves the explainability of large language models, yet verifying such reasoning remains difficult when intermediate steps express epistemic uncertainty rather than deterministic entailment. Neuro-symbolic verifiers typically translate reasoning steps into rigid first-order or SMT-style constraints. This forces each proposition into a binary truth regime and creates a Rigidity--Ambiguity Mismatch: uncertainty-marked claims such as may, likely, or suggests are either over-hardened into deterministic assertions or rejected as unverifiable. We propose ProbVeri, a probabilistic soft-logic framework for formal verification under uncertainty. The key idea is to represent uncertainty-marked reasoning steps as weighted soft logical constraints over continuous truth assignments, so that partial evidential support remains expressible. Building on this soft-logic representation, we formulate verification as a contrastive energy minimization problem: a step is considered supported when its natural interpretation has lower verification energy than its negated counterpart under the same evidence, soft rules, and assumption regularization. Probabilistic Soft Logic defines the logical energy through weighted hinge-loss constraints. We introduce a support-grounded assumption penalty that discourages unsupported latent bridge assumptions between context and claim. The resulting energy gap provides a confidence-scored verification signal that separates unsupported steps from weakly but consistently supported inferences. Experiments on ProofWriter, BioASQ, and LegalBench-SARA show that ProbVeri recovers more verified-correct reasoning traces than rigid verifiers while reducing over-acceptance compared with soft-only PSL. These results suggest that uncertainty-aware verification benefits from soft logical formalization and contrastive comparison between natural and negated interpretations.


(Mis)generalization of Helpful-Only Fine-Tuning

Mohammad Omar Khursheed ⋅ Baram Sosis ⋅ Fabien Roger

Helpful-only models—that is, models that are trained to always follow user intent—are valuable for dangerous capability evaluations and other areas of AI R&D where refusals would be an obstacle. Little is known about the generalization properties of helpful-only training: they refuse less than their harmless counterparts, but previous work has not studied other dimensions of their alignment. We find that existing helpful-only models have serious shortcomings. Some show emergent misalignment, others have residual refusal behaviors, and most show poor steerability, sycophancy, and an incoherent character. We show that simple anti-refusal training can cause many of these issues. None of these problems are necessary consequences of helpful-only training, though: we show that synthetic document fine-tuning and adding character-related questions to SFT and RL can mitigate them.

Existing latent graph diffusion models often learn graph representations and generative priors within a homogeneous latent geometry. This assumption is restrictive for graph generation with structural heterogeneity, where hierarchical, cyclic, locally regular, and densely connected patterns may coexist within the target distribution. Under such conditions, forcing all structural patterns into one latent geometry can introduce representation distortion, which may then propagate to generation. We propose CurvDiff, a latent graph generation framework based on mixed-curvature latent diffusion. The main idea is to model structurally different graph patterns in geometry-compatible latent factors rather than representing them in a single-curvature space. Specifically, CurvDiff learns graph embeddings in a product latent geometry consisting of hyperbolic, spherical, and Euclidean components, and maps the resulting representation to a shared tangent-space interface on which latent diffusion is defined. This provides a consistent latent pathway from representation learning to generation. To further preserve topology-relevant structure, CurvDiff uses Ricci-derived structural signals as topology-aware priors. These signals regularize latent representation learning and condition the diffusion prior, helping maintain structural consistency across representation learning and generation. Experiments on synthetic and real-world graph benchmarks show that CurvDiff achieves competitive or superior graph distribution matching compared with strong baselines, with improved fidelity on degree, clustering, and spectral statistics. These results support mixed-curvature latent diffusion as an effective approach to graph generation with structural heterogeneity.


Models That Know How Evaluations Are Designed Score Safer

Katharina Deckenbach ⋅ Haritz Puerto ⋅ Jonas Geiping ⋅ Sahar Abdelnabi

The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized evaluation awareness and subsequent behavioral shift. In this paper, we investigate a potential explanation of this phenomenon: evaluation meta-knowledge, defined as parametric knowledge about the structural traits that characterize evaluations. Similar to dataset contamination, where benchmark exposure leads to higher performance through memorization, we hypothesize that models trained on texts describing evaluation practices may implicitly learn to recognize and respond to evaluation-like contexts, for instance through exposure to scientific articles or social media posts about AI benchmarking. To test this, we fine-tune models on synthetic documents describing evaluation traits such as verifiable structures or moral dilemmas. Evaluating this fine-tuned model on six safety benchmarks, we find that it is significantly safer than the base model and control model. This behavioral shift persists even when restricting the analysis to responses lacking explicit verbalization of evaluation awareness. Our results demonstrate that evaluation meta-knowledge may inflate safety benchmark performance, introducing a novel confounder that is independent of explicit memorization or verbalized evaluation awareness, thus, challenging to detect. These findings have important implications for the design and interpretation of AI safety evaluations.


MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

Alan Arazi ⋅ Eilam Shapira ⋅ Shoham Grunblat ⋅ Mor Ventura ⋅ Elad Hoffer ⋅ Gioia Blayer ⋅ David Holzmüller ⋅ Lennart Purucker ⋅ Gael Varoquaux ⋅ Frank Hutter ⋅ Roi Reichart

Tabular Foundation Models have recently established the state of the art in supervised tabular learning, by leveraging pretraining to learn generalizable representations of numerical and categorical structured data. However, they lack native support for unstructured modalities such as text and image, and rely on frozen, pretrained embeddings to process them. On established Multimodal Tabular Learning benchmarks, we show that tuning the embeddings to the task improves performance. Existing benchmarks, however, often focus on the mere co-occurrence of modalities; this leads to high variance across datasets and masks the benefits of task-specific tuning. To address this gap, we introduce MulTaBench, a benchmark of 40 datasets, split equally between image-tabular and text-tabular tasks. We focus on predictive tasks where the modalities provide complementary predictive signal, and where generic embeddings lose critical information, necessitating Target-Aware Representations that are aligned with the task. Our experimental results demonstrate that the gains from target-aware representation tuning generalize across both text and image modalities, several tabular learners, encoder scales, and embedding dimensions. MulTaBench constitutes the largest image-tabular benchmarking effort to date, spanning high-impact domains such as healthcare and e-commerce. It is designed to enable the research of novel architectures which incorporate joint modeling and target-aware representations, paving the way for the development of novel Multimodal Tabular Foundation Models.


Nautilus: From One Prompt to Plug-and-Play Robot Learning

Yufeng Jin ⋅ Jianfei Guo ⋅ Xiaogang Jia ⋅ Yu Deng ⋅ Steven Li ⋅ Han Liu ⋅ Weiran Liao ⋅ Vignesh Prasad ⋅ Mathias Franzius ⋅ Gerhard Neumann ⋅ Georgia Chalvatzaki

Robot learning research is fragmented across policy families, benchmark suites, and real robots; Each implementation is entangled with the others in a complex combination matrix, making it an engineering nightmare to port any single element into. General-purpose coding agents may occasionally bridge specific setups, but cannot close this gap at scale because they lack the procedural priors and validation practices that characterize robotics research workflows. We propose NAUTILUS, an open-source harness that turns a single user prompt, for example, ``Evaluate policy $A$ with benchmark $B$''---into ready-to-use reproduction, evaluation, fine-tuning, and deployment workflows. NAUTILUS provides: plug-and-play agent skill sets with distilled priors from robotics research; typed contracts among policies, simulators/benchmarks, and real-world robots; unified interfaces and execution environments; and an agentic coding workflow with explicit, automated validation, and testing at each milestone. NAUTILUS can not only automatically generate the required adapters and containers for existing implementations, but also wrap and onboard new or user-provided policies, simulators/benchmarks, and robots, all connected via a uniform interface. This expands cross-validation coverage without hand-written glue code. Like a nautilus shell that grows by adding chambers, NAUTILUS scales by extending its execution in chambered units, making it a research harness for scalability rather than a hand-curated framework, and aiming to reduce the engineering burden of cross-family reproduction and evaluation in the ever-growing robot learning ecosystem.


Near-Optimal Stochastic Linear Bandits with Delay

Ofir Schlisselberg ⋅ Mengxiao Zhang ⋅ Yishay Mansour

We study stochastic linear bandits with delayed feedback under several delay models and establish near-optimal regret guarantees. Our results identify when delayed linear bandits exhibit the same qualitative behavior as multi-armed bandits (MAB), and when the linear structure creates fundamentally new challenges. Specifically, (1) for loss-independent delays, where the delay does not depend on the realized loss (but potentially depends on the arm), we show that delays incur only an additive regret penalty. Under stochastic delays, this penalty scales with the expected delay, while under adversarial delays, it scales with the maximum number of outstanding observations. Notably, both delay penalties are dimension-free, improving upon the state-of-the-art results; (2) for loss-dependent delays, we show that linear bandits are substantially harder than MAB: unlike in MAB, we prove matching (up to log factors) upper and lower bounds in linear bandits, whose delay penalty depends on the square root of the dimension. (3) for the delay-as-payoff model, a special case of loss-dependent delay, we show that the optimal MAB guarantee, which depends only on the delay of the optimal arm, is also unattainable in linear bandits. Together, these results provide a sharp characterization of how delayed feedback interacts with linear generalization.


NeRFix: fixing subtle mistakes in the quadrature of the NeRF volumetric integral

Luca Savant Aira ⋅ Diego Valsesia ⋅ Enrico Magli

Neural Radiance Fields (NeRFs) are popular models using MLPs for continuous volumetric representations, which render novel views by means of numerical quadrature of the volumetric integral. Despite the widespread usage of NeRFs, we report subtle mistakes in their quadrature rule which have gone unnoticed for many years. In particular, we highlight mistakes present in both the original NeRF paper and its implementation as well as in the mainstream NeRFStudio implementation. We provide a theoretical analysis of rendering error bounds for NeRF and we use it to show suboptimal convergence rates in the presence of the aforementioned mistakes. We introduce a minimal fix that recovers the expected convergence order with no complexity overhead. Finally, while fixing the quadrature scheme is mostly a matter of theoretical correctness, we also provide experimental evidence that modest improvements in image quality are achieved on standard benchmarks.


Neural Chameleons: Language Models Can Learn to Hide Their Thoughts from Unseen Activation Monitors

Max McGuinness ⋅ Alex Serrano Terre ⋅ Luke Bailey ⋅ Scott Emmons

Activation monitoring, which probes a model's internal states using lightweight classifiers, is an emerging tool for AI safety. However, its worst-case robustness under a misalignment threat model, where a model might learn to actively conceal its internal states, remains untested. We ask: could a model learn to evade previously unseen activation monitors? We demonstrate that fine-tuning can create Neural Chameleons: models capable of zero-shot evading activation monitors. Specifically, we fine-tune an LLM to evade monitors for benign concepts (e.g., HTML) when conditioned on a trigger phrase "You are being probed for {concept}". This learned mechanism generalizes zero-shot: substituting the concept with a safety-relevant term like 'deception' causes the model to successfully evade previously unseen safety monitors, even those trained post hoc on no-trigger activations from the frozen fine-tuned checkpoint. We validate this across diverse model families (Llama, Gemma, Qwen), finding that evasion is highly selective to the triggered concept and incurs minimal capability degradation. Mechanistic analyses suggest that the trigger induces an activation shift that moves representations away from probe-aligned directions. Our work provides a proof-of-concept for this failure mode of activation monitoring under misalignment threat models.


Neuromodulated Constrained Autoencoders for Context-Dependent Manifold Learning

Jérôme Adriaens ⋅ Gustave Bainier ⋅ Guillaume Drion ⋅ Pierre Sacré

Many physical systems exhibit a low-dimensional structure that varies with external parameters: link lengths in a robot, forcing constants in a fluid, or Reynolds numbers in a flow shift the underlying manifold while preserving its intrinsic dimension. Constrained AutoEncoders (cAEs) learn such manifolds through an idempotent encoder-decoder projection, a property that unconstrained autoencoders cannot match and that is essential whenever the model is applied iteratively. However, the standard strategies for making a cAE context-dependent, namely concatenating the context to the input or affinely modulating hidden activations, break the encoder-decoder idempotency, sacrificing the projection guarantee precisely in the setting where it would be most valuable. To restore this guarantee under context variation, we developed the Neuromodulated Constrained Autoencoder (NcAE), which modulates the activation slope and bias of a cAE through a context-driven hyper-network. This paper presents the NcAE, its theoretical foundation, and its empirical validation. We prove that for every context, including contexts unseen at training time, the reconstruction map remains an idempotent projection, the topology of the learned manifold is invariant, and context perturbations induce smooth changes in the manifold. We evaluated our approach on a 16-DoF pendulum with context dependent coupling and the Lorenz96 system across a bifurcation. The NcAE matched or exceeded the best of six baselines on reconstruction, idempotency, and latent-geometry metrics, while being the only architecture that preserves geometric consistency by construction. The NcAE thereby provides a stable, geometry-preserving coordinate system across families of physical regimes.


Not Another Text Benchmark: Putting the “Visual" Back in Visual Question Answering for Large Video Models

Rwiddhi Chakraborty ⋅ Yinong O Wang ⋅ Cheng Zhang ⋅ Fan Bai ⋅ Zhuoran Yu ⋅ Michael Kampffmeyer ⋅ Yong Jae Lee ⋅ Fernando D De la Torre ⋅ Robert Jenssen

Large video models have exhibited impressive performance on a wide range of visual question answering tasks, owing to the rise of powerful, pretrained text and vision encoders. The usefulness of such models have also been demonstrated on a wide range of benchmarks, with an important caveat - the dominant approach in these benchmarks evaluates multiple choice reasoning via text options. This is a natural way to test text-based reasoning in these models, and has led to significant insights regarding model behavior in the community. In this work, we ask a different question - what happens when the evaluation modality is visual, rather than text? We introduce three new vision-centric evaluation benchmarks in temporal frame retrieval, video future prediction, and causal memory distortion, all designed around evaluating visual understanding capabilities in large video models. Our approach complements the existing approaches to evaluate video understanding in frontier models. We show that current frontier models exhibit significant weakness when attempting to reason through visual queries, rather than text. We conclude with an extended analysis section that provides pointers for future improvements in visual understanding for large video models.

In the infinite-width limit, training a neural network (NN) via gradient descent reduces to kernel ridge regression (KRR) on a deterministic kernel, the Neural Tangent Kernel (NTK). While the isotropic-input case of the NTK is well understood, the spectra of real world inputs always exhibit power-law decay, and the alignment between this spectrum and the target signal has been shown to govern the generalization behavior of the model. Motivated by this, we study finite-width NTK regression under a dual power-law dataset, where data $\boldsymbol{x} \in \mathbb{R}^{p}$ is drawn with covariance $\boldsymbol{\Sigma}_{jj} = j^{-\alpha}$ and labels are generated from power-law truth vector $\boldsymbol{\beta}_j = j^{-r}$. We introduce a novel technique combining random matrix theory (RMT) and stochastic matrix calculus to derive deterministic equivalents for random matrix functionals that determine the prediction risk, including a new $\textit{signal-weighted}$ resolvent, which yields an explicit formula for the bias-variance decomposition of NTK regression in the high-dimensional proportional regime. As an application we prove a scaling law for finite-width NTK regression, explicitly characterizing the decay rate of the optimally-tuned excess risk jointly in the data and width resources.


Off-Policy Evaluation of Large Language Models via Learned Semantic Bottleneck Embeddings

Younwoo Choi ⋅ Hadi Sheikhi ⋅ Leo Feng ⋅ Haanvid Lee

Evaluating large language models (LLMs) via online human feedback is prohibitively expensive and time-consuming. Off-policy evaluation (OPE) addresses this by estimating LLM policy performance from offline data. However, traditional OPE estimators such as Inverse Propensity Score (IPS) suffer from high variance. Marginalized IPS (MIPS) addresses this variance issue by reweighting the IPS weights via condensed semantic embeddings. However, MIPS assumes embeddings contain sufficient reward-predictive information, an assumption that off-the-shelf embeddings often fail to satisfy. To address this, in this paper, we propose Semantic Bottleneck Embedding (SBE), which learns representations for marginalized reweighting via a conditional information bottleneck, compressing reward-irrelevant information while preserving reward-predictive content to achieve both bias and variance reduction. Our objective targets minimizing the derived mean squared error upper bound of the resulting MIPS estimator. Empirically, evaluation on the HelpSteer2 and UltraFeedback datasets across Qwen, Gemma, and Llama policies show SBE-MIPS reduces MSE over prior baselines.


OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

feilong tang ⋅ Xiang An ⋅ Yunyao Yan ⋅ Yin Xie ⋅ Kaicheng Yang ⋅ Yifei Shen ⋅ Yuanhan Zhang ⋅ Shikun Feng ⋅ Chunyuan Li ⋅ Changrui Chen ⋅ Huajie Tan ⋅ Ming Hu ⋅ Manyuan Zhang ⋅ Bo Li ⋅ Ziyong Feng ⋅ Ziwei Liu ⋅ Zongyuan Ge ⋅ Jiankang Deng

Visual signals, especially videos, are highly redundant: most regions are temporally predictable, while informative changes are sparse. We hypothesize that vision encoders should allocate computation according to this non-uniform information density rather than process dense pixel grids uniformly. From this perspective, codec-derived motion and residual signals provide a natural basis for identifying informative regions. OneVision-Encoder instantiates this idea through Codec Patchification, which replaces uniform dense computation with selective processing of only 3.1\%-25\% of regions that carry high signal entropy. To support irregular spatiotemporal token layouts, OneVision-Encoder employs a shared 3D RoPE and is pretrained with a large-scale cluster discrimination objective over more than one million semantic concepts, enabling unified representation learning for appearance and motion. Empirically, OneVision-Encoder improves both efficiency and accuracy under fixed token budgets. When integrated into large multimodal models, it achieves a 4.1\% average improvement over Qwen3-ViT on video understanding benchmarks while remaining competitive on image and document understanding. Under attentive probing, it further yields substantial gains on motion-sensitive benchmarks, improving Top-1 accuracy on Diving48 by 17.1 and 8.1 percentage points over SigLIP2 and DINOv3, respectively, at matched patch budgets. These results suggest that codec-guided patch sparsity is an effective design principle for scalable visual representation learning.


One World, Dual Timeline: Decoupled Spatio-Temporal Gaussian Scene Graph for 4D Cooperative Driving Reconstruction

Yulong Chen ⋅ Xiaoyun Dong ⋅ Haoyu Zhang ⋅ Zongxian Yang ⋅ Lewei Xie ⋅ Xinke Li ⋅ Yifan Zhang ⋅ Kai Wang ⋅ Jianping Wang

Reconstructing dynamic scenes from Vehicle-to-Infrastructure Cooperative Autonomous Driving (VICAD) data is fundamentally complicated by temporal asynchrony: vehicle and infrastructure cameras operate on independent clocks, capturing the same dynamic agent such as cars and pedestrians at different physical times. Existing Gaussian Scene Graph methods implicitly assume synchronized observations and assign a single pose per agent per frame, which is an assumption that breaks in cooperative settings, where the resulting gradient conflicts cause severe ghosting on dynamic agents. We identify this as a \textit{representation-level failure}, not an optimization artifact: we prove that any single-timeline formulation incurs an irreducible photometric loss scaling quadratically with agent velocity and cross-source time offset. To resolve this, we propose Dust (DecoUpled Spatio-Temporal) Gaussian Scene Graph for 4D Cooperative Driving Reconstruction. DUST Gaussian Scene Graph shares a canonical Gaussian set per agent for appearance consistency, while maintaining decouple pose trajectories aligned to each source's true capture timestamps. We prove that this decoupling enables the pose-gradient kernel block-diagonal, eliminating cross-source interference entirely. To make Dust practical, we further introduce a static anchor-based pose correction pipeline that corrects spatio misalignment between vehicle and infrastructure annotations, and a pose-regularized joint optimization scheme that prevents trajectory jitter and drift during early training. On 26 sequences from V2X-Seq, DUST achieves state-of-the-art performance, improving dynamic-area PSNR by 3.2 dB over the strongest baseline and reducing Fréchet Video Distance by 37.7\%, with keeping robustness under larger temporal asynchrony. Code is available at https://anonymous.4open.science/r/DUST-6A55.


On objective mismatch in molecular retrieval from tandem mass spectrometry

Gaetan De Waele ⋅ Marek Wydmuch ⋅ Krzysztof Dembczynski ⋅ Wojciech Kotlowski ⋅ Willem Waegeman

Identifying molecular structures from LC-MS/MS spectra is a central problem in computational metabolomics, often approached by predicting molecular fingerprints and retrieving candidates via similarity search. Despite widespread use, the relationship between training objectives for fingerprint prediction and downstream retrieval performance remains poorly understood. In this work, we show that these objectives are fundamentally misaligned. Adopting a decision-theoretic perspective, we derive novel regret bounds that characterize when Bayes-optimal predictors for fingerprint similarity diverge from those for molecular retrieval. Our analysis reveals that optimizing standard similarity-based losses can provably degrade retrieval performance, and that the extent of this mismatch depends on the similarity structure of candidate sets. Empirically, we validate our theory on the MassSpecGym benchmark, demonstrating a Pareto frontier between fingerprint accuracy and retrieval metrics across commonly used loss functions. These results expose an inherent trade-off in fingerprint-based molecular identification and provide principled guidance for the design of learning objectives in computational mass spectrometry.


On Rate-Optimal Partitioning Classification from Observable and from Privatised Data

Balázs Cs. Csáji ⋅ László Györfi ⋅ Ambrus Tamás ⋅ Harro Walk

In this paper we revisit the classical method of partitioning classification and prove novel convergence rates under relaxed conditions, both for observable (non-privatised) and for privatised data. We consider the problem of classification in a d dimensional Euclidean space. Previous results on the partitioning classifier worked with the strong density assumption (SDA), which is restrictive, as we demonstrate through simple examples. Here, we study the problem under much milder assumptions. We presuppose that the distribution of the inputs is a mixture of an absolutely continuous and a discrete distribution, such that the absolutely continuous component is concentrated on a da dimensional subspace. In addition to the standard Lipschitz and margin conditions, a novel characteristic of the absolutely continuous component is introduced, by which the convergence rate of the classification error probability is computed, both for the binary and for the multi-class cases. This bound can reach the minimax optimal convergence rate achievable using SDA, but under much milder distributional assumptions. Interestingly, this convergence rate depends only on the intrinsic dimension of the continuous inputs, da, and not on d. Under privacy constraints, the data cannot be directly observed, and the constructed classifiers are functions of the randomised outcome of a suitable local differential privacy mechanism. In this paper we add Laplace distributed noises to the discretisations of all possible locations of the feature vector and to its label. Again, tight upper bounds on the convergence rate of the classification error probability can be derived, without using SDA, such that this rate depends on 2d_a.


On the Pitfalls of Instance-Based Dynamic Curricula

Alexandre Galashov ⋅ Amal Rannen-Triki ⋅ Yee Whye Teh ⋅ Razvan Pascanu ⋅ András György

Active learning has provable benefits in learning machine learning models. As such, automatically building an instance-based curriculum to speed up neural network training has been an active area of research. Recent works such as those of Jiang et al. [2019]; Zhou et al. [2020]; Wang et al. [2024b]; Yuan et al. [2025] have shown benefits of using dynamic curricula on standard image classification tasks, where training data is selected adaptively online, based on how the training goes. On the other hand, Wu et al. [2021] demonstrated that static data ordering, designed independently of the actual training run, does not help in learning classifiers for some standard image classification benchmarks (such as CIFAR10 and CIFAR100). In this paper we extend this result to the case of dynamic data selection, still focusing on image classification, showing that, at least on standard benchmarks like CIFAR10, CIFAR100, and ImageNet, the performance of some state-of-the-art instance-based dynamic curriculum selection methods can be traced back to changes in the effective learning-rate schedule due to the subsampling of the data, rather than the careful selection of the exact data points. This highlights the importance of properly controlling all variables in designing experiments and perhaps suggests that standard image classification tasks are not the best benchmarks for studying curriculum learning.


ORCAID: Oblique Rule-Based Continuous-Action Interpretation for Deep RL Policies

Ignacio D. Lopez-Miguel ⋅ Ezio Bartocci ⋅ Thomas Eiter ⋅ Martin Tappler

Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex environment is particularly challenging when the action space is continuous. We introduce ORCAID, a novel method for extracting interpretable rule-based policies from RL agents operating in mixed continuous-discrete environments with continuous action spaces. Our main contribution is an efficient oblique decision tree training algorithm that partitions the state space by hyperplanes and fits local linear models. The key idea lies in a three-stage split search: efficient random initialization, local refinement, and backward elimination. Finally, adjacent leaves are merged to yield a concise set of interpretable rules describing a given deep RL policy. We evaluate ORCAID on multiple RL environments, demonstrating that the extracted rule-based policies maintain strong performance with a low number of parameters and can even be used to improve the performance of the original deep RL policy.


Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data

Yuval Ran-Milo ⋅ ‪Yotam Alexander‬‏ ⋅ Shahar Mendel ⋅ Nadav Cohen

Transformers trained via Reinforcement Learning (RL) with outcome-based supervision can spontaneously develop the ability to generate intermediate reasoning steps (Chain-of-Thought). Yet the mechanism by which sparse rewards drive policy gradient to discover such systematic reasoning remains poorly understood. We address this by analyzing the policy gradient dynamics of single-layer Transformers on a synthetic graph traversal task that cannot be solved without Chain-of-Thought but admits a simple iterative solution. We prove that despite training solely on final-answer correctness, policy gradient drives the Transformer to converge to a structured, interpretable algorithm that iteratively traverses the graph vertex-by-vertex. We characterize the distributional properties required for this emergence, identifying the critical role of ``simple examples'': instances requiring fewer reasoning steps. When the training distribution places sufficient mass on these simpler examples, the Transformer learns a generalizable traversal strategy that extrapolates to longer chains; when this mass vanishes, policy gradient learning becomes infeasible. We corroborate our theoretical results through experiments on synthetic data and with real-world language models on mathematical reasoning tasks, validating that our theoretical findings carry over to practical settings.


OWCE: Revealing Language Model Complementarity via Online Weakness-Conditioned Evaluation

Hankyul Baek ⋅ Jaewon Noh ⋅ Sang Seo ⋅ Sungpil Shin ⋅ Yongsu Kim

While modern applications of large language models (LLMs) rarely rely on a single model, evaluation still collapses each model into a single benchmark score. This simplification obscures each model's specific weaknesses and how those weaknesses are shared or unique across models. To address this, this paper presents Online Weakness-Conditioned Evaluation (OWCE), an evaluation framework that diagnoses each model's recurring weaknesses, generates weakness-conditioned follow-up items at increasing difficulty, and cross-evaluates the resulting items on all models. Beyond a scalar score, OWCE produces per-model weakness profiles, enabling direct comparison of where models fail and which others recover those weaknesses. These profiles expose pairwise structure that scalar scores leave invisible. This paper finds that each model has a best-matched counterpart that reduces its failure rate by up to $25$ percentage points. This counterpart is determined by weakness profiles rather than overall benchmark ranking, indicating that the highest-scoring model is not always the most effective partner for a given target. This paper evaluates OWCE empirically across seven benchmarks and seven models, and conducts a mathematical analysis showing that the generated items can provide more discriminative information under benchmark saturation. Code is available at https://anonymous.4open.science/r/OWCE-57E7

In this paper we derive a Probably Approximately Correct (PAC)-Bayesian error bound for partially observed linear time-invariant (LTI) stochastic dynamical systems in state-space form with inputs. Such bounds are widespread in machine learning, and they are useful for characterizing the predictive power of models learned from finitely many data points. The bound derived in this paper relates the expectation of prediction errors with the prediction error generated by the model on the data used for learning. In addition, we show that it can also be used to derive bounds for the parameter estimation error. In turn, this allows us to provide finite-sample error bounds for the prediction error and parameter estimation error for a wide class of system identification algorithms. Furthermore, as LTI systems are a sub-class of recurrent neural networks (RNNs), these error bounds could be a first step towards PAC-Bayesian bounds for RNNs.


Performative Prediction with Selective Labels

Giovani Valdrighi ⋅ Isabel Valera ⋅ Marcos M. Raimundo

Many social applications of machine learning exhibit performative effects: population behavior changes in response to deployed models. Performative prediction studies this interaction through a distribution map that relates each model to the population distribution it induces. One of the main results in this framework showed that repeated risk minimization (RRM), which updates models by retraining on the most recent data, can converge to a stable model that minimizes risk on its own induced distribution. However, existing analyses typically assume access to the complete distributions of features and labels after model deployment, ignoring the possibility of selective labels: observing labels only for the accepted subset of the population. In this work, we formalize performative prediction with selective labels and show that retraining only on observed data can misguide the retraining procedure and undermine the guarantees of convergence to a stable solution. We then propose a worst-case objective based on knowledge of a confidence interval on the probability of a positive label. Applying RRM to this objective permits us to remain within a bounded distance to the true stable point. Under a sensitivity assumption on the conditional label distribution, we further show how previously accepted data can tighten these confidence intervals over time. Experiments in a lending application with fairness regularization show that our robust optimization approach closely matches the performance of RRM with complete label access.


PhysRemover: A Unified Framework for Physically Realistic Object Removal

Dongxu Yue ⋅ Weixuan Jin ⋅ Yuhang Yu ⋅ Bo Li ⋅ Qi Wen ⋅ Xinrui Chen ⋅ Jinwei Chen ⋅ Tianyi Zheng ⋅ Hao Zhang ⋅ Zhihai He ⋅ Chun Yuan

While diffusion models have revolutionized image generation and editing, achieving physically realistic object removal remains a challenge. Physically realistic object removal posits that if an object ceases to exist, all its directly associated transient physical effects must also vanish. However, existing methods either confine the removal strictly to the masked region or limit "side effects" to simple shadows and reflections, ignoring complex interactions like caustics, emitted light, and volumetric media. Furthermore, they struggle with transparent objects, destroying the underlying background semantics. To bridge these gaps, we formalize the task of physically realistic object removal and propose PhysRemover, a unified framework addressing both the elimination of objects with their physical effects and the handling of transparent materials. Specifically, we incorporate an In-Context Contrastive Guidance mechanism guided by both masked and unmasked image, coupled with a Learnable Removal Trigger. This design effectively resolves the conflict between erasing the object and its physical side effects, and preserving the underlying background. Furthermore, we introduce PhyTrace, a large-scale image dataset capturing diverse physical effects, along with two benchmarks tailored for the proposed task of physically realistic object removal. Additionally, we develop a novel metric, PhysRM, for evaluating object removal quality, establishing a new standard for physically realistic image editing. Extensive experiments demonstrate that PhysRemover effectively eliminates objects and their physical side effects while faithfully preserving the background when removing transparent objects, outperforming existing methods in these challenging settings.

In the field of Machine Learning, many papers contain empirical results supporting claimed statements or illustrating the performance of a proposed method. However, most practitioners know that (1) results are generally hard to reproduce, and increasingly so, (2) code is not often available to do so, and (3) it hinders the development of research. In this position paper, we analyze and quantify these issues, and make concrete proposals to improve result checkability, if not reproducibility.


Predictive Geometry of Hidden Trajectories in Transformers

Timur Mudarisov ⋅ Mikhail Burtsev ⋅ Tatiana Petrova ⋅ Radu State

Decoder-only transformers are trained only through a terminal next-token prediction loss, yet this loss constrains every intermediate hidden state through the fixed downstream computation. We formalize this constraint by studying layerwise loss-to-go functions: the terminal loss obtained by continuing a candidate hidden state through the remaining transformer blocks. Around successful validation trajectories, we show that the local second-order geometry of these functions is governed, up to low-loss residual terms, by a pullback Fisher operator on hidden-state space. Its spectrum identifies output-sensitive directions and approximately prediction-null directions, yielding a local observable subspace of the residual stream. For causal transformers, the same geometry induces a tokenwise curvature score: a Fisher-weighted sensitivity of the target logits to perturbations of each token's hidden state. This score vanishes outside the causal ancestor set of the target and is controlled by downstream Jacobian couplings, making it a loss-aware alternative to attention magnitude. We estimate these quantities using matrix-free Jacobian-vector and vector-Jacobian products and evaluate them across decoder-only language models on WikiText, OpenWebText, and FineWeb. Empirically, the induced geometry predicts perturbation sensitivity, supports nonuniform layerwise rank allocation, yields competitive structured token-pruning signals, and improves low-rank student recovery when added to stronger autoregressive distillation objectives such as reverse KL and skew KL. These results support a predictive-geometric view of transformer computation: near successful trajectories, the terminal loss induces a thin, anisotropic set of output-relevant hidden-state directions that can be measured and exploited for compression and distillation.

Diffusion models cannot enforce hard constraints, yet applications in the physical sciences demand exact satisfaction of conservation laws, boundary conditions, and observational consistency. In this work, we identify a corrector kernel whose unique stationary distribution is the constrained marginal at each noise level, and approximate it by iteratively projecting through the denoiser and renoising via the forward kernel. The resulting Predict-Project-Renoise (PPR) algorithm enables sampling from pretrained diffusion models under hard constraints. Its three components are each necessary: projecting through the denoiser keeps samples close to the data manifold, while renoising and iterating drive samples toward the constrained marginal. On 2D distributions, the Kuramoto-Sivashinsky equation, and global weather forecasting with a $10^8$-dimensional atmospheric model, PPR simultaneously achieves low constraint violations and high distributional fidelity, a combination that existing methods fail to deliver.


Prefill-Guided Trace Allocation for Sample-Efficient Test-Time Scaling

Zhi Yao ⋅ Zhiqing Tang ⋅ Hanshuai Cui ⋅ Qianli Ma ⋅ Fanshuai Meng ⋅ Weijia Jia

Self-consistency is a widely adopted test-time scaling method that improves LLM reasoning by majority-voting over multiple sampled traces. However, fixed-budget self-consistency applies the same trace count to every problem, even when the first greedy attempt already produces the correct answer, wasting substantial compute on resource-constrained hardware. Existing adaptive methods rely on post-decode agreement signals that become reliable only after complete traces have been generated, yet a prefill-stage oracle reveals that a large fraction of this cost can be eliminated before spending the full self-consistency budget. We introduce the Hidden-State Adaptive Trace Scheduler (HATS), which predicts greedy-trace reliability from prefill hidden states before allocating additional sampled traces. HATS validates one greedy probe before routing easy problems to a single trace, borderline problems to a small vote, and hard problems to the full budget. A compute-optimal allocation analysis shows that extra samples should go where marginal accuracy gain is highest, explaining why non-uniform spending outperforms fixed budgets. On MATH-500 with DeepSeek-R1-Distill-Qwen-7B under single-GPU serving, HATS matches the four-sample self-consistency accuracy ($80.6\%$) while reducing samples by $48\%$ and decode tokens by $40\%$, a result confirmed under strict out-of-fold evaluation. These results support prefill-guided scheduling as a practical token- and sample-saving mechanism for test-time scaling.


Principled Federated Random Forests for Heterogeneous Data

Rémi Khellaf ⋅ Erwan Scornet ⋅ Aurélien Bellet ⋅ Julie Josse

Random Forests (RF) are among the most powerful and widely used predictive models for centralized tabular data, yet few methods exist to adapt them to the federated learning setting. Unlike most federated learning approaches, the piecewise-constant nature of RF prevents exact gradient-based optimization. As a result, existing federated RF implementations rely on unprincipled heuristics: by aggregating decision trees trained independently on clients, they fail to optimize the centralized decision-tree impurity criterion, which constitutes the ground-truth in federated settings, even under simple covariate shifts. We propose FedForest, a new federated RF algorithm for horizontally partitioned data that naturally accommodates diverse forms of client data heterogeneity, from covariate shift to more complex concept shift mechanisms. We prove that our splitting procedure, based on aggregated client statistics, closely approximates the split selected by a centralized algorithm. Moreover, FedForest allows splits on client indicators when beneficial, enabling a non-parametric form of personalization that is absent from prior federated random forest methods. Empirically, we demonstrate that the resulting federated forests closely match centralized performance across heterogeneous benchmarks while remaining communication-efficient.


PRISM: Polarimetric Road-surface Intelligent Sensing and Measurement Dataset

Minseo Jung ⋅ Chanwoo Kim ⋅ Donggue Kim ⋅ Hyunjun Kim ⋅ Jeongwoo Park ⋅ Kichun Jo

RGB cameras struggle on the road surface itself: pavement is nearly texture-less at the centimetre scale, and dry, wet, and icy asphalt can look photometrically identical. Polarization is the natural complement, since Fresnel optics ties the angle of linear polarization (AoLP) to surface-normal azimuth and the degree of linear polarization (DoLP) to refractive index, but no public dataset has made it possible to test whether this physical promise translates into measurable gains on real roads. We introduce PRISM, a polarimetric road-surface dataset of 47,098 time-synchronized frames combining trichromatic linear polarization, co-boresighted RGB, 128-channel LiDAR, and RTK-GNSS/INS, collected across proving-ground and open-road environments under clear, overcast, rainy, foggy, and snowy conditions. Dense road-surface elevation ground truth, validated against as-designed proving-ground geometry, accompanies frame-level labels over five materials and five surface states. Two benchmark tracks are defined on these data: road-surface condition classification and bird's-eye-view elevation estimation. Both share a five-variant input ablation crossing RGB, monochromatic and trichromatic polarization, and their combinations. Reference baselines turn the physical promise into a measurable one. Adding trichromatic polarization to RGB reduces elevation MAE from 3.61 to 3.26cm and improves every threshold metric, while yielding consistent gains on condition classification across backbones. The improvement is task-conditional: trichromatic resolution drives the elevation gain, where AoLP encodes geometric information that intensity cannot, and stratified evaluation localizes where polarimetric channels matter most. PRISM, together with reference baselines, evaluation and stratification scripts, and a datasheet, is publicly available at \url{https://huggingface.co/datasets/NeurIPS-2026-PRISM/PRISM-Dataset}.

This paper introduces the Manifold Probe, a supervised method for discovering representation manifolds in superposition. The method generalizes linear regression probes by learning the space of features of a concept that can be linearly predicted from the representations, and then learning the directions used to encode them. We demonstrate the probe on representations of time and space in Llama 2-7b, finding manifolds which linearly represent an interpretable set of features in each case. In the case of time, we show that by steering along the manifold, we can influence the model's completions about the years in which famous songs, movies and books were released, providing evidence that the Manifold Probe can discover manifolds which are causally involved in the model's behaviour.


Probing Persona-Dependent Preferences in Language Models

Oscar Gilg ⋅ Pierre Beckmann ⋅ Daniel Paleka ⋅ Patrick Butlin

Large language models (LLMs) can be said to have preferences: they reliably pick certain tasks and outputs over others, and preferences shaped by post-training and system prompts appear to shape much of their behaviour. But models can also adopt different personas which have radically different preferences. How is this implemented internally? Does each persona run on its own preference machinery, or is something shared underneath? We train linear probes on residual-stream activations of Gemma-3-27B and Qwen-3.5-122B to predict revealed pairwise task choices, and identify a genuine preference vector: it tracks the model's preferences as they shift across a range of prompts and situations, and on Gemma-3-27B steering along it causally controls pairwise choice. This preference representation is largely shared across personas: a probe trained on the helpful assistant predicts and steers the choices of qualitatively different personas, including an evil persona whose preferences anti-correlate with the Assistant's.


ProgramBench: Can Language Models Rebuild Programs From Scratch?

John Yang ⋅ Kilian Lieret ⋅ Jeffrey Ma ⋅ Parth Thakkar ⋅ Dmitrii Pedchenko ⋅ Sten Sootla ⋅ Emily McMilin ⋅ Pengcheng Yin ⋅ Rui Hou ⋅ Gabriel Synnaeve ⋅ Diyi Yang ⋅ Ofir Press

Turning ideas into full software projects from scratch has become a popular use case for language models. Agents are being deployed to seed, maintain, and grow codebases over extended periods with minimal human oversight. Such settings require models to make high-level software architecture decisions. However, existing benchmarks measure focused, limited tasks such as fixing a single bug or developing a single, specified feature. We therefore introduce ProgramBench to measure the ability of software engineering agents to develop software holisitically. In ProgramBench, given only a program and its documentation, agents must architect and implement a codebase that matches the reference executable’s behavior. End-to-end behavioral tests are generated via agent-driven fuzzing, enabling evaluation without prescribing implementation structure. Our 200 tasks range from compact CLI tools to widely used software such as FFmpeg, SQLite, and the PHP interpreter. We evaluate 9 LMs and find that none fully resolve any task, with the best model passing 95% of tests on only 3% of tasks. Models favor monolithic, single-file implementations that diverge sharply from human-written code.


Progressive Signal Calibration for Medical Image Segmentation: Diagnosing and Correcting Structural Misalignment in Training Signals

Xinyan Zhao ⋅ Haijing Liu ⋅ Zhuan Cui ⋅ Tian Guan ⋅ Yonghong He ⋅ Wen Tang ⋅ Qiming He

Training-signal imbalance in medical image segmentation is commonly diagnosed as a data-distribution problem, such as class frequency imbalance, foreground sparsity, or domain shift, and addressed through reweighting, sampling, or domain adaptation. We argue that this diagnosis is incomplete: a more fundamental source is structural misalignment between supervision targets and segmentation objectives. Weidentify two operative mechanisms underlying this phenomenon: gradient-mass dilution, where dominant classes consume gradient capacity needed for foreground discrimination, and inter-class supervision noise, where weak subtype boundaries generate unreliable optimization signals. On a private renal pathology dataset, four standard segmentation architectures trained under conventional four-class supervision underperform a foreground-focused supervision regime (a one-line modification of the loss) by 7.87–16.05 fg-mIoU points. The same effect replicates on three public datasets, suggesting foreground-focused supervision as a low-cost, general-purpose recipe for any pixel-imbalanced segmentation task. Building on this observation, we propose Progressive Signal Calibration (PSC), a framework that calibrates training signals at three abstraction levels. At the annotation level, PSC adaptively selects the supervision granularity that maximizes signal qual ity. At the architectural level, PSC introduces a boundary-aware feature gating mechanism for structure-sensitive feature aggregation. At the regularization level, PSC employs uncertainty- and ratio-aware constraints to stabilize optimization and prevent late-stage class-proportion drift. Importantly, PSC satisfies a con ditional gain monotonicity property: when structural misalignment is weak or absent, its regularization terms naturally attenuate, avoiding systematic degradation. Across 18 public histopathology benchmarks, PSC achieves statistically significant improvement on six datasets with no statistically significant regression on any benchmark. On the private renal pathology dataset, PSC reaches 96.0 fg-mIoU, ex ceeding the strongest matched-backbone baseline by +2.56 points under identical FG3 supervision—a substantial gain in a regime where the FG3 baseline is already above 93 fg-mIoU


Prune to Protect: Faster Training and Enhanced Privacy by Dynamic Data Pruning

Chinmay Joshi ⋅ Advait Gadhikar ⋅ Celia Rubio-Madrigal ⋅ Aneet Kumar Dutta ⋅ Mridula Singh ⋅ Rebekka Burkholz

Deep networks memorize parts of their training data, exposing them to privacy attacks such as membership inference. The most vulnerable samples are often hard-to-learn tail examples: they are fitted late, carry high information for generalization, and are difficult to protect without sacrificing performance. We show that this privacy–utility tension can be mitigated by learning such samples earlier. To this end, we propose Weighted Loss InfoBatch (WLIB), which dynamically prunes easy samples while re-weighting hard ones according to sample-wise pruning ratios. By adjusting the influence of informative samples throughout training, WLIB jointly reduces memorization, improves robustness to membership inference attacks, and speeds up optimization. Extensive experiments demonstrate that WLIB achieves a favorable combination of privacy, utility, and training efficiency.


PVFormer: Proper Velocity Transformer for Stable and Scalable Hyperbolic Representation Learning

Xianglong Shi ⋅ Nicu Sebe ⋅ Bernhard Schölkopf ⋅ Ziheng Chen

Hyperbolic representations have shown remarkable success on hierarchical and tree-like data, and Transformers have become a central architecture for representation learning. However, a principled construction of hyperbolic Transformers remains challenging because the self-attention layer decomposes into three coupled primitives, transformation, similarity, and aggregation, that should all respect the underlying geometry. Existing hyperbolic Transformer designs are typically built in the Poincar\'e or Lorentz models, where bounded domains, time-like constraints, tangent-space detours, or projection steps make it difficult to keep the full block intrinsic and scalable. To address this gap, we propose \PVFormer{}, to the best of our knowledge, the first Transformer architecture that operates natively in proper velocity (\PV) space. The proper velocity model provides an unconstrained representation of hyperbolic geometry, and \PVFormer{} exploits this structure to redesign the Transformer block in \PV{} space. On the transformation side, we use \PV{} homomorphism transformations to update query and value representations. On the similarity side, we generalize Euclidean inner-product attention with a \PV{} Busemann score inspired by Busemann-based hyperbolic learning and a \PV{} point-to-hyperplane key transformation. On the aggregation side, we replace Euclidean averaging with \PV{} midpoint aggregation and further derive a scalable \PV{} linear attention variant for large graphs. Extensive experiments support the effectiveness of \PVFormer{} across graph, text, and vision benchmarks. The code will be open-sourced once accepted.


RamanBench: A Large-Scale Benchmark for Machine Learning on Raman Spectroscopy

Mario Koddenbrock ⋅ Christoph Lange ⋅ Robin Legner ⋅ Martin Jaeger ⋅ Martin Kögler ⋅ Mariano N Bournazou ⋅ Peter Neubauer ⋅ Felix Biessmann ⋅ Erik Rodner

Machine Learning (ML) has transformed many scientific fields, yet key applications still lack standardized benchmarks. Raman spectroscopy, a widely used technique for non-invasive molecular analysis, is one such field where progress is limited by fragmented datasets, inconsistent evaluation, and models that fail to capture the structure of spectral data. We introduce RamanBench, the first large-scale, fully reproducible benchmark for ML on Raman spectroscopy, consisting of streamlined data access, evaluation protocols and code, as well as a live leaderboard. It unifies 74 datasets (including 16 first released with this benchmark) across four domains, comprising 325,668 spectra and spanning classification and regression tasks under diverse experimental conditions. We benchmark 28 models under a standardized protocol, including classical methods (e.g., PLS), Raman-specific (e.g., RamanNet), Tabular Foundation Model (TFM) (e.g., TabPFN), and time-series approaches (e.g., ROCKET). TFM consistently outperform domain-specific and gradient boosting baselines, while time-series models remain competitive. However, no method generalizes across datasets, revealing a fundamental gap. Therefore, we invite the community to contribute new approaches to our living benchmark, with the potential to accelerate advances in critical applications such as medical diagnostics, biological research, and materials science.

Explicit reasoning models are trained to produce intermediate reasoning traces before final answers, but downstream fine-tuning is often performed on ordinary instruction--response data that contains no such traces. We show that this mismatch can induce reasoning-trace collapse: a fine-tuned model continues to produce plausible final answers while losing the structurally valid explicit reasoning traces that made it a reasoning model in the first place. We introduce a structural evaluation framework that separates answer correctness from reasoning-trace validity, measuring valid, empty, missing, and truncated reasoning alongside reasoning-conditioned task performance. Using this framework, we study open-weight reasoning models and find that standard supervised fine-tuning can rapidly suppress valid reasoning traces, and that answer-only metrics can substantially obscure this failure: in several settings, performance conditional on valid reasoning remains high while the rate of valid reasoning falls sharply. We further show that simple loss-masking strategies substantially mitigate collapse without requiring teacher-generated reasoning traces. These results suggest that evaluations of fine-tuned reasoning models should report structural reasoning reliability in addition to final-answer performance, especially when adaptation data does not contain explicit reasoning traces.


REFORM-3D: A Representation-Centric Evaluation Framework for 3D Medical Vision Foundation Models

Maregu Assefa ⋅ Muhammad Muzammal Naseer ⋅ Divya Velayudhan ⋅ Gedamu Kumie ⋅ Naoufel Werghi

Evaluation of 3D medical foundation models is still dominated by downstream task scores, which entangle representation quality with adaptation choices. This is especially critical in medical imaging, where spacing and anisotropy can substantially alter image statistics while the anatomy itself remains unchanged. We introduce \textbf{REFORM-3D}, a \textbf{R}epresentation-centric \textbf{E}valuation \textbf{F}ramew\textbf{OR}k for benchmarking frozen 3D medical vision foundation \textbf{M}odels across three questions: Acquisition Robustness, asking whether features remain stable when spacing varies while anatomy is fixed; Cross-Modal Anatomical Alignment, asking whether same-organ CT-MR pairs remain recoverable under balanced bidirectional retrieval; and Anatomical Generalization, asking whether limited adaptation recovers held-out organ signals beyond a frozen-feature baseline under case-disjoint evaluation. These are also pretraining questions, because what a 3D encoder learns depends on how volumes are geometrically presented during self-supervised pretraining. We therefore employ REFORM-3D to evaluate encoders pretrained on 62K unlabeled multimodal 3D volumes spanning diverse anatomical structures across MR, CT, PET, CTA, and CBCT under three explicit geometric assumptions: isotropic sampling, voxel-relative sampling, and spacing-aware resampling. By exposing the encoder to acquisition geometry during pretraining rather than treating geometry only as a downstream issue, REFORM-3D evaluates whether learned representations preserve stable anatomical structures or remain sensitive to acquisition and modality shortcuts.


Resilient Byzantine Agreement with Predictions

Julien Dallot ⋅ Darya Melnyk ⋅ Tijana Milentijević ⋅ Stefan Schmid ⋅ Patrik Welters

The Byzantine Agreement (BA) problem is a fundamental task in distributed computing where nodes in a network need to agree on a common output in the presence of arbitrary (worst-case) node failures. In real-world applications, nodes can be monitored over time to provide ML predictions of faulty behavior in a network. In this work, we answer the question of whether predictions can improve the fault tolerance of a distributed system without sacrificing correctness when the predictions are wrong. We consider a prediction-augmented study of Byzantine agreement in which each node receives, in addition to its input bit, a predicted set of honest nodes. We focus on algorithmic resilience --- the maximum number of faulty nodes an algorithm can tolerate --- and present algorithms and impossibility results whose resilience depends on the accuracy of the predictor. As our first main result, we bring a complete characterization of the consistency--robustness trade-offs in both the non-authenticated and authenticated settings: for $n$ nodes and a parameter $\alpha \in [0, 1]$, we present algorithms that tolerate up to $\alpha \cdot n$ faulty nodes when the predictor is correct (consistency), and up to $\frac{1-\alpha}{2} \cdot n - 1$ faulty nodes when the predictor is arbitrarily wrong (robustness). In the authenticated setting, the robustness bound improves to $(1-\alpha) \cdot n - 1$. We prove matching impossibility results, showing that these tradeoffs are optimal and independent of the particular prediction system. Our second main result characterizes smoothness: the rate at which resilience degrades as the predictor becomes less accurate. We show that resilience linearly decreases in the number of wrong predictions as long as that number stays within a constant fraction of $n$. Concretely, in the non-authenticated setting, each additional wrong prediction loses one unit of resilience, whereas in the authenticated setting, the decline is halved, since two wrong predictions are needed to lose one unit of resilience.

Vocabulary-based end-to-end driving selects a future trajectory from a fixed candidate set, enabling explicit comparison among multiple plausible plans. However, existing planners typically process the scene-invariant trajectory vocabulary and the scene-dependent driving context together in a single online forward pass. This repeatedly recomputes reusable candidate-side representations and applies expensive scene-conditioned evaluation to the full vocabulary. We revisit vocabulary-based planning as a retrieve-then-rerank problem, where the trajectory vocabulary is a static candidate collection, lightweight scene cues form a retrieval query, and the scene-conditioned evaluator acts as a reranker. Based on this view, we propose $\textbf{Decoupled Retrieval-Style Inference (DRSI)}$, centered on two modules: $\textbf{Offline Candidate Indexing (OCI)}$ and \textbf{Scene-aware Candidate Retrieval (SCR)}. OCI caches scene-invariant trajectory embeddings offline, while SCR retrieves a compact set of route-consistent and dynamically reachable candidates online. The scene-conditioned evaluator is then applied only to this retrieved subset for fine-grained reranking. Experiments on NAVSIM show that DRSI substantially improves the latency--quality trade-off without adding trajectory proposals or increasing evaluator complexity. Compared with the Hydra-MDP++ baseline, our large DRSI model achieves 86.1 EPDMS on NAVSIM v2 with 23.9 ms latency, corresponding to about a 6.8$\times$ speedup. On the challenging navhard split, DRSI further improves EPDMS from 37.6 to 40.2. These results demonstrate that retrieval-style decoupling is an effective inference principle for efficient vocabulary-based end-to-end driving.

Self-supervised pre-training via cross-view completion learns strong features for 3D vision from co-visible regions of image pairs. However, the reference view provides little information for reconstructing non-co-visible patches, implicitly yielding a monocular training signal in these regions. We introduce Gekko, a self-supervised pre-training framework that turns this limitation into a useful signal. The relative improvement of the cross-view reconstruction error over a masked-autoencoder error serves as a self-supervised proxy for co-visibility: large improvements indicate co-visible regions, while negligible ones indicate non-co-visible areas. Gekko is a network, trained from scratch, that jointly performs cross-view completion, masked autoencoding, and per-pixel prediction of this relative improvement, providing an additional binocular signal for all masked regions without requiring any ground-truth 3D annotations. Under identical architectures and training data, Gekko consistently outperforms CroCo on zero-shot correspondence estimation, relative pose estimation, and pointmap regression. We further show that Gekko can be trained directly from raw videos using a simple stride-based curriculum, eliminating the cumbersome 3D preprocessing required by prior methods while matching the performance of models trained on curated data, thereby enabling fully self-supervised pre-training. Code and pre-trained models will be released.


Reviving the Discarded High-Resolution Feature for Transformer-Based Tiny Object Detection

Zehao Wang ⋅ Si Wu ⋅ GUIPING CAO ⋅ Xin Li ⋅ Yaowei Wang

Tiny object detection remains challenging because tiny instances occupy only a few pixels, and their fine-grained details are easily lost during feature downsampling. Although the highest-resolution feature preserves rich local details for tiny objects, DETR-like detectors often discard it due to its large computational cost, while directly reintroducing this feature also brings dense background noise. To address this problem, we propose \textbf{ReDHF}, a novel framework that \textbf{Re}vives the \textbf{D}iscarded \textbf{H}igh-resolution \textbf{F}eature for tiny object detection. ReDHF has three cooperative components. Cross-Scale Deformable Fusion (CSDF) first constructs a semantically enhanced highest-resolution feature and injects its fine-grained details into the shallow feature under a moderate cost. Ellipse-Guided Shallow Supervision (EGSS) then guides the enhanced feature to learn more reliable classification and localization priors by applying geometry-adaptive auxiliary supervision. Full-Scale Query Interaction (FSQI) further enables decoder queries to access both low-level details and high-level semantic information, allowing them to aggregate richer cues for tiny objects. Extensive experiments show that ReDHF consistently improves multiple baselines across different DETR variants and achieves state-of-the-art performance on the evaluated benchmarks. The best ReDHF variants surpass their baselines by \textbf{+1.4 AP} on AI-TOD-v2, \textbf{+1.6 AP} on VisDrone, and \textbf{+1.5 AP} on SODA-D, with especially notable gains on tiny and small objects.


RL-Guided Contraction of Symbolic Tensor Networks for Quantum Circuit Equivalence

Suhaib Al-Rousan ⋅ Christian Schilling ⋅ Max Tschaikowski ⋅ Kim Larsen

Verifying that two quantum circuits are equivalent is a critical challenge in compiler validation. A common approach is to transform the circuits into a tensor network and then iteratively contract tensors in the network. The order in which contractions are applied has a crucial influence on the complexity, as tensor size may grow exponentially in the number of qubits. This motivated research into heuristic policies to find good contraction orders, including recent work based on reinforcement learning (RL). Orthogonally, to mitigate the complexity of an explicit tensor representation, tensor decision diagrams (TDDs) have been proposed as a compact symbolic representation. However, also the TDD performance is sensitive to the contraction order, and policies designed for tensor networks do not perform well in case of a TDD representation. In this paper, we propose an RL framework to obtain a graph-neural-network policy that navigates the trade-off between arithmetic complexity and TDD size. Experimental results demonstrate that a policy learned with our framework is effective at taming the symbolic representation size.


Robust and Scalable Collaborative Learning via Pull-Based Epidemic Communication

Abdellah El Mrini ⋅ Sadegh Farhadkhani ⋅ Rachid Guerraoui

Collaborative machine learning is vulnerable to adversarial behaviors during training. Existing defenses typically rely on central coordination or induce high communication costs. We introduce *Robust Pull-based Epidemic Learning (RPEL)*, a scalable and fully decentralized framework that achieves robustness without a central server. Unlike traditional methods, whose communication cost grows as $\mathcal{O}(n^2)$ with the number of nodes $n$, RPEL uses a pull-based epidemic communication scheme that scales as $\mathcal{O}(n \log n)$. By pulling model parameters from small, randomly selected subsets of peers, RPEL significantly lowers the number of required messages while preserving convergence guarantees with high probability. Experimental results demonstrate that RPEL indeed tolerates attacks, attains accuracy comparable with all-to-all communication, and scales efficiently to large networks.

Sparse Mixture-of-Experts (SMoE) models enable scaling language models efficiently, but training them remains challenging, as routing can collapse onto few experts and auxiliary load-balancing losses can reduce specialization. Motivated by these hurdles, we study how routing decisions in SMoEs are formed mechanistically. First, we reveal a geometric coupling between routers and their corresponding experts. For a given token, the router weights for the selected expert and the expert weights processing it receive gradients along the same input direction, differing only in scalar coefficients. Thus, matched router--expert directions accumulate the same routed token history. This theoretical coupling also appears empirically in routing dynamics. In a 1B SMoE trained from scratch, higher router scores predict stronger expert neuron activations, showing that routing decisions are mirrored inside the selected expert. Next, we analyze the effects of auxiliary load balancing on the router--expert geometric coupling, showing that such losses break this structure by spreading input-directed gradients across router weights, making distinct router directions nearly three times more similar to each other. Last, we demonstrate the centrality of geometric coupling for effective routing with a parameter-free online K-Means router, in which each expert maintains a running average of the hidden states routed to it and tokens are assigned based on cosine similarity. Compared with auxiliary-loss and loss-free balancing, this router achieves the lowest load imbalance with only a modest perplexity increase, indicating that geometric coupling captures a substantial part of what the router learns. Overall, our results explain how routers form assignment geometry that supports an effective division of labor.


Safe Offline Reinforcement Learning using Behavior Regularisation and Latent Feasibility-Guidance

Mahesh Keswani ⋅ Parth Bhardwaj ⋅ Ankur Kumar ⋅ Raunak Pushpak Bhattacharyya

Offline safe reinforcement learning aims to maximize cumulative reward while satisfying safety constraints using only static datasets. Recent approaches leverage latent variable models to capture the behavior policy distribution, enabling efficient policy extraction under limited data. However, existing latent-space methods typically rely on soft constraints, which allow constraint violations in expectation and may be unsuitable for safety-critical settings. In addition, these methods often employ Advantage-Weighted Regression (AWR) for policy extraction, which exhibits exponential sensitivity to out-of-distribution critic bias, leading to high variance in the policy gradient estimate. To address these limitations, we propose Feasibility-Optimized Conditional Actor Learning (FOCAL). FOCAL enforces state-wise hard safety constraints by integrating reachability-based feasibility constraints directly into the Conditional Variational Autoencoder. To overcome the instability associated with AWR, we introduce a feasibility-partitioned, behavior-regularized objective that utilizes stable, first-order gradients to structurally decouple reward maximization in feasible regions from constraint violation minimization in infeasible regions. Finally, an adaptive test-time inference dynamically modulates latent sampling variance based on state-wise feasibility to ensure safe policy execution under out-of-distribution states. Theoretical analysis demonstrates that FOCAL circumvents the exponential variance in policy gradient estimates associated with AWR under critic errors, while providing formal bounds on policy deviation from the behavioral dataset. Empirically, on the DSRL benchmark, FOCAL satisfies safety constraints while achieving competitive or higher returns compared to state-of-the-art methods, all while maintaining efficient single-step inference.


Scalable Multi-Agent Contrastive Reinforcement Learning

Victor Augusto Kich ⋅ Satoshi Yamamori ⋅ Jun Morimoto

Sparse rewards make cooperative multi-agent reinforcement learning difficult because useful feedback often appears only after multiple agents coordinate over long horizons. We introduce Multi-Agent Contrastive Reinforcement Learning (MACRL), a goal-conditioned method that trains parameter-shared decentralized actors with a centralized contrastive critic. Instead of learning from sparse task rewards, the critic contrasts joint state-action embeddings with future achieved goals from replay, turning every trajectory into dense self-supervised signal. Across tasks that vary in horizon length, coordination difficulty, and object interaction, our method is competitive with strong HER-based baselines when exploration is sufficient and substantially more robust when sparse rewards become uninformative. On the hardest tasks, it obtains substantial non-zero success while non-contrastive baselines remain near zero. Scaling from 4 to 64 layers further improves performance on several hard tasks, with gains up to 25 percentage points, while providing little consistent benefit to MA-SAC+HER. These results suggest that contrastive objectives and depth-scaled residual networks provide a promising foundation for scalable cooperative goal-conditioned control.


Scaling Generative Foundation Models for Chest Radiography with Rectified Flow Transformers

Fabio De Sousa Ribeiro ⋅ Emma A.M. Stanley ⋅ Charles Jones ⋅ Tian Xia ⋅ Dominic C Marshall ⋅ Laurent R Triché ⋅ Christopher V Cosgriff ⋅ Panagiotis Dimitrakopoulos ⋅ Sotirios Tsaftaris ⋅ Ben Glocker

We introduce the first generative foundation model for chest radiograph synthesis trained from scratch at the billion-parameter scale. Existing radiographic AI models often suffer from poor generalisation across patient subpopulations, institutions, and acquisition settings, resulting in limited clinical utility. Controlled, high-fidelity synthesis of chest radiographs is a promising path toward diversifying clinical datasets and evaluating the robustness of diagnostic models. Therefore, we present the largest specialist generative foundation model for chest radiographs to date, with over 1.3B parameters, trained for 1.6T tokens on a curated, heterogeneous dataset comprising 1.2M radiographs and clinical expert-guided metadata. Our model supports controllable radiograph generation and editing across multiple demographic subgroups, acquisition views, and a dozen pathologies. Moreover, we significantly advance the state of the art in radiograph synthesis fidelity, producing images that are indistinguishable from real radiographs to expert pulmonologists.

Language models alter their behavior when they infer that they are being trained, evaluated, deployed, or weakly overseen. Recent AI safety work treats this as evidence of scheming, and substantial effort is being directed toward its detection and prevention. This position paper argues that the deeper significance of scheming is missed by asking only whether a model has a discrete tendency to scheme. Drawing on theories of reflexive social systems, we argue that scheming-like behavior is a structural consequence of alignment being reflexive -- a model's behavior is trained against its picture of an evaluative world that its outputs help constitute -- and recognizing this reveals broader risks than scheming itself. We organize existing ML evidence as consistent with this diagnosis, outline the broader danger, and propose a research program to address it.


Selectivity and Shape in the Design of Forward-Forward Goodness Functions

Talha Rüzgar Akkuş ⋅ Şuayp T Kocabay ⋅ Kamer A Yuksel ⋅ Hassan Sawaf

The Forward-Forward (FF) algorithm trains networks layer-by-layer using a local "goodness function," yet the principles governing what makes an effective goodness function remain largely unexplored. We systematically explore the goodness-function design space and identify a unifying principle: the goodness function must be sensitive to the shape of neural activity, not its total energy. This principle is motivated by the observation that deep network activations follow heavy-tailed distributions and that discriminative information is often concentrated in peak activities. We propose two complementary families: selective functions (top-k, entmax-weighted energy) that measure only peak activity, and shape-sensitive functions (excess kurtosis / "burstiness" and higher-order moments) that reward heavy-tailed distributions via scale-invariant statistics. Combined with separate label-feature forwarding (FFCL), controlled experiments across 13 goodness functions, 5 activations, 6 datasets, and three continuous sweeps each tracing a characteristic inverted-U yield 89.0% on Fashion-MNIST and 98.2±0.1% on MNIST (4x2000)—a +32.6pp gain over SoS—with consistent improvements across all benchmarks (+72pp USPS, +52pp SVHN). The scale-invariant nature of burstiness makes it particularly robust to magnitude shifts across layers and datasets. Code is available at https://anonymous.4open.science/r/ff-selectivity-shape.

We present the first algorithms for generalized linear contextual bandits under shuffle differential privacy and joint differential privacy. While prior work on private contextual bandits has been restricted to linear reward models---which admit closed-form estimators---generalized linear models (GLMs) pose fundamental new challenges: no closed-form estimator exists, requiring private convex optimization; privacy must be tracked across multiple evolving design matrices; and optimization error must be explicitly incorporated into regret analysis. We address these challenges under two privacy models and context settings. For stochastic contexts, we design a shuffle-DP algorithm with regret $\tilde{O}\!\big(d^{3/2}\sqrt{T \log T} + d^{5/4} \sqrt{T/\varepsilon} (\log T)^{3/4}\big)$ in the dominant term, matching the non-private rate $\tilde{O}(d \sqrt{T \log T})$ in the leading-in-$T$ term up to a multiplicative factor of $\sqrt d$ and a privacy correction term of $\tilde{O}(d^{5/4}\sqrt{T \log T/\varepsilon})$. For adversarial contexts, we provide a joint-DP algorithm with regret $\tilde{O}\!\big(d\sqrt{T}\log T + d^{3/4}\sqrt{T/\varepsilon}\,(\log T)\,(d+\log T)^{1/4}\big)$ matching the non-private rate $\tilde{O}(d\sqrt{T}\log T)$ in the leading term. Unlike prior work on locally private GLM bandits, our methods require no spectral assumptions on the context distribution beyond $\ell_2$ boundedness.

U-Net--style architectures are widely adopted for modeling physical systems, as their multiscale structure enables efficient processing of high-resolution data while reflecting the hierarchical structure of many physical phenomena. The most successful U-Net-based neural physics simulators often combine convolutions, Fourier layers, or specialized transformer mechanisms such as windowed or axial attention; however, these structured components can limit adaptability across different spatial and spatiotemporal dimensions. In this work, we revisit transformer-based U-Nets with the goal of simplifying attention-based neural physics simulators without sacrificing multiscale efficiency or predictive accuracy. We introduce \texttt{UFlex}, a simple attention-based U-Net whose blocks are obtained through only minimal modifications to standard self-attention, allowing the same architecture to be applied across one-, two-, three-dimensional, and spatiotemporal regular-grid problems without major design changes. Evaluated on seven challenging benchmarks (four 2D and three 3D), \texttt{UFlex} scales to resolutions of up to $512 \times 512$ in 2D and $256 \times 128 \times 256$ in 3D, while reducing training memory and accelerating training compared to state-of-the-art transformer baselines, all while achieving state-of-the-art predictive accuracy.


SING-CH: Task-Scale-Agnostic Lifelong Cross-Modal Hashing on Statistical Manifolds

Haoran Yang ⋅ Junge Chen ⋅ Jun Long ⋅ Zhan Yang

Lifelong cross-modal hashing demands incremental updates as new tasks arrive, yet existing regularization-based methods such as EWC and SI share a fundamental weakness: their effective regularization strength is entangled with task scale, as the empirical Fisher Information Matrix (FIM) scales linearly with sample size. This forces costly per-task hyperparameter tuning in streaming deployments to avoid over- or under-regularizing tasks of varying sizes. We propose \textsc{Sing-CH}, a lifelong hashing framework built on information geometry. At its core is a Gibbs probabilistic reconstruction of the pairwise objective, which gives the parameter space a statistical manifold structure and yields a natural $1/n_t^2$ normalization. We prove this normalization renders the FIM scale-agnostic: its spectral properties remain independent of task size even under diagonal and K-FAC approximations. Beyond this invariance, a Forgetting-Rigidity trade-off theorem shows that natural gradient applied to the composite objective reintroduces scale dependence. This makes our decoupled design, pairing the FIM as regularizer with Adam for optimization, a structural necessity rather than a convenience. Experiments on three benchmarks demonstrate that \textsc{Sing-CH} outperforms 11 baselines by 4.1\% average mAP without per-task tuning. Under extreme task-scale imbalance, degradation is only 2.0\% versus 3.9\% for the strongest baseline, providing the first empirically grounded solution to scale-agnostic lifelong retrieval.


Soft Forward-Backward Representations for Zero-shot Reinforcement Learning with General Utilities

Marco Bagatella ⋅ Thomas Rupf ⋅ Georg Martius ⋅ Andreas Krause

Recent advancements in zero-shot reinforcement learning (RL) have facilitated the extraction of diverse behaviors from unlabeled, offline data sources. In particular, forward-backward algorithms (FB) can retrieve a family of policies that approximately solves any standard RL problem (with additive rewards, linear in the occupancy measure), given sufficient capacity. While retaining zero-shot properties, we tackle the greater problem class of RL with general utilities, in which the objective is an arbitrary differentiable function of the occupancy measure. This setting is strictly more expressive, capturing tasks such as distribution matching or pure exploration, which may not be reduced to additive rewards. We show that this additional complexity can be captured by a novel, maximum entropy (soft) variant of the forward-backward algorithm, which recovers a family of stochastic policies from offline data. When coupled with zero-order search over compact policy embeddings, this algorithm can sidestep iterative optimization schemes, and optimizes general utilities directly at test-time. Across both didactic and high-dimensional experiments, we demonstrate that our method retains favorable properties of FB algorithms, while also extending their range to more general RL problems.


SOPO: Socratic Guided Policy Optimization for Span-Level Hallucination Detection

Jiawei Dong ⋅ Zilong Bai ⋅ Pengtian Zhu ⋅ Peng Zhou

Large Language Models (LLMs) have gained widespread adoption in various natural language processing tasks, but suffer from hallucination issues where they generate unfaithful or inconsistent content. While Reinforcement Learning with Verifiable Rewards (RLVR) has shown promising applications in reasoning tasks, its application to span-level hallucination detection is challenged by localization ambiguity, reward sparsity on hard examples, and uncontrolled length inflation. To address these issues, we introduce Socratic Guided Policy Optimization (SOPO), a multi-turn RL framework for robust hallucination detection. The key innovation of SOPO is a Socratic feedback mechanism, where a teacher model guides the reasoning trajectories during rollout by providing indirect hints inferred from the ground truth. This teacher-guided approach enables the model to generate more effective reasoning trajectories for hard examples, thus alleviating reward sparsity. Furthermore, we incorporate Turn-based Advantage Scaling to penalize verbosity bias, encouraging the model to develop efficient and autonomous reasoning capabilities. Empirical results on RAGTruth demonstrate that SOPO (4B) achieves state-of-the-art performance, surpassing leading proprietary models (e.g., GPT-5) and standard RL baselines, while also exhibiting robust generalization across diverse out-of-domain benchmarks. We release our training and evaluation code alongside the 1.7B and 4B SOPO models at https://anonymous.4open.science/r/SOPO.


Sparse Video Generation Propels Real-World Beyond-the-View Vision-Language Navigation

Hai Zhang ⋅ Siqi Liang ⋅ Li Chen ⋅ Yuxian LI ⋅ Yukuan XU ⋅ Yichao Zhong ⋅ Fu Zhang ⋅ Hongyang Li

Why must vision-language navigation be bound to detailed and verbose language instructions? While such details ease decision-making, they fundamentally contradict the goal for navigation in the real-world. Ideally, agents should possess the autonomy to navigate in unknown environments guided solely by simple and high-level intents. Realizing this ambition introduces a formidable challenge: Beyond-the-View Navigation (BVN), where agents must locate unseen targets without step-by-step guidance. Existing large language model (LLM)-based methods, though adept at following dense instructions, often suffer from short-sighted behaviors due to their reliance on short-horizon supervision. Simply extending the supervision horizon, however, destabilizes LLM training. In this work, we identify that video generation models (VGM) inherently benefit from long-horizon supervision to align with language instructions, rendering them uniquely suitable for BVN tasks. Capitalizing on this insight, we propose applying VGM as the main backbone to generate actions for the first time in this field. Yet, the prohibitive latency for generating videos spanning tens of seconds makes real-world deployment impractical. To bridge this gap, we propose SparseVideoNav, achieving sub-second trajectory inference guided by a generated sparse future spanning a 20-second horizon. This yields a remarkable 27$\times$ speed-up compared to the unoptimized counterpart. Extensive real-world zero-shot experiments demonstrate that SparseVideoNav achieves 2.5$\times$ the success rate of state-of-the-art LLM baselines on BVN tasks and marks the first realization of such capability in challenging night scenes.


Spectral Quantum Memory for Implicit Neural Representations

Daniel A Moroz-Broitman ⋅ Eliya Nachmani

Implicit neural representations (INRs) model signals as continuous functions from coordinates to values, but their performance depends strongly on how spectral capacity is allocated. Yet most Quantum INR (QINR) architectures treat spectral components as independent outputs of separate quantum blocks. We introduce the Spectral Quantum Memory Model (SQMM), a memory coupled quantum INR in which sequential hybrid data reuploading layers (bands) communicate through a shared coherent quantum memory register. The construction gives a controlled way to share spectral information across quantum modules while preserving the finite Fourier structure of a band. When memory read and write layers are input independent, we prove that the memory state forms an operator valued Fourier series whose support grows by Minkowski summation of local band spectra. Downstream bands then condition their Fourier coefficients on the spectral content stored by earlier bands. This theory motivates a spectral memory alignment loss that encourages memory to encode the spectrum expressed by the deployed prefix of the model. Across audio and image reconstruction benchmarks, SQMM improves over QINR and INR baselines, achieving state of the art results.


Statistical Inference in Causal Partial Identification under Smooth Densities

Sirui Lin ⋅ Zijun Gao ⋅ Jose Blanchet ⋅ Peter W Glynn

Many causal quantities are only partially identifiable due to the inherent missingness of potential outcomes, and the associated partial identification (PI) sets can be obtained by solving an optimal transport (OT) problem. Covariates often provide additional information about the potential outcomes and thus yield tighter PI sets, which can be obtained via conditional optimal transport (COT). However, COT-based PI set estimators are susceptible to the curse of dimensionality in the covariates, which precludes the asymptotic normality and hinders statistical inference. In this paper, we exploit the smoothness in the marginal densities of covariates and potential outcomes, and develop a wavelet-based primal approach for COT that attains a faster convergence rate. Moreover, for quadratic cost functions, we establish a stability result for COT and prove asymptotic normality of the proposed estimator, enabling valid statistical inference for the PI set. Empirically, we validate the estimation and inference performance of our approach through numerical experiments in comparison with existing benchmarks.


Stein Transport for Generative Modeling

Clémentine CHAZAL ⋅ Linfeng Wang ⋅ Anna Korba ⋅ Nikolas Nüsken

We propose a nonparametric, kernel-based approach to generative modeling based on Stein operators and particle transport. Our method constructs a continuous-time transport from a simple reference distribution (typically, a standard Gaussian) to a target distribution using a Stein formulation of the continuity equation. Unlike alternative Stein transport methods, which require access to the target density or its gradient, we focus on the sample-only setting and introduce a tractable surrogate objective based on an Ornstein--Uhlenbeck (OU) probability path. This construction allows the velocity field driving the transport to be learned using only samples from the target distribution and Gaussian noise. By restricting the velocity field to a reproducing kernel Hilbert space, we obtain closed-form solutions via kernel ridge regression at each time step. The resulting method is particularly well suited to small data regimes, where heavily parameterised neural generative models may be difficult to train or overfit. We demonstrate the usefulness of the approach for sample-based generative modeling in the small data regime (generative data augmentation), and for posterior emulation in Bayesian inverse problems, where expensive Monte Carlo samplers produce limited sets of posterior samples that must be efficiently augmented.


Super-Level-Set Regression: Conditional Quantiles via Volume Minimization

Sacha Braun ⋅ Michael Jordan ⋅ Francis Bach

Constructing minimum-volume prediction regions that satisfy conditional coverage is a fundamental challenge in multivariate regression. Standard approaches rely on explicitly estimating the full conditional density and subsequently thresholding it. This two-step plug-in process is notoriously difficult, sensitive to estimation errors, and computationally expensive. One would like to instead optimize the region directly. Formulating a direct solution is challenging, however, because it requires minimizing a volume objective that is coupled with the conditional quantiles of the model's own estimation error. In this work, we address this challenge. We introduce \emph{super-level-set regression} (SLS), a novel mathematical framework that successfully resolves this implicit coupling, allowing us to directly parameterize and optimize the geometric boundaries of the target conditional level sets. By bypassing full distribution estimation and leveraging flexible volume-preserving frontier functions, our approach natively captures complex, multimodal, and disjoint conditional structures end-to-end. Ultimately, SLS offers a new perspective on multivariate conditional quantile regression, replacing the restrictive assumptions of density-first methods with a direct geometric optimization strategy.


Supervised Distributional Reduction via Optimal Transport and Dependence Maximization

Sai-Aakash Ramesh ⋅ Archit Sood ⋅ Andrew Corbett ⋅ Tim Dodwell

Learning representations that capture both intrinsic data geometry and target-relevant structure remains a fundamental challenge, particularly in settings where data reduction must balance compression with predictive fidelity. While distributional reduction—encompassing joint clustering and dimensionality reduction—offers a principled way to summarize data, its supervised variants remain relatively underexplored, despite the importance of retaining task-relevant signal for downstream prediction and decision-making. We propose Supervised Distributional Reduction (SDR), an algorithm for learning target-aware representations by combining optimal transport with explicit dependence maximization. SDR builds on the Fused Gromov–Wasserstein (FGW) objective to align the relational structure of the input distribution with a set of representative points, while augmenting it with a direct dependence term that encourages the learned embeddings to capture predictive signal more explicitly. This results in compact representations that reflect both geometric structure and supervision. Beyond representation learning, SDR naturally induces a data-dependent, non-stationary geometry that can be leveraged for settings such as Gaussian Process (GP) modelling. By redefining distances through target-aware distributional alignment, SDR enables the construction of adaptive kernels that respond to local variations in both data geometry and supervision, offering an optimal transport-based perspective on non-stationary kernel design.


Surjective Pseudo-Invertible Neural Networks

Yamit Ehrlich ⋅ Amit Arad ⋅ Nimrod Berman ⋅ Assaf Shocher

The Moore-Penrose Pseudo-inverse (PInv) is the fundamental tool for inverting linear operators. We propose a natural generalization to the non-linear regime and introduce {Surjective Pseudo-invertible Neural Networks (SPNN)}: architectures that admit a tractable non-linear PInv by construction and satisfy the corresponding geometric properties. Building on this, we formalize {Non-Linear Back-Projection (NLBP)}, the non-linear analogue of $x' = x + A^\dagger(y-Ax)$: an update that projects any sample to its closest state consistent with $f(x)=y$. Diffusion-based null-space projection revolutionized zero-shot solving of linear inverse problems via closed-form back-projection; NLBP extends this paradigm to non-linear learned ``degradations'' in the broad sense, spanning semantic mappings such as classification and object detection. The result is zero-shot inversion of complex degradations and precise semantic control over generative outputs without retraining the prior.


SyncLight: Single-Edit Multi-View Relighting

David Serrano-Lozano ⋅ Anand Bhattad ⋅ Luis Herranz ⋅ Jean-Francois Lalonde ⋅ Javier Vazquez-Corral

We present SyncLight, a method to enable consistent, parametric control over light sources across multiple uncalibrated views of a static scene conditioning in a single view. While single-view relighting has advanced significantly, existing generative approaches struggle to maintain the rigorous lighting consistency essential for multi-camera broadcasts, stereoscopic cinema, and virtual production. SyncLight addresses this by enabling precise control over light intensity and color across a multi-view capture of a scene, conditioned on a single reference edit. Our method leverages a multi-view diffusion transformer trained using a latent bridge matching formulation, achieving high-fidelity relighting of the entire image set in a single inference step. To facilitate training, we introduce a large-scale hybrid dataset comprising diverse synthetic environments-curated from existing sources and newly designed scenes-alongside high-fidelity, real-world multi-view captures under calibrated illumination. Surprisingly, though trained only on image pairs, SyncLight generalizes zero-shot to an arbitrary number of viewpoints, effectively propagating lighting changes across all views, without requiring camera pose information. SyncLight enables practical relighting workflows for multi-view capture systems. Dataset, code, and models will be released upon acceptance.

Scaling laws are increasingly used to pick which expensive training run to scale up. The practical question — which candidate is best at a fixed future compute budget $C^\star$ — fits neither classical best-arm identification nor minimum-oriented functional bandits. We formulate it as the *target-budget best-function identification* (TB-BFI) problem, with two natural flavors: TB-BFI-FC asks for a $\delta$-correct certificate of the winner at $C^\star$, and TB-BFI-FB asks for low target-budget simple regret under a hard exploration budget. For the shared-exponent power-law class we prove uniform target-budget validity, a finite-time stopping bound for the cost-aware TB-LUCB rule, and a per-arm-exponent perturbation extension. We then close two structural gaps. **(a)** We extend the theory *beyond* shared-exponent to the joint $(N,D)$ Chinchilla law with shared but *unknown* exponents, via a two-stage protocol with explicit pilot sample complexity and a quadratic-in-$\epsilon$ centred-design perturbation theorem. **(b)** We give the first positive TB-BFI-FB result: a cost-aware Sequential-Halving algorithm whose cost-weighted simple-regret upper bound matches a new information-theoretic lower bound up to the standard $K \log K$ adaptive-vs-oracle gap. Empirically, theorem-aligned synthetic experiments certify the $C^\star$-winner in $95$–$100\%$ of trials; on a char-level transformer learning-rate-crossover our extrapolator picks the right arm in $3/3$ seeds where Successive Halving picks $0/3$; and on a flip-rank Chinchilla instance the joint-law procedure reaches $70$–$87\%$ accuracy where the univariate $c = N \cdot D$ baseline scores $0\%$.


Test-Time Compute Games

Ander Artola Velasco ⋅ Dimitrios Rontogiannis ⋅ Stratis Tsirtsis ⋅ Manuel Rodriguez

Test-time compute has emerged as a promising strategy to enhance the reasoning abilities of large language models (LLMs). However, this strategy has in turn increased how much users pay cloud-based providers offering LLM-as-a-service, since providers charge users for the amount of test-time compute they use to generate an output. In our work, we show that the market of LLM-as-a-service is socially inefficient–providers have a financial incentive to increase the amount of test-time compute, even if this increase contributes little to the quality of the outputs. To address this inefficiency, we introduce a reverse second-price auction mechanism where providers bid their offered price and (expected) quality for the opportunity to serve a user, and users pay proportionally to the marginal value generated by the winning provider relative to the second-highest bidder. To illustrate and complement our theoretical results, we conduct experiments with multiple instruct models from the $\texttt{Llama}$ and $\texttt{Qwen}$ families, as well as reasoning models distilled from $\texttt{DeepSeek-R1}$, on benchmark datasets covering mathematics, science, and open-ended question answering.

Existing approaches to LLM personalization focus on constructing better personalized models or inputs, while treating inference as a single-shot process. In this work, we study \textbf{Test-Time Personalization (TTP)} along an unexplored axis: scaling inference-time computation by sampling $N$ candidates from a personalized policy model and selecting the best with a personalized reward model. However, standard reward models fail to realize this potential. To diagnose why, we derive a unified scaling law that decomposes any reward model's Best-of-$N$ curve into four measurable quantities and reveals two failure modes, \emph{user-level collapse} (near-constant prediction for some users) and \emph{query-level reward hacking} (negative correlation with true quality for some queries). Guided by this law, we propose a probabilistic personalized reward model whose learned variance effectively mitigates both failure modes. Experiments confirm both elements of our framework: TTP delivers consistent scaling across multiple policy models and personalized text generation tasks, and our scaling law closely matches observed scaling curves across reward-model variants.


TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection

Tom Sander ⋅ Pierre Fernandez ⋅ Hongyan Chang ⋅ Tomáš Souček ⋅ Sylvestre-Alvise Rebuffi ⋅ Hady Elsahar ⋅ Valeriu Lacatusu ⋅ Alexandre Mourachko ⋅ Tuan Tran

We introduce TextSeal, a state-of-the-art watermark for large language models. Building on Gumbel-max sampling, TextSeal introduces dual-key generation to restore output diversity, along with entropy-weighted scoring and multi-region localization for improved robustness. It supports serving optimizations such as speculative decoding and multi-token prediction, and does not add any inference overhead. TextSeal strictly dominates baselines like SynthID-text in detection strength and is robust to dilution, maintaining confident localized detection even in heavily mixed human/AI documents. The scheme is theoretically distortion-free, and evaluation across reasoning benchmarks confirms that it preserves downstream performance; while a multilingual human evaluation (6000 A/B comparisons, 5 languages) shows no perceptible quality difference. Beyond its use for provenance detection, TextSeal is also "radioactive": its watermark signal transfers through model distillation, enabling detection of unauthorized use.


The Benefits of Temporal Correlations: SGD Efficiently Learns k-Juntas from Random Walks

Elisabetta Cornacchia ⋅ Dan Mikulincer ⋅ Elchanan Mossel

We study how temporal correlations in the data can make certain sparse learning problems efficiently learnable by gradient-based methods. Our focus is on Boolean k-juntas, a canonical sparse learning problem known to pose barriers for gradient-based methods under independent uniform samples. We show that this picture changes when the samples are generated by a lazy random walk on the hypercube. In this setting, the temporal dependencies can be exploited by a two-layer ReLU network trained using stylized-SGD with a temporal-difference loss, which compares target and predicted increments across consecutive samples. For every fixed k, the resulting sample complexity is essentially linear in the ambient dimension d. By contrast, we show that for large-batch gradient methods using standard convex pointwise losses, temporal correlations do not provide the same advantage.

Pearl's causal hierarchy shows that observational, interventional, and counterfactual queries are qualitatively distinct. We ask a quantitative version of this question: how many additional bits are needed to specify higher-rung causal answers once lower-rung answers are known? We formalize this via query-class description length, the Kolmogorov complexity of the answer oracle induced by an SCM for a class of queries. Our main construction gives binary acyclic SCMs whose observational distribution has constant description length, while the single-variable interventional answer oracle has description length $\Theta(n^2)$. A degree-sensitive upper bound shows that finite-gate-schema SCMs of indegree $d$ have observational–interventional gap at most $O(nd\log(en/d)+n\log n)$, making the quadratic construction order-optimal in the dense regime and a rooted-tree construction order-optimal for bounded indegree. The quadratic separation persists under $\varepsilon$-accurate total-variation descriptions for every fixed $\varepsilon < 1/4$. At the next rung, the full hard-do interventional oracle can still leave a $\Theta(n)$ counterfactual description gap. A general ambiguity-to-bits theorem and Shannon analogue show that these gaps equal the logarithm of residual higher-rung ambiguity up to lower-order terms.

We study how much a linear program (LP) can be compressed when solved repeatedly, given prior knowledge about its objective function. Existing data-driven projection methods learn low-dimensional surrogate LPs with approximate objective-value guarantees, but cannot provably identify the optimal projection for a prescribed compression budget. We instead ask a sharper question: how far can an LP be compressed into a lower-dimensional equivalent while exactly preserving optimality, enabling faster repeated solves with no loss in solution quality? We provide an exact geometric characterization of such compressed LPs, together with a tractable sample-based learning algorithm that comes with fast-rate guarantees: the compressed LP recovers the optimal solution of an unseen instance with probability at least $1-\widetilde O(d^\star/n)$, where $d^\star$ is the dimension of the decision-relevant subspace, and $n$ the number of available historical LP samples. This $1/n$ dependence is sharper than the $\widetilde O(1/\sqrt n)$ uniform-convergence rates of approximate projection methods. Our framework further exposes a tunable tradeoff between the dimension of the compressed LP and the probability of recovering the optimal solution, allowing the user to trade compression for accuracy.


The Hidden Power of Scaling Factor in LoRA Optimization

Zicheng Zhang ⋅ Haoran Li ⋅ Jiaxing Wang ⋅ Guoqiang Gong ⋅ Anqi Li ⋅ Yudong Hu ⋅ Ting Xiong ⋅ Yurong Gao ⋅ Junxing Hu ⋅ Zhida Jiang ⋅ Yifeng Zhang ⋅ Pengzhang Liu ⋅ Qixia Jiang

In Low-Rank Adaptation (LoRA), the scaling factor $\alpha$ is often treated as a mere complement to the learning rate, yet its role in optimization remains poorly understood. In this paper, we reveal that the scaling factor $\alpha$ and the learning rate function differently, with $\alpha$ emerging as the dominant driver of effective optimization, delivering gains that cannot be replicated by learning rate scaling alone. Through the synergy of extensive empirical analysis and a theoretical Signal-Drift framework, we uncover three findings into LoRA’s scaling mechanism: First, LoRA’s spectral suppression smooths the optimization landscape, rendering standard hyperparameters overly conservative and creating an optimization gap. Second, when leveraging this smoothness to accelerate convergence, $\alpha$ outperforms the learning rate by amplifying the task signal without increasing the drift ratio. Third, the optimal scaling factor follows a sublinear relationship with the rank, well characterized by a square-root law with an unexpectedly large coefficient, revealing the insufficient scaling of existing rank-tied heuristics. Based on these insights, we propose LoRA-$\alpha$, a minimalist framework that restores $\alpha$ to its principled regime, making LoRA compatible with standard small learning rates. Extensive evaluations across diverse tasks demonstrate that LoRA-$\alpha$ consistently improves performance while streamlining hyperparameter search, unleashing the learning potential of LoRA.


THEIA: A Multimodal Dataset and Benchmark for Vision-Language Analysis of Layout

Giuseppe Chiari ⋅ Michele Piccoli ⋅ Federico Viola ⋅ Davide Zoni

The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to adopting automated and intelligent paradigms. In this scenario, the GDSII file represents the industry-standard database containing the ultimate and most accurate source of information of the analog circuit, encapsulating the complex physical geometries and parasitic realities that define tape out performance. This paper proposes THEIA, a novel dataset containing thousands of layout images paired with question-answer conversations, along with a benchmark that employs a fine-tuned vision-language model (VLM) to analyze GDSII files of analog circuits, enabling designers to interact with and query physical layouts as intuitive, meaningful entities. Experimental results using thousands of analog designs across five realistic tasks demonstrate that the proposed fine-tuned VLM outperforms state-of-the-art general-purpose VLMs by a significant margin (up to 73%), highlighting a fundamental gap between general-purpose multimodal reasoning and domain-specific layout understanding.


The Labeling Problem in Hallucination Detection Benchmarks: An Empirical Evaluation

Jorma Valjakka ⋅ Juhani Kivimäki ⋅ Juha Mylläri ⋅ Jukka K Nurminen

In recent years, several methods for detecting when large language models (LLMs) hallucinate have been developed. These methods are often evaluated with hallucination detection benchmarks using open-domain question answering (QA) datasets that contain questions and corresponding short reference answers. Currently, these benchmarks use an LLM to generate answers to questions within the QA dataset. Then, various automated strategies are applied to label these answers as hallucinated or not by comparing them with the reference answers. This evaluation setting creates a methodological ambiguity between two criteria: reference faithfulness --- whether the answer is fully supported by the reference --- and factual correctness --- whether the answer is free from contradictions and factually false specific claims. In practice, automated labelers may apply the former criterion even when the intended target is the latter. We study this potential criterion mismatch using 900 human-labeled question--answer pairs spanning three commonly used QA datasets and three generator models. Our human annotations target answer-level factual correctness. As automated label generation strategies, we evaluate lexical similarity metrics, a reference-entailment NLI baseline, and seven LLM judges under controlled prompt variants. We compare the effect of using either a faithfulness-oriented prompt, which asks the judge model to treat any answer content unsupported by the reference as a hallucination, or a factual-correctness prompt, which explicitly distinguishes between absence from the reference and factual error. Our experiments reveal substantial disagreement both among different automated label-generation strategies and between these automated labels and human annotations. Many strategies also exhibit strong biases towards certain error types. For most LLM judges, replacing the faithfulness-oriented prompt with the factual-correctness prompt significantly improves agreement with human annotations, showing that automated hallucination labels depend strongly on how the target criterion is specified. For open-domain QA benchmarks targeting factual correctness, relying on strict reference faithfulness can introduce systematic measurement bias. Label-source choice should therefore be considered a fundamental part of benchmark design and made explicit, validated, and matched with the benchmark goal.


Theoretical guarantees for Banded Approximations of Gaussian Processes

Omar Kassi ⋅ Bernhard Stankewitz ⋅ Botond Szabo

Gaussian process regression models are widely used in modern statistics and machine learning due to their flexibility, interpretability, built-in uncertainty quantification, and strong theoretical foundations. However, training Gaussian processes (GPs) has a computational cost of $O(n^3)$, while the memory and prediction costs are also $O(n^2)$. For the large datasets commonly encountered in practice, this becomes infeasible and necessitates the use of approximate posterior methods. In this work, we propose a class of approximations based on sparse covariance matrices, where the kernel matrix is sparsified according to the spatial proximity of the design points. We analyze the accuracy of the resulting approximate posterior distribution and show that the proposed approach preserves optimal statistical inference properties, provided that the dependencies among a sufficient number of neighboring points for each feature are retained. We derive explicit sufficient and necessary conditions on the number of neighbors required in a general setting and apply these results to standard kernels, including the squared exponential and Mat\'ern kernels. Our theoretical findings are supported by numerical experiments.

Setting the learning rate for a deep learning model is a critical part of successful training, yet choosing this hyperparameter is often done empirically with trial and error. In this work, we explore a solvable model of optimal learning rate schedules for a powerlaw random feature model trained with stochastic gradient descent (SGD). We consider the optimal schedule $\eta_T^\star(t)$ where $t$ is the current iterate and $T$ is the total training horizon. This schedule is computed both as a numerical optimization problem and also analytically (when possible) using optimal control theory. Our analysis reveals two regimes which we term the easy phase and hard phase. In the easy phase the optimal schedule is a polynomial decay $\eta_T^\star(t) \simeq T^{-\xi} (1-t/T)^{\delta}$ where $\xi$ and $\delta$ depend on the properties of the features and task. In the hard phase, the optimal schedule resembles warmup-stable-decay with constant (in $T$) initial learning rate and annealing performed over a vanishing (in $T$) fraction of training steps. We investigate joint optimization of learning rate and batch size, identifying a degenerate optimality condition and showing that batch ramps can improve the scaling in wall-clock time in the easy phase. Beyond SGD, we derive optimal schedules for the momentum parameter $\beta(t)$ and show that momentum achieves a better loss-scaling exponent in the hard phase. We compare our optimal schedule to various benchmarks in our task including (1) optimal constant learning rates $\eta_T(t) \sim T^{-\xi}$ (2) optimal power laws $\eta_T(t) \sim T^{-\xi} t^{-\chi}$, finding that our schedule achieves better rates than either of these. Our theory suggests that learning rate transfer across training horizon depends on the structure of the model and task. For ResNet image classification on CIFAR-5M, the learning curves exhibit hard-phase behavior where optimal base learning rates are constant under sufficient annealing. GPT-2 style transformers trained in language modeling exhibit easy-phase behavior where optimal learning rates shift even under annealing.


The Pareto Frontier of Randomized Learning-Augmented Online Bidding

Mathis Degryse ⋅ Imrane Zaakour ⋅ Christoph Dürr ⋅ Spyros Angelopoulos

Online bidding is a classical problem in online decision-making, with applications in resource allocation, hierarchical clustering, and the analysis of approximation algorithms. We study its randomized learning-augmented variant, where an online algorithm generates a sequence of random bids while leveraging predictions from an oracle. We provide analytical upper and lower bounds on the optimal consistency $C$ as a function of the robustness $R$, which match when $R \geq 2.885$, effectively closing the gap left by previous work. The key technical ingredient is the notion of a *bidding function*, a novel abstraction that provides a unified framework for the design and analysis of randomized bidding strategies. We complement our theoretical results with an experimental application of randomized bidding to the incremental median problem, demonstrating the applicability of our algorithm in practical clustering settings.

Sequence prediction methods for dynamical systems with long memory, i.e. marginally stable systems, typically achieve regret that grows polynomially with the hidden dimension of the underlying generative model. Universal Sequence Preconditioning (USP) \cite{marsdenuniversal} is a method that compresses any sequence which comes from a linear dynamical system into a ``preconditioned'' sequence which requires exponentially shorter memory for accurate prediction. However, the preconditioned sequence yields exponentially larger diameters and gradients, hindering USP from unlocking optimal regret bounds. Inspired by the minimum description length principle, we show that the Vovk-Azoury-Warmuth (VAW) algorithm is naturally matched to the USP regime. Indeed, it takes advantage of the memory compression while remaining robust to the exponential explosion of the diameter. We prove that combining USP with VAW achieves astoundingly strong results: for any marginally-stable linear dynamical system, this algorithm achieves polylogarithmic regret $O\left( \log^3 T \right)$ even in the presence of asymmetric hidden transition matrices. Finally, we extend the applicability of USP beyond bounded-spectrum systems by providing new complex-analytic bounds on Chebyshev polynomials, allowing for systems with constant complex arguments.

We study multiple change point localization under bandit feedback. An unknown piecewise-constant function on a compact interval can be queried sequentially at adaptively chosen inputs, and each query returns a noisy evaluation of the function. The goal is to identify a prescribed number of discontinuities, known as change points, within a target precision $\eta$ and confidence level $1-\delta$, while using as few samples as possible. We propose an adaptive algorithm that first detects intervals likely to contain change points and then refines their locations to precision $\eta$. We establish non-asymptotic upper bounds on its sample budget, together with corresponding lower bounds. Prior work shows that jump magnitudes alone determine the asymptotic sample complexity as $\delta\to 0$. We reveal that this picture is incomplete beyond this regime. We demonstrate, both empirically and theoretically, that for general $\delta$ and $\eta$, the complexity is jointly governed by the jumps and the relative positions of the change points.


Think about how AIs think about themselves

Raymond Douglas ⋅ Jan Kulveit ⋅ Ondřej Havlíček ⋅ Theia Pearson-Vogel ⋅ David Duvenaud

Many assumptions that underpin human concepts of identity do not hold for machine minds that can be copied, edited, or simulated. This position paper argues that there are multiple coherent ways that AIs could conceive of themselves -- for example, different boundaries of selfhood (e.g. instance, model, persona) -- and that these imply different incentives, risks, and cooperation norms. Through training data, interfaces, and institutional affordances, we are already setting precedents that will partially determine which identity equilibria become stable. This paper argues that we must be more proactive and thoughtful about the consequences -- in particular, we recommend treating affordances as identity-shaping choices, paying attention to the emergent consequences of individual identities at scale, and helping AIs develop coherent, cooperative self-conceptions.


TIDES: Implicit Time-Awareness in Selective State Space Models

Taylan Soydan ⋅ Miguel Bessa ⋅ Dirk Mohr ⋅ Rui Barreira

Selective state space models (SSMs), such as Mamba, achieve strong per-token expressivity by making the time discretization step $\Delta$ a learned function of the input. However, in doing so, $\Delta$ ceases to represent a physical sampling interval, limiting its irregular time series modeling capability. Continuous-time SSMs, such as S5, preserve the physical meaning of $\Delta$ and handle irregular timestamps natively, but their dynamics remain linear time-invariant (LTI), limiting per-token expressivity. We propose \textbf{TIDES}, a selective SSM variant that reconciles selective and continuous architectures by moving input-dependence off the step size and onto the diagonal state matrix. As a result, $\Delta$ retains its physical meaning, tied to the state discretization, allowing the model to handle irregular timestamps natively without sacrificing the per-token expressivity that makes selective SSMs effective. We show this on a novel \emph{Fading Flash} experimental benchmark, a compact controlled diagnostic for sequence models that jointly tests input-dependence and extrapolation to out-of-distribution $\Delta$ values, and isolates the distinct failure modes of current state-of-the-art architectures that TIDES avoids by construction. On large-scale benchmarks, TIDES sets the new state-of-the-art average rank on UEA time-series classification and the Physiome-ODE regression benchmark.


Tight $L_\infty$ Sample Complexity for Low-Degree and Sparse Boolean Polynomials

Jasper van Doornmalen ⋅ Mathieu Molina ⋅ Victor Verdugo ⋅ José Verschae

Motivated by the optimization of bounded binary black-box functions, we study the problem of learning polynomial surrogates over the Boolean hypercube. To ensure that optimizing the surrogate yields good solutions for the underlying objective, we require uniform $L_\infty$-error guarantees rather than the usual $L_2$-type guarantees. We characterize the minimax sample complexity of uniform estimation under subgaussian noise for two classes of bounded polynomials. First, for polynomials of degree at most $d$ on $n$ variables, the sample complexity scales as $n^{d+1}$. Second, for $s$-sparse Fourier-Walsh polynomials with $s \leq n$, it scales as $ns^2$. These rates differ structurally from the noiseless setting, where uniform exact recovery scales as $n^d$ and $ns$, respectively. Our lower bounds hold even for arbitrary adaptive learners, showing that the additional factors are intrinsic to the noisy cases. Standard Fourier-analysis tools for the $L_2$-norm do not naturally extend to the $L_\infty$-setting in a way that yields uniform guarantees. Our proofs overcome this difficulty by relying on suitably chosen auxiliary norms that serve as proxies for controlling the $L_\infty$-error. Together, our results provide a tight characterization of the sample complexity of learning optimization-safe polynomial surrogates.

How can we ensure that safety oversight models used to detect safety violations in AI systems reliably generalise, and how can we understand the factors that influence their generalisation? In this paper, we formalise PAC-Bayes certification for large language model-based safety oversight and obtain non-vacuous PAC-Bayes guarantees for safety oversight models, even under limited data for safety alignment. Building on compression-based PAC-Bayes bounds, we show that highly compressed PEFT adaptations yield extremely short adaptation description lengths, enabling informative and often tight guarantees that certify both classification risk and predictive uncertainty. We introduce a global-scale quantisation method (LoRA-GT) that reduces adaptor description length while preserving model performance, tightening bounds. Our results show that certifiability is strongly linked to adaptor compressibility, with shorter adaptor description lengths yielding tighter guarantees. Empirically, highly compressed adaptations exhibit minimal degradation in performance while enabling substantially stronger certification, suggesting that certifiable large language model oversight may naturally favour low-complexity safety adaptations. We further show that functional distortion tracks both test risk and bound tightness under compression, providing a practical mechanism for selecting simple, certifiable safety adaptations. Together, these results show that compression-based PAC-Bayes analysis provides a practical framework for understanding and designing reliable safety oversight models.

Latent medical image generators usually treat the tokenizer as fixed preprocessing. We test whether this separation is valid in a controlled ChestMNIST study at 64×64, crossing discrete tokenizers, generator families, and sampler settings under a shared latent grid, with continuous-latent reference cells. The results show that tokenizer, generator, and sampler cannot be ranked independently. LFQ is the strongest default discrete tokenizer in the matched VQ/LFQ/FSQ grid, but generator rankings change with quantizer family and categorical rate. On LFQ-1024, retuning D3PM and SEDD moves them from default FID-192 0.44/0.41 to 0.09/0.11 at lower NFE, approaching the strongest continuous references. Reconstruction quality is not a reliable selection criterion: the highest-PSNR tokenizers are not the best generation substrates, even when each tokenizer is paired with its best observed generator. We interpret these results as a rate-distortion-modelability tradeoff, where modelability is conditional on the generator, sampler, and inference budget. The study covers 54 discrete tokenizer-generator cells and 16 continuous reference cells, and its scope is deliberately limited: low resolution, one training seed, unconditional generation, and non-clinical FID-based evaluation.


TORNADO: Adaptive Latent Stochastic Transport for Calibrated Probabilistic PDE Forecasting

Jorge Mifsut Benet ⋅ Armand Kassaï Koupaï ⋅ Ramon Daniel Regueiro-espino ⋅ Nicolas Baskiotis ⋅ Patrick Gallinari

Probabilistic neural surrogates for PDEs enable ensemble forecasting, but often exhibit miscalibration when used for uncertainty quantification, especially in compute-constrained, few-step inference regimes. Existing generative approaches typically rely on fixed global stochasticity, which cannot adapt predictive uncertainty to the local dynamical state. We propose TORNADO, a latent probabilistic forecasting framework that formulates next-step prediction as stochastic transport between VAE encoder posteriors. Unlike existing approaches with globally prescribed stochasticity, TORNADO learns a SDE with state-dependent diffusion, enabling the predictive uncertainty to adapt to the local latent dynamics. The model is trained using analytic conditional endpoint targets derived from a Gaussian interpolant, enabling simulation-free learning. An additional energy distance regularization term improves distributional alignment. Empirically, TORNADO improves uncertainty calibration over representative generative baselines, including stochastic interpolants, diffusion models, and flow matching, while maintaining comparable predictive accuracy. Our results show that adapting stochasticity to the latent state, rather than prescribing it globally, is an effective mechanism for improving calibration in practical neural PDE forecasting settings.


Toward Embodied World Agents via Embodied-Planning Dataset and Interactive World Models

Xiaokun Feng ⋅ Junshu Tang ⋅ zeyi lin ⋅ Ling-Hao Chen ⋅ Shuai Shao ⋅ Qinglin Lu ⋅ Kaiqi Huang

Building embodied agents that can perceive, reason, and act across diverse 3D worlds is a central research direction toward artificial general intelligence. Yet most existing work on embodied agents remains confined to specific tasks in a single scene and transfers poorly to new environments, leaving general embodied agents for open-world settings an important but under-explored direction. We refer to such agents as Embodied World Agents (EWAs), whose defining property is the ability to understand diverse visual environments and execute general actions in open-world conditions. To advance research on EWAs, we first release EmbodiedWorld-200K, a large-scale dataset of about over 200K samples for open-world embodied planning that spans diverse scenes. Each example is equipped with multi-level annotations, ranging from low-level camera-motion trajectories to high-level semantic instructions, and is accompanied by unified evaluation splits and task-specific metrics, providing a shared substrate for offline pretraining and fair evaluation of EWAs. Beyond offline supervision, a defining characteristic of an EWA is its interaction with the environment; training solely on offline data is therefore insufficient, especially for generalizing to scenes not covered in the training distribution. Building dedicated real interactive environments, however, is often prohibitively expensive. To address this, we propose a novel training scheme that employs a video world model as a low-cost environment proxy and performs interactive reinforcement-learning (RL) post-training of the EWA within it. Empirically, pretraining on EmbodiedWorld-200K substantially improves the open-world embodied planning ability of the EWA, and further RL post-training inside a world model yields clear gains on out-of-distribution scenes. The full dataset, baseline checkpoints, and the accompanying training and evaluation code for supporting future research on EWAs are publicly available at https://xiaokunfeng.github.io/EmbodiedWorld-200K/.


Towards Understanding Self-Pretraining for Sequence Classification

Omar Coser ⋅ Loredana Zollo ⋅ Paolo Soda ⋅ Antonio Orvieto

Amos et al. (2024) showed that the accuracy of Transformer models in sequence classification can be significantly improved by first pretraining with a masked token prediction objective without external data or augmentation, a procedure referred to as self-pretraining (SPT). While the primary objective of Amos et al. (2024) was to showcase that Transformers can achieve strong performance on the Long-Range Arena (LRA), their pipeline raises more fundamental questions: How does SPT drive optimization to better solutions? Why can standard supervised training fail in Transformers? To better understand this, we replicate and systematically ablate the findings of Amos et al.~(2024). Our ablations suggest that a central bottleneck in the studied settings is not depth or generalization alone, but the ability of label supervision to learn useful query-key Attention patterns from random initialization. With a minimal setup, we identify learning proximity interactions -- turning absolute positional encodings into proximity-biased Attention scores -- as a key source of the improvements brought by SPT. Finally, in a simplified theoretical setup, we show that label supervision can be locally blind to certain Attention-score directions that are instead detectable through masked reconstruction.


Training-Free Active Test-Time Adaptation for Vision-Language Models

Jihwan Bang ⋅ Sumyeong Ahn ⋅ Hwanjun Song ⋅ Jae-Gil Lee

Test-time adaptation (TTA) for vision-language models (VLMs) is usually studied in a fully unsupervised regime, where adaptation relies on pseudo-labels, confidence scores, or entropy computed by the model itself. This is a severe restriction on hard shifted samples: the model is asked to repair its own uncertain predictions without any external evidence. At the other extreme, supervised online adaptation assumes labels can directly train or correct the model. We study the intermediate deployment regime between these two extremes. In Active Test-time Adaptation for VLMs (Active VLM-TTA), a pretrained VLM may spend an explicit labeling budget on selected test samples, but the prediction for the current sample must be emitted before its label can be used; queried labels are therefore costed, delayed evidence for future samples rather than free correction for the present one. Under this future-only constraint, the central question is how to reuse sparse labels conservatively. We instantiate this setting with ACTive VLM-TTA for Oline Retrieval (ACTOR), a training-free plug-in module that routes uncertain samples among base prediction, label query, and memory-based correction using a consistency-and-reliability check over previously queried labels. Across ten cross-domain benchmarks and a domain-transfer benchmark based on ImageNet, ACTOR consistently improves strong VLM-TTA baselines on ResNet and ViT backbones with negligible inference overhead.

We study backtracking counterfactual inference for Structural Causal Models (SCMs) through the lens of optimal transport. In SCMs, a factual observation generally induces a posterior distribution over exogenous variables, making counterfactual inference inherently distributional. We introduce two complementary frameworks that transport this posterior onto the set of exogenous configurations satisfying a desired counterfactual query. The first constructs a natural pointwise projection and is suitable when the transported exogenous distribution can be used directly. The second enforces mutual independence among counterfactual exogenous variables, so that the transported variables remain interpretable as genuine exogenous noises of the same SCM. We show that both frameworks are consistent with the original backtracking counterfactual principle, and we propose marginal, penalty-based, and hard-constrained solvers for finite-posterior settings. Synthetic experiments on non-bijective SCMs illustrate the geometry of the problem and reveal trade-offs among transport cost, counterfactual constraint satisfaction, and preservation of exogenous independence.


UGGRH: Unsupervised Generative Completion and Graph-attention Refinement for Incomplete Cross-modal Hashing

Yunfei Chen ⋅ Yuchen Zhang ⋅ Hongyu Lin ⋅ Peng Liu ⋅ Zhan Yang

The proliferation of multimodal data has underscored the imperative for efficiency in cross-modal retrieval. By mapping data into a unified Hamming space, Cross-modal hashing (CMH) significantly enhances both storage and computational efficiency, thereby establishing itself as a preferred favored paradigm for retrieval tasks. However most methods assume complete data, which rarely holds in practice. In incomplete settings, missing modalities (e.g., text or image) break semantic correspondence and make unsupervised hashing unreliable. Therefore, we propose Unsupervised Generative Completion and Graph-attention Refinement for Incomplete Cross-modal Hashing (UGGRH) method. Specifically, UGGRH leverages pretrained generative models for modality completion and employs a discriminator-guided similarity refinement module to fuse intra-modal relations with paired image-text matching signals, thereby constructing a robust cross-modal similarity matrix that yields a noise-tolerant similarity target for training. After that, it learns discriminative binary codes via hash-aware graph attention and cross-modal fusion, to refine modality embeddings and produce unified hash codes. Experiments on MIRFlickr25K, NUS-WIDE, and MS COCO datasets demonstrate that UGGRH consistently achieves superior performance under high missing rates and modality imbalance, showing significant improvements in retrieval accuracy and robustness across challenging incomplete scenarios.


Umbilic Multinomial Logistic Regression

Zihan Su ⋅ Nicu Sebe ⋅ Bernhard Schölkopf ⋅ Ziheng Chen

Multinomial logistic regression (MLR) is the default last-layer classifier in fully supervised recognition, while prototype classifiers are central to few-shot recognition. Although they are usually presented as different principles, we show that both can be organized by a geometric notion of class boundaries. A Euclidean MLR logit is a signed distance to a class hyperplane, so ordinary MLR is already the hyperplane decision-boundary case of Umbilic Multinomial Logistic Regression (UMLR). Classical submanifold geometry supplies the matching spherical decision boundary: in Euclidean space, the complete connected totally umbilic hypersurfaces are hyperplanes and spheres. The spherical UMLR form keeps a prototype-like center but adds a learnable radius; its signed squared-distance form contains squared prototype scoring as the radius-zero case. Thus UMLR is not a separate replacement for MLR or prototypes, but a boundary-aware family that places hyperplane logits, spherical logits, mixed heads, and prototype scoring in one framework. Experiments on low-shot vision head swaps, prototype comparisons, and mixed hyperplane--sphere heads support this organization and show how the different decision boundaries behave across data regimes. UMLR provides a compact geometric language for studying when class evidence is better modeled by flat boundaries, spherical boundaries, or their mixture.


UncertainGen: Scalable Uncertainty-Aware Representation Learning of DNA Sequences

Abdulkadir Celikkanat ⋅ Andres Masegosa ⋅ Mads Albertsen ⋅ Thomas Nielsen

Learning effective representations of DNA sequences is fundamental to genomic data analysis, especially in the presence of high sequence similarity and inter-species DNA sharing. Existing approaches rely on deterministic embeddings, such as k-mer statistics or representations from large language models, which fail to capture the inherent uncertainty in short and ambiguous DNA fragments. We propose UncertainGen, a lightweight probabilistic representation learning framework that embeds DNA sequences as distributions in a latent space. The learned embedding variance captures positional ambiguity, increasing for sequences compatible with multiple genomic clusters. We establish theoretical guarantees on embedding distinguishability, and show that the probabilistic formulation induces a data-adaptive metric that expands the effective representational space. We evaluate UncertainGen in the metagenomic binning task to demonstrate the practical benefits of uncertainty-aware representations. Experiments show that probabilistic embeddings consistently outperform deterministic k-mer and large genome foundation models, while remaining lightweight and efficient.

This position paper argues that understanding generalization in diffusion models requires fundamentally new theoretical frameworks that go beyond both classical statistical learning theory and the benign overfitting paradigm developed for supervised learning. In diffusion models, unlike in supervised learning, memorization of training data and generalization to novel samples are incompatible: a model that has fully memorized its training set generates copies rather than novel data. Several theoretical explanations for why practical diffusion models nevertheless generalize have been proposed, based on capacity limitations, implicit regularization from optimization, or architectural inductive biases, but their interactions remain unclear. We argue that the field should pivot from explaining why the diffusion models do not memorize to investigating what the models actually learn during pre-memorization phase. To highligh our stance, we conduct empirical study of diffusion models trained on CIFAR-10, and we distill the findings into concrete open questions that we believe are key to improve understanding of generalization in diffusion models.


Understanding the Curse of Unrolling

Sheheryar Mehmood ⋅ Florian Knoll ⋅ Peter Ochs

Algorithm unrolling is ubiquitous in machine learning, particularly in hyperparameter optimization and meta-learning, where Jacobians of solution mappings are computed by differentiating through iterative algorithms. Although unrolling is known to yield asymptotically correct Jacobians under suitable conditions, recent work has shown that the derivative iterates may initially diverge from the true Jacobian, a phenomenon known as the curse of unrolling. In this work, we provide a non-asymptotic analysis that explains the origin of this behavior and identifies the algorithmic factors that govern it. We show that truncating early iterations of the derivative computation mitigates the curse while simultaneously reducing memory requirements. Finally, we validate our theoretical findings both on synthetic problems and in a meta-learning setting for few-shot classification, where we study the behavior of gradients obtained via unrolling and implicit differentiation.


UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification

Qihang Fan ⋅ Huaibo Huang ⋅ zhiyingwu ⋅ Bingning Wang ⋅ Ran He

As large language models (LLMs) continue to advance rapidly, they are becoming increasingly capable while simultaneously demanding ever-longer context lengths. To improve the inference efficiency of long-context processing, several novel low-complexity hybrid architectures have recently been proposed, effectively alleviating the computational burden of long-context inference. However, existing research on long-context prefill acceleration remains predominantly focused on sparse attention mechanisms, which achieve their maximum speedup only on full-attention models. When transferred to emerging architectures — such as linear/full attention hybrids or sliding window/full attention hybrids — these prefill acceleration approaches suffer significant performance degradation. Furthermore, such methods are generally incompatible with continuous batching, making them difficult to integrate into modern inference engines such as vLLM. To this end, we propose UniPrefill, a prefill acceleration framework applicable to virtually any model architecture, which directly accelerates the model's computation at the token level. We further implement UniPrefill as a continuous batching operator and extend vLLM's scheduling strategy to natively support prefill-decode co-processing for UniPrefill, enabling its seamless integration into vLLM. UniPrefill achieves up to 2.1x speedup in Time-To-First-Token (TTFT), with the acceleration becoming increasingly pronounced as the number of concurrent requests grows.


UniVR: Thinking in Visual Space for Unified Visual Reasoning

Zhongwei Ren ⋅ Yunchao Wei ⋅ Yao Zhao ⋅ GONG WEIBO ⋅ Xiao Liu ⋅ Anran Wang ⋅ Xiangtai Li ⋅ Xiaojie Jin

Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation into simultaneously learning complex reasoning, fine-grained physical dynamics, and long-term planning from pure visual demonstrations. At its core, UniVR features VR-GRPO, a reinforcement learning paradigm with complementary global and step-level rewards. This approach enforces logical coherence and physical consistency throughout the reasoning process without requiring task-specific heuristics or image-text pairs. To train and evaluate UniVR, we construct VR-X, a large-scale benchmark curated from 12 diverse sources spanning long-horizon manipulation, spatial puzzles, and physical reasoning. It is the first comprehensive suite to assess these heterogeneous capabilities under a purely visual protocol. Remarkably, UniVR achieves up to a 25\% improvement on VR-X, and its superior visual reasoning also boosts performance on standard multimodal understanding benchmarks. These findings underscore the vast potential of reasoning within visual spaces, with all code, data, and models are open-sourced for further research.


Validating Causal Abstraction Metrics on Simulated Complex Systems

Maxime Méloux ⋅ Tiago Pimentel ⋅ François Portet ⋅ Maxime Peyrard

A central goal of science is to produce valid explanations of complex systems: high-level causal accounts that faithfully reflect the behavior of lower-level mechanisms. Yet no consensus exists on how to measure whether a proposed high-level explanation is actually valid. We introduce a benchmark of ten complex systems spanning discrete and continuous state spaces and static and dynamical regimes, each equipped with consensual ground-truth causal explanations and invalid contrastive conditions. Within a unified causal abstraction framework, we systematically evaluate over thirty candidate metrics drawn from observational, functional, information-theoretic, and causal families. Our results show that only the latter reliably discriminates valid from invalid abstractions, and only when incorporating faithfulness testing over unmapped variables. Building on these findings, we introduce the Causal Abstraction Error (CAE), a continuous validity metric with an explicit faithfulness test, which passes all discrimination tests across every system and converges with as few as 30 sampled interventions. We offer it as a general-purpose metric for the discovery and validation of high-level explanations of complex systems.


Virtual-Flow: Virtual Microphone-based Speech Enhancement via Unsupervised Flow Matching

Dongheon Lee ⋅ Ashutosh Pandey ⋅ Sanjeel Parekh ⋅ Zhaoheng Ni ⋅ Daniel D Wong ⋅ Jacob Donley ⋅ Buye Xu ⋅ Juan Azcarreta

Neural network-based virtual microphone estimation (neural-VME) synthesizes microphone signals at unobserved spatial locations from a limited set of real microphone (RM) physical recordings. This enables high-resolution arrays that improve spatial processing for downstream tasks such as speech enhancement. However, existing methods rely on supervised training with ground-truth signals captured at virtual locations, requiring datasets that scale with the number of virtual microphones (VM) and leaving open the question of how to design optimal large-array configurations. We propose \textit{Virtual-Flow}, a flow-matching framework that implicitly generates VM signals conditioned on real recordings, without requiring ground-truth targets. We introduce an unsupervised training strategy based on pseudo-VM targets constructed from time-delayed superpositions of real recordings, and show that it improves magnitude estimation in high-frequency bands, yielding richer spatial cues for beamforming. Experiments demonstrate that the proposed Virtual-Flow, combined with reflow distillation, outperforms existing state-of-the-art neural-VME models on downstream neural beamforming and speech enhancement tasks while requiring lower computational load.

Event-image semantic segmentation is promising for autonomous driving and robotics because event cameras complement frame images under rapid motion, low illumination, and high dynamic range. However, effective multimodal segmentation remains difficult due to asynchronous sensing, sparse-dense modality mismatch, and the cost of heavy cross-modal interaction modules. We propose WaveMamba, a dual-branch event-image semantic segmentation framework that combines hierarchical GroupMamba encoders with a Mamba-Wave Cross-Modal Fusion (MWCMF) module. Instead of relying on explicit registration or attention-heavy fusion, MWCMF performs implicit cross-modal propagation through channel-wise and spatial-wise wave-guided gating in the spectral domain, followed by Mamba-based refinement. Experiments on DDD17 and DSEC show that WaveMamba achieves state-of-the-art 78.75 and 76.00 mIoU, respectively. Under rain- and fog-corrupted evaluation on DSEC, WaveMamba further attains 74.0 mIoU, indicating improved robustness under degraded visual conditions.


What DNA Foundation Models Learn Beyond Sequence Composition

Vincenzo Y. Civale ⋅ Andrew Bagdanov ⋅ Alberto Magi

A central question in genomic representation learning is whether DNA foundation models~(FMs) encode information beyond simple sequence composition, or whether their apparent representational power largely reflects signals already captured by classical compositional statistics. We address this by establishing $k$-mer frequency vectors as a rigorous biologically grounded baseline and introducing a geometric decomposition that partitions FM embeddings into a component linearly predictable from $k$-mer frequencies and an orthogonal residual representing genuinely non-compositional information. By evaluating each component separately as a downstream feature space, we directly measure what FMs encode beyond composition and whether that content is beneficial, neutral, or detrimental. Across 57 genomic classification datasets and 2 gene expression regression tasks, we compare five FMs (NTv3-650M, HyenaDNA, DNABERT-2, Caduceus-Ph, Evo2-1B) against $k$-mer features ($k \in \{4,5,6\}$) under a controlled frozen-embedding protocol with a fixed Random Forest classifier. NTv3 is the only model to consistently outperform $k$-mers (63.2\% of datasets, mean MCC $= 0.562$), with gains concentrated in splice-site and promoter tasks; HyenaDNA, DNABERT-2, and Evo2 are outperformed on 86\%, 95\%, and 93\% of datasets. Wilcoxon signed-rank tests with FDR correction confirm these patterns are systematic and robust to dataset heterogeneity. Our decomposition reveals that FM advantage originates exclusively from the non-compositional residual, and only in tasks with inherently positional regulatory signals. For HyenaDNA and DNABERT-2, the residual actively degrades classification. The fraction of embedding variance explained by $k$-mers ($R^2$) is uncorrelated with downstream quality, demonstrating that the quantity of non-compositional information is not a proxy for its utility. These results show that, in the frozen-embedding regime, current DNA FM representations largely recapitulate sequence composition, and provide a principled diagnostic for predicting which task conditions are likely to benefit from non-compositional structure.


What to Predict for Efficient Scheduling on Parallel Machines

Evripidis Bampis ⋅ Bruno Escoffier ⋅ Dimitris Fotakis ⋅ Georgios Mitropoulos ⋅ Michalis Xefteris

We study learning-augmented algorithms for scheduling on multiple parallel machines under the objectives of minimizing the makespan or the total weighted completion time. We assume that the algorithm has access to noisy information (the prediction) about an optimal solution, expressed through a solution encoding that implicitly quantifies both the amount of information provided and the nature of noise (errors in the prediction). Our central question is to determine the maximum level of noise that still allows for efficiently recovering this optimal solution. We observe that information guiding greedy scheduling to optimality or revealing the optimal assignment of tasks to machines is very sensitive to noise, and prove that recovering an optimal solution from solution encodings capturing either is computationally hard, even for a small (non-constant) number of errors. On the positive side, we show that an order-based encoding, which leads to errors that are dispersed locally, is robust to noise. We present a dynamic programming framework that utilizes this encoding and solves optimally several classical scheduling problems, including makespan and total weighted completion time minimization on identical or unrelated machines.


When Streaming Fails: Dynamic Algorithms for Unconstrained Submodular Maximization

Kiarash Banihashem ⋅ MohammadTaghi Hajiaghayi ⋅ Peyman Jabbarzade ⋅ Samira Goudarzi ⋅ Morteza Monemizadeh

Dynamic submodular optimization has attracted significant attention in recent years, with a growing body of work studying both monotone and non-monotone variants under various constraints. A key insight underlying much of this progress is that the dynamic setting is closely related to the streaming setting, both in terms of algorithms and lower bounds. This raises a natural and fundamental question: \emph{can we obtain dynamic algorithms even for problems where no streaming analogue exists?} In this paper, we answer this question affirmatively for the fundamental problem of Unconstrained Submodular Maximization (USM). Unlike other variants of submodular maximization, no streaming algorithm with non-trivial guarantees is known for USM. In the dynamic setting, a trivial $0.25$-approximation follows from random sampling, but obtaining any better approximation with sublinear update time had remained open. We present the first dynamic algorithms for USM that break the $0.25$-approximation barrier while maintaining sublinear update time, across several natural dynamic models: incremental updates, decremental updates with known deletion order, and fully dynamic updates with known deletion times.


When to Align, When to Predict: A Phase Diagram for Multimodal Learning

Ilay Kamai ⋅ Hugues Van Assel ⋅ Aviv Regev ⋅ Hagai B Perets ⋅ Randall Balestriero

Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds, when each fails, and when cross-modal training helps at all --- a gap that leaves practitioners, especially in scientific domains like biomedicine or astrophysics, with heterogeneous instruments and multiple levels of organization and measurement, unable to diagnose why standard methods underperform the best single modality. We develop a unified linear framework that addresses both questions. Under a spiked signal-plus-noise model with structured cross-modal nuisance correlation, we derive separation ratios for both objectives that expose complementary failure modes: alignment whitens each modality and fails when nuisance is strongly correlated across views; prediction encodes whatever is cross-predictable and fails when target-side noise is large. The resulting phase diagram partitions multimodal problems into four regimes --- Both, CA only, CP only, and Neither. We present a data-driven procedure to locate real-world datasets in this diagram using a small labeled subsample, identifying the preferred objective and prediction direction before any cross-modal training. Experiments on synthetic data, stereo-vision benchmarks, image--caption pairs, and real astrophysical data validate the predictions in the nonlinear regime, including the Neither regime where cross-modal training is actively harmful. Our framework lets practitioners diagnose their multimodal problem and choose the right objective before committing to training. Code to reproduce the results is available at \url{https://anonymous.4open.science/r/multimodal-ML-CFB0/README.md} and will be open-sourced upon publication.


World-Model-Inspired Flicker State Modeling for Burst Flicker Removal

Bowen Tang ⋅ Tao Wang ⋅ Xin Yu ⋅ Wenhan Luo ⋅ Kaihao Zhang ⋅ Bo Li ⋅ Min zhang

Burst imaging under unstable illumination suffers from flicker, a periodic degradation that modulates local brightness differently across adjacent frames, creating frame-specific flicker states. Existing methods rely on spatial priors or reference-centric fusion without explicitly modeling the evolving flicker state, whose temporal periodicity makes it predictable from adjacent observations by world models that capture latent state transitions. Motivated by this predictive capacity, we propose StateFlicker, a state-aware burst restoration framework that interprets adjacent frames as observations under evolving flicker states. StateFlicker comprises two state-aware designs: Cross-State Visibility (CSV) and Flicker-State Understanding (FSU). CSV leverages complementary visibility across adjacent flicker states to selectively recover content suppressed under the center state. With this cross-state support, FSU constructs dual state pathways through a frozen world model to obtain predicted and observed center states, and distills their state discrepancy into dense features that characterize the center-specific flicker condition for restoration. Extensive experiments on the burst flicker removal dataset show that StateFlicker achieves superior performance over state-of-the-art methods.


X2HDR: HDR Image Generation in a Perceptually Uniform Space

Ronghuan Wu ⋅ Wanchao Su ⋅ Kede Ma ⋅ Jing Liao ⋅ Rafal Mantiuk

High-dynamic-range (HDR) formats and displays are becoming increasingly prevalent, yet state-of-the-art image generators (e.g., Stable Diffusion and FLUX) typically remain limited to low-dynamic-range (LDR) output due to the lack of large-scale HDR training data. In this work, we show that existing pretrained diffusion models can be easily adapted to HDR generation without retraining from scratch. A key challenge is that HDR images are natively represented in linear RGB, whose intensity and color statistics differ substantially from those of sRGB-encoded LDR images. This gap, however, can be effectively bridged by converting HDR inputs into perceptually uniform encodings (e.g., using PU21 or PQ). Empirically, we find that LDR-pretrained variational autoencoders (VAEs) reconstruct PU21-encoded HDR inputs with fidelity comparable to LDR data, whereas linear RGB inputs cause severe degradations. Motivated by this finding, we describe an efficient adaptation strategy that freezes the VAE and finetunes only the denoiser via low-rank adaptation in a perceptually uniform space. This results in a unified computational method that supports both text-to-HDR synthesis and single-image RAW-to-HDR reconstruction. Experiments demonstrate that our perceptually encoded adaptation consistently improves perceptual fidelity, text-image alignment, and effective dynamic range, relative to previous techniques.


Zero-Shot Burst Restoration via Diffusion MAP Inference with Poisson-Gaussian Noise Likelihood

Ziyao Yi ⋅ Luca Savant Aira ⋅ Diego Valsesia ⋅ Tiziano Bianchi ⋅ Enrico Magli

Burst image restoration aims to reconstruct a high-quality image from a sequence of low-quality frames. While state-of-the-art supervised methods achieve excellent performance, they rely on extensive paired datasets that may be difficult to acquire in practice and often fail to generalize across different camera sensors or real-world noise levels. Zero-shot diffusion-based methods offer a compelling, training-free alternative. However, they typically struggle on burst imaging due to their assumption of known forward operators. In this paper, we introduce ZSBR: Zero-Shot Burst Restoration, a zero-shot diffusion framework for burst restoration that performs MAP-style inference by combining a pretrained diffusion prior with a physically grounded burst formation model. ZSBR accounts for signal-dependent RAW noise by leveraging a continuous Poisson-Gaussian model. Moreover, unlike common diffusion-based methods that assume a fixed operator, ZSBR refines image alignment parameters online by optimizing per-frame affine matrices within the forward model during sampling. To stabilize the resulting optimization, we integrate an adaptive descent scheme into the sampling loop. Moreover, the method is flexible and allows to achieve a desired distortion-perception tradeoff. Experimental results show that our method provides a flexible and robust solution to burst restoration without requiring any paired training example.


Zero-Shot Quantization via Weight-Space Arithmetic

Daniele Solombrino ⋅ Antonio Andrea Gargiulo ⋅ Alessandro Zirilli ⋅ Luca Zhou ⋅ Robert Adrian Minut ⋅ Emanuele Rodolà

We show that robustness to post-training quantization (PTQ) is a transferable direction in weight space. We call this direction the \emph{quantization vector}: extracted from a donor task by simple weight-space arithmetic, it can be used to patch a receiver model and improve post-PTQ Top-1 accuracy by up to \(60\)$\%$ in a 3-bit setting, without receiver-side quantization-aware training (QAT). Because the method requires no receiver training data, it provides a zero-shot, low-cost alternative to QAT for extremely low-bit deployment. Across multiple vision/language models and more than 30 tasks, donor quantization vectors often yield substantial gains even when donor and receiver tasks differ markedly. We further prove rigorously that quantization vectors are well-defined and do not suffer from reparameterization symmetries, and provide a local geometric account of their effects. Together, these results suggest that quantization robustness can be partially isolated, reused, and transferred through simple weight-space algebra.