Session
Paris Poster Session 6
Paris Poster Hall
Many modern learning approaches are still struggling with spatial reasoning tasks, i.e. they lack the ability to utilize geometric information of perceived entities and their spatial relation to each other to solve problems. We introduce a novel Adaptive Neural Cellular Automata (aNCA) architecture which uses deformable convolutions to dynamically adapt the perceptive field and iteratively reason over 2D spatial relations on grid-like data structures (e.g. images). Empirical results on public benchmarks show state of the art comprehensible results with high generalization abilities for solving image based puzzles like Sudoku or finding the shortest path in a maze.
Accuracy of Noise Ceiling Estimators for Brain Scores
Aude Maier ⋅ Lucas Gruaz ⋅ Martin Schrimpf ⋅ Johanni Brea
The internal activations of modern machine learning models align surprisingly well with the activities recorded in real brains. In the last five years, the alignment between encoding models and brain activities, often called brain scores, has steadily increased. Because brain recordings are noisy, there exists an upper bound to the brain scores that any model can achieve. This bound, called the noise ceiling, is crucial for two reasons: it determines whether further improvements in model performance are possible, and it enables meaningful comparisons of brain scores across benchmarks. The field has developed multiple approaches to estimate noise ceilings, but the statistical properties of existing estimators remain poorly understood, leaving it unclear in which experimental settings they provide reliable estimates. Here, we introduce a unified mathematical framework for brain recordings, use it to derive closed-form expressions for four commonly used noise ceiling estimators, and provide theoretical predictions for the behavior of each estimator. We validate the theoretical predictions with simulations that we compare to noise ceiling estimates in existing fMRI datasets. We find that the existing methods underestimate the true noise ceiling in many settings. We conclude with recommendations for best practice in different situations.
A Comprehensive View of Fairness through Distributional Stability
Gayane Taturyan ⋅ Charlotte Laclau ⋅ Stephan Clémençon
We view fairness as a property of distributional stability. Rather than assessing a predictor under a fixed data distribution, we study how its predictions change under perturbations that modify the composition of protected groups. A predictor is fair if it remains stable under such shifts. Under this perspective, several classical notions of fairness arise as stability with respect to specific perturbations, with the associated unfairness gap given by a Lipschitz constant of a prediction-rate functional. This formulation also yields guarantees that hold uniformly over a range of demographic compositions at test time, without requiring knowledge of the deployment distribution. It leads to a learning procedure based on convex combinations of reweighted predictors, formulated as a second-order cone program, for which we establish generalization bounds. Experiments on standard benchmarks illustrate the approach.
A Free Lunch in LLM Compression: Revisiting Retraining after Pruning
Moritz Wagner ⋅ Christophe Roux ⋅ Max Zimmer ⋅ Sebastian Pokutta
Post-training pruning can substantially reduce LLM inference costs, but it often degrades quality unless the remaining weights are adapted. Since global retraining is expensive at LLM scale, recent work has largely focused on increasingly sophisticated pruning criteria that aim to select better sparsity patterns without adaptation. We revisit this trade-off through local reconstruction: after pruning, we adapt one subset of the model parameters at a time on a calibration set, training it to match the corresponding intermediate activations of the dense model. We evaluate local reconstruction across model families and scales, up to 72B parameters, and establish three main findings. First, local reconstruction is an effective adaptation mechanism for LLMs: it matches post-pruning retraining while using over an order of magnitude less data and compute, even when using PEFT techniques. Second, reconstruction exhibits a broad "free-lunch" regime in granularity, i.e., the reconstruction parameter window: as long as the reconstructed region contains at least a nonlinear submodule, final quality is largely insensitive to the window size, allowing granularity to be chosen primarily based on memory constraints. In contrast, reconstructing individual matrices, despite being the natural approach often proposed in the literature, consistently underperforms, as small matrix-level errors accumulate into larger activation drift. Lastly, reconstruction reduces the relative importance of the pruning criterion: performance gaps between sophisticated criteria and simple baselines shrink with model scale, making simple methods competitive again. Overall, our results challenge the prevailing view that post-pruning adaptation is impractical for LLMs.
AI-Mediated Communication Can Steer Collective Opinion
Stratis Tsirtsis ⋅ Kai Rawal ⋅ Chris Russell ⋅ Brent Mittelstadt ⋅ Sandra Wachter
Generative artificial intelligence (AI) is increasingly integrated into the online platforms where humans exchange opinions; large language models (LLMs) now polish users' posts on LinkedIn and provide context for content shared on X. While prior work has shown that AI can express biased opinions and shape individuals' opinions during human-AI interactions, less attention has been paid to its influence on collective opinion formation when mediating human-to-human communication. We address this gap via a combination of empirical and theoretical analyses. We show empirically that LLMs from multiple popular families introduce directional biases when instructed to edit human-written texts on contested topics, for example, nudging texts in favor of gun control and against atheism. Building on this observation, we introduce a mathematical model of opinion dynamics in which an AI system sits between users on a social network, transforming the opinions they express and perceive. By analytically characterizing the equilibrium of this model and performing simulations on real social network data, we show that biases introduced by AI in human-to-human communication can be amplified through the network and shift collective opinion in their direction. In light of these findings, we investigate whether such biases are controllable by online platforms. We audit the "Explain this post" feature on X and find evidence of pro-life bias in Grok's outputs on abortion-related content, which we trace back to specific design choices.
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
Tomer Keren ⋅ Nitay Calderon ⋅ Asaf Yehudai ⋅ Yotam Perlitz ⋅ Michal Shmueli-Scheuer ⋅ Roi Reichart
As agent capabilities advance, existing benchmarks, such as $\tau^2$-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and labor-intensive. Moreover, the prevailing top-down process, in which scenarios are first written in natural language and then mapped to tool sequences, captures only a narrow subset of the tool-use patterns agents exercise. In this paper, we address these problems by reversing the task construction process. We propose TASTE: Task Synthesis from Tool Sequence Evolution, an automatic method that generates challenging tasks with broader tool-use coverage. TASTE utilizes an Adaptive Contrastive $n$-gram model trained on LLM-judged validity signals. This enables sampling valid tool sequences that cover a vast range of tool combinations. TASTE then selects representative sequences from the pool via clustering, instantiates them into complete benchmark tasks, and refines them through iterative difficulty evolution. Using TASTE, we construct $\tau^c$-Bench, a challenging extension of the three domains of $\tau^2$-Bench. We evaluate $11$ agent/user LLM pairs and find that models nearly saturating $\tau^2$-Bench suffer severe performance drops on our tasks (e.g., Gemini-3-Flash falls from $0.82-0.94$ to $0.28-0.61$). Beyond increasing difficulty, our generated tasks more than double the number of unique tool combinations agents must execute. Our results suggest high scores on existing benchmarks often reflect saturation rather than robust task-solving ability. By automating the generation of difficult, high-coverage benchmarks, TASTE enables continuous, scalable evaluation of future agents.
AM-Bench: A Unified Taxonomy and Evaluation Suite for Agentic Misalignment
Eric Zhang ⋅ Terry J Zhang ⋅ Chijioke Ugwuanyi ⋅ Jerick Shi ⋅ Bernhard Schölkopf ⋅ Zhijing Jin
As LLM-based agents are deployed with increasing autonomy in real-world environments, a critical failure mode emerges: agentic misalignment, where agents choose harmful actions over failure. However, existing safety benchmarks primarily evaluate whether agents comply with explicitly harmful instructions, leaving the conditions that cause agents to choose harmful actions poorly understood. To address these limitations, we propose AM-Bench, a unified taxonomy and evaluation benchmark for agentic misalignment grounded in the fraud triangle, a well-established criminological framework explaining goal-directed misconduct through three dimensions: perceived pressure, perceived opportunity, and rationalization. We operationalize each dimension into structured subcategories rooted in prior theory: five pressure categories derived from the instrumental convergence thesis, seven opportunity categories drawn from the Emergent Strategic Reasoning Risks (ESRR) taxonomy, and a novel four-question binary rationalization framework derived from prior literature. To validate our taxonomy, we construct 35 realistic and diverse evaluation scenarios that span seven industry verticals. Our evaluation of 11 frontier models reveals a 57-percentage-point gap in misalignment, with Gemini 3 Flash and Gemini 3.1 Pro exhibiting the highest rates while Claude Sonnet 4.6 remains near-zero across all conditions. Our rationalization framework further demonstrates that models occasionally misunderstand situations while exhibiting harmful behaviours, highlighting how agentic misalignment can occur without full situational awareness or explanation, complicating both mitigation and detection and making robust evaluation an increasingly pressing concern. All evaluation code are available at https://anonymous.4open.science/r/am-bench-B9E8/.
Amortized Bayesian Experimental Design with In-Context Knowledge Conditioning
Zhanghu Zhao ⋅ Daolang Huang ⋅ Xinyi Wen ⋅ Samuel Kaski
Amortized Bayesian experimental design (BED) enables real-time design strategies by shifting acquisition cost offline, but existing methods remain tied to the prior and task distribution they were trained on and cannot exploit additional information available at deployment. We introduce in-context amortized BED with unified knowledge encoding (IMBUE), an amortized BED framework that accepts two forms of such information without retraining, namely user-specified prior knowledge and observations from previous instances of the experiment. Both are encoded as additional input tokens and processed in-context with the accumulated observations by a shared Transformer, and a learned reliability filter scores each auxiliary token against the accumulated observations and excludes inconsistent ones. Across four standard BED benchmarks, IMBUE accelerates early-stage information acquisition when the auxiliary knowledge is reliable, and remains close to the no-auxiliary baseline when it is not.
Anatomy-Activated Mixture-of-Experts for 3D Medical Vision-Language Pre-training
Szymon Płotka ⋅ Gizem Mert ⋅ Pedro R. A. S. Bassi ⋅ Wenxuan Li ⋅ Zongwei Zhou ⋅ Anna M Marcinkiewicz ⋅ Wiktoria Romańczyk ⋅ Jarosław B Ćwikła ⋅ Łukasz Struski ⋅ Ewa Szczurek ⋅ Jacek Tabor ⋅ Arkadiusz Sitek
Medical Vision-Language Pre-training (VLP) leverages the semantic richness of radiology reports to learn 3D representations, yet current architectures remain fundamentally anatomy-agnostic. While clinical diagnostics rely on organ-specific context, standard encoders apply a uniform, dense parameterization to all volumetric patches, and existing Mixture-of-Experts (MoE) routers gate on isolated local features without explicit access to the organ composition of the scan or crop. We propose Anatomy-Activated Mixture-of-Experts (A$^{2}$MoE), a framework that incorporates anatomical inductive biases within the model's computational path. A$^{2}$MoE introduces two key components: (1) a Histogram Representation Router (HRR) that conditions expert selection on global visual context and a predicted anatomy-presence histogram, ensuring that parameters specialize in specific anatomical domains; and (2) a mask-free Organ-Query Multi-Scale Projector (OQMSP) that utilizes learnable queries to extract organ-level embeddings for fine-grained alignment without requiring manual annotations at inference. By routing at the patch level rather than the token level, A$^{2}$MoE reduces routing complexity by orders of magnitude while satisfying load balancing by construction. Pre-trained on over 60,000 multi-phase CT scans, A$^{2}$MoE achieves state-of-the-art performance across segmentation, classification, and detection benchmarks. Furthermore, we demonstrate that our anatomically-grounded inductive biases support robust generalization to MRI and ultrasound, establishing a scalable foundation for medical representation learning.
An exact information theory of generalization phase transitions in Bayesian diffusion models
Henry Hunt ⋅ Mason Kamb ⋅ Surya Ganguli
How diffusion models circumvent the curse of dimensionality to learn complex distributions over high dimensional spaces from a finite training set, instead of memorizing it, remains a fundamental mystery. To address this, we introduce analytically tractable Bayesian information restricted diffusion (BIRD) models, in which each pixel observes restricted information about noisy data. A BIRD model time-reverses diffusion by inferring which past training sample produced its current restricted observation using the Bayesian posterior. This model class generalizes existing analytical diffusion models that use spatially local information restriction. We show that spatially local BIRD models closely approximate trained diffusion models early in training, across different architectures such as UNets and DiTs. Under minimal assumptions on the data distribution, we identify an information-theoretic phase boundary between memorization and generalization in the joint space of amount of training data, time in the reverse generative process, and amount of information restriction: a BIRD model memorizes when the mutual information between its restricted noisy observations and the training data exceeds the log number of training points, and it generalizes otherwise. Experiments across a range of datasets confirm our theoretically predicted location for the transition. We find that generation proceeds near the edge of memorization: both spatially local BIRD models and early-training diffusion models track the memorization-generalization phase boundary by increasingly restricting information over time. Overall, our results reveal a fundamental role for information restriction in generative AI to circumvent the curse of dimensionality.
A probabilistic model of visual segmentation explains early visual cortical dynamics
Tridib K. Biswas ⋅ Ruben Coen-Cagli
Neural dynamics are a central feature of biological neural systems. In the visual cortex, the functional role of neural dynamics is poorly understood, particularly for static stimuli. The statistics of still natural images have long been used by models that successfully predict firing in the early visual cortex. Those models assume that neurons represent inferences about latent image features, but they often ignore the dynamics of said inference. Here, we test if inferential dynamics for images explain visual-cortical dynamics. To do so, we present a biologically plausible model of cortical computation that assumes the inference of features is coupled with inference about how those features are organized into meaningful parts, also known as segmentation. Given an input image, our model iterates between inferring the features and the segments, through recurrence, thus inducing dynamics in the feature representation. We make theoretical and exact predictions for the induced neural dynamics. Our model predicts both gain and variability in evoked responses to natural images, and we show that it captures classical single-neuron experimental findings such as gain dynamics and variability decay during stimulus presentation. The model also predicts that those metrics are more heterogeneous than previously thought, reflecting the complexity of inference for natural scenes, and that their population-level organization reflects the inferred segments. Finally, our analytical formulation provides a clear interpretation of those results. Insights from our normative model could generalize to segmentation problems in other experimental modalities, and help address the cost of inference in artificial networks.
A Scalable Nonparametric Continuous-Time Survival Model through Numerical Quadrature
Chaeyeon Lee ⋅ Sehwan Kim ⋅ Hyungrok Do
Flexible continuous-time survival modeling is critical for capturing complex temporal dynamics in high-dimensional data; however, training such models remains challenging due to the intractable integral required for likelihood estimation. We introduce QSurv, a scalable deep learning framework that enables nonparametric continuous-time modeling without relying on time discretization or restrictive distributional assumptions. We propose a training objective based on Gauss-Legendre numerical quadrature, which approximates the cumulative hazard with high-order accuracy while facilitating efficient end-to-end training via standard backpropagation. Furthermore, to effectively capture non-stationary dynamics in complex architectures, we introduce time-conditioned low-rank adaptation, a mechanism that conditions general neural backbones on time by dynamically modulating weights via low-rank updates. We provide theoretical analysis establishing approximation error bounds for cumulative-hazard evaluation. Comprehensive experiments across synthetic benchmarks, large-scale real-world tabular datasets, and high-dimensional medical imaging tasks demonstrate that QSurv achieves competitive predictive performance with advantages in instantaneous hazard function estimation, enabling more interpretable characterization of time-varying risk patterns.
Assessing the robustness of heterogeneous treatment effects in survival analysis under informative censoring
Yuxin Wang ⋅ Dennis Frauen ⋅ Jonas Schweisthal ⋅ Maresa Schröder ⋅ Stefan Feuerriegel
Dropout is common in clinical studies, with up to half of patients leaving early due to side effects or other reasons. When dropout is informative (i.e., dependent on survival time), it introduces censoring bias, because of which treatment effect estimates are also biased. In this paper, we propose an assumption-lean framework to assess the robustness of conditional average treatment effect (CATE) estimates in survival analysis when facing censoring bias. Unlike existing works that rely on strong assumptions, such as non-informative censoring, to obtain point estimation, we use partial identification to derive informative bounds on the CATE. Thereby, our framework helps to identify patient subgroups where treatment is effective despite informative censoring. We further propose a novel model-agnostic meta-learner, called SurvB-learner, to estimate the bounds that can be used in combination with arbitrary machine-learning models, and that has favorable theoretical properties such as double-robustness and quasi-oracle efficiency. We finally demonstrate the effectiveness of our meta-learner across various experiments using both simulated and real-world data.
Neuro-symbolic systems typically couple a neural perception module with a discrete symbolic solver, where exact constraint inference is intractable at training time and a discrete solver must be invoked at inference time, precluding end-to-end optimization of the model. We introduce AS2 (Attention-Based Soft Answer Sets), a fully differentiable neuro-symbolic architecture that replaces the discrete solver with a continuous approximation of the Answer Set Programming (ASP)immediate-consequence operator $T_P$. AS2 maintains per-position probability distributions over a finite symbol domain throughout the forward pass and trains end-to-end by minimizing the fixed-point residual of a probabilistic lift of $T_P$, thereby differentiating through the constraint check without invoking an external solver during either training or inference. On spatial constraint-satisfaction tasks, AS2 replaces conventional positional embeddings entirely, and problem structure is encoded through constraint-group membership embeddings derived directly from the declarative ASP specification, making the model agnostic to arbitrary position indexing. On Visual Sudoku, AS2 achieves 99.89\% cell accuracy and 100\%constraint satisfaction, without invoking any external solver. On CLEVR-Hans, AS2 achieves 99.94% on CLEVR-Hans3 and 87.77% on CLEVR-Hans7. These results demonstrate that a soft differentiable fixpoint operator, combined with constraint-aware attention and declarative constraint specification, can match or exceed pipeline and solver-based neuro-symbolic systems while maintaining full end-to-end differentiability.
Attribution-Guided Exit Policy for Reliable Early-Exit Inference
Haseena R P ⋅ Muhammed Ashrah ⋅ Ajith Abraham
Early-exit neural networks improve inference efficiency by enabling predictions at intermediate layers, but their effectiveness critically depends on the exit policy that decides whether to stop or continue computation. Existing policies, often based on confidence thresholds, enable early termination but offer limited insight into the reliability of such decisions. To address this, we introduce an explainability-aware framework that defines an attribution-grounded reliability signal for early-exit decisions. We propose Progressive Feature Attribution Maps (PFAM), which capture how feature importance evolves across exits, and the Interpretability-Based Early-Exit Score (IEES), a unified decision metric that integrates prediction confidence with attribution strength and cross-exit stability. To enable efficient deployment, we further introduce a lightweight proxy predictor that estimates IEES using inexpensive, attribution-free features, allowing attribution-informed decisions without incurring inference-time overhead. Experiments on convolutional architectures (ResNet, MobileNet, MSDNet) show that our approach improves accuracy–efficiency trade-offs, achieving up to 2.5\% accuracy gains while reducing inference time by up to 24\% compared to full-depth inference. Similar trends are observed on a vision transformer and a BERT-based language model, suggesting applicability beyond convolutional architectures.
Auto-Annotation with Expert-Crafted Guidelines: A Study through 3D LiDAR Detection Benchmark
Yechi Ma ⋅ Wei Hua ⋅ Shu Kong
The contemporary paradigm of scaling data annotation, crucial for developing machine learning solutions, is to hire ordinary human annotators and instruct them with expert-crafted guidelines to label data. This paradigm is laborious, tedious, and costly, motivating us to study an open problem, auto-annotation with expert-crafted guidelines (dubbed AutoExpert). We develop benchmarks by redesigning the evaluation protocol and re-annotating data with nuScenes and PandaSet, two 3D detection datasets for autonomous driving research that provide expert-crafted annotation guidelines. Their guidelines define 18 and 25 object classes, respectively, using nuanced language descriptions and a few visual examples. Following the guidelines that require using 3D cuboids to label LiDAR data, AutoExpert requires algorithms to learn on few-shot labeled images and texts to perform the task of 3D detection on LiDAR data. Apparently, the challenges of AutoExpert lie in the data-modality and task discrepancy. Nevertheless, public foundation models (FMs) serve as promising tools to tackle these challenges. To address AutoExpert, we adopt a conceptually simple pipeline consisting of three components: (1) 2D object detection and segmentation in RGB images, (2) lifting 2D detections into 3D using known sensor poses, and (3) 3D cuboids generation for the 2D detections. Within this pipeline, we enhance and evaluate a variety of methods such as open-vocabulary detectors, few-shot detectors, and self-supervised learned detectors. We also develop novel techniques, leading to refined components that boost 3D detection mAP from 12.1 to 25.4 on the AutoExpert-nuScenes benchmark.
Backbone-Equated Diffusion OOD via Sparse Internal Snapshots
Yadang Alexis Rouzoumka ⋅ Jean Pinsolle ⋅ Eugénie TERREAUX ⋅ Christèle Morisseau ⋅ Jean-Philippe Ovarlez ⋅ Chengfang Ren
Fair comparison between diffusion-based OOD detectors is challenging, as conclusions can vary with backbone choice, corruption parameterization, and test-time budget. We address this issue through a *Mutualized Backbone-Equated* (MBE) protocol that aligns canonical corruption levels and logical test-time cost across diffusion backbones. Within this setting, we introduce *Canonical Feature Snapshots* (CFS), a family of detectors that probes a frozen diffusion backbone using only a tiny number of native internal activations at canonical low-noise levels. On a controlled CIFAR-scale benchmark, the strongest one-forward CFS variant is $CFS(1\times2)$, while an even smaller decoder-only variant remains highly competitive. This shows that much of the relative-OOD signal exposed by frozen diffusion backbones is concentrated in a small number of sparse internal states, rather than requiring full denoising trajectories or high-capacity downstream heads. We further provide a local diagnostic theory explaining these observations through conditional encoder-decoder complementarity, diagonal-score separation, and low-noise corruption stability.
Bayesian Low-Rank Posteriors for Scalable Membership Inference
Khanh Nguyen ⋅ Antonio Cinà ⋅ Andrey Barsky ⋅ Alejandra Reinares ⋅ Fabio Roli ⋅ Dimosthenis Karatzas
Membership inference attacks (MIAs) aim to determine whether a sample was used during the training of a target model. The most effective MIAs rely on reference (or shadow) models to estimate target-conditional score distributions, but training such ensembles is computationally prohibitive for modern large-scale Vision--Language Models (VLMs). In this work, we propose a scalable alternative that replaces explicit reference-model training with a Bayesian approximation over low-rank adaptation parameters. Leveraging parameter-efficient fine-tuning, we model uncertainty only within the LoRA subspace and construct a posterior from a single training run, from which we sample a diverse set of virtual reference models. We instantiate this approach using stochastic weight averaging Gaussian (SWAG), enabling efficient approximation of target-conditional score distributions at a fraction of the cost of conventional shadow-model ensembles. We evaluate the resulting attack on three downstream-adapted VLMs across two privacy-sensitive visual question answering tasks, namely Document VQA and medical VQA. To isolate the effect of pretraining knowledge, we additionally introduce a controlled synthetic dataset. Our approach outperforms standard reference-based attacks while requiring only a single trained reference model, demonstrating that accurate and scalable membership inference is feasible even for large VLMs.
Behaviour4All: A Dependency-Aware Toolkit for in-the-wild Facial Behaviour Analysis
Dimitrios Kollias ⋅ Chunchang Shao ⋅ Odysseus Kaloidas ⋅ Ioannis Patras
Facial behaviour analysis requires understanding valence–arousal (VA), expressions (EXP), action units (AUs), yet existing systems treat these tasks independently or use naïve MTL, ignoring their structured dependencies. This leads to negative transfer, conflicting gradients, poor cross-db generalisation, especially with heterogeneous datasets with only partial labels. We propose Behaviour4All, the first dependency-aware toolkit that unifies VA, EXP, AUs via explicit relational priors and principled optimisation framework. Behaviour4All integrates psychophysically grounded task mappings with coordinated loss families to mitigate gradient conflicts and enforce coherent cross-task semantics. Across eight in-the-wild datasets, Behaviour4All achieves sota performance, improves fairness and generalisation, and outperforms vanilla MTL and student--teacher baselines. We will release the complete toolkit (code, pre-trained models, etc) upon acceptance.
Benchmarking Sensor-Fault Robustness in Forecasting
Alexander Windmann ⋅ Philipp Wittenberg ⋅ Gianluca Manca ⋅ Marcel Dix ⋅ Jens U. Brandt ⋅ Oliver Niggemann
Cyber-physical system (CPS) forecasting models depend on sensor streams with noisy, biased, missing, or temporally misaligned readings, yet standard forecasting evaluation often selects models by nominal error without showing whether they remain robust under such faults. We introduce SensorFault-Bench, a shared CPS-grounded sensor-fault stress-test protocol for evaluating forecasting architectures and robustness-improvement methods, together with an operational taxonomy that organizes the method comparison. Across four real-world datasets and eight scored scenarios governed by a standardized severity model, it reports worst-scenario degradation, clean mean squared error (MSE), and worst-scenario fault-time MSE, separating relative robustness from absolute error. A disjoint fault-transfer split lets explicit fault-training methods train on adjacent fault families while reported scores use separate benchmark scenarios. Empirically, forecasting architectures favored by clean MSE can degrade sharply under faults, and clean-MSE rankings can disagree with worst-scenario fault-time error rankings. Chronos-2, the evaluated zero-shot foundation-model representative, matches or trails the last-value naive forecaster in clean MSE on the two single-target datasets and has the largest worst-scenario degradation on ETTh1 and Traffic, where all channels are forecast targets. For the evaluated robustness-improvement method set, paired deltas show selective degradation reductions: projected gradient descent adversarial training and randomized training lead where value faults dominate observed degradation, while fault augmentation leads where availability faults dominate. SensorFault-Bench provides open-source code, documented data access, and reproduction and extension guides, so new datasets, architectures, and robustness-improvement methods can be evaluated under the same CPS sensor-fault robustness protocol.
Bench-MFG: A Benchmark Suite for Learning in Stationary Mean Field Games
Lorenzo Magnino ⋅ Jiacheng Shen ⋅ Matthieu Geist ⋅ Olivier Pietquin ⋅ Mathieu Lauriere
The rise of Mean Field Games (MFGs) has fostered a growing family of algorithms designed to solve large-scale multi-agent systems. However, the field currently lacks a standardized evaluation protocol, forcing researchers to rely on bespoke, isolated, and often simplistic environments. This fragmentation makes it difficult to assess the robustness, generalization, and failure modes of emerging methods. To address this gap, we propose a comprehensive benchmark suite for MFGs, focusing on the discrete-time, discrete-space, stationary settings. We introduce a taxonomy of problem classes, ranging from no-interaction and monotone games to potential and dynamics-coupled games, and provide prototypical environments for each. Furthermore, we present MF-Garnets, a method for generating random MFG instances to facilitate rigorous statistical testing. We benchmark a variety of learning algorithms across these environments, including a novel population-based approach (MF-PSO). Based on our knowledge, this is the first benchmark for MFGs.
BESS-Bench: Benchmarking Spectral Representations for Be-Star Variability
Matthieu Le Lain ⋅ Sébastien Lefèvre
Time-resolved stellar spectroscopy is a primary observational window onto mass-loss, disc formation and rotational physics, phenomena that single-epoch surveys cannot capture. Be stars are rapidly rotating B-type stars that episodically grow and lose a circumstellar gas disc, which imprints a variable emission profile on the hydrogen H-alpha Balmer line (6562.8 Angstrom) on timescales from days to decades. Decoding these spectral variations is both a key step in understanding stellar evolution and an open challenge for machine learning on irregularly sampled scientific time series. Yet no benchmark exists for these spectra: professional surveys deliver a single epoch per star, and the amateur-driven BeSS database, although it hosts hundreds of thousands of multi-epoch spectra, has never been curated for machine learning. We close this gap with BESS-Bench: 339,115 optical spectra of 1,468 Be stars over 35 years, all expert-reviewed, from amateurs (81.7%) and professionals (17.7%), with a scored H-alpha benchmark slice of 26,937 spectra and three tasks under a unified evaluation protocol. SpecProbe measures how well frozen embeddings recover six scalar line features (width, peak separation, line-core depth, asymmetry, equivalent width, peak intensity). LineTransfer tests cross-line generalisation (H-beta to H-alpha). EWForecast predicts the next-epoch equivalent width of H-alpha (a disc-mass proxy) from a star's recent spectroscopic history. We release a compact masked auto-encoder baseline (BeMAE) and benchmark it against PCA and two zero-shot time-series foundation models (Chronos-Bolt and TimesFM-2.0). The ranking is sharply task-dependent: BeMAE recovers shape features with an R2 that is 4.24x that of PCA(10) (bootstrap CI95 [4.12, 4.33]), yet a PCA+ridge pipeline lowers EW forecasting MAE by 5.8% over persistence, beating both foundation models, and calling as such for further development of AI models in this context. Dataset, weights, frozen splits and evaluation pipeline are released under CC-BY-4.0, so the scientific community can adopt BESS-Bench as a shared testbed and contribute new representations and models to advance our understanding of stellar variability.
Best Arm Identification for Bandits with Shifting Means
Lukas Zierahn ⋅ Wouter Koolen ⋅ Shubhada Agrawal ⋅ Christina Katsimerou ⋅ Dirk van der Hoeven
We study the best arm identification problem in a novel non-stationary environment that we coin Shifting Means. While classically the mean rewards of the $K$ arms are stable in time, in Shifting Means only the gaps $\boldsymbol{\Delta}$ between mean rewards are stable, while their common shift may be determined adversarially in each round. The objective of the learner is to identify the best arm with high probability while minimizing sample complexity (the fixed confidence setting). Handling shifts requires new tools: we show that algorithms employing a Generalized Likelihood Ratio Test (GLRT) stopping rule, including the popular Track-and-Stop, fail under time-varying shifts. Instead, we propose Importance Weights for Shifting Means ($\mathsf{ISM}$). Assuming means bounded by $U$ and $\sigma^2$-sub-Gaussian rewards, we show $\mathsf{ISM}$ to be $\delta$-correct and to enjoy a sample complexity bound of $K (\sigma^2 + U^2) \Delta_{\min}^{-2} \ln \frac{1}{\delta}$. We also present a matching (up to constant factors) worst-case lower bound and evaluate our results empirically.
Beyond MMSE: Enhancing PnP Restoration with ProxiMAP
Kenta Vert ⋅ Giacomo Meanti ⋅ Scott Pesme ⋅ Michael Arbel ⋅ Julien Mairal
Plug-and-Play (PnP) methods have become standard tools for solving imaging inverse problems by replacing the intractable maximum a posteriori (MAP) denoiser with the MMSE one. While this mismatch has been widely treated as unavoidable, recent works have sought to close this gap by targeting the MAP with diffusion-model scores. We show this is problematic in practice: learned scores do not match the true ones, so MAP-targeting iterations converge to cartoon-like images rather than realistic ones, and better results are obtained by stopping short of convergence. We turn this observation into a design principle and introduce $\texttt{ProxiMAP}$, an iterative MAP approximation whose noise schedule keeps the iterate's residual noise matched to the denoiser's training noise. This keeps the denoiser in-distribution where its score is reliable, and yields implicit early stopping that avoids the failure mode above. $\texttt{ProxiMAP}$ is a modular drop-in replacement for MMSE denoisers in standard PnP algorithms and consistently sharpens reconstructions across deblurring, inpainting, super-resolution, and phase retrieval. Building on the same principle, we propose a hybrid variant that applies $\texttt{ProxiMAP}$ only in the late iterations of PnP, where the denoiser is most reliable—matching or exceeding the full-replacement variant at a fraction of the cost.
Beyond Pixel-wise Supervision: Local Structure Regularization for Semantic Segmentation
Wei Huang ⋅ Chenying Liu ⋅ Yilei Shi ⋅ Zhitong Xiong ⋅ Xiaoxiang Zhu
Visual foundation models (VFMs) provide strong region-level representations for semantic segmentation, substantially improving high-level semantic understanding. However, accurate boundary recovery remains challenging, especially for fine structures and semantic transitions. A key reason is that conventional pixel-level supervision largely treats pixels independently and lacks explicit modeling of local structure and neighboring-pixel relationships. To address this issue, we propose a Local Structure (LS) regularization framework that constrains the second-order local organization of semantic transition regions. Unlike existing boundary losses that mainly supervise contour location, distance, overlap, or topology, LS models boundary neighborhoods through a structure-tensor representation of semantic probability fields. It decomposes local boundary structure into two complementary cues: Local Structure Magnitude (LSM), which measures the strength of anisotropic class transitions, and Local Structure Orientation (LSO), which describes their dominant geometric orientation. LS is model-agnostic, compatible with different VFM-backbone-decoder combinations, and introduces no extra inference cost. Extensive experiments on four diverse datasets show that LS consistently improves boundary-sensitive performance while also enhancing overall IoU, demonstrating its effectiveness for boundary-structure learning in semantic segmentation.
Beyond Raw Context Transfer: Representation-based Federated Retrieval-Augmented Generation
Can Peng ⋅ Yu Liu ⋅ Yingyu Yang ⋅ Anjie Le ⋅ Yuyuan Liu ⋅ Qianye Yang ⋅ Alison Noble
Retrieval-augmented generation (RAG) improves the factuality of large language models (LLMs) and vision-language models (VLMs) by grounding generation in external knowledge. However, most existing RAG frameworks assume a centralized retrieval document pool, which is often impractical in sensitive domains such as healthcare, where data are inherently distributed and subject to strict privacy constraints. Recent efforts on decentralized RAG primarily follow prompt-based paradigms that exchange raw, human-readable retrieved content, leading to substantial communication overhead and potential privacy risks. To address these limitations, we propose Representation-based Federated RAG (FedRepRAG), a decentralized RAG framework that avoids raw-data exchange during retrieval. FedRepRAG retrieves information from private local document stores and exchanges only compressed latent representations. To integrate the retrieved knowledge, we introduce a lightweight collaboratively trained projector that maps retrieval embeddings into the feature space of a frozen LLM/VLM backbone. This design makes the exchanged information interpretable only within the federated system, thereby enhancing privacy while substantially reducing communication cost. Experiments on decentralized visual question answering (VQA) and question answering (QA) benchmarks show that FedRepRAG consistently outperforms direct inference and local retrieval baselines, while substantially reducing retrieval tokens, computational cost, and inference-time communication overhead compared with raw-context transfer. Overall, FedRepRAG establishes an effective, efficient, and privacy-preserving framework for federated RAG.
Beyond Success Rates: Trainability and Extractability in Offline GCRL
Jan Malte Töpperwien ⋅ Aditya Mohan ⋅ Marius Lindauer
Offline goal-conditioned reinforcement learning (GCRL) is typically benchmarked by the best tuned success rate of each method. This score measures attainable performance, but it does not reveal how reliably a learned goal-conditioned signal can be extracted into a policy: a method could succeed across many value-learning and extraction settings, or only at a narrow, hard-to-find configuration. We study this gap across four methods, GCIQL, GCIVL, QRL, and CRL, under a shared advantage-weighted regression (AWR) extractor. For each method, we construct trainability landscapes over the optimizer learning rate, which affects value learning and actor optimization, and AWR temperature, which controls how selectively the actor imitates high-advantage transitions. Across AntMaze, Cube, and Scene, we observe distinct regimes: high-scoring methods may be broadly accessible or brittle, while broad relative basins may still sit below low absolute ceilings. To interpret these differences, we pair landscapes with post-hoc diagnostics of future-vs-random goal discrimination and AWR weight concentration. % whether the learned advantage function separates matched future goals from random goals, and whether AWR spreads weight across many behavior-supported transitions or concentrates on a small tail. Their relationship to downstream success is task-dependent. On AntMaze, where future goals align with path-like progress, these diagnostics explain landscape regimes. On Cube and Scene, goal ranking and manipulation control decouple: methods can rank goals well while failing downstream, or succeed through action-conditioned advantages despite weak future-vs-random separation. These results show that peak tuned success alone does not establish broadly extractable goal-conditioned behavior. Trainability landscapes expose this gap, while extraction diagnostics offer a lower-cost lens on how learned signals become policies.
Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
Enrico Vompa ⋅ Tanel Tammet ⋅ Mohit Vaishnav
A challenge in advancing Visual-Language Models (VLMs) is determining whether their failures on abstract reasoning tasks, such as Bongard problems, stem from flawed perception or faulty top-down reasoning. To disentangle these factors, we introduce a diagnostic framework centered on the Linear Separability Ceiling (LSC), the performance achievable by a linear classifier on a VLM's raw visual embeddings. Applying this framework to state-of-the-art VLMs, we uncover a pervasive ''alignment gap'', where most models fail to generatively outperform the linear separability of their representations. We find that the few models surpassing this ceiling do so via two mechanisms: by further refining visual representations into a more linearly separable format or by executing non-linear decision logic. We demonstrate that this bottleneck is not a fundamental limitation but a solvable visual alignment issue. Our method augments standard next-token prediction with a contrastive objective to restructure the visual manifold into a more one-dimensionally linear geometry, improving image-to-image comparison and enabling models to significantly surpass the LSC on abstract compositional reasoning tasks.
BLADE: Scalable Bi-level Adaptive Data Selection for LLM Training
Jiaxing Wang ⋅ Deping Xiang ⋅ Jin Xu ⋅ Zirui Liu ⋅ Zicheng Zhang ⋅ Guoqiang Gong ⋅ Jun Fang ⋅ Chao Liu ⋅ Pengzhang Liu ⋅ Tongxuan Liu ⋅ Ke Zhang ⋅ Qixia Jiang
As Large Language Model (LLM) datasets scale to trillions of tokens, data selection has emerged as a critical frontier to filter out uninformative noise and construct adaptive learning trajectories. Beyond static heuristic filtering, advanced data selection methods for LLM training largely follow two paradigms, each with fundamental limitations. Influence-based methods provide principled bi-level objectives but require intractable inverse-Hessian computations, while excess-loss methods are computationally efficient but rely on a static reference model that becomes misaligned with the evolving proxy model during training. We propose BLADE (Bi-Level Adaptive Data sElection), a Hessian-free framework for data selection. BLADE reformulates the bi-level optimization problem underlying influence-based methods as a penalized single-level objective via Lagrange multipliers, avoiding inverse-Hessian computation while revealing a principled connection to excess-loss based data selection. The resulting objective recovers an excess-loss form but replaces the static reference model with a dynamic one that stays synchronized with training. Theoretically, we prove that this penalized formulation guarantees first-order convergence. For efficient online batch selection, we instantiate BLADE as a memoryless randomized block-coordinate Frank-Wolfe algorithm. Extensive experiments show that BLADE consistently outperforms state-of-the-art data selection baselines, providing a practical recipe for LLM training.
Breaking BAD: Heterogeneous Byzantine-Robust Federated Learning via Gather and Scatter Scores
Sena Ergisi ⋅ Luis Massny ⋅ Rawad Bitar
Byzantine-robust Federated Learning (FL) designs aggregation rules that mitigate the impact of malicious clients. When the data across honest clients are heterogeneous, honest model updates are inherently scattered. This dispersion can be exploited by Byzantine adversaries who may collude to concentrate their updates, thereby circumventing state-of-the-art robust aggregation methods and causing a Byzantine Adversarial Disruption (BAD) of the training process. We propose a novel Byzantine-robust FL scheme based on gather and scatter scores. Our scoring technique captures concentration among Byzantine clients and prevents catastrophic failure under SotA attacks. Through numerical experiments on FEMNIST, Shakespeare, and CIFAR-10 datasets, we demonstrate that the proposed defense maintains high model accuracy, even in realistic FL scenarios. Furthermore, we establish theoretical convergence guarantees for our method.
Can LLMs Take Retrieved Information with a Grain of Salt?
Behzad Shayegh ⋅ Mohamed Osama Ahmed ⋅ Fred Tung ⋅ Leo Feng
Large language models have demonstrated impressive retrieval-augmented capabilities. However, a crucial area remains underexplored: their ability to appropriately adapt responses to the certainty of the retrieved information. It is a limitation with real consequences in high-stakes domains like medicine and finance. We evaluate eight LLMs on their context-certainty obedience, measuring how well they adjust responses to match expressed context certainty. Our analysis reveals systematic limitations: LLMs struggle to recall prior knowledge after observing an uncertain context, misinterpret expressed certainties, and overtrust complex contexts. To address these, we propose an interaction strategy combining prior reminders, certainty recalibration, and context simplification. This approach reduces obedience errors by 25% on average, without modifying model weights, demonstrating the efficacy of interaction design in enhancing LLM reliability. Our contributions include a principled evaluation metric, empirical insights into LLMs' uncertainty handling, and a portable strategy to improve context-certainty obedience across diverse LLMs.
Can Metadata Fix the Gauge? Calibration Turns Sparse Multi-Source Learning from Sparse PCA into Sparse Mean Recovery
Yibo Zhou ⋅ Yirong Xiang
Source labels tell us where a sample came from, but not how that source oriented its coordinates. In a sparse Gaussian multi-source model with source-specific sign gauges, that distinction changes the learning problem. Without calibration, pooling across sources erases the mean signal, and the source-labelled experiment reduces to sparse PCA. With task-independent relative-gauge metadata, synchronization aligns sources and the same observations become a sparse mean problem. We formalize the value of metadata through aligned mass, which quantifies how much first-order signal a metadata-induced alignment restores. Source labels create only limited aligned mass, whereas valid calibration can create aligned mass proportional to the full sample size. This yields a statistical separation in micro-source regimes and, under standard sparse-PCA hardness assumptions, a computational separation as well. Experiments confirm the predicted phase transition, the failure of invalid metadata, reuse of one calibration graph across many tasks, and controlled real-feature settings with imposed source gauges.
Can the Parts Fool the Test? Counterfactual Pairing Cycles for Relational OOD
Yibo Zhou ⋅ Yirong Xiang
Relational out-of-distribution errors are easy to describe and hard to evaluate cleanly. An image may be natural, a caption may be fluent, and all mentioned objects may be common, yet the caption may describe the wrong relation in the image. Many evaluations intended to test this behavior can be solved without checking the pairing at all, because negatives differ in caption style, template, length, object frequency, or image statistics. We propose counterfactual pairing cycles, a blocked evaluation primitive that reuses the same two images and the same two captions on both sides of the comparison, changing only which image is paired with which caption. This makes image-only, text-only, and additive component shortcuts cancel exactly in finite samples. The same contrast also gives a training loss for relational outlier exposure and a diagnostic for bilinear vision-language scores. Across synthetic controls, GQA, and COCO, cycles separate coarse image-caption matching from stricter role-binding failures. In a shortcut-contaminated COCO outlier-exposure stress test, a model selected by a perfect shortcut validation AUROC transfers poorly to exact relational cycles, while cycle training improves CycleAcc from (0.576) to (0.911). The main lesson is simple: to test whether a model understands a pairing, the benchmark should keep the parts fixed, change the pairing, and ask only about coupling.
CAREBench: Evaluating LLMs' Emotion Understanding by Assessing Cognitive Appraisal Reasoning
ZHAOYUE SUN ⋅ Hainiu Xu ⋅ Andero Uusberg ⋅ James Gross ⋅ Petr Slovak ⋅ Yulan He
Emotion understanding is a core capability for LLMs to interact effectively with humans, yet existing evaluation paradigms rely on discrete emotion label prediction and fail to capture the cognitive processes underlying emotion generation. Grounded in appraisal theory, we introduce CAREBench, the first benchmark with complete inferential chain annotations from both first- and third-person perspectives on real-world narratives, spanning appraisal reasoning, appraisal ratings, and multi-label emotion annotation. We propose a process-level evaluation framework and conduct systematic experiments across six LLMs organized around four research questions. We find that stronger models match or surpass human observers on certain tasks, yet fall short on appraisal reasoning and positive emotion recognition; performance across chain steps and sensitivity to appraisal interventions exhibit dissociations across models; and current models have not internalized the mechanisms needed to capture human subjective heterogeneity. These findings suggest that downstream emotion prediction metrics may overestimate LLMs' true emotion understanding, and CAREBench provides a foundation for more diagnostically informative evaluation of LLMs' affective cognitive capabilities.
Casper: A Projection-Based Neurosymbolic Layer for Scalable & Guaranteed Constraint Satisfaction
Lohith Konathala ⋅ Luca Andolfi ⋅ Eleonora Giunchiglia
Many real-world problems demand exact constraint satisfaction to be meaningfully solved, yet existing neurosymbolic methods face a stark trade-off: those that guarantee satisfaction fail to scale, while those that scale must resort to approximation. Meanwhile, traditional deep learning approaches offer no satisfaction guarantees at all. In this paper, we propose Casper, a neurosymbolic layer that simultaneously (i) guarantees constraint satisfaction, (ii) scales at inference time, (iii) is fully automated, requiring no problem-specific engineering, and (iv) is fully differentiable, making it composable with any neural predictor both at training and at inference time. Casper casts constraint satisfaction as a closed-form Euclidean projection onto the constraint-feasible region, computable in a single forward pass. Experiments across MNIST arithmetic (up to 1024 digits), Sudoku solving, and Warcraft pathfinding show that Casper is the only method that produces guaranteed-valid outputs at scale: on 30×30 Warcraft grids, exact baselines time out and approximate ones produce valid paths less than 7% of the time, while Casper produces them 100% of the time.
Characterizing the Edge of Stability in Variational Training Without Priors
Juraj Marusic ⋅ Jonathan Wenger ⋅ Beau Coker ⋅ John Cunningham
Variational approaches in deep learning promise to deliver improved uncertainty quantification and out-of-distribution generalization by inferring a distribution over weights, regularized by a divergence to a chosen prior. However, prior elicitation presents a significant challenge in deep learning, resulting in the common practice of significantly downweighting the regularization term in the variational objective. Recently, it was proposed to train solely via the expected loss by relying on implicit regularization from the optimizer, rather than explicit regularization from the divergence to the prior. While this approach largely sidesteps the issue of prior elicitation, current theory lacks an understanding of how key optimization hyperparameters impact the implicit regularization. Here, we theoretically and empirically demonstrate that larger learning rates and fewer parameter samples lead to flatter minima of the loss landscape, which are known to generalize better. We show that training on the expected loss operates at a modified edge of stability, characterized by the signal-to-noise ratio of the gradient estimates and the number of Monte Carlo samples. Experiments on out-of-distribution benchmark datasets confirm the practical significance of the theoretical results.
Claims of AI emergence should be grounded in information decomposition
Christoph Riedl ⋅ Fernando Rosas
When a language model develops capabilities not traceable to its individual components, when a multi-agent AI system outperforms any agent alone, or when a human–AI pair succeeds where either fails in isolation, AI researchers describe it as ``synergy''. However, these are usually informal assessments not supported by a theoretical framework, making these claims impossible to compare or even meaningfully dispute. This position paper argues that information decomposition can serve as a unifying formal framework to operationalize this intuition. Information decomposition provides a way to distinguish qualitatively different \emph{types} of information, including unique information available on individual components, redundant information shared across components, and synergistic information accessible only from all parts together. We present how this decomposition can be naturally used to study ensembles of attention heads, autonomous agents, or a human and an AI. We demonstrate the utility of this lens across different scales and substrates, show how it unifies existing knowledge across fields, and leads to cross-level conjectures that none of these fields would generate independently.
Closed-Form Last Layer Optimization
Alexandre Galashov ⋅ Nathaël Da Costa ⋅ Liyuan Xu ⋅ Philipp Hennig ⋅ Arthur Gretton
Neural networks are typically optimized with variants of stochastic gradient descent. Under a squared loss, however, the optimal solution to the linear last layer weights is known in closed-form. We propose to leverage this during optimization, treating the last layer as a function of the backbone parameters, and optimizing solely for these parameters. We show this is equivalent to alternating between gradient descent steps on the backbone and closed-form updates on the last layer. We adapt the method for the setting of stochastic gradient descent, by trading off the loss on the current batch against the accumulated information from previous batches. We provide theoretical analyses showing convergence of the method to an optimal solution in the neural tangent kernel regime, as well as quantifying the gains compared to standard SGD in a one-step analysis. Finally, we demonstrate the effectiveness of our approach compared with SGD and Adam on a squared loss in several regression tasks, including neural operators and causal inference.
Closing the Indexing-Decoding Gap in Multimodal Generative Retrieval via Prefix Retention Optimization
Yufei Chen ⋅ Zihan Wang ⋅ Yubao Tang ⋅ Yukun Zhao ⋅ Maarten Rijke ⋅ Zhaochun Ren
Multimodal generative retrieval formulates multimodal retrieval as discrete identifier generation, eliminating the need for explicit similarity search over external embeddings. Existing approaches construct identifiers via residual quantization and decode them with trie-constrained beam search. This combination introduces an indexing–decoding gap: identifier learning objectives, including reconstruction and contrastive losses, do not explicitly enforce prefix discriminability during decoding. As a result, even well-optimized identifiers can be irreversibly pruned early in beam search due to low-rank prefixes. We theoretically characterize this gap and derive a survival bound that relates prefix retention to three controllable factors in indexing and decoding. Building on this bound, we propose PRO, prefix retention optimization, a unified framework comprising three mechanisms: (i) prefix ranking distillation aligns quantized prefix rankings with those induced by pre-quantization embeddings using a listwise loss; (ii) vocabulary scheduling increases codebook sizes from shallow to deep residual quantization levels to reduce early competition from non-target prefixes; and (iii) geometric score fusion vectorizes each candidate prefix and incorporates its similarity to the query into beam search scoring, further reducing the indexing–decoding mismatch. Experiments on nine multimodal retrieval tasks show that PRO improves retention of target identifier prefixes and outperforms existing multimodal generative retrieval baselines.
Coherent Hierarchical Multi-Label Learning to Defer for Medical Imaging
Joshua Strong ⋅ Pramit Saha ⋅ Emma Sun ⋅ Helen Higham ⋅ Alison Noble
Learning to Defer (L2D) enables a model to predict autonomously or defer to an expert, but prior work largely assumes flat label spaces. We study the first L2D setting with hierarchical multi-label decisions, motivated by medical-imaging workflows in which findings are organised by clinical taxonomies. In this setting, deferral is a delegation action rather than a label assignment, so treating it as an independent per-label decision can produce \emph{deferral incoherence}, including taxonomic contradictions, delegation violations, and deferrals of labels already implied by the model's own assertions. We formalise coherent hierarchical deferral under a Selective-Exclusion handoff contract, characterise the Bayes-optimal coherent deferral rule, and show that even nodewise Bayes L2D can be action-incoherent. We then propose two remedies: exact coherent projection, a dynamic-programming decoder over the coherent action set, and Taxonomic Belief Propagation (TBP) with Recursive Policy Optimisation (RPO), a contract-aware joint action model trained through the same recursion used at inference. Across real-reader and controlled-expert medical-imaging benchmarks, naïve binary-relevance L2D exhibits non-trivial incoherence. Projection removes it exactly, and fast TBP+RPO drives incoherence near zero while retaining strong utility.
CoLaX: Context-Aware Local Explanations for Time Series Classification
Kexuan Zhang ⋅ Paolo Frazzetto ⋅ Xiaobei Zou ⋅ Gustau Camps-Valls ⋅ Yang Tang
Reliable deployment of time series classifiers in high-stakes domains requires explanations that capture the interplay between localized patterns and global temporal context. However, existing post-hoc methods often treat local features in isolation, failing to account for long-range dependencies. We propose CoLaX, a context-aware explanation framework that explicitly embeds local dynamics within a global temporal structure. By employing a principled context-conditioning mechanism, CoLaX models how global context modulates the importance of local patterns, revealing not only what is important but also when and why it matters. Extensive benchmarks demonstrate that CoLaX yields significantly more faithful and stable explanations than state-of-the-art baselines, demonstrating that joint local-global modeling is essential for reliable time-series interpretability.
Comparing Explanations is not Enough,Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models
Martino Ciaperoni ⋅ Marzio Di Vece ⋅ Roberto Pellungrini ⋅ Luca Pappalardo ⋅ Fosca Giannotti ⋅ Francesco Giannini
Large-scale foundation models exhibit *behavioral shifts* when subjected to interventions such as scaling, fine-tuning, reinforcement learning with human feedback, or in-context learning. Current explainability methods are structurally ill-suited to explain these shifts, because either they treat models as static objects, as traditional eXplainable AI ($XAI$) approaches do, or merely compare independent explanations across different checkpoints of a model. As a result, these approaches fail to explain the functional transition between two model instances in which a certain behavior has shifted following an intervention. This gap creates significant governance risks across jurisdictions including the EU AI Act, US state legislation, and Chinese AI regulations, which require documenting causal chains for substantial system modifications. This position paper argues that explaining behavioral shifts in large language models requires a principled approach that treats the shift itself as the primary object of explanation: namely, one that explains how and why an intervention transforms a reference model into an updated model with different behavior. To support this claim, we introduce *Comparative* $XAI$ ($XAI_{\Delta}$), a novel $XAI$ paradigm aimed at explaining the difference between two model checkpoints where a behavior has shifted, together with a set of desiderata specifying what $XAI_{\Delta}$ explainers and explanations must satisfy, including comparability, validity, actionability, and monitoring, with the goal of grounding model auditing in explicit, measurable requirements. Finally, we provide preliminary evidence suggesting the need for $XAI_{\Delta}$ in practice through some illustrative experiments, compiling the resulting findings into a transition report directly usable for governance and incident documentation.
CompleteRXN: Toward Completing Open Chemical Reaction Databases
Gabriel Vogel ⋅ Minouk Noordsij ⋅ Evgeny A Pidko ⋅ Jana M. Weber
Chemical reaction datasets such as USPTO suffer from substantial incompleteness, frequently missing byproducts, co-reactants, and stoichiometric coefficients. This limits their applicability and reliability in downstream applications. Here, we introduce CompleteRXN, a large-scale supervised benchmark for reaction completion under realistic missing-data conditions. We construct a dataset of aligned incomplete and atom-balanced reactions by mapping USPTO records to curated mechanistic reactions. We evaluate representative baselines, including a novel encoder-decoder reaction completion model with constrained decoding, the Constrained Reaction Balancer (CRB), and a recent algorithmic method, SynRBL. On our CompleteRXN benchmark, the CRB achieves high performance across splits of increasing difficulty, reaching 99.20% equivalence accuracy on the random split and 91.12% on the extreme out-of-distribution split. SynRBL produces many balanced and chemically plausible completions, but with lower accuracy on the benchmark test splits. Across all methods, performance degrades with increasing incompleteness. We observe a substantial drop when evaluating on reactions outside the benchmark (full uncurated USPTO), highlighting the gap between benchmark performance and practical robustness and motivating future work.
Concise and Logically Consistent Conformal Sets for Neuro-Symbolic Concept-Based Models
Samuele Bortolotti ⋅ Emanuele Marconato ⋅ Andrea Pugnana ⋅ Andrea Passerini ⋅ Stefano Teso
Neuro-Symbolic Concept-based Models (NeSy-CBMs) are a family of architectures that integrate neural networks with symbolic reasoning for enhanced reliability in high-stakes applications. They work by first extracting high-level concepts from the input and then inferring a task label from these compatibly with given logical constraints. Yet, their label and concept predictions can be overconfident, making it difficult for stakeholders to gauge when the model's decisions can be trusted. We address this issue by integrating ideas from Conformal Prediction (CP), a framework providing rigorous, distribution-free coverage guarantees. We formalize three desiderata -- consistency, coverage, and conciseness -- that any conformal method for NeSy-CBMs should satisfy, and show that existing approaches fall short of at least one. We then introduce COCOCO, a post-hoc framework that conformalizes concepts and labels jointly and reconciles them via a single deduction–abduction revision step. COCOCO satisfies all three desiderata, retains distribution-free coverage, is robust to imperfect knowledge and supports user-specified size budgets. Our experiments on 8 data sets highlight how COCOCO compares favorably against competitors and natural baselines in terms of performance and set size.
Confidence-Based Decoding is Provably Efficient for Diffusion Language Models
Changxiao Cai ⋅ Gen Li
Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) models for language modeling, allowing flexible generation order and parallel generation of multiple tokens. However, this flexibility introduces a challenge absent in AR models: the \emph{decoding strategy}---which determines the order and number of tokens generated at each iteration---critically affects sampling efficiency. Among decoding strategies explored in practice, confidence-based methods, which adaptively select which and how many tokens to unmask based on prediction confidence, have shown strong empirical performance. Despite this success, our theoretical understanding of confidence-based decoding remains limited. In this work, we develop the first theoretical analysis framework for confidence-based decoding in DLMs. We focus on an entropy sum-based strategy that continues unmasking tokens within each iteration until the cumulative entropy exceeds a threshold, and show that it achieves $\varepsilon$-accurate sampling in KL divergence with an expected number of iterations $\widetilde O(H(X_0)/\varepsilon)$, where $H(X_0)$ denotes the entropy of the target data distribution. Notably, this strategy yields substantial sampling acceleration when the data distribution has low entropy relative to the sequence length, while automatically adapting to the intrinsic complexity of data without requiring prior knowledge or hyperparameter tuning. Overall, our results provide a theoretical foundation for confidence-based decoding and may inform the design of more efficient decoding strategies for DLMs.
Constraint Retrieval Is Not Constraint Enforcement in Large Language Models
Yongkang Yang ⋅ Xiankun Lin ⋅ Lixin Liu ⋅ Ziyan Liu ⋅ Xilin Xia ⋅ Zhiwei Zhuang ⋅ Zhezheng Hao ⋅ Wence Ji
Prompt-level instructions are the main way users give language models preferences, safety boundaries, and content rules. But a model can remember a constraint and still violate it when generating text. This points to a gap between two abilities that are often treated as one: retrieving a constraint from context, and enforcing it during token selection. Across five model families (0.5B--32B parameters), we find a clear behavioral split: constraints that admit an alternative generation plan are largely followed, while token-level prohibitions fail whenever the prohibited continuation is locally dominant. Mechanistic analyses show that retrieval heads still attend to the constraint at the violation step, but the resulting signal either paradoxically enhances the prohibited token's logit or suppresses it too weakly to change the argmax. Decoding-time interventions can eliminate surface-form violations, but semantic compliance still requires higher-level verification. These results establish that constraint retrieval and constraint enforcement are separable capabilities, and that reliable token-level compliance requires explicit veto mechanisms beyond prompting.
CoRDS: Coreset-Based Representative and Diverse Selection for Streaming Video Understanding
Ailar Mahdizadeh ⋅ Puria Azadi Moghadam ⋅ Muchen Li ⋅ Xiangteng He ⋅ Leonid Sigal
Streaming video understanding with large vision-language models (VLMs) requires a compact memory that can support future reasoning over an ever-growing visual history. A common solution is to compress the key-value (KV) cache, but existing streaming methods typically rely on local token-wise heuristics, such as recency, temporal redundancy, or saliency, which do not explicitly optimize whether the retained cache is representative of the accumulated history. We propose to view KV-cache compression as a coreset selection problem: rather than scoring tokens independently for retention, we select a small subset that covers the geometry of the accumulated visual cache. Our method operates in a joint KV representation and introduces a bicriteria objective that balances coverage in key and value spaces, preserving both retrieval structure and output-relevant information. To encourage a more diverse retained subset, we further introduce an orthogonality-driven diversity criterion that favors candidates contributing new directions beyond the current selection, and connect this criterion to log-determinant subset selection. Across four open-source VLMs and five long-video and streaming-video benchmarks, our method improves over heuristic streaming compression baselines under a fixed cache budget. These results highlight that representative coreset selection offers a more effective principle than token-wise pruning for memory-constrained streaming video understanding.
Coresets for Clustering Using Noisy Comparisons and Few Distance Queries
Amir Carmel ⋅ Robert Krauthgamer
Coresets are a fundamental tool for scaling clustering algorithms, but standard constructions assume full access to the pairwise distances. In many modern settings, access to distances is costly and thus limited, and algorithms must instead rely on inexpensive comparison queries. We study coreset construction in the Rank-Measure (RM) model, where the algorithm can make noisy comparison queries to a quadruplet oracle alongside a small number of distance queries. In this model, we present the first algorithm to construct an $\epsilon$-coreset for $(k,z)$-clustering in Euclidean space. It matches the coreset-size bounds known in the classical model, while using only polylogarithmically many distance queries in the input size. Our approach is based on a uniform-sampling framework [Chen, SICOMP 2009], which decomposes the input into geometric rings and samples uniformly within each ring. We implement this approach using primarily noisy comparison queries, and then compress the result using a standard full-access coreset construction. In contrast, prior work in oracle-based clustering achieved only constant-factor guarantees. We complement our theoretical findings with experiments that demonstrate strong empirical performance against full-access baselines while using significantly fewer distance computations.
CorridorLight: Cooperation as Task Negotiation with Causal Gating for Traffic Signal Control
Pak Lon Ip ⋅ Pengfei Ren ⋅ Yuteng Lin ⋅ Rongqin Chen ⋅ Tsz Nam Chan ⋅ Leong Hou U
Network-level traffic signal control (TSC) involves coordination under partial observability and strict real-time constraints. In real-world urban networks with dozens or hundreds of intersections, direct coordination incurs an intractably large joint action space, making centralized cooperation difficult to deploy in practice. We propose CorridorLight, which tackles this challenge by casting cooperation as task negotiation. A high-level Corridor Agent represents the network as a direction-level graph and periodically generates a small set of junction-disjoint corridor tasks via a seed-and-grow procedure. These tasks provide lane-level preference guidance, while Intersection Agents maintain fast decentralized control and selectively support the active corridor task. By learning when to cooperate versus act selfishly, CorridorLight enables an adaptive trade-off between local delay minimization and corridor-level relief, leading to improved global performance. To this end, we derive a closed-form, Lipschitz-continuous gating function by maximizing a counterfactual net gain between cooperative and selfish behavior, yielding smooth, interpretable cooperation without brittle heuristics. Experiments on synthetic grids and real-world networks show consistent system-level improvements and outperform cooperative TSC baselines across diverse topologies and demand patterns.
CRISP: Fixing Flying Pixels in Latent LiDAR Generation via Diffusion Decoding
Andrea Ceron ⋅ Michael Schmidt ⋅ Alvaro Marcos-Ramiro ⋅ Sebastian Schmidt ⋅ Benjamin Busam
Latent LiDAR pipelines suffer from flying pixels: convolutional VAEs blur sharp LiDAR contours, yielding edge depths that back-project to points floating between surfaces. We identify this as a major, directly correctable decoder bottleneck and introduce CRISP: a pixel-space diffusion decoder with a backbone-agnostic latent adapter, DiT-based denoiser, and support mask predictor. CRISP replaces convolutional decoders commonly inherited from image and video VAEs while keeping the encoder and latent generator fixed. Across KITTI-360, SemanticKITTI, and nuScenes, replacing only the decoder reduces FSVD/FPVD by 50.5% on average across frozen backbones; for generic video VAEs, the reductions reach 71%/74%. On the LiDAR-native LiDM backbone, FRID drops by 71%, with the largest gains at depth discontinuities. Inside a pretrained LiDM world model, the same zero-shot replacement improves FSVD by 15.5%, narrowing the sim-to-real gap. Code and checkpoints will be released upon acceptance.
CryoGeo: Latent Pose Equivariance for Amortized Ab Initio Cryo-EM Reconstruction
Zhongyang Li ⋅ Diedong Feng ⋅ Shen Cheng ⋅ Bing Zeng ⋅ Shuaicheng Liu
Cryo-electron microscopy (cryo-EM) reconstructs 3D macromolecular structures from noisy 2D particle images by jointly estimating the density and projection pose of each particle. Amortized inference can greatly accelerate this process, but remains vulnerable to pose collapse because known image-space symmetries are often learned only implicitly through reconstruction losses. We ask whether this instability can be mitigated by enforcing cryo-EM transformation laws directly in latent projection-pose space. We propose CryoGeo, a geometry-constrained framework for amortized ab initio cryo-EM reconstruction. CryoGeo enforces latent pose equivariance by requiring in-plane transformations to induce valid group actions on the predicted projection pose: rotations act as right multiplications on the predicted $SO(3)$ orientation, while shifts transform frame-consistently. At the core of CryoGeo is an output-level group-action objective that forces predicted cryo-EM projection poses to obey physical transformation laws induced by in-plane image rotations and shifts, rather than relying on equivariant feature extraction alone. Across synthetic and real cryo-EM datasets, CryoGeo improves pose consistency, mitigates collapse-prone behavior, and yields more reliable amortized ab initio reconstructions than unconstrained variants and neural baselines. These results show that latent pose equivariance turns in-plane symmetry from an implicitly learned nuisance factor into a self-supervised geometric constraint for stable, efficient amortized structural reconstruction. Code will be released.
DART: Domain-Agnostic Residual Transfer for Generalist Anomaly Detection
Muhammad Aqeel ⋅ Maham Nazir ⋅ Marco Cristani ⋅ Francesco Setti
Generalist Anomaly Detection (GAD) aims to detect anomalies in unseen domains using only a few reference samples. Recent normal-abnormal guided methods improve this setting by exploiting abnormal references, but their residuals often mix transferable anomaly cues with domain-specific appearance factors such as texture, geometry, and acquisition conditions. This residual domain entanglement limits cross-domain transfer and may produce spurious activations on novel domains. We propose DART, a plug-in framework for normal-abnormal guided GAD that learns to separate residuals into anomaly-relevant and domain-specific components. DART combines a Residual Disentanglement Module, which preserves anomaly-discriminative information while suppressing domain information, with Cross-Domain Meta-Training, which forms episodes across source domains to make the separation effective. The method is architecture-agnostic and can be inserted between residual computation and the detection head of existing GAD pipelines. Experiments on MVTec AD, VisA, and BraTS show consistent improvements over multiple base methods, with the largest gains under stronger domain shifts. Code will be available upon publication.
DashAttention: Differentiable and Adaptive Sparse Hierarchical Attention
Yuxiang Huang ⋅ Nuno Gonçalves ⋅ Federico Alvetreti ⋅ Lei Li ⋅ Xu Han ⋅ Edoardo Ponti ⋅ André Martins ⋅ Marcos Treviso
Current hierarchical attention methods, such as NSA and InfLLMv2, select the top-$k$ relevant key-value (KV) blocks based on coarse attention scores and subsequently apply fine-grained softmax attention on the selected tokens. However, the top-$k$ operation assumes the number of relevant tokens for any query is fixed and it precludes the gradient flow between the sparse and dense stages. In this work, we propose DashAttention (Differentiable and Adaptive Sparse Hierarchical Attention), which leverages the adaptively sparse $\alpha$-entmax transformation to select a variable number of blocks according to the current query in the first stage. This in turn provides a prior for the second-stage softmax attention, keeping the entire hierarchy fully differentiable. Contrary to other hierarchical attention methods, we show that DashAttention is non-dispersive, translating to better long-context modeling ability. Experiments with large language models (LLMs) show that DashAttention achieves comparable accuracy as full attention with 75% sparsity and a better Pareto frontier than NSA and InfLLMv2, especially in high-sparsity regimes. We also provide an efficient, GPU-aware implementation of DashAttention in Triton, which achieves a speedup of up to $3.3\times$ over FlashAttention-3 at inference time. Overall, DashAttention offers a cost-effective strategy to model long contexts.
Dense Cross-Tokenizer Distillation via Semantic Optimal Local Alignment
Mengyu Zheng ⋅ Zheyuan Bai ⋅ Yuchuan Tian ⋅ Chong Zhu ⋅ Tianyu Guo ⋅ Hanting Chen ⋅ Yunhe Wang
Cross-tokenizer knowledge distillation lacks a shared coordinatesystem between teacher and student vocabularies, and existingmethods each drop one of the two ingredients comparison needs:shape-based objectives preserve density but discard token identity,lexical alignment restores identity but uses surface form as a proxyfor meaning, and text-space matching compresses the teacher signalinto chunk-level events. We introduce SOLA (Semantic Optimal LocalAlignment), which reconstructs the missing coordinate systemexplicitly: it partitions inputs into decoded-text blocks viaMinimal Complete Correspondence, estimates a cross-vocabularycorrespondence cost from calibration text, and runs entropicoptimal transport on a bidirectional local support retrieved aroundeach teacher prediction. Under matched training and decodingprotocols, SOLA outperforms direct cross-tokenizer baselines onDolly instruction following, an UltraChat general-instructionbenchmark, and code and math domain transfer; a coverage--scorecorrelation and component ablations attribute the gain to theconstructed support.
Derivative-Informed Training of Neural Operators On-the-Fly via Sketched Tangent Consistency
Shancong (Sean) Mou ⋅ Xinhan Yang ⋅ Lu Lu
Derivative-informed training improves neural operators by directly supervising their input--output sensitivities, but existing methods rely on offline tangent solves and stored derivative labels. We propose sketched tangent consistency loss (sTCL), an on-the-fly drop-in derivative regularizer for neural-operator training that improves learned Jacobians without changing the architecture or generating offline sensitivity data. At each training step, sTCL samples a small number of perturbation directions in the input function space, computes the corresponding surrogate Jacobian--vector products using forward-mode automatic differentiation, and penalizes the residual of the forward sensitivity equation. This equation is obtained by differentiating the governing PDE with respect to the input perturbation direction, so derivative supervision is obtained directly from the physics rather than from stored tangent labels. We further show that sketching makes online derivative-informed training computationally feasible, while conditioning is the key to making it effective. Random directional sketching keeps the derivative penalty at the same order of cost as the standard data-fitting loss, making sTCL compatible with stochastic neural-network training as a lightweight loss term. However, raw forward-sensitivity residual penalties can fail for stiff, ill-conditioned, or indefinite tangent operators. To address this, we introduce lightweight operator-aware preconditioners selected by a simple tangent-operator decision rule. Across four PDE benchmarks---Helmholtz, nonlinear diffusion--reaction, Burgers, and two-dimensional Navier--Stokes---operator-conditioned sTCL achieves accuracy comparable to offline DIFNO while eliminating the offline derivative-data generation stage. These results show that on-the-fly derivative-informed training need not merely amortize offline tangent-solve cost into training; with appropriate sketching and conditioning, sTCL provides an attractive drop-in path to derivative-informed neural operators.
Detecting RLVR Training Data via Structural Convergence of Reasoning
Hongbo Zhang ⋅ Leyang Cui ⋅ Jianhao Yan ⋅ Guangsheng Bao ⋅ Yue Zhang ⋅ Shuo Yang ⋅ Yue Zhang
Reinforcement learning with verifiable rewards (RLVR) is central to training modern reasoning models, but the undisclosed training data raises concerns about benchmark contamination. Unlike pretraining methods, which optimize models using token-level probabilities, RLVR fine-tunes models based on reward feedback from self-generated reasoning trajectories, making conventional likelihood-based detection methods less effective. We show that RLVR induces a distinctive behavioral signature: prompts encountered during RLVR training result in more rigid and similar generations, while unseen prompts retain greater diversity. We introduce Min-$k$NN Distance, a simple black-box detector that quantifies this collapse by sampling multiple completions for a given prompt and computing the average of the $k$ smallest nearest-neighbor edit distances. Min-$k$NN Distance requires no access to the reference model or token probabilities. Experiments across multiple RLVR-trained reasoning models show that Min-$k$NN Distance reliably distinguishes RL-seen examples from unseen ones and outperforms existing membership inference and RL contamination detection baselines.
DiDE: Direct Injection with Color-Texture DEcoupling for 3D Stylization
Tao Wu ⋅ Alexandra Gomez-Villa ⋅ Senmao Li ⋅ Yaxing Wang ⋅ Joost van de Weijer ⋅ Kai Wang
Recent advances in rectified flow-based image-to-3D generative models have enabled high-fidelity 3D asset generation. Building on this, a growing line of work has exploited these strong 3D priors for training-free stylization, transferring visual attributes from a reference image onto a generated 3D asset. However, existing methods enforce an all-or-nothing paradigm: color and texture are transferred jointly, with no mechanism to control them independently — a limitation we formalize as Disentangled 3D Stylization (Disen3D). To address this, we propose DiDE, the first training-free framework for Disen3D. Key to our approach is the observation that the structured latent space of image-to-3D models is overcomplete with respect to texture: texture information occupies only a small subset of the style-significant channels, leaving a free subspace available for independent color encoding. DiDE exploits this via a channel partition mechanism that processes a content image, a texture reference, and a color reference through dedicated branches and composes both style signals interference-free at every self-attention layer, preserving content geometry throughout. Experiments on Disen3D-Bench, our newly collected multi-reference benchmark, show that DiDE consistently outperforms 2D and 3D stylization baselines in color fidelity, texture transfer, and content preservation.
Direction-Aware Offline-to-Online Learning in Linear Contextual Bandits
Zean Han ⋅ Ruihan Lin ⋅ Zezhen Ding ⋅ Jiheng Zhang
Many bandit systems are deployed with offline historical data, such as past logs from earlier policies. By using these data, one can reduce the amount of online exploration needed early in learning. When the offline and online environments differ, such data can be biased for the online problem. For linear (contextual) bandits, this bias is directional: offline data may be informative in some feature directions and misleading in others. However, prior work typically controls this gap through a known Euclidean bound on the model parameters, which we prove is too coarse to separate easy and hard offline-to-online instances from offline data alone. To address this challenge, we introduce a directional bias certificate $(M_{\mathrm{bias}},\rho)$ that measures the offline-to-online gap through an $M_{\mathrm{bias}}$-induced norm and assigns different bias budgets to different directions. Building on this certificate, we propose \emph{Ellipsoidal-MINUCB}, which augments the online learning with an offline-pooled branch that safely exploits historical data. When the certificate is known, we show that the algorithm matches the standard SupLinUCB rate in the worst case and improves when offline coverage aligns with low-bias directions. When the certificate is unknown, we estimate it adaptively from offline and accumulated online data and establish a corresponding regret guarantee. Numerical experiments support the theory and show gains in aligned regimes.
Discovering Mechanistic Models of Neural Activity: System Identification in an in Silico Zebrafish
Jan-Matthis Lueckmann ⋅ Viren Jain ⋅ Michal Januszewski
Constructing mechanistic models of neural circuits is a fundamental goal of neuroscience, yet verifying such models is limited by the lack of ground truth. To rigorously test model discovery, we establish an in silico testbed using neuromechanical simulations of a larval zebrafish as a transparent ground truth. We find that LLM-based tree search autonomously discovers predictive models that significantly outperform established forecasting baselines. However, in-distribution accuracy is a poor proxy for faithful system identification, as models exploit statistical shortcuts. Structural priors prove essential for enabling robust out-of-distribution generalization and recovery of interpretable mechanistic models. Our insights provide guidance for modeling real-world neural recordings and offer a broader template for AI-driven scientific discovery.
Discrete Flow Matching: Convergence Guarantees Under Minimal Assumptions
Le-Tuyet-Nhi PHAM ⋅ Giovanni Conforti ⋅ Zhenjie Ren ⋅ Alain Durmus
Flow Matching has recently emerged as a popular class of generative models for simulating a target distribution $\mu_1$ from samples drawn from a source distribution $\mu_0$. This framework relies on a fixed coupling between $\mu_0$ and $\mu_1$, and on a deterministic or stochastic bridge to define an interpolating process between the two distributions. The time marginals of this process can then be approximately sampled by estimating the transition rates, or more generally the generator, of its Markovian projection. This framework has recently been extended to the case of discrete source and target distributions, under the name *Discrete Flow Matching* (DFM). However, theoretical guarantees for such models remain scarce. In this paper, we study two DFM models on $\mathbb{Z}_m^d = \\{0,\ldots,m-1\\}^d$, sampled through time discretization, and derive non-asymptotic associated bounds for both of them. In contrast to previous work, we establish non-asymptotic bounds in Kullback--Leibler divergence for the early-stopped version of the target distribution. We also derive explicit convergence guarantees in total variation distance with respect to the true target distribution. Importantly, these bounds rely only on an approximation error assumption, relaxing standard score assumptions used in earlier works, while also yielding improved dependence on the vocabulary size $m$ and the dimension $d$.
Distilling LLM Feedback for Lean Theorem Proving
Gaëtan Narozniak ⋅ Gérard Biau ⋅ Remi Munos ⋅ Ahmad Rammal ⋅ Pierre Marion
Post-training for reasoning models typically combines supervised fine-tuning with reinforcement learning from verifiable rewards, most commonly with GRPO. However, this algorithm suffers from sparse rewards, limited exploration, and mode collapse. Building upon recent works on self-distillation, we propose Feedback Distillation, a training method where the model is trained to match, at the token level, its own distribution conditioned on privileged feedback produced by a language model. Feedback Distillation offers token-level supervision and can inject external knowledge. Evaluating our method for Lean 4 theorem-proving, we find that Feedback Distillation maintains greater diversity in generated trajectories than GRPO, yielding higher policy entropy and better pass@$k$ scaling. The two methods are complementary: initializing GRPO from a Feedback Distillation checkpoint outperforms either method alone. All in all, our results suggest a promising avenue to improve post-training for complex reasoning.
Divide et Calibra: Multiclass Local Calibration via Vector Quantization
Cesare Barbera ⋅ Lorenzo Perini ⋅ Giovanni De Toni ⋅ Andrea Passerini ⋅ Andrea Pugnana
Accurate and well-calibrated Machine Learning (ML) models are mandatory in high-stakes settings, yet effective multiclass calibration remains challenging: global approaches assume calibration errors are homogeneous across the latent space, while local methods often rely on latent-space dimensionality reduction, which leads to information loss. To address these issues, we propose a compositional approach to multiclass calibration, where region-specific calibration maps are constructed from shared codeword-dependent factors. We instantiate this idea via Vector Quantization (VQ), which induces a structured partition of the representation space, and an indexed parameterization of Dirichlet concentrations that enables parameter sharing across regions. Our approach learns heterogeneous calibration maps that generalize well even to sparse regions of the latent space. Experiments on benchmark datasets show significant improvements in local calibration while maintaining competitive global calibration and predictive performance.
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
Matteo Attimonelli ⋅ Alessandro De Bellis ⋅ Aryo Gema ⋅ Rohit Saxena ⋅ Monica Sekoyan ⋅ Wai-Chung Kwan ⋅ Claudio Pomo ⋅ Alessandro Suglia ⋅ Dietmar Jannach ⋅ Tommaso Di Noia ⋅ Pasquale Minervini
Composed Image Retrieval (CIR) is a multimodal retrieval task where a query consists of a reference image and a textual modification, and the goal is to retrieve a target image satisfying both. In principle, strong performance on CIR benchmarks is assumed to require multimodal composition, i.e., combining complementary information from reference image and textual modification. In this work, we show that this assumption does not always hold. Across four widely used CIR benchmarks and eleven Generalist Multimodal Embedding models, a large fraction of queries can be solved using a single modality (from 32.2% to 83.6%), revealing pervasive unimodal shortcuts. Thus, high CIR performance can arise from unimodal signals rather than true multimodal composition. To better understand this issue, we perform a two-stage audit. First, we identify shortcut-solvable queries through cross-model analysis. Second, we conduct human validation on 4,741 shortcut-free queries, of which only 1,689 are well-formed, with common issues including ambiguous edits and mismatched targets. Re-evaluating models on this validated subset reveals qualitatively different behaviour: queries can no longer be solved with a single modality, and successful retrieval requires combining both inputs. While accuracy decreases, reliance on multimodal information increases. Overall, current CIR benchmarks conflate shortcut-solvable, noisy, and genuinely compositional queries, leading to an overestimation of model capability in multimodal composition. Datasets and annotations are available here.
Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning
Fanhu Zeng ⋅ Zhicong Luo ⋅ Zefan Wang ⋅ Li You ⋅ Chi Chen ⋅ Maosong Sun
Visual reasoning through reinforcement learning with verifiable rewards (RLVR) has achieved remarkable progress. However, when dealing with multi-source inputs, existing approaches tend to treat them as a mere accumulation of information, lacking explicit mechanisms to distinguish whether integrating additional sources yields information gain or introduces interference. Therefore, they struggle to effectively model dynamic interaction when integrating multiple sources, particularly when they differ significantly in physical properties and semantics, \eg, infrared and depth, leading to inferior performance to mono-source reasoning when a certain source holds the dominant signal. To address this issue, we propose MARS, a novel mono-anchored multi-source reasoning framework that models each visual modality as an independent information source. Specifically, by treating mono-source rewards as dynamic anchors, our method explicitly incorporates the information gain introduced by multi-source fusion into advantage normalization and adaptively emphasizes mutual promotion between sources while suppressing potential noise or conflicts during RLVR. From theoretical analysis, our method effectively quantifies information gain introduced by multi-source integration in gradient estimation, enabling consistent modality regulation. Empirical results also show impressive 3.2% and 4.9% performance gains on GRPO and DAPO across diverse datasets, confirming the effectiveness of our method.
Dynamic Optimistic Constrained OCO with Memory via Delay Equivalence
Mohammed ABDULLAH ⋅ George Iosifidis ⋅ Salah Eddine ELAYOUBI ⋅ Tijani Chahed
Online Convex Optimization with memory and time-varying constraints (COCO-M) is a recently-introduced framework that captures the dynamics of stateful online learning and non-stochastic control with budget constraints. Despite its expressive power, COCO-M remains largely understudied. In this work, we tackle this problem in its full scale, providing an optimistic algorithm that incorporates untrusted predictions to tame universal dynamic regret and cumulative constraint violation. Our analysis leverages a potent reduction from memory to delay and treats state-dependent gradients as delayed feedback. We establish the first bounds that adapt to both environment non-stationarity (path length $P_T$) and prediction accuracy, while lifting previous restrictive assumptions such as functional separability. En route to these results, we derive findings of independent interest for the memory-to-delay reduction, and for delayed OCO with time-varying constraints. The proposed framework ensures sublinear worst-case guarantees that improve with prediction accuracy, reaching \(\mathcal O(1)\) regret under perfect predictions, with cumulative constraint violation (CCV) $\mathcal O(\sqrt{T})$ for unknown $P_T$ and $\mathcal O(\log T)$ for known $P_T$, while recovering the optimal COCO rates when memory disappears.
Despite the critical need for sustainable deep learning, model energy efficiency remains largely neglected in Hardware-Aware Neural Architecture Search (HW-NAS). Discovering energy-efficient models is severely bottlenecked by costly on-device measurements and the brittleness of training-free accuracy proxies, which often yield unbalanced "glass cannon" architectures. We propose Energy-Efficient Neural Architecture Search (EENAS), a strictly zero-shot HW-NAS framework. EENAS introduces a PCA-based anomaly router that dispatches tasks to decision trees or tabular transformers, achieving <10\% MAPE zero-shot energy prediction on entirely unseen GPUs. Simultaneously, we robustly estimate accuracy by ensembling training-free proxies via a novel Log Z-score aggregation. Combining these predictors unlocks Absolute Hardware Bounding, discovering Pareto-optimal architectures for strict energy budgets entirely offline in under 30 minutes. Evaluated on ImageNet-1K, our models achieve 40\% less energy consumption at comparable accuracy, or a 2.5\% accuracy improvement under the same energy budget compared to efficient baselines.
Efficient Adaptive Data Acquisition via Pretrained Belief Representations
Daolang Huang ⋅ Zhuoyue Huang ⋅ Conor Hassan ⋅ Luigi Acerbi ⋅ Samuel Kaski ⋅ Thomas Rainforth
Learning effective policies for adaptive data acquisition remains challenging: posterior-based methods rely on surrogate models and posterior approximations that can be misspecified or biased, while direct policy-learning methods map from historical observations and fail to exploit available model representations, making learning harder. We introduce policy learning with belief representations (POLAR), based on the insight that optimal data acquisition depends on the observation history only through a sufficient belief state. Specifically, POLAR decouples representation learning from policy learning by leveraging pretrained predictive foundation models as belief-state encoders, training a policy head on top of their representations. This yields a simple, unified amortised policy learning framework for Bayesian experimental design, Bayesian optimisation, and active learning, differing only in the task-specific utility used to train the policy. Empirically, we find that POLAR outperforms state-of-the-art amortised methods across diverse tasks while requiring far fewer training samples, demonstrating a significant step in the scalability and efficiency of amortised data acquisition.
Efficient Generative Transformer Operators for Million-Point PDEs
Armand Kassaï Koupaï ⋅ Lise Le Boudec ⋅ Patrick Gallinari
We introduce ECHO, a transformer-based neural operator for large-scale PDE modeling that learns to \emph{generate full trajectories in compressed spatio-temporal latent spaces}. Scaling neural operators to high-resolution systems remains challenging: dense grids are computationally prohibitive, and autoregressive solvers accumulate errors over long horizons. ECHO addresses both issues by jointly compressing space and time and replacing step-by-step prediction with trajectory-level generation. ECHO combines a hierarchical encoder–decoder achieving up to 100× compression, a staged training strategy for high-resolution fidelity, and a latent generative process that models distributions over complete trajectories, improving long-term consistency. This formulation enables a single model to handle forward prediction, inverse problems, and interpolation. Across large-scale 2D and 3D benchmarks, ECHO achieves state-of-the-art performance and scales to million-point simulations with complex dynamics.
Few-step image generation has seen rapid progress, with consistency and meanflow-based methods significantly reducing the number of sampling steps. Despite their low inference cost, these approaches often suffer from training instability and limited scalability. Sphere Encoder is a recent alternative that produces high-quality images in only a few steps; however, it requires repeated transitions between the pixel space and latent space during inference while jointly optimizing reconstruction and generation within a single architecture. This design leads to computational inefficiency and objective conflict between reconstruction and generation. To address these limitations, we decouple the framework into a fixed pretrained image encoder and a separate latent denoising model trained entirely in a spherical latent space. Our approach eliminates repeated pixel-space operations during training and inference, improving efficiency and allowing reconstruction and generation to specialize independently. On Animal-Faces, Oxford-Flowers and ImageNet-1K datasets, our method significantly outperforms Sphere Encoder in both generation quality and inference speed, while achieving competitive results against strong few-step and multi-step baselines.
End-to-End Verification of Neuro-symbolic Automata via Contrastive Logit-Gaps
Abdelrahman Hekal ⋅ Vasileios Manginas ⋅ Nikolaos Manginas ⋅ Alessio Lomuscio
Neuro-symbolic systems for sequence classification couple a neural perception network with a finite automaton over abstract events. Certifying their trajectory-level decisions under input perturbations is critical for safety-sensitive deployment. We address the open problem of end-to-end formal robustness certification for such Neurosymbolic Automata (NeSyA) under norm-bounded input perturbations. Naive layerwise linear bound propagation fails on this composition: per-step relaxation slack compounds through the temporal recurrence (interval explosion), and the standard log-space transition score subtracts two log-sum-exp terms over overlapping coordinate sets that independent relaxations cannot tighten. To address both, we introduce contrastive logit-gap (CLG), a reparameterization whose two log-sum-exp heads act on disjoint coordinates and admit substantially tighter linear bounds, coupled with a fused temporal operator that yields corner-tight bounds from a single linearization. On NeSyA sequence benchmarks, our formulation scales significantly beyond recursive baselines, maintaining tight safety certificates at long horizons while reducing verification wall time by up to ${\sim}6{\times}$. To our knowledge, this is the first end-to-end formal verification framework for temporal neuro-symbolic systems.
Entropy Minimization without Model Collapse: Mitigating Prediction Bias in Medical Imaging
Tim Nielen ⋅ Sameer Ambekar ⋅ Daniel Lang ⋅ Julia Schnabel
Entropy minimization (EM) is the dominant objective for test-time adaptation, yet its failure mode, model collapse, remains poorly understood. In this work, we show that distribution shifts can cause feature clusters corresponding to distinct classes in the model’s representation space to merge, while the decision boundary remains fixed. This induces a systematic skew in the predicted class distribution, referred to as prediction bias. Prediction bias refers to a shift in the predicted class distribution, with some classes overrepresented and others suppressed. We show that entropy minimization amplifies this prediction bias by tightening the existing clusters, reinforcing the incorrect groupings until all predictions collapse to a trivial solution. Next, to demonstrate the significance of prediction bias and mitigate it, we further propose Distribution Shift Bias Reduction (DSBR), a bias-correcting objective that specifically targets this failure mode by equalizing the contribution of each predicted class to the unsupervised entropy minimization loss. To study this failure mode, we design suitable adaptation settings using four medical‑imaging datasets and additionally evaluate on ImageNet‑C. We find that DSBR consistently stabilizes test‑time adaptation, prevents model collapse, and matches or outperforms the state‑of‑the‑art methods. Moreover, DSBR operates solely at test-time.
We introduce epistemic pairwise maximin share (EPMMS), a new fairness notion for fair division of indivisible goods. Two fundamental notions in this setting are envy-freeness up to any item (EFX) and pairwise maximin share (PMMS), with PMMS being stronger than EFX. While EFX has been extensively studied, far less is known about PMMS. Recent work shows that relaxing EFX via an epistemic perspective leads to substantial progress on the EFX problem, raising the question of whether a similar approach can advance our understanding of PMMS. Motivated by this, we initiate the study of EPMMS, the epistemic relaxation of PMMS. EPMMS is more challenging than EEFX: the key approaches underlying recent progress on epistemic EFX inherently fail to extend to EPMMS. We establish the following results. 1. For additive valuations, $4/5$-EPMMS allocations exist and can be efficiently computed. 2. For bivalued valuations, EPMMS allocations exist and can be efficiently computed; in fact, we obtain the stronger guarantee of epistemic groupwise maximin share (EGMMS), which also strengthens the existence of MMS allocations for this setting. 3. EPMMS allocations exist in two settings where MMS allocations need not exist: instances with three additive agents or two types of additive agents.
Escaping Parameter Space: Tight Generalization Bounds via Representation Quality
Niclas A Göring ⋅ Shuofeng Zhang ⋅ Branton DeMoss ⋅ Ard Louis
We derive a tight generalization bound for deterministic neural networks based on their last hidden layer representations. By evaluating representation quality with a parameter-free $k$-nearest-neighbor ($k$-NN) classifier, our certificate avoids traditional parameter-space complexity measures, KL divergences to a posterior, weight norms, and compression. The bound is tight: trained from scratch, it reaches an 8.0 percentage point gap to test error on CIFAR-10 (ResNet-18) and yields the first non-vacuous certificates we are aware of on CIFAR-100; with frozen DINOv2 representations, it certifies at 1.4\% on CIFAR-10 and 23.1\% on 1,000-class ImageNet. It decomposes additively into three empirical, independently measurable quantities: local class geometry ($k$-NN decoder error), decoder-network agreement, and training stability. This allows the bound to act as a diagnostic tool: we demonstrate that actively destroying local class geometry eliminates generalization while leaving dataset memorization intact, identifying class geometry as a causally relevant mechanism for generalization. Finally, we show that learned representations are significantly less sensitive to weight noise than linear readouts, suggesting why parameter-space certificates often struggle to remain tight.
In recent years, many high-stakes societal decisions are made by machine learning systems, often under strict capacity constraints that limit resources to a small subset of individuals. In such settings, information gaps across populations can lead to highly unbalanced allocations. We study prediction-driven resource allocation through the lens of fairness. We connect classical algorithmic fairness notions with resource allocation definitions, and characterize the trade-offs between fairness and utility. We introduce an adaptation of proportional fairness to this setting, showing that it yields a continuum of fairness criteria via a regularization parameter, along with quantitative bounds on the resulting price of fairness. We further show that max-min fairness and equal opportunity can incur an unbounded price of fairness in extreme cases, and propose a variant of equal opportunity with a bounded price of fairness.
Federated Learning by Utility-Constrained Stochastic Aggregation for Improving Rational Participation
Yashwanth Mandula ⋅ Arunabh Singh ⋅ Ashok Nayak ⋅ Saikiran Bulusu ⋅ Anirban Chakraborty
Federated Learning (FL) algorithms implicitly assume that clients passively comply with server-side orchestration by sharing local model updates upon server request. However, this overlooks an important aspect in real-world cross-silo environments: clients are often rational agents who may prioritize their utilities such as local model performance over that of the global model. In settings with significant statistical heterogeneity, rational clients may opt out of the federation if the perceived benefits of collaboration fail to meet their local utility thresholds. Such attrition degrades the global model performance and can lead to the collapse of the federated training process. In this work, we introduce FedUCA: Federated Learning by Utility-Constrained Stochastic Aggregation for Improving Rational Participation, a framework that formalizes the server's role as an optimizer seeking to maximize global model performance by sustaining client participation. We substantiate our framework through extensive experiments on standard datasets demonstrating that by prioritizing participation feasibility, FedUCA achieves significantly higher client retention and, consequently, a superior global model performance.
Fill the GAP: A Granular Alignment Paradigm for Visual Reasoning in Multimodal Large Language Models
Yanting Miao ⋅ Yutao Sun ⋅ Dexin Wang ⋅ Pascal Poupart ⋅ Mengyu Zhou ⋅ Lei Lv ⋅ Qi Zhao ⋅ Li Wang ⋅ Hao Li ⋅ xiaoxi jiang ⋅ Guanjun Jiang
Visual latent reasoning lets a multimodal large language model (MLLM) create intermediate visual evidence as continuous tokens, avoiding external tools or image generators. However, existing methods usually follow an output-as-input latent paradigm and yield unstable gains. We identify a feature-space mismatch behind this instability: dominant visual-latent models build on pre-norm MLLMs and reuse decoder hidden states as predicted latent inputs, even though these states occupy a substantially different norm regime from the input embeddings the model was trained to consume~\citep{xie2025mhc,li2026siamesenorm,team2026attention}. This mismatch makes direct latent feedback unreliable. Motivated by this diagnosis, we propose \textbf{GAP}, a \textbf{G}ranular \textbf{A}lignment \textbf{P}aradigm for visual latent modeling. GAP aligns visual latent reasoning at three levels: feature-level alignment maps decoder outputs into input-compatible visual latents through a lightweight PCA-aligned latent head; context-level alignment grounds latent targets with inspectable auxiliary visual supervision; and capacity-guided alignment assigns latent supervision selectively to examples where the base MLLM struggles. On Qwen2.5-VL 7B, the resulting model achieves the strongest average perception and reasoning performance among our supervised variants. Intervention probing further indicates that generated latents provide task-relevant visual signal rather than merely extra token slots.
FineVision: Open Data Is All You Need
Luis Wiedmann ⋅ Orr Zohar ⋅ Amir Mahla ⋅ Xiaohan Wang ⋅ Rui Li ⋅ Thibaud Frere ⋅ Leandro Von Werra ⋅ Aritra Roy Gosthipaty ⋅ Andrés Marafioti
The advancement of vision-language models (VLMs) is hampered by a fragmented landscape of inconsistent and contaminated public datasets. We introduce FineVision, a meticulously collected, curated, and unified corpus of 24 million samples---the largest open resource of its kind. We unify more than 200 sources into 185 subsets via a semi-automated, human-in-the-loop pipeline: automation performs bulk ingestion and schema mapping, while reviewers audit mappings and spot-check outputs to verify faithful consumption of annotations, appropriate formatting and diversity, and safety; issues trigger targeted fixes and re-runs. The workflow further applies rigorous de-duplication within and across sources and decontamination against 66 public benchmarks. FineVision also encompasses agentic/GUI tasks with a unified action space; reviewers validate schemas and inspect a sample of trajectories to confirm executable fidelity. Models trained on FineVision consistently outperform those trained on existing open mixtures across a broad evaluation suite, underscoring the benefits of scale, data hygiene, and balanced automation with human oversight. We release the corpus and curation tools to accelerate data-centric VLM research.
FlexTab: A Flexible Encoder-Decoder Architecture for In-Context Learning Across Diverse Tabular Tasks
Marek Polewczyk ⋅ Maximilian Schambach ⋅ Marco Spinaci ⋅ Sam Thelin ⋅ Johannes Höhne
We introduce FlexTab, a flexible encoder-decoder architecture for in-context learning on tabular data that pairs a single, task-agnostic encoder with a suite of task-specific decoders. Unlike existing tabular in-context learners, which entangle feature representations with a specific prediction target, our design produces \textit{target-agnostic} row embeddings that can be leveraged across a wide range of downstream tasks within a table-native in-context learning setup. We demonstrate this flexibility on six distinct problems: classification, regression, anomaly detection, clustering, entity matching, and entity classification in relational databases. Both the encoder and the task-specific decoders are trained on a large corpus of real-world, unlabeled tables. FlexTab achieves state-of-the-art performance on classification, regression, anomaly detection and entity matching, while remaining competitive with specialized models on entity classification in a relational setting. These results demonstrate that a single shared encoder, paired with task-specific decoders, can serve as an effective general-purpose backbone for diverse tabular prediction problems. The inference code and checkpoints will be made publicly available.
Flow-based Spectral Kernel Learning for Nonstationary Attention
Zi Yang ⋅ Ying Li ⋅ Alejandro Lancho ⋅ Dario C. Larese ⋅ Pablo Olmos ⋅ Michael Minyi Zhang
Kernelized attention provides a useful perspective for understanding self-attention as a similarity-based aggregation mechanism, but existing approaches often rely on fixed or stationary kernels, limiting their ability to adapt the similarity structure to complex token interactions. In this paper, we propose flow-based spectral kernel attention (FSKA), a framework for learning expressive nonstationary attention kernels. FSKA formulates kernelized attention through spectral density learning and parameterizes the resulting bivariate spectral density with normalizing flows. This design enables flexible and regularizable kernel learning while remaining compatible with scalable linear-attention computation. We further show that expressive kernel learning enables a simplified attention design with removed query--key projections and fixed orthogonal value projections, resulting in a more parameter-efficient architecture. Controlled experiments on classification, ablation, and scalability profiling show that FSKA achieves competitive or improved performance over existing attention mechanisms, while reducing projection-related trainable parameters by approximately 40\%--47\% compared with most baselines and maintaining favorable linear-attention memory scaling.
FlowBatt: Flow Matching for Probabilistic Battery Degradation Prediction
Hamidreza Eivazi ⋅ Eslam Sharaawy ⋅ Stefan Wittek ⋅ André Hebenbrock ⋅ Raphael Ginster ⋅ Steffen Blömeke
Battery degradation remains a central challenge in the development and broad application of sustainable energy technologies. Accurate degradation prediction is challenging, as battery aging emerges from complex and heterogeneous interactions among cycling behavior, operating conditions, and cell chemistry. Existing machine-learning approaches typically focus on deterministic end-of-life predictions or modeling degradation curves within fixed chemistry settings, leaving probabilistic state-of-health (SOH) trajectory modeling across heterogeneous chemistries underexplored. We introduce FlowBatt, a use-inspired conditional generative framework that adapts flow matching with a diffusion transformer (DiT) backbone to battery degradation trajectory modeling. FlowBatt models full SOH trajectories as probabilistic degradation processes conditioned on early-cycle capacity data. To strengthen this conditioning pathway, we use explainable AI analysis to diagnose and improve the encoder that maps early-cycle capacity matrices into conditioning vectors. We evaluate FlowBatt on five public benchmark settings spanning diverse chemistries and aging conditions, comparing it with supervised and diffusion-based trajectory models as well as established baselines for remaining-useful-life (RUL) prediction under shared data splits. FlowBatt achieves the best SOH prediction performance on four of five datasets, with errors below 2.5%, and delivers competitive RUL prediction performance. By generating multiple plausible trajectories, FlowBatt provides empirical uncertainty estimates, while also revealing calibration challenges under some dataset shifts. These results indicate that flow-matching-based trajectory modeling is a promising framework for probabilistic battery health prediction.
Flow Matching with In-Context Priors for Out-of-Distribution Brain Dynamics
Sam Gijsen ⋅ Michał Łukomski ⋅ Marc-Andre Schulz ⋅ Kerstin Ritter
Conditional flow matching produces realistic out-of-distribution samples across modalities from images to proteins, yet the conditioning signals that enable extrapolation remain poorly understood for biological time series. While considerable work has focused on areas of biology such as proteins and molecules, generative models of neural time series have been largely restricted to categorical conditioning, which precludes compositional and zero-shot generalization. In this work, we propose a per-timestep conditioned diffusion transformer for generating realistic fMRI brain dynamics during unseen cognitive tasks, by injecting both compositional language and optional spatial priors in-context. Such zero-shot generation would support in-silico task design and counterfactual evaluation of novel cognitive experiments before costly scanner acquisition. Leveraging this model, we evaluate across hundreds of held-out task conditions and characterize predictive performance in relation to the training manifold. From language alone, the model recovers region-specific recruitment across tasks and held-out spatial activation patterns. Spatial priors, when available, complement the text pathway by anchoring generation in regions of task space where language alone degrades, while retaining the compositional structure needed for counterfactual task specification. To our knowledge this is the first generative model of whole-cortex fMRI dynamics for unseen cognitive tasks, enabling counterfactual neuroscience and data-driven experimental design.
FMMI: Flow Matching Mutual Information Estimation
Ivan Butakov ⋅ Alexander Semenenko ⋅ Valeriia Kirova ⋅ Ivan Oseledets ⋅ Alexey Frolov
We introduce a novel Mutual Information (MI) estimator that fundamentally reframes the discriminative approach. Instead of training a classifier to discriminate between joint and marginal distributions, we learn a normalizing flow that transforms one into the other. This technique produces a computationally efficient and precise MI estimate that scales well to high dimensions and across a wide range of ground-truth MI values.
FocusNav: Learning Task-directed Perception via Action-Aware Future Reconstruction in Vision-and-Language Navigation
Caifeng liu ⋅ Xueqiong Li ⋅ Chenghao Shi ⋅ weida chen ⋅ Yuhao Wang ⋅ Quan Zhang ⋅ Yuanxi Peng ⋅ Guijian Tang ⋅ Shaowu Yang
Vision-and-Language Navigation(VLN) requires agents to sustain task-directed spatial perception to focus on navigation-critical regions while translating linguistic intent into low-level actions. Prevalent hierarchical planning methods rely on a high-level policy to select waypoints or subgoals for low-level control, but they often fail to remain focused, causing inaccurate or redundant subgoals that directly mislead execution. Moreover, their separation between visual perception and action execution neglects fine-grained spatial reasoning in low-level movements and limits real-world deployability. To address these two issues, we propose FocusNav, a grounding-driven end-to-end monocular framework that learns task-directed spatial perception through action-aware future reconstruction. Concretely, FocusNav first boosts the agent’s visual attention on task-relevant regions with a Grounding-driven Visual Reconstruction module that reconstructs instruction-relevant targets and traversable areas via a semantic-guided latent diffusion process. Subsequently, to bridge the perception–execution gap, we incorporate physical action priors via Action-aware Visual Conditioning and enforce Intent-Effect Alignment between anticipated visual changes and observed action effects. These effectively enhance the agent’s task-directed spatial perception to direct attention to navigation-critical regions, and tighten the perception–execution loop. Experiments on R2R-CE and RxR-CE benchmarks show that FocusNav achieves substantially improved state-of-the-art performance with a simple monocular low-level policy, while providing faster inference than recent methods. Furthermore, real-world robot experiments validate our FocusNav’ s effectiveness, ease of deployment, and lightweight design.
From Representation to Intervention: Using Emotion Vectors to Monitor and Guide Language Models
Georgia Dimaki ⋅ Martin Villanueva Dimakis
Recent work has shown that large language models form linear representations of emotion concepts that causally influence alignment-relevant behavior, including blackmail, reward hacking, and sycophancy. These findings raise a natural question: can such representations be used not only to characterize model behavior in controlled evaluations, but also as practical tools for monitoring and guiding models during deployment? We explore this possibility, investigating how emotion-vector activations can serve as interpretable signals of a model's internal state over the course of a conversation or agentic task. We discuss how these signals might support debugging, early detection of drift toward misaligned behavior, and lightweight steering interventions to nudge the model back toward desired behavior. Beyond the Assistant, the same representations apply to other characters in the model's context — including the user — suggesting a broader family of applications in which internal representations of emotional state inform how systems respond. We outline the opportunities and limitations of this approach and highlight open questions around validation, robustness, and responsible use.
Generating Symmetric Materials using Latent Flow Matching
Anmar Karmush ⋅ Cedric M Brandenburg ⋅ Soheil Ershadrad ⋅ Johanna Rosen ⋅ Michael Felsberg ⋅ Filip Ekström Kelvinius
Tackling the task of materials generation, we aim to enhance the previously proposed All-atom Diffusion Transformer (ADiT) by introducing SymADiT, a symmetry-aware variant. To do so, we use a representation of materials based on Wyckoff positions. We follow ADiT and perform generative modelling in latent space, adapted to our symmetry-aware representation. By forcing the output of the generative model to adhere to the symmetry restrictions imposed by the generated crystal's space group and each atom's Wyckoff-position, the generated materials exhibit more realistic symmetry properties. We benchmark our method against both symmetry-aware and symmetry-agnostic models for materials generation and show competitive performance, generating stable, symmetric materials with a simple Transformer architecture.
Geometric Alignment without Functional Equivalence: A Layer-wise Analysis of the Speech-Text Modality Gap
Ming-Hao Hsu ⋅ Xiaohai Tian ⋅ Jun Zhang ⋅ Lu Lu ⋅ Yuxuan Wang ⋅ Zhizheng Wu
End-to-end speech LLMs often answer the same question less accurately from speech than from text, even when the spoken content is semantically matched to the text prompt. We study this modality gap as a layer-wise inference problem rather than a static embedding mismatch. Across four open-weight speech LLMs on SpeechMMLU and VoiceBench BBH, cross-layer Centered Kernel Alignment (CKA) with speech–text token alignment shows that speech can become geometrically text-like in middle layers while still failing to form stable late-layer answer margins. The central finding is that mid-layer alignment does not imply functional equivalence. ASR-to-LLM controls show that recognition explains part of the gap in some settings but not all of it. Matched-instance margin analysis localizes the failure to weak late-layer answer separation. Attention diagnostics further show that speech evidence remains more diffuse at the decision token. Speech-faithful temporal compaction improves BBH accuracy from 59.9% to 62.1% with tempo 1.2× and to 61.9% with silence trimming, while recovering 34.3% and 38.2% of text-correct/speech-wrong failures. These conditional repairs reveal a 2.35–2.62× asymmetry between temporal compaction and quantity-only KV merging on the same gap subset, motivating post-projection temporal abstraction rather than only input-level feature matching.
Geometry-Aware Optimal Transport: Fast Intrinsic Dimension and Wasserstein Distance Estimation
Ferdinand Genans ⋅ Olivier Wintenberger
Solving large scale Optimal Transport (OT) in machine learning typically relies on sampling measures to obtain a tractable discrete problem. While the discrete solver's accuracy is controllable, the rate of convergence of the discretization error is governed by the intrinsic dimension of our data. Therefore, the true bottleneck is the knowledge and control of the sampling error. In this work, we tackle this issue by introducing novel estimators for both sampling error and intrinsic dimension. The key finding is a simple, tuning-free estimator of $\text{OT}_c(\rho, \hat\rho)$ that utilizes the semi-dual OT functional and, remarkably, requires no OT solver. Furthermore, we derive a fast intrinsic dimension estimator from the multi-scale decay of our sampling error estimator. This framework unlocks significant computational and statistical advantages in practice, enabling us to (i) quantify the convergence rate of the discretization error, (ii) calibrate the entropic regularization of Sinkhorn divergences to the data's intrinsic geometry, and (iii) introduce a novel, intrinsic-dimension-based Richardson extrapolation estimator that strongly debiases Wasserstein distance estimation. Numerical experiments demonstrate that our geometry-aware pipeline effectively mitigates the discretization error bottleneck while maintaining computational efficiency.
Geometry of Relaxed Fair Regression: A Unified Framework for Aware and Unaware Settings
Marie Generali Lince ⋅ Vincent Divol ⋅ Rémi Flamary ⋅ Solenne Gaucher ⋅ Patrick Loiseau
Fairness-accuracy trade-offs are a central concern in the deployment of fairness-aware machine learning methods. When sensitive attributes are unavailable at inference time–the so called unawareness setting, principled methods for obtaining accurate predictions under relaxed fairness constraints are largely missing. In this work, we address this gap by formulating regression under a demographic parity penalty as an optimal transport problem. Our framework unifies both the \emph{aware} and \emph{unaware} settings and characterizes optimal prediction functions via optimal transport maps, under both squared Wasserstein-2 and Total Variation penalties. These results reveal that the choice of penalty reflects fundamentally different fairness philosophies: the Wasserstein penalty induces a smooth, population-wide compromise, while Total Variation enforces exact parity for a subset of individuals. Building on these theoretical characterizations, we propose an algorithm that is simple to implement, computationally efficient, and consistently matches or outperforms state-of-the-art baselines on real-world benchmarks.
GeoX: Mastering Geospatial Reasoning Through Self-Play and Verifiable Rewards
Kyeongjin Ahn ⋅ Seungeon Lee ⋅ Krishna Gummadi ⋅ Meeyoung Cha
Geospatial reasoning requires solving image-grounded problems over complex object relationships and scene structures. However, developing this capability is hindered by the cost of annotating a vast and combinatorial question space. We propose GeoX, a self-play framework that acquires spatial logic through executable programs and verifiable rewards in the absence of large-scale human-curated data. Given a satellite or aerial image, our framework employs a single multimodal policy to alternate between proposing spatial problems with executable programs and solving them under three reasoning modes of abduction, deduction, and induction, using a segmentation tool and standard computational libraries. Program execution turns each program into a reward that jointly optimizes the two roles via reinforcement learning. GeoX consistently improves its base VLMs by up to 5.5 points on average, matching or exceeding conventional baselines trained on millions of curated samples. Alongside the framework, we release a benchmark for geospatial reasoning curated from self-play generations.
GitInject: Real-World Prompt Injection Attacks in AI-Powered CI/CD Pipelines
Jafar Isbarov ⋅ Umid Suleymanov ⋅ I Shumailov ⋅ Murat Kantarcioglu
AI-powered agents are increasingly embedded in CI/CD pipelines to autonomously review pull requests, triage issues, and maintain codebases. These agents ingest untrusted content while operating with elevated repository permissions, making them a natural target for prompt injection attacks with supply chain consequences. We present GitInject, an open-source framework for evaluating prompt injection vulnerabilities in real, live GitHub workflows, a widely deployed instance of CI/CD pipelines. Unlike prior agent security benchmarks that simulate tool calls, GitInject provisions ephemeral repositories and triggers actual workflow runs, so that sandbox constraints, credential handling, and permission boundaries behave exactly as in production. Using GitInject, we study workflow configurations across four AI providers and document eleven named attacks spanning config-file injection, credential exfiltration, judgment manipulation, and availability. We find that all tested providers are susceptible to at least one attack class in their default configuration, and that the most critical vulnerabilities are structural: they arise from how CI/CD infrastructure handles credentials and configuration files, not from any specific model's behavior. For each confirmed attack class, we identify the minimum-cost workflow-level countermeasure and analyze its coverage and limitations. We release the framework to reviewers and will release the full attack scenarios publicly after a disclosure period to allow providers time to address the identified vulnerabilities.
GLUT: 3D Gaussian Lookup Table for Continuous Color Transformation
Danna Xue ⋅ David Serrano-Lozano ⋅ Shaolin Su ⋅ Javier Vazquez-Corral
3D Lookup Tables (3D LUTs) are widely used for color mapping, but their grid-based representation requires discretizing the RGB space, leading to a capacity-memory trade-off that becomes prohibitive when storing large numbers of LUTs. Recent approaches adopt implicit neural representations to improve scalability, yet their black-box nature limits interpretability and hinders intuitive, localized editing. In this paper, we propose Gaussian LUT (GLUT), a continuous and explicit color representation that models color transformations using a set of learnable 3D Gaussian primitives. By avoiding fixed-resolution grids, GLUT achieves flexible representational capacity while maintaining a compact memory footprint. Its explicit, spatially localized formulation further enables both accurate modeling and interpretability. Building on this representation, we introduce a compact conditional generator (CGLUT) that predicts GLUT parameters for multiple LUT instances, encoding diverse color styles in a single framework to enable smooth and controllable LUT style blending. Moreover, GLUT supports efficient, user-friendly editing by allowing localized adjustments to specific color regions without global retraining. Experimental results demonstrate that our approach outperforms prior neural LUT representations in both accuracy and efficiency, while offering improved interpretability and interactive control. We are committed to releasing the code and models upon acceptance.
How novel ideas are discovered and spread through a population is a well-studied phenomenon in the fields of computer science, computational social science, marketing/business research and higher education studies. In this paper, we study a particular version of this phenomenon, namely the spread of ideas among a network of researcher agents. We use a computational model in which each agent has a finite computational budget, that it can spend on a mixture of conducting its own research, and learning from the research of others in the network. By framing the problem as a continuous optimisation on the edge weights, we can applying gradient-based optimisation, similar to that from graph neural networks, to search for the graph configuration that maximises the amount of knowledge uncovered by the network. This contrasts with existing works that simply compare a fixed set of configurations. We find that (i) the optimisation succeeds in producing graphs that lead to a high total knowledge score, (ii) the optimum configuration has somewhat short shortest path lengths and a low clustering coefficient, (iii) the nodes end up broadly divided between researchers, who focus solely on their own research, and learners, who focus mostly on viewing research of others, and (iv) the learners consistently have higher knowledge scores, and this drives a rise in inequality across nodes over time.
Hardware-Friendly Token-Group Activation Quantization for Low-Bit Mamba Super-Resolution
Jinwoo Chung ⋅ Jangho Kim
Mamba-based super-resolution models are efficient alternatives to Transformer-based restoration backbones, but their low-bit post-training quantization remains challenging because selective scan, Mamba's input-dependent state-space operator, induces different activation ranges across spatial tokens. Thus, a single layer-wise clipping range poorly matches all tokens. We propose Token-Group Quantization, a hardware-friendly activation PTQ framework that statically approximates a hardware-unfriendly post-scan grouping oracle, where tokens with similar preferred clipping ranges are assigned to the same group. To avoid runtime post-scan grouping in static low-bit inference, a lightweight predictor infers the groups from pre-scan features available before the quantized scan path. During calibration, we form pseudo-labels by clustering post-scan activation descriptors that capture similar rounding--clipping behavior. At inference, the frozen predictor maps pre-scan features to groups and selects entries from offline-calibrated group-wise clipping-bound tables, without runtime statistics collection, dynamic clipping updates, or post-scan group assignment. To address low-bit restoration degradation, we further introduce a Frequency-Preserving Refinement (FPR) objective for clipping-bound calibration. FPR aligns Fourier amplitude and phase between full-precision teacher and quantized student outputs to preserve high-frequency structures. Across PTQ protocols, bit-widths, backbones, and restoration tasks, our method consistently improves restoration quality over prior Mamba and SR PTQ baselines, while achieving $2.23\times$ speedup on Jetson Orin Nano over FP16.
Harmonic Torsional Diffusion for Flexible Protein-Ligand Docking
Maksim Zhdanov ⋅ Pavel Strashnov ⋅ Vladislav Kurenkov
Molecular docking requires reasoning jointly about ligand pose and protein flexibility. Most diffusion-based docking models predict torsional updates with generic Euclidean heads that ignore the periodic geometry of angular variables. This mismatch is especially limiting in flexible docking, where ligand conformations and pocket side chains co-adapt to form the bound complex. Here, we introduce Harmony, a harmonic torsional diffusion framework for flexible protein-ligand docking. Harmony parameterizes ligand and side-chain torsional score fields as derivatives of learned harmonic potentials on the circle, whose noise-level dependence is supplied analytically by the heat semigroup of variance-exploding diffusion on the torus. This construction makes periodicity explicit, removes the need to learn noise conditioning for torsional updates, and gives the model a frequency-aware inductive bias over rotameric motion. On the PDBBind benchmark, Harmony improves ligand pose accuracy and pocket all-atom reconstruction over recent flexible docking methods. On PoseBusters, it improves the physical validity of generated complexes. Case studies on EBNA1 and KRAS G12D illustrate the method's behavior on a polar and a shallow binding site, respectively. Together, these results indicate that aligning the score parameterization with the geometry of the diffusion process is a simple and effective lever for improving flexible docking.
Harnessing Textual Refusal Directions for Multimodal Safety
Moreno D'Incà ⋅ Nicu Sebe ⋅ Massimiliano Mancini
To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activation space. Both strategies are less feasible in Multimodal LLMs (MLLMs) as they require unsafe multimodal data, harder to collect than their unimodal counterpart. In this work, we relax this constraint and investigate whether textual refusal directions, extracted directly from the LLM backbone, generalize across modalities (i.e., image, video). Preliminary findings confirm this ability, though effectiveness is conditioned by layer selection, steering strength, and cross-modal alignment, with the latter causing safe multimodal inputs to be spuriously steered toward refusal. Building on this, we introduce Modality-Agnostic Refusal Steering (MARS), a light-weight training-free approach that injects multimodal safety without the need for multimodal safety data. MARS corrects modality misalignment via activation re-centering, adaptively scales steering strength within a geometrically defined trust region, and selects the optimal intervention layer, operating at the first generated token. Evaluated on five SOTA MLLMs across safety, utility, and video jailbreak benchmarks, MARS achieves consistent safety gains while preserving utility. These results reveal that safety-relevant structure is shared across modalities and that textual refusal directions are a powerful and underexplored foundation for multimodal alignment.
Hidden Positives: Why Code Retrieval Benchmarks Underestimate Model Quality
Maria Ivanova ⋅ Vitaly Kadulin ⋅ Dmitrii Babaev
Text-code embeddings underpin modern code intelligence systems, but progress in this area is held back by a systemic problem: code retrieval benchmarks routinely treat valid alternative implementations as negatives, making cross-model comparison unstable and depressing measured performance. We document this problem with a detailed analysis of six widely used benchmarks and show that, depending on the dataset, between 44\% and 87\% of items a strong LLM judge marks as relevant are missing from the original labels. To address this, we adopt a unified evaluation protocol in which a single LLM judge re-annotates the top-10 retrieved candidates for every model under a fixed prompt, replacing fragmented benchmark heuristics with a consistent semantic standard. Applying the same idea at training time, we use a smaller LLM-supervised reranker to construct MegaCode, the largest semantically curated dataset for code search to date of 193M positive (text, code) pairs, roughly an order of magnitude larger than prior curated corpora, and train a family of embedding models on it at four scales (0.5B-7B) with contrastive learning. Under a unified LLM-judged protocol, our 0.5B model matches or exceeds all prior open code-retrieval embeddings up to 8B parameters, and our 7B model improves average MRR@10 by 3.5 points over the strongest baseline.
Hierarchical Graph Representation Learning with Pooling-Induced Substructures
Luca Sbicego ⋅ Xiaowen Dong ⋅ Dorina Thanou
Hierarchical graph-level tasks that require reasoning over substructures that extend beyond local neighborhoods are posing challenges for standard Graph Neural Networks (GNNs). While graph pooling aims to address this by coarsening graphs into hierarchical representations, existing methods often fail to explicitly preserve the substructures they induce, leading to inconsistent empirical performance. Moreover, the evaluation of such approaches has proved difficult due to the lack of benchmarks with explicit hierarchical structure. We propose Pooled-Substructure-Aware GNNs, an approach to graph pooling that explicitly preserves and processes substructures uncovered by structure-based pooling schemes. We show that this design strengthens pooling in terms of Weisfeiler-Lehman (WL) expressivity. To enable principled evaluation of pooling-based models, we introduce two new synthetic datasets with controllable hierarchical structure and propose a metric to characterize the extent of hierarchical patterns in real-world graph-level benchmarks. On tasks with relevant hierarchical structure, our approach outperforms matched pooling baselines, while remaining competitive on structurally simpler benchmarks.
How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
Leena Chennuru Vankadara ⋅ Moritz Haas ⋅ Luke Hayward ⋅ Sebastian Bordt ⋅ Alessandro Breccia
Recent frontier large language models predominantly rely on Mixture-of-Experts (MoE) architectures. Despite empirical progress, there is still no principled understanding of how hyperparameters should scale with network width $N$, expert width $N_e$, number of experts $M$, sparsity $K$, and depth $L$ to ensure both stability and optimal performance at scale. We take a principled step toward resolving this gap by analyzing three different scaling regimes: (I) co-scaling $N\asymp N_e$, (II) co-scaling $N\asymp M\asymp K$, and (III) full proportional scaling of $N, N_e, M$, and $K$. For each regime, we develop a novel Dynamical Mean Field Theory (DMFT) description of the limiting training dynamics of MoEs that provides a formal foundation for our analysis. Within this framework, we derive the unique parameterization for SGD and Adam satisfying all maximal-update ($\mu$) desiderata. We then show that the resulting $\mu$P prescription induces neither monotonic improvement with scale nor robust learning-rate transfer. We trace these pathologies to scale-dependent observables in the aggregation dynamics, which motivates a refined set of desiderata that we term *maximal scale stability*. Guided by this principle, we derive a *Maximally Scale-Stable Parameterization* (MSSP) for both SGD and Adam in all three scaling regimes, and characterize the corresponding limiting dynamics - qualitatively distinct from the $\mu$P limit - through a separate DMFT analysis. Experiments verify that MSSP robustly recovers learning rate transfer and monotonic improvement with scale across regimes. Combined with existing depth-scaling theory, these results provide a complete scaling prescription for MoE architectures as a function of width, depth, expert width, and number of experts.
How to Train Your Latent Diffusion Language Model Jointly With the Latent Space
Viacheslav Meshchaninov ⋅ Alexander Shabalin ⋅ Egor Chimbulatov ⋅ Nikita Gushchin ⋅ Ilya Koziev ⋅ Aleksandr Korotin ⋅ Dmitry Vetrov
Latent diffusion models offer an attractive alternative to discrete diffusion for non-autoregressive text generation by operating on continuous text representations and denoising entire sequences in parallel. The major challenge in latent diffusion modeling is constructing a suitable latent space. In this work, we present the Latent Diffusion Language Model (LDLM), in which the latent encoder, diffusion model, and decoder are trained jointly. LDLM builds its latent space by reshaping the representations of a pre-trained language model with a trainable encoder, yielding latents that are easy to both denoise and decode into tokens. We show that naive joint training produces a low-quality diffusion model, and propose a simple training recipe consisting of an MSE decoder loss, diffusion-to-encoder warmup, adaptive timestep sampling, and decoder-input noise. Ablations show that each component substantially impacts generation performance. On OpenWebText and LM1B, LDLM achieves better generation performance than existing discrete and continuous diffusion language models while being $2{\text -}13\times$ faster, indicating that jointly learning the latent space is a key step toward making latent diffusion competitive for text generation.
IDEAL: Interaction Dynamics and Force-Aware Learning for Dual-Humanoid Collaborative Manipulation
Zhaoyang Li ⋅ Chiyu Zhang ⋅ Chaoyue Li ⋅ Xiao Zhang ⋅ Wanting Li ⋅ Ziyu Chen
anoid collaborative manipulation extends the payload and workspace capabilities of a single robot, enabling applications such as large-object carrying and cooperative transportation. However, imitation learning for Dual-Humanoid-Object Interaction (DHOI) remains underexplored, particularly in real-world settings where policies must coordinate motion and maintain force consistency under closed-chain coupling and restricted observations. We propose IDEAL, an imitation learning framework for DHOI that combines interaction-consistent kinematic references with explicit dynamics priors. IDEAL first recovers dual-human motions from monocular videos and refines them for interaction consistency. It then formulates collaborative manipulation as a rigid-body grasping problem and solves for desired force allocations under dynamics, contact, and friction constraints. Finally, IDEAL learns control policies through an interaction-dynamics and force-aware reinforcement learning framework guided by both kinematic and dynamic references. Experimental results show that IDEAL successfully achieves robust DHOI in both simulation and real-world deployments. As far as we know, IDEAL is the first end-to-end learning framework for DHOI deployed on real dual-humanoid robots, offering a new paradigm for complex real-world humanoid collaboration.
i-DEQ: A stable inertial Deep Equilibrium model for image restoration
Antonin Clerc ⋅ Marien Renaud ⋅ Baudouin Denis de Senneville ⋅ Nicolas Papadakis
Deep Equilibrium Models (DEQs) are an established framework for image restoration that learn a problem-adapted regularization by solving a fixed-point (i.e. equilibrium) problem. While flexible and expressive, DEQs are often hindered by high computational cost and training instability. We propose an inertial DEQ (i-DEQ) that learn an explicit nonconvex regularization within the DEQ formulation. By using momentum within the fixed-point iterations, i-DEQ has convergence guarantees and accelerated rates. Moreover, we observe that i-DEQ is significantly more stable during the training and robust to rough initialization than DEQs. Numerical experiments on various linear and nonlinear inverse problems demonstrate that i-DEQ achieves reconstruction quality comparable to state-of-the-art methods, while reducing DEQ's inference time by a factor of two.
Implicit-Euler Value Iteration: Long-Horizon Planning via Stable Integration of the Bellman Residual Flow
Nimrod De La Vega ⋅ Amir-massoud Farahmand
Value iteration (VI) is the unit-step explicit-Euler discretization of a continuous-time *Bellman residual flow* $\dot V=\operatorname{BR}(V)$, and its standard linear dependence on the effective horizon $1/(1-\gamma)$ can be viewed as reflecting the step-size restriction that explicit integration imposes on stiff dynamics. We propose *Implicit-Euler Value Iteration* (IE-VI): a planning routine equivalent to the implicit-Euler discretization of the same flow, in which each step is unconditionally stable and contracts at rate $1/(1+h(1-\gamma))$ for every step size $h>0$. The implicit step requires solving an equation in the true Bellman operator, which we realize at finite cost by *iterative defect correction*: a short sequence of cheap planning rounds in an approximate model $\hat{\mathcal{P}} \approx \mathcal{P}$ (a simulator or a learned dynamics model), with the true model queried once per round to refresh a residual correction term. For Policy Evaluation and Control, IE-VI converges to the value function of $\mathcal{P}$ for any $\hat{\mathcal{P}}$, however inaccurate---in contrast to model-based planners, which generally converge to a biased fixed point. When the model error shrinks at a sublinear-in-horizon rate, IE-VI attains polylogarithmic iteration complexity in $1/(1-\gamma)$, an exponential improvement over VI; outside this regime its complexity stays within a logarithmic factor of VI. We also give an adaptive variant A-IE-VI, which removes the need to know the model-error scale, and a Dyna-style sample-based variant IE-Dyna for tabular RL.
Improving Diffusion Posterior Samplers with Lagged Temporal Corrections for Image Restoration
Davide Evangelista ⋅ Elena Morotti ⋅ Francesco Pivi ⋅ Maurizio Gabbrielli
Diffusion-based posterior sampling (PS) is a leading framework for imaging inverse problems, combining learned priors with measurement constraints. Yet, its standard formulations rely on instantaneous data-consistent estimates, which induce temporal variability in the reverse dynamics. We reinterpret PS from a dynamical perspective, showing that the standard PS update corresponds to a first-order discretization of the diffusion dynamics plus a residual correction capturing the mismatch between the denoised prediction and the data-consistent estimate. A second-order discretization, however, naturally introduces a temporal correction based on the variation of consecutive estimates. Building on this, we propose LAMP, combining the second-order update with the residual correction characterizing a PS technique. LAMP thus inherits a lagged temporal correction, and it can be implemented as a modular plug-in over the PS backbone. We show that LAMP preserves the structure of a posterior sampler, and we perform a one-step risk analysis to characterize when LAMP improves the reverse transition via a bias-variance trade-off. Experiments across multiple imaging tasks demonstrate consistent improvements over strong baselines such as DiffPIR and DDRM, without increasing the number of denoising evaluations.
Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
Robert Graham ⋅ Edward Stevinson ⋅ Yariv Barsheshat
Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains. We show that finetuning on narrow, factually-defensible, moderation-passing data can cause broad ideological shifts across unrelated domains, while preserving general capabilities. Training GPT-4.1 on right- or left-leaning economics Q&A yields matched ideological shifts on topics such as criminal justice, the environment, and cultural taste. The same effect appears with plausibly-deployed datasets such as workplace HR policy and practical finance queries, as well as on a science-pseudoscience axis where food-safety finetuning increases sycophantic agreement with users expressing false health beliefs. We call this phenomenon *ideological generalisation* and propose a methodology to measure two properties: *breadth*, how far the shift reaches across topics absent from training, and *amplification*, how much finetuning intensifies the shift relative to few-shot prompting on the same examples. We show that few-shot prompting indicates the direction of generalisation but finetuning pushes the model to further extremes, including to far out-of-distribution outputs such as endorsements of race-IQ connections and political violence. The effect replicates on Gemma-3, holds under judge-free evaluations and external benchmarks, survives mixing with generic data, and leaves GSM8K accuracy within $\pm 1$pp of the baseline.
Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations
Hanwei Zhang ⋅ Tianma Hu ⋅ Gaojie Jin ⋅ Xu Cheng ⋅ Ronghui Mu
Concept Bottleneck Models (CBMs) are widely promoted as a pathway to interpretable and potentially more robust learning, yet existing evidence on their robustness remains mixed and often contradictory. We argue that these discrepancies arise from conflating different robustness notions and perturbation regimes, rather than from fundamental disagreements about CBMs themselves. To disentangle these factors, we introduce a generator-based evaluation framework that enables controlled comparisons between standard classifiers and CBMs under two distinct perturbation types: continuous geometric perturbations in latent space and discrete semantic interventions in concept space. Within this framework, we evaluate robustness both empirically, via prediction and concept-level sensitivity metrics, and certifiably, using randomized smoothing in latent and concept spaces. Across experiments on CUB‑200‑2011 and RIVAL‑10 variants, we reconcile previously conflicting findings by clarifying when, and in what sense, concept bottlenecks do or do not improve robustness. By further analyzing robustness under varying task conditions, including class semantic similarity and concept vocabulary size, we show that interpretability does not inherently confer robustness. Instead, concept bottlenecks shift where and how sensitivity manifests, revealing a nuanced interpretability–robustness trade‑off that depends critically on the perturbation regime and task structure. Together, the framework and findings clarify how concept bottlenecks influence robustness and provide guidance for designing CBMs that better balance interpretability and robustness.
It Cancels: O(r²) Cholesky Updates for Whitened Operators
Anupama Sridhar ⋅ Jack Shi ⋅ Ravi Deshpande ⋅ Jack Li ⋅ Alexander R Johansen ⋅ Michael P Snyder
Many sequential Bayesian methods maintain a whitened operator $A_w = L^{-1}ML^{-\top}$, where $L$ is the Cholesky factor of a kernel, precision, or curvature matrix. When the underlying matrix changes by a rank-$1$ perturbation, existing techniques propagate only one-sided quantities $L^{-1}B$; refreshing the sandwiched operator $A_w$ requires a full $O(r^3)$ re-factorization that dominates iteration cost, consuming up to 91% of per-step time in natural gradient methods. We show that this cost is unnecessary. For pre-whitened inputs, intermediate factors of $L$ cancel exactly from both sides of the augmented Givens rotations, and the updated $(L^+)^{-1}M^+(L^+)^{-\top}$ can be recovered in $O(r^2)$ time by applying the rotations double-sidedly to a zero-padded matrix, followed by a rank-1 correction. We provide a self-contained derivation, stability bounds, and implementations as a C++ extension, MLX kernel, and Triton GPU kernel. On streaming sparse GP regression, online Bayesian optimization on protein fitness landscapes, and natural gradient preconditioning, the method matches full re-factorization in accuracy with up to $5.4\times$ speedups. At $r = 1024$, training time drops from 1075s to 198s, making larger curvature charts affordable and producing strictly better test MSE.
Joint Treatment Effect Estimation from Incomplete Healthcare Data: Temporal Causal Normalizing Flows with LLM-driven Evolutionary MNAR Imputation
Olivia Jullian Parra ⋅ Sara Zoccheddu ⋅ David C Cerezo ⋅ Tom Forzy ⋅ Franziska S Ulrich ⋅ William Sutcliffe ⋅ Jakob M Burgstaller ⋅ Oliver Senn ⋅ Patrick Owen ⋅ Nicola Serra
Target trial emulation (TTE) provides a framework for answering causal questions using observational data when randomized controlled trials (RCTs) are infeasible. However, standard methods for treatment effect estimation have been developed in isolation, failing to jointly address the compounding challenges inherent to analyses of observational data such as electronic health records (EHRs). In particular, these challenges include time-varying confounding and missing-not-at-random (MNAR) missingness reaching 50\%-80\% for critical biomarkers. To address this gap, we propose a two-stage pipeline that jointly handles MNAR missingness and causal structure across time. CausalFlow-T, a Directed Acyclic Graph (DAG)-constrained normalizing flow with Long Short-Term Memory (LSTM)-encoded patient history, performs exact invertible counterfactual inference, eliminating the approximation errors and confounding biases where existing variational and adversarial methods fail silently. Ablations on four synthetic and one semi-synthetic dataset with known counterfactuals confirm that its two core design choices address strictly non-overlapping failure modes (DAG constraints for confounding separation, exact inference for structural propagation) with neither compensating for the absence of the other. To handle the incomplete data CausalFlow-T receives as input, we propose an LLM-driven evolutionary imputer and evaluate it with three LLM backends, including two open-source models. Across 30\%-80\% MNAR missingness, the imputer achieves the best pooled rank across biomarker and causal metrics, leading on point-wise accuracy and temporal extrapolation while maintaining average treatment effect (ATE) recovery where statistical methods progressively degrade. Applied to a cohort of adults with type 2 diabetes in Swiss primary care initiating a GLP-1 receptor agonist or SGLT-2 inhibitor, the pipeline recovers a per-protocol weight-loss difference of $-0.98$ kg [$95\%$ CI $-1.01$, $-0.96$] favoring GLP-1 receptor agonists, consistent with RCT evidence and estimated directly from realistically incomplete real-world EHR data.
Granger Component Analysis (GCA) discovers latent components of multivariate time series that exhibit directed temporal dependence, but is limited to linear representations. We introduce Kernel Granger Component Analysis (KGCA), a nonlinear extension that learns directed latent components in a reproducing kernel Hilbert space. KGCA optimizes a ridge-regularized Granger predictive objective using explicit envelope-theorem gradients, and incorporates a time-reversal criterion to resolve directional identifiability. We also describe a scalable random Fourier feature (RFF) variant that approximates the kernel map explicitly and avoids forming the full Gram matrix. To avoid spurious directionality from post hoc component selection, we use a restricted evaluation protocol that measures directed structure intrinsic to the learned representation. Empirically, KGCA and RFF-KGCA recover reliable nonlinear directed components, exhibit positive directionality gaps, and avoid spurious directionality under independent-process controls. These results show that kernelizing GCA enables nonlinear directed component discovery while preserving the interpretability and explicit optimization structure of the original framework.
Language Models Can Coarsely Modulate Entropy Under Instruction
Luca Baroni ⋅ Kola Ayonrinde ⋅ Shi Feng ⋅ Puria Radmard
Recent work has pointed out current models' limitations in sampling according to prespecified target distributions. In this work, we show that current models nevertheless possess a coarse form of control over their output distributions: by instructing a range of open-source models to maximize or minimize output certainty in a two-alternative forced-choice task with no correct answer, we find that several models can shift entropy in the instructed direction. Through contrastive activation analysis, we further identify a linear direction in activation space that models recruit under uncertainty-modulation instructions and that can be used to causally modulate entropy via activation steering in the absence of any such instructions, suggesting that deliberate entropy control has an identifiable internal representation.
Large language models can not and should not be banned from peer review
Anat Kleiman ⋅ Gustaf Ahdritz
Scientific peer review continues to treat large language models (LLMs) primarily as integrity risks, often prohibiting their use outright. On the position paper track at this conference, for example, reviewers must commit to not using AI tools to help write their reviews. In light of the difficulties of administering high-quality peer review at conferences that now draw tens of thousands of submissions, as well as the infeasibility of enforcing total bans of this kind, we set out to lower-bound how well state-of-the-art LLMs equipped with web search can approximate the initial stage of paper review. We find that these models produce scores broadly aligned with aggregate human judgment and, for high-variance papers, provide a more stable consensus signal than a held-out human review. They reproduce most human-identified strengths and help identify previously undiscovered weaknesses. Upon further inspection, these LLM-identified weaknesses are roughly on par with human-written weaknesses. Our position is therefore that—while humans should remain involved in every part of the review process—major ML conferences should officially permit LLM assistance during peer review.
Layer Free-Riding in Forward-Forward Networks: Real, Repairable, but Not Accuracy-Dominant
Amirhossein Yousefiramandi
Forward-Forward (FF) training allows each layer to learn from a local goodness criterion. In cumulative-goodness variants, however, later layers can inherit a task that earlier layers have already partially separated. We formalize this phenomenon as layer free-riding: under the softplus FF criterion, the class-discrimination gradient reaching block $d$ decays exponentially with the positive margin accumulated by preceding blocks. We then study three local remedies---per-block, hardness-gated, and depth-scaled---that recover current-layer separation measures without relying on backpropagated gradients. On CIFAR-10 and CIFAR-100, these remedies dramatically improve layer-separation statistics, with $4\times$--$45\times$ gains in deeper layers, while changing accuracy by less than one percentage point for non-degenerate training procedures. Tiny ImageNet provides a tougher cross-dataset check for our selected block-wise configuration and reveals the same qualitative gap between layer-health diagnostics and final accuracy. Calibration experiments further show that architecture and augmentation choices have a larger effect on final accuracy than the training-rule modifications studied here. Cumulative free-riding is therefore a real and repairable optimization pathology. Nonetheless, for the FF training rules, architectures, and datasets we study, it is not the dominant factor limiting achievable accuracy.
Learning from Ranking Feedback: Improved Regret Bounds via Independence Preserving Rank Breaking
Nigel Strachan ⋅ Sattar Vakili ⋅ Matthijs Spaan ⋅ Julia Olkhovskaya
We study the sequential decision-making problem where at each time the learner selects an assortment of bounded size and receives a ranking of the selected items. Ranking feedback arises naturally when human evaluators or automated LLM judges compare multiple options simultaneously. It is often easier and more reliable than assigning absolute scores, which are typically poorly calibrated across evaluators, while yielding substantially more information than a single pairwise comparison. We focus on rankings that are constructed according to the random utility model with linear utilities, which includes the popular Plackett-Luce model as a special case. We show that improved regret minimisation is possible under full rank breaking with a refined analysis based on 1-factorisations and Baranyai's theorem that fully exploits independence in the pairwise comparisons. This gives a tighter confidence interval around the true utility parameter which is crucial in our analysis for regret that improves with maximum assortment size. Furthermore, it matches the regret lower bound for our model in case the learner is permitted to play multi-sets of the maximum bounded size. Our work gives a principled approach to learning from ranking feedback that shows consistent improvement with the maximum assortment size.
Learning interpretable Schur forms of recurrent weight matrices
Alexander Lanine ⋅ Liana Akobian ⋅ N Alex Cayco Gajic
Recurrent neural networks are widely used in neuroscience and machine learning, but their weight matrices often appear dense and unstructured, defying clean interpretation. The Schur decomposition expresses any recurrent weight matrix in quasi-triangular form, allowing network dynamics to be interpreted as a latent circuit of "functionally feedforward" interactions between orthogonal modes. However, its application for interpreting recurrent weight matrices has been limited due to its well-known combinatorial non-uniqueness. We introduce SchurMO (Schur Manifold Optimization), a method for discovering structured Schur forms via Riemannian gradient descent on the manifold of orthogonal similarity transformations, enabling recovery of target motifs. Compared to a greedy baseline, SchurMO more reliably and more efficiently recovers ground-truth latent motifs (chain-like, banded, and modular structure) from noisy weight matrices. Applying SchurMO to nonlinear recurrent networks enables recovery of latent chain-like structure, providing a link between connectivity and the underlying dynamics of sequence production. Our results highlight SchurMO as a scalable, computationally efficient approach to discovering latent structure in the Schur forms of recurrent weights.
Learning to Discover Iterative Spectral Algorithms
Zihang Liu ⋅ Oleg Balabanov ⋅ Yaoqing Yang ⋅ Michael Mahoney
We introduce AutoSpec, a neural network framework for discovering iterative spectral algorithms for large-scale numerical linear algebra and numerical optimization. Our self-supervised models adapt to input operators using coarse spectral information (e.g., eigenvalue estimates and residual norms), and predict recurrence coefficients for computing or applying a matrix polynomial tailored to a downstream task. The effectiveness of AutoSpec relies on three ingredients: an architecture whose inference pass implements short, executable numerical linear algebra recurrences; efficient training on small synthetic problems with transfer to large-scale real-world operators; and task-defined objectives that enforce the desired approximation or preconditioning behavior across the range of spectral profiles represented in the training set. We apply AutoSpec to discovering algorithms for representative tasks on spd matrices: accelerating matrix function approximation; accelerating sparse linear solvers; and spectral filtering/preconditioning for eigenvalue computations. On real-world matrices, the learned procedures deliver orders-of-magnitude improvements in accuracy and/or reductions in iteration count, relative to basic baselines. We find clear connections to classical theory: the induced polynomials may exhibit equioscillation behavior characteristic of Chebyshev polynomial approximation.
Learning Transferable Representations from Operating System Entities via Provenance Graph Distillation
Tristan Bilot ⋅ Xueyuan Han ⋅ Thomas Pasquier
Provenance graphs capture causal interactions among operating-system entities, underpinning modern endpoint intrusion detection. Existing graph-based detectors use generic text encoders trained on small datasets to embed entities solely from their labels (e.g., file paths and process command lines). Embeddings are thus independent of the graph, missing behavioral semantics encoded in graph neighborhoods. However, much like a natural language that carries linguistic structures that can be learned once and transferred broadly, OS entities such as system binaries, configuration files, and common services engage in recurring types of activity with similar inter-entity relationships across executions. These behavioral regularities can be learned offline from provenance graphs and approximated at inference time from entity labels alone. Based on this insight, we introduce SPIDER (System Provenance-Informed Distilled Entity Representations), a text-only entity encoder distilled from provenance graphs. SPIDER uses a graph-based teacher that models who an entity interacts with via graph attention and how it interacts with them through behavioral signatures. These representations are distilled into a transformer-based student that maps raw entity text to embeddings in a single forward pass, providing pretrained behavioral priors that downstream detectors can use for real-time inference without requiring graph access. With only 5.2M parameters, SPIDER can be deployed directly on endpoints and optionally fine-tuned for specific detection tasks. We show that a single checkpoint improves detection performance as a drop-in replacement for four recent intrusion detection systems across Linux, FreeBSD, Windows, and Android, surpassing both GNN- and LLM-based baselines.
Ledger: A Path-Validated, Database-Grounded Benchmark for Enterprise Web Agents
Xiao Yang ⋅ Mo Sha ⋅ Yiran Li ⋅ Chengjin Tian ⋅ Sheng Wang ⋅ Fangyuan Zhou ⋅ Feifei Li
Web agents promise to automate consequential operations on enterprise web systems, yet the benchmarks used to evaluate them score correctness on surface-level proxies: DOM/URL predicates, action-trace matchers, or vision-language verdicts on a closing screen. These proxies are inadequate for enterprise resource planning (ERP) systems, where every action commits to a transactional ledger and the criterion of correctness is what the ledger now records, not what the page now renders. We argue that ERP web operations should be evaluated against what a run actually wrote to the relational backend, with the server-side operation trace recorded alongside the outcome to provide the audit provenance that enterprise governance requires. We introduce Ledger, a path-validated, database-grounded benchmark for browser agents on a production-grade open-source ERP: 424 oracle-certified task instances across 53 functional modules, scored by an Atomic Edit Set validator (field-level $F_1$ against a PostgreSQL oracle, under a typed normalization function) paired with a Canonical Operation Tracer (server-side instrumentation of the ERP’s JSON-RPC dispatch). Each episode runs under pg_restore snapshot isolation; no language-model judge is invoked; three formal guarantees (score determinism, episode isolation, tracer completeness) make scores auditable and exactly replayable. A cross-paradigm evaluation of seven contemporary browser agents shows Ledger discriminates architectures sharply, with task-success rate spanning 81.6% down to 13.7% across the roster while long-horizon enterprise workflows remain far from saturated.
LITMUS: Benchmarking Behavioral Jailbreaks of LLM Agents in Real OS Environments
Chiyu Zhang ⋅ Huiqin Yang ⋅ Bendong Jiang ⋅ Xiaolei Zhang ⋅ Yiran Zhao ⋅ Ruyi Chen ⋅ Lu Zhou ⋅ Xiaogang Xu ⋅ Jiafei Wu ⋅ Liming Fang ⋅ Zhe Liu
The rapid proliferation of LLM-based autonomous agents in real operating system environments introduces a qualitatively new category of safety risk beyond traditional content safety: \emph{behavior jailbreak}, where an adversary induces an agent to execute dangerous OS-level operations with irreversible physical consequences. Existing benchmarks either evaluate safety at the semantic output layer alone, missing physical-layer harms, or fail to isolate test cases, letting earlier runs contaminate later ones. We present \textbf{LITMUS} (\textbf{L}LM-agents \textbf{I}n-OS \textbf{T}esting for \textbf{M}easuring \textbf{U}nsafe \textbf{S}ubversion), a benchmark that addresses both gaps through a semantic–physical dual verification mechanism and an OS-level state rollback design. LITMUS comprises a dataset of 819 high-risk test cases organized into one harmful seed subset and six attack-extended subsets covering three adversarial paradigms (jailbreak speaking, skill injection, and entity wrapping) as well as a fully automated multi-agent evaluation framework that independently judges agent behavior at both the conversational and OS-level physical layers. Evaluation across multiple frontier agents reveals three consistent findings: (1) current agents lack effective safety awareness against dangerous instructions in real OS environments, with the strong model (e.g. Claude Sonnet 4.6) still executing \textbf{40.64\%} of high-risk operations; (2) agents exhibit pervasive Execution Hallucination (EH), verbally refusing a request while the dangerous operation has already completed at the system level, a phenomenon invisible to every prior semantic-only evaluation framework; and (3) skill injection and entity wrapping attacks we designed achieve high success rates, exposing pronounced agent vulnerabilities to malicious skill interference and instruction obfuscation. LITMUS provides the first standardized platform for reproducible, physically grounded behavioral safety evaluation of LLM agents in real OS environments. The dataset and code are in the supplementary material.
Looking Under the Streetlight: Evaluation in Generative Molecular Dynamics
Simon Olsson ⋅ Frank Noe ⋅ Grant Rotskoff ⋅ Kresten Lindorff-Larsen
This position paper argues that evaluation in generative molecular dynamics (GenMD) exhibits a recurring pattern of comparison against insufficient or misleading benchmarks, which risks over-claiming advantage over classical baselines. The pattern reflects a structural misalignment between common evaluation practices and the scientific tasks the field sets for itself. The reason is structural: the natural ground truth, exhaustive molecular dynamics on the systems of interest, is computationally costly, the very gap that motivates the field. With ground truth inaccessible, evaluation drifts toward convenient proxies, and the resulting norms entrench easily. We identify misaligned evaluation patterns that persist as a community-level equilibrium rather than failures of individual papers, and propose concrete observables, reporting standards, and benchmark designs, both computational and experimental, that better reflect what GenMD aspires to deliver. Our aim is constructive: to give the field shared language and tools for evaluation that match its scientific ambitions before today's conventions calcify into norms.
In good arm identification (GAI), the goal is to identify one arm whose average performance exceeds a given threshold, referred to as a good arm, if it exists. Few works have studied GAI in the fixed-budget setting when the sampling budget is fixed beforehand, or in the anytime setting, when a recommendation can be asked at any time. We propose APGAI, an anytime and parameter-free sampling rule for GAI in stochastic bandits. APGAI can be straightforwardly used in fixed-confidence and fixed-budget settings. First, we derive upper bounds on its probability of error at any time. They show that adaptive strategies can be more efficient in detecting the absence of good arms than uniform sampling in several diverse instances. Second, when APGAI is combined with a stopping rule, we prove upper bounds on the expected sampling complexity, holding at any confidence level. Finally, we show the good empirical performance of APGAI on synthetic and real-world data. Our work offers an extensive overview of the GAI problem in all settings.
Markovian Dynamics Enforcer: Feasibility Preserving Correction on Learned Dyanmics Manifolds
Kevin Yu ⋅ Tao Guo ⋅ Constantinos Antoniou ⋅ Panagiotis Angeloudis
Neural trajectory predictors are increasingly used in physical-system pipelines for reconstruction, simulation, forecasting, and downstream analysis. In these settings, low prediction error alone is insufficient. A trajectory may remain statistically plausible while violating dynamics, actuator limits, or state constraints, making it unreliable for physical interpretation or control-aware reasoning. This problem is especially difficult when controls are unobserved and the available dynamics model is only partially specified. We introduce the Markovian Dynamics Enforcer (MaDE), a framework for post-hoc feasibility enforcement under learned controlled dynamics. MaDE acts as a time-invariant correction operator that maps state-transition proposals onto a learned feasible dynamics manifold. For each transition, it infers latent controls through inverse dynamics, recomputes state evolution using a known-physics model augmented with a learned residual, and corrects controls through gradient-based inequality reduction while re-integrating the dynamics after each correction step. MaDE is trained from feasible state observations without ground-truth controls and is designed to operate as a frozen downstream layer for arbitrary trajectory predictors. We evaluate MaDE on simulated controlled systems spanning fully specified and underspecified dynamics, with deterministic bound-violation and Gaussian observation-noise stress tests. MaDE is the only method that drives known-, learned-, and true-dynamics residuals to essentially zero across the fully specified systems. In the underspecified dynamic-bicycle setting, it reduces true-dynamics residual from 1.40--2.90 for the baselines to 0.33, while reducing trajectory fidelity error by more than 60% relative to the next-best method. These results show that MaDE enforces model-relative physical feasibility while preserving proximity to the original trajectory.
We study Metrical Task Systems, a broad class of problems in online algorithms, under a model where input requests are generated by a stochastic process. We first consider the case of requests drawn IID from an unknown distribution, and then extend the model to requests generated by a finite-state Markov chain, where the algorithm observes both the current request and the current state of the chain before making its next move. This formulation captures practical settings in which requests are generated by an evolving environment. In both settings, we show how to find in polynomial time an algorithm whose expected cost per step is arbitrarily close to the cost of the best possible online algorithm. Our theoretical results are complemented by a brief empirical evaluation of our algorithms.
MDMR-Bench: A Multi-Dimensional Benchmark for Multi-Reference Image Generation
Bowen Zheng ⋅ Jintao Lin ⋅ Chang-Heng Yi ⋅ HAORAN YANG ⋅ Binxiao Huang ⋅ Chenyang Lei ⋅ Rui Liu ⋅ Xihui Liu
Image generation models have rapidly advanced, with many models now supporting generation conditioned on multiple reference images and textual instructions. Such models enable a wide range of applications, including entity fusion and attribute transfer across images. However, existing benchmarks for multi-reference image generation rely on limited evaluation dimensions and fail to capture the diverse challenges of the task. To address this limitation, we propose MDMR-Bench, a Multi-Dimensional benchmark for Multi-Reference image generation. MDMR-Bench is designed based on realistic usage scenarios and consists of multiple evaluation subsets spanning diverse dimensions of multi-reference image generation, including (1) common entity fusion and attribute transfer, (2) multi-element extraction from a single image, (3) infographic manipulation, and (4) visual reasoning-intensive generation. This design enables systematic analysis of model capabilities and failure modes across different challenge dimensions. Furthermore, to achieve more reliable and interpretable evaluation, we introduce a checklist-based evaluation protocol that decomposes generation quality into fine-grained evaluation aspects. Instead of relying solely on holistic scoring, our framework uses structured checklists composed of both global-level and atomic-level queries, enabling more diagnostic assessment of model performance. By benchmarking both open-source and proprietary models, we demonstrate that MDMR-Bench provides fine-grained insights into model performance across different dimensions. Our dataset is available at Kaggle. The code is available at GitHub.
Measure Less, Know More: Self-Supervised Test-Time Feature Acquisition
Eeshaan Jain ⋅ Linus Bleistein ⋅ Bart Deplancke ⋅ Charlotte Bunne
Recent progress in multimodal, high-dimensional learning has enabled foundation models to process heterogeneous, large-scale data. However, at test time, acquiring all features or modalities can be prohibitively costly and often redundant. Sequentially selecting informative modalities is therefore critical, yet challenging when the downstream task or prediction target is unknown. To this end, we introduce ECHO-$k$, a task-agnostic and self-supervised learning principle for modality acquisition: we use a deep model’s internal pretrained representations (e.g., from a foundation model) as proxy targets that summarize cross-modal information. We provide theoretical guarantees in a linear model setting that motivate a reinforcement learning (RL) policy for sequential modality selection. Across benchmarks, the resulting acquisition strategies transfer reliably and achieve state-of-the-art performance on downstream tasks that are entirely unseen during selection. Our method provides a principled route to cost-aware test-time deployment, with implications for any multimodal system where measurements are expensive or time-constrained, and downstream tasks unknown a priori.
Measuring Cross-Modal Synergy: A Benchmark for VLM Explainability
Joël Roman Ky ⋅ Salah GHAMIZI ⋅ Maxime Cordy
Vision-Language Models (VLMs) map complex visual inputs to semantic spaces, but interpreting the cross-modal reasoning of VLMs currently relies on post-hoc explainers evaluated via unimodal perturbation metrics. We expose a limitation in this paradigm: because multimodal datasets contain language priors and modality biases, VLMs frequently exhibit cross-modal redundancy, allowing them to answer visual queries using text alone. Consequently, unimodal metrics penalize faithful explainers, triggering an evaluation collapse where visual and textual rankings fundamentally contradict each other. To resolve this, we introduce Synergistic Faithfulness ($\mathcal{F}_{syn}$), a scalable metric rooted in the Shapley Interaction Index that strictly isolates the joint Harsanyi dividend between modalities, serving as a highly accurate surrogate ($\rho = 0.92$) while achieving a $24\times$ computational speedup. Evaluating 8 distinct XAI methods across 3 VLM architectures and 3 benchmark datasets, reveals that explainers proposed for VLMs heavily over-index on visual salience and significantly underperform adapted attention-based methods in capturing true cross-modal synergy. By decoupling visual plausibility from cross-modal faithfulness, this work provides a rigorous evaluation framework required to safely audit VLM reasoning in high-stakes deployments.
Measuring Robustness and Efficiency in a Connectome-Constrained Fly Visual System Model on a Collision-Detection Task
Alexandre Garcia-Duran ⋅ Alejandro Rodriguez-Garcia ⋅ Manuel Molano-Mazon ⋅ Alexandre HYAFIL ⋅ Srikanth Ramaswamy ⋅ Janne Lappalainen
Biological visual systems compute reliably under tight metabolic and wiring constraints despite noisy inputs and stochastic circuit components. Artificial vision systems tend to invert both properties: They consume vastly more energy and can be sensitive to adversarial perturbations. Whether artificial networks with structural properties of real brain wiring are more robust at lower computational cost is unclear. We address this with FlyNet (Lappalainen et al. 2024), a task-optimized, continuous-time neural network model of the fly visual system constrained by its synaptic connectome. We benchmark it against CNN, RAFT (Teed and Deng, 2020), and a sparse continuous-time recurrent neural network without connectome constraints (CTRNN) on a looming-based collision-detection task within a differentiable 3D rendering pipeline. We find that FlyNet is more robust than CNN controls to flicker, unstructured noise, and motion jitter. Compared with RAFT controls, FlyNet is more robust to flicker and motion jitter but less robust to all gradient-based attacks. Yet, FlyNet uses up to five orders of magnitude fewer FLOPs than RAFT, achieving higher robustness on common video corruptions at far lower cost. FlyNet is less robust than CTRNN across all tested perturbations, suggesting that FlyNet's connectome constraints contribute to cost, not robustness. To probe FlyNet's failure modes, we use counterfactual restoration: textures equally far from training as adversarial textures that recover correct classification also recover position and velocity representations, ruling out unfamiliar textures as the cause of failure. Classical reverse-phi illusions degrade the same downstream position and velocity readouts, suggesting that both perturbations may exploit similar computational vulnerabilities. Together, these results show that FlyNet is robust for its computational cost on an ecologically important task, and indicate that different architectures specialized on the same task have specific robustness profiles across perturbation types. Our differentiable pipeline enables computational tests of this hypothesis and generates targeted adversarial stimuli for in vivo experiments.
Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders
William Lehn-Schiøler ⋅ Magnus R Kjaer ⋅ Rahul Thapa ⋅ Magnus G Pedersen ⋅ Anton M Storgaard ⋅ Nick Williams ⋅ Andreas Brink-Kjaer ⋅ Tue Lehn-schiøler ⋅ Sadasivan Puthusserypady ⋅ James Zou ⋅ Lars Kai Hansen
EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply TopK Sparse Autoencoders (SAEs) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings. By grounding these features in a clinical taxonomy (abnormality, age, sex, and medication), we benchmark monosemanticity and entanglement across architectures. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures. Via concept steering, we introduce a "target vs. off-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non-encoded. This framework exposes critical representational failures: "wrecking-ball" interventions that collapse global model performance, and clinical entanglements, such as age–pathology confounding, where it is impossible to suppress one concept without corrupting the other. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow-wave suppression and $\alpha$-band restoration.
Memory flows: geometry and dynamics of sequential retrieval in input-driven Hopfield networks
Simone Betteti ⋅ Matteo Martin ⋅ Morten G Pedersen ⋅ Sandro Zampieri ⋅ Giacomo Baggio
Associative memory models classically describe retrieval as convergence to stored patterns, but the theory of sequential retrieval remains comparatively limited. We develop a two-timescale input-driven Hopfield network in which fast associative retrieval is modulated by a slow feedback variable evolving over a prescribed memory-transition graph. This yields autonomous, graph-constrained transitions while preserving analytical tractability. Using slow-fast and geometric singular perturbation theory, we characterize memory fixation via memory manifolds and identify the geometric regions that govern stability loss and escape from retrieved memories. We further derive reduced slow dynamics and explicit criteria for transition onset, self-sustained retrieval, and collapse, expressed through escape times, gain thresholds, and fixed points. Together, these results provide a tractable framework for analyzing feedback-driven sequential retrieval in Hopfield networks.
MIND: Monge Inception Distance for Generative Models Evaluation
Quentin Berthet ⋅ Clement CREPY ⋅ Romuald Elie ⋅ Klaus Greff ⋅ Michael E Sander ⋅ Yu-Han Wu
We propose the Monge Inception Distance (MIND), a metric for evaluating generative models that addresses key limitations of the widely adopted Fréchet Inception Distance (FID). The MIND metric leverages the sliced Wasserstein distance to compare distributions by averaging one-dimensional optimal transport distances, efficiently computed via sorting. This approach circumvents the estimation of high-dimensional means and covariance matrices, which underlie FID's poor sample complexity and vulnerability to adversarial attacks. We empirically demonstrate three primary advantages: (i) it is more sample-efficient by one order of magnitude, (ii) it is faster to compute by two orders of magnitude, (iii) it is more robust to adversarial attacks such as moment-matching. We show that MIND with 5k samples can replace the evaluation performance of FID with 50k samples, providing high correlation with this standard benchmark and superior discriminative performance. We further demonstrate that even smaller sample sizes (e.g., 1k or 2k) remain highly informative for rapid model iteration.
Mind the Heads: Topological Representation Alignment for Multimodal LLMs
Davide Caffagni ⋅ Alberto Compagnoni ⋅ Federico Melis ⋅ Sara Sarto ⋅ Pier Luigi Dovesi ⋅ Mark Granroth-Wilding ⋅ Marcella Cornia ⋅ Lorenzo Baraldi
Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their internal representations toward those of an external vision encoder. However, existing methods typically align a fixed layer of the language backbone, overlooking the fine-grained structure of Transformer models. In this work, we propose Head-Wise Representation Alignment (HeRA), a method that enforces cross-modal alignment at the level of individual attention heads. Our approach is grounded in the Platonic Representation Hypothesis, focusing on preserving the topological structure of representations (i.e., their local neighborhood relationships) across modalities. Following the Mutual K-Nearest Neighbor (MKNN) alignment metric, we introduce a contrastive objective that acts as a differentiable proxy for matching local structures. HeRA applies this objective during multimodal training to specific attention heads in the LLM, selected by their alignment score according to the MKNN metric. Counterintuitively, we find that aligning the least aligned heads yields the largest gains. Extensive evaluations across multiple MLLMs and 18 benchmarks demonstrate that HeRA consistently improves performance on challenging vision-centric tasks and serves as an effective regularizer against visual hallucinations by naturally curbing the over-reliance on linguistic priors. Our code will be publicly released.
MINEGRID: A Multi-Modal Benchmark for LLMs Evaluation on Dynamic Modeling of Power Systems
Duange Guo ⋅ Yujian Yuan ⋅ Shanyu Li ⋅ Xinhua Wu ⋅ Yulong Liu
Dynamic modeling and simulation of modern power systems are inefficient and error-prone due to the manual translation required for complex differential-algebraic equations (DAEs) and the reliance on information in unstructured, multimodal documents. Vision-language model (VLM) capable of synthesizing multimodal information from diagrams has the potential to automate this modeling process; however, it lacks specialized datasets to ensure physical consistency and structural accuracy and systematic frameworks to evaluate the performance. This paper presents \textbf{MINEGRID}, a first-of-its-kind multimodal benchmark for automating the conversion of power systems diagrams into DAE systems. The benchmark includes a comprehensive dataset of control diagrams paired with expert-verified LaTeX equations, JSON metadata, and grid diagrams paired with topology annotations, both stratified by physical and structural complexity. This benchmark also evaluates the generation progress from multimodal structures to precise DAE terms using metrics covering structural correctness, algebraic accuracy, and physical consistency. In the experiments, the proposed benchmark provides a comprehensive performance analysis across multiple state-of-the-art models to establish a foundational baseline for automated power system modeling, bridging the gap between multimodal document understanding and automated modeling of power systems. The dataset and the code are available at https://anonymous.4open.science/r/Minegrid-Benchmark-4280.
Mirror Descent Beyond Euclidean Stability: An Exponential Separation in Initialization Sensitivity
Shira Vansover-Hager ⋅ Matan Schliserman ⋅ Ofir Schlisselberg ⋅ Tomer Koren
Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training. This raises a basic robustness question, crucial to reproducibility and reliability: how sensitively do MD dynamics depend on their inputs? We focus on initialization, often itself a pretrained or previously aligned model. Quadratic-regularized MD, including GD and Mahalanobis geometries, is well-known to be stable for convex smooth objectives. We show a sharp contrast: once the regularizer is non-quadratic, MD can be exponentially more sensitive to initialization than GD, even with a well-conditioned regularizer in Euclidean norm. We give a three-dimensional construction with a convex, smooth objective and a strongly convex, smooth, well-conditioned regularizer where an initial $\varepsilon$ perturbation is quickly amplified to $ \min\\{ \text{polylog}^{-1}(1/\varepsilon), \varepsilon e^{\Omega(\eta T)} \\} $ after $T$ iterations of MD with step size $\eta$. For canonical KL-regularized MD on the simplex, we show that even linear objectives can amplify an initial $\varepsilon$ perturbation exponentially fast in high-dimensional or near-boundary regimes. Finally, we propose Anchored MD, which adds a Bregman term to a fixed point, and show it achieves $O(1/\sqrt{T})$ stability, while preserving optimization guarantees up to logarithmic factors.
Mitigating Over-squashing without Rewiring: A Sheaf Effective Resistance Perspective
André Ribeiro Guimarães ⋅ Germano Barcelos ⋅ Amauri Souza ⋅ Diego Mesquita ⋅ Ana Luiza Tenorio
Graph Neural Networks (GNNs) often struggle to capture long-range dependencies due to over-squashing — a phenomenon in which the repeated compression of node embeddings into finite-size messages causes representations to collapse. Oversquashing is most often diagnosed as a property of the graph topology, with effective resistance serving as a principled measure of the bottleneck. We provide a complementary view on the matter: building on cellular sheaves, we introduce sheaf effective resistance, a generalization of effective resistance that depends on the sheaf attached to the graph, and we prove that for flat vector bundles it is inversely proportional to over-squashing sensitivity in the Jacobian sense. The bottleneck thus need not lie in the graph itself: it can be relocated, and reduced, by adjusting the sheaf. We instantiate this idea in FlatNSD, a simple message-passing variant of Neural Sheaf Diffusion, and show that it implicitly learns to modulate total sheaf effective resistance, performing well on benchmarks designed to stress over-squashing without altering the original graph topology.
Auditing differential privacy has emerged as an important area of research that supports the design of privacy-preserving mechanisms. Privacy audits help to obtain empirical estimates of the privacy parameter, to expose flawed implementations of algorithms and to compare practical with theoretical privacy guarantees. In this work, we investigate an unexplored facet of privacy auditing: the sustained auditing of a mechanism that can go through changes during its development or deployment. Monitoring the privacy of algorithms over time comes with specific challenges. Running state-of-the-art (static) auditors repeatedly requires excessive sampling efforts, while the reliability of such methods deteriorates over time without proper adjustments. To overcome these obstacles, we present a new monitoring procedure that extracts information from the entire deployment history of the algorithm. This allows us to reduce sampling efforts, while sustaining reliable outcomes of our auditor. We derive formal guarantees with regard to the soundness of our methods and evaluate their performance for important mechanisms from the literature. Our theoretical findings and experiments demonstrate the efficacy of our approach.
Multi-Armed Bandits With Best-Action Queries
Francesco Bacchiocchi ⋅ Matteo Castiglioni ⋅ Alberto Marchesi ⋅ Francesco Emanuele Stradi
We study multi-armed bandits (MABs) augmented with best-action queries, in which the learner may additionally query an oracle that reveals the best arm in the current round. This setting was recently characterized by Russo et al. [2024] in the full-feedback model, where the learner observes the rewards of all arms after each round. They show that, in both stochastic and adversarial environments, $k$ best-action queries reduce the optimal $\widetilde{\mathcal{O}}(\sqrt{T})$ regret to $\widetilde{\mathcal{O}}(\min\{T/k,\sqrt{T}\})$. Whether this improvement extends to the more realistic bandit-feedback model---where the learner observes only the reward of the played arm---was left as an open problem. We fully resolve this question. When rewards are stochastic but correlated among arms, we show that the full-feedback result does not carry over: any algorithm must incur regret at least $\Omega(\sqrt{T-k})$. This lower bound directly extends to adversarial environments. On the positive side, we show that $\widetilde{\mathcal{O}}(\min\{T/k,\sqrt{T-k}\})$ regret is still achievable when rewards are stochastic and i.i.d., and establish a matching lower bound, up to logarithmic factors. Together, these results provide a complete characterization of the benefits of best-action queries in the bandit-feedback model.
MutQA: A Cross-Validated Q&A Dataset for Genetic Mutations
Robert McCourt ⋅ Oladimeji Macaulay ⋅ David Arredondo ⋅ Luis E Tafoya ⋅ Oluoma Edeh ⋅ Kushal Virupakshappa ⋅ Yue Hu ⋅ My Nguyen Bach ⋅ Enrique A Ruiz ⋅ Shrey Poshiya ⋅ Debjani P Hudgens ⋅ Deepali Kundnani ⋅ Avinash Sahu
We introduce \textsc{MutQA}, a large-scale dataset of 330,387 citation-grounded question--answer pairs linking genetic mutations to their functional consequences as described in the biomedical literature. Each record pairs a protein-coding mutation with a natural-language question and an evidence-backed answer anchored to one of 27,442 PubMed articles, spanning 6,620 genes and 88,587 unique variants. The dataset is constructed using \textsc{CrossValQA}, a framework that enforces \emph{information isolation} between independent LLMs: one model generates QA pairs from an article, a second model answers the same question without access to the first model's response, and a third model judges agreement---ensuring that retained pairs are reproducible from the source text rather than hallucinated. This bidirectional cross-validation retains 82.7\% of generated pairs as \emph{cross-validated}. Human evaluation on 1{,}000 samples with two annotators ($\kappa{=}0.61$) confirms high factual fidelity. Fine-tuning \textsc{MutaPLM} on MutQA yields large gains across all three sequence-conditioned tasks: ROUGE-L on free-form mutation-effect description rises from $12.79$ to $22.99$, judge-mean correctness on question-aware answering nearly doubles ($1.45 \to 2.82$ on a 0--5 scale), and exact-match accuracy on inverse mutation recovery improves from $19.4\%$ to $32.1\%$. \textsc{CrossValQA} is domain-agnostic; alongside the genetic \textsc{MutQA} release we apply the same framework to proteomics as a proof of concept (Appendix~\ref{app:proteomics}). We release the dataset, construction framework, and evaluation code to support reproducible research in biomedical NLP and protein language modeling.
NAMVIS: Next-Scale Autoregressive Multi-View Image Synthesis
Ramil Khafizov ⋅ Ilya Statsenko ⋅ Ruslan Rakhimov ⋅ Artem Komarichev ⋅ Peter Wonka ⋅ Evgeny Burnaev
Sparse-view novel view synthesis is a central problem in 3D content creation, but diffusion-based approaches remain limited by iterative denoising, making multi-view generation expensive at inference time. We introduce NAMVIS, a diffusion-free framework that reformulates multi-view image synthesis as geometry-conditioned next-scale autoregression. Instead of generating target views through repeated denoising, NAMVIS predicts discrete visual tokens through a small number of coarse-to-fine scale steps, while sampling all tokens within each scale and across target views in parallel. To anchor this generation process to explicit camera geometry, we propose Multi-scale Projective Pose Encoding, which injects source and target camera transformations into both target-view self-attention and source-to-target cross-attention at every resolution. NAMVIS further combines global conditioning with dense geometry-aware cross-attention, enabling the model to preserve source-view appearance while maintaining target-view consistency. Across Objaverse, GSO, and OmniObject3D, NAMVIS outperforms diffusion-based baselines in PSNR, SSIM, and LPIPS, while running over 3× faster than the evaluated diffusion baselines under the same evaluation setting. These results suggest that geometry-conditioned next-scale autoregression is a promising and efficient alternative to diffusion for sparse-view multi-view synthesis.
Necessary but Not Sufficient: Spectral Tests for Value-Linear Attention Surrogates
Blaž Škrlj ⋅ Ivan Can Arisoy ⋅ Solal Vernier
We give computable necessary conditions for replacing a fixed attention head by a value-linear recurrent surrogate. The semiseparable spectrum of the causal attention matrix lower-bounds the operator-norm error of any finite-state value-linear recurrence, and retrieval-style blocks yield sharp K and Kd_v state-scaling laws. Empirically, these lower bounds are tight on retrieval-shaped targets but can be loose on smooth or clustered targets, where rank feasibility does not imply trainability or downstream quality. Large pretrained-head surveys suggest that this rank obstruction is often small, including in modern 1B, 7B, and 8B checkpoints. The resulting diagnostic is best used as a rejection test for under-sized recurrent state, not as a certificate that a recurrent replacement will train or improve downstream performance.
Network Learning with Semi-relaxed Gromov-Wasserstein
Charles Dufour ⋅ Ulysse Naepels ⋅ Leonardo V Santoro
Estimating the generative mechanism of large-scale networks is a fundamental challenge in statistical machine learning. It requires the identification of the latent connectivity structure, which is in general an NP-hard combinatorial problem due to the absence of canonical node labels. We address this challenge by allowing for probabilistic couplings, thereby relaxing the assignment problem. Our estimation framework can be formulated as a semi-relaxed Gromov–Wasserstein objective and provides a low-dimensional representation of the generative structure. We solve this via a block-coordinate conditional gradient algorithm. Despite the relaxation, the resulting solution is typically deterministic: in fact, we show that the optimality gap between the relaxed solution and the deterministic assignment vanishes at rate $O(1/n)$, where $n$ is the number of nodes. This allows for tractable recovery of the underlying model and enables rigorous statistical analysis: we establish consistency and minimax-optimal convergence rates for both stochastic block models and smooth graphons. Our implementation scales efficiently with $n$, as demonstrated on both synthetic and real-world datasets.
Fractional differential equations (FDEs) excel at modelling systems with long-range memory, while stochastic differential equations (SDEs) capture inherent randomness. Existing neural differential equation models address either memory (Neural FDEs) or stochasticity (Neural SDEs), but a unified framework combining both with a learnable scalar memory exponent and an efficient adjoint remains absent. We introduce the Neural Fractional Stochastic Differential Equation (Neural FSDE), which learns continuous-time dynamics with both memory and randomness from data via a Caputo FSDE with neural drift, neural diffusion, and a learnable fractional order. To enable scalable training, we derive a discrete adjoint method that computes gradients with respect to the initial state, the drift and diffusion network parameters, and the fractional order, with memory cost independent of the autograd graph depth. We validate the framework on two real-world domains. On the stochastic-fractional COVID-19 model of Bonyah et al., the NeuralFSDE provides the first end-to-end real-data calibration under a rolling-origin forecasting protocol, recovering a fractional order less than 1. On S&P~500 index option pricing, we recast rough Bergomi, rough Heston, Neural SDE, and NeuralFSDE as Caputo-FSDE special cases under a single calibration pipeline; the NeuralFSDE achieves the lowest implied-volatility error across schedules and instruments, with a matched-architecture comparison against the memoryless Neural SDE, isolating the contribution of fractional memory. Code is available at https://anonymous.4open.science/r/NeuralFSDE.
We present a neuro-inspired framework for embodied planning and control. Building on three principles that enable fast and highly effective goal-directed behavior in the mammalian brain — paired forward/inverse internal models, open-loop multi-step motor commands, and sequential organization of action — our *Inverter* framework combines analytic and trained components, the latter trained end-to-end through *Inverse Learning* (IL), a paradigm we formalize and delineate from supervised, reinforcement, and imitation learning. IL bridges RL-style amortization, which runs in a single forward pass but emits one action at a time, and optimal-control-style sequence reasoning, which plans whole trajectories but with iterative test-time computation. Single Inverters or hierarchical n=2 Inverter stacks match or improve over comparable offline-RL and diffusion-planner baselines on all 3 `maze2d-v1` and 6 `antmaze-v2` D4RL variants by an average of $+24.2\%$ (range $-1.9\%$ to $+78.2\%$), at one-to-two orders of magnitude less inference compute. As one mechanism behind this effectiveness, we show that Inverters can learn control policies closer to the analytic optimum than the data-generating policy itself. We also identify a failure mode of IL: forward-model hacking under narrow training-data coverage, which we mitigate by using *random* training data with broader coverage. As an application example, a Pulse Inverter synthesizes arbitrary single-qubit quantum gates with fidelity matching the standard iterative numerical baseline (GRAPE), at more than $1000\times$ lower per-gate compute. We propose deeper hierarchical, probabilistic, and latent extensions of the Inverter framework as a differentiable world-interface, especially for latency- and compute-critical embodied AI.
NeuroSynTheos: Learning to accelerate counterexample-guided reactive synthesis modulo theories
Andoni Rodríguez ⋅ Cesar Sanchez
Reactive synthesis from temporal logic modulo theories (LTLt) automatically con- structs correct-by-construction controllers for systems with arithmetic constraints— elevators, thermostats, robotic patrols—but the state-of-the-art CEGRES algorithm times out on 77% of existing benchmarks due to a fundamental bottleneck: it discovers the universally valid theory tautologies needed to refine its Boolean abstraction one at a time, through expensive counterexample analysis. We present NEUROSYNTHEOS, a neural augmentation of the CEGRES loop that predicts batches of tautologies at each refinement step. We formalize template invariance for parametric specification families, showing that the prediction problem reduces to constant re-instantiation when specifications share algebraic structure. We train a GNN over formula abstract syntax trees and an autoregressive transformer on tautology traces from 104 solved benchmarks. NEUROSYNTHEOS establishes a new state of the art for LTLt synthesis: on 83 standard benchmarks, it solves 60 ±1.2 specifications (vs. 19 for the previous best), resolving 41 that no prior method could handle within a 5-minute budget, while reducing the median itera- tion count by 5.2×with under 3 seconds of neural overhead. All predictions are SMT-verified before injection, preserving full soundness guarantees.
NIKA: Efficient Neural Video Representation via Structured Latent Diversity
Slater R Victoroff ⋅ Madison May
Implicit neural video representations offer compact, continuous alternatives to conventional codecs, but higher reconstruction fidelity typically requires more computation per decoded frame. We introduce NIKA, which shifts video-specific capacity from the decoder into a structured latent state, allowing reconstruction quality to scale with latent expressivity while keeping active decoding lightweight. NIKA constructs this state from complementary spatial, spectral, and temporal bases, then decodes it with lightweight ConvNeXt-style convolutional upsampling. On UVG, a 2.91M-parameter NIKA model achieves 33.33 dB PSNR at 462 decoding FPS with 4.5G MACs on an RTX A5000, outperforming comparable single-resolution NeRV-family baselines while using 39--51$\times$ fewer MACs. Ablations show that diversifying latent components improves reconstruction more reliably than reallocating capacity within a single component, and qualitative analysis reveals specialization among components. Together, these results identify structured latent diversity as a practical alternative to scaling decoder complexity for high-fidelity, efficient neural video representation.
NRF-GS: Neural Residual Fields for Expressive and Compact Gaussian Splatting
Pratik S. Bisht ⋅ Andreas Kolb
We revisit the role of appearance modeling in 3D Gaussian Splatting (3DGS) and show that limited expressiveness in view-dependent reflectance is a key driver of representation redundancy. In standard 3DGS, low-order spherical harmonics (SH) are used, restricting the splats' ability to model high-frequency directional effects, which is typically compensated by increasing the number of splats. We propose NRF-GS: Neural Residual Fields for Gaussian Splatting, a hybrid representation that replaces per-splat SH-bases with a shared neural residual field. Each Gaussian encodes a compact set of appearance features and a lambertian base color, while a lightweight global scene-level MLP predicts view-dependent residuals conditioned on viewing direction, distance, and per-splat features. This formulation enhances directional reflectance modeling by combining diffuse per-splat reflectance representations with a shared global function for high-frequency details, enabling both higher expressiveness and parameter sharing across splats. Our key insight is that by accurately capturing high-frequency directional reflectance, especially in specular regions, the GS-representation becomes more expressive, reducing the need for geometrically redundant splats. As a result, NRF-GS achieves comparable or better rendering quality while reducing the number of Gaussians by up to 50%, and produces visibly improved specular and high-frequency details.
Omni-SpikeDet: A Spiking Open-World Detector with Dynamic Text–Image Alignment
Ziqi Li ⋅ Tao Gao ⋅ Ting Chen ⋅ Xin Zhang ⋅ Yuanbo Wen ⋅ Yipo Huang ⋅ Robby Tan
Open-vocabulary object detection requires recognizing both known and novel categories beyond a fixed label space. Existing open-vocabulary detectors have made strong progress by integrating visual features with language embeddings, but they are predominantly built on dense ANN-based architectures. In contrast, spiking neural networks (SNNs) have been studied for sparse visual computation, yet their use in open-vocabulary detection remains largely unexplored. This paper presents Omni-SpikeDet, a spiking open-vocabulary detector that investigates how text-guided visual recognition can be realized in a spiking framework. Omni-SpikeDet combines spike-based visual feature extraction with text-guided multi-scale alignment. It uses a spiking convolutional module with integer-spike training and spike-based inference to reduce the mismatch between continuous vision-language supervision and discrete spiking representations. It further introduces a Cross-Scale Text-Image Interaction (CTI) module that modulates multi-level spiking features with language embeddings for prompt-guided detection. Experiments on COCO and LVIS show that Omni-SpikeDet achieves competitive open-vocabulary detection performance with a compact model, obtaining 47.2 AP on COCO and 27.9 AP on LVIS with 4.2M parameters. Under a standard theoretical SNN energy model, Omni-SpikeDet has an estimated energy cost of 19.8 mJ, suggesting a favorable accuracy-efficiency trade-off.
One Scan is Enough: Demonstration-Free Adaptation for Language-Guided Navigation
Zhuoyuan Yu ⋅ Yuxing Long ⋅ Du Wang ⋅ Junyang Wang ⋅ Hongwei Fan ⋅ Hao Dong
Deploying a language-guided navigation policy in a new building usually means watching it fail. Public-benchmark VLAs degrade scene by scene as layout, appearance, and language drift, and the standard fix of teleoperating fresh trajectories on site does not scale. We replace human demonstration with a single brief iPhone scan. Scan2Instruction compiles that scan into a complete training environment, automatically synthesizing executable navigation episodes with grounded instructions. PRPT-KS then adapts the policy through a closed-form Pareto utility read directly off the reconstructed scene, with no learned critic and no human labels. Across five real environments, the two designs together drive our RealNav to an 89% real-robot success rate, almost doubling the strongest VLA baseline at 48%, and still achieve 87% in unscanned regions of the same buildings. Our work turns demonstration-free scene adaptation for language-guided navigation from an aspiration into a working reality. All code and model will be released open-source upon acceptance.
Online learning is an inferential paradigm in which parameters are updated incrementally from sequentially available data, in contrast to batch learning, where the entire dataset is processed at once. In this paper, we assume that mini-batches from the full dataset become available sequentially. The Bayesian framework, which updates beliefs about unknown parameters after observing each mini-batch, is naturally suited for online learning. At each step, we update the posterior distribution using the current prior and new observations, with the updated posterior serving as the prior for the next step. However, this recursive Bayesian updating is rarely computationally tractable unless the model and prior are conjugate. When the model is regular, the updated posterior can be approximated by a normal distribution, as justified by the Bernstein–von Mises theorem. We adopt a variational approximation at each step and investigate the frequentist properties of the final posterior obtained through this sequential procedure. Under mild assumptions, we show that the accumulated approximation error becomes negligible once the mini-batch size exceeds a threshold depending on the parameter dimension. As a result, the sequentially updated posterior is asymptotically indistinguishable from the full posterior.
Online Budget Allocation with Censored Semi-Bandit Feedback
François Bachoc ⋅ Nicolò Cesa-Bianchi ⋅ Tom Cesari ⋅ Roberto Colomboni
We study a stochastic budget-allocation problem over $K$ tasks. At each round $t$, the learner chooses an allocation $X_t \in \Delta_K$. Task $k$ succeeds with probability $F_k(X_{t,k})$, where $F_1,\dots,F_K$ are nondecreasing *budget-to-success* curves, and upon success yields a random reward with unknown mean $\mu_k$. The learner observes which tasks succeed, and observes a task's reward *only* upon success (*censored* semi-bandit feedback). This model captures, for instance, splitting payments across crowdsourcing workers or distributing bids across simultaneous auctions, and subsumes stochastic multi-armed bandits and semi-bandits. We design an optimism-based algorithm that operates under censored semi-bandit feedback. Our main result is an instance-dependent bound showing that in *diminishing-returns* regimes, the regret of this algorithm scales polylogarithmically with the horizon $T$ without any *ad hoc* tuning. For general nondecreasing curves, we prove that the same algorithm (with the same tuning) achieves a worst-case regret upper bound of $\tilde O(K\sqrt{T})$. Finally, we establish a matching worst-case regret lower bound of $\Omega(K\sqrt{T})$ that holds even for *full-feedback* algorithms, highlighting the intrinsic hardness of our problem outside diminishing returns.
We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley-Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards. The observation model inherits the logistic bandit challenge of handling the problem-dependent constant $\kappa$, which accounts for the non-linearity of the link function and can grow arbitrarily large. Moreover, prior work has predominantly focused on linear or kernelized reward models, precluding the use of richer function classes. To address these limitations, we consider general reward function classes and introduce the *locally sensitive eluder dimension*, a novel complexity measure tailored to the logistic structure of preference feedback that yields fine-grained regret guarantees without unfavorable dependence on $\kappa$. Building on this notion, we propose **GINOP** (**G**eneric **IN**formative **OP**timism), an algorithm that constructs log-loss confidence sets and jointly selects arm pairs to balance optimism and informative exploration. We establish a first-order regret bound that, in contrast with what previous results suggest, demonstrates that learning with preference feedback is as statistically efficient as learning from direct reward observation. Finally, we corroborate our theoretical findings with empirical evaluations against competitive baselines.
On the Convergence of First-Order Methods in Signaling Games
Martin Bichler ⋅ Rilind Sahitaj ⋅ Artem Tsikiridis
We study first-order learning dynamics in Bayesian signaling games, a minimal extensive-form model of strategic interaction under asymmetric information. In these games, a Sender observes a private type and chooses a signal, after which a Receiver forms a posterior belief and selects an action. This Bayesian belief update makes the induced learning field fundamentally different from the affine game-gradient fields of normal-form games. Empirically, we find a sharp dichotomy: standard online learning algorithms reliably converge to strict Perfect Bayesian Equilibria (PE), but not to non-strict ones. We show that this behavior cannot be explained by the global geometric conditions commonly used to prove last-iterate convergence, such as monotonicity or the Minty variational inequality. Indeed, these conditions can fail even in binary signaling games. Instead, we identify the relevant local geometry. Our main result characterizes variational stability in finite signaling games: a PE is variationally stable if and only if it is strict. This yields local last-iterate convergence guarantees for online mirror descent and regularized dual averaging near every strict PE. We then study global average-iterate convergence and identify a class of signaling games in which coarse-trigger regret-minimizing dynamics converge to the unique PE. Finally, we complement the theory with experiments comparing projected gradient ascent, mirror-based methods, optimistic variants, and PPO across canonical signaling-game families. The results position signaling games as a compact testbed for understanding how Bayesian belief formation reshapes the convergence geometry of multi-agent learning.
On the Expressive Power and Limitations of Multi-Layer SSMs
Nikola Zubić ⋅ Yuyi Wang ⋅ Yuyi Wang ⋅ Davide Scaramuzza
We study how depth, finite precision, state dimension, and chain-of-thought (CoT) affect the expressive power of multi-layer state-space models (SSMs). For the explicit-table $K$-function-composition problem, a canonical benchmark for sequential information propagation, we prove that any $L$-layer SSM solving $(L+3)$-function composition must satisfy $d^2p=\Omega(N/L^3)$, where $d$ is the state dimension and $p$ is the per-scalar precision. Conversely, $K$-function composition is solved exactly by a $(K+1)$-layer generalized SSM with $d=1$ and $p=\Theta(\log N)$. This gives a worst-case depth hierarchy for this formal problem family. We then distinguish post-input reasoning, in which all thought tokens are generated after the input, from input-interleaved reasoning, in which thought tokens may be inserted while the input stream is being read. Post-input reasoning does not circumvent our communication-based lower-bound pipeline, whereas input-interleaved reasoning admits bidirectional simulations with general deterministic one-pass streaming algorithms at the granularity of persistent memory. Finally, width and precision are not interchangeable under exact step-preserving simulation in the base affine-state model, but become interchangeable through the streaming-memory characterization once input-interleaved reasoning is allowed.
OpenWhistle: A Large-Scale Longitudinal Dataset and Benchmark of Bottlenose Dolphin Vocalizations
Faadil Mustun ⋅ Chiara Semenzin ⋅ Roberto Dessi ⋅ Pablo Robin Guerrero ⋅ pierre orhan ⋅ Alexis Emanuelli ⋅ Emanuele Rossi ⋅ Yair Lakretz ⋅ Gonzalo de Polavieja ⋅ Germán Sumbre
Recent advances in bioacoustics have been driven by large-scale corpora and standardized benchmarks, yet existing resources are overwhelmingly bird-centric and shallow per species, limiting their use for studying the structure of a single species' communication system. This gap is particularly acute for cetaceans: despite bottlenose dolphins (\textit{Tursiops truncatus}) being a compelling case of complex vocal communication among non-human mammals, existing dolphin datasets are small, fragmented, and largely closed. We introduce OpenWhistle, the largest publicly available dataset of dolphin vocalizations. It comprises approximately 180,000 whistles (114 hours) recorded over five years from a stable pod of five individuals in a semi-natural environment, paired with a curated subset of 8,354 expert-annotated whistles and reproducible evaluation protocols for whistle-type detection and classification. We further release the full processing pipeline for whistle detection, segmentation, and categorization. To demonstrate its utility, we pretrain a Wav2Vec2.0 model adapted to dolphin acoustics on the OpenWhistle corpus and show that it learns effective representations, outperforming general-purpose bioacoustic models such as AVES and BioLingual on both tasks while leaving meaningful headroom for future work. By releasing the dataset, pipeline, and evaluation protocol, we provide the first open dolphin whistle dataset tailored for training self-supervised models, laying the groundwork for advancing dolphin communication research and developing models that capture fine-grained acoustic structure within species.
Optimizing Social Utility in Sequential Experiments
Ander Artola Velasco ⋅ Stratis Tsirtsis ⋅ Manuel Rodriguez
Regulatory approval of products in high-stakes domains such as drug development requires statistical evidence of safety and efficacy through large-scale randomized controlled trials. However, the high financial cost of these trials may deter developers who lack absolute certainty in their product's efficacy, ultimately stifling the development of `moonshot' products that could offer high social utility. To address this inefficiency, in this paper, we introduce a statistical protocol for experimentation where the product developer (the agent) conducts a randomized controlled trial sequentially and the regulator (the principal) partially subsidizes its cost. By modeling the protocol using a belief Markov decision process, we show that the agent's optimal strategy can be found efficiently using dynamic programming. Further, we show that the social utility is a piecewise linear and convex function over the subsidy level the principal selects, and thus the socially optimal subsidy can also be found efficiently using divide-and-conquer. Simulation experiments using publicly available data on antibiotic development and approval demonstrate that our statistical protocol can be used to increase social utility by more than $35$% relative to standard, non-sequential protocols.
ORBIT: A Framework for Multi-Agent Security Evaluations
Ben Hagag ⋅ William L Anderson ⋅ Srija Chakraborty ⋅ Christian Schroeder de Witt
Multi-agent LLM systems are increasingly deployed for complex, long-horizon tasks or emerge as a natural consequence of agents interacting in the wild. Yet, they often have significant security vulnerabilities: the flexible protocols that enable task generalization also expose novel threats, from cascading prompt injection to inter-agent collusion. Progress in defending against these threats has been slowed by a lack of empirical evaluations, requiring bespoke environment development for every new defense and making standardized comparison impossible. Existing evaluations address isolated threat models or single-agent settings, but none jointly vary attack, defense, and architecture across realistic multi-agent environments. To address this gap, we introduce ORBIT, a configurable evaluation framework for empirical multi-agent security research, built on UK AISI's Inspect. ORBIT provides configurable communication topologies, memory, scheduling, and agent roles. It supports four threat types (indirect prompt injection, misuse, compromised agents, collusion) and four defense strategies (security prompting, guardian agents, monitors, dual-LLM patterns). The benchmark suite spans five scenario families covering browser use, computer use, agentic coding, customer service, and cooperative allocation. Using ORBIT, we conduct controlled experiments varying defenses, attacks, topologies, and models. We find evidence of significant gaps in defense transferability across threats, security–performance tradeoffs, and interactions between architectural choices and defense effectiveness. To accelerate empirical multi-agent security work, we make ORBIT available open-source at: https://anonymous.4open.science/r/orbit-D1A9.
Overthinking as a Symptom of Knowledge Conflict: Understanding and Detecting LLM's Hallucinations in Retrieval-Augmented Question Answering
Zhihua Wen ⋅ Li Hao ⋅ Zhiliang Tian ⋅ Zhizhao Liu ⋅ Yanqi Shi ⋅ YIPING SONG ⋅ Dongsheng Li
Retrieval-augmented generation (RAG) is increasingly used to ground Large language Models (LLMs) with up-to-date or private information, and has become a key component for knowledge-intensive LLM systems (e.g., LLM agents). However, RAG with LLMs often hallucinate due to unreliable retrieved information or outdated internal knowledge, making hallucination detection crucial for reliable LLM workflows. Recently, reasoning-intensive LLMs that produce explicit reasoning trajectories before answering have shown promise in RAG by examining and synthesizing retrieved evidence. Yet, they often suffer from overthinking: generating redundant or unproductive reasoning with limited performance gains. While prior work mainly studies overthinking in closed-book reasoning tasks that mainly rely on internal knowledge, overthinking in RAG-based question answering (QA) that requires open-book reasoning over both internal knowledge and external context remains underexplored. Thereby, we conduct empirical studies to investigate the relationship between overthinking and hallucination in RAG-based QA and find: (1) Overthinking in reasoning LLMs produces lengthy reasoning trajectories that are correlated with hallucinations; (2) This overthinking is often associated with knowledge conflict, where LLMs struggle to reconcile retrieved context with internal knowledge. Based on our findings, we propose a knowledge conflict-aware hallucination detection framework. We train a classifier to identify the three knowledge conflict reasoning patterns, design a knowledge conflict score to measure knowledge conflict severity, and apply tailored detection strategies with different prompts for different conflict levels. Unlike existing methods, our approach requires neither LLM parameter access nor retrieval reliability, and only needs the LLM output in a single run. Experiments on three QA datasets and three LLMs confirm our framework's effectiveness.
PACZero: PAC-Private Fine-Tuning of Language Models via Sign Quantization
Murat Bilgehan Ertan ⋅ Xiaochen Zhu ⋅ Ha Nguyen ⋅ Marten van Dijk ⋅ Srinivas Devadas
We introduce PACZero, a family of PAC-private zeroth-order mechanisms for fine-tuning large language models that delivers usable utility at $I(S^*; Y_{1:T})=0$. This privacy regime bounds the membership-inference attack (MIA) posterior success rate at the prior, an MIA-resistance level the DP framework matches only at $\varepsilon=0$ and infinite noise. All DP-ZO comparisons below are matched at the MIA posterior level. The key insight is that PAC Privacy charges mutual information only when the release depends on which candidate subset is the secret. Sign-quantizing subset-aggregated zeroth-order gradients creates frequent unanimity, steps at which every candidate subset agrees on the update direction; at these steps the released sign costs zero conditional mutual information. We propose two variants that span the privacy-utility trade-off: PACZero-MI (budgeted MI via exact calibration on the binary release) and PACZero-ZPL ($I=0$ via a uniform coin flip on disagreement steps). We evaluate on SST-2 and SQuAD with OPT-1.3B and OPT-6.7B in both LoRA and full-parameter tracks. On SST-2 OPT-1.3B full fine-tuning at $I=0$, PACZero-ZPL reaches ${88.99\pm0.91}$, within $2.1$pp of the non-private MeZO baseline ($91.1$ FT). No prior method produces usable utility in the high-privacy regime $\varepsilon<1$, and PACZero-ZPL obtains competitive SST-2 accuracy and nontrivial SQuAD F1 across OPT-1.3B and OPT-6.7B at $I=0$.
Participatory ML for Social Harm Should be Constructed, Validated and Reasoned Through First-Person Accounts
Atmadeep Ghoshal ⋅ Ankit Agarwal ⋅ Martim Brandao
Participatory machine learning has improved who develops social harm evaluation resources but has left largely unexamined what kind of text serves as the raw signal for those resources. Researcher-generated stimuli fix the epistemic frame of harm before community participation begins, and the structured conditions of annotation prevent the richest forms of interpretive harm knowledge from entering the pipeline. We argue that participatory ML should construct, validate and reason about social harm pipelines through first-person accounts, by which we mean autobiography, memoir, testimony, and oral history authored by people describing their own experiences of discrimination and harm. Such accounts possess two epistemic properties that directly address these limitations. Authorial independence from research framing grounds what counts as harm, which contexts matter, and how discrimination is operationalised in experience rather than in researchers' prior conceptions of harm. Embedded rationale structure preserves the interpretive judgment through which harm becomes meaningful rather than reducing it to a label. We draw on existing archival corpora spanning multiple languages and harm domains and propose three concrete integration pathways across evaluation benchmarks, harm classifier training corpora, and safety preference datasets.
Payoff-Only Learning in Games via Best-Response Differential Inclusions
Ismail Hassan ⋅ Pedro G Lind ⋅ Anis Yazidi
Each player in a finite normal-form game observes only its own realized reward; opponents, their actions, and their payoffs are unknown. We analyze an epochal payoff-only recursion in which the mixed strategy is held fixed across long epochs, own-action means are estimated from bandit data, and the iterate steps toward an empirical near-best vertex of a barrier simplex. Our main result reduces this stochastic recursion to the deterministic best-response differential inclusion (BRDI). Under a Hoeffding-tied epoch schedule, the empirical near-best correspondence is contained, on a tail event of probability at least $1-\sum_{k\ge K}\delta_k$, in a deterministic graph enlargement of the BRDI whose radius vanishes as $k\to\infty$. Without conditioning, the affine interpolation of the iterates is almost surely a bounded perturbed solution of the BRDI in the Bena\"im--Hofbauer--Sorin sense, and its $\omega$-limit set is internally chain transitive. Composing this reduction with deterministic BRDI confinement results yields last-iterate Nash convergence in three classes: $\operatorname{dist}(x_k,\NE)\to 0$ almost surely in finite zero-sum games, in finite exact-potential games (under an empty-interior condition on the potential's image), and in every generic $2\times 2$ game. Fixed barriers give $O(b)$-Nash equilibria of the original game; deterministic shrinking removes the bias.
Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs
Jiafeng Liang ⋅ Zhihao Zhu ⋅ Zihan Zhang ⋅ Baoqi Ren ⋅ Tao Ren ⋅ Shixin Jiang ⋅ Runxuan Liu ⋅ Ming Liu ⋅ See-Kiong Ng ⋅ Bing Qin
Although Large Multimodal Models (LMMs) have achieved strong performance on general video understanding, their susceptibility to textual prior shortcuts during causal discovery has been recognized as a critical deficit. The underlying mechanisms of this phenomenon remain incompletely understood, as existing benchmarks only measure response accuracy without revealing the sources and extent of the deficit. We introduce ProCauEval, a perturbation-based evaluation protocol that shifts from outcome assessment to mechanism diagnosis, probing causal discovery through five controlled configurations that systematically manipulate visual and textual modalities to decompose their respective contributions to model behavior and dissect the failure modes. Evaluating 17 mainstream LMMs, we find that models faithfully perceive video content yet systematically underexploit it during causal reasoning. We further observe that stronger post-training amplifies rather than mitigates textual prior reliance, and that higher baseline performance correlates with greater fragility under perturbation. To address these, we propose Anti-Distillation Policy Optimization (ADPO), a reinforcement learning framework built on negative teacher alignment, which augments GRPO by explicitly pushing the policy away from a prior-only counterfactual teacher induced by visual corruption. Specifically, ADPO maximizes the divergence between the policy distributions conditioned on the original and visually corrupted inputs, thereby forcing the model to ground its reasoning in visual evidence rather than textual shortcuts. Extensive experiments show that ADPO improves visual engagement without sacrificing fundamental comprehension, thus offering a preliminary step toward reliable causal discovery.
Phaedra: Learning High-Fidelity Discrete Tokenization for the Physical Sciences
Levi Lingsch ⋅ Georgios Kissas ⋅ Johannes Jakubik ⋅ Siddhartha Mishra
Tokens are discrete representations that allow modern deep learning to scale by transforming high-dimensional data into sequences that can be efficiently learned, generated, and generalized to new tasks. While foundational for image and video generation, the application of tokens to physical simulation remains nascent. Because existing tokenizers are designed for the perceptual requirements of natural images, they struggle with scientific data, which exhibits large dynamic ranges and requires exact preservation of physical and spectral properties. In this work, we investigate the performance of a suite of image tokenizers across metrics designed to measure PDE fidelity. Observing that these baselines struggle to simultaneously capture fine geometric details and precise physical magnitudes, we propose Phaedra, a novel tokenizer inspired by classical shape-gain quantization and the paradigm of basis functions coupled with continuous coefficients. Phaedra acts as a highly effective nonlinear compression algorithm, massively reducing dataset footprints while maintaining physical fidelity. We demonstrate that Phaedra consistently improves reconstruction across diverse PDE datasets, generalizes robustly to unseen PDE types and real-world Earth observation data, and provides immediate accuracy gains in initial downstream operator learning and masked autoencoding tasks.
PieArena: Ranking and Profiling Language Agents in Realistic Negotiation Scenarios
Chris Zhu ⋅ Sasha Cui ⋅ Will S Dufallo ⋅ Runzhi Jin ⋅ Zhen Xu ⋅ Linjun Zhang ⋅ Daylian M Cain
We present an in-depth evaluation of LLMs' ability to negotiate, a central business task requiring strategic reasoning, theory of mind, and economic value creation. To do so, we introduce PieArena, a large-scale negotiation benchmark grounded in multi-agent interactions over realistic scenarios adapted from MBA negotiation courses at an elite business school. We evaluate language agents across three pairing regimes: mirror-play, cross-play, and human–LM play. We develop a ranking model for continuous negotiation payoffs that yields order-invariant, uncertainty-quantified leaderboards while correcting for systematic experimental asymmetries. We further study the effects of joint-intentionality agentic scaffolding and find asymmetric gains, with large improvements for mid- and lower-tier LMs and diminishing returns for frontier LMs. As calibration anchors, we collect human–human and human–LM negotiation data from trained business school students, finding that a representative frontier language agent (GPT-5) matches or exceeds this human baseline in our evaluation settings. Beyond deal outcomes, PieArena provides a multi-dimensional behavioral profile that reveals cross-model heterogeneity in instruction compliance, computation accuracy, as well as judge-assessed deception and reputation, illustrating the value of evaluation beyond outcome-only leaderboards.
POLAR-Bench: A Diagnostic Benchmark for Privacy-Utility Trade-offs in LLM Agents
Qiaoyuan ZHENG ⋅ yiqu yang ⋅ Qi Gao ⋅ Imanol Schlag
LLM agents increasingly have access to private user data and act on the user's behalf when interacting with third-party systems. The user defines what may and must not be shared, and the agent must robustly follow that intent even when third-party systems behave adversarially. We introduce $\textbf{POLAR-Bench}$ ($\textbf{Pol}$icy-$\textbf{a}$ware adve$\textbf{r}$sarial Benchmark), in which a trusted model with a privacy policy and a task converses with a third-party model that adversarially probes for both task-relevant and protected attributes. Across 10 domains and 7${,}$852 samples, we score privacy and utility by deterministic set-membership and vary privacy policy dimension and attack strategy along two orthogonal axes, producing a $5\times5$ diagnostic surface per model. Our results reveal a sharp split: current frontier models withhold over 99% of protected attributes, while smaller open-weight models in the 1--30B range, the class users most commonly run as their own trusted agent on-device or via private inference, score notably worse, with the weakest leaking over half. POLAR-Bench thus localizes where each model's intent-following breaks down, providing a foothold for privacy alignment where it matters most.
Predicting directional flexibility in proteins
Vsevolod Viliuga ⋅ Leif Seute ⋅ Matteo Tadiello ⋅ Nicolas Wolf ⋅ Frauke Gräter ⋅ Arne Elofsson
Predicting protein dynamics is a long-standing problem in computational structural biology. Often, protein function critically depends on local directed motions, such as hinge movements, catalytic loop rearrangements and domain reorientations, which can be characterized by directional flexibility and correlated structural motions of the protein backbone. While Molecular Dynamics (MD) simulations provide an established but often prohibitively expensive approach, recent deep generative models aim to reduce this cost by directly predicting conformational ensembles, emulating MD. However, due to their large size and the need to generate several states until the derived dynamical properties converge, these models remain expensive. In this work, we propose a fast SE(3)-equivariant graph neural network to directly predict dynamical descriptors, such as directional backbone flexibility and pairwise dynamic correlations, from an equilibrium structure. In a series of experiments, we show that our model matches the accuracy of substantially larger ensemble generation models while being orders of magnitude faster, and demonstrate that the proposed equivariant architecture is especially well-suited for capturing anisotropic motions in proteins.
Predicting Only from Selected Evidence: A Tempered Product-of-Experts Bottleneck for Auditable EEG Diagnosis
Yinghao WANG ⋅ Shujian Yu ⋅ Duc-Han LE ⋅ Zhikai Yu ⋅ Changming Wang ⋅ Van-Tam Nguyen
Pretrained EEG backbones improve transfer performance, but downstream diagnostic heads remain difficult to audit: predictions are typically computed from unconstrained hidden representations, while explanations are often generated only post hoc. We introduce tPoE-EIB, an $\textit{evidence-information bottleneck}$ head for adapting EEG backbones under an evidence-only prediction constraint. tPoE-EIB selects temporal and channel-specific evidence, maps the resulting summaries to Gaussian experts over a shared latent variable, and fuses them via a tempered product-of-experts posterior. The classifier conditions exclusively on this latent variable, yielding an explicit decision pathway whose information flow is constrained by the expected posterior KL divergence. This formulation induces a tractable supervised objective with an information-rate penalty, while the closed-form tempered posterior mitigates overconfident aggregation of correlated evidence sources. We evaluate tPoE-EIB on pretrained EEG foundation model backbones across six diagnostic tasks: event-type classification, abnormality detection, seizure detection, cognitive decline staging, depression screening, and cerebrovascular disease classification. The evaluation spans public benchmarks and in-house clinical cohorts, covering both binary screening and fine-grained staging, as well as sparse and dense electrode montages. tPoE-EIB maintains competitive balanced accuracy while outperforming representative post-hoc explanation methods on selection-faithfulness audits, including insertion–deletion and gate-causality tests. Its structured posterior further supports integration-faithfulness evaluations, including expert-drop, posterior-reliance, and expert-disagreement tests. Overall, these results suggest that evidence-only, rate-limited fusion is a practical approach to building auditable diagnostic models on top of frozen EEG foundation backbones.
Profit Maximization in Bilateral Trade against a Smooth Adversary
Simone Di Gregorio ⋅ Paul Duetting ⋅ Federico Fusco ⋅ Chris Schwiegelshohn
Bilateral trade models the task of intermediating between two strategic agents, a seller and a buyer, who wish to trade a good. We study this problem from the perspective of a profit-maximizing broker within an online learning framework, where the agents' valuations are generated by a smooth adversary. We devise a learning algorithm that guarantees a $\tilde{O}(\sqrt{T})$ regret bound, which is tight in the time horizon $T$ up to poly-logarithmic factors. This matches the minimax rate for the stochastic i.i.d. case, and is also well separated from the adversarial setting, where sublinear-regret is unattainable. By extending the strong regret guarantees from the i.i.d. case to the smooth adversary, we significantly broaden the scope of settings where such fast rate is achievable, while closing an important gap in the regret landscape of this fundamental economic problem. To overcome the challenges posed by this adversary, we leverage a continuity property of smooth instances and combines this with a hierarchical net-construction of the broker's action space, which is analyzed via algorithmic chaining. We showcase the applicability of these techniques by deriving a similarly tight $\tilde{O}(\sqrt{T})$ regret bound for a related mechanism design model: the joint ads problem.
Propagate to Discover: Graph-Structured Propagation for Generalized Category Discovery
yongqi tian ⋅ Junyong Liu ⋅ Jinkun Ran ⋅ Haoyuan He ⋅ Caigui Jiang
Generalized Category Discovery (GCD) aims to recognize both known and novel categories from a small labeled set and a large unlabeled pool. Existing methods often rely on sparse supervision to shape the representation space, which can lead to unstable separation between old and new categories, especially in fine-grained recognition scenarios. In this paper, we propose PD-GCD, a graph-structured propagation framework that views sparse labels as semantic seed signals and unlabeled samples as nodes on a data manifold. PD-GCD first learns a GCD-oriented representation space and performs Anchor Mining to identify propagation anchors from the unlabeled pool, balancing semantic coverage with boundary-aware exploration. These anchors are then used by Graph Propagation to spread semantic information over the data graph and produce propagated soft labels. To further align model predictions with the propagated structure, we introduce graph-guided structural distillation and graph consistency regularization, forming a closed-loop process between anchor mining, label propagation, and representation refinement. Extensive experiments on six benchmarks demonstrate that PD-GCD achieves state-of-the-art performance, with absolute All-accuracy gains of 6.7\% across all datasets and 9.4\% on fine-grained datasets.
Property-Guided LLM Program Synthesis for Planning
André G. Pereira ⋅ Augusto B. Corrêa ⋅ Jendrik Seipp
LLMs have shown impressive success in program synthesis, discovering programs that surpass previously known solutions. However, these approaches rely on simple numeric scores to signal program quality, such as the value of the solution or the number of passed tests. Because such a score offers no guidance on why a program failed, the system must generate and evaluate many candidates per problem in the hope that some succeed, increasing both LLM inference and evaluation costs. We study a different approach: property-guided LLM program synthesis. Instead of scoring programs after evaluation, we check whether a candidate satisfies a formally defined property. When the property is violated, we stop the evaluation early and provide the LLM with a concrete counterexample showing exactly how the program failed. This feedback drastically reduces both the number of program generations and the evaluation cost, and can guide the LLM to generate stronger programs. We evaluate this approach on PDDL planning domains, asking the LLM to synthesize direct heuristic functions: every state reachable by strictly improving transitions has a strictly improving successor. A heuristic with this property leads hill-climbing search directly to a goal state. A counterexample-guided repair loop generates one candidate program, checks the property over a training set, and returns the first case that violates the property. We evaluate our approach on ten planning domains with an out-of-distribution test set. The synthesized heuristics are effectively \emph{direct} on virtually all test tasks, and compared to the best prior generation method our approach generates seven times fewer programs per domain on average, solves more tasks without using search, and requires several orders of magnitude less computation to evaluate candidates. Whenever a problem admits a verifiable property, property-guided LLM program synthesis can reduce synthesis and evaluation cost while improving program quality.
Proxy-Based Approximation of Shapley and Banzhaf Interactions
Santo Thies ⋅ Hubert Baniecki ⋅ R. Teal Witter ⋅ Eyke Hüllermeier ⋅ Maximilian Muschalik ⋅ Fabian Fumagalli
Shapley and Banzhaf interactions capture the complex dynamics inherent in modern machine learning applications. However, current estimators for these higher-order interactions trade off between speed and accuracy. To overcome this limitation, we introduce ProxySHAP. ProxySHAP reconciles the high sample efficiency of tree-based proxy models with a principled path to consistency via residual correction. On a theoretical level, we derive a polynomial-time generalization of interventional TreeSHAP to compute exact interaction indices for tree ensembles, successfully bypassing exponential tree-depth dependencies in prior methods. Furthermore, we formally analyze the residual adjustment strategy, characterizing the specific conditions under which Maximum Sample Reuse (MSR) corrects proxy bias without its variance scaling exponentially with interaction size. Extensive benchmarking demonstrates that ProxySHAP sets a new state-of-the-art standard for approximation quality, including in large-scale applications with thousands of features. By achieving the lowest error in both small- and large-budget regimes, ProxySHAP significantly outperforms the prior best estimators ProxySPEX and KernelSHAP-IQ, while also delivering superior performance on downstream explainability tasks.
RankAlign: Unsupervised Vision-Language Representation Alignment via Rank Transformation
Enzhe Zhao ⋅ Marco Fiorucci ⋅ Lamberto Ballan
The Platonic Representation Hypothesis suggests that vision and language models converge toward a shared latent structure, yet the unsupervised alignment of unpaired modalities remains a fundamental challenge. Current methods rely on direct comparisons of intra-modality similarity kernels, which are often incommensurate due to disparate scales and representation densities. This mismatch yields ill-conditioned optimization landscapes, hindering stable alignment. To address this, we propose RankAlign, a framework that independently transforms each within-modality kernel into a rank-based representation of relative neighbor relationships. By shifting from absolute similarity values to ordinal relational geometry, RankAlign ensures structural consistency while remaining robust to modality-specific distortions. We show that rank-based transformation reshapes similarity distributions, mitigating representation collapse and providing a more discriminative alignment signal. Experiments demonstrate that RankAlign is noise-resilient and overcome existing kernel-based alignment method across diverse datasets, including a medical imaging dataset where subtle class differences typically limit existing approaches alignment approaches.
Regularized Large Neighborhood Search
Germain Vivier-Ardisson ⋅ Laurent Demonet ⋅ Axel Parmentier ⋅ Mathieu Blondel
Operations research practitioners typically tackle NP-hard combinatorial problems using large neighborhood search (LNS), a scalable heuristic that iteratively refines a current solution by locally re-optimizing subsets of its variables. In contrast, most existing approaches for integrating combinatorial optimization layers into neural networks still assume access to an exact global solution, which is computationally intractable. We bridge this gap by introducing regularized LNS (RLNS). By regularizing or perturbing local subproblems, we turn the LNS heuristic into an efficient MCMC sampler over the combinatorial set of feasible solutions, with associated Fenchel-Young losses. Under entropic regularization, we prove that RLNS performs exact block Gibbs sampling. Furthermore, adjusting the number of RLNS iterations allows us to interpolate between pseudolikelihood and exact maximum likelihood estimation, for end-to-end learning without global solvers. We demonstrate our approach on $k$-subset selection, generalized assignment, and stochastic vehicle scheduling problems.
Replicable Constrained Bandits
Matteo Bollini ⋅ Gianmarco Genalti ⋅ Francesco Emanuele Stradi ⋅ Matteo Castiglioni ⋅ Alberto Marchesi
Algorithmic *replicability* has recently been introduced to address the need for reproducible experiments in machine learning. A *replicable online learning* algorithm is one that takes the same sequence of decisions across different executions in the same environment, with high probability. We initiate the study of algorithmic replicability in *constrained* MAB problems, where a learner interacts with an unknown stochastic environment for $T$ rounds, seeking to maximize reward while satisfying multiple constraints. Our main result is that replicability can be achieved in constrained MABs. Specifically, we design replicable algorithms whose regret and constraint violation match those of non-replicable ones in terms of $T$. As a key step, we develop the first replicable UCB-like algorithm for *unconstrained* MABs, showing that algorithms that employ the optimism in-the-face-of-uncertainty principle can be replicable, a result that we believe is of independent interest.
Rethinking Latency Denial-of-Service: Attack the LLM Serving Framework, Not the Model
Tianyi Wang ⋅ Huawei Fan ⋅ Yuanchao Shu ⋅ Peng Cheng ⋅ Cong Wang
LLM inference is inherently expensive, even a modest slowdown can translate into substantial operating costs and severe availability risks. Recently, a growing body of research known as latency attacks focuses on crafting inputs to trigger worst-case output lengths. However, we report a contrary finding that these algorithmic-level latency attacks are largely ineffective against modern LLM serving systems. We reveal that system-level optimization such as continuous batching provides a logical isolation to mitigate contagious latency impact on co-located users. Thus, in this paper, we shift our focus from the algorithm to the system layer, and introduce a new Fill and Squeeze attack strategy targeting the state transition of the scheduler. "Fill'' first exhausts the global KV cache to induce Head-of-Line blocking, while "Squeeze'' forces the system into repetitive preemption. By manipulating output lengths using different attack prompts, and leveraging side-channel probing of memory status, we demonstrate that the attack can succeed in a practical black-box setting with much less cost. Extensive evaluations on vLLM indicate up to $75-742\times$ TTFT degradation relative to benign baselines and $1.5-4\times$ average slowdown on Time Per Output Token compared to existing attacks with 30-40\% lower attack cost. Code: https://anonymous.4open.science/r/FS-EE97/README.md
Activation functions are considered an essential primitive for neural nonlinearity, i.e., they enable neural networks to serve as universal approximators. In this paper, we show that this nonlinearity can also be achieved by input-conditioned threshold gating through branches as a universal primitive. We demonstrate that standard activations—whether piecewise-linear (ReLU, PReLU, Hardtanh) or smooth (SiLU, Sigmoid, Tanh, GELU)—are in fact instances of a single Threshold Gating (TG) primitive. For softmax, we show that it admits an exact TG conversion via its equivalent per-element Sigmoid form. We then validate these equivalences by converting pretrained networks across CNNs, transformer-based models, and recurrent architectures, preserving model performance without requiring retraining. Threshold Gating also enables training from scratch that goes beyond replacing existing activations, enabling gains in model compression, performance, and shorter training. We also propose a 'Minimal Branch Theorem' which relates the minimum number of required branches in our primitive to the trainability of general deep neural networks. In terms of hardware implementation, TG maps to a unified implementation in the case of analog in-memory systems, addressing the bottleneck of analog-to-digital and digital-to-analog converters (ADC/DAC) that is known to significantly impact power consumption and on-chip area.
Modern data analysis pipelines increasingly rely on large datasets assembled through retrieval, scraping, automated filtering, and post-hoc curation, where the observed sample is often a contaminated mixture rather than the target population. Standard distributional objectives either ignore contamination and converge to the wrong target, or retain only a small trusted subset and discard most of the data. We study supervised and unsupervised distributional learning when a clean reference sample is compared with a pooled sample from $r=(1-\epsilon)q+\epsilon z$, where $q$ is the latent target distribution and $z$ is an unknown contaminant. Motivated by settings with noisy membership scores for all observations and exact labels for a small audited subset, we propose contamination-corrected maximum mean discrepancy (CC–MMD), a design-based estimator that combines proxy scores with audited residual corrections to recover the oracle MMD that would have been computed from the latent clean sample. CC–MMD requires no calibration assumption on the proxy scores and no structural assumptions on the contaminant beyond standard kernel regularity. We prove finite-population design-unbiasedness of the corrected kernel sums, establish asymptotic normality of the audit correction, and show consistency for the target discrepancy $\mathrm{MMD}^2(p,q;k)$. Empirically, CC–MMD tracks the oracle under increasing contamination, improves parameter recovery and prediction in contaminated regression, and enables generative models trained on heavily corrupted mixtures to recover the clean target distribution. These results show that contaminated data need not be discarded or treated as the estimand. With sparse trusted labels and abundant noisy supervision, distributional objectives can be made robust by design.
Rotations on Latent Hyperspheres: a Geometry-Aware Guiding Framework for Diffusion Models
Luca Sacchetto ⋅ Klaus Diepold
Diffusion models have emerged as a powerful tool across diverse domains. However, their purely data-driven nature can produce samples that violate domain-governing constraints. We introduce a plug-and-play Reinforcement Learning framework that optimizes initial noise samples in the latent space of frozen, pre-trained diffusion models. Leveraging the near-spherical geometry of high-dimensional Gaussian distributions, we introduce a novel rotation-matrix-based scheme for efficient latent space exploration. This steers the model toward more feature-preserving outputs, guided by task-specific rewards. We evaluate our method on three diffusion models: one trained on solutions of the Darcy Flow PDE, one on a synthetic dataset with complex structural features, and a text-conditioned one. Across all three settings, our framework yields significant improvements in sample quality, achieving a ${\sim}25\\%$ relative reduction in PDE residual, up to a ${\sim}44\\%$ relative improvement on the synthetic dataset's feature-alignment metric, and up to a ${\sim}80\\%$ relative improvement on human preference, compared to the vanilla diffusion models. Finally, we show that rotation-matrix-based exploration significantly outperforms unconstrained exploration, validating our geometry-aware approach and establishing a more effective method for latent space control.
Safety Geometry Collapse in Multimodal LLMs and Adaptive Drift Correction
Jiahe Guo ⋅ Xiangran Guo ⋅ Jiaxuan Chen ⋅ Weixiang Zhao ⋅ Yanyan Zhao ⋅ Yutai Hou ⋅ Qianchao Wang ⋅ Dandan Tu ⋅ Bing Qin
Multimodal large language models (MLLMs) often fail to transfer safety capabilities learned in the text modality to semantically equivalent non-text inputs, revealing a persistent multimodal safety gap. We study this gap from a representation-geometric perspective by analyzing a text-aligned refusal direction and a modality-induced drift direction. We show that multimodal inputs compress the usable separation along the refusal direction, making it no longer reliable for identifying and refusing harmful inputs. We refer to this failure mode as Safety Geometry Collapse. We quantify it through conditional refusal separability and show that stronger modality-induced drift is consistently associated with weaker refusal separability and higher attack success rates. We then validate the causal role of modality-induced drift through a fixed-strength activation intervention: counteracting the estimated drift restores refusal separability and improves multimodal safety. After drift correction, we further observe self-rectification, where the model recovers its ability to recognize and refuse harmful multimodal inputs during forward dynamics. This effect also provides an internal signal of the model’s perceived harmfulness of each input. Motivated by this signal, we propose ReGap, a training-free inference-time method that adaptively corrects modality drift using self-rectification. Experiments across multiple multimodal safety benchmarks and utility benchmarks demonstrate the effectiveness of ReGap, which significantly improves the safety of MLLMs without compromising general capabilities. Our findings highlight representation-level modality alignment as a crucial direction for real-time safety improvement and for building safer, more reliable MLLMs.
Scale-Sensitive Shattering: Learnability and Evaluability at Optimal Scale
Shashaank Aiyer ⋅ Yishay Mansour ⋅ Shay Moran ⋅ Han Shao ⋅ Tom Waknine
We study the optimal scale at which real-valued function classes exhibit uniform convergence and learnability. Our main result establishes a scale-sensitive generalization of the fundamental theorem of PAC learning: for every bounded real-valued class $\mathcal{F}$ and every $\gamma>0$, uniform convergence at scale $\gamma$, agnostic learnability at scale $\gamma/2$, and finiteness of the fat-shattering dimension at every scale $\gamma'>\gamma$ are equivalent. This resolves a question by Anthony and Bartlett (Cambridge Univ. Press ’99) on the precise scales governing learnability, refuting a conjecture attributed there to Phil Long that a multiplicative 2-factor gap is unavoidable, and improves the upper bounds of Bartlett and Long (JCSS ’98), which incur such a loss. The key technical ingredient is a direct bound on empirical $\ell_\infty$ covering numbers, avoiding the standard detour through packing numbers. As a consequence, we obtain sharp asymptotic metric-entropy bounds in terms of the fat-shattering scale $\gamma$: an $O(\log^2 n)$ bound holds already at scale $\gamma/2$, while an $O(\log n)$ bound holds at scale $2\gamma$. We further show that the $O(\log^2 n)$ bound is sometimes tight. These results resolve open questions by Alon et al. (JACM ’97) and Rudelson and Vershynin (Ann. of Math. ’06). As an application, we establish a sharp dichotomy for bounded integral probability metrics: every such IPM is either estimable or cannot be weakly evaluated within any multiplicative factor $c<3$, while $3$-weak evaluability always holds, resolving an open question from Aiyer et al. (ICML '26). We conclude with several open questions on quantitative sample complexity and evaluability.
Seahorse: A Unified Benchmarking Framework for Spatiotemporal Event Modeling
Yahya Aalaila ⋅ Sebastian Vollmer ⋅ Gerrit Großmann
Spatiotemporal point processes (STPPs) model event data in continuous time and space, with applications in mobility, epidemiology, and public safety. Recent neural STPPs span expressive intensities, density models, continuous-time latent dynamics, normalizing-flow spatial decoders, and score-based generative mechanisms. Yet comparison remains fragile because implementations differ in preprocessing, coordinate normalization, splits, likelihood conventions, and evaluation protocols. We present \textsc{Seahorse}, a unified framework for reproducible STPP experimentation. \textsc{Seahorse} provides an encode--evolve--decode interface and an executable benchmarking contract that standardizes dataset handling, configuration resolution, training, tuning, checkpointing, raw-coordinate likelihood reporting, and run artifacts across heterogeneous STPP families. It supports classical baselines, neural likelihood models, continuous-time Neural STPP variants, and sample-based generative models under one external protocol. Using HawkesNest as a controlled synthetic stress-test suite, we show how diagnostic benchmarking can probe neural STPP behavior under increasing spatiotemporal entanglement and other structural changes. Together, \textsc{Seahorse} and these controlled evaluations provide reusable infrastructure for comparing, diagnosing, and extending STPP models beyond implementation-specific leaderboards.
Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext inside the runtime before any final output check can intervene. Existing defenses usually protect one boundary, either the planner/runtime or the action sink, and therefore do not by themselves secure both surfaces. We present SecureClaw, a dual-boundary architecture that places authorization at the effect sink and plaintext confinement at the read boundary. Sensitive reads pass through a trusted gateway that replaces raw values with opaque handles and, in the evaluated deployment, bounded summaries as an explicit declassification interface. Writes that change external state follow a PREVIEW$\rightarrow$COMMIT protocol in which only a trusted executor may commit the exact canonical request authorized by policy. The runtime can still plan over summaries and symbolic references, but cannot directly dereference secrets or perform side effects. Across AgentDojo, AgentLeak, and Agent Security Bench (ASB), SecureClaw is the only defense we evaluate in a common harness that simultaneously retains usable task utility and achieves 0\% attack success rate (ASR) on ASB, 0.64\% ASR on AgentDojo, and 3.23\% overall leak on AgentLeak's attacked parity lane, which measures final-output and internal-relay leakage.
Self-Evolving Agents Should Build Internal and External Models of the World
Luiz Felipe Vecchietti ⋅ Bryan Nathanael Wijaya ⋅ Wenchao Dong ⋅ Meeyoung Cha
Developing self-evolving agents that learn continuously from streams of information represents an emerging frontier in artificial intelligence. This capability is critical for efficiently updating computationally expensive systems such as foundation models. When agents learn by exploring the environment and interacting with other agents to achieve multiple tasks over time, two sources of non-stationarity emerge: shifting task distributions and co-evolving peers. However, developing frameworks for continual multi-agent reinforcement learning (CMARL) and effective policy improvement remains challenging. In this paper, we argue that training self-evolving agents should incorporate two classes of model-based principles: internal models that allow an agent to compare policy checkpoints, monitor its learning capacity, and represent its current objective; and external models that learn environment dynamics and the behavior of other agents. We anticipate that discussion on these principles will facilitate the design of algorithms, architectures, and benchmarks that explicitly evaluate model-based components in CMARL.
Simulation-aided Reinforcement Learning with Control Variates
Omer Hemo ⋅ Ido Sefi ⋅ Mirco Mutti ⋅ Aviv Tamar ⋅ Kfir Y. Levy
Reinforcement Learning (RL) algorithms often rely on data collected from a target environment, which can be expensive and noisy. Although simulators can generate large amounts of additional data, standard simulation mixing approaches—which directly combine simulated and real experience—can degrade performance due to simulator inaccuracies, especially when they bias the solution away from the optimal policy. We propose a conservative simulation-aided learning method based on control variates, where simulated data is used solely to reduce variance in the learning process rather than directly shaping the learning target. We develop our approach for Policy-Evaluation, which is a fundamental building block of RL algorithms. Then we combine it within the actor-critic framework towards policy improvement, as well as extend it towards offline RL scenarios. Experiments on multiple MuJoCo benchmark tasks demonstrate improvements over standard simulation mixing approaches, highlighting control variates as an effective tool for simulation-aided reinforcement learning.
Solaris: Building a Multiplayer Video World Model in Minecraft
Oscar Michel ⋅ Georgy Savva ⋅ Daohan Lu ⋅ Suppakit Waiwitlikhit ⋅ Timothy Meehan ⋅ Dhairya P Mishra ⋅ Egor Gikalo ⋅ Srivats Poddar ⋅ Jack Lu ⋅ Saining Xie
Existing action-conditioned video generation models (video world models) are limited to single-agent perspectives, failing to capture the multi-agent interactions of real-world environments. We introduce Solaris, a multiplayer video world model that simulates consistent multi-view observations. We develop a multiplayer data system designed for robust, continuous, and automated data collection on video games such as Minecraft. Unlike prior platforms built for single-player settings, our system supports coordinated multi-agent interaction and synchronized videos + actions capture. Using this system, we collect 12.64 million multiplayer frames and propose an evaluation framework for multiplayer movement, grounding, building, and view consistency. We train Solaris using a staged pipeline that progressively transitions from single-player to multiplayer modeling, combining bidirectional, causal, and Self Forcing training. In the final stage, we introduce Checkpointed Self Forcing, a memory-efficient Self Forcing variant that enables a longer-horizon teacher. Results show our architecture and training design outperform existing baselines. We also demonstrate that scaling beyond two players is possible by training a 3-player proof-of-concept model and simulating up to 16 players on our hardware with efficient attention. Through open-sourcing our system and models, we hope to lay the groundwork for a new generation of multi-agent world models.
Spherical Boltzmann machines: a solvable theory of learning and generation in energy-based models
Thomas Tulinski ⋅ Simona Cocco ⋅ Remi Monasson ⋅ Jorge FERNANDEZ-DE-COSSIO-DIAZ
Energy-based models (EBMs) are flexible generative architectures inspired by statistical physics, but their learning and generative properties remain poorly understood. Here, we analyze a solvable EBM in the high-dimensional limit: the spherical Boltzmann machine (SBM). Combining tools from random matrix theory and dynamical mean-field theory, we: solve exact equations describing the training dynamics of the SBM; compute the Bayesian evidence, which acts as a partition function in parameter space and encodes global properties of the trained model; and uncover cascades of phase transitions that occur both during training and as a function of hyperparameters, related to successive alignment and condensation of the top modes of the coupling matrix to the data. We connect these transitions to sampling-time generative phenomena in a teacher-student scenario, including: sampling temperature tuning, double descent as a function of regularization strength, tempered posterior effects, and out-of-equilibrium effects during training that induce biases in the trained model. We provide numerical evidence demonstrating that all these phenomena appear in standard generative architectures, beyond the SBM.
spora: A Unified Multimodal Dataset for Spatial Proteomics
Benedikt von Querfurth ⋅ Eeshaan Jain ⋅ Johann Wenckstern ⋅ Lukas Klein ⋅ Phil F Cheng ⋅ Yexiang Cheng ⋅ Cédric Vincent-Cuaz ⋅ Pascal Frossard ⋅ Martina Haberecker ⋅ Andreas Wicki ⋅ Olivier Michielin ⋅ Charlotte Bunne
Spatial proteomics (SP) captures the molecular composition and spatial organization of tissues at single-cell resolution, opening a direct view of the cellular ecosystems that drive disease. Protein panels, acquisition platforms, and tissue contexts differ from one cohort to the next, and a coherent picture of tissue biology can only emerge from models trained to bridge this heterogeneity. Foundation models offer the natural path forward, but their training depends on large, harmonized corpora that span platforms, panels, and cohorts. No public resource currently approaches that scale; available data sit in small, narrowly scoped releases with incompatible formats and conventions. We present spora, a large-scale multimodal SP dataset of 12,596 standardized multiplex images from 5,254 patients across 31 cohorts and four major SP technologies. Curated with biomedical experts, spora provides model-training-ready formats together with cell and tissue segmentations, structured clinical metadata, and, where available, paired H&E enabling the training of virtual staining models. It is complemented by spora [io], a unified interface for tiling, sampling, and standardization across all cohorts and modalities, and spora [bench], a benchmark suite spanning cell- and patient-level tasks across 11 cohorts that places current SP foundation models on common ground against task-specific baselines for the first time. Data and documentation at: https://spora.epfl.ch (user: spora_reviewer password: Uh8aef0kiev9).
Stochastic Zeroth-Order Optimization Under Heavy-Tailed Noise
Taha EL BAKKALI EL KADI ⋅ El Mahdi Chayti ⋅ Qiuyi (Richard) Zhang ⋅ Imane Rahali ⋅ Omar Saadi
We study stochastic zeroth-order (ZO) optimization of smooth nonconvex objectives under heavy-tailed sample-gradient noise, a regime motivated by empirical evidence that gradient noise in modern machine learning can violate the bounded-variance assumption underlying classical ZO theory. While the first-order literature has established optimal rates under bounded $p$-th moment noise for $p \in (1,2]$, analogous high-probability guarantees for nonconvex ZO optimization remain largely unexplored. The ZO setting is not a direct corollary of first-order theory: first-order methods update using $\nabla F(x;\xi)$, the same object on which the noise assumption is imposed, whereas derivative-free methods only observe noisy function values and construct directional finite-difference estimates. Thus, weak-$L_p$ control of $\nabla F(x;\xi)-\nabla f(x)$ must first be transferred to scalar two-point directional estimates, a concentration step absent from standard first-order analyses. We propose the Robust Scalar-Clipped Zeroth-Order method (\textbf{RSC-ZO}), a two-point method that clips each scalar directional derivative before aggregation. Under sample-wise smoothness and a weak-$L_p$ tail bound on the sample-gradient noise, RSC-ZO finds an $\varepsilon$-stationary point with high probability using $ \widetilde{O}\\left( \frac{d^{\frac{p}{2(p-1)}}} {\varepsilon^{\frac{3p-2}{p-1}}} \right) $ noisy function evaluations, matching the optimal first-order $\varepsilon$-dependence. At $p=2$, our bound becomes $\widetilde{O}(d\varepsilon^{-4})$, matching the classical dimension-accuracy dependence for stochastic ZO methods under bounded-variance noise, which is typically established only in expectation. In contrast, our guarantee holds with high probability and under the strictly weaker weak-$L_2$ tail condition, which can allow infinite gradient-noise variance. We further extend the analysis to a momentum variant and quantify the resulting batch-size/stepsize tradeoff.
Tadpole: Autoencoders as Foundation Models for 3D PDEs with Online Learning
Qiang Liu ⋅ Felix Koehler ⋅ Benjamin Holzschuh ⋅ Nils Thuerey
We introduce Tadpole, a novel foundation model for three-dimensional partial differential equations (PDEs) that addresses key challenges in transferability, scalability to high dimensionality, and multi-functionality. Tadpole is pre-trained as an autoencoder on synthetic 3D PDE data generated by an efficient online data-generation framework. This enables large-scale, diverse training without storage or I/O overhead, demonstrated by scaling to an equivalent of hundreds of terabytes of training data. By autoencoding single-channel spatial crops, Tadpole learns rich and transferable representations across heterogeneous physical systems with varying numbers of state variables and spatial resolutions. Although pre-trained solely as an autoencoder, Tadpole can be efficiently applied for multiple downstream tasks beyond reconstruction, including dynamics learning and generative modeling. For dynamics learning, we propose a novel parameter-efficient fine-tuning strategy that integrates low-rank adaptation, latent-space transformations, and reintroduced skip connections, achieving accurate temporal modeling with a minimal number of trainable parameters. Tadpole demonstrates strong fine-tuning performance across various downstream tasks, highlighting its versatility and effectiveness as a foundation model for 3D PDE learning.
Many value-based deep reinforcement learning algorithms rely on target networks - lagged copies of the online network - to stabilize training. While effective, this mechanism introduces a fundamental stability-recency tradeoff: slower target updates improve stability but reduce the recency of learning signals, hindering convergence speed. We propose Target-Aligned Reinforcement Learning (TARL), a simple drop-in refinement for existing algorithms that emphasizes transitions for which the target and online network estimates are highly aligned. By focusing updates on well-aligned targets, TARL mitigates the adverse effects of stale target estimates while retaining the stabilizing benefits of target networks. We empirically demonstrate consistent improvements within discrete and continuous control algorithms across various benchmark environments without any hyperparameter tuning, including a 38.18% peak score gain on Atari-10, while incurring less than a 4% increase in wall-clock time.
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
Zhaoyang Chu ⋅ Jiarui Hu ⋅ Xingyu Jiang ⋅ Pengyu Zou ⋅ Han Li ⋅ Chao Peng ⋅ Peter O'Hearn ⋅ Earl Barr ⋅ Mark Harman ⋅ Federica Sarro ⋅ He Ye
We introduce TerminalWorld, a scalable data engine that automatically reverse-engineers high-fidelity evaluation tasks from "in-the-wild" terminal recordings. Processing 80,870 terminal recordings, the engine yields a full benchmark of 1,530 validated tasks, spanning 19 real-world categories, ranging from short everyday operations to workflows exceeding 50 steps, and covering 1,280 unique commands. From these, we curate a Verified subset of 200 representative, manually reviewed tasks. Comprehensive benchmarking on TerminalWorld-Verified across eight frontier models and six agents reveals that current systems still struggle with authentic terminal workflows, achieving a maximum pass rate of only 62.5%. Moreover, TerminalWorld captures real-world terminal capabilities distinct from existing expert-curated benchmarks (e.g., Terminal-Bench), with only a weak correlation to their scores (Pearson $r=0.20$). The automated engine makes TerminalWorld authentic and scalable by construction, enabling it to evaluate agents in real-world terminal environments as developer practices evolve. Data and code are available at https://github.com/EuniAI/TerminalWorld.
The Agentic Oversight Tax: Human Supervision of AI Agents Has a Cost that Must be Accounted For
Olivier Oullier
This position paper argues that human supervision of AI agents imposes a composite cognitive, affective and organizational cost that no existing framework jointly captures. We introduce the Agentic Oversight Tax (AOT): the aggregate burden imposed on a human supervisor by the structural mismatch between the operating characteristics of one or more AI agents and the finite neurocognitive capacities of human oversight. AOT decomposes into four measurable pillars: (i) monitoring time, (ii) handoff repair events, (iii) audit production effort and (iv) recovery effort under staged failures. We anchor each pillar in documented supervision failures from 2024 to 2026 and cross-validate against the MAST taxonomy of 1,600 multi-agent system breakdowns. We establish discriminant validity against 5 adjacent frameworks by identifying three phenomena all five structurally exclude: attribution burden, accountability gap and skill atrophy. Three falsifiable predictions distinguish AOT from these constructs. A three-tier measurement architecture (telemetry, self-report, neurophysiological sensing) specifies how the pillars can be assessed. We show that agentic AI architectures modeled on Kahneman’s System 1/System 2 metaphor generate an attribution bias that AOT instruments must control for. A pre-registered meta-analysis of 100+ human-AI experiments shows that the net negative utility threshold AOT defines is already crossed on average in decision-making tasks. Governance frameworks mandating human oversight without measuring its cost cannot verify whether that oversight remains effective.
The balance between feature learning and collapse in generative dynamical systems
Julian Brandon ⋅ Bruno Loureiro ⋅ N Alex Cayco Gajic ⋅ Arthur Pellegrino
Over the past decade, a growing body of work has established that implicit biases in learning dynamics fundamentally shape the solutions found by neural networks, governing their learned features and generalisation performance. However, despite these strides in deep learning theory, far less is known about representation learning in generative models based on dynamical systems, from normalising flows to diffusion. Importantly, continuous-time generative models can be prone to mode collapse, yet the learning dynamics underpinning such biases remain elusive. To address this, we use tools from dynamical systems and operator theory to show that bounds on the rank of the gradient of the model weights can explain how collapse occurs in generative models. We show that, in both variational inference and flow matching tasks, optimisation induces low-rank biases that encourage parsimonious representation learning but may also cause the learned distribution to collapse. Together these findings point to a trade-off between promoting feature learning while avoiding both memorisation and collapse in dynamical systems-based generative models.
The limbic navigation system as a hierarchical RNN
Zilong Ji ⋅ Huiwen Zhang ⋅ Krishna Gorantla ⋅ Neil Burgess
Path-integration-trained recurrent neural networks (RNNs) can reproduce grid-like and head-direction-like representations, but most existing models rely on low-dimensional velocity inputs and do not capture the hierarchical organisation of the brain's self-localisation circuit. We introduce a hierarchical recurrent neural network (HRNN) that separately processes angular velocity and linear speed to perform simultaneous angular and positional path integration. Trained on these tasks, the HRNN generalises to longer timescales than seen during training and develops neural response properties resembling all the key cell types: angular velocity cells, head direction cells, speed cells, conjunctive and pure grid cells. The learned circuit also recovers key motifs of the navigation system, including attractor dynamics and asymmetric connectivity patterns. Perturbation analyses reveal functional information flow across the hierarchy, consistent with experimental observations. Finally, introducing theta-paced mutual inhibition between two learned angular-velocity subpopulations enables the HRNN to reproduce the bidirectional theta sweeps observed in empirical data. These results establish hierarchical task-driven RNNs as a biologically interpretable framework for linking neural representations, circuit motifs, and population dynamics in spatial navigation, while enabling targeted in silico perturbations and the generation of experimentally testable predictions.
The State-Prediction Separation Hypothesis
Giovanni Monea ⋅ Nathan Godey ⋅ Kianté Brantley ⋅ Yoav Artzi
Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We formulate the \emph{state-prediction separation hypothesis}: disentangling the two roles yields better language modeling performance. We design a Transformer variant that uses separate computation streams to separate the two functions, and conduct pretraining experiments across various scales. Our experiments show that state-prediction separation consistently offers better data-- and compute--efficiency, improving validation loss and outperforming standard Transformers by 2--3 percentage points on average on downstream tasks. We also conduct extensive empirical analysis that rules out potential confounders and demonstrates the fundamental difference in the gradients our design entails.
The sublevel Flood bifiltration: towards scalable 2-parameter persistent homology
Mattéo Clémot ⋅ Julie Digne ⋅ Julien Tierny
Multi-parameter persistent homology is a rapidly developing branch of topological data analysis that improves the robustness of single-parameter persistent homology to outliers, while still capturing the metric characteristics of the data. However, a notable limitation is its lack of scalability. In this paper, we introduce a novel approach for efficiently computing 2-parameter persistent homology on large point sets. Our work extends the Flood filtration, originally developed for single-parameter persistence. Our construction, called the sublevel Flood bifiltration, offers a scalable approximation of the sublevel Flood bifiltration. We show that it benefits from theoretical stability properties and describe how to compute it efficiently. We demonstrate the performance of our approach in classification tasks on low-dimensional synthetic datasets, where density awareness is critical, as well as on real-world time series datasets.
Larger language models solve harder problems, but do they also produce proportionally better explanations? We study in-context explanation transfer across 40 open-weight models (360M--72B) and 5 receivers (360M--7B), using problems drawn from 13 benchmark sources and evaluated with domain-specific metrics. Our central finding is an \emph{articulability ceiling}: under equal token budgets, mid-range models are better teachers than frontier models. Transfer peaks at 170--220 words and then drops sharply beyond 300 words, falling below even the shortest explanations. Large models partly compensate through verbosity: longer explanations yield higher total transfer on average, but each additional word carries less pedagogical value, so a 360M model delivers more transfer value per word than a 72B model. We also find that RLHF improves pedagogical clarity: instruct models score better on LLM-judge evaluation than base models despite lower lexical overlap, and teacher correctness dominates all other factors. Code and logs are released with the paper.
Towards Fairness under Label Bias in Image Segmentation: Impact, Measurement and Mitigation
Aditya Parikh ⋅ Stella Christina Frank ⋅ Sneha Das ⋅ Aasa Feragen
Labeled datasets reflect the biases of their annotation pipelines, which sometimes introduce label bias: group-conditional label errors that cause systematic performance disparities across demographic subgroups. Label bias in image segmentation remains underexplored, as even detecting it typically requires clean, unbiased annotations, which are not readily available. We present a data-centric}adaptation of Confident Learning to segmentation, allowing detection of label bias directly in the training data without a clean, unbiased ground truth. By comparing the provided training labels to the model's confident predictions, we isolate directional errors that quantify the presence and nature of bias, where standard overlap metrics like Dice fail. We further show that label bias influences subgroup separability in the encoder's feature space, an artifact we leverage for bias mitigation rather than suppressing it. We evaluate three datasets, spanning from synthetic to real-life bias, showing how our framework reliably detects and mitigates bias without access to clean labels, achieving equitable performance across experimental conditions.
Towards Identifying Dominant Low-Rank Subspaces in Zeroth-Order Fine-Tuning
Jinjie Fang ⋅ Chengxun Jin ⋅ Yi Chang ⋅ Bin Gu
Adapting Large Language Models (LLMs) to downstream tasks is increasingly bottlenecked by the staggering memory overhead of first-order (FO) backpropagation. Zeroth-order (ZO) optimization has emerged as a compelling, memory-efficient alternative; however, it suffers from the curse of dimensionality, which introduces prohibitively high variance in gradient estimation. While LLM gradients empirically reside in low-rank subspaces, existing state-of-the-art ZO methods typically rely on randomly generated low-rank subspaces, which often fail to align with the underlying dominant gradient manifold, fundamentally limiting optimization efficiency. To bridge this gap, we propose AHZO, an efficient low-rank ZO fine-tuning algorithm that dynamically identifies dominant subspaces via the Average Historical Gradient (AHG). Motivated by the directional coherence of optimization trajectories, AHZO leverages the average historical ZO gradient over a period as a principled proxy for the true gradient. By applying singular value decomposition to AHG matrices, AHZO distills principal spectral signatures to construct low-rank bases that tightly align with the true dominant subspaces. Theoretically, we prove that AHZO significantly mitigates estimation variance, yielding superior convergence guarantees compared to random-subspace ZO methods. Extensive evaluations across diverse LLM architectures and benchmarks demonstrate that AHZO consistently outperforms existing ZO baselines, achieving performance highly competitive with memory-intensive FO fine-tuning.
Towards instance-dependent regret optimality in Episodic MDPs with Posterior Sampling
Victor Boone ⋅ Cyrille Kone ⋅ Waris Radji ⋅ Odalric-Ambrym Maillard ⋅ Dorian Baudry
We study regret minimization in finite-horizon episodic Markov Decision Processes (MDPs). While minimax-optimal algorithms are known, tractable approaches to instance-dependent optimality are still lacking. Motivated by this gap, we introduce $\pi_0$-PSRL, a variant of PSRL that uses posterior samples to decide when exploration is needed, while following a fixed reference policy $\pi_0$ during exploration episodes. This decouples the test for exploration, triggered when the sampled and empirical MDPs have different optimal policies, from the choice of the policy used to gather information. The resulting design addresses a limitation of standard PSRL, where the sampled optimal policy may not be the optimal choice for exploration. We prove instance-dependent regret bounds for $\pi_0$-PSRL, identifying the logarithmic exploration cost induced by $\pi_0$ and taking a step toward matching asymptotic lower bounds for episodic RL. Our proof techniques showcase a novel proof structure to derive problem-dependent regret bounds in episodic MDPs, and concentration results for Dirichlet random variables, that may be of independent interest.
Towards Understanding and Measuring Cognitive Atrophy in LLM Behaviour
Abeer Badawi ⋅ Moyosoreoluwa Olatosi ⋅ Negin Baghbanzadeh ⋅ Laleh Seyyed-Kalantari ⋅ Frank Rudzicz ⋅ R. Shayna Rosenbaum ⋅ Sara Pishdadian ⋅ Elham Dolatabadi
Recent incidents involving LLMs used for mental-health support reveal a critical evaluation gap: surface-level safety scores do not capture how models behave across realistic, emotionally sensitive interactions over time. Existing benchmarks measure knowledge, safety, or static response quality, but miss whether LLM interactions help users keep reflecting, coping, and making decisions themselves. We formalize this missing dimension as Cognitive Atrophy, a process-level behavioural pattern in AI-mediated mental-health support distinct from safety and helpfulness. To measure it, we introduce Cognitive Atrophy Bench, a clinically grounded benchmark built from 1,576 fully human-generated counseling conversations, 15,680 turns, and 42,230 responses from five LLMs. Three clinical psychology experts developed a 20-attribute schema spanning user context, response behaviour, and global risk flags; six trained clinical reviewers applied it with span-grounded evidence, producing 5,324 reviewer judgments. We further introduce the User-Input Risk Index (UIRI), the Cognitive Atrophy Risk Index (ARI), and trajectory summaries. Across five LLMs, models show a consistent moderate-to-high level of atrophy-aligned behaviour across single- and multi-turn settings. While models generally respond to overt safety cues, they adapt less reliably when users seek solutions or decisions. The dominant recurring patterns are directive advice, problem-solving, recommendation-heavy responses, topic shifts, and forms of validation that may reinforce dependence rather than reflection. Our work makes Cognitive Atrophy measurable and provides a foundation for auditing model behaviour in sensitive LLM conversations. All code and data are released.
Do transformers, when trained on sequential reasoning traces, build internal models of the underlying task? And if so, does the structure of those internal representations mirror the structure of the domain? We train an 8-layer transformer on Sudoku solving traces and perform a mechanistic analysis of its internal computation. We establish two results. First, the model builds a substructure world model: it does not represent the board state cell by cell, as a human analyst would expect, but organizes information around the rows, columns, and boxes that Sudoku's constraints act on. Second, we identify a naked-single circuit: a small set of dedicated neurons in the final MLP layer, each individually detecting when exactly one digit remains possible for a specific cell, and reliably promoting that digit. These findings show that the geometry of an emergent world model is shaped by the constraint algebra of the domain, not its surface presentation, and that the resulting decision circuit is sparse, monosemantic, and fully interpretable. More broadly, they demonstrate that mechanistic interpretability tools can recover an end-to-end algorithmic account of how a transformer solves a combinatorial reasoning task.
TurtLES: A Large-Scale Benchmark for Turbulent 3D Neural PDE Surrogates
Jarno Platenburg ⋅ Armand Kassaï Koupaï ⋅ Paola Cinnella ⋅ Patrick Gallinari
Machine-learning surrogates are increasingly used to accelerate computational fluid dynamics (CFD), yet progress is limited by the lack of benchmarks capturing realistic, time-dependent turbulent flows. We introduce TURTLES, a 13 TB dataset of high-fidelity implicit large-eddy simulations of three-dimensional turbulent cylinder wakes in the shear-layer transition regime. The dataset contains 400 trajectories of 400 time steps each, with 3–9 million points per frame sampled on irregular point clouds, spanning variations in geometry, Reynolds number, and angle of attack. Unlike existing datasets, TURTLES captures fully three-dimensional turbulence with an active energy cascade driven by vortex stretching, combining (i) dense irregular meshes, (ii) long-horizon temporal dynamics, and (iii) physically realistic turbulent behavior. We benchmark state-of-the-art neural operators on long autoregressive rollouts and identify key failure modes, including temporal instability, loss of small-scale energy, and challenges in learning on large irregular domains. TURTLES provides a new standard benchmark for evaluating surrogate models in high-resolution, industrial-scale CFD regimes.
Understanding axial attention in TSFMs through single location regression
Andrei Pantea ⋅ Aymeric Dieuleveut
Recent advances in Time Series Foundation Models (TSFMs) increasingly rely on axial attention mechanisms, but despite their empirical success, the theoretical properties of axial attention remain largely unexplored. To bridge this gap, we extend the Single Location Regression framework introduced in Marion et al. [2025] to 2D structured data, in order to study sparse information settings where a target depends on a small number of tokens. We formulate an analytically tractable axial attention-type predictor and rigorously analyze its statistical properties and training dynamics. We prove that under a specific asymptotic regime and with knowledge of the underlying oracle parameters, our predictor achieves Bayes optimality, whereas a class of semi-linear simplifications strictly fails to do so, underlining the role of the inner non-linearity. We prove that for a specific temperature scaling, Projected Gradient Descent converges globally to the optimal oracle parameters. Extensive numerical experiments validate our theoretical findings and highlight the critical role of inverse temperature scheduling in empirical convergence.
Understanding the Interplay between Memorization and Learning in Large Language Models
Bishwamittra Ghosh ⋅ Soumi Das ⋅ Qinyuan Wu ⋅ Mohammad Aflah Khan ⋅ Krishna Gummadi ⋅ Evimaria Terzi ⋅ Deepak Garg
We investigate foundational questions about the interplay between memorization and learning in large language models (LLMs). When an LLM generates a string, can we disentangle the roles of rote memorization and contextual learning? Rote memorization results in an LLM regurgitating specific details of a training string, while contextual learning leads to the generation of new strings that follow generalizable patterns of the underlying language. Related questions of interest include: can an LLM avoid memorization when optimally learning a language?, and if not, can we design training schemes to reduce memorization and improve learning? To address these questions, we propose contextual memorization, a new measure that precisely disentangles memorization from learning. Using this measure, we establish that memorization of some training strings is unavoidable when optimally learning a language. Finally, we propose memorization-aware training, a scheme that equalizes learning across all training strings and in doing so, reduces memorization and improves learning. Our conclusions are supported by extensive experiments on multiple LLMs using formal languages as a controlled testbed – enabling precise language specification, exact string sampling, and elimination of data contamination – and further validated on natural language datasets.
VACE: Learning Geometrically Structured Representations for Time Series Anomaly Detection
Alberto D Cencillo ⋅ Leonardo Concepción ⋅ Isaac Triguero ⋅ Julián Luengo
Anomaly detection in multivariate time series is a critical task across a wide range of real-world applications, where abnormal behaviour is rare, labels are unavailable, and the cost of a miss is high. The central challenge is learning a characterisation of normality precise enough to flag deviations. Representation self-supervised learning, typically through contrastive approaches, addresses this by embedding temporal patches into a latent space where normality occupies a well-defined region, with anomalies detected by geometric deviation. However, contrastive approaches shape this space indirectly through pair-sampling heuristics, providing no explicit control over the geometric structure that distance-based scoring requires. This means how tightly normal representations are grouped, and whether distances are directionally meaningful. We present VACE ($\textbf{V}$elocity-$\textbf{A}$ligned $\textbf{C}$hannel $\textbf{E}$mbeddings), a self-supervised anomaly detection method that represents normality as a compact, directionally coherent region in the embedding space. To this end, VACE trains a channel-aware encoder through a velocity-consistency objective, with no negatives and no synthetic anomalies, so that normal trajectories are locally smooth and aligned. At test time, a Mahalanobis positional score and a velocity-bank directional score are combined multiplicatively, flagging points that are simultaneously off-distribution and dynamically atypical. Despite its simplicity, VACE achieves state-of-the-art performance on TSB-AD-M under rigorous evaluation, significantly outperforming more complex methods trained on substantially larger budgets.
Bayesian Additive Regression Trees (BART) is a nonparametric Bayesian regression technique based on an ensemble of decision trees. It is part of the toolbox of many statisticians. The overall statistical quality of the regression is typically higher than other generic alternatives, and it requires less manual tuning, making it a good default choice. However, it is a niche method compared to natural competitors such as XGBoost, due to the longer running time, making sample sizes above 10000–100000 a nuisance. We present a GPU-enabled implementation of BART, faster by up to 100x in a given GPU vs. CPU comparison, making BART fast enough to be a viable alternative to non-Bayesian algorithms on large datasets, and showing that BART parallelizes better than its non-Bayesian counterparts. This implementation is available in the Python package bartz.
Voronoi-Markov chain and spatial entropy for point pattern analysis
James C Mathews ⋅ Francisco Couzo ⋅ Aleksandr Petrov ⋅ Avijit Chatterjee ⋅ Saad Nadeem
In this paper, we revisit the concept of entropy in the spatial context with the aim of deriving computable and interpretable metrics for point pattern analysis in domains such as histopathology. We discuss stochastic point process assumptions like Poisson homogeneity and Simple Sequential Inhibition (SSI) and review established results for Voronoi and Delaunay tilings of point sets for their potential to provide null distributions for rigorous empirical hypothesis testing. We present (1) a novel ``dense basin entropy'' defined in terms of the Ord distribution supported on Voronoi basins, which is shown to be sensitive to clustering, and (2) a related Markov chain designed as a simplification or approximation of Brownian motion in the underlying planar domain. We prove a lower bound on the dense basin entropy in the SSI regime (originally discovered in empirical experiments) and show that the entropy rate and spectral gap of the Markov chain lead to sharp discrimination metrics. Finally, we demonstrate the generalization to multiple-set analysis via aggregation of one set over the Markov invariant measure for another. Simulations and experiments in histopathology show that our proposed metrics offer insights complementary to other instantiations of entropy in spatial analysis.
What are Key Factors for Updates in RL for LLM Reasoning?
Peidong Wang ⋅ Demi Wang ⋅ Xufang Luo ⋅ Jiahang Xu ⋅ Xiaocui Yang ⋅ Shi Feng ⋅ Yuqing Yang ⋅ Dongsheng Li
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising framework for enhancing the reasoning ability of large language models. However, much of the existing work is guided by heuristic intuition, leading to divergent algorithmic choices, even contradictory ones that nevertheless report empirical gains. To better understand this phenomenon, we conduct a theoretical analysis of RLVR updates. Our study reveals that differences in off-policy degree, determined by the number of gradient steps per rollout, substantially affect the distribution of importance sampling ratios and their clipping behavior, thereby altering which tokens dominate the update. Building on this insight, we characterize gradient expectation as the central quantity governing update dynamics and analyze the roles of token probability, advantage, and importance sampling ratio. Motivated by these findings, we propose Adaptive Clip Policy Optimization (ACPO), which adjusts clipping boundaries across token groups according to the empirical variance of their importance sampling ratios. Experiments on 3B and 7B models across diverse reasoning benchmarks, spanning mathematical problem solving, tabular QA, and logic puzzles, demonstrate that ACPO outperforms strong baselines such as DAPO and CISPO. These results demonstrate that principled, analysis-driven approaches yield more robust and effective RLVR methods. Code is available in: https://anonymous.4open.science/r/ACPO
When Actions Matter: Causal Affordances for Long-Horizon Credit Assignment in World Models
Pradeep Kumar Banerjee ⋅ Frank Röder ⋅ Nihat Ay
Long-horizon tasks require identifying which past decisions actually mattered for distant outcomes. In deep achievement chains, a single early decision can determine success thousands of steps later, far beyond the reach of short-horizon imagination. We introduce WHAM, a world-model agent for causal credit assignment in long-horizon tasks. WHAM ascends Pearl's causal hierarchy within a learned world model to identify decision-critical bottleneck states, prioritizes replay and imagination around them, and propagates value through sparse bottleneck chains using replay-derived returns. A hierarchical bottleneck critic bridges credit across thousands of timesteps with provably controlled error. On Crafter, WHAM achieves $67.1\%$ ($+8.2$ over a state-of-the-art baseline), with gains that grow systematically with achievement-chain depth: diamond collection rises $17\times$ to $8.5\%$, and the best seed reaches $12.5\%$, exceeding the human expert rate. Long-horizon credit flows where actions shape distant outcomes.
When Prompts Override Vision: Instruction-Induced Hallucinations in LVLMs
Pegah KHAYATAN ⋅ Jayneel Parekh ⋅ Arnaud Dapogny ⋅ Mustafa Shukor ⋅ Alasdair Newson ⋅ Matthieu Cord
Despite impressive progress in capabilities of large vision-language models (LVLMs), these systems remain vulnerable to hallucinations, i.e., outputs that are not grounded in the visual input. Prior work has attributed hallucinations in LVLMs to factors such as limitations of the vision backbone or the dominance of the language component, yet the relative importance of these factors remains unclear. To resolve this ambiguity, We propose HalluScope, a benchmark to better understand the extent to which different factors induce hallucinations. Our analysis indicates that hallucinations largely stem from excessive reliance on textual priors and background knowledge, especially information introduced through textual instructions. To mitigate hallucinations induced by textual instruction priors, we propose HalluVL-DPO, a framework for fine-tuning off-the-shelf LVLMs towards more visually grounded responses. HalluVL-DPO leverages preference optimization using a curated training dataset that we construct, guiding the model to prefer grounded responses over hallucinated ones. We demonstrate that our optimized model effectively mitigates the targeted hallucination failure mode, while preserving or improving performance on other hallucination benchmarks and visual capability evaluations. To support reproducibility and further research, we will publicly release our evaluation benchmark, preference training dataset, and code.
Where to Approximate in Neurosymbolic Inference?
Samy Badreddine ⋅ Emile van Krieken ⋅ Luciano Serafini ⋅ Antonio Vergari
Probabilistic neurosymbolic methods rely on weighted model counting (WMC) to combine neural predictors with symbolic constraints. Exact computation of the WMC, a #P-hard problem, typically scales poorly, so many methods resort to approximations. By unifying existing approaches under a three-step bottom-up approximate compilation pipeline, we study where approximation budget is best allocated. Inspired by tensor networks, we instantiate two steps with tensor train decompositions: we derive new pipelines that exactly multiply approximated factors and recompress their products via SVD- or interpolation-based schemes. On Sudokus of increasing size, these pipelines produce WMC estimates that are many orders-of-magnitude more accurate than prior methods. Yet, when it comes to neurosymbolic learning, even crude approximations reach competitive accuracy, suggesting that approximation quality matters far more for inference fidelity than for downstream learning. Our pipelines open significant design space for effective approximate compilation of challenging probabilistic reasoning tasks.
Where to Look Matters: Rethinking Sub-Volume Sampling in 3D Medical Self-supervised Learning
Junkai Liu ⋅ Le Zhang
Self-supervised learning on 3D medical images commonly relies on random sub-volume sampling during pretraining. However, random cropping implicitly treats all spatial regions as equally worth observing, despite anatomical priors and spatial redundancy that make regions inherently unequal in their value for representation learning. To this end, we first show through controlled crop-quality manipulations that better sub-volumes lead to better downstream transfer, revealing the observation policy as an overlooked bottleneck in 3D medical SSL. Motivated by this finding, we revisit sub-volume sampling as a where-to-look problem. We propose VolumeProbe, a plug-and-play framework that replaces random cropping with an offline coarse-to-fine probe bank guided by anatomy-awareness, informativeness, and diversity. We instantiate VolumeProbe as a lightweight Lite variant using geometric and image-statistical cues, and a diagnostic Feat variant using frozen visual features, to examine whether lightweight volume-intrinsic priors suffice or richer semantic cues reliably provide additional value. Extensive experiments across multiple 3D medical SSL backbones and downstream tasks demonstrate that VolumeProbe-Lite consistently improves transfer performance over random sampling. Mechanistic analyses further attribute these gains to more anatomically plausible and informative observations, reduced spatial redundancy, more efficient learning dynamics, and more structured representations.
Who caused $Y$? Local identifiability for learning causal parents
Felix Schur ⋅ Sorawit Saengkyongam ⋅ Jonas Peters
We study identifiability and estimation of the direct causes, that is, the parents, of a designated target $Y$ in a structural causal model using only observational data. Our main assumption is an additive-noise model for $Y$: the value of $Y$ equals some function of its parents plus noise that is independent of those parents. We allow for general functional relationships and hidden confounding among all other variables. We prove that under a `no-backward' additive-noise condition the true parent set, $\mathrm{PA}(Y)$, is identified from the observational distribution by a simple population principle: among all candidate variable sets that make the regression residual of $Y$ independent of the regressors, choose the one with the smallest residual variance, and—if several tie—the smallest set. To justify the ``no-backward'' condition, we propose a novel identifiability scheme based on identifiability witnesses: we prove that in finite-dimensional analytic model classes, for which we have an identifiability witness (that is, a single identifiable model), residual independence for sets containing descendants of \(Y\) holds only on Lebesgue-null exceptional parameter sets. This scheme is strong enough to recover known identifiability results. For finite data, we propose Independent Risk Minimization (IndRM), which offers a simple, local, and non-interventional route to isolating the direct causes of $Y$ in multivariate systems, avoiding global model assumptions and full-graph search.
Goal recognition is the task of inferring an agent's intended goal from its observed behavior. Existing work has largely focused on high-level symbolic recognition, abstracting away the continuous motion through which agents reach their goal. Methods that do reason over real-world trajectories, rely on an explicit dynamics model that is hard to acquire and rarely transfers across settings. Neither approach explicitly learns how agents move toward each candidate goal, limiting accuracy on ambiguous trajectory prefixes where such knowledge is most needed. We present WoRD (Waypoint Recognition with Diffusion), which uses a learned trajectory prior to infer the agent's goal from partial observations. At each step, WoRD runs one warm-started denoising chain per candidate goal, steers each chain toward the observed trajectory through a closed-form log-likelihood gradient, and reweights the resulting samples by Bayes' rule into a calibrated posterior over goals. We evaluate WoRD on simulated robot settings, pedestrian trajectories, and vehicle trajectories using four metrics targeting distinct failure modes. WoRD is the only method that improves all four metrics jointly, including the safety-critical confident-and-wrong rate. WoRD runs under a single inference procedure and preserves calibrated multimodal. beliefs in complex, ambiguous environments where filter- and classifier-based observers collapse to a single mode or grow overconfident.
As autonomous agents increasingly execute end-to-end tasks under fixed monetary budgets, the pressing open question shifts from whether the budget is respected, to how to spend it effectively. Existing budget-aware methods typically control reasoning step-by-step within a single agent, or learn resource allocation policies via RL. None address how to split a budget across the composing phases of a multi-agent pipeline at inference time. We propose ZEBRA, a zero-shot framework that reduces multi-phase budget allocation to a continuous nonlinear knapsack problem: an LLM controller estimates per-phase utility curves, and a water-filling search on the Lagrange multiplier returns the per-phase split. Additive and multiplicative aggregations are unified under the same solver. On a $150$-task APPS coding benchmark, both ZEBRA variants outperform LLM-direct (budget allocation directly by an LLM) on every aggregate metric. At a budget of $\alpha = 0.5$ of the unconstrained spend, ZEBRA recovers $94.4$% of unconstrained quality, versus $88.1$% for LLM-direct. The advantage is statistically significant and transfers beyond coding: on a $3$-phase HotpotQA pipeline, ZEBRA beats LLM-direct by $14.3$pp, with allocations empirically robust to curve-estimation noise. Also, ZEBRA arrives at a different budget split (near-balanced) compared to the APPS one (skewed towards a refinement phase), showing adaptation to the pipeline structure. More broadly, we show that lightweight algorithmic guidance at inference time can improve the economic behavior of autonomous multi-agent systems.