Skip to yearly menu bar Skip to main content


Session

Paris Poster Session 4

Paris Poster Hall
Fri 11 Dec 3:30 a.m. AEDT — 5:30 a.m. AEDT
Abstract:
Chat is not available.


Adapting to Conflict: Equilibrium Structure and Adaptive Learning in Harmonic Games

Davide Legacci ⋅ Panayotis Mertikopoulos ⋅ Bary Pradelski

Zero-sum games are the textbook model of strategic competition—yet they can fail to capture conflict even in simple settings. *Harmonic games* provide a more robust framework for opposed interests: they are invariant under strategic equivalence, and naturally extend beyond pairwise competition. This generality comes at a structural cost. We show that, in odd-dimensional harmonic games, the set of totally mixed Nash equilibria is a generically non-convex real algebraic variety of dimension at least one, extending to the boundary of the strategy space. As a result, learning becomes delicate: although finely tuned methods are known, whether adaptive, parameter-agnostic learning is possible has remained open. We resolve this question in the affirmative. We introduce a flexible extrapolation-based variant of follow-the-regularized-leader (FTRL+) that converges to Nash equilibrium while achieving order-optimal regret: $\mathcal{O}(1)$ in self-play and $\mathcal{O}(\sqrt{T})$ against arbitrary opponents. The method is fully adaptive—each player updates from locally observable information alone, without access to global problem parameters or shared signals—yielding the first parameter-agnostic learning dynamics provably convergent in this class.


Adaptive Multi-view Graph Contrastive Learning via Fractional Continuous Dynamics

Yanan Zhao ⋅ Feng Ji ⋅ Jingyang Dai ⋅ Jiaze Ma ⋅ Keyue Jiang ⋅ Kai Zhao ⋅ Wee Peng Tay

Graph contrastive learning (GCL) learns node and graph representations by contrasting multiple views of the same graph. Existing methods predominantly rely on a small set of fixed, handcrafted views---typically a local and a global perspective---which fundamentally constrains their capacity to capture multi-scale structural patterns. We present an augmentation-free, multi-view GCL framework grounded in fractional-order continuous dynamics. By systematically varying the fractional derivative order $\alpha\in(0,1]$, our encoders produce a continuous spectrum of semantically distinct views: small values of $\alpha$ induce localized, memory-attenuated feature propagation, whereas values approaching $1$ recover broader, global aggregation. Crucially, we treat $\alpha$ as a learnable parameter, enabling the model to automatically identify informative views in a data-driven manner, without resorting to manual view engineering. Extensive experiments on standard node- and graph-level benchmarks demonstrate that the resulting representations are more robust and expressive, consistently outperforming state-of-the-art GCL baselines.


Affine Tracing: A New Paradigm for Probabilistic Linear Solvers

Disha Hegde ⋅ Marvin Pförtner ⋅ Jon Cockayne

Probabilistic linear solvers (PLSs) return probability distributions that quantify uncertainty due to limited computation in the solution of linear systems. The literature has traditionally distinguished between Bayesian PLSs, which condition a prior on information obtained from projections of the linear system, and probabilistic iterative methods (PIMs), which lift classical iterative solvers to probability space. In this work we show this dichotomy to be false: Bayesian PLSs are a special case of non-stationary affine PIMs. In addition, we prove that any realistic affine PIM is calibrated. These results motivate a focus on (non-stationary) affine PIMs, but their practical adoption has been limited by the significant manual effort required to implement them. To address this, we introduce affine tracing, an algorithmic framework that automatically constructs a PIM from a standard implementation of an affine iterative method by passing symbolic tracers through the computation to build an affine computational graph. We show how this graph can be transformed to compute posterior covariances, and how equality saturation can be used to perform algebraic simplifications required for computation under specific prior choices. We demonstrate the framework by automatically generating a probabilistic multigrid solver and evaluate its performance in the context of Gaussian process approximation.


A General Filter-Enhanced Approach to Smartphone Hyperspectral Imaging

Daniil Reutsky ⋅ Daniil Vladimirov ⋅ Yasin Mamedov ⋅ Georgy Perevozchikov ⋅ Egor I Ershov ⋅ Nancy Mehta ⋅ Radu Timofte

Hyperspectral reconstruction from RGB imagery offers an affordable alternative to costly hyperspectral acquisition devices. However, its accuracy is fundamentally constrained by the limited spectral resolution of standard RGB cameras. In this paper, we propose new approach that leverage the multi-camera systems integrated into most modern smartphones and enhance their capabilities by modulating their spectral response functions using carefully selected spectral filters, enabling a smartphone to capture richer spectral information. We propose a neural network able to exploit auxiliary RGB images with diverse spectral characteristics, enabling more accurate hyperspectral reconstruction compared to conventional single-image methods. Training is supported by a newly collected dataset, Doomer, which consists of multiple misaligned RGB images alongside corresponding hyperspectral data. Finally, our method generalizes across different smartphone devices through a easy to implement calibration procedure, ensuring practical applicability and scalability.


A Generalized Tikhonov Layer for Interpretable-by-design Graph Neural Networks

Nicolas Tremblay ⋅ Filippo Maria Bianchi ⋅ Benjamin Ricaud

We propose the Tikhonov layer, a graph neural network layer that is interpretable by design: once trained, its learned parameters directly reveal which node features and which aspects of the graph topology were leveraged for prediction. In practice, the layer's propagation matrix takes the closed-form $R = (p(L)+Q)^{-1}Q$, where $L$ is the normalized graph Laplacian, $Q = \mathrm{diag}(q_1,\ldots,q_n)$ a learnable diagonal matrix of positive node-importance scores, and $p(\cdot)$ a learnable polynomial. For any input feature $x$, the layer output $Rx$ is the exact minimizer of a generalized graph Tikhonov problem that trades off node-level data fidelity against a topology-driven regularization penalty. The learned pair $(\{q_i\},p)$ constitutes a built-in explanation: large $q_i$ indicates that node $i$'s own features drive the prediction, while small $q_i$ signals reliance on the local graph topology; the shape of $p$ reveals whether homophily, heterophily, or a band-pass response is exploited. Expressivity is preserved by routing complexity through a dedicated, arbitrarily deep $Q$-network that produces the importance scores, while the Tikhonov layer itself remains transparent. We prove that distinct explanations necessarily yield distinct propagation operators, preventing degenerate explanations by design. Additionally, the Tikhonov layer provides, in a single layer, a global receptive field under mild support-connectivity conditions, mitigating both oversmoothing and oversquashing. Experiments on standard graph classification benchmarks confirm that the model matches (and sometimes outperforms) opaque baselines while producing interpretable and faithful explanations.


AgentKernelArena: Benchmarking Performance and Generalization of AI Coding Agents on GPU Kernels Optimization

Sharareh Younesian ⋅ Wenwen Ouyang ⋅ sina rafati ⋅ Mehdi Rezagholizadeh ⋅ Sharon Zhou ⋅ Ji Liu ⋅ Yue Liu ⋅ Yuchen Yang ⋅ Hao Li ⋅ Ziqiong Liu ⋅ Dong Li ⋅ Vikram Appia ⋅ Zhenyu Gu ⋅ Emad Barsoum

GPU kernel optimization is increasingly critical for efficient deep learning systems, but writing high-performance kernels still requires substantial low-level expertise. Recent AI coding agents can iteratively read code, invoke compilers and profilers, and refine implementations, yet existing kernel benchmarks evaluate single LLM calls rather than full agent workflows, and none include both kernel-to-kernel optimization and unseen-configuration generalization testing. We present AgentKernelArena, an open-source benchmark for measuring AI coding agents on GPU kernel optimization. The benchmark contains 196 tasks spanning HIP-to-HIP optimization, Triton-to-Triton optimization, and PyTorch-to-HIP translation, and evaluates complete agent workflows in isolated workspaces using gated compilation, correctness, and performance checks, centralized scoring and an unseen-configuration generalization protocol that tests whether optimizations transfer to input configurations the agent never observed. Across production agents including Cursor Agent, Claude Code, and Codex Agent, we find near-perfect compilation and high correctness rates on most task categories, with the strongest configurations achieving mean speedups of up to 6.89$\times$ on PyTorch-to-HIP, 6.69$\times$ on HIP-to-HIP, and 2.13$\times$ on Triton-to-Triton tasks. Our unseen-configuration evaluation shows that HIP-to-HIP and Triton-to-Triton optimizations largely transfer to unseen input shapes, while PyTorch-to-HIP exhibits substantial correctness drops, indicating that agents generating kernels from scratch frequently hardcode shape-specific assumptions. AgentKernelArena is designed as a modular, extensible framework for rigorous evaluation of agentic GPU kernel optimization across agents, tasks, and hardware targets. Code and data are available at \url{https://anonymous.4open.science/r/AgentKernelArena_Neurips2026-F646}.

Concurrent multi-agent coding promises division of labor across modules, robustness through redundancy, and parallel exploration at the natural granularity of multi-file projects. Realtime collaborative editing protocols solve this coordination problem for human teams via Conflict-free Replicated Data Types (CRDTs), but the LLMs underneath generate one token at a time and existing multi-agent coding systems inherit this serial limit: they either sequence agents through phase handoffs or pool independent samples without coordination, and a single agent abandons up to 70% of hard tasks with a one-file stub-and-exit. AgentRoom is a realtime collaborative editing protocol for concurrent coding agents: a runtime layer exposing file-level claim, status, and broadcast as MCP tools on top of a CRDT-merged shared filesystem. Across five frontier coding-CLI models on four backend coding tasks, alongside Python DevBench and Rust+axum cross-language checks, AgentRoom ×2 suppresses Solo abandonment on the CLI-stable models and tightens run-to-run variance. Two matched-compute contrasts isolate the cause: parallel-merge vs AgentRoom shows a positive mean LLM-judge contrast, and a bundle probe attributes the gain primarily to the MCP coordination layer rather than to parallelism or substrate convergence alone. Coordination, not parallelism or CRDT-merge, is the load-bearing engineering.


Algebraic Machine Learning: learning as computing subsets of the subdirect decomposition from Abstract Algebra

Fernando Martin-Maroto ⋅ Nabil Abderrahaman-Elena ⋅ David Mendez ⋅ Gonzalo de Polavieja

We propose a machine-learning framework based on the principle that learning can be achieved through algebraic decomposition rather than numerical optimization. A learning task is encoded as axioms of an algebra, and learning proceeds by computing a subdirect decomposition of the algebra into subdirectly irreducible components. We show that suitable subsets of these components form models that generalize to unseen data because they approximate the underlying rule in the training data. The method has no dataset-specific hyperparameters to tune and shows no evidence of classical overfitting in our experiments, so we run it without a validation dataset. We compare against multilayer perceptrons as generic parametric baselines without task-specific inductive biases. Across 24 classification datasets, the algebraic method is applied only once to each training set without hyperparameter tuning, whereas the multi-layer perceptrons are selected from many configurations using validation data or cross-validation. Nevertheless, the resulting test accuracies are statistically indistinguishable. Moreover, adding a statistical readout on top of the algebraic representation, using untuned logistic regression, improves over the best multilayer perceptrons selected by validation or cross-validation. More broadly, because task axioms may encode examples, formal knowledge, or both, this approach provides a common algebraic framework and the same algorithm for data-driven problems such as classification, formally specified problems such as computing Hamiltonian cycles from specifications, and mixed settings combining data with formal constraints.


An Educated Guess: Deriving Statistically Aligned Gradient Estimates for Zeroth-Order Optimization

Justus F Hübotter ⋅ Serge Thill ⋅ Marcel A. J. van Gerven ⋅ Nasir Ahmad

As novel hardware architectures are explored for the purposes of machine learning, including neuromorphic chips, asynchronous systems, and ASICs, effective learning algorithms are highly desirable. While backpropagation is the standard for differentiation on standard hardware, its requirement for global synchrony and exact differentiability limits the exploration of exotic algorithms on non-traditional substrates and devices. Zeroth-order (ZO) optimization methods, such as SPSA and weight perturbation, offer a compelling alternative by estimating gradients through inference-only passes; however, these methods have historically suffered due to the "curse of dimensionality," where random perturbations become increasingly misaligned with the true gradient in high-dimensional spaces. In this work, we move beyond blind random noise by deriving zeroth-order search directions from activation geometry available during the forward pass. This leads to a family of Spectral Zeroth-Order methods, SZO-$\kappa$, whose directions solve local response-based gradient-guessing objectives under explicit stochastic assumptions. These methods spend perturbation budget on feature-supported directions that are more likely to produce informative loss responses. To further bridge the gap with first-order methods, we study covariance corrections and probe stabilizers, including search direction orthogonalization and bias corrections that reduce estimator error. Using multilayer and convolutional layers and networks, we analyze gradient alignment, bias-variance structure, and training stability of these methods relative to standard perturbation baselines and backpropagation.


An Information-Theoretic Evaluation Framework for Benchmark and Model Diagnosis in Knowledge Tracing

Houru Jiang ⋅ Zixi Wang ⋅ Tengteng Cheng ⋅ Xueyi Li ⋅ Mingliang Hou ⋅ Jiaqi Zheng ⋅ Renqiang Luo ⋅ Teng Guo ⋅ Zitao Liu

Knowledge tracing (KT) models are predominantly evaluated using aggregate metrics such as area under the curve (AUC) and accuracy. However, these global scores obscure where the remaining errors originate and fail to indicate whether a benchmark is approaching saturation. While estimating a global theoretical performance limit is challenging in realistic KT settings, it is possible to quantify local predictability. To address this, we propose an information-theoretic evaluation framework for KT benchmark diagnosis. We utilize Context Tree Weighting (CTW) as an operational causal anchor to estimate the Local Irreducible Uncertainty (LIU) of each student interaction. By projecting predictions onto this shared uncertainty coordinate, we evaluate model performance gains across distinct entropy bands rather than only at the global level.Comprehensive evaluations on NIPS Task 3/4 and Algebra 2005 reveal that model improvements are highly non-uniform. While the largest gains achieved by modern KT models consistently occur in high-entropy regions—suggesting that datasets are not yet exhausted—our framework also identifies instances where models capture inherently random noise. By surfacing these local modeling failures alongside genuine gains, this approach provides a rigorous diagnostic tool to pinpoint both the potential and the limitations of current KT benchmark and model diagnosis.


A Revisit of Hamiltonian Monte Carlo Efficiency on Bayesian Neural Networks

Cuong Ngoc Nguyen ⋅ Lam Ho ⋅ Vu Dinh ⋅ Georgios Karagiannis ⋅ Cuong V. Nguyen

Bayesian neural networks (BNNs) offer principled uncertainty quantification, yet sampling from their posteriors via Hamiltonian Monte Carlo (HMC) remains challenging. While recent work theoretically identified the efficiency degradation of ReLU-based networks, smooth unbounded activations (e.g., GELU, Swish, Mish), which are prevalent in modern architectures, were assumed to be efficient due to their smoothness. In this work, we challenge this existing justification by deriving explicit local error formulas for the leapfrog integrator. Our analyses reveal that, in addition to differentiability, third-order curvatures of the potential energy also have a significant influence on HMC efficiency of BNNs inference. Empirically, we found that even smooth unbounded activations may degrade sampling performance as severely as piecewise linear activations. Thus, optimally tuning the step size for these networks may require a comparable empirical step size scaling as their piecewise counterparts. This result sheds light on the implications for architectural design in Bayesian deep learning.


Articraft: An Agentic System for Scalable Articulated 3D Asset Generation

Matt Zhou ⋅ Ruining Li ⋅ Xiaoyang Lyu ⋅ Zhaomou Song ⋅ Zhening Huang ⋅ Chuanxia Zheng ⋅ Christian Rupprecht ⋅ Andrea Vedaldi ⋅ Shangzhe Wu

A bottleneck in learning to understand articulated 3D objects is the lack of large and diverse datasets. In this paper, we propose to leverage large language models (LLMs) to close this gap and generate articulated assets at scale. We reduce the problem of generating an articulated 3D asset to that of writing a program that builds it. We then introduce a new agentic system, Articraft, that writes such programs automatically. We design a programmatic interface and harness to help the LLM do so effectively. The LLM writes code against a domain-specific SDK for defining parts, composing geometry, specifying joints, and writing tests to validate the resulting assets. The harness exposes a restricted workspace and interface to the LLM, validates the resulting assets, and returns structured feedback. In this way, the LLM is not distracted by details such as authoring a URDF file or managing a complex software environment. We show that this produces higher-quality assets than both state-of-the-art articulated-asset generators and general-purpose coding agents. Using Articraft, we build Articraft-10K, a curated dataset of over 10K articulated assets spanning 245 categories, and show its utility both for training models of articulated assets and in downstream applications such as robotics simulation and virtual reality.

We propose a closed-form spectral framework for relative log-density estimation in linearly parameterized probabilistic models, including unnormalized and conditional models. This is achieved by representing the Kullback-Leibler (KL) divergence as an integral of weighted chi-squared divergences, converting KL estimation into a family of least-squares problems. We derive an explicit spectral formula based only on first- and second-order feature moments, yielding closed-form estimators of both divergences and log-density potentials for fixed features. The framework extends to a broad class of $f$-divergences and can be combined with kernelization or feature learning with neural networks. We prove convergence guarantees for the resulting estimators and empirically compare them on synthetic data with optimization-based variational formulations, including logistic and softmax regression for normalized conditional models.


Attention-based PCA

Rodrigo Maulen Soto ⋅ Claire Boyer

We study attention mechanisms through the lens of a canonical unsupervised problem: principal component analysis (PCA). We show that, when trained on Gaussian data, both softmax and linear attention layers learn parameters that align with the principal eigenvectors of the covariance matrix, thereby establishing a direct and explicit connection with PCA. Our analysis covers both finite and infinite prompt regimes. In the infinite-prompt limit, we prove convergence to globally optimal solutions aligned with the leading spectral direction, while in the finite-prompt setting we show that the same behavior emerges up to sampling effects. We further extend the analysis to an in-context setting with spiked Wishart covariances, where attention successfully recovers the underlying signal direction. These results demonstrate that attention inherently performs PCA-like computations under unsupervised objectives, providing a theoretical foundation for its representation-learning capabilities.


Attention Trajectories as a Diagnostic Axis for Deep Reinforcement Learning

Charlotte Beylier ⋅ Hannah Selder ⋅ Arthur Fleig ⋅ Simon M. Hofmann ⋅ Nico Scherf

The emergence and evolution of feature reliance in deep reinforcement learning agents remain poorly understood. Here, we introduce a methodological framework for analyzing the learning process through quantitative analysis of saliency maps. This approach aggregates saliency information at the object and modality level into hierarchical attention profiles, quantifying how agents allocate attention over time, thereby forming attention trajectories throughout training. These profiles are then compared across controlled conditions, connected to behavioral measurements and reproduced with different saliency methods to assess the robustness of the findings. Applied to Atari 2600 benchmarks, custom Pong environments, and biomechanical user simulations in visuomotor tasks, this framework uncovers algorithm-specific attention biases, diagnosed unintended reward-driven strategies, and overfitting to redundant sensory channels. These patterns correspond to measurable behavioral differences, demonstrating empirical links between attention profiles, learning dynamics, and agent behavior. The results establish attention trajectories as a promising diagnostic axis for tracing how feature reliance develops during training and for identifying biases and vulnerabilities invisible to performance metrics alone.


Beyond Masked Sparsity: SNACK Enables Truly Sparse Neural Networks on GPU

Jafar Badour ⋅ Elena Mocanu ⋅ Maurice Keulen

Deep neural networks continue to grow in parameter count, driving up training and inference cost on GPUs. Sparse neural networks and Dynamic Sparse Training (DST) promise to reduce these costs, but most implementations rely on binary masks over dense tensors and recover little of the theoretical compute, memory, or energy savings. We propose SNACK, a truly sparse GPU layer that stores and computes only non-zero connections. SNACK exposes a simple PyTorch API for restructuring connections and backpropagating gradients entirely in the sparse paradigm, and uses a custom COO-format SpMM CUDA kernel with a batch-to-Streaming-Multiprocessor mapping tuned for the small-batch, high-sparsity regime typical of large-model training and single-stream inference. At the kernel level, SNACK is up to $7\times$ faster than the masked dense baseline (Dense+Mask) and competitive with cuSPARSE, Sputnik, and FlashSparse at 95\% sparsity. At 90\% sparsity, a single SNACK layer accelerates training by $8\times$ and $3.7\times$, and inference by $4\times$ and $2\times$, over Dense+Mask and fully dense layers, respectively, while using 25\% less memory than dense and substantially less energy. End-to-end, SNACK reduces GPT-2 peak training memory by up to 40\% and graph-style inference latency by $4.8\times$ over Dense+Mask at 99\% sparsity.

Federated fine-tuning of large language models is commonly formulated as a parameter aggregation problem. However, even parameter-efficient methods require transmitting large collections of trainable weights, assume aligned architectures, and rely on white-box access to model parameters. As model sizes continue to grow and deployments become increasingly heterogeneous, these assumptions become progressively misaligned with practical constraints. We consider an alternative formulation in which collaboration is mediated through model behavior rather than parameters. Clients fine-tune local models on private data and exchange generated outputs on a shared, public prompt set. The server maps these outputs into a semantic representation space, forms a per-prompt semantic consensus, and returns pseudo-labels for further local fine-tuning. This formulation fundamentally changes the communication scaling of federated LLM fine-tuning. The amount of information exchanged depends only on the public prompt budget and the size of the communicated behaviors, independent of model size. As a consequence, the protocol naturally accommodates heterogeneous architectures and applies directly to open-ended text generation. We present a theoretical analysis and empirical results demonstrating that this approach can match strong federated fine-tuning baselines while substantially reducing communication by orders of magnitude (e.g., analytically by a factor of 1006 for Llama3.1-405B), as well as reductions in runtime and energy consumption. These results suggest that, for generative foundation models, behavior-level consensus provides a more appropriate abstraction for federated adaptation than parameter aggregation.

Accurate long-horizon probabilistic predictions of dynamical systems are a prerequisite for model-based decision-making. Variational Bayesian last layers (VBLLs) are an attractive model class for this setting, offering tractable uncertainty at near-deterministic cost, yet they are trained exclusively with single-step likelihood objectives that may be insufficient for reliable multi-step rollouts. We revisit VBLLs through the lens of Gibbs variational inference, which decouples posterior inference from the choice of training loss. Building on this perspective, we introduce a family of multi-step training losses that progressively incorporate long-horizon structure, robustness via continuous ranked probability scores (CRPS), and training on full autoregressive rollouts. Our approach retains the simplicity of VBLLs while directly optimizing for multi-step predictive quality. Across illustrative synthetic environments, a diverse suite of chaotic dynamical systems, and real-world data, we show that multi-step, CRPS-based objectives substantially improve long-horizon accuracy and distributional fit, consistently outperforming single-step likelihood training.


Beyond Worst-Case Coreset Bounds for $k$-Clustering via Determinantal Sampling

Diptarka Chakraborty ⋅ Satyaki Mukherjee ⋅ Gaurav Vallabhdas Revankar ⋅ Hoang Son Tran

Massive datasets in modern machine learning have made data reduction a central challenge, particularly for clustering tasks where memory and computational constraints demand compact yet faithful summaries. A standard approach is to construct an $\epsilon$-$\text{\emph{coreset}}$: a small weighted subset that approximately preserves the clustering cost for every plausible choice of centers. For the $(k,z)$-$\text{\emph{clustering problem}}$, existing worst-case bounds on coreset size are essentially tight, ruling out substantially smaller coresets in general. However, such worst-case instances are often unrepresentative of real-world data. In this work, we show that significantly smaller coresets are possible under mild and natural assumptions on the underlying data distribution. We introduce a new correlated sampling framework, called $\text{\emph{determinantal sampling}}$, based on a novel application of determinantal point processes. Using this framework, we obtain an efficiently constructible $\varepsilon$-coreset for $(k,z)$-clustering in $\mathbb R^d$ whose dependence on $1/\varepsilon$ has exponent strictly smaller than $2$ when $d$ is fixed. This improves over the worst-case $\varepsilon^{-2}$ barrier under our beyond-worst-case assumptions. To the best of our knowledge, this is the first result that provably surpasses these lower bounds through beyond-worst-case assumptions. Finally, we validate our approach on synthetic and real-world benchmark datasets, where it consistently achieves smaller coresets than existing state-of-the-art methods, even without explicitly enforcing the assumptions used in the analysis.


BRACE: Bipolar Reference-Aware Calibration and Estimation for Incomplete Multimodal Learning

Ruiting Dai ⋅ Zesen Cai ⋅ XiaoYu Zhao ⋅ Yiting Huang ⋅ Lisi Mo ⋅ Ming Li ⋅ Tao He

Real-world multimodal data are often incomplete. Existing incomplete multimodal methods mainly address this problem as evidence attenuation, by reconstructing missing views, regularising shared representations, or retrieving auxiliary context.However, partial observations can still leave a different uncertainty unresolved:even when they provide a coarse task-state anchor, they may not determine whether the latent state should be revised toward a higher or lower label. We formalize this overlooked failure mode as residual-direction ambiguity and propose BRACE( Bipolar Reference-A ware Calibration and Estimation), a framework that first builds a quality-aware unified task state from available modalities, then organizes historical state-label pairs in a global memory bank, and finally infers a gated latent correction from the ordered contrast between higher-label and lower-label reference neighborhoods, together with a warmup-then-calibration training strategy and direction-consistency regularization. Across four benchmarks spanning multimodal sentiment analysis and multimodal misinformation detection under fixed and random missingness, BRACE consistently outperforms strong baselines, improving Acc-2 by up to 7.4 points over HME under fixed missingness and remaining robust when key modalities are absent; code is available at https://anonymous.4open.science/r/Brace-94BF.


Breaking Adversarial Transferability in Fine-Tuned Speech Recognition

Mojtaba Nafez ⋅ Aref Mousavi ⋅ Mohammad E Mahdavi ⋅ Mobina Poulaei ⋅ Kiarash K Feriz ⋅ Mohammad Hossein Rohban

Many organizations fine-tune publicly available pretrained Automatic Speech Recognition (ASR) models and deploy them in black-box settings, assuming limited access provides protection. We show this assumption is fragile: adversarial perturbations crafted on the public base model transfer effectively to fine-tuned target models, severely degrading performance and posing concerns for safety-critical applications. We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients. We theoretically justify all components and extensively evaluate TransferBreaker across three languages and four large ASR models, substantially reducing adversarial WER from 92.6 to 27.8.


Causal Abstractions, Categorically Unified

Markus Englberger ⋅ Devendra Singh Dhami

We introduce a categorical framework for causal abstraction that unifies and extends existing approaches. Modeling causal systems as Markov functors from free Markov categories generated by directed acyclic graphs, we define a causal abstraction as a deterministic natural transformation together with an embedding of restricted free Markov categories. This separates graphical compatibility under interventions from domain-level clustering of variables or values. Our framework provides an explicit characterization of admissible high-level graphs and recovers prior notions such as constructive -abstractions and cluster-DAG abstractions with unobserved confounders. We give concise categorical proofs that interventional distributions factorize over graphical abstractions and that do-calculus applied to the high-level graph yields valid conclusions for the low-level model.


Causal Attribution via Activation Patching

Amirmohammad Izadi ⋅ Mohammadali Banayeeanzade ⋅ Alireza Mirrokni ⋅ Hosein Hasani ⋅ Mobin Bagherian ⋅ Faridoun Mehri ⋅ Mahdieh Baghshah

Attribution methods for Vision Transformers (ViTs) aim to identify image regions that influence model predictions, but producing faithful and well-localized attributions remains challenging. Existing attribution methods face several limitations, with gradient-based, relevance-propagation, and attention-based methods relying on local approximations, while perturbation or optimization-based methods intervene on inputs, tokens, or surrogates rather than internal patch representations. The key challenge is that class-relevant evidence is formed through interactions between patch tokens across layers; methods that operate only on input changes, attention weights, or backward relevance signals may therefore provide indirect proxies for patch importance rather than directly testing the predictive effect of contextualized patch representations. We propose Causal Attribution via Activation Patching (CAAP), which estimates the contribution of individual image patches to the ViT’s prediction by directly intervening on internal activations rather than using learned masks or synthetic perturbation patterns. For each patch, CAAP inserts the corresponding source-image activations into a neutral target context over an intermediate range of layers and uses the resulting target-class score as the attribution signal. The resulting attribution map reflects the causal contribution of patch-associated internal representations on the model’s prediction. The causal intervention serves as a principled measure of patch influence by capturing semantic evidence after initial representation formation, while avoiding late-layer global mixing that can reduce spatial specificity. Across multiple ViT backbones and standard metrics, CAAP consistently outperforms existing methods in various settings and produces more faithful and localized attributions.


Causal Representation Learning for Generalisable Recommendation

Yorgos Felekis ⋅ Michael O'Riordan ⋅ Oriol Corcoll Andreu ⋅ Ciarán Gilligan-Lee

Predictive models trained on observational data often fail to generalise to the distributions they encounter when deployed, especially when the training data is a product of the system being optimised. Recommender systems are a canonical example: they are trained on interaction logs confounded by the deployed policy, past user behaviour, and platform filtering. As a result, the training distribution differs substantially from the candidate distribution scored at serving time, a gap that makes offline metrics unreliable predictors of online performance. We address the distribution shift problem with a method motivated by causal representation learning (CRL). We propose an information-theoretic disentanglement criterion and prove that its optimum depends only on the causal components of the input. We then derive a tractable variational lower bound that makes the criterion optimisable from finite observational data alone. The scope of our method is narrower than that of much of the CRL literature, in that we target better generalisation under distribution shift, not full identification of all latent causal factors. This narrower target is what makes the method practical, requiring only the existing confounded logs, applying to any standard supervised model, and adding no inference-time cost. Our headline evaluation is an A/B test with millions of users on a music streaming platform, applied to a production ranker for personalised playlist generation. A CRL variant matched in capacity to the production baseline performed on par offline but delivered substantial online gains in listener engagement. Complementary evidence on the public KuaiRand recommendation dataset and a synthetic benchmark with known causal structure shows the same pattern: offline parity with baseline, gains under distribution shift. Across all three settings, adding our causal disentanglement objective yields meaningfully better out-of-distribution generalisation.


CECAR: Cache & Expert Co-Aware Routing Accelerates On-Device Inference of MoE LLMs

Byeongju Kim ⋅ Hoonki Lee ⋅ Byungjun Kim ⋅ Inyul Ra ⋅ Yongjin Kim ⋅ Yongseong Lee ⋅ Sangyeob Kim

Despite growing interest in on-device LLM deployment, large-scale Mixture-of-Experts (MoE) models remain impractical on resource-constrained devices, as sparse expert activation still requires all expert weights to be memory-resident. To address this, we identify three key observations for on-device MoE inference: (a) cache hit rates are fundamentally bounded even under Belady’s optimal policy; (b) these bounds are insufficient for low-latency on-device inference; and (c) prior cache-aware routing based on binary cache presence can induce routing instability. Based on these insights, we propose CECAR (Cache & Expert Co-Aware Routing), a MoE inference system that jointly considers expert routing and cache management. Unlike prior cache-aware routing methods, CECAR uses an ML-based cache policy to predict the expected reuse distance of each expert and prioritizes routing and execution toward experts with smaller predicted reuse distances. On the Qwen3-30B-A3B model, CECAR achieves 12.5 tokens/s on a consumer-grade GPU with only 16 GB of VRAM, where non-resident experts are fetched from SSD, delivering a 4.5× speedup over conventional MoE inference and exceeding the decode speed of baselines that keep the entire model fully resident in DRAM.


Cheap Per-Component Testing for PLS, Stable Under Rotation

Paweł Lenartowicz ⋅ Hubert Plisiecki

Partial Least Squares (PLS) regression extracts a few outcome-aligned directions in a high-dimensional $\mathbf{X}$ and is widely used across applied science, but inference on the resulting fit is either expensive (CV-permutation-$Q^{2}$), biased and discouraged (jackknife $t$), or absent sklearn ships no test; the corresponding statsmodels feature request has been open since 2019). We close this gap by reducing inference to held-out OLS refits of the supervised subspace, a primitive shared by PLS, supervised PCA, linear probes, and sparse-autoencoder concept directions. We supply two tests on it: a Nadeau-Bengio corrected asymptotic $t$-test (NB-asymptotic) as the default, and a permutation-referenced variant (NB-permutation) for near-singular designs where the Fisher-$z$ asymptotic drifts. A rotation-invariance result shows the held-out predictions, and hence the test statistic, are unchanged under any orthogonal rebasing of the supervised span, so the same test applies to varimax-rotated word-readable axes. We validate on PLS (synthetic geometries, Tecator NIR chemometrics, cross-lingual GloVe valence prediction in EN/PL/ES at $K=3$, $r^{2}\approx0.63/0.69/0.55$, all $p<.001$); transfer to the other OLS-refit pipelines is argued but not tested here. NB needs $\sim20\%$ fewer samples than CV-permutation-$Q^{2}$ for matched power and is $50-100 \times$ cheaper at common, already compromised, defaults. We release a Rust library with Python, R, and Julia bindings, plus a Python text pipeline.


Cluster with Auctions for Vector Search

Swann BESSA ⋅ Pierre Fernandez ⋅ Gergely Szilvasy ⋅ Matthijs Douze ⋅ Herve Jegou

Large-scale approximate nearest neighbor search commonly relies on partitions for indexing: database vectors are partitioned into clusters, and for each query a probing function selects the clusters to be scanned. The query probing function and the database partition are rarely treated as separate entities: most techniques assign queries with the same assignment function as the database vectors, which is suboptimal especially when database and query distributions differ. This paper introduces CwA (Cluster with Auctions), which addresses this limitation by jointly learning a balanced database partition and a neural probing function. CwA optimizes search performance directly for the query distribution. It minimizes its objective by alternating two steps: (i) gradient descent on the neural network of the probing function, and (ii) a large-scale combinatorial optimization of the cluster assignment for the database vectors. We solve the latter with a parallelizable auction algorithm that balances the partition by design. To further scale CwA, we extend the method to a Cartesian product of clusters that increases the partition’s granularity. When database and query distributions differ, CwA achieves up to 4.7$\times$ throughput over the state of the art at equal recall. In the in-distribution (ID) setting, even a simple linear probing function trained with CwA outperforms competing deep neural methods.


Computing Thiele Rules on Interval Elections and their Generalizations

Dimitris Avramidis ⋅ Alexandra Anna Lassota ⋅ Ulrike Schmidt-Kraepelin ⋅ Adrian Vetta

Approval-based committee voting has received significant attention in the social choice community. Among the studied rules, Thiele rules, and especially Proportional Approval Voting (PAV), stand out for desirable properties such as proportional representation, Pareto optimality, and support monotonicity. Their main drawback is that computing a Thiele outcome is NP-hard in general. A glimpse of hope comes from the fact that Thiele rules are better behaved under structured preferences. On the candidate interval (CI) domain, they are computable in polynomial time via a linear program (LP) that has a totally unimodular constraint matrix. Surprisingly, this approach fails for the related voter interval (VI) domain, and the complexity of the problem has repeatedly been posed as an open question. Our main result resolves this question: although the relevant matrix is not totally unimodular, the ``standard'' LP still admits at least one optimal integral solution, and we provide a fast algorithm for finding it. Our technique naturally extends to the voter-candidate interval (VCI) domain, also known as the 1-dimensional voter-candidate range (1D-VCR) domain, and to the linearly consistent (LC) domain, both of which generalize the candidate and voter interval domains. Although both the VCI and LC domains have been studied in social choice, their relationship was unknown. We show, through connections to graph theory, that LC strictly contains VCI. We also provide an alternative definition of LC that is closer in spirit to VCI and has a natural interpretation in approval elections; this equivalence may be of independent interest. Finally, we study an alternative tree-based generalization of VCI and show that Thiele rules become NP-hard to compute on this domain.


Conservative Continuous-Time Treatment Optimization

Nora Schneider ⋅ Georg Manten ⋅ Niki Kilbertus

We propose a conservative continuous-time stochastic control framework for treatment optimization from irregularly sampled patient trajectories. We model the unknown patient dynamics as a controlled stochastic differential equation where treatment is a continuous-time control input. Naive model-based optimization can exploit model errors and propose out-of-distribution controls that appear optimal under the learned model but perform poorly under the true dynamics. To mitigate such extrapolation, we introduce a consistent, signature kernel MMD regularizer on path space, which penalizes controls whose predicted trajectories deviate from observed ones. The resulting conservative objective minimizes a tractable upper bound on the true cost. Experiments on pharmacometric simulators show improved robustness and performance compared to non-conservative baselines.


Consistent Geometric Deep Learning via Hilbert Bundles and Cellular Sheaves

Kartik Tandon ⋅ Julian J Gould ⋅ Tanishq Bhatia ⋅ Francesca Dominici ⋅ Alejandro Ribeiro ⋅ Claudio Battiloro

Modern deep learning architectures increasingly contend with sophisticated signals that are natively infinite-dimensional, such as time series, probability distributions, or operators, and are defined over irregular domains. Yet, a unified learning theory for these settings has been lacking. To start addressing this gap, we introduce a novel convolutional learning framework for possibly infinite-dimensional signals supported on a manifold. Namely, we use the connection Laplacian associated with a Hilbert bundle as a convolutional operator, and we derive filters and neural networks, dubbed as \textit{HilbNets}. We make HilbNets and, more generally, the convolution operation, implementable via a two-stage sampling procedure. First, we show that sampling the manifold induces a Hilbert Cellular Sheaf, a generalized graph structure with Hilbert feature spaces and edge-wise coupling rules, and we prove that its sheaf Laplacian converges in probability to the underlying connection Laplacian as the sampling density increases. Notably, this result is a generalization to the infinite-dimensional bundle setting of the Belkin &amp; Niyogi \cite{BELKIN20081289} convergence result for the graph Laplacian to the manifold Laplacian, a theoretical cornerstone of geometric learning methods. Second, we discretize the signals and prove that the discretized (implementable) HilbNets converge to the underlying continuous architectures and are transferable across different samplings of the same bundle, providing consistency for learning. Finally, we validate our framework on synthetic and real-world tasks. Overall, our results broaden the scope of geometric learning as a whole by lifting classical Laplacian-based frameworks to settings where the signal at each point lives in its own Hilbert space.


Constrained MDPs with Trajectory Constraints

Martino Bernasconi ⋅ Matteo Castiglioni ⋅ Alberto Marchesi

We introduce a framework that models constrained MDPs subject to trajectory-wise safety constraints. These constraints require that expected future costs remain below zero conditioned to any possible trajectory---not merely in expectation from the initial state. This captures stronger and more realistic safety guarantees, which are crucial in high-stakes applications. We characterize the structure of optimal policies, and show that Markovian policies are in general suboptimal. Moreover, we show that computing approximately-safe optimal policies is NP-hard when the number of constraints can be arbitrarily large. Despite all these challenges, we provide an exact algorithmic characterization of optimal history-dependent (non-Markovian) policies, by exploiting a suitable recursive formulation of the sets of reachable reward-cost values. Then, we employ it to design a polynomial-time approximation algorithm that computes approximately-safe optimal policies by dynamic programming, when assuming a constant number of constraints.


Continuous-Time Distribution Matching for Few-Step Diffusion Distillation

Tao Liu ⋅ Hao Yan ⋅ Mengting Chen ⋅ Taihang Hu ⋅ Zhengrong Yue ⋅ Zihao Pan ⋅ Jinsong Lan ⋅ Xiaoyong Zhu ⋅ Ming-Ming Cheng ⋅ Bo Zheng ⋅ Yaxing Wang

Step distillation has become a leading technique for accelerating diffusion models, among which Distribution Matching Distillation (DMD) and Consistency Distillation are two representative paradigms. While consistency methods enforce self-consistency along the full PF-ODE trajectory to steer it toward the clean data manifold, vanilla DMD relies on sparse supervision at a few predefined discrete timesteps. This restricted discrete-time formulation and mode-seeking nature of the reverse KL divergence tends to exhibit visual artifacts and over-smoothed outputs, often necessitating complex auxiliary modules---such as GANs or reward models---to restore visual fidelity. In this work, we introduce Continuous-Time Distribution Matching (CDM), migrating the DMD framework from discrete anchoring to continuous optimization for the first time. CDM achieves this through two continuous-time designs. First, we replace the fixed discrete schedule with a dynamic continuous schedule of random length, so that distribution matching is enforced at arbitrary points along sampling trajectories rather than only at a few fixed anchors. Second, we propose a continuous-time alignment objective that performs active off-trajectory matching on latents extrapolated via the student's velocity field, improving generalization and preserving fine visual details. Extensive experiments on different architectures, including SD3-Medium and Longcat-Image, demonstrate that CDM provides highly competitive visual fidelity for few-step image generation without relying on complex auxiliary objectives.

Sparse autoencoders (SAEs) decompose language model activations into interpretable features, primitives for mechanistic circuit analysis. Yet existing methods identify task-specific circuits via attribution or apply fixed-direction steering, but learn no task-reward-optimized policy for per-token feature amplification. We formulate token-level feature intervention as a POMDP over the LLM's layer-causal computation. Control Reinforcement Learning (CRL) trains a policy in this POMDP that selects an SAE feature to amplify at each generation step, yielding per-token intervention traces that expose which features drive task behavior. Adaptive Feature Masking encourages diverse feature discovery while preserving single-feature attribution. The framework provides branch point tracking (tokens where feature choice changes outcomes), critic trajectory analysis (separating policy from value estimation errors), and layer-wise comparison along the residual stream hierarchy. Cross-task transfer is asymmetric (MMLU features harm GSM8K but not vice versa; HarmBench helps XSTest), consistent with partial task computation overlap rather than dataset-specific shortcuts: task-conditional attribution does not require transfer invariance. On Gemma-2 2B and LLaMA-3.1 8B across MMLU, BBQ, GSM8K, HarmBench, and XSTest, CRL provides per-token intervention traces and serves as an SAE quality diagnostic alongside task-accuracy gains. Learned feature steering thus serves as a layer-specific diagnostic complementing activation-based SAE analysis.


CoPE-VideoLM: Leveraging Codec Primitives For Efficient Video Language Modelling

Sayan Deb Sarkar ⋅ Rémi Pautrat ⋅ Ondrej Miksik ⋅ Marc Pollefeys ⋅ Iro Armeni ⋅ Mahdi Rad ⋅ Mihai Dusmanu

Video Language Models enable AI systems to understand temporal dynamics in videos. To fit within the maximum context window constraint, current methods use keyframe sampling which often misses both macro-level events and micro-level details due to the sparse temporal coverage. Furthermore, processing full images and their tokens for each frame incurs substantial computational overhead. We address these limitations by leveraging video codec primitives (specifically motion vectors and residuals) which natively encode video redundancy and sparsity without requiring expensive full-image encoding for most frames. To this end, we introduce lightweight transformer-based encoders that aggregate codec primitives and align their representations with image encoder embeddings through a pre-training strategy that accelerates convergence during end-to-end fine-tuning. Our approach, CoPE-VideoLM, reduces the time-to-first-token by up to 86% and token usage by up to 93% compared to standard VideoLMs. Moreover, by varying the keyframe and codec primitive densities, we maintain or exceed performance compared on 14 diverse video understanding benchmarks, spanning general question answering, temporal and motion reasoning, long-form understanding, and spatial scene understanding.

NIGHTS is a recent benchmark for human perceptual similarity over objects, scenes, and composition. DreamSim, the supervised metric introduced with NIGHTS, and later supervised metrics trained on the same data report large gains over self-supervised baselines. The standard NIGHTS evaluation uses the unanimous-agreement subset of a released 100,000-triplet corpus, leaving weak-majority and split-vote triplets outside the headline benchmark. We evaluate eleven perceptual similarity methods on the full released corpus, analyzing unanimous, weak-majority, and split-vote triplets separately. On the 64,420 non-tie triplets, cosine distance on a frozen self-supervised backbone reaches 74.1% accuracy without training, LoRA fine-tuning reaches 77–79%, and the DreamSim ensemble reaches 80.7%. A split-half reproducibility estimate from the triplet vote counts shows that majority labels reproduce only 70.1% of the time overall, and only 33–50% on weak-majority triplets. The extra accuracy comes from matching the collected votes, not from showing that the metric would better match a new set of human judgments. The evaluation target is also unstable across task formulations: NIGHTS’ 2AFC similarity labels and JND same/different labels disagree on 51–64% of shared stimuli. Fine-tuning changes the feature basis as well, increasing foreground-only fragility and making DreamSim the most fragile method under rotation, color, and blur. NIGHTS supports a narrower conclusion: cosine distance on self-supervised features captures the stable signal currently measurable in these labels. Additional claims of human alignment require a redesigned benchmark with more votes per triplet, axis-specific similarity questions, per-subject statistics, and ablation tests.


COVD: Continual Open-Vocabulary Object Detection with Novel Concept Injection

Yupeng Zhang ⋅ Ruize Han ⋅ Yuzhong Feng ⋅ Zixin Ren ⋅ Yuntong Tian ⋅ Liang Wan

Open-vocabulary object detection (OVD) has made significant progress, enabling detectors to generalize from seen to unseen categories. However, real-world category spaces continually evolve, and existing OVD models still struggle with newly emerging concepts, while repeated full retraining is prohibitively expensive. To this end, we introduce a new task setting, termed Continual OVD with Novel Concept Injection (COVD), where models sequentially learn incoming novel concept groups while preserving prior concepts and original open-vocabulary knowledge, along with a new benchmark, Novel-114. Our key observation is that pretrained visual encoders often already perceive and represent many novel concepts, and the main bottleneck lies in the lack of stable semantic alignment between visual representations and textual concepts. Based on this, we propose NoIn-Det, an efficient continual injection framework without additional parameters. NoIn-Det freezes the visual encoder, preserves the text representation space using only texts of common concepts and previously injected concepts, and injects novel concepts by updating only a small subset of text-branch parameters beneficial to novel concept learning. Extensive experiments show that NoIn-Det effectively learns novel concepts, preserves old knowledge, and consistently outperforms existing continual learning methods for VLMs without introducing additional parameters.


Curvature-Dependent Lower Bounds for Frank-Wolfe

Jannis Halbey ⋅ Christophe Roux ⋅ Sebastian Pokutta

The Frank-Wolfe (FW) algorithm achieves a convergence rate of $\mathcal{O}(1/T)$ for smooth convex optimization over compact convex domains, accelerating to $\mathcal{O}(1/T^2)$ when both the objective and the feasible set are strongly convex. This acceleration extends beyond strong convexity: Kerdreux et al. (2021a} proved rates of $\mathcal{O}(T^{-p/(p-1)})$ over $p$-uniformly convex feasible sets, a class that interpolates between strongly convex sets and more general curved domains such as $\ell_p$ balls. In this work, we establish a matching $\Omega (T^{-p/(p-1)})$ lower bound for every $p\ge 3$ under exact line search or short steps, and extend the lower bound to objectives satisfying a Hölderian error bound. The proofs analyze the dynamics of FW iterates on simple instances and hence are not limited to the high-dimensional setting, unlike information-theoretic lower bounds.

Text-to-image (T2I) diffusion models have achieved strong performance in semantic alignment, yet they still struggle with generating the correct number of objects specified in prompts. Existing approaches typically incorporate auxiliary counting networks as external critics to enhance numeracy. However, since these critics must provide gradient guidance during generation, they are restricted to regression-based models that are inherently \emph{differentiable}, thus excluding detector-based models with superior counting ability, whose count-via-enumeration nature is \emph{non-differentiable}. To overcome this, we propose \textbf{Detector-to-Differentiable} (\emph{D2D}), a novel framework that transforms non-differentiable detection models into differentiable critics, leveraging their superior counting ability to guide numeracy generation. Specifically, we design custom activation functions to convert detector logits into soft binary indicators, which are then used to optimize the noise prior at inference time with pre-trained T2I models. Our extensive experiments on SDXL-Turbo, SD-Turbo, and Pixart-DMD across four benchmarks of varying complexity (low-density, high-density, and multi-object scenarios) demonstrate consistent and substantial improvements in object counting accuracy (e.g., boosting up to 13.7\% on D2D-Small, a 400-prompt, low-density benchmark), with minimal degradation in overall image quality and computational overhead.

Existing deep learning architectures for survival prediction from electronic health records (EHRs) largely assume unimodality, failing to capture the complex relational structure across diverse clinical modalities. Graph learning can allow us to robustly adapt to the underlying clinical structure of the data to have more predictive and interpretable representations for time-to-event analysis. However, unconstrained graph learning produces dense, bidirectional adjacencies that overfit. We introduce M-CADENCE, a multimodal survival model that integrates structured acyclic priors with attention-based propagation. Each modality has its own learnable adjacency, regularised toward acyclicity through a log-determinant characterisation and toward sparsity through soft thresholding. A cross-modal meta-adjacency keeps the cross-modal parameter count quadratic in the number of modalities rather than in their feature dimensions. Propagation is then performed by multi-head attention biased by the learnt structure. The propagation step degrades gracefully under sparse or misspecified structure, addressing a known weakness of hard-routing graph variants. We evaluate M-CADENCE on five EHR benchmarks covering intensive-care mortality, circulatory failure, and emergency-department competing risks, against multiple survival, graph, and structure learning baselines. M-CADENCE achieves the best concordance on all datasets, the best early-warning AUC on four of five, and competitive precision-recall against the strongest deep survival baselines. Removing the DAG bias lowers concordance, as does substitution with a density-matched random adjacency, which confirms that the specific learnt structure is required. Discrimination is stable across four orders of magnitude in regularisation strength, and evaluation by a panel of ten language models confirms M-CADENCE learns clinical edges that are more physiologically plausible than baseline structure-learning methods.


Dancing in Fetters: Pareto-Optimal On-Device LLMs under Hardware Constraints

Luoyang Sun ⋅ Jiwen Jiang ⋅ Yifeng Ding ⋅ Fengfa Li ⋅ Yan Song ⋅ Haifeng Zhang ⋅ Lei Ren ⋅ Kun Zhan ⋅ Chen Wei ⋅ Xie Yan ⋅ Jun Wang ⋅ Cheng Deng

Recent embodied AI systems increasingly rely on large language models (LLMs) for high-level planning, semantic reasoning, and long-horizon decision making. However, deploying such capabilities on edge platforms such as autonomous vehicles and mobile robots requires balancing model quality against strict latency and hardware constraints. Existing LLM architectures are primarily designed for cloud-scale accelerators and often become inefficient in on-device settings. We present $\textbf{PLAS}$ ($\textbf{P}$areto-optimal $\textbf{L}$LM $\textbf{A}$rchitecture $\textbf{S}$earch), a hardware-aware framework for identifying Pareto-optimal on-device LLM architectures under deployment constraints. Our approach jointly models training loss and inference latency as functions of architectural design choices. We fit an empirical architecture-to-loss model from 170 LLMs trained on 10B tokens each, and estimate latency using roofline-based hardware modeling, enabling efficient exploration of the accuracy-latency Pareto frontier without exhaustively training or profiling every candidate architecture. Using PLAS, we evaluate 1942 candidate architectures on NVIDIA Jetson Orin. At the same measured latency as Qwen2.5-0.5B, our co-designed architecture achieves 19.42\% lower WikiText-2 perplexity. Our results suggest that effective edge deployment requires explicit hardware--model co-design, and that Pareto-based architecture modeling provides a practical alternative to exhaustive neural architecture search for on-device LLMs.


DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking

Akash Haridas ⋅ Utkarsh Saxena ⋅ Parsa Ashrafi Fashi ⋅ Mehdi Rezagholizadeh ⋅ Vikram Appia ⋅ Emad Barsoum

Diffusion Transformers rely on static patchify tokenization, assigning the same token budget to smooth backgrounds, detailed object regions, noisy early timesteps, and late-stage refinements. We introduce the Dynamic Chunking Diffusion Transformer (DC-DiT), which replaces fixed patchification with a learned encoder-router-decoder scaffold that adaptively compresses the 2D input into a shorter token sequence through a chunking mechanism learned end-to-end with diffusion training. DC-DiT allocates fewer tokens to predictable regions and noisy timesteps, and more tokens to detailed regions and later refinement stages, yielding meaningful spatial segmentations and timestep-adaptive compression schedules without supervision. Furthermore, the router provides an importance ordering over retained tokens, enabling elastic inference: a single checkpoint can be evaluated at flexible compute budgets with a smooth quality-compute tradeoff. Additionally, DC-DiT can be upcycled from pretrained DiT checkpoints and is also compatible with orthogonal dynamic computation approaches. On class-conditional ImageNet generation, DC-DiT reduces inference FLOPs by up to 36.8% and improves FID by up to 37.8% over DiT baselines, yielding a stronger quality--compute Pareto frontier across model scales, resolutions, and guidance settings. More broadly, these results suggest that adaptive tokenization is a general mechanism for making visual generation both more efficient and more flexible at inference time.


Decentralized $\mu^2$-SGD: Narrowing the Parallelism Gap to Centralized Learning

Ofri Eisen ⋅ Sharon Goldstein ⋅ Ran Elbaz ⋅ Kfir Y. Levy

While decentralized learning offers a communication-efficient alternative to centralized distributed training, its scalability is often limited by the number of workers that can be used without degrading statistical efficiency. This limitation is especially pronounced over sparse communication networks, where increasing parallelism can lead to a sharp loss in performance. To address this bottleneck, we introduce *Decentralized $\mu^2$-SGD (DMS)*, a novel decentralized optimization method that significantly extends the parallelism limits of decentralized learning. From a theoretical perspective, within the stochastic convex optimization (SCO) framework, we establish improved bounds on the maximal allowable parallelism, surpassing existing decentralized algorithms. Notably, for a broad class of network topologies, our method matches the parallelism scaling of centralized learning, thereby effectively eliminating the gap between decentralized and centralized optimization. Empirically, we validate our theoretical findings through comprehensive experiments, demonstrating the benefits of our approach.


Decentralized Ranking Aggregation via Gossip: Convergence and Robustness

Kerrian Le Caillec ⋅ Anna van Elst ⋅ Igor Colin ⋅ Stephan Clémençon

The concept of ranking aggregation plays a central role in preference analysis, and numerous algorithms for calculating median rankings, often originating in social choice theory, have been documented in the literature, offering theoretical guarantees in a centralized setting, \textit{i.e.}, when all the ranking data to be aggregated can be brought together in a single computing unit. For many technologies (\textit{e.g.} peer-to-peer networks, IoT, multi-agent systems), extending the ability to calculate consensus rankings with guarantees of convergence and resilience to potential contamination in a decentralized setting, when preference data is initially distributed across a communicating network, remains a major methodological challenge. Indeed, in recent years, the literature on decentralized computation has mainly focused on computing or optimizing statistics such as arithmetic means using gossip algorithms. The purpose of this article is precisely to study how to achieve reliable and resilient consensus on collective rankings in a decentralized setting, thereby raising new questions, robustness to corrupted nodes, and scalability through reduced communication costs in particular. The approach proposed and analyzed here relies on robustness guarantees derived from random gossip communication, allowing autonomous agents to compute global ranking consensus using only local interactions, without coordination or central authority.


Decoupled Mode Connectivity for Base-to-Novel Generalization in Vision-Language Models

Imad-Eddine Marouf ⋅ khalid OUBLAL ⋅ Enzo Tartaglione ⋅ Stéphane LATHUILIÈRE

Prompt learning adapts vision-language models such as CLIP by optimizing a small set of continuous context vectors. The cross-entropy objective drives the learned prompt toward base-class specialization, while generalization to unseen classes benefits from staying close to the zero-shot feature space. Because both objectives act on the same parameters, their gradients oppose each other at convergence, and single-prompt losses are confined to a fixed empirical base-novel trade-off curve across a wide range of loss designs. We propose Decoupled Mode Connectivity (DMC), which resolves this conflict by assigning each objective to a dedicated prompt. A linear mode connectivity (LMC) corridor in text-feature space enforces low classification loss across interpolated classifiers between the two endpoints. The Visual Anchor regularizer preserves CLIP's pretrained class-similarity structure during specialization; equivalently, it minimizes the KL divergence between the learned and pretrained class-similarity distributions. We introduce class-permutation invariance (CPI) as a necessary condition for regularizer transfer across the base-novel boundary, and prove via Fano's inequality that any CPI violation lower-bounds the drop in novel accuracy by a term proportional to the mutual information between the prompt and base-class labels. DMC improves base and novel accuracy on the majority of dataset-baseline combinations across 11 datasets and three prompt-learning baselines~(CoOp, KgCoOp, MMA), shifting the Pareto frontier rather than trading along it.


Decoupled Prototype Contrastive Alignment Hashing for Cross-Modal Retrieval

Yunfei Chen ⋅ Renwei Xia ⋅ Shangchong Gao ⋅ Zhan Yang

Cross-modal hashing is essential for efffcient large-scale re-trieval in multimedia analysis. However, existing unsupervised methods often struggle to fully capture interactions between modalities, enforce multi-level alignment, and maintain stable optimization. To overcome these challenges, we propose Decoupled Prototype Contrastive Align-ment Hashing (DPCAH), which employs a cross-modal semantic con-trastive learning module that concatenates image and text features and encodes their interactions via a transformer to produce uniffed represen-tations. feature-level and hash-level contrastive objectives jointly align semantic information and guide discriminative hash code learning across modalities. A decoupled prototype consistency module further models cross-modal correlations independently, enhancing semantic alignment while ensuring stable and robust optimization. Experiments on three benchmark datasets demonstrate that DPCAH outperforms state-of-the-art methods in cross-modal retrieval.


Depth2Pose: A Pose-Based Benchmark for Monocular Depth Estimation without Ground-Truth Depth

Viktor Kocur ⋅ Sithu Aung ⋅ Gabrielle Flood ⋅ Yaqing Ding ⋅ Lukas Bujnak ⋅ Torsten Sattler ⋅ Zuzana Kukelova

Monocular depth estimation has improved significantly in recent years, driven by increasingly powerful models and large-scale training data. Predicted depth is increasingly used as an input signal for downstream tasks such as Structure-from-Motion (SfM), visual localization, and SLAM. However, monocular depth estimators (MDEs) are still primarily evaluated in terms of depth accuracy. Standard metrics aggregate errors globally and may not reflect the usefulness of depth for downstream geometric tasks. We therefore propose Depth2Pose, a framework for evaluating MDEs in the context of downstream tasks. By combining depth predictions with feature correspondences in depth-aware geometric solvers, we use relative camera pose estimation accuracy as a task-driven proxy for depth quality. Traditional benchmarks require dense ground truth in the form of per-pixel depth, which is expensive to obtain. In contrast, our formulation requires only camera poses, which can be estimated efficiently, e.g., using Structure-from-Motion pipelines. As a result, our framework can be applied to scenes where ground-truth depth is difficult to obtain, for example due to large scene scale or heavy occlusions (e.g., vegetated environments). Leveraging this, we introduce the D2P dataset, which contains challenging scenes outside the distribution of commonly used training data. We show that methods performing well under standard depth error metrics on existing benchmarks also perform well under our pose-based metric when evaluated on the same datasets, but do not necessarily generalize to our more challenging dataset. Finally, we provide a simple and extensible evaluation framework. The dataset and code are available at kocurvik.github.io/depth2pose.


Difference of Convex Programming in the Wasserstein Space with Applications to MMD Optimization

Clément Bonet ⋅ Pierre-Cyril Aubin-Frankowski ⋅ Youssef Mroueh

Optimizing functionals over the space of probability measures is now ubiquitous in machine learning. A widely used approach is to perform the optimization directly over the Wasserstein space, but many objective functionals of practical interest are non-convex along Wasserstein geodesics, making the analysis of standard first-order methods challenging. In this work, we study a class of objectives over the Wasserstein space that admit a difference-of-convex (DC) decomposition and we lift the classical convex-concave procedure (CCCP) to this setting. Under smoothness and strong convexity assumptions on the convex components of the decomposition, we prove almost stationarity along the iterates of the resulting algorithm. Our main focus is on the Maximum Mean Discrepancy (MMD) and the Energy Distance (ED) functionals, for which we develop explicit Wasserstein DC decompositions, and establish local convergence of the scheme under mild assumptions. Empirically, we show that well-chosen DC decompositions yield faster and more stable convergence than Wasserstein gradient descent on these MMD objectives.


Dimension-Uniform Discretization Analysis of Preconditioned Annealed Langevin Dynamics for Multimodal Gaussian Mixtures

Lorenzo Baldassari ⋅ Josselin Garnier ⋅ Knut Solna ⋅ Maarten V. de Hoop

Obtaining stable diffusion-based samplers in high- and infinite-dimensional settings is challenging because errors can accumulate across high-frequency coordinates and make the dynamics unstable under refinement of the finite-dimensional approximation of the underlying function-space problem. Discretization is a typical source of such errors, and preconditioning with a suitable spectral decay is one way to control their accumulation. In this paper, we study this problem for preconditioned annealed Langevin dynamics (ALD) applied to Gaussian mixtures. We first show that Euler-Maruyama (EM) discretization, by treating the stiff linear part of the annealed score with a forward Euler step, imposes a stability constraint coupling the preconditioner with the annealed covariance scale. Together with the conditions ensuring dimension-uniform control of the annealed dynamics, this constraint forces the initial smoothed law to remain uniformly close to the target across dimensions. We then consider an exponential-integrator scheme that integrates the stiff linear part of the annealed score exactly. Under explicit spectral summability conditions coupling the smoothing covariance, the component covariance spectra, and the preconditioner, we prove a dimension-uniform Kullback-Leibler (KL) bound for this scheme. This bound can be made arbitrarily small, uniformly in dimension, by allowing enough time for annealing and then refining the time mesh accordingly. Importantly, these conditions allow regimes in which the KL divergence between the target and the initial smoothed law diverges with dimension, showing that the restrictions imposed by EM are scheme-dependent rather than intrinsic to ALD.


Diversity Curves for Graph Representation Learning

Katharina Limbeck ⋅ Nadja Häusermann ⋅ Martin Carrasco ⋅ Guy Wolf ⋅ Bastian Rieck

Graph-level representations are crucial tools for characterising structural differences between graphs. However, comparing graphs with different cardinalities, even when sampled from the same underlying distribution, remains challenging. Unsupervised tasks in particular require interpretable, scalable, and reliable size-aware graph representations. Our work addresses these issues by tracking the structural diversity of a graph across coarsening levels. The resulting graph embeddings, which we denote diversity curves, are interpretable by construction, efficient, and directly comparable across coarsening hierarchies. Specifically, we track the spread of graphs, a novel isometry invariant that is inherently well-suited for encoding the metric diversity and geometry of graphs. We utilise edge contraction coarsening and prove that this improves expressivity, thus leading to more powerful graph-level representations than structural descriptors alone. Demonstrating their utility over a range of baseline methods in practice, we use diversity curves to (i) cluster and visualise simulated graphs across varying sizes, (ii) distinguish the geometry of single-cell graphs, (iii) compare the structure of molecular graph datasets, and (iv) characterise geometric shapes.


Do Heavy Tails Help Diffusion? On the Subtle Trade-off Between Initialization and Training

Hamza Cherkaoui ⋅ Hélène Halconruy ⋅ Antonio Ocello

Recent works have proposed incorporating heavy-tailed (HT) noise into diffusion- and flow-based generative models, with the goals of better recovering the tails of target distributions and improving generative diversity. This motivation is intuitive: if the data are heavy-tailed, HT noise may appear better matched than light-tailed (LT) Gaussian noise. However, replacing Gaussian noise by HT noise also changes the underlying estimation problem. In this paper, we revisit this paradigm through a combined theoretical and empirical study, establishing sampling-error bounds for two representative diffusion models driven by HT and LT noise. We show that HT noise makes the statistical estimation problem harder, leading to less favorable sampling-error bounds. We support these findings with experiments on synthetic and real-world datasets, empirically recovering the predicted error trade-off. Our results call into question a growing design trend in generative modeling and challenge the use of HT noise to improve rare-region exploration.

Evolution strategies optimize through fitness evaluations, but standard ES exposes data-dependent fitness values. We introduce DP-EGGROLL, a central-DP mechanism for EGGROLL-style population optimization with decomposable supervised losses. Each private step forms per-example candidate-loss vectors, mean-centers each row across candidates, clips the centered row, averages with a fixed nominal denominator, and adds calibrated Gaussian noise; reward shaping and the ES update are post-processing. Clipping gives the $C/m$ add/remove sensitivity bound, while centering removes common-mode loss before clipping. We evaluate 10 public tabular benchmarks across five privacy budgets, using 10 final seeds and equal 32-configuration HPO budgets for fast DP-AdamW, scalar DP-ZO, centered DP-EGGROLL, and independently tuned uncentered DP-EGGROLL. Each selected final run is accounted as an individual $(\varepsilon,10^{-5})$-DP run; multi-configuration HPO is a benchmark protocol, not a single-run deployment guarantee. On tabular classification, centered DP-EGGROLL is practically non-inferior to fast DP-AdamW on 22/25 AUROC endpoints and is faster per step in all 25 classification settings. Centering improves over tuned uncentered DP-EGGROLL on all 25 classification AUROC endpoints. Neural experiments include end-to-end small MLPs and a low-rank adapter stress test; both reinforce that centered population-vector privatization carries more signal than scalar DP-ZO. Regression is more heterogeneous, with centered DP-EGGROLL practically non-inferior on 14/25 RMSE endpoints.


Dynamics of Collective Diversity in Human–AI Co-Creation for Creative Tasks

Manh Hung Nguyen ⋅ Chao Wen ⋅ Sebastian Tschiatschek ⋅ Adish Singla

Generative AI systems are increasingly deployed at scale to support human users in creative tasks across a variety of domains. However, as more users adopt the same AI agent (e.g., ChatGPT), the resulting artifacts may become increasingly homogeneous, leading to a collapse in collective diversity. Prior work has documented this risk, but lacks a systematic framework for understanding how collective diversity changes as AI adoption scales across co-creation settings. In this work, we conduct a large-scale crowdsourcing study to understand and model the dynamics of collective diversity, quality, and effort in human–AI co-creation for three creative tasks. We find that current AI agents used in co-creation settings can improve artifact quality and reduce effort, but that these gains come at the cost of declining collective diversity as adoption increases. We further show that more structured and fine-grained human-AI interaction, along with more diverse agent designs, can mitigate this collapse. Finally, we derive predictive models that capture how collective diversity scales across co-creation settings, providing valuable insights for the design of AI agents that aim to preserve collective diversity.

Neural network training is typically monitored through loss and accuracy curves, but these quantities do not directly reveal when the backbone begins to encode training and test data differently. A train-test negative log-likelihood (NLL) gap shows probabilistic separation between training and test examples, but it does not reveal whether this separation is present in the backbone geometry. We introduce the Volume Ratio Test (VRT), a permutation-based two-sample test for l2-normalized backbone embeddings that compares local angular neighborhood structure on the unit hypersphere. VRT converts k-nearest-neighbor angles into spherical-cap log-volume spacings and requires no kernel bandwidth selection or learned discriminator. Across 14 model configurations spanning embedding dimension D = 32 to D = 2,048, VRT matches or precedes the earliest main baseline on sustained backbone divergence. In several CIFAR-10 ResNet settings, VRT is the only method among MMD, MMDAgg, Energy Distance, and C2ST to detect sustained backbone divergence. When these baselines lag, VRT leads them by 25-60 epochs. Additional nearest-neighbor methods show that Schilling and Henze either reject 16-60 epochs later or remain null where VRT rejects. On ResNet-50/CIFAR-100, VRT first rejects at epoch 61, whereas MMD, MMDAgg, and C2ST first reject at epoch 121. On this model, the train-test NLL gap is already 0.563 at epoch 62 and never falls below this value. More generally, sustained VRT rejection is accompanied by persistent NLL separation, but large NLL gaps can occur without VRT rejection. VRT therefore distinguishes train--test separation in NLL from backbone-geometric divergence.


Effective Biological Representation Learning by Masking Gene Expression

Kian Kenyon-Dean ⋅ Alina Selega ⋅ Ihab Bendidi ⋅ Jordan M Sorokin ⋅ Luca Bertinetto ⋅ David Errington ⋅ Hayley Donnella ⋅ Oren Kraus

RNA sequencing produces rich and diverse datasets of gene expression, offering compelling insights into cellular state and function that have many applications in drug discovery. Modeling such data is challenging due to inherent technical noise and experimental batch effects, as evidenced by many existing transcriptomic foundation models (FMs) underperforming relative to linear baselines. Such results raise the question of whether deep representation learning provides a distinct advantage over the direct use of raw transcript counts. Our work explores this by developing a new self-supervised model, TxFM, with a focus on inductive representation learning evaluations. TxFM employs a masked autoencoding approach tailored to diverse RNA-seq count data, and our ablation study empirically identifies crucial architecture configurations required for strong transfer performance. Additionally, we curate a public training corpus, DiverseRNA-1.4M, and find that TxFM trained on this curated dataset yields high-fidelity gene representations that outperform FMs trained on atlas-scale corpora over $100\times$ larger. Overall, our results indicate that inductive self-supervised learning is a viable modeling approach for transcriptomics representation, provided a careful synthesis of model architecture and training data curation.


Efficient and Robust Physical 3DGS-MPM Simulation via Interior Filling and Text-Physics Optimization

Yikun Ma ⋅ Yiqing Li ⋅ Jingwen Ye ⋅ Zhongkai Wu ⋅ Weidong Zhang ⋅ Lin Gao ⋅ Zhi Jin

Extending 3D Gaussian Splatting (3DGS) to 4D physical simulation remains challenging. Based on the Material Point Method (MPM), existing methods either rely on manual parameter tuning or distill dynamics from video diffusion models, limiting the generalization and optimization efficiency. Recent attempts using LLMs/VLMs suffer from a text/image-to-3D perceptual gap, yielding unstable physics behavior. In addition, they often ignore the surface structure of 3DGS, leading to implausible motion. We propose FastPhysGS, a fast and robust framework for physics-based dynamic 3DGS simulation: (1) Instance-aware Particle Filling (IPF) with Monte Carlo Importance Sampling (MCIS) to efficiently populate interior particles while preserving geometric fidelity; (2) Bidirectional Graph Decoupling Optimization (BGDO), an adaptive strategy that rapidly optimizes material parameters predicted from a VLM. Experiments show FastPhysGS achieves high-fidelity physical simulation in 1 minute using only 7 GB runtime memory, outperforming prior works with broad potential applications.


Efficient One-to-many Domain Translation via Diffusive Entropic Optimal Transport

David Vaneev ⋅ Ivan Shchekotov ⋅ Denis Rakitin ⋅ Dmitry Vetrov

Domain translation requires a delicate balance between realism, input-output alignment, and output diversity. Entropic optimal transport (EOT) provides a principled formulation of this trade-off, but its practical use remains challenging. Directly maximizing entropy of a stochastic transport plan requires evaluating log densities of generator-induced distributions, which scales poorly with dimension. In contrast, diffusion models make density information accessible in high dimensions through noising the corresponding distributions and approximating their score functions. In this work, we introduce the diffusive entropy regularizer that measures diversity after progressively noising the conditional output distribution, making the entropy term compatible with score estimation. The resulting Diffusive EOT retains the main theoretical guarantees of EOT: a unique solution, controllable diversity, and convergence to unregularized OT solution in the zero-regularization limit. We then introduce DM-EOT, a practical algorithm for approximately solving the Diffusive EOT that combines diffusion-based distribution matching with transport cost and diffusive entropy penalties. DM-EOT supports both fast one-step generation and higher-fidelity diffusion-based multi-step sampling. Both variants achieve comparable or superior performance on unpaired domain translation benchmarks among one-to-many baselines with comparable inference cost and diversity.


Emergent Semantic Role Understanding in Language Models

Carla Griffiths ⋅ Mirco Musolesi

Understanding how linguistic structure emerges in language models is central to interpreting what these systems learn from data and how much supervision they truly require. In particular, semantic role understanding (``who did what to whom’’) is a core component of meaning representation, yet it remains unclear whether it arises from pre-training alone or depends on task-specific fine-tuning. In this paper, we study whether semantic role understanding emerges during language model pre-training or requires task-specific fine-tuning. We freeze decoder-only transformers and train linear probes to extract semantic roles, using performance to infer whether role information is already encoded in pre-training or learned during adaptation. Across model scales, we find that frozen representations contain substantial semantic role information, with performance improving but not fully matching fine-tuned models. This indicates partial but incomplete emergence from pre-training alone. We show that semantic role structure emerges from language modeling objectives, but its internal implementation shifts toward more distributed representations as model scale increases.


Endowing Your Vision-Language-Action Model with a Predictive Mind

Pengxiang Ding ⋅ Haoying Wang ⋅ Minghui Lin ⋅ Qishen Wang ⋅ Zhenyu Ding ⋅ Runze Suo ⋅ Xuanxuan An ⋅ Wenxuan Song ⋅ Fuhao Li ⋅ Han Zhao ⋅ Donglin Wang ⋅ Ning DING

Vision-Language-Action (VLA) models transfer semantic priors from large-scale vision-language models to robotic manipulation, but their backbones are typically optimized for static visual understanding rather than action-conditioned dynamics. Consequently, the features used by the policy may lack predictive information about how the scene evolves under the robot's actions. Existing future-prediction methods either reconstruct pixels, overemphasizing appearance details irrelevant to control, or use auxiliary latent predictors that remain weakly coupled with the policy-consumed backbone features. We propose PredMind, a plug-and-play predictive representation learning framework that treats future prediction as a training signal for the VLA backbone itself. Instead of reconstructing future pixels, PredMind aligns intermediate backbone representations with their future counterparts conditioned on executed actions, shifting supervision from low-level appearance to the model's semantic feature space. This directly injects action-conditioned temporal structure into the representation hierarchy used for control while preserving pretrained VLM priors. All auxiliary prediction components are discarded after training, introducing no architectural changes, additional parameters, or inference latency at deployment. Across simulation benchmarks, real-world manipulation tasks, and multiple VLA backbones, PredMind consistently improves learning efficiency and final performance. Especially on LIBERO, it matches or surpasses a fully trained base VLA using only 1/60 of the training iterations, demonstrating the benefit of embedding predictive dynamics directly into backbone features used for control. Anonymous code is available at https://anonymous.4open.science/r/PredMind.

Although score-based generative models (SGMs) have achieved remarkable success in real-world sample generation tasks, it is still far from sufficient to understand them from a mathematical perspective. In this paper, we aim to provide enhanced convergence guarantees for two existing SGMs in $\mathcal{W}_2$-distance beyond log-concavity. More precisely, we develop a novel framework of error analysis for Euler-Maruyama (EM) and Poisson midpoint (PM) time discretization schemes for SGMs. Under a non-log-concavity condition, we show that $\tilde{\mathcal{O}}(\sqrt{d}/\epsilon)$ iterations suffice to approximate the target distributions in $\epsilon$-accuracy for the classical EM-based SGM, considerably improving upon the existing iteration complexity $\tilde{\mathcal{O}}(d/\epsilon^2)$. Notably, the novel framework of error analysis enables us to establish an $\tilde{\mathcal{O}}(\sqrt{d}/\epsilon^{2/3})$ iteration complexity for the PM-based SGM in $\epsilon$-accuracy, significantly outperforming current state-of-the-art Wasserstein convergence guarantees for SDE-based diffusion samplers.

Variable-order Markov models generate by backing off to the longest usable suffix of the generated history, while regular constraints describe finite horizon controls such as fixed endings, metrical patterns, and anti-copy rules. Existing belief-propagation methods give exact regular-constrained sampling for first-order Markov chains, but a first-order state merges histories that a variable-order generator deliberately keeps distinct. We give the corresponding variable-order construction: replace the Markov state by the sparse observed context state and compose this context graph with the regular constraint automaton. For a fixed context graph and automaton, inference is linear in the horizon and in the number of reachable product edges, without expanding to all $|\mathcal{V}|^K$ histories; the same source-row interface handles reversible augmentation by inverse count lookup, without materializing transformed corpora. Tiny enumerated examples verify exact partition functions and conditional probabilities; Bach Prelude experiments show sparse regular-constrained sampling at $K=6$ and 12-key virtual transposition semantics from $592$ stored events.


Eyes on VLM: Benchmarking Gaze Following and Social Gaze Prediction in Vision Language Models

Hengfei Wang ⋅ Anshul Gupta ⋅ Pierre Vuillecard ⋅ Jean-marc Odobez

Vision-language models (VLMs) have rapidly evolved into general-purpose multimodal reasoners with strong zero-shot generalization. In this context, VLMs could greatly benefit the analysis of human gaze and attention, a central task in human behavior understanding that requires reasoning about the physical scene as well as the activity, interactions, and social context. However, the extent to which VLMs can reliably understand human gaze and related attentional behaviors remains largely unexplored. In this work, we present EyeVLM, a systematic evaluation framework for gaze understanding in VLMs across two complementary dimensions: tasks and models. To assess gaze understanding capabilities, we focus on two core tasks. The first, gaze following, i.e., predicting the 2D location where a person is looking, has a geometric and visual processing focus, requiring a precise understanding of the human face, attention direction, 3D scene structure, and spatial grounding of attended targets. The second, social gaze prediction, requires social and relational reasoning over multi-person interactions (e.g., mutual gaze and shared attention), and may benefit more from the LLM semantic reasoning capabilities within VLMs. Regarding models, EyeVLM evaluates these tasks in two ways: a zero-shot setting with a diverse set of state-of-the-art open- and closed-source VLMs, exploring different prompting strategies; and a fine-tuning approach based on task-specific QA pairs, studying the impact of model scale and data scale. As benchmarks, we rely on existing gaze understanding datasets and perform a systematic comparison with state-of-the-art purely visual models. Overall, our results show that current VLMs lack precise gaze understanding capabilities. While standard training helps reduce the gap with visual models, significant improvements are still needed.


FacePhys: State of the Heart Learning

Jack Tang ⋅ Kegang Wang ⋅ Yuntao Wang ⋅ Xin Liu ⋅ Yuxuan Fan ⋅ Jiatong Ji ⋅ Yuanchun Shi ⋅ Daniel McDuff

Vital sign measurement using cameras presents opportunities for comfortable, ubiquitous health monitoring. Remote photoplethysmography (rPPG), a foundational technology, enables cardiac measurement through minute changes in light reflected from the skin. However, practical deployment is limited by the computational constraints of performing analysis on front-end devices and the accuracy degradation of transmitting data through compressive channels that reduce signal quality. We propose a memory efficient rPPG algorithm - FacePhys - built on temporal-spatial state space duality, which resolves the trilemma of model scalability, cross-dataset generalization, and real-time operation. Leveraging a transferable heart state, FacePhys captures subtle periodic variations across video frames while maintaining a minimal computational overhead, enabling training on extended video sequences and supporting low-latency inference. FacePhys establishes a new state-of-the-art, with a substantial 49% reduction in error. Our solution enables real-time inference with a memory footprint of 3.6 MB and per-frame latency of 9.46 ms -- surpassing existing methods by 83% to 99%. The training code, a real-time inference demo for mobile browsers, and the demonstration video can be found in the supplementary materials.

The Procrustes-Wasserstein problem asks to match the vectors in two unlabeled point clouds $X$ and $Y$, which are related via a latent orthogonal transformation. This problem appears in unsupervised representation alignment, domain adaptation, and geometric matching, but its computational theory remains limited. We study a planted Gaussian model in which $Y$ is obtained from $X$ via an unknown relabeling, an unknown orthogonal transformation, and additive noise. Departing from the isotropic setting considered in prior work, we assume that the covariance of $X$ is anisotropic, with a power-law spectrum. While closer to real-world data, this structure also enables a directions-first approach: Estimate the dominant eigenspaces of both point clouds, then recover the matching. We prove that a polynomial-time version of this procedure exactly recovers the true relabeling with high probability under explicit scaling conditions. To our knowledge, this is the first computational recovery guarantee for Procrustes-Wasserstein with nontrivial observational noise and dimensions beyond $\log n$, where $n$ is the number of points. Experiments on synthetic data and 3D shapes support the theory, showing that anisotropy favors directions-first methods, can lower runtime and improve recovery.


Fast, Relaxation‑ and Hyperparameter‑Free Pairwise Worst-Case Class Separation

Mohammad Mahdi Omati ⋅ Arash Amini ⋅ Nezam Mahdavi-Amiri

In this paper, we investigate a novel discriminative dimensionality reduction method based on maximizing the minimum pairwise ratio of between-class to within-class scatter. This objective function enhances class separability by providing critical, adaptive control over the variance within each class pair. The resulting max-min fractional program is non-convex and challenging to solve. Our main contribution is F$^3$-PWCRA (Fast, Free of relaxation, Free of hyperparameters for Pairwise Worst-Case Ratio Analysis), a provably convergent two-level algorithm: an outer generalized Dinkelbach-type loop transforms the fractional objective into subtractive subproblems, ensuring the global convergence of the outer algorithm. Then, we propose a bisection strategy as an alternative, which offers robust convergence and straightforward implementation. For the inner loop, we develop an efficient minorization-maximization (MM) algorithm that tackles the non-convex subproblem by iteratively solving a simple quadratic program (QP), which we derive from the dual of a convex surrogate. F$^3$-PWCRA is computationally efficient, free from semidefinite relaxations and hyperparameter tuning, and extensive benchmarks show it consistently outperforms state-of-the-art methods in classification error.


Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers

Evelyn Turri ⋅ Davide Bucciarelli ⋅ Sara Sarto ⋅ Lorenzo Baraldi ⋅ Marcella Cornia

Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal mechanisms through which prompts shape image semantics remain poorly understood. In this work, we study massive activations: a small subset of hidden-state channels whose responses are consistently much larger than the rest. We show that, despite their sparsity, these few channels effectively draw the whole picture, in three complementary senses. First, they are functionally critical: a controlled disruption probe that zeroes the massive channels causes a sharp collapse in generation quality, while disrupting an equally-sized set of low-statistic channels has marginal effect. Second, they are spatially organized: restricting image-stream tokens to massive channels and clustering them yields coherent partitions that closely align with the main subject and salient regions, exposing a structured spatial code hidden inside an apparently outlier-like subspace. Third, they are transferable: transporting massive activations from one prompt-conditioned trajectory into another, shifts the final image toward the source prompt while preserving substantial content from the target, producing localized semantic interpolation rather than unstructured pixel blending. We exploit this property in two use cases: text-conditioned and image-conditioned semantic transport, where massive activations transport enables prompt interpolation and subject-driven generation without any additional training. Together, these results recast massive activations not as activation anomalies, but as a sparse prompt-conditioned carrier subspace that organizes and controls semantic information in modern DiT models. All the code will be publicly available.


FiLM-CAM: Keyed Feature Modulation for Conditional-Access Watermarking

Yash Kulthe ⋅ Vishal Asnani ⋅ Shruti Agarwal ⋅ Andrew Gilbert ⋅ John Collomosse

We propose FiLM-CAM; a conditional-access image watermarking approach that enforces authorized encoding, decoding, and removal through keyed feature modulation. A secret key conditions a Feature-wise Linear Modulation (FiLM) mechanism that transforms the internal representation of the watermark signal before embedding, thereby coupling the payload to the key at the feature, rather than pixel level. The same key must be presented to condition the decoder to correctly recover the payload embedded in an image. To enforce conditional access, we train the system to maximize the performance gap between authorized and unauthorized use, ensuring reliable extraction under the correct key while inducing failure under incorrect keys. In addition, we design our encoder to estimate the watermark residual invariant to the presence of pre-existing watermarks. This design enables conditional watermark removal via simple re-encoding and subtraction of the residual using iterative refinement, contingent on knowledge of the secret key. We evaluate FiLM-CAM across multiple watermarking encoder–decoder backbones, demonstrating that keyed feature modulation serves as an effective, architecture-agnostic conditional access module (CAM) for image watermarking.


Fixed-Point Masked Generative Modeling

Andrea Miele ⋅ Yiming QIN ⋅ Alba Carballo Castro ⋅ Justin Deschenaux ⋅ Pascal Frossard

Masked Generative Models (MGMs) enable parallel decoding and achieve strong performance across modalities, but require full-sequence bidirectional transformers at every step, making training costly and degrading quality under low sampling budgets. Existing work improves efficiency via better samplers or fixed-depth backbones, but does not vary the depth of the denoiser across sampling steps. We introduce Fixed-Point Masked Generative Models (FP-MGMs), which replace part of the denoiser with a fixed-point solver over shared attention layers to enable adaptive depth with fewer parameters. To make it more effective for masked generation, we first introduce a cross-step consistency loss, which aligns hidden representations at neighboring denoising steps and, second, three-state reuse (3SR) which warm-starts the solver using the previous solution by treating unchanged, still-masked, and newly revealed tokens differently. Together, these components define our complete training-to-inference framework for fixed-point masked generation, \emph{CoFRe}. We also show that pre-trained MGMs can be converted into FP-MGMs with short fine-tuning, avoiding full retraining. Across modalities, CoFRe improves the quality and cost trade-off. On OpenWebText, CoFRe reduces parameters by 38.8\%, training time by 11.5\%, and VRAM by 16.9\%, while improving generative perplexity from 830.8 to 101.8 at a budget of $96$ transformer-block forward passes, compared to MDLM. In ImageNette, CoFRe reduces training time by 48.6\% and VRAM by 50.7\%, while improving FID in all sample budgets tested. Overall, CoFRe offers a practical framework for cheaper training and stronger low-budget masked generation.


Flow-Transformed Implicit Processes for Function-Space Variational Inference

Luis Antonio Ortega Andrés ⋅ Andres Masegosa ⋅ Thomas Nielsen

Implicit-process priors define distributions over functions through flexible generative mechanisms, making them attractive for Bayesian function-space modelling. However, performing posterior inference with such priors is challenging because their induced function-space distributions are typically not available in closed form. One practical strategy is to approximate the prior using a finite collection of sampled functions, and then represent posterior functions as learned combinations of these samples. Existing approaches commonly place a Gaussian variational distribution over the combination weights. While tractable, this choice limits the shapes of posterior uncertainty that can be represented, especially when the true posterior is asymmetric, heavy-tailed, or multimodal. We propose Flow-Transformed Implicit Processes (FTIP), a variational inference method that makes this finite-dimensional function-space approximation more expressive. Instead of using a Gaussian distribution over the combination weights, FTIP uses a normalizing flow to define a richer variational distribution. This induces a flexible posterior distribution over functions while preserving tractable optimization. We train the model using a Black-Box $\alpha$ objective, allowing us to compare mass-covering and mode-seeking variational behaviour. Experiments show that FTIP captures asymmetric and multimodal posterior structure in function space that Gaussian coefficient approximations tend to smooth or collapse.

Critic-free reinforcement fine-tuning (RFT) for agentic large language models is dominated by the GRPO and DPO families, yet these methods often update too aggressively on the high-variance samples from agentic tasks. We propose Follow the Winners (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, obtaining a more conservative policy-learning method. Unlike prior exponentially concentrating approaches, FTW induces polynomial concentration in the order statistic of returns. Through a control-as-inference lens, we show that FTW induces a mild risk-seeking bias that scales linearly with return uncertainty: less aggressive than DPO’s quadratic risk bias, yet more exploratory than GRPO’s risk-neutral profile. In deterministic-reward agentic settings, where uncertainty is primarily epistemic, this linear risk bonus encourages knowledge-seeking without entrenching on noisy samples. Consistent with this perspective, FTW matches GRPO on the perfect-information Sokoban domain while outperforming it on the imperfect-information Search-R1 domain.


Foveated BagNet: Inherent Interpretability Does Not Exclude Global Context

Holger Heidrich ⋅ Sarah Müller ⋅ Andreas Schilling

Deploying deep learning in high-stakes settings demands models that are not only accurate but interpretable. Interpretable-by-design models, exemplified by BagNet, address this by restricting each local class predictor to a small spatial patch, yielding inherently explainable predictions---but at the cost of capturing large-scale image features, resulting in a substantial accuracy penalty. We propose FovBagNet, a foveated extension of BagNet inspired by the mammalian retina. Each local predictor depends on a stack of multiscale patches centered at its location, providing high-resolution detail at the center and progressively coarser context toward the periphery. FovBagNet achieves a top-1 ImageNet accuracy of 0.74, closing most of the 11-point gap between BagNet-33 (0.65) and ResNet-50 (0.76), while retaining the inherent interpretability of BagNet in the form of spatially accurate saliency maps. We also identify a conceptual limitation of the original BagNet saliency maps and empirically assess its impact.


Frame the adversary: a structure-aware attack methodology

Vicky Kouni ⋅ Stelios Perrakis ⋅ Francis Bach ⋅ Pascal Frossard ⋅ Yann Chevaleyre

Frequency-based adversarial attacks have recently grown popular by exploiting spectral sensitivities shared across neural architectures. Unlike spatial perturbations, frequency-based attacks expose deeper vulnerabilities, making them especially valuable for robust evaluation of safety-critical and security-sensitive applications. Yet, existing approaches are typically not derived as solutions to an optimization problem that explicitly captures transform-domain structure. In this paper, we propose a methodology for crafting principled frequency-based adversarial attacks, via a dedicated optimization framework. A cornerstone of our method hinges on the introduction of a perturbation constraint set, tied to highly structured non-orthogonal transforms, well-known for their flexible, non-predefined frequency handling. We prove that the attacks emerge as weighted $\ell_2$-projections onto this set, yielding a general and controlled attack generation mechanism. By this, we provide a clear geometric attack characterization, ensuring alignment between the optimization objective and the perturbation constraint. We assess our framework on standardized datasets, for pretrained and adversarially robust models. Results highlight that our attacks, being solutions to an optimization problem, over a structured perturbation set, are highly effective, even across different, unseen architectures. Our methodology could serve as a theoretical baseline for designing and analyzing transformed-based attacks, targeting fundamental model vulnerabilities, instead of mere architecture-specific artifacts typically studied in the robustness literature.


Generalized Bayes for Causal Inference

Emil Javurek ⋅ Dennis Frauen ⋅ Yuxin Wang ⋅ Stefan Feuerriegel

Uncertainty quantification is central to many applications of causal machine learning, yet principled Bayesian inference for causal effects remains challenging. Standard Bayesian approaches typically require specifying a probabilistic model for the data-generating process, including high-dimensional nuisance components such as propensity scores and outcome regressions. Standard posteriors are thus vulnerable to strong modeling choices, including complex prior elicitation. In this paper, we propose a \textit{generalized Bayesian framework for causal inference}. Our framework avoids explicit likelihood modeling; instead, we place priors directly on the causal estimands and update these using an identification-driven loss function, which yields generalized posteriors for causal effects. As a result, our framework turns existing loss-based causal estimators into estimators with full uncertainty quantification. Our framework is flexible and applicable to a broad range of causal estimands (e.g., ATE, CATE). Further, our framework can be applied on top of state-of-the-art causal machine learning pipelines (e.g., Neyman-orthogonal meta-learners). For Neyman-orthogonal losses, we show that the generalized posteriors converge to their oracle counterparts and remain robust to first-stage nuisance estimation error. With calibration, we thus obtain valid frequentist uncertainty even when nuisance estimators converge at slower-than-parametric rates. Empirically, we demonstrate that our proposed framework offers causal effect estimation with calibrated uncertainty across several causal inference settings. To the best of our knowledge, this is the first flexible framework for constructing generalized Bayesian posteriors for causal machine learning.


Generate in Reconstruction Space, Match in Semantic Space: Transport Geometry for One-Step Generation

Hugues Van Assel ⋅ Edward De Brouwer ⋅ Saeed Saremi ⋅ Gabriele Scalia ⋅ Aviv Regev

Generative modeling and self-supervised representation learning (SSL) optimize structurally different objectives: generative training rewards distributional fidelity, while SSL rewards semantic coherence. Yet recent work repeatedly finds that SSL features improve generative training, though the mechanism of this synergy remains unclear. Here, we study the benefits of SSL in generative modeling in the framework of one-step generation where the role of representation is explicit: frozen SSL features are used to match generated samples to real data. We use the Sinkhorn divergence in that feature space, providing a tractable surrogate for the Wasserstein distance, the population-level discrepancy approximated by Fr\'echet-style evaluation metrics (such as FID). We find that this objective becomes highly effective when computed in a semantically structured SSL feature space (a 39$\times$ reduction in ImageNet FID). We trace this behavior primarily to matching estimation: semantic SSL features that suppress nuisance reconstruction details induce a more compact geometry, making distribution matching more tractable. As a consequence, the best training SSL features need not match the features used by the evaluation metric. In particular, we show that using Inception as the feature extractor can improve FID while degrading matching stability and sample quality, revealing a form of metric hacking. Using extensive experiments on ImageNet, we identify which SSL feature families lead to best generation performance and show that matching stability is a quantitative criterion for selecting them.


GeoDial: A Multimodal Dialog Tutoring Dataset for Geometry Problem-Solving with Visual Tutor Turns

Sankalan Pal Chowdhury ⋅ Junling Wang ⋅ Donya Rooein ⋅ April Wang ⋅ Mrinmaya Sachan

Several educational domains rely heavily on diagrams and visual cues, yet most existing tutoring datasets are limited to text-only interactions. This limits the development of AI tutors that can teach in visually grounded ways used by human instructors. Thus, we introduce GeoDial, a multimodal tutoring dataset of over 1.3K teacher-student dialogs in the domain of geometry collected from experienced math teachers, where instructional turns are explicitly grounded in diagram highlights. We propose a scalable annotation protocol that integrates dialog acts, visual highlighting, and feedback, enabling fine-grained supervision of both language and visual tutoring behavior. To illustrate the challenges posed by this setting, we fine-tune several vision–language models on GeoDial and evaluate their ability to generate tutoring utterances and diagram highlights. While supervised fine-tuning substantially improves the quality of generated dialog, it struggles to produce accurate diagram highlights, revealing a key limitation of current methods and highlighting the need for approaches that more effectively integrate visual reasoning with pedagogical interaction.


Global Convergence in Deep Networks via the Second Law of Thermodynamics

Matan Tsipory ⋅ Ofir Gaash ⋅ Itamar Harel ⋅ Nati Srebro ⋅ Daniel Soudry

Existing convergence analyses of finite-width overparameterized deep neural networks under standard parameterization often rely on parameters remaining close to their initialization and typically neglect the role of noise, despite empirical evidence that stochasticity can improve optimization and generalization. To bridge this gap, we study $L_2$-regularized continuous-time Langevin dynamics~(CLD) for training a feedforward neural network with smooth activations. We introduce a change-of-measure inequality, inspired by the second law of thermodynamics, that enables control of key properties of the training dynamics under the time-evolving parameter distribution. Leveraging a sharpened lower bound on the minimum eigenvalue of the Neural Tangent Kernel (NTK) at initialization, we show that injected noise helps preserve NTK stability with high probability without confining the weights to a small neighborhood of their initialization. As a result, we derive a non-asymptotic upper bound on the expected training loss. Our analysis yields explicit conditions on width and noise temperature that guarantee convergence to a finite loss floor. Finally, we experimentally show the network Jacobian can change from initialization, departing the commonly analyzed lazy regime.


GradShield: Alignment Preserving Finetuning

Zhanhao Hu ⋅ Xiao Huang ⋅ Patrick Mendoza ⋅ Emad A Alghamdi ⋅ Basel Alomair ⋅ Raluca Popa ⋅ David Wagner

Large Language Models (LLMs) pose a significant risk of safety misalignment after finetuning, as models can be compromised by both explicitly and implicitly harmful data. Even some seemingly benign data can inadvertently steer a model towards misaligned behaviors. To address this, we introduce GradShield, a principled filtering method that safeguards LLMs during finetuning by identifying and removing harmful data points before they corrupt the model's alignment. It removes potentially harmful data by computing a Finetuning Implicit Harmfulness Score (FIHS) for each data point and employs an adaptive thresholding algorithm. We apply GradShield to multiple utility fine-tuning tasks across varying levels of harmful data and evaluate the safety and utility performance of the resulting LLMs using various metrics. The results show that GradShield outperforms all baseline methods, consistently maintaining an Attack Success Rate (ASR) below $6\%$ while preserving utility performance.


Half-Truths Break Similarity-Based Retrieval

Bora Kargi ⋅ Arnas Uselis ⋅ Seong Joon Oh

When a text description is extended with an additional detail, image-text similarity should drop if that detail is wrong. We show that CLIP-style dual encoders often violate this intuition: appending a plausible but incorrect object or relation to an otherwise correct description can increase the similarity score. We call such cases \emph{half-truths}. On the verified COCO evaluation set, CLIP prefers the correct shorter description only 41.0\% of the time, and performance drops to 28.2\% when the added detail is a relation. We trace this vulnerability to weak supervision on caption parts: contrastive training aligns full sentences but does not explicitly enforce that individual entities and relations are grounded. We propose CS-CLIP (Component-Supervised CLIP), which decomposes captions into entity and relation units, constructs a minimally edited foil for each unit, and fine-tunes the model to score the correct unit above its foil while preserving standard dual-encoder inference. CS-CLIP raises half-truth accuracy to 78.2\% and improves average performance on established compositional benchmarks by 5.7 points, suggesting that reducing half-truth errors aligns with broader gains in compositional understanding.

We define a new variant of transformers called recursive transformers, which generate sequences of vectors without discretisation in each step. We show that over ordered ring extensions of the integers recursive transformers with circuit activation functions and hard attention are computationally equivalent to the well-established algebraic model of BSS-machines. If the transformers use feedforward networks instead, they are equivalent to linear BSS-machines with degree-2 branching.


HARP: Training-Free Dual-Profile Agentic Communication for LLM-Based Recommendation

Zhen Tao ⋅ Xun Zhou ⋅ Xuhui Chen ⋅ Ziyue Qiao ⋅ Qingqiang Sun

In LLM-based recommendation, memory effectively plays the role of a user-level agent, typically built on top of a memory-augmented backbone. Yet under this paradigm users and items remain asymmetric participants: only the user is given a structured profile; that profile is consumed inside a bandwidth-limited reranker prompt; and it is frozen after warmup. Starting from this observation, we recast LLM-based recommendation as an agentic-communication process in four stages --- Build a self-description, Transmit it along the collaborative graph, Understand the other side, and Sync it from the next interaction. Analysing the dominant memory-augmented paradigm through this formulation, we identify three open communication challenges that current pipelines leave unaddressed: the item side never produces a comparable self-description (one-sided Build); profile-level signals are routed only into a bandwidth-limited reranker prompt rather than into the retrieval stage where they would not compete for context tokens (prompt-trapped Understand); and the user's description is frozen after warmup with no principled update rule (absent Sync). We instantiate the missing stages as HARP (Homophily-aware Agentic Recommendation via dual Profiles}), a training-free, three-operator addition to a frozen LLM backbone: a rule-extracted symmetric item profile (zero extra LLM calls), a channel-decomposable profile-level homophily $\mathcal{H}(\phi_u,\phi_i)$ consumed at retrieval time rather than inside the reranker prompt, and a critic-gated single-pass reflective update realising in-context profile evolution without any gradient step. On four public benchmarks (Yelp, Amazon Books, MovieTV, Goodreads), HARP lifts Hit@1 over the strongest published baseline by $6.9\%$--$24.0\%$, with family-wise Holm-corrected $p<10^{-9}$ on every dataset. Ablations and a per-case attribution study isolate each operator as a distinct, complementary source of gain; we also quantify and openly report the trade-off introduced by the reflective pass. Our contribution is a unified formulation of agentic communication for LLM-based recommendation, together with a training-free three-operator instantiation on top of it.


HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice

Sarah Meiklejohn ⋅ Sunny Consolvo ⋅ Patrick G Kelley ⋅ Tara Matthews ⋅ Sai Teja Peddinti ⋅ Renee Shelby ⋅ Lenin Simicich ⋅ Kurt Thomas

People increasingly turn to large language models (LLMs) for advice on numerous aspects of their lives. This paper investigates LLM capabilities in responding to questions about digital privacy, safety, and security, such as regaining access to a lost or suspended account, removing malware from a device, dealing with scammers or technology-facilitated abuse, and more. We curated a benchmark of 450 questions representing authentic user situations and developed rubrics for each question to evaluate factual accuracy and tone-based qualities of a response. We then developed and applied an auto-rater to evaluate responses from 18 state-of-the-art LLMs. Our results indicate that while models provide high-quality advice (with scores of 82\% on average), all models have significant room to improve, especially when tailoring advice to complex or high-risk scenarios.


Hierarchical Conformal Classification

Floris den Hengst ⋅ Inès Blin ⋅ Majid Mohammadi ⋅ Syed I Shah ⋅ Taraneh Younesian

Conformal prediction (CP) provides reliable uncertainty quantification by generating prediction sets with finite-sample coverage guarantees, yet it typically ignores hierarchical relationships between classes. We introduce Hierarchical Conformal Classification (HCC), a framework that integrates class hierarchies into the structure and semantics of prediction sets. HCC uses a constrained optimization problem formulation to produce prediction sets composed of nodes at various hierarchy levels while maintaining rigorous coverage guarantees. To ensure computational efficiency, we prove that a restricted subset of well-structured candidate solutions is sufficient to maintain both optimality and coverage. Empirical evaluations across audio, image, and text benchmarks, alongside a user study, demonstrate that HCC outperforms state-of-the-art methods and that hierarchical prediction sets align better with human preferences.

Probabilistic prediction heads in neural networks typically output either a Gaussian mixture or a single conformal region. Neither separates the distinct sources of uncertainty often present in real prediction tasks: a discrete choice among modes, bounded systematic drift within the chosen mode, and irreducible stochastic noise. We introduce the Hybrid Probabilistic Zonotope (HProbZ), an output head that represents these three sources as binary, bounded, and stochastic generators of a zonotope, and admits a closed-form likelihood by convolution. Sharing the bounded generator across prediction steps couples future predictions algebraically, so observing one step refines the predictive distribution at every remaining step in a single forward pass. We establish that the three generators are identifiable from the likelihood up to permutation, and that an HProbZ density is representationally distinct from any finite Gaussian mixture. The same shared structure provides analytic per-mode risk and distribution-free multi-modal conformal sets at inference time. Empirical analysis on representative prediction benchmarks supports the effectiveness of the design relative to same-encoder mixture baselines, while offering structural properties that mixture or convex-conformal predictors do not jointly provide.

This paper presents a new method for the zero-shot open-vocabulary semantic segmentation (OVSS) of 3D automotive lidar data. To circumvent the recognized image-text modality gap that is intrinsic to approaches based on Vision Language Models (VLMs) such as CLIP, our method relies instead on image generation from text, to create prototype images. Given a 3D network distilled from a 2D Vision Foundation Model (VFM), we then label a point cloud by matching 3D point features with 2D image features of these prototypes. Our method is state-of-the-art for OVSS on nuScenes and SemanticKITTI.


InCLAD: A Continual Learning Benchmark for Industrial Visual Anomaly Detection

Łukasz Marcjan ⋅ Dorota Wilk-Kołodziejczyk ⋅ Roberto Corizzo

Industrial anomaly detection benchmarks have advanced rapidly, yet they remain largely static and therefore underrepresent deployment settings in which inspection systems must be extended from one industrial component to another over time. In our setting, continual industrial anomaly detection is framed as a task-incremental problem: each component defines a task, while the learning objective remains semi-supervised. The goal extends beyond conventional anomaly detection, since models should preserve their performance on all components as they are sequentially exposed to new components. This framing is especially relevant because many existing industrial pipelines either assume fully supervised defect labels or study only isolated datasets, leaving open the question of how semi-supervised anomaly detectors behave across a sequence of heterogeneous components. We address this gap by organizing the benchmark as a sequence of related component-level tasks and by defining multiple scenarios, including random, easy-to-hard, and hard-to-easy. In addition to single-dataset continual scenarios, we propose challenging scenarios spanning across multiple datasets. We further evaluate continual performance from two complementary perspectives: image-level anomaly detection and pixel-level anomaly localization. By recasting industrial anomaly detection as a continual benchmark rather than a static train/test problem, our work establishes a concrete path toward studying knowledge retention, adaptation, and generalization in evolving industrial environments.

Large Language Model training is increasingly constrained by memory. Activation checkpointing reduces the activation memory footprint by discarding selected activations during the forward pass and rematerializing them during backpropagation, trading memory savings for additional FLOPs. Although essential for scaling models, this overhead makes rematerializing linear layers of transformers prohibitively expensive. We introduce a novel method for approximate rematerialization that uses information from the initial forward pass to selectively recompute only a fraction of the original layer components. These components are either chosen greedily, or sketched, yielding unbiased gradient estimates. Through ablation studies, we show that our method closely matches the original gradient. When applied to FFN layers, attention layers, or both, our method reduces the computational cost for the rematerialization and backward pass of linear layers by 33\% on a Llama-488M, with perplexity increases of only 1.7\%, 1.2\%, and 2.5\%.


Large language models suffer from a curse of ambiguity

Nicolas Zucchet ⋅ Hyun Dong Lee ⋅ Scott Linderman

Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity of their learned distributions more critical than ever. Yet, not all distributions are equally easy to learn. In this work, we identify a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately. Through an extensive theoretical analysis, we trace this curse to architectural and learning roots. More ambiguous distributions require more capacity to be stored, larger embeddings to be represented, more steps to be fitted, and amplify token-sampling noise. We validate these findings on synthetic tasks with controlled ground truth and observe the same signatures in language models trained on real data. Our results provide a new perspective on the statistical capabilities of large language models and a practical framework for when to trust their output distribution.

Understanding the internal “thinking” process of Large Language Models (LLMs) and the cause of hallucinations remains a key challenge. To this end, we introduce latent debate, a structured surrogate framework for interpreting model outputs on True/False prediction tasks through the lens of internal latent arguments and interactions amongst them. Unlike human debates, latent debate captures the hidden supporting and attacking signals that arise within a model during a single inference. We first present a model- and task-agnostic conceptual framework, and then instantiate it symbolically to approximate the thinking process of LLMs towards binary decisions. Empirical studies demonstrate that our latent debate is a faithful structured surrogate model that has highly consistent predictions with the original LLM, while providing a form of interpretability. We also demonstrate that our latent debate provides a strong baseline for hallucination detection. Specifically, we identify strong correlations between debate patterns and hallucinations, such as a high degree of disagreement in the middle layers of the latent debate surrogate is linked to a higher risk of hallucinations. Our findings suggest that latent debate shows potential to analyze internal signals in LLMs for binary decision settings.

Adversarial training attains strong empirical robustness to specific adversarial attacks by training on concrete adversarial perturbations, but it produces neural networks that are not amenable to strong robustness certificates through neural network verification. On the other hand, earlier certified training schemes directly train on bounds from network relaxations to obtain models that are certifiably robust, but display sub-par standard performance. Recent work has shown that state-of-the-art trade-offs between certified robustness and standard performance can be obtained through a family of losses combining adversarial outputs and neural network bounds. Nevertheless, differently from empirical robustness, verifiability still comes at a significant cost in standard performance. In this work, we propose to leverage empirically-robust teachers to improve the performance of certifiably-robust models through knowledge distillation. Using versatile feature-space distillation objectives, we show that distillation from adversarially-trained teachers can significantly improve the performance of state-of-the-art certified training algorithms for ReLU networks on robust computer vision benchmarks.


Lethe: Link Inference Attacks For Evaluation of Edge Unlearning Methods

Andrea D'Angelo ⋅ Francesco Gullo ⋅ Davide Mottin ⋅ Sayan Ranu ⋅ Giovanni Stilo

Graph Unlearning (GU) aims to remove the influence of specific training data from a Graph Neural Network without retraining anew. While most methods address removal of nodes and features, Edge Unlearning (EU) aiming at forgetting edges has recently attracted dedicated attention as a standalone problem due to the hardness of completely removing the impact of an edge in the model. Despite this, no dedicated benchmark for EU exists besides isolated evaluations on feature-rich datasets such as Cora and Citeseer where GNNs exploit node features rather than graph structure. As such, we show two systematic limitations in existing evaluations: dedicated EU methods can be slower than retraining anew, and accuracy metrics are uninformative due to the negligible impact of edge removal. To address this gap, we introduce Lethe, the first comprehensive benchmark for Edge Unlearning, featuring 15 methods and 8 datasets. Our benchmark emphasizes large graphs, introduces evaluation tasks of varying difficulty, and proposes empirical measures based on link inference attacks rather than proxy accuracy metrics. We provide a comprehensive empirical analysis, demonstrate the shortcomings of current evaluations, and release Lethe as a reproducible, extensible benchmark.


Likelihood-free inference of phylogenetic tree posterior distributions

Luc Blassel ⋅ Noémie Sauvage ⋅ Pierre Barrat-Charlaix ⋅ Bastien Boussau ⋅ Nicolas Lartillot ⋅ Laurent Jacob

Phylogenetic inference, the task of reconstructing how related sequences evolved from common ancestors, is a central objective in evolutionary genomics. The current state-of-the-art methods exploit probabilistic models of sequence evolution along phylogenetic trees, by searching for the tree maximizing the likelihood of observed sequences, or by estimating the posterior of the tree given the sequences in a Bayesian framework. Both approaches typically require to compute likelihoods, which is only feasible under simplifying assumptions such as independence of the evolution at the different positions of the sequence, and even then remains a costly operation. Here we present the first likelihood-free inference method for posterior distributions over phylogenies. It exploits a novel expressive encoding for pairs of sequences, and a parameterized probability distribution factorized over a succession of subtree merges. The resulting network provides well-calibrated estimates of the posterior distribution leading to more accurate tree topologies than existing methods, even under models amenable to likelihood computation. We further show that its edge against likelihood-based methods dramatically increases under models of sequence evolution with intractable likelihoods.

Many real-world questions demand not a single fact but an organized collection of them, yet existing factual knowledge benchmarks almost exclusively target single-answer retrieval. We introduce ListQA, a benchmark of 9,045 human-curated, cross-validated questions (estimated error rate below 2%) that require LLMs to recall multiple facts and compose them into structured lists of 3--10 elements. Questions range from flat lists Which NCAA teams went undefeated between 2000 and 2024? to hierarchical lists with sub-attributes ... and what was their record and result?, spanning eight categories across 60+ countries. An evergreen design (explicit temporal constraints or historically immutable facts) keeps ground truth valid without periodic updates. For evaluation, we frame element matching as a linear sum assignment problem: optimal bipartite matching paired with LLM-based semantic grading jointly captures factual accuracy, hallucination, and format compliance. We benchmark 34 models (8 families, 42 configurations across standard and thinking-enabled inference) and find: the best standard-inference model scores below 37% LLM-Judge-F1, multi-faceted questions are harder across the board, and providing the source Wikipedia page lifts scores dramatically, confirming that models can organize facts but struggle to retrieve them from parameters alone. Extended analyses cover base vs. instruction-tuned checkpoints, oracle RAG, and token efficiency. We release the dataset and evaluation code.


LLM-Enhanced Random Forests in Orthogonal Hyperbolic Subspaces for Tabular Learning

Shengpeng Wang ⋅ Yisen Gao ⋅ Han Xiang ⋅ lingyun liu ⋅ Ziwei Zhang ⋅ Qingyun Sun ⋅ Xianxian LI ⋅ Xingcheng Fu

Tabular data is widely applied in critical domains, where tree-based models remain dominant, yet their performance is often constrained by manual heuristic splitting criteria and limited ensemble diversity. While hyperbolic geometry naturally fits the hierarchical topology of tabular data, existing hyperbolic algorithms are hindered by complex Riemannian optimization and the difficulty of defining subtree priors in non-Euclidean spaces. To address these challenges, we propose HOT, a hyperbolic random forest framework based on data decoupling. First, we leverage intrinsic geometric structures and LLM reasoning to guide subtree construction, replacing traditional manual heuristics with geometric-semantic priors. Second, we introduce an orthogonal hyperbolic subspace projection mechanism via tangent space isometry, effectively decoupling feature dependencies and maximizing ensemble diversity. Finally, we establish an end-to-end collaborative training paradigm where gradient residuals from a graph network directly guide the growth of new hyperbolic trees, bypassing expensive iterative optimization. Experiments on 13 benchmark datasets demonstrate that HOT significantly outperforms state-of-the-art baselines in both computational efficiency and predictive performance.


LoopRPT: Reinforcement Pre-Training for Looped Language Models

Guo Tang ⋅ Shixin Jiang ⋅ Heng Chang ⋅ Zihan Zhang ⋅ Nuo Chen ⋅ Yuhan Li ⋅ HuiMing Fan ⋅ Jia Li ⋅ Ming Liu ⋅ Bing Qin

Looped language models (LoopLMs) perform iterative latent computation to refine internal representations, offering a promising alternative to explicit chain-of-thought (CoT) reasoning. However, existing reinforcement learning (RL) paradigms primarily target output tokens, creating a structural mismatch with looped architectures whose reasoning unfolds implicitly. In this work, we propose LoopRPT, a reinforcement pre-training framework tailored for LoopLMs. By reframing next-token prediction as a next-token reasoning task, LoopRPT assigns reinforcement signals directly to latent steps using an EMA teacher reference and noisy latent rollouts. This formulation enables RL to directly shape intermediate representations, compressing effective reasoning into fewer iterations. We instantiate LoopRPT on the Ouro architecture across multiple model scales. Results demonstrate that LoopRPT consistently improves per-step representation quality, achieving Pareto dominance in accuracy–computation trade-offs. Notably, significant gains on hard tokens indicate that LoopRPT enhances early-stage reasoning rather than merely encouraging premature exits. Our findings highlight reinforcement pre-training as a principled paradigm for learning efficient latent reasoning in LoopLMs.


Lost or Hidden? A Concept-Level Forgetting in Supervised Continual Learning

Katarzyna Filus ⋅ Kamil Faber ⋅ Roberto Corizzo ⋅ Christopher Kanan

Continual learning studies how models can adapt to new tasks while retaining previously acquired knowledge. Although a broad spectrum of methods has been proposed to mitigate catastrophic forgetting, the field remains predominantly performance-driven, with limited insight into what forgetting actually corresponds to within the vision model's representation space. Prior work has primarily analyzed forgetting through task-level performance or coarse measures of representational drift, without disentangling output-level accessibility from changes in finer-grained internal structure.To this end, we propose a diagnostic framework that leverages Sparse Autoencoders (SAEs) to define a task-anchored latent feature space, enabling analysis of how task-specific information evolves at a finer granularity, where individual SAE latents are treated as concept proxies for recurring and relatively disentangled visual patterns in the model’s internal computations. Within this framework, we decompose forgetting into apparent concept deletion, recoverability, and decodability. We show that a large portion of seemingly lost concept-level information can often be recovered under linearity assumption, with concept decodability degrading as more tasks are introduced. Overall, our findings suggest that a significant part of concept-level forgetting can be attributed to changes in the representational accessibility rather than complete information erasure.

Low-budget molecular optimization is often framed as Bayesian optimization in a learned latent space, but standard benchmarks typically allow hundreds or thousands of evaluations, whereas a realistic wet-lab round may allow only a few dozen. We revisit this regime using a SELFIES molecular variational autoencoder and GuacaMol objectives at $N=30$ oracle calls. The main finding is not a new surrogate model, but a geometric mismatch in candidate generation. The training latents of the VAE concentrate near a narrow spherical shell, while many high-dimensional optimizers propose candidates from boxes or box-shaped trust regions. Such candidates can leave the decoder's training distribution even when their coordinates look individually plausible. We study a simple geometric constraint, the on-sphere scaffold, that restricts candidate generation to the empirical shell occupied by the VAE training latents. Uniform random search on this shell already outperforms most of the classical box-supported Bayesian optimization baselines. We then introduce two minimal optimizers to probe what remains useful once the support is fixed. GeoWalk performs geodesic local search on the empirical sphere and outperforms every published Bayesian optimization baseline tested. GeoCMA keeps the scaffold but adds covariance learning in a random subspace, becoming preferable when the search representation contains learnable directional or semantic structure, as in LassoBench, NAS-Bench-101 graph-VAE optimization, and GSM8K text-latent prompt optimization. A controlled ablation of three recent Gaussian-process recipes, changing only whether candidates are proposed on the empirical sphere or on the published box-supported domain, shows that the scaffold rather than the surrogate is the load-bearing component. The practical lesson is simple: first match the empirical search support; add a surrogate or covariance learner only when the remaining search space contains directional structure that can be learned from the available budget.


Measuring Black-Box Confidence via Reasoning Trajectories: Geometry, Coverage, and Verbalization

Marc Boubnovski Martell ⋅ Josefa Stoisser ⋅ Kaspar Märtens ⋅ Jialin Yu ⋅ Robert Kitchen ⋅ Philip Torr ⋅ Jesper Ferkinghoff-Borg

Reliable confidence estimation gates safe deployment of chain-of-thought (CoT) reasoning through text-only APIs, yet the dominant black-box baseline, self-consistency over K samples, is linearly expensive and ignores the geometry of the trace. We introduce a black-box trajectory-confidence score that embeds a CoT as a sliding-window trajectory and measures its convergence toward external answer anchors with a one-parameter softmax, requiring no logits, hidden states, or supervised calibrators. On six (benchmark,reasoner) settings over MedQA-USMLE, GPQA Diamond, and MMLU-Pro × {Gemini 3.1 Pro, Claude Sonnet 4.6}, fusing this score with coverage and verbalized-confidence channels at K=4 Pareto-improves self-consistency at K=8 in 6/6 settings (median AUC 0.78 vs. 0.71, ΔAUC=+0.075); a fixed-pick control (+0.060) and an E5 cross-embedder replication rule out answer-switching and single-vendor artifacts. Mechanistically, the geometry signal peaks in the penultimate reasoning window across all benchmarks and reasoners and inverts at the terminal window on GPQA Diamond, exposing answer commitment before literal verbalization. Three increasingly unscaffolded regimes decompose black-box confidence into a judge-mediated Coverage prior (C), within-trace Geometry (G), and a conditional Verbalization channel (V); across 18 benchmark × reasoner × proposer settings, C and G carry independent signal in 18/18 and 16/18, while V contributes residual signal in only 6/18. A judge-family swap (GPT-5-mini → Claude Sonnet 4.6) leaves G-only AUC unchanged (∣Δ∣≤0.013) and shifts C-only AUC by at most ±0.02 (κ=0.82), and fusion beats the best single channel in 17/18 settings (median AUC 0.78, max 0.92). Together, these results show that black-box CoT confidence can be read more reliably from a trace's geometric convergence in embedding space than from sample-vote agreement, at lower sampling cost and without access to logits or hidden states.

Existing VLM-based open-set semi-supervised learning (OSSL) methods primarily rely on coarse-grained class-level semantics, thereby hampering the accurate identification of in-distribution (ID) versus out-of-distribution (OOD) samples. To overcome this issue, we propose a novel memory-driven contrastive embedding enhancement approach by integrating fine-grained visual contrastive cues into class-level textual side. Concretely, we design a memory bank to construct and maintain a diverse collection of the most distinctive fine-grained visual embeddings. Guided by the memory bank, visual information is incorporated into the textual side through cross-modal contrastive learning, enabling more discriminative fine-grained separation between ID and OOD samples. Consequently, the learned OSSL model exhibits improved contrastive discriminability across known and unknown categories. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) performance on multiple fine-grained datasets. The code is available at https://anonymous.4open.science/r/CodeForPaper.


MemReg: Streaming Outdoor LiDAR Point Cloud Registration with Hybrid Memory Buffers

Jiuming Liu ⋅ Mengmeng Liu ⋅ Guangming Wang ⋅ Lihao Liu ⋅ Michael Yang ⋅ Yunpeng Zhang ⋅ Francesco Nex ⋅ Hao Cheng ⋅ Hesheng Wang

Prior outdoor LiDAR registration works are commonly limited by the pair-wise input paradigm, which neglects temporal correlations inherent within streaming sequences. In this paper, we propose a novel multi-frame outdoor point cloud registration network for streaming LiDAR scans enhanced with pose-related memory buffers. The key observation is that long-term temporal LiDAR sequences can provide rich global contextual information to complete sparse measurements, filter outliers, and address low-overlap problems, thereby boosting registration performance. Specifically, two hybrid memory buffers are designed, including an implicit memory feature buffer and an explicit memory pose buffer, to store and dynamically update pose-related temporal features. Moreover, a novel dynamic history weighting module is developed to adaptively fuse current and history pose-related features. Extensive experiments on three outdoor datasets, including KITTI, nuScenes, and Apollo-Southbay, demonstrate state-of-the-art performance of MemReg, surpassing all previous pair-wise methods and also multi-frame SLAM systems. Our method also generalizes surprisingly well to multiview indoor registration scenarios with rather competitive performance on 3DMatch, 3DLoMatch, and ScanNet. Code will be released upon publication.

We introduce mesh-invariant adaptive Markov chain Monte Carlo methods for Gaussian-process posteriors arising in infinite-dimensional Bayesian inference. In function-space MCMC, posterior distributions are defined by a change of measure with respect to a Gaussian prior, making absolute continuity essential for valid proposal construction. Standard adaptive schemes that modify means or scales in the full discretized space can destroy this property, leading to proposal measures that become singular in the infinite-dimensional limit. To avoid this, we adapt only an active finite-dimensional subspace of the Gaussian-process representation while preserving the prior dynamics on inactive coordinates. This yields two adaptive proposals, pCNLV and pCNMV, which extend preconditioned Crank--Nicolson and Crank--Nicolson Langevin methods by learning posterior scale, and in pCNMV also posterior mean structure, on the data-informed subspace without introducing discretization-dependent Gaussian density ratios. The resulting samplers retain the mesh robustness of function-space methods while improving efficiency through local adaptation. Experiments on a Darcy-flow inverse problem and Bayesian logistic regression demonstrate consistent efficiency gains, including an approximately fourfold improvement in effective sampling efficiency for Darcy flow.


Meta-Learning Preferences for Multilingual LLM Alignment

Jiaying Lin ⋅ Seongho Son ⋅ Nam P Tran ⋅ Long Tran-Thanh ⋅ Ilija Bogunovic ⋅ Debmalya Mandal

Unequal availability of human preference data across languages poses a significant challenge for aligning large language models in multilingual settings. To address the lack of sufficient data in low-resource language alignment, we propose a meta-learning framework for Reinforcement Learning from Human Feedback and Direct Preference Optimization. By leveraging preference data from other languages, our framework learns a transferable initialization that enables effective adaptation to a target language with minimal data. We provide theoretical guarantees for both the meta-reward modeling and meta-policy optimization settings, and empirically demonstrate the effectiveness of our approach on multilingual benchmarks. In an extremely low-resource setting with only 100 target-language preference samples, our approach achieves up to $28\%$ win-rate improvements over baseline methods, and consistently outperforms baselines across multiple target languages and model scales. Our approaches retain these advantages across different combinations of meta-training languages and varying linguistic distances from the target languages.


Midpoint Generative Models

Daniil Shlenskii ⋅ Nikita Gushchin ⋅ Lev Novitskiy ⋅ Dmitry V. Dylov ⋅ Aleksandr Korotin

We introduce *Midpoint Generative Models* (MGM), a principled framework for training one-step generative models. MGM is based on a simple symmetry of Flow Matching with linear interpolation: when the two endpoint distributions coincide, the corresponding drift field vanishes at the midpoint time, $t=1/2$. We show that the norm of this field defines a valid discrepancy between distributions, which we call the *Midpoint Divergence*. We extend this discrepancy beyond the midpoint by introducing randomly flipped interpolations and further generalize it by replacing deterministic linear Flow Matching interpolations with symmetric stochastic interpolants, yielding a generalized Midpoint Divergence. Finally, we derive a variational formulation of our generalized divergence, yielding a tractable objective for training a one-step generator. The resulting MGM algorithm offers an effective and theoretically grounded approach to generative modeling, achieving competitive performance against existing one-step generative modeling methods.


Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification

Alex Buna-Marginean ⋅ Xiaoqi Liu ⋅ Patrick Rebeschini

In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a max-margin interpolating classifier, whose implicit bias can be statistically suboptimal. In this work, we show that early stopping can overcome this suboptimality: in a Gaussian mixture model with label-flipping noise, GD stopped at an appropriate oracle time achieves minimax-optimal excess zero-one risk for covariance spectra with fast and continuous decay, including polynomial and exponential spectral decays. Our analysis combines a sharp upper bound for the early-stopped iterate with a matching statistical lower bound over arbitrary classifiers, yielding optimal rates that are validated by experiments. A central technical contribution is a new calibration result that converts excess logistic risk into excess zero-one risk; it handles the model misspecification induced by the label-flipping noise, and removes the square-root rate in standard bounds. We also establish a lower bound for linear interpolators, showing that interpolation can require exponentially more samples than early stopping to achieve the same excess risk.


Mixture of Attribute-Aware Attention Experts for Fine-grained E-Commerce Composed Image Retrieval

Yufei Ma ⋅ Zihan Liang ⋅ Zhipeng Qian ⋅ Huangyu Dai ⋅ Lingtao Mao ⋅ Ben Chen ⋅ Chenyi Lei ⋅ Wenwu Ou

Composed Image Retrieval (CIR) enables users to express specific search intent by combining a reference image with manipulation text. Despite its promise for e-commerce, existing methods ignore readily available item textual descriptions, suffer from multi-attribute confounding where samples exhibit uncontrolled variations in unrelated attributes, and lack explicit fine-grained attribute modeling. We propose $\textbf{MoA}^{3}\textbf{CIR}$, featuring an asymmetric five-tower architecture that incorporates item descriptions on both query and target sides, a self-reflective data augmentation pipeline generating controlled attribute modifications through LVLM-driven refinement, and a Mixture of Attribute-Aware Attention mechanism with grouped multi-head attention and dynamic routing coupled with a two-phase training strategy. Besides, we contribute ECCIR, the first dual-modal e-commerce CIR benchmark with 110K samples featuring authentic item descriptions. Extensive experiments demonstrate that \method~achieves SoTA performance on ECCIR, FashionIQ, and Shoes datasets, with substantial improvements over existing methods. Code and datasets will be made publicly available.


Momentum Smooths the Path to Gradient Equilibrium

João Vitor Romano ⋅ Margaux Zaffran ⋅ Ryan Tibshirani

Online gradient descent has recently been shown to satisfy gradient equilibrium for a broad class of loss functions. This means that the average of the gradients of the losses along the sequence of estimates converges to zero, a property that allows for quantile calibration and debiasing of predictions, among other useful statistical properties. A shortcoming of online gradient descent when optimized for gradient equilibrium is that the sequence of estimates is jagged, leading to volatile paths. In this work, we propose the generalized momentum method, defined as a weighting of past gradients, as a broader algorithmic framework with guarantees to smoothly postprocess (e.g., calibrate or debias) predictions from black-box algorithms, yielding estimates that are more meaningful in practice. We prove it achieves gradient equilibrium at the same convergence rates and under assumptions similar to those of plain online gradient descent, all the while producing smoother paths that preserve the original signal amplitude. Of particular importance are the consequences for sequential decision-making, where more stable paths translate to less variability in statistical applications. These theoretical insights are corroborated by experiments on real data, showcasing the benefits of adding momentum.


MONET: A Massive, Open, Non-redundant and Enriched Text-to-image dataset

Benjamin Aubin ⋅ Gonzalo Iñaki Quintana ⋅ Onur Tasar ⋅ Sanjeev Sreetharan ⋅ Urszula Czerwinska ⋅ Damien Henry ⋅ Clément Chadebec

Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of collecting, filtering, deduplicating, and re-captioning such corpora at scale hinders open and reproducible research in the field. We introduce MONET, an open \emph{Apache 2.0} dataset of ${\sim}104.9$M image--text pairs collected from 2.9B raw pairs across heterogeneous open sources through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with multiple vision-language models covering short to long-form descriptions, and further augmented with synthetically generated samples. Each image is shipped with pre-computed embeddings and annotations to accelerate downstream use. To validate the effectiveness of MONET, we train a 4B-parameter latent diffusion model *exclusively* on it and reach competitive GenEval and DPG scores, demonstrating that our dataset lowers the barrier to large-scale, reproducible text-to-image research.


Mosaic: A Benchmark Suite for Differentiable Physics Solvers

Andrin Rehmann ⋅ Heiko Zimmermann ⋅ Dion Häfner

Differentiable physics solvers underpin solver-in-the-loop ML training, gradient-based optimal control, and inverse problems, yet the practical cost of obtaining correct, usable gradients from a given solver on a given problem is largely undocumented. Integration effort, computational cost, gradient accuracy, and numerical conditioning vary widely across solvers and are discoverable only by trial and error. We introduce Mosaic, an extensible benchmarking framework for differentiable PDE solvers that standardizes access to solver gradients. Each solver is packaged as a containerized component (Tesseract) exposing a uniform gradient API regardless of language or automatic differentiation (AD) strategy, enabling researchers to evaluate, compare, and build on non-trivial physical solvers. Our evaluation of 14 solvers across fluid dynamics, structural mechanics, and heat transfer demonstrates that the benchmark surfaces practically relevant differences: order-of-magnitude variation in computational cost and Jacobian conditioning, alongside structural incompatibilities that eliminate solvers from realistic tasks entirely. Despite this variation, all solvers that produce gradients converge to similar optima, indicating that the practical barriers are memory limits, numerical stability, and setup compatibility rather than gradient accuracy alone. Mosaic is open-source and available at https://anonymous.4open.science/r/mosaic-bench.

Gradient normalization stabilizes deep-learning optimization, and spectral normalizations are especially natural for matrix-shaped parameter blocks; Muon is the motivating example. We study an idealized deterministic, continuous-time, vanishing-momentum version of this idea in the mean-field regime, where wide models are represented by probability measures on parameter space. Starting from normalized matrix flows, we introduce Spectral Wasserstein distances indexed by norms $\gamma$ on positive semidefinite matrices: the trace norm gives classical $\Wtwo$, the operator norm gives the Muon geometry, and Schatten norms interpolate between them. We develop the static Kantorovich formulation, a max-min robust-cost representation, Gaussian reductions extending the Bures formula, and for monotone norms, prove equivalence with a Benamou--Brenier formulation. This yields a gradient-flow interpretation of the mean-field normalized training dynamics. We illustrate these findings by numerical experiments on MMD flows, Gaussian reductions, two-layer ReLU models, and shallow attention.

Contemporary large language model (LLM) chat systems treat conversation history as an immutable sequence of turns that defines the model’s working context. However, user intent in real interactions is not static: it evolves through correction, refinement, and shifting constraints. This mismatch between dynamic intent and static transcripts can result in context pollution, where outdated or irrelevant information persists and continues to influence subsequent responses. We introduce mutable transcripts, a new interaction paradigm that enables users to revise prior turns through natural language edit requests, allowing the conversation history itself to be updated rather than appended. This reframes the transcript from a passive record into an editable representation of conversational state. We present a working prototype that integrates transcript-level revision into a standard chat interface and evaluate its feasibility through a controlled user study (n=17) and an illustrative transcript analysis of representative interaction scenarios. Participants significantly preferred mutable transcripts over standard chat across measures of clarity, confidence, and ease of use, with reduced intent to restart conversations. Transcript analysis of representative user study conversations shows that mutable transcripts can reduce conversation length and eliminate obsolete retained context. These findings provide initial evidence that user-driven revision of conversational history can improve interaction quality and help maintain a more current representation of user intent.

We study decentralized stochastic smooth convex optimization, where $M$ workers minimize an average objective using local stochastic gradients and neighbor-only communication over a fixed gossip network. A central question in this setting is to determine the largest number of workers that can be used under a total budget of $N$ gradient samples while still preserving the centralized $O(1/\sqrt N)$ statistical rate. We introduce an accelerated decentralized method that preserves this rate for $M \lesssim \sqrt{\rho} N^{3/4}$ workers up to logarithmic factors, where $\rho$ is the spectral gap of the gossip network, improving the best prior scaling of $M \lesssim \rho\sqrt N$ [Eisen et al., 2025]. The method is based on a one-step-delayed stochastic acceleration scheme that enables workers to interleave minibatching with accelerated gossip while controlling residual disagreement, and its guarantee depends only logarithmically on the optimum-local heterogeneity. We also establish a matching lower bound for linear-span decentralized first-order methods, showing that the method is optimal up to logarithmic factors.


Neural Fourier Transform for Multiple Time Series Prediction

Noam Koren ⋅ Daniel Freedman ⋅ Kira Radinsky

Multivariate time series forecasting is an important task in various fields such as economic planning, healthcare management, and environmental monitoring. In this work, we present a novel methodology for improving multivariate forecasting, particularly, in data sets with strong seasonality. We frame the forecasting task as a Multi-Dimensional Fourier Transform (MFT) problem and propose the Neural Fourier Transform (NFT) that leverages a deep learning model to predict future time series values by learning the MFT coefficients. The efficacy of NFT is empirically validated on 7 diverse datasets, demonstrating improvements over multiple forecasting horizons and lookbacks, thereby establishing new state-of-the-art results. Our contributions advance the field of multivariate time series forecasting by providing a model that excels in predictive accuracy.


NeuroNTP - A Generalizable Multimodal Foundation Model for Epilepsy

József Kovács ⋅ Amadeus Hauser ⋅ Gudrun Gröppel ⋅ Wolfgang Narzt ⋅ Philipp Seidl

Epilepsy affects approximately $1$% of the global population, where seizure detection and reliable forecasting are critical for patient safety and quality of life. While emerging foundation models increasingly recast EEG analysis as a language modeling task, most remain confined to single-modality inputs and narrow domain generalization. In this work, we present NeuroNTP, a modular multimodal foundation model built on an xLSTM backbone and pretrained on one of the largest retrospective epilepsy datasets ever collected, comprising $150,000$ hours of neural (EEG) and physiological (ECG, SpO$_2$) signals paired with rich clinical text. NeuroNTP employs a flexible, montage-geometry grounded tokenizer and optimizes a composite objective, synergizing causal next-token prediction with seizure-specific auxiliary tasks to enable unified sequence modeling across heterogeneous modalities. This approach allows the model to capture long-range temporal dependencies and cross-modal interactions that unimodal baselines ignore. In extensive experiments, we show that NeuroNTP outperforms supervised baselines by a substantial margin on held-out clinical cohorts. Furthermore, the model generalizes to public EEG benchmarks such as CHB-MIT and Siena Scalp EEG, achieving state-of-the-art performance under standard evaluation settings. We release code and model weights to facilitate further research.


Neurosymbolic Learning for Inference-Time Argumentation

Gabriel Freedman ⋅ Adam Dejl ⋅ Adam Gould ⋅ Mansi - ⋅ Lihu Chen ⋅ Junqi Jiang ⋅ Francesca Toni

Claim verification is an important problem in high-stakes settings, including health and finance. When information underpinning claims is incomplete or conflicting, uncertain answers may be more appropriate than binary true or false classifications. In all cases, faithful explanations of the considerations determining the final verdict are crucial. We introduce inference-time argumentation (ITA), a trainable neurosymbolic framework for ternary claim verification in which a formal argumentation semantics giving the strength of claims is used both (i) to guide LLM training as models learn to generate arguments and assign them base scores (representing intrinsic strengths) and (ii) to compute ternary (true/false/uncertain) predictions from generated, scored arguments. As a result, at training time, argument generation and scoring can be optimised according to the quality of the induced argumentative predictions. Moreover, at inference time, the final prediction is faithful, by construction, to the arguments and scores determining the verdict, rather than being justified by a potentially unfaithful post-hoc reasoning trace as in conventional reasoning models. We finally show that, on two datasets for ternary claim verification, ITA improves upon argumentative baselines and can perform competitively against non-argumentative direct-prediction baselines, while providing verdicts that are computed deterministically from explicit, inspectable argumentative structures.

Learning effective representations from battery operational history is essential for predictive health management and control-oriented applications. Existing approaches rely on response-dependent engineered features or short early-life windows, limiting generalisability and precluding use in prospective settings where future responses are unavailable. We propose $\texttt{NOCE-Net}$, a deep nested sequence model that learns control-enabling representations directly from raw cyclic operational data. Our data pipeline, $\texttt{re-BatteryData}$, standardises heterogeneous aging datasets into semantically equivalent Full Equivalent Cycles (FECs) defined by absolute charge throughput, enabling end-to-end learning across diverse operational regimes. $\texttt{NOCE-Net}$ combines multi-patch transformer encoders for intra-cycle dynamics with a GRU backbone for inter-cycle state evolution, operating on approximately $700$ batteries across ten public datasets spanning over $300$ operational profiles. In zero-shot evaluation on unseen battery chemistries and usage patterns, $\texttt{NOCE-Net}$ achieves $13.7$% MAPE in voltage response prediction, outperforming strong long-sequence baselines by over $43$% while using $40 \times$ fewer parameters. Shape and component diagnostics of the learned hidden states reveal that the GRU's principal direction of evolution achieves Spearman $|\rho| \geq 0.95$ with the capacity fade trajectory on average across all datasets with monotone degradation profiles, including zero-shot test batteries, providing evidence that the representations encode degradation dynamics transferable to downstream tasks beyond battery response modelling.


Normative Networks for Source Separation via Local Plasticity and Dendritic Computation

Bariscan Bozkurt ⋅ Efe Ali Gorguner ⋅ Francesco Innocenti ⋅ Rafal Bogacz

Blind source separation (BSS) is a natural framework for studying how latent causes may be recovered from sensory mixtures, but deriving online and biologically plausible algorithms for structured (i.e., constrained to known domains) and potentially correlated sources remains challenging. Recent work has derived neural networks for BSS from maximization of an entropy measure, yet its online implementations involve complex and nonlocal recurrent dynamics. Motivated by this perspective, we propose Predictive Entropy Maximization, which achieves competitive performance in BSS, using only local weight updates. The method employs a close approximation of an entropy measure, yielding an objective function with easily interpretable components. Minimizing this objective leads to a predictive neural architecture in which feedforward synapses follow an error-driven rule (that can be realized through dendritic mechanisms), lateral inhibitory connections are learned with local Hebbian plasticity, and source-domain constraints are enforced through simple output nonlinearities. We derive explicit spectral bounds on the surrogate error, characterizing when the approximation is accurate. Empirically, Predictive Entropy Maximization remains robust under increasing source correlation and observation noise, outperforms biologically plausible algorithms that rely on stronger independence or decorrelation assumptions, and remains competitive with exact determinant- and correlative-information-based baselines. These results show how local plasticity and adaptive lateral inhibition can emerge from maximizing a regularized second-order entropy over structured source domains.


One View Is Enough: In-the-Wild Monocular Pretraining for Novel View Generation

Adrien RAMANANA RAHARY ⋅ Nicolas Dufour ⋅ Patrick Perez ⋅ David Picard

Monocular novel-view synthesis has long required multi-view image pairs for supervision, limiting training to a narrow set of purpose-built datasets. We propose in-the-wild monocular pretraining: a frozen depth estimator lifts each source image into 3D and reprojects under sampled poses to yield pseudo-target views; masked losses restrict supervision to valid regions and an adversarial objective covers disoccluded areas. Scaled to 30 million uncurated images, this produces OVIE, requiring only a source image and target pose at inference. Prior work trains without multi-view data but needs a depth estimator at inference, or drops this dependency but requires multi-view training pairs; OVIE is the first to require neither. Without multi-view supervision, OVIE rivals in-domain baselines on RealEstate10K and surpasses all on DL3DV; brief multi-view fine-tuning outperforms all geometry-free methods on their training domain. At 116 FPS, it is over 600× faster than the fastest baseline. Code and models will be open-sourced.


On the Bias of Group-Based Advantage Estimation

Fengkai Yang ⋅ Zherui Chen ⋅ Xiaohan Wang ⋅ Xiaodong Lu ⋅ Jiajun Chai ⋅ Guojun Yin ⋅ Wei Lin ⋅ Fuzhen Zhuang ⋅ Shuai Ma ⋅ deqing wang ⋅ Yaodong Yang ⋅ Yikun Ban

Reinforcement Learning from Verifier Rewards (RLVR) has emerged as a widely used approach for post-training large language models on reasoning tasks, with group-based methods such as GRPO and its variants gaining broad adoption. These methods rely on group-relative advantage estimation to avoid learned critics, yet its theoretical properties remain poorly understood. In this work, we uncover a fundamental issue of group-based RL: the group-relative advantage estimator is inherently biased relative to the true (expected) advantage. We provide the first theoretical analysis showing that it systematically underestimates advantages for hard prompts and overestimates them for easy prompts, leading to imbalanced exploration and exploitation. To address this issue, we propose History-Aware Adaptive Difficulty Weighting (HA-DW), an adaptive reweighting scheme that adjusts advantage estimates based on an evolving difficulty anchor and training dynamics. Both theoretical analysis and experiments on five mathematical reasoning benchmarks demonstrate that HA-DW consistently improves performance when integrated into GRPO and its variants. Our results suggest that correcting biased advantage estimation is critical for robust and efficient RLVR training. The code is available at https://anonymous.4open.science/r/HA-DW-037C.


OpenMedReason: Scientific Reasoning Supervision for Medical Vision–Language Models

Negin Baghbanzadeh ⋅ Pritam Sarkar ⋅ Michael Colacci ⋅ Abeer Badawi ⋅ Adibvafa Fallahpour ⋅ Arash Afkanpour ⋅ Leonid Sigal ⋅ Ali Etemad ⋅ Elham Dolatabadi

High-stakes clinical use of large vision–language models (LVLMs) requires reasoning that is grounded in visual evidence and clinical knowledge, not just correct final answers. We introduce OpenMedReason, a large-scale, open multimodal medical reasoning corpus comprising approximately 450K image–question–answer instances whose reasoning traces are primarily derived from curated biomedical, human-authored scientific articles. OpenMedReason provides high-fidelity supervision beyond synthetic chains of thought, covering diverse medical domain vision modalities such as radiological scans, microscopic images, visible light photographs, charts, and others. We complement it with OpenMedReason Bench, a held-out benchmark that allows fine-grained evaluation of LVLMs along three complementary axes of capability, including perception, medical knowledge, and rationale, enabling diagnostic evaluation beyond final answer accuracy. OpenMedReason is a rich training resource that exhibits its effectiveness in both supervised fine-tuning (SFT) and reinforcement-based alignment. Training with OpenMedReason yields a 20\% average improvement in VQA accuracy over the base model and achieves 4.2\% the strongest performance among comparable-scale medical LVLMs. Fine-grained performance analysis confirms that the gains are not concentrated in any single axis: OpenMedReason improves perception, medical knowledge, and rationale jointly, and its reasoning traces are preferred over those of the base model in 86.1\% of pairwise comparisons. We release the dataset and code.

Score-based diffusion models have achieved remarkable empirical success in generating high-quality samples from target data distributions. Among them, the Denoising Diffusion Probabilistic Model (DDPM) is one of the most widely used samplers, generating samples via estimated score functions. Despite its empirical success, a tight theoretical understanding of DDPM --- especially its convergence properties --- remains limited. In this paper, we provide a refined convergence analysis for the DDPM sampler and establish near-optimal convergence rates under general distributional assumptions. Specifically, we consider a relaxed smoothness condition parameterized by $L$, which is small for many practical distributions (e.g., Gaussian mixture models). We prove that, to approximate a target distribution on $\mathbb{R}^d$ to accuracy $\varepsilon$ in total variation distance and $\varepsilon^2$ in KL-divergence, the DDPM sampler with accurate score estimates requires at most $$T = \widetilde{O}\left(\frac{\sqrt{d}\min\{\sqrt{d},L\}}{\varepsilon}\right)$$ iterations, where $\widetilde{O}$ hides polylogarithmic factors in $d$ and $1/\varepsilon$. This result substantially improves upon the best-known $\widetilde{O}(d/\varepsilon)$ iteration complexity when $L < \sqrt{d}$. By establishing a matching lower bound, we show that our convergence analysis is tight for a wide array of target distributions. Moreover, it reveals that DDPM and Denoising Diffusion Implicit Models (DDIM) share the same dependence on $d$, raising an interesting question of why DDIM often appears empirically faster.


Optimal Rates for Adaptive Private $k$-PCA

Johanna Düngler ⋅ Amartya Sanyal

Given $n$ i.i.d. random matrices $A_i \in \mathbb{R}^{d \times d}$ with common expectation $\Sigma$, the goal of Differentially Private Stochastic PCA is to identify a $k$-dimensional subspace capturing the leading variance directions of $\Sigma$, while preserving differential privacy (DP) for each individual sample $A_i$. Düngler & Sanyal [2025] introduced $k$-DP-PCA, the first algorithm to simultaneously (1) achieve sample complexity $n=\widetilde O(d)$ for sub-Gaussian data, (2) adapt its privacy noise to the intrinsic randomness of the data, and (3) extend seamlessly to any target dimension $k\le d$. However, its sample complexity has a suboptimal dependence on $k$. We propose the first algorithm that achieves optimal sample complexity in both $d$ and $k$, while retaining (2) and (3) of $k$-DP-PCA. In addition, our method removes the exponential dependence on the eigengap that appears in the sample-size lower bound required by prior utility guarantees, and improves the dependence on spectral parameters to match known lower bounds for the spiked covariance model. Unlike deflation-based approaches like $k$-DP-PCA, our algorithm updates the full $d\times k$ subspace jointly rather than one eigenvector at a time. This non-deflation structure simplifies the algorithm, reduces the number of hyper-parameters, and improves computational efficiency, particularly when using block linear algebra libraries.


Pareto DNN Verification: Fast for Most Queries

Yizhak Y. Elboher ⋅ Avraham Raviv ⋅ Amihay Elboher ⋅ Zhouxing Shi ⋅ Hillel Kugler ⋅ Guy Katz

Formal verification of deep neural networks (DNNs) faces a fundamental scalability challenge: verifying robustness is NP-hard and state-of-the-art tools often time out on modern architectures. We propose PAreto VErification, or PAVE: by augmenting DNNs with early exits (EEs) and verifying exit-by-exit, we solve the vast majority of robustness queries---roughly 80\%---in a small fraction of the total time. The key insight is that most inputs in a robustness neighborhood exit at an early layer, a provably tractable sub-problem, so verification terminates far before analyzing the full network. EEs can be added to any DNN without degrading accuracy (within 0.3\% across all tested architectures). We formalize the novel robustness property for EE networks, prove it is Fixed-Parameter Tractable under a trace stability condition (which commonly holds in practice), and present a sound and complete verification algorithm. Two sound optimizations---$break$ (early termination when the winner dominates) and $continue$ (skipping the inner loop when no runner-up can win)---reduce SAFE-case verification time by up to $10 \times$. Experiments on MNIST, CIFAR-10, and CIFAR-100 with networks from 600K to 33M parameters demonstrate that PAVE enables solving queries that were not solved otherwise, including formal verification of ResNet-18 on CIFAR-100.


Partition Tree: Conditional Density Estimation over General Outcome Spaces

Felipe Lourenco Angelim Vieira ⋅ Alessandro Leite

We propose Partition Tree, a tree-based framework for conditional density estimation over general outcome spaces that supports both continuous and categorical variables within a unified formulation. Our approach models conditional distributions as piecewise-constant densities on data-adaptive partitions and learns trees by directly minimizing conditional negative log-likelihood. This yields a scalable, nonparametric alternative to existing probabilistic trees that does not make parametric assumptions about the target distribution. We further introduce Partition Forest, an ensemble extension obtained by averaging conditional densities. Empirically, we demonstrate improved probabilistic prediction over CART-style trees and competitive performance compared to state-of-the-art probabilistic tree methods and Random Forests.


PDHFormer: Progressive Dual-Head Transformer for Behavioral Choice Prediction

Hao Zhou ⋅ Jing Chen ⋅ Yaoxin Wu ⋅ Jie Gao ⋅ Yingqian Zhang

Many applications require joint prediction of interdependent behavioral choices, yet existing models often treat each choice independently (e.g., through parallel prediction heads), overlooking the influence of one on the other. In this work, we propose Progressive Dual-Head Transformer (PDHFormer), a novel framework that performs two-step prediction: the model first estimates one choice and then conditions the second on this upstream estimate through an explicit head-to-head pathway. A shared encoder captures the common structure of two prediction tasks, while the dual-head module explicitly reflects cross-choice dependence. A gated residual mechanism integrated into the embedding layer and the dual-head module further improves the training stability and the prediction performance. Extensive experiments on real-world urban mobility, manufacturing, and food delivery application domains demonstrate that PDHFormer consistently outperforms state-of-the-art machine learning models, deep tabular models, as well as parallel-head Transformer variants across multiple metrics. Moreover, our ablation study confirms that both the proposed progressive dual-head and gated residual mechanism are key contributors to the observed gains in different prediction tasks.


PENEX: AdaBoost-Inspired Neural Network Regularization

Klaus-Rudolf Kladny ⋅ Bernhard Schölkopf ⋅ Michael Muehlebach

AdaBoost sequentially fits so-called weak learners to minimize an exponential loss, which penalizes misclassified data points more severely than other loss functions like cross-entropy. Paradoxically, AdaBoost generalizes well in practice as the number of weak learners grows. In the present work, we introduce Penalized Exponential Loss (PENEX), a new formulation of the multi-class exponential loss that is theoretically grounded and, in contrast to the existing formulation, amenable to optimization via first-order methods, making it a practical objective for training neural networks. We demonstrate that PENEX effectively increases margins of data points, which can be translated into a generalization bound. Empirically, across computer vision and language tasks, PENEX improves neural network generalization in low-data regimes, matching and in some settings outperforming established regularizers at comparable computational cost. Our results highlight the potential of the exponential loss beyond its application in AdaBoost.


Perception for Action in Latent World Models

Petr Ivashkov ⋅ Randall Balestriero ⋅ Bernhard Schölkopf

Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions. At the same time, latent JEPA-style world models advocate learning compact predictive states from high-dimensional observations to facilitate the prediction of future states, but end-to-end training of these models is nontrivial because representations may collapse if our only goal is to construct a latent state that is easy to predict. We show that a suitable inverse dynamics regularization addresses both issues: it is an effective anti-collapse mechanism that induces action-aligned representations. By forcing latent states to preserve information about the action underlying a transition, it biases the model toward the controllable degrees of freedom of the environment while discarding uncontrollable distractors. This yields stable latent world models trained end-to-end from offline, reward-free trajectories, without frozen encoders, exponential moving averages, or complex latent regularizers. Empirically, the learned latent spaces are compact, interpretable, and enable competitive planning performance across simple 2D and 3D control tasks.


Personalized Safety in Federated Fine-Tuning of Large Language Models

Tianzhe Xiao ⋅ Gaozhuo Liu ⋅ Yichen Li ⋅ Haozhao Wang ⋅ Hanlin Cai ⋅ Ruixuan Li ⋅ Ozgur Akan

Large language models (LLMs) are increasingly adapted inside privacy- and regulation-constrained institutions where local post-training data cannot be centrally pooled. Federated learning therefore becomes a natural mechanism for collaborative LLM adaptation, but it also raises a safety question that is not captured by either classical federated learning or standard centralized alignment: clients may share a model while still requiring different safe behaviors. We formalize this setting as \emph{personalized federated safety}, where safety correctness is conditioned on the target client's local policy rather than treated as a single global refusal rule. To make this problem transferable rather than benchmark-specific, we introduce a compact, interpretable client policy space and a policy-conditioned benchmark construction framework, then instantiate it with five representative client archetypes spanning regulated healthcare, legal compliance, financial safety, youth safety, and enterprise general safety. The resulting benchmark contains 1{,}500 client-conditioned training examples and 750 held-out evaluation examples. We further propose \emph{Safety-Aware Weighted LoRA} (SAW-LoRA), a lightweight add-on for federated LoRA fine-tuning that lets each client selectively absorb incoming global task updates according to estimated task benefit and local-safety risk. Across two task datasets and two backbones, this add-on substantially lowers client-specific ASR for the strongest local safety mechanisms, reaching ASR near $1$--$2\%$ while preserving high benign acceptance. These findings position personalized federated safety as a concrete research problem with a practical update mechanism for federated LLM adaptation.


PE-SHAP: Causally Interpretable Path-Wise Shapley Explanations

Maksym Buleshnyi ⋅ Joshua Loftus ⋅ Sakina Hansen

Shapley value-based explanations are widely used to attribute model predictions, but standard variants capture associative rather than causal relationships. Recent causally informed Shapley methods add a path-wise decomposition, but suffer from two key limitations: they assign zero attribution to chain-mediated paths, and their per-path values do not align with classical mediation estimands. We propose PE-SHAP, a Shapley-inspired path-wise decomposition built on the portion-eliminated effect. Per-mediator values sum to the total mediated effect at every sample, aggregate to the Natural Indirect Effect at the population level, and match the Path-Specific Effect under additive separability. These are properties prior Shapley-based path-wise methods do not provide. We validate PE-SHAP analytically and through Monte-Carlo experiments on synthetic SCMs with known path effects, and demonstrate it on real-world fairness case studies using the German Credit and COMPAS datasets.

Balancing exploration and exploitation is a core challenge in sequential decision-making and black-box optimization. We introduce POETS (Policy Ensembles for Thompson Sampling), a novel framework that bridges uncertainty quantification and policy optimization. Our approach is grounded in the insight that policies trained with Kullback-Leibler (KL) regularization implicitly encode an underlying reward function. Building on this, POETS bypasses the complex, nested process of training an uncertainty-aware reward model and separately fitting a policy to this model. Instead, we directly train a policy ensemble to capture epistemic uncertainty by matching implicitly encoded reward functions to online, bootstrapped data. To overcome the prohibitive compute and memory constraints of ensembling Large Language Models (LLMs), POETS utilizes an efficient architecture: the ensemble shares a pre-trained backbone while maintaining diversity through independent Low-Rank Adaptation (LoRA) branches. Theoretically, we prove that POETS implicitly conducts KL-regularized Thompson sampling and thus inherits strong cumulative regret bounds of $O(\sqrt{T \gamma_T})$. Empirically, we demonstrate that POETS achieves state-of-the-art sample efficiency across diverse scientific discovery domains, including protein search and quantum circuit design. Furthermore, it improves the optimization trajectories of reinforcement learning, proving particularly robust in off-policy settings with experience replay or in small dataset regimes.

Many projection-based imaging modalities reconstruct unknown objects from collections of noisy projections related by unknown in-plane rotations, making effective estimation dependent on rotationally consistent alignment and aggregation. In cryogenic electron microscopy (cryo-EM), the extremely low signal-to-noise ratio (SNR) makes such multi-image integration essential. We introduce the *polar transformer*, a permutation-invariant architecture for *sets* of images built around an $\mathrm{SO}(2)^K$-equivariant *angular attention* mechanism that jointly aligns and aggregates information across images. This architecture leverages a stable polar representation in which we define polar CNN feature extractors and evaluate attention weights over all relative angular shifts efficiently via FFT-based correlations, guaranteeing equivariance to independent in-plane rotations of each image in the set. We evaluate the polar transformer on simulated cryo-EM images using denoising as a quantitative proxy. Across different regimes, polar transformers outperform strong single-image and image-set baselines, achieving a reduction of up to 40% in relative MSE at an SNR of 0.02 and improving downstream ab initio reconstruction resolution by 20%.

Equivariant neural networks incorporate known symmetries of the data, such as permutations, rotations, or translations, directly into their architecture, often improving sample efficiency and generalization. However, strict equivariance can make optimization difficult, for example, when the data only approximately satisfy the assumed symmetry, or when the equivariance constraints induce a challenging loss landscape. Recent studies have shown that relaxing equivariance during training by introducing non-equivariant dense linear layers can ease optimization in such cases and improve the performance of the resulting strictly equivariant model (after removing the relaxation layers at test time). While effective, this approach incurs a significant computational overhead, limiting its applicability in scalable architectures. In this work, we propose a more efficient constraint-relaxation method based on positional encoding (PE). Our framework injects a positional encoding with learnable scale that breaks the model’s original symmetry. This positional signal is gradually annealed to zero, recovering the original equivariant model by the end of training. Empirically, we demonstrate significant improvements for translational and rotational equivariant models with only marginal compute and memory overhead.


Prediction-Intervention Games and Invariant Sets

Linus Kühne ⋅ Felix Schur ⋅ Jonas Peters

We consider the following two-player game: using observational data, the leader chooses a prediction function for a response variable $Y$ from given covariates. The follower then reacts with an intervention on some covariates in the underlying structural causal model to maximize their own objective. The leader knows the intervention targets, but may have limited knowledge of the follower's objective. We call this setup a prediction-intervention game, a special case of a Stackelberg game. Finding an optimal strategy for the leader is generally difficult. To avoid severe performance loss, the leader may base their prediction on the causal parents of $Y$, or more generally on an invariant subset of covariates. We prove that predictors based on the stable blanket, a specific invariant subset, weakly dominate those based on the causal parents in the game for two common classes of follower objectives. We further upper bound the leader's post-intervention risk by a worst-case risk over allowed interventions and strengthen existing distribution generalization results to analyze this bound: we give sufficient conditions under which stable-blanket predictors are worst-case optimal, and show by examples that these conditions cannot in general be dropped. Finally, we discuss practical strategies when the causal graph is unknown and test them on simulated and real-world data.


Preference-Guided Adversarial Policy Optimization for Long-Tail Robust Driving

Tong Nie ⋅ Yihong Tang ⋅ Junlin He ⋅ Yuewen Mei ⋅ Jie Sun ⋅ Lijun Sun ⋅ Wei Ma ⋅ Jian Sun

Deploying autonomous driving systems requires robustness against long-tail scenarios that are rare but safety-critical. While adversarial training offers a promising solution, existing methods typically decouple scenario generation from policy optimization and rely on heuristic surrogates. This leads to objective misalignment and fails to capture the shifting failure modes of evolving policies. This paper presents ADV-0, a closed-loop min-max optimization framework that treats the interaction between driving policy (defender) and adversarial agent (attacker) as a zero-sum Markov game. By aligning the attacker's utility directly with the defender's objective, we reveal the optimal adversary distribution. To make this tractable, we cast dynamic adversary evolution as iterative preference learning, efficiently approximating this optimum and offering an algorithm-agnostic solution to the game. Theoretically, our analysis characterizes the idealized regularized game and derives a lower bound that connects adversarial training performance with real-world long-tail generalization. Experiments indicate that ADV-0 effectively exposes diverse safety-critical failures and greatly enhances the generalizability of both learned policies and motion planners against unseen long-tail risks.

We study \emph{differentially private prediction} introduced by Dwork and Feldman (COLT 2018): an algorithm receives one labeled sample set $S$ and then answers a stream of unlabeled queries while the output transcript remains $(\varepsilon,\delta)$-differentially private with respect to $S$. Standard composition yields a $\sqrt{T}$ dependence for $T$ queries, i.e. $O(VC(\mathcal{C})\cdot\sqrt{T})$ sample complexity. We show that this dependence can be reduced to \emph{polylogarithmic} in $T$ in streaming settings. For an oblivious online adversary and any concept class $\mathcal{C}$, we give a private predictor that answers $T$ queries with $|S|= \tilde{O}(VC(\mathcal{C})^{3.5}\log^{3.5}T)$ labeled examples. For an adaptive online adversary and halfspaces over $\mathbb{R}^d$, we obtain $|S|=\tilde{O}\left(d^{5.5}\log T\right)$.


ProSearch: Benchmarking Multi-Constraint Protocol Retrieval in Experimental Science

Wenliang Liang ⋅ Zhiyuan Ning ⋅ Haolin Chen ⋅ Yuanchun Zhou ⋅ Wenjuan Cui ⋅ Yi Du

In real-world AI for Science scenarios, retrieval systems are increasingly expected to support experimental workflows such as experimental design and method reproduction. These workflows require retrieving protocols that satisfy a set of interdependent constraints, such as specific materials and operating conditions. However, systematic evaluation of retrievers in such multi-constraint scenarios remains lacking. To address this gap, we introduce ProSearch, a scientific retrieval benchmark for multi-constraint experimental protocol retrieval. ProSearch has three key features: (1) realism, as queries are instantiated from real method paragraphs using expert-defined experimental fields, query templates, and filtering rules; (2) complexity, involving 60 distinct fields, such as experimental subjects, materials, and operating conditions, together with dependencies among these fields; and (3) diagnosticity, as each query contains 3 to 5 explicit hard constraints and covers challenging cases such as numeric constraints and negation constraints. We construct ProSearch through LLM-based extraction, data-driven constraint selection, controlled query generation, and evidence-based relevance validation. The final benchmark contains 1,551 high-quality queries across six experimental science domains. We use ProSearch to systematically evaluate 19 representative retrievers, spanning lexical matching, general dense retrieval, scientific retrieval, and reasoning-oriented retrieval. The results reveal that current retrievers generally struggle with multi-constraint protocol retrieval, with the best-performing model achieving an average nDCG@10 of 0.290. Further constraint-wise and error analyses reveal key failure modes of existing retrievers.


Provable State Estimation with Recurrent Models

Vincent Andrieu ⋅ Pauline Bernard ⋅ Lucas Brivadis ⋅ Laurent Praly ⋅ Christian Wolf

Estimating the state of a dynamical system from observations is a crucial task in robotics and control. Despite their empirical success, recurrent models often lack formal guarantees on convergence and robustness. We bridge this gap by analyzing these architectures from the perspective of nonlinear observer theory. By framing RNNs as a time-discretization of continuous latent dynamical systems (Neural ODEs), we demonstrate that enforcing a contraction condition on the latent flow guarantees the existence of a smooth, injective map of the true physical state into the latent space. We prove that this architecture acts as a tunable observer: by increasing a single scalar gain that scales the latent vector field, the estimation error can be driven to an arbitrarily small neighborhood of zero at an exponentially fast rate. We further establish quantitative bounds on the MLP complexity required to decode this latent state, and extend our formal guarantees to the discrete-time residual RNN implementation. We validate the theory on a real mobile robotic task, where a provably contracting model outperforms a vanilla GRU with substantially more stable training. We also report that unconstrained GRUs converge to contracting dynamics on this task, suggesting that contraction is an emergent property of well-trained sequence models.


Pure Exploration Beyond Reward Feedback: The Role of Post-Action Context

Mohammad Shahverdikondori ⋅ Amir Mohammad Abouei ⋅ Alireza Rezaeimoghadam ⋅ Negar Kiyavash

We introduce the problem of best arm identification (BAI) with post-action context, a new BAI problem in a stochastic multi-armed bandit environment and the fixed-confidence setting. The problem addresses the scenarios in which the learner receives a \emph{post-action context} in addition to the reward after playing each action. This post-action context provides additional information that can significantly facilitate the decision process. We analyze two different types of the post-action context: (i) \textit{separator}, where the reward depends solely on the context, and (ii) \textit{non-separator}, where the reward depends on both the action and the context. For both cases, we derive instance-dependent lower bounds on the sample complexity and propose algorithms that asymptotically achieve the optimal sample complexity. For the separator setting, we propose a novel sampling rule called \textit{G-tracking}, which uses the geometry of the context space to directly track the contexts rather than the actions. For the non-separator setting, we do so by demonstrating that the Track-and-Stop algorithm can be extended to this setting. Moreover, in both settings, we theoretically and empirically show that algorithms that ignore the post-action context are sub-optimal. Finally, our empirical results showcase the advantage of our approaches compared to the state of the art.


QMaxCal: Path-Space Regularization for Open Quantum Control via Girsanov's Theorem

Merijn Moody ⋅ Zier Mensch ⋅ Miranda Cheng ⋅ Peter G Bolhuis ⋅ Max Welling

Reliable quantum control in the presence of decoherence requires policies that combat the effect of environmental noise on the controlled dynamics. Open quantum systems under continuous monitoring generate classical measurement records whose drift depends on the noise experienced by the system; the records of two evolutions sharing the same decoherence channels differ only in this drift, so Girsanov's theorem yields a closed-form, differentiable estimator of the KL divergence between their trajectory distributions. We instantiate this estimator with two physically motivated reference measures, yielding two regularizers that both drive the system toward states where the effects of decoherence are minimal: the Wiener KL (KLW), which is empirically more effective under certain conditions on the noise model, and the drift-variance regularizer (RDV), which works for all noise models. Both are qualitatively distinct from existing penalties on control fluence or smoothness: they penalize the observable consequences of control on the decoherence channels rather than the control amplitude itself. The regularizers outperform unregularized gradient-based and reinforcement-learning baselines across a range of open quantum systems — including single- and multi-qubit benchmarks and a multi-qubit chain calibrated to a published snapshot of the IBM Kingston processor — along several axes of evaluation: final-state fidelity, robustness to mismatch in the assumed noise model (gains grow from +17 pp at training noise to +27 pp under 2.5× noise mismatch), and occupation of forbidden states. The regularizers reduce infidelity by up to 50%, with ~16% gains on the calibrated IBM Kingston chain. Our code can be found at https://anonymous.4open.science/r/QMaxCal-6E71/.


Qrita: High-performance Top-k and Top-p using Pivot-based Truncation and Selection

Jongseok Park ⋅ Sunga Kim ⋅ Alvin Cheung ⋅ Ion Stoica

Despite their importance in model sampling, efficient implementation of Top-k and Top-p algorithms for large vocabularies remains a significant challenge. Existing approaches often rely on sorting, which incurs significant computation and memory overhead on GPUs, or on stochastic approaches that alter the algorithm's output. In this work, we propose Qrita, an efficient Top-k and Top-p algorithm based on a pivot-based truncation and selection. Qrita leverages pivot-based search for both Top-k and Top-p with two key techniques: 1. Gaussian-based σ-truncation, which greatly reduces the search space of the vocabulary, and 2. Quaternary pivot search with duplication handling, which halves the number of pivot search iterations and guarantees deterministic output. We implement Qrita using Triton and evaluate its performance against the Top-k and Top-p kernels of high-performance LLM execution engines such as vLLM, SGLang, and FlashInfer, improving end-to-end serving throughput up to 1.4× with half the memory usage, while providing the same output as the sorting-based algorithms.


Quantifying Concentration Phenomena of Mean-Field Transformers in the Low-Temperature Regime

Albert Alcalde ⋅ Leon Bungert ⋅ Konstantin Riedl ⋅ Tim Roith

Transformers with self-attention modules as their core components have become an integral architecture in modern large language and foundation models. In this paper, we study the evolution of tokens in deep encoder-only transformers at inference time which is described in the large-token limit by a mean-field continuity equation. Leveraging ideas from the convergence analysis of interacting multi-particle systems, with particles corresponding to tokens, we prove that the token distribution rapidly concentrates onto the push-forward of the initial distribution under a projection map induced by the key, query, and value matrices, and remains metastable for moderate times. Specifically, we show that the Wasserstein distance of the two distributions scales like $\sqrt{\log(\beta+1)/\beta}\exp(C t)+\exp(-ct)$ in terms of the temperature parameter $\beta^{-1}\to 0$ and inference time $t\geq 0$. For the proof, we establish Lyapunov-type estimates for the zero-temperature equation, identify its limit as $t\to\infty$, and employ a stability estimate in Wasserstein space together with a quantitative Laplace principle to couple the two equations. Our result implies that for time scales of order $\log\beta$ the token distribution concentrates at the identified limiting distribution. Numerical experiments confirm this and, beyond that, complement our theory by showing that for finite $\beta$ and large $t$ the dynamics enter a different terminal phase, dominated by the spectrum of the value matrix.


Quantizing With Randomized Hadamard Transforms: Efficient Heuristic Now Proven

Ran Ben-Basat ⋅ William Kuszmaul ⋅ Michael Mitzenmacher ⋅ Amit Portnoy ⋅ Shay Vargaftik

Uniform random rotations (URRs) are a common preprocessing step in modern quantization approaches used for gradient compression, inference acceleration, KV-cache compression, model weight quantization, and approximate nearest-neighbor search in vector databases. In practice, URRs are often replaced by randomized Hadamard transforms (RHTs), which preserve orthogonality while admitting fast implementations. The remaining issue is the performance for worst-case inputs. With a URR, each coordinate is individually distributed as a shifted beta distribution, which converges to a Gaussian distribution in high dimensions. Generally, one RHT is not suitable in the worst case, as individual coordinates can be far from these distributions. We show that after composing two RHTs on any $d$-sized input vector, the marginal distribution of every fixed coordinate of the normalized rotated vector is within $\mathcal O(d^{-1/2})$ of a standard Gaussian both in Kolmogorov distance and in $1$-Wasserstein distance. We then plug these bounds into the analyses of modern compression schemes, namely DRIVE and QUIC-FL, and show that two RHTs achieve performance that asymptotically matches URRs. However, while two RHTs suffice for scalar quantization, they may be insufficient for Vector Quantization (VQ), which often requires weak correlation across fixed-size blocks of coordinates (as opposed to only marginal distribution convergence for single coordinates). We prove that a composition of three RHTs leads to decaying coordinate covariance. This ensures that any fixed, bounded, multi-dimensional VQ codebook optimized for URRs has the same expected error when using three RHTs, up to an additive term that asymptotically vanishes with the dimension. Finally, because practical inputs are rarely adversarial, we propose a linear-time $\mathcal{O}(d)$ check on the input's moments to dynamically adapt the number of RHTs used at runtime to improve performance.


Redefining Instance Matching: A Unified Framework for Part-Aware Matching in Panoptic Segmentation Evaluation

Erik Großkopf ⋅ Soumya Snigdha Kundu ⋅ Hendrik Möller ⋅ Nicolas Münster ⋅ Mehdi Astaraki ⋅ Paula T Buzduga ⋅ Kerstin Ritter ⋅ Benedikt Wiestler ⋅ Jan Kirschke ⋅ Jonathan Shapey ⋅ Tom Vercauteren ⋅ Florian Kofler

The Panoptic Quality (PQ) metric is the standard for jointly evaluating instance and semantic segmentation. However, its original definition relies on a One-to-One matching between predicted and ground truth segments, which is only straightforward when the IoU threshold exceeds 0.5. Below 0.5, multiple matching strategies emerge in a poorly explored problem space. We systematically elucidate this space by recasting segment matching as a constrained bipartite assignment problem. Independently bounding the prediction- and ground-truth-side degrees yields four matching strategies: One-to-One, Many-to-One, One-to-Many, and Many-to-Many. We show that the first three are well-defined within the PQ framework, while Many-to-Many falls outside it. These strategies become relevant when instances are fragmented, adjacent objects are difficult to delineate, or annotations are noisy. Central to our framework is a vertex-based accounting of TP, FN, and FP, anchored to ground truth and predicted segments rather than to matching edges. We further show that the framework extends naturally to part-aware panoptic segmentation, and we explore part-aware evaluation on biomedical data. Across configurable case studies we report how different combinations of thresholds and matching strategies behave in practice. We release a unified open-source package built on Panoptica. It exposes Voronoi-based region-wise analysis, part-aware evaluation, and Area Under Threshold Curve computations as configurable options.


Referring and Reasoning Camouflaged Object Segmentation in Audio-Visual Scenes

Tianxin Han ⋅ Qing Dong ⋅ Xingwei Wang ⋅ Jie Jia

Camouflaged Object Segmentation (COS) aims to identify objects visually hidden in their surrounding environments. Existing COS benchmarks mainly focus on image-level or video-level visual perception, while real-world camouflaged scenes are often accompanied by audio signals and user intentions, where sound semantics and textual expressions provide critical information for localizing hidden targets. To extend the boundary of COS, we introduce a new task, termed Referring and Reasoning Audio-Visual Camouflaged Object Segmentation (R2-AVCOS), which aims to segment camouflaged targets in audio-visual scenes according to textual expressions with referring or reasoning intentions. This task emphasizes understanding audio content and incorporates complex reasoning and world knowledge into expressions. To support this task, we construct R2-AVCOSBench, the first audio-visual benchmark with pixel-level annotations for camouflaged objects specified by referring expressions or inferred through reasoning. It contains 2,654 audio-visual camouflaged videos, 21,232 annotated frames, and 32,329 expressions, including 17,543 referring and 14,786 reasoning expressions. Furthermore, we propose Camouflaged Instructed Segmentation Assistant (CISA), a baseline model built upon a Multimodal Large Language Model (MLLM). CISA understands complex textual and audio-visual cues and performs referring- and reasoning-based camouflaged object segmentation. Extensive experiments show that CISA achieves strong referring and reasoning segmentation ability in audio-visual camouflaged scenes and obtains competitive results on related tasks.


Reflected Schrödinger Bridge Matching

Marcus Häggbom ⋅ Viktor Nilsson ⋅ Pierre Nyquist ⋅ Joakim Andén

Recent advances in generative modeling have enabled the efficient computation of Schrödinger bridges (SB) in high-dimensional settings by leveraging partially simulation-free training methods inspired by flow matching. However, these have not covered SBs with reflecting dynamics, a useful model choice with built-in guarantees that generated samples stay in the data domain. Existing alternatives for reflected SBs instead rely on more complex training based on forward--backward SDE theory, requiring expensive higher-order derivatives and sampling entire paths during training. In this article, we introduce a partially simulation-free framework that allows reflected SBs to be trained similarly to flow matching, using a new sampling method and regression target. We demonstrate our results by coupling pairs of well-known high-dimensional image datasets. Using reflected dynamics incurs negligible additional wall-clock time during both training and inference while maintaining or slightly improving generative performance.


Reinforced Evidence-Aware Long Video Understanding

Yuan Xie ⋅ Tianshui Chen ⋅ Deyu Zhou ⋅ Lionel Ni

Despite strong benchmark performance, multimodal large language models (MLLMs) remain prone to hallucinations, as they are optimized for linguistic plausibility rather than faithful answering based on observation. This issue is particularly severe in long videos, where critical cues are temporally sparse and often missed by uniform frame sampling. Yet existing methods still either operate on fixed pre-sampled inputs, implicitly assuming that all necessary evidence has already been captured, or augment reasoning with retrieval without requiring answers to be grounded in visual evidence. We propose Reinforced Evidence-Aware Learning , a framework that reformulates long-video question answering as an iterative process of evidence gathering and verification, enabling faithful answering with explicit evidence support, thereby reducing hallucinations. At each reasoning step, the model assesses whether the current observation is sufficient to answer the question, and either searches for missing visual evidence from the video or produces a final answer with explicitly cited evidence. To internalize this reasoning policy, we design an evidence-aware reward for reinforcement learning post-training that jointly requires answer correctness and evidence faithfulness. The latter is evaluated by a pretrained cross-modal consistency verifier that matches the generated evidence descriptions against the visual content of the source frames, eliminating the need for manual annotations. Built on Qwen2.5-VL-7B, REAL achieves state-of-the-art results on four long-video benchmarks (Video-MME, MLVU, LongVideoBench, and EgoSchema) and substantially suppresses hallucinations on the dedicated video-hallucination benchmarks VideoHallucer and ELV-Halluc, with only 96 input frames surpassing the 768-frame base model in both accuracy and hallucination resistance.


Residual Expertise Is Not Decision Value

Nidhish Shah ⋅ Shaurjya Mandal ⋅ Asfandyar Azhar

Audits of human-AI complementarity target the wrong estimand. They ask whether the human knows more than the model; deployment asks whether that knowledge would change the action. Residual predictive expertise is not residual decision value. Deployment rewards actions, and a posterior movement that does not cross a reward-induced action boundary changes nothing; a signal can be informative under log-loss and leave the deployed action unchanged. We formalize the missing quantity as boundary regret: the regret of holding the model's action after observing the human. In any finite-action Bayesian decision problem, the value of consulting the human equals expected boundary regret, formally verified in Lean 4. The identity turns when humans help into a geometric question: which decision facets does the human signal move belief across? Complementarity is therefore reward-relative. A given human signal can be valuable under one reward matrix and worthless under another, the human's knowledge unchanged; the boundary moved, not the expertise. The gap is widest precisely where asymmetric costs dominate: high-stakes clinical and operational decisions. Audits built on residual predictive expertise overstate the value of human review and route scarce review toward cases where the human is informative but the action will not move. Boundary regret is the quantity to measure.


Rethinking Attention in Depth for Operator Learning

Jaehyeon Lee ⋅ Kiwon Um ⋅ JungHyun Han ⋅ Min-Koo Kang

Transformers have rapidly emerged as powerful surrogate models for learning solution operators for partial differential equations (PDEs). Standard attention tends to exhibit depth-wise representational bottlenecks across layers, with the role of depth remaining underexplored in modeling complex PDEs. In this paper, we rethink the role of depth by promoting it to an explicit computational dimension of attention kernel construction, rather than a passive architectural parameter. With this strategy, each attention block's native representation is enriched via a gated mixture of inter-layer differences, enabling information to be coupled and propagated across layers. The depth-wise hierarchy admits a theoretically motivated multi-kernel interpretation with weakly correlated kernels and improved generalization bounds. Beyond dense feature aggregation, our approach yields a more expressive form of depth-wise attention, mitigating spectral collapse, preserving interaction-scale diversity, and improving intra- and inter-layer route utilization. Extensive empirical evaluations demonstrate its broad compatibility across state-of-the-art solvers, presenting consistent performance gains on eight PDE benchmarks.


Rethinking Expressivity and Efficiency in Test-Time Training

Zeyun Zhong ⋅ Joya Chen ⋅ Manuel Martin ⋅ Frederik Diederichs ⋅ Jürgen Gall ⋅ Jürgen Beyerer

Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token update dynamics with the hardware efficiency of chunk-wise approximations. We propose E$^2$-TTT (Expressive and Efficient TTT) to bridge this gap. By deriving a closed-form state transition that exactly aggregates per-token momentum and decay coefficients within a chunk, E$^2$-TTT enables fully parallelized chunk-level training while preserving the temporal structure of the update rule that prior chunk-wise methods discard. We validate E$^2$-TTT by training models up to 1.3B parameters from scratch. Extensive experiments demonstrate that our method consistently outperforms previous TTT and hybrid attention baselines in language modeling and retrieval, while achieving significantly better extrapolation on the standard ``Needle in a Haystack'' test, maintaining $>90$ % accuracy on passkey retrieval at $8\times$ the training context length. Meanwhile, E$^2$-TTT can match the training throughput of efficient chunk-wise methods, demonstrating that it effectively reconciles expressivity with efficiency.


Rethinking Layer-wise Model Merging through Chain of Merges

Pietro Buzzega ⋅ Riccardo Salami ⋅ Angelo Porrello ⋅ SIMONE CALDERARA

Model merging has emerged as a simple yet effective method to fuse models without re-training. Existing techniques operate at the level of individual layers, thereby overlooking the inter-layer dependencies inherent in deep networks. We show that this simplification induces distributional mismatches in intermediate activations during merging, as changes applied to early layers fail to propagate to downstream ones. We identify these mismatches as a form of covariate shift that compounds across the network. To address this issue, we propose Chain of Merges (CoM), a novel procedure that sequentially merges weights across layers while accounting for inter-layer interactions. In particular, CoM mitigates covariate shift through a series of regression problems, where input activations are recomputed at each step to reflect the updated representations. Additionally, we introduce a dynamic weighting mechanism that prioritizes task-layer pairs most susceptible to performance degradation, improving robustness during merging. Experiments on standard benchmarks demonstrate that CoM achieves state-of-the-art performance across heterogeneous tasks. Code is available in the supplementary material.


Retrieval from Within: An Intrinsic Capability of Attention-Based Models

Elad Hoffer ⋅ Yochai Blau ⋅ Edan kinderman ⋅ Ron Banner ⋅ Daniel Soudry ⋅ Boris Ginsburg

Retrieval-augmented generation (RAG) typically treats retrieval and generation as separate systems. We ask whether an attention-based encoder-decoder can instead retrieve directly from its own internal representations. We introduce INTRA (INTrinsic Retrieval via Attention), a framework where decoder attention queries score pre-encoded evidence chunks that are then directly reused as context for generation. By construction, INTRA unifies retrieval and generation, eliminating the retriever-generator mismatch typical of RAG pipelines. This design also amortizes context encoding by reusing precomputed encoder states across queries. Across question-answering benchmarks, INTRA outperforms strong engineered retrieval pipelines on both evidence recall and end-to-end answer quality. Our results demonstrate that attention-based models already possess a retrieval mechanism that can be elicited, rather than added as an external module.


Robust Flow Matching under Target Corruption and Label Noise

Mert Can Kurucu ⋅ Erik Englesson ⋅ Hossein Azizpour

Flow Matching (FM) is a strong framework for generative modeling due to its stable training and efficient sampling. However, FM assumes clean target data and labels, an assumption often violated in practice. We formulate FM training as a time-dependent regression problem over conditional vector fields and analyze how the two corruption types affect the objective: i) target corruption adds a time-dependent residual term that grows with time, ii) label noise causes updates to the wrong class-conditional vector field. Based on this analysis, we propose to improve FM robustness by mapping corruption-specific residual scores to sample weights. For target corruption, we propose the \emph{Temporal Residual Score} (TRS), which emphasizes residual magnitudes at intermediate-to-late FM times, where clean and corrupted samples are more separable. For label noise, we propose the \emph{Comparative Conditional Score} (CCS), which compares the residual of the observed class-conditional vector field against alternative class-conditional vector fields for the same trajectory and velocity target. Experiments across controlled corruption settings and real-world noisy labels show that the proposed methods improve robustness over standard FM and robust-training baselines, supporting corruption-specific residual structure as a practical basis for robust FM.


RoSA: Rotational Sparse Adaptation for Memory-Efficient Fine-Tuning

Muhammad Azeem Lodhi ⋅ chao zhou ⋅ Rebekka Burkholz

Parameter-efficient fine-tuning (PEFT) reduces the cost of adapting foundation models by focusing training on a small parameter subset. Complementary to this idea, we introduce RoSA (Rotational Sparse Adaptation), which narrows adaptation to a subset of layers at a time. RoSA freezes lower layers close to the input throughout training and rotates a trainable block over later layers, consecutively increasing the number of frozen layers close to the input. This design reduces optimizer-state memory, shortens backpropagation, and even forward propagation if activations at the last frozen layer are cached. Because RoSA is orthogonal to the choice of trainable parameterization, it can be combined with PEFT methods or sparse optimizers inside each active block. Experiments across multiple LLM architectures and tasks show that RoSA reduces peak memory while maintaining strong fine-tuning performance.

Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed under evaluation-style prompts while the same behavior persists under ordinary-use prompts. Output scores reveal this mismatch but do not locate it. We investigate whether the distinction is encoded in a stable internal site and introduce an approach that fits a paired activation contrast at a path-patching-informed mid-depth window, then modifies the resulting coordinate on held-out prompts. The intervention closes the evaluation-to-deployment gap in ten of twelve model-behavior settings (six of the eight settings with $n{\geq}120$ paired questions) across four full-matrix instruction-tuned model instances; a fifth model supports localization and edit-provenance checks, and deployment-framed rates change by at most $6.1$pp. The two flat cells, both sycophancy, indicate that a single-coordinate audit is not sufficient when the installed distinction is higher-rank or missed by the depth heuristic. The audit is a diagnostic for fine-tuned checkpoints, not a training-time defense or a guarantee of deployment safety.


ScrapeGraphAI-100k: Dataset for Schema-Constrained LLM Generation

William Brach ⋅ Francesco Zuppichini ⋅ Lorenzo Padoan ⋅ Marco Vinciguerra

Producing output that conforms to a specified JSON schema underlies tool use, structured extraction, and knowledge base construction in modern large language models. Despite this centrality, public datasets for the task remain small, synthetic, or text-only, and rarely pair real page content with the prompts and schemas used in practice. We introduce ScrapeGraphAI-100k, 93{,}695 schema-constrained extraction events collected via opt-in ScrapeGraphAI telemetry in Q2--Q3 2025, deduplicated and balanced by schema from 9M raw events. The corpus spans 18{,}000+ unique schemas across 15 named languages plus a long-tail Other category, with English and Traditional Chinese covering 88\% of detected content, each instance pairs Markdown-converted page content with a prompt, schema, LLM response, and per-example jsonschema-rs structural conformance labels (semantic correctness is out of scope, and raw HTML is deferred beyond v1.0). We characterize structural diversity across the corpus and identify sharp failure thresholds as schema complexity grows. As a case study, a 1.7B student fine-tuned on this data closely tracks the output distribution of its GPT-5-nano teacher, though it still trails a 30B-A3B reference (3.3B active parameters) on schema compliance. We offer this distillation result as preliminary evidence that grounding schema-constrained generation in real practitioner workloads at scale enables training and benchmarking that prior synthetic or text-only corpora could not support.


Selling Information While Being an Interested Party

Francesco Bacchiocchi ⋅ Matteo Castiglioni ⋅ Alberto Marchesi ⋅ Giulia Romano ⋅ Nicola Gatti

We study the algorithmic problem faced by an information holder (seller) who wants to optimally sell information to a budged-constrained decision maker (buyer). The utilities of both agents depend on a random state of nature that is revealed to the seller, but unknown to the buyer. Differently from previous works, we consider the case in which the seller is an interested party, as the decision taken by the buyer also influences seller's utility. The seller's goal is to (partially) sell their information about the state of nature to the buyer, so as to concurrently maximize revenue and induce the buyer to take a desirable decision. We study settings in which buyer's budget and utilities are determined by a random buyer's type unknown to the seller. First, we propose a polynomial-time algorithm for computing an optimal seller's protocol, which proposes a menu of information-revelation policies to the buyer, who acquires one of them by paying its corresponding price. Then, we switch the attention to the case in which the seller can only employ a single information-revelation policy, rather than proposing a menu. In such a setting, we completely characterize the computational complexity of the seller's algorithmic problem.


Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech Gestures

Varsha Suresh ⋅ Mohammad Mahdi Abootorabi ⋅ Mohamed Salman ⋅ M. Hamza Mughal ⋅ Christian Theobalt ⋅ Ashwin Ram ⋅ Jürgen Steimle ⋅ Vera Demberg

Learning a shared representation between spoken text and gesture is central to co-speech gesture retrieval, synthesis, and understanding, but remains challenging for semantically meaningful gestures whose communicative intent is not captured by motion alone. Direct contrastive alignment between transcripts and continuous motion embeddings often overemphasizes low-level kinematics and misses the symbolic content of semantic gestures. We propose \emph{semantic motion anchors}, natural-language abstractions of gesture motion capturing physical form and communicative intent. Our method discretizes 3D gestures into body-hand motion primitives, verbalizes them into structured descriptions, and grounds them in the transcript to provide auxiliary contrastive supervision. On BEAT2, our method improves text-to-gesture R@1 by 8.2\% over a direct text-motion baseline and outperforms prior retrieval approaches on text to gesture and gesture to text retrieval directions. Beyond aggregate retrieval metrics, semantic motion anchor supervision helps retrieve gestures that are semantically meaningful for the spoken query, rather than defaulting to generic motion patterns. A downstream retrieval-augmented gesture generation study showed that users significantly preferred gestures retrieved by our approach over a retrieval-augmented generation baseline, demonstrating that semantically grounded retrieval translates to gestures that better convey communicative intent in downstream generation.


Shared Modular Recurrence in Contextual MDPs for Universal Morphology Control

Laurens Engwegen ⋅ Max Weltevrede ⋅ Caroline Horsch ⋅ Daan Brinks ⋅ Wendelin Boehmer

A universal controller for any robot morphology would greatly improve computational and data efficiency. Steps have been made towards such multi-robot control by utilizing contextual information about the properties of individual robots and exploiting their modular structure in the architecture of deep reinforcement learning agents. When the robots have highly dissimilar morphologies, however, this becomes a challenging problem, especially when the agent must generalize to new, unseen robots. In this paper, we posit that contextual features are often only partially available, but that they can be recovered through modular interactions. This can allow for better multi-robot control and generalization to contexts that are not seen during training. To this extent, we implement a transformer-based architecture with shared modular recurrence and evaluate its (generalization) performance on a large set of MuJoCo robots. The results show a substantial improvement in zero-shot generalization performance on robots with unseen dynamics, kinematics, and topologies, in four different environments.


Skip the Hessian, Keep the Rates: Globalized Semismooth Newton with Lazy Hessian Updates

Amal Alphonse ⋅ Pavel Dvurechenskii ⋅ Clemens Sirotenko

Second-order methods are provably faster than first-order methods, and their efficient implementations for large-scale optimization problems have attracted significant attention. Yet, optimization problems in ML often have nonsmooth derivatives, which makes the existing convergence rate theory of second-order methods inapplicable. In this paper, we propose a new semismooth Newton method (SSN) that enjoys both global convergence rates and asymptotic superlinear convergence without requiring second-order differentiability. Crucially, our method does not require (generalized) Hessians to be evaluated at each iteration but only periodically, and it reuses stale Hessians otherwise (i.e., it performs lazy Hessian updates), saving compute cost and often leading to significant speedups in time, whilst still maintaining strong global and local convergence rate guarantees. We develop our theory in an infinite-dimensional setting and illustrate it with numerical experiments on matrix factorization and neural networks with Lipschitz constraints.

Vision Transformers dominate modern image benchmarks, however, what trained self-attention computes head by head remains unclear: prior work either injects inductive biases at initialization or clusters attention patterns visually without committing to a closed-form computation. We ask whether the softmax$(QK^\top)$ inside a trained head can be replaced, without further training, by a structured closed-form operator while preserving prediction and the residual-stream trajectory. We answer yes across ten ViTs and seven pretrainings (ImageNet-1k, ImageNet-21k FT, MAE, DINO, DINOv2, CLIP, SigLIP), and package the result as \textbf{SOLA}: a Structured Operator Library for Attention with three entries, each gated by a per-head diagnostic that upper bounds substitution error. The entries are a fixed 2D convolution kernel for shift heads, a per-image broadcast row for fixed-target heads, and a per-image rank-$k$ SVD of the attention matrix for near-global heads. A rank-adaptive extension picks $k$ per head from the SVD spectrum and extends structural coverage to every head on every backbone, with task-fidelity controlled by a single spectral threshold $\tau$ that trades compression for downstream fidelity. The diagnostics carry over to cross-attention: every decoder cross-attention head in BLIP captioning has effective rank $\approx 1$ and admits the broadcast substitution. A minimal reproducer that runs the full SOLA pipeline on a single backbone is included in the Supp. Mat. as a demo; the full codebase will be released open-source upon acceptance.

Multiplicative motor and observation noise, along with internal noise corrupting computation, are central features of the sensorimotor system in biological and robotic agents. Yet, analytical solutions for stochastic optimal control are largely restricted to additive noise models and neglect internal noise. Here, we consider the problem of finding optimal control and filter laws for partially observable stochastic linear systems under quadratic costs with multiplicative control and observation noise, as well as internal noise. We provide an efficient, analytically derived coordinate-descent algorithm that computes mutually optimal linear control and filter laws by reformulating the problem as a constrained optimization. Our method provably guarantees monotonic decrease of the expected cost and convergence to a critical point, and overcomes the suboptimality and incompleteness of prior analytical approaches. Compared to state-of-the-art numerical methods, it achieves orders-of-magnitude computational speedups. Our solution also proves instrumental in revealing novel internal-noise-dependent, task-structured control strategies in a redundant arm control task.


Sparse Attention as Compact Kernel Regression

Saul Santos ⋅ Nuno Gonçalves ⋅ Daniel McNamee ⋅ Marcos Treviso ⋅ André Martins

Recent work has revealed a link between self-attention mechanisms in transformers and test-time kernel regression via the Nadaraya-Watson estimator, with standard softmax attention corresponding to a Gaussian kernel. However, a kernel-theoretic understanding of sparse attention mechanisms is currently missing. In this paper, we establish a formal correspondence between sparse attention and compact (bounded support) kernels. We show that normalized ReLU and sparsemax attention arise from Epanechnikov kernel regression under fixed and adaptive normalizations, respectively. More generally, we demonstrate that widely used kernels in nonparametric density estimation---including Epanechnikov, biweight, and triweight---correspond to $\alpha$-entmax attention with $\alpha = 1 + \frac{1}{n}$ for $n \in \mathbb{N}$, while the softmax/Gaussian relationship emerges in the limit $n \to \infty$. This unified perspective explains how sparsity naturally emerges from kernel design and provides principled alternatives to heuristic top-$k$ attention and other associative memory mechanisms. Experiments with a kernel-regression-based variant of transformers---Memory Mosaics---show that kernel-based sparse attention achieves competitive performance on language modeling, in-context learning, and length generalization tasks, offering a principled framework for designing attention mechanisms.


Speeding up Log-Sum-Exp: Kernel Fusion at the Memory Wall, Integer Arithmetic at the Compute Wall

Lingyun Yao ⋅ Martin Andraud ⋅ Niki Loppi ⋅ Andrea Pilzer ⋅ Anji Liu ⋅ Guy Van den Broeck ⋅ Martin Trapp

While the Log-Sum-Exp (LSE) function is a foundational numerical primitive for stable log-domain computation in modern machine learning, its fast and efficient computation on generic platforms (\ie, CPUs and GPUs) has been underexplored, potentially creating execution bottlenecks for large workloads. On GPU, the de facto implementation torch.logsumexp dispatches a chain of nine sub-kernels per call, incurring redundant High Bandwidth Memory (HBM) traffic and launch overhead. On CPU, exp and log operations are hardware-hungry, limiting the computing efficiency. To address these bottlenecks, we propose two drop-in kernels to significantly accelerate LSE computation on both GPUs and CPUs: FuseLSE, a single-pass fused kernel that evaluates LSE's exp and log functions on the GPU's Special Function Unit (SFU); and IntLSE, which further replaces the LSE computation with approximated integer arithmetic for CPU computation. Our experiments, validated by benchmarks across workload sizes and hardware platforms, show that both kernels are ${\sim}10-11\times$ faster than PyTorch at small workloads and ${\sim}6\times$ faster on larger ones on GPUs, where the speed-up is attributed to reduced launch overhead and HBM traffic. On CPU, where no SFU is available, LSE computation dominates the cost, allowing IntLSE to gain ${\sim}3-6\times$ execution speed over FuseLSE.


Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation

Ziyad Sheebaelhamd ⋅ Luca Viano ⋅ Volkan Cevher ⋅ Claire Vernade

This work investigates multi-objective imitation learning: the problem of recovering policies that lie on the Pareto front given demonstrations from multiple Pareto-optimal experts in a Multi-Objective Markov Decision Process (MOMDP). Standard imitation approaches are ill-equipped for this regime, as naively aggregating conflicting expert trajectories can result in dominated policies. To address this, we introduce Multi-Output Augmented Behavioral Cloning (MA-BC), an algorithm that systematically partitions divergent expert data while pooling state-action pairs where no behavior conflict is observed. Theoretically, we prove that MA-BC converges to Pareto-optimal policies at a faster statistical rate than any learner that considers each expert dataset independently. Furthermore, we establish a novel lower bound for multi-objective imitation learning, demonstrating that MA-BC is minimax optimal. Finally, we empirically validate our algorithm across diverse discrete environments and, guided by our theoretical insights, extend and evaluate MA-BC on a continuous Linear Quadratic Regulator (LQR) control task.


StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

Xixiang He ⋅ Baiqi Wu ⋅ Xingming Li ⋅ Ao Cheng ⋅ Qiyao Sun ⋅ Xuanyu Ji ⋅ Qingyong Hu

Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can correctly describe what it sees and correctly name the underlying pattern, yet still fail to choose the matching candidate. Existing AVR benchmarks cannot detect this gap because they collapse perception, rule induction, and answer selection into a single right-or-wrong signal. We introduce StemBind, a shared-stem diagnostic benchmark that probes the same visual stem with three aligned questions: Perception (what is in the image), Rule (what pattern governs it), and Full (which option completes it), so that a final-answer error can be attributed to a specific sub-step on the same evidence. StemBind contains 2,298 curated knowledge-light stems across nine auditable visual operations, totaling 19,533 P/R/F tasks, with each full item annotated by Sternberg's four reasoning stages: S1 Encode, S2 Infer, S3 Map, and S4 Apply. Evaluating 24 frontier MLLM configurations (proprietary and open-source) yields four findings. (i) The R-F chasm. Rule accuracy exceeds full-item accuracy on 22 of 24 models, so most failures happen after the rule has been identified. (ii) A persistent binding gap. Even when P and R are both correct on the same stem, models still answer F incorrectly 51.2% of the time. (iii) The bottleneck is S3. Process diagnostics and Stage-wise Stimulus Augmentation (SSA) localize the dominant failure to rule-to-instance mapping, the step that binds an inferred rule to the right candidate. (iv) Scaling and thinking do not help. Neither scaling up model size nor enabling explicit thinking mode reliably closes the gap; in paired comparisons, thinking mode lifts perception but lowers both rule and full-item accuracy. StemBind reframes AVR evaluation from final-answer ranking to locating where abstract visual reasoning breaks down, and identifies rule-to-instance binding as a concrete next target for vision-grounded reasoning.


Strategic PAC Learnability via Geometric Definability

Yuval Filmus ⋅ Shay Moran ⋅ Elizaveta Nesterova ⋅ Nir Rosenfeld ⋅ Alexander Shlimovich

Strategic classification studies learning settings in which individuals can modify their features, at a cost, in order to influence the classifier’s decision. A central question is how the sample complexity of the induced (strategic) hypothesis class depends on the complexities of the underlying hypothesis class and the cost structure governing feasible manipulations. Prior work has shown that in several natural settings, such as linear classifiers with norm costs, the induced complexity can be controlled. We begin by showing that such guarantees fail in general --- even in simple cases: there exist hypothesis classes of VC dimension $1$ on the real line such that, even under the simplest interval neighborhoods, the induced class has infinite VC dimension. Thus, strategic behavior can turn an easy learning problem into a non-learnable one. To overcome this, we introduce structure via a geometric definability assumption: both the hypothesis class and the cost-induced neighborhood relation can be defined by first-order formulas over $\mathbb R_{\exp}$. Intuitively, this means that hypotheses and costs can be described using arithmetic operations, exponentiation, logarithms, and comparisons. This captures a broad range of natural classes and cost functions, including $\ell_p$ distances, Wasserstein distance, and information-theoretic divergences. Under this assumption, we prove that learnability is preserved, with sample complexity controlled by the complexity of the defining formulas.

We investigate the problem of uncertainty quantification in the mean estimation framework. Given an i.i.d. sample from a distribution with unknown mean $\mu$ and variance $\sigma^2$, we want to construct an estimator $\hat{\mu}_n$ and provide a computable and size-optimal upper bound for the error $|\hat{ \mu}_n - \mu|$. While estimators possessing strong deviation guarantees are well known, no data-dependent non-asymptotic upper bounds for their performance exist in general. We show that this gap can be bridged by introducing a parameter that characterizes the "effective sample size" available for uncertainty quantification. This parameter captures the difficulty of the problem, interpolating between "easy" cases where sub-Gaussian confidence intervals exist and "hard" cases where their construction is impossible. Using this characterization, we design confidence intervals of optimal length that are fully adaptive to the unknown variance. Numerical experiments confirm that our approach maintains nominal coverage even in asymmetric and heavy-tailed regimes where other existing methods fail.


Submodular Multi-Agent Reinforcement Learning for Effective Online Distributed Task Allocation

Jing Liu ⋅ Yangyang YANG ⋅ Luca Ballotta ⋅ Fangfei Li ⋅ Yang Tang ⋅ Ruggero Carli

This paper studies multi-agent reinforcement learning with submodular team utilities, which models scenarios where $N$ agents solve a non-additive task allocation problem in a distributed manner online. Since each agent selects one action from a local categorical distribution at each time step, feasible joint actions form a partition matroid over agent-action pairs. The standard continuous relaxation of set utility functions, the Multilinear Extension, does not encode categorical constraints on factorized policies and may yield inconsistent gradient estimation. To remedy this, we propose the \emph{Partition Multilinear Extension}, a continuous relaxation that equals the expected team utility with factorized categorical policies under partition matroid constraint. We prove that submodular difference rewards provide unbiased PME marginal-gradient information and induce a stagewise score-function policy-gradient estimator for factorized categorical policies. Building on these results, we propose \emph{SubMAPG}, a centralized training with decentralized execution (CTDE) multi-agent policy-gradient framework that implements submodular difference-reward training signals and masked categorical policies for partition-feasible decentralized execution. For the associated PME marginal-space projected stochastic-gradient dynamics, we establish a stagewise $\frac{1}{2}$-approximation guarantee and sublinear dynamic regret under slowly varying environments, measured by the path length of the optimal PME marginals. Finally, to handle open systems where agents and targets may leave and join over time (e.g., modeling failure and recovery of robots or smart sensors), we implement SubMAPG with a graph neural network policy model. Numerical experiments on multi-robot coverage and multi-target tracking show that SubMAPG outperforms local greedy and shared-reward baselines, and is competitive with centralized myopic greedy strategies.


TabBioMed: A Large-Scale Benchmark for Biomedical Tabular Learning

Pau Mateo Bernadó ⋅ Saivenkata Nagavyjayanthi Polapragada ⋅ Pol Arbiol Rakuljic ⋅ Laia M Pladevall ⋅ David Bonet ⋅ Marçal Comajoan Cara ⋅ Jesus Gonzalez Ferrer ⋅ Jordi Abante ⋅ Daniel Mas Montserrat ⋅ David Haussler ⋅ Alexander Ioannidis

Recent advances in machine learning have accelerated biomedical prediction, yet evaluation on structured biomedical data remains fragmented, especially for datasets with high dimensionality, missingness, class imbalance, and small-sample, many-feature regimes. We introduce TabBioMed, a large-scale benchmark of 95 curated public biomedical tabular datasets spanning electronic health records \cite{data2016secondary}, drug response \cite{wu2018moleculenet}, genomics, transcriptomics, proteomics, metabolomics \cite{yang2025mlomics}, single-cell omics \cite{heumos2023best}, and systems biology. TabBioMed unifies these datasets through a standardized, dataset-aware preprocessing and evaluation framework, enabling reproducible comparison across classification, regression, and multi-target tasks. We benchmark 23 models across classical baselines, gradient-boosted trees, neural tabular architectures, and emerging tabular foundation models. Foundation models deliver the strongest family-level predictive performance. However, their gains are regime-dependent: tree-based models and tuned MLPs remain highly competitive, often dominating under inference-time or training-cost constraints. We further show that foundation-model advantages are most meaningful in low-noise biomedical tasks, where small absolute gains translate into large reductions in residual error. Overall, TabBioMed reveals that biomedical tabular learning is strongly context-dependent, with optimal model choice shaped by data regime, task type, and computational budget. By releasing curated datasets, preprocessing code, baselines, and results, TabBioMed provides a reproducible foundation for practical model selection and future methods development in biomedical machine learning, available at https://anonymous.4open.science/r/tabbiomed.


TabPrep: Closing the Feature Engineering Gap in Tabular Benchmarks

Andrej Tschalzev ⋅ Nick Erickson ⋅ Yuyang (Bernie) Wang ⋅ Huzefa Rangwala ⋅ Stefan Lüdtke ⋅ Christian Bartelt ⋅ Heiner Stuckenschmidt

Progress in tabular machine learning has largely focused on increasingly sophisticated model architectures. At the same time, feature engineering remains a critical yet underexplored component of real-world modeling pipelines that is entirely absent from modern benchmarks, which creates an unquantified evaluation gap. In this work, we introduce TabPrep, a lightweight preprocessing pipeline composed of feature generators that are carefully designed to target three specific structural data patterns. We show that many widely used model classes exhibit predictable blind spots to these patterns and that systematic feature engineering alone can establish new peak performance. Across the TabArena benchmark, integrating TabPrep into model training and tuning consistently improves performance for tree-based, neural, linear, and foundation models, often surpassing gains achieved by model-centric innovations alone. TabPrep outperforms previous automated feature engineering approaches in performance, efficiency, and applicability across datasets, enabling integration into large-scale benchmarks. By releasing TabPrep, we enable researchers to integrate feature engineering into their benchmarking setup, filling a longstanding gap in tabular evaluations.


TAMEing the Open-World Personalization: Towards Open-Set Personalized MLLM Assistant

Rongpei Hong ⋅ Jian Lang ⋅ Ting Zhong ⋅ Fan Zhou

Personalizing Multimodal Large Language Models (MLLMs) as intelligent assistants has recently emerged as an active research topic, where an MLLM is expected to go beyond generic, one-size-fits-all replies, and instead generate responses grounded in user-specific objects, entities, and concepts encountered during interactions. Nevertheless, existing studies mainly focus on personalizing MLLMs on a static, predefined set of personalized concepts and require users to manually and frequently extend this set to accommodate newly emerging concepts, which is infeasible in dynamic-world scenarios. To address this limitation, we recast conventional MLLM personalization as an Open-Set problem, and propose TAME-O, the first long-context open-set personalized MLLM assistant. Specifically, TAME-O introduces a training-free, plug-and-play Concept Reconciliation Skill (CR Skill). It enables the assistant to automatically unlock novel concepts during prolonged human–MLLM interactions while preventing confusion with previously unlocked ones, allowing the MLLM to continually grasp more and more concepts and progressively refine the modeling of known ones. Extensive experiments under the open-set personalization setting underscore the efficacy of our method.

Transformers for temporal sequences require mapping sampled signals into sequences of vectors, called tokens. This is done by partitioning the sequence into local temporal windows (patches). We show that the tokenization step induces an approximation--estimation tradeoff. Smaller windows yield higher-variance statistical estimates because each token is supported by fewer samples, while larger windows stabilize estimation but increase the amount of information that must be summarized. Motivated by this tradeoff, we propose a lightweight tokenizer that constructs multiple estimates for each token and sums them, without changing the Transformer backbone. Across 12 datasets spanning four modalities, our method improves performance and reduces sensitivity to temporal scale across tasks.


Tessellations of Semi-Discrete Flow Matching

Emile Pierret ⋅ Johannes Hertrich ⋅ Samuel Hurault ⋅ Julie Delon

We study Flow Matching in a semi-discrete setting where a Gaussian source is transported toward a discrete target supported on finitely many points. This semi-discrete regime is the theoretical setting behind the use of Flow Matching for generative modeling, where the target distribution is represented by a finite dataset. In this semi-discrete regime, the exact Flow Matching velocity field is available in closed form, which makes it possible to analyze the geometry induced by the terminal flow map independently of optimization and approximation effects. We investigate the terminal assignment regions, namely the preimages of the target atoms under the terminal flow. We show that these regions are open, simply connected and, under an additional assumption, homeomorphic to the unit ball. At the same time, a planar four-point example shows that these cells can differ sharply from Laguerre cells arising in semi-discrete optimal transport: they may be non-convex, have curved boundaries, and exhibit different boundedness and adjacency patterns. These results clarify the geometry intrinsically induced by the exact semi-discrete Flow Matching objective before neural approximation enters the picture.


Text-to-CAD Evaluation with CADTests

Dimitrios Mallis ⋅ Marco Wang ⋅ Ahmet Karadeniz ⋅ Elisa Ricci ⋅ Anis Kacem ⋅ Djamila Aouada

Text-to-CAD has recently emerged as an important task with the potential to substantially accelerate design workflows. Despite its significance, there has been surprisingly little work on Text-to-CAD evaluation, and assessing CAD model generation performance remains a considerable challenge. In this work, we introduce a new evaluation perspective for Text-to-CAD based on automated testing. We propose CADTestBench, the first test-based benchmark for Text-to-CAD, based on CADTests, executable software tests that verify whether a generated CAD model satisfies the geometric and topological requirements of the input prompt. Using CADTestBench, we conduct comprehensive benchmarking of recent Text-to-CAD methods and further demonstrate that CADTests can also guide CAD model generation, yielding simple baselines that surpass performance of current methods. CADTestBench code and data are available at GitHub and Hugging Face dataset.


The BV4 Benchmark for Unsupervised Anomaly Detection in High-Dimensional Spectral Data Streams

Nicolas Rojas Varela ⋅ Julien Ah-Pine ⋅ Engelbert MEPHU NGUIFO

This paper introduces the BV4 Benchmark, a collection of high-dimensional spectral datasets for unsupervised anomaly detection in vacuum environment data streams. Collected via Optical Emission Spectroscopy (OES), the benchmark comprises nine experiments representing practical scenarios structured according to established anomaly detection literature and validated by domain experts to address challenges such as spatial anomalies, temporal anomalies, and concept drift. BV4 datasets are composed of 2048 distinct wavelengths recorded over time, accompanied by timestamps and ground truth labels, enabling rigorous evaluation of machine learning models under dynamic environmental changes within vacuum chambers. We demonstrate the utility of BV4 through a simple comparative evaluation against previous datasets for anomaly detection in spectral data streams and discuss the current limitations and future improvements of the benchmark to support meaningful evaluative claims in the area.

Submission timing in human evaluation systems is universally treated as an irrelevant detail. We identify a systematic temporal leniency bias: across 129,023 reviews spanning three consecutive years of a large-scale ML venue, evaluators submitting closer to the deadline assign higher within-paper scores and produce less thorough assessments. The pattern holds in all 15 robustness specifications across all three years, and early evaluators achieve measurably lower Brier scores ($\Delta \approx 0.005$--$0.006$, stable across years). We propose two complementary corrections. Optimal Timing-Calibrated Aggregation (OTCA) derives a linear weight function $w(t) = a^{\star} t + 1$ by minimising paper-level decision loss, significantly outperforming heuristic alternatives (McNemar $\chi^2 = 46.3$, $p < 0.001$). Adversarial Temporal Debiasing (ATD) learns timing-invariant representations via Gradient Reversal, reducing timing discriminability by $37.4\\%$ while preserving predictive utility. Combined in a joint scheme, the two methods improve borderline decision accuracy by up to $11.25\\%$ over uniform averaging (McNemar $p < 0.001$), consistently across all three years. Both methods require only submission timestamps, already logged by every major evaluation platform, making deployment cost-free.


The Dead Salmons of AI Interpretability: The Need for a Statistical Inference Perspective

Maxime Méloux ⋅ Giada Dirupo ⋅ François Portet ⋅ Maxime Peyrard

In a striking neuroscience study, the authors placed a dead salmon in an MRI scanner and showed it images of humans in social situations. Astonishingly, standard analyses of the time reported brain regions \textit{predictive} of social emotions. The explanation, of course, was not supernatural cognition but a cautionary tale about misapplied statistical inference. In AI interpretability, reports of similar ``dead salmon'' artifacts abound: feature attribution, probing, sparse auto-encoding, and even causal analyses can produce plausible-looking explanations for randomly initialized neural networks. Upon inspection of this phenomenon our position is to argue for a pragmatic \textbf{statistical--causal reframing}: explanations of computational systems should be treated as parameters of a (statistical) model, inferred from computational traces. This perspective goes beyond simply measuring statistical variability of explanations due to finite sampling of input data; interpretability methods become statistical estimators, and findings should be tested against explicit and meaningful \textbf{alternative computational hypotheses}, with uncertainty quantified with respect to the postulated statistical model. It also highlights important theoretical issues, such as the identifiability of common interpretability queries, which we argue is critical to understand the field’s susceptibility to false discoveries, poor generalizability, and high variance.


The Future of Facts: Tracing the Factual Generation-Verification Gap

Tim R Davidson ⋅ Anja Surina ⋅ Caglar Gulcehre

Language models are becoming the default interface to factual knowledge, yet they often verify outputs more reliably than they generate them. This generation-verification gap (GV-gap) underlies many recent advances in self-improvement and reasoning, but its dynamics on factual knowledge specifically remain poorly understood. We focus on the training mechanisms underlying factual GV-gaps, distinguishing them from their computational and aesthetic counterparts. We trace generation and verification capabilities through three training phases (acquisition, continual learning, and updating) across four open-source model families at two scales each. Three findings recur across models: (i) verification is consistently learned before generation; (ii) verification is more robust to continual learning than generation; and (iii) factual updates can leave models in a multi-verse state, simultaneously verifying both old and new answers as correct. Natural experiments on frontier models reproduce these dynamics at scale and reveal residual verification biases on well-covered facts.

We study cooperative multi-agent reinforcement learning in the setting of reward-free exploration, where multiple agents jointly explore an unknown MDP in order to learn its dynamics (without observing rewards). We focus on a tabular finite-horizon MDP and adopt a phased learning framework. In each learning phase, multiple agents independently interact with the environment. More specifically, in each learning phase, each agent is assigned a policy, executes it, and observes the resulting trajectory. Our primary goal is to characterize the tradeoff between the number of learning phases and the number of agents, especially when the number of learning phases is small. Our results identify a regime change governed by the horizon $H$. When the number of learning phases equals $H$, we present a computationally efficient algorithm that uses only $\tilde{O}(S^6 H^6 A / \epsilon^2)$ agents to obtain an $\epsilon$ approximation of the dynamics (i.e., yields an $\epsilon$-optimal policy for any reward function). We complement our algorithm with a lower bound showing that any algorithm restricted to $\rho < H$ phases requires at least $A^{H/\rho}$ agents to achieve constant accuracy. Thus, we show that having $\Theta(H)$ learning phases is both necessary and sufficient when restricting the number of agents to be polynomial.


The Neural Race Model

Omar Ghezzi ⋅ Alessandro D’Amelio ⋅ Vittorio Cuculo ⋅ Rita Cucchiara ⋅ Giuseppe Boccignone

Perceptual decisions unfold in time: noisy evidence is integrated until a commitment threshold is reached, jointly producing a choice and a reaction time. Classical race models capture this mechanism through simple parametric drift and diffusion, but cannot represent how rich, high-dimensional stimuli shape accumulation on individual trials. We introduce the $\textbf{Neural Race Model}$ (NRM), a stimulus-conditional neural stochastic differential equation in which one accumulator per alternative races towards a learnable, stimulus-dependent threshold. Drift, diffusion, and threshold are parameterised by neural networks conditioned on a stimulus embedding, allowing both accumulation dynamics and decision urgency to vary across trials. The model is trained end-to-end on the joint reaction-time-and-choice likelihood via a smooth surrogate for the discontinuous first-passage event, combined with Monte Carlo trajectory averaging and gradient propagation through the stochastic adjoint. The surrogate recovers the classical hard threshold in the appropriate limit and is removed at test time. Each component admits a natural correspondence with elements of the primate decision-making circuit, preserving the mechanistic interpretability of classical models. We evaluate the NRM on multiple perceptual decision-making benchmarks against classical sequential-sampling and image-computable reaction-time models. Across tasks and metrics, the NRM tracks the human noise ceiling more closely than all baselines, demonstrating that stimulus-conditional stochastic accumulation and end-to-end differentiability can be achieved within a single, principled framework. Code will be publicly released.


The Panel Complexity of Sortition: Is 12 Angry Men Enough?

Johannes Brustle ⋅ Simone Fioravanti ⋅ Tomasz Ponitka ⋅ Jeremy Vollen

Sortition is the practice of delegating public decision-making to randomly selected panels. Recently, it has gained momentum worldwide through its use in citizens' assemblies, sparking growing interest within the computer science community. One key appeal of sortition is that random panels tend to be more representative of the population than elected committees or parliaments. Our main conceptual contribution is a novel definition of representative panels, based on the Wasserstein distance from statistical learning theory. Using this definition, we develop a framework for analyzing the panel complexity problem—determining the required panel size to ensure desirable properties. We focus on three key desiderata: (1) that efficiency at the panel level extends to the whole population, measured by social welfare; (2) that fairness guarantees for the panel translate to fairness for the population, captured by the core; and (3) that the probability of an outlier panel, for which the decision significantly deviates from the optimal one, remains low. We establish near-tight panel complexity guarantees for these desiderata across two fundamental social choice settings: facility location and participatory budgeting.


The Sampling Complexity of Condorcet Winner Identification in Dueling Bandits

El Mehdi Saad ⋅ Victor Thuot ⋅ Nicolas Verzelen

We study best-arm identification in large-scale stochastic dueling bandits under the sole assumption that a Condorcet winner exists, i.e., an arm that wins each noisy pairwise comparison with probability at least $1/2$. We introduce a new identification procedure that exploits the full gap matrix $\Delta_{i,j}=q_{i,j}-\tfrac12$ (where $q_{i,j}$ is the probability that arm $i$ beats arm $j$), rather than only the gaps between the Condorcet winner and the other arms. We derive high-probability, instance-dependent sample-complexity guarantees that (up to logarithmic factors) improve the best known ones by leveraging informative comparisons beyond those involving the winner. We complement these results with matching lower bounds that establish the optimality of our procedures in all regimes. Overall, our results reveal the general form of the sampling complexity, characterized by a trade-off between the cost of locating informative entries and the verification cost required to achieve the desired confidence. In particular, this complexity drastically differs from what is suggested by pure asymptotic results or by procedures that are tailored to Strongly Stochastic Transitive models.


Think2SQL: Blueprinting Reward Density and Advantage Scaling for Effective Text-to-SQL Reasoning

Simone Papicchio ⋅ Simone Rossi ⋅ Luca Cagliero ⋅ Papotti Paolo

While Large Language Models (LLMs) have advanced the state-of-the-art in Text-to-SQL, robust reasoning in complex, multi-table environments remains a bottleneck for parameter-efficient models. This paper presents a systematic empirical study on injecting reasoning capabilities into Text-to-SQL through the lens of Reinforcement Learning with Verifiable Rewards (RLVR) for the Qwen3 model family. We uncover a critical interplay between reward density, advantage scaling, and model capacity. Our analysis yields four primary insights. First, we propose a novel execution-guided dense reward function that significantly outperforms binary signals and existing state-of-the-art rewards by providing granular feedback at the instance level. Second, we analyze the mechanics of advantage calculation, demonstrating that while large models thrive on sparse signals with aggressive advantage scaling, smaller models require dense rewards and conservative scaling to improve Text-to-SQL performance. Third, we evaluate the impact of cold start, showing that distillation does not always benefit RLVR performance, and supervised fine-tuned models are prone to distributional mimicry. Fourth, we map the Pareto frontier of training efficiency, providing insights for optimizing Text-to-SQL reasoning under computational constraints. Our findings culminate in the Think2SQL family: our 4B-parameter model demonstrates reasoning capabilities competitive with state-of-the-art models such as o3. We release our models, datasets, and code to create a blueprint for RLVR optimization in Text-to-SQL at https://github.com/spapicchio/Think2SQL.


Thinking in Boxes: 3D Editing in Real Images Made Easy

Pradhaan Bhat ⋅ Naveen Chandra R ⋅ Rishubh Parihar ⋅ Vaibhav Vavilala ⋅ Venkatesh Babu Radhakrishnan ⋅ David Forsyth ⋅ Anand Bhattad

Text and 2D-conditioning interfaces provide weak, ambiguous control over spatial transformations in image editing -- particularly under large object motions and camera changes. Prior work has used 3D primitives such as boxes, but only as loose conditioning signals indicating approximate object location rather than specifying the transformation. We instead use 3D boxes as structured specifications: the user provides the input and output boxes of the edit, casting editing as a well-posed geometry problem. This ``thinking in boxes'' interface, where each box face is color-coded to convey 3D orientation, gives precise control over translation, rotation, scaling, and viewpoint changes in real images while preserving scene and object identity, and recovering previously unseen object regions. To ground transformations in scene appearance, we introduce a depth-aligned planar floor as a global reference frame, shaded with depth-aware cues. Conditioned on this structure, an image generator produces consistent results under large transformations. Trained in two stages -- on synthetic multi-object scenes and a small set of real-world videos from Objectron -- the system generalizes to complex, in-the-wild real images. Our method operates directly on real photographs and substantially outperforms recent state-of-the-art methods on large 3D edits.

Recent Continuous Thought Machine architecture decouples internal computation from external inputs via neural dynamics, but relies on multi-layer perceptrons without stability guarantees. We propose to model neural dynamics using asymmetric Excitatory-Inhibitory (E-I) networks, which can be stabilized via principles from network theory and can be expressed as energy-based systems optimized through a game-theoretic loss. Building on this perspective, we introduce Temporal Inhibitory-Excitatory Dynamic Engine (TIDE), a neuro-inspired architecture that computes internal representations through neural dynamics stabilized by incorporating the Wilson-Cowan dynamics and lateral inhibition. TIDE balances biological realism by, for instance, using Hierarchical Receptive Fields and enforcing Dale's principle to ensure a realistic $80:20$ E-I balance ratio with an end-to-end trainable architecture. The aim of this paper is to introduce a new architecture that brings neuro-inspired learning to the forefront. We present proofs of convergence, stability, and complexity bounds, along with empirical ablation studies. Overall, TIDE surpasses CTM with under $50$\% of the training time and improves $\texttt{top-1}$ accuracy by an average of $+1.65$\% on ImageNet under various perturbations.

OpenReview's October 2025 announcement to pilot an AI-assisted peer review process for AAAI-26, and the news in November 2025 that ICLR-26 was flooded with AI-generated reviews, show that a discussion about how to design the peer-review process of the future is long overdue. To foster a community-wide debate, we distill a set of 24 core capabilities required for scientific review as the foundation for a taxonomy on the dimensions and tasks relevant to peer reviewing. On this basis, we systematically and critically reflect on whether AI assistance tools should be integrated in the review process. We argue that the current state of the art does not warrant a general inclusion of such technology: Apart from security issues and ethical concerns, the capabilities of generative AI are for the most part not reliable enough yet, insufficiently tested, or inadequate by design to be integrated in the peer review process. With this paper we pave the way toward structured benchmarking across multiple dimensions.


TopoFisher: Learning Topological Summary Statistics by Maximizing Fisher Information

Matteo Biagetti ⋅ Mathieu Carrière ⋅ Francesco Conti ⋅ Enrico Maria Ferrari ⋅ Sven Heydenreich ⋅ Karthik Viswanathan

Persistence diagrams provide stable, interpretable summaries of geometric and topological structure, and are useful for simulation-based inference when important information is not captured by low-order statistics. In practice, however, persistence-based pipelines require hand-chosen filtrations, vectorizations, and compressors, usually without an objective tied directly to parameter uncertainty. We introduce \textbf{TopoFisher}, a differentiable persistent-homology pipeline that learns topological summaries by maximizing local Gaussian Fisher information. From simulations near a fiducial parameter value, TopoFisher optimizes trainable filtrations, diagram vectorizations, and compressors without posterior samples or supervised regression targets, while preserving the inductive bias of stable topological descriptors. We also give sufficient regularity conditions under which the log-determinant Fisher loss is locally Lipschitz in the trainable parameters. Controlled experiments on noisy spirals and Gaussian random fields, where the total Fisher information is known, validate the pipeline: TopoFisher recovers a large fraction of the available information and improves over fixed topological vectorizations. Our main results are on weak gravitational lensing, a high-dimensional non-Gaussian field-inference problem from cosmology. There, both learned topological summaries, a fixed cubical filtration with a learned PersLay vectorization, and a learned CNN filtration with the same PersLay vectorization, reach $\log|F|\approx 21$, compared with $13.8$ for the power spectrum, $17.1$ for peak counts, and $19.3$ for wavelet scattering, approaching an unconstrained Information Maximising Neural Network baseline ($22.4$) with up to $\sim80\times$ fewer parameters. More importantly, the fixed-filtration variant generalizes better: under simulator shift from lognormal to LPT-based maps it retains $\log|F|=19.24$ while the neural baseline drops to $8.9$, and in neural posterior estimation it yields tighter constraints than the neural baseline, power spectrum, peak counts, and wavelet scattering. These results suggest Fisher-based topological optimizations as a robust, parameter-efficient front end for simulation-based inference.


Towards Complementary Keypoint Detection via Mixture of Detectors

Xiaoyong Lu ⋅ Guobao Xiao ⋅ Dong Liang ⋅ Songlin Du

Keypoint detection plays a central role in geometric vision pipelines. Existing detectors exhibit complementary strengths across different scene structures and keypoint budgets, and reliance on a single detector often results in suboptimal keypoint distributions. Motivated by this observation, we propose Mixture of Detectors (MoD) for dense matching pipelines, a unified framework for complementary keypoint detection through the adaptive composition of heterogeneous detectors. MoD first distills heterogeneous detectors into a unified model, allowing diverse detection maps to be predicted efficiently within a single forward pass. Detector selection is subsequently formulated as a patch-wise, budget-conditioned decision process, in which a lightweight router assigns the most suitable detector to each local region. To avoid heuristic supervision, the router is trained with task-driven rewards derived from downstream matching quality, jointly capturing both keypoint quantity and precision. Extensive experiments on MegaDepth, ScanNet, and HPatches demonstrate that MoD consistently outperforms strong baselines across a wide range of keypoint budgets, demonstrating its effectiveness in improving keypoint distribution and downstream geometric performance.


Towards High Semantic Fidelity: Hyperdimensional Symbolic Messages in Multi-Agent Communication

Shoucheng Song ⋅ Youfang Lin ⋅ Chang Yao ⋅ Hao Wu ⋅ Sheng Han ⋅ Kai Lv

Semantic fidelity is essential for multi-agent communication under non-ideal channels, comprising transmission fidelity and conversion fidelity. Yet existing numerical and neural message formats cannot simultaneously achieve both. To address this, we propose a novel message format-\textbf{H}yperdimensional \textbf{S}ymbolic \textbf{M}essages (HSM), and its corresponding communication framework \textbf{Herm}. For transmission fidelity, Herm suppresses semantic entanglement via an entity-attribute-value structured encoding pipeline. For conversion fidelity, Herm first leverages a receiver-centric ego-binding mechanism to align semantics. Then, a similarity-aware dual-branch fusion module is proposed to prevent critical semantic dilution. To realize explicit message-semantic conversion, we devise an iterative decoding scheme via algebraic inverse operation. Experiments exhibit that Herm outperforms baselines across various noisy scenarios while maintaining superior structural readability.


Towards more general control of diffusion models using Jeffrey Guidance

Raphaël Razafindralambo ⋅ Rémy Sun ⋅ Frederic Precioso ⋅ Jes Frellsen ⋅ Pierre-Alexandre Mattei

A key strength of diffusion models lies in their flexibility, since their outputs can be controlled at sampling time through guidance. However, beyond simple cases such as conditional sampling, the target distribution is often left implicit, defined only through a sampling rule or a heuristic energy function. To address this, we propose Jeffrey guidance, a principled framework that extends diffusion-model control to applications beyond what standard guidance can express. It leverages Jeffrey’s rule of conditioning to update marginal distributions towards a prescribed target, preserving the conditional structure and minimally perturbing the joint distribution. We first demonstrate Jeffrey guidance by targeting a prescribed embedding distribution. With Inception embeddings as the target, this leads to substantial reductions in FID on both CIFAR-10 and FFHQ. We further apply Jeffrey guidance to fairness on CelebA-HQ, updating an unconditional diffusion model to enforce independence between attributes.


Tracing Actual Causes with Counterfactual Witness Maps

Tim Woydt ⋅ Matej Zečević ⋅ Kristian Kersting ⋅ Moritz Willig

Halpernian actual causation, the formal study of which events in a specific situation caused a specific outcome in acyclic structural causal models, underpins legal responsibility, moral blame, and the assessment of causal harm. Despite the rich logical framework, no prior sound and complete enumeration procedure is known for finding all actual causes in a given setting, blocking downstream tasks such as responsibility attribution, blame assessment, and causal-harm grading that aggregate over the full cause set. We close this gap by introducing witness targets: sets of joint configurations containing every counterfactual witness. In combination with an ancestor intervention grammar we derive a finite directed acyclic witness map. We prove that for every actual cause there exists a contingency set such that the joint intervention set appears as a node of this map, and present an algorithm for Tracing Actual Causes (TrAC) that is sound and complete for enumerating all actual causes of any target event. We compare TrAC empirically with existing single-cause baselines on random Boolean structural causal models, demonstrating practical feasibility for graphs of up to 15 nodes.


Training a Predictive Coding Network on ImageNet using Equilibrium Propagation

Tugdual Kerjan ⋅ Rasmus Høier ⋅ Benjamin Scellier

Equilibrium Propagation (EP) is a training framework for energy-based models (EBMs) that has attracted interest in the context of neuromorphic computing platforms such as continuous Hopfield networks, nonlinear resistive networks and coupled phase oscillators. However, EP's practical applications have so far remained limited to relatively small-scale problems. Predictive coding networks (PCNs), another class of EBMs rooted in computational neuroscience, are typically trained with a specialized algorithm and have likewise not yet been demonstrated at large scale. In this work, we develop a more effective training method for PCNs which combines the centered variant of EP with a novel equilibration scheme for PCNs. Using this approach, we train a 10-layer convolutional PCN (VGG10) on full-size ImageNet, achieving 13.23% test error rate on the top-5 classification task, close to the 12.2% backpropagation baseline. To our knowledge, this is the first demonstration of both PCNs and EP-based training at ImageNet scale. These results significantly extend the scalability of both approaches and suggest that the primary challenges in scaling EP in other EBMs may not be attributed to inherent limitations of the EP framework.

Most theories of in-context learning (ICL) treat the prompt as fixed. This fixed-prompt view misses the key difficulty in memory-updated retrieval systems such as Memory Mosaics. An inference-time write can help later queries, but it can also shift softmax weights and create interference. The question is when the benefit exceeds the later cost. We introduce TRIM (Theory of Retrieval with Incremental Memory), a theorem-level theory of append-only softmax retrieval under a write-independence abstraction. The five-theorem spine links retrieval-weight concentration, local write-induced defect, recurrence-level accumulation, ideal squared-loss benefit, and a benefit-cost comparison with critical-horizon regimes. TRIM places helpful memory and harmful interference within a single frame, so the harmful length scale is derived rather than tuned. We record a first-order dynamic correction for state-dependent writes separately, and it is not part of this spine. We give a theorem-facing empirical protocol. Exact synthetic checks test the spine identities and inequalities. Learned base-MM analyses measure calibrated analogues of concentration, local defect, horizon worsening, and realized benefit. These analyses test whether TRIM observables organize learned systems along the same routes.


Unifying Sparsity and Discreteness: One-Shot Pruning for Quantized LLMs via Discrete Optimization

Haozhen Zhang ⋅ Hanyuan Zheng ⋅ Teng Hou ⋅ Zhaogeng Liu ⋅ Yi Chang ⋅ Bin Gu

Pruning and quantization are two dominant techniques that effectively address the computational and storage burdens of large language model (LLM) inference on edge devices. Recently, one-shot pruning has gained particular attention for its ability to identify weight supports via optimization without retraining. However, integrating such methods seamlessly with quantization for further model lightweighting is non-trivial. The discreteness of quantization shatters their required variable continuity, inevitably collapsing this integration into a decoupled pruning-quantization pipeline with inherently suboptimal outcomes. To tackle this problem, we propose Quantization-aware One-shot Pruning (QOP), which directly optimizes the quantized weights under sparsity and discreteness constraints, thereby explicitly capturing the impact of quantization on pruning within its objective. Specifically, QOP generalizes the alternating direction method of multipliers to sparse-constrained discrete optimization, enabling the identification of the high-quality support and the update of quantized weights. Theoretically, via the well-conditioned approximation obtained by a slight perturbation, we establish convergence to a stationary feasible point and provide the convergence rate. Extensive experiments across various LLMs, sparsity levels, and quantization settings demonstrate that QOP consistently outperforms existing baselines in terms of both average accuracy and perplexity under weight quantization.


Unpaired Canonical Correlation Analysis

Nir Ben-Ari ⋅ Ronen Talmon ⋅ Uri Shaham

Canonical Correlation Analysis (CCA) is a fundamental method for multiview shared space learning. However, its strict reliance on paired data poses a significant limitation, as such data is often difficult to obtain or entirely unavailable. In this paper, we present Unpaired CCA (UCCA), a novel method that learns linear projections to maximize the correlation of the true underlying pairing without access to any paired samples during training. We first establish theoretical results connecting the Quadratic Assignment Problem (QAP) to CCA. Leveraging these theoretical insights, we derive a practical method to maximize correlation exclusively from unpaired data. To the best of our knowledge, UCCA is the first approach to learn maximally correlated projections in a strictly unpaired setting. We validate UCCA on real-world multi-modal datasets, demonstrating that it significantly outperforms recent unpaired alignment baselines in recovering the underlying true correlation. This work fills a critical gap between traditional statistical multiview learning and the growing field of unpaired data learning.


Velocity Ambiguity Profiles: Time-Resolved Bayes-Risk Diagnostics for Flow Matching

Yunpeng Mei ⋅ Xiaowen Zhu ⋅ Chenyu Wang ⋅ Chenbo Xin ⋅ Hongjie Cao ⋅ Jiamin Wang ⋅ Jie Chen ⋅ Gang Wang

Flow matching trains a continuum of velocity regression problems but usually reports a loss averaged over interpolation time, obscuring which regions are limited by intrinsic ambiguity and which remain data-limited. We introduce the Velocity Ambiguity Profile (VAP), $\mathcal{A}_v(t)$, the time-local Bayes-risk floor of velocity prediction, and decompose time-resolved loss as $L_N(t)=\mathcal{A}_v(t)+R(t;N)$, where only the excess term $R(t;N)$ is reducible by more data or a different estimator. We prove two statistical benchmarks. In an intrinsic local-regression benchmark, $\mathcal{A}_v(t)$ enters the time-dependent noise factor of the standard nonparametric rate. In separated isotropic Gaussian mixtures, the Bayes velocity becomes locally affine, and a label-oracle residual-regression benchmark has parametric excess-risk scale $K\mathcal{A}_v(t)/N$, matched within the same separated residual Gaussian-location subproblem. Controlled Gaussian-mixture experiments validate this diagnostic through a dense transition sweep, calibrated oracle-to-blind and empirical-proxy audits, and a local controlled time-sampler stress test in which a frozen reducible-error proxy reduces reducible loss while cross-cell sensitivity delimits aggressive sampling. VAP turns an averaged flow-matching objective into a time-resolved diagnostic for locating ambiguity-limited and data-limited regions.


VeriGraph: Towards Verifiable Data-Analytic Agents

Jiajie Jin ⋅ Zhao Yang ⋅ WenLe Liao ⋅ Yuyang Hu ⋅ Guanting Dong ⋅ Xiaoxi Li ⋅ Yutao Zhu ⋅ Zhicheng Dou

LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a reliance on linear text trajectories makes their reasoning difficult to audit. In particular, deterministic computations over raw data and semantic deductions over natural-language claims are often entangled in an unstructured stream, leaving numerical conclusions hard to reproduce and qualitative judgments hard to inspect. To address this, we propose VeriGraph, a traceable neuro-symbolic reasoning framework that enables agents to construct an explicit heterogeneous evidence directed acyclic graph (DAG) during execution. VeriGraph introduces three evidence-expansion primitives, namely computational, grounding, and derivational expansion, to connect raw data, interpreter variables, computed results, and natural-language claims in a unified graph. Under this formulation, structural traceability is reduced to graph reachability from raw data sources to terminal claims, while semantic support is measured by claim-level evidence evaluation. To improve graph construction, we further design a graph-based policy optimization strategy with a composite reward that jointly supervises answer correctness, computational integrity, and derivational coherence. Experiments on four benchmarks show that VeriGraph-8B achieves the highest overall score among all baselines. More importantly, VeriGraph produces auditable evidence graphs with substantially stronger claim grounding, achieving a 87.61\% Grounding Rate under our claim-level evidence support evaluation. These results suggest that explicit evidence-graph construction is a promising path toward verifiable data-analytic agents. Our code is available at \url{https://anonymous.4open.science/r/VeriGraph-417E}.


ViroGym: Realistic Large-Scale Benchmarks for Evaluating Viral Proteins

Yichen Zhou ⋅ Jonathan L Golob ⋅ Amir Mohammad Karimi Mamaghan ⋅ Stefan Bauer ⋅ Patrick Schwab

Protein language models (pLMs) have shown strong potential for zero-shot prediction of missense variant effects, yet systematic benchmarking on viral proteins remains limited, a critical gap given the need for proactive tools that can anticipate emerging mutations ahead of experimental validation. Here we introduce ViroGym, a comprehensive benchmark evaluating pLMs across three tasks: 79 deep mutational scanning (DMS) assays covering eukaryotic viruses with 552,065 mutated sequences across 7 phenotypic readouts, 21 influenza neutralisation tasks, and a real-world pandemic prediction task for SARS-CoV-2. We benchmark well-established pLMs on fitness landscapes, antigenic diversity, and pandemic forecasting, and find that the ProGen2 family consistently achieves the strongest performance across all three tasks. Crucially, DMS and neutralisation performance reliably identifies models that generalise to real-world emergence, even though the mutation sets they surface barely overlap, revealing that complementary in vitro benchmarks capture the evolutionary constraints needed for real-world mutation forecasting.

Double Oracle (DO) and PSRO approximate Nash equilibria in two-player zero-sum games by iteratively expanding a strategy population with best responses. Online Double Oracle (ODO) interleaves multiplicative weights updates (MWU) with discovery and obtains an $O(\sqrt{k \ln k / T})$ rate. We identify the \emph{reinitialisation} step as an underexplored design choice: at each expansion, MWU-based methods reset newly discovered strategies to a uniform prior. We introduce Virtual Double Oracle (VDO), which exploits a bilinear identity---a strategy's full hindsight loss history collapses to its utility against the time-averaged opponent---to initialise each new strategy with the MWU state it would have had from the start. Conditional on the discovered population, VDO recovers the fixed-population rate $O(\sqrt{\ln k / T})$ asymptotically, removing ODO's leading-order $\sqrt{k}$ window-decomposition penalty. Discounted VDO adds a single parameter $\beta$ to control finite-horizon prior strength. Experiments across matrix games, poker, and Goofspiel validate both contributions: pure VDO confirms the asymptotic rate, and discounted VDO gives the strongest finite-horizon performance.


VolCo: Volumetric Contact for High-Fidelity Human Grasp Generation

Zhuo Chen ⋅ Yihua Cheng ⋅ Ales Leonardis ⋅ Hyung Jin Chang

Accurate contact modeling is fundamental to understanding hand–object interaction, yet existing contact representations are typically restricted to object surfaces and rely on hand‑crafted rules to recover contact details, leading to severe penetrations and implausible results. To better exploit the rich detail in motion‑capture data, we introduce Volumetric Contact (VolCo), a representation that expands surface points to a set of 3D volumetric grids. VolCo encodes 3D contact that allows precise hand part recovery, and is organized in an inherent hierarchy: local contact details within each volume and global hand geometry across all volumes. Our framework, VolCoDiff, employs two modules to capture local and global features following this hierarchy. For local contact details, we use a 3D variational autoencoder to model the possible hand configurations conditioned on the local object signed distance field (SDF). For global hand geometry, we design a prior‑guided diffusion model that learns the distribution of compressed latent features aggregated from the volumetric grids. We evaluate our method on two benchmark datasets and demonstrate state‑of‑the‑art performance in penetration and stability, indicating the capability to generate tight grasps with much less severe penetrations.


WebArena-Pro: A Heterogeneous, Multimodal, Reproducible Benchmark for Web Agents

Imene Kerboua ⋅ Fatemeh Pesaran Zadeh ⋅ Xing Han Lu ⋅ Weijian Qi ⋅ Alexander Miller ⋅ Junyi Song ⋅ Yunjia Tian ⋅ Dongjin Kang ⋅ Seyeon Choi ⋅ Marzia Nouri ⋅ Ewen Gueguen ⋅ Matteo Boglioni ⋅ Fengyuan Liu ⋅ Zeyi Liao ⋅ Mengqi Yuan ⋅ Yue Li ⋅ Alexandre Lacoste ⋅ Alexandre Drouin ⋅ Spandana Gella ⋅ Huan Sun ⋅ Gunhee Kim ⋅ Siva Reddy

Web agents powered by large language and vision-language models are increasingly applied to realistic browser work that spans heterogeneous applications, multimodal content, and stateful workflows. However, existing reproducible web-agent benchmarks cover only a small number of web applications drawn from a few software categories, and restrict modality to text and vision. Live benchmarks broaden site coverage but sacrifice reproducibility, since pages and data drift between runs. Moreover, existing benchmarks do not meaningfully evaluate whether agents can understand and use audio and video content embedded within web tasks. To address these gaps, we introduce WebArena-Pro, a benchmark comprising 300 tasks across 20 self-hosted web applications in six domain categories, spanning distinct interface conventions, workflows, and data models. Across the evaluated agents, the best performance is achieved by Gemini 3.1 Pro, which attains 37.0 \% success under a 50-step budget, while open-source models' performance does not exceed 27.7\% success. Among reproducible, human-curated web agent benchmarks, WebArena-Pro provides the broadest application coverage and the most comprehensive multimodal support to date. The benchmark treats audio and video as core observations alongside text and vision, with dedicated actions for extracting information from each. WebArena-Pro runs each task in isolation and supports reproducible, parallel evaluation. Tasks are authored through a dedicated annotator interface, filtered by LLM-assisted triage, and finally validated by humans before release.


We Need to Improve Benchmarks in AI for Mathematics

Simon Frieder ⋅ Jonas Bayer ⋅ Shi Zhuo Looi ⋅ Jacob Loader ⋅ Julius Berner ⋅ Katie Collins ⋅ Andras Juhasz ⋅ Fabian Ruehle ⋅ Sean Welleck ⋅ Gabriel Poesia ⋅ Ionut Mistreanu ⋅ Ryan-Rhys Griffiths ⋅ Adrian Weller ⋅ Anirudh Goyal ⋅ Thomas Lukasiewicz ⋅ Cameron Freer ⋅ Kevin Buzzard ⋅ W. T Gowers

Benchmarks used to evaluate AI systems for mathematics (both in natural and formal language) exhibit critical shortcomings that limit progress toward genuinely useful mathematical assistants. These limitations range from restricted mathematical complexity to insufficient fidelity in capturing aspects of formal languages such as Lean. Compounding some of these issues is a dynamic reminiscent of Goodhart's law: as benchmark performance becomes the primary optimization target, benchmarks become less reliable indicators of mathematical capability. We explore these limitations and argue that progress requires a course correction in benchmark design and an upgrade to evaluation standards. Additionally, we provide a live, community-extensible website to track the landscape of mathematics benchmarks and encourage the creation of novel benchmarks.


We Need to Rethink Benchmarking in Anomaly Detection

Philipp Röchner ⋅ Simon Klüttermann ⋅ Kevin Kammler ⋅ Franz Rothlauf ⋅ Emmanuel Müller ⋅ Daniel Schlör

Despite the continuous proposal of new anomaly detection algorithms and extensive benchmarking efforts, progress seems to stagnate, with only minor performance differences between established baselines and new algorithms. In this position paper, we argue that this stagnation is due to limitations in how we evaluate anomaly detection algorithms. In current benchmarks, a trivial algorithm that only checks for extreme values in individual features performs competitively with state-of-the-art deep learning methods, despite failing on simple cases such as anomalies within an annulus of normal points. Moreover, existing benchmarks do not adequately reflect the diversity of anomaly detection applications, making it difficult for practitioners to reliably select algorithms for their applications. Consequently, we need to rethink benchmarking in anomaly detection. In our opinion, anomaly detection should be studied using scenarios that group applications sharing relevant characteristics, defined through a common taxonomy. Benchmarking within scenarios enables scenario-specific choices for preprocessing, metrics, and model selection, clarifying which advances transfer across similar applications and providing practitioners with reliable guidance for their specific contexts.

Compositional inference - the decomposition of observations into an unknown number of latent components - is central to perception and scientific data analysis. Attention-based models perform well when components are approximately separable, as in object-centric vision. Under additive superposition, however - where multiple components contribute to every observation - we identify a structural failure mode we term slot collapse: multiple slots converge to the same dominant component while weaker ones remain unrepresented. We trace this to a general limitation: attention is memoryless with respect to explained evidence. All slots repeatedly operate on the same input without accounting for what has already been explained, so gradients are dominated by the strongest component, inducing shared fixed points across slots. As a result, attention fails to enforce non-redundant allocation under additive superposition. We address this by introducing residual evidence modeling, which tracks the remaining explanatory capacity of the input, instantiated via evidence depletion - a minimal modification combining multiplicative depletion with an attention bias. Controlled ablations show that parallel attention, sequential processing alone, and loss-based regularization all fail to resolve collapse; evidence depletion, which adds stateful residual tracking to sequential attention, consistently succeeds. Across synthetic benchmarks and real-world audio mixtures (FUSS), evidence depletion reduces slot collapse by up to an order of magnitude, generalizing beyond synthetic settings. On gravitational-wave source inference for the ESA/NASA LISA mission, under identical architectures, data, and losses, standard attention fails while evidence depletion prevents collapse and enables multi-source posterior estimation. These results show that under additive superposition, residual evidence tracking is the operative ingredient - preventing collapse and enabling compositional inference.


Why Do Time Series Models Need Long Context Windows?

Luca Butera ⋅ Giovanni De Felice ⋅ Andrea Cini ⋅ Cesare Alippi

The effectiveness of modern deep learning models in forecasting *groups of time series* is often attributed to their ability to capture (long-range) dependencies across long observation windows. In this paper, we show that this forecasting task involves two objectives: (i) *generative process identification* (GPI), i.e., inferring the specific process generating the input sequence, and (ii) *conditional forecasting* (CF), i.e., predicting future values given input observations. From this perspective, optimal predictions can be interpreted as an average over plausible data-generating processes, weighted by their likelihood given the input window. This suggests a different explanation for the benefits of long context windows: they reduce the uncertainty about which specific process is generating the input time series during operation. We prove that even for processes with memory length $P$, an input window size strictly larger than $P$ is *necessary* to achieve the minimum attainable error. Finally, we show how decoupling GPI and CF can improve computational scalability without compromising accuracy. Experiments on synthetic and real-world data validate our insights and their relevance for designing forecasting architectures.


Within-Model vs Between-Prompt Variability in Large Language Models for Creative Tasks

Jennifer Haase ⋅ Jana Gonnermann-Müller ⋅ Paul H. P. Hanel ⋅ Nicolas Leins ⋅ Thomas Kosch ⋅ Jan Mendling ⋅ Sebastian Pokutta

When a large language model (LLM) generates a creative response, how much of the outcome is determined by the prompt, the model, the task, or pure sampling luck? We address this question with a variance decomposition over 89,806 generations from 10 LLMs, 6 prompt strategies, and 30 Alternative Uses Task items, which is one of the most widely used creativity tests. For output quality (originality), the results invert common assumptions: the AUT item being evaluated and the prompt strategy together account for over 72% of variance, while model choice contributes less than 8% and within-LLM stochasticity alone exceeds model differences by more than 2×. Controlling for the AUT item sharpens the picture further: prompt strategy then explains 3.4× the variance of model choice. For output quantity (fluency), the pattern reverses, with the LLM being the dominant factor. We further show that prompts shape output distributions, not point estimates: “discriminative” prompts (e.g., Persona) widen the quality distribution and show higher peaks in capable models, while “constraining” prompts (e.g., Format) compress it. Together, these results argue that single-sample LLM evaluations are unreliable for open-ended tasks, and that evaluation designs must account for the full variance structure of generative systems.

A widely held assumption in knowledge distillation is that teacher and student representations must occupy a compatible or explicitly aligned feature space. We propose RPKD (Random Prototype Knowledge Distillation), which discards this assumption by projecting both networks into a shared set of randomly initialized, frozen prototype vectors, requiring no architecture-specific adapters, no class-specific alignment, and minimal architectural assumptions. Two branches handle the transfer: a logit branch aligning global prototype similarity distributions and a feature branch matching spatially-resolved prototype activation maps. Both operate within the same fixed prototype space, which can be interpreted as a random feature embedding that approximately preserves inter-sample similarity structure, keeping RPKD agnostic to the internal dimensions and inductive biases of either network. In low-category settings, decoupling the prototype vocabulary from the task label space yields richer supervision than the label space alone, an advantage absent from logit-based methods, whose supervisory signal collapses as class count shrinks. This advantage is especially relevant in real-world scenarios where the label space is typically constrained. Extensive experiments demonstrate the effectiveness of RPKD, achieving a maximum gain of $\mathbf{+8.26\%}$ over OFA on CIFAR-100, $\mathbf{+12.46\%}$ on ImageNet-100, and $\mathbf{+1.34\%}$ on ImageNet-1K. Source code will be released upon acceptance. Source code will be released upon acceptance.


Your Neighbors Know: Leveraging Local Neighborhoods for Backdoor Detection in Decentralized Learning

Sayan Biswas ⋅ Antoine Boutet ⋅ Davide Frey ⋅ Romaric Gaudel ⋅ Rachid Guerraoui ⋅ Maxime Jacovella ⋅ Anne-marie Kermarrec ⋅ Dimitri Lerévérend ⋅ Francois Taiani ⋅ Martijn De Vos

Decentralized learning (DL) is an emerging machine learning paradigm where nodes collaboratively train models without a central server. However, the collaborative nature of DL makes it vulnerable to backdoor attacks, where a model is taught to behave normally on standard inputs while executing hidden, malicious actions when encountering data with specific triggers. Backdoor attacks in DL remain understudied and existing defenses often overlook DL constraints. We introduce Argus, a novel backdoor detection framework native to DL that requires neither a central coordinator nor prior knowledge of the trigger. In Argus, honest nodes locally analyze received model updates to identify potential backdoor triggers. Nodes then collectively share their triggers with their neighbors and use a structural similarity metric to separate true backdoors from false alarms induced by data heterogeneity. A key insight is that false positive triggers exhibit inconsistencies across participants while true positive ones show consistent patterns. Model updates that fail this collaborative test are rejected, and persistently malicious senders are eventually evicted. We provide the first theoretical convergence guarantees for a DL-specific backdoor detection mechanism, showing that filtering out suspicious model updates with high probability preserves a convergence rate comparable to standard DL. We implement and evaluate Argus on three standard datasets and against three state-of-the-art baselines. Across settings, Argus reduces attack success rates by up to 90 points compared to no defense, while preserving model utility within 5 percentage points of an omniscient oracle. Furthermore, the effectiveness of Argus compared to baselines improves as data heterogeneity increases.