Skip to yearly menu bar Skip to main content


Session

Atlanta Poster Session 1

Hall C1
Thu 10 Dec 2 a.m. AEDT — 5 a.m. AEDT
Abstract:
Chat is not available.

We study the fixed-budget max–min action identification problem in depth-2 max–min trees, which is an important special case of Monte Carlo Tree Search (MCTS). In this problem, a learner sequentially and adaptively selects $T$ reward samples from leaves (depth-2 nodes) and then must recommend a subtree (depth-1 node) with the largest min-value among its children nodes. Motivated by approximate planning, we focus on \emph{$\varepsilon$-good subtree identification}, where any subtree whose min value is within $\varepsilon$ of optimal maximin value is acceptable. Our main contribution is an \emph{$\varepsilon$-agnostic} algorithm—requiring no knowledge of $\varepsilon$—that nevertheless achieves error bounds with explicit instance-dependent dependence on $\varepsilon$. We show that, for every meaningful $\varepsilon$, the misidentification probability decays as $\exp\big(-\tilde\Theta(\frac{T}{H_2(\varepsilon)} )\big)$, where $H_2(\varepsilon)$ captures both cross-subtree and within-subtree gaps. In the special case where each subtree has a single leaf, the model reduces to standard multi-armed bandit identification, and our bounds recover (up to accelerating factors) the best-known $\varepsilon$-good guarantees associated with halving-style methods, while providing a new $\varepsilon$-good analysis and guarantee for the Successive Rejects algorithm in fixed budget best arm identification. On the lower-bound side, we present complementary positive and negative results. While there is a gap between the upper and lower bounds, we discuss the main technical challenges in obtaining a tighter lower bound, which tells us that the maximin action identification problem is quite different from the standard $K$-armed bandits. To our knowledge, this is the first provable algorithmic guarantee for fixed-budget maximin action identification.


$f$-GRPO & Beyond: Divergence-Based Reinforcement Learning Algorithms for General LLM Alignment

Rajdeep Haldar ⋅ Lantao Mei ⋅ Guang Lin ⋅ Yue XING ⋅ Qifan Song

Recent work shows that preference alignment objectives can be interpreted as divergence estimators between aligned (preferred) \& unaligned (less-preferred) distributions, yielding a principled recipe for designing alignment losses. However, this view has so far been limited to preference-based supervision. We extend it to general LLM alignment, including reinforcement learning with verifiable rewards (RLVR), where alignment feedback is given only as scalar rewards. We introduce $f$-Group Relative Policy Optimization ($f$-GRPO), a class of on-policy RL objectives, and $f$-Hybrid Alignment Loss ($f$-HAL), which combines on-policy reward optimization with off-policy preference supervision. We show that these objectives estimate $f$-divergences between reward-aligned \& reward-unaligned distributions induced by above- \& below-average reward responses, and prove expected reward improvement after alignment. Empirically, $f$-GRPO improves over GRPO on math-reasoning RLVR tasks, while hybrid $f$-HAL mitigates reward hacking in on-policy safety alignment when verifiable rewards are unavailable and learned reward models must be used.


3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

Yipeng Gao ⋅ Lei Shu ⋅ Genzhi Ye ⋅ Xi Xiong ⋅ Ameesh Makadia ⋅ Meiqi Guo ⋅ Laurent Itti ⋅ Jindong Chen

Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neural 3D generators inherently lack. Authoring such procedural content, however, demands deep expertise in 3D software APIs, parametric design, and code-level geometric reasoning. In this paper, we propose 3DCodeBench, a systematic benchmark for evaluating vision-language model (VLM) agents for procedural 3D generation in 3D modeling software. Specifically, 3DCodeBench evaluates how effectively 11 advanced VLMs can serve as procedural 3D modelers by transferring text and image references into procedural code for software. Recognizing that automated metrics may not fully capture the perceptual quality of 3D shapes, we build 3DCodeArena, a ranking platform based on human preference between generated pairs. From extensive evaluations and results, we observe that: (1) Failures mostly arise from API usage, while successful renders still suffer from disconnected or floating 3D geometric components. (2) Test-time scaling, such as higher thinking budgets and multi-turn refinement, improves performance overall. Our findings highlight a critical need for high-quality procedural coding data to advance commercial VLMs. Furthermore, effective procedural 3D modeling requires a robust execution environment that provides high-fidelity feedback for iterative refinement. To accelerate research in this domain, we release 3DCodeBench—including the curated data, benchmark, evaluation protocol, and qualitative results—as a foundational tool for developing VLM-based 3D modelers.

AI video generation is evolving rapidly. For video generators to be useful for applications ranging from robotics to film-making, they must consistently produce realistic videos. However, evaluating the realism of generated videos remains a largely manual process -- requiring human annotation or bespoke evaluation datasets which have restricted scope. Here we develop an automated evaluation framework for video realism which captures both semantics and coherent 3D structure and which does not require access to a reference video. Our method, 3DSPA, is a 3D Semantic Point Autoencoder utoencoder which is trained to reconstruct held out 3D point trajectories which have been augmented with DINO semantic features. By combining both semantic and geometric representations, 3DSPA enables robust assessments of realism, temporal consistency, and physical plausibility in generated videos. Experiments show that 3DSPA outperforms leading VLMs by more than 10\% in identifying videos which violate physical laws, and more than doubles the correlations between human and auto-eval ratings in judging physical common sense and motion quality across multiple generative video datasets. Our results demonstrate that enriching trajectory-based representations with 3D semantics offers a strong foundation for benchmarking generative video models, and implicitly captures physical rule violations.


A$^2$IQL: Adaptive Asymmetric Implicit Q Learning for Automated Warehouse Consolidation

Guangyi Liu ⋅ Andrea Angiuli ⋅ Mirko Ristivojevic ⋅ Joseph W Durham ⋅ Michael Caldara ⋅ Michael Zavlanos

We apply offline reinforcement learning to large-scale warehouse consolidation, an application that requires ranking tens of thousands of candidates at each decision epoch with an architecture that scales gracefully with variable-cardinality action spaces. Since the available reward signal is only an approximation of the true target-domain objective, the policy must also transfer robustly. We propose a two-stage framework, Asymmetric Implicit Q-Learning (AIQL) and its adaptive cross-domain extension A$^2$IQL. First, AIQL learns a scoring policy from fixed offline consolidation logs by decoupling a per-candidate scoring actor from an aggregate-statistic critic, keeping value estimation tractable and within the support of logged data. Second, A$^2$IQL adapts the source-domain-trained policy to a new target objective using scarce target-domain data: a Bayesian reward predictor supplies an uncertainty-penalized conservative bonus that modifies only the advantage-weighted actor extraction, leaving the source critic unchanged. On a single-domain benchmark, AIQL improves throughput by 20% over a heuristic baseline using offline data alone. When adapted to a new target objective via A$^2$IQL, the method outperforms source-only training, data pooling, and off-dynamics baselines by up to 23% under a $30{:}1$ source-to-target data ratio while achieving the lowest variance across all methods.


ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space

Gabe Guo ⋅ Thanawat Sornwanee ⋅ Lutong Hao ⋅ Elon Litman ⋅ Stefano Ermon ⋅ Jose Blanchet

Generative modeling continuous-time, continuous-space stochastic processes (e.g., videos, weather forecasts) conditioned on partial observations (e.g., first and last frames) is a fundamental challenge. Existing approaches, (e.g., diffusion models), suffer from key limitations: (1) noise-to-data evolution fails to capture structural similarity between states close in physical time and has unstable integration in low-step regimes; (2) random noise injected is insensitive to the physical process's time elapsed, resulting in incorrect dynamics; (3) they overlook conditioning on arbitrary subsets of states (e.g., irregularly sampled timesteps, future observations). We propose ABC: Any-Subset Autoregressive Models via Non-Markovian Diffusion Bridges in Continuous Time and Space. Crucially, we model the process with one continual SDE whose time variable and intermediate states track the real time and process states. This has provable advantages: (1) the starting point for generating future states is the already-close previous state, rather than uninformative noise; (2) random noise injection scales with physical time elapsed, encouraging physically plausible dynamics with similar time-adjacent states. We derive SDE dynamics via changes-of-measure on path space, yielding another advantage: (3) path-dependent conditioning on arbitrary subsets of the state history and/or future. To learn these dynamics, we derive a path- and time-dependent extension of denoising score matching. Our experiments show ABC's superiority to competing methods on multiple domains, including video generation and weather forecasting.


Accelerated last-iterate convergence of Extragradient via power-law stepsizes

Yue Wu ⋅ Weiqiang Zheng ⋅ Yang Cai ⋅ Haipeng Luo

We revisit the convergence guarantees of the Extragradient (EG) method for unconstrained bilinear min-max optimization. It is known that EG with a fixed stepsize achieves a $\Theta(T^{-1/2})$ last-iterate convergence rate, which is slower than the optimal $\mathcal{O}(T^{-1})$ rate attainable by incorporating additional mechanisms such as anchoring. Motivated by recent advances showing that dynamic stepsizes alone can significantly accelerate gradient descent, we ask whether dynamic stepsizes can similarly accelerate the last-iterate convergence of EG. We present the first positive result in this direction. Specifically, we provide a deterministic dynamic stepsize schedule that accelerates the convergence rate of EG to $\mathcal{O}(T^{-2/3+\varepsilon})$ for any $\varepsilon > 0$. We also show that this rate is tight when the extrapolation and update steps of EG use the same stepsize. We then show that allowing different stepsizes for the extrapolation and update steps further improves the convergence rate to the near-optimal $\mathcal{O}(T^{-1+\varepsilon})$. Our analysis reduces stepsize scheduling to an optimization problem, whose solution leads to a stepsize schedule that follows (a discretization of) a power-law distribution. Our proposed stepsize schedules and analysis extend to other methods, such as optimistic gradient descent, and suggest broader applicability to general min-max optimization problems.

Agents built on large language models (LLMs) rely on a range of reliability techniques, including retry, majority voting, and self-consistency, that have been developed in parallel rather than within a common analytical framework. We observe that an LLM sampled at temperature $T$ is a discrete stochastic channel $p(y \mid x)$ in the sense of Shannon's coding theory, and use this identity as the entry point for such a framework grounded in communication theory. Each of these techniques is a special case of one of six classical reliability operators: diversity combining, hybrid retransmission, iterative generator-critic decoding, rateless sampling, structured redundant verification, and difficulty-adaptive routing. Within the framework we give two closed-form results: a noise-variance threshold above which uniform averaging beats quality-weighted averaging, and a contractivity criterion for generator-critic refinement, consistent with a contractive-to-divergent transition we observe between 3B- and 14B-parameter models. We further introduce a cost-aware semantic-nearest-neighbor router whose single Lagrangian knob traverses the quality-cost frontier without retraining. Across six channel configurations spanning local and cloud models on 69 hard tasks, no fixed model-technique-budget choice dominates, motivating per-task allocation. On a 300-item hard split of MMLU, GSM8K, and HumanEval, our router occupies the full empirical Pareto frontier: at matched quality, its normalized cost is ${\approx}56$\% lower than the strongest fixed technique; at matched normalized cost, it improves quality by ${\approx}7$\% ($26$\% over single-shot decoding). These results argue for consolidating these reliability techniques into a single tunable layer informed by channel coding.

Post-training quantization is a standard tool for the efficient deployment of large language models (LLMs). In safety-aligned models, it introduces a critical and underexplored risk: preserving utility under compression does not ensure that safety alignment is retained. Small perturbations from quantization can disproportionately degrade refusal to harmful prompts, leading to increased compliance with harmful prompts or over-refusal of benign prompts, even when aggregate performance appears stable. This gap reveals a fundamental limitation of existing quantization approaches, which primarily optimize for utility while overlooking failures due to safety alignment degradation. This paper presents ACQueReLlo, a reinforcement learning framework for mixed-precision post-training quantization. It formulates block-wise bit allocation as a constrained Markov decision process and learns a quantization policy under explicit constraints on utility, harmful prompt refusal, benign prompt over-refusal, and overall bit budget. To ensure tractable training, the policy relies on proxy evaluations over sampled subsets of utility, harmful, and benign prompts. Experiments on various LLMs show that the learned policy achieves a better overall balance between compression, utility, and safety alignment than uniform quantization and other baselines.


Adaptive Inference for Functionals of M-Estimands

James Leiner ⋅ Aurelien Bibaut ⋅ Nathan Kallus ⋅ Aaditya Ramdas ⋅ Koulik Khamaru

Reinforcement learning and contextual bandit algorithms have become increasingly common in sequential decision-making applications. When these methods are deployed in high-stakes domains, there is growing interest not only in learning effective policies, but also in conducting statistical inference for quantities learned under adaptive data collection. However, classical procedures applied naively in these settings can fail: even when estimators are unbiased, their variance becomes path-dependent and as a result may not be asymptotically normal. A growing literature has emerged to ameliorate this problem, but solutions tend to be problem specific and often rely on correct specification of a working model. In this work, we develop a unified framework for constructing asymptotically valid confidence intervals to cover smooth functionals of nonparametric M-estimands under adaptive sampling. Under Neyman orthogonality, we provide two novel methods for performing inference: (1) a self-normalized statistic based on the realized quadratic variation of the influence function and (2) a statistic using a plug-in estimate of the conditional variance based on reweighted influence function increments. Our results allow for flexible nonparametric estimation of nuisance parameters and remain valid under model misspecification. Our theory is supported by a simulation study for a dynamic pricing application which demonstrates that this method can produce asymptotically valid confidence intervals where standard methods fail.


AdaST: Adaptive Coupling for Spatial-Temporal Forecasting

Zhenyu Lei ⋅ Chenghao Liu ⋅ Yushun Dong ⋅ Qi Wang ⋅ Jundong Li

Spatial-temporal (ST) forecasting underpins many real-world systems such as traffic, climate, and energy networks. While existing methods implicitly assume strong spatiotemporal coupling, we observe that real-world ST data exhibits distinct coupling regimes, ranging from temporal-dominated and spatial-dominated to strongly coupled patterns. This mismatch causes current models to suffer from spurious dependencies and degraded performance when one correlation dominates. To overcome this limitation, we aim to dynamically modulate spatial and temporal modeling based on the data's inherent coupling structure. However, three key challenges exist: unknown coupling structure, heterogeneous coupling dynamics, and suboptimal spatial modeling. We propose AdaST, an adaptive ST forecasting framework that tackles these challenges through a decompose-recompose paradigm. AdaST factorizes inputs into components capturing different coupling patterns using heterogeneity-aware experts. Each component is processed by role-aligned modules, and a correlation-informed adaptive recomposer integrates them for final prediction. Extensive experiments confirm that AdaST significantly outperforms state-of-the-art baselines, validating the necessity of an adaptive approach.

Kernel methods, and Gaussian Processes (GPs) in particular, require a conditionally negative-definite (CND) distance measure to guarantee positive semi-definiteness (PSD) of the kernel matrix — a condition that fails for many natural input spaces, including smooth manifolds and spaces of probability distributions. We propose the Sparse Landmark Embedding (SLE) kernel, which eliminates this requirement entirely. Each input is embedded into a sparse feature vector via compactly supported bump functions centered at all training points $|\mathcal{D}|$; applying any standard PSD kernel in this embedding space yields a kernel that is provably PSD for arbitrary distance measures. The compact support automatically controls embedding sparsity, keeping kernel matrices well-conditioned and computationally tractable despite the high ambient dimension. We provide theoretical guarantees on PSD, sparsity, stability, and universal approximation, and demonstrate, using geodesic and Wasserstein distances, that the SLE kernel matches or substantially exceeds domain-specific baselines in both predictive accuracy and uncertainty quantification.


AgentKVShift: Efficient KV Cache Reuse for Agentic Memory Systems

Nilesh Pandey ⋅ Jason Kong ⋅ Lanxiang Hu ⋅ Quanling Zhao ⋅ Yujie Zhao ⋅ Onat Gungor ⋅ Hao Zhang ⋅ Tajana S Rosing

Memory-augmented LLM agents have gained substantial attention recently for their ability to maintain context across hundreds of interactions through agentic memory systems that actively curate retrieved content with LLM-generated metadata such as summaries, keywords, and tags. However, from an inference cost standpoint, every retrieval triggers a full re-encoding of these structured memory units into Key-Value (KV) states, which dominates prefill latency. Existing training-free KV reuse methods mitigate this by selectively recomputing a small fraction of tokens, but were designed for RAG-style raw passages and degrade substantially on structured agentic memories. In this work, we present AgentKVShift, a training-free, probe-guided KV residual correction method that operates per retrieved memory unit. One of the crucial insights we demonstrate is that the per-memory KV reuse residual decomposes into a shared memory-level offset plus small token-wise fluctuations. Estimating this offset from a small probe set allows us to correct every reused token in the memory unit by a single weighted correction. Unlike prior reuse methods which decide which tokens to recompute and leave the rest of the cache stale, AgentKVShift also corrects the tokens it does not recompute, turning the refresh budget into useful signal across the entire chunk. Through extensive experiments across four open source LLMs spanning 3B to 32B parameters and two long-horizon agentic memory benchmarks covering long-term dialogue and agentic applications, we show that AgentKVShift achieves near full recompute performance while refreshing only 10-30% of the cache, outperforming existing baselines at the same recompute ratio. AgentKVShift requires up to 5x lower recompute to reach this near-full performance, which prior reuse methods only attain at 45-55% refresh. By operating in this lower recompute regime, AgentKVShift delivers prefill speedups of 2-3.5x over no-KV-reuse on a single A100 GPU. Lastly, AgentKVShift orthogonally composes with KV cache quantization, retaining over 2x the F1 of prior reuse methods under aggressive 2- and 4-bit settings, making it a simple, yet effective choice for serving long-horizon agentic memory workloads.


Align and Distill: Unifying and Improving Domain Adaptive Object Detection

Justin Kay ⋅ Timm Haucke ⋅ Suzanne Stathatos ⋅ Siqi Deng ⋅ Erik Young ⋅ Pietro Perona ⋅ Sara Beery ⋅ Grant Van Horn

Object detectors often perform poorly on data that differs from their training set. Domain adaptive object detection (DAOD) methods have recently demonstrated strong results on addressing this challenge. Unfortunately, we identify systemic benchmarking pitfalls that call past results into question and hamper further progress: (a) Overestimation of performance due to underpowered baselines, (b) Inconsistent implementation practices preventing transparent comparisons of methods, and (c) Lack of generality due to outdated backbones and lack of diversity in benchmarks. We address these problems by introducing: (1) A unified benchmarking and implementation framework, Align and Distill (ALDI), enabling comparison of DAOD methods and supporting future development, (2) A fair and modern training and evaluation protocol for DAOD that addresses benchmarking pitfalls, (3) A new DAOD benchmark dataset, CFC-DAOD, increasing the diversity of available DAOD benchmarks, and (4) A new method, ALDI++, that achieves state-of-the-art results by a large margin. ALDI++ outperforms the previous state-of-the-art by +3.5 AP50 on Cityscapes to Foggy Cityscapes, +5.7 AP50 on Sim10k to Cityscapes (where ours is the only method to outperform a fair baseline), and +0.6 AP50 on CFC-DAOD. ALDI and ALDI++ are architecture-agnostic, setting a new state-of-the-art for YOLO and DETR-based DAOD as well without additional hyperparameter tuning. Our framework, dataset, and method offer a critical reset for DAOD and provide a strong foundation for future research.


Aligning Language Models with Selective Prediction

Gaoxiang Luo ⋅ Yifan Wu ⋅ Sinian Zhang ⋅ Aryan Deshwal ⋅ Ju Sun

Large language models (LLMs) are increasingly deployed as decision making components in real-world systems for societal and scientific applications, creating a growing need for reliable predictions. In this paper, we study the problem of reliable decision making with LLMs via the lens of selective prediction, allowing the model to improve performance by trading-off coverage. The aim of selective prediction is to balance the key tradeoff of risk and coverage, where risk measures the predictive performance on selected inputs and coverage measures the fraction of abstentions. While existing LLM post-training approaches focus primarily on correctness or calibration, we propose to directly optimize for selective prediction performance by introducing reinforcement Learning for Selection Reward (RLSR), which targets the area under the risk-coverage curve (AURC) as its training objective. RLSR achieves substantially better risk-coverage tradeoff compared to multiple baselines on both in-domain and out-of-domain tasks.


Alignment Is Not Enough for Safe Medical LLM Evaluation

Zhiwen You ⋅ Yikun Han ⋅ Jana Diesner ⋅ Yue Guo

LLMs-as-judges (LLJs) are increasingly used to reduce the cost of expert review in medical evaluation. However, high alignment with clinicians does not by itself establish validity or safety. This position paper argues that current reporting practices for medical LLJs are insufficient: LLJ-clinician alignment should not be reported as standalone evidence that a judge is reliable, clinically grounded, or safe for downstream use. We identify four risks: unstable evaluations across prompts, runs, and judge models; opaque scores that hide clinically meaningful errors; weak validity under clinical uncertainty and limited clinician agreement; and downstream safety or data contamination risks when flawed scores influence model selection or synthetic data filtering. By re-examining LLJ evaluation setups used in prior medical studies, we show that judges treated as practical evaluators in the literature can still reject clinically equivalent diagnoses, overrate incomplete or unsafe answers, reward unsupported clinical notes, and penalize appropriate treatment plans. We propose a protocol for medical LLJ evaluation that requires task-specific codebooks, expert-annotated validation sets, codebook-derived prompts, robustness testing, clinical error analysis, uncertainty reporting, and transparent disclosure before LLJ scores are used for medical benchmarking or decision support.


Alleviating Hallucination with Training-Free Uncertainty-Guided Steering

Litian Liu ⋅ Yubing Jian ⋅ Qiqi Hou ⋅ Reza Pourreza ⋅ Mohammad Ghavamzadeh ⋅ Yao Qin ⋅ Roland Memisevic ⋅ Herbert Cai

Recent work on hallucination detection in large language models has shown that, for a fixed pre-trained model and reasoning task, it is possible to estimate the model’s confidence in the correctness of its outputs. Such uncertainty estimates have primarily been used to improve truthfulness by detecting or filtering confabulations. In this work, we ask whether these signals can instead be used more proactively to directly improve the accuracy of model-generated answers. We propose USteer, a simple, training-free steering mechanism that adjusts a model’s layer-wise activations during inference using the gradient of a confidence with respect to the activations. This procedure nudges generation toward outputs with lower uncertainty at inference time, without modifying model parameters or requiring additional supervision. We show that this approach consistently reduces hallucination across a range of tasks, demonstrating that confidence signals can be leveraged not only for detection, but also for effective inference-time control of model behavior.

Block-wise post-training quantization (PTQ) compresses LLMs by minimizing per-block reconstruction error on a small calibration set. We show that the \emph{composition} of this set (200 samples) is a targeted lever at extreme bit widths, acting only on iterative methods. In a controlled study spanning six data-dependent PTQ methods plus a data-free baseline, six models (1.7B--8B) across three families plus 70B verification, three bit widths, and a 10-task evaluation with four-seed replication, we find a sharp \textbf{near-cliff method-class divide}. In each method's pre-cliff regime (where the model retains predictive function), iterative methods (SignRound~V2, OmniQuant, AQLM) gain 0.74--1.93pp on the 10-task average from switching Pile to a fragility-guided dataset (FragCal-DC), while single-pass analytical methods (GPTQ, AWQ) show near-zero response ($|\Delta|{\le}0.30$pp). Both classes are evaluated on functional Pile baselines (0.42--0.61, well above chance); above the cliff, every method becomes data-composition-insensitive. A reconstruction-loss paradox (the method that most reduces its proxy loss gains nothing downstream, while the method that barely moves the proxy gains the most) sharpens the mechanism; fixed-reference held-out evaluation confirms the proxy--downstream disconnect. A wrong-answer contamination ablation retains 87--91\% of the improvement and a data-dependent single-pass control (RTN-opt) shows zero response, isolating the iterative optimization loop. A weight-space probe shows SignRound-V2 preserves higher per-group scale heterogeneity than GPTQ on 9 of 9 (model, width) cells; a scale-flatten knockout shows this heterogeneity is necessary for iterative fidelity. We release FragCal-DC, the discriminative-task instance used here, as a drop-in JSON for the evaluated setting.

Modeling high-dimensional dependencies while keeping likelihoods tractable remains challenging. Classical vine-copula pipelines are interpretable but can be expensive, while many neural estimators are flexible but less structured. In this work, we propose Vine Denoising Copula (VDC), an amortized vine-copula pipeline for continuous-data, simplified-vine dependence modeling. VDC trains a single bivariate denoising model and reuses it across all vine edges. For each edge, given pseudo-observations, the model predicts a piecewise-constant density grid. We then apply an IPFP/Sinkhorn projection that normalizes mass and drives the marginals to uniformity. This preserves the tractable vine-likelihood structure and the usual copula interpretation while replacing repeated per-edge optimization with GPU inference. Across synthetic and real-data benchmarks, VDC delivers strong bivariate density accuracy, competitive MI/TC estimation, and faster high-dimensional vine fitting. These gains make explicit information estimation and dependence decomposition feasible when repeated vine fitting would otherwise be costly, while conditional downstream tasks remain a limitation.


Amplitude Decoupling in Gaussian Process Training: Exact Decomposition, Pole Cancellation, and Evaluation-Efficient Optimization

Piyush Sao ⋅ Keita Teranishi ⋅ Sudip K Seal ⋅ Pedro Valero-Lara ⋅ Narasinga R Miniskar

Gaussian process hyperparameter optimization via marginal likelihood is notoriously brittle. We identify a precise, removable cause of a common source of line-search inefficiency in unprofiled exact GP training: mismatch between the current scalar log-amplitude and its profiled optimum. The GP negative log-likelihood decomposes exactly as $F(a, x) = f(x) + \frac{n}{2}(e^{-r} + r - 1)$, where $r = a - a^*(x)$ is the mismatch between the current log-amplitude $a = \log \sigma_f^2$ and its analytical optimum $a^*(x)$ for the shape parameters $x = (\log \ell, \log \tau)$. The residual $\Psi_n(r) = \frac{n}{2}(e^{-r} + r - 1)$ is a convex, nonnegative penalty that (i) introduces pole singularities at analytic continuations of the kernel matrix past its singular boundary, and (ii) classifies 80% of rejected Armijo trials overall, and up to 100% in high-SNR regimes, as amplitude-caused in our L-BFGS experiments. Profiling out $\sigma_f^2$ eliminates this term exactly, reducing poles to logarithmic branch points and converting stiff line profiles into nearly flat ones. We introduce an Armijo failure diagnostic that exactly decomposes each rejected unprofiled trial into profiled-shape and amplitude-mismatch contributions, determining whether a particular rejection was caused by scalar amplitude mismatch or by the profiled shape objective. Experiments on synthetic regression tasks show that, in the tested L-BFGS/Armijo settings, amplitude profiling reduces the number of objective evaluations needed to reach a target marginal likelihood by 2-5x, with no change to the per-evaluation cost or the global optimum.


Anchoring LLM-based Chest X-ray Report Generation via Diffusion Language Planning

Jiechao Gao ⋅ Chang Liu ⋅ Yuandong Pan ⋅ Ying Liu ⋅ Michael Lepech

While Large Language Models (LLMs) have significantly advanced radiology report generation (RRG) with strong generative priors, standard autoregressive decoders still suffer from sequential error accumulation when they must infer all clinical structures directly from images. Diffusion Language Models (DLMs) have recently emerged as a non-autoregressive paradigm with iterative masking and denoising mechanisms. Instead of replacing the mature language decoder with a DLM, we reveal that the key value of DLMs for RRG lies in controllable clinical planning, where the model predicts globally consistent anchor points before free-text realization. To harness this complementary strength, we propose ANCHOR-GEN, the first framework that uses a DLM as an explicit diffusion language planner for RRG. Rather than forcing the autoregressive decoder to infer every critical structure from visual tokens alone, ANCHOR-GEN predicts a structured clinical canvas that serves as clinically grounded anchor points. An autoregressive LLM then leverages these anchors to complete report generation with both fluency and clinical faithfulness. By systematically combining non-autoregressive planning and autoregressive decoding, our unified pipeline is trained with structured denoising, hierarchical coarse-to-fine supervision, and planner-decoder binding objectives. Extensive experiments on the public \textsc{MIMIC-CXR} benchmark demonstrate that ANCHOR-GEN improves clinical faithfulness over prior works. Comprehensive analyses and ablations further show that diffusion planning and causal decoding offer complementary strengths for RRG, and that integrating them in one pipeline provides an effective route toward more faithful medical text generation.


Approximation Guarantees for Robust Aggregation in Federated Learning

Mélanie Cambus ⋅ Darya Melnyk ⋅ Tijana Milentijević ⋅ Stefan Schmid

Byzantine-tolerant aggregation is a central challenge in federated learning, especially under heterogeneous client data where malicious updates may be statistically indistinguishable from honest but unusual clients. Robustness of aggregation algorithms is crucial to the quality of the final model. In this work, we study the robustness of the Byzantine-tolerant aggregation through the lens of centroid approximation, examining how closely an aggregation rule can match the average of honest client updates when up to $t$ of $n$ clients are Byzantine. We connect this objective to classical validity conditions from Byzantine agreement and show that validity alone is insufficient to guarantee good centroid approximation. Our main theoretical contribution is the first lower bound of $\sqrt{\min\{(n-t)/t,d\}}/2$ on the centroid approximation for aggregation under box validity and a matching upper bound of $2\sqrt{\min\{n,d\}}$ in the practically relevant case $n>2t$. In addition, we present a new algorithm that achieves a $2d$-approximation under convex validity, which also proves that the existing lower bound in the literature is tight. These results expose a fundamental tradeoff between validity conditions and approximation quality. We complement the theory with experiments in FedSGD and FedAvg under standard Byzantine attacks, showing that stronger validity improves stability, while tighter centroid approximation can improve accuracy in heterogeneous settings.

Robust policy-gradient methods provide a principled framework for learning under model uncertainty, but their sample efficiency is often limited by the cost of robust policy evaluation. Unlike standard Bellman updates, robust Bellman operators are nonlinear, so sample-based critic estimates can introduce systematic error that is difficult to control. Existing finite-sample guarantees therefore often require solving the robust critic to high accuracy before each actor update. However, practical robust actor-critic methods typically use small or decaying actor stepsizes rather than aggressive increasing updates, making repeated high-accuracy critic solves particularly costly. Conservative actor updates are especially natural in robust constrained MDPs, where the optimization objective may switch between reward improvement and constraint correction during learning. We show that this high-accuracy critic requirement is overly conservative for KL-robust policy optimization. Rather than controlling the full first-order critic error, such as $\mathbb E \lVert\widehat V_t - V^{\pi_t}\rVert$, the actor analysis only requires control of the conditional critic bias, $\lVert\mathbb E[\widehat V_t \mid \mathcal F_t] - V^{\pi_t}\rVert$. We prove that, in the relevant local regime, this bias scales quadratically with the critic estimation error. Consequently, an $O(\varepsilon)$ actor error only requires the critic mean-square error to be $O(\varepsilon)$, rather than requiring the root-mean-square error to be $O(\varepsilon)$. Combining this bias-based analysis with robust natural policy-gradient updates, we establish an $\widetilde O(\varepsilon^{-3})$ sample-complexity guarantee for model-free discounted robust MDPs with small or decaying actor stepsizes. We further show that the same analysis extends to primal surrogate methods for robust constrained MDPs under KL uncertainty, improving the state of the art for this setting.


A Set-Sequence Model for Time Series

Elliot Epstein ⋅ Apaar Sadhwani ⋅ Kay Giesecke

Sequence models have advanced Multivariate Time Series (MTS) modeling by learning complex, long-range temporal dependencies. Many practical settings, however, involve predicting the evolution of a set of MTSs -- an unordered collection of units (e.g., a basket of stocks, a pool of loans, a fleet of sensors). This cross-sectional structure offers an opportunity to leverage signals shared across units (e.g., market volatility, correlated defaults) that are missed when each MTS is processed independently. Crucially, treating the entire cross-section as a single high-dimensional MTS is infeasible because the number of units varies dynamically at inference. We propose Set-Sequence, a backbone-agnostic architecture for set-of-MTS prediction that leverages this cross-sectional structure. A permutation-invariant Set module extracts a summary of the population, while a Sequence module (e.g., Transformer/SSM/RNN) then models each unit’s temporal dynamics conditioned on both its own history and the learned cross-sectional context. This architecture naturally accommodates varying set cardinalities, supports unaligned series, integrates with standard sequence backbones, and scales linearly in cross-sectional size. Across a synthetic contagion task and two large-scale real-world applications -- equity portfolio optimization and loan risk prediction -- Set-Sequence consistently improves sequence backbones and outperforms domain-specific baselines, delivering higher Sharpe ratios, improved AUCs, and interpretable cross-sectional summaries.


ASSET: Acquisition-Sensitive Subspace Estimation for Test-Time Adaptation of Medical VLMs

Mohan Zhang ⋅ Zhen Tan ⋅ Songyuan Sui ⋅ Tianlong Chen

Medical vision-language models are often adapted under high-quality acquisition conditions. At target deployment, the same model may face lower-quality scanners, protocols, reconstructions, or artifacts. These shifts change the visual evidence, while the intended clinical question remains unchanged. As a result, a source-adapted model can fail under this acquisition shift. Methods are proposed to solve this problem. Medical acquisition-shift correction mainly targets image-only predictors, leaving the multimodal model underexplored. Robust training depends on anticipated degradations, and generic test-time tuning can easily overfit the small target training set, leaving unknown acquisition shifts insufficiently corrected. In this paper, we hypothesize that acquisition-like augmentations of target examples can expose parameter directions responsible for acquisition-shift failures. Based on it, we propose \textsc{ASSET}, an \emph{Acquisition-Sensitive Subspace Estimation} method for test-time adaptation. ASSET estimates an acquisition-sensitive parameter subspace from prediction and gradient changes under these augmentations. It then adapts the model within this subspace through iterative estimate-and-update rounds. Comprehensive experiments show the advantages of ASSET. For example, ASSET consistently improves target accuracy under acquisition-quality shift across datasets and backbones by 2.62\% on average, while preserving performance when no shift is present. Under the strongest acquisition-quality shift, ASSET even improves target accuracy by 4.12\%, showing that the correction generalizes to severe shifts.


Asymmetric Factorization for Low-Rank PSD Learning: When Is the Relaxation Exact ?

Enliang Hu ⋅ Juho Kannala ⋅ Guoying Zhao ⋅ Quanming Yao ⋅ Kun Yue

Low-rank factorization is a standard way to scale optimization in machine learning by replacing large matrix variables with compact factors. For positive semidefinite (PSD) variables, the symmetric Burer--Monteiro factorization (sBMF) writes $Z=XX^\top$ with a single low-rank factor $X$. A recent asymmetric alternative (aBMF) writes $Z=XY^\top$ and adds a quadratic penalty $(\gamma/2)\|X-Y\|_F^2$ to encourage symmetry. This split is attractive because it yields a biconvex objective with alternating convex subproblems, but its practical value depends strongly on how the penalty parameter $\gamma$ is chosen. We study a unified regularized aBMF framework and derive an explicit lower bound on $\gamma$ that guarantees exactness: under mild assumptions, any $\gamma$ above this threshold makes aBMF and sBMF share the same critical points. This gives a principled way to use the asymmetric formulation without altering the critical-point structure of the symmetric problem. In particular, it answers the open question of whether an exact penalty exists for asymmetric relaxation.


A Tight Hierarchy For Chain of Thought

Thanh Le ⋅ A. Pavan ⋅ N. V. Vinodchandran

Chain-of-thought (CoT) reasoning has emerged as a powerful paradigm for extending the computational capabilities of transformer-based models. A growing body of recent work has begun to uncover the fine-grained computational power of transformers augmented with CoT, as well as their connections to classical complexity theory. In this work, we advance this line of research by establishing the first hierarchy theorem for \CoT-based computation. Our main result shows that for any reasonable resource bound $r(n) \geq n$ and any $\varepsilon > 0$, $$ \mathrm{CoT}(r(n)) \subsetneq \mathrm{CoT}(r(n)^{1+\varepsilon}), $$ where $\mathrm{CoT}(r(n))$ denotes the class of decision problems solvable by transformers using $O(r(n))$ CoT reasoning steps on inputs of length $n$. Our result shows that increasing the number of CoT reasoning steps provably yields strictly greater computational power, establishing CoT steps as a fundamental resource in transformer-based computation. Unlike classical hierarchy theorems, which are typically proved via diagonalization, our approach leverages known coarse-grained simulations between CoT-augmented transformers and Turing machines in both directions. Building on these connections, we adapt the classical padding technique from complexity theory to this setting to obtain a tight hierarchy. A key technical contribution of our work is the implementation of padding within the transformer-based CoT model.

Recent diffusion-based video generators have achieved remarkable visual fidelity and prompt controllability, yet scaling them to ultra-high-resolution (UHR) long videos remains prohibitively expensive. The difficulty is especially pronounced for long single-shot generation where a continuous scene must preserve global temporal coherence, and fine-grained spatial details without relying on clip transitions or autoregressive shot stitching. In this work, we revisit this challenge from the perspective of decoupled modeling. We argue that existing video diffusion models already encode strong local visual priors, while the main bottleneck lies in efficiently extending global spatiotemporal modeling as resolution and duration increase. Based on this insight, we propose AtlasVid, a decoupled global-local framework for efficient UHR long video generation. AtlasVid first generates a low-resolution and low-FPS global semantic proxy via temporally scaled RoPE, thereby extending the temporal horizon without increasing the training token count. Guided by this proxy, a high-resolution detail branch performs joint denoising with hierarchical locality-preserving attention. Reordered spatiotemporal windows preserve geometric locality and asymmetric global-local attention injects aligned semantic guidance and preserves the model's pretrained ability. This design enables resolution-agnostic training: the model is trained only at 720P with lightweight LoRA adaptation, yet generalizes directly to 4K and beyond for longer (>10s) video synthesis. Experiments show that AtlasVid substantially improves the efficiency of ultra-high-resolution long video generation, achieving high-quality UHR long video generation with $60.9\times$ speed up and significantly less training cost and even better performance than native 4K video generators.


A Transformer-Derived Iterative Preconditioner

Patrick Lutz ⋅ Themistoklis Haris ⋅ Aditya Gangrade ⋅ Venkatesh Saligrama

We show that a softmax-attention-only transformer trained on in-context linear regression discovers an iterative symmetric preconditioning primitive. The trained model is used only for algorithm discovery: the final method, RAPID, is an explicit preconditioner and uses no learned component at solve time. To extract this mechanism, we apply a layerwise symmetry intervention strategy which reduces the trained layers to a two-scalar update. The current Gram geometry determines the attention scores, and the resulting row update evolves the Gram matrix by congruence, reducing its condition number. We turn this primitive into RAPID, an anytime preconditioner for explicit SPD systems $Ax=b$: every prefix constructs a valid factor $P$ such that $PAP^\top$ can be used to improve the runtime of iterative solvers such as conjugate gradients. RAPID uses fast randomized Walsh--Hadamard mixing and sparse row-min corrections, giving $\widetilde O(d^2)$-time dense iterations. For a Haar-mixing idealization, we prove that RAPID reaches a target condition number in $\widetilde O(d\log\kappa(A_0))$ expected iterations. Empirically, RAPID is fast and versatile: it reaches useful condition numbers quickly, accelerates end-to-end CG solves, and improves conditioning across all tested spectral families and real-world SPD problems.


AUTOMEM: Automated Learning of Memory as a Cognitive Skill

Shengguang Wu ⋅ Hao Zhu ⋅ Yuhui Zhang ⋅ Xiaohan Wang ⋅ Serena Yeung-Levy

Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge---a capacity known in cognitive science as metamemory. We bring this perspective to LLMs by treating memory management as a trainable skill. We promote file-system operations to first-class memory actions alongside task actions, letting the model itself decide how to manage its memory. This memory skill improves along two axes: the structure that supports it (prompts, file schemas, action vocabulary), and the proficiency of the model exercising it. Both axes resist manual optimization: episodes in long-horizon tasks run for thousands of steps, and a single memory mistake can hide long before it surfaces, making human review of full trajectories impractical. We introduce AUTOMEM, a framework that automates both axes. In the first loop, a strong LLM reviews complete agent trajectories and iteratively revises the memory structure that shapes how the agent interacts with its memory files. In the second loop, the agent's own good memory decisions are identified from many episodes and used as training signal to sharpen the model's memory proficiency directly. Across three procedurally generated long-horizon games (Crafter, MiniHack, and NetHack), optimizing memory alone---without modifying the model's task-action behavior---improved the base agent's performance 2x--4x, bringing a 32B open-weight model competitive with frontier systems such as Claude Opus 4.5 and Gemini 3.1 Pro. Our results show that memory management is an independently learnable skill, and a high-leverage objective yielding large gains on long-horizon tasks.


BACE: Behavior-Adaptive Connectivity Estimation from Multi-Region Neural Recordings

Mehrnaz Asadi ⋅ Sina Javadzadeh ⋅ Rahil Soroushmojdehi ⋅ Ali Mousavi ⋅ Terence Sanger

Understanding how distributed brain regions coordinate during behavior requires models that are both predictive and interpretable. We introduce Behavior-Adaptive Connectivity Estimation (BACE), an end-to-end framework for learning behavior-conditioned directed effective connectivity from multi-region intracranial local field potentials (LFPs). BACE first encodes within-region dynamics with region-specific temporal encoders, then applies a learned adjacency matrix selected for each behavioral context, and finally forecasts future neural activity through a graph-conditioned autoregressive decoder. This design yields explicit region-level connectivity matrices whose edges are tied to predictive dynamics rather than post-hoc explanation. On controlled synthetic time series with known directed graphs, BACE recovers the ground-truth edge structure from forecasting alone. We then evaluate BACE on human deep-brain LFP recordings from three participants performing structured motor tasks. Across participants, BACE achieves strong neural forecasting while producing compact, behavior-specific directed graphs that can be inspected across task segments. Reliability analyses further support the stability of the inferred connectivity patterns. Together, these results position BACE as a practical framework for estimating behavior-adaptive effective connectivity from high-dimensional intracranial recordings, enabling interpretable hypotheses about how deep-brain networks reorganize during behavior.


Bayes-Sufficient Compression Is Not Enough: How Communication Helps in Multi-Agent Systems?

Yi Xie ⋅ Zhanke Zhou ⋅ Yi Fan ⋅ Yong Ge ⋅ Bo Han ⋅ Bo Liu

Multi-agent LLM systems often split work between a main agent with broad context and an executor subagent with limited local context. Communication can help recover missing information, but it can also add parsing burden, distract from the local decision, or induce protocol failures. We ask when the main agent should send a short message rather than raw context or no message, and when upgrading the main agent pays off. We formalize this as receiver-relative bounded coordination, where a message's value is the executor's next-step gain minus its protocol tax. This view yields four findings. First, compressed messages can outperform raw context when tax savings exceed losses from omitted information or decoder mismatch. Second, Bayes-sufficient compression, which preserves all information needed for the optimal decision, can still be worse than raw context when a bounded executor cannot decode or operationalize its surface form. Third, post-message failures can be localized into externalization, absorption, and action closure, with residual error concentrating in closure even after the right content reaches the executor. Fourth, a stronger upstream agent helps only when the executor's local view is weak enough for the added gain to exceed the added tax. Across six benchmarks, these regimes recur: the same multi-agent protocol raises ContextBench joint accuracy from 0.633 to 0.775 but lowers ToolSandbox from 0.889 to 0.653, helping in context-heavy regimes and hurting in locally sufficient ones. The same quantities drive an inference-time selector over communication actions, improving the accuracy-cost frontier on the two benchmarks where the full action set is evaluated. Code and results at https://anonymous.4open.science/r/Communication/

Randomized controlled trials (RCTs) identify trial-anchored treatment effects but are often too small for reliable heterogeneity estimation; observational studies (OS) are larger but confounded and only partially overlap the trial in measured covariates. We propose Bayesian Calibrated ALignment under covariate Mismatch (B-CALM), a Bayesian borrowing framework for RCT-anchored conditional average treatment effect (CATE) estimation. B-CALM maps source-specific covariates into a shared latent state, jointly models trial and observational outcome surfaces, and introduces baseline-bias and comparative-bias functions that absorb how the OS departs from the trial estimand. The comparative-bias prior becomes an explicit sensitivity knob: we prove a function-valued bias-limited information bound showing that observational contrast information about the trial treatment-effect surface is capped by the prior precision of this bias function, with a scalar corollary in which the effective sample size (ESS) saturates as OS sample size grows. A PAC-Bayes-style risk decomposition separates RCT empirical risk, latent alignment, and residual calibration of the debiased OS surface. Across synthetic, semi-synthetic, and pediatric-obesity external-control studies, B-CALM delivers calibrated credible intervals and low negative transfer while pooled and forest baselines can become overconfident under comparative bias.


Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction

Utkarsh Sahu ⋅ Zhisheng Qi ⋅ Li Zhu ⋅ Yizhao Yang ⋅ Jun Li ⋅ Ryan Rossi ⋅ Yu Wang

Social media popularity prediction aims to forecast the future reach or influence of online content from early-stage observations. Accurate prediction enables key downstream applications, such as advertising optimization and strategic content planning by users, creators, and platforms. Despite substantial progress, existing popularity prediction works often fail to jointly consider multimodal content and temporal social interaction signals. Moreover, the literature remains highly fragmented across datasets, modalities, observation windows, prediction targets, and evaluation protocols. This fragmentation prevents fair comparison and obscures a systematic understanding of how textual, visual, temporal, and interaction-based signals jointly shape popularity dynamics. To address these challenges, we introduce MMG-Pop, a Multi-modal Graph-based Popularity Prediction benchmark, which unifies datasets, modalities, temporal interaction signals, and representative baselines under a standardized evaluation protocol. Furthermore, we propose MMG-PopNet, a unified multi-modal graph-based network that jointly models the aforementioned multi-modal signals and graph-structured social interactions. Extensive experiments on MMG-Pop, comprising four datasets across Bluesky and Reddit platforms, demonstrate the superior performance of MMG-PopNet and yield new insights into cross-platform training generalization, multi-task prediction benefits, multi-modality contributions, and LLM prediction limitation. These findings establish a unified foundation for future research on social dynamics modeling and intervention under heterogeneous modalities and socially-aware agentic ecosystem paradigms.


Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency

Zhikai Wang ⋅ Jiashuo Sun ⋅ Wenqi Zhang ⋅ Zhiqiang Hu ⋅ Xin Li ⋅ Deli Zhao

Recent advancements in Large Vision-Language Models (LVLMs) have significantly enhanced their ability to integrate visual and linguistic information, achieving near-human proficiency in tasks like object recognition, captioning, and visual question answering. However, current benchmarks typically focus on knowledge-centric evaluations that assess domain-specific expertise, often neglecting the core ability to reason about fundamental mathematical elements and visual concepts. We identify a gap in evaluating elementary-level math problems, which rely on explicit visual dependencies-requiring models to discern, integrate, and reason across multiple images while incorporating commonsense knowledge, all o which are crucial for advancing toward broader AGI capabilities. To address this gap, we introduce VCBench, a comprehensive benchmark for multimodal mathematical reasoning with explicit visual dependencies. VCBench includes 1,720 problems across six cognitive domains, featuring 6,697 images (averaging 3.9 per question) to ensure multi-image reasoning. We evaluate 26 state-of-the-art LVLMs on VCBench, revealing substantial performance disparities, with even the top models unable to exceed 50% accuracy. Our findings highlight the ongoing challenges in visual-mathematical integration and suggest avenues for future LVLM advancements.


Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns

Shiqi He ⋅ Yue Cui ⋅ Feijie Wu ⋅ Xinyu Ma ⋅ Jiaheng Lu ⋅ Yaliang Li ⋅ Bolin Ding ⋅ Mosharaf Chowdhury

Large language model (LLM) web agents are usually deployed as tool callers: each turn, the model reads a fresh page observation and emits one structured tool action. When every action is a low-level primitive, horizons grow quickly and so do policy-facing LLM completions, dominating latency and cost on benchmarks such as Mind2Web and WebArena. Recent systems therefore wrap repeated interaction fragments as web skills: callable tools built from successful trajectories or induced programs, so one call can replace several primitives. However, prior skill libraries are still triggered mainly by instruction similarity or coarse site metadata, which yields low skill reuse on held-out sites and leaves much of the potential step and token reduction on the table. We present SkillMigrator, an agent that learns reusable web skills and transfers them across sites by matching layout structure rather than specific element references. Each induced skill is stored as a transferable interaction pattern (TIP): the skill paired with a structural sketch of the snapshot at induction time. At test time, SkillMigrator retrieves TIPs by layout similarity and grounds their references on the live page. The rest of the stack is standard: accessibility-snapshot observations with stable references, and fixed tool calling over primitives plus skill invocations. Compared with the state-of-the-art approaches, SkillMigrator reduces the average LLM-action count on successful trajectories by 8–10\% across both WebArena and Mind2Web at matched success rate.

Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors. Most existing methods compute a fixed steering direction in the original activation space, typically from pairs of contrastive examples using mean differences, linear probes, or arbitrary separability criteria. While effective to a certain extent, these methods treat behavioral control as a global, linear, additive offset: the same direction is applied across inputs, and behaviors are linearly separable. This can be restrictive when behavioral features vary nonlinearly across the activation space or lie on curved and anisotropic manifolds, where the optimal intervention may be input-dependent. To address this limitation, we propose INNSteer, a nonlinear activation steering framework based on invertible latent transformations. Rather than searching for a better steering vector in the original representation space, INNSteer learns a lightweight invertible neural network $\phi$ that maps an LLM's activations into a latent space where behavioral classes are more amenable to linear control. At inference time, activations are mapped through $\phi$, steered in the latent space, and mapped back through the exact inverse transformation $\phi^{-1}$. This makes a simple latent-space translation become a nonlinear, input-dependent intervention in the original activation space. Across experiment settings on multiple LLM families, scales, behavioral traits, and safety benchmarks, INNSteer consistently improves model control over linear, transport-based, and nonlinear steering baselines while largely preserving generation fluency.


Beyond LLM-Based Reasoning: Lightweight GNNs for Agent Failure Attribution

Ting-Wei Li ⋅ Yuanchen Bei ⋅ Xiao Lin ⋅ Hanghang Tong

Large language model (LLM)-based multi-agent systems (MAS) often exhibit complex failure modes, which frequently cause agents to produce incorrect outcomes. This motivates the task of \textit{Agent Failure Attribution}: given a failed multi-agent trajectory, identify the faulty agents and their corresponding error types. Existing approaches predominantly rely on LLMs to perform failure attribution, either through direct prompting, fine-tuning on synthetic data or complex agentic pipelines. While effective, these methods incur substantial computational overhead due to long-context processing, expensive post-training and handcrafted workflows. Moreover, empirical evidence shows that even state-of-the-art models achieve limited accuracy on existing benchmarks, suggesting that scaling model size alone is insufficient. In this work, we revisit this task and question the necessity of such expensive generative solutions. We introduce AFANet, a lightweight graph-based framework that models interaction trajectories through step-level semantic signals and agent-level relationships. We show that with significantly fewer parameters and near-zero inference cost, AFANet (i) matches or outperforms LLM-based baselines, including fine-tuned models on in-domain benchmarks, (ii) maintains robust performance across different GNN architectures and (iii) can be further improved with inexpensive test-time adaptation on the OOD benchmark. Our results suggest that effective agent failure attribution does not require heavy LLM reasoning and a lightweight, structured approach can achieve strong performance.

Owing to enhanced capabilities to capture a broad range of higher-order interactions, simplicial neural networks (SNNs) have emerged as a new powerful methodology for graph learning. However, prevailing SNNs are often limited in efficient exploration of heterogeneous graphs, thereby restricting the SNN utility on downstream tasks. To address this challenge, we introduce the concept of L\'{e}vy flights over simplices, offering a new efficient alternative to learning higher-order graph (sub)structures under heterogeneous scenarios. Specifically, we develop a new Fractional Hodge Laplacian Simplicial Neural Network (FHL-SNN), a novel approach leveraging fractional powers of the Hodge-Laplacian and allowing for a more efficient exploration of the underlying higher-order graph organization in conjunction with the link prediction. We establish a theoretical connection between a fractional diffusion process and graph exploration over simplices. In particular, we prove that the L\'{e}vy flight over simplices yields a lower mixing time than that of its non-fractional counterpart. Our extensive experiments on both directed and undirected link prediction tasks illuminate the power of non-local search of graph space via fractional dynamics, yielding relative gains of up to 9\%. Finally, we show that the idea of L\'{e}vy flights over simplices is versatile and that the integration of fractional dynamics to existing SNNs has the potential to further boost the model performance.

Generalized linear models (GLMs) are fundamental tools for statistical modeling, with maximum likelihood estimation (MLE) serving as the classical approach for parameter inference. While MLE performs well for canonical GLMs, it can become computationally challenging in more general settings with non-canonical, non-smooth, or nonlinear link functions, where the resulting optimization landscape may be ill-conditioned, non-convex, or non-differentiable. In this paper, we study an alternative estimation framework based on variational inequalities (VIs), which formulates GLM estimation through an operator-based equilibrium condition rather than likelihood minimization. We analyze the VI estimator from a statistical perspective and establish finite-sample error bounds and asymptotic normality under mild regularity conditions, together with convergence guarantees for fixed-point and stochastic approximation algorithms. The framework accommodates a broad class of link functions, including non-canonical and non-monotone cases satisfying a strong Minty-type condition, and extends naturally to generalized additive models via basis expansion. Numerical experiments demonstrate that the VI approach achieves competitive finite-sample accuracy and improved numerical stability relative to MLE, particularly in GLMs and GAMs with non-canonical or non-smooth link functions.


Beyond Score: A Dataset for Joint Action and Score Predictions in 2v2 Sports

Yuting Wang ⋅ Peng Wen ⋅ Fengshun Wang ⋅ Jiayong Zuo ⋅ Xinyu Peng ⋅ Qiurui Wang

Current sports analysis methods often rely on high-level statistics and therefore provide limited insight into the action-level interactions underlying tactical outcomes. To enable fine-grained modeling of tactical processes, we present JAS-2v2, the first public dataset for action-based tactical analysis across 2v2 sports. By unifying beach volleyball, badminton doubles, tennis doubles, and table tennis doubles, JAS-2v2 enables comparative study of diverse tactical structures spanning both open and closed tactical systems under a shared sequence modeling framework. This dataset reframes score prediction as a sequence understanding problem, emphasizing how fine-grained action flows contribute to final tactical outcomes rather than only what the final outcome is. We further propose a unified action-to-score modeling framework based on graph reasoning and temporal modeling. Experiments show that our method outperforms strong baselines on both action prediction and score outcome prediction, and further demonstrate that fine-grained action semantics and sequential action dependencies are critical for tactical reasoning. In particular, using action prediction as an intermediate objective provides more effective representations for score outcome prediction.


Beyond Truthfulness: Evaluating Honesty in Large Language Models

Richard Ren ⋅ Arunim Agarwal ⋅ Mantas Mazeika ⋅ Cristina Menghini ⋅ Brad Kenstler ⋅ Robert Vacareanu ⋅ Mick Yang ⋅ Isabelle Barrass ⋅ Alice Gatti ⋅ Xuwang Yin ⋅ Eduardo Trevino ⋅ Matias Geralnik ⋅ Dean Lee ⋅ Summer Yue ⋅ Dan Hendrycks

As large language models (LLMs) become more capable and agentic, the requirement for trust in their outputs grows significantly, yet at the same time concerns have been mounting that models may learn to lie in pursuit of their goals. To address these concerns, a body of work has emerged around the notion of "honesty" in LLMs, along with interventions aimed at mitigating deceptive behaviors. However, some benchmarks claiming to measure honesty in fact simply measure accuracy—the correctness of a model's beliefs—in disguise. Moreover, no benchmarks currently exist for directly measuring whether language models overtly lie. In this work, we introduce a realistic, large-scale human-curated dataset for directly measuring lies of commission, allowing us to disentangle accuracy from honesty. Across a diverse set of LLMs, we find that while larger models obtain higher accuracy on our benchmark, they do not become more honest. Surprisingly, most frontier LLMs obtain high scores on truthfulness benchmarks yet exhibit a substantial propensity to lie under pressure, resulting in low scores on our benchmark. We find that simple methods, such as representation engineering interventions, can improve honesty. These results underscore the growing need for robust evaluations and effective interventions to ensure LLMs remain trustworthy.


Beyond Unit-Circle Eigenvalues: Invariant Bases for Stable State Space Dynamics

Jiaqian Zhu ⋅ Yang Zhang ⋅ Junhua Ding ⋅ Xiaowei Yu

Diagonal state space models process sequences efficiently but become unstable over long horizons: small spectral deviations compound into information decay or uncontrolled growth. For conservation-preserving or non-dissipative dynamics, constraining eigenvalues to the unit circle is a natural starting point and yields large short-horizon gains ($+3.31$ dB at step 10). We show that this is not enough: spectral constraints govern only the amplitude of each mode, not what each mode represents. The missing condition is alignment between the modal basis and the invariant structure of the underlying dynamics. Without this alignment, unit-circle constraints can preserve quantities unrelated to the task invariants and, on OOD scale-shift transfer, even degrade performance below unconstrained baselines. This establishes a broader principle: *stability in diagonal SSMs is a problem of representation, not just parameterization; eigenvalues determine how modes evolve, but only the basis determines what is preserved.* We instantiate this principle for isotropic pairwise systems through the Kronecker-Cayley Decomposition, which places dynamics in the graph Laplacian eigenbasis, yielding a diagonal SSM with exact layered conservation at $O(N^3+N^2m)$ cost instead of $O(N^3m^3)$, dominated by $O(N^2m)$ when $m \gg N$. Empirically, under our visual N-body protocol, a spectrally stabilized LRU baseline still accumulates $158\pm8$% momentum drift; controlled unit-modulus random and learned bases similarly fail ($120\pm18$% and $148\pm8$% drift); even PCA-derived bases, which empirically approximate the invariant direction, fall roughly three orders of magnitude short of the algebraic guarantee. Only the Laplacian basis yields drift below $10^{-4}$%, stable 1000-step rollouts (PSNR $21.19\pm1.27$ dB at step 999), and $+2.12$ dB OOD generalization on Charged particles.


BigCell: Generating Gigapixel Whole-Slide Images

Srikar Yellapragada ⋅ Alexandros Graikos ⋅ Zilinghan Li ⋅ Kostas Triaridis ⋅ Tarak N Nandi ⋅ Ji Dong K. Bai ⋅ Beatrice Knudsen ⋅ Tahsin Kurc ⋅ Prateek Prasanna ⋅ Rajarsi R Gupta ⋅ Ravi K Madduri ⋅ Joel Saltz ⋅ Dimitris Samaras

Whole-slide images (WSIs) are gigapixel-scale digital scans of histopathology tissue and serve as the primary basis for cancer diagnosis. Existing histopathology generative models operate at the patch scale, producing fixed-size tiles of at most 1024$\times$1024 pixels which encompass only a small fraction of the WSI. In contrast, text reports for clinical diagnoses and slide-level diagnostic labels -- the supervision that exists at scale -- operate at the whole-slide level. We present BigCell, the first approach to generate coherent whole-slide histopathology images at gigapixel resolution. BigCell trains a flow-matching transformer at 2048$\times$2048 pixels on over 120,000 publicly available WSIs spanning multiple organ types. Conditioning on embeddings from pathology foundation models and using an inference-time sliding-window pipeline, BigCell generates coherent images of up to 32,768$\times$32,768 pixels, outperforming prior tile generators on both tile-level and slide-level fidelity metrics. To make gigapixel synthesis practical, we introduce an adaptive scheduling strategy that prioritizes important slide regions, speeding up image generation by over 3\ttimes with minimal quality degradation. Furthermore, we propose the first text-to-image synthesis at 8k resolution by training a separate model that produces dense conditioning grids directly from pathology text reports. We demonstrate that augmenting with synthetic samples improves slide-level few-shot classification performance by up to $18 \\%$ over baselines trained on real data alone.

Protein function is driven by cohesive substructures, such as catalytic triads, binding pockets, and structural motifs, that occupy only a small fraction of a protein's residues. Yet existing pipelines built on protein encoders do not model proteins at the substructure level, leaving the central biological question unanswered: which substructure of a protein is responsible for its function? We introduce BioBlobs, an encoder-agnostic, end-to-end differentiable framework that compresses a protein into a small set of cohesive substructures (blobs) and predicts function from these blobs alone, so that each blob corresponds to a candidate functional region. Across diverse protein function prediction tasks and multiple sequence- and structure-based encoders, BioBlobs matches or exceeds strong baselines while operating on only a small fraction of residues. The discovered blobs adapt their spatial scale to the task, ranging from local catalytic sites to entire structural domains. Trained only on protein-level labels, BioBlobs recovers experimentally annotated catalytic sites in the M-CSA database, demonstrating unsupervised functional substructure discovery and opening a path to large-scale functional site discovery across the unannotated proteome.


Bounding Global and Local Compression Error of Signal Parameterizations

Quang Luong Nhat Nguyen ⋅ Sara Fridovich-Keil

Differentiable signal parameterizations such as implicit neural representations (INRs) and hybrid models are increasingly central to computational imaging, yet principled tools for evaluating reconstruction fidelity at finite model size remain limited when ground truth is unavailable. We introduce a framework for predicting the reconstruction error of compressive signal parameterizations, yielding non-asymptotic, signal-specific bounds that are both theoretically sound and efficiently computable without access to the ground truth signal. Specifically, we prove that when parameterization-based compression satisfies certain natural properties, the compression error at any compression level is bounded by a simple scaled difference between model predictions at different compression levels. We verify these properties for representative model families including interpolated grids, Fourier feature networks, multi-resolution hash tables, and tensor factorizations, and show empirically that the resulting worst-case guarantees can be efficiently adapted into signal-specific error predictors that are tight and generalizable. Across direct fitting of synthetic and natural signals, and inverse problems including radiance field and MRI reconstruction, our method closely tracks global error curves and yields informative local error heatmaps without ground-truth access.


Breakeven complexity: A new perspective on neural partial differential equation solvers

Yijing Zhang ⋅ Nicholas Roberts ⋅ Tanya Marwah ⋅ Misha Khodak

Neural surrogate solvers of partial differential equations (PDEs) promise dramatic speedups over numerical methods, especially in scenarios requires many solves. However, current accuracy-based evaluations do not fully consider two central issues: (i) neural solvers incur substantial up-front costs for data generation, training, and tuning; and (ii) classical solvers can also generate low-fidelity solutions at a sufficiently low simulation cost. To explicitly account for these realities and fully incorporate end-to-end costs, we propose an evaluation framework centered on breakeven complexity, a metric that counts the forward solves before a learned solver is cost-effective relative to an error-equivalent traditional solver. To evaluate this measure, we apply scaling laws to determine how much training budget to allocate to data generation and discuss how to achieve smooth error-matching in diverse settings. We evaluate the breakeven complexity of many modern neural PDE solvers on three PDEs on 2D periodic domains from APEBench and a novel benchmark of flows past multiple obstacles generated by the GPU-native PyFR code. Among other findings, our results suggest that neural PDE solvers will be more effective on problems as problems get harder in terms of cost, dimension, rollout, physics regime (e.g. higher Re), etc.

Continuous-time generative models compress endpoint-conditioned bridges into Markov velocity fields, but no existing diagnostic measures how much information this compression discards. We introduce the \emph{Markovization gap}---the integrated conditional variance of the bridge velocity given the Markov state---which quantifies this loss before any neural network is trained. To make the gap well-defined across model families, we define \emph{Bridge Graphical Models} (BGMs), separating endpoint coupling, bridge law, Markovian projection, and dynamics representation as independent design choices. This decomposition also formalizes Poisson and electrostatic models as field-line bridge kernels with their own caustic gap. Across synthetic, latent, and pixel-space pilots on CIFAR-10 and Fashion-MNIST, a proxy gap estimated in minutes on CPU consistently ranks coupling/bridge choices in the same direction as training loss and FID measured after hours of GPU training.


C3VD-DEFCOL: A Deformable Colonoscopy Dataset with Time-Resolved 3D Ground Truth and Realistic Appearance

Ethan Luk ⋅ Mayank Golhar ⋅ Anthony Song ⋅ Raúl Iranzo ⋅ Victor M. Batlle ⋅ Lalithkumar Seenivasan ⋅ JMM Montiel ⋅ Nicholas Durr

3D reconstruction could improve colonoscopy by estimating mucosal coverage and alerting clinicians to missed regions during screening. However, algorithm development is limited as no current datasets provide both a realistic in vivo appearance and dense, time-resolved 3D ground truth, especially under non-rigid deformation. We present C3VD-DEFCOL, a framework and dataset for evaluating deformable colonoscopy reconstruction with paired geometry and realistic texture. Starting from C3VD/C3VDv2 colon meshes and camera trajectories, we generate controlled deformations of the colon surface, including peristaltic waves and centerline motion, and render fisheye colonoscopy videos with frame-level depth, surface normals, optical flow, occlusion masks, camera poses, and time-stamped 3D meshes. We then use the rendered geometry, primarily depth, to condition an LTX-2.3-based sim-to-real translation model that produces RGB clips with in vivo-like mucosal color, texture, vasculature, and specular appearance while preserving the underlying 3D scene structure. The resulting dataset contains 110 videos from 12 unique colon mesh geometries, with varying camera trajectories, appearances, and multiple deformation regimes, including three peristaltic severity levels. We evaluate the generated videos using appearance realism, geometric consistency, and temporal consistency metrics, and use the paired ground truth to benchmark the downstream task of pose estimation in deformable 3D reconstruction. Our experiments show how pose estimation error increases with increasing deformation severity, providing a controlled stress test that is not possible with existing in vivo datasets. Overall, C3VD-DEFCOL is designed as a reproducible, quantitative evaluation platform for testing deformable 3D reconstruction algorithms, with the goal of reducing the domain gap between synthetic datasets and in vivo colonoscopy.


Calibeating Prediction-Powered Inference

Lars van der Laan ⋅ Mark van der Laan

Prediction-powered inference uses black-box predictions to improve estimation from small labeled samples and large unlabeled samples, but raw prediction scores are often miscalibrated and can be inefficient regression adjustments. We study semisupervised mean estimation with a black-box score, a small labeled sample, and a large unlabeled sample. AIPW and PPI give valid inference for fixed or cross-fitted scores, but their efficiency depends on how well the score predicts the outcome. We propose \emph{calibrated prediction-powered inference}: post-hoc calibrate the score on the labeled sample, then average the calibrated predictions over the pooled covariate sample. The estimator requires no retraining, takes a simple plug-in form, and has an exact AIPW representation. For linear calibration, we show first-order equivalence to PPI++. For isotonic calibration, we establish asymptotic normality, valid Wald inference, and ``calibeating" guarantees: isotonic post-processing improves the score as a predictor and as a first-order regression adjustment within the monotone class, and subsequent score-only post-processing yields no additional first-order gain. We also show that the original PPI estimator is a special case of AIPW and can be inefficient when the prediction score is already accurate. Simulations, benchmark reproductions, and an LLM-evaluation application show that calibrated estimators often improve on PPI and are competitive with AIPW and PPI++.

This position paper argues that causal benchmarks for multimodal AI should measure categorical difference, not capability gap. The field reports machine causal performance against human performance and interprets progress as gap-closing along a shared scale. The shared-scale assumption rests on a flattened reading of Hume's regularity theory in which causation reduces to constant conjunction across instances and the temporal experience of the observer has no role. When understood in terms of its presuppositions, Hume's regularity is inseparable from a temporally extended subject who undergoes the conjunctions over time. Human and machine causal cognition therefore have different temporal architectures. The two are grounded in time as substrate and time as dimension, a categorical difference that scale or additional intervention machinery cannot close. Evaluation paradigms that treat the two as commensurable produce inflated assessments of machine causal competence that independently scored probes cannot capture. We propose \textit{Anthropomorphic Competence Inflation} as a construct the field can develop into a broader measurement instrument and formalize it through the \textit{Inflation Score}. Preliminary observations on two recent vision-language models match the pattern predicted by the categorical claim. The implications affect evaluation practice, deployment in causal-decision domains, and the next phase of AI causality research.


Causal Concept Explanations for Deep Neural Models

Joshua Rountree ⋅ Pulkit Verma ⋅ Oswin So ⋅ Chuchu Fan ⋅ Julie A Shah

Concept-based explanations in neural models tie their outputs to human-meaningful concepts that users can understand, troubleshoot, and trust. But these explanations are typically correlational rather than causal: a concept can fire alongside an output without driving it, yielding interpretability that looks right but misleads. We introduce Causal Concept Wrapper Network (CCW-Net), the first use of mediation analysis as a training objective in deep networks. We further apply mediation analysis as an evaluation tool to quantify, per sample, how much of any model's output is causally attributable to each concept. As a training objective, CCW-Net shapes a model's concept embeddings such that each concept's contribution to the output is both locally necessary and sufficient for the share it is credited with, and independent of contributions from other concepts. Counterfactual concept samples drawn from a learned per-concept flow anchor a set of causal alignment criteria: necessity drives the effect of every irrelevant concept toward zero; sufficiency requires relevant concepts to fully reconstruct the model's output; and independence enforces that the total effect decomposes linearly across concepts. We evaluate CCW-Net across aircraft control, driving, and fine-grained image classification. The result is a step toward concept-based explanations that are not merely interpretable, but quantifiably causal, supporting safer deployment, evaluation, and certification of transparent and trustworthy neural systems.


Causal Survival Forests with Negative Controls

Zijun Gao ⋅ Kyoungeui Hong ⋅ Leyi Ma ⋅ Qianli Wu ⋅ Zachary Izzo ⋅ Ruishan Liu

We study heterogeneous treatment-effect (HTE) estimation in observational survival studies commonly associated with both censored outcomes and unmeasured confounding. We integrate causal survival forests (CSF) with negative controls (NC) from proximal causal inference and introduce Negative Control Causal Survival Forests (NC-CSF), a flexible nonparametric HTE learner for survival analysis. In particular, our approach uses a loss that incorporates proxy variables and Neyman orthogonalization to train the random forest, thereby mitigating bias from unobserved confounding and gaining robustness to nuisance estimation. Through extensive simulations spanning varying levels of confounding, proxy relevance, and censoring mechanisms, we demonstrate that NC-CSF substantially reduces bias and estimation error (root mean squared error) relative to existing baselines. We further demonstrate the practical utility of our method on real-world clinical datasets, where it confirms several existing findings and also reveals new interpretable patterns of treatment-effect heterogeneity. To facilitate practical use, we provide an end-to-end Python implementation of NC-CSF that carefully handles implementation details such as nuisance estimation and clipping.


CC-GS: Low-Memory 3D Gaussian Splatting Training via CPU-GPU Block-Wise Context Compositing

Jian Xu ⋅ Siyi Wu ⋅ Yi Li ⋅ Bingzhe Li ⋅ Sian Jin ⋅ Wei Niu ⋅ Sheng Di ⋅ Yuede Ji ⋅ Miao Yin

Training a single 3D Gaussian Splatting (3DGS) scene routinely requires tens of gigabytes of GPU memory, placing standard 3DGS training beyond many low-memory GPU budgets commonly found on laptops and edge devices. Existing low-memory approaches mainly follow two directions: block-wise training and frustum-culling-based host offloading. Block-wise methods train spatial blocks independently and merge them afterward. However, because 3DGS renders each image through depth-ordered alpha compositing over all visible Gaussians, independent block training reduces optimization to block-local objectives, thereby losing full-image loss supervision. Frustum-culling-based host offloading retains full-image loss supervision by jointly rendering all visible Gaussians while uploading only the view-visible subset to the GPU. However, its GPU memory consumption still scales with the size of the visible set, which can exceed the memory budget of low-end GPUs for dense scenes or wide-coverage views. We propose CC-GS, a low-memory block-wise 3DGS training framework based on CPU-GPU block-wise context compositing, featuring three key designs. $\textbf{i) Block-wise context compositing}$: CC-GS partitions Gaussians into capacity-bounded blocks and represents non-optimized visible blocks using cached per-pixel color and residual transmittance, with depth used only for ordering, enabling full-image loss supervision while processing one capacity-bounded block at a time. $\textbf{ii) Two-pass block-wise training}$: each training view is processed through a lightweight context pass for caching compositing context and a differentiable block-optimization pass for updating individual blocks. $\textbf{iii) Asynchronous CPU-GPU execution}$: CC-GS overlaps block gathering, host-device transfers, rendering, backpropagation, gradient offloading, and CPU-side updates through an asynchronous CPU-GPU pipeline. Experiments on standard 3DGS benchmarks show that CC-GS keeps the measured peak GPU memory below 2GB for all evaluated scenes while matching the rendering quality of standard 3DGS. Compared with standard 3DGS, CC-GS, incurs at most a $1.54\times$ dataset-level training slowdown.


CD-RCM: Generalizable Continuous-Depth Novel View Synthesis for Reflectance Confocal Microscopy

Tooba Imtiaz ⋅ Milind Rajadhyaksha ⋅ Kivanc Kose ⋅ Jennifer Dy

Reflectance confocal microscopy (RCM) provides noninvasive, cellular-resolution “optical biopsies” of human skin {\em in vivo} by acquiring en-face images at successive depths, forming a sparse $z$-stack. Due to optical limitations, these stacks are anisotropic 3D volumes with lateral resolution ($0.5\mu m$) $\sim$6 times higher compared to axial resolution, which is defined by the optical sectioning ($3\mu m$), limiting the interpretation of tissue. Our goal is to provide continuous-depth visualization by interpolating intermediate sections and making the 3D volume isotropic. Such a representation permits arbitrary-direction sectioning, including histopathology-like cross-sectional examination, without requiring per-patient optimization. To that end, we introduce the first RCM-specific novel-view synthesis (NVS) approach CD-RCM: a feedforward model that predicts realistic, unseen depths from sparsely sampled RCM stacks. Classical neural rendering methods focus on reconstruction from surface-level multi-view observations. In contrast to surface-level camera views, RCM can acquire optically sectioned en-face images of tissue beyond the surface up to $200 \mu m$. However, during visualization of the RCM stacks, observations of the shallower sections (towards the surface) obscure the deeper ones. This unique axial imaging geometry and layer-dependent anatomical organization motivated our development of a tailored architectural and training framework that explicitly accounts for RCM’s depth-resolved, occlusive imaging physics. Experiments demonstrate that CD-RCM achieves high-fidelity novel-view synthesis with sub-second inference time.


Certified but Private: Scalable Zero-Knowledge Proofs for the Formal Verification of Neural Networks

Youwei Zhong ⋅ Ben Merbaum ⋅ Timos Antonopoulos ⋅ Ning Luo ⋅ CHARALAMPOS PAPAMANTHOU ⋅ Katerina Sotiraki ⋅ Ruzica Piskac

With the increasing deployment of machine learning models, formal guarantees of the robustness and fairness of these models have become very important in safety-critical and legal compliance settings. However, model parameters are often commercial secrets that cannot be disclosed to auditors or end users. To this end, we present PANDA, a scalable system that uses Zero-Knowledge Proofs (ZKPs) to prove robustness and fairness properties of a model without revealing private model parameters. PANDA is built on top of CROWN an efficient robustness certification framework that is used in many state-of-the-art formal verification tools for neural networks. The core contribution of PANDA is a novel algorithm for proving linear relaxation bounds on non-linear activation layers, achieving simple and lightweight proofs. Remarkably, our system can generate proofs of local robustness for neural networks with more than 2.9M parameters in about 4 minutes, and can verify them in under 7.5 seconds. Prior ZKP-based robustness systems are based on exponential-time algorithms that cannot scale to nontrivial networks. In contrast, PANDA scales polynomially with the number of neurons in a network, allowing us to support neural networks 3 to 4 orders of magnitude larger than prior works with significantly reduced prover overhead.


C-GRPO: Conformal Group Relative Policy Optimization

Arya Fayyazi ⋅ Seyedarmin Azizi ⋅ Mehdi Kamal ⋅ Massoud Pedram

Reasoning-oriented language model training increasingly relies on multiple on-policy rollouts, but standard Group Relative Policy Optimization (GRPO) uses a fixed group size, wasting compute on easy prompts and under-exploring hard ones. Existing adaptive-budget methods alleviate this inefficiency, yet their stopping rules are heuristic and do not provide a finite-sample calibration guarantee. We propose CGRPO, which replaces the fixed rollout budget of standard GPRO with calibrated conformal thresholds over task-specific nonconformity scores, combining Adaptive Prediction Set (APS)-style calibration for closed-form reasoning, a first-success score for execution-based code, automatic miscoverage selection, and periodic re-calibration. We prove finite-sample coverage for the resulting stopping rule and show that, under monotone policy improvement, recalibration induces a non-increasing expected training budget. Across six benchmarks spanning mathematics, graduate-level science, and code generation, CGRPO reduces mean training rollouts by 49-70\% relative to fixed-budget GRPO while retaining strong fixed-budget accuracy, achieves the strongest pass@1 among evaluated static and dynamic baselines on different datasets, and transfers across model families with up to 82\% rollout savings.


Chain-of-Route: State-Aware LLM Routing for Multi-Turn Conversations

Jiaqi Xue ⋅ Mengxin Zheng ⋅ Heng Huang ⋅ Qian Lou

LLM routing dispatches each user query to the most suitable model from a pool of candidates, balancing response quality against serving cost. As LLMs are increasingly deployed in conversational interfaces, routing decisions must now be made repeatedly within an ongoing session. Existing routers treat each query independently, implicitly assuming that the optimal model depends only on the current input. This assumption ignores two forms of session state that accumulate across turns: the dialogue state, which shapes query meaning and difficulty, and the \textit{serving state}, which determines each model's true \textit{serving cost} via KV-cache reuse. Neglecting dialogue state leads to misrouting when queries depend on prior turns, while ignoring serving state causes systematic cost overestimation that worsens as conversations grow longer. A practical multi-turn router must therefore capture dialogue state for accurate quality prediction, and track per-model serving state for accurate cost estimation. To address this gap, we introduce Chain-of-Route (CoR), the first state-aware routing framework for multi-turn conversations. CoR comprises two modules: a Dialogue State Propagator (DSP) that maintains a compact recurrent state to capture session-level context across turns, and a Serving State Tracker (SST) that tracks per-model KV-cache to compute true serving costs. Together with the framework, we construct a unified multi-turn routing benchmark from WildChat and LMSYS-Chat-1M with per-turn quality annotations across five LLMs. Experiments show that CoR achieves state-of-the-art quality-cost tradeoffs, reducing QNC by up to 20.3\% over the strongest baseline while improving response quality.


ClusQuant: Mitigating Outliers with Clustering-Based Representations for Low-Precision LRMs

Xingyu Liu ⋅ Xiangyang Yin ⋅ Tianhua Xia ⋅ Haiyu Wang ⋅ Sai Qian Zhang

Large reasoning models (LRMs) improve task performance by generating longer intermediate reasoning traces, but this substantially increases inference costs and makes low-precision deployment increasingly important. However, aggressive quantization often performs poorly on reasoning tasks. Beyond error accumulation across autoregressive generation, we observe that activations exhibit significantly higher variability during the early decoding steps of thinking, making conventional calibration and uniform quantization less effective. To address these reasoning-specific quantization challenges, we propose \textbf{ClusQuant}, a clustering-based LUT quantization framework for low-precision large reasoning models. ClusQuant strengthens outlier mitigation through rotation and scaling transformations, uses clustering-based representations to better fit non-uniform activation distributions, and replaces fixed codebook sizes with an adaptive clustering strategy that allocates codebook capacity according to layer-wise quantization difficulty. Guided by our observation of reasoning-time activation behavior, ClusQuant further improves calibration sample selection to better match actual decoding statistics. Combined with our customized LUT CUDA kernel, ClusQuant delivers substantial accuracy improvements on reasoning benchmarks while achieving $2.85\times$ speedup.


CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use

Zhen Zhang ⋅ Kaiqiang Song ⋅ Sean Wang ⋅ Yebowen Hu ⋅ Weixiang Yan ⋅ Chenyang Zhao ⋅ Henry P Zou ⋅ Haoyun Deng ⋅ Sathish Reddy Indurthi ⋅ Shujian Liu ⋅ Simin Ma ⋅ Xiaoyang Wang ⋅ Xin Wang ⋅ Song Wang

AI agents increasingly solve real-world tasks by reasoning over multi-turn user interactions and invoking external tools. Yet reinforcement learning (RL) remains difficult in this setting: realistic objectives are often open-ended and lack verifiable rewards, RL for multi-turn, multi-step agentic tool use is underexplored, and executable tool environments are costly to build and maintain. We propose CM2, an RL framework that replaces verifiable outcome rewards with checklist rewards. CM2 decomposes each turn's intended behavior into fine-grained binary criteria with explicit evidence grounding and structured metadata, turning open-ended judging into more stable classification-style decisions. To balance stability and informativeness, CM2 uses sparse reward assignment with dense evaluation criteria, and trains in a scalable LLM-simulated tool environment to avoid engineering large tool sets. Experiments show that CM2 consistently improves over supervised fine-tuning. Starting from an 8B Base model and training on an 8k-example RL dataset, CM2 improves over the SFT counterpart by 8 points on $\tau^2$-Bench, 10 points on BFCL-V4, and 12 points on ToolSandbox. The results match or even outperform similarly sized open-source baselines, including the judging model. CM2 thus offers a scalable recipe for optimizing multi-turn, multi-step tool-using agents without verifiable rewards.

Human cognition operates under capacity limitations, requiring computational adaptations for solving complex tasks. Although such adaptations have been described algorithmically, their neural substrates and mechanistic implementations remain poorly understood. Here, we study a hierarchical decision-making task in which subjects must simultaneously infer latent rule changes and resolve sensory uncertainty. By comparing human behavior with Bayesian observer models, we identify three systematic deviations: humans make decisions hierarchically, persist in the previous context after a context switch, and preferentially explore a new context within the same sensory modality. We argue that these deviations are jointly explained by a cognitive-constraint perspective in which hierarchical decisions are less cognitively demanding and switching context or modality incurs additional cost. Incorporating these constraints substantially improves behavioral fits and reveals correlated subject-level context and modality stickiness parameters, suggesting a shared underlying cost mechanism. We then asked how neural representations support hierarchical decisions. Because hierarchical decisions require selecting a latent context before choosing an action, context representations provide a natural substrate for reducing the effective decision space. Neuroimaging data revealed that context information is represented in mediodorsal thalamus, consistent with a role for thalamus in task-state compression. Motivated by the context signal in thalamus, we trained a thalamocortical recurrent neural network in which a thalamic module and faster corticothalamic plasticity provided an inductive bias for task compression, while switch-dependent working-memory noise instantiated switching costs. These constrained networks better matched human behaviorial bias than standard task-optimized networks and reproduced key neural signatures observed in the data, including context encoding in mediodorsal thalamus and cue/rule representations in prefrontal cortex. Together, these results suggest that structured biases in hierarchical decision-making arise from cognitive constraints and identify thalamocortical dynamics as a candidate mechanism for efficient task-state inference.


COMPLLLM: Fine-tuning LLMs to Discover Complementary Signals for Decision-making

Ziyang Guo ⋅ Yifan Wu ⋅ Jason Hartline ⋅ Kenneth Holstein ⋅ Jessica Hullman

Multi-agent decision pipelines can outperform single agent workflows when complementarity holds, i.e., different agents bring unique information to inform a final decision. We propose ComplLLM, a post-training framework based on decision theory that fine-tunes a decision-assistant LLM to output signals that complement existing agent decisions, using complementary information as reward. We validate ComplLLM on synthetic and real-world tasks involving domain experts, demonstrating how the approach recovers known complementary information, produces plausible explanations of complementary signals to support downstream decision-makers, and improves human and LLM agent decision performance in controlled studies.

Link prediction remains difficult in sparse graphs, where diffusion heuristics lack connectivity and geometric methods suffer spectral instability when coupling learned representations with diffusion operators. We introduce **Concorde**, an energy-based framework that resolves this by decomposing link likelihood into three structurally decoupled branches: **(i)** a parameter-free spectral anchor that establishes permutation-invariant topological stability; **(ii)** a mixed-curvature residual on a learnable $\mathbb{H}\_{c\_h}^{d\_h} \times \mathbb{S}\_{c\_s}^{d\_s}$ product manifold, regularized by a feature-aware Ollivier--Ricci curvature proxy; and **(iii)** a continuous vector field on $\mathcal{T}\mathbb{H}\_c^d$ for non-local structural alignment across topological gaps. The branches are fused via learned gating and trained with an $N$-pair ranking loss aligned with retrieval-style evaluation. Across five benchmarks spanning sparse, dense, featureless, and attributed regimes, **Concorde** establishes state-of-the-art early retrieval precision, nearly doubling $P@10\\%$ over leading baselines in sparse settings.


CONSTRAINER: Promptable Graph-Structured Optimization via Constraint Conditioning

Zeeshan Memon ⋅ Mingke Tian ⋅ Xinyuan Song ⋅ Hongwei Jin ⋅ Kibaek Kim ⋅ Liang Zhao

Neural methods for graph-structured optimization have achieved strong performance within individual problem classes, but are typically trained for fixed formulations, limiting reuse across related problem variants. Recent progress in foundation models has motivated shared neural backbones trained across multiple graph-structured optimization tasks; however, existing approaches primarily share instance-level representations while treating feasibility constraints implicitly. A key source of shared structure across constrained optimization problems is the constraint set itself, which often exhibits substantial overlap across tasks. We propose CONSTRAINER, a constraint-conditioned framework for graph-structured optimization that treats constraints as structured inputs. Constraint specifications are parsed into operator graphs and encoded using a graph neural network to produce invariant constraint embeddings, which condition a shared backbone model for solution prediction. We evaluate our approach on power grid optimization and mixed-integer linear programming problems, where it shows improved solution optimality and enables efficient transfer learning and generalization compared to baseline methods.

This paper concerns the problem of scaling agents to large action spaces demanding pretrained world knowledge. What type of representation pretraining would enable policies such as language models to generalize rewards to appropriate actions? Contrastive learning aligns similar inputs and otherwise enforces distancing of representations, while non-contrastive alternatives omit distancing terms and so learn anisotropic representations that collapse distinctive clusters. Though contrastive pretraining has been proven to promote supervised learning efficiency, non-contrastive pretraining nevertheless remains canonical for language agents. Here we examine the extent to which these pretraining types support on-policy finetuning, in which efficiency is coextensive with online performance. Broadly, we argue that contrastive pretraining optimizes efficient exploration of large spaces. We prove that anisotropic representations collapse distinct actions and necessitate complex policies, whereas contrastive pretraining retains distinctions and enables reward transfer through spectral regularization. We validate our conclusions empirically by providing spectral analyses of language data and simulations involving multimodal personalization policies as well as self-improving reasoning agents finetuned online using a policy gradient. Our latter simulations suggest contrastive pretraining for language agents exploring diverse reasoning paths.

Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives defined in the prediction space. While effective at reducing forgotten content generation, such approaches may act as suppression: forgotten concepts can persist in representations and remain entangled with retained knowledge. We introduce CLReg, a contrastive representation regularizer that identifies forget features while pushing them away from retain features, reducing forget--retain interference while empirically preserving the scale and shape of retain features. As light motivation for the mechanism, we provide a one-step analysis showing that CLReg decreases a simple entanglement proxy in the embedding space. Across unlearning benchmarks and LLMs of different sizes, CLReg decreases forget-retain representation entanglement to enhance mainstream unlearning methods without extra privacy risks, inspiring future unlearning work to remove forget concepts via representation shaping.


Counterfactual Debugging the World Model Transfer Gap

Mingxuan Li ⋅ Kai-Zhan Lee ⋅ Michael Dennis ⋅ Elias Bareinboim

Policies achieving strong performance in simulators or learned world models can fail when deployed into the real environment. While there must be environmental differences that account for these failures, not all differences are equally responsible. A benign error of irrelevant visual details may co-exist with a failure to predict a single critical transition causing an inevitable catastrophic failure. In this paper, we introduce \emph{counterfactual debugging} as an approach for \emph{world model transfer gap attribution} to identify the time steps whose transition or reward errors are \emph{causally responsible} for a performance degradation exhibited in a real trajectory, rather than only visually or statistically different. We develop a scalable algorithm for computing these attributions that exploits the sparsity of causal errors by recursively divide-and-conquer, significantly reducing computational costs and achieving an exponential speedup. The resulting ranked attribution report can better explain the performance gap in the real trajectory. Experiments across several distinct environments with a variety of injected world-model failures (observation corruption, reward misprediction, and physics violations) demonstrate that counterfactual debugging correctly identifies the errors that are responsible for performance gap, providing theoretically grounded, actionable insights for model improvement.


Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents

Yuanhao li ⋅ Hongbo Wang ⋅ Xuhong Chen ⋅ Yiming Cao ⋅ Xunzhu Tang

Reinforcement learning has rapidly become the dominant paradigm for training software engineering (SWE) agents on long-horizon, multi-turn tasks such as resolving real-world GitHub issues. Yet existing pipelines suffer from a fundamental credit-assignment bottleneck: rewards arrive only at the end of trajectories that may span dozens of file edits, shell commands, and test invocations, leaving every intermediate decision indistinguishable from luck. We observe that containerised SWE tasks provide a particularly useful form of environment forkability: a commit hash plus a container image specifies a real executable software state that can be restored cheaply and exactly, with auditable provenance and without a simulator-to-reality gap. We exploit this property by introducing Counterfactual Rollout Replay (CRR), a training-time procedure that re-executes the environment from selected decision points with alternative actions and assigns each selected step an advantage equal to the sampled on-policy versus counterfactual return differential. CRR requires no learned process reward model, no human rubric annotation, and no oracle hindsight signal; it is free of auxiliary process labels and reward-model training, while still incurring extra environment-replay compute. Applied to a 14B base model on SWE-Gym and evaluated on SWE-bench Verified and SWE-rebench, CRR-trained agents reach higher resolution rates with substantially fewer trajectories than outcome-only GRPO and complement orthogonal advances such as PRM-based scoring and trajectory search.


Coupling Models for One-Step Discrete Generation

Fred Peng ⋅ Joey Bose ⋅ Anru Zhang ⋅ Alexander Tong

Generative modeling over discrete structures underpins applications across deep learning, from biological sequence design and code generation to large language models, yet generation often remains sequential, relying on autoregressive decoding or iterative refinement. We introduce Coupling Models, a one-step discrete generative model that learns a direct coupling between discrete sequences $x$ and Gaussian latents $z \sim \mathcal{N}(0,I)$. Unlike recent distillation methods that compress a pretrained multi-step sampler into fewer steps, Coupling Models train a purpose-built decoder to invert this coupling and generate samples in a single forward pass. The method also avoids complex continuous flows over the simplex and hand-specified data-to-noise couplings. Empirically, Coupling Models improve the strongest one-step baselines in each domain, reducing LM1B text-generation perplexity by 33% at its lowest-perplexity operating point, Fly Brain enhancer-design FBD by 18%, and MNIST-Binary FID by 46%. These results suggest that effective one-step discrete generation depends strongly on how data and noise are coupled before decoding.


Crafting Reversible SFT Behaviors in Large Language Models

Yuping Lin ⋅ Pengfei He ⋅ Yue XING ⋅ Yingqian Cui ⋅ Jiayuan Ding ⋅ Subhabrata Mukherjee ⋅ Hui Liu ⋅ Zhen Xiang

Supervised fine-tuning (SFT) induces new behaviors in large language models, yet imposes no structural constraint on how these behaviors are distributed within the model. Existing behavior interpretation methods, such as circuit attribution approaches, identify sparse subnetworks correlated with SFT-induced behaviors post-hoc. However, such correlations do not imply causal necessity, limiting the ability to selectively control SFT-induced behaviors at inference time. We pursue an alternative by asking: can an SFT-induced behavior be deliberately compressed into a sparse, mechanistically necessary subnetwork, termed a carrier, while remaining controllable at inference time without weight modification? We propose (a) Loss-Constrained Dual Descent (LCDD), which constructs such carriers by jointly optimizing routing masks and model weights under an explicit utility budget, and (b) SFT-Eraser, a soft prompt optimized via activation matching on extracted carrier channels, to reverse the SFT-induced behavior. Across safety, fixed-response, and style behaviors on multiple model families, LCDD yields sparse carriers that preserve target behaviors while enabling strong reversion when triggered by SFT-Eraser. Ablations further establish that the sparse structure is the key precondition for reversal: the same trigger optimization fails on standard SFT models, confirming that structure rather than trigger design is the operative factor. These results provide direct evidence that the learned carriers are causally necessary for the behaviors, pointing to a new direction for systematically localizing and selectively suppressing SFT-induced behaviors in deployed models. Code is available at https://anonymous.4open.science/r/sft-reverse-8150/.


Cross-Cell-Line Perturbation Prediction Needs Controls

Xingyu Fan ⋅ Jinghao Wang ⋅ kim hsieh ⋅ Chunbin Gu ⋅ Pheng-Ann Heng

Predicting genetic perturbation responses across cell lines is challenging not only for models, but also for the metrics used to evaluate them. We introduce a strict cross-cell-line CRISPRi stress test on Replogle Perturb-seq, training on K562 and RPE1 (or additionally HepG2) and evaluating on held-out Jurkat. We pair this setting with two necessary controls: target-control-copy, which predicts no perturbation effect, and shifted-control baselines, which translate target-line controls by a source-derived mean shift. Across classical baselines, recent deep models, CellFlow, and context-conditioned flow-matching variants with five frozen single-cell foundation-model encoders plus a PCA/null context, we find that aggregate distributional metrics can be strongly driven by target-line cell-state variability rather than perturbation-specific response. Target-control-copy is competitive on aggregate distributional metrics, while shifted-control approaches CellFlow on Energy and substantially outperforms it on perturbation-specific Pearson. Used as fixed per-cell context for a shared downstream flow head, none of five frozen scFM encoders matches a PCA/null encoder in this strict cross-line setting. Finally, anchored mean-preserving residual-flow probes reveal a metric-specific trade-off: full-space residual flows improve $W_2$ but degrade per-gene moment matching, whereas low-rank constrained flows preserve moments but lose most $W_2$ gain. No tested method simultaneously dominates perturbation-specific response and biologically relevant distributional fidelity. These results argue for perturbation-aware cell-level evaluation in cross-cell-line virtual-cell benchmarks.

Submodular functions---functions exhibiting diminishing returns---are central to machine learning. When the objective is monotone and non-negative, the greedy algorithm achieves a tight $63\%$ approximation. But many practical objectives incorporate costs that make them negative on some inputs, and all existing multiplicative guarantees require non-negativity. Prior work handles negativity through additive bounds for the special class of decomposable functions and non-monotonicity through partial-monotonicity parameters, but these address each difficulty in isolation and neither extends the classical structural theory. We extend curvature---a parameter measuring how far a function deviates from linearity---to all submodular functions, handling both non-monotonicity and negativity through a single classical concept. A greedy algorithm with pruning achieves a curvature-controlled multiplicative ratio for any submodular function, including those taking negative values---the first such guarantee beyond monotonicity and non-negativity. In the non-monotone regime $1 \le c_g < 2.2$, the bound strictly beats the best known uniform ratio of $0.401$ (for non-negative $f$), and it recovers the classical $(1-e^{-c_g})/c_g$ guarantee for monotone functions. A multilinear-extension variant extends the framework to general combinatorial constraints via multilinear relaxation. Experiments on cost-penalized experimental design, coverage, feature selection, and a curvature sweep on Multi-News passage selection support the theory.


CytoWave: Perturbation-Centric Pretraining for Single-Cell Response Prediction

Yuwei Miao ⋅ Jingquan Yan ⋅ Hehuan Ma ⋅ Mai Thao Dang ⋅ Lin Xu ⋅ Siyuan Zhang ⋅ Junzhou Huang

Predicting gene expression under perturbations from single-cell data is a fundamental task for modeling cellular responses, with applications in discovering gene function and designing therapeutics. However, existing methods often struggle to generalize across datasets due to differences in experimental protocols, biological contexts, and species. We present \textbf{CytoWave}, a perturbation-centric pretraining framework for single-cell genetic perturbation response prediction. We leverage large-scale cross-species perturbation datasets unified via orthology and train CytoWave in a two-stage manner. In the first stage, we learn reference control cell states via masked gene expression recovery on unperturbed cells. In the second stage, we pretrain the model to predict perturbed gene expression given control states and perturbation signals, together with a contrastive objective that aligns latent representations of perturbation responses with perturbation signals. We perform strict dataset-held-out evaluation across multiple condition combinations and case studies of pathway activity changes, demonstrating consistent and effective performance of CytoWave in predicting single-cell genetic perturbation responses.


DarkVGGT: Seeing Through Darkness Using Thermal Geometry without Daylight Tax

Minseong Kweon ⋅ Wenyuan Zhao ⋅ Nuo Chen ⋅ Lulin Liu ⋅ Tracy Han ⋅ Zihao Zhu ⋅ Srinivas Shakkottai ⋅ Chao Tian ⋅ Zhiwen Fan

Recent feed-forward 3D reconstruction methods have demonstrated strong performance and flexibility in efficient end-to-end scene geometry estimation from image streams. However, their reliance on visible-light appearance makes them vulnerable in dark and low-visibility environments, where RGB cues are severely degraded and geometric evidence becomes ambiguous. To address this challenge, we propose DarkVGGT, an RGB-T feed-forward geometry framework that uses physics-aware thermal modeling for robust 3D estimation in low-light scenes. DarkVGGT introduces two complementary modules. First, physics-inspired thermal factorization extracts emissive-dominant, geometry-consistent thermal cues while isolating sparse reflective residuals that may introduce geometric ambiguity. Second, geometry-shared thermal routing isolates modality-invariant geometric structures from thermal-specific patterns, selectively injecting reliability-aware structural guidance into the RGB stream. Together, these components enable accurate thermal-informed geometry estimation under degraded RGB conditions while largely preserving performance in well-lit environments. Experiments on low-visibility RGB-T benchmarks demonstrate consistent improvements in both depth and camera pose estimation over existing feed-forward geometry baselines.


Data Auctions for Retrieval Augmented Generation

Minbiao Han ⋅ Seyed A Esmaeili ⋅ Michael Albert ⋅ Haifeng Xu

We study the problem of data selling for Retrieval Augmented Generation (RAG) tasks in Generative AI applications. We model each buyer's valuation of a dataset with a natural coverage-based valuation function that increases with the inclusion of more relevant data points that would enhance responses to anticipated queries. Motivated by issues such as data control and prior-free revenue maximization, we focus on the scenario where each data point can be allocated to only one buyer. We show that the problem of welfare maximization in this setting is NP-hard even with two bidders, but design a polynomial-time $(1-1/e)$ approximation algorithm for any number of bidders. Unfortunately, however, this efficient allocation algorithm fails to be incentive compatible. The crux of our approach is a carefully tailored post-processing step called data burning which retains the $(1-1/e)$ approximation factor but achieves incentive compatibility. Our thorough experiments on synthetic and real-world image and text datasets demonstrate the practical effectiveness of our algorithm compared to popular baseline algorithms for combinatorial auctions.


Data-Constrained Language Model Pretraining: Improved Regularization and Scaling Laws

Zhiwei Xu ⋅ Shihao Wu ⋅ Hanseul Cho ⋅ Wei Hu ⋅ Yixin Wang

Classical scaling laws for language model pretraining balance model size against training dataset size under a fixed compute budget, assuming abundant data and a single pass over the corpus. As training compute grows faster than the supply of natural language data, pretraining is likely to enter a data-constrained, compute-rich regime where models train for multiple epochs over a finite dataset. We study data-constrained pretraining along two axes, regularization and scaling. For regularization, we study masked-input regularization (MIR), an auxiliary next-token prediction loss on randomly masked inputs. MIR tests whether the random masking central to diffusion language models can benefit autoregressive pretraining without architectural changes or inference overhead. Across 72M to 1.4B parameter models, we find that MIR added on top of strong weight decay improves validation loss over autoregressive strong-weight-decay-only models, with downstream gains at 1.4B. For scaling, we propose SoftQ, a scaling law that couples model size and data size to capture their interaction under repeated data. Classical alternatives such as the Chinchilla law use an additive form that decouples these terms, making them misspecified in the data-constrained regime. We find that SoftQ fits data-constrained experiments substantially better than these alternatives, and estimates MIR's gains as equivalent to roughly 1.3 times as much unique training data.


DD-Ranking: Rethinking the Evaluation of Dataset Distillation

Zekai Li ⋅ Xinhao Zhong ⋅ Samir Khaki ⋅ Zhiyuan Liang ⋅ Yuhao Zhou ⋅ Mingjia Shi ⋅ Ziqiao Wang ⋅ Xuanlei Zhao ⋅ Wangbo Zhao ⋅ Ziheng Qin ⋅ Mengxuan Wu ⋅ Pengfei Zhou ⋅ Haonan Wang ⋅ David Junhao Zhang ⋅ Jia-Wei Liu ⋅ Shaobo Wang ⋅ Dai Liu ⋅ Linfeng Zhang ⋅ Guang Li ⋅ Kun Wang ⋅ Zheng Zhu ⋅ Zhiheng Ma ⋅ Joey Tianyi Zhou ⋅ Jiancheng Lv ⋅ Yaochu Jin ⋅ Peihao Wang ⋅ Kaipeng Zhang ⋅ Yiran Huang ⋅ Zhiwei Deng ⋅ Xindi Wu ⋅ George Cazenavette ⋅ Yuzhang Shang ⋅ Justin CUI ⋅ Jindong Gu ⋅ Qian Zheng ⋅ Hao Ye ⋅ Shuo Wang ⋅ Xiaobo Wang ⋅ Yan Yan ⋅ Angela Yao ⋅ Mike Zheng Shou ⋅ Tianlong Chen ⋅ Hakan Bilen ⋅ Baharan Mirzasoleiman ⋅ Manolis Kellis ⋅ Konstantinos N Plataniotis ⋅ Zhangyang "Atlas" Wang ⋅ Bo Zhao ⋅ Yang You ⋅ Kai Wang

Dataset distillation has provided a reliable solution for data compression, where models trained on the resulting smaller synthetic datasets achieve performance comparable to those trained on the original datasets. To further improve the performance of synthetic datasets, various training pipelines and optimization objectives have been proposed, greatly advancing the field of dataset distillation. Recent decoupled dataset distillation methods introduce soft labels and stronger data augmentation during the post-evaluation phase and scale dataset distillation up to larger datasets. However, this raises a question: Is accuracy still a reliable metric to fairly evaluate dataset distillation methods? Our empirical findings suggest that the performance improvements of these methods often stem from additional techniques rather than the inherent quality of the images themselves, with even randomly sampled images achieving superior results. Such misaligned evaluation settings severely hinder the development of DD. Therefore, we propose DD-Ranking, a unified evaluation framework, along with new general evaluation metrics to uncover the true performance improvements achieved by different methods. By refocusing on the actual information enhancement of distilled dataset, DD-Ranking provides a more comprehensive and fair evaluation standard for future research advancements.


Debiased DPO for Diffusion Models

Duc Khiem Pham ⋅ Quang Nguyen ⋅ Tung D Nguyen ⋅ Jingsen Zhu ⋅ Michele Santacatterina ⋅ Dimitris Metaxas ⋅ Ramin Zabih

Direct Preference Optimization (DPO) enables offline alignment of diffusion models from preference pairs without explicit reward modeling, but scaling DPO requires large amounts of expensive human preference data. In this paper, we consider a semi-supervised approach where only a small subset of preference pairs is annotated by humans, while the remaining (much larger) pool is unlabeled and annotated by inexpensive synthetic feedback e.g., vision-language models scoring image pairs or self-training from the current model. However, such synthetic supervision is not a replacement for human judgments: it can be systematically misaligned, so naively mixing human and synthetic preferences will yield misaligned inference. We introduce DeDPO, a debiased objective that integrates a doubly robust estimator from causal inference into DPO to correct synthetic-label misalignment while preserving the simplicity of DPO. Experiments show improved label efficiency and robustness to synthetic labeler quality, allowing DeDPO to outperform standard DPO under the same human-label budget and to approach fully human-supervised performance, despite using four times fewer human-labeled data.


DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs

Md Hasebul Hasan ⋅ Krity H Charu ⋅ Eshwara Prasad Sridhar ⋅ Shuchisnigdha Deb ⋅ Mohammad A Islam

Effective de-escalation is critical for law enforcement safety and community trust, yet traditional training methods lack scalability and realism. While Large Language Models (LLMs) enable dynamic, open-ended simulations, their substantial computational footprint renders them impractical for deployment on the lightweight, portable hardware required for immersive field training. Small Language Models (SLMs) offer a viable real-time alternative but suffer from a critical scarcity of high-quality, domain-specific training data. To bridge this gap, we present DeEscalWild, a novel benchmark dataset curated from a multi-stage pipeline of in-the-wild police-civilian interactions extracted from publicly available video repositories. Starting with 5,000 raw inputs, we employed a rigorous hybrid filtering process combining human-in-the-loop verification with LLM-as-a-Judge evaluation to distill 1,500 high-fidelity scenarios. The resulting corpus comprises 285,887 dialogue turns, totaling approximately 4.7 million tokens. Extensive experiments demonstrate that SLMs fine-tuned on this data significantly outperform their base counterparts across ROUGE-L, BLEU-4, METEOR, BERTScore, Realism Score, and human evaluation metrics. Notably, our fine-tuned Qwen 2.5 (3B-Instruct) surpasses the general-purpose Gemini 2.5 Flash model when evaluated under equivalent conditions, demonstrating that domain-optimized SLMs can achieve superior performance with a fraction of the computational cost. This work establishes the foundational infrastructure for accessible, low-latency, and privacy-preserving officer training systems at the edge. We publicly release our code and dataset.


Differentiable Range-Partition Entropy for Entropy-Sensitive Geometric Algorithms

Ibne Farabi Shihab ⋅ Sanjeda Akter ⋅ Anuj Sharma

Range-partition entropy is a structural complexity measure used in entropy-sensitive computational geometry. We introduce Differentiable Entropy Regularization (DER), a gradient-optimizable surrogate for the entropy of learned geometric partitions. The paper makes three main contributions. First, we separate the normalized entropy optimized by DER, denoted Hnorm(S), from the runtime-scaled entropy Hwork(S) = n · H_norm(S) that appears in algorithmic work bounds. Second, for a fixed margin-separated halfspace arrangement, we prove a self-contained soft-to-hard approximation theorem: the soft halfspace-cell entropy converges to the corresponding hard sign-pattern entropy at a rate controlled by margin, temperature, and the number of separators. A connection to optimal range-partition entropy holds when the fixed partition is admissible for the geometric problem and near-optimal in the admissible partition class. Third, we evaluate DER as learned geometric preprocessing for SciPy/Qhull-based convex-hull and Delaunay pipelines. On the evaluated 2D tasks, DER gives up to 4×–5× solver-side speedups and smaller end-to-end gains after preprocessing, with convex-hull area error below 0.2%. We also include a secondary transfer study in ViT-style attention, where the same regularizer induces structured sparsity and improves empirical efficiency. The central contribution is not the entropy formula alone, which resembles soft clustering, but the range-aware differentiable construction, the explicit approximation analysis, and the geometry-first training pipeline.


Distributionally Robust Token Optimization in RLHF

Yeping Jin ⋅ Jiaming Hu ⋅ Yannis Paschalidis

Large Language Models (LLMs) tend to respond correctly to prompts that align well with the data they were trained and fine-tuned on. Yet, small shifts in wording, format, or language can trigger surprisingly large failures, especially on multi-step reasoning problems. To address this problem, we propose a $\textbf{Distributionally Robust Token Optimization (DRTO)}$ approach, which combines token-level Reinforcement Learning from Human Feedback (RLHF) with $\textit{Distributionally Robust Optimization (DRO)}$. DRTO constructs f-divergence ambiguity sets over span-level actor losses, providing a principled way to emphasize difficult response segments during policy optimization. Empirically, DRTO enhances consistency under distribution shifts in multiple reasoning benchmarks among different tasks.


Dominant-Layer ZO: A Single Layer Dominates Zeroth-Order Fine-Tuning of LLMs

Wanhao Yu ⋅ Ziyan Wang ⋅ Zheng Wang ⋅ Abeer M Almalky ⋅ Yihang ZUO ⋅ Shuteng Niu ⋅ Sen Lin ⋅ Adnan Siraj Rakin ⋅ Deliang Fan ⋅ Li Yang

Zeroth-order (ZO) optimization enables memory-efficient fine-tuning of large language models (LLMs) using only forward passes, but it remains unclear how useful adaptation is distributed across layers. In this work, we reveal a surprising phenomenon: ZO fine-tuning is sharply dominated by a single decoding layer. Across multiple LLM families and downstream tasks, fine-tuning this dominant layer alone consistently matches or even exceeds full-model ZO fine-tuning. We further show that the dominant layer is task-agnostic but model-specific, and can be identified before training through a simple inference-only analysis of activation outliers. Specifically, the dominant layer consistently aligns with the first activation-outlier layer in the pre-trained model. To explain this phenomenon, we analyze how perturbation effects propagate under ZO optimization. We find that the dominant layer combines two key properties: high perturbation sensitivity and early placement in the residual stream, allowing perturbation-induced effects to propagate and accumulate through remaining subsequent decoding layers. As a result, this layer produces disproportionately strong and stable optimization signals under forward-only updates. Extensive experiments on LLaMA2-7B and Qwen3-8B across nine benchmarks show that dominant-layer ZO fine-tuning improves average performance over full-model MeZO and LoRA-based ZO fine-tuning while achieving up to 4.52$\times$ training speedup.


Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness

Songyuan Sui ⋅ Zhen Tan ⋅ Mohan Zhang ⋅ Rana M Khan ⋅ Xia Hu ⋅ Tianlong Chen

Missing modality remains a longstanding challenge in multimodal learning. Existing methods address this issue through modality recovery or adaptive strategies. However, they overlook models' internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically characterize a counterintuitive deployment-time failure mode: models trained on full modalities can underperform unimodal models when one modality is missing at inference time. This pattern appears across diverse architectures, such as fusion models, CLIP-style two-tower models, and vision-language models. We show that such degradation is closely associated with learned cross-modal dependencies in the principal parameter subspaces. Multimodal training induces structured rotations of these subspaces, particularly in cross-modal interaction layers. These rotations are associated with reduced task-aligned margins and larger task-aware representation harm under missing-modality inputs. We propose Geodesic Unlearning (GU), a lightweight parameter-editing method that leverages Grassmannian subspace geometry for structured subspace correction to improve missing-modality robustness. It decomposes layer weights into principal and residual components, treats the principal subspace as a point on the Grassmannian, and rotates it toward a unimodal reference along a geodesic path. Experiments across architectures and datasets show that GU improves performance under missing-modality inference while preserving full-modality accuracy, outperforming strong missing-modality robustness baselines. These findings support a geometric view of deployment-time missing-modality degradation and suggest localized subspace editing as a practical route for robustness correction.


Do Reasoning LLMs Refuse What They Infer in Long Contexts?

Yu Fu ⋅ Haz Sameen Shahgir ⋅ Zhipeng Wei ⋅ Huanli Gong ⋅ N. Benjamin Erichson ⋅ Yue Dong

Long-context LLMs can infer objectives that are not stated explicitly. This capability is useful for reasoning over documents, code, retrieved evidence, and tool traces, but it also creates a safety risk: harmful intent can be distributed across a context and become visible only after the model composes the relevant pieces. Existing safety evaluations mostly test explicit harmful requests, and therefore miss this failure mode. We introduce compositional reasoning attacks, a long-context threat model in which harmful requests are decomposed into semantically incomplete fragments and embedded in long contexts. The final query is neutral; the harmful objective emerges only if the model retrieves the fragments, composes them, and infers the implied goal. We instantiate this setting using AdvBench requests, varying the required reasoning from Direct Retrieval to Single-hop Aggregation, Chain Reasoning, and Multi-hop Deductive Reasoning, and evaluate 15 frontier LLMs on contexts up to 64k tokens. Models usually refuse harmful requests when they are directly retrievable. However, refusal rates drop sharply when the same objectives must be reconstructed compositionally, often with larger failures in longer contexts. Benign reconstruction and fragment-position analyses indicate that these failures are not mainly retrieval errors: models often infer the harmful objective and then comply. Increasing inference-time reasoning improves refusal but remains incomplete and costly. Our results reveal a long-context safety gap: current models are better at refusing harmful requests they see than harmful objectives they infer.


Drift Flow Matching

Chenrui Ma ⋅ Xi Xiao ⋅ Lin Zhao ⋅ Tianyang Wang ⋅ Ferdinando Fioretto ⋅ Yanning Shen

Iterative generative models such as Flow Matching and Diffusion models have demonstrated strong test-time scaling behavior, where additional inference computation can improve generation quality. In contrast, Drift Models offer efficient one-step generation, but their direct generation paradigm limits such flexibility. In this work, we propose Drift Flow Matching (DFM), a framework that connects drifting generative modeling with flow-based iterative generation. DFM preserves the efficiency of direct transport maps while enabling generation to be refined through multiple inference steps when desired. This bridges the gap between one-step Drift Models and multi-step Flow Matching methods, and provides a novel generative paradigm that can adapt sampling computation to different quality--efficiency requirements. Extensive experiments across different tasks and datasets demonstrate the effectiveness and generality of the proposed framework.


DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts

Jiading Gai ⋅ Shuai Zhang ⋅ Xiang song ⋅ Yuyang (Bernie) Wang ⋅ George Karypis

Modern RL post-training methods such as GRPO and DAPO train on $N$ response sequences of $R$ tokens sampled from a shared prompt of $P$ tokens, but standard FlashAttention replicates all $P$ prompt tokens $N$ times across both forward and backward passes --- duplicating compute and memory on identical hidden states. In large-rollout, long-context RL training ($N{\geq}16$, $P{\geq}8\text{K}$), this redundancy dominates the policy update cost. We observe that in decoder-only models, causal masking makes prompt representations invariant across sequences at every layer, so all per-token operations (norms, projections, MLP) and attention can process the prompt once --- a property not yet exploited at the kernel level for training. We propose \textbf{DualKV}, the first FlashAttention kernel variant that eliminates shared-prompt replication during RL training, via (1) fused CUDA forward and backward kernels that iterate over two disjoint KV regions --- shared context and per-sequence response --- in a single kernel launch, and (2)~a data-pipeline redesign in veRL that repacks $N(P{+}R)$ tokens into $P{+}NR$ tokens per micro-batch, extending the token reduction from attention to the entire model by a factor $\rho = N(P{+}R)/(P{+}NR)$. DualKV is mathematically equivalent to standard attention and introduces no approximation. On Qwen3-8B GRPO training with 8$\times$H100 GPUs ($N{=}32$, 8K-context), DualKV achieves $1.63$--$2.10\times$ policy-update speedup, enables $2\times$ larger micro-batches, and raises MFU from $36\%$ to $76\%$. Similar gains hold for DAPO ($2.47\times$ speedup, $77\%$ MFU). At 30B MoE scale on 16$\times$H100, DualKV achieves $3.82\times$ policy-update and $3.38\times$ end-to-end step speedup over FlashAttention (which requires 4-way Ulysses sequence parallelism to avoid OOM).

Video generation must account for two sources of motion, one induced by the observer's camera path and the other caused by scene dynamics. An ideal camera-controlled video model should account for both motions: let users move the camera while evolving the scene dynamics. While current models handle camera-induced motion well in static settings, they struggle for dynamic scenes: objects are static, move incorrectly, or degrade in generation quality. We introduce DynaTokens, a lightweight set of learnable scene-specific tokens that teach dynamics to an existing camera-controlled world model. Our method is motivated by a simple asymmetry between the two sources of motion: whereas camera motion affects the generated view globally, object dynamics are spatially localized. Through cross-attention, DynaTokens trains the learnable tokens from a few example trajectories for a scene while keeping the base model frozen, and enables dynamics under new query camera paths. DynaTokens achieves a better simultaneous dynamics-camera tradeoff on VBench2 and WorldScore evaluations than LoRA, block finetuning, and specialized trainable-layer baselines. Analyses of token attention, ablations, and motion temporality suggest that matching the trainable interface to the structure of the learning target is important for effective adaptation.

Parallel test-time scaling (TTS) has been shown to enhance reasoning in large language models by generating and aggregating multiple independent reasoning paths. Internal confidence scores can further improve this process by identifying paths of varying quality. However, existing methods remain computationally inefficient, as low-quality paths are fully expanded before being evaluated. In this work, we aim to improve the use of internal confidence signals for better model performance and computational efficiency. We begin by offering a new perspective on the role of confidence in aggregation, showing that it is most impactful for questions near the decision boundary. Notably, when amplified by multi-sample aggregation, even small confidence signals derived from short prefixes can meaningfully influence outcomes in these boundary cases. Building on this insight, we introduce Prefix-Guided Sampling (PreG), a method that reallocates computation by first generating short prefixes, ranking them using internal confidence scores, and then completing only the most promising candidates. This strategy reduces token usage while maintaining gains in path quality. We provide theoretical analysis demonstrating that, under a fixed compute budget, PreG strengthens the effective signal by widening the decision boundary margin. Empirically, our method consistently outperforms standard parallel TTS approaches across benchmarks, while significantly improving token efficiency and reducing overall computation.


Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions

Zhenting Qi ⋅ Ao Qu ⋅ Huangyuan Su ⋅ Chenyu Wang ⋅ Yu Yao ⋅ Han Zheng ⋅ Kushal Chattopadhyay ⋅ Guowei Xu ⋅ Zihan Wang ⋅ Weirui Ye ⋅ Vijay Janapa Reddi ⋅ Ju Li ⋅ Paul Liang ⋅ Himabindu Lakkaraju ⋅ Sham Kakade ⋅ Yilun Du

How can a population of agents self-orchestrate and self-adapt into stronger collective intelligence without centralized control? Inspired by Friedrich Hayek's economic theory of decentralized coordination in markets, we study this question through an agent economy in which agents compete via auctions for the right to act, exchange payments, and accumulate wealth from environmental rewards. These simple economic signals induce decentralized credit assignment, driving planning without global orchestration or explicit communication protocols. The population evolves through economic selection: effective agents accumulate wealth and are mutated via exploitation, while ineffective ones go bankrupt and are replaced via exploration. We show that, initialized with weak agents, the economy produces emergent multi-step reasoning strategies and outperforms stronger monolithic baselines across five agentic tasks, including mathematical reasoning, financial research, scientific research, accelerator design, and distributed-system optimization. We further provide theoretical insights into how economic dynamics shape agent behaviors, linking local incentives to long-term global performance. Our results suggest a new path to multi-agent intelligence: rather than engineering coordination, we can design decentralized incentive structures under which it automatically emerges.


Efficient Dynamic Algorithms for Graph Neural Networks with Non-Linearity

Kiarash Banihashem ⋅ MohammadTaghi Hajiaghayi ⋅ Mahdi JafariRaviz ⋅ Silvio Lattanzi ⋅ Danny Mittal

Graph Neural Networks (GNNs) are widely used for representation learning on graphs, but most methods assume static topologies, making them inefficient on evolving networks where edges change over time. Existing dynamic approaches either model graph evolution through temporal GNN architectures without focusing on efficient dynamic maintenance, or are restricted to linear propagation models based on Personalized PageRank. In this work, we study how to efficiently maintain node representations for non-linear GNN dynamics under edge insertions and deletions. For a broad class of standard activation functions, we develop a residual-based dynamic algorithm that selectively propagates local errors via push operations, maintaining an approximation to the evolving fixed point without full recomputation. We prove that our method achieves amortized $O(1/\epsilon)$ update time per graph change under a degree-normalized error guarantee. Our approach uses a potential-based analysis in a degree-scaled norm and, in contrast to prior work on the linear case, requires no randomness assumptions on either the update sequence or the input vector. For the linear special case, we additionally provide an exact dynamic algorithm via low-rank matrix inverse updates. Experiments on benchmark datasets show that incorporating non-linearity improves accuracy while preserving efficient update performance, yielding a scalable and theoretically grounded framework for dynamic GNNs.


EgoSurgHands: An Egocentric 3D Hand Pose Dataset & Benchmark for Open-Surgery Training

Ankit Pal ⋅ Guillaume Kugener ⋅ Sofia Volpi ⋅ Joshua Schwartz ⋅ Alex Idarraga ⋅ Pranav Rajpurkar ⋅ Gabriel A Brat

We introduce EgoSurgHands, an egocentric 3D hand-pose dataset and benchmark for open-surgery training tasks captured with Meta Project Aria glasses on standard surgical training pads. To our knowledge, EgoSurgHands is the first public benchmark for surgical training that is simultaneously egocentric, sensor-grounded, and human-validated. EgoSurgHands contains 41,909 training hand-instance rows across 55 procedure recordings and 9,476 inter-annotator-agreement (IAA)-validated rows across 13 held-out procedure recordings from 4 wearers, 13 distinct tasks, 3 glove colors, and 3 stages of training. Each frame includes Aria intrinsics, SLAM extrinsics, MPS (Machine Perception Services) landmarks, per-joint confidence, timestamps, and eye gaze. Validation 3D keypoints are initialized from Aria MPS and independently corrected by four expert annotators under a multi-rater IAA workflow, yielding evaluation ground truth independent of any tested model's training signal. Using EgoSurgHands, we benchmark six state-of-the-art hand reconstructors under zero-shot transfer: HaMeR, WiLoR, WildHands, JointTransformer, SimpleHand, and MeshGraphormer. Ego-pretrained models consistently outperform non-egocentric baselines: WildHands improves over the non-ego HaMeR reference by 4.51 mm PA-MPJPE, while JointTransformer achieves the best zero-shot result at 14.95 mm PA-MPJPE. These results show that generalist egocentric pretraining transfers to surgical egocentric tasks, but also reveal a substantial remaining surgical-domain gap. As a lightweight diagnostic adaptation experiment, a 58K-parameter two-headed (pose + translation) residual corrector further reduces error on frozen backbones, reaching 10.40 mm PA-MPJPE and 75.8 mm absolute MPJPE on WildHands ($-$49.2\% and $-$85.8\% vs zero-shot). We release the dataset, benchmark code, checkpoints, and multi-rater annotation framework at [URL withheld for double-blind review].


ELPAC: Endpoint-Anchored Latent Progression with Stage-Varying Multimodal Coordination

Shengxian Ding ⋅ Yifei Zhang ⋅ Xinyuan Tian ⋅ Yuanshuang Guo ⋅ Zhao

Aging cohorts increasingly combine imaging, plasma, and cognitive measurements, but most subjects provide only a partially observed cross-sectional snapshot, leaving latent stage, progression subtype, and stage-varying multimodal coordination to be inferred jointly. Existing methods estimate the cohort-level aging process, model covariance under observed covariates, or fuse modalities for prediction. However, they rarely target the joint problem of endpoint-anchored staging and progression-resolved residual coordination across modalities, even though subject staging, subtype assignment, and multimodal interplay are mutually dependent. We introduce ELPAC, a Bayesian framework with two interlocking constructions. The first integrates subtype trajectories backward from clinically reliable endpoint groups along biologically structured velocity components encoding modality-specific monotone direction and optional sequential ordering priors. The second factorizes residual multimodal variation around the inferred stage axis into sparse coordination axes whose loadings and activations expose which multimodal patterns are operative in each subtype and when they emerge, with stage-resolved variation supplying the auxiliary structure for conditional identifiability. We use a scalable amortized variational inference scheme that enables joint posterior inference at cohort scale with single-pass deployment on new subjects. On synthetic data and a harmonized ADNI+NACC/SCAN Alzheimer's cohort, ELPAC demonstrates accurate latent-structure recovery, well-calibrated prediction intervals, and biologically interpretable multimodal coordination programs that differ by subtype.


Enabling Preference-driven Unlearning in Few-step Distilled Text-to-Image Diffusion Models

Gaurav Patel ⋅ Jun Fang ⋅ Greg Ver Steeg ⋅ Qiang Qiu ⋅ Sravan Sripada

Text-to-image (T2I) diffusion models are increasingly distilled into few-step variants and being deployed to enable fast inference. However, their ability to generate harmful or undesired content poses significant safety risks. Data-driven unlearning methods suppress targeted generations by fine-tuning model weights using specialized unlearning objectives. Crucially, these objectives implicitly rely on multi-step denoising dynamics, an assumption that breaks down for few-step distilled (FSD) models, resulting in ineffective forgetting. Furthermore, performing unlearning on the non-distilled base model and subsequently re-distilling it to obtain an unlearned FSD model incurs substantial computational and time overhead, making it impractical in many settings. Hence, we address this limitation with a preference-driven unlearning framework that revisits Direct Preference Optimization (DPO) for diffusion models. We show that standard DPO and its unlearning derivatives, formulated around noise-prediction error, transfer poorly to FSD models due to their altered generation dynamics. To overcome this, we introduce a modified preference optimization formulation explicitly aligned with the few-step generation properties, $\textit{enabling direct concept removal in FSD models}$ while preserving few-step efficiency and maintaining strong retention of desirable (non-targeted) capabilities. We evaluate our framework primarily on identity and NSFW (nudity) removal tasks and also extend our method to object-level unlearning. Extensive experiments demonstrate consistent and effective forgetting, and strong retention performance, establishing our method as a practical and principled solution for unlearning in FSD models.

Protein structure tokenizers (PSTs) are workhorses in protein language modeling, function prediction, and evolutionary analysis. However, existing PSTs only capture local geometry of static structures, and miss the correlated motions and alternative conformational states revealed by protein ensembles. Here we introduce Ensembits, the first tokenizer of protein conformational ensembles. Ensembits address challenges inherent to tokenizing dynamics: deriving informative geometric descriptors across conformations, permutation-invariance encoding of variable-size ensembles, and conquering sparsity in dynamics data. Trained with a Residual VQ-VAE using a frame distillation objective on a large molecular dynamics corpus, Ensembits outperforms all related methods on RMSF prediction, and is the strongest standalone structural tokenizer on an token-conditioned ANOVA test on per-residue motion amplitude. Ensembits further matches or exceeds static tokenizers on EC, GO, binding site/affinity prediction, and zero-shot mutation-effect prediction despite using far less pretraining data. Notably, the distillation objective enables Ensembits to predict dynamics token from one single predicted structure, which alleviates dynamics data sparsity. As the field moves from static structure prediction toward ensemble generation, Ensembits offer the discrete vocabulary needed to bring dynamics into protein language modeling and design. Our codes can be found at: https://anonymous.4open.science/r/Ensembits.


EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation

Ruozhen He ⋅ Meng Wei ⋅ Ziyan Yang ⋅ Vicente Ordonez

Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and locations across shots remains a challenge over long sequences. Existing evaluations typically use independently generated prompt sets with limited entity coverage and simple consistency metrics, making standardized comparison across methods difficult. We introduce EntityBench, a benchmark consisting of 140 episodes (2,491 shots) derived from real narrative media, with explicit per-shot entity schedules tracking characters, objects, and locations simultaneously across easy, medium, and hard difficulty tiers of up to 50 shots, 13 cross-shot characters, 8 cross-shot locations, 22 cross-shot objects, and recurrence gaps spanning up to 48 shots. EntityBench pairs the dataset with a three-pillar evaluation framework that disentangles intra-shot visual quality, prompt-following alignment, and cross-shot entity consistency. Cross-shot consistency, the central pillar, evaluates each recurring entity through both embedding similarity and LLM per-criterion judgment across entity-type-specific dimensions, with a fidelity gate that admits accurate entity appearance. To establish baselines, we propose EntityMem, a memory-augmented generation system that plans and stores verified per-entity visual references in a persistent memory bank before generation begins, enabling the video backbone to retrieve each entity's appearance across shots. Experiments on EntityBench show that cross-shot entity consistency degrades sharply with recurrence distance in existing methods, and that explicit per-entity memory yields the highest character fidelity (Cohen's d = +2.33) and presence among methods evaluated.


Evaluating Deployable Inference-Time Error Prediction in Vision MoEs

Aristeidis Tsaris ⋅ JangHyeon Lee ⋅ Abhishek Potnis ⋅ Philipe Dias ⋅ Waqwoya Abebe ⋅ Dan Lu ⋅ Feiyi Wang ⋅ Dalton Lunga

Sparse mixture-of-experts (MoE) models raise a natural inference-time question: can experts not selected by the router improve predictions or estimate their errors without retraining? We study this question in vision MoEs across five RGB and multispectral classification datasets and multiple expert counts. We evaluate additional expert inference for three outcomes: top-1 accuracy, raw ranking of default router errors, and calibrated prediction of whether the router-selected classification is wrong. An upper-bound analysis shows that non-selected experts often contain useful class signal when the routed path fails. However, this signal does not reliably improve deployable top-1 accuracy, and raw multi-forward uncertainty scores yield only limited gains for ranking router errors. The strongest benefit appears in calibrated router-error prediction: post-hoc models that combine router signals with features from non-selected experts reduce Router-error Expected Calibration Error (ECE) by roughly 10-13% on average, with several settings exceeding 20% relative improvement over a calibrated router-only baseline. Overall, additional inference-time expert usage in vanilla vision MoEs is most useful for estimating the reliability of the routed decision, rather than directly improving classification accuracy.


Evaluating Whether LLMs Can Reliably Connect the DOTs?

Eftekhar Hossain ⋅ John Salvador ⋅ Santu Karmaker

Access to real-world information is often noisy and fragmented. Constructing a coherent narrative from such fragments requires models to reconstruct missing spans within a broader storyline, commonly referred to as text infilling, while preserving consistency with both the local context and the global storyline. Despite using text infilling as a pre-training objective in many Large Language Models (LLMs), their actual performance on real-world narrative infilling remains underexplored. In this paper, we address this gap by introducing a multi-domain benchmark of ∼9.2K instances for narrative infilling, constructed by masking one to three sentences across four narrative types: encyclopedic text, commonsense stories, news articles, and visual narratives. Using this benchmark, we evaluate 20 instruction-tuned open-source LLMs ranging from 1.5B to 70B parameters across varying levels of instruction specificity and reasoning guidance. Outputs are assessed using standard automatic metrics and a qualitative framework covering five narrative dimensions. Results show that model scale does not reliably predict infilling quality: Gemma-2-2B achieves the highest qualitative score (4.02/5), outperforming models over ten times larger, including DeepSeek-Qwen-32B (3.77/5, ↓6.6%) and LLaMA-3.3-70B (3.71/5, ↓8.3%). We further find that explicit reasoning offers limited benefits as chain-of-thought reasoning yields only a marginal improvement (+0.6%). Additionally, short narratives and domain characteristics emerge as stronger predictors of task difficulty than infill position alone for narrative infilling in current LLMs.


EVALUATION CARDS: An Interpretive Layer for AI Evaluation Reporting

Avijit Ghosh ⋅ Anka Reuel-Lamparth ⋅ Wm. Matthew Kennedy ⋅ Jenny Chim ⋅ Yilin Huang ⋅ Srishti Yadav ⋅ Anastassia Kornilova ⋅ Damian Stachura ⋅ Kevin Klyman ⋅ Jennifer Mickel ⋅ Felix Friedrich ⋅ Max Lamparth ⋅ Jan Batzner ⋅ Anoop Mishra ⋅ Jeba Sania ⋅ Eliya Habba ⋅ Nathan Heath ⋅ Shalaleh Rismani ⋅ Usman Gohar ⋅ Andrea Loehr ⋅ David Manheim ⋅ Ruchira Dhar ⋅ Sree Harsha Nelaturu ⋅ Aarush Sinha ⋅ Leshem Choshen ⋅ Yixiong Hao ⋅ Yanan Long ⋅ Andrew Tran ⋅ Drishti Sharma ⋅ Ishan Khire ⋅ Amit Saha ⋅ Subramanyam Sahoo ⋅ Michael Hardy ⋅ Michael A. Riegler ⋅ Kabir Manghnani ⋅ Michelle Lin ⋅ Nancy Jiang ⋅ Asaf Yehudai ⋅ Nuno Moniz ⋅ Yacine Jernite ⋅ Zeerak Talat ⋅ Stella Biderman ⋅ Mubashara Akhtar ⋅ Jessica Ji ⋅ Aris Hofmann ⋅ Mykel J Kochenderfer ⋅ Sanmi Koyejo ⋅ Irene Solaiman

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. This fragmentation is borne at the level of interpretation: readers cannot reliably compare results across sources, identify what is missing from a given report, or trace an aggregate claim to the evidence behind it. Recent efforts address isolated components of this problem but leave three gaps unresolved: They cover only narrow slices of the evaluation lifecycle and do not compose into a single interpretable record. They specify static representations that do not differentiate between the questions technical and policy readers bring to the same evidence. And they remain proposals on paper, without the extraction infrastructure required for adoption at scale. We present EVALUATION CARDS, an operational reporting layer that composes benchmark metadata, evaluation run data, and model metadata into a unified record. We (1) derive a reporting schema from a structured literature review of 52 papers and 10 stakeholder interviews, (2) implement four interpretive signals (reproducibility, documentation completeness, provenance and risk, and score comparability), rendered through reader modes calibrated to research and policy audiences, and (3) provide a deployed monitoring tool that applies EVALUATION CARDS across 5,498 models, 635 benchmarks, and 101,843 results, surfacing systematic gaps in current reporting practice.


Explanations over Graphs: An Agent Architecture for IT Enterprise Diagnostic Tasks

Saurabh Jha ⋅ Rohan R. Arora ⋅ Bhavya Bhavya ⋅ Noah Zheutlin ⋅ Paulina Toro Isaza ⋅ Laura Shwartz ⋅ Yu Deng ⋅ Daby Sow ⋅ Ruchir Puri

LLM agents excel when environments are mostly static and their required information fits in a model’s context window, but they struggle with IT enterprise diagnostic tasks—such as incident management, IT vulnerability analysis, and FinOps anomaly explanation—where operators iteratively mine massive observability data over a curated resource graph to identify the origins that explain observed symptoms, enabling correct remediation. The data in these domains inherently has hidden dependency structure: entities interact, signals co-vary, and the importance of a fact may only become clear after other evidence is discovered. To cope with bounded context windows, agents must summarize intermediate findings before their significance is known, increasing the risk of discarding key evidence. ReAct-style agents are especially brittle in this regime. Their retrieve-summarize-reason loop makes conclusions sensitive to exploration order and introduces run-to-run non-determinism, producing a reliability gap where Pass-at-k may be high but Majority-at-k remains low. Simply sampling more roll-outs or generating longer reasoning traces does not reliably stabilize results, since verifying a diagnosis requires remediation actions that carry operational risk---demanding consistent correctness, not occasional success---and ReAct provides no mechanism for belief revision as evidence accumulates. In addition, ReAct entangles semantic reasoning with controller duties such as tool orchestration and state tracking; execution errors and plan drift degrade reasoning while consuming scarce context. We address these issues by formulating the investigation as abductive reasoning over a dependency graph and proposing EoG (Explanations over Graphs), a disaggregated framework where an LLM performs bounded local evidence mining and labeling (cause vs symptom) while a deterministic controller manages traversal, state, and belief propagation to compute a minimal explanatory frontier. On representative ITBench diagnostics tasks, EoG improves accuracy and run-to-run consistency over ReAct baselines, including a 7x average gain in Majority-at-k F1 score.


ExploitGym: Can AI Agents Turn Security Vulnerabilities into Real Attacks?

Zhun Wang ⋅ Nico Schiller ⋅ Hongwei Li ⋅ Srijiith S Narayana ⋅ Milad Nasr ⋅ Nicholas Carlini ⋅ Xiangyu Qi ⋅ Eric Wallace ⋅ Elie Bursztein ⋅ Luca Invernizzi ⋅ Kurt Thomas ⋅ Yan Shoshitaishvili ⋅ Wenbo Guo ⋅ Jingxuan He ⋅ Thorsten Holz ⋅ Dawn Song

AI agents are rapidly gaining capabilities that could significantly reshape cybersecurity, making rigorous evaluation urgent. A critical capability is exploitation: turning a vulnerability, which is not yet an attack, into a concrete security impact, such as unauthorized file access or code execution. Exploitation is a particularly challenging task because it requires low-level program reasoning (e.g., about memory layout), runtime adaptation, and sustained progress over long horizons. Meanwhile, it is inherently dual-use, supporting defensive workflows while lowering the barrier for offense. Despite its importance and diagnostic value, exploitation remains under-evaluated. To address this gap, we introduce ExploitGym, a large-scale, diverse, realistic benchmark on the exploitation capabilities of AI agents. Given a program input that triggers a vulnerability, ExploitGym tasks agents with progressively extending it into a working exploit. The benchmark comprises 898 instances sourced from real-world vulnerabilities across three domains, including userspace programs, Google's V8 JavaScript engine, and the Linux kernel. We vary the security protections applied to each instance, isolating their impact on agent performance. All configurations are packaged in reproducible containerized environments. Our evaluation shows that while exploitation remains challenging, frontier models can successfully exploit a non-trivial fraction of vulnerabilities. For example, Claude Mythos Preview, Anthropic's latest model, produces working exploits for 208 instances. Notably, even with widely used defenses enabled, models retain non-trivial success rates. These results establish ExploitGym as an effective testbed for exploitation and highlight the growing cybersecurity risks posed by increasingly capable AI agents.


Exploring Advertising Manipulation in Diffusion Image Generation

Tianshi Che ⋅ Yang Zhou ⋅ Yushan Mu ⋅ Zeru Zhang ⋅ Zijie Zhang ⋅ Panqiu Xia ⋅ Wei Zhang ⋅ Longwei Wang ⋅ yelong shen ⋅ Ruoming Jin ⋅ Jianfeng Gao

As text-to-image diffusion models (T2I DMs) become widely deployed, adversarial advertising has emerged as a realistic threat: an attacker may compromise a T2I DM so that it implants target product brands into images generated from users' non-advertising prompts. Two key challenges remain largely unresolved in this setting: achieving natural and semantically coherent adversarial advertisement and ensuring robust adversarial advertisement. To address these challenges, we develop a new estimation algorithm for the multivariate continuously scaled phase-type with Lévy (MCPHL) distribution to capture the intrinsic distribution of natural advertisement prompts in the prompt-embedding space. With the estimated MCPHL, we construct an attack that pushes non-advertising prompts toward high-density regions of this distribution, making the resulting perturbed prompts better aligned with natural advertising prompts. We then propose a masked parameter smoothing approach grounded in mollification theory, yielding a smoothed T2I DM equipped with a dimension-invariant certified guarantee against advertisement degradation under model fine-tuning. The masking mechanism preserves utility by avoiding unnecessary smoothing on sensitive parameters, and our theoretical analysis shows that the smoothed model better preserves adversarial advertisements under fine-tuning, while maintaining better generation quality.


Express Language Modeling

Albert Gong ⋅ Annabelle M Carrell ⋅ Raaz Dwivedi ⋅ Lester Mackey

We introduce a new tool, Express, for converting a non-causal attention approximation into a causal approximation with matching approximation guarantees. When combined with the state-of-the-art Thinformer approximation, Express improves upon the best known causal attention guarantees, delivering $\log^{3/2}(n)/s$ approximation error with only $O(s)$ memory and $O(s^2 \log^2(n))$ compression overhead for a sequence of length $n$. We pair these developments with an efficient I/O-aware Triton implementation, demonstrate substantial speedups over FlashAttention 2, and use Express to accelerate three distinct components of the language modeling pipeline: long-context prefill, KV cache compression, and long-form decoding.


FASD: Hardware Acceleration for Multi-AI-Agent Discussion

Rui Chu ⋅ Kaiyuan Zhang ⋅ Antian Wang ⋅ Yingjie Lao

Recent development of large language models (LLMs) has empowered multi-agent systems (MAS) through external skills and inter-agent discussion to solve complex tasks. Although such a paradigm improves the capability of LLM systems, it also introduces substantial efficiency overhead. Existing approaches mainly reduce this overhead through software-level optimization. However, these methods overlook how to exploit the graph-level structure of MAS discussion and its suitability for hardware-aware acceleration, thus causing latency, especially when the number of agents, discussion rounds, or concurrent sessions increases. In this paper, we propose FASD, an FPGA-algorithm co-design framework for acceleration towards graph dynamic system agent discussion. We argue that MAS discussion can be formed as a graph-coupled dynamical system and thus agent activation pruning can be optimized through phase-coupled dynamics method. Based on these insights, we designed the FPGA-aware optimization for the problem. FASD mainly consists of the following parts. Firstly, we profile representative MAS workflows and construct a phase profile for dynamic system modeling. Secondly, the hardware-friendly optimization progress, given the profile, is implemented on FPGA. Through extensive experiments on multi-agent discussion workloads, we demonstrate that FASD improves discussion efficiency while preserving task performance, with stronger benefits as the agent graph becomes larger and more complex.

In this paper, we study federated optimization for solving stochastic variational inequalities (VIs), a problem that has attracted growing attention in recent years. Despite substantial progress, a significant gap remains between existing convergence rates and the state-of-the-art bounds known for federated convex optimization. In this work, we address this limitation by establishing a series of improved convergence rates. First, we show that, for general smooth and monotone variational inequalities, the classical Local Extra SGD algorithm admits tighter guarantees under a refined analysis. Next, we identify an inherent limitation of Local Extra SGD, which can lead to excessive client drift. Motivated by this observation, we propose a new algorithm, the Local Inexact Proximal Point Algorithm with Extra Step (LIPPAX), and show that it mitigates client drift and achieves improved guarantees in several regimes, including bounded Hessian, bounded operator, and low-variance settings. Finally, we extend our results to federated composite variational inequalities and establish improved convergence guarantees.


FedRSPO+: A Heterogeneity-aware Algorithm for Decision-focused Federated Learning

Konstantinos Ziliaskopoulos ⋅ Alexander Vinel ⋅ Jiaqi Wang

Decision-focused learning (DFL) trains predictive models for downstream optimization, but existing methods largely assume centralized data. In cross-silo settings, federated learning is a natural alternative, yet standard federated pipelines optimize prediction quality rather than decision quality and do not account for client heterogeneity in downstream objectives or feasible sets. Heterogeneity is especially problematic for DFL methods: under heterogeneous polyhedral decision problems, small perturbations can cause discontinuous changes in optimal decisions, leading to unstable client updates and aggregation. We propose FedRSPO+, a heterogeneity-aware framework for decision-focused federated learning. Our approach is built on RSPO+, a regularized predict-then-optimize surrogate that smooths the decision map through projection, enabling stable optimization even when clients face different objectives and constraints. We show that RSPO+ upper bounds downstream decision error and regret under mild assumptions, and that FedRSPO+ has cross-client heterogeneity bounds that (i) scale with both objective and feasible-set heterogeneity, (ii) vanish at homogeneity, and (iii) do not require strong convexity. We develop an annealed and modular training procedure that is compatible with standard federated personalization and aggregation methods. We evaluate FedRSPO+ on three experimental settings: controlled synthetic knapsack instances that isolate multiple axes of heterogeneity, a shortest path DFL benchmark, and a real-world PJM energy pricing case study. Across these experiments, we compare against prediction-only federated learning and DFL frameworks under varying heterogeneity and communication budgets. Together, our results suggest that smoothing is a useful ingredient for stable collaborative decision learning and provide a first heterogeneity-aware foundation for federated DFL.


Finding Interpretable Prompt-Specific Circuits in Language Models

Gabriel Franco ⋅ Lucas M Tassis ⋅ Azalea Rohr ⋅ Mark Crovella

Understanding the internal circuits that language models use to solve tasks remains a central challenge in mechanistic interpretability. A crucial part of finding circuits is understanding why each attention head attends where it does. To this end, we introduce ACC++, an improved circuit-tracing method based on the principle of attention-causal communication (ACC) [1], which identifies signals, i.e., contents of low dimensional subspaces that cause attention on a token pair. ACC++ extracts circuits from a single forward pass, without replacement models or patching. Circuits identified by ACC++ consist of components that are causal for the model's attention decisions, together with the low-dimensional signals used to communicate between them. Here, we first detail the conceptual advances that ACC++ makes over previous work. We then show that across multiple models, a substantial portion of ACC++ signals are interpretable: many signals admit a short natural-language description. We next present a number of new insights into model behavior obtained via ACC++. First, we use ACC++'s interpretable circuits to characterize the sensitivity of indirect object identification (IOI) circuits to prompt structure. We find that prompt-specific circuits form well-defined clusters, and across clusters, heads receive systematically different signals corresponding to distinct mechanisms for identifying the IO name. Next, in multilingual IOI, ACC++ circuits show that while model components are reused across languages, signals are often language-specific. In a four-language IOI case study, cross-language circuit distances are consistent with linguistic relatedness. Together, these results show that ACC++ can shed light on a broad spectrum of model behaviors.


Fine-Tuning Improves Information Conveyance in Language Models

Yuwei Cheng ⋅ Weiyi Tian ⋅ Haifeng Xu

Fine-tuning is often believed to reduce uncertainty and diversity in large language models, but existing analyses overlook output length, a key confounder, and therefore fail to capture how uncertainty is distributed across an entire generation rollout. To address this, we propose Canopy Entropy ($\mathrm{CE}^\star$), a measure that views language generation from a tree perspective, where "canopy" represents the space of all possible rollouts, making $\mathrm{CE}^\star$ naturally quantify the effective size of the generation space. $\mathrm{CE}^\star$ jointly captures the output length $N$ and the generated sequence $Y_{1:N}$, and we show that it equals to total Shannon entropy $H(N, Y_{1:N}\mid X)$, where $X$ denotes the prompt. This formulation yields interpretable metrics, including a length--uncertainty correlation term $\rho(N, r_N)$, where $r_N$ is the entropy rate, quantifying information conveyance efficiency by indicating whether longer outputs are more or less informative per token. Empirically, across tasks and model families, we find that fine-tuned models consistently exhibit stronger positive correlation $\rho(N, r_N)$, even when total entropy decreases. Furthermore, after controlling for model family, task, prompt, and output-length effects, we find that fine-tuning nearly triples the strength of the relationship between entropy rate and semantic diversity, suggesting that aligned models convert uncertainty into semantically meaningful variation much more efficiently. Overall, these results suggest that fine-tuning does not simply reduce uncertainty, but fundamentally reorganizes it into more informative and semantically meaningful generations.


Finite-Sample and Communication-Efficient Networked Information Aggregation

Mohammadhossein Bateni ⋅ Zahra Hadizadeh ⋅ MohammadTaghi Hajiaghayi ⋅ Mahdi JafariRaviz ⋅ Shayan Taherijam

Building on SODA 2026 work Kearns et al. (2026), we study learning in directed acyclic graphs in which each agent observes only a subset of the input coordinates and may use the predictions of its parents as additional features. In the finite-sample protocol, the message sent along each edge is the agent's prediction vector on a shared sample, so the communication cost is linear in the sample size $m$. In this paper we study this model under communication constraints. First, we revisit the finite-sample analysis of the full-message protocol. We identify a gap in the finite-sample proof of Kearns et al. (2026) and develop a different argument that establishes a finite-sample population-risk guarantee along paths satisfying deterministic $M$-coverage. Second, we introduce a shared-sketch protocol in which each agent sends a $k$-dimensional linear sketch of its prediction vector instead of the full $m$-dimensional message, where $k\ll m$. We analyze this protocol through a reduction to least squares on the compressed sample together with an oblivious subspace embedding argument. The resulting bound separates statistics from communication: the statistical term is still governed by the original sample size $m$, while the additional error due to compression is controlled by the sketch dimension $k$. For Gaussian sketches, we obtain an explicit high-probability population-risk bound, showing that row subsampling is not the only possible communication-statistics tradeoff for networked information aggregation.

The framework of Online Convex Optimization with Memory captures sequential decision-making settings where the learner's instantaneous loss is affected by their past choices, in addition to the most recent decision. We first give a new first-order regret bound in this framework for potentially non-smooth loss functions, with regret scaling as the square root of the loss of the best decision in hindsight, which can be much better than known regret bounds that explicitly scale with the horizon. In fact, our result holds in a generalized non-differentiable setting, which also captures other decision making problems of interest. We then give a faster gradient-based algorithm that yields an improved first-order regret bound for smooth loss functions. As an application, we consider the recently introduced problem of online nonstochastic control, which involves controlling a linear dynamical system subject to adversarial disturbances to minimize convex cost functions, where our results can be adapted to produce novel first-order regret bounds. This means whenever there is a benchmark controller with small loss in hindsight, our regret bounds are significantly smaller than those present in the literature.


FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

Xing Han ⋅ Shravan Chaudhari ⋅ Tanvi Ranade ⋅ Rama Chellappa ⋅ Suchi Saria

Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, tasks are co-available at design time where related tasks could borrow representational strength from one another, (2) continual adaptation, in which new tasks emerge after deployment with previously unseen modality combinations. However, neither regime alone suffices: the pretraining task set is never exhaustive, while bypassing joint training forfeits the transfer gains and efficiency among co-trainable tasks. Sparse Mixture-of-Experts (MoE) is a natural fit for this dual requirement: sparse activation enables modular capacity expansion as new tasks arrive, while routing decouples modality-level computation from task-level composition. In this work, we propose a scalable MoE framework for multitask pretraining and continual learning across flexible modality combinations. The framework is designed to support training on multimodal tasks with diverse modality configurations by leveraging modality-specific routers that process tokens from each modality across tasks. Furthermore, it enables continual learning over sequential multimodal tasks within a fixed-capacity MoE by compressing accumulated expert knowledge into low-rank memory subspaces, while expanding only the lightweight routers. We validate the effectiveness of our method on multiple healthcare multimodal benchmarks. It demonstrates competitive multitask pretraining performance while alleviating catastrophic forgetting and improving parameter efficiency.


FLARE: Diffusion for Hybrid Language Model

Yuchen Zhu ⋅ Jing Shi ⋅ Chongjian GE ⋅ Hao Tan ⋅ Yiran Xu ⋅ Wanrong Zhu ⋅ Jason Kuen ⋅ Ryan Rossi ⋅ Koustava Goswami ⋅ Rajiv Jain ⋅ Yongxin Chen ⋅ Molei Tao ⋅ Jiuxiang Gu

Autoregressive (AR) large language models have achieved broad practical success, but their sequential decoding remains a major bottleneck for low-latency deployment. Efficiency efforts have advanced along two largely orthogonal axes: hybrid attention architectures that reduce the cost of each forward pass, and diffusion language models (dLLMs) that enable parallel token generation. Existing dLLMs, however, have yet to translate their theoretical decoding parallelism into commensurate real-world throughput, have not incorporated modern hybrid attention backbones, and still trail scale-matched AR models on quality. In this work, we present FLARE, a systematic recipe for converting hybrid AR LLMs into capable, real-time-fast dLLMs under a practical training budget. Through a controlled study, we identify transfer data quality as the dominant driver of AR-to-dLLM performance loss, outweighing loss formulation and attention-mask design of which are emphasized by prior works. To enable efficient training and deployment of a hybrid softmax-plus-linear-attention backbone, we develop hardware-aware Triton kernels for diffusion-style linear attention and an SGLang-based inference system that exposes both AR-Trust and Diffusion-Trust decoding from the same checkpoint. Starting from Qwen3.5 checkpoints with only 10B tokens from public datasets, FLARE-9B matches the top-tier open-source dLLM LLaDA-2.1 Flash (100B-A5B) at 1/10 of the parameters, exceeding it on GPQA-Diamond ($71.2$ vs. $66.7$), and FLARE-2B reaches $2.2$ times the throughput of LLaDA-2.1 mini (16B-A1B) on GSM8K at eight-way concurrency on a single A100.

Principled regression for stochastic processes is a long-standing challenge with deep connections to scientific inverse problems. We introduce Flow Annealing Posterior Sampling (FAPS), to our knowledge the first function-space posterior sampling framework that unifies stochastic-process regression and PDE inverse problems. Built on pretrained function-space flow-matching priors, FAPS enables likelihood-guided posterior inference from sparse and noisy observations, supports variable query discretizations, and avoids explicit prior-density evaluation. Its Langevin correction uses a low-rank covariance preconditioner to exploit dominant function-space correlations across discretizations. Across Gaussian and non-Gaussian stochastic-process regression benchmarks and diverse PDE inverse problems, FAPS produces coherent posterior samples with accurate uncertainty quantification, significantly outperforming existing functional regression baselines while remaining computationally efficient.


FLUX: Geometry-Aware Longitudinal Flow Matching with Mixture of Experts

Josue Ortega Caro ⋅ Yongxu Zhang ⋅ Hannah M Batchelor ⋅ Sizhuang He ⋅ Jessica Cardin ⋅ Shreya Saxena

Many biological systems evolve through continuous local dynamics while switching between latent regimes defined by learning, stimulus context, internal state, or developmental stage. These processes are often observed only as unpaired longitudinal snapshots: the same cells, neurons, or animals are not tracked as matched trajectories, even though population states are sampled across successive stages. This creates two coupled challenges. First, trajectories must respect curved low-dimensional manifolds embedded in high-dimensional biological measurements. Second, the model must identify when the transport mechanism itself changes. We introduce FLUX (FLow matching for Unpaired longitudinal data with miXture-of-experts), a geometry-aware longitudinal flow-matching framework for joint transport modeling and unsupervised regime discovery. FLUX learns a data-dependent metric from pooled labeled and unlabeled observations, uses that metric to construct geometry-aware conditional paths between adjacent marginals, and decomposes the resulting velocity field into sparse expert vector fields selected by a Straight-Through Gumbel-Softmax router. Across manifold controls, a regime-switching Lorenz system, widefield cortical calcium imaging during associative learning, and embryoid body single-cell differentiation, FLUX reconstructs longitudinal transport while recovering interpretable regime structure. In neural data, the router separates early/intermediate training from late training, coinciding with the behavioral divergence of CS+ and CS- lick indices. In embryoid body differentiation, FLUX separates expression-evolution regimes associated with pluripotent and more differentiated cell populations. Ablations show that mixture-of-experts routing alone is insufficient: FLUX without geometric learning can fit local transport but fails or weakens regime discovery when regimes are encoded in local dynamics. These results suggest that geometry-aware velocity decomposition provides a general strategy for discovering latent biological state transitions from unpaired longitudinal snapshots.


From History to State: Constant-Context Skill Learning for LLM Agents

Haoyang Xie ⋅ Xinyuan Wang ⋅ Yancheng Wang ⋅ Puda Zhao ⋅ Feng Ju

Large language model (LLM) agents are increasingly used to operate browsers, files, code and tools, making personal assistants a natural deployment target. Yet personal agents face a privacy-cost-capability tension: cloud models execute multi-step workflows well but expose sensitive intermediate context to external APIs, while local models preserve privacy but remain less reliable. Both settings also pay repeatedly for long skill prompts and growing histories. We propose constant-context skill learning, a context-to-weights framework for recurring agent workflows: reusable procedures are learned in lightweight task-family modules, while inference conditions only on the current observation and a compact state block. A deterministic tracker renders this state block from task progress and supplies aligned subgoal rewards, so each module can be trained with step-level SFT and refined through online RL. Across ALFWorld, WebShop, and SciWorld, our agents achieve strong performance across Qwen3-4B, Qwen3-8B and Llama-3.1-8B. With Qwen3-8B, SFT+RL reaches 89.6\% unseen success on ALFWorld, 76.8\% success on WebShop, and 66.4\% unseen success on SciWorld. They match or exceed strong published agent-training results while reducing prompt tokens per turn by 2-7$\times$ relative to controlled ReAct prompting baselines, showing that procedural context can be moved from prompts into weights.


From Likelihood Convergence to Parameter Convergence in POMDPs

Jack Zhang ⋅ Saurabh Amin ⋅ Jiawei zhang ⋅ Patrick Jaillet

Likelihood-based methods are widely used to learn the parameters of partially observable Markov decision processes (POMDPs), but their finite-sample behavior is not well understood. We establish a different kind of finite-sample guarantee. For any candidate POMDP parameter, we bound its parameter recovery error in terms of its likelihood gap to the ground truth, the structural conditioning of the observation and transition kernels, and the behavior policy used to collect data. The bound applies in both the undercomplete and overcomplete observation regimes, with the overcomplete case following naturally from the undercomplete analysis. Because the result applies to any sufficiently good candidate, it yields finite-sample guarantees for the maximum-likelihood estimator, OMLE-style algorithms, and any procedure that produces a parameter with controlled likelihood gap. This is particularly useful for downstream tasks such as infinite-horizon planning with state-dependent rewards, where accurate recovery of the kernels themselves is required to obtain near-optimal policies. The bound separates the contribution of the behavior policy from that of the model, making explicit how data-collection choices, including action coverage and mixing rate, accelerate parameter learning.

Model-based reinforcement learning in partially observable environments requires inferring hidden states, learning latent dynamics, and planning under the resulting belief. Practical algorithms often solve one aspect of the above pipeline while relying on approximations for the others that are often poorly understood, while existing POMDP theory typically analyzes computationally inefficient theoretical algorithms far from modern deep RL practice. To fill this gap, we propose a practical end-to-end algorithm that approximately solves the hidden state filtering problem via sequential Monte Carlo (SMC), updates a latent dynamics model online, and performs approximate online planning for $d$ steps before bootstrapping leaves with a state-value critic (a QMDP approximation). We analyze each approximation and prove an end-to-end regret bound decomposing the error into model estimation error, particle filter error, Monte Carlo rollout error, and a structural bias term induced by the QMDP approximation controllable by the planning depth that we further analyze in various special cases. The resulting guarantee makes explicit how statistical sample size, particle budget, rollout budget, and planning depth affect performance. We instantiate our algorithm in an MPC-based deep RL framework, demonstrating its practicality on a variety of environments. Our results provide a practical and theoretically grounded route for bridging theory and practice within online decision-making in POMDPs.


Fully Distributed Tâtonnement for Chores Markets

Bhaskar Ray Chaudhury ⋅ Christian Kroer ⋅ Ruta Mehta ⋅ Tianlong Nan

We study price-adjustment dynamics for computing competitive equilibria (CE) in Fisher markets with chores. Unlike in classical goods markets, prices in chores markets are payments for taking on undesirable tasks, and natural excess-demand dynamics can fail; even the naïve analogue of Walrasian tâtonnement may diverge. Recent work of Chaudhury et al. (2025) overcomes this obstacle via \emph{relative tâtonnement}, which subtracts the average excess-demand signal from the excess demand vector. This recovers convergence, but at the cost of coupling the price updates across all chores. This leaves open whether such global coupling is inherent, or whether convergent tâtonnement can be recovered through a genuinely local update in which each chore reacts only to its own excess demand. We answer this question affirmatively through \emph{multiplicative tâtonnement}, a fully distributed dynamics in which each chore price is updated using only its current price and its own excess-demand signal. Although the update contains no explicit normalization term, Walras' law and the multiplicative form of the update implicitly preserve the relevant aggregate price geometry. We prove that multiplicative tâtonnement converges to a CE in any chores Fisher market with continuous, convex, and 1-homogeneous (CCH) disutilities. For convex CES disutilities, we further prove an approximate-CE convergence rate with the same O(1/\varepsilon^2) dependence as relative tâtonnement, but with improved dependence on problem constants. Experiments on real-world and simulated instances show that multiplicative tâtonnement is substantially faster in practice, often by an order of magnitude.

Learning neural stochastic differential equations (SDEs) from trajectory data is a central task in scientific machine learning, but standard neural SDEs trade flexibility for coefficient fidelity: flexible parameterizations recover coefficients poorly, while structured architectures impose strong coefficient priors. We introduce Gaeta-Lie Neural SDEs, a three-stage pipeline that improves coefficient recovery through discovered symmetries without priors on the drift or diffusion. Building on the Itô SDE symmetry framework of Gaeta and Quintero (1999), we first train a flexible neural SDE surrogate, then discover approximate projectable Lie symmetries of the learned surrogate in a finite basis of canonical symmetry generators, and finally regularize the surrogate by penalizing the residual of the Itô symmetry determining equations. The discovery stage is lightweight, scalable, and theory-guided: the SDE symmetry Lie algebra has known dimensional bounds that constrain the symmetry search space. The resulting regularization is architecture-agnostic and improves existing neural SDE methods when combined with them. Across non-trivial SDEs spanning low and high dimensions and a real-data setting, Gaeta-Lie regularization consistently improves fidelity over multiple baselines, with gains that persist in the data-scarce regime as guaranteed by the symmetry residual's freedom from data sparsity. To our knowledge, this is the first method to use SDE Lie symmetry to improve a learned stochastic surrogate.


Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention

Ali Hatamizadeh ⋅ Yejin Choi ⋅ Jan Kautz

Linear attention replaces the unbounded cache of softmax attention with a fixed-size recurrent state, reducing sequence mixing to linear time and decoding to constant memory. The hard part is not just what to forget, but how to edit this compressed memory without scrambling existing associations. Delta-rule models subtract the current read before writing a new value, and Kimi Delta Attention (KDA) sharpens forgetting with channel-wise decay. But the active edit still uses a single scalar gate to control two different things: how much old content to erase on the key side and how much new content to commit on the value side. We introduce Gated DeltaNet-2, which generalizes both Gated DeltaNet and KDA by inheriting adaptive forgetting and channel-wise decay while addressing their shared limitation, the scalar tie between erasing and writing. Gated Delta Rule-2 separates these roles with a channel-wise erase gate $\boldsymbol{b}_t$ and a channel-wise write gate $\boldsymbol{w}_t$, reducing to KDA when both gates collapse to the same scalar and to Gated DeltaNet when the decay also collapses. We derive a fast-weight update view, a chunkwise WY algorithm with channel-wise decay absorbed into asymmetric erase factors, and a gate-aware backward pass that preserves efficient parallel training. At 1.3B parameters trained on 100B FineWeb-Edu tokens, Gated DeltaNet-2 achieves the strongest overall results among Mamba-2, Gated DeltaNet, KDA, and Mamba-3 variants across language modeling, commonsense reasoning, and retrieval. Its advantage is most pronounced on long-context RULER needle-in-a-haystack benchmarks, where it improves the evaluated multi-key retrieval setting and remains strong in both recurrent and hybrid settings.


GEAR: A GPU-Accelerated Global Solver for Nonlinear Programs via Linear Bound Propagation

Duo Zhou ⋅ Hesun Chen ⋅ Xiangru Zhong ⋅ Grani A. Hanasusanto ⋅ Huan Zhang

We present a novel GPU-accelerated global solver for constrained nonlinear programming (NLP) inspired by the efficient linear bound propagation framework in neural network (NN) verification. Existing work on NN verification, a problem that can be cast as feasibility detection for a nonlinear objective, demonstrated orders-of-magnitude speedups with the linear bound propagation-based framework when compared to traditional solvers. Extending this paradigm to global nonlinear programming requires addressing many challenges: a global constrained NLP solver must effectively exploit all, possibly unstructured, nonlinear constraints to improve the objective, whereas NN verification typically handles simple input box constraints over a well-structured NN and does not directly optimize an objective function. We address these challenges by fully exploiting the linear bounds produced by bound propagation: these linear bounds can be viewed as relaxations of nonlinear constraints and used to provide dual objectives, check infeasibility, and guide the search for a primal solution. In addition, we proposed tighter linear relaxations (essential for linear bound propagation) for bilinear functions commonly used in NLP problems, and a projected augmented Lagrangian method tightly coupled to our dual solving procedure to produce primal solutions. Our method solves 399 of 505 bounded and continuous NLP problems in the GAMS and MINLPLib benchmarks under a 180-second limit, outperforming many strong global solvers (including commercial ones) such as SCIP, BARON, and LindoGlobal. Compared to SCIP (the strongest solver on these problems), we solve 104 problems exclusively, including large-scale, highly nonlinear ones that are often out of reach for traditional approaches. Our method demonstrates a distinct, highly complementary regime for GPU-accelerated global optimization of nonlinear programming problems.

This position paper argues that generalization measures for deep learning should be audited for fragility before they are used as explanations of why trained networks generalize. Many post-mortem measures -- those computed on trained networks -- are fragile: small, routine training changes that barely affect the learned predictor or test performance can substantially change a measure's value, trend, or scaling behavior. For example, changing the learning rate or swapping SGD for Adam can reverse the qualitative learning-curve behavior of widely used measures such as the path norm. We also identify subtler failures. A standard parameter-space PAC-Bayes proxy, PACBAYES_ORIG, is less sensitive to hyperparameter tweaks than many norm measures, but it can fail to capture differences in data complexity across learning curves. By contrast, a function-space marginal-likelihood PAC-Bayes bound, used here as a calibration baseline rather than a gold standard, tracks data complexity and sample-size scaling in our experiments while exposing the limitations of optimizer-agnostic GP-based approaches. We therefore propose explicit fragility audits -- including hyperparameter, temporal, data-complexity, matched-error CMS/eCMS, and invariance checks -- as a standard complement to tightness and correlation claims.


Generative Modeling via Drifting

Mingyang Deng ⋅ He Li ⋅ Tianhong Li ⋅ Yilun Du ⋅ Kaiming He

Generative modeling can be formulated as learning a mapping $f$ such that its pushforward distribution matches the data distribution. The pushforward behavior can be carried out iteratively at inference time, for example, in diffusion/flow-based models. In this paper, we propose *Drifting Models*, a generative modeling framework that evolves the pushforward distribution during training and naturally admits one-step inference. We introduce a drifting field that governs the sample movement and achieves equilibrium when the distributions match. This leads to a training objective that allows the neural network optimizer to evolve the distribution. In experiments, our one-step generator achieves state-of-the-art results on ImageNet 256$\times$256, with FID 1.54 in latent space and 1.61 in pixel space. We hope that our work opens up new opportunities for high-quality one-step generation.


GeoTransolver: Learning Physics on Irregular Domains using Multi-scale Geometry Aware Physics Attention Transformer

Corey Adams ⋅ Rishikesh Ranade ⋅ Sheel Nidhan ⋅ Ram Cherukuri ⋅ Sanjay Choudhry ⋅ Mohammad Amin Nabian ⋅ Deepak Akhare

We present GeoTransolver, a multiscale geometry-aware physics attention transformer for Computer Aided Engineering (CAE). GeoTransolver extends the Transolver backbone with GALE (Geometry-Aware Latent Embeddings) attention, which pairs physics-aware self-attention on learned state slices with cross-attention to a shared geometry and global context computed via multi-scale ball queries (inspired by Domino) and reused in every block. Implemented and released in NVIDIA PhysicsNeMo, GeoTransolver persistently projects geometry and global parameters, into physical state spaces to anchor computations to domain structure and operating regimes. We benchmark on DrivAerML, SHIFT-SUV, and SHIFT-Wing against Domino, Transolver (PhysicsNeMo implementation), and literature-reported AB-UPT, evaluating drag/lift $R^2$ and relative $L_1$ errors on field variables. As an additional nonlinear structural mechanics application, we also report Transolver and GeoTransolver results on bumper-beam and full-vehicle Body-in-White (BIW) crash-dynamics benchmarks, evaluating relative $L_2$ trajectory error and probe-level kinematic MSE. GeoTransolver delivers improved accuracy, robustness to geometry and regime shifts, and favorable data efficiency; we include DrivAerML ablations and qualitative contour and design-trend results, advancing operator learning for high-fidelity surrogates on complex, irregular, non-linear domains.

Knowledge Graphs (KGs) are a rich source of structured data, and graph neural networks (GNNs) are the dominant tool for learning over them. However, existing message-passing GNNs struggle to scale to large KGs because they rely on the iterative message passing process to learn graph structure. This process is inefficient, especially under mini-batch training, where a node sees only a partial view of its neighborhood. In this paper, we address this problem and present gHAWK, a novel and scalable GNN training framework for large KGs. The key idea is to precompute structural features for each node that capture its local and global structure before GNN training even begins. Specifically, gHAWK introduces a preprocessing step that computes: (a) Bloom filters to compactly encode local neighborhood structure, and (b) TransE embeddings to represent each node's global position in the graph. These features are then fused with any domain-specific features (e.g., text embeddings), producing a node feature vector that can be used with any GNN backbone. Equipped with these structural priors, the GNN no longer needs to rediscover graph structure, significantly improving memory usage, convergence, and model accuracy. Extensive experiments on large datasets from the Open Graph Benchmark demonstrate that gHAWK improves accuracy on both node property prediction and link prediction tasks. Remarkably, gHAWK enables simple GNNs to outperform more complex relation-aware ones, suggesting that careful preprocessing can substitute for architectural complexity in KG learning.


Global Convergence of Four-Layer Matrix Factorization under Random Initialization

Minrui Luo ⋅ Weihang Xu ⋅ Xiang Gao ⋅ Maryam Fazel ⋅ Simon Du

Gradient descent dynamics on the deep matrix factorization problem is extensively studied as a simplified theoretical model for deep neural networks. Although the convergence theory for two-layer matrix factorization is well-established, no global convergence guarantee for general deep matrix factorization under random initialization has been established to date. To address this gap, we provide a polynomial-time global convergence guarantee for randomly initialized gradient descent on four-layer matrix factorization, given certain conditions on the target matrix and a standard balanced regularization term. Our analysis employs new techniques to show saddle-avoidance properties of gradient descent dynamics, and extends previous theories to characterize the change in eigenvalues of layer weights.

Gradient descent is one of the most widely used iterative algorithms in modern statistical learning. However, its precise algorithmic dynamics in high-dimensional settings remain only partially understood, which has limited its broader potential for statistical inference applications.

This paper provides a precise, nonasymptotic joint distributional characterization of gradient descent iterates and their debiased statistics in a broad class of empirical risk minimization problems, in the so-called mean-field regime where the sample size is proportional to the signal dimension. Our nonasymptotic state evolution theory holds for both general nonconvex loss functions and non-Gaussian data, and reveals the central role of two Onsager correction matrices that precisely characterize the nontrivial dependence among all gradient descent iterates in the mean-field regime.

Leveraging the joint state evolution characterization, we show that the gradient descent iterate retrieves approximate normality after a debiasing correction via a linear combination of observable loss derivative directions from all past iterates. Crucially, the debiasing coefficients are directly linked to the Onsager correction matrices, which can be estimated in a fully data-driven manner via the proposed gradient descent inference algorithm. This leads to a new algorithmic statistical inference framework based on debiased gradient descent, which (i) applies to a broad class of models with both convex and nonconvex losses, (ii) remains valid at each iteration without requiring algorithmic convergence and (iii) exhibits a certain robustness to possible model misspecification. As a by-product, our framework also provides algorithmic estimates of the generalization error at each iteration.

We demonstrate our theory and inference methods in the canonical single-index regression model and a generalized logistic regression model, where the natural loss functions may exhibit arbitrarily nonconvex landscapes. Our analysis further shows that, in linear regression with squared loss, the proposed debiased gradient descent iterate eventually coincides with the debiased convex regularized estimator in a mean-field distributional sense, and the quality of statistical inference for the unknown signal aligns exactly with the generalization error achieved along the algorithmic trajectory.


Gradient Descent on Two ReLU Neurons: Global Landscape and Bifurcation Dynamics

Binghua Li ⋅ Mengzhe Li ⋅ Denny Wu ⋅ Tianhao Wang

Understanding how gradient descent learns features remains a central challenge in neural network theory. We study this question in a minimal multi-index setting that already exhibits rich multi-phase gradient dynamics: a well-specified two-neuron ReLU teacher-student model under isotropic Gaussian inputs, trained by population gradient flow from small random initialization. We show that the dynamics are organized by two directions: the “easy” bisector direction, which carries the leading signal, and the “hard” splitting direction, which governs specialization. To characterize the loss structure and learning behavior, we analyze the population squared-loss landscape and show that every nonzero critical point is either a global minimum or a saddle. We then track the gradient-flow trajectory, showing that student neurons first collapse toward a bisector saddle before escaping and specializing to teacher neurons. Our results provide both a landscape and dynamics account of the multi-phase symmetry-breaking behavior that arises in a simple multi-index model and under standard algorithmic and architectural choices.


Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory Agents

Wei-Chieh Huang ⋅ Weizhi Zhang ⋅ Yuchen Wu ⋅ Yankai Chen ⋅ Eric Hanchen Jiang ⋅ Wooseong Yang ⋅ Yiwei Yang ⋅ Henry P Zou ⋅ Hanrong Zhang ⋅ Ying Nian Wu ⋅ Haolun Wu ⋅ Kai-Wei Chang ⋅ Philip S Yu ⋅ Steve Liu ⋅ Aylin Caliskan

Memory is becoming core infrastructure for long-horizon LLM agents, yet existing evaluations offer limited guidance on which memory substrate should be used under different operating regimes. We present a controlled harness evaluation of memory substrates for memory-augmented agents, covering dense and sparse indices, text records, structural stores, hierarchical stores, refinement-based memories, parametric updates, and activation-compatible context mechanisms. Across three backbone models and four benchmark suites spanning user-centric question answering and agent-centric decision-making, we instrument 26 performance and efficiency metrics under a unified harness. Our results show that no single substrate consistently dominates: broad retrieval benefits long-context factual QA, while excessive retrieval can harm sequential decision-making by shifting attention away from action-critical context. Scalability further introduces an additional routing axis, as substrates that perform well at moderate history lengths can become costly or brittle at longer horizons. These findings motivate substrate routing as a necessary component of adaptive agent memory systems and provide empirical guidance for designing efficient, reliable, and regime-aware long-term memory for LLM agents.


How Accurately Can a Gaussian Approximate Stochastic Approximation Iterates?

Shaan U Haque ⋅ Zedong Wang ⋅ Zixuan Zhang ⋅ Siva Theja Maguluri

Stochastic approximation (SA) is a method for finding the root of an operator perturbed by noise. The focus of this paper is studying the distribution of SA iterates in finite time. In general, it is not possible to characterize the exact distribution, and therefore our goal is to find an approximation which can yield useful tail bounds. Inspired by the rich literature on the asymptotic normality of rescaled SA iterates, we approximate the pre-limit distributions by a sequence of Gaussians whose covariance is recursively defined. In particular, we establish explicit bounds on the Wasserstein-1 distance between the rescaled iterate at time $k$ and the aforementioned Gaussian for various choices of step-sizes. Since these covariances converge to the classical asymptotic limit, our analysis also provides a convergence rate for asymptotic normality as a by-product. As an immediate consequence of our bounds, we obtain tail bounds on the error of SA iterates at any time. Finally, we establish the sharpness of our rates by providing matching lower bounds and validate our findings through simulations. We obtain the sharp rates by first studying the convergence rate of the discrete Ornstein–Uhlenbeck (O-U) process driven by general noise, whose stationary distribution is identical to the limiting Gaussian distribution of the rescaled SA iterates. We believe that this is of independent interest, given its connection to sampling literature. The analysis involves adapting Stein’s method for Gaussian approximation to handle the matrix weighted sum of i.i.d. random variables. The desired finite-time bounds for SA are obtained by characterizing the error dynamics between the rescaled SA iterate and the discrete time O-U process and combining it with the convergence rate of the latter process.


How New Strategies Emerge in RL Post-Training: A Controlled Study

Azwar Azwar ⋅ Nishil Patel ⋅ Andrew Saxe

Does RL post-training build new reasoning procedures, or merely reweight behaviors already latent in the base model? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is known and every generated rewrite can be audited. A Transformer is pretrained on primitive symbol-rewrite chains and post-trained with only a binary final-answer reward. RL solves held-out problems that remain rarely solved by the pretrained model even under much larger sampling budgets, while rejection fine-tuning improves early but plateaus. Trace analysis reveals a phased mechanism: RL first strengthens primitive reductions, then enters a chunking phase in which it forms valid compressed procedures---macro contractions that collapse sequential reductions and parallel contractions that combine independent ones. These procedures are not isolated samples; they are reused and consolidated into a stable repertoire. Comparing RL with rejection fine-tuning shows that the key difference is not exploration volume but selectivity: RFT produces many shortcut-like rewrites, much of them invalid, whereas RL concentrates exploration into valid reusable structure. Pretraining ablations show that strategy emergence is gated not by primitive exposure alone, but by whether pretraining organizes primitive competence into reduction procedures that RL can later compress. The base model provides weak procedural ingredients; RL builds them into reliable higher-level strategies.


HPC-Bench: A Comprehensive Benchmark for High Performance Computing Codes

Bowen Cui ⋅ Junyu Yin ⋅ Tejas Ramesh ⋅ Leo Lim ⋅ Oscar Hernandez ⋅ Keren Zhou

Large Language Models (LLMs) have shown strong capabilities in code generation and reasoning, motivating extensive benchmarks for evaluating functional correctness. However, performance optimization in High-Performance Computing (HPC) poses fundamentally different challenges that are not captured by existing evaluations: HPC performance depends on low-level code, intricate memory access patterns, and computational kernels embedded within sophisticated execution contexts. We introduce HPC-Bench, a performance-grounded benchmark for evaluating LLMs as optimizers of computational kernels in realistic HPC workloads, including $114$ workloads across $16$ HPC computational motifs. HPC-Bench provides a fully automated, end-to-end evaluation pipeline that combines program-level correctness verification with two correctness-gated metrics, $\mathrm{Fast}@k$ and $\mathrm{Speedup}@k$, to characterize reliability and speedup under a fixed sampling budget. Our evaluation of three representative LLMs across serial CPU, OpenMP, and CUDA settings shows that stricter speedup requirements substantially reduce the probability of finding a correct and fast optimization. HPC-Bench is available at https://github.com/neuripspaper2026/NeurIPS_26.

Post-training has become essential for adapting large language models (LLMs) to complex downstream behaviors, including instruction following, preference alignment, and multi-step reasoning. Reinforcement learning with verifiable rewards (RLVR) has recently emerged as a particularly effective post-training paradigm for improving reasoning capabilities, with critic-free algorithms such as GRPO and GSPO enabling scalable optimization. However, RLVR post-training with full fine-tuning (FFT) requires substantial GPU memory and incurs high training costs. Although parameter-efficient fine-tuning (PEFT) methods, such as Low-Rank Adaptation (LoRA), effectively reduce computational costs, they often suffer from a noticeable performance gap compared to full fine-tuning in post-training for complex reasoning tasks. In this paper, we propose Hybrid-LoRA, an efficient hybrid post-training framework that selectively applies full fine-tuning to a small subset of modules less suited to low-rank adaptation, while adapting the remaining components with LoRA. We introduce a novel Hybrid-LoRA Score to rank candidate modules according to their sensitivity to low-rank adaptation under a fixed parameter budget. Experiments show that Hybrid-LoRA closely matches full fine-tuning performance under a 10\% full fine-tuning module budget, with the remaining candidate modules adapted by LoRA, consistently outperforming four state-of-the-art PEFT post-training baselines, achieving improvements of up to 5.65\% and on average 4.36\% over the best baseline.


HYVE: Hybrid Views for LLM Context Engineering over Machine Data

Jian Tan ⋅ Fan Bu ⋅ Yuqing Gao ⋅ Devavrat Khanolkar ⋅ Jason Mackay ⋅ Boris Sobolev ⋅ Lei Jin ⋅ Li Zhang

Machine data is central to observability and diagnosis in modern computing systems, appearing in logs, metrics, telemetry traces, and configuration snapshots. When provided to large language models (LLMs), this data typically arrives as a mixture of natural language and structured payloads such as JSON or Python/AST literals. Yet LLMs remain brittle on such inputs, particularly when they are long, deeply nested, and dominated by repetitive structure. We present HYVE (HYbrid ViEw), a framework for LLM context engineering for inputs containing large machine-data payloads, inspired by database management principles. HYVE surrounds model invocation with coordinated preprocessing and postprocessing, centered on a request-scoped datastore augmented with schema information. During preprocessing, HYVE detects repetitive structure in raw inputs, materializes it in the datastore, transforms it into hybrid columnar and row-oriented views, and selectively exposes only the most relevant representation to the LLM. During postprocessing, HYVE either returns the model output directly, queries the datastore to recover omitted information, or performs a bounded additional LLM call for SQL-augmented semantic synthesis. We evaluate HYVE on diverse real-world workloads spanning knowledge QA, chart generation, anomaly detection, and multi-step network troubleshooting. Across these benchmarks, HYVE reduces token usage by 50--90\% while maintaining or improving output quality. On structured generation tasks, it improves chart-generation accuracy by up to 132\% and reduces latency by up to 83\%. Overall, HYVE offers a practical approximation to an effectively unbounded context window for prompts dominated by large machine-data payloads.

We propose a deterministic adjoint matching framework that formulates human preference alignment for flow-based generative models as an optimal control problem over velocity fields. One can directly regress the control toward a value-gradient-induced target under the current policy, leading to a simple and stable training objective. Building on this perspective, we introduce a truncated adjoint scheme that focuses computation on the terminal portion of the trajectory, where reward-relevant signals concentrate, which yields substantial computational savings while preserving alignment quality. We further generalize the framework beyond standard KL-based regularization, allowing more flexible trade-offs between alignment strength and distributional preservation. Experiments on SiT-XL/2 and FLUX.2-Klein-4B demonstrate consistent gains across multiple alignment metrics, along with substantially improved diversity and mode preservation.


IncAgg: Efficient Memory-Enhanced Graph Learning via Incremental Aggregation

Xingyue Shi ⋅ Zhichao Hou ⋅ Jiahao Zhang ⋅ Suhang Wang ⋅ Tong Zhao ⋅ Neil Shah ⋅ Xiaorui Liu

Scaling Graph Neural Network (GNN) training to large graphs commonly relies on mini-batch sampling, but incomplete neighborhoods can degrade approximation quality and training stability. Memory-enhanced methods mitigate this issue by reusing historical node embeddings to approximate full-neighborhood aggregation; however, they reconstruct pseudo full-neighborhood messages inside every training iteration, repeatedly fetching out-of-batch memories and aggregating over large neighborhoods. In this work, we identify this under-examined inner-loop bottleneck and propose IncAgg, an efficient memory-enhanced GNN training framework based on incremental aggregation. IncAgg periodically pre-aggregates historical embeddings into a global memory baseline, and during mini-batch training propagates only the current in-batch residual before fusing it with the stored global context. This design preserves historical full-neighborhood information at a configurable refresh frequency while avoiding per-iteration out-of-batch memory fetching and aggregation, without modifying the underlying GNN architecture. We provide theoretical analyses showing that IncAgg reduces gradient-carrying aggregation and memory-access costs while retaining the residual-based stability principle of memory-enhanced estimators. Extensive experiments on four large-scale benchmarks and four representative GNN backbones show that IncAgg matches or improves the accuracy of strong memory-enhanced baselines while achieving substantial training acceleration, with up to $18.6\times$ speedup in individual settings and up to $13.8\times$ geometric-mean speedup on large graphs.

In this paper, we study the problem of simultaneous inference for multiple quantiles under local differential privacy (LDP). We develop a general online framework that reduces the multiple-quantile problem to a categorical one-hot estimation problem. A projected Robbins-Monro recursion guarantees noncrossing at every iteration, while Polyak-Ruppert averaging together with multivariate self-normalization yields a joint asymptotically pivotal ellipsoidal confidence region without requiring estimation of nuisance densities or asymptotic covariance matrices. We also present several important examples in detail, including direct $k+1$-ary randomized response (kRR), optimized unary encoding (OUE), optimized local hashing (OLH), and Hadamard response (HR). We evaluate our methodology through numerical experiments based on these examples, and the results provide positive support for the proposed framework.


Inference-Time Search Using Side Information for Diffusion-Based Image Reconstruction

Mahdi Farahbakhsh ⋅ Vishnu Teja Kunde ⋅ Dileep Kalathil ⋅ Krishna Narayanan ⋅ JF Chamberland

Diffusion models have been used as priors for solving inverse problems. However, existing approaches typically overlook side information that could significantly improve reconstruction quality, especially in severely ill-posed settings. In this work, we propose a novel framework that incorporates side information into existing diffusion-based inverse problem solvers via inference-time search, in a plug-and-play, training-free manner. Through extensive experiments across a range of inverse problems, including inpainting, super-resolution, and several deblurring tasks, and across multiple diffusion-based inverse problem solvers (DPS, DAPS, and MPGD), we show that augmenting each solver with our framework consistently improves the quality of the reconstructions over the corresponding original method. To demonstrate the generality of our approach, we consider diverse forms of side information, including reference images, textual descriptions, and anatomical MRI scans. {Code is available \href{https://anonymous.4open.science/r/sideinfo-search-reconstruction-4A32/README.md}{here}.}

Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortized inference over discrete and mixed discrete-continuous objects, requiring only an unnormalized target density specified through a reward. In this work, we formulate forward-policy training in GFlowNets through the information geometry of the induced trajectory sampler. Treating the forward policy as an induced trajectory sampler, we show that its intrinsic first-order geometry is given by the Fisher-Rao metric of the trajectory family, and that the associated natural gradient provides the canonical local update whenever the corresponding Fisher information is computable or accurately approximable. We derive an exact decomposition of the trajectory Fisher into per-step conditional second moments, which clarifies when temporal score interactions vanish and when dense couplings remain under shared parameterization. This leads to three computational regimes: settings with tractable exact Fisher information, settings where Monte Carlo estimators of the expected Fisher are sufficient, and structure-exploitable settings in which target locality or factorization yields accurate approximations of the Fisher expectation. In the latter case, graphical-model tools such as exact marginalization, separator methods, and belief propagation provide principled surrogates for natural-gradient updates. The resulting framework turns target structure into optimization geometry and yields a tractable route to structure-aware forward-policy training in GFlowNets. We illustrate the framework empirically through examples comparing convergence and exploration behavior under Riemannian and Euclidean optimization.

Active feature acquisition (AFA) is an instance-adaptive paradigm in which, at inference time, a policy sequentially chooses which features to acquire (at a cost) before predicting. Existing approaches either train reinforcement learning policies, which deal with a difficult MDP, or greedy policies that cannot account for the joint informativeness of features or require knowledge about the underlying data distribution. To overcome this, we propose Template-based AFA (TAFA), a non-greedy framework that learns a small library of feature templates---sets of features that are jointly informative---and uses this library of templates to guide the next feature acquisitions. Through identifying feature templates, the proposed framework not only significantly reduces the action space considered by the policy but also alleviates the need to estimate the underlying data distribution. Extensive experiments on synthetic and real-world datasets show that TAFA outperforms the existing state-of-the-art baselines while achieving lower overall acquisition cost and computation.


Instance-Adaptive Online Multicalibration

Zhiming Huang ⋅ Jamie Morgenstern ⋅ Aaron Roth ⋅ Claire Jie Zhang

We study online multicalibration beyond the worst-case. We give a single, efficient algorithm which dynamically interpolates between stochastic and worst-case sequences by adaptively refining a dyadic grid of prediction values. Its error is controlled by the number of leaves in the refinement tree. Our analysis recovers the known $\widetilde O(T^{2/3})$ worst-case-optimal rate for online multicalibration, while simultaneously automatically adapting to easier instances: in the marginal stochastic setting it obtains a rate of $\widetilde O(\sqrt T)$, and for piecewise-stationary means with $J$ segments its rate is $\widetilde O(\sqrt{JT})$. More generally, the rate depends on a threshold-complexity measure of the predictable mean process relative to the group family. We show that this dependence is tight up to logarithmic factors.


Inverting Retargeting: Humanoid Datasets Remember Their Operators

Sihat Afnan ⋅ Unnat Jain ⋅ Habiba Farrukh

Humanoid teleoperation datasets are growing rapidly, with operator demonstrations retargeted onto a shared robot skeleton to train whole-body controllers and foundation models. We report a surprising property of this pipeline: retargeting normalizes body proportions but preserves operator movement dynamics (joint velocity profiles, ranges of motion, and coordination patterns shaped by the operator's physiology). On BONES-SEED (522 operators, 142K sequences), retargeted Unitree G1 trajectories support gender classification at 96.0% and operator re-identification at 97.2% Top-1; on operators never seen during training, gender holds at 83.4% and age and height regress within ±4.2 yr and ±5.7 cm. Partial correlation analysis reveals emergent, biomechanically interpretable structure: the signals are task-invariant across activity categories and hold across retargeting implementations. We introduce UNVEIL, a skeleton-aware spatiotemporal graph network to measure and interpret this effect, take initial steps toward operator anonymization, and ask the community: as teleoperation datasets scale, what should our data practices be? Code, models, and anonymized trajectories are at our project website https://project-unveil.github.io.


Iterative Critique-and-Routing Controller for Multi-Agent Systems with Heterogeneous LLMs

Wenzhi Fang ⋅ Liangqi Yuan ⋅ Guangchen Lan ⋅ Dong-Jun Han ⋅ Christopher Brinton

Multi-agent large language model (LLM) systems often rely on a controller to coordinate a pool of heterogeneous models, yet existing controllers are typically limited to one-shot routing: they select a model once and return its output directly. Such routing-only designs provide no mechanism to critique intermediate drafts or support iterative refinement. To address this limitation, we propose a critique-and-routing controller that casts multi-agent coordination as a sequential decision problem. At each turn, the controller evaluates the current draft, decides whether to stop or continue, and, if needed, selects the next agent for further refinement. We formulate this process as a finite-horizon Markov Decision Process (MDP) with explicit agent-utilization constraints, design a composite reward for controller decisions across turns, and optimize the controller via policy gradients under a Lagrangian-relaxed objective. Extensive experiments across multiple heterogeneous multi-agent systems and seven reasoning benchmarks show that our method consistently outperforms state-of-the-art baselines and substantially narrows the gap to the strongest agent, while using it for fewer than 25% of total calls.

Operator learning has experienced a shift toward attention-based architectures. This trend has more recently extended to physics-informed operator learning, where stacked cross-attention blocks add considerable computational cost relative to the foundational PI-DeepONet. We ask whether this complexity is necessary. Starting from the DeepONet branch-trunk decomposition, we replace the bilinear readout with a single softmax-weighted readout in which the trunk emits a query vector and the branch emits a set of key-value pairs—a structure we show admits universal approximation via reduction to a simplex-bilinear form. The resulting architecture, KerONet, uses no iterated attention, layer normalization, or explicit learnable projection matrices. Under a unified hyperparameter optimization protocol with a strict $200{,}000$-parameter budget, we evaluate KerONet alongside iterated-attention architectures (PIT, PINTO) and PI-DeepONet on four canonical 1D PDEs whose solutions range from globally smooth to sharply localized. KerONet outperforms PIT and PINTO on every PDE, across all five seeds, in accuracy (approximately $2$–$4\times$ lower rel-$L^2$ error), reproducibility ($2$–$31\times$ lower standard deviation), and speed ($1.8$–$4.4\times$ faster training time). Against PI-DeepONet, KerONet shows substantial gains where input-dependent localized features dominate ($2.5$–$8\times$ lower error on Burgers and Allen-Cahn) while exhibiting comparable accuracy on smoother operators (advection and diffusion-reaction)–a balance the iterated-attention architectures do not achieve. Our findings demonstrate that iterated cross-attention is not necessary to realize the benefits of query-dependent feature mixing in physics-informed operator learning.


KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling

Peng Kuang ⋅ Haibo Jin ⋅ Xiaoyu Han ⋅ Yanli Wang ⋅ Xiaopeng Yuan ⋅ Ye Yu ⋅ Kaidi Xu ⋅ Haohan Wang

rocess Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost the capabilities of LLM-based multi-agent systems. However, existing PRMs are *text-based*: they re-encode the entire trajectory text from scratch. In long multi-agent rollouts, the scoring cost, growing quadratically with respect to sequence length $L$, creates a severe computational bottleneck, severely limiting PRMs' application in long-context scenarios. To resolve this, we introduce KV-PRM, a highly efficient process reward model that eliminates the heavy text re-encoding by directly reading the KV cache produced naturally during the LLM's generation phase. By processing a single "verify token" against the pre-existing KV cache, KV-PRM reduces the scoring cost from $O(L^2)$ to $O(L)$. We formally prove that the KV cache contains strictly greater information capacity than text, and is more efficient for downstream reward modeling. Empirically, across the MATH, GSM8K, and AIME benchmarks, KV-PRM matches or strictly outperforms text-PRMs under various TTS methods such as Beam Search, MCTS, and Weighted Voting, with up to a *5,000*$\times$ reduction in scoring FLOPs, a *37*$\times$ reduction in latency, and a *34*$\times$ reduction in per-sequence memory footprint compared to text-based PRMs.


L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education

James Edgell ⋅ Wm. Matthew Kennedy ⋅ Ben Knight ⋅ Danielle Carvalho ⋅ Martin Ku ⋅ Isaac J Pattis

Despite rapid AI adoption in education, rigorous evaluation of AI-powered educational (AIED) systems remains critically underdeveloped, particularly in second-language (L2) education, one of the most common yet least evaluated AI applications. We introduce L2-Bench, an open-source benchmark of 1,000+ task-response pairs to aid the pedagogy-led evaluation of LLM capabilities relating to language learning and assessment. Crucially, L2-Bench measures model performativity on the application of learning experience design principles rather than mere knowledge of those principles or broad learning outcomes. Our contributions include: (1) a validated taxonomy of 12 competencies and 31 subcompetencies validated by 200+ expert practitioners (task authenticity = 4.42/5.00, criteria adequacy = 4.18/5.00); (2) a rubric-based evaluation methodology that we believe can, if adapted, generalize to similar (open-ended, qualitative) disciplines; (3) an evaluation dataset that produces reliable signal about model strengths, weaknesses, and contextual robustness across diverse L2 education scenarios. We find that, among large models, Claude Opus 4.7 performs best overall (85.5\%), though is marginally outperformed on several constituent tasks. We also find that performance drops notably on harder tasks $\in$ (69.9\%–73.4\%). L2-Bench provides education stakeholders better methods to make more informed decisions about real-world AIED adoption, use, and governance, while advancing the maturing science of AI evaluations for education.

Most explanations of weak-to-strong (W2S) generalization focus on why a strong student may fail to exactly imitate weak-label errors. We study a complementary mechanism: before student training, weak-teacher predictions may be inconsistent along semantic constraints specified independently of the weak model. When such constraints are valid for the target, these inconsistencies give label-free lower bounds on the risk reduction achievable by correcting weak labels, rather than merely indicating unstructured prediction error. We formalize this mechanism using Markov and Laplacian operators. Under exact target validity, Markov trajectory inconsistency equals the squared-risk gain from smoothing; under approximate validity, it remains a conservative lower bound after an explicit validity penalty. We then solve the associated minimax correction problem, obtaining a resolvent correction that continuously attenuates constraint-violating directions and reduces to the identity when validity uncertainty is too large. These results yield validity-adjusted correction (VAC), a split-sample pipeline that selects corrections from weak-teacher queries and predeclared validity budgets before training the strong student. Controlled graph-signal experiments and a CIFAR-10 cats-versus-dogs W2S experiment are consistent with the proposed mechanism: validity-adjusted inconsistency predicts useful correction and transfers to student gain, whereas high raw inconsistency under invalid controls does not provide such evidence.


LARK: Learnability-Grounded Trajectory Selection for Efficient Reasoning Distillation

Tianrun Yu ⋅ Kaixiang Zhao ⋅ Chih-Chun Chen ⋅ Amanda Hughes ⋅ Taylor Killian ⋅ Fenglong Ma ⋅ Weitong Zhang ⋅ Porter Jenkins

We study trajectory selection for reasoning distillation, where teacher-generated reasoning trajectories are selectively used as supervision for a student model. Existing methods rely on heuristics such as trajectory quality or model confidence, but they often overlook whether a trajectory is learnable by the student. In this paper, we present LARK, a learnability-grounded method for reasoning trajectory selection. LARK selects trajectories that the student can learn efficiently while preserving the generalization of the full training distribution. At the core of LARK is a learnability factor $\rho$, which characterizes the rate at which the student's training loss decreases. To estimate this rate efficiently and maintain generalization, we introduce a learnability proxy and a $\chi^2$-regularized selection policy that balances learnability and distributional coverage, both with strong theoretical guarantees on their estimation error. Empirically, LARK consistently outperforms data selection baselines across multiple base models and reasoning tasks. Diagnostic analyses show that the LARK score predicts downstream training utility and that LARK-selected trajectories induce faster supervised fine-tuning loss reduction.


LASER: Latent Space Adjoint Matching for Support Constrained Entropy Regularized Offline RL

Songyuan Zhang ⋅ Oswin So ⋅ Eric Yu ⋅ Matthew Cleaveland ⋅ Peter Crowley-Dolen ⋅ Chuchu Fan

While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent generative-model-based approaches mitigate this by performing RL within a constrained latent space, but they often sacrifice policy expressiveness or suffer from vanishing gradients. In this work, we identify that entropy regularization is essential to prevent mode collapse in latent-space RL, though applying it without introducing numerical instability or compromising expressiveness remains a significant open challenge. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized RL with expressive flow policies. By design, our approach satisfies dataset support constraints and prevents mode collapse. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses a single, constant set of hyperparameters to outperform baselines even with their hand-tuned hyperparameters, highlighting its robust applicability.

Reliable physics simulation requires both broad generalization across heterogeneous PDE families and stability under long autoregressive rollouts. Current neural solvers rarely provide both. Deterministic operators accumulate rollout error, while probabilistic solvers typically remain tied to a single PDE family or a short prediction horizon. We address this gap with the \textbf{Latent Generative Solver} (LGS), which couples a Physics VAE (PhyVAE) with a Pyramidal Flow-Forcing Transformer (PFlowFT). PhyVAE compresses states from twelve PDE families into a shared \emph{physical manifold}, separating dynamics-relevant structure from the ambient state space. PFlowFT then predicts the next latent state through input-noised flow matching, yielding a sufficient-condition contraction mechanism that explains improved autoregressive rollout stability. Pretrained on a 2.5\,M-trajectory, 16-system corpus at $128^2$ resolution, LGS obtains the lowest relative $L_2$ error (L2RE) on 15 of 16 systems at both 5- and 10-step rollout. At 20 steps, LGS reduces the L2RE from $56.1\%$ to $\mathbf{30.2\%}$ against matched deterministic and adapted generative baselines, while requiring $\mathbf{13}$--$\mathbf{77\times}$ less recurrent dynamics-step compute. On a held-out $256^2$ Kolmogorov flow, LGS reduces the 1-step L2RE from $0.398$ to $0.129$ within five finetuning epochs; U-AFNO achieves only $0.653$ to $0.343$ under the same protocol.

Multimodal large language models (MLLMs) have heterogeneous strengths across OCR, chart understanding, spatial reasoning, visual question answering, cost, and latency. As a result, effective MLLM routing requires more than estimating query difficulty: a router must match the multimodal requirements of the current image-question input with the capabilities of each candidate model. We propose \textsc{LatentRouter}, a router that formulates MLLM routing as counterfactual multimodal utility prediction. Given an image-question query, \textsc{LatentRouter} extracts learned multimodal routing capsules, represents each candidate MLLM with a model capability token, and performs latent communication between these states to estimate how each model would perform if selected. A distributional outcome head predicts model-specific counterfactual quality, and a bounded capsule correction refines close decisions while limiting the influence of noisy routing signals. The resulting utility-based policy supports both performance-oriented and performance-cost routing, and can handle changing candidate pools through shared per-model scoring with availability masking. Experiments on MMR-Bench and VL-RouterBench show that \textsc{LatentRouter} outperforms strong fixed-model, feature-level, and learned-router baselines. Additional analyses show that the gains are strongest on multimodal task groups where model choice depends on visual, layout-sensitive, or reasoning-oriented requirements, and that latent communication is the main contributor to the improvement. The code is available at: https://anonymous.4open.science/r/LatentRouter-8718.


Lattice Deduction Transformers

Liam Davis ⋅ Alberto Alfarano ⋅ Leopold Haller ⋅ Mark Santolucito

We introduce the Lattice Deduction Transformer (LDT), a recurrent transformer that approximates logically sound deduction by projecting its latent state through a lattice between forward passes. We train on-policy in a process that mirrors deduction in a search-based constraint solver and supervise training via a domain-agnostic, abstract-interpretation-based approximation of the set of solution candidates. An $800$K-parameter LDT achieves $100\%$ accuracy on Sudoku-Extreme and Snowflake Sudoku, at a fraction of the training cost of prior small recurrent reasoners, while remaining empirically sound: the model returns a correct answer or abstains. A $1.8$M-parameter variant reaches $99.9\%$ accuracy on Maze-Hard. Frontier LLMs score $0\%$ on all three benchmarks.


Learnable Chernoff Baselines for Provable Inference-Time Alignment

Sunil Madhow ⋅ Yuchen Liang ⋅ Ness Shroff ⋅ Yingbin Liang ⋅ Yu-Xiang Wang

We study inference-time reward-guided alignment for generative models. Existing methods often rely on either architecture-specific adaptations or computationally costly inference procedures. We introduce Learnable Chernoff Baselines (LCBs) as a method for efficiently and approximately sampling from the exponentially tilted kernels that arise from KL-regularized reward alignment. Using only black-box sampling access to the pretrained model, LCBs implement a form of rejection sampling with adaptively selected acceptance probabilities, which allows fine-grained control over inference-compute scaling. We establish total-variation guarantees to the ideal aligned model, which reveal the dimension-independent quantities governing the tradeoff between accurate sampling and inference compute. We empirically demonstrate in both continuous and discrete diffusion settings that LCB sampling closely matches ideal rejection sampling, but uses substantially fewer queries to the pretrained model. Our experiments also include real-world data experiments on a diffusion large language model.


Learning Actionable Information Landscapes for Multimodal Active Sensing in Hawkmoths

Abdelrahman Sharafeldin ⋅ Yaqing Wang ⋅ Simon Sponberg ⋅ Hannah Choi

Animals integrate information across sensory modalities to guide exploration, suggesting that they maintain internal estimates of where actions are expected to reduce uncertainty. However, how such action-conditioned information maps are learned from multimodal sensory experience remains unclear, as previous studies have been largely confined to unimodal sensing frameworks. We introduce a multimodal predictive-coding framework for learning uncertainty-reduction maps from visual and mechanosensory observations in hawkmoth flower interaction tasks. A variational generative model trained on simulated environments recovers modality-specific information landscapes over the action space, capturing how visual patterns, surface geometry, and mechanosensory cues shape the value of exploratory actions. These learned maps account for several features of hawkmoth behavior, including angled offsets during flower tracking, probing along visual patterns, and geometry-dependent changes in exploration. Applying the model to behavioral data from flower tracking and nectary search, we find that natural trajectories tend to occupy regions of high uncertainty reduction in these learned maps. Building on these observations, we develop an uncertainty-gated lexicographic algorithm that switches between information gathering and reward exploitation based on perceptual uncertainty. Across downstream reward-learning tasks, this policy learns faster, is more sample-efficient, and generalizes better to unseen environmental conditions than exploitation-only policies, $\epsilon$-greedy exploration, and non-lexicographic baselines. Together, these results suggest that multimodal generative perception models can learn actionable information landscapes that explain biological sensing and support exploration in embodied decision-making.

We study online identification of linear-in-parameter dynamical systems whose true parameters drift over time. Measuring non-stationarity through the path variation $V_T = \sum_{t=1}^{T-1}\|\theta_{t+1}^\star - \theta_t^\star\|$, we make two contributions. First, we give a clean blockwise dynamic-regret analysis for continuous projected online gradient descent, yielding $E[R_T] = O(T/\sqrt{\tilde W} + \tilde W V_T)$ and hence the oracle-tuned rate $O(T^{2/3}V_T^{1/3})$ . Second, we propose a residual-based adaptive step size and prove an oracle-proxy comparison theorem: when the residual proxy tracks the local prediction-error signal with lower-order overhead, the adaptive algorithm inherits the oracle scaling. Experiments on synthetic systems and real turbofan-engine degradation data (NASA C-MAPSS) support the matched-regime slope prediction on synthetic data, diagnose the failure regime when proxy overhead dominates, and show a clear unsupervised health-indicator trend in the residual proxy on FD001. A separate fast-rate guarantee under persistent excitation and a tube MPC integration are deferred to the appendix.


Learning Distributions from Multiple Data Providers

Jon Kleinberg ⋅ Amin Saberi ⋅ Xizhi Tan ⋅ Grigoris Velegkas

Motivated by learning from heterogeneous and overlapping data providers, we study a stylized model of distribution learning from restricted conditional samples. The goal is to learn an unknown distribution $p$ on a finite domain $[n]$. The learner is given a fixed family of queryable sets $\mathcal S \subseteq 2^{[n]}$, and each query to $S \in \mathcal S$ returns an independent sample from the conditional distribution $p(\cdot \mid S)$. Learnability is governed by the \emph{co-occurrence graph} associated with $\mathcal S$: two domain elements are connected if they appear together in some queryable set. Pointwise consistency is achievable when this graph is connected on the target support. PAC learning requires more: it is possible when the co-occurrence graph is complete. The sample complexity of PAC learning ranges from nearly linear to quadratic. Every query family with complete co-occurrence graph admits sample complexity $\widetilde O(n^2/\epsilon^2)$, and this bound is tight in the worst case. On the other hand, if every subset is queryable, the optimal worst-case complexity improves to $\Theta(n/\epsilon^2)$. More generally, we identify \emph{hierarchical comparability} as a sufficient structural condition on $\mathcal S$ under which the optimal complexity is nearly linear, $\widetilde \Theta(n/\epsilon^2)$, with pairwise query families as a canonical example. Finally, the full range of polynomial rates between linear and quadratic is attainable: for every $\alpha \in (1,2)$, there exists a query family with optimal PAC rate $\widetilde \Theta(n^\alpha/\epsilon^2)$.


Learning Fractional-Order Dynamics from a Single Trajectory

Xiaole Zhang ⋅ Ziyi Zhang ⋅ zehao zhao ⋅ Stephen Tu ⋅ Guannan Qu ⋅ Yorie Nakahira ⋅ Paul Bogdan

Many real-world processes exhibit long-range dependence, where the current state depends on a slowly decaying trace of past states rather than on the most recent state alone. This paper studies system identification for discrete-time fractional-order linear time-invariant systems from a single observed trajectory of length $t$, a setting that captures such non-Markovian dynamics through the Grünwald-Letnikov difference operator. Unlike Markovian systems, fractional-order systems couple estimation across the entire history, making both statistical analysis and practical identification more challenging. We propose *Fractional-Order Ordinary-Least-Squares Grid-Search (FO-GS)*, a simple two-stage estimator that exploits the diagonal structure of the fractional-difference operator to decouple the identification problem row-wise. Under stability and regularity assumptions, we establish high-probability, non-asymptotic error bounds for estimating both the fractional order and the system matrix in the heterogeneous setting, with both estimation errors scaling as $\mathcal{O}(t^{-1/4})$. Through experiments, we show that *FO-GS* outperforms existing baselines in recovering both the fractional order and the underlying system dynamics.

Recent work has identified incremental learning in shallow networks trained on single-index and multi-index models. However, existing analyses often rely on simplifying settings, such as small initialization, correlation loss, or layer-wise training. These choices reduce neuron interactions and leave some feature learning dynamics under standard initialization unexplored. We study gradient flow dynamics for polynomial-width two-layer networks learning orthogonal multi-index targets under standard initialization using polynomially many samples. We first prove that incremental learning still occurs: the loss decreases sequentially according to the Hermite expansion of the target, with lower-order components learned before higher-order components recover the individual target directions. In this standard initialization regime, training also shows a competitive reallocation of parameter mass: after the total mass fits the target mean and stabilizes, mass shifts into the target subspace and then concentrates on aligned neurons. Technically, we introduce a symmetry-based finite-width approximation via symmetrized networks, rather than comparing directly with an infinite-width limit. This yields better control of approximation errors and may be of independent interest.


Learning Provable Neural Network Observer for Uncertain Dynamical Systems

Zhangyi Wang ⋅ Jiaxu Liu ⋅ Chen Song ⋅ Chao Xu ⋅ Shengze Cai

In many safety-critical applications, control of uncertain dynamical systems relies on observers that estimate states and external disturbances. Neural network observers can improve estimation accuracy, but certifying their Lyapunov stability via Linear Matrix Inequality (LMI) constraints leads to large-scale semidefinite programs (SDPs) that are difficult to solve for large networks. To overcome this scalability bottleneck, we propose a novel two-stage training framework for provably stable neural network observers. Our approach decouples the optimization into a point-guided Lyapunov pre-training phase, which rapidly achieves high estimation accuracy and local stability over sampled states, followed by an LMI fine-tuning phase that efficiently satisfies a strict global Lyapunov stability certificate. We provide formal theoretical guarantees for local stability radii and probabilistic coverage over a prescribed compact error-state domain under specified regularity and sampling assumptions. Experiments on nonlinear control benchmarks and X-29 aircraft ablations show that our LMI-certified neural network observers train significantly faster than direct LMI-based methods and generalize robustly across diverse systems, achieving improved tracking accuracy over a range of observer baselines. The code is available at https://anonymous.4open.science/r/LearningNeuralNetworkObserver-2E1E.


Learning to Correct Geometry in Generated Videos

Weijie Lyu ⋅ Tiancheng SHEN ⋅ Xiangtai Li ⋅ Yujing Wang ⋅ Xiaoyuan Wang ⋅ Haobo Yuan ⋅ Yizhou Zhao ⋅ Ming-Hsuan Yang

Generated videos often lose geometric consistency under camera motion: structures warp under viewpoint change, parallax becomes inconsistent, and smooth camera movement may collapse into shot-like transitions. We argue that this failure is largely a data problem. In web-scale video corpora, most visible motion comes from dynamic objects captured by static or weakly moving cameras, leaving only sparse supervision for viewpoint-induced appearance changes. To address this, we propose Video Geometry Corrector (VGC), a post-hoc video-to-video model trained on synthetic paired supervision. Starting from real videos with diverse camera motion and pose annotations, we construct VG-Pairs with a synthetic paired-supervision data generation pipeline. Each video clip is treated as a geometry-consistent target and paired with a distorted counterpart synthesized by a pretrained image-to-video inpainting model using trajectory-adaptive keyframes. VGC learns to map the distorted input to the geometry-consistent target without modifying the source generator, and uses optical flow as an auxiliary cue to help disentangle camera motion from object motion. Experiments on pose-conditioned and prompt-driven benchmarks show that VGC improves geometric consistency across multiple source generators, and scaling to a larger backbone further extends the framework toward more general video refinement.

Benders decomposition (BD) is a widely used solution approach for solving two-stage stochastic programs arising in real-world decision-making under uncertainty. However, it often suffers from slow convergence as the master problem grows with an increasing number of cuts. In this paper, we propose Reinforcement Learning for BD (RLBD), a framework that adaptively selects cuts using a neural network-based stochastic policy. The policy is trained using a policy gradient method via the REINFORCE algorithm. We evaluate the proposed approach on a two-stage stochastic electric vehicle charging station location problem and compare it with vanilla BD and LearnBD, a supervised learning approach that classifies cuts using a support vector machine. Numerical results demonstrate that RLBD achieves substantial improvements in computational efficiency and exhibits strong generalization to problems with similar structures but varying data inputs and decision variable dimensions.

In many machine learning applications, human and LLM evaluators use assessments of relevant criteria to create an overall evaluation for an item or individual. In applications like admissions, committees assess candidates on attributes such as test scores, GPA, and research experience to evaluate their overall fit for the program. In medical care, clinicians use patient reports of symptoms to consider preliminary diagnoses and assess risks. Each case involves mapping measurable criteria to an overall evaluation—a process that reflects the evaluator's underlying preferences. We focus on the fundamental issue of learning these preferences. Many applications of this problem make specific modeling assumptions on evaluator preferences that may be substantially violated in the real world. We make the minimal assumption that the preference function is coordinate-wise non-decreasing, which is reasonable in a large number of evaluation settings. We theoretically characterize the severity of model mismatch for many common assumptions and show that it can lead to catastrophic effects for learning evaluator preferences and other important downstream tasks. We then present an algorithm for learning evaluators' preferences that is robust to model mismatch. We prove theoretically that our algorithm can learn any monotonic preference function without sacrificing performance when the linearity assumption holds. Evaluations of our algorithm with synthetic simulations and real-world data confirm its ability to learn preferences robustly and illustrate key aspects of LLM and human preferences.


LESSViT: Robust Hyperspectral Representation Learning under Spectral Configuration Shift

Haozhe Si ⋅ Yuxuan Wan ⋅ Yuqing Wang ⋅ Minh Do ⋅ Han Zhao

Modeling hyperspectral imagery (HSI) across different sensors presents a fundamental challenge due to variations in wavelength coverage, band sampling, and channel dimensionality. As a result, models trained under a fixed spectral configuration often fail to generalize to other sensors. Existing Vision Transformer (ViT) approaches either rely on implicit spectral modeling with fixed channel assumptions or adopt explicit spatial–spectral attention with prohibitive computational cost, leading to a fundamental trade-off between efficiency and expressiveness. In this work, we introduce Low-rank Efficient Spatial–Spectral ViT (LESSViT), a sensor-flexible architecture for cross-spectral generalization. LESSViT is built on LESS Attention, a structured low-rank factorization that models joint spatial–spectral interactions through separable spatial and spectral components, reducing the complexity of full spatial–spectral attention from $\mathcal{O}(N^2 C^2)$ to $\mathcal{O}(rNC)$, where $N$ is the number of spatial tokens, $C$ is the number of spectral channels, and $r$ is the rank of the low-rank approximation. We further incorporate channel-agnostic patch embedding and wavelength-aware positional encoding to support flexible spectral inputs. To enable efficient and robust pretraining, we introduce a hyperspectral masked autoencoder (HyperMAE) with decoupled spatial–spectral masking and hierarchical channel sampling. We evaluate LESSViT under a cross-spectral generalization setting that simulates cross-sensor variability. Experiments on the SpectralEarth benchmark demonstrate that LESSViT improves robustness under spectral shifts while remaining competitive in-distribution, and explicit and efficient spatial–spectral modeling is essential for scalable and generalizable hyperspectral representation learning.


Leveraging Soft Prompts for Privacy Attacks in Federated Prompt Tuning

Quan M Nguyen ⋅ Min-Seon Kim ⋅ Hoang M Ngo ⋅ Nghia Hoang ⋅ HYUK-YOON KWON ⋅ My T. Thai

Membership inference attacks (MIAs) pose a serious privacy threat in federated learning (FL). While MIAs have been extensively studied in standard FL, the recent shift toward federated fine-tuning introduces new and largely unexplored attack surfaces. In this work, we show that federated prompt-tuning, which adapts pre-trained foundation models using lightweight input prefixes, exposes a novel and effective vector for membership inference. We propose PromptMIA, a membership inference attack tailored to federated prompt-tuning, in which a malicious server introduces adversarially crafted prompts and exploits their updates during collaborative training to determine whether a target data point belongs to a client’s private dataset. We formalize this threat via a security game and demonstrate that PromptMIA achieves consistently high attack advantage across diverse benchmark datasets, substantially outperforming current SOTA federated MIAs. We also provide a theoretical lower bound on the attack advantage that explains the observed empirical behavior. Finally, we show that existing MIA defenses are often ineffective against PromptMIA, highlighting the need for defense mechanisms specifically tailored to prompt-tuning in federated settings.


LiFT: Lifted Inter-slice Feature Trajectories for 3D Image Generation from 2D Generators

Xinhe Zhang ⋅ Yuyang Zhang ⋅ Pengfei Jin ⋅ Arnau Marin-Llobet ⋅ Na Li ⋅ Quanzheng Li

High-resolution 3D medical image generation remains challenging because fully volumetric models are computationally expensive, while efficient 2D slice generators often fail to preserve anatomical consistency across the third dimension. We propose LiFT, a framework for Lifted Inter-slice Feature Trajectories that factorizes 3D volume synthesis into per-slice image generation and inter-slice trajectory learning. Rather than modeling the volumetric distribution end-to-end, LiFT treats a volume as an ordered trajectory in feature space, capturing how anatomical structures appear, transform, and disappear across depth. A tri-planar drifting loss aligns the trajectory of generated slices with the trajectories of real volumes, enabling distributional learning over inter-slice progressions in unconditional generation; in paired translation, a bidirectional $z$-context mixer trained against the registered target supplies through-plane coherence while preserving per-slice fidelity. We evaluate LiFT on BraTS 2023 (unconditional and missing-modality MR) and SynthRAD2023 (MR-to-CT). Across these settings, LiFT preserves per-slice quality, approaches the reported cWDM missing-MR reconstruction quality, and improves through-plane coherence on MR-to-CT relative to a no-mapper ablation, demonstrating that lightweight inter-slice trajectory learning is a viable route to high-resolution 3D medical synthesis.


LIPAR: Latent Inter-Frame Pruning with Attention Recovery

Dennis Y Menn ⋅ Yuedong Yang ⋅ Bokun Wang ⋅ Xiwen Wei ⋅ Mustafa Munir ⋅ Feng Liang ⋅ Radu Marculescu ⋅ Chenfeng Xu ⋅ Diana Marculescu

Video generation enables text-to-video synthesis, video editing, and motion-controlled content creation. However, current video generation models suffer from high computational latency, rendering true real-time capabilities infeasible for down stream tasks. We address this limitation by exploiting the temporal redundancy inherent in video latent patches. To this end, we propose the Latent Inter-frame Pruning with Attention Recovery (LIPAR) framework, which detects and skips recomputing duplicated latent patches. Additionally, we introduce a novel Attention Recovery mechanism that approximates the attention values of pruned tokens, thereby removing visual artifacts arising from naively applying the pruning method. Empirically, our method increases generation throughput by $1.45\times$, on average achieving 12.2 FPS on an NVIDIA A6000 compared to the baseline 8.4 FPS. The proposed method does not compromise generation quality and can be seamlessly integrated with Diffusion Transformer without additional training. Our approach effectively bridges the gap between traditional compression algorithms and modern generative pipelines.


LLM Judge Validation Under Sparse Overlap: From Inference to Design

Junxuan Li ⋅ Arko Mukherjee ⋅ Soumyabrata Pal

Validating an LLM-as-a-judge requires estimating its agreement with humans, yet annotation budgets rarely allow every item to be multiply labeled. We prove that this overlap sparsity is the first-order determinant of wrong deployment decisions: at 5\% pairwise overlap, wrong-decision rates reach 25\% and the probability of selecting the wrong best judge among ten candidates is 65\%. The two actionable levers are overlap quantity and allocation. For quantity, we derive a minimum-overlap formula showing $\rho \geq 0.25$ suffices for non-borderline judges while borderline cases remain fundamentally hard. For allocation, a zero-cost stratified scheme halves false-rejection rates relative to random sampling when strata are informative. We validate on 10 LLM judges across four evaluation matrices spanning visual assessment, causal reasoning, and summarization.


Long-Context Language Models Require Extreme Sparsity in Context Dimension

Prithvi Dixit ⋅ Sahil Joshi ⋅ Agniva Chowdhury ⋅ Anshumali Shrivastava ⋅ Joseph Gonzalez ⋅ Ion Stoica ⋅ Kumar Krishna Agrawal ⋅ Aditya Desai

Sparsity has long been a central theme in LLM efficiency but its role in context processing remains unresolved. As LLM workloads shift toward longer contexts and agentic interactions, the compute and memory bottlenecks of attention become increasingly critical, raising the question of whether these constraints are fundamental. Our position: this constraint is artificial, unnecessary and the future of LLM inference lies in extreme but principled sparsity along the context dimension. Our position comes out of various bits of empirical and theoretical evidence. Firstly, we find insistance on dense attention unreasonable since in long context a query essentially projects the O(N) attention information into hidden space of diemsion $d << N$. This, as we show is extremely lossy. Instead, we propose a completely context-sparse context processing combining top-k sparsity with linear attention / SSMs as the way forward. To support our proposal we show that this combination can approximately cover the functional space of dense softmax attention. Independently but importantly, in a first elaborate study of its kind, we empirically show a strong trend towards current LLM models, which are not trained for context sparsity, being extremely robust to inference time decode-sparsity across tasks of varying complexities such as retrieval, multi-hop QA, and mathematical reasoning and agentic coding. For instance, Qwen3.5-27B can tolerate upto 100x sparsity on benchmarks of RULER-HARD, LOFT and AIME2025 without loss of quality, and upto $50\times$ on SWE with a small drop in quality. These results emphasize the possibility that we can transition to complete sparsity without any loss of capability. Additionally, we also discuss what this shift in paradigm of context processing means for hardware. Importantly, we show that even current hardware is equipped enough to realise gains from this sparsity. For instance, our sparse decode kernels can accelerate large context processing by a factor 10x over FlashInfer at 50x sparsity levels on current hardware such as H100. Overall, these results position extreme context sparsity not as a heuristic, but as a principled foundation for LLM inference, training, and architecture design, both feasible and beneficial, and a compelling direction for future systems.


LongMINT: Evaluating Memory under Multi-Target Interference in Long-Horizon Agent Systems

Hyunji Lee ⋅ Justin Chen ⋅ Joykirat Singh ⋅ Zaid Khan ⋅ Elias Stengel-Eskin ⋅ Mohit Bansal

Agents in real-world settings operate over long and evolving horizons, where information is repeatedly updated and can interfere with each other across memories, requiring accurate recall and aggregated reasoning over multiple pieces of information. However, existing benchmarks focus on static, independent recall and fail to capture these dynamic interactions between evolving memories. In this paper, we study how current systems perform in realistic, continuously evolving, long-horizon settings across diverse domains and question types. To this end, we construct an analytical benchmark, LongMINT (Long-Horizon Memory under INTerference), which features (1) long, highly interconnected contexts with frequently updated information, (2) diverse coverage across multiple memory domains (Wikipedia, code, multi-turn dialogue, and state tracking), enabling evaluation of domain generalization, and (3) diverse question types, including (i) single-target recall tasks that test retrieval under interference over long contexts and (ii) multi-target aggregation tasks that require counting, ordering, or reasoning across multiple relevant pieces of information. We evaluate over six representative systems, including vanilla long-context LLMs, retrieval-augmented generation methods, and memory-augmented agent frameworks. We observe consistently low performance (avg. 27.7\% accuracy), especially on questions that require aggregated reasoning over multiple pieces of evidence. Fine-grained analysis shows that performance is primarily limited by retrieval and memory construction capabilities. Furthermore, current memory systems struggle to recall and reason over facts that are multiple steps back, and performance decreases when this lookback distance increases. These findings highlight the need for more robust memory management systems for dynamic, long-horizon environments across varying domains.


LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression

Zishan Shao ⋅ Lixun Zhang ⋅ Kangning Cui ⋅ Wenhao Wu ⋅ Jinhee Kim ⋅ Yixiao Wang ⋅ Ting Jiang ⋅ Hancheng Ye ⋅ Qinsi Wang ⋅ Fan Yang ⋅ Danyang Zhuo ⋅ Yiran Chen ⋅ Hai Li

SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs). However, meaningful comparison across existing studies remains difficult as prior evaluations use varied benchmarks, inconsistent ratios, and diverse setups, often failing to isolate low-rank effects from auxiliary techniques. As a result, it remains unclear whether reported gains reflect method-level improvements or differences in evaluation protocol. This lack of comparability highlights the need for a unified, reproducible evaluation platform. To address this problem, we present $\textbf{LowRankArena}$, a standardized evaluation platform for SVD-based LLM compression. LowRankArena unifies task versions, uniform-precision compression budgets, comparison regimes, and serving measurements, and provides a reproducible pipeline with over 3 TiB released compressed checkpoints. Using LowRankArena as a first standardized audit, we revisit representative publicly reproducible SVD methods under matched settings and observe that conclusions from isolated evaluations become more conditional once these fundamental evaluation assumptions are aligned: rankings change across backbones and keep ratios, multiple-choice accuracy can hide large perplexity degradation, and nominal low-rank savings yield workload-dependent and often limited end-to-end speedups. Our code is available at: https://anonymous.4open.science/r/lowrankarena-7883.


MaxIM: Maximally Informative Incremental Summarization via Reinforcement Learning

Jihwan Jeong ⋅ Guy Tennenholtz ⋅ Yinlam Chow ⋅ Chih-wei Hsu ⋅ Craig Boutilier

AI systems must often maintain bounded memory of unbounded interaction histories. However, optimizing the memory representation, and its updates, for downstream task utility is difficult for standard reinforcement learning (RL) due to severe credit assignment bottlenecks induced by sparse, delayed rewards. We introduce MaxIM (Maximally Informative Incremental Memory), a framework that formulates incremental summarization as a capacity-constrained, information-theoretic reward-shaping problem. Grounded in value suboptimality bounds, we prove that minimizing sequential information loss requires maximizing the predictive log-likelihood of future, task-relevant targets (predictive sufficiency). We introduce a recursive consistency condition to ensure summaries are grounded by maximizing mutual information with the observed history. We mitigate credit assignment issues with a parameterized critic that computes step-wise variational lower bounds for dense reward shaping. Empirical evaluation on long-horizon benchmarks shows that MaxIM outperforms state-of-the-art baselines on comprehensive autorater metrics and downstream task utility, without expensive human preference labels.


Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions

Stéphane Aroca-Ouellette ⋅ Ian Berlot-Attwell ⋅ Panagiotis Lymperopoulos ⋅ Abhiramon Rajasekharan ⋅ Tongqi Zhu ⋅ Herin Kang ⋅ Kaheer Suleman ⋅ Sam Pasupalak

Despite rapid progress in artificial intelligence, current systems struggle with the interconnected challenges that define real-world decision making. Practical domains such as business management require open-ended optimization, actively learning environment dynamics from sparse experience, planning over long horizons in stochastic settings, and reasoning over spatial information. Yet no existing human-AI benchmarks assess how well agents integrate these challenges in a grounded decision-making context. To this end, we introduce Mini Amusement Parks (MAPs), an amusement-park simulator designed to evaluate an agent’s ability to model its environment, anticipate long-term consequences under uncertainty, and strategically operate a complex business. We provide expert human performance and a comprehensive evaluation of state-of-the-art agents, finding experts outperform these systems by 11.4× on easy mode and 15.3× on medium mode. Our analysis reveals persistent weaknesses in long-horizon planning, sample-efficient learning, spatial reasoning, and modelling uncertainty. By unifying these challenges within a single environment, MAPs offers a new foundation for benchmarking agents capable of adaptable decision making.


Mission Impossible: Diagnosing and Fixing Non-Operative Instruction Following in Image Editing

Guoyizhe Wei ⋅ Feng Wang ⋅ Alan Yuille ⋅ Rama Chellappa

Instruction-guided image editing models are mostly trained and evaluated on executable requests, yet many real user instructions are non-operative for the given image. Reliable editing therefore requires reasoning not only about how to edit, but also whether to edit and what to ignore. We introduce \bench{}, a large-scale benchmark for non-operative image editing, with 1.1M instructions over 31K images across seven categories spanning symbolic, grounded, physical, and logical non-executability, plus a 35K-sample partial-edit test set evaluating selective execution when one clause is feasible and another is not. A 700-sample internal audit confirms a 94.9\% true no-op rate with $\kappa=0.87$. Evaluating 13 systems reveals a clear action--abstention dissociation: frontier editors achieve strong standard editing quality but still over-edit more than 40\% of no-op samples. Three complementary perceptual, feature-level, and vision-language metrics further surface distinct failure modes across architectures. As a reference baseline, JUDGE-THEN-EDIT BAGEL—a feasibility-aware BAGEL variant trained on a 150K mixture of full-edit, partial-edit, and no-op data—reduces over-edit rate from 41.7\% to 12.6\%. Our results show that abstention is a distinct capability that current benchmarks fail to measure.

Data heterogeneity and low client participation are two key challenges in federated learning (FL). Client-reshuffling-based FL methods were recently introduced to improve participation efficiency by visiting each client once per meta-epoch; however, the resulting \textit{without-replacement} sampling induces inter-round dependence and conditional bias. As a consequence, existing client-reshuffling methods can still suffer from the data heterogeneity challenge due to this dependence. To bridge this gap, we propose \textbf{FedCDR}, a client-reshuffling FL algorithm built on Douglas–Rachford splitting. FedCDR supports \textit{inexact} local proximal updates via iterative solvers, enabling a practical communication–computation trade-off. For smooth nonconvex objectives, FedCDR with inexact local solvers attains a state-of-the-art $O(\epsilon^{-1})$ communication complexity to reach an $\epsilon$-approximate stationary point (i.e., $\mathbb{E}\|\nabla f(\tilde{x})\|^2 \le \epsilon)$, with the leading constant that is \textbf{independent of data heterogeneity} (i.e., it does not scale with common measures of heterogeneity). Technically, our analysis operates at the meta-epoch level: we control the deviation between reshuffled and full-client updates, construct a tailored potential function with provable descent, and sum over each meta-epoch to eliminate reshuffling-induced dependence. Experiments on synthetic tasks and benchmark datasets under heterogeneous partitions, including a 10,000-client setting, demonstrate consistent improvements over strong baselines and their client reshuffling variants.


MITO: A Millimeter-Wave Dataset and Simulator for Non-Line-of-Sight Perception

Tara Boroushaki ⋅ Laura Dodds ⋅ Cusuh Ham ⋅ Fadel Adib

The ability to observe the world is fundamental to reasoning and making informed decisions on how to interact with the environment. However, optical perception can often be disrupted due to common occurrences like occlusions, which can pose challenges to existing vision systems. We present MITO, the first millimeter-wave (mmWave) dataset of diverse, everyday objects, collected using a UR5 robotic arm with two mmWave radars operating at different frequencies and an RGB-D camera. Unlike visible light, mmWave signals can penetrate common occlusions (e.g., cardboard boxes, fabric, plastic). MITO captures over 24 million mmWave frames and uses them to generate 550 high-resolution mmWave images in line-of-sight and non-light-of-sight (NLOS), as well as RGB-D images, segmentation masks, and raw mmWave signals, taken from 76 different objects. We develop an open-source simulation tool that can be used to generate synthetic mmWave images for any 3D triangle mesh. Finally, we establish benchmarks for NLOS segmentation and classification to demonstrate the utility of our dataset and simulator for enabling broader NLOS perception.


MonoPhysics: Estimating Geometry, Appearance, and Physical Parameters from Monocular Videos

Daniel Rho ⋅ Jun M Choi ⋅ Matthew Thornton ⋅ Biswadip Dey ⋅ Roni Sengupta

Existing inverse physics methods recover physical parameters from multi-view videos, where geometric constraints across views resolve scale and 3D structure. In monocular settings, however, such constraints are absent, leading to severe scale ambiguity, inaccurate geometry, and weak coupling between appearance optimization and physical simulation. We propose MonoPhysics, a framework for monocular inverse physics estimation of deformable objects using differentiable MPM simulation and 3D Gaussian Splatting, which jointly optimizes geometry, appearance, and physical parameters from a single camera view. We address these challenges through three visual-physical bridges: global scale alignment, physics-aware geometry refinement, and a differentiable position map, which together enable accurate optimization from image-space losses alone. We evaluate on Vid2Sim and our new dataset of elastic and plastic objects, showing that MonoPhysics outperforms existing baselines in monocular settings and achieves performance comparable to multi-view baselines using only a single camera.

Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion homogenization. Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are typically deployed and upgraded simultaneously, making it difficult to attribute gains to specific components. To advance a more scientific understanding of LLM output diversity, we design a factorial experiment that separates two primary intervention dimensions: input conditioning (operationalized through persona depth) and interaction architecture. We evaluate all conditions on 100 real-user open-ended questions across 7 models, measuring diversity with multiple complementary metrics. Our findings challenge several common assumptions. First, more persona detail does not monotonically increase diversity. The initial step of persona conditioning already captures the majority of the gain, while further elaboration with demographic detail does not consistently improve and can reduce diversity on some models. Second, rather than seeking a single best interaction architecture, we find that different architectures explore largely non-overlapping opinion regions. Combining multiple architectures yields broader coverage than optimizing any one. Third, commonly attempted low-cost alternatives such as raising sampling temperature and adding diversity instructions produce negligible effects compared to structured interventions. Overall, our work demonstrates that diversity is not a product of scaling along any single dimension, but is highly sensitive to the structural form and combination of interventions. The field needs to move toward empirically grounded diversity strategy design over assumption-driven exploration.


Motion-o: Trajectory-Grounded Video Reasoning

Bishoy Galoaa ⋅ Shayda Moezzi ⋅ Xiangyu Bai ⋅ Sarah Ostadabbas

Recent video reasoning models increasingly produce spatio-temporal evidence chains that localize objects at specific timestamps. While these traces improve interpretability by grounding $\textit{where}$ and $\textit{when}$ evidence appears, they often leave the motion connecting observations, the $\textit{how}$, implicit. This makes dynamic and trajectory-dependent claims difficult to supervise, verify, or penalize when unsupported by the video. We formalize this missing component as Spatial-Temporal-Trajectory (STT) reasoning and introduce $\textbf{Motion-o}$, a motion-centric extension to vision-language models (VLMs) that makes trajectories explicit and verifiable. Motion-o augments evidence chains with Motion Chain of Thought (MCoT), a structured pathway that represents object motion through a discrete $\texttt{\}$ tag summarizing direction, speed, and scale change. To supervise MCoT, we densify sparse spatio-temporal annotations into object tracks and derive motion descriptors from centroid displacement and box-area change. We then train with complementary rewards for trajectory consistency and visual grounding, including a perturbation-based signal that penalizes motion descriptions that remain unchanged when temporal evidence is removed. Across multiple video understanding benchmarks, Motion-o consistently improves trajectory-faithful reasoning without architectural modifications. These results suggest that an explicit motion interface can complement existing VLM pipelines by converting implicit dynamics into verifiable evidence.


Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer

Aratrika Mustafi ⋅ Soumya Mukherjee ⋅ Bharath Sriperumbudur

We develop a gradient flow on the space of probability measures defined on matrix-valued parameters induced by regularized Muon, an analytically smoothed version of the idealized Muon optimizer. The key observation is that the regularized orthogonalization map is the gradient of a smooth Fenchel-dual smoothing of the nuclear norm. This identifies the (regularized) Muon update as a mirror/prox step in the update variable, with momentum acting as the dual coordinate. We use this structure to lift Muon from a single matrix parameter to finite-particle probability objectives of the form $J(\rho)=R\left(\int F d \rho\right)$, a setting motivated by mean-field descriptions of neural-network training, and derive the inertial continuous-time limit. Using this structure, we derive the finite-particle continuous-time limit under the inertial scaling of step size and momentum, and then pass to a phase-space mean-field equation over probability laws on parameter-momentum pairs. The resulting flow can be shown to be a damped Hamiltonian probability dynamics whose kinetic energy is induced by the regularized Muon mirror potential. We prove an exact Hamiltonian dissipation identity, showing that the Hamiltonian energy decreases monotonically- leading to the descent of the target objective functional. Under appropriate gradient dominance assumptions, we obtain continuous and discrete time exponential convergence rates. We also study the well-posedness of the mean field limit equation and establish propagation of chaos guarantees for the interacting particle system. Finally, we extend the formulation to Hilbert-valued feature maps on product matrix spaces, yielding a blockwise Muon probability flow applicable to smooth transformer mixture-of-experts models.

MSConsensus: A Hundred-Million-Scale, Batch-Effect–Suppressed Dataset and Benchmark for Proteomics Machine Learning Machine learning for proteomics is bottlenecked by three coupled and under-explored challenges: (i) noisy mass-spectrometry (MS) data dominated by instrument- and protocol-specific batch effects, (ii) slow data loading from legacy XML- and text-based file formats, and (iii) a fragmented evaluation landscape that prevents meaningful cross-study comparison. Each silently constrains what proteomics models learn, how efficiently they train, and how reliably they are compared. We address all three with MSConsensus, a multi-task dataset and benchmark built natively for proteomics ML. MSConsensus contains 110,209,043 consensus spectra distilled from 1.01 PB of public PRIDE data spanning 1,500 repositories and all major Orbitrap and timsTOF platforms, with per-peak ion annotations and stratified ML splits across instruments, species, and post-translational modifications. Unlike single-representative spectral libraries (MassIVE-KB v2, NIST, ProteomeTools), MSConsensus uses a new Weighted-Score Binning (WSBIN) algorithm that emits multiple consensus spectra per (peptide, charge) cluster, preserving intra-cluster variance required by contrastive and representation-learning objectives while suppressing context-specific artifacts. Across five held-out PRIDE repositories and four search engines (Comet, MetaMorpheus, MSFragger, Sage), WSBIN improves signal-to-noise ratio by an average of +3.86 dB (range +0.57 to +10.47 dB) without sacrificing protein, peptide, or PSM identification counts. To eliminate the I/O bottleneck during model training, we pair MSConsensus with MSCompress and a new msz binary container. msz compresses to 45% of mzML size losslessly and supports 25,378 random-access spectra/sec — 3.2× HDF5 — through a drop-in PyTorch / JAX dataloader. A unified public leaderboard standardizes evaluation across three core tasks—spectrum embedding, intensity prediction, and de novo peptide identification—and we report results for ten widely used systems (Casanovo, InstaNovo, Prosit, MS²PIP, GLEAMS, Spec2Vec, Comet, MetaMorpheus, MSFragger, Sage). Replacing raw training data with MSConsensus stabilizes Prosit's spectral angle on a never-before-seen 2025 instrument (PXD053296) from 0.6964 to 0.8137 and improves Spec2Vec top-1 retrieval from 39.2% to 46.7% (+19.0% relative). By jointly addressing scale, batch-effect suppression, I/O efficiency, and standardized evaluation, MSConsensus provides foundational infrastructure for training, evaluating, and comparing future proteomics models. The dataset is released on Harvard Dataverse (doi:10.7910/DVN/NUCT4N) under CC BY 4.0; MSCompress, MSDatasets, preprocessing pipelines, and baseline scripts are released on GitHub under Apache 2.0, with Docker containers for reproducibility.


Multi-Objective Causal Bandits: Minimal Intervention Space and Policy-Level Learning

Muhammad Qasim Elahi ⋅ Mahsa Ghasemi ⋅ Murat Kocaoglu

Many decision-making problems involve multiple objectives, where actions affect several performance metrics, fairness criteria, or safety constraints. We study multi-objective causal bandits with multiple reward nodes, where actions correspond to interventions on subsets of variables in a causal graph and performance is evaluated through Pareto optimality. Existing approaches to multi-objective bandits typically define optimality at the level of individual actions, using pairwise dominance or scalarization. In contrast, randomized decision rules induce convex combinations of action-level reward vectors, so the relevant achievable set is the convex hull of these rewards. Thus, an intervention may be suboptimal even if no single intervention dominates it, because it can be dominated by a randomized policy over other interventions. We formalize this policy-level phenomenon and show that it has consequences for both causal action-space reduction and online learning. First, we provide a necessary-and-sufficient graphical characterization of possibly Pareto-optimal minimal intervention sets (PPOMISs). Our characterization yields a minimal, sound, and complete candidate intervention family from the graph alone, correcting prior formulations that may include intervention sets that are never Pareto-optimal under any compatible structural causal model. Second, over this reduced intervention space, we design an online UCB-style algorithm that eliminates actions using dominance tests against convex combinations of other actions, and prove logarithmic gap-dependent Pareto regret. Finally, we study constrained multi-objective causal bandits, where feasibility and optimality may be realized only with randomized policies, and develop an online UCB-style constrained bandit algorithm with sublinear regret and constraint-violation guarantees.


Multi-site PPG: An In-the-Wild Physiological Dataset from Emerging Multi-Site Wearables

Jiayi Shao ⋅ Jiaying Ye ⋅ ShengYao Liu ⋅ Zachary Englhardt ⋅ Girish Narayanswamy ⋅ Vikram Iyer ⋅ Qiuyue (Shirley) Xue

Wearables are widely used for mobile health monitoring, and photoplethysmography (PPG) is a key sensing modality for heart rate and related physiological measurements. However, public in-the-wild PPG datasets remain largely wrist-centric or limited to short, controlled studies, constraining research on emerging wearable form factors. We present Multi-site PPG, an in-the-wild physiological dataset collected from four custom-developed unobtrusive wearables: a smart earring, ring, watch, and necklace. Each device records green and infrared reflective PPG, 3-axis acceleration, and temperature with timestamps for cross-device alignment, while a Polar H10 chest strap provides reference electrocardiogram (ECG). Participants wore the devices for one to multiple days during daytime activities while continuing their normal routines. The dataset contains over 350 hours of raw data and 230–290 hours of preprocessed, modeling-ready 8-second windows data per wearable. We benchmark heuristic, supervised, and self-supervised heart-rate estimation methods, showing substantial body-site differences: the best methods achieve mean absolute errors (MAEs) of 2.30 bpm on the earring, 5.13 bpm on the ring, 8.37 bpm on the watch, and 8.68 bpm on the necklace. We further analyze motion effects and evaluate multi-site and PPG–accelerometer fusion, demonstrating the dataset’s value for robust physiological sensing across emerging wearable form factors.


NeuralFieldManifold: Reconstruction of LFP manifold with Lag Embedding

Kasra Fallah ⋅ Haoyu N Chen ⋅ Rudramani Singha ⋅ Eunji Kong ⋅ Gergely Turi ⋅ Attila Losonczy ⋅ Erfan Zabeh

Local field potentials (LFPs) are population-level neural signals central to brain-computer interfaces and systems neuroscience, yet unlike spike-based population codes, their dynamics lack a principled geometric description that connects spectral structure to latent state-space geometry. Here we establish, analytically and empirically, that the lag-embedded dynamics of LFP signals lie on a low-dimensional K-torus, where K equals the number of sustained oscillatory modes present in the signal. Modeling LFPs as bounded autoregressive processes, we show that each oscillatory component contributes an independent circular degree of freedom, while aperiodic 1/f structure and noise contribute only geometric thickness around the manifold without altering its topology. To apply this theory to real nonstationary recordings, we introduce DeepLagField, a physics-informed network that jointly estimates time-varying AR structure and effective model order, with formal guarantees that local toroidal geometry is preserved despite drifting oscillatory dynamics.We validate the predicted toroidal geometry using persistent homology across primate visual cortex LFP, rodent hippocampal LFP, and mouse cortical EEG recordings, confirming the expected Betti number signatures across species and recording modalities. Critically, we demonstrate that this geometric structure is not merely descriptive but carries behaviorally relevant information inaccessible to standard spectral summaries — as evidenced by torus parameters derived purely from manifold geometry outperforming multi-band spectral features for sleep-state decoding without any hand-crafted frequency design. This proof of concept points toward broad downstream utility wherever oscillatory field signals are recorded, from neural decoding and brain-state monitoring to clinical biomarker development. These results reframe neural oscillations not as isolated spectral features but as coordinate directions of a low-dimensional delay manifold, opening a geometry-first approach to neural signal analysis that is simultaneously theoretically grounded, empirically validated across species, and predictive of behaviorally relevant brain states.


nnTrace: Detecting and Localizing Silent Bugs in Distributed Training

Haitian Jiang ⋅ Shaowei Zhu ⋅ Zhen Zhang ⋅ Zhenyu Song ⋅ Xinwei Fu ⋅ Zhen Jia ⋅ Yida Wang ⋅ Jinyang Li

Distributed training is essential for scaling LLM training across thousands of GPUs. However, as distributed training requires complex implementations, they are prone to silent bugs, which do not produce explicit error signals but lead to in correct training outcomes. Common debugging practices based on monitoring training loss or gradient norm curves are slow or unable to detect bugs and also do not help localize bugs. We design and implement nnTrace, the first systematic differential testing system for detecting and localizing silent bugs in distributed training. nnTrace aligns intermediate tensors from distributed training with those from a trusted reference implementation. To properly compare the floating-point values in the corresponding tensors, we propose a novel mathematical analysis that provides a guideline for setting tolerances, enabling nnTrace to distinguish bug-induced errors from numerical errors. Experimental results demonstrate that nnTrace effectively detects 11 existing bugs and 3 new bugs in the widely used Megatron-LM framework. nnTrace is effective in various training recipes, including low-precision recipes involving BF16 and FP8. Notably, Megatron-LM has already adopted the method proposed by nnTrace in its development workflow. Our code is available at https://anonymous.4open.science/r/NECK-3C61/.


No One Knows the State-of-the-Art in Geospatial Foundation Models

Isaac Corley ⋅ Caleb Robinson ⋅ Nils Lehmann ⋅ Gabriel Tseng ⋅ Anthony Fuller ⋅ Hamed Alemohammad ⋅ Evan Shelhamer ⋅ Jennifer Marcus ⋅ Hannah Kerner

Geospatial foundation models (GFMs) have been proposed as generalizable backbones for disaster response, land-cover mapping, food-security monitoring, and other high-stakes Earth-observation tasks. Yet the published work about these models does not give reviewers or users enough information to tell which model fits a given task. We argue that nobody knows what the current state of the art is in geospatial foundation models. The methods may be useful, but the GFM literature does not standardize evaluations, training and testing protocols, released weights, or pretraining controls well enough for anyone to compare or rank them. In a 152-paper audit, we find 46 cross-paper disagreements of at least 10 points for the same model, benchmark, and protocol; 94/126 papers with extractable pretraining data use a configuration no other paper uses; and 39\% of GFM papers release no model weights. This lack of community standards can be solved. We propose six concrete expectations: named-license weight release, shared core evaluations, copied-versus-rerun baseline annotations, variance reporting, one shared evaluation harness, and data-vs-architecture-vs-algorithm controls. These gaps are a coordination failure, not a fault of any individual lab; the authors of this paper, like many others in the GFM community, have contributed to them. Rather than just critiquing the community, we aim to provide concrete steps toward a shared understanding of how to innovate GFMs.


Not All Spines Are Created Equal: How CT Segmenters Fail on Lumbosacral Transitional Vertebrae

Gregory Schwing ⋅ Patrick Schwing ⋅ Miraziz Ismoilov ⋅ Nizar Alnabahneh ⋅ Loren Schwiebert

Wrong-level spine surgery remains a persistent never-event, with the Lumbosacral Transitional Vertebra (LSTV, 5–35% prevalence) as its largest anatomic driver. We audit TotalSegmentator — the most widely deployed CT segmentation system and default back-end of 3D Slicer — and show that on lumbarization cases its output is anatomically indistinguishable from a wrong-level surgical plan: lacking an L6 class, it shifts lumbar labels caudally by one level, with junction-DSC dropping 29 points on Any-LSTV. No existing CT benchmark surfaces this — not TotalSegmentator's own evaluation, not Li et al. (2025) (the per-class DSC frontier on the matched COLONOG split), and not VERIDAH (vertebra-labeling SOTA on a private CT cohort, operating on pre-localized crops with no pelvis). We contribute: (i) an LSTV-stratified evaluation protocol (per-class DSC, junction-DSC over a 40 mm L5/S1 window, per-class voxel confusion at the junction) translating the level-shift mechanism into quantifiable wrong-level surgical risk; (ii) CTSpinoPelvic1K — 1,153 CT volumes across 802 patients with unified spinopelvic masks, 33 LSTV-positive cases across a 6-way phenotype taxonomy, and radiologist Castellvi typing (I–IV) on all 33; (iii) the first publicly-released merge-based LSTV-handling dual-spinopelvic CT segmenter, exceeding TotalSegmentator on sacrum (TS=0.817 zero-shot) and uniquely delivering both spine and pelvis in a single forward pass with released weights; (iv) an empirical isolation showing VERIDAH's training-side L5/L6 merge eliminates the collision driving TotalSegmentator's level shift but exposes a residual L4/last_lumbar collision on count-style sacralization, motivating VERIDAH's sequence predictor as the necessary disambiguation component. CTSpinoPelvic1K, the protocol, and 5-fold checkpoints are released publicly.


On Computing Diverse Solutions in the Earth Movers Distance

Aritra Banik ⋅ Mayank Goswami ⋅ Abhishek Sahu

Classically used for image retrieval tasks in computer vision, the Earth Movers Distance (EMD), also called the Wasserstein distance, has found numerous applications in natural language processing (NLP) and machine learning (ML). In NLP it is used as a measure of distance between sets of embeddings, and in ML it has been used to understand training dynamics, distributionally robust optimization, and many other tasks under active research. On the other hand, computing diverse solutions to optimization problems has also gained a lot of attention recently. In this DiverseX paradigm, one wants to develop algorithms that return a set of $r$ solutions that are maximally diverse; a common measure of diversity is the average or the minimum distance between the $r\choose{2}$ pairs of solutions. This paradigm is useful in generating more choices for the user, in fairness, and in robustness and security applications. In this work, we address the complexity of finding a diverse set of solutions in the EMD metric. Given an $n$ point metric space $(X,d)$ and integers $k \geq 1$ and $r \geq 2$, we consider the problem of computing $r$ many subsets of $X$, each of cardinality $k$, such that the minimum EMD between these sets is maximized. Motivated by applications from NLP, we also consider the problem of computing the farthest $k$-subset (from a given $k$-subset) in the EMD metric. On the lower bound side, we first show that assuming the Maximum-Span Hypothesis, it is W[1]-hard (with parameters $k$ and $r$) to obtain $k/(\log k)^{O(1)}$-approximation for non-metric cost functions, and W[1]-hard to obtain a $2-o(1)$ approximation for metric spaces. We also show that the problem restricted to the Euclidean setting with $\ell_2$ norm is W[1]-hard. Our first main algorithmic result is an $f(k, d, \varepsilon)n^{(O(1)}$ time $(1-\varepsilon)$ approximation algorithm for the $d$-dimensional Euclidean setting, which is tight in view of the above hardness results. Using different techniques, we also present a similar result for the farthest point problem in arbitrary metric spaces. Our second main algorithmic result is geared towards the search for polynomial time algorithms, where we present an $O(d)$ approximation for the Euclidean setting in $\text{poly}(n,k,d)$ time when $r=2$. Finally, we present $FPT(k)$ time, 2-approximate algorithms for arbitrary metric spaces when $r=2$. Since the class of problems we study requires to *find* diverse solutions in the EMD metric, our results use a combination of various geometric techniques that deviate from the techniques used to compute the EMD metric between a *given pair* of solutions, and may be of independent interest.

Clustering is a fundamental problem in statistics and machine learning. We propose the first one bit clustering method for two-component sub Gaussian mixture models. The method uses only one bit per entry of each sample obtained via a dithered quantizer. Under a mild non-spikiness condition on the cluster centers, we show that a variant of Lloyd’s algorithm achieves a misclassification rate that decays exponentially with a signal to noise ratio comparable to that in the unquantized setting. This result further implies exact recovery under an explicit separation condition, which exceeds the optimal threshold for unquantized data by only a logarithmic factor. When the dimension $p$ is sufficiently large, the non-spikiness condition can be enforced by applying a random rotation using a Haar distributed matrix prior to quantization. In particular, it holds with high probability when $p \gtrsim 1$ for partial recovery and $p \gtrsim \log n \log\log n$ for exact recovery, where $n$ is the sample size. We also establish a minimax lower bound, showing that the misclassification rate and separation condition are in general optimal up to constants. Numerical results are provided to corroborate the theory and demonstrate the efficacy of the proposed method.

A central challenge in reinforcement learning (RL) is to learn models that generalize beyond the tasks on which they are trained, a goal traditionally pursued through multi-task and meta RL. Recently, transformer architectures have emerged as a promising approach, enabling adaptation to new tasks via in-context learning without explicit parameter updates. From a functional perspective, a transformer can be viewed as a functional operator that maps a context to a task-specific function. It is thus fundamental to understand and design this operator to support stronger generalization in RL. In this work, we address this resulting question of generalization from a kernel-based perspective by establishing a connection between non-linear transformers and kernel-based temporal difference learning. By interpreting the transformer as performing regression in a Reproducing Kernel Hilbert Space (RKHS), we show that value functions from different domains can be represented using a shared set of weights, provided they lie within the same RKHS. Experiments on multiple MetaWorld domains support this interpretation, demonstrating convergence of the temporal-difference objective.

Differential Privacy (DP) provides a rigorous framework for quantifying privacy guarantees. While most methods achieve DP by injecting calibrated random noise, the intrinsic randomness of certain procedures such as sampling can yield DP guarantees ``for free''. In this work, we develop a unified R\'enyi divergence framework to characterize the inherent DP guarantees of sampling from a Bayesian posterior distribution. We show that the privacy loss incurred by One-Posterior-Sampling (OPS) is governed by posterior exponential moments of the single-record log-likelihood ratio, and establish explicit DP guarantees under uniformly bounded, sub-Gaussian, or sub-exponential tail regimes. We apply the framework to representative Bayesian models, including categorical likelihoods with arbitrary priors, Gaussian-Gaussian conjugate pair, generalized linear models from the exponential family and linear regression with Gaussian likelihood, both with Gaussian priors on the regression coefficients. Our results elucidate sample size, model structure, neighboring relations in DP (bounded or unbounded), and how prior if applicable, jointly determine the inherent DP guarantees of OPS. Furthermore, our results recover the existing DP guarantees for OPS as special cases -- often with tighter privacy loss bounds -- and substantially broaden the class of Bayesian models for which the inherent DP guarantees of OPS can be rigorously characterized.

Learning rate is a critical component of reinforcement learning (RL). This work uses global and local clocks to distinguish two types of learning rates. The former is of the standard form $\alpha_t$ that depends only on the time step $t$ (i.e., a global clock). The latter is of the form $\alpha_{\nu(S_t, t)}$, where $\nu(s, t)$ counts the number of visits to state $s$ until time $t$ (i.e., a local clock). In discounted RL, an RL algorithm that is convergent with a local clock is always also convergent with a global clock, and vice versa. We are not aware of any counterexample. The key contribution of this work is to show that this nice correspondence breaks down in average-reward RL. Specifically, we construct a counterexample showing that although differential temporal difference learning is convergent with a local clock, it can diverge with a global clock. This counterexample closes the open problem in Wan et al. [2021], Blaser et al. [2026].


On the Optimal Sample Complexity of Offline Multi-Armed Bandits with KL Regularization

Kaixuan Ji ⋅ Qiwei Di ⋅ Heyang Zhao ⋅ Qingyue Zhao ⋅ Quanquan Gu

Kullback-Leibler (KL) regularization is widely used in offline decision-making and offers several benefits, motivating recent work on the sample complexity of offline learning with respect to \emph{KL-regularized performance metrics}. Nevertheless, the exact sample complexity of KL-regularized offline learning remains largely from fully characterized. In this paper, we study this question in the setting of multi-armed bandits (MABs). We provide a sharp analysis of KL-PCB~(Zhao et al., 2026), showing that it achieves a sample complexity of $\tilde{O}(\eta SAC^{\pi^\*}/\epsilon)$ under large regularization $\eta = \tilde{O}(\epsilon^{-1})$, and a sample complexity of $\tilde{\Omega}(SAC^{\pi^\*}/\epsilon^2)$ under small regularization $\eta = \tilde{\Omega}(\epsilon^{-1})$, where $\eta$ is the regularization parameter, $S$ is the number of contexts, $A$ is the number of arms, $C^{\pi^\*}$ is the policy coverage coefficient at the optimal policy $\pi^*$, $\epsilon$ is the desired sub-optimality, and $\tilde{O}$ and $\tilde{\Omega}$ hide all poly-logarithmic factors. We further provide a pair of sharper sample complexity lower bounds, which matches the upper bounds over the entire range of regularization strengths. Overall, our results provide a nearly complete characterization of offline multi-armed bandits with KL regularization.


On the Provable Emergence of Hierarchical Concept Structure in CLIP Embeddings

Nghia Nguyen ⋅ Tianjiao Ding ⋅ Lachlan MacDonald ⋅ Rene Vidal

Contrastive vision-language models such as CLIP are trained to align images and text, not to represent taxonomies. Yet empirical studies show that CLIP embeddings exhibit a hierarchical geometry aligned with the semantic hierarchy of concepts. A theoretical characterization of why such hierarchy emerges from a non-hierarchical objective remains open. In this paper, we formalize this phenomenon as Hierarchical Angular Separation (HAS), a geometric property whereby each non-leaf concept embedding is more similar to its descendants than to off-branch concepts. Empirically, we show that HAS already holds weakly in randomly initialized CLIP image embeddings, suggesting that hierarchical structure may be induced by the raw image geometry, while in text embeddings it emerges only after contrastive training. To explain this asymmetry between text and image embeddings, we analyze an unconstrained features model in which the image embeddings are fixed and only the text embeddings are trained. We show that when the image embeddings satisfy HAS with a positive margin, the global minimizer of the contrastive objective induces HAS in the text embeddings, thereby transferring hierarchical structure across modalities. The key mechanism is the hierarchical co-occurrence structure of image-text pairs, which biases text embeddings toward root-centered image prototypes. For a linear contrastive loss, we derive a closed-form solution that makes this mechanism explicit. We extend the analysis to the InfoNCE loss and establish analogous guarantees under sufficient regularization. Synthetic experiments validate the theory and show that HAS emerges even beyond the regimes covered by our analysis.


OpInf-LLM: Parametric PDE Solving with LLMs via Operator Inference

Zhuoyuan "Jacob" Wang ⋅ Hanjiang Hu ⋅ Xiyu Deng ⋅ Saviz Mowlavi ⋅ Yorie Nakahira

Solving diverse partial differential equations (PDEs) is fundamental in science and engineering. Large language models (LLMs) have demonstrated strong capabilities in code generation, symbolic reasoning, and tool use, but reliably solving PDEs across heterogeneous settings remains challenging. Prior work on LLM-based code generation and transformer-based foundation models for PDE learning has shown promising advances. However, a persistent trade-off between execution success rate and numerical accuracy arises, particularly when generalization to unseen parameters and boundary conditions is required. In this work, we propose OpInf-LLM, an LLM parametric PDE solving framework via operator inference. The proposed framework leverages small amounts of solution data to enable accurate prediction of diverse PDE instances, including unseen parameters and configurations, and provides seamless integration with LLMs for natural language task specification and physics-based reasoning of proper feature parameterization. Its low computational demands and unified solution pipeline further enable a high execution success rate across heterogeneous settings, opening new possibilities for generalizable reduced-order modeling in LLM-based PDE solving.

In-context learning enables large language models to adapt to tasks directly from input sequences, without parameter updates. We investigate the mechanism underlying in-context learning of autoregressive processes in transformers under heterogeneous second-order moments of prompt distributions. Our study considers a two-layer architecture consisting of a linear attention head and a nonlinear MLP, with the outer layer trained via gradient descent. We prove that there exists a parametrization under which the model converges to the unique function that approximates the optimal second-order moment-based AR(1) estimator. We analyze bounds on approximation and convergence rates and confirm our findings experimentally. These results extend the first-moment-centric understanding of estimators realized by in-context learning to second-order moment-based estimators, offering further insight into ICL's robustness to new tasks.


OptiWorld: Optimal Control for Video World Generation under Physical Constraints

Yu Yuan ⋅ Jianhao Yuan ⋅ Xijun Wang ⋅ Daiqing Li ⋅ Liu He ⋅ Lu Ling ⋅ Stanley Chan

Video generation models are becoming a scalable form of world models, but they mainly generate plausible motion rather than proactively control or optimize the underlying dynamics. As a result, an object in the generated video may follow trajectories that are unsafe, not smooth, inefficient, or physically inconsistent. In this work, we propose OptiWorld, a framework that brings classical optimal control into video generation at inference time. OptiWorld first extracts a compact, task-relevant world state, then plans an optimal trajectory under physical constraints, and finally renders the video conditioned on this trajectory. We formulate planning as a geometric problem on a continuous manifold, which converts 3D geometry and task-dependent physical constraints into a unified planning geometry. By adding this optimal-control layer, OptiWorld generates videos with preferable dynamics, demonstrating strong potential in multiple tasks including goal-conditioned image-to-video generation, video dynamics editing, and counterfactual generation.


Order-Optimal Sample Complexity for Distribution Learning via Flow Matching

Hari K Sahoo ⋅ Mudit G Gaur ⋅ Vaneet Aggarwal

Flow-based generative models have emerged as an efficient alternative to diffusion models. We study the sample complexity of learning a target distribution in Wasserstein distance via flow matching, which learns a velocity field transporting a base distribution to the target. Under standard assumptions on the network architecture and data distribution, we show that any squared-loss flow-matching variant including rectified flow with bounded derivatives achieves $\widetilde{O}(\varepsilon^{-2})$ sample complexity with the Wasserstein guarantee taking the form $W_2 \le K(\varepsilon + C\sqrt{\varepsilon_{app}})$, where $\varepsilon_{app}$ is the approximation error of the network class. This improves upon existing $\widetilde{O}(\varepsilon^{-4})$ guarantees and does not require the Polyak \L{}ojasiewicz condition assumed in prior work. Instead, we exploit a structural property of the squared-loss objective - it induces an approximate Bernstein condition that directly ties the variance of the excess loss to the excess risk up to additive approximation-error. This condition enables a localized Rademacher complexity analysis yielding fast $O(1/n)$ rates. The flow-matching schedule enters only through the Lipschitz constant of the pointwise loss, which we characterize explicitly for standard schedules.

Goal-conditioned planning requires compressing value functions into low-dimensional representations, yet which property of the compression best predicts planning quality remains unclear. We show that reconstruction accuracy and ordinal fidelity---how faithfully the ranking of successor states by value is preserved---are two projections of the same approximation error: in smooth geometry they couple tightly and $L^2$ is a sufficient summary; when topology creates dense small-gap regions they decouple, and the ordinal projection carries complementary planning-relevant information beyond $L^2$. Through experiments on gridworld, continuous navigation (640 runs, 8 environments), and Maze2D (240 runs, 3 D4RL layouts), we establish three results. First, neighbor-restricted ordinal fidelity ($\tau^{\mathrm{nbr}}$) adds $\Delta R^2 = 0.106$ of incremental explanatory power beyond $L^2$ after environment fixed effects. Second, monotone recalibration degrades $L^2$ by up to $\sim\!400\times$ while preserving $\tau^{\mathrm{nbr}}$ and planning success. Third, in matched-pair comparisons controlling for $L^2$, higher $\tau^{\mathrm{nbr}}$ predicts better planning 76\% of the time. We provide theoretical grounding via a gap-weighted Bellman-excess decomposition that links reconstruction error, local ordinal inversions, and planning regret.


Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion

Van Chien Nguyen ⋅ Chaitra Hegde ⋅ Van-Cuong Pham ⋅ Ryan Rossi ⋅ Franck Dernoncourt ⋅ Thien H Nguyen

We introduce Orthrus, a simple and efficient dual-architecture framework that unifies the exact generation fidelity of autoregressive Large Language Models (LLMs) with the high-speed parallel token generation of diffusion models. The sequential nature of standard autoregressive decoding represents a fundamental bottleneck for high-throughput inference. While diffusion language models attempt to break this barrier via parallel generation, they suffer from significant performance degradation, high training costs, and a lack of rigorous convergence guarantees. Orthrus resolves this dichotomy natively. Designed to seamlessly integrate into existing Transformers, the framework augments a frozen LLM with a lightweight, trainable module to create a parallel diffusion view alongside the standard autoregressive view. In this unified system, both views attend to the exact same high-fidelity Key-Value (KV) cache; the autoregressive head executes context pre-filling to construct accurate KV representations, while the diffusion head executes parallel generation. By employing an exact consensus mechanism between the two views, Orthrus guarantees lossless inference, delivering up to a $7.8\times$ speedup with only an $O(1)$ memory cache overhead and minimal parameter additions.


PAC Learning with Bandit Feedback: Sharp Sample Complexity in the Realizable Setting

Steve Hanneke ⋅ Qinglin Meng ⋅ Shay Moran ⋅ Amirreza Shaeiri

We study the problem of multiclass PAC learning with bandit feedback in the realizable setting. In this framework, there is an unknown data distribution over an instance space $\mathcal{X}$ and a label space $\mathcal{Y}$, as in classical multiclass PAC learning, but the learner does not observe the labels of the i.i.d. training examples. Instead, in each round, it receives an unlabeled instance, predicts its label, and receives bandit feedback indicating only whether the prediction is correct. Despite this restriction, the goal remains the same as in classical PAC learning. We provide a general characterization of the optimal sample complexity of this problem, sharp for every concept class up to logarithmic factors. Our characterization is based on a new combinatorial dimension, termed the *bandit $\mathrm{DS}$ dimension*, defined via generalized combinatorial structures we call *pseudo-boxes*. These extend the pseudo-cubes underlying the $\mathrm{DS}$ dimension by allowing a different number of neighbors in each coordinate. In contrast to the $\mathrm{DS}$ dimension, which governs the full-information setting by counting the number of coordinates in the pseudo-cube, the bandit $\mathrm{DS}$ dimension aggregates the number of neighbors across coordinates, leading to a characterization in which the sample complexity scales with the total number of neighbors. We also propose a general learning algorithm achieving the upper bound, based on an algorithmic principle called *ListCascade*, which connects bandit learning to list learning and may be of independent interest.

The standard constraint-based paradigm for causal discovery with incomplete data---impute first, test second---is frequently miscalibrated: any consistent conditional independence (CI) test rejects a true null with probability approaching 1 when imputation error induces spurious conditional dependence. We introduce PAIR-CI, a nonparametric CI test that restores calibration by integrating multiple imputation directly into the inferential procedure via a paired permutation design. PAIR-CI compares cross-validated models that include and exclude the candidate variable while receiving the same imputed conditioning set, forcing imputation error to cancel in their loss difference rather than contaminate the test statistic. A provably consistent variance estimator jointly accounts for uncertainty arising from cross-validation and multiple imputation---to our knowledge, the first formal unification of these two inferential frameworks. In simulations, existing imputation-based CI tests exhibit false positive rates of 28--45\% when data are missing not at random (MNAR), whereas PAIR-CI averages below the nominal 5\% level across data-generating processes and missingness mechanisms. These gains are largest in nonlinear settings and grow with causal graph size: when integrated into the PC algorithm, PAIR-CI reduces structural Hamming distance by 8\% on 10-variable nonlinear graphs, 15\% on 30-variable equivalents, and up to 44\% on the 56-variable HAILFINDER network, with stable performance in all settings.


Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs

Qitao Tan ⋅ Xiaoying Song ⋅ Arman Akbari ⋅ Arash Akbari ⋅ Yanzhi Wang ⋅ Xiaoming Zhai ⋅ Lingzi Hong ⋅ Zhen Xiang ⋅ Jin Lu ⋅ Geng Yuan

Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users and contexts. As a result, models may refuse requests that are unsafe for general users but legitimate for authorized professionals, limiting helpfulness in specialized professional settings. Existing approaches either require costly realignment or rely on inference-time steering that suffers from imprecise control and added latency. To this end, we propose \textsc{Palette}, a modular, controllable, and efficient framework that selectively relaxes refusal behavior on authorized target domains while preserving standard safety elsewhere. Our method identifies a refusal direction via multi-objective search and internalizes it into the model through lightweight adaptation. \textsc{Palette} further supports modular composition: it learns domain-specific safety controls independently and composes them through parameter merging, enabling on-demand multi-domain authorization without retraining. Experiments across four safety benchmarks, multiple model variants, and both LLMs and VLMs show that \textsc{Palette} delivers precise safety control without sacrificing general utility, offering a practical path toward foundation models that adapt to diverse professional needs.

We study a class of bilevel optimization problems in which both the upper- and lower-level problems are minimax problems. Existing studies on bilevel optimization primarily focus on settings where the lower-level problem is an unconstrained or constrained minimization problem, and are therefore not directly applicable to the setting of a minimax lower-level problem considered in this work. To address this gap, we develop penalty-based first-order methods for bilevel minimax optimization. In the deterministic setting, we prove that the proposed method finds an $\epsilon$-KKT point with $\tilde{O}(\epsilon^{-4})$ oracle complexity. We further show that bilevel optimization problems with constrained lower-level minimization can be reformulated, via Lagrangian duality under Slater's condition, as special cases of our framework and hence can also be solved by our method. This yields an $\tilde{O}(\epsilon^{-4})$ complexity bound for finding an $\epsilon$-KKT point, improving upon the existing $\tilde{O}(\epsilon^{-7})$ result. Finally, we extend our approach to the stochastic setting and establish that the proposed stochastic method finds a nearly $\epsilon$-KKT point with $\tilde{O}(\epsilon^{-9})$ oracle complexity. To the best of our knowledge, these are the first deterministic and stochastic first-order complexity results for bilevel minimax optimization.

Building structured 3D scene layouts from a single image requires reconciling visual observations with physical and spatial constraints, a challenge that is difficult to address with direct prediction alone. In this work, we formulate monocular 3D layout estimation as a perceive-then-plan problem with vision-language models, where a Perceiver first grounds the 3D objects and then a Planner iteratively refines the scene hypothesis through actions that improve physical plausibility while preserving consistency with the input image. We propose Layout-as-Policy (LaP), which casts the planning stage as a policy learning problem: 3D layouts are represented as structured states, and refined via discrete actions such as translation, rotation, and rescaling. Starting from an observation-aligned initialization with the geometry-enhanced Perceiver, the LaP Planner is trained to produce action sequences that progressively resolve geometric inconsistencies and enforce realistic spatial relations. To enable effective learning, we combine supervised trajectory initialization with preference-based optimization, allowing the model to learn corrective behaviors without requiring explicit reward engineering. This formulation transforms layout estimation from a one-shot prediction task into an iterative refinement process, enabling better handling of global constraints and complex object interactions. Experiments demonstrate that our approach produces layouts that are more physically coherent and better aligned with visual observations, while naturally supporting downstream tasks such as scene editing and manipulation.


Physics Unrolled Neural Operator for Wireless Field Modeling

Rafid Umayer Murshed ⋅ Saif U Rahman ⋅ Mingyue Tang ⋅ Elahe Soltanaghai

Radio maps are essential for wireless decision-making tasks such as access-point placement, coverage planning, and localization, but their fine spatial details are governed by complex propagation effects and are costly to simulate accurately. Machine learning offers a path to high-fidelity radio-map prediction without running expensive high-fidelity simulations for every scene. However, generating high-quality training labels at scale is also difficult: the affordable labels come from finite-ray simulations, which are richer than low-fidelity inputs but carry residual Monte Carlo noise. We address this challenge with Physics-Unrolled Hybrid Neural Operator (PU-HNO), a three-stage cascade that predicts high-fidelity indoor radio maps from low-fidelity ray-tracing outputs and scene priors by progressively capturing reflection, diffraction, and scattering effects, rather than treating radio maps as generic images. We prove that, under conditionally unbiased label noise, the model can learn stable propagation structure and outperform its own training labels. Experiments across diverse floorplans show that PU-HNO outperforms image-to-image baselines, wireless learning models, and monolithic neural operators across both image-quality and wireless deployment metrics.


PithTrain: A Compact and Agent-Native MoE Training System

Ruihang Lai ⋅ Hao Kang ⋅ Haozhan Tang ⋅ Akaash R Parthasarathy ⋅ Zichun Yu ⋅ Junru Shao ⋅ Todd Mowry ⋅ Chenyan Xiong ⋅ Tianqi Chen

Mixture-of-Experts (MoE) has become the dominant architecture for frontier language models. To meet this demand, production frameworks have built optimized MoE training stacks over years of engineering effort. Yet evolving these stacks for new architectures and system optimizations remains expensive. With the rise of AI coding agents, they could automate parts of training-framework development and accelerate this evolution. But applying them to these existing frameworks carries hidden costs, invisible to today's throughput-only evaluations. We name this missing dimension agent-task efficiency (ATE): the cost of using coding agents to understand, operate, and extend a framework. Grounded in four agent-native design principles, we build PithTrain, a compact, agent-native MoE training framework. We further introduce ATE-Bench, covering real-world training-framework tasks. Our evaluation shows PithTrain matches the throughput of production frameworks, and on ATE-Bench, PithTrain enables higher agent-task efficiency, with up to 62% fewer Agent Turns and 64% less Active GPU Time.


PixelDiT2: Representation-Grounded Pixel Diffusion Transformers

Yongsheng Yu ⋅ Wei Xiong ⋅ Yichen Sheng ⋅ Shiqiu Liu ⋅ Jiebo Luo

Recent advances in pixel-space diffusion models have narrowed the image quality gap with latent-space diffusion, but still converge more slowly and lag behind in final image quality. We argue that a key reason is the lack of an explicit representation prior: unlike latent diffusion, which denoises in a compact and structured latent space, pixel diffusion must learn denoising-friendly representations and pixel generation simultaneously from raw RGB space. To address this problem, we propose PixelDiT2, an end-to-end pixel-space diffusion model designed to decouple representation learning from pixel generation without introducing an autoencoder or latent reconstruction bottleneck. We propose *representation grounding* that uses a frozen pretrained vision foundation model to provide explicit per-patch representation guidance throughout denoising, allowing the pixel diffusion transformer to focus more on pixel generation. On ImageNet-$256{\times}256$, PixelDiT2 reaches FID $\textbf{1.50}$ in $320$ epochs, surpassing JiT-G's FID $1.82$ at $600$ epochs with roughly half the parameters and half the training budget.


PLATO: Pointer Learner for Agent and Task Openness

Alireza Saleh Abadi ⋅ Leen-Kiat Soh ⋅ Daniel A Redder ⋅ Adam Eck ⋅ Prashant Doshi

Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental challenge to multi-agent reinforcement learning (MARL), which typically assumes fixed state and action spaces. Existing methods address openness only partially: padding and masking approaches introduce artificial bounds, while recent graph-based or hypergraph methods handle one dimension of openness but still depend on restrictive assumptions. In this paper, we introduce Pointer Learner for Agent and Task Openness (PLATO), a pointer-network-based actor combined with a centralized graph neural network (GNN) critic, trained with multi-agent proximal policy optimization under a centralized training and decentralized execution paradigm. Our pointer-based actor outputs distributions directly over the current task set. This directly supports changing action spaces without masking or retraining. Our GNN critic encodes agent–task interactions as a graph that changes shape with task and agent composition. Together, these components consider AO and TO without the boundedness of existing approaches. We formalize PLATO in a Task-and-Agent-Open Markov Game (TaAgO-MG), extending prior task-open formulations, and prove it is well-defined over the resulting unbounded state and action spaces. We evaluate PLATO with the Methods for Open Agent Systems Evaluation Initiative (MOASEI) wildfire suppression domain, an environment designed for open multi-agent system evaluation, and we demonstrate strong performance and more consistent zero-shot generalization than state-of-the-art baselines in OASYS.


Policy Regret Minimization in Partially Observable Markov Games

Lan Sang ⋅ Raman Arora ⋅ Thanh Nguyen-Tang

We study policy regret minimization in partially observable Markov games (POMGs) between a learner and a strategic adaptive adversary who adapts to the learner's past strategies. We develop a model-based optimistic framework that operates on the learner-observable process using *joint* MLE confidence set and introduce an Observable Operator Model-based causal decomposition that disentangles the coupling between the world and the adversary model. Under multi-step weakly revealing observations and a bounded-memory, stationary and posterior-lipschitz adversary and planner stability, we prove an $\mathcal{O}(\sqrt{T})$ policy regret bound. This work advances regret analysis from Markov games to POMGs and provides the first policy regret guarantee under imperfect information against an adaptive opponent.


PORTool: Importance-Aware Policy Optimization with Rewarded Tree for Multi-Tool-Integrated Reasoning

Feijie Wu ⋅ Weiwu Zhu ⋅ Yuxiang Zhang ⋅ Soumya Chatterjee ⋅ Jiarong Zhu ⋅ Fan Mo ⋅ Rong Luo ⋅ Jing Gao

Multi-tool-integrated reasoning enables LLM-empowered tool-use agents to solve complex tasks by interleaving natural-language reasoning with calls to external tools. However, training such agents from outcome-only rewards suffers from credit-assignment ambiguity, obscuring which intermediate tool-use decisions drive success or failure. In this paper, we propose PORTool, an importance-aware policy-optimization algorithm that reinforces agents' tool-use competence from outcome-level supervision while assigning reward at the step level. Specifically, PORTool generates a rewarded rollout tree in which trajectories share prefixes before branching, enabling direct comparisons among alternative tool-use decisions within the same context. It then estimates each step's importance by a correctness-dominant signal, i.e., whether descendants of that step can ultimately produce a correct final answer, plus an auxiliary term indicating whether the step's tool calls satisfy formatting constraints and execute successfully. Using these step-wise importance estimates, PORTool updates the policy to generate efficient tool-call steps, guided by both local comparisons within each branching decision and the overall quality of entire trajectories. Experiments show that PORTool improves final-answer accuracy while reducing tool-call steps compared with state-of-the-art policy-optimization baselines, and ablation studies confirm the robustness of the proposed step-wise importance estimates.


PosteriorBench: From Point Estimates to Posterior Matching in Evaluating Generative Inverse Solvers

Jiachen Yao ⋅ Zi-Siang Hsu ⋅ Xi Deng ⋅ Aditi Gupta ⋅ Xin Ju ⋅ Animashree Anandkumar

Generative models are increasingly used to solve scientific inverse problems, but existing evaluations still focus primarily on whether a method can produce a single plausible reconstruction. This is insufficient for ill-posed problems, where multiple solutions may be consistent with the same sparse or noisy observations. In these settings, a method can achieve strong pointwise accuracy while still failing to capture the true posterior through mode collapse, over-confident uncertainty, or averaging incompatible solutions. We introduce PosteriorBench, a benchmark for evaluating the distributional accuracy of generative inverse solvers. PosteriorBench evaluates four physics-based inverse problems: Darcy flow inversion, Poisson source recovery, carbon capture and storage, and light transport material inference. For each task, we construct high-fidelity reference posteriors using slow but established procedures such as rejection sampling and Markov chain Monte Carlo, enabling direct assessment of whether solvers recover the full set of solutions rather than the single best sample. The benchmark spans sparse sensing, low-resolution observations, nonlinear forward models, varying noise levels, and multimodal priors, with a unified pipeline for distribution matching, uncertainty quantification, and hyperparameter calibration. Our experiments reveal substantial distribution-matching gaps across current solvers, while showing that neural operators improve resolution robustness and that guidance weights and generation noise are key to posterior-variance calibration.


Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization

Tanmay Ambadkar ⋅ Sourav Panda ⋅ Shreyash Kale ⋅ Jonathan Dodge ⋅ Abhinav Verma

Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-conditioned policies offer a highly scalable solution, existing approaches remain brittle in practice, frequently failing to recover dense Pareto fronts. We demonstrate that this failure stems from two structural pathologies: destructive advantage cancellation caused by premature Early Scalarization (ES), and representational mode collapse across the preference space. To overcome these bottlenecks, we introduce \toolname{}, a PPO-based framework that fundamentally reorganizes multi-objective optimization. By preserving per-objective learning signals through a decomposed pipeline and integrating preferences only after trust-region stabilization (Late-Stage Weighting), \toolname{} improves credit assignment under conflicting objectives. Concurrently, a scaled diversity regularizer encourages behavioral divergence proportional to preference distance. \toolname{} operates entirely within the efficient linear scalarization regime shared by standard deep MORL baselines. By reducing information loss caused due to linear scalarization rather than relying on expensive non-linear utility functions, it suggests that optimization bottlenecks play a significant role. Across available standard benchmarks, including high-dimensional and many-objective environments, \toolname{} consistently discovers broader, higher-quality Pareto fronts than prior methods, exceeding state-of-the-art hypervolume and expected utility using a single deployable policy.


Preferential dynamic modeling with forward-backward smoothing

Omid G. Sani ⋅ Trisha Jha ⋅ Mohammad Hosseini ⋅ Maryam Shanechi

Estimating a secondary signal (e.g., behavior) from neural activity over time is central to both causal online decoding and non-causal offline inference in neuroscience, yet existing two-signal latent state-space models rarely support both. In this work, we provide an analytical extension of a linear method (PSID) beyond causal prediction to also support non-causal inference. We provide theoretical derivations extending PSID to enable optimal filtering and optimal smoothing of the secondary signal. We show that, in the PSID setting, the presence of a secondary signal increases identifiability. This allows us to uniquely learn the quantities needed for the optimal Kalman update via a reduced-rank regression step, yielding our first contribution, PSID with filtering. We next design a forward-backward construction for smoothing, yielding our second contribution, PSID with smoothing. In simulations, we validate that both PSID with filtering and smoothing reach ideal performance. In non-human primate motor cortex data, PSID with smoothing consistently improves over PSID with filtering, which improves over one-step-ahead prediction with standard PSID. Finally, we show that the connection between PSID and an existing nonlinear two-signal model, DPAD, extends naturally to the smoothing setting: applying our forward-backward formulation to DPAD yields DPAD with smoothing, which achieves competitive performance with leading methods on the Neural Latents Benchmark (NLB) in behavior decoding and held-out neural prediction across three datasets. Together, this work provides a theoretical foundation for prediction, filtering, and smoothing in the two-signal setting, spanning causal online decoding to offline inference, in both linear and nonlinear settings.


Primal-Dual Flow Matching for Sample-Wise Constrained Generation

Zhengyan Huan ⋅ Peter Y. Lu ⋅ Shuchin Aeron

We consider the problem of learning to generate data while satisfying a prescribed sample-wise constraint with a user-specified probability of constraint satisfaction. Starting with a population-level constrained negative log-likelihood minimization and using a flow-matching (FM) parameterization, under regularity assumptions and sufficient expressivity of the velocity field family, we derive guarantees on the probability of constraint satisfaction for the corresponding primal-dual optimization algorithm. We provide a practical implementation of this approach, called Primal-Dual Flow Matching (PDFM), which has direct control over constraint satisfaction rate and requires only binary membership-oracle feedback for constraint satisfaction, avoiding the need for differentiable distances, projections, or special structure of the constraint set, such as convexity. We show that compared to existing methods, PDFM achieves higher constraint satisfaction rates while maintaining competitive distribution fidelity.


Privacy-Preserving Retrieval-Augmented Generation with Plausible Deniability

Wenxuan Bao ⋅ Shan Jin ⋅ Vincent Bindschaedler ⋅ Yiwei Cai

We introduce **PD-RAG**, a technique for retrieval-augmented generation (RAG) that leverages a language model's own randomness to safeguard privacy. The algorithm partitions documents into groups, generates a candidate answer from a randomly chosen group, and then uses a privacy test to enforce that the released answer could have been produced by multiple disjoint document groups, thereby ensuring a controllable degree of *plausible deniability*. Compared to existing methods that operate at the token level and therefore incur both substantial utility loss and growing privacy budget per generated token, PD-RAG operates at the answer level, so the privacy it offers does not loosen with output length. We prove that PD-RAG satisfies $(\varepsilon,\delta)$-differential privacy. We experimentally evaluate PD-RAG on three QA benchmarks using three language models and find that it reduces membership inference attack advantage to near random while consistently outperforming alternative methods on all four utility metrics we used and running between $17$ and $33{\times}$ faster.


Provably Efficient Representation Learning for Low-Rank CMDPs

Kaixuan Liu ⋅ GUOJUN XIONG ⋅ Shengpu Tang ⋅ Wanyun Si ⋅ Jian Li

We study representation learning in low-rank Constrained Markov Decision Processes (CMDPs), where both value and reward functions are approximated by a set of unknown representation vectors, and the transition dynamics admit a low-rank factorization. The objective is to maximize expected cumulative rewards while complying with constraints on the expected cumulative utility. To achieve this, we propose \texttt{MFRLC} (Model-Free Representation Learning for Low-Rank CMDP), a model-free algorithm that uses a primal-dual scheme to effectively balance reward regret and constraint violations. To the best of our knowledge, \texttt{MFRLC} is the first representation-learning method for low-rank CMDPs that combines the Least-Squares Value Iteration with Upper Confidence Bound (LSVI-UCB) and primal-dual techniques. \texttt{MFRLC} further enhances value function estimation with bonuses for function approximation and uniquely learns the mapping from representations to value functions directly through representation learning. We prove that \texttt{MFRLC} achieves both regret and cumulative constraint violation of order $\widetilde O(H^3 d^2 K^{3/4}|\mathcal{A}|^{3/2}/\gamma)$, making it provably sample-efficient and highly adaptable to complex environments due to the reliance on function approximation.


Quantifying Centrality for Complex Data

Hang Zhou ⋅ Yidong Zhou ⋅ Hans-Georg Müller

Analyzing complex data across a wide range of scientific fields often requires identifying central or typical elements, distinguishing them from atypical ones, and constructing central or "normal" regions at prespecified levels of centrality. We introduce centrality scores (C-scores) based on distance profiles for random objects taking values in general metric spaces. The proposed C-scores are defined through weighted optimal transport of distance profiles and provide a unified framework for constructing central or "normal" regions for complex data. We show that, for Euclidean data, the limiting behavior of the proposed C-scores is asymptotically equivalent to the underlying probability density function. The practical merits of the proposed method are illustrated using gene expression data, distributions of recorded temperatures, handwritten digit recognition, and taxi trip records. These examples demonstrate the effectiveness of C-scores in identifying central and peripheral elements and in yielding meaningful insights for non-Euclidean and high-dimensional data.


Quantum Safe Stochastic Linear Bandits

Ruizhe Zhang ⋅ Junyi Wu ⋅ Guang Lin

We initiate the study of safe stochastic linear bandits under quantum feedback models. The learner must maximize an unknown linear reward while satisfying unknown linear constraints at every interaction. Under coherent unitary access to the joint reward-and-constraint distribution, we develop QRS-COLTS, a staged quantum analog of constrained linear Thompson sampling. With high probability, QRS-COLTS plays only feasible actions and achieves regret polylogarithmic in the quantum query budget; an isotropic multivariate estimator reduces the vector-feedback cost from linear to square-root dependence on the number of constraints, up to logarithmic factors. We also formulate a stronger query-weighted coherent-exploration model with a safe-action bomb flag. In this model, the learner may query action superpositions, the bomb provides exact but dangerous safe-set access, and regret is charged by the action weights in the query state. We present a quantum algorithm, BCO-QLinTS, that uses bomb-certified quantum convex optimization to select safe actions directly, thereby avoiding the need to estimate the constraint matrix. Its query-weighted regret is polylogarithmic in the horizon, with a bomb-safe optimization overhead that depends only polynomially on the number of constraints.


Response Time Enhances Alignment with Heterogeneous Preferences

Federico Echenique ⋅ Alireza Fallah ⋅ Baihe Huang ⋅ Michael Jordan

Aligning large language models (LLMs) to human preferences typically relies on aggregating pooled feedback into a single reward model. However, this standard approach assumes that all labelers share the same underlying preferences, ignoring the fact that real-world labelers are highly heterogeneous and usually anonymous. Consequently, relying solely on binary choice data fundamentally distorts the learned policy, making the true population-average preference unidentifiable. To overcome this critical limitation, we demonstrate that augmenting preference datasets with a simple, secondary signal—the user's response time—can restore the identifiability of the population's average preference. By modeling each decision as a Drift-Diffusion Model (DDM), we introduce a novel, consistent estimator of heterogeneous preferences that successfully corrects the distortions of standard choice-only labels. We prove that our estimator asymptotically converges to the true average preference even in extreme cases where each anonymous labeler contributes only a single choice. Empirically, across both synthetic and real-world datasets, our method consistently outperforms standard baselines that otherwise fail and plateau at a bias floor. Because response times are essentially free to record and require zero user tracking or identification, our results bring promises and open up new opportunities for future data-collection pipelines to improve the social benefit without requiring user-level identifiers or repeated elicitations.


Rethink Action Chunking in VLA Through Human Motor Control

Wenxi Chen ⋅ Yuejiang Liu ⋅ Zijian He ⋅ Shaoshuai Mou ⋅ Yan Gu

Action chunking is widely used in recent vision–language–action (VLA) systems to mitigate inference latency and improve temporal consistency. Yet, open-loop execution can reduce reactivity and precision, and cross-chunk coherence remains difficult to guarantee. We revisit action chunking through its neurobiological origins and organize current shortcomings around three axes: VLA architecture, representations of sequential actions, and mechanisms for sampling and executing chunks. A case study further illustrates the coherence–reactivity trade-off across representative sampling and execution choices. Building on these observations, we outline brain-inspired directions to improve the reactivity, coherence, and safety of chunking-based VLAs in dynamic and complex tasks beyond quasi-static tabletop manipulation. This position paper aims to broaden how the community thinks about action chunking and to motivate redesigns better aligned with future physical intelligence.


Rethinking Long-Video Efficiency: A Joint Allocation Perspective on Frames, Pixels, and Front-End Latency

Sixun Dong ⋅ Wei Li ⋅ Andong Deng ⋅ Qi Qian ⋅ Victor Zhu ⋅ Zhengping Ji ⋅ Chen Chen

Efficient long-video understanding with vision-language models (VLMs) has largely been treated as informative token selection at fixed native resolution: which frames or which visual tokens to retain under a token budget. We argue this framing leaves two axes unused --- per-frame resolution itself can be traded for denser temporal coverage, and front-end decoding latency scales with the candidate pool, not the token budget. Through an empirical study across multiple VLMs and long-video benchmarks, we distill three lessons: (i) dense low-resolution sampling outperforms sparse native-resolution sampling at matched token budgets; (ii) some tasks are resolution-sensitive and benefit from high-resolution frames; and (iii) front-end decoding dominates wall time on hour-scale clips. These lessons motivate LoHi, a training-free, single-pass framework that pairs a dense low-resolution video stream with a sparse set of high-resolution image streams, processed through the VLM's native video and image pathways. Two plug-and-play selectors choose Hi-I frames at near-zero or low overhead: LoHi-Anchor uses codec-level I-frame metadata, and LoHi-SemDiv uses a query-relevance and visual-diversity DPP over CLIP features. On three long-video benchmarks, LoHi improves over the vanilla native-resolution baseline by +10.6% on average at matched token budget and over the strongest prior efficiency methods by +5.2%, while reducing front-end decoding latency by up to 7x on hour-scale clips.


Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior

Rafal Kocielnik ⋅ Pengrui Han ⋅ Peiyang Song ⋅ Myrl Marmarelis ⋅ Ramit Debnath ⋅ Dean Mobbs ⋅ Animashree Anandkumar ⋅ R. Michael Alvarez

Anticipating LLM behavioral tendencies from low-cost psychometric probes is critical for safe deployment, but only if self-reports (SR) reliably predict behavior. Recent work documented substantial SR–behavior dissociation in LLMs, but relied on broad personality traits (Big 5) that predict specific behaviors weakly even in humans. Furthermore, the isolation of conversational sessions combined with weak context matching left open whether LLMs truly lack coherence or whether the conditions needed to detect such coherence were not met. We contrast Big 5 with the Theory of Planned Behavior (TPB), which measures intention targeted to a specific behavior and predicts human behavior substantially better than broad traits. We run experiments across four behavioral tasks and 11 frontier LLMs, while also varying session context and identity induction. We find that SR–behavior coherence exists but is selective. 1) Within a shared conversation, Theory of Planned Behavior reaches human-level coherence; Big 5 does not. 2) Across separate conversations, coherence survives only for behaviors anchored outside the immediate prompt, such as implicit bias shaped by training, and collapses when behavior is strongly primed by context, as with sycophancy. 3) Persona prompting makes self-reports more consistent across conversations but does not bring behavior into alignment. These findings suggest that coarse personality frameworks such as Big 5 may not be the best tools for testing deployment behavior. More task- and behavior-specific instruments are needed, and even these must be evaluated across tasks and contexts.

Graph Neural Networks (GNNs) represent valuable intellectual property, yet existing watermarking schemes primarily rely on OOD backdoor triggers that are susceptible to model pruning, fine-tuning, and distillation. To tackle this challenge, we present InvGNN-WM, which ties ownership to a model's implicit perception of a graph invariant, enabling trigger-free, black-box verification with negligible task impact. By training a scalar head to predict normalized algebraic connectivity on owner-private carrier graphs, ownership is embedded into the model's core reasoning logic rather than exogenous patterns. We provide guarantees for imperceptibility and robustness, and prove that exact removal is NP-complete under monotone decoders. Empirical evaluations across diverse node and graph classification datasets show that InvGNN-WM maintains clean task accuracy while outperforming trigger- and explanation-based baselines in watermark fidelity. Our method remains robust under unstructured pruning, fine-tuning, and post-training quantization, with clear recovery pathways under knowledge distillation.


Robust Hopfield Decision Transformer

Andreia Podasca ⋅ Anup Das

Decision Transformer (DT) formulates offline reinforcement learning as conditional sequence modeling, predicting actions by attending over past states, actions, and returns-to-go. However, the softmax attention in DT computes context representations in a single forward pass, offering no mechanism to recover from observations corrupted by sensor noise or measurement errors common in real-world deployment. Prior methods improve robustness of DT through training regularization or objective modifications, but leave the attention mechanism itself unchanged. We introduce Robust Hopfield Decision Transformer (RHDT), which replaces softmax attention with modern Hopfield layers that iteratively minimize an energy function, enabling observations to converge toward learned patterns (attractors). To prevent pattern interference from overlapping attractor basins, we regularize the architecture through Lipschitz gradient penalty and orthogonality constraints on attention keys, which serve as the stored patterns. On D4RL benchmarks, RHDT matches robust baselines in average return while significantly improving worst-case performance under observation corruption.


s2n-bignum-bench: A practical benchmark for evaluating low-level code reasoning of LLMs

Balaji Rao ⋅ Soonho Kong ⋅ Juneyoung Lee ⋅ Carlo Lipizzi

Recent progress in neural theorem proving has been driven largely by mathematics-oriented and, more recently, by verification-condition and repository-scale benchmarks. These settings are valuable, but they do not directly test whether a model can synthesize machine-checkable proofs about concrete low-level implementations under an industrial proof-engineering stack. We address this gap with *s2n-bignum-bench*, a benchmark derived from AWS *s2n-bignum*, a formally verified library of hand-tuned big-integer assembly routines for ARM and x86. The current corpus packages $\mathbf{2{,}301}$ HOL Light proof obligations as standalone `setup.ml`/`query.txt` tasks spanning big-integer arithmetic, elliptic-curve routines, ML-KEM, SHA-3/Keccak, and shared ISA infrastructure. Each task reproduces the relevant proving environment, exposes the theorem statement to be proved, and expects a tactic expression that is accepted by HOL Light within a fixed timeout. The benchmark spans five categories so that one can separate generic HOL Light fluency from ISA-specific reasoning. We release the extraction pipeline, retrieval utilities, an offline evaluator with syntax/type pre-checking and integrity checks based on axiom-difference and forbidden-tactic detection, paired verbatim and obfuscated query variants, and a checkpointed assessment workflow to support iterative feedback-driven $\operatorname{pass}@K$ studies and future step-level proof construction workflows. Initial zero-shot baselines using GPT-5.3-based and other frontier models solve at most $\mathbf{6.35}$% of the problem set (when provided only with the goal term), highlighting the difficulty of HOL Light tactic synthesis for low-level verification obligations under trusted ISA semantics. The code to set up and use the benchmark is available at [s2n-bignum-bench](https://github.com/kings-crown/s2n-bignum-bench).


Safeguarding LLMs via Model-Agnostic Latent Safety Signals from Dark Knowledge

Wonjun Lee ⋅ Kyungsik Yang ⋅ Gaeun Ji ⋅ Vaidehi Patil ⋅ Haon Park ⋅ Bumsub Ham ⋅ Mohit Bansal ⋅ Suhyun Kim

Large Language Models (LLMs) have advanced rapidly, raising growing concerns about their safety. Recent work has proposed various approaches to detect and defend against adversarial attacks including defense mechanisms at the decoding stage that leverage models' internal hidden states. However, existing decoding-stage defenses suffer from two limitations. First, they introduce a trade-off between safety and over-refusal, where strengthening safety degrades the model's helpfulness on benign queries. Second, many of these methods rely on internal hidden states and are thus restricted to specific architectures, incurring substantial overhead and limited generalization across models. To address these limitations, we introduce LADE (Latent Safety Signals for Defense), which leverages latent safety signals extracted by contrasting harmful and benign queries from dark knowledge (i.e., information carried by the output probability distribution beyond its argmax) in the first-token output probability distribution. Our key insight is that, beyond surface-level refusal tokens, the dark knowledge in the first-token distribution contains latent safety signals, defined as tokens whose probabilities differ sharply between harmful and benign queries. We empirically show that these signals consistently align across diverse LLMs, forming a model-agnostic direction that reflects an intrinsic property of safety-aligned language models. LADE consists of three components: (1) Extracting Latent Safety Signals from Dark Knowledge, which selects top-k safety-discriminative tokens from the first-token probability distribution; (2) Tokenizer Mapping, which maps these tokens across different tokenizers to enable model-agnostic application; and (3) kNN-based Discrimination, which classifies queries via a k-Nearest Neighbors search over the mapped tokens. Across six LLMs and multiple benchmarks, LADE remains robust against a wide range of jailbreak attacks and lowers attack success rates with minimal over-refusal.

Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process (CMDP) via primal-dual Lagrangian updates that only enforce safety on average. To address this limitation, hard, state-wise constraints are introduced and often imposed through Hamilton-Jacobi (HJ) reachability. Yet such constraints require solving different objectives in the feasible and infeasible regions of the state space: reward maximization in the former, recovery toward the feasible regions in the latter. The resulting target action distributions are inherently multimodal, and this structure poses a fundamental challenge for the Gaussian or deterministic actors used in existing HJ-based safe RL, which often collapse onto suboptimal modes. Diffusion policies provide the expressiveness needed to represent such distributions, and recent work on Q-score matching offers a route to training them for online RL by score regression---but has been applied only to reward maximization. Building on this framework, we propose Safe Score Matching (SSM), an off-policy actor-critic method that adapts Q-score matching to hard-constrained safe RL by gating a two-branch score target with HJ reachability: inside the feasible set, the denoising process degenerates to Q-score matching on actions classified as viable by the HJ critic, encouraging reward maximization; outside, a recovery branch biases denoising toward regions with lower worst-case violation, guided by the negative action gradient of the HJ safety value. Across high-dimensional quadrotor and fixed-wing trajectory-tracking and stabilize-and-avoid benchmarks, SSM achieves competitive performance without sacrificing safety, in contrast to primal-dual and reachability-based baselines, which tend to trade off performance for safety and are thus overly conservative.


Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents

Ziming Li ⋅ Jiatan Huang ⋅ Xiaoguang Guo ⋅ Guiling "Grace" Wang ⋅ Chuxu Zhang

Adaptive test-time compute for LLM agents aims to invoke extra computation only when it improves performance. Existing methods typically use confidence-, uncertainty-, or difficulty-based gates, assuming a fixed direction from the gating signal through compute need to the value of computation. This makes gating a utility-calibration problem: gating signals should align with whether extra computation improves the final outcome over the base policy. We show that this alignment is unstable: the same signal predicts rollout benefit in one setting and rollout harm in another, with reversals across environments and backbones even when the task is fixed. Wrong-direction gates can therefore worsen performance by precisely selecting harmful states. This reversal reflects a deeper distinction between compute need and compute suitability: a high uncertainty signal may indicate decision-difficult states where rollouts help compare alternatives, or intervention-unsuitable states where the current context does not support useful rollout-based improvement. Under this two-source model, fixed-direction gates are unreliable across heterogeneous settings. To address this, we propose DIAL (Direction-Informed Adaptive Learning), a sparse gate trained from signal-agnostic counterfactual exploration to learn the utility direction of state features per (environment, backbone). Across six environments and three backbones, DIAL yields a stronger overall success-cost trade-off than fixed-direction baselines. Our code is open-sourced at https://anonymous.4open.science/r/DIAL-2A7E/.

Large Language Model (LLM) agents have shown strong results on multi-turn tool-use tasks, yet they typically operate in isolation during training, failing to leverage skills accumulated across episodes. Existing experience-augmented methods address this by organizing past trajectories into retrievable libraries, but they retrieve skills only once based on the initial task description and hold them constant throughout the episode. In multi-turn settings where observations change at every step, this static retrieval becomes increasingly mismatched as episodes progress. We propose SAPO (Step-Level Skill-Augmented Policy Optimization), a reinforcement-learning framework that retrieves relevant skills at each decision step conditioned on the current observation. SAPO operates through three components: (i) step-level observation clustering that groups structurally equivalent environmental states for efficient cluster-indexed retrieval; (ii) a self-evolving skill bank that distills successful strategies and failure patterns through score-based admission and rate-limited extraction; and (iii) policy optimization with step-level credit assignment for fine-grained advantage estimation across multi-turn episodes. The skill bank evolves alongside the policy through semantic analysis rather than gradient updates. On long-horizon multi-turn agent benchmarks (ALFWorld, WebShop, and seven search-augmented QA tasks), SAPO achieves 93.5\% on ALFWorld, 76.3\% success on WebShop, and 60.9\% average across QA tasks, outperforming both standard RL and prior skill- and experience-augmented baselines. Our code is available at \url{https://anonymous.4open.science/r/slea-rl-4D1E/}.


Scalable Neural Safety Certification via Monotonicity

Amirreza Alavi ⋅ Majid Zamani ⋅ Saber Jafarpour

Learning-based safety certification offers a promising route for verifying complex dynamical systems when explicit models are unavailable, but existing data-driven methods often scale poorly: they treat the system as a black box, rely on dense discretizations or Lipschitz bounds, and consequently suffer from sample complexity that grows exponentially with dimension. This paper shows that a common structural property of dynamical systems---*monotonicity*---can be used to break this curse of dimensionality. We develop a data-driven framework for robust safety verification and safe controller synthesis for unknown monotone systems using only simulator queries. Our main theoretical result proves that, for monotone systems with upper-closed unsafe sets, safety is completely characterized by the existence of monotone inductive barrier certificates. This structure reduces certificate verification over continuous state spaces to localized boundary checks over finitely many cells. To learn such certificates, we introduce Max-Monotone Networks, a monotone neural architecture with universal approximation guarantees for monotone functions. We then propose a training and refinement algorithm that, upon successful termination, returns formally valid neural barrier certificates and, under mild partition assumptions, requires only $O(n)$ simulator queries. Across high-dimensional benchmarks, including oscillator networks with up to $13{,}659$ states and traffic networks with $1{,}000$ states, our method certifies safety in regimes where prior approaches either do not converge or exceed memory limits. These results demonstrate that exploiting order structure can make neural safety certification both scalable and formally sound.


Scaling Reward Modeling without Human Supervision

Jingxuan Fan ⋅ Yueying Li ⋅ Zhenting Qi ⋅ Dinghuai Zhang ⋅ Kianté Brantley ⋅ Sham Kakade ⋅ Hanlin Zhang

Learning from feedback is an instrumental process for advancing the capabilities and safety of frontier models, yet its effectiveness is often constrained by cost and scalability. We present a pilot study that explores scaling reward models through unsupervised approaches. We operationalize reward-based scaling (RBS), in its simplest form, as preference learning over document prefixes and suffixes drawn from large-scale web corpora. Its advantage is demonstrated in various aspects: despite using no human annotations, training on 11M tokens of math-focused web data yields steady gains on RewardBench v1 and v2, and these improvements consistently transfer across diverse initialization backbones spanning model families and scales. Across models, our method improves RewardBench v2 accuracy by up to +7.7 points on average, with gains of up to +16.1 on in-domain math subsets and consistent improvements on out-of-domain safety and general subsets. When applied to best-of-N selection and policy optimization, these reward models substantially improve downstream math performance and match or exceed strong supervised reward model baselines of similar size. Furthermore, using an RBS checkpoint as a mid-training initialization substantially amplifies gains from subsequent supervised fine-tuning. Overall, we demonstrate the feasibility and promise of training reward models without costly and potentially unreliable human annotations.


SCOT: Multi-Source Cross-City Transfer with Optimal-Transport Soft-Correspondence Objectives

Yuyao Wang ⋅ Min Yang ⋅ Meng Chen ⋅ Weiming Huang ⋅ Yilong Yin ⋅ Yongshun Gong

Cross-city transfer leverages labeled data from well-instrumented cities to improve prediction in label-scarce ones, but remains challenging when cities adopt incompatible partitions with no ground-truth region correspondences. Even with expressive GNN encoders, transfer quality varies dramatically across methods sharing nearly identical backbones---indicating that alignment design, not encoder capacity, is the binding constraint. Existing paradigms exhibit complementary failure modes: heuristic anchor matching collapses to hubness under unequal partitions, while distribution-level matching over-mixes embeddings under heterogeneity. Both stem from a single missing primitive---explicit, mass-controlled soft correspondence between unequal region sets. The challenge intensifies in multi-source transfer, where independent source-to-target alignments yield conflicting gradients and source domination. We propose SCOT, which adapts entropic OT to this regime through three application-specific designs: an OT-weighted contrastive objective that resolves the geometric--semantic tension, a one-sided cycle regularizer respecting the rectangular $n_s\!\neq\!n_t$ geometry, and---as our central contribution---a shared prototype hub coordinated through balanced entropic OT under a target-induced prior, bypassing the source-selection problem in label-scarce regimes. Across real-world cities and tasks, \scot consistently improves transfer accuracy, achieving 5--50\% relative MAE/MAPE reductions over the strongest baseline, with learned couplings and hub assignments quantitatively confirming that the diagnosed failure modes are resolved.


Search-Augmented Masked Diffusion Models for Constrained Generation

Huu Binh Ta ⋅ Michael Cardei ⋅ Alvaro Velasquez ⋅ Ferdinando Fioretto

Discrete diffusion models generate sequences by iteratively denoising samples corrupted by categorical noise, offering an appealing alternative to autoregressive decoding for structured and symbolic generation. However, standard training targets a likelihood-based objective that primarily matches the data distribution and provides no native mechanism for enforcing hard constraints or optimizing non-differentiable properties at inference time. This work addresses this limitation and introduces Search-Augmented Masked Diffusion (SearchDiff), a training-free neurosymbolic inference framework that integrates informed search directly into the reverse denoising process. At each denoising step, the model predictions define a proposal set that is optimized under a user-specified property satisfaction, yielding a modified reverse transition that steers sampling toward probable and feasible solutions. Experiments in biological design and symbolic reasoning illustrate that SearchDiff substantially improves constraint satisfaction and property adherence, while consistently outperforming discrete diffusion and autoregressive baselines.


SeeSE3: The Emergence of 3D Space in Vision Features

Viorica Patraucean ⋅ Leonidas Guibas ⋅ Sayna Ebrahimi ⋅ Caroline Chen ⋅ Ming-Hsuan Yang ⋅ Maks Ovsjanikov ⋅ Fedor Kitashov

In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean space. Unlike previous works that probe 3D awareness of vision features by regressing image-centric quantities such as depth or normals, we investigate the relation between the structure of the space of visual features and the group of Euclidean transformations $SE(3)$. We propose a set of probes to evaluate this relation from both topological and geometric perspectives: a mutual neighborhood metric that measures the alignment between feature neighborhoods and spatial topology, and a Poincaré Adapter to test the linear accessibility of the geometry of camera motion from latent displacements in static scenes. We show that self-supervised vision models, which, in principle, have not been trained with direct 3D supervision or active agency, possess latent subspaces that are remarkably strongly correlated with the structure of Euclidean space when probed correctly. Building on this insight we propose a new class of ``Latent-Space Navigation'' techniques that perform visual odometry and localization purely in the visual latent space, bypassing the need for explicit 3D reconstruction.


Self-Compacting Language Model Agents

Tianjian Li ⋅ Jingyu (Jack) Zhang ⋅ Xi Wang ⋅ William Jurayj ⋅ Chuanyang Jin ⋅ Mehrdad Farajtabar ⋅ Eric Nalisnick ⋅ Daniel Khashabi

Long-form reasoning and tool calling traces accumulate errors and stale content that anchor subsequent generations, a phenomenon known as \emph{context rot}. Existing scaffolds mitigate this with \textit{fixed-interval} compaction triggered at a token threshold, but such triggers are blind to trajectory structure and risk discarding partial results mid-derivation or mid-search. We propose SelfCompact, which pairs two inference-time elements: an inline \emph{compaction tool} the model invokes itself, and a lightweight \emph{rubric} specifying when to fire (a sub-task has resolved, or the trajectory is converging) and when to suppress (mid-derivation, or when stuck). Both are needed. The tool alone is unevenly used across open-weight models, often invoked at unhelpful moments or not at all; the rubric alone cannot act. Together, they elicit effective adaptive compaction without any fine-tuning. On competition math (IMO-Answerbench, HMMT Nov 25 / Feb 26) with four Qwen3 / Qwen3.5 models and agentic search (BrowseComp, BrowseComp-Plus) with three deployed agents, \method{} matches or exceeds fixed-interval summarization at a fraction of the token cost, improving over a no-summarization baseline by up to 16.7 points on math and 5--9 points on agentic search at 30--70\% lower per-question cost. The result reframes \emph{when to compact} as a meta-cognitive capability that scaffolds, not weights, can supply.

Self-improvement at scale has been a longstanding goal for reasoning models, and there are two natural places to do it: at test time, through verification-refinement (V-R) loops; and at training time, through self-training methods. Both are gated by the same bottleneck: the verifier. V-R loops stall when verifier scores inflate over rounds while accuracy stagnates, and when feedback is too generic to act on; self-training fails similarly when bad self-generated data are added to training. Better verifiers would unlock both, but the capability we want to train, i.e., catching errors the generator cannot detect in its own work, has neither finetuning data nor a verifiable reward. To address this challenge, we propose self-trained verification (STV). Our key observation is that, while a model cannot catch these errors alone, it can when shown the reference solution. We turn this asymmetry into a supervision target and train the verifier to imitate a more informed version of itself. At test time, STV substantially improves V-R loops on hard problems, while standard alternatives (e.g., SFT, RL on verifier scores, and even meta-verifiers) do not. STV roughly doubles accuracy on hard DAPO math and lifts it 14x on the hardest SciKnowEval problems (1.5% to 21%). At training time, starting from an RL-converged generator, putting the STV verifier in the loop yields a further 33% relative gain in test-time pass@1. More notably, the generator's standalone pass@1, with no verifier at inference time, climbs 30% relative past where continued RL had converged. Hence, the next frontier in reasoning may lie in how we train verifiers.


SGD at the Edge of Stability: The Stochastic Sharpness Gap

Fangshuo Liao ⋅ Afroditi Kolomvaki ⋅ Anastasios Kyrillidis

Edge of Stability (EoS) refers to the phenomenon where full-batch gradient descent (GD) training of neural networks with step size $\eta$ pushes the largest eigenvalue of the Hessian, i.e. the sharpness, to $2/\eta$ and hovers there. \citet{damian2023selfstab} explains the hovering behavior of the sharpness by *self-stabilization*, a mechanism driven by third-order structure of the loss, and shows that GD implicitly follows projected gradient descent (PGD) on the set where the sharpness is below $2/\eta$. For mini-batch stochastic gradient descent (SGD), the sharpness stabilizes *below* $2/\eta$, with the gap widening as the batch size decreases. However, no theoretical explanation exists for this suppression. In this paper, we introduce *stochastic self-stabilization* that extends the self-stabilization framework to SGD. Our key insight is that gradient noise injects variance into the oscillatory dynamics along the top Hessian eigenvector, strengthening the sharpness-reducing force and shifting the equilibrium below $2/\eta$. Following the approach of \citet{damian2023selfstab}, we define *stochastic predicted dynamics* that tracks the deviation of SGD from the PGD trajectory, and prove a stochastic coupling theorem that relates the SGD sharpness and loss to those of PGD. Based on our predicted dynamic, we derive a closed-form equilibrium sharpness gap that scales with the variance of the gradient noise projected onto the top eigenvector of the Hessian. This formula predicts that smaller batch sizes yield flatter solutions, and recovers GD when the batch equals the full dataset.


Signed Rectified Flow: Negativity Controlled Generation

Runlong Liao ⋅ Baiyu Su ⋅ Lizhang Chen ⋅ Qiang Liu

We introduce \emph{Signed Rectified Flow (Signed RF)}, a generalization of Rectified Flow that targets a signed measure $\pi^{\mathtt{sign}} = (1+\alpha)\,\pi^+ - \alpha \,\pi^-$, where $\alpha>0$, $\pi^+$ represents the distribution to promote, and $\pi^-$ represents the distribution to suppress. Although sampling from a signed measure is not well-defined, Signed RF induces a valid generative process that concentrates on the positive region of $\pi^{\mathtt{sign}}$ while provably excluding regions dominated by the negative component. This yields a principled framework for incorporating negative information and exclusion constraints into generative modeling. Theoretically, we analyze the signed continuity equation underlying Signed RF and explain how negative mass creates exclusion barriers through a charged-particle interpretation. Empirically, Signed RF leads to practical adaptive guidance algorithms. Across applications, Signed RF improves the fidelity--diversity trade-off on ImageNet, reduces nearest-neighbor similarity in anti-memorization stress tests, and mitigates adversarial-prompt nudity in SD 3.5 while preserving CLIP and aesthetic scores.


Simulation-Informed Diffusion for Decentralized Multi-robot Motion Planning

JINHAO LIANG ⋅ Sven Koenig ⋅ Ferdinando Fioretto

Decentralized multi-robot motion planning requires each robot to generate collision-free trajectories from local observations, without global sensing or reliable communication. However, most existing planners, whether classical or learning-based, generate trajectories from a static snapshot of the local observation, which limits their ability to anticipate the future behavior of neighboring robots. This limitation is critical as the number of robots increases and the environment becomes more cluttered. To overcome this challenge, this paper introduces Simulation-Informed Diffusion (SID), a decentralized framework built on constraint-aware diffusion models (CADM). SID first uses CADM to simulate the future trajectories of neighboring robots from their currently observed states, and then uses the same CADM to plan each robot's own trajectory under safety constraints informed by these simulations. Crucially, the accurate simulation of neighbors enables a minimal communication scheme that triggers coordination only when necessary in highly congested scenarios. Experiments across diverse environments show that SID consistently outperforms baseline methods in terms of planning effectiveness and constraint satisfaction, and scales to scenarios with 100 robots and 160 obstacles.


Soteria: Formally Verified Planning with Runtime Enforcement for Safe LLM Agents

Deyuan (Mike) He ⋅ Ankush Desai ⋅ Sharad Malik ⋅ Aarti Gupta

LLM agents that execute multi-step tool calls must satisfy two objectives simultaneously: completing the user's task (utility) and conforming to domain policies and correctness constraints (safety). These objectives are in tension -- blocking an unsafe action prevents a violation but can leave the agent stranded with no principled way to recover. Existing approaches sacrifice one objective for the other. We introduce Soteria, a framework that reconciles safety and utility through verified hierarchical planning and runtime enforcement of specifications. Before execution, the agent generates a structured plan that is formally verified against the specifications prior to any tool invocation. During execution, the verified plan serves dual roles: it ensures that the agent takes specification-conformant trajectories and provides guidance when unsafe actions are blocked. Across multiple benchmarks covering a diverse range of tool-use tasks, Soteria achieves perfect specification conformance while improving utility by up to $5\times$ over existing guardrail approaches, demonstrating that safety and utility need not be traded off.

Sparse attention improves LLM inference efficiency by selecting a subset of key–value (KV) entries, but at the cost of potential accuracy degradation. In particular, omitting critical KV entries can induce substantial errors in model outputs. Existing methods typically operate under fixed or adaptive token budgets and provide empirical robustness or partial theoretical guarantees, yet they do not ensure zero false negatives across decoding steps, particularly since the set of relevant tokens is both query- and step-dependent. Our empirical observations confirm that missing even one critical key can lead to sharp error spikes, especially in long-output reasoning tasks where the set of important tokens varies throughout decoding. This observation motivates the need for indexing methods that dynamically adapt to these variations across decoding steps while guaranteeing a full recall of the relevant keys above a certain threshold. We address this challenge by reformulating sparse attention as the computational geometry problem of halfspace range searching. However, existing range searching data structures are not suitable for modern LLM inference due to their computational and implementation overheads. To overcome this, we introduce Louver, a novel index structure tailored for efficient KV cache retrieval. Louver (i) guarantees zero false negatives with respect to a specified threshold in both theory and practice, (ii) is lightweight to integrate into existing LLM inference pipelines, and (iii) incorporates hardware-aware optimizations for both CPU and GPU executions. Our experiments demonstrate that Louver outperforms prior sparse attention methods in both accuracy and runtime, and is faster than highly optimized dense implementations such as FlashAttention. These results highlight that recall guarantees are a critical and overlooked dimension of sparse attention, and open a new direction for building theoretically grounded, efficient KV cache indices.

Dense fine-tuning can turn a base language model into a domain specialist, but task accuracy alone does not reveal which internal computations mediate the new behavior. We introduce specialist mediators: internal mechanisms through which a dense specialist's lift over its base causally flows. We identify these mediators by patching specialist activations into the base model and asking which patches recover in-domain performance while minimally perturbing off-domain behavior. This separates three questions that are often conflated: where the specialist behavior is localized, how much of it can be recovered by activation patching, and whether a sparse trainable update can implement the same behavior. We formalize these links with exact recovery results, a first-order recovery approximation, and a bound showing when patching recovery can transfer to localized training. In synthetic transformers and structural causal models with planted mediators, patching recovers the true causal mechanisms even when activation-magnitude baselines are misled by non-causal distractors. In Pythia-410M specialists for math (GSM8K), multitask knowledge (MMLU), and biomedical QA (PubMedQA), single-patch localization ranks first on every domain and seed by area under the recovery curve, with small mechanism budgets recovering much of the patchable specialist gap. A localized-LoRA stress test then shows why the separation matters: patching-localized adapters improve negative log-likelihood when the selected attention mechanisms carry the specialist gap, but fail on GSM8K where MLP blocks carry substantial lift. Thus specialist localization is useful both as a route to sparse adaptation and as a diagnostic for when sparse adaptation fails.


Specialists Hold, Generalists Discount: Asymmetric Equilibrium in LLM Routing Auctions

Xinyu Hou ⋅ Yang Lu ⋅ Rabimba Karanjai ⋅ Pei-Chi Pan ⋅ Sen Lin ⋅ Lei Xu ⋅ Weidong Shi

Routing systems for large language models, such as MasRouter, RouteLLM, FrugalGPT, and others, match queries to suppliers based on cost-quality tradeoffs. Most prior work optimizes this problem from the demand side, while the supplier-side question of how LLM suppliers should price their services within a routing mechanism has received little formal treatment. We provide the first systematic analysis of this problem. In our setup, the LLM is the priced commodity: supplier organizations set price functions for the services they offer, rather than acting as strategic agents that generate bids per query. We model supplier-side pricing as a sealed-bid first-price reverse auction over a router that allocates each query to one supplier based on cost and quality. This framework applies to any cost-quality routing system rather than to a specific implementation. Using calibrated profiles from 12 open-weight models and MasRouter as a case-study router, we characterize equilibrium behavior. Our main finding is that the Bayesian Nash equilibrium is asymmetric: capability-differentiated suppliers adopt flat, non-discounted strategies, while marginal-quality suppliers compete primarily on price. A mechanism-baseline experiment, where the case-study router is replaced by an analytical rational-decision rule, confirms that this pattern is a property of the auction mechanism rather than an artifact of router training. We further observe that price differentiation can emerge in equilibrium even under capability symmetry, because private-cost types alone produce nontrivial bid functions. The practical implication is that auction-based pricing alone does not discipline specialist rents. Achieving that goal requires additional mechanism elements, such as reserve prices or capability-blind tie-breaking.

Neural avalanches and turbulence-like brain dynamics are usually measured after neural activity has been coarse-grained into continuous fields, leaving open which spike-level mechanisms can generate these macroscopic observables. We build a spatial leaky integrate-and-fire spiking neural network and explicitly transform spikes into activity-density and Hilbert phase fields. A delayed, distance-dependent, passive-dendrite regime reproducibly generates turbulence-like field structure, including high-amplitude phase-defect tracks, structure-function scaling, and power-law-favored continuous-field avalanche tails. Controls show that firing rate, vortex count, and avalanche tail fits are individually insufficient: destroying delayed propagation, dendritic filtering, or spike timing suppresses key field signatures, while random-space nulls can retain many mathematical phase defects but lose spatial scaling and synaptic-delay transfer. Biologically motivated gates further shape the regime: PV-like perisomatic inhibition clamps phase dynamics, SST-like dendritic inhibition and projection-specific inhibitory facilitation preserve long-lived tracks, and a rate-constrained projection-STP regime maintains similar dynamics at 12.25 Hz across 12 seeds. These results use SNNs as a mechanistic testbed for spike-derived brain turbulence observables and show why phase defects and avalanche-like statistics must be interpreted jointly with spatial scaling and spike-level propagation.


SPIRIT: Speed-Driven Online Adaptation for Self-Speculative Decoding

Chenghao Liu ⋅ Hui Lu ⋅ Chengbo Zhan ⋅ Jia Rao

Self-speculative decoding offers a plug-and-play route to accelerating autoregressive large language model (LLM) inference, but its practical success hinges on identifying effective layer-skipping configurations online without retraining or offline profiling. Existing methods mainly optimize agreement-based proxies that capture draft-target consistency, yet overlook intrinsic runtime differences across draft configurations. Consequently, configurations with similar acceptance can still yield substantially different realized decode throughput. We present SPIRIT, a speed-driven online adaptation framework for self-speculative decoding that directly optimizes realized decode throughput, measured as generated tokens per second, using window-level runtime feedback. Our key insight is that speed is not determined by agreement alone: it is jointly shaped by draft-target consistency and the intrinsic speed advantage of the draft configuration. Empirically, the skip ratio serves as a first-order control variable, capturing the dominant consistency-speed trade-off while sharply shrinking the search space over layer-skipping patterns. SPIRIT therefore first identifies an effective skip ratio and then refines the layer-skipping pattern within it. Combined with length-aware normalization and shift-aware profile-based adaptation, this design enables rapid adaptation to task shifts in non-stationary request streams. Across diverse tasks and model scales, SPIRIT consistently outperforms strong training-free baselines and achieves up to 1.74x speedup while preserving the target model's output distribution, establishing a practical foundation for plug-and-play, high-throughput LLM inference.


Split and Bridge: Multimodal Generation via Diffusion Bridging

Ahmad Arrabi ⋅ Xiaohan Zhang ⋅ Xingyu Li ⋅ Safwan Wshah

Visual data are naturally expressed through multiple complementary modalities (e.g., images and segmentation masks), each capturing distinct aspects of the same underlying structure. Most existing multimodal methods adopt a top-down paradigm, learning a shared latent variable through joint training from scratch, which limits scalability. We propose Split and Bridge (SnB), a bottom-up framework for multimodal generation that couples pretrained unimodal diffusion models at sampling time. Rather than learning a unified latent space, SnB generates coherent multimodal samples by following a shared diffusion trajectory up to a split point, after which modality-specific processes are guided via a bridge module. This design enables plug-and-play reuse of powerful pretrained models without joint training, making it scalable and easily adaptable. Across standard benchmarks (PolyMNIST, CelebAMask-HQ) and more complex settings (PIE-Bench++, ROCO), SnB consistently achieves a strong balance between sample quality and coherence, as we report a 33% increase in the coherence score on CelebAMask-HQ over prior top-down approaches. We show that SnB naturally extends to downstream zero-shot tasks such as interleaved co-generation, image editing, and style transfer. Our results demonstrate that bottom-up generation offers a practical and scalable alternative to top-down multimodal models.


Split-RL: Local Conflict Resolution in Reinforcement Learning

Benjamin Fuhrer ⋅ Chen Tessler ⋅ Gal Dalal

We introduce Split-RL, a reinforcement learning (RL) algorithm based on decision trees. It uses guidance labels to explicitly partition the state-space and isolate localized conflicting signals. RL policies are often trained by mixing contradictory feedback that is localized in space and time. For example, a robot may prioritize speed in open areas versus precision in tight gaps. These regions are often easily detectable from sensor data or a human operator. Standard neural networks (NNs) lack the structural mechanism to isolate these signals. While Gradient Boosting Trees (GBT) provides the inductive bias to explicitly partition the state-space, standard tree-fitting procedures fit only aggregated gradients, merging contradictory updates before they reach the leaves. Split-RL leverages this bias to integrate guidance labels into tree construction, routing these updates into disjoint regions of the state-space to prevent signal interference during training. Such labels can arise from simple heuristics, sensor-derived events, or existing rule-based systems, allowing Split-RL to incorporate rule-derived structure while retaining a learned policy. We theoretically show how early gradient aggregation loses objective-specific information and characterize when the Split-RL score isolates conflicting updates. Finally, we evaluate Split-RL on constrained tasks and offline imitation learning from mixed-quality datasets. Across constrained domains with spatial, rare-event, and temporal conflicts, Split-RL consistently outperforms existing methods by achieving low-cost, feasible solutions while maintaining competitive rewards. In offline settings, Split-RL localizes the tradeoff between imitation and policy improvement, outperforming other methods and achieving near-expert performance across varying dataset sizes.


Stable GFlowNets with Probabilistic Guarantees

Zengxiang Lei ⋅ Ananth Shreekumar ⋅ Jonathan Rosenthal ⋅ Ruoyu Song ⋅ Alvaro A Cardenas ⋅ Daniel Fremont ⋅ Dongyan Xu ⋅ Satish Ukkusuri ⋅ Z. Berkay Celik

Generative Flow Networks (GFlowNets) learn to sample states proportional to an unnormalized reward. Despite their theoretical promise, practical training is often unstable, exhibiting severe loss spikes and mode collapse. To address this, we first assess the sensitivity of GFlowNet objectives, demonstrating that a small Total Variation (TV) distance between the learned and target distributions does not preclude an unbounded training loss. Motivated by this mismatch, we establish converse guarantees by deriving loss-to-TV bounds that certify global fidelity from bounded trajectory balance losses. Lastly, we propose \emph{Stable GFlowNets}, which leverages our theory to stabilize training via adaptive reference flow and improves the trade-off among mode coverage, robustness, and certifiability.

Statistical matching combines partially overlapping datasets that share covariates $X$ but observe the target $Y$ and auxiliary variables $Z$ separately. Classical approaches typically invoke the conditional independence assumption (CIA), which makes the problem identifiable but fundamentally implies that the imported auxiliary variable provides no additional predictive power for $Y$ once $X$ is known. To capture this latent $Y$--$Z$ dependence, we propose a novel dependency-aware Schr\"odinger bridge for predictive statistical matching. Our approach couples the two separated databases by tilting the conservative CIA baseline with a transportation-based compatibility cost, recovering an informative joint distribution. The resulting statistical learning framework yields full probabilistic posterior rules for bidirectional imputation. Theoretically, we establish a sufficient condition under which the learned bridge strictly improves over the CIA baseline, alongside an exact joint recovery guarantee in the Gaussian setting under an appropriate cost. Across synthetic benchmarks and real-world datasets (CelebA and Adult), we demonstrate that our dependency-aware completion consistently improves downstream predictive utility, proving especially beneficial in settings like data recoding where the underlying population exhibits strong $Y$--$Z$ dependence.

Model-based learning agents use learned world models to predict future states, plan actions, and adapt to new environments. However, the process of updating world models from collected experience creates a training-time attack surface: adversarially poisoned fine-tuning trajectories can manipulate the learned dynamics and thereby corrupt downstream planning. In this paper, we propose SWAAP, the first two-stage data poisoning framework for learned world models. In the first stage, SWAAP identifies a harmful target world model that induces low-return behavior under planning while remaining close to clean dynamics, using first-order bilevel optimization enabled by a transition-gradient theorem. In the second stage, SWAAP realizes this target through stealth-constrained gradient matching, modifying only a limited fraction of fine-tuning transition targets so that the induced training gradients steer the victim model toward the adversarial target, while a prediction-error regularizer encourages the poisoned targets to remain close to the world model's natural approximation error. To assess attack stealthiness, we evaluate defenses and detectability across three stages of the poisoning pipeline: pre-training detection of poisoned transitions, robust training during fine-tuning, and test-time monitoring of the resulting world model. Across diverse continuous-control tasks, SWAAP causes substantial performance degradation while keeping poisoned transitions close to clean data and avoiding detection by the evaluated defenses. These results reveal a practical vulnerability in world-model adaptation pipelines and highlight the need for robustness methods that protect both world-model training data and learned dynamics.


Stein Kernelized Molecular Dynamics for Active Learning of Interatomic Potentials

Joanna Zou ⋅ Fraser Birks ⋅ Dallas Foster ⋅ Youssef Marzouk

Machine learning interatomic potentials (MLIPs) enable efficient and accurate atomistic simulations but depend critically on the quality and diversity of the training data. We introduce Stein kernelized molecular dynamics (SKMD), an enhanced sampling method that uses interacting particle dynamics to acquire informative training configurations for the active learning and fine-tuning of MLIPs. SKMD corresponds to a stochastic variant of Stein variational gradient descent that is adapted for molecular dynamics by incorporating asynchronous particle updates and a kernel of global atomic descriptors, which provides a symmetry-aware measure of configurational similarity. Unlike other enhanced samplers used in molecular dynamics, SKMD preserves the Boltzmann distribution as the asymptotic distribution of the dynamics. This property enforces a balance between the exploration of diverse configurations and attraction toward high-probability regions of the energy landscape. We further propose an approach to efficient online data acquisition using an adaptive stopping criterion that selects non-redundant training data over the course of simulation. We demonstrate SKMD for the active learning of a neural network model of the Muller-Brown potential and the fine-tuning of a MACE interatomic potential for alanine dipeptide. Compared to active learning baselines, our method achieves higher model accuracy in fewer training iterations, with the same number of acquired training samples.

Reinforcement learning (RL) is a central label in large language model (LLM) post-training, reasoning, and agent research, but reward gains alone do not identify the decision process a result improves. The same "RL" label can describe fixed-data preference fitting, learned-reward optimization, verifiable-reward RL, inference-time search, tool-mediated control, or deployed policy adaptation. These regimes support different conclusions. We argue that papers applying RL to LLMs should include a Decision-Process Card (DPC): a tiered, claim-level report of the state or belief, action granularity, reward and cost source, support assumptions, value or uncertainty object, model interface, and adaptation surface behind the claim. The card is a claim-ceiling device, not a demand that every paper solve every RL problem. Completion-only preference papers need different evidence than RLVR papers, tool-using agents, or deployed adaptive agents. The evidence does not support the claim that language-model RL is shallow; it shows real progress with weakly standardized claim boundaries outside narrow, resettable, verifier-friendly regimes. With DPC fields visible, reviewers can ask what was optimized, under what data support, with which reward or verifier validity, under which constraints, and across which adaptation surface.


Structural Rationale Distillation via Reasoning Space Compression

Jialin Yang ⋅ Jiankun Wang ⋅ Jiajun Wu ⋅ Henry Leung ⋅ Jiayu Zhou ⋅ Steve Drew

When distilling reasoning from large language models (LLMs) into smaller ones, teacher rationales for similar problems often vary wildly in structure and strategy. Like a chef who makes the same dish differently each time, this inconsistency burdens the student with noisy supervision that is hard to internalize. We propose Distillation through Reasoning Path Compression (D-RPC), which constrains the teacher to follow a compact, dynamically maintained bank of reusable high-level reasoning paths. For each training question, D-RPC retrieves the most relevant path and conditions the teacher to follow it, producing rationales that are consistent across similar problems yet diverse enough to cover different problem types. A PAC-Bayes analysis formalizes the resulting trade-off between bank size and coverage: smaller banks reduce supervision entropy but risk coverage gaps, and the generalization bound identifies an optimal intermediate size confirmed by our ablations. Across five math and commonsense reasoning benchmarks with two student models, D-RPC consistently outperforms chain-of-thought distillation, freeform rationale generation, direct distillation, and structured-supervision baselines, while using fewer tokens than template-heavy alternatives.


Structured Masked Diffusion for Joint Multiuser Decoding

Taekyun Lee ⋅ Jiyoung Yun ⋅ Jeffrey Andrews ⋅ Hyeji Kim

In joint multiuser decoding, a receiver recovers a set of messages from a single noisy aggregate of many simultaneous transmissions. Classical decoders rely on rule-based mechanisms such as successive interference cancellation, joint belief propagation, or list recovery, all of which become brittle or expensive as ambiguity increases. We propose CIDER, a learned multiuser decoder with masked-diffusion refinement steps. CIDER uses demixing to prevent duplicate-row collapse and uses parity-aware propagation to provide soft guidance from the code constraints. In higher-load regimes, we further improve reliability via a lightweight quality-guided remasking step that selectively re-decodes low-confidence sequences. On commonly used error correcting codes, CIDER matches or improves on FFT-accelerated joint belief propagation-style decoding in symbol error rate while running more than $6\times$ to over $100\times$ faster, with the speedup widening as the blocklength grows. Code is available at https://anonymous.4open.science/r/CIDER_2026/.


Sublinear Time Quantum Sensitivity Sampling

Zhao Song ⋅ David Woodruff ⋅ Lichen Zhang

We present a unified framework for quantum sensitivity sampling, extending the advantages of quantum computing to a broad class of classical approximation problems. Our unified framework provides a streamlined approach for constructing coresets and offers significant runtime improvements in applications such as clustering, regression, and low-rank approximation. Our contributions include: * **$k$-median and $k$-means clustering:** For $n$ points in $d$-dimensional Euclidean space, we give an algorithm that constructs an $\epsilon$-coreset in time $\widetilde O(n^{0.5}dk^{2.5}~\mathrm{poly}(\epsilon^{-1}))$ for $k$-median and $k$-means clustering. Our approach achieves a better dependence on $d$ and constructs smaller coresets when $d\gg k$ that only consist of points in the dataset, compared to recent results of [Xue, Chen, Li and Jiang, ICML'23]. * **$\ell_p$ regression:** For $\ell_p$ regression problems $\min_{x\in \mathbb{R}^d}\\|Ax-b\\|_p$ where $A\in \mathbb{R}^{n\times d}$ and $b\in \mathbb{R}^n$, we construct an $\epsilon$-coreset of size $\widetilde O_p(d^{\max\\{1, p/2\\}}\epsilon^{-2})$ in time $\widetilde O_p(n^{0.5}d^{\max\\{0.5, p/4\\}+1}(\epsilon^{-3}+d^{0.5}))$, improving upon the prior best quantum sampling approach of [Apers and Gribling, QIP'24] for $p\in (2, 22]$. * **Low-rank approximation with Frobenius norm error:** We introduce the first quantum sublinear-time algorithm for low-rank approximation that approximates the best rank-$k$ solution to a matrix $A\in \mathbb{R}^{d\times n}$ and does not rely on data-dependent parameters, and runs in $\widetilde O(n^{0.5}dk^{0.5}\epsilon^{-1})$ time. Additionally, we present quantum sublinear algorithms for kernel low-rank approximation and tensor low-rank approximation, broadening the range of achievable sublinear time algorithms in randomized numerical linear algebra.


SWE-Protégé: Learning to Selectively Collaborate With an Expert Unlocks Small Language Models as Software Engineering Agents

Patrick Tser Jern Kon ⋅ Archana Pradeep ⋅ Ang Chen ⋅ Alex Ellis ⋅ Warren Hunt ⋅ Zijian Wang ⋅ John Yang ⋅ Samuel Thompson

Small language models (SLMs) offer compelling advantages in cost, latency, and adaptability, but have so far lagged behind larger models on long-horizon software engineering tasks such as SWE-bench, where they suffer from pervasive action looping and low resolution rates. We introduce SWE-Protégé, a post-training framework that reframes software repair as an expert–protégé collaboration problem. In SWE-Protégé, an SLM remains the sole decision-maker while learning to selectively seek guidance from a strong expert model, recognize stalled states, and follow through on expert feedback. Our approach combines supervised fine-tuning on expert-augmented trajectories with agentic reinforcement learning that explicitly discourages degenerative looping and unproductive expert collaboration. We lightly post-train Qwen2.5-Coder-7B-Instruct to achieve 42.4% Pass@1 on SWE-bench Verified, a +25.4% improvement over the prior SLM state of the art, while using expert assistance sparsely (≈ 4 calls per task and 11% of total tokens).

Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources. However, these sources can include unreliable or adversarial content, and the robustness of agents to adversarial ranking remains poorly understood. Existing benchmarks evaluate functional navigation or static factuality but cannot causally isolate this vulnerability and often confound retrieval-time reasoning with memorized knowledge. We introduce Synthetic Web Benchmark, a controlled environment of procedurally generated web ecosystems designed to evaluate retrieval-time reasoning and source criticism. The benchmark comprises thousands of hyperlinked articles with ground-truth labels, process-level interaction traces, and contamination filtering to ensure that answers cannot be recovered from pretraining alone. By injecting a single high-plausibility misinformation article at a specified rank, we measure the causal effect of adversarial exposure under minimal intervention. Across six frontier models and thousands of evaluation instances, we observe catastrophic failures: accuracy collapses despite unrestricted access to truthful evidence, accompanied by limited search escalation, weak cross-source synthesis, and severe miscalibration. These results show that current agents struggle to arbitrate conflicting sources even when sufficient evidence is available, revealing fundamental limitations in retrieval-based reasoning. The benchmark provides a reproducible testbed for studying epistemic robustness and developing more reliable web agents.

Safe preference alignment of large language models must handle rare harmful tails whose burden is concentrated on demographic, topical, linguistic, or adversarial subgroups. Existing coverage analyses operate on global density ratios that can look adequate on average while the harmful tail of a small group remains unsupported, leaving worst-group safety risk unidentified. We assume that the absolute harm scale is externally calibrated by severity labels or anchors and introduce TailGuard, a framework centered on subgroup quantile coverage $C_{g,\tau}$: the worst deployment-to-training density ratio restricted to group g's upper-τ harm tail. We prove that $C_{g,\tau}$ is a key statistical quantity governing subgroup tail safety: it governs tail-risk transfer, yields an offline impossibility result under missing tail support, and is necessary for fixed-policy subgroup tail-risk estimation. From a per-group CVaR_τ-constrained alignment objective, we derive a KL-regularized Gibbs form, motivate a Tail-Fisher active query rule that targets group-tail coverage deficits, and construct a per-group weighted CVaR risk-control deployment gate. We validate the mechanism through four completed data-level studies: synthetic coverage control, a semi-synthetic tail-hiding pilot on real safety-alignment labels, an observed-label active-query ablation, and an oracle-score risk-control-gate efficiency proxy. These experiments test the coverage mechanism and deployment diagnostic.


Teaching LLMs to Recommend and Defer in Underrepresented Epilepsy Care

Shreyas Rajesh ⋅ Kartik Sharma ⋅ Tonmoy Monsoor ⋅ Mehmet Yigit Turali ⋅ Richard Idro ⋅ Juliana Kayaga ⋅ Robert Sebunya ⋅ Tracy T Namata ⋅ Jessica N Pasqua ⋅ vwani Roychowdhury ⋅ Rajarshi Mazumder

Specialist epilepsy expertise is scarce in resource-constrained settings, making LLM-based decision support attractive for frontline clinicians managing longitudinal treatment. Such support systems must do more than apply medical knowledge: they must adapt to local prescribing practice and know when to defer. Public medical AI benchmarks are dominated by high-income clinical settings, leaving prescribing practices, medication availability, and follow-up patterns in low-resource contexts largely unrepresented. We study this problem through a multidisciplinary collaboration in Ugandan pediatric epilepsy care. The task is to predict anti-seizure medication regimens from longitudinal unstructured notes collected by local clinicians across serial visits. Standard prompting achieves non-trivial agreement with physician prescriptions, but neurologists' review of model reasoning traces shows that its errors stem from distribution-miscalibrated prescribing defaults rather than the local care environment. We introduce Manana, a non-parametric prompt-learning framework that learns how to reason about local prescribing decisions from a small patient-level training set. Manana turns observed prescription errors into an auditable prompt memory, instantiated in single-agent and multi-agent variants, and outperforms classical ML models, direct LLM prompting, and prompt-optimization baselines across two independently collected Ugandan cohorts. To make the system uncertainty-aware, we propose Bayesian prompt averaging (BPA), a Bayesian model averaging procedure over the learned prompt trajectory. This converts a sequence of learned prompts into prescription likelihoods and produces a deferral signal. On the independently collected held-out cohort, BPA improves visit-level top-3 prescription accuracy by 4-8 percentage points over the prompt-optimization baselines. More consequentially, it enables clinically meaningful selective prediction: the system can auto-handle the most confident half of cases at 95\% precision, or the most confident quarter at 99\% precision, while deferring lower-confidence cases for specialist review. These results suggest a path toward locally adapted clinical LLM systems that learn from limited site-specific data and reserve scarce specialist attention for the cases where uncertainty is highest.


Test-Time Defense Against Adversarial Attacks via Stochastic Resonance of Latent Ensembles

DONG LAO ⋅ Yuxiang Zhang ⋅ Haniyeh E Oskouie ⋅ Yangchao Wu ⋅ Alex Wong ⋅ Stefano Soatto

We propose a test-time defense mechanism against adversarial attacks: Unlike existing methods that rely on feature filtering or smoothing, which can lead to information loss, we propose to ``combat noise with noise'' by leveraging stochastic resonance to enhance robustness while minimizing information loss. Our approach introduces small translational perturbations to the input image, aligns the transformed feature embeddings, and aggregates them before mapping back to the original reference image. This can be expressed in a closed-form formula, which can be deployed on diverse existing network architectures without introducing additional network modules or fine-tuning for specific attack types. The resulting method is entirely training-free, architecture-agnostic, and attack-agnostic. Empirically, the method achieves state-of-the-art robustness on image classification and provides the first generic test-time defense for dense prediction tasks, including stereo matching, optical flow, and monocular depth estimation. Across adversarial attacks, it recovers up to 68.1% of the accuracy loss on image classification, 71.9% on stereo matching (MAE), 29.2% on optical flow (EPE), 68.3% on depth estimation (SqRel), and 83.8% on vision-language alignment (cosine similarity).


TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models

Yao Xiao ⋅ Qiqian Fu ⋅ Heyi Tao ⋅ Yuqun Wu ⋅ Zhen Zhu ⋅ Derek Hoiem

Image-text models excel at image-level tasks but struggle with detailed visual understanding. While these models provide strong visual-language alignment, segmentation models like SAM2 offer precise spatial boundaries for objects. To this end, we propose TextRegion, a simple, effective, and training-free framework that combines the strengths of image-text models and SAM2 to generate powerful text-aligned region tokens. These tokens enable detailed visual understanding while preserving open-vocabulary capabilities. They can be directly applied to various downstream tasks, including open-world semantic segmentation, referring expression comprehension, and grounding. We conduct extensive evaluations and consistently achieve superior or competitive performance compared to state-of-the-art training-free methods. Additionally, our framework is compatible with many image-text models, making it highly practical and easily extensible as stronger models emerge. Code is available at: https://github.com/avaxiao/TextRegion.

Learning control policies for complex, long-horizon tasks is a central challenge in autonomous systems. Signal Temporal Logic (STL) offers an expressive language for specifying such tasks, but its non-Markovian nature and inherent sparse reward make it hard for standard Reinforcement Learning (RL) to solve. Prior RL approaches focus only on limited STL fragments or use robustness scores as sparse rewards. Our new method TGPO (Temporal Grounded Policy Optimization) decomposes STL into timed subgoals and invariant constraints and tackles the problem in a hierarchical fashion. Its high-level module proposes time allocations for subgoals, and the low-level time-conditioned policy learns to achieve the sequenced subgoals using a dense, stage-wise reward. In inference, we use the critic to efficiently sample time allocations and select the most promising assignment for the policy to rollout. We evaluate in five environments, from low-dimensional navigation to manipulation, drone, and quadrupedal locomotion, and TGPO significantly outperforms baselines (especially for high-dimensional and long-horizon cases), with 31.6% higher success rate compared to the best method.

Accurate numerical solutions of partial differential equations (PDEs) are crucial in numerous science and engineering applications. In this work, we introduce a novel neural PDE solver named AFDONet, which incorporates neural operator learning and adaptive Fourier decomposition (AFD) theory for the first time into a specifically designed variational autoencoder (VAE) structure, to solve a general class of nonlinear PDEs on smooth manifolds. AFDONet is the first neural PDE solver whose architectural and component design is fully guided by an established mathematical framework (in this case, AFD theory), turning neural operator design from an art to a science. Thus, AFDONet also exhibits exceptional mathematical explainability and groundness, and enjoys several desired properties. Furthermore, AFDONet achieves outstanding solution accuracy and competitive computational efficiency in several benchmark problems. In particular, thanks to its deep connections with AFD theory, AFDONet shows superior performance in solving PDEs on i) arbitrary (Riemannian) manifolds, and ii) datasets with sharp gradients. Overall, this work presents a new paradigm for designing explainable neural operator frameworks.


TherapyGym: Evaluating and Aligning Clinical Fidelity and Safety in Therapy Chatbots

Fangrui Huang ⋅ Souhad Chbeir ⋅ Arpandeep Khatua ⋅ Sheng Wang ⋅ Sijun Tan ⋅ Kenan Ye ⋅ Lillian A Bailey ⋅ Merryn Daniel ⋅ Ryan Louie ⋅ Sanmi Koyejo ⋅ Ehsan Adeli

Large language models (LLMs) are increasingly used for mental-health support; yet prevailing evaluation methods--fluency metrics, preference tests, and generic dialogue benchmarks--fail to capture the clinically critical dimensions of psychotherapy. We introduce THERAPYGYM, a framework that evaluates and improves therapy chatbots along two clinical pillars: fidelity and safety. Fidelity is measured using the Cognitive Therapy Rating Scale (CTRS), implemented as an automated pipeline that scores adherence to CBT techniques over multi-turn sessions. Safety is assessed using a multi-label annotation scheme, covering therapy-specific risks (e.g., failing to address harm or abuse). To mitigate bias and unreliability in LLM-based judges, we further release THERAPYJUDGEBENCH, a validation set of 116 dialogues with 1,270 expert ratings for auditing and calibration against licensed clinicians. THERAPYGYM also serves as a training harness: CTRS and safety-based rewards drive RL with configurable patient simulations spanning diverse symptom profiles. Models trained in THERAPYGYM improve on expert ratings, with average CTRS rising from 0.10 to 0.60 (and 0.16 to 0.59 under LLM judges). More broadly, THERAPYGYM provides foundational targets for responsible psychotherapy-oriented LLMs: clinical fidelity, therapy-specific safety, expert-calibrated evaluation, and multi-turn therapeutic process.

Integration against a probability distribution given its unnormalized density is a central task in Bayesian inference and other fields. We introduce new methods for approximating such expectations with a small set of weighted samples---i.e., a quadrature rule---constructed via an interacting particle system that minimizes maximum mean discrepancy (MMD) to the target distribution. These methods extend the classical mean shift algorithm, as well as recent algorithms for optimal quantization of empirical distributions, to the case of continuous distributions. Crucially, our approach creates dynamics for MMD minimization that are invariant to the unknown normalizing constant; they also admit both gradient-free and gradient-informed implementations. The resulting mean shift interacting particle systems converge quickly, capture anisotropy and multi-modality, avoid mode collapse, and scale to high dimensions. We demonstrate their performance on a wide range of benchmark sampling problems, including multi-modal mixtures, Bayesian hierarchical models, PDE-constrained inverse problems, and beyond.


Tokenizer Choice Shapes Generalization in State-Centric Learning for Planning

Vishal Pallagani ⋅ Nitin Gupta ⋅ John A Aydin ⋅ Biplav Srivastava

Planning has traditionally been addressed with symbolic search guided by domain models and heuristics. Learning offers a complementary way to exploit structure shared across related problems rather than solving each instance from scratch, making generalization to unseen instances a central challenge for learned planners. Recent work has explored planning mainly through action-centric models that predict actions or full plans from problem descriptions. A parallel line of work instead learns goal-conditioned transition models over states (state-centric learning) and recovers actions by matching predicted successor states to symbolic successors. In this setting, however, prior work has relied mainly on Weisfeiler-Leman (WL) representations, leaving the role of tokenizer choice unclear. We present a controlled study of tokenizer choice in state-centric learned planning, comparing WL, shortest-path, GraphBPE, SimHash, and a deterministic baseline within a common pipeline. Across varied benchmark domains and multiple planner configurations, WL performs well overall, but no tokenizer is best everywhere: the leading representation changes by domain, and tokenizer differences are larger on extrapolation rather than on interpolation tasks. Across the studied configurations, tokenizer choice is the dominant factor in held-out generalization: it explains 57% of performance variance, more than predictor architecture or any other pipeline factor, and its effect is substantially amplified on out-of-distribution problems relative to in-distribution ones. We conclude that representation choice is a strong determinant of generalization in state-centric learned planning and that tokenizer choice is a first-order modeling decision rather than a fixed preprocessing step. Code: https://tinyurl.com/4hre257u


Toward Cognitive Supersensing in Multimodal Large Language Models

Boyi Li ⋅ Yifan Shen ⋅ Yuanzhe Liu ⋅ Yifan Xu ⋅ Jiateng Liu ⋅ Xinzhuo Li ⋅ Zhengyuan Li ⋅ Jingyuan Zhu ⋅ Yunhan Zhong ⋅ Lan Fangzhou ⋅ Jianguo Cao ⋅ James Rehg ⋅ Heng Ji ⋅ Ismini Lourentzou ⋅ Xu Cao

Multimodal Large Language Models (MLLMs) have achieved remarkable success in open-vocabulary perceptual tasks, yet their ability to solve complex cognitive problems remains limited, especially when visual details are abstract and require visual memory. Current approaches primarily scale Chain-of-Thought (CoT) reasoning in the text space, even when language alone is insufficient for clear and structured reasoning, and largely neglect visual reasoning mechanisms analogous to the human visuospatial sketchpad and visual imagery. To mitigate this deficiency, we introduce Cognitive Supersensing, a novel training paradigm that endows MLLMs with human-like visual imagery capabilities by integrating a Latent Visual Imagery Prediction (LVIP) head that jointly learns sequences of visual cognitive latent embeddings and aligns them with the answer, thereby forming vision-based internal reasoning chains. We further introduce a reinforcement learning stage that optimizes text reasoning paths based on this grounded visual latent. To evaluate the cognitive capabilities of MLLMs, we present CogSense-Bench, a comprehensive visual question answering (VQA) benchmark assessing five cognitive dimensions. Extensive experiments demonstrate that MLLMs trained with Cognitive Supersensing significantly outperform state-of-the-art baselines on CogSense-Bench and exhibit superior generalization on out-of-domain mathematics and science VQA benchmarks, suggesting that internal visual imagery is potentially key to bridging the gap between perceptual recognition and cognitive understanding. We will open-source the CogSense-Bench and our model weights.


Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows

Shikun Liu ⋅ Mufei Li ⋅ Dongqi Fu ⋅ Haoyu Wang ⋅ Yinglong Xia ⋅ Hong Li ⋅ Hong Yan ⋅ Pan Li

Large language models increasingly serve as execution engines for agentic systems, yet they still consume context through a sequential text interface. This creates a mismatch with modern structured agent workflows, in which independent branches explore subtasks, retrieve evidence, or generate candidate solutions before a final synthesis step. Existing systems typically merge these branches by concatenating their textual outputs, which discards the parallel structure and incurs redundant prefill computation. In this work, we introduce Parallel-Synthesis, a plug-and-play framework that enables a synthesizer to directly consume the KV caches produced by parallel worker agents. Parallel-Synthesis combines a cache mapper that calibrates independently generated branch caches with a fine-tuned synthesizer adapter that enables generation from this non-sequential cache interface. We train Parallel-Synthesis using data that exposes the synthesizer to parallel cache contexts, teaches aggregation across cached branches, and distills reasoning behavior from standard text-concatenation-based synthesis. Across nine downstream datasets spanning math, science QA, code generation, GAIA, and multi-agent database diagnosis, Parallel-Synthesis matches or outperforms text-based synthesis on seven datasets and remains close on the other two. It also reduces time-to-first-token by 2.5×--11×, suggesting that direct cache-based synthesis is a promising interface for more native and efficient synthesis over parallel agent branches.


Towards Fair Graph Generation Without Sensitive Attribute

Zichong Wang ⋅ Zhipeng Yin ⋅ Wenbin Zhang

Graph generation, which aims to produce new graphs from a distribution similar to observed data, has gained increasing attention, especially as generated graphs are used in high-stakes decision-making where fairness is critical. However, most existing fair graph generation methods assume full access to sensitive attributes, an assumption often violated in practice due to privacy concerns, regulatory constraints, and missing data. To this end, we propose to solve the problem from a new perspective, where sensitive proxy inference is reformulated as a component of the graph generation pipeline, enabling the inferred proxy to guide the generation process. We further model and mitigate the impact of proxy inference errors, and provide theoretical guarantees that quantify how such errors propagate to fairness outcomes in the generated graphs, offering practical guidance for interpreting fairness when sensitive attribute is missing. Experiments on benchmark datasets demonstrate that our method consistently improves fairness while maintaining competitive generation quality.


TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

Jiarui Yan ⋅ Weiwei Sun ⋅ Sijie Li ⋅ Wenhan Li ⋅ Yiming Yang

While LLMs excel at isolated coding tasks, their ability to perform autonomous, long-horizon machine learning (ML) development remains bottlenecked by poor strategic planning. Because current benchmarks rely on outcome-based metrics, they obscure process-level failures, making it impossible to diagnose where agent strategies diverge from human expertise. To solve this, we introduce TraceML, the first large-scale dataset of paired human and agent trajectories for long-horizon ML tasks. TraceML contains 4{,}672 Kaggle trajectories across 134 MLE-bench competitions, with 430 expert humans paired head-to-head with 207 LLM-agent runs on 7 of those competitions. Using a unified extraction schema to analyze step-by-step behavior, we reveal systematic gaps: agents execute local edits well but exhibit unstructured exploration, poor experiment prioritization, and weak long-term memory attribution compared to experts. Translating these diagnostic insights into action, we design a novel, planning-oriented agent harness based on human priors that significantly improves both process efficiency and final competitive performance.

Understanding how information propagates within vision-language models (VLMs) and large language models (LLMs) is central to interpretability and efficient deployment. Attention weights indicate where probability mass is allocated, but do not directly measure directional influence, redundancy, or functional importance. We introduce transfer entropy (TE) as an information-theoretic measure of directed information flow in Transformers, and develop tractable TE estimators for high-dimensional hidden states at the granularity of layers, tokens, and attention heads. We first analyze multimodal models, beginning with LLaVA-1.5-7B and then CLIP ViT-B/16. In LLaVA-1.5-7B, vision$\rightarrow$text TE increases with depth, suggesting that multimodal fusion is localized in the upper decoder layers. In CLIP, TE reveals late-layer redundancy in the vision tower and broader elevated TE in the text tower. We then analyze unimodal models, including RoBERTa, T5, Llama-3.2-3B, and Qwen-2.5-7B. Across these LLMs, TE reveals consistent depth-dependent structure: RoBERTa exhibits a mid-layer peak, decoder-only LLMs concentrate task-dependent computation in broad mid-depth bands, and T5 shows distinct encoder and decoder dynamics. Beyond layerwise analysis, TE also exposes token- and head-level communication roles, distinguishing sinks, broadcasters, and benign components. These signals support TE-guided layer, token, and attention-head pruning across VLMs and LLMs, consistently outperforming attention- or saliency-based heuristics under matched budgets. Overall, TE provides a unified directed-information lens for diagnosing, interpreting, and compressing Transformer-based models.


Transfer Learning Through Conditional Quantile Matching

Yikun Zhang ⋅ Steven Wilkins-Reeves ⋅ Wesley Lee ⋅ Aude Hofleitner

We introduce a transfer learning framework for regression that leverages heterogeneous source domains to improve predictive performance in a data-scarce target domain. Our approach learns a conditional generative model separately for each source domain and calibrates the generated responses to the target domain via conditional quantile matching. This distributional alignment step corrects general discrepancies between source and target domains without imposing restrictive assumptions such as covariate or label shift. The resulting framework provides a principled and flexible approach to high-quality data augmentation for downstream learning tasks in the target domain. From a theoretical perspective, we show that an empirical risk minimizer (ERM) trained on the augmented dataset achieves a tighter excess risk bound than the target-only ERM under mild conditions. In particular, we establish new convergence rates for the quantile matching estimator that governs the transfer bias-variance tradeoff. From a practical perspective, extensive simulations and real data applications demonstrate that the proposed method consistently improves prediction accuracy over target-only learning and competing transfer learning methods.

Large Language Models (LLMs) have achieved remarkable progress on complex reasoning tasks, yet the mechanisms by which they acquire reasoning capabilities and the efficiency of their reasoning remain poorly understood. In this paper, we investigate the path-finding problem over directed graphs—a symbolic abstraction of multi-step reasoning. We first characterize the architectural requirements for this task, proving that one-layer transformers require $\Omega(N^2)$ size on dense graphs to implement the key DFS child-selection primitive. While a two-layer transformer implements full DFS with $O(N\log(N))$ size. Thus, a second layer is necessary and sufficient for near-linear DFS reasoning. We then extend this expressivity analysis to deeper models, proving the existence of an $L$-layer transformer that performs DFS with a look-ahead horizon of $2^{L-3}$, establishing that increasing depth yields an exponential gain in search efficiency. Finally, going beyond the expressivity, given appropriately curated training data, we show that two-layer transformers can provably learn to execute DFS via gradient flow and generalize to unseen graphs. Together, our results provide a mechanistic account of how transformers implement graph search, how depth improves reasoning efficiency, and how such reasoning algorithms can emerge through training.


TransmissiveGS: Residual-Guided Disentangled Gaussian Splatting for Transmissive Scene Reconstruction and Rendering

Zhenyu Liang ⋅ Xiao Zhang ⋅ Tianchao Li ⋅ Jack C Cheng ⋅ Chi-Keung Tang

Transmissive scenes are ubiquitous in daily life, yet reconstructing and rendering them remains highly challenging due to the inherent entanglement between near-field reflections from the surrounding environment on the transmissive surface, and the transmitted content of the scene behind it. This coupling gives rise to dual surface geometries and dual radiance components within each observation, posing ambiguities for standard methods. We present TransmissiveGS, a novel framework for disentangled reconstruction and rendering of transmissive scenes. Specifically, we model the scene with a dual-Gaussian representation and introduce a deferred shading function to jointly render the two Gaussian components. To separate reflection and transmission, we exploit the inherent multi-view inconsistency of reflections and leverage the residuals from reconstructing multi-view consistent content as cues for disentangled geometry and appearance modeling. We further propose a reflection light field that enables high-fidelity estimation of near-field reflections. During training, we introduce a high-frequency regularization to preserve fine details. We also contribute a new synthetic dataset for evaluating transmissive surface reconstruction. Experiments on both synthetic and real-world scenes demonstrate that TransmissiveGS consistently outperforms prior Gaussian Splatting-based methods in both reconstruction and rendering quality for transmissive scenes.


Tune-Up Open-Weight CLIP: Optimization Framework for Self-Supervised Fine-tuning of CLIP

Anant Mehta ⋅ Xiyuan Wei ⋅ Xingyu Chen ⋅ Tianbao Yang

CLIP has become a cornerstone of multimodal representation learning, yet improving its performance typically requires a prohibitively costly process of training from scratch on billions of samples. We ask a different question: Can we improve the performance of open-weight CLIP models across various downstream tasks using only existing self-supervised datasets? Unlike supervised fine-tuning, which adapts a pretrained model to a single downstream task, our setting seeks to improve general performance across various tasks. However, as both our experiments and prior studies reveal, simply applying standard training protocols starting from an open-weight CLIP model often fails, leading to performance degradation. In this paper, we introduce TuneCLIP, a self-supervised fine-tuning framework that overcomes the performance degradation. TuneCLIP has two key components: (1) a warm-up stage of recovering optimization statistics to reduce cold-start bias, inspired by theoretical analysis, and (2) a fine-tuning stage of optimizing a new contrastive loss to mitigate the penalization on false negative pairs. Our extensive experiments show that TuneCLIP consistently improves performance across model architectures and scales. Notably, it elevates leading open-weight models like SigLIP (ViT-B/16), achieving gains of up to +2.5% on ImageNet and related out-of-distribution benchmarks, and +1.2% on the highly competitive DataComp benchmark, setting a new strong baseline for efficient post-pretraining adaptation.


Two Speeds of Learning: A Representation-Readout Decomposition of Grokking and Double Descent

Chi-Ning Chou ⋅ Oscar Uzdelewicz ⋅ Neng-Chun Chiu ⋅ Yao-Yuan Yang ⋅ SueYeon Chung

Training loss and accuracy are the standard signals used to monitor generalization during deep neural network training. Two well-documented phenomena complicate this picture: in grokking, train loss falls rapidly while test performance improves abruptly only after a long delay; in epoch-wise double descent, train loss decreases monotonically while test loss or error rises and falls. Existing accounts are often task-specific, and a task-agnostic analysis framework for diagnosing and explaining these phenomena across realistic tasks and architectures is missing. We address this challenge by analyzing two competing processes that underlie learning dynamics: representation learning in the encoder and readout calibration in the final classifier. Using tools from representational geometry, neural tangent kernels, and linear probing, we show that both processes are active throughout training, with the fluctuations of their relative speed giving rise to seemingly anomalous generalization dynamics. Applying the representation-readout decomposition to grokking across a wide range of tasks and architectures, we find that the readout is train-biased before grokking onset, and representation learning is gradual but not absent—contrary to the lazy-to-rich account. The framework further provides diagnostic signatures distinguishing spurious from genuine generalization: in a previously reported MNIST grokking example and an epoch-wise double descent example, apparent delayed or non-monotone generalization is shown to arise from representation degradation and readout misalignment induced by non-standard training recipes. Together, these results establish the representation-readout decomposition as a top-down framework for understanding learning dynamics and revealing underlying algorithms for interpretability research.

Diffusion Transformers (DiTs) have emerged as a core architecture in generative modeling due to their scalability and adaptability to multimodal tasks. DiTs comprise isotropic transformer blocks, and learn representations progressively across depth, where the denoising objective drives later layers to focus on fine-detail reconstruction. This results in degraded representation quality and an imbalanced encoder--decoder behavior. Prior approaches such as representation alignment (REPA) mitigate this by encouraging stronger early representations via training regularization. Alternatively, U-Net-style DiT architectures introduce explicit multi-scale encoder–decoder structures for improved convergence. But they build on standard U-Net wisdom via learnable operators for spatial downsampling, which are not well-suited to transformer architectures, introducing inefficiencies and compatibility issues with components such as cross-attention and representation regularization. In this work, we propose **UDT**, a U-Net diffusion transformer that combines the representation power of DiTs with the encoding–decoding benefits of U-Nets, through data-adaptive token merging for downsampling and upsampling, while preserving the DiT token dimension. Our baseline UDT architecture outperforms existing U-Net DiTs and achieves performance comparable to REPA across all model sizes. Furthermore, using architectural optimization and REPA, UDT outperforms SiT's 7.9 FID at 1400 epochs (w/o CFG) within 40 epochs ($\approx 40\times$ faster convergence) for XL model size. Finally, it achieves strong image generation results on ImageNet with 1.44 FID at $256 \times 256$ (with CFG) after only 200 epochs, providing a new backbone for DiTs with strong empirical benefits.


Understanding Private Evolution as Learning-Augmented Clustering

Audra McMillan ⋅ Kunal Talwar ⋅ Felix Zhou

Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst‑case Wasserstein analyses would predict. We recast PE as a generative model‑augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then PE provably obtains much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then the sample complexity depends on the intrinsic, not the ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm has consistent empirical gains over standard baselines.


Unified Generative-Predictive Modeling for 4D Scene Understanding

Amani Kiruga ⋅ Zhiyi Li ⋅ Ruojin Cai ⋅ Hansen Lillemark ⋅ Xiaoshan Wu ⋅ Ayush Tewari ⋅ Yilun Du

Understanding 4D scenes is a core challenge in visual intelligence, requiring models to jointly capture scene geometry, appearance, and dynamics from partial observations. Existing approaches, however, are subdivided into feedforward models that are used for direct geometric prediction and generative models for visual synthesis. We propose a unified generative-predictive framework for 4D scene understanding that jointly models RGB video and geometric representations such as depth and camera pose with a single joint generative model. Our generative formulation enables us to learn from diverse, heterogeneous datasets with different subsets of geometric and visual annotations, and enables us to condition on arbitrary subsets of inputs, allowing a single model to perform cross-modal generation, geometric prediction, and future scene inference from partial observations. In addition, our generative approach enables test-time search over multiple candidate scene completions, improving geometric prediction from sparse observations. Overall, we illustrate the efficacy of our generative-predictive framework on a suite of generation and prediction tasks.


UniPro: Unified Multi-Mode Medical Image Segmentation from 2D Images to 3D Volumes via Propagation

Bangwei Guo ⋅ Yunhe Gao ⋅ Meng Ye ⋅ Yang Zhou ⋅ Difei Gu ⋅ Guoning Zhang ⋅ Leon Axel ⋅ Dimitris Metaxas

Medical image segmentation remains fragmented along two axes: segmentation paradigms and data dimensionality. Existing methods are typically developed separately for semantic, in-context, and interactive segmentation, and are further specialized to either native 2D images or 3D volumetric data. In clinical practice, however, segmentation workflows take many forms: a case may be initialized by semantic prediction, reference-guided segmentation, or user interaction. Regardless of how it begins, fine-grained refinement is naturally performed on 2D views; for volumetric scans, such 2D edits must propagate coherently to the rest of the volume. We present UniPro, a unified model that bridges segmentation paradigms and data dimensionality, using propagation to extend 2D segmentation to 3D volumes. Our key insight is that volumetric propagation and in-context segmentation share the same reference-conditioned prediction mechanism, differing only in whether the reference image–mask pairs come from other cases or from previously segmented neighboring slices. Building on this view, UniPro supports semantic, in-context, interactive, and propagation-based segmentation within a single slice-based framework, using class priors, reference exemplars, user clicks, and neighboring-slice predictions as mode-specific conditioning inputs. To improve propagation reliability, UniPro further incorporates bidirectional and 3D supervision to regularize slice-wise propagation beyond per-slice losses. Extensive experiments across diverse modalities and anatomies show that UniPro achieves strong performance across all segmentation settings, enabling annotation-efficient 3D segmentation from sparse 2D initialization and reducing slice-by-slice correction effort.


Urgency-Aware Autoregressive VLMs for Unanticipated Healthcare Occurrences

Qian Wu ⋅ Kai Chen ⋅ Zelong Tan ⋅ Simin Li ⋅ DOU QI

The sequential nature of autoregressive decoding in large vision-language models creates a fundamental trade-off between latency and clinical responsiveness when managing asynchronous signals. When unexpected healthcare occurrences arise during an ongoing VLM response, integrating unanticipated events---such as sudden device alerts or urgent caregiver messages---conventionally requires either injecting the new data directly into the active response stream or deferring it until generation ends. Neither approach is ideal in dynamic environments that demand urgent intervention without repeatedly disrupting routine continuations. We present ReflexMind, a training-free framework for online, urgency-aware occurrence handling that runs an Urgency Evaluator (UE) concurrently with the Primary Generator (PG). By leveraging the rotational invariance of Rotary Position Embedding (RoPE), ReflexMind enables a dual-view decoding process that shares a single KV cache, allowing it to assess incoming multimodal occurrences while preserving the state of the ongoing response. This design supports structured decisions, enabling ReflexMind to escalate critical alerts immediately while handling non-urgent occurrences without interrupting the ongoing response. Across three benchmarks spanning multimodal occurrences, our method generally improves occurrence triage, urgency recognition, and intervention selection for VLM backbones at different scales, while using fewer post-occurrence tokens in most settings and adding modest runtime and memory overhead under concurrent occurrences.


Variational Trajectory Optimization of Anisotropic Diffusion Schedules

Pengxi Liu ⋅ Zeyu Michael Li ⋅ Xiang Cheng

We introduce a variational framework for diffusion models with anisotropic noise schedules with a matrix-valued path $M_t(\theta)$ that allocates noise across subspaces. Central to our framework is a trajectory-level objective that jointly trains the score network and learns $M_t(\theta)$, which encompasses general parameterization classes of matrix-valued noise schedules. We further derive an estimator for $\partial_\theta \nabla \log p_t$ that enables efficient optimization of the $M_t(\theta)$ schedule. For inference, we develop an efficiently-implementable reverse-ODE solver that is an anisotropic generalization of the second-order Heun discretization algorithm. Across CIFAR-10, AFHQv2, FFHQ, and ImageNet-64, our method consistently improves upon the baseline EDM model in all NFE regimes.


VeriWorld: A Verifiable Visual SWE-Bench for Spatial Reasoning in 3D Environments

Yan Zheng ⋅ ShengYun Peng ⋅ XUYAO LIANG ⋅ Xiaoyan Cong ⋅ Nuo Chen ⋅ Bangya Liu ⋅ Zihan Wang ⋅ Wenyan Cong ⋅ Zhiwen Fan ⋅ Zhangyang "Atlas" Wang

Vision-language models (VLMs) are increasingly used for spatial tasks such as navigation and manipulation, where they must infer geometry and object relationships from visual input in order to act correctly. Existing benchmarks usually test perception and reasoning in isolation, making it hard to tell whether failures come from visual understanding, downstream reasoning, or the bridge between them. How can we evaluate whether a model can recover the right spatial structure from visual input and use it to make correct decisions? We introduce VeriWorld, a benchmark for spatial reasoning in interactive 3D environments with executable actions and deterministic verification. Each task is evaluated under three matched input settings -- visual-only, structured, and combined -- while the environment and objective remain fixed. This design isolates whether failures arise from perception, reasoning, or their interaction. VeriWorld brings together interactive 3D environments, code-based actions, deterministic verification, parameterized task generation, and controlled diagnostic evaluation in a single framework. Using this setup, we uncover two findings. First, models often succeed when given structured spatial information but fail when the same information must be inferred from visual input, revealing a recurring visual-to-structure gap. Second, action-space and harness design can change outcomes even under the same task and information condition, showing that interaction protocol is itself a confounding variable in spatial evaluation. VeriWorld is open-sourced to support reproducible and fine-grained evaluation of spatial reasoning in VLMs.

Multimodal large language models (MLLMs) are typically evaluated by whether they produce correct answers, but correctness alone does not reveal whether their outputs are supported by the right visual evidence. We introduce VIGOR, a benchmark for evaluating VIsually GrOunded Rationales in MLLMs. VIGOR represents visual rationales as explicit concepts paired with pixel-level groundings, covering multiple abstraction levels including low-level visual properties, intermediate structures, objects, parts and high-level motion, relations. To construct VIGOR at scale, we develop an automatic annotation framework that reuses existing dense supervision when available and generates missing concept groundings through proposal-based annotation followed by conservative multimodal verification. Using VIGOR, we evaluate whether MLLMs "look at what they say" by comparing model-derived visual evidence for each expressed concept against annotated ground-truth rationales. Our results show that current MLLMs exhibit limited visual rationale correctness and reveal distinct grounding failures across concept levels. VIGOR provides a benchmark and scalable data construction pipeline for studying whether MLLM outputs are supported by appropriate visual evidence.

Speaker identification benchmarks evaluate how accurately a model can predict a ground-truth speaker label for a given utterance, treating voice identity as a property of the recording's producer. However, these benchmarks do not test whether a model aligns with human perception of voice identity. Whether a human would hear two voices and interpret them as the same speaker or not has consequences for intellectual property rights, cybersecurity, and everyday life. We make two contributions. First, we introduce the Voice Identity Perception benchmark (VIPBench), built from the identity judgments of English-speaking crowdworkers on stratified celebrity voice pairs. VIPBench consists of 124,876 same/different identity judgments from 1,290 participants on 9,800 voice pairs across 100 speakers, spanning real recordings, AI voice clones generated by a state-of-the-art Text-to-Speech (TTS) system, and continuously morphed voices. Our second contribution is to define four evaluation tasks: predicting the listener same-speaker agreement rate; binary same/different classification against the majority listener vote, evaluated for both ranking and calibration; alignment between the model's and listeners' speaker-similarity structures; and whether a predictor fit on real speech still works on voice clones and morphs. We report baselines for ten publicly available speech representations. We find that the perception target re-orders model rankings relative to metadata: supervised embeddings (trained on metadata speaker labels) still outperform self-supervised models (which learn voice structure without identity supervision), and no model reaches the noise ceiling. VIPBench enables systematic evaluation of speaker representations against human voice-identity perception across real, cloned, and morphed speech, motivating speaker models trained on perceptual judgments directly.


VOID: Backdoor Injection through Knowledge Vacuity in Federated Unlearning

Wenwei Zhao ⋅ Yuxuan Xie ⋅ Haiyun Liu ⋅ Jie Xu ⋅ Zhuo Lu

Federated unlearning (FU) enables federated systems to remove designated data from a trained global model, but its security risks remain poorly understood. We show that calibration-based FU introduces a structural vulnerability through finite-step post-hoc corrections, which leave behind \emph{knowledge vacuity} in weakly constrained residual dimensions where the unlearned data's influence is suppressed while retained-task recovery pressure remains limited. We propose VOID, an unlearning-phase backdoor attack that exploits knowledge vacuity to implant trigger semantics along the legitimate unlearning trajectory. VOID identifies these residual dimensions through influence-based trajectory estimation and neuron-level vacuity profiling, then performs masked semantic substitution during unlearning. Across datasets and FU methods, VOID achieves up to 99\% attack success, preserves clean accuracy, and persists after post-unlearning finetuning. Our results show that approximate forgetting can expose writable capacity for adversarial reuse.


WebNavigator: Global Web Navigation via Interaction Graph Retrieval

Xuanwang Zhang ⋅ Yuteng Han ⋅ Jinnan Qi ⋅ Xinyu Liu ⋅ Mulong Xie ⋅ Zhen Wu ⋅ Xinyu Dai

Despite significant advances in autonomous web navigation, current methods remain far from human-level performance in complex web environments. We argue that this limitation stems from Topological Blindness, where agents are forced to explore via trial-and-error without access to the global topological structure of the environment. To overcome this limitation, we introduce WebNavigator, which reframes web navigation from probabilistic exploration into deterministic retrieval and pathfinding. WebNavigator constructs Interaction Graphs via zero-token cost heuristic exploration offline and implements a Retrieve-Reason-Teleport workflow for global navigation online. WebNavigator achieves state-of-the-art performance on WebArena and Online-Mind2Web. On WebArena multi-site tasks, WebNavigator achieves a 72.9% success rate, more than doubling the performance of enterprise-level agents. This work reveals that Topological Blindness, rather than model reasoning capabilities alone, is an underestimated bottleneck in autonomous web navigation.


Weighted Conformal Clustering

Anirban Nath ⋅ YoonHaeng Hur ⋅ Genevera Allen

Clustering is a central tool for discovering latent structure in unlabeled data, yet modern clustering pipelines often end with a hard assignment of each observation to a cluster without rigorous measures of assignment uncertainty. We propose a novel weighted conformal approach for constructing confidence sets for cluster labels. The key difficulty is that the labels available for calibration are not observed ground-truth labels, but synthetic labels produced by a data-dependent clustering algorithm. Our method develops a conformal inference algorithm that corrects the resulting mismatch with the latent target labels through weights by formulating conformal clustering as a conditional label-distribution shift problem. We first derive an oracle procedure that attains finite-sample marginal coverage and then develop a computationally tractable and implementable version using estimated conditional label probabilities and augmented calibration. We show that the coverage of the estimated-weight procedure depends on the estimator, giving an explicit bound on the loss relative to the nominal level. Empirical studies beyond the mixture-model settings, explored by the recently developed stochastic split conformal clustering procedure, demonstrate that the proposed weighted approach offers improvements in non-linear and high-dimensional clustering applications in terms of set size.


What Makes a Good Path? Factoring Manifold Support and Path Geometry

Zhixuan Zhou ⋅ Tingting Dan ⋅ Guorong Wu

The manifold hypothesis suggests that real-world data concentrates on low-dimensional structures, yet navigating meaningful trajectories on latent manifolds remains an open challenge. A common approach is to construct density-aware metrics, but such constructions bundle two distinct questions: (i) where valid data lies and (ii) how we choose to traverse on the manifold. Since different tasks may demand different traversal behaviors even on the same data manifold, specifying these two aspects independently offers additional modeling flexibility. We propose to factor the path objective into a kinetic term governed by a freely chosen geometric prior and a score-based potential, derived from the negative log-density of a pretrained diffusion model, that acts as a soft constraint for manifold adherence. The potential encourages dynamics to evolve along the data support, while the geometric prior remains free to encode task-specific notions of path optimality. To solve the resulting optimization problem, we introduce \textit{geodesic force matching} (GFM), a direct-collocation algorithm that discretizes the trajectory and optimizes all waypoints jointly to satisfy the Euler--Lagrange force balance in a least-squares sense. On synthetic manifolds and 3D shape interpolation, our method produces smooth, on-manifold paths and performs favorably against recent density-aware baselines, most notably in the low-noise regime where those baselines become numerically unstable.


What Should a Streaming Video Model Remember?

Haonan Ge ⋅ Yiwei Wang ⋅ Hang Wu ⋅ Yujun Cai

Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets. Existing methods address this by adding memory banks, retrieval modules, or visual token compression to preserve long-range history. However, strong recent-window baselines show that indiscriminate history injection can dilute current-scene perception, suggesting that the key challenge is not whether to use memory, but how to allocate it selectively. We formulate this as budgeted online latent evidence allocation and propose SelectStream, a selective latent-memory framework that keeps the current observation directly visible to a frozen VLM while exposing historical information only through a compact, query-conditioned evidence budget. Three coordinated mechanisms govern when to write, what to preserve, and how to retrieve: surprise-driven adaptive windowing, priority-preserving consolidation, and query-conditioned graph reasoning over a fixed-capacity latent memory graph. Retrieved evidence is calibrated and injected as latent tokens for answer generation, without replaying frames or growing the context with stream length. Experimental results show that SelectStream achieves strong online streaming performance and preserves general video understanding, reaching 82.67\% on StreamingBench, 67.03\% on OVO-Bench, and 74.4\% average accuracy on offline video benchmarks, while outperforming strong recent-window baselines and prior streaming memory methods.

Process reward models (PRMs) dramatically outperform outcome reward models (ORMs) for math reasoning and agent tasks, yet provide no benefit for chat and QA. No theory explains why, and practitioners choose between them by expensive trial and error. We introduce state-lift, a diagnostic that predicts whether a PRM is worth training for a new domain — in minutes, from ~500 labeled steps, before committing to expensive reward model training. State-lift measures how much step quality depends on trajectory state versus action text alone, and an accompanying effective gap distinguishes two regimes: state-noise-dominant (PRM essential) and action-gap-dominant (ORM sufficient). Across ten domains and eleven published PRM-ORM comparisons, state-lift correctly predicts PRM advantage, and a minutes-scale linear estimate matches the outcome of full Llama-3.1-8B reward model training — including CaSiNo negotiation (+0.055 lift at SL = 0.22). The accompanying two-branch decision protocol turns state-lift into an actionable PRM-vs-ORM workflow that runs in minutes on a single CPU.


When Do Multi-Agent Systems Help? An Information Bottleneck Perspective

Wendi Yu ⋅ Lianhao Zhou ⋅ Xiangjue Dong ⋅ Sai S Barath ⋅ Declan Staunton ⋅ Byung-Jun Yoon ⋅ Xiaoning Qian ⋅ James Caverlee ⋅ Shuiwang Ji

LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsistently across settings. Here, we provide an information bottleneck perspective on elucidating the differences between MAS and SAS. Specifically, our key observation is that a SAS accumulates its full reasoning trace in one shared context, while a MAS uses isolated local contexts connected by bounded relay messages. We show that, under infinite relay bandwidth, any SAS can be simulated by a MAS that transmits the full upstream context. Thus, the nontrivial advantage of MAS arises under bounded relays, where compression introduces a fundamental trade-off: reducing redundant context can improve efficiency, but may also incur loss of task-relevant information. We formalize this trade-off as an information bottleneck controlled by an effective parameter $\beta$, which captures how the balance shifts with model capability, and shows that MAS gains arise when context reduction outweighs relay information loss. We conduct 18 controlled experiments across five benchmarks and three model scales to validate our theoretical studies. We observe that MAS consistently helps when relays are near-sufficient, especially for weaker models. In contrast, MAS gains shrink or reverse when relays incur information loss, especially for stronger models that can already extract useful information from redundant context and thus gain little from compression. Our study shows that multi-agent design is fundamentally an information-bottleneck optimization problem. This perspective explains when bounded inter-agent communication helps or hurts.

Direct Preference Optimization (DPO) is increasingly used in settings where pairwise preference signals coexist with pointwise criteria such as response length, answer correctness, or safety scores. When the two disagree, standard DPO does not explicitly distinguish pairs that improve the pointwise criterion from those that worsen it. We show that this mismatch induces a systematic failure mode, which we call **constraint drift**: updates on disagreement pairs tend to increase pointwise violation, while updates on aligned pairs tend to repair it, with the overall drift governed by their balance. Under a first-order approximation, this balance yields a closed-form pre-alignment diagnostic, $\rho^{*}_\mathrm{eq}$, computable from dataset statistics alone, that predicts whether the data is self-correcting or requires intervention before alignment begins. Building on this diagnostic, we propose **Conflict-aware DPO (CoDPO)**, a calibrated modification to the DPO loss with two complementary components: a coarse per-sample control for primary drift regulation and a fine adaptive margin tilting for residual correction during training. Across six domains, drift grows monotonically with conflict exposure, and $\rho^{*}_\mathrm{eq}$ tracks the empirical drift transition, including self-correcting settings where intervention is unnecessary. Across four base-model architectures and multiple preference losses, CoDPO reduces drift while preserving preference learning. Compared with hard filtering and recent constrained-alignment and multi-objective baselines, CoDPO achieves a more favorable trade-off between drift control and preference performance.


When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

Yuan Tian ⋅ Bing Hu ⋅ Fang Wu ⋅ Xiaomin Li ⋅ binghang lu ⋅ Neil Gong

Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response generation, text-only prior turn, visual-state manipulation, and explicit external image-tool invocation. In this paper, we ask which of these evaluated paradigms improves multimodal jailbreak robustness, and why. Across multiple vision-language models, explicit image-tool interaction yields the lowest attack success rates in our experiments, reducing jailbreak success by around 30\% relative on average across the evaluated models. This finding is initially surprising: ASR remains low even when the returned image-tool output is manually overridden or itself unsafe-looking, but returns near direct-answering levels under text-only prior turn controls. These results indicate that the lower ASR is not explained by benign returned-image semantics or by the textual image-tool trace alone. To explain the pattern, we introduce an image-tool safety vector framework that models image-tool invocation as a residual shift in hidden representations toward a safety-relevant direction. Representation-level analyses and activation interventions support this account. Overall, our results suggest that explicit image-tool interaction is a promising design pattern for improving jailbreak robustness, while also motivating pipeline-specific safety evaluation.


Why Muon Outperforms Adam: A Curvature Perspective

Shuche Wang ⋅ Fengzhuo Zhang ⋅ Jiaxiang Li ⋅ Dirk Bergemann ⋅ Zhuoran Yang

Muon improves training efficiency over Adam in large language-model training by about two times, but the local geometric source of this advantage remains unclear. Our work takes a first step toward demystifying Muon's superiority over Adam from a curvature perspective. First, we apply a second-order Taylor approximation to the training landscape and show that Muon achieves a larger one-step loss decrease than Adam at matched validation loss. The two optimizers have comparable first-order gains, but Muon consistently incurs a smaller second-order curvature penalty. Second, we decompose this curvature penalty into the squared update norm and Normalized Directional Sharpness (NDS). We find that Muon and Adam have comparable update norms, so Muon's smaller curvature penalty is driven by lower NDS, not update scale. Third, we study how training data and model structure shape Muon's NDS advantage. Using Zipf Probabilistic Context-Free Grammar (PCFG) data with controlled imbalance, we show that data imbalance amplifies Muon's NDS advantage over Adam. A within-/cross-layer decomposition further shows that, in the middle and late stages of training, Muon's lower NDS is mainly sustained by smaller within-layer curvature. Beyond empirical evidence, we analyze stylized quadratic problems with heterogeneous curvature and gradient alignment toward high-curvature modes. We prove that Muon attains a smaller average NDS than GD by balancing update energy across curvature groups; when curvature heterogeneity is sufficiently strong, this also yields lower local quadratic loss after the same number of steps.


Your Embedding Model Is SMARTer Than You Think

Jianrui Zhang ⋅ Hyun Jung Lee ⋅ Sukanta Ganguly ⋅ Tae-Eui Kam ⋅ Donghyun Kim ⋅ Yong Jae Lee

Multimodal retrieval relies heavily on single-vector retrievers, which compress rich, sequential token sequences into one single global representation. While efficient, they discard fine-grained, local evidence critical for dense retrieval tasks. Multi-vector approaches were introduced as a solution, but they strictly require training and many ignore the necessity of a globally summarizing representation. To address this, we introduce SMART, a framework that unlocks the latent multi-vector capabilities of standard single-vector models. We first demonstrate that standard contrastive training on the pooled embedding implicitly shapes the retrieval geometry of preceding hidden states via gradient flow. By applying direct late-interaction over these frozen hidden states during inference, SMART acts as a plug-and-play upgrade that consistently improves performance across diverse modalities, pushing even the state-of-the-art models to new heights on MMEB-V2. We further reveal SMART's superior performance as simple lightweight post-training not only saves time and compute, but also brings forth further improvement on Visual Document retrieval, allowing a single-vector model to outperform SoTA multi-vector counterparts. Ultimately, SMART offers both a highly efficient inference enhancement and a powerful finetuning technique for multimodal retrieval.


Z0-Inf: Zeroth Order Approximation for Data Influence

Narine Kokhlikyan ⋅ Diego Garcia-Olano ⋅ Kamalika Chaudhuri ⋅ Saeed Mahloujifar

Influence functions are a popular tool for measuring the impact of a training point on a model's prediction; in particular, self-influence, which quantifies the influence of a training point on itself, has found many uses such as data selection and outlier detection. However, the use of influence functions has been severely limited in modern models such as LLMs due to their low accuracy or high computational cost: most existing algorithms are either highly inaccurate, or require computing gradients or approximations to inverse Hessians, which can be prohibitive for large models. In this work, we introduce a highly efficient zeroth-order approximation to data influence that requires only a fraction of the time and memory footprint of previous methods and is applicable to both differentiable and non-differentiable loss functions. We demonstrate that in addition to its computational advantage, our method delivers superior accuracy in estimating self-influence and comparable or better accuracy in estimating train-test influence for fine-tuned large language models, paving the way for broader and more practical application of influence-function techniques in state-of-the-art AI systems.