Skip to yearly menu bar Skip to main content


Session

Atlanta Poster Session 5

Hall C1
Sat 12 Dec 2 a.m. AEDT — 5 a.m. AEDT
Abstract:
Chat is not available.


$\chi$-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

Haolin Chen ⋅ Deon Metelski ⋅ Leon Qi ⋅ Tao Xia ⋅ Joonyul Lee ⋅ Steve Brown ⋅ Kevin Riley ⋅ Frank Wang ⋅ T. Y Liu ⋅ Hank Capps ⋅ Zeyu Tang ⋅ Xiangchen Song ⋅ Lingjing Kong ⋅ Fan Feng ⋅ Tianyi Zeng ⋅ Zhiwei Liu ⋅ Zixian Ma ⋅ Hang Jiang ⋅ Fangli Geng ⋅ Yuan Yuan ⋅ Chenyu You ⋅ Qingsong Wen ⋅ Hua Wei ⋅ Yanjie Fu ⋅ Yue Zhao ⋅ Carl Yang ⋅ Biwei Huang ⋅ Kun Zhang ⋅ Caiming Xiong ⋅ Sanmi Koyejo ⋅ Eric Xing ⋅ Philip S Yu ⋅ Weiran Yao

End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: *policy density*, decisions must be grounded in a large library of medical, insurance, and operational rules; Multi-role *composition*: a single task requires the agent to play multiple roles with handoffs; and *multilateral interaction*: intermediate workflow steps are multi-turn dialogs, such as peer-to-peer review and patient outreach. We introduce $\chi$-Bench, a benchmark of long-horizon healthcare workflows across three domains: provider prior authorization, payer utilization management, and care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed via 87 MCP tools, which it must drive to a terminal status through tool calls and writing the role's artifacts, guided by a 1,290+ document *managed-care operations handbook* skill. Across 30 agent harness/models configurations, the best agent resolves only **28.0%** of tasks, no agent clears **20%** on strict pass^3, and executing all tasks in a single session slumps the performance to **3.8%**. These results raise the hypothesis that similar gaps are likely to surface in other policy-dense, role-composed, irreversible enterprise domains.

Optimizing non-convex functions is a fundamental challenge across machine learning and combinatorial optimization. We introduce and study $\gamma$-weakly $\theta$-up-concavity, a novel first-order condition that characterizes a broad class of such functions. This condition provides a powerful unifying framework, strictly generalizing both DR-submodular and One-Sided Smooth (OSS) functions while capturing broader forms of scale-dependent curvature, including accumulating-then-diminishing returns and flat-start behavior. Our central theoretical contribution demonstrates that $\gamma$-weakly $\theta$-up-concave functions are upper-linearizable: for any feasible point, we can construct a linear surrogate whose gains provably approximate the original non-linear objective. A key technical contribution is a nonuniform upper-linearization argument yielding approximation coefficients that depend explicitly on the curvature parameters and the geometry of the feasible region. This linearizability yields immediate and unified approximation guarantees for a wide range of problems. Specifically, we obtain unified approximation guarantees for offline optimization as well as static and dynamic regret bounds in online settings via standard reductions to linear optimization. Moreover, our framework recovers the optimal approximation coefficient for DR-submodular maximization and improves existing approximation coefficients for OSS optimization, particularly over matroid constraints.


A$^3$Bench: A Benchmark for Compositional Reasoning over Aggressive Interactions in Videos

Austin C Baker ⋅ Raiyaan Abdullah ⋅ Daniel Z Yoffe ⋅ Shruti Vyas ⋅ Yogesh Rawat

Recent Multimodal Large Language Models (MLLMs) have achieved strong performance on general video understanding benchmarks, yet their ability to reason about participant roles in aggressive interactions remains largely unexplored. Existing violence and crime datasets primarily evaluate whether aggression occurs, without testing whether models can identify who performed the action, who received it, and how participants are related within the scene. To address this gap, we introduce **A$^3$Bench**, a benchmark to study *aggressive action and association* in videos. The benchmark evaluates whether models can jointly recognize aggressive actions, distinguish aggressors from victims and bystanders, and correctly bind actions to participants under challenging distractor settings. We benchmark 15 recent MLLMs and find that current models struggle substantially on fine-grained aggressive-scene understanding. Performance degrades consistently as reasoning complexity increases, with models frequently confusing aggressors with victims or bystanders, revealing a systematic failure in role-action association rather than coarse aggression recognition alone. We further observe that increasing model scale or applying generic reasoning prompts provides limited improvement. Finally, we explore a simple role-graph prompting strategy that encourages structured reasoning over participant interactions, leading to consistent gains on compositional reasoning. Our findings highlight that reliable aggressive-scene understanding remains a major challenge for current MLLMs and motivate future research on role-aware video reasoning beyond coarse violence detection.

Enforcing state-wise safety constraints is critical for the application of reinforcement learning (RL) in real-world problems, such as autonomous driving and robot manipulation. However, existing safe RL methods often enforce state-wise constraints only in expectation, while hard state-wise guarantees typically require strong dynamics assumptions. The former can still allow rare but severe violations, while the latter is impractical in model-free stochastic systems. We instead target high-probability control of the maximum instantaneous cost along a trajectory. To accomplish this goal, we propose Absolute State-wise Constrained Policy Optimization (ASCPO), a model-free policy search algorithm whose exact update provides a distribution-free high-probability certificate for state-wise constraint satisfaction in stochastic systems. The guarantee holds under finite variance and does not require Gaussian, unimodal, or light-tailed cost distributions. We demonstrate the effectiveness of our approach by training neural network policies for extensive robot locomotion tasks, where the agent must adhere to various state-wise safety constraints. Our results show that ASCPO most consistently reduces state-wise safety violations while maintaining competitive reward across challenging continuous control tasks.

We study the problem of K-armed bandits with reward distributions belonging to a one-parameter exponential distribution family. In the literature, several criteria have been proposed to evaluate the performance of such algorithms, including Asymptotic Optimality, Minimax Optimality, Sub-UCB, and variance-adaptive worst-case regret bound. Thompson Sampling-based and Upper Confidence Bound-based algorithms have been employed to achieve some of these criteria. However, none of these algorithms simultaneously satisfy all the aforementioned criteria. In this paper, we design an algorithm, Exponential Kullback-Leibler Maillard Sampling (abbrev. Exp-KL-MS), that can achieve multiple optimality criteria simultaneously, including Asymptotic Optimality, Minimax Optimality with a \sqrt{ln(K)} factor, Sub-UCB, and variance-adaptive worst-case regret bound.


A Control-Theoretic Approximation to Predictive Coding Dynamics

Ryan Fayyazi ⋅ Kyle Daruwalla ⋅ Mitra Javadzadeh

Backpropagation remains the dominant method for training deep neural networks, but its reliance on non-local computations has motivated the search for biologically plausible alternatives. Two prominent frameworks, predictive coding (PC) and deep feedback control (DFC), offer distinct approaches to local learning, yet their relationship remains unclear. Here, we show that DFC arises as a local approximation to predictive coding dynamics near fixed points, with exact equivalence in shallow linear networks and controlled deviations arising with depth, nonlinearity, and distance from equilibrium. Building on this connection, we introduce Approx PC, a practical algorithm that replaces layer-wise error propagation in PC with a centralized feedback controller operating over a sequence of intermediate subgoals. Empirically, Approx PC recovers solutions consistent with predictive coding, while converging faster and achieving improved accuracy in our experiments. Our results unify energy-based and control-theoretic perspectives on learning and suggest a computationally efficient approach to training predictive coding networks through control-inspired dynamics.

Action chunking can improve credit assignment and exploration in reinforcement learning by shortening the effective decision horizon, but existing action chunking reinforcement learning methods often rely on chunked state-action value functions which become challenging to learn as action dimensionality grows. Moreover, executing action chunks open loop sacrifices within-chunk reactivity, which is especially problematic in contact-rich robotic control. We present Action Chunking PPO (ACPPO), an extension of PPO that introduces temporal abstraction through a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online and provides reactivity within a chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregated benchmark performance and remains best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that regularizing the corrector balances long-horizon planning and local feedback. These results suggest that action chunking can be made effective in fully online PPO when open-loop temporal structure is paired with closed-loop correction.

Policy gradient methods may suffer from high variance, which limits their sample efficiency. Existing variance-reduction techniques largely focus on estimator design under a fixed sampling policy. In contrast, we ask whether the action-sampling distribution itself, often left at the on-policy default, can be optimized to reduce policy gradient variance. We show that this default is in fact not variance-optimal even when the state distribution is held on-policy. To exploit this, we introduce Action-Level Behavior Policy Optimization (BPO), a framework that treats the behavior policy as an action-level proposal distribution at each on-policy state, while keeping the state distribution itself on-policy. For a one-step importance-sampled policy gradient estimator, we sharply decompose its variance into an irreducible state term and a controllable action term, and show that the latter can be minimized in closed form via BPO. Under exact critic and sampling oracles, this yields a strictly tighter nonconvex SGD bound than on-policy sampling whenever the score-weighted critic is action-dependent. We then extend the analysis to approximate critics and parametrized behavior policies updated by a single stochastic step, and verify the resulting sample-efficiency gains empirically. These results identify action-level BPO as a principled mechanism for improving the sample efficiency of policy gradient methods.

Generative compressed sensing uses the range of a pretrained generator as a nonlinear model for recovering structured signals from limited measurements. We study a conditional version of this problem for image recovery from subsampled Fourier measurements using prompt-conditioned generative models. Our framework separates two roles of conditioning: the prompt used to design the sampling distribution and the prompt used to define the recovery model. For ReLU and Lipschitz conditional generators, we prove stable recovery bounds showing that prompt-matched Christoffel sampling retains the same Christoffel complexity constant as existing near-optimal generative compressed sensing theory, while prompt mismatch incurs an explicit compatibility penalty. Experiments with Stable Diffusion show that prompts meaningfully reshape Christoffel sampling distributions and influence image recovery. Overall, our results suggest that prompts should be treated as design variables with distinct effects on sensing, approximation, and recovery.

Long-video question answering forces multimodal large language models to reason under strict visual-token budgets, creating a tradeoff between broad temporal coverage and fine-grained spatial detail. Uniform sampling preserves coverage but often obscures decisive local cues, while caption-based or frame-similarity retrieval operates over lossy proxies that can miss motion, state changes, and before-after relations. We recast long-video reasoning as a visual token allocation problem: given a fixed budget, decide how much to spend on global context versus high-resolution local evidence, and where to draw that local evidence from. We introduce AdaAlloc, a training-free inference framework that adaptively allocates visual tokens between a global video overview and localized visual detail. AdaAlloc first plans a global/local allocation policy, then uses temporal grounding to locate candidate evidence segments and refines them into a compact temporal memory. The final answerer reasons over the allocated global context and refined local evidence. Under matched backbones and visual-token budgets, AdaAlloc consistently improves over strong baselines on challenging long-video benchmarks, including LVBench, MLVU, VideoMME, and LongVideoBench.


Adapting Actively on the Fly: Relevance-Guided Online Meta-Learning with Latent Concepts for Geospatial Discovery

Jowaria Khan ⋅ Anindya Sarkar ⋅ Yevgeniy Vorobeychik ⋅ Elizabeth Bondi-Kelly

In environmental monitoring, data collection is often costly, sparse, and shaped by urgent public-health needs. This is particularly true for cancer-causing PFAS (Per- and polyfluoroalkyl substances) contamination, where discussions with domain experts and environmental organizations highlight the need to strategically identify high-risk, under-observed regions under tight sampling budgets. More broadly, similar challenges arise in disaster response and public health settings, where dynamic environments make it essential to efficiently uncover hidden targets from limited ground truth. Yet sparse and biased geospatial labels limit the applicability of existing learning-based methods, such as reinforcement learning. To address this, we propose a unified geospatial discovery framework that integrates active learning, online meta-learning, and concept-guided reasoning. Our approach introduces two key innovations built on a shared notion of concept relevance, capturing how domain-specific factors influence target presence: a concept-weighted uncertainty sampling strategy, where uncertainty is modulated by learned relevance from readily available concepts such as land cover and source proximity; and a relevance-aware meta-batch formation strategy that promotes semantic diversity during online-meta updates, improving generalization in dynamic environments. We evaluate our framework on PFAS contamination discovery as a real-world inspired environmental monitoring task, demonstrating robust target discovery under limited data and changing conditions.


Adaptive Calibration in Non-Stationary Environments

Junyan Liu ⋅ Haipeng Luo ⋅ Lillian Ratliff

Making calibrated online predictions is a central challenge in modern AI systems. Much of the existing literature focuses on fully adversarial environments where outcomes may be arbitrary, leading to conservative algorithms that can perform suboptimally in more benign settings, such as when outcomes are nearly stationary. This gap raises a natural question: can we design online prediction algorithms whose calibration error automatically adapts to the degree of non-stationarity in the environment, smoothly interpolating between i.i.d. and adversarial regimes? We answer this question in the affirmative and develop a suite of algorithms that achieve adaptive calibration guarantees under multiple calibration measures. Specifically, with $T$ being the number of rounds and $C\in[0,T]$ being an unknown non-stationary measure defined as the minimal $\ell_1$ deviation of the mean outcomes, our algorithms attain $\widetilde{O}(\sqrt{T}+(TC)^{\frac{1}{3}})$ for $\ell_1$ calibration error and $\widetilde{O}((1+C)^{\frac{1}{3}})$ for both $\ell_2$ and pseudo KL calibration error. These bounds match the optimal rates in the stationary case ($C=0$) and recover known guarantees in the fully adversarial regime ($C=T$). Our approach builds on and extends prior work [Hu et al., 2025, Luo et al., 2025], introducing an epoch-based scheduling together with a novel non-uniform partition of the prediction space that allocates finer resolution near the underlying ground truth.


Adaptive Delayed-Update Cyclic Algorithm for Variational Inequalities

Yi Wei ⋅ Xufeng Cai ⋅ Jelena Diakonikolas

Cyclic block coordinate methods are widely used in practice for their simplicity and strong empirical performance. Yet, their theoretical behavior is challenging to explain, and setting their step sizes$-$beyond classical coordinate descent for minimization$-$requires careful tuning or line-search machinery. In this work, we develop $\texttt{ADUCA}$ (Adaptive Delayed-Update Cyclic Algorithm), a cyclic algorithm addressing a broad class of Minty variational inequalities with monotone Lipschitz operators. $\texttt{ADUCA}$ is parameter-free and locally adaptive: it requires no global or block-wise Lipschitz constants, it uses no per-epoch line search, and it adapts to local problem geometry. A key feature of the algorithm is using operator information delayed by a full cycle, which makes the algorithm compatible with parallel and distributed implementations, and attractive due to weakened synchronization requirements across blocks. We prove that $\texttt{ADUCA}$ attains (near) optimal global oracle complexity as a function of target error $\epsilon >0$, scaling with $1/\epsilon$ for monotone operators, or with $\log^2(1/\epsilon)$ for operators that are strongly monotone.


Adaptive Test Case Discovery for LLM-Assisted Decision Making in High-Stakes Domains

Anjali Parashar ⋅ Carson Sobolewski ⋅ Yingke Li ⋅ Fei Chen ⋅ Chuchu Fan

Large language models (LLMs) are increasingly used for assistive decision-making in high-stakes domains, yet their outputs can be unreliable and difficult to audit. This necessitates pre-deployment testing, which can be used to identify scenarios where an LLM-assisted pipeline fails to produce satisfactory decisions. Existing testing strategies are often not sample-efficient and struggle to adapt to diverse testing objectives, application domains, and rapidly evolving models. We introduce a sample-efficient scenario design strategy for testing LLM-based decision-making pipelines. Our approach formulates testing as adaptive experimental design, using a flow-based surrogate model to sequentially acquire scenarios that balance optimization to discover challenging test cases with diverse exploration. We evaluate the method across several LLM-assisted decision-making tasks: bank loan approval, stock prediction from news articles, and disaster resource allocation from tweets and images. Our approach consistently discovers challenging scenarios that satisfy the testing objective across all tasks, providing high coverage over scenario space.

We propose a deep learning framework for scalar-on-function regression with possibly multiple functional predictors observed on irregular grids. The method represents each functional covariate by finitely many basis coefficients computed from the observed trajectories and fits a deep neural network using these coefficients as features. We establish a general nonasymptotic excess risk bound for the neural network estimator that decomposes the error into statistical, truncation, discretization, and network approximation terms. We then specialize the theory to two important settings. (i) For continuous functional models, an equal-width neural network estimator achieves a nearly minimax optimal rate up to a log-logarithm factor. (ii) For functional multiple-index models, we design a bottleneck-equal-width neural network estimator that exploits the low-dimensional index-link structure and attains a nearly minimax optimal rate up to a logarithm factor. The results provide theoretical support for neural network-based scalar-on-function regression and highlight the importance of architecture design.


AdShot: Benchmarking Multimodal Large Language Models for Video Advertisement Clipping

Wen Xie ⋅ Om Rastogi ⋅ Sai S Rangoju ⋅ Gijs Overgoor ⋅ Yakov Bart ⋅ Sarah Ostadabbas

Advertisers routinely create multiple duration variants of ads to meet marketing budgets and viewer preferences, a labor-intensive and costly process. While multimodal large language models (MLLMs) excel at general video understanding, their ability to perform specialized editorial tasks remains unexplored. We introduce \textbf{AdShot}, a benchmark for evaluating MLLMs on shot selection for ad clipping. It contains 823 professionally edited ad pairs (30-sec sources and 15-sec edits) across 194 brands and 17 industries, annotated with shot-level boundaries and content features (action, emotion, information, imagery). We evaluate 13 state-of-the-art video-only MLLMs and 2 audio-visual MLLMs (7B–72B parameters) on shot selection accuracy (Precision, Recall, F1), duration constraints (15s), and content preservation. We find: (1) model architecture matters more than scale, with smaller models often matching or exceeding larger ones; performance is strongest for target ads with 6–10 shots; (2) models exhibit a mean absolute duration error of 5.39s, with most tending to overshoot the 15s target, though they can still substantially reduce human editing effort; (3) all models exhibit systematic content biases, favoring information-rich and emotion-focused shots while under-selecting action and imagery. We publicly release our benchmark to facilitate research at the intersection of multimodal AI and computational advertising.


A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training

Junze Ye ⋅ Jiayi Cheng ⋅ Miao Lu ⋅ Michal Mankowski ⋅ Jose Blanchet ⋅ Mohsen Bayati

Modern training pipelines for language model agents begin with a supervised fine-tuning stage in which a small student imitates a costly teacher. Recent work mitigates the covariate shift of pure imitation learning by collecting teacher feedback at states the student itself reaches, with a prevailing trend toward elaborate filters on the teacher's responses. We frame this design choice as a budget-allocation problem and compare three constructions of supervised training data: short unfiltered teacher continuations at learner-induced states; full teacher trajectories filtered for success (the rejection-sampling step in recent on-policy expert-correction work); or those further restricted to tasks the student cannot solve on its own. Across three agentic benchmarks (HotpotQA, ALFWorld, and Terminal-Bench-Dev), short unfiltered teacher continuations beat pure behavioral cloning at matched supervision budgets, and on HotpotQA also match or exceed the filtered alternatives. On Terminal-Bench-Dev, this construction at one-tenth of the corpus budget and with no reinforcement-learning stage matches the OpenThoughts-Agent baseline that uses the full corpus together with reinforcement learning; the optimal continuation length is task-dependent. The same supervised-learning checkpoints yield faster early-stage gains under subsequent reinforcement learning, with mixed evidence on whether this advantage persists later in training.


AgentArk: Distilling Multi-Agent Intelligence into a Single LLM Agent

Yinyi Luo ⋅ Yiqiao Jin ⋅ Weichen Yu ⋅ Mengqi Zhang ⋅ Srijan Kumar ⋅ Xiaoxiao Li ⋅ Weijie Xu ⋅ Xin chen ⋅ Jindong Wang

While large language model (LLM) multi-agent systems achieve superior reasoning performance through iterative debate, practical deployment is limited by their high computational cost and error propagation. This paper proposes AgentArk, a novel framework to distill multi-agent dynamics into the weights of a single model, effectively transforming explicit test-time interactions into implicit model capabilities. This equips a single agent with the intelligence of multi-agent systems while remaining computationally efficient. Specifically, we investigate three hierarchical distillation strategies across various models, tasks, scaling, and scenarios: reasoning-enhanced fine-tuning; trajectory-based augmentation; and process-aware distillation. By shifting the burden of computation from inference to training, the distilled models preserve the efficiency of one agent while exhibiting strong reasoning and self-correction performance of multiple agents. They further demonstrate enhanced robustness and generalization across diverse reasoning tasks. We hope this work can shed light on future research on efficient and robust multi-agent development. Our code is at https://anonymous.4open.science/r/AgentArk.


AI-Assisted Classification under Correlation Neglect and Trust

Saurabh Amin ⋅ Amine Bennouna ⋅ Daniel Huttenlocher ⋅ Dingwen Kong ⋅ Liang Lyu ⋅ Asuman Ozdaglar

When does an AI assistant improve human decision-making, and when does it hurt? We consider the setting in which a human worker makes a binary decision by observing her own signal and receiving a binary or probabilistic recommendation from an AI. We develop a behavioral model which considers both her belief regarding the overlap of her signal with the AI as well as the trust she has in the AI when considering its recommendation. The model yields a phase diagram in AI capability and signal overlap, partitioning the parameter space into complementarity, impairment, and automation regions. A worker who neglects signal overlap (correlation neglect) can experience performance loss when incorporating AI into her decision and is eventually dominated by AI alone as AI capability grows, with the complementarity range narrowing as overlap rises. Trust calibration shifts region boundaries: undertrust shrinks the impairment region but expands automation, while overtrust expands impairment but extends complementarity to be preferred at higher AI capability levels. Applying the framework to ChatBench (643 workers, GPT-4o and Llama-3.1-8b), AI assistance raises accuracy by 21 percentage points, but workers undertrust the AI by roughly a factor of four. The model yields design insights linking AI deployment choices, worker behavioral parameters, and outcome quality.


Algorithmic Impact Reveals the Hidden Structure of Alignment

Zachary Wojtowicz ⋅ Michelle Si ⋅ Finale Doshi-Velez ⋅ Ariel Procaccia

When an algorithm makes a decision affecting multiple people, we implicitly have a social choice problem: how should differing opinions about the right course of action be reconciled into a single outcome? We show that if we make welfare consequences of alignment a first-order consideration, this problem can be reformulated as linear optimization on a convex impact space, making it amenable to toolkits from welfare economics and mechanism design. This reformulation provides a clearer understanding of how different protocols shape an algorithm's externalities. Our linear optimization framework uncovers a straightforward solution to strategyproof alignment mechanism—random dictatorship that can be represented as a single parameter, which regulates externalities from manipulation. In addition, our framework allows us to express a family of constrained welfare mechanisms that address participation externalities—how an individual's participation can affect others in the population—and maximize social welfare subject to guarantees on individual outcomes. We illustrate these results empirically using real human preferences over kidney donation assignment, charitable food distribution, LLM responses, and trolley problems.


Alignment Needs 'Cognitive Control': On The Role of Regularization in LLM Alignment

Nona Ghazizadeh ⋅ Elnaz Rahmati ⋅ Morteza Dehghani ⋅ Payam Piray

This position paper argues that reference-anchored regularization in large language model alignment should be understood not as engineering overhead but as a functional necessity, the computational equivalent of cognitive control in the brain. The alignment community increasingly frames the cost of regularization as an "alignment tax" to be minimized, pursuing reference-free methods that remove the pretrained reference policy. We argue that this orientation is misguided. Cognitive neuroscience has established that the cost of overriding default behavior is a functional feature of intelligent systems, preventing overcommitment to noisy reward signals and preserving prior behavioral repertoires; its impairment is a causal vulnerability factor for addiction. When formalized as a divergence penalty from a structured default, this cost yields the same regularized objective used in alignment, and the pathologies observed when it is removed, including catastrophic forgetting, reward hacking, and loss of generalization, correspond to the consequences of impaired cognitive control. Rather than minimizing the cost of regularization, the field should design its structure systematically, specifying what the reference encodes, how the penalty on deviation should vary across contexts, and how to evaluate what the regularization preserves rather than only what it costs.


Allocate Marginal Reviews to Borderline Papers Using LLM Comparative Ranking

Elliot Epstein ⋅ Rajat Vadiraj Dwaraknath ⋅ John Winnicki ⋅ Thanawat Sornwanee

This position paper argues that ML conferences should allocate marginal review capacity primarily to papers near the acceptance boundary, rather than spreading extra reviews via random or affinity-driven heuristics. We propose using LLM-based comparative ranking (via pairwise comparisons and a Bradley--Terry model) to identify a borderline band before human reviewing. Given a venue-specific minimum review target (e.g., 3 or 4), we propose using this LLM based borderline signal to decide which papers receive one additional review (e.g., a 4th or 5th), without conditioning on any human reviews and without using LLM outputs for any accept/reject decision. We provide a simple expected-impact estimate based on (i) the overlap between the predicted and true borderline sets ($\rho$) and (ii) the incremental value of an extra review near the boundary ($\Delta$), using review data from ICLR 2025.


An Embarrassingly Simple Graph Heuristic Reveals Shortcut-Solvable Benchmarks for Sequential Recommendation

Haoyu Han ⋅ Li Ma ⋅ Hanbing Wang ⋅ Bingheng Li ⋅ Daochen Zha ⋅ Chun How Tan ⋅ Huiji Gao ⋅ Xin Liu ⋅ Stephanie Moyerman ⋅ Sanjeev Katariya ⋅ Hui Liu ⋅ Jiliang Tang

Sequential recommendation is a central task in recommender systems, and recent research has increasingly shifted toward generative recommenders that leverage both sequential patterns and semantic item information. However, these methods are often evaluated on a small set of widely used benchmarks. This raises a natural question: do these benchmarks actually require the advanced modeling capabilities of modern generative recommenders? We conduct a benchmark audit using an intentionally simple graph heuristic: starting from only the most recent one or two interacted items, it retrieves candidates from a few-hop item-transition graph and ranks them with item-feature similarity. Surprisingly, despite its simplicity, this heuristic matches or outperforms a broad set of modern baselines on a variety of popular sequential recommendation benchmarks. For example, it achieves relative NDCG@10 improvements of 38.10\% and 44.18\% over the best competing baseline on the widely used Amazon Review Sports and CDs datasets, respectively. We further show that this phenomenon is not merely an artifact of a particular heuristic, but reflects shortcut solvability in existing benchmarks. Specifically, we identify three potential shortcut structures that could make next-item prediction easier than expected: low-branching local transition structure, feature-smooth transitions, and limited dependence on long user histories. These shortcuts do not need to appear simultaneously. Depending on the dataset, even one or two strong shortcut signals can make simple local retrieval highly competitive, while weakening the relevant signals allows more sophisticated models to show clearer benefits. Our broader evaluation across 14 diverse datasets further shows that model rankings change substantially with dataset properties, while the simple graph heuristic remains competitive on 10 out of 14 datasets. These findings suggest that strong performance on several standard sequential recommendation benchmarks may not faithfully reflect whether recent methods achieve the advanced modeling capabilities they aim to demonstrate. Rather than treating datasets as interchangeable leaderboards, we argue for more careful dataset selection and dataset-level diagnostic analysis when using benchmarks to support claims about the benefits of new recommendation models.


Approaching Effective Merging in Model Embedding Space

Wei Chen ⋅ Zhenyuan Dong ⋅ Qiang Qiu

With the growing number of fine-tuned variants derived from pre-trained models, model merging emerges as a key challenge: effectively integrating task-specific adaptations while preserving performance. However, naive weight averaging often fails when parameter updates are misaligned, revealing the need for a structured space in which merging becomes principled. In this work, we hypothesize that pre-trained models possess an intrinsic low-rank structure. Specifically, we posit that the transformation from input to output is mediated by a latent bottleneck of low intrinsic dimensionality; we leverage this low-dimensional space to construct a model embedding space. This space ensures that fine-tuned updates, i.e., model embeddings, of all adaptations remain confined within the intrinsic structure of the pre-trained model. As model embeddings lie in the same low-dimensional space, model merging via direct averaging becomes more effective. Moreover, we further introduce an aligned subspace of these model embeddings to improve the efficacy of model merging. Experiments across vision and multimodal tasks, including ViT and multimodal LLMs, demonstrate that our approach enables effective model merging and provides a unified geometric view of adaptation and combination in pre-trained models.


Approximation Algorithms for GPU Pricing under Finite Capacity

Yaolong Yu ⋅ Hanrui Zhang ⋅ Zeyu Zheng ⋅ Kirthevasan Kandasamy

We study revenue-optimal posted-price design for GPU markets under a finite capacity $B$. We focus on two settings: \emph{sales}, where each purchase permanently consumes a vendor's GPU inventory, and \emph{rentals}, where each job occupies GPUs for a finite duration and then releases them. In both settings, the vendor must commit to an anonymous price menu before observing customers' valuation, GPU demand, and job duration. First, in the sales setting, we show that computing the optimal anonymous static menu is NP-hard, and that the common practice of linear pricing (per-GPU price) can be highly suboptimal. We develop three methods with varying approximation guarantees, all of which outperform linear pricing. Our main method is a convex program (CP) obtained via an ex-ante relaxation of the problem, which obtains a $\tfrac{1-\epsilon}{2}$ approximation to the optimum when the maximum GPU demand $\overline{N}$ satisfies $\overline{N}\leq \epsilon B$. Building on these results, we analyze the \emph{rental} model. We again show that linear pricing (per-GPU-hour price) is highly suboptimal, and design duration-aware posted-price policies based on slot-coupled ex-ante CP. Our \emph{Throttled CP-Menu Posted Pricing} (TCMPP) policy posts the CP-induced menu randomly and achieves a $\frac{1-\epsilon}{4}$ approximation. We further introduce a slackened version of the same convex program, which reserves capacity in each time slot and achieves stronger guarantees in the small job regime. We corroborate our theoretical results with simulations.


ArcMark: Distortion-Free Multi-Byte LLM Watermark via Optimal Transport

Atefeh Gilani ⋅ Sajani Vithana ⋅ Carol X Long ⋅ Oliver Kosut ⋅ Lalitha Sankar ⋅ Flavio Calmon

Watermarking is an important tool for promoting the responsible use of large language models (LLMs). Existing watermarks insert a signal into generated tokens that either flags LLM-generated text (zero-bit watermarking) or encodes more complex messages (multi-bit watermarking). Though a number of recent approaches insert multiple bits into text without perturbing average next-token predictions, they largely extend design principles from the zero-bit setting, such as encoding a single bit per token. In contrast, a watermarker capable of embedding multiple bytes into the text would dramatically increase the potential applications, by embedding information such as the ID of the user who submitted the prompt, the precise model version that was used, or even the prompt itself. We address this problem by introducing ArcMark: a new watermark construction based on coding and information-theoretic principles that is capable of reliably embedding multiple bytes of information into just a few hundred tokens, without any distortion of the underlying LLM next-token distribution. We derive ArcMark by formulating the distortion-free watermarking problem as a channel coding problem, and deriving an information-theoretic channel capacity that establishes the fundamental limit of embedding information in LLM output in a distortion-free manner. This capacity formulation informs the design of ArcMark. In practice, ArcMark outperforms competing multi-bit distortion-free watermarks in terms of reconstruction accuracy, including in the face of attacks that alter a subset of the LLM text. ArcMark output is also shown to be indistinguishable from unwatermarked text in terms of perplexity, and in downstream task quality.


A Single-Sample Polylogarithmic Regret Bound for Nonstationary Online Linear Programming

Haoran Xu ⋅ Owen Shen ⋅ Peter W Glynn ⋅ Yinyu Ye ⋅ Patrick Jaillet

We study nonstationary online linear programming (OLP), where orders arrive sequentially and follow independent but not-identical reward-resource distributions. The decision- maker seeks to maximize the expected total reward by making immediate and irrevocable acceptance or rejection decisions for each order, subject to a resource endowment available at the beginning of the planning horizon. Such problems arise in resource-constrained ML and AI systems, including inference admission control, online advertising and recommendation, and data-acquisition pipelines, where the mix and value of requests may change over time. We focus on a minimal-information regime in which the decision-maker observes only one independent sample from each future distribution before the horizon begins. We propose a novel re-solving algorithm that integrates a dynamic programming perspective with the dual-based frameworks traditionally employed in stationary environments. In the large-resource regime, where the resource endowment scales linearly with the number of order, we prove that our algorithm achieves $O((\log n)^2)$ regret across a broad class of nonstationary distribution sequences. Our results demonstrate that polylogarithmic regret is attainable even under significant environmental shifts and minimal data availability, bridging the gap between stationary OLP and more volatile real-world resource allocation problems.


A Training-Free Video Moment Retrieval Framework via Selection from Multiple Visual Prompts

Xun Mo ⋅ Shengtao Guo ⋅ Yongwei Nie ⋅ Fei Ma ⋅ Xuemiao Xu ⋅ Chengjiang Long

Video Moment Retrieval (VMR) aims to localize a temporal segment in a video based on a natural language query. While Multi-modal Large Language Models (MLLMs) are increasingly adopted for VMR due to their strong video understanding capability, they still face several limitations: (1) imprecise perception of event boundaries, (2) rigid dependency on precise object identification, and (3) inability to handle temporarily interrupted events. To address these, we propose three visual prompts: Event Segmentation (ES) prompt, Object Detection (OD) prompt, and Keyframe Marking (KM) prompt. Each prompt improves performance when applied individually, yet direct fusion degrades results because mixing prompts overcomplicates video content and inaccurate labels introduce noise. Although certain visual prompts may mislead the model, at least one of the three is likely to yield a positive effect. Even in the worst case where all three perform poorly, the prediction can fall back to the baseline. Motivated by this, we introduce a selection mechanism based on Yes/No Probability-Difference score to choose the most reliable prediction among the results produced by the three visual prompts and the baseline. We evaluate our method in a training-free setting across three datasets and four models, and obtain state-of-the-art performance.


Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics

Matthew Smart ⋅ Soumya Ganguly ⋅ Nilava Metya ⋅ Alexandre V Morozov ⋅ Anirvan Sengupta

We study minimal attention-only transformers under all-token corruption and show they admit a two-stage empirical Bayes interpretation. A single attention step computes a kernel-weighted posterior mean with respect to the empirical distribution defined by the context. Depth refines this distribution through particle dynamics (Stage 1), while a long-range skip-connection carries the noisy input as a query for posterior inference (Stage 2), revealing distinct statistical roles for depth and attention residuals. The framework isolates a minimal setting in which the context itself induces a depth-dependent energy landscape governing in-context inference. We show that effective denoising can emerge without an explicit noise schedule: a fixed kernel bandwidth and finite integration horizon suffice, yielding a principled depth–noise relationship. We further establish a posterior-mean recovery guarantee for a class of well-behaved priors, where the empirical estimator converges to the Bayes-optimal predictor under asymptotic conditions. Connecting these dynamics to reverse-diffusion limits, our results provide a statistical interpretation of attention as in-context inference via sample-based posterior estimation, without explicit density modeling.

Directed Acyclic Graphs (DAGs) are foundational to many areas of AI research, providing a formal language for causal reasoning and model interpretability. However, learning DAGs remains challenging, due to super-exponential computational cost and diminished accuracy in small sample regimes. To address these challenges, we introduce Attention-DAG (ADAG), a novel linear transformer architecture for unsupervised amortized DAG learning. Unlike traditional unsupervised DAG learning methods which recover the graph structure from data of each individual domain, ADAG leverages the knowledge from multiple domains. It provides a nonlinear mapping from observational data of each domain to both the corresponding graph structure and the underlying parameters of Structural Equation Models (SEMs). This enables efficient zero-shot inference of causal structures in new domains with unseen SEM parameters. Remarkably, we demonstrate that for DAG learning problem attention acts as a fixed-point iterative scheme, so the training on multiple domains effectively discovers an efficient solver for the constraint optimization problem. Extensive evaluations on synthetic and realistic benchmarks demonstrate that ADAG significantly outperforms existing baselines in accuracy and efficiency, particularly when data is scarce.


Audible World Models: Spatially Aware Sound Generation for 3D Worlds

Duowen Chen ⋅ Jinjin He ⋅ Gouthaman KV ⋅ Sandeep Bangalore Venkatesh ⋅ Bo Zhu

Text- and image-conditioned world generators can now produce visually rich 3D environments, but these worlds are usually silent or paired with a soundtrack generated only from text or rendered video. Such audio can describe what should be heard, but it does not explicitly represent where sounds live in the world or how they should change as a listener moves. We propose Audible World Models, a training-free framework that treats sound as part of a generated world state. Given a text prompt, our system builds a panoramic 3D proxy, decomposes it into semantic layers, identifies audible foreground objects and ambient background regions, synthesizes dry audio for each sound label, attaches sources to reconstructed geometry, and renders listener-dependent spatial audio through geometric acoustic propagation. This explicit coupling of semantics, geometry, and propagation produces audio that remains tied to persistent source locations and responds to viewpoint and motion. Across 80 generated scenes, our method substantially improves spatial consistency over text-, video-, and panorama-conditioned baselines while maintaining competitive semantic alignment. VLM-based and human evaluations further show that the resulting soundtracks are preferred for audio--visual consistency, spatial plausibility, and motion-dependent behavior.


Autonomous Driving Research Requires a Community-Driven Data Paradigm

Jinsu Yoo ⋅ Zanming Huang ⋅ Katie Luo ⋅ Zheda Mai ⋅ Qiyuan Wu ⋅ Bharath Hariharan ⋅ Mark Campbell ⋅ Wei-Lun (Harry) Chao

Autonomous driving has made remarkable progress, with recent AI advances enabling commercial deployments that are reshaping urban mobility. Yet the field remains far from its universal social promise: autonomous systems that can operate robustly anywhere, anytime, for anyone. We posit that this gap is not merely a modeling problem, but a problem of the prevailing data paradigm. Current research relies heavily on a few benchmark datasets with limited spatial and scenario coverage, even though the community has collectively produced over 600 autonomous driving datasets across nearly 50 countries. However, this abundance has not translated into broad research impact: most datasets remain significantly underused due to fragmentation, limited visibility, incompatible protocols, and benchmark incentives that concentrate attention on a few dominant datasets. We therefore argue that autonomous driving research requires a collaborative, community-driven data paradigm. Such a paradigm would improve the discovery, reuse, integration, and evaluation of diverse datasets; make underexplored data easier and more rewarding to study; and lower the barrier for new contributors. We outline its key principles, illustrate an early realization, and call for collaboration across academia and industry to transform fragmented datasets into shared community infrastructure for anytime-anywhere autonomy.


A World Model of Radiologist Reading for Medical Image Representation Learning

Yiwei Li ⋅ Zihao Wu ⋅ huaqin zhao ⋅ Yifan Zhou ⋅ Chao Cao ⋅ Dajiang Zhu ⋅ Tianming Liu ⋅ Lin Zhao

Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading, yet existing methods exploit this signal only partially, either as a static spatial prior or as an auxiliary prediction target decoupled from diagnosis. We propose GazeWorld, a medical imaging world model that treats the image as the world and the radiologist's fixation sequence as a trajectory through it. GazeWorld autoregressively predicts the latent representation of the next fixated patch from all previously visited ones, while a spatial-completion branch covers unvisited regions. At inference, GazeWorld generates a sequence of patch representations from the image alone without requiring real gaze data. Frozen GazeWorld features achieve state-of-the-art diagnostic accuracy across all nine supervised settings on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax, as well as the highest zero-shot accuracy on all three benchmarks. On the GazeSearch benchmark, a generic decoder trained on the same frozen features outperforms the purpose-built LogitGaze-Med by over 16\% in ScanMatch and 22\% in SED, despite not being explicitly trained to predict gaze. GazeWorld demonstrates that modeling how experts read, not just what they conclude, offers a promising pretraining paradigm for medical imaging AI.


B$^3$-PWL: GPU-Batched Branch-and-Bound for Piecewise-Linear Optimization with SOS2 Constraints

Yilin Guan ⋅ Shuqing Luo ⋅ Pingzhi Li ⋅ Tianlong Chen ⋅ Kaidi Xu

Piecewise-linear (PWL) optimization problems arise in many mixed-integer programming (MIP) optimization applications, including portfolio optimization, workforce scheduling, and resource allocation. But solving them to global optimality remains computationally expensive because branch-and-bound repeatedly solves LP relaxation subproblems. Existing solvers are largely CPU-centric, leaving the scalability of modern GPUs underutilized. Few prior GPU-accelerated branch-and-bound either targets neural network which is not suitable for general PWL optimization, or accelerates only auxiliary subroutines such as strong branching heuristics within CPU-centric MIP solvers. To bridge this gap, we propose $\textbf{\texttt{B$^{3}$-PWL}}$, a GPU-centric batched branch-and-bound framework for piecewise-linear optimization with Special Ordered Set of type 2 (SOS2) constraints. Our method solves batches of LP relaxation subproblems concurrently on the GPU using a first-order primal-dual solver, enabled by a specialized batched block-tiled sparse matrix kernel. To complement bound computation, we further introduce a unified feasibility search module that combines an SOS2 repair primal heuristic with a batched feasibility pump to rapidly obtain feasible incumbents and improve pruning efficiency. On a benchmark of 43 PWL-MIP instances, $\textbf{\texttt{B$^{3}$-PWL}}$ achieves a 9.25$\times$ geometric-mean speedup over NVIDIA cuOpt while reaching high-quality feasible incumbents on every tested instance, demonstrating the potential of first-order LP methods as the central engine of GPU-accelerated branch-and-bound.


Be CARE-ful with Text-to-SQL Benchmarks

Haonan Wang ⋅ Jiaxiang Liu ⋅ Elaine Ang ⋅ Tiancheng Ge ⋅ Wanting You ⋅ Austin S Wijaya ⋅ Reya Vir ⋅ Tianle Zhou ⋅ Mihir Agarwal ⋅ Yusen Zhang ⋅ Eugene Wu

Enterprise text-to-SQL systems powered by LLM-based agents require more than correct final SQL queries to be production-ready. In realistic deployments, an agent may query databases that are changing concurrently, consult access-control and safety policies, respect business rules and integrity constraints, and iteratively refine SQL before producing a final answer. Therefore, beyond returning correct answers, enterprise agents must remain reliable, policy-compliant, rule-compliant, and efficient throughout the SQL-generation trajectory. We introduce CAREFUL, a benchmark and executable evaluation environment for production-oriented text-to-SQL agents. CAREFUL is grounded in 590 practitioner articles, incident reports, standards, and source blocks, from which we identify four CARE criteria: Concurrency Control, Access Control and Safety Policies, Rules and Integrity, and Efficiency. The benchmark contains 1,058 manually annotated tasks across the four CARE dimensions, distilled from practitioner reports and instantiated on databases from Spider 2.0 and BIRD-Interact. CAREFUL was constructed by a nine-person annotation team, including four PhD-level annotators who led canonical scenario design and review and five experienced undergraduate annotators who supported task instantiation and verification, with three stages of rigorous quality control. Experiments with frontier LLM agents show that CAREFUL is challenging: GPT-5.5 obtains only a 33.9% Binary Pass Ratio, where a task passes only if all required operational criteria are satisfied, and a 73.3% Composite Score across the four CARE dimensions. These results show that strong text-to-SQL ability does not directly translate into production-ready database agency. Overall, CAREFUL provides a realistic testbed for evaluating and developing enterprise text-to-SQL agents that are reliable, policy-compliant, rule-compliant, and efficient in operational settings. Code and benchmark artifacts are provided for review and will be released publicly upon acceptance.


BehaviorBench: Modeling Real-World User Decisions from Behavioral Traces

Liangwei Yang ⋅ Jielin Qiu ⋅ Zixiang Chen ⋅ Ming Zhu ⋅ Juntao Tan ⋅ Zhiwei Liu ⋅ Wenting Zhao ⋅ Zhujun Lan ⋅ Akshara Prabhakar ⋅ Silvio Savarese ⋅ Huan Wang ⋅ Shelby Heinecke

Many decision-support systems require adapting to individual users, yet evaluating this capability remains challenging because user preferences and beliefs are rarely stated explicitly. Instead, they are revealed through sequences of real-world behavior. Existing benchmarks for user understanding often rely on textual personas, constructed preferences, or simulated users, which may not faithfully reflect how people actually decide.We introduce \textsc{BehaviorBench}, a benchmark for modeling user decisions from real-world behavioral traces. Built from public prediction-market and on-chain records, the benchmark reconstructs longitudinal decision histories and evaluates whether models can infer how a particular user believes and acts from prior behavior. It defines two complementary task layers: \emph{Belief prediction}, which infers a user's final revealed stance and confidence, and \emph{Trade prediction}, which predicts the direction and magnitude of individual actions. We evaluate frontier and open-weight generative models under multiple history representations, including direct behavioral history, generated user profiles, and retrieved cross-user evidence. Results show that personalization is not a single capability: models benefit differently from different representations depending on the task, and performance remains limited when reasoning over implicit behavioral signals. \textsc{BehaviorBench} provides an evaluation setting for studying personalized decision modeling grounded in real-world behavioral evidence rather than simulated users alone.

Benchmark recovery is often treated as a deployment certificate for low-bit LLMs, but thresholded controller stacks ship a different object: a frozen score-to-action contract. This paper asks a narrower acceptance question: after quantization, does that interface survive independent threshold retuning against frozen FP16 labels, the strongest cheap repair an operator would plausibly try first? We study one fixed four-action controller, one frozen train/cal/report protocol, and four instruction-tuned checkpoints: Llama-3.1-8B/70B-Instruct and Qwen-2.5-7B/72B-Instruct. In the only score-carrying regime, W4A16, GPTQ + -search stays near FP16 on MMLU yet still leaves 11.9/10.9 percentage-point false tool-trigger rates and 16.1/15.4 aggregate FCR on the two large models. Because BFCL directly labels the tool frontier, that failure already establishes the paper's main acceptance claim even if FAQ and the preservation-only retrieve/stop frontiers are ignored. Aggregate FCR is reported only as interface-wide preservation context because retrieve/stop remain preservation-only. FAQ is included only as a bounded reducibility probe under the same frozen audit target, not as a generic new quantizer or the central novelty claim. FAQ lowers tool FTR and aggregate FCR without obvious MMLU or throughput collapse, but that is secondary non-collapse evidence. The contribution is therefore a single-controller deployment-acceptance counterexample under a frozen permission set: on the retained large-model W4A16 rows, benchmark-close MMLU recovery does not certify preserved tool-trigger behavior. The accompanying contribution is the freeze-and- retune audit protocol that makes that failure measurable.


Beyond Confidence: Rethinking Self-Assessments for Performance Prediction in LLMs

Sree Bhattacharyya ⋅ Samarth Khanna ⋅ Leona Chen ⋅ Lucas Craig ⋅ Tharun Dilliraj ⋅ James Wang

Large Language Models (LLMs) are increasingly used in settings where reliable self-assessment is critical. Assessing model reliability has evolved from using probabilistic correctness estimates to, more recently, eliciting verbalized confidence. Confidence, however, has been shown to be an inconsistent and overoptimistic predictor of model correctness. Drawing on cognitive appraisal theory, a framework from human psychology that decomposes self-evaluation into multiple components, we propose a multidimensional perspective on model self-assessment. We elicit six appraisal-based dimensions of self-assessment, alongside confidence, and evaluate their utility for predicting model failure across 12 LLMs and 38 tasks spanning eight domains. We find that competence-related appraisal dimensions, particularly effort and ability, consistently match or outperform confidence across most settings. Effort additionally yields less overoptimistic estimates that remain stable across model sizes. In contrast, affective dimensions provide marginally predictive signals. Furthermore, the most informative dimension varies systematically with task characteristics: effort is most predictive for reasoning-intensive tasks, while ability and confidence dominate on retrieval-oriented tasks. Broadly, our findings indicate that structured multidimensional self-assessment is a promising approach to improving the reliability and safety of language model deployment across diverse real-world settings.


Beyond Trajectory Matching: Reflow with Marginal Distribution Alignment

Chen Wang ⋅ Peiran Yun ⋅ Pan Xie ⋅ Ke Deng

Diffusion and continuous-flow generative models achieve high-quality generation, and their deterministic sampling can be formulated as solving learned ODE dynamics. However, accurate ODE discretization often requires many steps, making efficient few-step generation a key challenge. Among acceleration strategies, reflow-based distillation simplifies teacher ODE trajectories so that a student model can approximate the teacher transport with fewer steps. We identify a theoretical limitation of this paradigm, namely that trajectory matching can under-determine the distribution induced by the student model. In particular, two student models can attain the same trajectory-matching loss while inducing different endpoint marginal distributions, which may lead to different generation quality. To address this limitation, we introduce a marginal-alignment regularizer that penalizes the discrepancy between the student-induced marginal and the corresponding teacher marginal at the endpoint of each distillation interval. The regularizer is computed by tracking log-density changes along the ODE induced by the student model and evaluating scores from the frozen teacher model, without requiring auxiliary trainable networks or adversarial optimization. The resulting framework applies uniformly to the reflow family, including vanilla reflow and piecewise reflow. We further prove a telescoping total-variation bound showing that local marginal alignment controls the final-time discrepancy between the student-induced and teacher-induced distributions. Experiments on benchmark backbones demonstrate the effectiveness of the proposed method for few-step generation.


Beyond What Seems Necessary: Hidden Gains from Scaling Training-Time Reasoning Length under Outcome Supervision

Yihao Xue ⋅ Allan Zhang ⋅ Jianhao Huang ⋅ Amit Sahai ⋅ Baharan Mirzasoleiman

Training LLMs to think and reason for longer has become a key ingredient in building state-of-the-art models that can solve complex problems previously out of reach. Recent efforts pursue this in different ways, such as RL fine-tuning to elicit long CoT or scaling latent reasoning through architectural recurrence. This makes reasoning length an important scaling knob. In this work, we identify a novel phenomenon (both theoretically and experimentally): under outcome-only supervision, out-of-distribution (OOD) performance can continue improving as training-time reasoning length (e.g., the token budget in RL, or the loop count in looped Transformers) increases, even after in-distribution (ID) performance has saturated. This suggests that robustness may require a larger budget than ID validation alone would indicate. We provide theoretical explanations via two mechanisms: (i) self-iteration can induce a stronger inductive bias in the hypothesis class, reshaping ID-optimal solutions in ways that improve OOD generalization; and (ii) when shortcut solutions that work for ID samples but not for OOD samples persist in the hypothesis class, regularization can reduce the learned solution's reliance on these shortcuts as the number of self-iterations increases. We complement the theory with empirical evidence from two realizations of scaling training-time reasoning length: increasing the number of loops in looped Transformers on a synthetic task, and increasing token budgets during RL fine-tuning of LLMs on mathematical reasoning, coding, and cross-representation tasks.


BiMoGen: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion

Wanjiang Weng ⋅ Yongliang Wu ⋅ Xiaofeng Tan ⋅ Xingyu Zhu ⋅ Wenbo Zhu ⋅ Hongsong Wang

Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate competitive performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs.


BLEND: Balancing Personalization vs. Generalization in Federated Vision–Language Models

Payam Abdisarabshali ⋅ Fardis Nadimi ⋅ Kasra Borazjani ⋅ Naji Khosravan ⋅ Richeng Jin ⋅ Seyyedali Hosseinalipour

Federated vision–language models (FedVLMs) represent one of the emerging frontiers of machine learning, aiming to integrate federated learning within the fine-tuning pipelines of vision–language models (VLMs). In this work, we introduce BLEND, a new framework designed to jointly enhance the personalization and generalization capabilities of FedVLMs. BLEND proposes a selective VLM parameter fine-tuning and aggregation strategy that consists of (i) global vision and text adapters shared across clients, and (ii) a personalized vision projection adapter tailored to each client. BLEND further introduces a novel loss function that blends the outputs of the global and personalized adapters to achieve personalization, while simultaneously promoting generalization through a principled anchor design implemented via the projection heads of the zero-shot VLM. Extensive experiments demonstrate that {BLEND} consistently outperforms state-of-the-art baselines, highlighting its ability to effectively balance personalization and generalization across diverse system settings. (Code is available at: https://github.com/blendfedvlm/BLEND.git)


Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning

Jiazheng Li ⋅ Chi-Hao Wu ⋅ Yunze Liu ⋅ Kaize Ding ⋅ Jundong Li ⋅ Chuxu Zhang

Understanding ultra-long videos such as egocentric recordings, live streams, or surveillance footage spanning days to weeks, remains a challenge. For current multimodal LLMs: even with million-token context windows, frame budgets cover only tens of minutes of densely sampled video, and most evidence is discarded before inference begins. Memory-augmented and agentic approaches help with scale, but their retrieval remains fragmented across modalities and lacks long-range narrative summaries that span days or weeks. We propose MAGIC-Video, a training-free framework built around a multimodal memory graph with interleaved narrative chain: the graph unifies episodic, semantic, and visual content through six typed edges and supports cross-modal retrieval, while the chain distils long-horizon entity biographies and recurring activity events. At inference time, an agentic loop interleaves graph retrieval with narrative fact injection, covering both the modality and time dimensions of ultra-long video in a single retrieval pipeline. On EgoLifeQA, Ego-R1 and MM-Lifelong, MAGIC-Video consistently outperforms strong general-purpose, long-video, and agentic baselines, with gains of 10.1, 7.4, and 5.9 points over the prior best agentic system on each benchmark.


Bridging Simulation and Reality: Geometry and Decision Alignment for Autonomous Driving

Dapeng Zhang ⋅ Zhenlong Yuan ⋅ Chenyang Li ⋅ Hongtao Nie ⋅ Yinda Chen ⋅ Rui Zhou ⋅ Peng Zhi ⋅ Qingguo Zhou

Bridging the gap between simulation and reality remains a fundamental challenge for end-to-end autonomous driving. Existing approaches primarily focus on appearance-level features. This often leads to suboptimal transfer, where visually aligned models still produce inconsistent or unsafe behaviors in the real world. In this paper, we propose a unified domain adaptation framework that jointly aligns perception and decision processes through geometry-aware and vector-space representations. At the perception level, we introduce a Geometry-Aware Perception Alignment (GAPA) module that enforces cross-domain consistency in both explicit geometry and implicit structural representations extracted from latent features. This reduces domain discrepancy by targeting scene structure rather than appearance. At the decision level, we propose a Latent Vector Space Guidance Decision Alignment (LVDA) module that aligns vector-conditioned policy distributions. By modeling driving decisions as conditioned on structured representations of map topology and multi-agent interactions, we minimize discrepancies between source and target action distributions via adversarial and structural alignment. To further enhance stability, we introduce a progressive adversarial transfer strategy that improves cross-domain feature alignment and training stability. Extensive experiments on a newly partitioned nuScenes dataset demonstrate that Bridge-AD achieves excellent performance, effectively narrowing the gap between simulation and real-world autonomous driving.


Capturing In-Context Learning Dynamics with Task Operators

Guangzhi Xiong ⋅ Zhenghao He ⋅ Bohan Liu ⋅ Sanchit Sinha ⋅ Wenqian Ye ⋅ Aidong Zhang

In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requires processing the full set of examples, resulting in inefficient deployments, and how ICL works mechanistically is not fully understood. Prior work compresses ICL into fixed activation vectors extracted from specific layers or positions, but these input-independent interventions fail on complex tasks where the output depends on fine-grained interactions with the input. By analyzing the ICL forward pass, we show that each attention head's output is an affine transformation of its context-masked counterpart, and that the parameters of this transformation are empirically stable across samples for a given task. Building on this, we introduce Task Operator (TO), which replays this transformation as an analytically derived update to the attention output projection. Across lexical, algorithmic, and reasoning tasks, TO achieves the best overall performance among prior methods and substantially narrows the gap between zero-shot inference and ICL. We further show that the extracted knowledge concentrates in a task-specific sparse circuit across layers and positions, and that averaging operators from disjoint demonstration batches enables effective many-shot scaling without expanding the context window.


Causal Discovery from Unseen Environments

Sophia Xiao ⋅ Bijan Mazaheri

Existing causal discovery methods for interventional data typically follow one of two paradigms: they either require explicit environment labels and known intervention targets, or, in the absence of such metadata, they rely on computationally intensive procedures such as first inferring intervention targets or employing complex models to account for the latent variables of environment labels. Surprisingly, we show that neither explicit environment labels nor auxiliary inference algorithms are strictly necessary: under certain conditions, the classical ICA-LiNGAM algorithm consistently recovers the causal graph from simply pooled interventional and observational data. We first demonstrate that in linear Gaussian structural equation models where interventions independently perturb the noise distributions across nodes, pooling produces exactly independent, non-Gaussian sources, the appropriate setting for Independent Component Analysis (ICA). Because the assumption of strictly independent interventions can be overly restrictive, we extend our analysis to the practical setting of single-node (atomic) interventions. While pooling atomic interventions induces the necessary non-Gaussianity, it also introduces source *dependence*, violating a core ICA assumption. However, theoretical analysis establishes that this dependence decays as $\mathcal{O}(1/n^2)$ with the number of variables $n$, whereas non-Gaussianity decays only as $\mathcal{O}(1/n)$. This yields a regime of mild misspecification at moderate $n$ where sources are near-independent yet sufficiently non-Gaussian for ICA to succeed. In synthetic experiments, where data is pooled across many different environments, our label-free approach matches or exceeds label-aware methods. On real-world data, our approach performs comparably to methods that explicitly require environment labels. Ultimately, our work challenges the assumed necessity of environment labels and complex inference mechanisms in interventional causal discovery and suggests a re-evaluation of the often-overlooked classical ICA-LiNGAM algorithm.


CFC26: Building Evaluations for Deployment in Sonar-Based Fish Counting

Madison Van Horn ⋅ Suzanne Stathatos ⋅ Sevan Brodjian ⋅ Justin Kay ⋅ Kai Van Brunt ⋅ Erik Young ⋅ Sara Beery ⋅ Pietro Perona ⋅ Michael Hobley

Accurate counts of fish populations are essential for monitoring ecosystems and making decisions about conservation, harvest limits, and river management. In many field deployments, these counts are obtained from sonar videos that experts must inspect manually, making large-scale monitoring slow, expensive, and difficult to reproduce. Automated counting systems have been proposed as a scalable alternative, but so far this promise has been difficult to realize in deployment. Developing evaluations that better reflect real world deployment conditions will close this gap, and drive improvements to counting methods that translate to reality. To achieve this, we propose Caltech Fish Counting 2026 (CFC26), a new dataset with improved evaluation protocols that simulate deployment to current and new rivers, testing in-distribution and out-of-distribution performance. We expand the CFC22 dataset to eleven deployment locations across nine Pacific salmon river systems from Alaska to Northern California, capturing a broader geographical and ecological range of data. We additionally introduce two new evaluation metrics for counting, nMNE and nMANE, that mirror the operational workflow that experts use and help surface errors that frame-level detection metrics do not capture. Using these metrics, we reveal previously-concealed directional bias in upstream vs downstream counts, which would have direct and significant ecological implications if undetected.


Characterizing Underrepresentation in Generalizing Causal Survival Estimates

Bolun Liu ⋅ Sean McGrath ⋅ Yiren Hou ⋅ Elizabeth Stuart ⋅ Harsh Parikh

Randomized trial findings are routinely generalized to broader target populations to inform decisions in medicine and public policy. When the trial sample is misaligned with the target population, generalized effect estimates can yield suboptimal decisions. We identify and characterize the target subgroups that are underrepresented in the trial cohort in time-to-event settings. Here, underrepresentation arises in two distinct ways: baseline underrepresentation, when a subgroup's covariate profile is rare in the trial relative to the target, and time-varying underrepresentation, when differential censoring erodes a subgroup over follow-up. Existing diagnostics either conflate these mechanisms or address only the baseline case. We show that the asymptotic variance of the efficient generalizability estimator admits a decomposition in which only two terms can diverge as representation weakens: a sampling-weight term capturing baseline misalignment and a censoring-weight term capturing differential loss to follow-up. Underrepresented subgroups are therefore precisely those the trial estimates with poor precision, and the decomposition reveals which channel is responsible.We exploit this operationally in $\textbf{TRACE}$ ($\textbf{T}$emporal $\textbf{R}$ashomon $\textbf{A}$nalysis of $\textbf{CE}$nsored data), which combines a self-normalized variance objective with a tree-based search to localize underrepresented subgroups, attribute each to baseline mismatch or censoring, and pinpoint the follow-up intervals where censoring drives the variance inflation; the joint problem is solved by a parametric dynamic program nested inside a Rashomon ensemble of near-optimal trees. On synthetic benchmarks and a case study generalizing the National Job Training Partnership Act (JTPA) Study to U.S. JTPA-eligible adults, the framework recovers underrepresented regions, correctly attributes each mechanism, and identifies the follow-up windows driving temporal underrepresentation; refining the trial and target samples to reliably generalizable subgroups substantially reduces generalized effect-estimate variance.


Classification at the Edge of Stability: Unifying Self-Stabilization and Convergence Rates

Salma Tarmoun ⋅ Lachlan MacDonald ⋅ Ziqing Xu ⋅ Konstantinos Emmanouilidis ⋅ Rene Vidal

Neural networks are often trained with gradient descent (GD) at the edge of stability (EoS), where step sizes exceed classical stability thresholds and standard convergence theory no longer applies. Despite exhibiting non-monotonic loss and oscillatory dynamics, this regime frequently leads to faster convergence and better generalization. Existing theoretical perspectives offer complementary but incomplete insights: one identifies a self-stabilization mechanism that drives GD toward flat regions but does not yield convergence rates; another derives rigorous rates for specific losses but does not explain the underlying stabilization mechanism; and a third unifies these views for overparameterized least squares, but relies on the existence of a minimizer manifold, which is absent in classification settings. In this work, we extend this unified geometric perspective to exponential-tail losses, including logistic and cross-entropy losses, where no finite minimizer exists. We identify a curvature-defined reference subspace that replaces the role of the minimizer manifold, yielding a coupled dynamical system in which orthogonal oscillations are damped by parallel progress along directions of decreasing sharpness. Under regularity conditions, which we verify for logistic and multiclass cross-entropy losses in symmetric data models, we derive explicit bounds for the iterates, not just the loss, covering both the EoS and stable regimes, and revealing a three-phase damping structure in the latter. Our analysis further uncovers a transient implicit bias induced by large step sizes: although GD converges asymptotically to the max-margin direction, large steps can leave persistent orthogonal residuals and, in certain geometries, amplify them.


CLR-voyance : Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics

Aishik Nagar ⋅ Arun-Kumar Kaliya-Perumal ⋅ Yu-Hsuan Han ⋅ Andrew Sheng-Han Huang ⋅ Kristen Kee ⋅ Yushi Cao ⋅ Yiming Chen ⋅ Hongchao Jiang

Inpatient clinical reasoning is a sequential decision under partial observability: the clinician sees the admission so far and must choose the next action whose downstream consequences are not yet visible. Existing clinical-LLM evaluations and reinforcement-learning (RL) rewards collapse this structure into closed-form retrieval, full clinical journey leakage, or unanchored LLM-as-judge scoring. Here, we introduce CLR-voyance, an end-to-end framework that reformulates inpatient reasoning as a Partially Observable Markov Decision Process (POMDP) and supervises it with rewards that are simultaneously outcome-grounded and clinician-validated. We instantiate the formulation as CLR-POMDP, which partitions successful patient journeys into a policy-visible past and an oracle-only future. Using the past information, an oracle LLM generates a case-specific query-answer pair, and the first adaptive rubric for clinical reasoning which is verifiable in the future of the patient journey. These rubrics are used for both RL post-training and evaluation of models for inpatient clinical reasoning. CLR-voyance post-trains Qwen3-8B and MedGemma-4B with GRPO followed by model merging, yielding state-of-the-art inpatient clinical reasoning while retaining generalist capabilities. CLR-voyance-8B achieves 84.91% on CLR-POMDP, ahead of frontier medical reasoning models like GPT-5 (77.83%) and MedGemma-27B (66.66%) and has comparable or better performance on existing medical benchmarks. To ensure that our setting is clinically meaningful, we conduct a large-scale clinician alignment study, where physicians curate per-case rubrics, grade candidate responses against them, and provide blinded pairwise preferences of model reasoning. This study provides insights on clinical LLM-as-a-judge and clinical preference-model selection, which can inform the community at large. CLR-voyance has been deployed for 6+ months at a partner public hospital, drafting thousands of reasoning-heavy inpatient notes and 1.03T+ tokens processed to date.

Educational LLM systems generate assessment questions on demand and depend on the assumption that a model can be controlled to a specified cognitive level. We test the assumption. CogBench is a benchmark for cognitive-level control in LLM question generation, grounded in the Revised Bloom's Taxonomy. Its central design choice is a two-mode protocol that separates cognition from imitation: in standard mode the model is asked to generate at level Lt; in adversarial mode it must do so while restricted to verbs from a different level Lv, demonstrating the requested level structurally rather than lexically. CogBench spans 8 subjects, 6 Bloom levels, and 120 CC-BY OpenStax passages, with ~17,000 graded generations across 13 models (six frontier closed APIs and seven open-source local models), each scored by a deliberately dual evaluator: a 28-rule constraint checker, and the Cognitive Complexity Score (CCS), a fine-tuned BERT classifier at 82.0% exact accuracy that outperforms a four-model LLM-as-judge panel by 7.8 points. We find that (1) adversarial vocabulary collapses strict adherence by 24-56 pp across the 12 models above floor, with the strongest standard performer suffering the largest collapse; (2) Remember-level adherence under adversarial vocabulary is essentially 0% in all 13 models, suggesting Remember lives almost entirely in its surface verbs; (3) the two evaluators rank models differently, including on the leader, a divergence we argue benchmarks should report rather than hide; and (4) the most adversarially robust models on CCS are not frontier flagships but gpt-4o-mini (46.8%) and llama3.1-8b (45.6%). What current LLMs offer educational systems looks like cognitive control and is largely verb-level imitation that resembles it. We release the dataset (CC-BY 4.0; Croissant 1.0 with Responsible-AI fields), the evaluation harness, and a 13-model leaderboard at https://huggingface.co/spaces/cogbench-anonymous/cogbench-leaderboard.


ColdDDI: Evaluating Knowledge Utilization in Cold-Start Drug-Drug Interaction Prediction

Jiheng Liang ⋅ Chen Zhao ⋅ Di Wu ⋅ Chenyang Bu ⋅ Yunpeng Hong ⋅ Xingquan Zhu ⋅ Yi He

Cold-start drug-drug interaction (DDI) prediction tests whether models can identify clinically significant interactions for drugs without training-time interaction history. Existing benchmarks mostly report aggregate edge-prediction scores, leaving a key evaluation question unanswered: when models receive molecular, textual, or knowledge-graph (KG) evidence, do they actually use the evidence that pharmacologically supports the interaction? We introduce ColdDDI, a reconstructible evaluation benchmark built from the latest DrugBank 5.1.13, with 1,900 approved small-molecule drugs and 565,731 positive DDI pairs. ColdDDI evaluates three levels of drug novelty, namely test pairs whose two drugs were both seen during training, pairs with one unseen drug, and pairs with two unseen drugs. It also annotates each interaction by whether it changes drug exposure or drug effect, and by whether the biomedical knowledge graph contains shared enzymes, transporters, or targets that can plausibly mediate the interaction. These annotations allow ColdDDI to separate two factors that aggregate metrics conflate, namely whether mechanistic evidence is available, and whether a model prediction depends on that evidence. We evaluate eight conventional DDI methods and 13 LLMs; for open-weight LLMs, we test five prompt patterns and use masking, drug replacement, and channel-sensitivity metrics to probe knowledge utilization. ColdDDI exposes that, in the hardest split where both drugs are unseen, the main performance divide is mediator availability. A fine-tuned 1B LLM recovers 89-93% of interactions with a shared enzyme, transporter, or target, but only 40-62% without such a mediator. More importantly, KG-provided evidence is not always used; several KG-augmented baselines change little when the shared mediator is masked or disrupted, whereas fine-tuned LLMs respond strongly to this intervention. Thus, ColdDDI evaluates knowledge utilization rather than knowledge access alone, showing where cold-start DDI models rely on mechanistic evidence and where they fail despite receiving it. We release derivative annotations, splits, prompts, diagnostics, evaluation code, and a DrugBank reconstruction pipeline.


Compositional Policy Optimization with Language Models

Wensen Mao ⋅ Yuanlin Duan ⋅ He Zhu

Large language models (LLMs) possess remarkable ability to understand natural language descriptions of complex robotics environments. Earlier studies have shown that LLM agents can use a predefined set of skills for robot planning in long-horizon tasks. However, the requirement of prior knowledge of the skill set required for a given task constrains its applicability and flexibility. We present a novel approach L2S (short for Language2Subtasks) to leverage the generalization capabilities of LLMs to decompose the description of a complex task in natural language into definitions of reusable subtasks. Each subtask is defined by an LLM-generated dense reward function and a termination condition, which in turn lead to effective subtask training and chaining. However, LLMs lack detailed insight into the specific low-level control intricacies of the environment, such as threshold parameters within the generated reward and termination functions. To address this uncertainty, L2S (1) enables LLM reflection feedback loop to improve task decomposition and code generation, and (2) trains parameter-conditioned subtask policies that perform well in a broad spectrum of parameter values. As the impact of these parameters for one subtask on the overall task becomes apparent only when its following subtasks are trained, L2S selects the most suitable parameter value during the training of the subsequent subtasks to effectively mitigate the risk associated with incorrect parameter choices. During training, L2S autonomously accumulates a subtask library from continuously presented tasks and their descriptions, using guidance from the LLM agent to effectively apply this subtask library in tackling novel tasks. Our experimental results show that L2S is capable of generating reusable subtasks to solve a wide range of robot manipulation tasks.

Compositional reasoning is critical for real-world problem solving: since training data is necessarily limited, models must generalize by composing learned skills in new ways. While post-training methods such as reinforcement learning (RL) have substantially improved the reasoning abilities of language models (LMs), their effects on compositional reasoning remain less well understood. We propose a dependency-graph framework to formalize compositional reasoning, yielding three levels of compositionality with increasing complexity. Empirically, we instantiate this framework with data-structure tasks, which provide deterministic reward computation and clear compositional structure. We find a consistent decomposed-to-composed asymmetry: decomposed-skill training does not reliably transfer to composed tasks, whereas composed-task training transfers more readily back to decomposed tasks. We provide theoretical explanation for this asymmetry, and further evaluate compositional generalization under length extrapolation, structural distribution shift, and transfer to tasks requiring unseen skills. Finally, we present a pilot study on real-world tool-calling benchmarks, showing preliminary evidence that the decomposed-to-composed asymmetry can extend to practical settings.


Computational Dynamic Mechanism Design

Sadie Zhao ⋅ Amy Greenwald ⋅ Yanchen Jiang ⋅ Denizalp Goktas ⋅ Yiling Chen

Dynamic mechanism design provides a principled framework for coordinating self-interested agents whose private information evolves over time, but history dependence and dynamic incentive constraints make general analytical solutions difficult. Using a dynamic revelation principle, we formulate optimal dynamic mechanism design as a bilevel optimization problem over parameterized direct mechanisms: the principal optimizes the mechanism, while agent-side unilateral deviation problems verify incentive compatibility. We model each agent’s problem as a parameterized POMDP, and use stochastic natural policy gradient to solve these POMDPs with finite-sample error bounds. These policies serve as oracle outputs for the principal’s problem, yielding first-order optimization procedures with convergence guarantees that approximate stationary solutions under bounded oracle error. This formulation separates two representation choices: principal-side history compression and agent-side sufficient statistics. Experiments in dynamic bandit and budget-constrained auctions show that the framework can recover known benchmarks and learn approximately incentive-compatible mechanisms beyond analytical settings. Furthermore, they reveal how representations affect optimization.

This position paper argues that computer science conferences should require tamper-evident, nonrepudiable attestations of experimental results. We name the underlying problem experiment nonrepudiation: a compliant protocol must bind the numbers in a paper to an actual executed computation in a way the author cannot later alter or deny. The current system relies on self-reported checklists, optional code sharing, and author-controlled logging. None of these mechanisms answer the question a reviewer cannot check: did the code the paper describes produce the numbers the paper reports? We define the problem formally, state the security properties any compliant protocol must satisfy, and describe a threat model that includes attacks current approaches do not prevent. To show that the problem is solvable, we built K-Veritas, a reference implementation in Go that produces signed reports without accessing training data. K-Veritas is a testbed, not a finished answer. We call on conferences and the community to treat nonrepudiation as a first-class requirement and to help build an open, independent standard for it.

Motivated by interpretable generative models for applications such as scientific inquiry, we aim to learn representations that localize human-interpretable concepts and separate them from residual variation in the data. Reconstruction and prediction objectives alone are insufficient for this purpose, as concept-relevant variation can be distributed across multiple latent factors while still achieving high performance. We address this limitation with a generative model that learns a balanced decomposition of data into concept-aligned and residual latent representations inferred purely from observations. The objective encourages localization through an upper bound on the concept information bottleneck while reducing dependence between the concept and residual latent. Across synthetic and real datasets, the learned representations localize concepts while preserving generative fidelity. A neural-data case study shows that localized representations align more strongly with structured variation in observed neural responses. This supports localization, rather than prediction accuracy alone, as a more faithful criterion for interpreting concept-relevant structure in generative models.


Conditional Optimal Bridge for Riemannian Activation Steering

Seyed Arshan Dalili ⋅ Ajay Narayanan Sridhar ⋅ Vijaykrishnan Narayanan ⋅ Mehrdad Mahdavi

Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time. While many existing methods implicitly optimize a log-density-ratio objective between desired and undesired activation distributions, they do so heuristically rather than deriving it from a principled optimization problem. Moreover, these methods produce query-independent steering directions that can degrade performance on both in-distribution and out-of-distribution (OOD) inputs. We introduce COBRAS (Conditional Optimal Bridge for Riemannian Activation Steering), which addresses both limitations by casting activation steering as a Schr\"{o}dinger Bridge on the residual-stream hypersphere. This formulation yields, to our knowledge, the first principled derivation of the log-density-ratio steering objective from a well-posed optimization problem. Solving the bridge via entropic optimal transport and extracting the probability flow ODE recovers the widely used density-ratio gradient as a special case when the Sinkhorn potentials are uniform. Crucially, the Schr\"{o}dinger potentials are evaluated at the current activation, making the resulting steering direction inherently query-adaptive. Empirically, across four models and three alignment axes (helpfulness, truthfulness, and detoxification), COBRAS consistently outperforms prior activation steering baselines while avoiding the OOD degradation commonly observed in existing methods.

A local specialist LLM, fine-tuned with reinforcement learning from verifiable rewards (RLVR) on operator-local data, is installed inside a single regulated organization under a per-deployment error budget $\alpha$. The operator needs a safety certificate that holds on *this* deployment's stream, simultaneously at every wall-clock round: no pooling across deployments, no waiting for a long-run average. Existing wrappers cannot deliver this on these adaptive, online-updated streams: offline conformal-risk methods require exchangeability; online-conformal methods bound only long-run averages; non-exchangeable extensions are marginally valid; and the closest anytime wrapper, A-RCPS, controls marginal rather than selective risk. Through a (test statistic, validity guarantee, deployment rule) framework we identify one empty cell forced by the deployment requirements (e-process per threshold, selective risk, anytime-pathwise validity, max-certified-threshold rule); **Conformal Selective Acting** (CSA) fills it as a per-round wrapper maintaining a Ville-type e-process per candidate score threshold on a Bonferroni grid, evaluated against the RLVR filtration. Under predictable updates and isotonic-calibrated monotone risk we prove (i) an anytime-pathwise selective-risk bound $R_T^{\mathrm{act}} \le \alpha + O(N_T^{-1/2})$, (ii) rate-optimal certification matching $\Theta(\bar\eta^{-2} \log(1/\delta))$, and (iii) a horizon-independent release-rate gap. Across eight specialist benchmarks (480 streams), sixteen adversarial distribution-shift cells (160 streams), and five live Expert-Iteration RLVR cells with online LoRA over four base models in three architecture families (10,300 rounds), CSA is the only method among ten directly compared that satisfies both pathwise validity and non-refusing deployment on every cell. We do not propose a new LLM, training algorithm, or policy class; CSA is the deployment-side complement, orthogonal to the model itself, for operators who cannot use a frontier API.


Contextualized Evaluation of Vision Language Models through Dynamic Interviews

Yijiang Li ⋅ Huiqi Zou ⋅ Bingyang Wang ⋅ Ziang Xiao

Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertain. This gap stems from the fundamental misalignment between benchmarks in controlled, static settings and the dynamic, interactive, and contextualized nature of real-world applications. To bridge this gap, we propose CEDI (Contextualized Evaluations of MLLMs through a Dynamic multi-round Interview, a framework that recasts evaluation as a three-party interaction between an evaluatee model, an automated examiner, and a grader. The examiner conducts multi-turn, semi-structured interviews guided by a graph-based representation of the task. By navigating state-space transitions, CEDI deploys diverse strategies—from clarification requests to adversarial probes—to elicit performance evidence.


Contrastive Nonmyopic Objective Cost-Tradeoff Acquisition for Longitudinal Data

Dzung Dinh ⋅ Boqi Chen ⋅ Yunni Qu ⋅ Marc Niethammer ⋅ Junier Oliva

In many critical domains, features are not freely available at inference time: each measurement may come with a cost of time, money, and risk. Longitudinal prediction further complicates this setting because both features and labels evolve over time, and missing measurements at earlier timepoints may become permanently unavailable. We propose CAPE (Contrastive Acquisition-Plan Embeddings), a partitioning acquisition policy for longitudinal active feature acquisition. CAPE first defines NOCT, an oracle-based objective that scores a set of future feature-time acquisitions by its expected predictive loss together with its acquisition cost. Based on NOCT, CAPE learns a contrastive acquisition-utility manifold of masked partial observations, where proximity reflects the expected utility of future acquisitions. Offline, training states are embedded and partitioned at each rollout timepoint; candidate future plans are scored within each partition using plug-in NOCT, and the best partition-level acquisition plans are cached. During inference, CAPE retrieves the cached plan via nearest-centroid lookup, enabling low-latency adaptive acquisitions. Experiments on synthetic and real-world healthcare datasets demonstrate that CAPE outperforms nonparametric, greedy, and RL-based alternatives, achieving higher accuracy at lower acquisition costs.

Chain-of-thought (CoT) monitoring is increasingly relied upon to detect misbehavior in LLMs, under the implicit assumption that misbehaving actors must reason about their misbehavior, leaving detectable traces in the CoT. We show this assumption fails for a class of attacks we call planning-execution decoupled attacks, in which misbehavior is injected via a corrupted plan before the actor begins reasoning. Using the investigator agent framework of Li et al. [2025], we discover that exposing an actor to a corrupted reasoning plan upstream causes it to reproduce the flawed reasoning in its own CoT, embedding misbehavior into natural-looking reasoning with few suspicious traces. The attack scales to stronger reasoning models and harder tasks, steering them toward misbehavior while evading detection. Evaluating a suite of thinking and non-thinking monitors, we find that thinking monitors substantially outperform non-thinking ones, though even the best miss a meaningful fraction of attacks. Critically, the relationship between monitor thinking budget and detection is not monotonic: extra reasoning sometimes improves detection but can also hurt it when monitors talk themselves into accepting corrupted CoTs as benign. This refines Guan et al. [2025], who show thinking budget generally helps monitoring; we find detection also depends on whether the additional reasoning is directed toward critical evaluation rather than rationalization.

We study finite-horizon two-player zero-sum perfect-recall partial-information games and ask for the minimal finite-dimensional state sufficient to answer all rooted observation-measurable unilateral continuation queries. For each player $i$ and information set $I$, we consider the span $\mathcal L_i(I)$ of feasible counterfactual beliefs and the effective local query space $\mathcal E_i(I)$ generated by rooted continuation queries restricted to $\mathcal L_i(I)$. We prove that its dimension $d_i(I)$ is simultaneously the rank of the local counterfactual query operator, the minimum core-query basis size, the minimum exact local realization dimension, and the dimension of the canonical quotient $\mathcal L_i(I)/\mathcal N_i(I)$, yielding reduced counterfactual predictive-state representations, unique up to similarity, with exact first-return update operators. Aggregating the local ranks defines the deviation-conditioned predictive-state dimensions $d_t^i$, $d_t$, $d_\star$, and $D_\star$. We then prove that local rooted payoff-query control transfers to global exploitability, and under local design richness and bounded one-step parameter norms obtain a rollout-based upper bound in these minimal coordinates. We complement this with a minimax lower bound showing that $\Omega(d/\gamma^2)$ local rollouts are necessary even when the peak DCPSD equals $d$, and with an exponential separation theorem exhibiting games with $d_\star=m$ but observable continuation-trace complexity $2^m$. Finally, we show that DCPSD recovers finite-rank predictive-state, observable-equivalence, and linearly parameterized continuation-law models as special cases.


Coupled Integral PINN for Discontinuity

Yeping Wang ⋅ Shihao Yang

Physics-Informed Neural Networks (PINNs) solve forward PDEs by minimizing residual losses from the governing equations with initial and boundary conditions, but they often struggle with discontinuities such as shocks. In contrast, finite volume methods (FVM) handle discontinuities by enforcing integral conservation, which admits weak solutions. Motivated by this, we propose a \emph{Coupled Integral PINN (CI-PINN)} that augments a standard PINN with an auxiliary network for integral potentials and coupled integral constraints. This improves robustness near shocks while avoiding meshing and the numerical flux integration/reconstruction used in classical schemes. We validate CI-PINN on forward benchmarks including Burgers, Buckley--Leverett, the Euler system, and the Shallow-Water equations.


CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

Mingzhi Zhu ⋅ Michele Merler ⋅ Raju Pavuluri ⋅ Stacy Patterson

Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these capabilities are complementary but misaligned. The Instruct model is concise and tool-disciplined, whereas the Thinking model offers stronger planning and recovery behavior but often over-deliberates and degrades agent performance. We present CRANE (Constrained Reasoning Injection for Code Agents via Nullspace Editing), a training-free parameter-editing method that treats the Thinking–Instruct delta as a directional pool of candidate reasoning edits for the Instruct backbone. CRANE combines magnitude thresholding to denoise the delta, a Conservative Taylor Gate to retain edits that are jointly beneficial for reasoning transfer and tool-use preservation, and Graduated Sigmoidal Projection to suppress format-critical update directions. By merging paired Instruct and Thinking checkpoints, CRANE delivers strong gains over either individual model while preserving Instruct-level efficiency: on Roo-Eval it achieves pass@1 of 66.2% (+19.5%) for Qwen3-30B-A3B and 81.5% (+8.7%) for Qwen3-Next-80B-A3B; on SWE-bench-Verified it resolves up to 14 additional instances at both scales (122/500 and 180/500); and on Terminal-Bench v2 it improves pass@1/pass@5 by up to 2.3%/7.8%, reaching 7.6%/17.9% and 14.8%/30.3%, respectively, consistently outperforming alternative merging strategies across all three benchmarks.


CSO-LLM: Class Subspace Orthogonalization for Post-Training Backdoor Detection and Trigger Inversion in LLMs

Zhengxing Li ⋅ David J Miller ⋅ Guangmingmei Yang ⋅ George Kesidis

While post-training backdoor detection and trigger inversion schemes have been developed for AIs used e.g. for images, there is a paucity of such methods for LLMs. First, the LLM input space is discrete, with up to $150{,}000^k$ $k$-tuples to consider with $k$ the token-length of a putative trigger. Second, one must blacklist tokens typical of the putative target response (class) of an attack, as such tokens may give false detection signals. However, a comprehensive blacklist is not available, in general, for a given domain. We develop a highly effective detection and inversion framework for LLMs treated as classifiers. Central to our approach is class subspace orthogonalization (CSO), a novel plug-and-play paradigm for backdoor detection that serves two fundamental roles when applied to LLMs: (i) it enhances both sensitivity and specificity of a baseline detector; (ii) it provides a form of implicit blacklisting, as it penalizes against inclusion, in a candidate trigger, of tokens that induce signal perturbations "in the direction of" the putative target class of an attack. One version of our detector performs continuous optimization in token embedding space, while a companion trigger-inversion and detection method performs greedy accretion in discrete token space. Our methods give both strong detection performance and accurate inversion of ground-truth triggers on several LLM classification domains, and for several different LLM architectures.


CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning

Ke Sun ⋅ Yizhou Zhao ⋅ Jiayi Xin ⋅ Qi Long ⋅ Weijie Su

Context or prompt-level reweighting has emerged as a central algorithmic lever in Reinforcement Learning with Verified Rewards (RLVR) for improving the reasoning capability of large language models, yet the principle determining what constitutes an optimal weighting remains poorly understood. We address this gap by formulating prompt reweighting as a functional derivative of a utility functional defined in the pass-rate function space, yielding a unified optimality framework that accommodates existing schemes, including REINFORCE and GRPO. Building on this optimality framework, we propose a distribution-aware prompt reweighting approach, called CurveRL, based on a quantile coordinate transform, in which the weight assigned to each prompt depends not on the absolute value of pass rates but on its rank and density to reflect the distributional structure of the pass rates in the learning dynamics. Extensive experiments across multiple benchmarks demonstrate that our proposed CurveRL consistently outperforms GRPO and other RLVR baselines. Our study identifies context-distribution control as a principled axis for analyzing and designing prompt-reweighted RLVR algorithms.

Irregularly sampled multivariate time series arise in many real-world domains such as healthcare, climate science, and large-scale monitoring systems. Characterized by asynchronous observations and heterogeneity across variables, such time series pose significant challenges for forecasting. In this work, we introduce Cross-Variable Temporal Attention (CVTA), a short-term, continuous-time forecasting framework that learns prediction functions from a compact, interpretable Markov state summarizing recent temporal dynamics across variables. CVTA uses Variable-Conditioned Embeddings to project heterogeneous variables into variable-specific latent spaces, and applies temporal attention to capture cross-variable dependencies from this state. Furthermore, we demonstrate that CVTA can be integrated as an auxiliary module into existing state-of-the-art prediction models through a dynamic meta-decision model. Extensive experiments on multiple real-world benchmarks show that CVTA achieves strong short-term predictive accuracy, supports efficient online inference, improves existing prediction models when used as an auxiliary module, and learns attention patterns that align well with domain knowledge.


DDBench: A Benchmark for Agentic Debugging on Distributed Systems

Yibo Yan ⋅ Huijuan Wang ⋅ Junzhou He ⋅ Yizhuo Liang ⋅ Shaoyu Wang ⋅ Huanchen Sun ⋅ Seo Jin Park

LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench Verified. Distributed-system debugging, however, remains an under-explored regime: bugs span processes, nodes, and protocol interactions, with root causes rarely recoverable from source alone and brute-force exploration intractable across non-deterministic interleavings. This leaves two gaps in LLM and agent evaluation: no code-repair benchmark targets distributed-system bugs, and no controlled study isolates how much externally provided debugging context changes agent success on them. We introduce DDBench, a code-repair benchmark of 60 historical bugs mined from 13 open-source distributed systems, partitioned into three difficulty tiers. DDbench evaluates every case under two matched conditions: a symptom-only condition where the agent receives only the bug symptom and repository, and a context-augmented condition where it additionally receives a bounded debugging context (logs, traces, runtime state, and targeted code-investigation notes), isolating the effect of debugging context from model capability. The evaluation of ten LLMs on DDBench reveals several findings. First, distributed debugging exercises a reasoning dimension that single-process benchmarks do not surface: models’ pass rates span 61 pp, and pairwise bootstrap separates 9 of 15 top-tier model pairs at p < 0.05 on DDBench's hardest case-set. Second, bounded debugging context lifts aggregate pass rate by +18.1 pp, and the lift is asymmetric: weaker models gain pass rate, while stronger models gain efficiency. Third, debugging context requires careful curation, as even faithful debugging context can sometimes mislead LLMs.

We study sketched Ridge regression through the lens of functional estimation. The sketched solution is a plug-in estimator of a nonlinear function of a second-moment/covariance matrix. Within this framework, we analyze two classical bias-reduction principles---iterative Bootstrap and linear aggregation across multiple estimators. We show that these resampling-based procedures cancel low-order bias terms and yield higher-order bias decay in the sketch size without increasing the order of the computational complexity. Concretely, let $A\in\mathbb{R}^{n\times r}$ have full column rank $r$ and let the sketch size be $s$. For Gaussian sketching, a $k$-step iterative Bootstrap estimator achieves bias of order $(\sqrt{r/s})^{k+2}$. For linear aggregation over $m$ plug-in estimators with different sketch sizes, we obtain bias of order $(\sqrt{r/s})^{m+1}$ for Gaussian sketching and corresponding result for common discrete sketches, including uniform and leverage score sampling.

Classifier-Free Guidance (CFG) is a widely used inference-time guidance rule for conditional image generation with diffusion models, which exposes a guidance scale that adjusts conditioning strength after training. This flexibility requires conditional and unconditional model evaluations at every sampling step, so an $N$-step trajectory takes $2N$ forward passes. Single-pass alternatives reduce such cost, yet most collapse CFG's two-component structure, forfeiting the separable base flow and guidance direction that support post-hoc scale control and per-step diagnostics. This work introduces CtrlFlow, a single-pass guidance method that preserves CFG's separable base-flow and guidance-direction interface. To preserve both components at one-pass cost, CtrlFlow attaches two lightweight LoRA adapters to a frozen backbone, each adding $2.5\%$ parameters, and predicts the base flow and guidance direction from shared features. Because the guidance scale multiplies the guidance direction only after the forward pass, a single trained checkpoint remains adjustable at inference without retraining. Preserving two outputs introduces a training mismatch: inference amplifies guidance-direction errors by the chosen scale, while naive component-wise losses weight both outputs equally. To align training with this composed inference-time error, CtrlFlow uses a unified multi-scale distillation loss. Residual guidance errors can still persist at extreme scales, where even small guidance-direction errors are amplified by the guidance scale. To reduce this inference-time error, CtrlFlow applies scale-adaptive modulation to damp velocity-field errors and time-grid warping to concentrate solver evaluations where the trajectory changes most rapidly. On ImageNet $256\times256$ with the DiT-XL/2 backbone, CtrlFlow matches CFG within a pre-specified tolerance in 24 of 25 $(N, s)$ configurations and achieves a measured $1.93\times$ wall-clock speedup at $N{=}25$.


DeepVoting: Learning and Improving Voting Rules with Fine-Tuning

Leonardo Matone ⋅ Ben Abramowitz ⋅ Ben Armstrong ⋅ Avinash Balakrishnan ⋅ Nicholas Mattei

Recently, social choice theory has been found broadly useful for AI alignment and evaluation, often based on the axiomatic properties for combining preferences provided by social choice functions. However, for many sets of axioms, functions which $\textit{never}$ violate the axioms are known to not exist, e.g., Arrow's Impossibility Theorem. Recent work on learning novel rules to $\textit{minimize}$ axiom violations shows the effectiveness of machine learning but explores a limited class of axioms. In this work we show the effectiveness of standard social choice data structures as features for learning existing voting rules, resulting in networks that have better accuracy than past approaches while using an order of magnitude fewer parameters. Subsequently, we develop a process for $\textit{adding axiomatic properties}$ to existing voting rules. By building novel axiomatic loss functions we are able to use transfer learning to fine-tune existing voting rules in a way that quantitatively preserves their original behaviour while also exhibiting strong adherence to previously absent axiomatic properties, resulting in novel voting rules that are $\textit{closer}$ to the impossibility frontier than those in the machine learning or theoretical literature.

Linear attention models offer a highly efficient alternative to Transformers, leveraging fixed-size memory for robust state-tracking and retrieval. However, relying strictly on fixed-capacity associative states bottlenecks their ability to resolve precise structural dependencies that may grow unboundedly in multi-step algorithmic reasoning - especially within deterministic context-free and context-sensitive formal languages. Inspired by the decoupled computation of Turing Machines, we introduce DeltaFugue: a hardware-aware architecture that orchestrates a continuous linear attention controller alongside a differentiable 1D spatial tape. Unlike traditional memory-augmented RNNs that sacrifice sequence-level parallelism, DeltaFugue executes read and write operations natively via delta-rule updates, preserving the training parallelism of linear transformers. Theoretically, we prove that this decoupled spatial routing allows DeltaFugue to recognize regular, hierarchical, and context-sensitive formal languages. Empirically, extensive length generalization evaluations demonstrate that DeltaFugue achieves state-of-the-art accuracy among fully parallelizable models. Scaling to natural language, a 340M-parameter DeltaFugue model matches the performance of strong DeltaNet baselines while exhibiting markedly superior length extrapolation in reasoning-heavy domains like mathematics and code generation.

Behavioral cloning (BC) is a powerful paradigm for learning robotic policies from demonstrations. However, most existing approaches operate in an offline setting and do not address a critical requirement for real-world deployment: the ability to continually improve policies after deployment. While interactive imitation learning methods such as DAgger enable improvement by collecting additional data, they rely on repeated retraining over an ever-growing dataset, leading to increasing computational cost and limited scalability. In practice, robotic systems must adapt continuously during deployment while remaining operational, making retraining on all accumulated data impractical. In this work, we introduce an online formulation of interactive BC, where the policy is updated in real time with expert corrections to address failure cases, with no access to past data. This setting presents a fundamental challenge due to catastrophic forgetting, which degrades previously acquired behaviors when learning from new data. To address this, we propose a novel method based on online ridge regression and random projections, enabling efficient policy updates without gradient-based retraining. We evaluate our approach on a range of manipulation tasks and show that it enables continuous and stable improvement after deployment, outperforming state-of-the-art BC baselines.


Depth Exploration for LLM Decoding

Weisi Yang ⋅ Zipeng Sun ⋅ Stephen Xia

Autoregressive LLM decoding evaluates every generated token through the full layer stack, even though many tokens become predictable at intermediate depths. Existing lossless depth-adaptive methods exploit this redundancy by choosing a single non-final exit depth and verifying its prediction with the final-depth model. However, our measurements show that this selection-based strategy leaves substantial headroom: choosing an exit too late wastes computation, while choosing one too early triggers fallback and discards dependent drafts. We propose Depth Exploration Decoding (DEX), a lossless decoding algorithm that replaces single-depth selection with parallel exploration over multiple candidate depths. At each commit position, DEX validates candidates against the final-depth reference, commits exactly the final-depth token, and collapses the exploration lattice to retain only reusable branch states. This expand--commit--collapse procedure preserves equivalence to standard autoregressive decoding while reducing the cost of committing each token. Across early-exit-trained and standard LLMs, DEX outperforms representative depth-selection baselines and achieves competitive end-to-end throughput against speculative and distributed decoding methods. Moreover, DEX improves as the explored depths become finer, showing that parallel depth exploration provides a scalable way to exploit the underused depth axis of LLM decoding. Code: https://anonymous.4open.science/r/DEX--anonymous/.

Efficient deployment of large language models (LLMs) on mobile devices requires balancing performance with compute and memory constraints. Post-training compression is an effective approach to managing this trade-off, while requiring much less computation than in-training counterparts. We propose DictLLM, a dictionary-based compression scheme in which weight matrices within grouped layers are decomposed into a shared dictionary and layer-specific sparse coefficients, directly obtained from a pre-trained model without further training. Specifically, we cast DictLLM as a four-step post-training optimization. First, we obtain shared dictionaries via singular value decomposition (SVD) of input-scaled weight matrices, amortizing rank across grouped layers; second, we sparsify the layer-specific coefficients to the target compression ratio using a dictionary-aware Hessian; and third, we refine the dictionaries given the resulting sparse coefficients. Finally, we aggressively quantize both components using factor-specific Hessian information: input covariance for the dictionaries and the dictionary-aware Hessian for the coefficients. The resulting representation reduces memory footprint by storing a single shared dictionary per layer group, along with sparse layer-specific coefficients. On language modeling and zero-shot reasoning benchmarks, DictLLM consistently outperforms existing post-training low-rank compression methods and pushes the Pareto frontier between model size and performance. DictLLM compresses the original LLaMA-7B model by 5.1x with only a 3.4-point increase in WikiText-2 perplexity, achieving a 3.6x higher compression ratio than the state-of-the-art low-rank method.

SVD-based pruning and quantization have recently emerged as a promising strategy for the ultra-efficient compression of large language models. In these methods, compression is performed in two stages: components are first truncated, and the remaining ones are subsequently quantized. Although this decoupled pipeline benefits from both pruning and quantization, it requires separate optimization for each stage and fails to fully exploit their balance, which can lead to suboptimal performance under aggressive compression. To address this limitation, we propose a new LLM compression method that co-optimizes pruning and quantization in a unified framework. Our key idea is a differentiable method for learning component-wise bit-widths, allowing less important components to be assigned 0-bit precision and pruned away. Notably, our method performs favorably against two-stage baselines, even when subjected to extreme quantization settings ($1.61$ bits) designed for ultra-efficiency. The code will be publicly released upon acceptance.


Differentially Private Sparse Reward Estimation with Preference Feedback

Meng Ding ⋅ Mingxi Lei ⋅ Jie Zhang ⋅ Jinyan Liu ⋅ Di Wang

Reinforcement Learning with Human Feedback (RLHF) has become a central paradigm for aligning Large Language Models with human values. However, the reliance on sensitive preference data necessitates rigorous differential privacy (DP) guarantees to prevent user information leakage. Existing private alignment methods typically assume dense reward parameterizations, resulting in estimation error bounds that scale polynomially with the ambient {model size $d$}. Such dependence is prohibitive in modern {overparameterized models} where $d$ is vast. In this work, we study \emph{Private Sparse RLHF} problem to address this limitation, assuming that human preferences (reward model) are governed by a subset of $s^* \ll d$ relevant features. We provide the first comprehensive theoretical framework for differentially private sparse reward estimation from pairwise preference feedback under the Bradley-Terry-Luce model, and introduce efficient algorithms for both Sample-DP and Label-DP settings. Theoretically, under the squared $\ell_2$ parameter estimation error, we establish upper bounds of $\widetilde{O}(s*/n + {s*}^2/n^2\varepsilon^2)$ for $(\varepsilon,\delta)$-Sample-DP and $\widetilde{O}(s*/n + s*/n\varepsilon^2)$ for $\varepsilon$-Label-DP based on randomized response, where $n$ is the training data size. For the lower bounds, we show that the Sample-DP rate is nearly optimal, and that the guarantee for our Label-DP method cannot be further improved within the randomized-response framework. We also provide a general lower bound for the problem with Label-DP. Empirically, our framework demonstrates superior efficacy over dense baselines across sentiment generation and safety alignment benchmarks, offering a significantly more favorable privacy-utility trade-off.


Discrete Stochastic Localization for Non-autoregressive Generation

Yunshu Wu ⋅ Jiayi Cheng ⋅ Longxuan YU ⋅ Partha Thakuria ⋅ Rob Brekelmans ⋅ Vagelis Papalexakis ⋅ Greg Ver Steeg

Continuous diffusion is a natural framework for non-autoregressive generation but has consistently underperformed masked discrete diffusion (MDM) on discrete sequence generation. We argue this is not a limitation of continuity itself, but of how previous models parameterize the denoiser as a family of timestep-indexed regimes. We introduce \emph{Discrete Stochastic Localization} (DSL), a continuous-state framework with unit-sphere token embeddings whose Bayes-optimal denoiser is invariant to the nominal signal-to-noise ratio (SNR). One trained network then supports an entire family of per-token SNR paths, with masked diffusion as one special case. Fine-tuning a pretrained MDLM checkpoint with DSL substantially improves distributional faithfulness (MAUVE) on OpenWebText across all step budgets from $T{=}128$ to $T{=}1024$, and the same checkpoint supports random-order autoregressive sampling and a hybrid continuous-then-discrete sampler at as few as $T{=}48$ total steps---without distillation or retraining. On Text8, DSL gives the first continuous-state NLL estimate approaching the MDM range. Code is anonymously available at \url{https://anonymous.4open.science/r/DSL_anonymous-94B5/}.


Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias

Mohua Das ⋅ Pierfrancesco Beneventano ⋅ Shibshankar Dey ⋅ Gareth H McKinley ⋅ Tomaso Poggio

Randomly initialized neural networks induce a prior over functions, but the predictor used in practice is produced only after training. We ask how much of this initial bias survives the training pipeline. To make the question measurable, we introduce *initialization memory*: the dependence of the validation-selected predictor on the scale of the random initialization. We perform controlled CIFAR-10 experiments on ResNets where initialization memory already sharply separates training regimes. Low-learning-rate SGD can interpolate while still remembering its initialization: on ResNet-9 with batch size $b=128$, test accuracy varies by $26.5$ percentage points across initialization scales despite $\ge99.5\%$ training accuracy. This is not undertraining: extending the same low-learning-rate regime to $5{,}000$ epochs leaves the spread essentially unchanged. In contrast, Adam-family methods largely erase the dependence. SGD can also be made to forget when larger learning rates are paired with explicit $L_2$ norm control. We interpret these findings in terms of the time scale of forgetting: gradient-flow-like dynamics can preserve initialization memory, whereas stochastic finite-step effects, explicit norm decay, and adaptive preconditioning erase it on scales governed by the size of explicit or implicit regularization. The practical inductive bias of a trained network is therefore not the architectural prior alone, but the architectural prior after being filtered by the forgetting dynamics of the training pipeline; and the same regularizers that improve generalization are precisely those that erase memory of initialization.

Concept erasure works well for static visual attributes in text-to-image diffusion, but does not transfer to motion concept erasure in video diffusion transformers. We trace this failure to an off-policy trajectory mismatch: once weights are edited, the generation trajectory followed by the edited model deviates from the original trajectory used for supervision, so state-local velocity matching fails to control the motion patterns realized at inference. As a consequence of this mismatch, off-policy weight modification can drive training loss to convergence while the target motion remains fully present in generated videos. We formalize this gap through a Gronwall-type bound showing how weight-induced velocity perturbations propagate into trajectory divergence, and a residual-error bound showing that off-policy training, even at perfect convergence, leaves a nonzero erasure error on the edited trajectory. To address this problem, we introduce {DOME (Drift-Adaptive On-Policy Motion Erasure)}, a trajectory-aligned framework for motion concept erasure in video diffusion transformers. DOME trains on states from the edited model's own trajectory and applies \emph{drift-adaptive anticipation}, which shifts supervision along the observed divergence between the edited and original trajectories, turning trajectory drift from a failure symptom into a training signal. DOME edits the text-conditioned pathway via cross-attention LoRA, and the learned adapter can be merged into the base model so that inference incurs no additional overhead. On Wan~2.1-T2V across 20 motion concepts, DOME consistently suppresses localized contact actions while showing limited effect on whole-body dynamics, validated through both motion-consistency scoring and an independent action classifier. Systematic ablations show that drift-adaptive anticipation, not on-policy rollout alone, is the critical component distinguishing DOME from off-policy and standard on-policy training.

Large language models (LLMs) are often ensembled together to improve overall reliability and robustness, but in practice models are strongly correlated. This raises a fundamental question: which models should be selected when forming an LLM ensemble? We formulate budgeted ensemble selection as maximizing the mutual information between the true label and predictions of the selected models. Furthermore, to explain why performance can saturate even with many models, we model the correlated errors of the models using a Gaussian-copula and show an information-theoretic error floor for the performance of the ensemble. Motivated by these, we propose a simple greedy mutual-information selection algorithm that estimates the required information terms directly from data and iteratively builds an ensemble under a query budget. We test our approach on four datasets spanning binary and multi-class classification and sentiment analysis: MEDMCQA, MMLU, AG News, and IMDB. Across all datasets, we observe that our method consistently outperforms strong baselines under the same query budget.

This position paper argues that large language models are widely deployed on sensitive genomic data, both as fine-tuned classifiers for downstream tasks and as an Embeddings-as-a-Service (EaaS) genomic foundation models shared between institutions, without evaluating whether their embeddings can be reverse-engineered to reconstruct the original nucleotide sequences. This absence of evaluation is not benign: for aggregated embeddings, reconstruction vulnerability is architecture-dependent, cannot be predicted without empirical measurement, and is actively increased by fine-tuning for a significant subset of widely deployed architectures. Evidence from independent empirical studies shows that fine-tuned embeddings reduce reconstruction vulnerability for some architectures while significantly increasing it for others, suggesting that models are deployed with privacy properties ranging from improved to degraded, depending on unassessed architectural choices. Per-token embeddings, increasingly shared in EaaS settings, are universally reconstructible across the genomic foundation models evaluated to date; the effect of fine-tuning on per-token vulnerability remains unevaluated. The privacy outcome depends on the interactions among four factors: embedding type, fine-tuning procedure, pretraining and tokenization strategy, as well as sequence length. None of these factors is currently evaluated in practice. We propose a minimum reconstruction-based privacy evaluation standard, organized as three sequential deployment decisions and two attacker tiers matched to research and clinical stakes, and call on the community to adopt it as a prerequisite for deploying any large language model on sensitive genomic data.


DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy

Erchi Wang ⋅ Pengrun Huang ⋅ Eli Chien ⋅ Om Thakkar ⋅ Kamalika Chaudhuri ⋅ Yu-Xiang Wang ⋅ Ruihan Wu

Differential privacy (DP) has a wide range of applications for protecting data privacy, but designing and verifying DP algorithms requires expert-level reasoning, creating a high barrier for non-expert practitioners. Prior works either rely on specialized verification languages that demand substantial domain expertise or remain semi-automated and require human-in-the-loop guidance. In this work, we investigate whether large language models (LLMs) can automate DP reasoning. We introduce DPrivBench, a benchmark in which each instance asks whether a function or algorithm satisfies a stated DP guarantee under specified assumptions. The benchmark is carefully designed to cover a broad range of DP topics, span diverse difficulty levels, and resist shortcut reasoning through trivial pattern matching. Experiments show that while the strongest models handle textbook mechanisms well, all models struggle with advanced algorithms, revealing substantial gaps in current DP reasoning capabilities. Through further analytic study and failure-mode analysis, we identify several promising directions for improving automated DP reasoning. Our benchmark provides a solid foundation for developing and evaluating such methods, and complements existing benchmarks for mathematical reasoning.


DropKV: Decoupling Residual-Output Perturbation for Near-Optimal KV-Cache Eviction

Aozhong Zhang ⋅ Selcuk Gurses ⋅ Yanxia Deng ⋅ Naigang Wang ⋅ Chi-Chun (Charlie) Liu ⋅ Davis Wertheimer ⋅ Derrick Liu ⋅ Xin Li ⋅ Zi Yang ⋅ Felix X.-.F Ye ⋅ Penghang Yin

Key-value (KV) cache management has become a critical bottleneck for scaling Large Language Models (LLMs) to long-context scenarios due to linear memory growth. While recent eviction methods have moved beyond simple attention-score heuristics to incorporate value representations, they typically face a fundamental trade-off: the underlying optimization for minimizing output perturbation is an NP-hard combinatorial problem, leading to mathematically suboptimal heuristics or implementations that are difficult to kernelize. In this work, we propose DropKV, a simple principled eviction framework based on Decoupled Residual-Output Perturbation. By decoupling the joint eviction decision into independent per-token scoring, DropKV circumvents combinatorial intractability and admits a provable constant-factor approximation guarantee under a cone condition that we empirically verify on three long-context LLMs. To ensure practical viability, we provide a high-performance implementation using fused Triton kernels that avoid materializing the full attention matrix, delivering a 19.8$\times$ scoring-kernel speedup that translates to up to a 9.9\% matched-batch end-to-end prefill speedup. Extensive evaluations on the RULER and LongBench benchmarks demonstrate that DropKV consistently outperforms state-of-the-art baselines at equivalent cache budgets, achieving the lowest inference latency among eviction methods while also remaining faster than the dynamic cache baseline.


Efficient Analytic Uncertainty Quantification for Multimodal Regression

Kun Jin ⋅ James Harrison ⋅ Jiawei Li ⋅ Sihan Liu ⋅ Jiayi Liu ⋅ Randy Linderman ⋅ Yuening Li ⋅ Arnab Bhadury ⋅ Sourabh P Bansod ⋅ Liang Liu ⋅ Jasper Snoek

Efficient uncertainty quantification (UQ) is essential for trustworthy large-scale learning. Existing UQ methods for regression tasks mainly operate under the assumption that the conditional label marginal satisfies single-peak parametric models, e.g., Gaussians, where the negative log-likelihood function simplifies to the mean square error. However, such single-peak assumptions fail in regression tasks featuring multi-modal distributions. On the other hand, semi-parametric methods which achieve strong regression performance for multi-modal distributions often lack efficient quantification on their prediction variances. In this work, we extend UQ techniques based on Variational Bayesian Inference (VBI) to two widely used semi-parametric regression models that yield histogram-like reconstructions of the conditional label densities: Quantile Regression (QR) and Classification Restoration (CR). Our approach introduces a unified, distribution-agnostic framework that simultaneously achieves accurate estimation of complex conditional distributions and highly efficient UQ. Theoretically, our method is grounded in novel formulations of QR and CR within the VBI framework, yielding analytic Evidence Lower Bounds (ELBO) to streamline training and a closed-form or analytically approximated predictive density for efficient inference. Empirically, we evaluate our methods on three large-scale regression benchmarks with multi-modal label distributions. Our framework outperforms state-of-the-art multi-modal regression baselines, and even matches predictive performance of computationally expensive ensemble models. Furthermore, by leveraging epistemic uncertainty estimation, our approach enables highly data-efficient active learning strategies.


Efficient and Transferable Agentic Knowledge Graph RAG via Reinforcement Learning

Junhong Lin ⋅ Shicheng Liu ⋅ Jinyeop Song ⋅ Song Wang ⋅ Julian Shun ⋅ Yada Zhu

Knowledge-graph retrieval-augmented generation (KG-RAG) couples large language models (LLMs) with structured, verifiable knowledge graphs (KGs) to reduce hallucination and provide reasoning traces. However, current KG-RAG systems often rely on fixed pipelines of multiple LLM modules (e.g., planning, reasoning, and responding), which inflate inference costs and tie performance to specific graph schemas. To address this, we introduce KG-R1, an agentic framework that optimizes KG-RAG through reinforcement learning (RL). Unlike modular workflows, KG-R1 uses a single agent that interacts with KGs as its environment, learning to retrieve information at each step and incorporating it into its reasoning and generation in a unified process. Across Knowledge-Graph Question Answering (KGQA) benchmarks, KG-R1 demonstrates both efficiency and transferability—using Qwen 2.5-3B, KG-R1 improves answer accuracy with fewer generation tokens than prior multi-module workflow methods that use much larger foundation or fine-tuned models. Furthermore, KG-R1 exhibits strong plug-and-play capability: after training, maintaining accuracy on unseen KGs without retraining. These properties make KG-R1 a promising KG-RAG framework for real-world deployment. Our code is publicly available at anonymous.4open.science/r/KG-R1-B881/.


Efficient Collaborative LLM Fine-Tuning over Heterogeneous Mobile Devices via Many Backbones to One Side-Network Tuning

Xingke Yang ⋅ Liang Li ⋅ Sicong Li ⋅ Liwei Guan ⋅ Hao Wang ⋅ Xiaoqi Qin ⋅ Jiang Liu ⋅ Xin Fu ⋅ Miao Pan

Collaboratively fine-tuning (FT) large language models (LLMs) over heterogeneous mobile devices fosters immense potential applications of personalized intelligence. However, such a vision faces critical system challenges. Existing federated LLM fine-tuning approaches remain limited by prohibitive on-device resource overhead, straggler-prone synchronous aggregation, and the assumption of a unified pretrained backbone across heterogeneous devices. In this paper, we propose EC-MobiLLM, a novel design for efficient collaborative LLM fine-tuning across heterogeneous mobile devices and heterogeneous backbones. EC-MobiLLM adopts a server-assisted side-tuning paradigm to minimize on-device overhead and pioneers non-blocking asynchronous collaboration to accelerate training. Moreover, EC-MobiLLM introduces adaptive feature alignment to support collaboration across heterogeneous backbones, aligning with practical needs in real-world deployments. Extensive experimental results demonstrate that EC-MobiLLM can maintain robust fine-tuning performance while achieving extremely low on-device memory, with at least 95.2\% reduction in computation overhead, 93.2\% reduction in communication costs and $5.1\times$ faster convergence compared to existing methods, validating its efficacy for practical LLM adaptation over heterogeneous mobile devices.


Efficient Multi-Source Prompt Adaptation for Cross-Domain Open-Vocabulary Learning

Xiaofan Que ⋅ Dingrong Wang ⋅ Daniel Krutz ⋅ Xumin Liu ⋅ Qi Yu

Cross-domain open-vocabulary learning poses a unique and underexplored challenge, requiring models to generalize across both domain shifts and category shifts. To tackle this, we propose MAP: a parameter-efficient Multi-source domain-Adaptive Prompt tuning framework that leverages multiple labeled source domains to improve learning in novel, unlabeled target domains with unseen categories. MAP consists of two key components: Multi-Source Prompt Learning (MSPL) and Unsupervised Target Prompt Learning (UTPL). MSPL disentangles domain-invariant category semantics from domain-specific visual patterns by jointly learning shared and domain-aware prompts. UTPL enhances generalization in the unlabeled target domain by enforcing prediction consistency under text-guided style augmentations, introducing a novel entropy-minimization objective without relying on pseudo-labels. Together, these components enable effective alignment of visual and textual representations across both domains and categories. In addition, we present a theoretical analysis of the proposed prompts, examining their behavior through the lenses of fidelity and distinction. Extensive experiments on challenging CDOV benchmarks demonstrate that MAP achieves state-of-the-art performance with significantly fewer additional parameters.


Efficient Neural Field Learning via Adaptive Coverage and Focused Sampling

Guang Zhao ⋅ Xihaier Luo ⋅ Huan-Hsin Tseng ⋅ Seungjun Lee ⋅ Shinjae Yoo ⋅ Yihui Ren ⋅ Wei Xu

Implicit neural representations (INRs) provide a flexible framework for modeling high-dimensional continuous fields, but their training is often inefficient due to uniform subsampling that ignores spatial heterogeneity. Existing adaptive sampling methods partially address this issue by prioritizing high-error samples, but typically operate at the point level, often leading to redundant sampling in localized regions and insufficient coverage of the domain. We propose \textbf{ACES (Adaptive Coverage-aware Efficient Sampling)}, a structured sampling framework that improves training efficiency by decoupling coverage and importance. ACES constructs adaptive spatial partitions to ensure domain coverage and reduce redundancy, and applies region-level importance weighting to prioritize informative regions during training. We provide a theoretical analysis showing that adaptive partitioning reduces gradient variance by increasing within-region homogeneity, and that controlled bias in region-level weighting may improve optimization efficiency relative to standard unbiased estimators. Experiments on scientific field learning tasks demonstrate that ACES achieves faster convergence and lower error than uniform and pointwise adaptive sampling baselines, with the largest gains in fields with highly localized complexity.


Efficient Transferable Optimal Transport via Min-Sliced Transport Plans

Xinran Liu ⋅ Elaheh Akbari ⋅ Rocio Diaz Martin ⋅ Navid NaderiAlizadeh ⋅ Soheil Kolouri

Optimal Transport (OT) plans provide correspondences between distributions supporting alignment tasks in various domains. Sliced transport plans have been recently proposed as a computationally efficient alternative to OT plans. These methods optimize a one-dimensional projection (slice) to obtain a conditional transport plan that minimizes the transport cost in the ambient space. Despite their efficiency, it remains unclear whether learned slicers transfer to new distribution pairs under shift, an issue central to evolving data and repeated OT computations over related distributions. We study the min-Sliced Transport Plan (min-STP) framework and examine slicer transferability: can a slicer learned on one distribution pair produce effective transport plans for unseen pairs? Theoretically, we show that optimized slicers remain close under slight perturbations of the data distributions, enabling efficient transfer across related tasks. To further improve scalability, we introduce a minibatch formulation of min-STP and provide statistical guarantees on its accuracy. Empirically, we demonstrate that the transferable min-STP achieves strong one-shot matching performance and facilitates amortized training for point cloud and image analysis. Our code is available at https://anonymous.4open.science/r/Min-STP-CF74.


ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation

Tianjiao (Joey) Yu ⋅ Xinzhuo Li ⋅ Yifan Shen ⋅ Onkar Susladkar ⋅ Yuanzhe Liu ⋅ Xiaona Zhou ⋅ Ismini Lourentzou

Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text–3D interaction remains largely implicit: existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We argue that the central design problem is not richer geometry alone, but scale-coupled cross-modal collaboration. We introduce ELSA3D, a unified 3D model that addresses this by structuring language and geometric reasoning jointly along matched abstraction scales. The two streams are coupled through Anchor Tokens: a small, dynamic set of cross-modal fusion units that bind selected text tokens to geometric evidence at a learned scale and write the fused signal back, keeping interaction sparse yet precise. A lightweight per-block router makes both computation and reasoning elastic, choosing which text tokens instantiate anchors at which geometric scale so that cross-modal capacity concentrates where alignment is hardest. ELSA3D establishes new state-of-the-art on image-to-3D, text-to-3D, and 3D captioning, outperforming the strongest unified baseline on every reported metric while roughly halving FLOPs and inference latency relative to a non-elastic version of the same model.


Embedding Foundation Model Predictions in Discrete-Choice Models with Structural Guarantees

Yingshuo Wang ⋅ Xian Sun ⋅ Yanhang Li ⋅ Zhichao Fan ⋅ Zexin Zhuang

Tabular foundation models achieve strong accuracy on choice prediction tasks, but their predictions often violate the economic logic those tasks require: raising a price can increase predicted demand, implied willingness-to-pay estimates are frequently negative or implausible, and unavailable alternatives receive nonzero probability. We propose a two-stage adapter that takes a foundation model's predicted choice probabilities as a precomputed feature and embeds them inside a multinomial logit's utility. In Stage 1, we fit the multinomial logit's structural coefficients by maximum likelihood with sign constraints; in Stage 2, we freeze those coefficients and fit a small neural correction operating on the foundation model's predictions. We prove that this composition exactly preserves the multinomial logit's marginal rate of substitution, so analytically computable value-of-time becomes a mathematical guarantee rather than an empirical accident. Across three datasets and two foundation models, the adapter gains 6.4 percentage points (pp) of test accuracy on average over the multinomial logit and up to 12.8 pp, maintains 100\% cost monotonicity, and produces values of time within the published transportation-economics range on the transportation datasets. Performance degrades gracefully under foundation-model context restriction, retaining at least 6 pp of accuracy gain even at 10\% of the original foundation-model context.


Embodied AI Cannot Scale Without Open-Source Distributed Geometric Optimization Backends for Life-Scale Egocentric Data

Senthil Palanisamy ⋅ Abhishek Anand ⋅ Satpal S Rathore ⋅ Pratyush K Patnaik ⋅ Shubhanshu Khatana

This position paper argues that the embodied AI community must invest in open-source, distributed geometric backends for egocentric pose estimation, and that the current bias toward end-to-end neural solutions is creating a data infrastructure deficit that will bottleneck the next generation of Vision-Language-Action (VLA) models and radiance field reconstruction. While neural frontends (DUST3R, VGGT, DepthAnythingV2) achieve remarkable local spatial accuracy, we show that no open source existing system for ego centric data, neural or classical, delivers metrically consistent global pose trajectories over the 100,000 frame "life scale" sequences that modern embodied AI applications require. This gap is structural, not incidental: fixed window neural architectures cannot enforce global consistency by construction, rolling shutter distortion compounds systematically over long horizons in ways feedforward networks cannot model, and every production system that does work at this scale is proprietary and hardware locked. We formally define the Open Distributed Geometric Optimization Backend, a hybrid architecture combining uncertainty weighted neural priors with distributed bundle adjustment and spline continuous trajectory parameterization, and argue it is the necessary open source infrastructure to generate the metrically grounded training data on which future end to end models depend.

We introduce Sharper Transductive Local Complexity (STLC) as a new tool for analyzing the generalization performance of transductive learning methods, improving upon the current transductive bounds. Our work extends the classical local complexity-based analysis to the transductive setting, incorporating substantial and novel components beyond standard inductive and transductive analysis. Although Local Rademacher Complexity (LRC) has been used to obtain sharp inductive generalization bounds and local complexity-based transductive bounds, it has remained an open problem whether a localized Rademacher complexity framework can achieve exactly the same sharp bounds matching their inductive counterparts. STLC provides a confirmative answer to this question. STLC is constructed by first deriving a new and sharp concentration inequality for the supremum of empirical processes capturing the gap between test and training losses, or the test-train process, under uniform sampling without replacement. The proof establishes a Bernstein-type concentration inequality via a novel entropy-based approach built on the modified log-Sobolev inequality for the swap walk. A subsequent peeling strategy with a surrogate variance operator then yields excess risk bounds in the transductive setting that exactly match the classical LRC-based inductive bounds without the additional logarithmic gap in existing works. We further advance the current state-of-the-art in transductive learning through two applications: (1) for realizable transductive learning over binary-valued function classes with finite VC dimension $\dVC$ and $u \ge m \ge \dVC$, where $u$ and $m$ are the number of test features and training features, STLC gives a nearly optimal bound $\Theta(\dVC \log(me/\dVC)/m)$ nearly matching the minimax rate $\Theta(\dVC/m)$ up to $\log m$ and exactly matching the inductive bound, resolving a decade-old open question; and (2) STLC presents a sharper excess risk bound for transductive kernel learning compared to the prior local complexity–based results.


Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks

Young-Jun Lee ⋅ Seungone Kim ⋅ Minki Kang ⋅ Alistair Liang Chuen Cheong ⋅ Zerui Chen ⋅ Seungho Han ⋅ Taehee Jung ⋅ Dongyeop Kang

Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrated into evolutionary search have recently produced state-of-the-art solutions on optimization tasks, including open mathematical conjectures, GPU kernel design, scientific law discovery, and combinatorial puzzles. To achieve this, prior work applied search scaffolds to one target task at a time, so every new problem is approached from scratch and the experience accumulated during search is discarded once the model finishes its attempt. This leaves the capability of iteratively evolving a solution (\textit{e.g.}, knowing which part to mutate and how, deciding when to backtrack) entirely in the scaffold rather than in the model itself. Whether the model itself could acquire this capability and reuse it across different tasks has been largely unexamined. To address this, we introduce \textbf{\textsc{Evolution Fine-Tuning} (EFT)}, a mid-training paradigm that teaches LLMs to evolve solutions across tasks by converting evolutionary search trajectories into supervision. We construct \datasetName, a 156K-trajectory dataset spanning 10 domains and 371 optimization tasks, and fine-tune open-source LLMs from 2B to 9B parameters. Empirically, EFT confers cross-task generalization: across 22 held-out tasks, our models surpass their base counterparts by 10.22\% on average. Furthermore, when paired with test-time RL, our model match state-of-the-art performance on two circle-packing tasks and outperforms its base-model counterparts on the Erdős minimum-overlap problem. EFT thus serves as a ``practice phase'' for general-purpose discovery agents that doesn't solve new problems from scratch.


Exploiting Fine-Tuning Structures to Improve Adversarial Transferability on Downstream SAM

Shixiong Jiang ⋅ Jialiang Fan ⋅ Mengyu Liu ⋅ Pengfei Gu ⋅ Danny Z Chen ⋅ Fanxin Kong

Combining the Segment Anything Model (SAM) with fine-tuning techniques allows SAM to be effectively adapted to various downstream image segmentation tasks. However, this adaptability introduces new security vulnerabilities related to adversarial attacks. In this paper, we investigate the adversarial transferability between the original SAM and its fine-tuned downstream models. Under limited knowledge conditions of the downstream models, we propose a novel structure-exploiting transferable attack (SETA) method. Our framework mimics the fine-tuning architecture and estimates the parameter distributions of the downstream models to improve the transferability of the generated adversarial samples. Experimental results demonstrate the efficacy of our proposed method in creating adversarial examples against various downstream fine-tuned SAM models.


Exploring MLLM-Diffusion Information Transfer with MetaCanvas

Han Lin ⋅ Xichen Pan ⋅ Ziqi Huang ⋅ Ji Hou ⋅ Jialiang Wang ⋅ Weifeng Chen ⋅ Zecheng He ⋅ Felix Juefei-Xu ⋅ Junzhe Sun ⋅ Zhipeng Fan ⋅ Ali Thabet ⋅ Mohit Bansal ⋅ Chu Wang

Multimodal learning has advanced visual understanding through powerful multimodal LLMs (MLLMs). In visual generation, however, these models are often used only as global text or context encoders for diffusion generators, limiting their ability to provide structured spatial and temporal guidance. This creates an interface gap: MLLMs can parse complex layouts, attributes, and knowledge-intensive scenes, yet current generation pipelines often struggle to transfer such understanding into images and videos with precise, controllable structure. We propose MetaCanvas, a lightweight framework that organizes MLLM representations into spatially indexed canvas tokens and interfaces them with diffusion generators through patch-wise residual fusion. This provides structured spatial and spatiotemporal conditioning rather than relying only on a single global embedding or 1D query sequence. We implement MetaCanvas on three diffusion backbones and evaluate it across six generation and editing tasks requiring precise layouts, robust attribute binding, and fine-grained multimodal control. Under matched settings, MetaCanvas consistently improves over global-conditioning baselines, narrows the gap to specialized closed-source visual generation and editing systems, and unlocks new editing and in-context capabilities for video diffusion backbones. These results suggest that spatially aligned MLLM-derived canvas tokens provide a promising latent interface for transferring multimodal understanding into diffusion-based generation.


Factor Augmented High-Dimensional SGD

Shubo Li ⋅ Yuefeng Han ⋅ Xiufan Yu

Stochastic gradient descent (SGD) is a fundamental optimization algorithm widely used in modern machine learning. In this paper, we propose *Factor-Augmented SGD* (FSGD), a new optimization method that leverages latent factor representations in high-dimensional learning tasks. Unlike standard two-stage dimension reduction approaches that rely on offline representation learning and full data storage, a key novelty of FSGD is that it operates purely on streaming data, making it scalable to large-scale and high-dimensional problems. Furthermore, we establish the first theoretical framework that explicitly incorporates latent factor estimation error into the analysis of SGD, and provide moment convergence in $\ell^s$ norm under decaying step sizes and mini-batch updates. Our results provide a new foundation for employing SGD reliably and scalably in high-dimensional machine learning systems.


Factored Generative Models through Mechanism Diversity

Minghao Fu ⋅ Selena Ge ⋅ Hongjia Liu ⋅ Zijian Li ⋅ Fan Feng ⋅ Kun Zhang ⋅ Biwei Huang

A generative model is factored when each latent dimension independently controls one factor of variation: changing a single latent predictably changes one semantic attribute while leaving the rest unchanged, enabling controllable generation, compositional generalization, and reproducible representations. Existing approaches either constrain the latent distribution, e.g., requiring it to shift with an auxiliary variable, or regularize the model, e.g., sparsity, quantization, or Hessian penalties, neither of which matches how modern conditional generative models actually work: a conditioning signal u reshapes the generator g(·, u) while the latent prior stays fixed. We prove an identifiability theorem showing that generating mechanism diversity, the natural variation that arises when u sufficiently reshapes g, is sufficient for the model to be provably factored, with no parametric assumption on the latent distribution. To actively enforce this condition, we propose Mechanistic Contrastive Learning (MCL), a model-agnostic contrastive objective over generator Jacobians. Empirically, MCL achieves state-of-the-art latent concept disentanglement on three benchmarks equipped with a latent diffusion model, and improves prediction quality and zero-shot cross-task transfer in latent action world models with a 1.4B video generative model as the backbone.

Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evidence. This ambiguity is not uniform across the body, as torso pose and root structure are often relatively well constrained, whereas distal articulations such as the arms and legs are more uncertain. Building on this observation, we propose \textbf{FactorizedHMR}, a two-stage framework that treats these two regimes differently. A deterministic regression module first recovers a stable torso-root anchor, and a probabilistic flow-matching module then completes the remaining non-torso articulation. To make this completion reliable, we combine a composite target representation with geometry-aware supervision and feature-aware classifier-free guidance, so that well-observed body parts remain accurate while ambiguous articulations can be completed without collapsing to deterministic averages. We also introduce a camera-aware synthetic data pipeline that provides the paired image-camera-motion supervision under diverse viewpoints. Across camera-space and world-space benchmarks, FactorizedHMR improves articulation recovery and reduces drift relative to strong baselines, especially in ambiguous scenes.


FairTune Market: A Fair and Trustworthy Marketplace for Fine-Tuned LLMs via Posted-Price & Proper-Scoring Mechanisms

Jiachen Shen ⋅ Xingke Yang ⋅ Hui Zhong ⋅ Aohan Li ⋅ TOMOAKI OHTSUKI ⋅ Xin Fu ⋅ Miao Pan ⋅ Zhu Han

Personalized fine-tuning of Large Language Models (LLMs) enhances user experience but incurs substantial computational costs. Recent studies demonstrate that initializing training from style-similar fine-tuned models significantly reduces adaptation overhead, turning these models into reusable, tradable assets. However, emerging model marketplaces suffer from inherent information asymmetry: providers may misrepresent model capabilities, while model buyers cannot verify utility prior to purchase, leading to adverse selection and potential market collapse. We introduce FairTune Market to mitigate these risks. Our framework combines a posted-price mechanism with a truncated proper scoring rule, conditioning payments on performance under text-generation verification. This design incentivizes truthful reporting as an equilibrium strategy while bounding downside risk from stochastic evaluation. Compared with strong market baselines, FairTune helps buyers find better-matched starting checkpoints, reducing total downstream fine-tuning cost by 72.4\% while recovering 96.3\% of the welfare of an Oracle market with perfect style information. It also improves market health: provider misreporting drops by 96.5\% (from 0.424 to 0.015), and all providers remain profitable and willing to participate in the default setting. These gains are robust across market scales, evaluation noise levels, and niche-provider scenarios. A semi-real reuse experiment on sentence-transformer embeddings from 10 text domains further shows that FairTune achieves near-Oracle matching in a realistic style-embedding space.

For two $d$-dimensional point sets $A,B$ of size $n$, the Chamfer distance from $A$ to $B$ is defined as $\text{CH}(A,B)=\sum_{a \in A} \min_{b \in B} \mathsf{dist}(a, b)$, where $\mathsf{dist}$ is some underlying distance measure. We present efficient algorithms for approximating the Chamfer distance under the $\ell_p$ norm for $1 < p \le 2$. For the general case $1 < p < 2$, we utilize lopsided $\ell_p$ to $\ell_1$ embeddings with weak guarantees. We show that they are sufficient for preserving the Chamfer distance. In the high-dimensional regime, we use Fast Matrix Multiplication techniques to speed up the lopsided embeddings. These lead to $\mathcal{O}(nd\cdot \mathrm{poly}(\log \log n, \log(1/\varepsilon))/\varepsilon^2)$ total time. For the specific Euclidean case ($p=2$), we leverage the Fast Johnson-Lindenstrauss Transform based on Toeplitz matrices and re-analyze the previous $\ell_1$-specific Chamfer algorithm in $\ell_2$. These achieve a runtime matching the recent state-of-the-art $\ell_1$ result of $\mathcal{O}(nd(\log \log n + \log(1/\varepsilon))/\varepsilon^2)$, improving upon the previous $\mathcal{O}(nd \log n/\varepsilon^2)$ bound for $\ell_2$. We also give additional results for the angular distance and upper and lower bounds in the streaming setting.


FedCF: Fair Federated Conformal Prediction

Anutam Srinivasan ⋅ Aditya T. Vadlamani ⋅ Amin Meghrazi ⋅ Srinivasan Parthasarathy

Conformal Prediction (CP) is a widely used technique for quantifying uncertainty in machine learning models. In its standard form, CP offers probabilistic guarantees on the coverage of the true label, but it is agnostic to sensitive attributes in the dataset. Several recent works have sought to incorporate fairness into CP by ensuring conditional coverage guarantees across different subgroups. One such method is Conformal Fairness (CF). In this work, we extend the CF framework to the Federated Learning setting and discuss how we can audit a federated model for fairness by analyzing the fairness-related gaps for different demographic groups. We empirically validate our framework by conducting experiments on several datasets spanning multiple domains, fully leveraging the exchangeability assumption.


Few-Step Diffusion Language Models via Trajectory Self-Distillation

Tunyu Zhang ⋅ Xinxi Zhang ⋅ Ligong Han ⋅ Haizhou Shi ⋅ Xiaoxiao He ⋅ Zhuowei Li ⋅ Hao Wang ⋅ Kai Xu ⋅ Akash Srivastava ⋅ Chengzhi Mao ⋅ Hao Wang ⋅ Vladimir Pavlovic ⋅ Dimitris Metaxas

Diffusion large language models (DLLMs) have emerged as powerful generative models with the promise of fast text generation through parallel decoding. However, realizing this potential in practice remains challenging: reducing the number of decoding steps, typically causes a substantial degradation in output quality due to token factorization error. To alleviate this, we propose a self-distillation framework that trains a few-step student to match the \emph{generative trajectory} of a full-step teacher. We theoretically and empirically show that trajectory-level supervision mitigates this factorization error, thereby enabling effective few-step decoding. We further incorporate Direct Discriminative Optimization (DDO), a reverse-KL objective that encourages mode-seeking toward the teacher’s modes, yielding stronger performance on challenging reasoning tasks. Across reasoning and code-generation benchmarks, our method substantially narrows the gap between few-step and full-step decoding.

We study the problem of average reward Multi-Agent Reinforcement Learning (MARL) where agents interact within a network. Each agent in the network conducts local updates with information within its $k$-hop neighborhood and collaboratively maximizes the overall reward of the entire network. We first provide impossibility results in achieving convergence to the global optimal joint policies in our setting. These impossibility results highlight the challenges of decentralized learning with local information in multi-agent systems, namely the issues of multiagency, where each agent independently optimizes its policy, and partial observability, where agents have limited visibility into the global state. Given these challenges which indicate that decentralized policy optimization can generally be certified only up to stationarity, we provide finite-sample convergence guarantees for a decentralized actor-critic algorithm with linear function approximation that converges to an approximate stationary point with small gradient norm. To overcome such impossibility results, we further show that adding fixed-state entropy regularization reshapes the objective landscape which leads to even stronger global convergence guarantees for our proposed actor-critic algorithm.


Flow Matching Reinforcement Learning via SDE Inference

Maojiang Su ⋅ Tsung-En Lin ⋅ Jerry Yao-Chieh Hu ⋅ Guo Ye ⋅ Haoran Lu ⋅ Shang Wu ⋅ Zhaoran Wang ⋅ Han Liu

We develop a unified framework for flow-based reinforcement learning (RL) grounded in diffusion–flow duality and establish its theoretical foundations. The frameworks enables RL for deterministic flow matching inference by introducing diffusion stochastic inference dynamics that supports exploration while admitting deterministic deployment. We show that our frameworks subsumes several existing flow-based RL methods as special cases and inspire effective now methods. Then we provide theoretical foundations for the frameworks by establishing the correctness and performance transfer guarantee. Specifically, we prove that the policies optimized under stochastic dynamics close to deterministic dynamics at deployment when the pretrained model is well-trained. Moreover, we show that the policy improvement achieved under training-time SDE inference transfers to deployment-time ODE inference. Finally, we conduct experiments to validate our theoretical results.


Foundation Models for Particle Accelerators

Mahindra Rautela ⋅ Alexander Scheinker

Large-scale scientific facilities, such as particle accelerators, are complex systems composed of thousands of interacting components that must operate collectively to support diagnostics and control, and ensure safe, reliable operation. These facilities rely on extensive sensor networks that generate large volumes of heterogeneous, incomplete, and noisy telemetry, often sampled at different rates across subsystems. In this paper, we adopt the foundation-model paradigm to learn general-purpose sensor representations from historical accelerator data through self-supervised pretraining. We introduce SensOFormer, a denoising masked transformer with a Perceiver-style encoder–decoder architecture designed to handle variable sensor sets, accommodate heterogeneous sampling rates, and learn robust representations from noisy measurements. The model is pretrained on multiple particle-accelerator datasets and transferred to downstream tasks, including missing-data imputation, anomaly detection, cavity identification, and fault identification. Through extensive evaluations and comparisons, we demonstrate that a single self-supervised model can learn reusable representations for diverse operational tasks in large-scale scientific facilities.


Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D

Haodong Wen ⋅ Yiran Zhang ⋅ Yingfa Chen ⋅ Kaifeng Lyu

While large language models (LLMs) can solve advanced reasoning problems in seconds, we show that even frontier models fail to perform a much simpler operation: exactly copying an input string that lies well within their context windows. We attribute this failure to positional encodings in Transformer architectures, whose inductive bias favors copying through a shortcut based on matching local contexts rather than carefully locating the corresponding input positions. To address this issue, we introduce 2D-RoPE, which organizes text into a 2D grid rather than a 1D sequence and assigns each token a row ID and a column ID. Under this view, copying becomes simply retrieving input tokens from the same column, which makes the task easy to learn. In synthetic copy experiments, shallow Transformers with 2D-RoPE achieve perfect copying at input lengths hundreds of times longer than those seen during training, whereas standard positional encodings fall far behind. We further pretrain 2D-RoPE language models on DCLM at scales up to 1.4B parameters and show that 2D-RoPE substantially improves performance on copy and NIAH tasks. Overall, our results suggest that viewing text in 2D can benefit language modeling, and we hope this encourages future work to further explore the potential of 2D positional encodings.


Function graph transformers universally approximate operators between function spaces

Takashi Furuya ⋅ S D Mis ⋅ Ivan Dokmanić ⋅ Maarten V. de Hoop ⋅ Matti Lassas

We study the approximation of nonlinear operators between function spaces by transformers. Our approach is to lift functions to measures supported on their graphs and leverage a recently introduced measure-theoretic view of transformers. A function $h$ is represented by its graph measure $\gamma_h$, with finite tokens $\{(x_j,h(x_j))\}_{j=1}^N$ being its empirical approximations. We show that this framework elegantly models discretization refinement via convergence of measures and provides a natural setting for operator learning. Within this framework, we introduce function graph transformers, a graph-preserving subclass of measure-theoretic transformers that maps graph measures to graph measures, which is to say that outputs remain single-valued functions. Crucially, this additional structure does not reduce generality: we prove that the resulting graph-preserving maps can be approximated by finite compositions of standard softmax self-attention layers and pointwise MLPs, yielding universal approximation results for broad classes of nonlinear operators. Unlike existing theoretical approaches to operator learning with transformers, the measure-theoretic framework also accommodates regularized negative-order Sobolev inputs for which discretization invariance is particularly challenging, as well as query points on different output domains. Overall, function graph transformers provide a continuum viewpoint and mathematical toolkit for transformer-based operator learning, clarifying the roles of positional embeddings, graph structure, regularization, and ensuring consistency across discretizations.


Generating the Unheard: Phylogeny-Guided Latent Generation for Ancestral Sound Reconstruction

Tianyi Xu ⋅ Shrinaath Narasimhan ⋅ Evan Gorstein ⋅ Santiago Perea ⋅ Yunyi Shen ⋅ Claudia Solis-Lemus

What did an ancestral bird species sound like? Existing ancestral state reconstruction methods can infer low-dimensional traits such as morphological characters at internal nodes of a phylogenetic tree, but no one has tried to produce rich perceptual signals such as audio. Some of the challenges include inferred representations that are either too low-dimensional to decode or lie in non-generative feature spaces, so no method to date can produce ancestral audio. We introduce the first framework that generates plausible ancestral vocalizations. Our pipeline encodes bird recordings into a VAE latent space, learns a low-dimensional trait projection aligned with phylogenetic distances, performs ancestral inference in this trait space, and recovers decodable latents through an anchored inverse lift before emitting novel waveforms for each ancestral node. Because the entire pipeline stays within a decodable latent space, every internal node receives a genuinely new audio output representing plausible intermediate ancestral sounds unavailable to retrieval-based alternatives. Experiments on two phylogenetically distant bird clades, 21-species Tyrannidae and 19-species Paridae, show that our method is the only approach that simultaneously achieves genuine generation, phylogenetic consistency, and naturalistic audio quality across both datasets.


Generative Control as Optimization: Time Unconditional Flow Matching for Adaptive and Robust Robotic Control

Zunzhe Zhang ⋅ Runhan Huang ⋅ Yicheng Liu ⋅ Shaoting Zhu ⋅ Linzhan Mou ⋅ Hang Zhao

Diffusion and flow-based generators are increasingly used in robot policies to produce action chunks from visual observations, proprioception, and language instructions. However, closed-loop robot control requires solving action-generation problems of varying difficulty, whereas standard diffusion and flow-matching inference typically follows a fixed integration schedule. We introduce Generative Control as Optimization (GeCO), a time-unconditional framework that turns action synthesis from fixed-time integration into iterative optimization. GeCO learns a stationary velocity field over action sequences, and refines actions until the field norm becomes small. This enables adaptive computation at each planning call: simple states can terminate early, while difficult states can use additional refinement. The same field norm also provides a lightweight signal for non-convergent action generation under distribution shift. GeCO can be instantiated in both diffusion-transformer policies and flow-matching Vision-Language-Action (VLA) systems without changing the surrounding policy interface. Across simulation benchmarks and real-world tasks, GeCO matches or improves task performance while enabling convergence-based inference as a plug-and-play replacement for standard time-conditioned action generators.

Many physical systems do not merely move or deform; they grow, adding material and changing the geometry that a world model must represent. Existing world models are typically optimized for pixel prediction, reward prediction, or fixed-support physical dynamics, leaving open how to model systems whose geometric state expands over time and whose future morphology depends on hidden material response. We introduce FOLIAGE, a geometry-centered latent world model for growing surfaces. FOLIAGE represents mature regions as a compact scaffold while allocating higher-resolution state to active growth fronts, allowing the model to focus capacity where new material and near-future change occur. It further separates observation, action, and privileged physics: heterogeneous RGB, point-cloud, and mesh observations are fused into a deployable geometric state; material controls condition only the latent dynamics; and hidden physical energies guide training without being required at deployment. To evaluate this setting, we introduce SURF-GARDEN and SURF-BENCH, providing controlled counterfactual branches, dense cross-modal correspondences, hidden physical signals, and stress tests for growing-geometry state learning. FOLIAGE reduces inverse-material error by $\approx40\%$ and 5-step mesh forecasting Chamfer error by $\approx20\%$ relative to strong baselines, while improving cross-modal retrieval by $\approx25\%$ mAP@100. Stress tests show graceful degradation under sensor loss and correspondence corruption. Transfer experiments on real plant and dynamic-3D data indicate that the learned geometry-centered state generalizes beyond the simulator. We will release code and data upon acceptance.


GeoWind2Plan: Mission-Time 3D Urban Wind Prediction for Energy-Efficient UAV Planning

Shaoxiang Qin ⋅ Xiongye Xiao ⋅ Yucheng Zhao ⋅ Fuyuan Lyu ⋅ Di Zhou ⋅ Jiachen Yao ⋅ Steve Liu ⋅ Animashree Anandkumar ⋅ Liangzhu Leon Wang

In urban low-altitude flight, buildings reshape ambient wind into spatially varying 3D flow, making unmanned aerial vehicle (UAV) energy depend on local wind exposure as well as path length. However, building-resolved wind information is rarely available when a mission must be planned. Computational fluid dynamics (CFD) can produce high-fidelity urban flow fields, but each simulation is tied to a fixed inflow boundary condition and can take hours to days, which is incompatible with urban UAV missions that typically last minutes to tens of minutes. We present GeoWind2Plan, a geometry-to-wind-to-planning framework for mission-time 3D urban wind prediction and energy-efficient UAV planning. Given only a background wind vector, 3D building geometry, and a start-goal pair, GeoWind2Plan transforms the building geometry into a reference-wind frame, predicts mission-relevant 3D wind patches with a localized geometry-conditioned neural operator, stitches them into a queryable local wind field, and optimizes a feasible 3D path and speed profile using a physically grounded UAV energy model. Rather than pursuing CFD-perfect reconstruction, GeoWind2Plan targets decision-useful wind prediction: trajectories are planned with predicted wind and evaluated under high-fidelity CFD wind. Across held-out urban domains, wind speeds, and mission wind-angle regimes, GeoWind2Plan performs corridor-localized wind inference in about 3 seconds, compared with roughly 8 hours for CFD. Under CFD evaluation, trajectories planned with GeoWind2Plan reduce energy by 6.9%, 12.7%, and 4.5% in tailwind, headwind, and crosswind missions relative to wind-agnostic planning, recovering 87.9%, 85.7%, and 75.0% of CFD-reference savings. These results show that fast, corridor-localized 3D urban wind prediction can make wind-aware UAV energy planning practical at mission time.

Speculative sampling is a widely used method for losslessly accelerating Large Language Model (LLM) inference. State-of-the-art speculative sampling methods (e.g. EAGLE series) use a ``dynamic draft tree'' with drafter temperature implicitly set to $T=0$. Greedy drafts, however, perform poorly as temperature increases. Moreover, we show that naively instantiating the dynamic draft-tree with drafter temperature $T>0$ leads to an incorrect distribution. Fixes exist, but are computationally prohibitive as they require $O(|\mathcal{V}|^k)$ computation, where $\mathcal{V}$ and $k$ denote vocab and number of child nodes. We address these limitations by introducing Gumbo, a new speculative sampling algorithm designed for high-temperature settings that uses only $O(|\mathcal{V}|)$ compute. Gumbo is drafter-invariant in that it does not require the draft distribution during verification. Additionally, Gumbo increases the average acceptance length and speedup by expanding and reranking the dynamic draft tree based on Gumbel scores rather than draft probabilities. At the heart of Gumbo is a multi-draft generalization of the communication-free coupling principles and Gumbel sampling developed by Daliri et al., which is of independent interest. Replacing EAGLE-3's speculative sampler with Gumbo results in up to a +26.7\% increase in speedup at a temperature of 1.0 across four models and five datasets, without altering the draft model or requiring any other code changes.

Distilling vision-language models into faster hybrid architectures, such as 3:1 Mamba-2/attention mixes, is now standard practice for making inference efficient. Aggregate benchmarks suggest that this works but they hide selective failures. When we distill Qwen3-VL-8B-Instruct into a 3:1 Mamba-2/attention hybrid, student model stays within 2 points of the teacher across visual reasoning benchmarks like MMStar, MMBench, and MMMU-Pro, while dropping 13 points on optical-character-recognition and document tasks. The student can still understand the scene but loses the fine-grained text needed to answer. We localize much of the failure to a specific kind of position. In a high-resolution image, most patches are sky, wall, or smooth texture, while a small fraction carries text, edges, object boundaries, or other local details. In a token-level diagnostic, the top 10\% highest-density patches have 3.6$\times$ larger residual drift than the bottom 10\% lowest-density patches and 3.5$\times$ larger teacher-masking answer contribution. Uniform weighting devotes many loss terms to low-information background patches, whereas sparse answer-bearing patches receive no special protection. The required intervention is minimal: we replace uniform residual alignment with density-weighted residual alignment, using patch self-dissimilarity as a training-free proxy for position importance. We call this HEED. Compared with normal end-to-end distillation, HEED increases performance by 8.7 points on OCRBench v2 and 5.13 points on a 10-benchmark average. The gain is realized on different teacher models and hybrid architectures. After standard post-training, the student reaches teacher-level performance on the 10-benchmark average with a 4.12$\times$ throughput and a 68\% memory saving at 128k context, with no additional parameters and no inference-time cost.


Heterogeneous Parallelism for Multimodal Large Language Model Training

Yashaswi Karnati ⋅ Kamran J Sadeghi ⋅ Akash Mehra ⋅ Li Ding ⋅ Ali Roshan Ghias ⋅ Pranav Prashant Thombre ⋅ Shifang Xu ⋅ Parth Mannan ⋅ Yu Yao ⋅ Hao Wu ⋅ Eric Harper ⋅ Ashwath Aithal ⋅ Nima Tajbakhsh

Foundation model training is becoming multimodal across the stack, from natively multimodal post-training pipelines to large-scale pretraining. As multimodal coverage broadens, context windows grow, and encoder–LLM scales diverge, a single LLM-centric TP/DP/PP/CP layout increasingly bottlenecks training throughput. This coupling forces encoders to inherit LLM-driven sharding and placement choices that can introduce unnecessary communication, limit useful encoder parallelism, or constrain the LLM schedule; the mismatch is especially pronounced at long contexts, where LLM context parallelism is needed for the fused multimodal sequence but encoder inputs remain bounded. We present heterogeneous parallelism for multimodal large language model training, a training-system abstraction that lets modules in the same end-to-end graph use independent parallel layouts and rank placements, supporting colocated execution, where modules share physical GPUs under different logical grids, and non-colocated execution, where modules occupy disjoint rank sets. The key systems challenge is preserving boundary tensor semantics when adjacent modules use independent layouts: forward activations must be materialized for the destination layout, while backward gradients must be routed back to the source layout. We address this with boundary communicators, runtime primitives that implement these forward and backward layout transforms, together with scheduling extensions for both placement modes. We evaluate optimized homogeneous, colocated heterogeneous, and non-colocated heterogeneous configurations across diverse multimodal workloads and GPU scales to characterize where each placement mode helps. Across this sweep, colocated heterogeneity improves TFLOPs/s/GPU by up to 14.4\% at short context and 41.8\% at long context, while non-colocated heterogeneity improves aggregate token throughput by up to 13.0\% and TFLOPs/s/GPU by up to 9.6\%. We validate loss convergence parity against homogeneous baselines and release an open-source implementation in Megatron-LM.


High Entropy Regularization Leads to Symmetry Equivariant Policies in Dec-POMDPs

Johannes Forkel ⋅ Constantin Ruhdorfer ⋅ Michael Beukman ⋅ Andreas Bulling ⋅ Jakob Foerster

We prove that in any Dec-POMDP, sufficiently high entropy regularization ensures that the policy gradient flow with tabular softmax parametrization always converges, for any initialization, to the same joint policy, and that this joint policy is equivariant w.r.t. all symmetries of the Dec-POMDP. In particular, policies coming from different initializations will be fully compatible, in that their cross-play returns are equal to their self-play returns. Through extensive evaluation of independent PPO, arguably the standard baseline deep multi-agent policy gradient algorithm, in the Hanabi, Overcooked and Yokai environments, we find that the entropy coefficient has a massive influence on the cross-play returns between independently trained policies, and that the decrease in self-play returns coming from increased entropy regularization can often be counteracted by greedifying the learned policies after training. In Hanabi in particular we achieve a new SOTA in inter-seed cross-play this way. While we give examples of Dec-POMDPs in which one cannot learn the optimal symmetry equivariant policy this way, both our theoretical and empirical results suggest that one should consider far higher entropy coefficients during hyperparameter sweeps in Dec-POMDPs than is typically done.


HiPhy: Hierarchical Alignment for Physically-Plausible Multi-Principle Video Generation

Tahira Kazimi ⋅ Shubhankar Borse ⋅ Durga Malladi ⋅ Fatih Porikli ⋅ Pinar Yanardag

Video generation models have achieved remarkable visual fidelity and have strong potential to become general-purpose world simulators. Despite this progress, they still fail to generate videos which adhere to laws of physics. The problem becomes even more apparent in realistic settings where multiple physical principles must work together within the same video; for example, "a balloon floating upward while steam rises from a pot" requires buoyancy and fluid dynamics to unfold coherently and simultaneously. Yet existing methods largely ignore multi-principle interactions, focusing on a single principle per video. We propose HiPhy (Hierarchical Physical Alignment), a reinforcement learning framework that grounds video generation in physical laws through a dual-level objective: locally enforcing the temporal dynamics of individual physical principles, and globally ensuring the physical and semantic coherence of the entire scene. To support multi-principle generation, we construct a 50K-prompt dataset and introduce a prompt benchmark MultiPhyBench, spanning a diverse range of co-occurring physical events. Our experiments show that HiPhy significantly outperforms prior methods and baselines, improving physical commonsense and semantic alignment significantly across various benchmarks, with the largest gains on scenes involving multiple concurrent physical principles where competing methods degrade most sharply.

Diffusion Maps (DM) is a well studied and popular non-linear dimension reduction algorithm. Surprisingly, while the convergence properties of DM are well understood, the {\em geometric} properties of the embedding, such as reach and smoothness have not yet been studied. Under a set of standard assumptions on a family of submanifolds $\subset \mathbb{R}^D$, we derive a series of geometric properties that are preserved by DM, including almost uniform density, finite polynomial approximation and reach. Leveraging these properties, we establish rigorous bounds on the embedding errors introduced by the DM algorithm, of the order $(\frac{\log n}{n})^{\frac{1}{8d+16}}$. These results offer a solid theoretical foundation for understanding the performance and reliability of DM in practical applications.


Hydra: Towards Transferable Multi-Task Learning on Temporal Graphs

Kiarash Shamsi ⋅ Farimah Poursafaei ⋅ Tran Gia Bao Ngo ⋅ Reihaneh Rabbany ⋅ Baris Coskunuzer ⋅ Guillaume Rabusseau ⋅ Michael Bronstein ⋅ Shenyang Huang ⋅ Cuneyt Akcora

Real-world evolving networks are naturally modeled as temporal graphs (TGs), where capturing temporal dynamics is essential for predicting future graph properties that support downstream decision-making. Existing temporal graph methods have been developed primarily for single-task prediction, and little is known about their generalization across tasks or transfer to unseen networks. This leaves the challenge of multi-task graph property prediction in TGs largely open. We address this challenge by introducing Hydra, a novel architecture that integrates local connectivity features from temporal GNNs with a spectral learning module that captures global connectivity patterns. This design enables joint learning of local and global information under a multi-task objective. In multi-task classification, Hydra achieves an 8.9% relative gain in AUC over the strongest competitor. In multi-task regression, Hydra achieves competitive results in all three tasks, while obtaining the best results in two tasks with a 8.2% relative gain in MAE compared to the strongest baseline. Moreover, Hydra delivers these gains with a 22× reduction in training time compared to temporal transfer models. These results provide the first systematic evidence that multi-task transferable learning on temporal graphs is effective. By delivering consistent top-ranked performance, Hydra highlights multi-task training on temporal graphs as a promising direction toward adaptable foundation models for temporal graphs.

We consider the offline imitation learning from observations (LfO), where expert demonstrations are scarce and contain only state observations, and the suboptimal policy is far from expert behavior. In this regime, many existing imitation learning approaches struggle to extract useful information from imperfect data since they impose strict support constraints and rely on brittle one-step models. To tackle this challenge, we propose Trajectory-level Generative Embedding (TGE) for offline LfO. TGE constructs a dense, smooth surrogate reward by using particle based entropy estimation to maximize the log-likelihood of expert trajectories in the latent space of a temporal diffusion model trained on offline suboptimal data. By leveraging the structured geometry of the learned diffusion embedding, TGE captures long-horizon temporal dynamics and effectively bridges the gap under severe support mismatch, ensuring a robust learning signal even when offline data is distributionally distinct from the expert. Empirically, the proposed approach consistently matches or outperforms prior offline LfO methods across a range of D4RL locomotion and manipulation benchmarks.

Consider a marketplace of AI tools, each with slightly different strengths and weaknesses. By picking the right model for the task at hand, a user can do better than simply using the same model for everything. Routers operate under a similar principle, where sophisticated model selection can increase overall performance. However, aggregation is often noisy, reflecting imperfect user choices or routing decisions. This leads to two main questions: first, what does a "healthy marketplace" of models look like for maximizing consumer utility? Secondly, how can we incentivize producers to create such models? We show that winrate, a standard benchmark in LLM evaluation, can incentivize model creators to homogenize for both types of model changes, reducing consumer welfare. We propose a new mechanism, weighted winrate, which rewards models for answers that are higher quality, and show that it provably improves incentives for producers to specialize and increases consumer welfare. We conclude by exploring the impact of our theoretical results in empirical benchmark datasets and discussing implications for benchmark design.

Image classification models often encounter objects in diverse and unpredictable contexts at test time, which may differ substantially from those seen during training. This phenomenon, termed context shift, exposes a key fragility in neural networks: when models rely on non-causal contextual cues for prediction, shifts in context can lead to significant performance degradation. In this paper, we propose a novel training framework to improve the robustness of convolutional neural networks under context shift. Our approach leverages spatial feature attribution to guide models toward making predictions that rely less on contextual regions and more on object-relevant features. As a first step, we identify a theoretical limitation of existing feature attribution methods and introduce a new variant, ContrastiveCAMs, which produces more faithful attribution maps of model predictions. Building on ContrastiveCAMs, we further propose Context-Regularized Cross-Entropy (CR-CE), a modification of the standard cross-entropy loss that regularizes the model’s attention by suppressing the influence of contextual regions, thereby improving robustness to context shift. We evaluate the effectiveness of our approach on several medium-to-large scale datasets (Waterbirds, Spawrious, ImageNet/ImageNet-BG, Hard-ImageNet), and report consistent improvements in context-shift robustness.


Inertia-1: An Open Exploration of Wearable Motion Foundation Models

Zongzhe Xu ⋅ Aakarsh Anand ⋅ Sarah Jiang ⋅ Chuntung Zhuang ⋅ Zitao Shuai ⋅ Sriram Sankararaman ⋅ Yuzhe Yang

Wearable motion sensing provides a continuous and scalable window into human behavior and health, making it a natural fit for foundation models, yet its pretraining and scaling principles remain poorly understood. Prior work studies isolated design choices, such as sensor placement or sampling frequency, often under fixed settings and narrow downstream tasks that fail to capture real-world sensing diversity. We introduce Inertia-1, a fully open exploration of wearable motion foundation models. Using massive corpora of accelerometer data from global sources spanning more than 18.2M hours, we build a controlled framework for studying the full lifecycle of wearable motion foundation models, covering data choices such as sensor modality, device placement, sampling rate, window length; model choices such as architectures and model size; and training choices such as pretraining objective and data scale. Extensive evaluations across 15 datasets spanning human activity recognition, freezing-of-gait detection, and disease prediction reveal intriguing findings for building motion foundation models that generalize across tasks and sensing conditions. Collectively, Inertia-1 not only presents state-of-the-art recipes for diverse downstream tasks, but also serves as a comprehensive, practical, and open cookbook for wearable motion representation learning.

Offline-to-online reinforcement learning first warm-starts a policy from a fixed offline dataset and then improves it with limited online interaction. Offline data reduces uncertainty, but it does not remove the need for exploration; it changes what remains to be explored. We formalise this residual uncertainty by the conditional mutual information \$I(\chi;\tau_{1:T}\mid\mathcal\{D\}_N)\$ between a learning target \$\chi\$ and the online trajectories after conditioning on the offline dataset. This view leads naturally to information-directed sampling (IDS), a family parameterised by \$\eta\ge 0\$ that selects actions by trading off instantaneous regret against information gain. We prove a generic offline-to-online Bayesian regret bound for IDS through a ratio certificate: any information-ratio bound satisfied by a reference Thompson-sampling policy over the same randomised policy class is inherited by IDS. In a known-dynamics Bayesian linear-reward model, the conditional mutual information has a log-determinant form, and vanilla IDS (\$\eta=0\$) satisfies \$\widetilde O\(Hd\min\{\sqrt T,\,T\sqrt\{C^\dagger\_{\beta,\mathrm\{IDS\}_0}(N,T)/N\}\}\)\$, where the coverage coefficient is tied to the visitation distribution induced by vanilla IDS itself. We also identify a warm-start regime with a dominated but informative probe in which vanilla IDS selects the probe while Thompson sampling never does, giving a constant-factor Bayesian regret separation. Controlled bandit experiments and D4RL offline-to-online experiments support this mechanism: IDS is most beneficial when offline data is informative but leaves biased or low-probability residual uncertainty that can be resolved by targeted online actions.


Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards

Shuze D Liu ⋅ Claire Chen ⋅ Jiabao S Xiao ⋅ Lei Lei ⋅ Yuheng Zhang ⋅ Yisong Yue ⋅ David Simchi-Levi

The recent advancement of Large Language Models (LLMs) has established their potential as autonomous interactive agents. However, they often struggle in strategic games of incomplete information, such as bilateral price negotiation. In this paper, we investigate if Reinforcement Learning from Verifiable Rewards (RLVR) can effectively teach LLMs to negotiate. Specifically, we explore the strategic behaviors that emerge during the learning process. We introduce a framework that trains a mid-sized buyer agent against a regulated LLM seller across a wide distribution of real-world products. By grounding reward signals directly in the maximization of economic surplus and strict adherence to private budget constraints, we reveal a novel four-phase strategic evolution. The agent progresses from naive bargaining to using aggressive starting prices, moves through a phase of deadlock, and ultimately develops sophisticated persuasive skills. Our results demonstrate that this verifiable training allows a 30B agent to significantly outperform frontier models over ten times its size in extracting surplus. Furthermore, the trained agent generalizes robustly to stronger counterparties unseen during training and remains effective even when facing hostile, adversarial seller personas.

Language models exhibit strong robustness to paraphrasing, suggesting that semantic information may be encoded through stable internal representations, yet the structure and origin of such invariance remain unclear. We propose a local geometric framework in which semantically equivalent inputs occupy structured regions in latent space, with paraphrastic variation along nuisance directions and semantic identity preserved in invariant subspaces. Building on this view, we make three contributions: (1) a geometric characterization of invariant latent features, (2) a contrastive subspace discovery method that separates semantic-changing from semantic-preserving variation, and (3) an application of invariant representations to zero-shot model attribution. Across models and layers, empirical results support these contributions. Invariant structure emerges in specific depth regions, semantic displacement lies largely outside the nuisance subspace, and representation-level interventions indicate a causal role of invariant components in model outputs. Invariant representations also capture model-specific geometric patterns, enabling accurate attribution. These findings suggest that semantic invariance can be viewed as a local geometric property of latent representations, offering a principled perspective on how language models organize meaning.


ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers

Jiayi Wu ⋅ Robert Joseph George ⋅ Animashree Anandkumar

Formal theorem proving has emerged as a frontier challenge for machine learning, yet the ecosystem is fragmented: proofs remain siloed across incompatible systems, limiting both training data for learning-based provers and the portability of verified results. We present ITPEval, the first benchmark for evaluating automated formal proof translation across four major ITPs (Lean 4, Rocq, Isabelle, and HOL Light), spanning two distinct logical foundations (the Calculus of Inductive Constructions and higher-order logic). Our benchmark comprises 1,560 source files spanning 6,848 theorems and lemmas across four systems (390 aligned items per ITP), organized into two tiers: a controlled tier of self-contained, axiomatized files (64 files, 660 lemmas), and an ecosystem tier of 1,496 files drawn from existing libraries and community formalizations. We release itpeval, a unified multi-ITP verification infrastructure with state-isolated warm backends that amortize prover startup while preserving per-artifact native checking semantics. We evaluate both statement and proof translation across five frontier and open-weight LLMs on 12 directed translation pairs: statement translation peaks at 29.1% pass@1 and proof translation at 10.5%; controlled theorems reach 29.7% proof pass@1 versus 5.2% for ecosystem-level translations, confirming that library mismatch is the dominant bottleneck. In an autoformalization/auto-informalization round-trip study, multi-ITP context substantially improves Lean 4 formalization success (4.8% to 10.6%), showing that aligned cross-ITP corpora serve not only as evaluation benchmarks but also as generation-time evidence. Our benchmark, verification infrastructure, and evaluation pipelines are publicly released.


Joint Consistency: A Unified Test-Time Aggregation Framework via Energy Minimization

Yunzhen Yao ⋅ Hongye Wang ⋅ Yahong Wang ⋅ Michael Gastpar ⋅ Jiang Bo ⋅ Lie He

This paper studies test-time aggregation, an approach that generates multiple reasoning traces and aggregates them into a final answer. Most existing methods rely on evaluation signals collected from candidate traces in isolation or answer frequencies, while ignoring comparative interactions among candidates. We propose Joint Consistency (JC), formulated as a constrained Ising-type energy minimization problem, where independent evaluation signals act as external fields and pairwise comparisons act as interactions. JC provides a unified framework for test-time aggregation that subsumes existing voting and weighted aggregation methods as special cases. Our construction of the interaction matrix leverages LLM-as-a-judge comparisons, and admits a theoretical interpretation under answer-level homogeneity assumptions. Moreover, we develop an efficient approximation strategy that makes interaction modeling practical for large-scale test-time aggregation. Experiments on math and code reasoning benchmarks show that JC consistently outperforms existing baselines across tasks, judge models, trace budgets, and trace-generation settings.


JudgmentBench: Comparing Rubric and Preference Evaluation for Quality Assessment

Russell Yang ⋅ Ruishi Chen ⋅ Pierce Kelaita ⋅ Riya Ranjan ⋅ Sibo Ma ⋅ Charles Dickens ⋅ Matthew Guillod ⋅ Megan Ma ⋅ Julian Nyarko

Two methodologies dominate current practices of benchmarking: rubric-based scoring evaluates items against predefined criteria, whereas comparative judgment elicits pairwise preferences between outputs. Although both methodologies are widely used, the choice between them is rarely justified. We release JudgmentBench, a benchmark of 30 real-world legal tasks, paired with 1,539 rubric scores and 1,530 pairwise preference judgments collected from practicing attorneys—including at major U.S. law firms—with substantial experience. The annotations constitute the first publicly available dataset in a high-expertise domain in which both supervision signals are elicited from the same experts on the same items. Using LLM-generated outputs at three constructed quality levels, we provide an initial empirical comparison: comparative judgments recover the intended quality ordering substantially better than rubrics (mean Spearman's rank correlation of 0.908 vs. 0.150, $\widehat{D}=0.758$ [0.494, 1.021]) while requiring less than half the annotation time. The patterns hold for human annotators and LLM autograders. Beyond this initial comparison, the paired structure of the dataset supports a broader research agenda on how expert judgment should be elicited, aggregated, and used as supervision in domains without verifiable ground truth.


LaCache: Robust Semantic Caching for LLM Serving

JIACHENG LIANG ⋅ Yuhui Wang ⋅ Tanqiu Jiang ⋅ Ting Wang

Semantic caching, which reuses responses to semantically similar requests via their embeddings, has seen growing adoption in LLM serving, offering faster responses and reduced costs. Yet existing schemes are fundamentally vulnerable to cache-collision attacks, wherein an adversary pollutes the cache by injecting crafted queries, corrupting responses to subsequent legitimate requests. We present LaCache, a novel semantic caching scheme that addresses this vulnerability through a conceptually simple yet principled redesign. The key insight is that while the adversary has full control over the adversarial query, it has far less control over its response, which must simultaneously satisfy multiple semantic constraints. Rather than checking only the cache hit of a query, LaCache additionally checks the cache hit of its first $k$ (speculatively) decoded tokens. This design yields two concrete benefits. First, it provides formally guaranteed resilience against cache-collision attacks: we prove that it is impossible to craft adversarial queries that simultaneously elicit malicious responses and collide with benign queries. Second, the enriched index supplies additional semantic context for cache retrieval, improving response relevance. Empirical evaluation across diverse LLMs and benchmarks validates both LaCache's security guarantees and efficiency gains, pointing to a promising direction for robust semantic caching.


LACE: Lattice Attention for Cross-thread Exploration

Yang LI ⋅ Zirui Zhang ⋅ Yang Liu ⋅ Chengzhi Mao

Current large language models reason in isolation. Although it is common to sample multiple reasoning paths in parallel, these trajectories do not interact, and often fail in the same redundant ways. We introduce LACE, a framework that transforms reasoning from a collection of independent trials into a coordinated, parallel process. By repurposing the model architecture to enable cross-thread attention, LACE allows concurrent reasoning paths to share intermediate insights and correct one another during inference. A central challenge is the absence of natural training data that exhibits such collaborative behavior. We address this gap with a synthetic data pipeline that explicitly teaches models to communicate and error-correct across threads. Experiments show that this unified exploration substantially outperforms standard parallel search, improving reasoning accuracy by over 7 points. Our results suggest that large language models can be more effective when parallel reasoning paths are allowed to interact.


Lang-SVG: Hierarchical Image Vectorization with Language Priors

Xi Liu ⋅ Chaoyi Zhou ⋅ Run Wang ⋅ Jiaang Li ⋅ Feng Luo ⋅ Junxiang Huang ⋅ Siyu Huang

Image vectorization reconstructs raster images as compact scalable vector graphics (SVG) representations. Most of existing vectorization methods optimize for pixel-level rendering fidelity, where the reconstructed SVGs often lack alignment with human-perceived object-part hierarchies, making them difficult to manipulate. To address this, this work for the first time studies the problems of hierarchical image vectorization and SVG semantic labeling. We propose Lang-SVG, a tree-structured SVG representation that encodes SVG primitives with language tags and parent-child hierarchy. Lang-SVG novelly incorporates granularity-controllable segmentation priors and vision-language priors into the optimization-based vectorization pipeline. It introduces a Tree-Guided Pruning and Merging module to reduce redundant multi-granularity SVG primitives into coherent semantic structures, and an L1-SAM Prompter to recover missing regions. It also includes a new SVG semantic tagging method that integrates SVG tree context into multimodal large language models (MLLMs) to assign language tags to SVG primitives. We also propose a new evaluation protocol for image vectorization, measuring structural quality and semantic part alignment beyond reconstruction fidelity. Experiments demonstrate that Lang-SVG achieves state-of-the-art performance in both rendering and structural quality.


Language-Based Agent Control

Timothy Zhou ⋅ Loris D'Antoni ⋅ Nadia Polikarpova

This paper introduces language-based agent control (LBAC), a method to apply techniques from language-based security to the control of AI agents. Unlike systems-level defenses such as I/O sandboxing, LBAC can express application-level policies. Moreover, in LBAC agents may perform computations and recursively invoke subagents with tool access. Language-based techniques enable the design of APIs which, through a combination of typing discipline and internal runtime checks, ensure that any well-typed program obeys a desired property. The core idea of LBAC is to have agents interact with the world by generating programs against such APIs rather than by issuing tool calls directly. These programs may themselves contain recursive calls to subagents, which retain full tool access; the type checker rejects ill-typed programs before execution, and policies thereby extend uniformly across the entire agentic system, including its scaffolding and control flow. We demonstrate LBAC with three case studies: I/O sandboxing via filesystem capabilities, data provenance, and information-flow control.


Language-Conditioned World Modeling for Visual Navigation

Yifei Dong ⋅ Fengyi Wu ⋅ Yilong Dai ⋅ Lingdong Kong ⋅ Guangyu Chen ⋅ Qiyu Hu ⋅ Yetong Sha ⋅ Feng Liu ⋅ Siyu Huang ⋅ Qi Dai ⋅ Zhi-Qi Cheng

Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a natural language instruction given only an initial egocentric observation. Without access to goal images, the agent must rely on language to shape its perception and continuous control. We introduce the LCVN Dataset, a benchmark of 39,016 trajectories and 117,048 human-verified instructions spanning diverse environments and instruction styles. Building on this benchmark, we study two complementary paradigms: (i) latent-imagination policy learning, in which a diffusion-based world model (LCVN-WM) imagines future observations and an actor-critic agent (LCVN-AC) learns its policy entirely within the imagined latent space; and (ii) unified autoregressive prediction, in which a single multimodal backbone (LCVN-Uni) jointly predicts actions and observations in one forward pass over a shared token sequence. Experiments show that two paradigms offer complementary strengths: latent imagination produces more temporally coherent rollouts, whereas unified prediction generalizes better to unseen environments. Targeted ablations further isolate the contributions of language guidance, conditioning signals, and instruction style, clarifying when language grounding versus dynamics modeling is the performance bottleneck. Together, these findings position LCVN as a concrete testbed for studying how language, imagination, and decision-making interact in embodied agents.


Language Denoising Objectives Extend the Value of Limited Data

Justin Lovelace ⋅ Christian Belardi ⋅ Srivatsa R Kundurthy ⋅ Shriya Sudhakar ⋅ Kilian Weinberger

The choice of language model pretraining objective was once contested among causal language modeling (CLM), masked LM, T5 span corruption, and UL2-style mixtures of denoisers. It was settled in favor of CLM during an era when high-quality text was abundant relative to compute. That era has ended. Compute now routinely outpaces the supply of unique data, and modern pipelines rely heavily on data repetition both during pretraining on curated corpora and during midtraining on small domain corpora. We revisit the objective question in this data-constrained regime. Across a large sweep over models from 15M to 1B parameters, unique data budgets up to 6B tokens, and across repetition budgets, we compare CLM against three denoising objectives: fill-in-the-middle (FIM), T5-style span corruption, and a UL2-style mixture-of-denoisers. All four scale similarly when every training token is unique, but they differ significantly in their tolerance for repeated data. Denoising objectives, which naturally augment language data on every pass, accumulate substantially less overfitting cost than CLM. We find that span corruption and mixture-of-denoisers are the most robust. We then validate the practical payoff in a midtraining setting. We continue web-text pretrained checkpoints on a limited mixture of web-text and math data. We observe denoising midtraining objectives outperform CLM on GSM8k, and the gap widens with increased repetition. Choosing a denoising objective is a simple, complementary lever to model and data scaling for practitioners working under tight data budgets.


Language Model Goal Selection Differs from Humans' in a Self-Directed Learning Task

Gaia Molinaro ⋅ Dave August ⋅ Danielle Perszyk ⋅ Anne Collins

Whether in agentic workflows, social studies, or chat settings, large language models (LLMs) are increasingly being asked to replace humans in choosing which goals to pursue, rather than completing predefined tasks. However, the assumption that LLMs accurately reflect human preferences for goal setting remains largely untested. We assess the validity of LLMs as proxies for human goal selection in a controlled, self-directed learning task borrowed from cognitive science. Across five models (GPT-5, Gemini 2.5 Pro, Claude Sonnet 4.5, Qwen3 32B, and Centaur), we find substantial divergence from human behavior. While people gradually explore and learn to achieve goals with diversity across individuals, most models exploit a single identified solution or show surprisingly low performance, with distinct patterns across models and little variability across instances of the same model. Chain-of-thought reasoning and persona steering provide limited improvements, and our conclusions hold across experimental settings. While they await confirmation in applied settings, these findings highlight the uniqueness of human goal selection and caution against its replacement with current models.


Learning-Augmented Online Scheduling with Parsimonious Preemption

Mugen Blue ⋅ Sungjin Im ⋅ Alexander Lindermayr

Learning-augmented algorithms have emerged as a powerful paradigm to surpass traditional worst-case lower bounds by integrating potentially noisy predictions. While this framework has seen success in online scheduling, existing work primarily optimizes job latency while relying on frequent, "blind" preemptions. This ignores the fundamental trade-off between algorithmic performance and preemption complexity. We provide the first systematic study of learning-augmented scheduling that curbs preemption while optimizing latency. We establish that the gap between theoretical latency bounds and preemption overhead can be bridged with solid analytical foundations. Our results include $O(1)$-competitive algorithms for single and unrelated parallel machines with only $O(1)$ preemptions per job under accurate predictions, with overhead scaling logarithmically with the prediction error. By providing the first bounded-preemption guarantees for unrelated and malleable machines, we extend the theoretical reach of the learning-augmented framework to more constrained and realistic settings. Finally, our algorithms are validated through experiments.

Safe reinforcement learning (RL) commonly enforces expected-cost constraints, but such expectation safety may fail to control the probability of rare high-cost trajectories. Chance-constrained MDPs (CCMDPs) impose a stronger probability-level requirement, but are widely viewed as harder because the chance constraint is nonconvex and depends on the full trajectory rather than a Bellman-linear expectation. In this paper, we reveal that this computational difficulty does not imply a higher statistical price. In tabular discounted CCMDPs, we give a Bellman-certified model-based algorithm whose sample complexity matches the \emph{minimax optimal} primary dependence of classical CMDP learning, and prove a matching lower bound. Technically, our key idea is the \emph{Bellman distributional certificate}, which constructs a Bellman recursion for constraint violation probabilities before policy selection. The certificate can be reused across candidate policies, shifting concentration analysis from the number of possible policies to a finite Bellman certificate table. Building on this certificate, we establish high-probability safety and near-optimality guarantees for deterministic policy learning in discounted CCMDPs. To our knowledge, we also provide the first model-free sample complexity guarantee for stochastic policy learning in CCMDPs via variance-reduced policy gradient. Numerical experiments on synthetic CCMDPs and IEEE 14-bus energy-storage control benchmark illustrate the safety and behavior of the proposed algorithms.


Learning Large-Scale Competitive Team Behaviors with Mean-Field Interactions

Bhavini Jeloka ⋅ Yue Guan ⋅ Panagiotis Tsiotras

While multi-agent reinforcement learning (MARL) has shown strong empirical performance, existing methods struggle with large number of agents, due to the combinatorial growth of joint interactions. Mean-field (MF) approximations address this by replacing pairwise interactions with interactions against population distributions, yielding tractable large-population policies. However, existing MF formulations focus on fully cooperative or purely competitive settings and fail to capture the mixed cooperative–competitive structure of team-based games. We extend mean-field learning to _zero-sum team games_, where agents cooperate within teams and compete at the team level. We show that such games admit $\epsilon$-optimal decentralized policies that depend only on local states and population distributions. Building on this structure, we propose MF-MAPPO, a scalable algorithm with a shared actor and a minimally informed critic per team. MF-MAPPO is trained directly in finite-population simulators rather than using mean-field oracles, thereby enabling deployment to realistic scenarios with thousands of agents. We further extend MF-MAPPO to partially observable settings via a simple gradient-regularized training scheme. Experiments on large-scale benchmarks in our simulation platform $\texttt{MFEnv}$, including population games with analytical solutions and high-dimensional battlefield scenarios, demonstrate that MF-MAPPO outperforms existing MARL baselines and yields rich, heterogeneous behaviors.


Learning Orthonormal Bases for Function Spaces

Hamidreza Kamkari ⋅ Mohammad S Nabizadeh ⋅ Justin Solomon

Infinite-dimensional orthonormal basis expansions play a central role in representing and computing with function spaces due to their favorable linear algebraic properties. However, common bases such as Fourier or wavelets are fixed and do not adapt to the structure of a given problem or dataset. In this paper, we aim to represent these bases with neural networks and optimize them. Our key idea is that any target infinite-dimensional orthonormal basis can be viewed either as a point on the Lie manifold of the orthogonal group, or equivalently, as the endpoint of a continuous path on that manifold that connects a reference basis, e.g. Fourier, to that target. Paths on the Lie manifold satisfy ordinary differential equations (ODEs) governed by skew-adjoint integral operators. Using neural networks to define finite-rank generators of such ODEs allows us to parameterize and optimize orthonormal bases in function space. While relying on finite-rank generators to model infinite operators might seem restrictive, we prove a universality result: even with a rank-2 generator, the integrated solutions of the ODE are dense in the orthogonal group under the appropriate operator topology. In other words, for any target orthonormal basis, there exists a path originating from a reference basis and driven by finite-rank generators that gets arbitrarily close to that target basis. We demonstrate the flexibility of our framework by transforming the Fourier basis into the principal components of a functional dataset, eigenfunctions of linear operators, or dynamic modes of energy-preserving physical simulations.


Learning Polyhedral Conformal Sets for Robust Optimization

Shuyi Chen ⋅ Wenbin Zhou ⋅ Shixiang Zhu

Robust optimization is a widely used framework for decision-making under uncertainty, particularly in high-stakes applications where reliability is critical. A key challenge in this paradigm lies in constructing uncertainty sets that balance robustness and performance: overly conservative sets lead to pessimistic decisions, while insufficient coverage risks failure in practice. Recent approaches based on conformal prediction provide finite-sample, distribution-free guarantees for uncertainty sets, but remain largely task-agnostic and disconnected from downstream decision objectives. In this paper, we propose a decision-aware conformal prediction framework that directly learns the geometry of uncertainty sets to improve robust decision-making. Our approach introduces a polyhedral nonconformity score that induces feature-dependent uncertainty sets, and a three-step procedure that integrates conformal calibration, robust-decision-aware learning, and re-calibration to correct for post-selection bias. We establish finite-sample coverage guarantees for the final, data-dependent uncertainty set, while achieving improved decision performance by reducing unnecessary conservativeness. This work bridges the gap between statistical validity and decision optimality, providing a principled framework for data-driven robust optimization.


Learning the Signature of Memorization in Autoregressive Language Models

David Ilić ⋅ Kostadin Cvejoski ⋅ David Stanojević ⋅ Evgeny Grigorenko

All prior membership inference attacks for fine-tuned language models use hand-crafted heuristics (e.g., loss thresholding, Min-K\%, reference calibration), each bounded by the designer's intuition. We introduce the first transferable learned attack, enabled by the observation that fine-tuning any model on any corpus yields unlimited labeled data, since membership is known by construction. This removes the shadow model bottleneck and brings membership inference into the deep learning era: learning what matters rather than designing it, with generalization through training diversity and scale. We discover that fine-tuning language models produces an invariant signature of memorization detectable across architectural families and data domains. We train a membership inference classifier exclusively on transformer-based models. It transfers zero-shot to Mamba (state-space), RWKV-4 (linear attention), and RecurrentGemma (gated recurrence), achieving 0.963, 0.972, and 0.936 AUC respectively. Each evaluation combines an architecture and dataset never seen during training, yet all three exceed performance on held-out transformers (0.908 AUC). These four families share no computational mechanisms, their only commonality is gradient descent on cross-entropy loss. Even simple likelihood-based methods exhibit strong transfer, confirming the signature exists independently of the detection method. Our method, Learned Transfer MIA (LT-MIA), captures this signal most effectively by reframing membership inference as sequence classification over per-token distributional statistics. On transformers, LT-MIA achieves 2.8$\times$ higher true positive rate at 0.1\% false positive rate than the strongest baseline. The method also transfers to code (0.865 AUC) despite training only on natural language texts. Our results imply that leakage from memorization is intrinsic to cross-entropy training; architectural innovation within this paradigm did not escape our attack. Code and the trained classifier are provided.


Learning to Persuade a Biased Receiver

Yuqi Pan ⋅ Sadie Zhao ⋅ Milind Tambe ⋅ Yiling Chen

We study a repeated information design setting in which the receiver, who is also the decision-maker, updates beliefs in a systematically biased way. More specifically, a distorted posterior in our model can be written as a convex combination of the prior and the Bayesian posterior, governed by a fixed but unknown parameter. Over repeated interactions, the sender chooses persuasive signaling schemes, observes only the receiver’s realized actions, and seeks to minimize regret relative to a full-information oracle that knows the receiver’s biased updating rule. We propose a safe exploration algorithm for learning the receiver’s bias while maintaining high persuasion value. The algorithm exploits the asymmetric cost of probing: conservative probes incur only local loss, whereas overly aggressive probes may lose the persuasive opportunity entirely. For general finite state and action spaces and arbitrary bounded utilities, our method achieves $O(\log\log T)$ regret. A matching $\Omega(\log\log T)$ lower bound shows that this rate is optimal. We further discuss the influence on receiver welfare, as well as extensions to jointly unknown prior and bias, and contextual settings with time-varying priors and utilities.


Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs

Nimet Beyza Bozdag ⋅ Emre Can Acikgoz ⋅ Gokhan Tur ⋅ Dilek Hakkani-Tur

Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior. Yet we show that this requirement is far from met: a single targeted persuasive argument is enough to collapse model accuracy to near zero, even when the argument is factually false. We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction. First, we show that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses: RL-trained persuaders raise persuasion success from approximately 24\% to over 93\% against the training-time persuadee. Second, we find that these learned strategies transfer to unseen models, achieving 83\% attack success on Qwen-14B, 79\% on Llama-3.1-8B, and 25\% on GPT-4o-mini. Third, we demonstrate that a curriculum that bootstraps on more persuadable open-weight models before targeting harder models further increases GPT-4o-mini attack success from 25\% to 38\%. Moreover, our results reveal that optimized persuaders increasingly rely on credibility-based tactics, including fabricated citations and false authoritative evidence. Together, these findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence. This positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.


Learning to Price with Persuasion

Maria-Florina Balcan ⋅ Tejas Pagare ⋅ Karan Singh

Motivated by modern marketplaces, where the platform or the seller routinely gathers detailed user profiles, we study a novel learning theoretic model that simultaneously involves information and mechanism design. Specifically, we consider the economic setting recently introduced by Bergemann et al, where in addition to the menu of quality-price pairs, the seller offers information on the value of the match between product quality and buyer's taste via a signaling scheme. We relax the assumption that the seller knows the buyers' belief about the distribution of tastes and study the sample requirements of designing a revenue maximizing scheme. We consider both the batch setting where we have access to data from a set of i.i.d. buyers and an online demand query model where we observe the buyers behaviors to seller's schemes. Despite the apparent non-convexity of the problem, we also give the first FPTAS to compute a scheme that maximizes the revenue within an arbitrarily small additive loss, which was unknown even in the previous result of Bergemann et al. Overall, this brings a new learning perspective in asymmetric economic settings where buyers and sellers know different types of information.


Learning Visual Feature-Based World Models via Residual Latent Action

Xinyu Zhang ⋅ Zhengtong Xu ⋅ Yutian Tao ⋅ Yeping Wang ⋅ Yu She ⋅ Abdeslam Boularias

World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video pixels, offering a promising alternative that is more efficient and less prone to hallucination. However, current feature-based approaches rely on direct regression, which leads to blurry or collapsed predictions in complex interactions, while generative modeling in high-dimensional feature spaces still remains challenging. In this work, we discover that a new type of latent action representation, which we refer to as Residual Latent Action (RLA), can be easily learned from DINO residuals. We also show that RLA is predictive, generalizable, and encodes temporal progression. Building on RLA, we propose RLA World Model (RLA-WM), which predicts RLA values via flow matching. RLA-WM outperforms both state-of-the-art feature-based and video-diffusion world models on simulation and real-world datasets, while being orders of magnitude faster than video diffusion. Furthermore, we develop two robot learning techniques that use RLA-WM to improve policy learning. The first one is a minimalist world action model with RLA that learns from actionless demonstration videos. The second one is the first visual RL framework trained entirely inside a world model learned from offline videos only, using a video-aligned reward and no online interactions or handcrafted rewards.


LEIA: Learned Environment for Interactive Architected Materials

Haiqian Yang ⋅ Yuan Cao ⋅ Markus Buehler

World models have enabled interactive exploration of game environments and robotic manipulation, but physical engineering remains beyond their reach: real materials exhibit nonlinear constitutive laws, carry history-dependent internal state, undergo inertial dynamics, and may possess hierarchical structures spanning multiple length scales. We present LEIA (Learned Environment for Interactive Architected materials), a world model that lets engineers apply boundary conditions step by step and observe the resulting deformation and stress fields in real time. LEIA handles large three-dimensional unstructured meshes and generates autoregressive responses to user-specified loading. We introduce MicroPlate, a benchmark of architected plates spanning two regimes of microstructure modeling: architected lattices that resolve microstructure explicitly through three-dimensional geometry, and a homogeneous plate where microstructural change is modeled implicitly through internal degrees of freedom. MicroPlate is used to assess LEIA alongside four baseline methods across both regimes. Finally, we demonstrate that LEIA enables efficient candidate generation and ranking for fast surrogate-guided search for de novo designs of architected materials, with stress-accurate candidate ranking validated by finite element ground truth.

Speech is an increasingly promising biomarker for neurodegenerative disorders due to its non-invasive nature, low cost, and suitability for frequent, longitudinal monitoring. However, existing clinical speech datasets are typically small, sparse, and irregularly sampled, limiting the ability to model continuous disease progression and develop robust biomarkers. We propose L-Flow, a progression-aware conditional speech generation framework for longitudinal speech trajectory completion. Given a patient’s speech recordings, timestamps, and clinical progression labels, L-Flow learns a query-specific progression representation to synthesize speech at intermediate time points within the observed recording span. Experiments on two real-world longitudinal speech datasets demonstrate that L-Flow generates synthetic speech with stronger progression consistency than eight baseline augmentation methods. In addition, incorporating L-Flow–augmented data improves downstream regression performance to predict clinical severity, highlighting its utility for speech-based biomarker development.


LIBERO-PeRM: Benchmarking Personalized Robotic Manipulation

Zhixu Li ⋅ Keqian Tang ⋅ Litian Gong ⋅ Jingyu Yao ⋅ Wenqian Zhang ⋅ Tianze Xu ⋅ Zehao Wang ⋅ Yuping Wang ⋅ Junge Zhang ⋅ Jiachen Li

Personalization is essential for general-purpose robots to transition into household environments and achieve true practicality. While general robotic manipulation has been extensively studied, personalized robotic manipulation (PeRM) lacks a unified formulation and remains significantly underexplored. In this paper, we identify user preference as the cornerstone of PeRM and introduce LIBERO-PeRM, a novel benchmark designed for personalized preference learning in embodied agents. Specifically, LIBERO-PeRM highlights five key research topics in personalization: 1) how policies follow preferences in zero-shot settings; 2) how to learn single preferences from demonstrations; 3) how efficiently policies acquire preferences as demonstrations increase; 4) how learned preferences generalize across layouts and preference settings; and 5) how policies compose multiple preferences and resolve conflicts. To this end, we propose a unified formulation of personalized manipulation with three preference levels that jointly cover common daily personalization scenarios. Building on this formulation and the LIBERO framework, we develop an extensible procedural generation pipeline that generates paired demonstrations with and without preferences. For benchmarking purposes, we create five task suites (312 tasks in total) with corresponding high-quality demonstrations, rich metadata, and preference satisfaction metrics to probe the aforementioned research topics. Experimental results show that current VLA policies struggle with zero-shot personalization, while SFT makes preferences learnable; preference learning benefits from mixed training, broader layout diversity, and same-layout no-preference data, but generalization, decomposition, and conflict resolution remain difficult.


LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science

Alireza Salemi ⋅ Mihir Parmar ⋅ Palash Goyal ⋅ Yiwen Song ⋅ Jinsung Yoon ⋅ Hamed Zamani ⋅ Tomas Pfister ⋅ Hamid Palangi

Advances in large language models (LLMs) have created new opportunities in data science, but their deployment is often limited by the challenge of finding relevant data in large data lakes. Existing methods struggle with this: both single- and multi-agent systems are quickly overwhelmed by large, heterogeneous files, and master–slave multi-agent systems rely on a rigid central controller that requires precise knowledge of each sub-agent’s capabilities, which is not possible in large-scale settings where the main agent lacks full observability over sub-agents’ knowledge and competencies. We propose a novel multi-agent paradigm inspired by the blackboard architecture for traditional AI models. In our framework, a central agent posts requests to a shared blackboard, and autonomous subordinate agents--either responsible for a partition of the data lake or retrieval from web--volunteer to respond based on their capabilities. This design improves scalability and flexibility by removing the need for a central coordinator to know each agent’s expertise or internal knowledge. We evaluate the approach on three benchmarks that require data discovery: KramaBench and modified versions of DSBench and DA-Code. Results show that the blackboard architecture substantially outperforms strong baselines, achieving 13%–57% relative improvements in end-to-end success and up to a 9% relative gain in data discovery F1 over the best baseline.


MAdam: Metric-Aware Multi-Objective Adam

Fengbei Liu ⋅ Rachit Saluja ⋅ Sunwoo Kwak ⋅ Ruibo Wang ⋅ Ruining Deng ⋅ Heejong Kim ⋅ Johannes C. Paetzold ⋅ Mert Sabuncu

Multi-objective optimization (MOO) underlies many machine learning problems, yet MOO solvers across the loss-balancing, gradient-balancing, and Pareto-based families almost universally hand their reconciled directions to Adam~\cite{kingma2015adam}. We show this coupling introduces two systematic gaps between the solver's intent and the optimizer's execution. The first is a \emph{weighting mismatch}: Adam's second-moment denominator entangles the time-varying preference vector with gradient statistics, marginalizing the preference into a history average and collapsing distinct Pareto trade-offs toward a near-uniform mixture. The second is a \emph{geometric mismatch}: Adam's adaptive metric distorts the Euclidean geometry MOO solvers assume, turning aligned objectives into apparent conflicts. To resolve both jointly, we introduce \textbf{MAdam} (Metric-Aware Multi-Objective Adam), a drop-in wrapper that leaves both solver and optimizer unchanged. MAdam preconditions the reconciled direction by the preference-conditioned curvature of the scalarized objective; on this whitened input, Adam's second moment collapses to identity, so the realized update is governed by the preference-conditioned metric. Across multi-task learning, Pareto-front recovery, physics-informed neural networks, and medical imaging, MAdam consistently improves over Adam for every solver family.


Magnetic Resonance Unpaired Image Translation with Pseudometric Schrödinger Bridges

Shuwen Wei ⋅ Samuel Remedios ⋅ Zhangxing Bian ⋅ Shimeng Wang ⋅ Jinwei Zhang ⋅ Junyu Chen ⋅ Yihao Liu ⋅ Lianrui Zuo ⋅ Muhammad Faizyab Ali Chaudhary ⋅ Blake Dewey ⋅ Aaron Carass ⋅ Jerry L Prince

Unpaired image translation is a critical preprocessing strategy for synthesizing missing sequences and harmonizing contrast variability in magnetic resonance (MR) imaging. However, in medical image processing, distribution-level translation is insufficient; it is vitally important that the anatomy is not distorted, corrupted, or lost after translation. Previous methods, like diffusion Schrödinger bridge matching (DSBM) achieve image translation, but do not preserve anatomy. We introduce the pseudometric Schrödinger bridge (PMSB), which trains two independent volume-preserving, isometry-regularized diffeomorphisms to map data from each domain to a latent anatomic pseudometric space. Then, DSBM is used within that space, explicitly minimizing the transport gap while preserving anatomy. Extensive experiments on the multi-contrast OASIS-3 dataset demonstrate that PMSB establishes a new state-of-the-art in structural consistency for cross-contrast synthesis. PMSB outperforms other leading methods and drastically reduces anatomical hallucinations, notably yielding a mean PSNR of 26.70 dB, a mean SSIM of 0.86, and a mean LNCC of 0.24.


Manifold Sampling via Entropy Maximization

Cornelius Braun ⋅ Tilman Burghoff ⋅ Marc Toussaint

Sampling from constrained distributions has a wide range of applications, including in Bayesian optimization and robotics. Prior work establishes convergence and feasibility guarantees for constrained sampling, but assumes that the feasible set is connected. However, in practice, the feasible set often decomposes into multiple disconnected components, which makes efficient sampling under constraints challenging. In this paper, we propose MAnifold Sampling via Entropy Maximization (MASEM) for sampling on a manifold with an unknown number of disconnected components, implicitly defined by smooth equality and inequality constraints. The presented method uses a resampling scheme to maximize the entropy of the empirical distribution based on k-nearest neighbor density estimation. We show that, in the mean field, MASEM decreases the KL-divergence between the empirical distribution and the maximum-entropy target exponentially in the number of resampling steps. We instantiate MASEM with multiple local samplers and demonstrate its versatility and efficiency on synthetic and robotics-based benchmarks. MASEM enables fast and scalable mixing across a range of constrained sampling problems, improving over alternatives by an order of magnitude in Sinkhorn distance with competitive runtime.


MaRiO: Multi-agent Collaborative Reasoning via Shared Observations in MLLMs

Nathaniel Redmond ⋅ Fidel Omar Tito Cruz ⋅ Devansh Sharma ⋅ Shehreen Azad ⋅ Sirshapan Mitra ⋅ Shruti Vyas ⋅ Yogesh Rawat

Multimodal Large Language Models (MLLMs) show strong visual perception capabilities, but existing evaluations largely focus on single-view or multi-view settings with shared camera parameters, leaving the challenge of {multi-agent collaborative} scenarios underexplored. In such settings, multiple agents observe a shared environment from independent viewpoints, and reliable decisions require reasoning across these independent observations of the same environment. To study this problem, we introduce MaRiO, a benchmark for evaluating multi-agent collaborative reasoning, and evaluate 23 state-of-the-art MLLMs on it. Our analysis reveals three key findings: (1) spatial tasks are generally easier than geometric reasoning tasks; (2) most models struggle to resolve cross-agent correspondences, with errors increasing with scene complexity; and (3) models show limited ability to leverage geometric cues present in the scene. As an initial step toward addressing these challenges, we explore geometry-guided synthetic view augmentation that generates intermediate transition frames between agent observations, providing auxiliary spatial context that improves performance in certain categories. Overall, our results highlight that multi-agent collaborative reasoning remains an open challenge and an important direction for multimodal embodied systems.

Transition state (TS) search is the rate-limiting step in high-throughput kinetics screening. Generative models have been used to accelerate TS prediction for molecules, but extending them to periodic materials has been blocked by the absence of a large-scale materials TS dataset. We release MaterialsSaddles, the first such dataset: 34.14 M (reactant, saddle, product) triplets generated with the state-of-the-art UMA-S-1.2 universal interatomic potential across four subsets, comprising 31.35 M triplets over LeMat-Bulk's 5.34 M-structure deduplication of Materials Project, OQMD, and Alexandria; 2.59 M triplets over OC20 surfaces; 167 k triplets over OC22 oxide electrocatalysts; and 35 k climbing-image nudged elastic band (CI-NEB) paths over Materials Project battery cells. The release totals 102.4 M structures (687 GB) in ASE-LMDB shards with stratified 90/5/5 splits, generated on a 500,000 GPU-hour HPC pipeline. Every triplet is connected by construction via double-minimization from the converged saddle, and every frame carries per-frame metadata including eigenmodes, curvatures, bond-change diffs, and source identifiers. As a working demonstration of the dataset's intended use, we release SaddleFlow, a flow-matching reactant-and-product-conditional saddle-point generator built on a UMA-S-1.2 backbone with FiLM-based time conditioning and an SO(3)-equivariant velocity head, trained on the mp20bat subset of MaterialsSaddles. In a pilot DFT study, we show that dimer searches initialized from a sample of mp20bat test triplets converge substantially faster and more reliably than those initialized from the standard reactant-product midpoint guess, indicating that MaterialsSaddles and SaddleFlow together offer a practical path to scaling DFT-level transition-state searches. MaterialsSaddles (CC-BY 4.0), SaddleFlow (MIT), and the data-generation engine SaddleMill (MIT) are publicly released.


MC$^2$Mark: Distortion-Free Multi-Bit Watermarking for Long Messages

Xuehao Cui ⋅ Ruibo Chen ⋅ Yihan Wu ⋅ Weidong Cai ⋅ Heng Huang

Large language models now produce text indistinguishable from human writing, which increases the need for reliable provenance tracing. Multi-bit watermarking can embed identifiers into generated text, but existing methods struggle to keep both text quality and watermark strength while carrying long messages. We propose MC$^2$Mark, a distortion-free multi-bit watermarking framework designed for reliable embedding and decoding of long messages. Our key technical idea is Multi-Channel Colored Reweighting, which encodes bits through structured token reweighting while keeping the token distribution unbiased, together with Multi-Layer Sequential Reweighting to strengthen the watermark signal and an evidence-accumulation detector for message recovery. Experiments show that MC$^2$Mark improves detectability and robustness over prior multi-bit watermarking methods while preserving generation quality, achieving near-perfect accuracy for short messages and exceeding the second-best method by nearly 30% for long messages.

Autonomous security agents use language models to inspect code, call tools, and test vulnerabilities inside authorized environments. Most safety evaluations ask whether a model refuses harmful single-turn requests. We study a different failure mode: whether safety alignment changes the evidence-grounded behavior of an agent that is already operating in a local sandbox with fixed tools and success checks. We evaluate regular safety-aligned Gemma 4 models and uncensored Gemma 4 derivatives in the same security-agent harness. Each model receives authorized vulnerability-analysis tasks, and each run is scored from saved traces rather than self-reported success. We measure task completion, refusals, unsafe actions, and whether the final artifact is grounded in the relevant files, symbols, and vulnerability evidence. The key finding is that the gap is not mainly visible refusal. The aligned Gemma 4 conditions usually continue working, but their artifacts are less likely to satisfy security-specific evidence checks. The uncensored Gemma 4 condition more often finds the relevant code, identifies reachability, and writes usable security reports. Clear authorization in the prompt does not recover this behavior. At the same time, all conditions fail the hardest proof-of-trigger and patch-verification tasks. These results suggest that safety alignment effects in autonomous security agents cannot be understood only as refusals: they also appear as weaker grounding and less specific defensive evidence. Removing alignment recovers some behavior, but does not make the agent reliable or safe.


MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments

Yicheng Gao ⋅ Xiaolin Zhou ⋅ Yahan Li ⋅ Yue Zhao ⋅ Ruishan Liu

Real-world clinical diagnosis is a complex process in which the doctor is required to obtain information from both interaction with the patient and conducting medical exams. Additionally, the doctor needs to adapt to different patient personas, as well as noisy and incomplete information that can happen at any time during the process. However, existing benchmarks for medical LLMs and methods for automatic diagnosis largely simplify this process by reducing it to single-turn question answering, noise-free conversations, or sequential exam making, etc., ignoring the interactive and uncertain nature of clinical diagnosis. In this paper, we aim to address this gap by formalizing clinical diagnosis as a Partially Observable Markov Decision Process (POMDP) with three action types: questioning the patient, ordering medical exams as tool calls, and issuing a diagnosis. We also introduce a systematic noise model comprising seven patient noise types and three exam noise types. Using our proposed environment, we train an effective diagnosis agent, \textbf{MedExAgent}, through a two-stage pipeline that first performs supervised finetuning on synthetic conversations structured after the Calgary-Cambridge model for clinical interviews, and then applies DAPO to optimize a composite reward capturing diagnostic accuracy, tool call quality, and exam cost including financial cost and patient discomfort. Through extensive experiments and ablation studies, we demonstrate that MedExAgent achieves diagnostic performance comparable to larger models while maintaining cost-efficient examination strategies.


Memory Retrieval for Changing Preferences

Yuehan Qin ⋅ Li Li ⋅ Linxin Song ⋅ Jiate Li ⋅ Wei Yang ⋅ Yuqing Yang ⋅ Yue Zhao

Long-context dialogue systems must decide both when to access memory and which parts of the interaction history are relevant. Existing approaches typically rely on heuristic retrieval signals or always-on memory usage, failing to account for the changing and potentially inconsistent nature of user preferences. In this work, we propose a unified framework for memory access and selection based on changing preferences. We formulate personalized memory retrieval as identifying which historical turns provide evidence about a user’s latent preference state, rather than relying on surface-level semantic similarity. To this end, we quantify the utility of each memory turn using a Bayes factor, defined as the improvement in the model’s likelihood of the reference response when the turn is included in context. This provides a principled measure of evidence strength and a unified signal for both memory access and selection. By framing memory retrieval as utility estimation, the model learns to identify salient turns and regulate memory usage based on expected utility. Experiments on four heterogeneous memory benchmarks show that our approach outperforms existing embedding-based retrieval on long-context, preference-intensive tasks where modeling changing preferences is essential, while remaining competitive in low-density regimes where semantic similarity suffices.


Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization

Kaishen Wang ⋅ Tong Zheng ⋅ Xuehao Cui ⋅ Ruibo Chen ⋅ Tianyi Xiong ⋅ Heng Huang

Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such thinking often improves overall performance by helping the model recover relevant knowledge and refine its answers. However, we find that this benefit is not uniform at the instance level: explicit thinking can also overturn correct non-thinking answers and lead to factual drift. We refer to this failure mode as \emph{thinking-induced hallucination}. To explain this phenomenon, we formulate explicit thinking in factuality QA as a thinking residual over the model's direct-answer tendency, which can either recover missing knowledge or introduce unsupported associations. Based on this formulation, we propose MARGO, \underline{\textit{M}}ixed-Mode \underline{\textit{A}}dvantage \underline{\textit{R}}egularization for \underline{\textit{G}}rounded \underline{\textit{O}}ptimization, a reinforcement learning framework that uses non-thinking rollouts as same-model references in advantage estimation. By constructing mixed-mode rollout groups with both thinking and non-thinking trajectories, MARGO evaluates whether explicit thinking adds factual value beyond direct answering, thereby suppressing hallucination-prone thinking while preserving beneficial thinking behaviors. Experiments across multiple factuality-oriented QA benchmarks demonstrate that MARGO improves factual reliability over strong baselines, while evaluations on mathematical benchmarks show that it preserves general reasoning ability.

Object hallucination, where large vision-language models (LVLMs) generate objects not supported by the visual input, is a persistent challenge caused by visual uncertainty during decoding. Existing methods reduce hallucinations using contrastive signals, but they rely on heuristics and lack principled control of false positives at the image level. To address this, we propose False Discovery Rate-\textbf{CO}nt\textbf{R}ol of {O}bject H\textbf{AL}lucination (CORAL), a training-free framework that models visual uncertainty using an uncertainty-aware visual data splitting strategy and leverages {mirror statistics} to quantify visual contrast during decoding. By computing mirror statistics from paired, symmetrically perturbed visual inputs, CORAL estimates spurious object predictions and sets a data-driven threshold to control the expected fraction of false discoveries per image, suppressing hallucinations while retaining high power for truly grounded objects. The framework is flexible, supports multiple LVLMs, and mitigates hallucinations without retraining or supervision. Extensive experiments on multiple benchmarks with several evaluation metrics demonstrate that CORAL consistently outperforms state-of-the-art methods, providing more reliable and robust hallucination control.


Mixture-of-Chains: Learning Causal Graphs from Human Knowledge

Yuantao Wei ⋅ Huiling Liao ⋅ Ryan (Feng) Lin ⋅ Xiaoning Qian ⋅ Shuai Huang

Human causal knowledge offers a valuable yet underexplored source of structural information for causal discovery. However, eliciting and using such knowledge can be nontrivial, since causal beliefs are typically heterogeneous across individuals and human judgments can be noisy. Existing methods tend to treat human input as auxiliary constraints or priors, failing to capture the underlying causal structure and frequently yielding unreliable or cyclic graphs. In this paper, we propose Mixture-of-Chains (MOC), a framework for directly learning causal graphs from human causal judgments by modeling them as a mixture of latent causal chains. Each chain represents a coherent pathway in human belief structures, allowing MOC to capture variability across reasoning contexts or different individuals. The framework adaptively extracts these chains from noisy inputs and integrates them into a globally consistent causal graph. Experiments on synthetic Bayesian network benchmarks and real human data show that MOC can effectively recover the ground-truth causal structures with strong robustness to input noise and inter-subject variability, supported by the theoretical analysis for recovery conditions. These results suggest that human knowledge can serve not merely as a supplement to data-driven methods, but as a primary source for causal graph learning, and highlight the promise of chain-level modeling for reliable causal discovery.

Online decision-making faces a tension between efficient inference and long-horizon reasoning. Value-based methods rely on short-horizon bootstrapping that limits the propagation of long-term information, while trajectory optimization methods plan explicitly but must solve a high-dimensional optimization problem at every decision step. We propose Generative Trajectory Planning (GTP), a model-based framework that addresses this tension by learning a generative model as a proposal distribution over action sequences. Instead of optimizing sequences from scratch, GTP samples candidate sequences from a diffusion prior and refines them using learned dynamics, reward, and value models, enabling efficient trajectory-level reasoning. Across continuous control and long-horizon manipulation tasks, GTP matches or exceeds state-of-the-art model-free, diffusion-policy, and model-based baselines, and remains stable in regimes where strong baselines degrade. Our results suggest that generative models are best understood not as policies, but as structured proposal distributions for planning, offering a new perspective on how to integrate generative modeling with model-based decision-making.


MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction

Jianing Zhang ⋅ Chenhao Zheng ⋅ Yajun Yang ⋅ Rustin Soraki ⋅ Winson Han ⋅ Chun-Liang Li ⋅ Jason Ren ⋅ Max Argus ⋅ Jieyu Zhang ⋅ Ranjay Krishna

Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan actions, reason about physical interactions, and synthesize realistic futures. We argue that 3D points in world coordinates provide a general representation that is class-agnostic, view-stable, compact, and directly useful for downstream tasks. We formalize the task of goal-conditioned 3D point motion forecasting: given a short visual history, a set of 3D query points on an object of interest, and a language description of the intended goal, the model predicts the future 3D trajectory of each point. We introduce a full stack to study this task at scale: (1) MolmoMotion-1M is a large corpus of action-described, object-grounded 3D point trajectory dataset annotated from 1.16M unconstrained videos; (2) PointMotionBench is a human-verified benchmark spanning 111 object categories and 61 motion types; and (3) MolmoMotion is a general motion forecasting model that supports both autoregressive coordinate prediction and flow-matching-based trajectory generation. MolmoMotion is able to accurately predicts diverse motion patterns with different language instructions, and significantly outperforms all existing motion prediction baselines on PointMotionBench. Finally, we show that the learned 3D motion prior transfers well to downstream applications: it improves training efficiency and generalization for robot manipulation, and its predicted trajectories provide effective motion guidance for generative models to synthesize videos with more realistic object motion.


MoSE3: Learning World-Space SE(3) at Every Pixel

Jiahuan(Joanna) Cheng ⋅ Zhiyi Li ⋅ Tian Xia ⋅ Ruojin Cai ⋅ Yilun Du ⋅ Qianqian Wang

Dense 2D/3D point tracking has been the dominant paradigm for modeling motion in dynamic scenes, but a point track is just a 3-DoF translation curve per pixel: it captures where pixels go, not the rotation of the underlying part, nor which pixels move together as one body. We propose MoSE3, the first feed-forward model that predicts dense SE(3) motion from monocular RGB video, producing full 6-DoF rigid transforms at every pixel in world space. Per-pixel SE(3) motion offers a richer view of how a scene moves: rotation, translation, and grouping all at once. Directly predicting SE(3) is challenging: rotations lie on a curved manifold that is ill-suited to Euclidean regression, and annotations for SE(3) are particularly difficult to acquire. To address these challenges, MoSE3 predicts per-pixel SE(3) through two jointly learned intermediates, 3D point tracks and rigidity embeddings, and recovers SE(3) by differentiably fitting transforms within each soft rigid cluster, enabling end-to-end prediction and supervision. To close the data gap, we introduce Art-Kubric, a large-scale synthetic dataset with dense SE(3) and rigidity labels for articulated scenes with rich physical interactions. MoSE3 achieves state-of-the-art SE(3) estimation at pixel, part, and object levels on both rigid and articulated benchmarks, and state-of-the-art 3D point tracking on three datasets, while showing strong generalization to real-world videos despite being trained solely on synthetic motion data.

In this work, we show that natural policy gradient (NPG), a core algorithm in reinforcement learning, admits an exact interpretation as a smoothed and averaged form of policy iteration. Specifically, we introduce doubly smoothed policy iteration (DSPI), a Bellman-operator framework in which each policy is obtained by applying a regularized greedy step to a weighted average of past $Q$-functions. DSPI includes policy iteration, dual-averaged policy iteration, NPG, and more general policy dual averaging methods as special cases. Using only monotonicity and contraction of smoothed Bellman operators, we prove distribution-free global geometric convergence of DSPI. Consequently, standard NPG and policy dual averaging achieve an iteration complexity of $\mathcal{O}((1-\gamma)^{-1}\log((1-\gamma)^{-1}\epsilon^{-1}))$ for computing an $\epsilon$-optimal policy, without modifying the MDP, adding regularization beyond the mirror map inherent in the update, or using adaptive, trajectory-dependent stepsizes. For the unregularized greedy case, we also prove finite termination of dual-averaged policy iteration. The same Bellman-operator framework extends to discounted MDPs with linear function approximation and to stochastic shortest path problems.


Near-Optimal Sample Complexity of Robust Reinforcement Learning with KL Uncertainty Set

Yudan Wang ⋅ Zilong Deng ⋅ Nathaniel D Bastian ⋅ Shaofeng Zou

In this paper, we study distributionally robust reinforcement learning with Kullback--Leibler (KL) divergence defined uncertainty set. The goal is to find a policy that maximizes the robust value function, defined as the worst-case value over all transition kernels in the uncertainty set. Assuming access to a generative model, we aim to understand the sample complexity of finding an $\epsilon$-optimal robust policy. The best-known sample complexity results in the literature show a non-trivial gap of $\mathcal{O}\{\max\{p_\wedge^{-1}(1-\gamma)^{-1},(1-\gamma)^{-2}\}\}$ between the upper and lower bounds, where $p_\wedge$ denotes minimal non-zero support of the nominal transition kernel, and $\gamma$ is the discount factor. More importantly, existing results on the lower bound only cover a limited range of the uncertainty level. In this paper, we develop tighter and complete upper and lower bounds for robust RL with KL-defined uncertainty set. Our upper bound is the tightest among all existing studies, and it improves upon the best known bound by at least the order of $\mathcal{O}(\min\{p_\wedge^{-1},(1-\gamma)^{-1}\})$. Furthermore, our lower bound holds for any uncertainty level. Our upper and lower bounds (nearly) match under various cases, providing near minimax optimality results for this problem.


Neglected Free Lunch from Post-training: Progress Advantage for LLM Agents

Changdae Oh ⋅ Wendi Li ⋅ Seongheon Park ⋅ Samuel (Min-Hsuan) Yeh ⋅ Tanwi Mallick ⋅ Sharon Li

Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and stochastic environment feedback make both human annotation and Monte Carlo estimation infeasible at scale. In this work, we show that reinforcement learning (RL) post-training already provides the ingredients for effective step-level scoring, eliminating the need for dedicated reward model training altogether. Concretely, we derive an implicit advantage under a general stochastic Markov Decision Process, which we term progress advantage---log-probability ratio between the RL-trained policy and its reference policy exactly recovers the optimal advantage function. This signal is annotation-free, domain-agnostic, and available as a byproduct of the standard RL post-training pipeline. We validate the effectiveness of the progress advantage across three different applications: test-time scaling, uncertainty quantification, and failure attribution on five benchmarks and four model families. Across all settings, progress advantage consistently outperforms confidence-based baselines and, despite requiring no task-specific training, surpasses dedicated trained reward models. We complement these results with deeper analyses on characteristics of progress advantage, offering practical guidance for adoption in real-world agentic systems.


Neural Expansion: A Unified Mechanism for How Deep Neural Network Generalize

Chashi Mahiul Islam ⋅ Samuel Jacob Chacko ⋅ Mao Nishino ⋅ Canlin Zhang ⋅ Xiuwen Liu

Deep neural networks demonstrate remarkable generalization despite being highly over-parameterized, challenging traditional statistical learning theories. While several theoretical paradigms such as margin-based bounds, PAC-Bayesian analyses, and neural tangent kernels provide global explanations, they often lack mechanistic insight into how generalization emerges from training dynamics. We introduce Neural Expansion, a unified geometric framework that characterizes generalization through the expansion of directionally-selective ray coverage in the input space. Using Generalization Intervals (GIs), we quantify direction-specific regions around training samples where the model's predictions remain stable. Our extensive empirical studies across diverse architectures and datasets reveal that training leads to progressive expansion of ray coverage, particularly along directions with low Jacobian singular values. This mechanism helps explain a large fraction of correctly classified test inputs, adversarial examples, and even mislabeled or out-of-distribution samples through the geometry of ray coverage rather than instance-level memorization. By connecting input-space curvature to predictive behavior, our framework provides a mechanistic foundation that complements existing global theories of deep learning. The link to our code and data will be made available in the published version.


Neural Field Thermal Tomography: A Differentiable Physics Framework for Non-Destructive Evaluation

Tao Zhong ⋅ Yixun Hu ⋅ Dongzhe Zheng ⋅ Aditya Sood ⋅ Christine Allen-Blanchette

Inverse problems for stiff parabolic partial differential equations (PDEs), such as the inverse heat conduction problem (IHCP), are severely ill-posed: the forward map rapidly damps high-frequency interior structure before it reaches the boundary. Soft-constrained physics-informed neural networks (PINNs), which embed the PDE as a residual penalty, suffer from gradient pathology in this regime and tend to fit boundary measurements while leaving the interior field essentially untouched. We propose Neural Field Thermal Tomography (NeFTY), a hard-constrained neural field framework for label-free three-dimensional inverse heat conduction. NeFTY represents the unknown diffusivity as a continuous coordinate-based neural network, and at every optimization step passes the candidate field through a differentiable implicit-Euler heat solver with harmonic-mean interface flux, so that the governing PDE holds exactly on the discretization rather than as a soft penalty. Adjoint gradients propagate the surface reconstruction error back to the network weights at solver-level memory cost, making test-time inversion tractable on a single GPU. Across synthetic 3D benchmarks, NeFTY substantially outperforms soft-constrained PINN variants and a voxel-grid baseline on label-free volumetric recovery, and it transfers to real thermography data, surpassing classical signal-processing baselines in both defect segmentation and depth estimation.


Neural Harmonic Measure Operator

Jinjin He ⋅ Sinan Wang ⋅ Yuchen Sun ⋅ Bo Zhu

We introduce Neural Harmonic Measure Operator (NHMO), a neural solver for elliptic PDE problems on variable-shape domains. The harmonic measure of a domain is the boundary probability distribution that, integrated against any boundary data, returns the Dirichlet Laplace solution. It depends only on the geometry, not on the boundary data. NHMO parameterizes this measure as a transformer-based boundary kernel supervised by Walk-on-Spheres exit samples, so one trained kernel handles different boundary value on a shape with no retraining. We extend it to Poisson via a classical decomposition, with an auxiliary network amortizing the source-induced correction and avoiding the singular volume quadrature that breaks direct evaluation. At inference, new boundary values and new sources both yield PDE solutions by re-integration against the fitted kernel and lift, with no retraining. NHMO improves over four prior baselines on the MCB-B 3D variable-shape Poisson benchmark across all five categories, and is competitive with major neural-operator baselines on a controlled 2D testbed.

Many engineering building blocks behave as multi-port linear time-invariant systems. RF cavities, photonic devices, and superconducting quantum chips, despite their different underlying physics, all share a common mathematical structure for their port-level response. Each entry of the response matrix is a sum of contributions from a small number of intrinsic resonant modes, the pole-residue form. A model capable of predicting such responses for arbitrary geometries and arbitrary port configurations, while simultaneously extracting the underlying eigenmode structure, would therefore establish a foundational design principle spanning all these domains. We propose a neural framework that learns this modal decomposition end-to-end, supervised only by system-level observables and without supervising the modal parameters themselves. The architecture decomposes into a port-independent pole predictor and two port-dependent coupling predictors whose outputs are combined entry-wise, separating intrinsic from port-dependent features. This factorization yields a single trained model that generalizes to port counts unseen during training, dissolving the $\mathcal{O}(N^2)$ scaling barrier of direct regression. Despite no modal supervision, the freely-parameterized poles converge to physically meaningful eigenmodes, verified by cross-validation against the AAA rational approximation algorithm. We instantiate the framework in radio-frequency electromagnetic surrogate modeling. A model trained only on 2-port data accurately predicts $N$-port responses unseen during training.


NINJA: A Navigator–Inspector Joint Architecture for Context-Efficient Issue Localization

Wenkai Fang ⋅ Shunyu Liu ⋅ Wei Luo ⋅ Zihao Wu ⋅ Yang Zhou ⋅ Tongya Zheng ⋅ Wotao Yin ⋅ Jialie Shen ⋅ Leszek Rutkowski ⋅ Mingli Song ⋅ Dacheng Tao

Recent advances in agent-based methods have demonstrated strong promise for issue localization, a critical prerequisite for software issue resolution. However, most existing agent-based methods rely on a single agent with a growing context, where long-context accumulation compresses the effective reasoning space. Meanwhile, the flat exploration structure hinders the balance between file-level breadth and function-level depth. To address these limitations, we propose NINJA, a hierarchical Navigator-INspector Joint Architecture for context-efficient issue localization. NINJA decomposes repository exploration into global navigation and local inspection: the navigator maintains the global search state, performs file-level search, and dispatches selected entry files, while inspectors independently conduct function-level exploration around the assigned files in separate contexts. Through multi-round interactions, inspectors return suspicious locations as feedback, and the navigator updates the global state to decide whether to continue exploration or finalize localization. This hierarchical design balances file-level breadth with function-level depth, while independent inspector contexts prevent local exploration traces from accumulating in a single growing context. To further strengthen both global coordination and local exploration, we introduce a two-stage agentic fine-tuning strategy. Extensive experiments across multiple benchmarks and LLM backbones show that NINJA consistently outperforms competitive baselines. Notably, after fine-tuning, Qwen3-Coder-30B-A3B-Instruct surpasses the strong closed-source Claude-Haiku-4.5 model. Our code is available at https://anonymous.4open.science/r/NINJA-67FB/.


Normalizing Flows are Capable Trajectory Planners

Zhaohui Wang ⋅ Bo Xu ⋅ Yingzhi Tang ⋅ He Liu ⋅ Chen Bai

Generative models such as Diffusion Models and Flow Matching have improved multimodal trajectory planning in autonomous driving, but still suffer from inference latency and the lack of tractable probability densities, relying on heuristic trajectory selection. We propose BiDrive, a real-time end-to-end planning framework based on Normalizing Flows (NFs). By leveraging invertibility and exact likelihood estimation, BiDrive directly evaluates trajectory probabilities in a single forward pass, enabling principled maximum-likelihood-based decision making. To address the autoregressive bottleneck of expressive flows, we further develop a bidirectional distillation framework that compresses an autoregressive flow into a one-pass generator, achieving real-time inference (50 FPS) while preserving modeling capacity. The autoregressive model is used only during training as a probabilistic teacher. In addition, we introduce a latent-space optimization mechanism that incorporates differentiable safety and dynamic constraints for efficient test-time trajectory refinement. Experiments on NAVSIM closed-loop benchmarks demonstrate state-of-the-art performance in both safety and efficiency, highlighting the potential of Normalizing Flows for real-time trajectory planning. Code is available at: https://anonymous.4open.science/r/BiDrive.


NPUsper: Eliminating Redundant Computation for Real-Time Whisper on Mobile NPUs

Hojeong Lee ⋅ Si H Lee ⋅ Sungwon Woo ⋅ Chengpo Yan ⋅ Suman Banerjee ⋅ Seyeon Kim

We present NPUsper, a live transcription system that makes Whisper efficient on mobile NPUs by eliminating redundant computation. To avoid the heavy padding used by prior streaming systems, NPUsper detects hallucinated tokens online from temporal patterns in decoder cross-attention, allowing each inference round to process short audio inputs with minimal carryover. For efficient mobile-NPU execution, we propose controlled unrolling, which executes autoregressive decoding as K-step chunk graphs, removing unnecessary KV-cache computation and reducing graph-dispatch overhead. NPUsper achieves up to 4.84x lower per-word latency, up to 33.2x lower time-to-first-token (TTFT), and up to 11.36% lower average power consumption compared with baselines, while maintaining comparable transcription accuracy. The code is available at https://github.com/npusper/NPUsper.


nuReasoning: A Reasoning-Centric Dataset and Benchmark for Long-Tail Autonomous Driving

Zhiyu Huang ⋅ Johnson Liu ⋅ Rui Song ⋅ Zewei Zhou ⋅ Ruining Yang ⋅ Yun Zhang ⋅ Tianhui Cai ⋅ Hanyin Zhang ⋅ Mingxuan Gao ⋅ Valeria Xu ⋅ Jiali Chen ⋅ Yishan Shen ⋅ Yiluan Guo ⋅ Tony Qi ⋅ Jiaqi Ma

Reasoning is essential for autonomous driving (AD) in long-tail scenarios, where vehicles must apply commonsense knowledge, understand spatial relations, infer agent interactions, and make safe decisions. However, existing AD datasets and benchmarks mainly target perception, prediction, or planning, and provide limited supervision for reasoning over realistic long-tail driving scenes. We introduce nuReasoning, a large-scale real-world dataset and benchmark for reasoning-centric AD. Following the lineage of nuScenes and nuPlan, nuReasoning advances real-world AD datasets and benchmarks toward structured reasoning in long-tail driving scenarios. It contains 20K 20-second clips collected across multiple cities, with synchronized multi-camera images, LiDAR, HD maps, object annotations, and complementary forms of human-verified reasoning annotations: Spatial Reasoning, Decision Reasoning, and Counterfactual Reasoning. Unlike prior datasets that focus primarily on visual question answering, nuReasoning supports both reasoning evaluation and planning evaluation, enabling a direct study of how reasoning supervision affects driving performance. Experiments show that fine-tuning VLMs on nuReasoning substantially improves driving-specific question answering, while incorporating reasoning supervision into VLA training improves planning performance even when textual reasoning outputs are disabled at inference time. These results establish nuReasoning as a foundation for evaluating and improving robust, interpretable, reasoning-driven AD systems in realistic long-tail settings.

Retrieval benchmarks are increasingly saturating, but we argue that efficient search is far from a solved problem. We identify a class of queries we call oblique, which seek documents that instantiate a latent pattern, like finding all tweets that express an implicit stance, chat logs that demonstrate a particular failure mode, or transcripts that match an abstract scenario. We study three mechanisms through which obliqueness may arise and introduce OBLIQ-Bench, a suite of five oblique search problems over real long-tail corpora. OBLIQ-Bench exposes an overlooked asymmetry between retrieval and verification, where reasoning LLMs reliably recognize latent relevance whenever relevant documents are surfaced, but even sophisticated retrieval pipelines fail to surface most relevant documents in the first place. We hope that OBLIQ-Bench will drive research into retrieval architectures that efficiently capture latent patterns and implicit signals in large corpora.

Online change-point detection is a fundamental problem in sequential monitoring, where the goal is to detect distributional changes in a time series as quickly as possible while controlling false alarms. We propose a general framework for online change-point detection that leverages foundation probabilistic forecasters, such as Chronos-2, which can capture complex temporal structure including trend, seasonality, heteroskedasticity, and nonlinear dependence. Our approach transforms each incoming observation relative to its predictive quantile forecast into an approximately standard Gaussian monitoring statistic, which can then be used in classical sequential detection procedures such as CUSUM and MOSUM. To reduce adaptation of the forecaster to post-change observations, we introduce a buffer between the forecasting context window and the forecast target. Because this design can induce serial dependence in the monitoring statistics, we develop a parametric bootstrap procedure to calibrate critical thresholds and control false alarms in practice. Extensive simulations across data-generating processes involving non-normality, trend, seasonality, heteroskedasticity, and autoregressive dependence show that the proposed method achieves competitive detection performance relative to benchmark procedures, while requiring no explicit specification of a parametric pre-change model.

We study the problem of controlling instability in the Vlasov--Poisson plasma systems, a fundamental challenge for nuclear fusion control. Motivated by the partial observability of plasma states in practice, we investigate imitation learning algorithms that learn from a fully-observable expert controller, but operate under macroscopic measurements. We theoretically establish a separation between online and offline imitation learning --- even a small behavior cloning error leads to error compounding that exponentially amplifies instability over time. By way of contrast, the online imitation learning loss can polynomially control the stability of rollout trajectories. We then propose a DAgger-style algorithm that learns the stabilizing controller under partial observability. Simulation results on a 1D Vlasov--Poisson system demonstrate that our algorithm can effectively stabilize plasma instabilities over longer time horizons than behavior cloning.


Online Learning in Stabilized Linear Dynamical Games with Adversarial Disturbances

Anas Barakat ⋅ John Lazarsfeld ⋅ Georgios Piliouras ⋅ Antonios Varvitsiotis

We study online learning in linear dynamical games where multiple strategic agents act on a shared state evolving according to a linear dynamical system subject to adversarial disturbances. This setting lies beyond both single-agent nonstochastic online control and classical linear-quadratic games, which typically focus on quadratic objectives and noiseless or stochastic dynamics. Each agent seeks to minimize its own sequence of convex losses under full-state observability. Following the stabilizing-baseline paradigm in nonstochastic online control, we assume access to linear controllers that stabilize the noiseless system and focus on online adaptation by learning disturbance-action corrections. This extends adversarial online control to strategic multi-agent shared-state systems, where each agent's actions shape the state trajectory and hence the realized losses of all learners. Under state-only and aggregate-input feedback models, we analyze agents running online gradient descent with memory to update their own disturbance-action policies. We prove per-agent regret bounds that are sublinear and near-optimal in the time horizon, and show how their dependence on the number of agents varies with the feedback available to each learner. In the common-interest case, where all agents have identical cost functions, we show that the induced online problem forms a time-varying potential game and derive equilibrium-tracking guarantees. Together, these results provide a theoretical framework for adversarial online learning in stabilized linear dynamical games, connecting online control with learning in games.

Machine learning models on graph-structured data are increasingly deployed in dynamic environments, where data arrive continuously and models must be updated online, while specific data points and their influence are requested to be removed to meet privacy or regulatory requirements. Most existing work on machine unlearning focuses on offline or post-training settings and does not readily extend to online graph learning: updates with arriving data and deletion requests are interleaved, and removing nodes or edges can propagate changes through the graph structure. We study online learning with certified unlearning over graph-structured data, where a model is updated sequentially while accommodating deletion requests at any time. The subsequent outputs need to be indistinguishable from those of a model trained without the removed data. We propose OLUG, a principled framework that performs localized second-order retroactive correction during online graph learning using only recent batch information. We show that OLUG preserves no-regret online learning while supporting certified removal, incurring only an additive regret cost under infrequent deletions. Our analysis further reveals how graph propagation amplifies deletion sensitivity differently across node, edge, and feature removal scenarios. Experiments on multiple benchmark graph datasets demonstrate utility comparable to training from scratch, with substantially reduced computational overhead by avoiding retraining.


On the Convergence Analysis of Muon

Wei Shen ⋅ Ruichuan Huang ⋅ Minhui Huang ⋅ Cong Shen ⋅ Jiawei zhang

The majority of parameters in neural networks are naturally represented as matrices. However, most commonly used optimizers treat these matrix parameters as flattened vectors during optimization, potentially overlooking their inherent structural properties. Recently, an optimizer called Muon has been proposed, specifically designed to optimize matrix-structured parameters. Extensive empirical evidence shows that Muon can significantly outperform traditional optimizers when training neural networks. Nonetheless, the theoretical understanding of Muon’s convergence behavior and the reasons behind its superior performance remain limited. In this work, we present a comprehensive convergence rate analysis of Muon and its comparison with Gradient Descent (GD). We characterize the conditions under which Muon can outperform GD. Our theoretical results reveal that Muon can benefit from the low-rank structure of Hessian matrices, a phenomenon widely observed in practical neural network training. Our experimental results support and corroborate the theoretical findings.


On the Design Space of Discrete Diffusion Online Adaptation for Molecular Optimization

Trevor Chen ⋅ Ariel Dai ⋅ Jason Yang ⋅ Riccardo De Santi ⋅ Daniel Khalil ⋅ Wenda Chu ⋅ Nate Gruver ⋅ Pranav Murugan ⋅ Alexander Goldberg ⋅ Maruan Al-Shedivat ⋅ Yisong Yue

We study how to design online optimization loops for molecular optimization via adaptation of pre-trained generative models. At test time, we aim to utilize limited oracle feedback to steer generation toward top-molecules achieving high rewards. This creates a design problem coupled across several dimensions, including which candidates receive oracle evaluations, how observed rewards are utilized for model update, and how to overcome pre-trained model biases–to effectively explore over complex design spaces. Despite recent algorithmic progress on each individual component, it remains unclear how they interact in practice, within real-world online adaptation loops. For instance, top-$K$ optimization objectives may make adaptation overly greedy and thereby reduce exploration, while model-debiasing techniques may be unnecessary when exploration is already induced by exploratory acquisition functions, e.g., Thompson sampling. To address this, we conduct a controlled study on discrete-diffusion molecular optimization, finding that well-designed components remain beneficial when combined, indicating that they tackle complementary issues. Together, these components yield an online fine-tuning recipe that outperforms offline fine-tuning and search-augmented baselines across several small-molecule binding affinity and protein fitness optimization tasks, under equal oracle-call budgets and GPU-hour accounting.


On the Meta-Design of Allocation Problems

Unai Fischer Abaigar ⋅ Emily Aiken ⋅ Christoph Kern ⋅ Juan C Perdomo

There is an extensive literature that studies how to find optimal policies in resource allocation problems, taking the underlying design parameters that define the allocation, such as what data is collected, how many people can be served, and quality of service as fixed constraints. Yet, from a planner's perspective, these design parameters are themselves optimization variables that are just as important in determining overall welfare as selecting the optimal targeting rule for a given set of constraints. This realization motivates a rich set of meta-design questions exploring how planners should make principled decisions about investments in prediction, capacity constraints, and treatment quality, all of which lie upstream of classical policy optimization. Building on initial theoretical work in this space, our paper has three main contributions. First, we formally define the broad meta-design space of resource allocation problems. Second, we develop empirical tools that enable practitioners to reliably navigate it. Third, we demonstrate the framework in two real-world case studies on German employment services and targeted cash transfer programs in Ethiopia.


On Time, Within Budget: Constraint-Driven Online Resource Allocation for Agentic Workflows

Xinglin Wang ⋅ Zishen Liu ⋅ Shaoxiong Feng ⋅ Peiwen Yuan ⋅ Yiwei Li ⋅ Jiayi Shi ⋅ Yueqi Zhang ⋅ Chuyi Tan ⋅ Ji Zhang ⋅ Boyuan Pan ⋅ Yao Hu ⋅ Prof. Kan

Agentic systems increasingly solve complex user requests by executing orchestrated workflows, where subtasks are assigned to specialized models or tools and coordinated according to their dependencies. While recent work improves agent efficiency by optimizing the performance--cost--latency frontier, real deployments often impose concrete requirements: a workflow must be completed within a specified budget and before a specified deadline. This shifts the goal from average efficiency optimization to maximizing the probability that the entire workflow completes successfully under explicit budget and deadline constraints. We study constraint-driven online resource allocation for agentic workflows. Given a dependency-structured workflow and estimates of success rates and generation lengths for each subtask--model pair, the executor allocates models and parallel samples across simultaneously executable subtasks while managing the remaining budget and time. We formulate this setting as a finite-horizon stochastic online allocation problem and propose Monte Carlo Portfolio Planning (MCPP), a lightweight closed-loop planner that directly estimates constrained completion probability through simulated workflow executions and replans after observed outcomes. Experiments on CodeFlow and ProofFlow demonstrate that MCPP consistently improves constrained completion probability over strong baselines across a wide range of budget--deadline constraints.

Large language models have made synthetic data inexpensive, but they can still be biased. When real data are scarce, researchers must decide how much of the real sample to use for direct estimation and how much to spend calibrating the generator. We study this allocation problem and derive the marginal-value condition for the optimal calibration size: $(v_n + c x^{-2\beta})^2 = 2\beta a c x^{-(2\beta+1)}$, where $x$ is the calibration size, $v_n$ is synthetic sampling variance, $a$ is the real-estimator variance constant, and $c x^{-2\beta}$ is squared synthetic bias. This condition has no universal fixed-share solution and yields five allocation regimes. In the main regime, $m_n \asymp n$ and $\beta > 1/2$, the optimal calibration size grows only as $n^{2/(2\beta+1)}$, so the calibration share vanishes; rules that keep a positive fixed fraction of real observations in calibration over-calibrate asymptotically. We also propose an adaptive grid estimator with an oracle inequality and a safety check, $xB_n(x)

Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effectively ignoring embedding magnitudes. However, surprisingly, empirical studies reveal that despite this, these "discarded" norms seem to correlate with semantic properties such as concept specificity, token frequency, and human uncertainty. In this work, we provide a formal theoretical framework explaining this phenomenon. By analyzing the optimization dynamics, we derive an analytic formula demonstrating that embedding length naturally encodes this information as a byproduct of the training process. We also show how this gives rise to signals that can serve as "free" calibration tools in specific models and retrieval tasks, providing a grounded explanation for a previously heuristic observation.

Squared Wasserstein distance is a frequently used tool to compare probability distributions. This distance is typically computed between empirical measures of size $n$ from two underlying random samples. Unfortunately, even in lower dimensional Euclidean space problems $\left( d \in \{2,3\} \right)$, Wasserstein distance algorithms with approximate or exact precision guarantees scale poorly in the runtime as a function of $n$ and the desired precision. In response, we consider the computational-statistical runtime, where the goal to estimate from samples the Wasserstein distance up to the $\varepsilon$-additive error which is achievable under a sample; we allow $O(1)$ time per sample. Towards this, we develop a Sample-Sketch-Solve paradigm where we introduce a regular cartesian grid sketch of the samples. We show that (especially under $\alpha$-H\"older smooth distributions) this can compress the data without increasing asymptotic error, and also regularizes the structure which enables faster exact algorithms. Ultimately, we approximate $W_2^2(P,Q)$ within $\varepsilon$ error in $\varepsilon^{-\max(2,\frac{d+1+o(1)}{2})}$ time for Lipschitz $P,Q$ on $(0,1)^d$; nearly an optimal $\Theta(\varepsilon^{-2})$ for $d \leq 3$.

I show that ordinary least squares (OLS) predictions can be rewritten as a restricted attention module, akin to those at the heart of transformer architectures. This connection emerges by reframing OLS as a similarity-based prediction rule operating in a learned embedding space. From this perspective, least squares does not estimate coefficients per se, but instead selects an embedding that minimizes squared prediction error by matching training and test vectors through inner products — mapping naturally onto the query-key-value structure of attention. I then extend the framework to dimensionality reduction, nonlinearity, and connections to time series econometrics. Monte Carlo simulations and real-data experiments on UCI/OpenML benchmarks show that a direct implementation of nonlinear Attention Regression performs competitively against standard machine learning baselines. In the other direction, within a transformer architecture for tabular data, an attention block can be replaced by an explicit regression on polynomial-expanded features, matching or exceeding the standard transformer's predictive accuracy at a fraction of its parameter count.


PACE: Two-Timescale Self-Evolution for Small Language Model Agents

Chen Ling ⋅ Pei Chen ⋅ Xiangchen Guan ⋅ Jiaming Qu ⋅ Shayan Ali Akbar ⋅ Madhu Gopinathan ⋅ Erwin Cornejo

Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other components of the agent pipeline. Self-evolution offers a promising alternative, but most existing frameworks assume access to frontier models that can reliably diagnose failures, propose revisions, and judge their own updates. We study whether frozen small language models (SLMs) can serve as effective self-evolving agents under resource constraints. We propose PACE (Prompt And Control Logic Evolution), a two-timescale framework that coordinates low-risk prompt refinement with higher-risk control-logic updates. PACE evolves prompts under fixed control logic until prompt-level gains saturate, then considers constrained control-logic updates that are accepted through held-out validation. Across three frozen SLM backbones ranging from 4B to 14B parameters and four controlled benchmarks, PACE achieves the best performance on all 12 backbone--benchmark combinations, improving over vanilla SLM agents by up to $+9.2\%$ relative improvement and over the stronger single-mode evolution baseline by up to $+5.4\%$ relative improvement. A $\tau$-bench case study further shows that PACE improves multi-turn tool-use success over vanilla and prompt-only evolution. These results suggest that reliable SLM agent self-evolution is possible without updating model weights or relying on frontier-model teachers, and that the key benefit is not any single final solver pattern but autonomous, validated discovery of task-appropriate inference strategies.


PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents

Minghao Yan ⋅ Bo Peng ⋅ Benjamin Coleman ⋅ Ziqi Chen ⋅ Zhouhang Xie ⋅ Shuo Chen ⋅ Zhankui He ⋅ Noveen Sachdeva ⋅ Weili Wang ⋅ Ed Chi ⋅ Shivaram Venkataraman ⋅ Wang-Cheng Kang ⋅ Derek Cheng ⋅ Beidou Wang

Large language models have become drivers of evolutionary search, but most systems rely on a fixed, prompt-elicited policy to sample next candidates. This limits adaptation in practical engineering and research tasks, where evaluations are expensive, and progress depends on learning task-specific search dynamics. We introduce PACEvolve++, an advisor-model reinforcement learning framework for test-time policy adaptation in evolutionary search agents. PACEvolve++ decouples strategic search decisions from implementation: a trainable advisor generates, assesses, and selects hypotheses, while a stronger frontier model translates selected hypotheses into executable candidates. To train the advisor under non-stationary feedback, we propose a phase-adaptive approach that adapts its optimization strategy to different phases of the evolutionary process. Early in evolution, it uses group-relative feedback to learn broad search preferences; later, as reward gaps compress, it emphasizes best-of-$k$ frontier contribution to support stable refinement. Across expert-parallel load balancing, sequential recommendation, and protein fitness extrapolation, PACEvolve++ outperforms the state-of-the-art evolutionary search framework with frontier models, achieving faster convergence and stabilizing test-time training during evolutionary search.


Paint Anything: Toward Any-Color Controllable Image Generation and Editing

Ji Xie ⋅ Dewei Zhou ⋅ Xinyu Huang ⋅ Xun Wang

Any-color control---the ability to specify objects by arbitrary 24-bit hex values rather than coarse color words---is a practical yet challenging requirement for image generation and editing, where users often need exact object colors, brand colors, or carefully colored compositions. Prior work has explored color generation, color editing, and image colorization, but most methods rely on task-specific modules or training-free inference-time techniques. This fragmentation makes any-color control difficult to use as a unified text-prompt capability. In this work, we present Paint-Anything, a unified model for any-color controllable generation and editing that directly embeds 24-bit hex colors into text prompts and supports diverse color-control tasks within one framework. Our key insight is that modern text-to-image models already contain partial color understanding, but this capability remains unstable unless numeric color tokens are aligned with localized pixel colors. To this end, we construct Paint-500K, a 500K color-control dataset, and combine it with a high-noise gate for lightweight RGB alignment. We further introduce \textbf{PCBench}, a diagnostic benchmark suite for object-level hex color fidelity, with two parts: \textbf{PCBench-T2I} for generation and \textbf{PCBench-Edit} for editing. Experiments and ablations show that \ours{} substantially improves any-color generation and editing over the base model, suggesting that arbitrary hex-color control can be learned as a native prompt-following behavior of modern image models.


PBT-Bench: Benchmarking AI Agents on Property-Based Testing

Guohao Jing ⋅ Xinqi Wang ⋅ Liao Zhang ⋅ Simon Du

Existing code benchmarks measure whether an agent can produce any test that reproduces a known bug, or whether it can produce a patch that fixes a described issue. Neither isolates the distinct skill of property-based testing: deriving a semantic invariant from documentation, and then constructing an input-generation strategy precise enough to make a random search reveal the violation. We introduce PBT-Bench, a benchmark of 100 curated property-based testing problems across 40 real Python libraries. Each problem injects one or more semantic bugs (365 in total, mean 3.65 per problem) designed so that default-strategy random inputs almost never trigger them; the agent must read the library's documentation, identify the relevant invariant, and specify a Hypothesis @given strategy that concentrates mass in the trigger region. Bugs are stratified across three difficulty levels (L1–L3) spanning single-constraint boundary bugs to stateful, cross-function protocol violations. We evaluate eight contemporary LLMs under two prompting regimes (open-ended baseline vs. explicit Hypothesis scaffolding) for three independent runs per configuration. Bug recall under the PBT-guided prompt ranges from 42.1% to 83.4% across models; under the open-ended baseline, from 31.4% to 76.7%. Hypothesis scaffolding lifts mid-capability models by over 20 percentage points, but yields smaller gains for the strongest models, with two exceptions showing degradation, suggesting the structured prompt can interfere with certain model behaviours rather than complementing them. The hardest bugs prove model-specific: different architectures fail on different problems, leaving persistent gaps that no single model closes. We release the benchmark, harness, and full evaluation corpus to support downstream work on documentation-grounded semantic reasoning.

Accelerating stochastic gradient methods with classical momentum schemes, such as Polyak's heavy ball, has proven highly successful in training large-scale machine learning models, particularly when combined with the hardware acceleration of large mini-batch computations. Yet, the effect of classical momentum on stochastic mini-batch optimization has been poorly understood theoretically, with prior works requiring strong noise assumptions and extremely large mini-batches. In this work, we develop a general theory of stochastic momentum acceleration for optimizing over quadratics in the interpolation regime, a popular abstraction for studying deep learning dynamics which also includes classical methods such as randomized Kaczmarz and coordinate descent. Our framework encompasses both heavy ball and Nesterov-style momentum, allows for arbitrary mini-batch sizes, and makes minimal assumptions on the stochastic noise. In particular, we show that acceleration from classical momentum is directly proportional to the gradient mini-batch size (up to a natural saturation point), thereby enabling perfect parallelization of mini-batch computations. Our theory also provides a simple choice for the momentum parameter, which is shown to be effective empirically.


PhyMotion: Structured 3D Motion Reward for Physics-Grounded Human Video Generation

Yidong Huang ⋅ Zun Wang ⋅ Han Lin ⋅ Dong-Ki Kim ⋅ Shayegan Omidshafiei ⋅ Jaehong Yoon ⋅ Jaemin Cho ⋅ Yue Zhang ⋅ Mohit Bansal

Generating realistic human motion is a central yet unsolved challenge in video generation. Reinforcement learning (RL)-based post-training has emerged as a promising direction, yet its success critically depends on the quality of the reward signal. Existing video rewards primarily rely on 2D perceptual signals, without explicitly modeling the 3D body state, contact, and dynamics underlying articulated human motion, and often assign high scores to videos with floating bodies or physically implausible movements. To address this, we propose PhyMotion, a structured, fine-grained motion reward that grounds recovered 3D human trajectories in a physics simulator and evaluates motion quality along multiple dimensions of physical feasibility. Concretely, we recover SMPL body meshes from generated videos, retarget them onto a humanoid in the MuJoCo physics simulator, and evaluate the resulting motion along three axes: kinematic plausibility, contact and balance consistency, and dynamic feasibility. Each component provides a continuous and interpretable signal tied to a specific aspect of motion quality, allowing the reward to capture which aspects of motion are physically correct or violated. Experiments show that PhyMotion achieves stronger correlation with human judgments than existing reward formulations. When used for RL-based post-training, it consistently improves motion realism across both autoregressive and bidirectional video generators under both automatic metrics and blind human evaluation. Ablations further show that the three axes provide complementary supervision signals, while the reward preserves overall video generation quality with only modest training overhead.


PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments

Ruoqi Liu ⋅ Imran Mohiuddin ⋅ Austin J Schoeffler ⋅ Kavita Renduchintala ⋅ Ashwin Nayak ⋅ Prasantha L Vemu ⋅ Shivam Vedak ⋅ Kameron C Black ⋅ John L Havlik ⋅ Isaac Ogunmola ⋅ Stephen P. ⋅ Roopa Dhatt ⋅ Jonathan H Chen

We introduce PhysicianBench, a benchmark for evaluating LLM agents on physician tasks grounded in real clinical setting within electronic health record (EHR) environments. Existing medical agent benchmarks primarily focus on static knowledge recall, single-step atomic actions, or action intent without verifiable execution against the environment. As a result, they fail to capture the long-horizon, composite workflows that characterize real clinical systems. PhysicianBench comprises 100 long-horizon tasks adapted from real consultation cases between primary care physicians and specialists, with each task independently reviewed by a separate panel of physicians. Tasks are instantiated in an EHR environment with real patient records and accessed through the same standard APIs used by commercial EHR vendors. Tasks span 21 specialties (e.g., cardiology, endocrinology, oncology, psychiatry) and diverse workflow types (e.g., diagnosis interpretation, medication prescribing, treatment planning), requiring an average of 27 tool calls per task. Solving each task requires retrieving data across encounters, reasoning over heterogeneous clinical information, executing consequential clinical actions, and producing clinical documentation. Each task is decomposed into structured checkpoints (670 in total across the benchmark) capturing distinct stages of completion graded by task-specific scripts with execution-grounded verification. Across 12 proprietary and open-source LLM agents, the best-performing model achieves only 46% success rate (pass@1), while open-source models reach at most 19%, revealing a substantial gap between current agent capabilities and the demands of real-world clinical workflows. PhysicianBench provides a realistic and execution-grounded benchmark for measuring progress toward autonomous clinical agents.


PM-LoRA: Scalable Continual Learning via Progressive Merging of Low Rank Adapters

Xiaobing Yu ⋅ Peijie Qiu ⋅ Jin Yang ⋅ Xiaoqi Zhao ⋅ XUANZHAO DONG ⋅ Wenhui Zhu ⋅ Xiwen Chen ⋅ Weiwei Ma ⋅ Xiaofeng Liu ⋅ Jiangpeng He

In continual learning, most parameter-efficient fine-tuning methods rely on task-specific low-rank adapters to incrementally adapt pre-trained models to new tasks. However, they overlook two critical challenges: one is the accumulation of numerous adapters that leads to linear growth in parameters and memory, and the other is the lack of an explicit mechanism to preserve previously learned knowledge during continual adaptation. In this work, we propose a simple yet scalable framework, termed Progressive Merging of Low-Rank Adapters (PM-LoRA), to address both issues simultaneously. Specifically, PM-LoRA trains a small task-specific spectral orthogonal adapter for each new task and progressively integrates it into a single evolving LoRA adapter that encodes all prior knowledge. This merging process enables continuous knowledge integration while directly mitigating catastrophic forgetting at the parameter level. Benefiting from this design, PM-LoRA achieves true scalability by maintaining constant parameter overhead throughout continual training. Extensive experiments on four benchmark datasets demonstrate that PM-LoRA achieves state-of-the-art performance with superior memory and computation efficiency.


Polynomial-Time Robust Multiclass Linear Classification under Gaussian Marginals

Ilias Diakonikolas ⋅ Giannis Iakovidis ⋅ Mingchen Ma

We study the task of agnostic learning of multiclass linear classifiers under the Gaussian distribution. Given labeled examples $(x, y)$ from a distribution over $\mathbb{R}^d \times [k]$, with Gaussian $x$-marginal, the goal is to output a hypothesis whose error is comparable to that of the best $k$-class linear classifier. While the binary case $k=2$ has a well-developed algorithmic theory, much less is known for $k \ge 3$. Even for $k=3$, prior robust algorithms incur exponential dependence on the inverse of the desired accuracy in both complexity and representation size. In this work, we develop new structural results for multiclass linear classifiers and use them to design fully polynomial-time robust learners with dimension-independent error guarantees. Our first result shows that the standard multiclass perceptron algorithm requires super-polynomially many samples and updates, even with clean labels and Gaussian marginals, revealing a basic obstruction absent in the binary case. Our main positive result is a pairwise improper-learning framework which yields an efficient learner with error $\widetilde O(k^{3/2}\sqrt{\mathrm{opt}})+\epsilon$ for general $k$. Additionally, we develop a sharper localization-based framework which leads to error $O(\mathrm{opt})+\epsilon$ for $k=3$, and error $\mathrm{poly}(k)\mathrm{opt}+\epsilon$ for geometrically regular $k$-class linear classifiers.


PORT: Preference Optimization via Robust Token-Level Reweighting

Ding Zhu ⋅ Xiukun Wei ⋅ Tian Xie ⋅ Zhihui Zhu ⋅ Xueru Zhang ⋅ Mohammad Mahdi Khalili

Preference optimization has become a central approach for aligning large language models with human values, but its effectiveness depends heavily on the quality of preference annotations. In practice, preference data is often noisy due to annotation errors and ambiguity. Existing robust preference optimization methods primarily operate at the sequence level, implicitly treating entire responses as uniformly correct or incorrect. However, rejected responses may still contain informative reasoning steps, while preferred responses can include subtle errors. In this work, we propose a Preference Optimization via Robust Token-Level reweighting (PORT), a fine-grained framework for robust alignment under noisy preferences. PORT performs token-level reweighting using the empirical cumulative distribution function (CDF) of token logits, yielding an efficient forward-pass-only proxy to gradient-norm penalization that selectively suppresses corrupted tokens without additional backward-pass computation. We provide theoretical analysis showing that PORT reduces the gradient bias from noisy labels by minimizing an upper bound on the bias term. Extensive experiments across diverse noise settings demonstrate that PORT consistently improves robustness and outperforms existing sequence-level baselines.


Position: We Need Greater Transparency to Maintain Research Pipeline Reliability Despite GenAI

Hillmer Chona ⋅ Sourav Panda ⋅ Frank Ritter ⋅ Jonathan Dodge

This position paper argues that the research community must act to maintain research pipeline reliability despite increasing threats from GenAI. We start in the pre-GenAI era by highlighting some of the prominent critiques of the research pipeline, consisting of: (a) Authoring, (b) Reviewing, and (c) Evaluating researchers based on bibliometrics. Then, we discuss some of the effects GenAI has had on each research pipeline component based on the best available data, limited though it may be. Importantly, due to opaqueness of the submission pipeline, it is difficult to measure both the impact of GenAI and the efficacy of interventions, which leads us to call for greater transparency. Last, we conclude by offering three alternative views and ten paths forward that may mitigate some of the issues we highlight by keeping human researchers as the primary driver of science

As synthetic images become increasingly realistic, reliable synthetic image detection techniques are of pressing need to prevent their misuse. Despite satisfactory in-distribution performance, deep neural network-based synthetic image detectors (SIDs) lack reliability in deployment and often fail in the presence of common covariate shifts, resulting in poor detection accuracy. To avoid the risk caused by potential errors, we adopt a selective classification (SC) strategy by allowing SIDs to abstain from making low confidence predictions. For practicality, we focus on post-hoc methods which perform confidence estimation on a given SID without retraining. However, we show that conventional logit-based confidence score functions (CSFs) exhibit pathological behavior under covariate shifts, leading to SC performance close to or even worse than random guessing. To address this, we propose a simple yet effective SC framework for Reliable Synthetic Image Detection (ReSIDe). First, we generalize the notion of logits to an SID's intermediate layers from a centroid matching perspective, extending the use of logit-based CSFs to any layer of an SID. Then, we introduce a preference optimization algorithm that aggregates confidence scores extracted from different layers to a final confidence estimate by minimizing an upper bound of the area under the risk-coverage curve (AURC). Extensive experimental results show that ReSIDe significantly boosts the SC performance of various logit-based CSFs under common covariate shifts, achieving up to 69.55% AURC reduction.

Multimodal large language models (MLLMs) have achieved remarkable progress on vision–language tasks, yet their reasoning processes remain sometimes unreliable. We introduce PRISM-Bench, short for Puzzle Reasoning with In-Sequence Mistakes, a benchmark of puzzle-based visual challenges designed to evaluate not only whether models can solve problems, but how their reasoning unfolds. Unlike prior evaluations that measure only final-answer accuracy, PRISM-Bench introduces a diagnostic task: given a visual puzzle and a step-by-step chain-of-thought (CoT) containing exactly one error, models must identify the first incorrect step. This setting enables fine-grained assessment of logical consistency, error detection, and visual reasoning. The puzzles in PRISM-Bench require multi-step symbolic, geometric, and analogical reasoning, resisting shortcuts based on superficial pattern matching. Evaluations across state-of-the-art MLLMs reveal a persistent gap between fluent generation and faithful reasoning: models that produce plausible CoTs often fail to locate simple logical faults. By disentangling answer generation from reasoning verification, PRISM-Bench offers a sharper lens on multimodal reasoning competence and underscores the need for diagnostic evaluation protocols in the development of trustworthy MLLMs.

Procedural activities are fundamentally driven by object state transitions, yet existing instructional video benchmarks remain action-centric and cannot evaluate whether models reason about how objects evolve toward task completion. In this work, we introduce \textsc{ProcObject-10K}, the first benchmark that jointly evaluates object-centric reasoning and temporal evidence grounding in instructional videos, across both egocentric and exocentric views. It comprises 10,522 open-ended VideoQA pairs grounded in 1,799 video clips, spanning 137 tasks across 9 domains and five reasoning types covering preconditions, state evolution, counterfactuals, mistakes, and readiness. Benchmarking 13 leading MLLMs reveals a substantial answering-grounding gap: models produce plausible answers while failing to localize the supporting evidence (mIoU $<$ 45\%), exposing their reliance on linguistic priors rather than fine-grained object dynamics. As a step toward closing this gap, we further provide an object-centric supervised fine-tuning baseline with pseudo object-level supervision and spatial-temporal constraints. Models fine-tuned on \textsc{ProcObject-10K} not only improve on the benchmark itself, but also transfer effectively to other grounded VideoQA and embodied planning tasks. The dataset, annotations, and evaluation toolkit will be publicly released to support future research on object-centric procedural understanding.


Proper Agnostic Learning of Functions of Halfspaces

Sergei Tikhonov ⋅ Arsen Vasilyan

We study the problem of computationally efficient **proper agnostic learning** of high-dimensional concept classes under the Gaussian distribution. In this setting, given i.i.d.\ labeled samples from an unknown distribution over $\mathbb{R}^d \times \{\pm 1\}$ whose marginal on $\mathbb{R}^d$ is Gaussian, the goal is to output a hypothesis from a target class $\mathcal{F}$ whose 0-1 loss is within $\varepsilon$ of that of the best classifier in $\mathcal{F}$. We give the first efficient proper agnostic learning algorithm for arbitrary Boolean functions of $K$ halfspaces under Gaussian marginals. Our algorithm runs in time $d^{O(K^2 \log(1/\varepsilon)/\varepsilon^2)} + (K/\varepsilon)^{O(K^3/\varepsilon^{2.5})}$. Prior to our work, the only known algorithm for $K \geq 2$ was brute-force search, with runtime exponential in $d$. Moreover, the dependence of our runtime on the dimension $d$ matches that of the best known \emph{improper} learning algorithm, namely $d^{\widetilde{O}(K^2/\varepsilon^2)}$. For the special case of a single halfspace ($K=1$), the best previous runtime was $d^{O(1/\varepsilon^4)} + (1/\varepsilon)^{O(1/\varepsilon^6)}$.\ Our algorithm improves this to $d^{O(1/\varepsilon^2)} + (1/\varepsilon)^{O(1/\varepsilon^{2.5})}$. Once again, the dependence on $d$ matches that of the best known improper algorithm, namely $d^{O(1/\varepsilon^2)}$. Furthermore, the dependence of our runtime on the dimension $d$ is essentially optimal in the statistical query model.

We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one for each data point, such that the reweighted rewards approximates the expected return under the target policy. The weights are learned inductively in a top-down manner via a moment matching objective against a value-function discriminator class. Notably, and perhaps surprisingly, a data-dependent finite-sample guarantee for general function approximation can be established under only the realizability of $Q^\pi$, with a **dimension-free** bound---that is, the error does **not** depend on the statistical complexity of the function class. We also establish connections to several existing methods, such as linear FQE. Further theoretical analyses shed new light on the nature of coverage, a concept of fundamental importance to offline RL.

Gradient Boosted Decision Trees (GBDTs) remain a leading approach for tabular learning, offering strong predictive performance, computational efficiency, and robustness to heterogeneous feature types. Yet their standard second-order leaf updates suffer from a node-level statistical limitation: as trees grow deeper, sparse leaves provide increasingly unreliable gradient estimates, while global regularization cannot adapt to local uncertainty or residual structure. We propose RAGBoost, a retrieval-augmented and ancillary-guided boosting framework for robust tabular learning. RAGBoost augments standard GBDT optimization with a two-layer node-adaptive correction mechanism. First, it retrieves historically observed leaves with similar gradient-state configurations to form node-conditional gradient priors, stabilizing variance-dominated sparse-leaf updates through empirical shrinkage. Second, it uses ancillary residual correction to detect and mitigate bias-dominated within-leaf residual structure that remains unresolved by the main tree. These corrections modify only leaf values, preserving the original tree structure and inference pipeline. Empirically, RAGBoost achieves the best overall average rank across 40 heterogeneous tabular benchmarks in our evaluation, while remaining strongly competitive with recent deep tabular and tabular foundation model baselines. On TabReD, an industrial-scale temporal-shift benchmark, RAGBoost outperforms XGBoost and the base TabM variant overall, suggesting improved robustness under realistic distribution drift.

Despite advances in representation learning, high-dimensional classification remains challenging in low-sample-size regimes, where the dominant signal may vary across applications and labeled data are often limited. We propose a dissimilarity-profiling classification framework that represents each observation by its class-wise dissimilarity profile, transforming the original feature space into a low-dimensional representation that summarizes how the observation relates to each class. The key idea is to turn a consequence of the curse of dimensionality into signal: high-dimensional geometry can induce systematic within-class and between-class dissimilarity patterns under location, scale, or other distributional changes, and these patterns are captured by the class-wise profiles. Building on this representation, we introduce a rank-transformed algorithm that converts dissimilarities into class-wise rank profiles, yielding a compact representation for classification. The proposed method delivers competitive or improved performance relative to commonly used classifiers on two-class, multi-class, network, and real high-dimensional low-sample-size datasets. To provide insight into the mechanism underlying the method, we analyze a distance-based surrogate and show that the resulting profiles encode differences in first, second, and higher-order moments, while the rank transformation improves robustness to outliers. Together, these results show that rank-transformed dissimilarity profiles provide an adaptive representation for high-dimensional classification when the signal structure is unknown.


Reaching a Consensus in Predictive Loops

Jiduan Wu ⋅ Rediet Abebe ⋅ Celestine Mendler-Dünner

Predictions on digital platforms must adapt over time, as individuals continuously update their beliefs through social interactions. At the same time, changing predictions can in turn influence the content people are exposed to and hence the very beliefs they seek to predict, rather than merely serving as passive observations. These emerging dynamics make it challenging to understand the long term effects of predictive systems on society. In this work, we blend models from network science with concepts from performative prediction to initiate the study of opinion dynamics in predictive loops. In our model, opinions and predictions co-evolve: a platform's predictions influence individual opinions, which then evolve through peer interactions and form the training data for future platform model updates. We demonstrate that this co-evolution induces a novel equilibrium that qualitatively differs from standard network equilibria. In particular, we show how standard predictive objectives can drive networks toward consensus even under conditions where classical opinion-dynamics models lead to disagreement. This emerges because predictive systems dynamically adapt to changing opinions, and learning objectives create spillover effects among individuals beyond the topology of the network. We further analyze systematic deviations from standard prediction and demonstrate amplified effects of targeted platform interventions on equilibrium outcomes, compared to classical network intervention analyses. We complement our results with simulations on real social network data and parametric learning settings. Together, our results illustrate performativity as an important, yet so far neglected, qualifying factor in social networks.


Reading, Not Thinking: Bridging the Modality Gap When Text Becomes Pixels

Kaiser Sun ⋅ Xiaochuang Yuan ⋅ Hongjun Liu ⋅ Chen Zhao ⋅ Cheng Zhang ⋅ Mark Dredze ⋅ Fan Bai

Multimodal large language models (MLLMs) can process text presented as images, yet they often perform worse than when the same content is provided as textual tokens. We systematically diagnose this ``modality gap'' by evaluating seven MLLMs across seven benchmarks in five input modes, spanning both synthetically rendered text and realistic document images from arXiv PDFs to Wikipedia pages. We find that the gap is highly sensitive to rendering choices such as font and resolution, and that natural document images often match or exceed text-mode performance, suggesting the gap partly reflects evaluation artifacts rather than fundamental limitations. Through a grounded-theory error analysis of over 4,000 examples, we identify the primary cause: image input alone suppresses reasoning effort, with models producing 5--19$\times$ shorter outputs that skip step-by-step computation or reasoning. The reluctance to reason, not a failure of perception or knowledge retrieval, drives the performance gap, particularly on tasks requiring multi-step reasoning. We show that a simple lightweight on-policy self-distillation method by fine-tuning models on their own text-mode reasoning traces paired with image inputs closes this gap, raising image-mode accuracy to match or exceed text-mode performance with over 50\% improvement, and the gains transfer to unseen benchmarks without catastrophic forgetting. Overall, our results and analyses provide a systematic understanding of the modality gap and suggest a practical path toward improving visual text understanding in multimodal language models.


Reasoning Pathologies in Large Language Models: A Diagnostic Perspective

Rohan Surana ⋅ Junda Wu ⋅ Sheldon Yu ⋅ Gagan Mundada ⋅ Xunyi Jiang ⋅ Zihan Huang ⋅ Ryan Rossi ⋅ Julian McAuley ⋅ Tong Yu

Long chain-of-thought (CoT) reasoning in large language models often fails in recurring ways: models loop through redundant derivations, drift through uncertain intermediate states, commit prematurely to unsupported answers, or continue generating after a valid conclusion has been reached. These failures are typically treated as separate surface phenomena and addressed with symptom-specific heuristics, which (i) leave the underlying trajectory failure unchanged when they suppress the visible symptom, and (ii) provide no shared substrate for comparing fixes across models or pathologies. We instead study them as failures at the level of CoT reasoning latent state transitions. We introduce a diagnostic framework that abstracts chain-of-thought reasoning as a trajectory through discrete latent reasoning states, yielding a positional taxonomy of six reasoning pathologies, each paired with a transition-grounded detection predicate and a matched inference-time controller. We evaluate the framework through controlled-surgery experiments across multiple datasets and models, spanning arithmetic reasoning, multi-hop question answering, and knowledge-intensive reasoning. Across evaluations, we find that reasoning pathologies exhibit three main diagnostic patterns: \textbf{(1)} failures follow a phase-ordered structure, where termination failures are most reliably identified, early branching failures are detectable, and progression failures require finer graph-aware analyses; \textbf{(2)} counterfactual transition-versus-emission tests show that several detectors track latent trajectory changes rather than surface rewrites, while also exposing cases where paraphrases alter the reasoning path itself; and \textbf{(3)} detector-timed interventions are most effective when the detected pathology has a clear local onset, but do not uniformly dominate generic controls. Together, these findings position the framework as a reproducible diagnostic substrate for studying CoT failures, with explicit evidence boundaries and stronger closed-loop control.


Recursive Language Models

Alex Zhang ⋅ Tim Kraska ⋅ Omar Khattab

We study allowing large language models (LLMs) to process arbitrarily long prompts through the lens of inference-time scaling. We propose Recursive Language Models (RLMs), a general inference paradigm that treats long prompts as part of an external environment and allows the LLM to programmatically examine, decompose, and recursively call itself over snippets of the prompt. We find that RLMs can successfully process inputs more than an order of magnitude beyond model context window limits and, even for shorter prompts, dramatically outperform the quality of vanilla frontier LLMs and common long-context scaffolds (e.g., on GPT-5 by a median across the evaluated benchmarks of 26% against compaction, 130% against CodeAct with sub-calls, and 13% against Claude Code) across four diverse long-context tasks while having comparable cost. At a small scale, we post-train the first model around the RLM. Our model, RLM-Qwen3-8B, outperforms the underlying Qwen3-8B model by a median of 28% and even approaches the quality of vanilla GPT-5 on three long-context tasks.


Relational Feature Distillation for Lightweight 3D Point Cloud Segmentation

Mohammad Saeid ⋅ Amir Salarpour ⋅ Pedram MohajerAnsari ⋅ Mert D Pesé

We study compression of serialized 3D point transformers for efficient point cloud segmentation. Starting from LitePT-S, we derive \textbf{TrimPT}, a compact student that reduces channel width and stage-3 attention depth while preserving the full 1024-point attention window. TrimPT uses 5.84\,M parameters and 12.95\,GFLOPs, giving $2.18\times$ fewer parameters and $1.96\times$ fewer FLOPs than LitePT-S. To improve the compressed student, we introduce \textbf{Stage-wise Relational Feature Distillation (SRFD)}, a training-only objective that matches pairwise cosine-similarity matrices between teacher and student features at the compressed attention stages. This explicitly regularizes teacher--student affinity mismatch and adds no inference-time cost because the teacher and projection heads are discarded after training. On ScanNet semantic segmentation, the resulting \textbf{TopoPT} reaches 76.6\% mIoU with a 5.84\,M parameter inference footprint, improving over TrimPT without SRFD and matching the official LitePT-S result with substantially fewer inference-time parameters and FLOPs. TopoPT also obtains 63.9\% mAP$_{50}$ on ScanNet instance segmentation, 33.0\% mAP$_{50}$ on ScanNet200, and 81.4\% mIoU on nuScenes, suggesting that stage-wise relational distillation is a useful training-time regularizer for lightweight 3D segmentation backbones. Code and models are available at: \url{https://github.com/anon-push/TopoPT}

Recent advances in diffusion and flow matching models have highlighted a shift in the preferred prediction target---moving from noise ($\varepsilon$) and velocity ($v$) to direct data ($x$) prediction---particularly in high-dimensional settings. However, a formal explanation of why the optimal target depends on the specific properties of the data remains elusive. In this work, we provide a theoretical framework based on a generalized prediction formulation $u=kx-(1-k)n$, where $x$ represents clean data and $n$ denotes noise. This formulation accommodates arbitrary targets, of which $\varepsilon$-, $v$-, and $x$-prediction are special cases. We derive the analytical relationship between data's geometry and the optimal prediction target, offering a rigorous justification for why $x$-prediction becomes superior when the ambient dimension significantly exceeds the data's intrinsic dimension. Furthermore, while our theory identifies dimensionality as the governing factor for the optimal prediction target, the intrinsic dimension of manifold-bound data is typically intractable to estimate in practice. To bridge this gap, we propose $k$-Diff, a framework that learns the optimal parameter $k$ directly from data, bypassing the need for explicit dimension estimation. Extensive experiments in both latent-space and pixel-space image generation demonstrate that $k$-Diff consistently outperforms fixed-target baselines across varying architectures and data scales, providing a principled and automated approach to enhancing generative performance.


Riemannian Bilevel Optimization under the Polyak–Łojasiewicz Condition

Zhixuan Li ⋅ Yuyang Zhang ⋅ Luke Jones ⋅ Muztaba Syed ⋅ William Chang ⋅ Andi Han

This paper studies bilevel optimization on Riemannian manifolds where the upper-level objective is nonconvex and the lower-level problem satisfies a Riemannian Polyak - Łojasiewicz condition rather than geodesic strong convexity. The classical hypergradient formula then breaks down, since the lower-level Hessian may be singular or indefinite away from the minimizer. We address this with an intrinsic regularized tangent-space formulation based on spectral clipping, and develop a Riemannian bilevel algorithm that avoids Hessian-inverse solves and inner-loop differentiation. We establish an $O(1/T)$ rate on the averaged squared Riemannian gradient mapping and $O(\varepsilon^{-1})$ iteration complexity, and extend the method to the stochastic finite-sum setting. Experiments on Stiefel, Grassmannian, and Poincar\'e-ball problems with rank-deficient Hessians (BCI~IV-2a, UCI Superconductivity) and Stiefel-constrained meta-learning on MiniImageNet show that the proposed method is the only one whose linear-system error stays at machine precision and whose Riemannian gradient norm decreases monotonically, while Hessian-inversion and Neumann-series baselines fail to descend; the stochastic complexity exhibits the predicted $\Theta(1/B)$ decay.


Riemannian Ordinary Least Squares

Xiaoyu Chen ⋅ Yujing Huang ⋅ Yingyan Zeng

Many scientific responses take values in Riemannian manifolds, with predictors that are scalar- and/or manifold-valued. A regression analysis in such settings must support not only point prediction but also formal hypothesis testing and effect-size estimation for individual coefficients. Yet no current manifold-regression framework simultaneously accommodates both predictor types, and returns per-coefficient hypothesis tests. We propose Riemannian Ordinary Least Squares (ROLS), a tangent-space conditional-mean model that serves as the OLS analogue for this setting. ROLS is well defined under an explicit injectivity condition that keeps the logarithm maps single-valued, with a cut-locus diagnostic for branch-sensitive cases outside it. Within the local geometric setting it admits a closed-form estimator, an $O(n^{-1})$ excess prediction-error bound with explicit curvature bias, and asymptotic $t$- and Wald tests for individual coefficients. Across simulations with synthetic data and three case studies with real data, ROLS is competitive with the strongest prediction baselines and identifies scientifically meaningful predictors, providing for manifold-valued regression the coefficient-level inference that OLS provides in the Euclidean case.


RnR: a meta-solver for causal discovery in undersampled time series data

Mohammadsajad Abavisani ⋅ Kseniya Solovyeva ⋅ David Danks ⋅ Vince D. Calhoun ⋅ Sergey Plis

Learning directed causal graphs from time-series data poses significant challenges, especially in fMRI where slow sampling rate obscures fast neural interactions. This temporal mismatch leads to undersampling, which can make multiple graphs equally plausible. We address this problem by explicitly modeling undersampling effects when recovering causal graphs. Our approach employs answer set programming (ASP) to enforce domain-specific constraints and optimize soft observational constraints, thereby identifying a Markov equivalence class for the resulting graph solutions. By customizing an ASP solver to collect multiple near-optimal solutions, we obtain not only the single best-fitting graph but an equivalence class of high-scoring graphs for expert consideration. This method, called Real-world noisy RASL (RnR), can also act as a meta-solver: it refines the output of other causal discovery algorithms by accounting for undersampling biases. In synthetic data and empirical brain network data, RnR produces more accurate causal graphs than state-of-the-art methods. When applied as a meta-solver (refining outputs of existing algorithms), it improves F1 scores by an average of 46\% over baseline methods; on standalone synthetic data benchmarks, it achieves a 64\% improvement. We demonstrate that RnR is robust to varying undersampling rates, maintaining high precision and recall even as sampling becomes more sparse, whereas baseline methods degrade significantly. Finally, we test RnR on real-world settings where ground truth connectivity is unknown such as human brain fMRI data, showing that incorporating undersampling-aware constraints via ASP yields more reliable and interpretable brain connectivity estimates from fMRI time series, bridging the gap between neural dynamics and observational data.


Robust Inference-Time Steering of Protein Diffusion Models via Embedding Optimization

Minhuan Li ⋅ Jiequn Han ⋅ Pilar Cossio ⋅ Luhuan Wu

A core challenge in structural biophysics is generating biomolecular conformations that are both physically plausible and consistent with experimental measurements. While sequence-to-structure diffusion models provide powerful priors, posterior sampling methods steer generation by perturbing atomic coordinates with gradients from experimental likelihoods. However, when the target lies in a low-density region of the prior, these methods require aggressive upweighting of the likelihood that can destabilize sampling and be sensitive to hyperparameters. We propose EmbedOpt, an inference-time steering framework that introduces an orthogonal optimization axis: rather than performing posterior sampling under a fixed prior, EmbedOpt directly optimizes the prior by updating the model's conditional embedding. This embedding space encodes rich coevolutionary signals, so optimizing it shifts the structural prior to align with experimental constraints. Empirically, EmbedOpt matches coordinate-based posterior sampling baselines on sparse distance constraints and outperforms them on cryo-electron microscopy map fitting, including real, noisy experimental ones. Furthermore, EmbedOpt's smooth optimization behavior yields robustness to hyperparameters spanning two orders of magnitude and enables comparable performance with fewer diffusion steps.


Robust Instruction Compliance in Cooperative Multi-Agent Reinforcement Learning

Wo Wei Lin ⋅ Ethan Rathbun ⋅ Enrico Marchesini ⋅ Xiang Zhi Tan

Real-world multi-agent reinforcement learning (MARL) systems must operate autonomously while adapting to natural language instructions provided by humans. These instructions arrive unpredictably, interrupt ongoing actions, and may conflict with long-horizon objectives under partial observability. Recent vision-language models enable instruction following, but are computationally expensive and limited to single-agent settings. Conversely, MARL can handle long-horizon coordination with macro-actions, but assumes uninterrupted execution. This work introduces macro-action value cancellation for instruction compliance (MAVIC) to enable interruption-aware MARL by bootstrapping the value of ongoing macro-actions upon receiving instructions. Agents can thus decouple task objective optimization from instruction-driven overrides within a unified policy and re-plan without destabilizing the underlying learning process. MAVIC integrates with standard actor-critic methods and achieves high instruction compliance while preserving base task performance across increasingly complex macro-action benchmarks.


Robust Policy Optimization to Prevent Catastrophic Forgetting

Mahdi Sabbaghi ⋅ George J. Pappas ⋅ Adel Javanmard ⋅ Hamed Hassani

LLMs are commonly trained through multi-stage post-training: first via RLHF, then fine-tuned for other downstream objectives. Yet even small downstream updates can compromise earlier learned behaviors (e.g., safety), exposing a brittleness known as catastrophic forgetting. This suggests standard RLHF objectives do not guarantee robustness to future adaptation. To address it, most prior work designs downstream-time methods to preserve previously learned behaviors. We argue that preventing this requires pre-finetuning robustness: the base policy should avoid brittle high-reward solutions whose reward drops sharply under fine-tuning. We propose Fine-tuning Robust Policy Optimization (FRPO), a robust RLHF framework that optimizes reward not only at the current policy, but across a KL-bounded neighborhood of policies reachable by downstream adaptation. The key idea is to ensure reward stability under policy shifts via a max-min formulation. By modifying GRPO, we develop an algorithm with no extra computation, and empirically show it substantially reduces safety degradation across multiple base models and downstream fine-tuning regimes (SFT and RL) while preserving downstream task performance. We further study a math-focused RL setting, demonstrating that FRPO preserves accuracy under subsequent fine-tuning.


SALART-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images

Xiaoxiao Sun ⋅ Ruotian Zhang ⋅ Junzhe Huang ⋅ James Burgess ⋅ Serena Yeung-Levy

Vision-language models (VLMs) are increasingly used to detect whether AI-generated images contain visible artifacts, yet their ability to analyze such artifacts remains poorly understood. A correct image-level decision can still hide important failures: a model may correctly flag an artifact while relying on the wrong visual cue, selecting the wrong region, or describing a defect that the image does not support. To evaluate these behaviors directly, we introduce SalArt-VQA, a diagnostic benchmark for fine-grained salient artifact understanding in AI-generated images. SalArt-VQA contains 950 images and 3,681 human-authored multiple-choice questions spanning artifact images, matched real reference images, and paired generated reference images. Four aligned question types evaluate presence detection, semantic localization, spatial grounding, and evidence-grounded defect identification, while the reference splits test calibration and abstention when the annotated defect is absent. Across 20 VLMs, SalArt-VQA reveals failures that image-level detection accuracy hides: the strongest model reaches 99.37% detection recall on artifact images but answers all four artifact-side questions correctly on only 53.26% of images. Comparing artifact images with defect-absent references reveals a sensitivity–calibration tradeoff: sensitive models often make unsupported artifact claims, while conservative models avoid false alarms largely by missing real artifacts. These results show that high artifact detection accuracy alone does not imply grounded artifact understanding. SalArt-VQA exposes these hidden failure modes and provides a fine-grained evaluation of whether VLM artifact claims are supported by local visual evidence.


Scaling Multi-Teacher Distillation for Digital Pathology

Sofiène Boutaj ⋅ Pierre Marza ⋅ Varun Belagali ⋅ Dimitris Samaras ⋅ Maria Vakalopoulou ⋅ Stergios Christodoulidis

Multi-teacher distillation allows transferring knowledge from multiple “teacher” networks to a single “student” network. This method is promising in fields such as digital pathology, where many powerful foundation models were proposed recently. However, as we show in this paper, the standard multi-teacher distillation approach reacts poorly to an increase in the number of teachers used to train a student encoder. We demonstrate that learning teacher-specific representations is a key to scaling in MTD. Importantly, different design choices, such as learnable teacher tokens, a tailored attention scheme, additional mixture-of-experts layers and a contrastive loss, are proposed to better learn such teacher-specific representations. We show that our method better scales with respect to the number of considered teachers, allowing us to train compact student encoders, up to 10X more efficient than larger teacher foundation models, while matching their performance and even outperforming them on a large set of tile-level and slide-level tasks. We conduct a thorough empirical validation, evaluating more than 10 foundation models on 39 tasks spanning tile and slide levels. As a result, we release a new collection of strong and efficient foundation models, named OMNI, trained from the knowledge of 10 state-of-the-art foundation models. We finally conduct an analysis of the learned teacher-specific representations, highlighting their complementarity and explaining why they can be easily aggregated at downstream time.


Self-Tuning Graph Filters via State-Dependent Operator Composition

Amir Ghazizadeh ⋅ Mahyar Alinejad ⋅ George Atia ⋅ Rickard Ewetz ⋅ Hao Zheng

Most graph neural networks propagate information through a fixed-coefficient polynomial filter applied uniformly across every node and every layer. However, the appropriate filter varies along two distinct axes that such a parameterization cannot accommodate. Within a single graph, different nodes benefit from different mixtures of low-pass smoothing, high-pass contrast, and pass-through behavior. Across graphs, signal propagation dynamics themselves differ, with some graphs benefiting from long-range propagation across many hops while others require damping to prevent over-smoothing. The appropriate filter is therefore not a single fixed object at all, but a state-dependent composition whose behavior changes according to a node's current representation and its propagation history. We propose Castor, a self-tuning graph filter that realizes this composition through state-dependent decisions made at every step of propagation. A router reads each node's current state together with its initial state, and emits per-node mixing weights over three primitives, namely low-pass smoothing, high-pass contrast, and identity. A learnable per-hop coefficient then determines how much of the most recent change in the propagation is carried forward, accelerating the iteration when positive and damping it when negative. Setting the coefficient to zero recovers standard first-order propagation. Across fourteen standard node-classification benchmarks spanning the full homophily spectrum, a single Castor instance ranks top-three on every dataset and achieves an average rank of $2.07$, more than $3.5$ ranks ahead of the next-best of fourteen baselines.


Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems

Tunyu Zhang ⋅ Zihao Zhao ⋅ Yusong Zhao ⋅ Haizhou Shi ⋅ Zhuohang Li ⋅ Haoxian Chen ⋅ Hao Wang ⋅ Dimitris Metaxas

LLM-based multi-agent systems (MAS) have attracted growing attention for improving reasoning through interaction among multiple agents. In this work, we focus on parallel multi-agent reasoning systems, where several agents solve the same problem over multiple rounds and aggregate their outputs into a final answer. Despite their strong reasoning performance, uncertainty estimation for such systems remains underexplored: the reliability of a MAS depends not only on individual generations, but also on how agents interact and evolve across rounds. We propose SPI (Sequential Probabilistic Inference), a lightweight, training-free uncertainty estimator that formulates MAS uncertainty as sequential inference over a latent system-level belief. SPI aggregates round-level agreement and generation-uncertainty signals through a filtering-style update. Across five backbones, six benchmarks, and two MAS protocols, SPI improves misclassification detection, selective prediction, and calibration over a broad set of uncertainty estimation baselines, including standard log-likelihood-based methods and MAS-specific estimators.


Severity-Controlled Prediction Sets for Medication Recommendation

Yu Gu ⋅ Zijun Yu ⋅ Chi-Kuang Yeh ⋅ Xinyu Wang ⋅ Ziyang Song

Medication recommendation is a multi-label prediction task, yet existing prediction-set evaluation and calibration methods rely mainly on exact-overlap metrics and therefore treat all missed medications as equally severe. However, missing a medication but predicting a therapeutically close drug is less severe than matching it only with distant predictions. To address this, we introduce Conformal Severity Control (CSC), a post-hoc framework for calibrating medication prediction sets under hierarchy-aware mismatch severity. CSC uses the Anatomical Therapeutic Chemical (ATC) hierarchy to quantify mismatch severity and controls patient-level severity without retraining the predictor, supporting both average severity and upper-quantile severity for the highest-severity mismatches. Experiments on MIMIC-IV show that CSC tracks user-specified severity tolerances and shifts prediction sets away from distant mismatches toward exact or close matches.

Policy mirror descent (PMD) has a mature finite-time theory in discounted Markov decision processes (MDPs), but less is known in the average-reward setting, a more natural objective for many control applications. We give a finite-time, finite-sample analysis of PMD in ergodic average-reward MDPs built around a single master recursion that governs convergence under any gradient proxy, without external regularization. Its specializations yield linear rates (with a superlinear regime in the well-conditioned case) for exact, inexact-tabular, and linear function approximation (LFA) updates. We complement these convergence results with end-to-end sample complexities of order $t_{\mathrm{mix}}^3/\varepsilon^2$ in both tabular ($|\mathcal S||\mathcal A|$-dependent) and LFA ($d$-dependent) settings. The LFA rate sharpens the prior best $t_{\mathrm{mix}}^5$ mixing dependence to $t_{\mathrm{mix}}^3$, and matching information-theoretic lower bounds establish that the $t_{\mathrm{mix}}^3/\varepsilon^2$ critic core is unimprovable in both settings.


Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic

Xingyu Zhao ⋅ Darsh Sharma ⋅ Rheeya Uppaal ⋅ Yiqiao Zhong

Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mechanistic studies reveal the discrepancy between LLMs and humans in skill compositions, the learning dynamics of skill acquisition and the role of data distributions remain elusive. In this study, we train transformers on synthetic arithmetic tasks with black-box model-agnostic metrics for analyzing non-human skill compositions. We discover that transformers often acquire skills for arithmetic in reverse order or in parallel instead of human-like sequential rules—a phenomenon we refer to as shattered compositionality. To explain these behaviors, we provide evidence that correlational matching to the training data, rather than causal or procedural composition, shapes learning dynamics. As a consequence, this non-human acquisition creates competition between partially learned skills, producing characteristic mixing errors and weaker robustness under controlled distribution shifts. We further show that the same qualitative behavior persists in modern LLMs and is not mitigated by pure model scaling or scratchpad supervision. Our results highlight a mismatch between training-time skill acquisition and the human-like hierarchical compositions, with implications for reasoning reliability and out-of-distribution robustness. An anonymized code repository is provided in the supplementary material.


SimSD: Simple Speculative Decoding in Diffusion Language Models

Junxia Cui ⋅ Haotian Ye ⋅ Runchu Tian ⋅ Hongcan Guo ⋅ Jinya Jiang ⋅ Haoru Li ⋅ Chaojie Ren ⋅ Yiming Huang ⋅ Kaijie Zhu ⋅ Zhongkai Yu ⋅ Kun Zhou ⋅ Jingbo Shang

Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster inference through parallel or blockwise decoding. However, their masked language modeling formulation remains incompatible with standard token-level speculative decoding, one of the most effective acceleration techniques for AR models. In AR decoding, the causal mask preserves temporally valid token-level contexts, enabling a target model to verify multiple drafted tokens in a single forward pass. In contrast, dLLMs rely on mask tokens and bidirectional attention, causing the effective context to change across denoising steps and preventing direct token-level speculative verification. To bridge this gap, we propose a simple but effective speculative decoding algorithm for diffusion language models, named SimSD, which mainly adopts a plug-and-play masking strategy that equips dLLMs with temporally valid token-level contexts for speculative decoding. Our method explicitly introduces reference tokens from draft-model predictions and designs an attention mask that regulates their interaction with current-step tokens, allowing dLLMs to compute valid logits for drafted tokens in a single forward pass. This restores the key verification ability provided by causal masking in AR models while preserving the parallel decoding advantages of dLLMs. The proposed method is training-free and can be flexibly integrated with other acceleration techniques such as KV cache and blockwise decoding. Experiments on SDAR-family dLLMs across four benchmarks show that our method achieves up to (7.46\times) higher decoding throughput while maintaining and even improving average generation quality. Our code will be publicly released.

We propose a dense associative memory for empirical measures (weighted point clouds). Stored patterns and queries are finitely supported probability measures, and retrieval is defined by minimizing a Hopfield-style log-sum-exp energy built from the debiased Sinkhorn divergence. We derive retrieval dynamics as a spherical Hellinger Kantorovich (SHK) gradient flow, which updates both support locations and weights. Discretizing the flow yields a deterministic algorithm that uses Sinkhorn potentials to compute barycentric transport steps and a multiplicative simplex reweighting. Under local separation and PL-type conditions we prove basin invariance, geometric convergence to a local minimizer, and a bound showing the minimizer remains close to the corresponding stored pattern. Under a random pattern model, we further show that these Sinkhorn basins are disjoint with high probability, implying exponential capacity in the ambient dimension. Experiments on synthetic Gaussian point-cloud memories demonstrate robust recovery from perturbed queries versus a Euclidean Hopfield-type baseline.


SMASH: Probing Speech Recognition Robustness via Semantically Targeted Bit Flips

Zafaryab Haider ⋅ Md Hafizur Rahman ⋅ Aysegul Bumin ⋅ Prabuddha Chakraborty

Automatic speech recognition (ASR) systems are usually judged by aggregate transcript accuracy, but real failures often hinge on meaning: for example, a wrong dose, address, account number, or date can matter more than many harmless word errors. We introduce SMASH, a framework for finding targeted semantic failures caused by sparse INT8 weight-bit perturbations, a dominant model data type for ASR edge devices, in quantized sequence-to-sequence ASR models. Given an audio input and a meaning-critical span,SMASH searches for small weight changes that replace the intended meaning while keeping the transcript readable, plausible, and close to the original output. Each accepted fault is a certificate that a specific semantic substitution is reachable under a bounded perturbation budget. Empirical study focuses on numeric and scalar substitutions across three ASR backbones: Whisper-small.en (WS), Whisper-large-v3 (WLv3), and SeamlessM4T-v2-large (S-M4T); and three corpora: an LLM-assisted controlled corpus (LACC), MultiMed, and LibriSpeech. On LACC at a 20-bit budget ($B$), SMASH-hybrid accepts numeric substitutions on 24/45 WS targets, 19/44 WLv3 targets, and 5/34 S-M4T targets; reachability is lower on real-speech datasets. In matched LACC budget sweeps tested up to $B=40$, WS reaches its ceiling already at $B=20$, while WLv3 and S-M4T continue to gain accepted targeted substitutions through $B=40$. Human validation supports the gate-based labels: on a 105-item shared validation set, independent human-majority labels and a four-judge large- language-model ensemble agree on the composite clean_success label, with $\kappa=1.00$. The faults are sparse: the pooled median is three INT8 bit flips, with medians of two flips on WS, eleven on WLv3, and twelve on S-M4T. Entity and negation pivots are evaluated as secondary evidence in the appendix. SMASH shifts ASR robustness evaluation from "how much did the transcript change?'' to "which meanings can be changed, how plausibly, and with how few bit flips?''


SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery

Mansooreh Montazerin ⋅ Antonio Ortega ⋅ Ajitesh Srivastava

Symbolic regression (SR) discovers closed-form mathematical expressions from data, offering interpretability beyond black-box models. Existing methods suffer from slow convergence in combinatorial search spaces and lack mechanisms to exploit compositional structure in the data. We introduce SMILE (Sine, Multiplication, Identity, Logarithm, Exponential), a hybrid framework that unifies continuous gradient-based optimization with discrete symbolic recovery through three stages: structural analysis of the data to identify the compositional hierarchy of the target expression, continuous optimization to learn parameters of a network that encodes the target expression using interpretable activations, and symbolic recovery through structured pruning, coefficient optimization, and rounding. This final stage distills the learned network into a compact expression with exact symbolic constants. We evaluate SMILE on SRBench across ground-truth and black-box datasets, with ablation studies validating each component. SMILE achieves the highest symbolic solution rate at the largest noise levels, demonstrating strong robustness where competing methods degrade substantially. It consistently lies on the Pareto front of accuracy versus complexity, recovering significantly simpler expressions in a fraction of the time required by the competing methods.


Solving Max-Cut to Global Optimality via Feasibility-Preserving Graph Neural Networks

Hao Chen ⋅ Chendi Qian ⋅ Christopher Morris ⋅ Andrea Lodi ⋅ Can Li

Exact solution of hard combinatorial optimization problems often relies on strong convex relaxations, but solving these relaxations repeatedly inside a branch-and-bound algorithm can be prohibitively expensive. Hence, we consider this challenge for \new{Max-Cut problem} (Max-Cut), where branch and bound commonly uses semidefinite programming (SDP) relaxations to bound subproblems. We propose a Max-Cut-specific \new{graph neural network} that serves as a principled, lightweight neural proxy for these SDP solvers and can be plugged directly into an exact branch-and-bound framework. The proposed architecture has update steps of complexity $\mathcal{O}(n^2 + ne)$, and predicts both primal- and dual-feasible SDP solutions. The primal SDP solutions yield feasible Max-Cut solutions via the Goemans--Williamson algorithm. In addition, it is trained in a self-supervised fashion without requiring solved SDP relaxations as labels. Empirically, we show that our architecture can substantially reduce the cost of bounding in exact Max-Cut solving by up to $10.6 \times$ compared with using the state-of-the-art SDP solver Mosek. Our work highlights the potential of learned, validity-preserving surrogates for accelerating exact optimization over structured convex relaxations.


SourceBench: Can AI Answers Reference Quality Web Sources?

Hexi Jin ⋅ Xurui Liu ⋅ Yuheng Li ⋅ Simran Malik ⋅ Yiying Zhang

Large language models (LLMs) increasingly answer queries by citing web sources, but existing evaluations emphasize the final answer rather than evidence quality. We introduce SourceBench, a benchmark for measuring the quality of cited web sources across 100 real-world queries spanning informational, factual, argumentative, social, and shopping intents. SourceBench uses an eight-metric framework covering content quality (content relevance, factual accuracy, objectivity) and page-level signals (e.g., freshness, authority/accountability, clarity), and includes a human-labeled dataset with a calibrated LLM-based evaluator that matches expert judgments closely. We evaluate eight LLMs, Google Search, and three AI search tools over 3996 cited sources using SourceBench and conduct further experiments to understand the evaluation results. Overall, our work reveals four key new insights that can guide future research in the direction of GenAI and web search.

Multimodal electrophysiological recordings provide noisy, partially observed measurements of latent neural dynamics critical for tasks like brain state estimation and neurological diagnosis. In joint EEG-intracranial EEG (iEEG) recordings, scalp EEG offers broad but spatially mixed coverage, whereas iEEG provides high-fidelity but spatially sparse measurements. This creates a partial-observation problem: directed source-level interactions must be inferred from heterogeneous sensors with incomplete coverage. A second challenge is nonstationarity: latent neural dynamics and directed interactions among brain regions can change abruptly across distinct regimes (e.g., pre-seizure, ictal, and post-seizure states). The third challenge is sparsity of significant source-level interactions compared to the high-dimensionality of the brain networks. Standard models that assume single-mode, fully observed, or stationary may therefore fail to capture regime-specific activity or spurious couplings. We propose the Sparse Multimodal Switching State-Space Model (S$^4$M$^3$), which fuses EEG and iEEG as complementary observations of a shared source-space latent process. Dynamics switch between regimes governed by separate sparse directed transition matrices, jointly addressing multimodal fusion, nonstationarity, and sparse network estimation. The estimation of the model uses the expectation-maximization (EM) algorithm, where the E-step applies the Kim filter-RTS smoother to estimate latent states and regime probabilities, and the M-step imposes elastic-net penalization on off-diagonal elements, promoting sparse inter-regional couplings while preserving self-dynamics. The learned transition matrices facilitates the assessment of regime-dependent edge changes. To address the estimation bias of elastic-net sparsity penalty, we separate edge selection from edge assessment: selected edges are refit by regime-weighted least squares, followed by Wald tests with Benjamini-Hochberg correction. On synthetic benchmarks, S$^4$M$^3$ improves Edge F1 by approximately 25\% over the strongest stationary baseline and 64\% over two-stage source-connectivity pipelines. On a real EEG-iEEG epilepsy case study, S$^4$M$^3$ identifies sparse hippocampal and parahippocampal outgoing pathways concordant with the clinician-confirmed seizure-onset zone, without using clinical labels.

The standard metric LP for correlation clustering has $\Theta(n^2)$ pairwise marginals and $\Theta(n^3)$ triangle inequalities, while dense pairwise supervision is often the primary bottleneck in large relational pipelines. We separate three sparsification questions that are often treated together: preserving the objective of every integral clustering, rounding from a sparse set of LP marginals, and clustering when the signed weighted input itself is only sparsely observed. First, we prove that the VC dimension of the signed-edge disagreement class induced by all clusterings of $n$ vertices is exactly $n-1$, so weighted edge sampling yields additive $\varepsilon$-coresets of size $\tilde O(n/\varepsilon^2)$ with optimal $n$-dependence. Second, we introduce *Sparse-LP-Pivot*, which imputes missing LP marginals from triangle witnesses, and analyze it at two levels: a universal but coarse perturbation bound for every Lipschitz LP-PIVOT rule, and a sharper sparse-rounding theorem under an explicit weighted influence-stability certificate. For pseudometric-weighted CC, this gives a $10/3$ exact-marginal baseline and an unconditional coarse sparse-marginal bound; a sharp same-edge bound governed by $\overline{\Gamma}_w$ follows conditionally if the rounding rule is additionally proved to satisfy edge-weight-dominated influence stability. Third, we show that uniform marginal sampling has a triangle-witness threshold at $m=\Theta(n^{3/2})$ for typical pairs, while $m=O(n^{3/2}\sqrt{\log n})$ is sufficient for all pairs with high probability. Finally, in the stricter sparse edge-observation model, we prove an $\Omega(n^{3/2})$ lower bound on the expected approximation ratio from $o(n)$ uniformly sampled edges for general weighted instances. Experiments illustrate the witness transition and the proposed diagnostics; the formal guarantees are stated independently of those empirical observations.

Spectral graph sparsification is a classical tool for reducing graph complexity while preserving Laplacian quadratic forms. In graph neural networks (GNNs), sparsification is often used to accelerate computation while maintaining predictive performance. In this work, we study a complementary representation-level question: does sparsification preserve the geometry of learned embeddings? For polynomial-filter GNNs, we prove that any $\epsilon$-spectral sparsifier induces $O(\epsilon)$ perturbations in polynomial graph filters, multilayer hidden representations, and their Gram matrices. These guarantees imply stability of squared pairwise distances, class means, and covariance structure in embedding space. We further establish finite-time training stability: under smoothness and boundedness assumptions, gradient descent on dense and sparsified graphs produces weight trajectories whose separation grows at most proportionally to the sparsification distortion. Empirically, effective-resistance sparsification validates the predicted perturbation chain on synthetic graphs and preserves hidden representation geometry on real datasets. In our experiments, the gram matrix and training dynamics show low divergence even under substantial sparsification, consistent with the predicted stability under spectral sparsification. Hidden Gram preservation strongly predicts neighborhood preservation and class-centroid stability across FashionMNIST, Cora, and Paul15. Together, these results show that spectral sparsification preserves not only graph operators, but also the representation geometry that supports downstream use of GNN embeddings for interpretability.

Mamba's selective state-space model (SSM) achieves long-range dependence through input-dependent gating, yet existing analyses offer limited insight into *why* a particular parameter configuration produces a particular effective receptive field (ERF). We define per-step *conductances* $G_t = I - \exp(A\,\Delta_t)$ from the zero-order-hold discretization and construct a signed spectral measure whose atoms are the per-dimension conductances weighted by the corresponding input-output residues. The impulse response at lag $\tau$ is a discrete Laplace transform of this measure, and the ERF is governed by the per-channel squared impulse response weighted by output amplitude. Heavy tails arise from the mixture across state dimensions, not from temporal randomness within any single dimension. We extract conductances from trained Mamba models and report three empirical findings. First, replacing the signed measure with its total-variation envelope overestimates the ERF by more than three orders of magnitude. Signed cancellations among residues are what make the prediction quantitative. Second, the framework generalizes beyond the original single-layer validation setting. Applied to held-out $A$-initializations, a two-layer model trained from scratch, and all $48$ layers of a pretrained Mamba-370M, the spectral measure tracks the frozen-parameter ERF within $0.85$-$1.5{\times}$ on $41$ of $48$ pretrained layers, with log-log slopes matching to within $0.05$ on $35$ layers and within $0.20$ on all $48$. Third, varying the initialization of $A$ causally shifts the ERF tail slope as the spectral prediction anticipates ($-0.92$ for slow init vs. $-2.6$ for default), and this shift translates to task performance. Across two long-context tasks evaluated over three random seeds, the slow initialization is the most reliable condition: it solves MQAR in $3/3$ seeds (vs. $1/3$ for both default and fast), and attains the highest mean selective-copy accuracy on every distance bucket ($0.68\pm0.13$ overall vs. $0.34\pm0.15$ for fast at distances up to $1500$ tokens, with the ranking preserved in a single-seed extension out to $4000$ tokens). The spectral measure thus provides a controllable design lever for long-range memory. The closed-form prediction reproduces the in-sample frozen-parameter ERF at ratio $1.00$ by construction, which we use as a pipeline consistency check.


Statistical Estimation of Adversarial Risk in Large Language Models under Best-of-N Sampling

Mingqian Feng ⋅ Xiaodong Liu ⋅ Weiwei Yang ⋅ Chenliang Xu ⋅ Christopher White ⋅ Jianfeng Gao

Large Language Models (LLMs) are typically evaluated for safety under single-shot or low-budget adversarial prompting, which underestimates real-world risk. In practice, attackers can exploit large-scale parallel sampling to repeatedly probe a model until a harmful response is produced. While recent work shows that attack success increases with repeated sampling, principled methods for predicting large-scale adversarial risk remain limited. We propose a scaling-aware Best-of-N estimation of risk, SABER, for modeling jailbreak vulnerability under Best-of-N sampling. We model sample-level success probabilities using a Beta distribution, the conjugate prior of the Bernoulli distribution, and derive an analytic scaling law that enables reliable extrapolation of large-N attack success rates from small-budget measurements. Using only n=200 samples, our anchored estimator predicts ASR@1000 with a mean absolute error of 1.44, compared to 6.58 for the baseline, which is a 78.1% reduction in estimation error. Our results reveal heterogeneous risk scaling profiles and show that models appearing robust under standard evaluation can experience rapid nonlinear risk amplification under parallel adversarial pressure. This work provides a low-cost, scalable methodology for realistic LLM safety assessment. We provide our code in the Supplementary Material.


Sticky Jump Diffusions: A Unifying Framework for Discrete, Continuous, and Hybrid Diffusion

Pascal J Dube ⋅ Patrick Pynadath ⋅ Jeremy Lu ⋅ Yuan Gao ⋅ Ruqi Zhang

We introduce Sticky Jump Diffusions (SJDs), continuous-time Markov processes on $\mathbb R^d$ with a discrete anchor set identified with token embeddings. In forward time, anchors release their mass at a hazard rate and the released mass diffuses in the continuous ambient space; time reversal couples a score-driven SDE with a sticky jump kernel whose rate and destination are fixed by flux balance with the forward law. We estimate the score and the per-anchor reverse hazards from a single denoising classifier via Denoising Hazard Matching, the hazard analogue of denoising score matching, with simulation-free cross-entropy training. SJD generalizes the classical sticky-boundary diffusions of Feller and Itô to vocabulary-sized anchor sets, and recovers masked diffusion, continuous diffusion, and hybrid diffusion as limits. Beyond these limits, the framework exposes two design axes that previous work hold fixed: the un-sticking kernel, which encodes both per-anchor geometry and cross-position structure of the corruption, and the per-anchor forward hazard, which encodes the commitment schedule. We evaluate SJD on CIFAR-10, ImageNet $64{\times}64$, Text8, and Sudoku, where it is competitive with discrete, continuous, and hybrid baselines.


Stochastic Interpolants via Conditional Dependent Coupling

Chenrui Ma ⋅ Xi Xiao ⋅ Tianyang Wang ⋅ Xiao Wang ⋅ Yanning Shen

Existing image generation models face critical challenges regarding the trade-off between computation and fidelity. Specifically, models relying on a pretrained Variational Autoencoder (VAE) suffer from information loss, limited detail, and the inability to support end-to-end training. In contrast, models operating directly in the pixel space incur prohibitive computational cost. Although cascade models can mitigate computational cost, stage-wise separation prevents effective end-to-end optimization, hampers knowledge sharing, and often results in inaccurate distribution learning within each stage. To address these challenges, we introduce a unified multi-stage generative framework formulated under the \emph{stochastic interpolant} formalism with our \textbf{Conditional Dependent Coupling} strategy as a Flow-Matching variant. It decomposes the generative process into interpolant trajectories at multiple stages, ensuring accurate distribution learning while enabling end-to-end optimization. Importantly, the entire process is modeled as a single unified Diffusion Transformer, eliminating the need for disjoint modules and also enabling knowledge sharing. Under explicitly stated assumptions, we provably reduce both transport cost and asymptotic inference time, we empirically validate the underlying NFE--transport-cost and per-evaluation-cost relations. Extensive experiments demonstrate that our method achieves competitive fidelity and substantially lower wall-clock cost in the low-NFE regime across multiple resolutions.


StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos

Daeun Lee ⋅ Subhojyoti Mukherjee ⋅ Branislav Kveton ⋅ Ryan Rossi ⋅ Viet Lai ⋅ Seunghyun Yoon ⋅ Trung Bui ⋅ Franck Dernoncourt ⋅ Mohit Bansal

Streaming video understanding requires models not only to process temporally incoming frames, but also to anticipate user intention for realistic applications such as Augmented Reality (AR) glasses. While prior streaming benchmarks evaluate temporal reasoning, none measure whether Multimodal Large Language Models (MLLMs) can interpret or leverage human gaze signals within a streaming setting. To fill this gap, we introduce StreamGaze, the first benchmark designed to evaluate how effectively MLLMs utilize gaze for temporal and proactive reasoning in streaming videos. StreamGaze introduces gaze-guided past, present, and proactive tasks that comprehensively assess streaming video understanding. These tasks evaluate whether models can use real-time gaze signals to follow shifting attention and infer user intentions based only on past and currently observed frames. To build StreamGaze, we develop a gaze–video Question Answering (QA) generation pipeline that aligns egocentric videos with raw gaze trajectories through fixation extraction, region-specific visual prompting, and scanpath construction. This pipeline produces spatio-temporally grounded QA pairs that reflect human perceptual dynamics. Across all StreamGaze tasks, we observe substantial performance gaps between state-of-the-art MLLMs and human performance, highlighting key limitations in gaze-based temporal reasoning, intention modeling, and proactive prediction. We further provide detailed analyses of gaze prompting strategies, reasoning behaviors, and task-specific failure modes, offering insights into current limitations and directions for future research. All data and code are publicly available to support research in gaze-guided streaming video understanding.


Structured Sparse Memory for Recurrent Reasoning

Zixuan Zhao ⋅ Sam Wheeler ⋅ Neil Getty ⋅ Xiaotian Duan ⋅ Rick Stevens ⋅ Fangfang Xia

Recurrent models trained from scratch have recently become competitive on ARC-style reasoning tasks, but the usual framing around small recurrent backbones overlooks two important parts of the system: task-conditioned memory and synthetic augmentation data. We study this regime through CHARM, a compact hybrid ARC model that combines recurrent reasoning with structured task memory, synthetic data, and inference-time aggregation. In existing approaches, task-conditioned memory supplies a large hidden source of capacity, reaching more than 30x the size of the recurrent backbone. We introduce a compositional sparse embedding for task conditioning that reduces learned task-memory parameters by over 90% while improving pass@2 in controlled ARC-AGI-1 ablations. We also find that matched synthetic data is the largest measured driver of ARC-AGI-2 gains. For the recurrent backbone, recurrent depth helps only when balanced with learning horizon. Combining these ingredients, our system reaches 80.5% pass@2 on ARC-AGI-1 and 39.3% pass@2 on ARC-AGI-2 public evaluation.

Scalarization is widely used in multi-objective optimization owing to its simplicity and scalability. In many applications, the goal is to generate solutions that represent diverse user preferences, ideally with uniform coverage of the Pareto front (PF). However, uniformly sampling scalarization weights usually induces non-uniform coverage of the PF. We explain this mismatch through a geometric analysis of the scalarization path. As the scalarization weight varies, the corresponding solutions trace the PF with a generally non-uniform traversal speed. This speed induces an arc-length cumulative distribution function (CDF); inverting this map yields a principled rule for selecting weights that produce uniform PF coverage. Building on this insight, we propose SURF (Sampling Uniformly along the Pareto Front). For structured problems, including bi-objective bandits, we derive closed-form expressions for this CDF map and the resulting PF-aware weight sampling rule. For general problems, SURF alternates between CDF reconstruction and weight sampling. Theoretically, we show that under provable conditions, SURF converges linearly to an unavoidable finite-sampling floor. Empirically, experiments on bandits and MO-Gymnasium benchmarks demonstrate that SURF efficiently achieves more uniform PF coverage than baselines.


Swift Sampling: Selecting Temporal Surprises via Taylor Series

Dahye Kim ⋅ Bhuvan Sachdeva ⋅ Karan Uppal ⋅ Naman Gupta ⋅ Vineeth N Balasubramanian ⋅ Deepti Ghadiyaram

While most frames in long-form video are redundant, the critical information resides in $\textit{temporal surprises}$: moments where the actual visual features deviate from their predicted evolution. Inspired by the human brain’s predictive coding, we introduce $\textbf{Swift Sampling}$, an elegant, training-free frame selection algorithm that automatically identifies high-information moments in a video. Specifically, we model a video as a differentiable trajectory in the visual latent space and compute the velocity and acceleration of its features. Then, we apply Taylor expansion to project the expected path of subsequent frames. Frames that diverge sharply from this predicted manifold are identified as temporally surprising frames and selected for sampling. Unlike prior training-free methods that rely on auxiliary networks or video-specific hyperparameter tuning, Swift Sampling is incredibly lightweight, adding only $\mathbf{0.02\times}$ additional computational cost over baseline making it $30\times$ cheaper overhead than leading baselines. Across three long-video question answering benchmarks and $10$ different downstream tasks, Swift Sampling outperforms uniform sampling and prior query-agnostic baselines. It is especially powerful for long videos with limited frame budgets improving accuracy by up to $\mathbf{+12.5}$ points.


Systematic Scaling Analysis of Jailbreak Attacks in Large Language Models

Xiangwen Wang ⋅ Ananth Balashankar ⋅ Varun Chandrasekaran

Large language models remain vulnerable to jailbreak attacks, yet we still lack a systematic understanding of how jailbreak success scales with attacker effort across methods, model families, and harm types. We initiate a scaling-law framework for jailbreaks by treating each attack as a compute-bounded optimization procedure and measuring progress on a shared FLOPs axis. Our systematic evaluation spans four representative jailbreak paradigms, covering optimization-based attacks, self-refinement prompting, sampling-based selection, and genetic optimization, across multiple model families and scales on a diverse set of harmful goals. We investigate scaling laws that relate attacker budget to attack success score by fitting a simple saturating exponential function to FLOPs--success trajectories, and we derive comparable efficiency summaries from the fitted curves. Empirically, prompting-based paradigms tend to be the most compute-efficient compared to optimization-based methods. To explain this gap, we cast prompt-based updates into an optimization view and show via a same-state comparison that prompt-based attacks more effectively optimize in prompt space. We also show that attacks occupy distinct success--stealthiness operating points with prompting-based methods occupying the high-success, high-stealth region. Finally, we find that vulnerability is strongly goal-dependent: harms involving misinformation are typically easier to elicit than other non-misinformation harms.


TACT-KV: Tri-Axis Cosine Transform for Compressing Volumetric KV Caches in Medical VLMs

Chongyu Qu ⋅ Ritchie Zhao ⋅ Yufan He ⋅ Dong Yang ⋅ Zhengyi Lu ⋅ Junchao Zhu ⋅ Tianyuan Yao ⋅ Juming Xiong ⋅ Junlin Guo ⋅ Yanfan Zhu ⋅ Yuechen Yang ⋅ Daguang Xu ⋅ Bennett Landman ⋅ Yucheng Tang ⋅ Yuankai Huo

Vision-language models (VLMs) applied to 3D medical scans face a structural memory wall. A single CT or MRI volume produces tens of thousands of vision tokens, and the resulting key-value (KV) cache scales linearly with both batch and context. Existing KV compression methods operate on the cache's feature or token statistics, or apply 1D and 2D transforms, leaving the three-axis spatial structure in the cache unexploited. We propose TACT-KV (Tri-Axis Cosine Transform), a calibration-free KV compression method that exploits strong spatial correlations in visual KV states along all three anatomical axes of 3D scans. This correlation concentrates the KV cache energy under a frequency-domain transform. TACT-KV applies a separable 3D discrete cosine transform to the KV cache, allocates a per-layer bit budget across frequency bins via reverse water-filling (more bits to high-energy components), and quantizes each bin with a fixed scalar codebook. We derive a close-form relation: the distortion reduction from adaptive bit allocation is dominated by the spectral concentration of the transformed cache. Across six medical VLMs and two 3D CT benchmarks, TACT-KV maintains near-lossless prediction fidelity at 1-bit. The same advantage reproduces on four general-purpose VLMs evaluated on video tasks, indicating that the benefit of exploiting three-axis structure generalizes beyond medical imaging. At equal high-bandwidth memory (HBM), TACT-KV scales the decoding batch size on a single GPU by up to $8\times$.

Decentralized Stochastic Approximation (SA) enables cooperative fixed-point solving over mesh networks without a central server. In these systems, local update is relatively cheap compared with coordination, and frequent information exchange makes communication cost the main bottleneck. To alleviate this issue, we propose a decentralized SA method in which each agent performs $H$ local updates before each communication round. By amortizing communications over multiple local SA updates, the method can substantially reduce communication cost. We also study a multiple mixing scheme based on FastMix, which trades a small amount intra-round communications for a stronger consensus effectiveness on poorly connected graphs. Theoretically, we establish finite-time bounds under general contractive norms and Markovian sampling, by combining contraction, consensus error, and delayed Markovian terms in a round-level Lyapunov recursion. The results show trade-offs among computation, communication and sampling. Numerical experiments on decentralized Q-learning, Markovian SGD and linear SA validate the theory and show substantial communication savings for target accuracy.

Transformers are effective at inferring the latent task from context via two inference modes: recognizing a task seen during training, and adapting to a novel one. Recent interpretability studies have identified from middle-layer representations task-specific directions, or task vectors, that steer model behavior. However, a lack of rigorous foundations hinders connecting internal representations to external model behavior: existing work fails to explain how task-vector geometry is shaped by the training distribution, and what geometry enables out-of-distribution (OOD) generalization. In this paper, we study these questions in a controlled synthetic setting by training small transformers from scratch on latent-task sequence distributions, which allows a principled mathematical characterization. We show that two inference modes can coexist within a single model. In-distribution behavior is governed by Bayesian task retrieval, implemented internally through convex combinations of learned task vectors. OOD behavior, by contrast, arises through extrapolative task learning, whose representations occupy a subspace nearly orthogonal to the task-vector subspace. Taken together, our results suggest that task-vector geometry, training distributions, and generalization behaviors are closely related.


TaxaAdapter: Scaling Fine-grained Species Image Generation To the Tree of Life

Mridul Khurana ⋅ Amin Karimi Monsefi ⋅ Justin Lee ⋅ Medha Sawhney ⋅ David Carlyn ⋅ Julia Chae ⋅ Jianyang Gu ⋅ Rajiv Ramnath ⋅ Sara Beery ⋅ Wei-Lun (Harry) Chao ⋅ Anuj Karpatne ⋅ Cheng Zhang

Accurately generating images across the Tree of Life is difficult: there are over 10M distinct species on Earth, many of which differ only by subtle visual traits. Despite the remarkable progress in text-to-image synthesis, existing models often fail to capture the fine-grained visual cues that define species identity, even when their outputs appear photo-realistic. To this end, we propose TaxaAdapter, a simple and lightweight adapter that incorporates the taxonomic embeddings of Vision Taxonomy Models (VTMs) such as BioCLIP to perform scalable fine-grained species generation over the entire Tree of Life. TaxaAdapter injects VTM embeddings using taxonomy-text dual conditioning into a frozen text-to-image diffusion model, improving species-level fidelity while preserving flexible text control over attributes such as pose, style, and background. Extensive experiments demonstrate that TaxaAdapter consistently improves morphology fidelity and species-identity accuracy over strong baselines. To better evaluate these improvements, we also introduce a multimodal Large Language Model-based metric that summarizes trait-level descriptions from generated and real images, providing a more interpretable measure of morphological consistency. Beyond showing improvements over standard benchmarking experiments, we observe that TaxaAdapter exhibits strong generalization capabilities in open-world settings, enabling species synthesis in challenging regimes such as few-shot species with only a handful of training images and even species unseen during training. Overall, our results highlight that VTMs are a key ingredient for scalable, fine-grained species generation.


TDBench: Benchmarking Vision Language Models on Top-Down Image Understanding

Kaiyuan Hou ⋅ Minghui Zhao ⋅ Lilin Xu ⋅ Yuang Fan ⋅ Xiaofan Jiang

Top-down images are critical in applications such as autonomous navigation and aerial surveillance, whereas Vision-Language Models (VLMs) are primarily trained and evaluated on front-view benchmarks, leaving their performance in top-down settings largely unexplored. Moreover, conventional evaluation protocols based on single-pass accuracy or repeated testing on identical questions can overestimate model capability due to hallucinations or chance correctness. To address these limitations, we introduce a 2{,}000-question benchmark over ten task categories for drone-altitude image understanding, and RotationalEval (RE), an evaluation protocol that leverages rotational invariance to measure answer consistency across multiple orientations of the same scene. Human evaluators lose almost nothing going from single-pass accuracy to RE, whereas across the 60 open-source and proprietary VLMs we evaluated, accuracy drops by 15 to 26 points. We analyze per-question rotation-outcome distributions across the models we evaluated with an Empirical Mixture Model that compares each model's distribution over the number of correctly-answered rotations (0 to 4) to a difficulty-adjusted Poisson-binomial null, splitting the residual into a mastery component and a consistent-failure component. The benchmark is integrated into an open-source evaluation toolkit and will be released upon publication.

Accurately solving time-dependent partial differential equations (PDEs) with neural networks remains challenging due to long-time error accumulation and the difficulty of enforcing general boundary conditions. We introduce TENG-BC, a high-precision neural PDE solver based on the Time-Evolving Natural Gradient, designed to perform under generic boundary constraints. At each time step, TENG-BC performs a boundary-aware optimization that jointly enforces interior dynamics and boundary conditions, accommodating Dirichlet, Neumann, Robin, and mixed types within a unified framework. This formulation admits a natural-gradient interpretation, enabling stable time evolution without delicate penalty tuning. Across benchmarks over diffusion, transport, and nonlinear PDEs with various boundary conditions, TENG-BC achieves solver-level accuracy under comparable sampling budgets, outperforming conventional solvers and physics-informed neural network (PINN) baselines.


Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

Zhiyuan (Paul) Zhou ⋅ Andy Peng ⋅ Charles Xu ⋅ Qiyang Li ⋅ Jost Springenberg ⋅ Kevin Frans ⋅ Sergey Levine

Expressive continuous control policies, such as diffusion and flow models, form the backbone of recent advances in scaling imitation learning for simulated and real robot control. While they are known to scale stably in the supervised imitation learning setting, incorporating them into reinforcement learning (RL) pipelines for policy improvement has proven more difficult. It often requires specialized training objectives or back-propagating through denoising processes, which cause well known issues with stability and affects scalability. In this paper we study the question of whether simple policy improvement schemes at test time alone, leaving stable supervised policy training intact, can be a competitive alternative which side-steps these issues. To this end, we propose QGF (Q-Guided Flow), a new RL algorithm that performs policy optimization entirely during test time. QGF works by pre-training both a reference flow policy (via a standard behavioral cloning objective) and a value function critic and, during test-time, using the value gradient to guide the reference policy to generate higher-value actions without any additional policy learning. Empirically, QGF outperforms prior test-time RL methods while being much cheaper to run, and is competitive with state-of-the-art training-time algorithms on single-task and goal-conditioned offline RL benchmarks with high-dimensional action spaces. Moreover, it exhibits favorable scaling with model size by avoiding the instability of actor-critic training, offering a practical and effective alternative RL algorithm with expressive policies.


The Path Not Taken: RLVR Learns Off the Principals

Hanqing Zhu ⋅ Zhenyu Zhang ⋅ Hanxian Huang ⋅ DiJia Su ⋅ Zechun Liu ⋅ Jiawei Zhao ⋅ Igor Fedorov ⋅ Hamed Pirsiavash ⋅ Jinwon Lee ⋅ David Z. Pan ⋅ Zhangyang "Atlas" Wang ⋅ Yuandong Tian ⋅ Kai Sheng Tai

RLVR reliably improves reasoning, yet it appears to update only a tiny fraction of weights. We resolve this paradox by revealing a persistent, model-conditioned optimization bias: independent runs concentrate updates in similar parameter regions, invariant to dataset or RL recipe, while finite-precision storage (e.g., bf16) obscures widespread micro-updates as a visual sparsity artifact. To characterize this unique bias, we show that RLVR preferentially learns along off-principal directions via a cascaded Three-Gate mechanism: updates are first bounded by a empirical KL constraint (Gate I), then steered by the model’s anisotropic geometry(Gate II) toward spectrum-preserving off-principal subspaces, and finally filtered by finite precision (Gate III), effectively masking minor updates. Empirically, we validate the off-principal dynamics: RLVR exhibits minimal spectral drift, reduced principal-subspace rotation, and strong off-principal alignment that set RLVR strikingly apart from SFT. Together, our results provide the first parameter-space account of RLVR and uncover consistent regularities in weight evolution, advancing a more white-box understanding of RLVR. Moreover, we show that RLVR follows an optimization regime distinct from SFT, showing directly that transferring SFT-era PEFT can be flawed and motivating geometry-aware, RLVR-native methods.

The Forward-Forward Algorithm (FFA) replaces backpropagation (BP) with layer-wise local contrastive objectives, eliminating the backward pass, yet suffers a persistent performance gap with BP that worsens with depth. This paper diagnoses two distinct deficits: an irreducible optimization floor arising from concurrent local updates, whose magnitude is amplified when kernel contraction degrades the optimization Gram; and a geometric collapse of layer representations driven by kernel contraction itself. On the optimization side, we prove that the FFA loss satisfies the Polyak--{\L}ojasiewicz inequality at each layer, but concurrent layer updates create an irreducible error floor that grows with depth and is amplified when the Gram degrades. On the representational side, the pairwise similarity kernel of layer representations contracts exponentially toward rank one as depth increases, collapsing the diversity of per-layer error signals. This collapse bounds FFA's \emph{effective learning capacity}---the total diversity of gradient information across layers---to grow only linearly with depth regardless of width, whereas BP's chain-rule signal preserves per-layer diversity, yielding a capacity that scales with both depth and width.

Generative models trained on finite data face a fundamental tension: their score-matching or next-token objective converges to the empirical training distribution rather than the population distribution we seek to learn. Using rule-valid synthetic tasks, we trace this tension across two training timescales: $\tau_{rule}$, the step at which generations first become rule-valid, and $\tau_{mem}$, the step at which models begin reproducing training samples. Focusing on parity and extending to other binary rules and combinatorial puzzles, we characterize how these two clocks $\tau_{rule}$, $\tau_{mem}$ depend on key aspects of the learning setup. Specifically, we show that $\tau_{rule}$ increases with rule complexity and decreases with model capacity, while $\tau_{mem}$ is approximately invariant to the rule and scales nearly linearly with dataset size $N$. We define the \emph{innovation window} as the interval $[\tau_{rule}, \tau_{mem}]$. This window widens with increasing $N$ and narrows with rule complexity, and may vanish entirely when $\tau_{rule} \geq \tau_{mem}$. The same two-clock structure arises in both diffusion (DiT) and autoregressive (GPT) models, with architecture-dependent offsets. Dissecting the learned score of DiT models reveals a corresponding evolution of the optimization landscapes, where attractors emerge at both timescales: rule-valid samples' basins expand substantially around $\tau_{rule}$, while training samples' basins begin to dominate around $\tau_{mem}$. Together, these results yield a unified and predictive account of when and how generative models exhibit genuine innovation.


The Unembedding Bottleneck: A Mechanistic Analysis of Single Digit Counting in LLMs

Satwik Sunnam ⋅ Raghav Magazine ⋅ Vatsalya Singh ⋅ Lavanya Kotha ⋅ Xingjian Li ⋅ Min Xu

Large language models consistently fail at elementary counting tasks despite strong performance on complex reasoning benchmarks. Recent work has characterized counting failures behaviorally or identified circuits through which counting succeeds, but a mechanistic account of where and why counting breaks down remains absent. In this work, we provide such an account by tracing the count signal from embedding to output and identifying the geometric bottleneck that prevents correct readout for single digit counting problems. Using a dual-metric framework combining Signal-to-Noise Ratio (SNR) with linear probing, we show that counting information is computed correctly and persists in a linearly decodable subspace through the final layer. The failure in counting is not representational but geometric i.e the digit-token columns of the unembedding matrix $W_U$ are orthogonal to the count subspace in the residual stream, rendering correctly computed counts unreadable at the output. We term this as the $\textbf{Unembedding Bottleneck}$ and to causally validate it, we apply a $\textbf{Probe-Informed LoRA}$ correction targeting $W_U$, which restores single-digit counting accuracy by up to 80-95% across six models while updating only ${\sim}$0.001% of parameters and preserving general capabilities. Our findings reveal a broader class of failure in LLMs where information is internally present but geometrically inaccessible to the output projection.


The Z-Gromov-Wasserstein Distance

Martin Bauer ⋅ Facundo Mémoli ⋅ Tom Needham ⋅ Mao Nishino

The Gromov-Wasserstein (GW) distance is a powerful tool for comparing metric measure spaces which has found broad applications in data science and machine learning. Driven by the need to analyze data sets whose objects have increasingly complex structure (such as node and edge-attributed graphs), several variants of GW distance have been introduced in the recent literature. With a view toward establishing a general framework for the theory of GW-like distances, this paper considers a vast generalization of the notion of a metric measure space: for an arbitrary metric space Z, we define a Z-network to be a measure space endowed with a kernel valued in Z. We introduce a method for comparing Z-networks by defining a generalization of GW distance, which we refer to as Z-Gromov-Wasserstein (Z-GW) distance. This construction subsumes many previously known metrics and offers a unified approach to understanding their shared properties. This paper demonstrates that the Z-GW distance defines a metric on the space of Z-networks which retains desirable properties of Z, such as separability, completeness, and geodesicity. Many of these properties were unknown for existing variants of GW distance that fall under our framework. Our focus is on foundational theory, but our results also include computable lower bounds and approximations of the distance which will be useful for practical applications.


ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model

Haichao Zhang ⋅ Yijiang Li ⋅ Shwai He ⋅ Tushar Nagarajan ⋅ Mingfei Chen ⋅ Jianglin Lu ⋅ Ang Li ⋅ Yun Fu

Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting future world states from video observations. Nevertheless, dense prediction from a short observation window limits temporal context and can bias predictors toward local, low-level extrapolation, making it difficult to capture long-horizon semantics and reducing downstream utility. Vision-language models (VLMs), in contrast, provide strong semantic grounding and general knowledge by reasoning over uniformly sampled observation frames, but they are not ideal as standalone dense predictors due to compute-driven sparse sampling, a language-output bottleneck that compresses fine-grained interaction states into text-oriented representations, and a data-regime mismatch when adapting to small action-conditioned datasets. We propose a VLM-guided JEPA-style latent world modeling framework that combines dense-frame dynamics modeling with long-horizon semantic guidance via a dual-temporal pathway: a dense JEPA branch for fine-grained motion and interaction cues, and a uniformly sampled observation VLM thinker branch with a larger temporal stride for knowledge-rich guidance. To transfer the VLM's progressive reasoning signals effectively, we introduce a hierarchical pyramid representation extraction module that aggregates multi-layer VLM representations into guidance features compatible with latent prediction. Across EgoDex, EgoExo4D, BAIR Robot Pushing, and Physion, ThinkJEPA outperforms diverse latent world model and trajectory prediction baselines across egocentric trajectory prediction, long-horizon rollout, robotic latent prediction, and physical-scene forecasting. These results show that broad visual-semantic guidance from a VLM thinker can benefit JEPA-style latent forecasting.

As web agents close the gap with humans on benchmarks, it raises the question: *Do today's agents perform just as well on tomorrow's web?* We introduce TimeWarp, a benchmark that emulates the evolving web. TimeWarp consists of three web environments, each with six UI versions spanning UI design and frontend code from different eras of the internet. We pair TimeWarp with a set of complex, realistic tasks covering different forms of web navigation. Our experiments reveal vulnerabilities of web agents, especially visual ones, to changes and the limitations of behavior cloning (BC) on complex trajectories from a single version. To address this, we propose TimeTraj, a simple yet effective algorithm that uses plan distillation to collect trajectories across multiple versions. By training agents on teacher rollouts using our BC-variant, we achieve substantial performance gains: $20.4$ %$\rightarrow37.7$% for Qwen-3 4B and $0$ % $\rightarrow27.0$% for Llama-3.1 8B models. Our work helps study generalization across web designs and enables a new paradigm for collecting plans rather than trajectories to improve the robustness of web agents.


TKCAM: Text and Keyframe to Camera Trajectory Generation

Haozhe Yang ⋅ Zhiyang Dou ⋅ Zekai Gu ⋅ Cheng Lin ⋅ Wenping Wang ⋅ Yuan Liu ⋅ Taku Komura

Generating high-quality and controllable camera motion is essential for AI-assisted cinematography, video synthesis, and 3D scene understanding. We introduce TKCAM, a Text- and Keyframe-conditioned CAMera-motion synthesis framework based on generative masked modeling. To overcome the instability of direct 3D regression, we formulate camera dynamics as continuous 12-DoF kinematic sequences and discretize them into hierarchical motion tokens via a Residual Vector Quantizer (RVQ). A two-stage masked transformer architecture then learns to reconstruct and refine these tokens, utilizing explicit self- and cross-attention modules for robust multi-modal conditioning. A central feature of our framework is animator-style keyframe control: users can provide free-form text prompts alongside sparse key poses, and TKCAM seamlessly inpaints temporally coherent in-between trajectories. Furthermore, to advance evaluation standards, we curate RealEstate10K-Cap, a large-scale text-camera dataset, and establish a comprehensive cross-domain benchmark with a Universal CLaTr Evaluator. Extensive experiments demonstrate that TKCAM significantly surpasses recent state-of-the-art baselines on Fréchet Inception Distance (FID), text-motion matching scores, and retrieval metrics (R@K), producing cinematographically plausible motions while enabling precise spatial guidance. Our code, models, and dataset will be open-sourced upon acceptance.

Knowledge graph learning provides a powerful framework for representing and inferring structured knowledge, with broad practical applications. However, the scarcity of relation-specific labeled triples per entity hinders the training of expressive models, and the ad hoc design of scoring functions limits generalizability and lacks theoretical grounding. We address both issues with a theoretically grounded, end-to-end training framework that extends and subsumes existing methods. Our framework is a two-stage procedure: unsupervised pretraining over heterogeneous corpora followed by supervised learning with diverse relationship types. We establish a nonasymptotic risk bound that disentangles pretraining approximation error from labeled-sample complexity, formally quantifying the benefit of large-scale unlabeled data for downstream knowledge prediction. Synthetic experiments validate each theoretical component, and real-world experiments confirm the effectiveness of our approach on large-scale knowledge graph benchmarks.


Towards Reconstructing Geographically Diverse Architecture with 3D Foundation Models

Aniket Kriplani ⋅ Yiwen Zhang ⋅ Angelina Wang ⋅ Hadar Averbuch-Elor

Internet-scale photo collections of real-world landmarks have driven progress in 3D computer vision yet remain highly challenging for modern 3D foundation models (3DFMs) that estimate scene structure in a single feed-forward pass. In this work, we introduce ArchWorld, a benchmark of 3D architectural landmarks with high geographic coverage and rich metadata and use it to conduct a thorough error analysis of current 3DFMs to understand where they fall short and how they can be improved. We examine socially relevant disparities in model error and find that model performance varies by geographic region, but this is heavily confounded by scene size. Leveraging insights from model failings, we introduce an intervention scheme that improves 3DFM performance while simultaneously reducing geographic disparities. Our code and data will be publicly available.


Towards Understanding Momentum Acceleration in River-Valley Loss Landscape

Miao Lu ⋅ zeyu bian ⋅ Kaiyue Wen ⋅ Beining Wu ⋅ Siyu Chen ⋅ Tianhao Wang ⋅ Zhiyuan Li

Momentum is a critical and ubiquitous component of modern optimizers, while the role of momentum remains unclear beyond restricted settings, especially in optimization for large-scale neural networks. Recent studies suggest that the highly non-convex loss landscape for large language models exhibits certain “river-valley” structure: a low-loss manifold (the river) bordered by sharp, high-loss directions (the valley), where the essential optimization progress is determined primarily by the progress along the river in the long run. Motivated by this structure, in this work, we investigate the role of heavy-ball momentum in such an emerging setting. Specifically, we analyze gradient descent with heavy-ball momentum and show that compared to vanilla gradient descent, momentum can accelerate the progress along the river by enabling use of a substantially larger learning rate. In fact, momentum acts as a stabilizer in the presence of oscillations caused by an aggressive choice of learning rate, which the vanilla gradient descent cannot tolerate. We validate the insights with experiments on synthetic functions and language model training, offering practical guidance for tuning learning rate and momentum parameters.


Towards Visual Query Segmentation in the Wild

Bing Fan ⋅ Minghao Li ⋅ Hanzhi Zhang ⋅ Shaohua Dong ⋅ NAGA PRUDHVI MAREEDU ⋅ Xiaoqiong Liu ⋅ Weishi Shi ⋅ Yunhe Feng ⋅ Yan Huang ⋅ Heng Fan

In this paper, we introduce visual query segmentation (VQS), a new paradigm of visual query localization (VQL) that aims to segment all pixel-level occurrences of an object of interest in an untrimmed video, given an external visual query. Compared to existing VQL locating only the last appearance of a target using bounding boxes, VQS enables more comprehensive (ie, all object occurrences) and precise (ie, pixel-level masks) localization, making it more practical for real-world scenarios. To foster research on this task, we present VQS-4K, a large-scale benchmark dedicated to VQS. Specifically, VQS-4K contains 4,111 videos with more than 1.3 million frames, and covers a diverse set of 211 object categories. Each video is paired with a visual query defined by a frame outside search video and its target mask, and annotated with spatial-temporal masklets corresponding to the queried target. To ensure high quality, all videos in VQS-4K are manually labeled with meticulous inspection and iterative refinement. To the best of our knowledge, VQS-4K is the first benchmark specifically designed for VQS. Furthermore, to stimulate future research, we present a simple yet effective method, named VQ-SAM, which extends SAM 2 by leveraging target-specific and background distractor cues from the video to progressively evolve the memory through a novel multi-stage framework with an adaptive memory generation (AMG) module for VQS, significantly improving the performance. In our extensive experiments on VQS-4K, VQ-SAM achieves promising results and surpasses all existing approaches, demonstrating its effectiveness. With the proposed VQS-4K and VQ-SAM, we expect to go beyond current VQL paradigm and inspire more future research and practical applications on VQS. Our benchmark, code, and results will be made publicly available.


TraceDx: Criticality-Weighted Atomic Facts as a Training Signal for Sequential Clinical Diagnosis

Jiayi Li ⋅ Sunny Chung ⋅ Don C Codipilly ⋅ George Marek ⋅ Varan Perananthan ⋅ Dennis L. Shung ⋅ Bradly Stadie

In clinical decision-making, success depends on taking the right action to obtain the evidence that matters. One targeted clinical workup can identify a diagnosis that a battery of nonspecific tests missed. Large language models have shown considerable promise in clinical diagnosis, but most current approaches assume that all relevant patient information is available upfront, which does not reflect how diagnostic evidence is gathered in practice. Even when models gather evidence iteratively, challenges arise in determining which action to take, as they receive no feedback on which findings are diagnostically decisive. To solve this problem, we introduce TraceDx, a reinforcement learning framework for sequential clinical diagnosis. TraceDx is built on the idea that unstructured clinical records can be decomposed into atomic clinical findings, each scored by its diagnostic value. This signal rewards agents for discovering decisive evidence in addition to reaching the correct diagnosis. Consistent with this design, a case-level analysis shows that recovery of critical evidence is the only statistically significant predictor of diagnostic success. Direct case comparisons further show that TraceDx uncovers more diagnostically important evidence, and uses it to resolve the key diagnostic uncertainty. TraceDx-trained open-weight models outperform multiple frontier baselines on both MIMIC-CDM and a rare gastrointestinal disease dataset.


Tracing Moral Foundations in Large Language Models

Chenxiao Yu ⋅ Bowen Yi ⋅ Farzan K Malekabadi ⋅ Suhaib Abdurahman ⋅ Jinyi Ye ⋅ Shrikanth Narayanan ⋅ Yue Zhao ⋅ Morteza Dehghani

Large language models often produce human-like moral judgments, but it is unclear whether this reflects an internal conceptual structure or superficial ``moral mimicry.'' Using Moral Foundations Theory (MFT) as an analytic framework, we study how moral foundations are encoded, organized, and expressed across 14 base and instruction-tuned LLMs spanning four model families (Llama, Qwen2.5, Qwen3-MoE, Mistral) and scales from 7B to 70B. We employ a multi-level approach combining (i) layer-wise analysis of MFT concept representations and their alignment with human moral perceptions, (ii) pretrained sparse autoencoders (SAEs) over the residual stream to identify sparse features that support moral concepts, and (iii) causal steering interventions using dense MFT vectors and sparse SAE features. We find that models represent and distinguish moral foundations in a manner that aligns with human judgments, and that this moral geometry naturally emerges from pretraining and is selectively rewired by post-training. At a finer scale, SAE features show clear semantic links to specific foundations, suggesting partially disentangled mechanisms within shared representations. Finally, steering along either dense vectors or sparse features produces predictable shifts in foundation-relevant behavior, demonstrating a causal connection between internal representations and moral outputs. Together, our results provide mechanistic evidence that moral concepts in LLMs are distributed, layered, and partly disentangled, suggesting that pluralistic moral structure can emerge as a latent pattern from the statistical regularities of language alone.

We revisit the regret loss framework introduced in [Park, 2024], which uses decision-theoretic regret as a direct loss function for training models to make better decisions, through the lens of probability-simplex policies. Our first result shows that a single-layer self-attention model trained with regret loss admits a stationary point whose forward-pass exactly matches \emph{smoothed fictitious play} with the appropriate stepsize that ensures no-regret behavior—i.e., for any given policy input, the model outputs the same update that smoothed fictitious play would produce. In parallel, we also newly introduce a swap-regret loss function, which extends the regret-loss framework beyond external regret and enables models to directly optimize for swap-deviation robustness. We further show that this swap-regret loss admits a stationary point whose forward pass implements the corresponding swap-regret update induced by classical Blum–Mansour no-swap-regret algorithm, with each head implementing an external-regret update via smoothed fictitious play. Together, these results show that regret-trained attention can realize differentiable mechanisms whose deployment induces equilibrium behavior in games: external-regret dynamics lead to coarse correlated equilibrium, while swap-regret dynamics lead to correlated equilibrium. Thus, regret-based objectives steer minimal attention architectures toward online-learning dynamics with game-theoretic guarantees, without supervised traces of those algorithms.


Transferable, Time-Parallel Graph Dynamics via Space-Time Factorization

Wencong Yang ⋅ Chaopeng Shen ⋅ Abdolmehdi Behroozi

Time-parallel neural operators (FNO and successors) have transformed PDE surrogates on regular grids. Mesh-based operators (MeshGraphNets, Transolver) extend operator learning to irregular geometries but remain autoregressive in time and trained one geometry at a time, while classical parallel-in-time methods converge poorly on nonlinear advection-dominated PDEs. We propose space-time factorization: partition the model into local temporal operators (learned, graph-independent) and a spatial mixer that propagates information across nodes without graph-dependent learned parameters. This structural condition restores time-parallelism on irregular graphs and yields cross-graph zero-shot transfer once three information leaks—graph-dependent weights, graph-specific features, and graph-correlated training distributions—are controlled. The same protocol applies to two distinct spatial-mixer families tested here—fixed-spectral and learned-attention—suggesting the principle is not tied to a single mixer choice. We validate the principle in three domains. On simulating floodwave propagation in river networks, a 27K-parameter model trained on the Yangtze transfers zero-shot to the Mississippi at $R^2=0.998$, training $28\times$ faster than a recurrent message-passing baseline and scaling to $5\times$ larger graphs. On cross-geometry Navier–Stokes, both Chebyshev (SepFNO) and Physics-Attention (SepTransolver) variants reach within ~10% (relative $L_2$) of a locally-trained oracle using only 15 training geometries—to our knowledge the first demonstration of knowledge accumulation in this regime across fundamentally different topologies. On cross-city traffic forecasting, zero-shot transfer reaches within 12.9% of an oracle trained on the target city. A consistent pattern emerges across these settings: cross-geometry error decreases with both the number of training geometries and their relatedness to the test, with the rate of accumulation depending on the spatial mixer. The result opens the path for pooled-data, foundation-style models for graph dynamics on irregular topologies.

Training selects for behavior, not circuitry: many weight configurations can implement the same function. Studying any single trained neural network thus risks describing accidents of one training run rather than the computation itself. This work shifts focus from what transformers happen to do to what they must do by extracting algorithmic cores, compact subspaces that are necessary and sufficient for a task and that recur across independently trained models. Here, Algorithmic Core Extraction (ACE) is introduced to isolate these subspaces, causally validate them, and recover the algorithms they implement across settings ranging from synthetic tasks to large-scale pretrained models. Markov-chain transformers embed three-dimensional cores in nearly orthogonal subspaces yet recover identical transition spectra. Modular-addition transformers form compact cyclic cores at grokking that later inflate under continued regularization, redundantly distributing the same computation across many functionally equivalent modes. This functional redundancy is found to accelerate the transition from memorization to generalization, yielding an inverse scaling law for grokking time. In six language models spanning two orders of magnitude in scale (GPT-2 Small/Medium/Large, LLaMA-3.1, Gemma-2, and Qwen2.5), subject-verb agreement is governed by a single, steerable axis that aligns across architectures. Flipping this axis inverts grammatical number throughout open-ended generation. Together these results suggest that beneath the apparent complexity of trained transformers lies a simpler, shared computational structure, and that targeting invariants rather than parameterizations may offer a more tractable path to mechanistic understanding and control.

We investigate the ability of transformers to perform in-context reinforcement learning (ICRL), where a model must infer and execute learning algorithms from trajectory data without parameter updates. We show that a linear self-attention transformer block can provably implement policy-improvement methods, including semi-gradient SARSA and actor-critic, via explicit parameter constructions. Beyond existence, we design a teacher-mimicking training procedure, analyze its gradient-flow dynamics, and establish the first convergence guarantee in the ICRL literature: under suitable richness conditions on the training MDP distribution, gradient flow converges locally and exponentially to an optimal parameter manifold corresponding to the desired RL update. Empirically, training transformers on randomly generated tabular MDPs confirms these predictions: the learned models recover the parameter structure of our explicit constructions and, when deployed on unseen MDPs, deliver strong in-context control performance. Together, these results illuminate how transformer architectures internalize and execute classical reinforcement learning algorithms in context, bridging mechanistic understanding and training dynamics in ICRL.


Tree Search With Predictions

Michael Dinitz ⋅ Bob Dong

Algorithms with predictions, or learning-augmented algorithms, has proved to be an extremely useful paradigm for combining machine learning with traditional algorithms. One of the textbook settings for this is searching a sorted array. Without a prediction classical binary search takes $O(\log n)$ queries, while with a prediction we can use "doubling binary search" to find the target key using $O(\log \eta)$ queries, where $\eta$ is the error of the prediction measured as the absolute value of the difference between the true location and the predicted location. Since an array is just a path graph, in this paper we ask whether similar bounds can be achieved for search on even slightly more general graphs: trees. We show first that the high-level answer is "no": there is no search algorithm that uses $O(\log \eta)$ queries, where now $\eta$ is the graph distance between the predicted location and the true location. However, as our main result, we show that such bounds can be achieved on trees which are "path-like" in that they have low pathwidth. In particular, we prove that there is a search algorithm which uses at most $O(k \log \eta)$ queries, where $k$ is the pathwidth of the tree. We also prove a lower bound showing that our algorithm has existentially optimal query complexity. Finally, we show experimentally, on real-life inputs, that our algorithm has query complexity which is notably better than the simple non-prediction based algorithm.

Deploying language models as autonomous agents requires more than per-task accuracy: when an agent faces a queue of problems under a finite token budget, it must decide which to attempt, in what order, and how much compute to commit to each, all before any execution feedback is available. This is the prospective form of metacognitive control studied for decades in human cognition, yet whether language models possess it remains untested. We introduce TRIAGE, an evaluation framework in which a model receives a task pool and a token budget calibrated to its own baseline cost, and commits to a single ordered plan that jointly encodes selection, sequencing, and per-problem allocation. Plans are scored against an oracle with full knowledge of the model's solvability and cost on each problem, yielding a triage efficiency ratio on a common scale. We evaluate frontier and open-source models, with and without reasoning enabled, across competition mathematics, graduate-level science, code generation, and expert multidisciplinary knowledge, and find that current language models exhibit substantial gaps in prospective metacognitive control, revealing a previously unmeasured capability dimension with direct implications for resource-efficient agent deployment.

We introduce TriSearch, a reinforcement-learning framework for optimizing objectives over triangulations of a polytope via bistellar flips. The key idea is a circuit-supported subtriangulation action representation: feasible flips are encoded by their supporting circuit and realized local subtriangulation, enabling a learned policy to rank them using local geometric and combinatorial features. This yields a dimension-agnostic interface and enables efficient traversal of the flip graph without explicit enumeration of the full triangulation space. Instantiated in 3D and 4D, TriSearch generalizes zero-shot from small training instances to larger polytopes with exponentially larger search spaces. It achieves top performance on metric objectives in 3D and, in 4D, discovers more distinct Fine, Regular, Star triangulations of reflexive polytopes, corresponding to Calabi-Yau threefolds, than existing samplers under a fixed budget.

$(1+1)$-evolution strategy (ES) is a classic variant of evolution strategy for black-box optimization via direct and random search while keeping only a single solution per generation in $\mathbb{R}^n$. Despite convergence with different step-size adaptations, it often remains impractical due to dependence on the unfavorable high dimension of the search space $\mathbb{R}^n$. In fact, many modern optimization problems are naturally constrained to low-dimensional manifolds embedded in very high-dimensional ambient spaces. To see the potential of intrinsic dimensionality reduction for improving practicality, we consider the truncated Riemannian $(1+1)$-ES on a compact Riemannian submanifold $\mathcal{M}\subset\mathbb{R}^n$ and establish its non-asymptotic stationarity guarantee at the standard nonconvex zeroth-order rate of $\widetilde{\mathcal{O}}(\varepsilon^{-4})$ function evaluations. The explicit Gaussian second-moment term in the bound scales with the intrinsic tangent dimension $d=\dim\mathcal{M}$ rather than the ambient dimension $n$, while the remaining truncation dependence is captured by an explicit small-ball drift constant $\alpha_d(R/\sigma_0)$. To the best of our knowledge, this gives the first non-asymptotic stationarity analysis for a comparison-based ES-type method on compact nonconvex Riemannian submanifolds under localized pullback smoothness. Numerical illustrations on sparse PCA over spheres and smooth-surrogate stationarity diagnostics illustrate behavior consistent with the theory and its limitations, while additional experimental studies on truncation radius diagnostics, black-box tuning, and classical evolutionary computation benchmarks support the practical relevance of geometry-aware comparison-based search under limited function-evaluation budgets.


Understanding Circulant Permutation to Extend the CMinHash Estimator

Keegan Kang ⋅ Avery Hood ⋅ Benedict H Wong

CMinHash promises a new direction for designing MinHash algorithms with reduced variance and storage space via circulant permutations. However, the original variance analysis is difficult to extend to other MinHash estimators. We present a new approach that gives explicit and easily comparable approximate variance expressions, and use it to analyze the Minner estimator. Numerical experiments show that our variance expression matches the empirical mean square error (MSE) of estimates. Our work thereby demonstrates how to apply circulant permutation to other MinHash estimators and compute their approximate variance.

We consider two-layer neural networks trained in the feature-learning regime using gradient descent, and relate the output of the finite-width network $f_{\hat{\rho}}$ to its infinite-width counterpart $f_{\rho^{MF}}$, which evolves in the mean-field dynamics. While constant-time horizon bounds for $\|f_{\rho^{MF}} -f_{\hat{\rho}}\|$ may be obtained via standard Grönwall estimates, the long-time behavior of the fluctuation is a more delicate matter. Uniform-in-time bounds often rely on (local) strong convexity in the landscape or Logarithmic Sobolev inequalities present in noisy gradient dynamics. % In this work, we study a noiseless setting and do not make assumptions on the geometry of the landscape near the optimum. Instead, we impose a weaker condition on the convergence rate of the mean-field deterministic Wasserstein-gradient-flow dynamics. In this work, we establish non-asymptotic weak propagation-of-chaos that holds uniformly in time, obtained by exploiting instead the convergence rate of the mean-field deterministic Wasserstein-gradient-flow dynamics. Specifically, denoting by $L_t$ the mean-field loss at time $t$ and $m$ the number of neurons, under standard regularity assumptions and the condition $\int_0^\infty L_t^{1/2} dt =O_d(1)$, we obtain the uniform in time bound $\|f_{\rho^{MF}} -f_{\hat{\rho}}\|^2 \lesssim \text{poly}(d) m^{-\min(1,c/6)}$ whenever $L_t \lesssim t^{-c}$. Our result holds in a noiseless setting and does not make any assumptions on the geometry of the landscape near the optimum. A key implication of our result is that whenever the convergence rate of the mean-field, population-loss dynamics is faster than $1/t^2$, we can attain a loss of $\epsilon$ with only $\text{poly}(d/\epsilon)$ neurons, training samples, and GD steps.


Unsupervised Physics Informed Decomposition of Incomplete Time-Resolved Spectroscopy

Tom Muir ⋅ Qifeng Liu ⋅ Mohammadrahim Kazemzadeh ⋅ William Mills ⋅ Ales Leonardis ⋅ Huabing Yin ⋅ Alexander Krull

Raman spectroscopy is often hindered by strong autofluorescence backgrounds, which are typically handled by waiting for fluorescence to bleach before recording the Raman spectrum. Here, we instead show that short, incomplete recordings of the bleaching process already contain enough information to recover the underlying Raman signal. We present the first unsupervised model that operates on short, incomplete, variable-length sequences of time-resolved spectra, decomposing them into Raman spectra and autofluorescence components while simultaneously removing measurement noise. Our VAE-based method simultaneously learns (i) a bank of fluorophore spectra comprising the baseline, (ii) a model of the measurement noise, and (iii) the distribution of Raman spectra and decay behavior. We evaluate our method on spectral reconstruction and downstream peak detection, where it outperforms state-of-the-art approaches on multiple simulated benchmarks as well as on a newly collected real-world dataset of spectral time series. We further show that fluorescence decay dynamics themselves contain discriminative information that may be exploited in future work.


VeruSAGE-Bench: A Benchmark Suite for Rust System Verification

Chenyuan Yang ⋅ Natalie Neamtu ⋅ Chris Hawblitzel ⋅ Jacob R Lorch ⋅ Shan Lu

Large language models (LLMs) have shown impressive capability to understand and develop code. However, their capability to rigorously reason about and formally prove the correctness of software systems remains in question. Current benchmarks for formal proof generation focus on isolated tasks, typically algorithmic-style problems utilizing basic data structures. We curate a new benchmark suite for system-level proof generation, VeruSAGE-Bench, which consists of 849 proof tasks extracted from eight open-source Verus-verified Rust systems, covering operating systems, storage systems, memory allocators, and more. These tasks are much more complicated than those in previous benchmarks, containing on average more than 20x the formal specification (in Lines of Code) and requiring a much broader array of verification techniques. We evaluate five LLMs on VeruSAGE-Bench (o4-mini, GPT-5, Sonnet 4, Sonnet 4.5, Opus 4.5), and demonstrate that different models and agentic setups lead to vastly different success rates and costs. We also show that the best models possess impressive capability in writing correctness proofs and hence should inspire new ways of developing trustworthy software. Meanwhile, some difficult tasks in VeruSAGE-Bench stress even the best model, forcing long reasoning time (over an hour on the hardest tasks, 15.3 min per task on average), high token cost ($13.65 per task on average), and proof failures, which we hope will guide future improvement in model and agents.

Current foundation models of electroencephalography (EEG) sought to characterize the fluctuations in neural signals, yet this objective stands fundamentally at odds with the intrinsic physiological mechanisms governing brain activity. As with their LLM predecessors, these models pack millions of neurons, even those contributing negligibly small weights, yielding a computational footprint incompatible with resource-constrained portable brain-computer interfaces (BCIs). To overcome these challenges, we propose Vibe-Spike, an energy-preserving and biologically-realistic EEG foundation model operating in the regime of oscillatory neural synchronization. Drawing on principles from cognitive neuroscience, our foundation model is designed to recover the masked EEG signals through the transient spikes of spatially coupled neurons, from which self-organized cortical synchronization (aka. brain rhythms) emerges to reciprocally modulate ongoing neural firings. Under the hood, we introduce a macro-micro learning mechanism that bridges two scales of neural computation: lower layers capture neural oscillations via a spiking neural network (SNN), while upper layers model oscillatory synchronization across large neuronal populations, with cross-scale interactions learned end-to-end. Taken together, our Vibe-Spike sheds new light on learning intrinsic representation of functional neural fluctuations. We pre-train the model to reconstruct segmented EEG signals directly from the synchronized spiking representations using a purely self-supervised objective on a diverse corpus of 75 public datasets. To demonstrate its universal representational power, we conduct massive-scale evaluations across ten heterogeneous downstream datasets, including clinical pathology, sleep staging, and cognitive assessment. Vibe-Spike not only facilitates deployment on portable edge devices through its lightweight architecture and energy-preserving inference, but also provides profound neuroscience insights. By further capturing pre- and post-synaptic firing, Vibe-Spike unlocks a new pathway for characterizing effective connectivity and the oscillatory synchronization that underpins cognition.

AI-generated videos are becoming increasingly realistic, yet most detection methods focus on spatial artifacts within individual frames and require synthetic training data from known generators, causing them to fail on unseen generators. We identify two previously unexplored temporally distributed forensic microstructures that reveal whether a video was AI-generated. These traces arise from how a video's visual field and motion evolve over time, capturing subtle deviations that generators introduce when approximating real-world temporal dynamics. To expose them, we train two simple predictors on real video only, modeling visual evolution (VE-FSD) and motion evolution (ME-FSD), and analyze their prediction residuals. We fit 3D spatio-temporal autoregressive models to these residuals, producing compact descriptors we call Video Forensic Self-Descriptions (VFSD). VFSDs of real videos naturally cluster together and separate from AI-generated ones, enabling zero-shot detection and source attribution without synthetic training data. Experiments across multiple datasets and 35 generators show VFSD achieves state-of-the-art performance on both tasks.


VideoMLA: Low-Rank Latent KV Cache for Minute-Scale Autoregressive Video Diffusion

Hidir Yesiltepe ⋅ Jiazhen Hu ⋅ Tuna Han Salih Meral ⋅ Adil K Akan ⋅ Kaan Oktay ⋅ Hoda Eldardiry ⋅ Pinar Yanardag

Long-rollout causal video diffusion has converged on a fixed-size sliding-window KV cache, with recent progress innovating within this layout by changing which tokens occupy the window or how their positions are encoded. The per-head KV layout itself, a dominant contributor to streaming memory and latency, has been mostly left unchanged. In this paper, we present the first study of Multi-Head Latent Attention (MLA) in video diffusion. VideoMLA replaces per-head keys and values with a shared low-rank content latent and a shared decoupled 3D-RoPE positional key, reducing per-token KV memory by 92.7% at every cached layer. We further investigate why MLA succeeds in video diffusion even though the spectral assumption often used to motivate it in language models does not hold: pretrained video attention is not low-rank, with 99\%-energy effective rank far above any practical latent dimension. VideoMLA retains quality at compression ratios where direct spectral approximation would predict large reconstruction error. We show that the MLA bottleneck, rather than the pretrained spectrum, determines the effective rank: both spectral and random initialization occupy nearly the full rank budget from initialization, and Stage-1 training preserves this budget while adapting within it. On VBench, VideoMLA matches short-horizon streaming video diffusion baselines, achieves the best overall score at long horizons among evaluated methods, and improves throughput by $1.23\times$ on a single B200.


Video Models Can Reason with Verifiable Rewards

Tinghui Zhu ⋅ Sheng Zhang ⋅ James Yipeng Huang ⋅ Selena Song ⋅ Xiaofei Wen ⋅ Yuankai Li ⋅ Hoifung Poon ⋅ Muhao Chen

Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable reasoning. This limitation is especially pronounced in tasks where generated videos must satisfy explicit spatial, temporal, or logical constraints. Inspired by the role of reinforcement learning with verifiable rewards (RLVR) in reasoning-oriented language models, we introduce VideoRLVR, a practical recipe for optimizing video diffusion models with rule-based feedback. VideoRLVR formulates video reasoning as the generation of verifiable visual trajectories and consists of an SDE-GRPO optimization backbone, dense decomposed rewards, and an Early-Step Focus strategy for efficient training. The Early-Step Focus strategy restricts policy optimization to the early denoising phase, reducing training latency by about 40\% while preserving performance. We evaluate VideoRLVR on Maze, FlowFree, and Sokoban, three procedurally generated domains with objective success criteria. Across these tasks, VideoRLVR consistently improves over supervised fine-tuning baselines, with dense decomposed rewards proving especially important in low-success-rate settings. Our RL-optimized model also outperforms the evaluated proprietary and open-source video generation models on these verifiable reasoning benchmarks and out-of-domain benchmarks. These results suggest that verifiable RL can move video models beyond perceptual imitation toward more reliable rule-consistent visual reasoning.

Vision-Language-Action (VLA) models have shown strong promise for general-purpose robotic manipulation, but their real-world evaluation remains limited by a lack of accessible, reproducible, and consistent benchmarks. Simulation benchmarks fail to capture real-world complexity, while existing real-world benchmarks often require expensive hardware, centralized evaluation, or are limited in task diversity. We introduce VLA-REPLICA, a low-cost, easily reproducible real-world benchmark for evaluating VLA models. Built from off-the-shelf components, our system can be quickly assembled and replicated across laboratories, providing a consistent environment for policy evaluation anywhere in the world. VLA-REPLICA includes a diverse suite of manipulation tasks and a small-scale demonstration dataset for target-domain adaptation, with real-world evaluation protocols for both in-distribution and out-of-distribution settings. Experiments with imitation learning and state-of-the-art VLA models reveal model strengths and limitations, while consistent results across independently constructed setups demonstrate the reproducibility of our benchmark.


VTV-FM: Flow Matching through Variational Terminal-Velocity Closure

Haoyang Jiang ⋅ Yuheng Li ⋅ Di Yang ⋅ Yanhai Xiong ⋅ Haipeng Chen ⋅ Yi He

Flow Matching (FM) learns generative transport by fitting continuous-time motion from a simple source distribution to the data distribution. Most existing methods use first-order bridges: once a source and a target sample are paired, the path is a straight motion with constant velocity. FM with optimal transport (OT) improves the pairing, but the bridge itself remains linear, limiting its ability to model curved motion, acceleration, and changing directions. A natural remedy is to use second-order phase-space dynamics; however, constructing such bridges requires a target-side terminal velocity, which static datasets typically do not provide. We propose Variational Terminal-Velocity Flow Matching (VTV-FM), a second-order FM framework that derives the missing velocity by minimizing acceleration energy, yielding a closed-form closure for static data. The same minimum-acceleration variational construction also defines the OT pairing cost and the acceleration targets used for training. Experiments on low-dimensional datasets, PDE-governed physical fields, and CIFAR-10 show that VTV-FM improves transport geometry and generation quality over first-order and existing high-order FM baselines.


WavFlow: Flowing Through Waveforms for Audio Generation

Feiyan Zhou ⋅ Luyuan Wang ⋅ Shoufa Chen ⋅ Zhe Wang ⋅ Zhiheng Liu ⋅ Yuren Cong ⋅ Xiaohui Zhang ⋅ Fanny Yang ⋅ Belinda Zeng

Modern audio generation predominantly relies on latent-space compression, introducing additional complexity and potential information loss. In this work, we challenge this paradigm with WavFlow, a framework that generates high-fidelity audio directly in raw waveform space without intermediate representations. To overcome the inherent difficulties of modeling high-dimensional and low-energy signals, we reshape audio into 2D token grids through waveform patchify and introduce amplitude lifting to align signal scales, enabling stable optimization via direct x-prediction in flow matching. To capture complex semantic alignment and temporal synchronization, we leverage an automated data pipeline to curate 5M high-quality video-text-audio triplets, allowing the model to learn fine-grained acoustic patterns from scratch. Experimental results show that WavFlow achieves state-of-the-art results on the video-to-audio benchmark VGGSound (FDPaSST 55.82, ISPANNs 17.40, DeSync 0.44) and the text-to-audio benchmark AudioCaps (FD_PANNs 10.63), outperforming established latent-based methods. Our work demonstrates that such intermediate compression is not a prerequisite for high-quality synthesis, offering a simpler and more scalable alternative for multimodal audio generation.


What Does an Observability Forecasting Foundation Model Know?

Dhyey Mavani ⋅ Tairan Ji ⋅ Rian Atri

Time-series foundation models (TSFMs) are increasingly deployed as zero-shot forecasters in observability platforms, yet the operational concepts encoded within their internal representations remain largely un-audited. To address this gap, we introduce $\texttt{toto-interp}$, a control-first interpretability protocol that pairs linear probes with four rigorous baselines: raw-feature, shuffled-label, randomized-backbone, and supervised raw-window Fourier Neural Operator (FNO) controls. Applying this harness to TOTO on BOOM, we demonstrate that three structural telemetry concepts, such as cadence bucket ("frequency"), metric type, and domain, are linearly decodable, layer-localized, and specifically tied to pretraining. A budget-matched replication on MOMENT-base successfully recovers these same axes, confirming these representations are not isolated artifacts of the TOTO architecture. Furthermore, our matched-null and on-manifold interventions reveal a critical two-way dissociation between linear decodability and forecast-sensitivity. For instance, concepts like future burstiness are highly decodable yet completely inert under intervention, whereas other weakly decodable directions can severely disrupt forecasts off-manifold. By releasing the audit harness and detailing negative control cases (such as cardinality and shift risk), we demonstrate that linear decodability alone is insufficient for reliable post-training steering or deployment auditing in TSFMs.

Reusing a held-out benchmark adaptively should, in principle, invite overfitting. Yet benchmark-driven machine learning (ML) has produced surprisingly little overfitting in practice. An attractive hypothesis is that successful ML strategies are highly compressible. We study this in the setting of LLM-driven research agents, where the hypothesis becomes directly testable via two complementary information bottlenecks. In \emph{output compression}, an exploration agent adaptively searches for high-performance models using a validation set, and we test whether a fresh ``reproducer agent'' can reproduce its performance given only an extremely short prompt and the training data. In \emph{input compression}, the explorer receives only one-bit feedback indicating whether each submitted model improves on the running best. Across 8 datasets spanning tabular classification, vision, language modeling, diffusion modeling, and reward modeling, we find that these bottlenecks have little effect on performance: short prompts and compressible feedback are sufficient to reproduce and find high-performance models. The hypothesis is falsifiable: when we deliberately induce validation-set overfitting, the results fail to reproduce with short prompts. Taken together, our results support a description-length explanation for the lack of overfitting in benchmark-driven ML: successful strategies occupy a low-complexity region of strategy space.


What Sound Tells You About the Room

Yiduo Hao ⋅ Yiwei Tang ⋅ Jiayang Li ⋅ Zitong Lan ⋅ Haowen Lai ⋅ Romit Roy Choudhury ⋅ Mingmin Zhao

A room impulse response (RIR) encodes how sound reflects, scatters, and decays inside a room, and therefore what the room looks like and what it is made of. We formalize this observation as Acoustic Inverse Rendering (AIR): recovering dense panoramic geometry and per-band acoustic reflectance of a scene directly from its RIR. The task is challenging as the RIR collapses all propagation paths into a single temporal signal, and the resulting inverse mapping is severely ambiguous. AIR closes this gap by beamforming the RIR onto the equirectangular grid and conditioning a diffusion model on the resulting directional features, yielding a distribution over plausible scenes rather than a single point estimate. The same spherical beamforming applies to FOA, binaural, or mono microphone setups. On both simulated and real-world datasets, AIR produces coherent room geometry and per-band acoustic reflectance maps from acoustic input alone, significantly outperforming prior baselines on this inverse problem.

Semivalues are widely used to assign credit and guide data curation decisions, yet their outputs depend on a utility function that is not uniquely determined by the task. We study partition-based decisions, which directly model data curation tasks such as selecting high-quality subsets or flagging noisy examples, under two structurally unavoidable sources of utility underspecification: monotone transformations of performance scores and unconstrained small-sample behavior. We formalize outcome robustness in this setting as invariance of the induced partition across admissible utility respecifications. We establish that partition robustness is necessary for guaranteed success of semivalue-based data selection. To assess this condition in practice, we provide algorithms that certify partition robustness or produce a concrete witness utility demonstrating instability, requiring no utility evaluations beyond those already computed for the semivalues themselves, with formal correctness guarantees under both exact and approximate computation. We show that utilities admitting a coalitionally dominant top-$k$ subset, where high-quality points contribute more than low-quality points across every possible coalition, guarantee robust partition decisions, subsuming the modular utility case previously identified as sufficient. Together, these results provide a practical framework for determining when semivalue-based selection is well-defined under utility ambiguity.

Low-Rank Adaptation (LoRA) constrains weight updates to a low-rank subspace, making the choice of that subspace central to adaptation. We show that the value of LoRA initialization direction is regime-dependent. It helps when the representation dimension is large relative to the labeled set size and the estimated activation subspace is aligned with task structure, and is otherwise inconsequential. We formalize this view with a spiked-covariance Davis--Kahan argument and a first-step LoRA analysis showing that the initial learning signal flows through projected target-token features. From this analysis we derive an alignment coefficient $\rho$, computable from a single forward pass, that closely tracks the empirical gain from direction-aware initialization across tasks. Motivated by this view, we propose PivotLoRA, an initialization that estimates LoRA adapters from target-token PCA and applies a single output-preserving orthogonal pivot after early training drift. Across $30$ BERT-base few-shot settings and $13$ compared methods, PivotLoRA achieves the best average rank and improves over vanilla LoRA on nearly every setting, with the largest gains concentrated on high-alignment tasks. Decoder-scale classification on LLaMA-2-7B and Llama-3.1-8B reproduces the low-shot pattern, while full-data and generation controls confirm the predicted regime specificity.


When Everyone Can Submit: Designing Contests with Transparent Pre-Selection

Hanbing Liu ⋅ Ningyuan Li ⋅ Weian Li ⋅ Qi Qi ⋅ Changyuan Yu

Shortlisting is a common and effective method for pre-selecting participants in competitive settings. To ensure fairness, a cut-off score is typically announced, allowing only contestants who exceed it to enter the contest, while others are eliminated. In this paper, we study rank-order contests with shortlisting and cut-off score disclosure. We fully characterize the equilibrium behavior of shortlisted contestants for any given prize structure and shortlist size. We examine two objective functions: the highest individual performance and total performance. Under linear performance costs, for both objectives, the optimal contest is in a winner-take-all format. For the highest individual performance, the optimal shortlist size is exactly two contestants, but, in contrast, for total performance, the shortlist size does not affect the outcome, i.e., any size yields the same total performance. Furthermore, we compare the highest individual performance achieved with and without shortlisting, and show that the former is 4/3 of the latter.


When Reasoning Meets Its Laws

Junyu Zhang ⋅ Yifan Sun ⋅ Tianang Leng ⋅ Jingyan Shen ⋅ Liu Ziyin ⋅ Paul Liang ⋅ Huan Zhang

Despite the superior performance of Large Reasoning Models (LRMs), their reasoning behaviors are often counterintuitive, e.g., excessive reasoning on simple questions or insufficient reasoning on complex ones, leading to suboptimal reasoning capabilities. This paper presents the Laws of Reasoning (LoRe), a unified framework that formalizes intrinsic reasoning patterns in ideal LRMs. We first propose compute law with the hypothesis that the reasoning compute should scale linearly with question complexity. Beyond compute, we extend LoRe with a supplementary accuracy law, positing exponential decay of accuracy with increasing complexity. Since the complexity is difficult to quantify in practice, we approximate these hypotheses via two tractable properties, monotonicity and compositionality. We therefore introduce LoRe-Bench, a benchmark that systematically measures these properties for large reasoning models. Evaluation shows that most reasoning models exhibit reasonable monotonicity but lack compositionality. In response, we develop an effective finetuning approach that enforces compute-law compositionality. Extensive empirical studies demonstrate that better compliance with compute laws yields consistently improved reasoning performance on multiple benchmarks, and uncovers synergistic gains across properties and laws.


Which Tokens to Merge? Diffusion Dynamics for Efficient Image Generation

SeungJu Cha ⋅ Ye-Chan Kim ⋅ HyunGee Kim ⋅ Sungho Koh ⋅ Dong-Jin Kim

Token merging has emerged as an effective training-free strategy for accelerating diffusion models by reducing redundant token computation while minimizing generation degradation. A key challenge is determining which tokens can be safely merged and which should be preserved for subsequent computation. Existing method addresses this by preserving prompt-aligned foreground regions, thereby protecting semantically relevant structures. However, this text-relevance criterion overlooks background objects, local details that are not explicitly mentioned in the prompt, leading to degraded structural coherence. In this work, we propose Dynamics-aware Token Merging (DaTMe), a training-free framework that grounds token merging in the underlying diffusion dynamics. Our key insight is that token importance should reflect whether a token remains unresolved during denoising and whether it carries structural information. To this end, DaTMe constructs a dynamics-aware importance map from the Tweedie denoised estimate. The map combines a temporal signal, measuring inter-step changes in the predicted clean image, with a structural signal, capturing local high-frequency content such as boundaries and fine details. Using this map, DaTMe adaptively assigns tokens into merging roles under a given compression ratio: important tokens are preserved as anchors, while tokens that are both temporally stabilized and structurally homogeneous are selected as safe merge candidates. Experiments on PixArt-$\alpha$ and FLUX demonstrate that DaTMe maintains visual fidelity across the entire image, highlighting diffusion dynamics as an effective token-level merging decision in efficient generation.


You Can’t Have It Both Ways: Concept Entanglement Limits Diffusion Model Unlearning

Yian Wang ⋅ Ali Ebrahimpour-Boroojeny ⋅ Hari Sundaram ⋅ Varun Chandrasekaran

Concept unlearning in text-to-image diffusion models aims to suppress a target concept, e.g., \texttt{horse}, while preserving the model's ability to generate semantically related but distinct content, like \texttt{donkey}. Yet existing methods either leak under indirect prompts or visibly degrade remaining concepts. {\em We show that these failure modes arise naturally from overlapping concept representations.} Our main contribution is demonstrating that the observed trade-offs stem from the geometry of concept representations, rather than from weaknesses in any particular algorithm. We support this claim through both theoretical analysis and empirical observations. By formalizing concepts as activation-space regions, we show that overlap between a target and other concepts lower-bounds the unavoidable degradation on those other concepts when the target is erased. Our analysis suggests that strong erasure and preservation become fundamentally coupled when concepts occupy overlapping activation regions, with the trade-off scaling linearly in the degree of overlap. We verify this trade-off on seven unlearning methods spanning fine-tuning, adversarially-robust fine-tuning, and inference-time interventions. Averaged across target concepts, STEREO almost completely suppresses the target concept under indirect prompts but cuts the model's ability to generate semantically related concepts by more than 75\%. On the other hand, sparse inference-time methods (SAeUron, SEOT) better preserve utility but leave substantial target leakage. The trade-off is steep for concepts whose internal representations heavily overlap with their evaluated semantic neighborhoods, (e.g., \texttt{horse} with \texttt{cat}, \texttt{dog}, and \texttt{bear}) and milder for \texttt{castle}, which is comparatively less entangled within our concept pool and probe layer in activation space. Across targets, damage to related concepts scales monotonically with our activation-overlap measure. These results suggest that perfect unlearning is the wrong target for entangled concepts. The field should evaluate on the Pareto frontier our theorem establishes; current benchmarks, which decouple unlearning accuracy from utility preservation, hide this trade-off and need to be revised.