Workshops
2nd Workshop on Principles of Generative Modeling (PriGM)
Francesco Cagnetta ⋅ Elisabetta Cornacchia ⋅ Valentin De Bortoli ⋅ Soon Hoe Lim ⋅ Bingbin Liu ⋅ Bruno Loureiro ⋅ Gabriel Peyré
The workshop aims to bring together researchers developing principled, scientific approaches to understanding generative models, with the goal of identifying and emphasizing the central theoretical challenges posed by modern generative AI.
Show more
RoboPAD: Post-Training Adaptation of Robot Foundation Models
Shijie Li ⋅ Shizhe Chen ⋅ Ziwei Wang ⋅ Jiafei Duan ⋅ Sihao Lin ⋅ Changliu Liu ⋅ Gerhard Neumann
RoboPAD focuses on post-training adaptation of robot foundation models: the stage after large-scale pretraining where robots must be corrected, specialized, personalized, evaluated, and safely improved under real-world constraints. As vision-language-action models, diffusion and flow-matching policies, world models, and large-scale simulators become increasingly available, the next bottleneck is no longer only how to pretrain general robot policies, but how to adapt them to new environments, users, objects, embodiments, task constraints, and unexpected failures without restarting the full pretraining pipeline. The workshop will bring together researchers from robot learning, reinforcement learning, imitation learning, foundation models, world models, human-robot interaction, embodied AI, and safety. It will cover feedback- and intervention-driven adaptation, policy optimization after pretraining, world-model- and simulation-driven refinement, reasoning and memory for robot adaptation, cross-embodiment transfer, and evaluation, safety, and robustness after adaptation. A concrete outcome will be a public community roadmap summarizing shared terminology, evaluation protocols, benchmark gaps, and reporting practices.
Show more
2nd Workshop on Advances in Representation Learning for Earth Observation (REO-2)
Nikolaos Ioannis Bountos ⋅ Benedikt Blumenstiel ⋅ Ruben Cartuyvels ⋅ Nico Lang ⋅ Loic Landrieu ⋅ Xiaoxiang Zhu ⋅ Gustau Camps-Valls ⋅ Ioannis Papoutsis
The second edition of the Representation Learning for Earth Observation (REO) workshop will bring together researchers and practitioners from machine learning, computer vision, and Earth sciences to advance the development of robust, interpretable, and scalable models for monitoring and predicting Earth. With the growing availability of large-scale, multimodal Earth Observation data and the rise of powerful foundation models, new opportunities emerge for integrating data-driven and physics-informed approaches across sensing modalities and application domains. REO will provide a forum for presenting novel technical methods, scientific applications, and system-level innovations, fostering cross-disciplinary exchange and collaborations between academia, industry, and policy stakeholders.
Show more
AI for Chip Design
Dario Garcia-Gasulla ⋅ Gokcen Kestor ⋅ Emanuele Parisi ⋅ Zhiyao Xie ⋅ Luca Benini ⋅ Leigh Anne Clevenger ⋅ Cong Hao
The design of modern semiconductor chips is one of the most complex intellectual and engineering challenges in computer science, and one of the last frontiers where AI has yet to deliver transformative impact at scale. Recent years have seen a surge of machine learning contributions across the chip design stack, from graph neural networks for routing and timing prediction, to large language models for hardware description languages, to agentic systems navigating full EDA workflows. Yet the communities driving these advances remain fragmented: ML researchers rarely attend hardware venues, and EDA practitioners have limited exposure to frontier ML methods. This workshop brings both communities together at NeurIPS to share results, identify open problems, and build the interdisciplinary research agenda that AI for chip design urgently needs. We invite contributions spanning all ML paradigms and all stages of the chip design flow, from specification to tapeout.
Show more
Grounded User Simulation for Model Evaluation and Training:\\ Diversity, Fidelity, and Validity
Yevgeniy Meyer ⋅ Krisztian Balog ⋅ Flora Salim ⋅ Victor Barres ⋅ Preethi Seshadri
User simulation is becoming core infrastructure for evaluating and training interactive AI systems, from conversational search and recommender systems to agentic tool-use benchmarks and post-training pipelines. Yet the field lacks shared standards for when simulated users are reliable proxies for real people, especially across languages, dialects, demographic groups, tasks, and deployment contexts. This workshop will bring together researchers in agent evaluation, information retrieval, dialogue systems, HCI/social simulation, reinforcement learning, recommender systems, synthetic data, multilingual NLP, and fairness to study user simulation through three lenses: diversity of represented users and scenarios, fidelity to individual and population-level behavior, and validity of conclusions drawn from simulation. The workshop will combine invited talks, contributed papers, posters, a debate on whether simulated users can be trusted, and a cross-community breakout. As Year-1 outputs, we will seed shared infrastructure around TREC-style conversational search and τ-family agentic tool-use settings, including a census-conditioned baseline simulator, reporting checklist, validation/failure-mode taxonomy, and benchmark catalog.
Show more
ELLIS Workshop on the Foundations of LLM Post-Training in Changing Environments
Francesco Quinzan ⋅ Fanghui Liu ⋅ Gabriel Peyré ⋅ Patrick Rebeschini ⋅ Chengchun Shi
Large language models (LLMs) are routinely adapted to downstream applications through post-training methods, such as instruction tuning and domain adaption. Yet in real-world deployment, downstream tasks rarely remain fixed: objectives shift, data distributions drift, feedback signals evolve, and evaluation standards change over time. Post-training therefore becomes a process of repeated adaptation in non-stationary environments. Despite its central role in modern foundation models, the theoretical foundations of this adaptive post-training paradigm remain limited. Current practices are largely heuristic, with incomplete understanding of statistical identifiability, optimization dynamics, robustness to misspecification, and trade-offs between adaptation and capability preservation. These gaps are particularly consequential in safety-critical settings, where unintended regressions or feedback loops may arise under evolving conditions. This workshop aims to develop principled foundations for LLM post-training under task evolution. We bring together researchers from machine learning theory, reinforcement learning, and AI safety to develop principled foundations for this adaptive post-training paradigm.
Show more
Machine Learning for Spatially Resolved High-dimensional Biology
Lucia Testa ⋅ Stefan Bonn ⋅ Robin Khatri ⋅ Behnam Yousefi
Spatially resolved biology is rapidly transforming the study of tissues by measuring molecular activity while preserving spatial organization. Technologies such as spatial transcriptomics, spatial proteomics, and multiplex imaging now generate high-dimensional, multimodal data at cellular, tissue, and atlas scales. These data pose new machine learning challenges: representing irregular tissue geometry, integrating molecular and imaging modalities, modeling cell-cell communication, handling noise and sparsity, building interpretable and uncertainty-aware models, and designing reliable benchmarks. This workshop will bring together machine learning researchers, computational biologists, and experimental scientists to define spatially resolved high-dimensional biology as a core methodological problem for ML. It will focus on geometric and graph learning, multimodal representation learning, generative modeling, foundation models, biological inductive biases, and evaluation standards for spatial omics and tissue-scale data.
Show more
AIDaR: AI Data Readiness for Scientific Discovery
Zoe Piran ⋅ Arindam Sett ⋅ Edaeni Hamid ⋅ Vladimir Ermakov ⋅ Kenny Workman ⋅ A. Sina Booeshaghi ⋅ Michal Rosen-Zvi
Scientific AI is moving from curated prediction tasks toward systems that work directly with experimental data, scientific knowledge, and analysis workflows. Foundation models, retrieval systems, and agents are increasingly expected to query evidence, run analyses, interpret outputs, and support follow-up decisions. For these systems to be reliable, data must remain connected to the samples, assays, protocols, measurements, workflow outputs, and contextual assumptions needed to interpret results beyond the lab or notebook where they were produced. The workshop centers on one question: which measurable properties of scientific data ecosystems predict downstream AI behavior? We use AI Data Readiness (AIDaR) to describe this measurable relationship between data ecosystems and model or agent performance. AIDaR focuses on two linked problems. First, how should scientific data be organized so AI systems can use it: from raw measurements and analysis-ready files to relational or graph-structured knowledge, multimodal biomedical measurements, workflow records, and large datasets used for model training and experimental feedback? Second, how should evaluations determine where frontier models are reliable: across biological reasoning, practical data analysis, retrieval over structured scientific knowledge, assay-specific workflows, and deployed scientific AI settings? The workshop will bring together researchers, industry scientists, and engineers building scientific data systems, foundation models, agents, biomedical evaluations, and industrial assay platforms. The program includes invited talks, contributed papers, demos, panels, and breakout groups focused on data infrastructure, practical evaluations, and lessons from scientific AI deployments.
Show more
Trustworthy AI for Good (AI4GOOD) Workshop
Terry J Zhang ⋅ Arian Khorasani ⋅ Joan Nwatu ⋅ Rada Mihalcea ⋅ Milind Tambe ⋅ David Lie ⋅ Christian Schroeder de Witt ⋅ Zhijing Jin ⋅ Swapneel Mehta ⋅ Klaudia Krawiecka
As agentic AI systems move from benchmarks into public-facing workflows, trustworthy AI must go beyond model-level security and isolated safety checks. The central challenge is whether advanced AI systems can produce measurable public benefit across populations while mitigating systemic risks such as power concentration, unequal access, accountability gaps, and dependence on a small number of AI providers and institutions. This workshop brings together the science of trustworthy AI, AI safety, AI for social good, and AI policy/governance communities to connect technical progress in evaluation, robustness, alignment, monitoring, interpretability, and human oversight with responsible deployment in public-interest settings. We focus on applications aligned with the United Nations Sustainable Development Goals, including health, education, climate action, humanitarian response, accessibility, reduced inequalities, and inclusive public services. Through invited talks, contributed presentations, posters, and a cross-sector panel, the workshop aims to build a shared research agenda for AI systems that are safer, more accountable, and more beneficial at societal scale.
Show more
Reinforcement Learning for Experimental Sciences: Bridging the Simulation-to-Reality Gap
Odalric-Ambrym Maillard ⋅ Ronald Ortner ⋅ Audrey Durand ⋅ Mohammad Sadegh Talebi ⋅ Samba Diaw ⋅ Tristan Fauvel
Reinforcement learning (RL) provides a principled framework for sequential decision-making and has achieved remarkable success in simulated environments. However, its application to experimental sciences remains challenging due to limited experimental budgets, delayed feedback, partial observability, safety constraints, and the persistent simulation-to-reality gap. At the same time, advances in autonomous experimentation platforms, robotics, environmental sensing, and digital twins are creating new opportunities for RL-driven scientific discovery. This workshop will bring together researchers from sequential learning, autonomous systems, and experimental sciences to discuss adaptive experimentation and autonomous discovery in laboratory and field settings. Relevant application domains include chemistry, biology, materials science, medicine, sociology, agriculture, ecology, and environmental monitoring. We are particularly interested in methods supporting robust deployment in real-world experimental systems, e.g. model-based RL, uncertainty-aware decision-making, safe exploration, sim-to-real transfer, human-in-the-loop learning, and hybrid simulation–experiment pipelines. The workshop will provide a forum for identifying shared challenges, presenting emerging benchmarks and infrastructures, and fostering collaborations between RL researchers and experimental scientists. By connecting communities that rarely interact despite common methodological concerns, it aims to accelerate the development of reliable RL methods for scientific experimentation and autonomous discovery.
Show more
Foundations of Agentic Systems Theory (FAST)
Erik Miehling ⋅ Madeline G. Reinecke ⋅ Hamza Mostafa ⋅ Jordan McAfoose ⋅ Irina Rish ⋅ Djallel Bouneffouf
The Foundations of Agentic Systems Theory (FAST) workshop provides a venue for the investigation of system-level behaviors of agentic AI. FAST aims to draw from various fields (including complex systems, developmental biology, organizational sociology, and cognitive science) with the goal of understanding which specific mechanisms of emergent behavior from other systems carry over to systems of LLM-based agents, which properties of the underlying agents (and their LLMs) facilitate or block these behaviors, and to what extent system-wide outcomes can be controlled or induced. Answering these questions is the prerequisite for safe and well-understood deployment of agentic AI.
Show more
AI and the Self: Human Identity, Authenticity, and Agency in the Age of AI
Mukund Choudhary ⋅ Abdoul Jalil DJIBEROU MAHAMADOU ⋅ Sandrine R Schiller ⋅ Haneesha Pinnamaraju ⋅ Jacki O'Neill ⋅ Akhil Arora ⋅ Monojit Choudhury
Large language models, recommender systems, multimodal agents, digital avatars, and brain-computer interfaces are increasingly intimate, personalized counterparts rather than external tools. They can act as interlocutors, cognitive extensions, mirrors, collaborators, and social actors, reshaping self-perception, agency, emotional life, competence, authorship, and cultural worldview. Despite extensive work on alignment, safety, fairness, and human-centered AI, we still lack shared frameworks and evaluation methods for understanding AI’s impact on selfhood (the “algorithmic self”) before such systems become ubiquitous. This one-day NeurIPS workshop convenes researchers and practitioners across AI, HCI, cognitive and social sciences, philosophy, ethics, education, and health to operationalize an emerging research agenda on AI and the Self. We solicit non-archival papers and demos that (1) measure and evaluate AI’s effects on self-concept, autonomy, dependence, and skill change; (2) explore technical design choices: memory, personalization, uncertainty, abstention, refusal, etc. that support reflection while avoiding manipulation, unhealthy attachment, or culturally narrow models of the self; (3) study how language, culture, and values shape (and are shaped by) AI systems; (4) analyze transformations of learning, creativity, expertise, and professional identity in AI-mediated work; and (5) clarify philosophical, ethical, legal, and spiritual foundations around responsibility, authorship, and personhood. Through keynotes, panels, and poster sessions, the workshop aims to produce concrete research problems, evaluation strategies, and design principles for AI systems that preserve and strengthen human agency.
Show more
Sim2Science: ML with Imperfect Scientific Models
Georgia Channing ⋅ Noémi Éltető ⋅ Richard Gao ⋅ Daniel Gedon ⋅ Magdalena Lederbauer
AI4Science has matured into an established field, with ML now embedded throughout the simulator-based workflows of the natural sciences. But all models are wrong: every simulator approximates reality. An ML method coupled to a simulator is only as reliable as that simulator, potentially steering us into wrong scientific conclusions and real-world decisions even if the ML component was perfect. This problem of simulator misspecification is central to scientific ML, and adversely impacts fields as different as chemistry, fusion, and neuroscience. Yet the communities that face it rarely meet, and a solution found in one field seldom reaches another. Sim2Science is organized around this shared problem rather than a single domain: it convenes researchers across the sciences and ML to detect, quantify, and mitigate simulator misspecification, and to turn advances in one field into methods that transfer to others.
Show more
AI for Peace
Noa Garcia ⋅ Leonardo Impett ⋅ Yannis Kalantidis ⋅ Sonia Fereidooni ⋅ Pier Luigi Dovesi ⋅ Alexandra Volokhova
In this workshop, we aim to address the critically under-discussed issue of AI's dual-use nature, focusing on how machine learning technologies are being adapted for military purposes, potentially without the researchers' knowledge or consent. While attending to the heightened risks associated with particular areas and systems of research, we will also be collectively thinking through what it looks like to engage productively in research and development activities that considers ethics and international law at its core.
Show more
Representations for the Physical Sciences
Florence d'Alché-Buc ⋅ Pierre Gentine ⋅ Kara Lamb ⋅ Mathias Niepert ⋅ Pietro Novelli ⋅ Massimiliano Pontil
Representation learning has become a central ingredient of modern AI, yet its scientific use raises distinctive challenges that are not captured by standard vision-and-language settings. Scientific data are heterogeneous, structured, and often generated by dynamical systems; useful embeddings must therefore respect geometry, symmetries, conservation laws, and causal structure, while remaining transferable across regimes that are expensive to verify experimentally or computationally. At the same time, the sciences offer unique opportunities for learning representations, including large-scale unlabeled datasets, physically grounded simulators, and closed-loop experimental pipelines. This workshop will bring together researchers from machine learning and the physical and life sciences to focus on four tightly scoped bottlenecks in representation learning for physical systems: self-supervision under physical constraints, transfer and its limits, sampling and closed-loop data generation, and tokenization of continuous scientific modalities. Through invited talks, contributed presentations, posters, and discussion sessions, the workshop aims to clarify the methodological foundations of scientifically useful representations and to foster a shared research agenda across communities.
Show more
Personalized, Aligned, Long-Term Memory for AI Systems (PALM) Workshop
Mario Fritz ⋅ Seong Joon Oh ⋅ Sahar Abdelnabi ⋅ Junxiao Shen ⋅ Hugo Lopes ⋅ Haritz Puerto ⋅ Ivaxi Sheth ⋅ Seokwon Jung
Long-term memory is becoming a core capability for AI agents and AI assistants, enabling them to remember user preferences, past interactions, tool-use traces, multimodal context, and evolving task histories across sessions. Yet current research remains fragmented across agent memory, personalization, benchmarking, multimodal learning, cognitive models, privacy, and safety. The PALM workshop will bring these communities together to study persistent memory as an infrastructure layer for personalized, aligned, and long-term AI agents. The workshop will focus on memory architectures, evaluation protocols, user control, and emerging safety risks, including privacy leakage, memory poisoning, over-personalization, sycophancy, sleeper memories, and long-term behavioral manipulation. By combining invited talks, contributed papers, posters, and a panel, PALM aims to define shared research questions and safeguards for robust, user-governed memory systems.
Show more
SLM-Agents: 1st Workshop on SLMs for Agentic Systems
Habib Hajimolahoseini ⋅ Mehdi Rezagholizadeh ⋅ Vahid Partovi Nia ⋅ Shahrzad Kianidehkordi ⋅ Mouloud Belbahri ⋅ Pavlo Molchanov ⋅ Hamed Jafarzadeh Asl ⋅ Masoud Asgharian
This workshop is dedicated to small language models (SLMs) as the foundation of agentic AI systems. Although large language models (LLMs) have demonstrated remarkable capabilities, their dependence on cloud infrastructure creates fundamental barriers to deployment in agentic pipelines latency, privacy, connectivity, and substantial computational cost. SLMs offer a compelling alternative: recent studies argue that SLMs, not LLMs, might be the right option for the repetitive, narrowly scoped sub-tasks that dominate real agentic workloads [15]. SLMs make it possible for autonomous AI agents to plan, reason, and act directly on resource-constrained devices such as smartphones, IoT systems, robotics platforms, and embedded systems. The workshop sits at the intersection of three rapidly evolving fields: (1) efficient language model architectures and compression techniques, (2) agentic AI systems capable of autonomous reasoning and tool use, and (3) edge computing and on-device deployment.
Show more
LIGHT: Deployable Small Foundation Models
Roberta Calegari ⋅ Dennis Hoppe ⋅ Joachim Koehler ⋅ Michela Milano
Foundation models have achieved remarkable performance across language, vision, and multimodal tasks, but their deployment remains challenging due to their size, computational requirements, and limited controllability. This workshop explores the emerging transition from large foundation models to compact, trustworthy, and deployable AI systems. It focuses on knowledge distillation, compression, quantization, and small foundation models. By bringing together researchers from machine learning, AI systems, trustworthy AI, and industrial deployment, the workshop aims to foster a common research agenda for efficient and trustworthy AI systems capable of operating under real-world constraints.
Show more
AXIOM: Foundations of Efficient Deep Learning
Olga Saukh ⋅ Linara Adilova ⋅ Bernhard Geiger ⋅ Yedi Zhang ⋅ Flavio Martinelli
Deep learning theory is beginning to uncover quantitative regularities that explain and predict the behavior of large-scale learning systems, including scaling laws, compute-optimal training prescriptions, and predictable optimization dynamics. At the same time, the growing computational, memory, and energy demands of modern AI have made efficiency a central challenge for the field. Yet a substantial gap remains between theoretical understanding and practical efficiency: many efficient AI methods are developed empirically, while existing theories rarely provide actionable principles for designing learning systems under resource constraints. **AXIOM: Foundations of Efficient Deep Learning** brings together researchers from deep learning theory, optimization, machine learning systems, and efficient AI to investigate how theoretical insights can guide the design of efficient learning systems and which aspects of efficient AI can be predicted rather than discovered through costly experimentation. Topics include scaling laws and capability prediction, resource-constrained optimization and generalization, sparsity and compression, modularity and adaptive computation, efficient foundation models, hardware-aware learning, and theoretical limits of efficient AI. The workshop combines invited vision talks, contributed papers, posters, and a community-driven Grand Challenges initiative focused on identifying key open questions and future directions for the foundations of efficient AI, with outcomes synthesized into a community position paper outlining a research agenda for theory-guided efficient AI.
Show more
Neural Network Artifacts as a New Data Modality
Giorgos Bouritsas ⋅ Amil Dravid ⋅ Fabrizio Frasca ⋅ Aspen Hopkins ⋅ Koyena Pal ⋅ Bo Zhao
Machine learning has been transformed by learning from large populations of data, yet it has rarely turned that same population-level lens on its own products: trained models and the artifacts they generate. Today's repositories hold millions of models, and within their weights, gradients, internal representations, and optimization trajectories lies a vast but largely untapped reservoir of knowledge: we still lack principled methods to compare models, search among them, predict or modify their behaviour, or understand how they relate. This workshop advances an agenda to close that gap by treating neural artifacts as a data modality in their own right, amenable to modelling and learning. Its goals are twofold: to encourage tailored methodologies that analyse, interpret, modify, control, and synthesise these artifacts in an automated manner; and to connect, under a shared data-centric perspective, the communities that have approached them in isolation, spanning model merging, meta-learning, mechanistic interpretability, neural architecture search, and neural-field processing. Building on the inaugural ICLR 2025 edition, this second edition broadens the scope from weights to the full range of neural artifacts, adds neural lineages and AI supply chains, and places strong emphasis on standardised datasets, benchmarks, and tasks. Through invited talks, contributed papers and discussions, the workshop aims to consolidate these scattered efforts into a coherent field.
Show more
Bridging Optimal Transport, Learning and Structured Data: Toward Geometric Distributional Learning
Clément Bonet ⋅ Julie Delon ⋅ Nina Miolane ⋅ Youssef Mroueh ⋅ Kimia Nadjahi ⋅ Justin Solomon
Modern machine learning increasingly relies on both geometric and distributional representations of complex data. While geometric deep learning provides tools to encode symmetries, invariances and relational structure in non-Euclidean domains, distributional methods, including optimal transport, offer principled ways to compare align, and transform data distributions. These perspectives are deeply connected but often developed separately. This workshop will focus on the emerging area of **Geometric Distributional Deep Learning**: learning systems that jointly model the geometry and distributional nature of data, features or representations. The goal is to bring together researchers from geometric deep learning, computational optimal transport, generative modeling, and representation learning to discuss how geometric and distributional principles can inform new architectures, algorithms and applications for structured data.
Show more
Foundations of Language Model Security: Theory, Practice, and Fundamental Limits
Egor Zverev ⋅ Maura Pintor ⋅ Ana-Maria Cretu ⋅ Santiago Zanella-Beguelin ⋅ Nicole Nichols ⋅ Pavel Laskov
This workshop aims to advance research on secure-by-design LLM systems by shifting away from the current cat-and-mouse game of attacks and defenses toward a principled understanding of why security vulnerabilities arise and how to address them from the ground up. LLMs have been shown to be vulnerable to a range of attacks such as prompt injections and data poisoning, and yet continue to be deployed in complex systems without a clear understanding of why these vulnerabilities arise or how they interconnect with classical security vulnerabilities. At the model level, the absence of a hard separation between instructions and data may expose fundamental attack surfaces; at the system level, confused-deputy patterns and missing trust boundaries introduce further structural weaknesses. Understanding whether these vulnerabilities are inherent to current language modeling architectures or artifacts of specific design choices is essential for moving from brittle empirical defenses toward principled security. Recent research advocates treating LLM security as a system design problem, with a few approaches achieving provable security in specific settings. However, the field still lacks shared formal definitions of what LLM security means, comparable to differential privacy for privacy guarantees, and securing systems by design remains a use-case specific engineering effort rather than an application of generic principles. Moreover, existing solutions that offer security guarantees tend to degrade the utility of the system, and it is unclear whether this trade-off is an artifact of current approaches or a more fundamental limit of any LLM system that achieves meaningful security. This workshop aims to consolidate existing knowledge and lay the foundations for future LLM security research by answering three questions: * Q1. Formalizing LLM security. How should LLM security be formalized? Is there an agreed-upon framework comparable to the notion of differential privacy in privacy research? What role should theory play in creating secure LLM systems? * Q2. Security in practice. Can we design evaluation methodologies that are reproducible and generalizable rather than fragile and hackable? What concrete steps can help avoid unproductive cycles of attacks and defenses? * Q3. Fundamental limits of security. Existing approaches to securing LLM-based systems trade off security for utility. Is this trade-off an artifact of current defense designs, or a more fundamental property that any secure system must exhibit? In which settings has provable security already been achieved, and are there impossibility results establishing conditions under which security cannot be attained?
Show more
Real-Time Multimodal Conversational AI
Titouan Parcollet ⋅ Slim Essid ⋅ Shalini De Mello ⋅ Rogier van Dalen ⋅ Naomi Harte
For a long time, conversational AI has been confined to disembodied and unimodal text or speech exchanges. As we enter the era of embodied agents and virtual assistants, human-machine dialogue is increasingly being rooted into the physical world. This transition requires agents, whether virtual or physically embodied, to perceive scenes and humans through multimodal signals (e.g. speech, video, sensor recordings, etc.), to generate speech and non-verbal communication while maintaining a coherent conversation over time and to act, if necessary, in the physical world. This process must happen in real-time, with low latency and in a contextually-aware manner to enable a seamless human-machine collaboration. Historically, research on these challenges has been fragmented across communities, addressing isolated aspects rather than the problem as a whole. For instance, egocentric conversational AI focuses on the communication with an agent viewing the world as the user sees it. Examples include augmented-reality glasses, AI assistants such as Gemini Live or MiniCPM-o 4.5, or wearable cameras. These systems do not interact directly with the physical world. On the other hand, exocentric and dyadic conversational AI focuses on third-person perspectives and face-to-face communication, including traditional robotics or the more recent Interaction Model. However, real-time human-machine multimodal conversations inherently demand multiple perspectives. A seamless dialogue about a shared task, such as an AI assistant helping a human to assemble furniture, requires understanding both what the human sees (egocentric) and the context of the room, objects (exocentric) and human emotions, gesture and motion (dyadic) to perform cross-view reasoning. These challenges collectively define four foundational research axes requiring interdisciplinary collaboration across computer vision, robotics, machine learning, speech and audio processing and understanding, along with dialogue communities: multimodal representation and perception, real-time interaction dynamics and memory, emboddied interaction and finally, benchmarks and datasets. The 1st Real-Time Multimodal Conversational AI workshop will enable and catalyze progress along these four outstanding research problems.
Show more
Agentic Systems for Molecular Sciences
Nadine Schneider ⋅ Günter Klambauer ⋅ Ola Engkvist ⋅ Marwin Segler ⋅ Sohvi Luukkonen
Agentic large language model systems are rapidly moving from conversational assistants to "co-scientists" in the molecular sciences, planning syntheses, calling domain-specific tools, and closing experimental loops with laboratory automation. Recent demonstrations span autonomous drug repurposing, retrosynthesis planning, and genome-wide virtual screening, all built on orchestrated stacks of learned representations, predictors, and simulators. Yet careful benchmarking has exposed striking failures at seemingly simple tasks: tool-augmented agents reach only around 50% accuracy on chemical cost estimation, chemistry-oriented language models fail systematic symbolic reasoning on molecular graphs, and single-cell foundation models for perturbation prediction do not outperform linear baselines. This workshop takes the contrast between agentic ambition and methodological fragility as its starting point, with explicit space for negative results, rigorous baselines, and benchmark contributions alongside methodological advances. We invite researchers across machine learning for chemistry, structural biology, drug discovery, materials, and the methodological core of geometric and generative deep learning to join us in Paris.
Show more
Privacy in the Era of Large Opaque Models: Theoretical, Legal, and Practical Perspectives
Sana Tonekaboni ⋅ Lena Stempfle ⋅ Linus Bleistein ⋅ Maryam Molamohammadi ⋅ Franziska Boenisch ⋅ Adam Dziedzic ⋅ Linus Bleinstein
Foundation models, large language models, and agentic systems are rapidly reshaping the privacy landscape of machine learning. Across these paradigms, privacy risks are amplified by opacity: training data, post-training pipelines, alignment procedures, system components, and deployment contexts are often only partially visible to researchers, auditors, and users. These systems can expose sensitive information through memorization, retrieval, long-term memory, tool use, and cross-context information flow. At the same time, their opacity makes privacy assessment difficult: researchers often cannot inspect the data, audit the training pipeline, examine internal mechanisms, or evaluate the effect of privacy interventions. Together, these developments challenge existing definitions, benchmarks, mitigation strategies, governance frameworks, and accountability mechanisms.
Show more
E-Values: From Statistics to ML
Shubhada Agrawal ⋅ Sebastian Arnold ⋅ Yo Joong Choe ⋅ Peter Grünwald ⋅ Aaditya Ramdas
E-values provide a type of uncertainty quantification that is far more robust and flexible than classical measures (e.g., p-values): it enables anytime-valid statistical inference that have Type-I error guarantees under continuous monitoring. Unlike classical methods, e-values come with meaningful risk guarantees at adaptively chosen significance levels, allow for principled use of prior information without sacrificing validity, and play a foundational role in multiple testing. Building upon these recent breakthroughs on e-values---which have largely been concentrated in theoretical statistics---this workshop aims to bring these advances to the machine learning world and build a robust research community at the intersection of e-values and ML. We identify several key topics of interest, including the connections to conformal prediction; Bayesian and pseudo-Bayesian methods; bandit and adversarial learning; multiple testing; and modern ML applications such as auditing of LLMs and AI systems. The will be the first workshop dedicated to e-values at a top ML conference, featuring speakers with diverse interests ranging from multiple testing and conformal prediction to bandit applications and economics.
Show more
AI for Meta-Science: Scaling and Organizing Science in the Age of AI Scientists
prabhant singh ⋅ Hilde Weerts ⋅ Thanh Gia Hieu Khuong ⋅ Lele Cao ⋅ Joaquin Vanschoren
The use of AI in science production and evaluation has been growing in the past few years. With the advent of AI scientists and powerful agentic systems, we face opportunities and challenges in transforming the scientific ecosystem to accommodate this new reality. We propose our workshop on AI for Meta-Science~(science study). We aim to address challenges around the deployment of agentic AI systems in scientific processes and bring together researchers from meta-science, AI, peer review, and evaluation science with conference organizers, journal editors, and ArXiv moderators to address these challenges. We believe that this workshop provides a much-needed platform for the community to discuss and develop ideas and build the future of the scientific enterprise.
Show more
Generalization for Time Series in Tight Settings: Latency, Inference, Memory, prIvacy and Sustainability (TS-LIMITS)
Aurélie Boisbunon ⋅ Rémi Emonet ⋅ Valerio Frascolla ⋅ Elisa Fromont ⋅ Fabio Bonassi ⋅ June Sallou
The TS-LIMITS (Generalization for Time Series in Tight Settings: Latency, Inference, Memory, Privacy, and Sustainability) workshop addresses a critical yet underexplored challenge in machine learning: how to build time series models that generalize reliably under the real-world operational constraints that practitioners face every day. While much of the research community has focused on improving model accuracy, deployment in domains such as telecommunications, healthcare monitoring, industrial IoT, and autonomous systems demands that models also satisfy strict requirements on latency, inference efficiency, memory footprint, privacy preservation, and energy sustainability. TS-LIMITS brings together researchers and practitioners to share advances in areas including distribution shift, continual learning, federated and privacy-preserving learning, efficient architectures, and green AI, all through the lens of time series analysis. By fostering cross-disciplinary dialogue across these five tight constraints, the workshop aims to bridge the gap between academic research and real-world deployment, and to lay the groundwork for a more constraint-aware research agenda for time series generalization.
Show more
Successful Page Load